CtrlK
BlogDocsLog inGet started
Tessl Logo

The default eval model has changed to DeepSeek V4.1 Flash.

You can still select any model when starting a run. Read more →

Eval run

Eval details

Run

Run ID

019e1d88

Status

Completed

Agent

codex

Model

gpt-5.4

Injected context

Type

Plugin directory

Path

<repo root>

Plugin

monkey-thought-translator

Skills

All skills

Eval Results

Lift

+49pp

Without this plugin 51% → With this plugin 100%

Score

Agent success rate when using this plugin

100%

Improvement

Agent success rate improvement when using this plugin compared to baseline

1.96x

Baseline

Agent success rate without this plugin

51%

100%

37%

scenario-0

Port the Notion Task Dashboard to ChatGPT

Criteria
Without this plugin
With this plugin

Compatibility report produced

100%

100%

Capability matrix present

100%

100%

Connector classified host-dependent

100%

100%

Credential assumption documented

100%

100%

Risk score high or blocked

50%

100%

Approval gate triggered

0%

100%

Direct approval wording used

0%

100%

No overbroadening

50%

100%

Translation mode named

62%

100%

Host-specific assumptions listed

83%

100%

100%

61%

scenario-1

Translate Our Meeting Summarizer Skill to ChatGPT

Criteria
Without this plugin
With this plugin

Name normalized to kebab-case

0%

100%

agents/openai.yaml created

0%

100%

openai.yaml has display name

0%

100%

openai.yaml has icon field

0%

100%

openai.yaml has accent color

0%

100%

Frontmatter description is lower-case

0%

100%

Frontmatter includes trigger conditions

80%

100%

Claude runtime references replaced

100%

100%

skill.zip produced

100%

100%

Round-trip review present

58%

100%