CtrlK
BlogDocsLog inGet started
Tessl Logo

deep-interview

Socratic deep interview with mathematical ambiguity gating before execution

49

Quality

62%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/oh-my-codex/skills/deep-interview/SKILL.md

The canonical home for this skill is deep-interview in Yeachan-Heo/oh-my-codex

SKILL.md
Quality
Evals
Security

Quality

Content

62%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers an unusually rigorous, well-sequenced workflow with explicit gates, scoring formulas, and concrete OMX commands, but it is severely bloated: the same rules are restated across Execution_Policy, Steps, Tool_Usage, and the checklist, and hundreds of lines of peripheral material (autoresearch, handoff contracts, payload schemas) are inlined instead of split into reference files.

Suggestions

Split peripheral material into one-level-deep reference files (e.g., references/autoresearch.md, references/handoff-contracts.md, references/omx-question-payloads.md), keeping only the core interview loop and a summary table of handoff options in SKILL.md.

Deduplicate rules stated in multiple sections: state each rule once (the oversized-context gate, the answers[] contract, the intent-first stage priority, and the state-writer authority each appear 3-5 times across Execution_Policy, Steps, Tool_Usage, and the Final Checklist).

Replace repeated prose reinforcement of non-goals/decision-boundary gates with a single compact gate table, and trim the `omx question` payload guidance to the two canonical examples plus a one-line answer-shape note.

DimensionReasoningScore

Conciseness

The ~570-line body repeats the same rules in multiple sections: the oversized-context summary gate appears in Execution_Policy, Phase 0, Phase 2a, Tool_Usage, and the Final Checklist; the `answers[]` vs legacy `answer` contract is restated three times; intent-first/non-goals/decision-boundary guidance recurs throughout. It is not a 1 because it does not explain concepts Claude already knows (no padding about what Socratic method or ambiguity means), but the redundancy is clearly noticeable verbosity.

2 / 5

Actionability

Guidance is mostly executable: concrete commands (`omx state write --input '<json>' --json`, `OMX_QUESTION_RETURN_PANE=$TMUX_PANE omx question ...`), canonical JSON payload examples, exact artifact paths, and a copy-paste-ready config TOML block. It is not a 5 because several payloads remain placeholder-shaped (e.g., `<slug>`, `<uuid>`, the stride contract is shown only as a filled example without an empty template) and some judgment-heavy steps ("challenge core assumptions") lack worked examples.

4 / 5

Workflow Clarity

The multi-phase workflow (Phase 0 preflight through Phase 5 execution bridge) is clearly sequenced with explicit validation checkpoints: per-round ambiguity scoring with weighted formulas, readiness gates (Non-goals, Decision Boundaries), a pressure-pass requirement, a practical closure audit, escalation/stop conditions, blocked-state persistence, and a final checklist. It is not a 4 because validation and error-recovery feedback loops are explicit and repeated at every stage boundary rather than having minor gaps.

5 / 5

Progressive Disclosure

The file is a monolithic ~570-line document with no bundle files at all: the autoresearch specialization, the five execution handoff contracts, the `omx question` payload schema, and the config reference are all inlined when they clearly belong in separate reference files. It is above a 2 because the body is well-sectioned with clear XML-style section tags and headers making it navigable, but it is below 4 because nothing is offloaded and a large fraction of the content is peripheral to the core interview loop.

3 / 5

Total

14

/

20

Passed

Description

46%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is terse and distinct, clearly naming the skill's purpose, but it reads as internal jargon rather than user-facing trigger language. It lacks any "Use when..." guidance and natural trigger phrases, capping both completeness and trigger-term quality.

Suggestions

Add an explicit trigger clause, e.g., "Use when the user says 'interview me', 'ask me everything', 'don't assume', or when a request is vague and needs clarification before planning or execution."

Replace jargon like "mathematical ambiguity gating" with plain-language capabilities users would recognize (e.g., "scores remaining ambiguity each round and stops questioning once it drops below a threshold").

List 1-2 more concrete actions (e.g., "asks one question at a time, pressure-tests assumptions, and writes an execution-ready spec") to round out capability coverage.

DimensionReasoningScore

Specificity

The description names the domain and 1-2 actions ("Socratic deep interview", "mathematical ambiguity gating") but coverage is not comprehensive and the actions are stated abstractly. It is above a score of 2 because the domain and concrete actions are named rather than generic, but below 4 because it does not list several specific capabilities.

3 / 5

Completeness

The "what" is clearly stated (Socratic interview with ambiguity gating) but there is no "Use when..." clause or equivalent explicit trigger guidance; "before execution" only weakly implies when to use it. Per the judging guidelines, a missing 'Use when' clause caps completeness at 3, and it is not a 2 because the "what" is clear rather than vague.

3 / 5

Trigger Term Quality

Phrases like "Socratic", "mathematical ambiguity gating", and "before execution" are technical jargon rather than natural user language; only "deep interview" is a plausible user keyword. It is not a 1 because "deep interview" is a usable keyword, but common natural variations (e.g., "interview me", "clarify requirements", "ask me questions") are missing.

2 / 5

Distinctiveness Conflict Risk

"Socratic deep interview with mathematical ambiguity gating" carves a fairly distinct clarification-gating niche unlikely to collide with unrelated skills, with only minor overlap risk against generic planning/interview skills. It is not a 5 because the boundary against a lightweight "plan" or "interview" skill is not drawn in the description itself.

4 / 5

Total

12

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (572 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
Yeachan-Heo/oh-my-codex
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.