CtrlK
BlogDocsLog inGet started
Tessl Logo

gstack-openclaw-office-hours

Use when asked to brainstorm, evaluate whether an idea is worth building, run office hours, or think through a new product idea or design direction before any code is written.

65

Quality

77%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./openclaw/skills/gstack-openclaw-office-hours/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An unusually actionable, well-sequenced conversational workflow: verbatim question scripts, GOOD/BAD pushback exemplars, explicit approval gates, and loop-back feedback on premise disagreement. The main structural weakness is that everything — templates, question banks, closing scripts — is inlined in one 370-line file with no progressive disclosure to reference files, and a modest amount of rule repetition.

Suggestions

Move the two design doc templates and the verbatim Phase 6 closing scripts into references/ files (e.g. references/templates.md, references/closing.md) and link them from the phase sections, keeping SKILL.md as the workflow overview.

Deduplicate the repeated 'questions ONE AT A TIME / STOP after each question' instruction into the Important Rules section only, and state the mode mapping once in Phase 1.

Tighten the Phase 2A red-flag lists, which overlap heavily with the corresponding 'Push until you hear' lines (e.g. Q2's workaround red flag restates the push target).

DimensionReasoningScore

Conciseness

The body is dense and purposeful — operating principles, GOOD/BAD pushback pairs, exact question scripts, and templates with almost no explanation of concepts Claude already knows. There is some redundancy that could be trimmed: "STOP after each question" appears in both Phase 2A and 2B and again in Important Rules, and the mode mapping is effectively stated twice. This matches 'efficient; minor instances of over-explanation that could be trimmed' rather than the lean every-token-earns-its-place level 5.

4 / 5

Actionability

For an instruction-only skill the guidance is fully executable: verbatim questions to ask ("What's the strongest evidence you have that someone actually wants this..."), concrete GOOD/BAD reply contrasts per pushback pattern, fill-in design doc templates, and word-for-word closing scripts keyed to signal counts. Per the rubric's code_vs_instruction note, absence of code is not penalized when guidance is this actionable.

5 / 5

Workflow Clarity

Phases 1-6 are clearly sequenced with explicit validation checkpoints and feedback loops: STOP-and-wait after every question, premise confirm/disagree with 'revise understanding and loop back', 'Do NOT proceed without their approval' after alternatives, and an Approve/Revise/Start-over gate on the design doc. Stage-based question routing and an escape hatch handle branching. This matches the top anchor.

5 / 5

Progressive Disclosure

No references/, scripts/, or assets/ directories exist, so all ~370 lines live in SKILL.md. Section organization is good, but content that clearly belongs in separate files is inlined — the two full design doc templates, the verbatim closing scripts, and the six-question bank are all candidates for one-level-deep reference files. This matches 'some structure but could be better organized; content that should be separate is inline'; it is not a 2 because the phase headers do provide real navigational structure.

3 / 5

Total

17

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A well-constructed description with an explicit and natural 'Use when' trigger clause and good distinctiveness. Its main weakness is that the action list is a cluster of near-synonyms for one ideation activity and never names the concrete output (a design document), leaving the 'what' side thinner than the 'when'.

Suggestions

State the concrete output in the description, e.g. 'Produces a design document capturing demand evidence, premises, and recommended approach — no code is written.'

Diversify the action verbs beyond ideation synonyms (e.g. 'challenges premises, compares 2-3 alternative approaches, and assigns one concrete next action') and add common trigger synonyms like 'validate an idea' or 'is this worth building'.

DimensionReasoningScore

Specificity

The description lists several action phrases — "brainstorm", "evaluate whether an idea is worth building", "run office hours", "think through a new product idea or design direction" — but these are largely synonymous variations of one ideation activity rather than distinct concrete capabilities, and the actual deliverable (a design document) is never stated. This matches the anchor 'names domain and 1-2 concrete actions, but not comprehensive' better than the level-4 'several specific actions with minor gaps', because the actions collapse into a single activity with no output artifact named.

3 / 5

Completeness

Both halves are explicitly present: an explicit "Use when asked to..." trigger clause plus a stated set of activities. It falls short of a 5 because the 'what' side never states what the skill actually produces (design docs) or the hard no-code boundary, so the 'what' is thinner than the clearly explicit 'when'.

4 / 5

Trigger Term Quality

Natural user phrasings are present: "brainstorm", "run office hours", "evaluate whether an idea is worth building", "think through a new product idea or design direction". Common synonyms a user might actually say — "validate an idea", "is this idea worth building", "product feedback", "pitch feedback" — are only partially covered, matching 'good keyword coverage; a few natural terms missing' rather than the comprehensive synonym coverage of a 5.

4 / 5

Distinctiveness Conflict Risk

The pre-code ideation framing ("before any code is written", "run office hours", "design direction") carves out a distinct niche that would not fire for implementation or code-review skills. Minor overlap risk remains with generic brainstorming/planning skills, matching 'mostly distinct; minor overlap risk with closely related skills'.

4 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
garrytan/gstack
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.