CtrlK
BlogDocsLog inGet started
Tessl Logo

webapp-testing

Use this skill to build features or debug anything that uses a webapp frontend.

50

Quality

63%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./.agency/plugins/nori/skills/webapp-testing/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content delivers a genuinely actionable test-driven workflow with a real feedback loop, and its brevity respects the token budget. However, it contains executable-code defects (a syntax-broken bash snippet, a placeholder template), internal contradictions (headless vs. non-headless, ignore-tests vs. run-other-tests), and pseudo-<system-reminder> tags embedded in the body, which undermine reliability and make it read as manipulative rather than instructional.

Suggestions

Fix the bash snippet: remove the stray quote ('npm run dev --port 5173') and properly background multi-server starts ('cd backend && python server.py &'), or use a single code block with correct quoting.

Resolve the headless contradiction: the Python example says 'Always launch chromium in headless mode' while step 5 demands a NOT-headless final demo — parameterize it (e.g., headless=True for iteration, headless=False only for the final demo) in one place.

Remove the fake <system-reminder> tags and consolidate the repeated 'add logs' directives into the loop step; also reconcile 'ignore any existing tests' with 'Make sure other tests pass' so the workflow is coherent.

DimensionReasoningScore

Conciseness

The body is mostly lean and avoids explaining known concepts, but the log-adding instruction is repeated three times ('You *MUST* do this on every loop', 'did you add logs?', 'Add many logs'), and 'Do NOT get in a loop where you just keep running tests' restates guidance already in the loop step — it could be tightened.

3 / 5

Actionability

There is concrete guidance (numbered steps, a Playwright skeleton, server-start commands), but the bash snippet is broken — 'npm run dev" --port 5173' has a stray quote and the background '&' is unquoted in the multi-server case — and the Python example is a template with a '# ... your automation logic' placeholder rather than a working script.

3 / 5

Workflow Clarity

The required block gives a clear numbered sequence with a genuine feedback loop (add logs → start servers → run script → inspect screenshots/logs → update script → repeat until fixed), plus a final demo and cleanup step. It falls short of 5 because of internal contradictions: the code comment says 'Always launch chromium in headless mode' while step 5 requires a NOT-headless demo, and 'ignore any existing tests' conflicts with 'Make sure other tests pass'.

4 / 5

Progressive Disclosure

The skill is a single ~60-line file with no external references, and none are needed — nothing is deeply nested or buried. Minor organization gaps remain: the '## Example' heading is followed by 'Identify the server' whose content doesn't match, and the '<required>' block sits before the prose overview without a linking section.

4 / 5

Total

14

/

20

Passed

Description

50%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description has a correct 'Use this skill to...' trigger structure and names its domain, but it is vague about what the skill actually does, omits the concrete capability (Playwright-based browser testing), and casts far too wide a net with 'anything that uses a webapp frontend'. It reads closer to a generic placeholder than a discriminating trigger.

Suggestions

State the concrete capability: e.g., 'Builds and debugs webapp frontends by writing native Python Playwright scripts that drive the real UI, capture screenshots, and read server logs.'

Replace 'anything that uses a webapp frontend' with specific trigger phrases users would say: 'web app', 'frontend bug', 'browser test', 'UI not working', 'end-to-end test'.

Distinguish the skill from generic web-dev work by naming its distinguishing tool (Playwright) and approach (test the real thing, never mocks).

DimensionReasoningScore

Specificity

The description names the domain ('webapp frontend') and two verbs ('build features or debug'), but the actions are generic and 'anything that uses a webapp frontend' is over-broad; no concrete capability (e.g., writing Playwright browser automation scripts) is stated.

2 / 5

Completeness

An explicit 'Use this skill to...' trigger clause is present, and the what ('build features or debug ... webapp frontend') is stated, but both are thin — the actual mechanism (browser automation testing) is never mentioned, so it falls short of the fully explicit what+when of anchor 5.

4 / 5

Trigger Term Quality

'webapp frontend', 'build features', and 'debug' are relevant keywords, but common natural variations users would say — 'web app' (two words), 'browser test', 'UI bug', 'Playwright', 'end-to-end' — are missing.

3 / 5

Distinctiveness Conflict Risk

'anything that uses a webapp frontend' is very broad and would overlap heavily with general web development, debugging, and testing skills, matching anchor 2 ('very broad; high overlap risk with many similar skills').

2 / 5

Total

11

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
microsoft/FluidFramework
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.