CtrlK
BlogDocsLog inGet started
Tessl Logo

mflux-model-tiny-test

Write a hermetic "tiny" model-saving test for an mflux model — a fast twin of the slow save/load test that runs the real ModelSaver/WeightLoader/WeightApplier seam on real component classes at toy dimensions, with no downloads or image generation. Use when asked to "make tiny test for <model>".

76

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality, expert-level skill body: concise, highly actionable, with a well-sequenced workflow and strong validation/feedback loops. The only minor gap is that all content lives in one inlined file with no progressive split, which is appropriate given the absence of a bundle.

DimensionReasoningScore

Conciseness

Lean and dense throughout: every section carries non-obvious, task-specific knowledge (quantization_group_size, the skip_quantization helper gap, interlocking dimension constraints) with no padding about what models or quantization are, assuming Claude's competence.

5 / 5

Actionability

Fully actionable — concrete file paths, exact shrink targets (layers to 2, hidden/intermediate to 128, heads to 2, vocab to 128), runnable commands (`grep -n ...`, `uv run --no-sync python -m pytest ...`, `ls tests/model_saving/`), and a symptom/cause troubleshooting table.

5 / 5

Workflow Clarity

A clear numbered 1–6 sequence with an explicit Verify section, a 'Constraints that will bite you' checklist, and a Troubleshooting table providing error-recovery feedback loops (e.g. lower tensors_per_shard, report a real save/load bug rather than weakening the assertion).

5 / 5

Progressive Disclosure

Well-organized into clear sections with clearly signaled one-level references to source-of-truth repo files (the three Reference implementations); no bundle files exist, so all content is inline in a single dense file with minor organization gaps.

4 / 5

Total

19

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that concretely states both the capability and an explicit natural-language trigger. It is specific, complete, and distinct, with only minor room to broaden trigger-term synonyms.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'runs the real ModelSaver/WeightLoader/WeightApplier seam on real component classes at toy dimensions' and 'no downloads or image generation' — giving comprehensive, specific coverage of what the skill produces.

5 / 5

Completeness

Clearly answers both 'what' (a hermetic fast twin save/load test running the real seam at toy dimensions) and 'when' ('Use when asked to "make tiny test for <model>"') with a concrete trigger phrase.

5 / 5

Trigger Term Quality

The trigger 'Use when asked to "make tiny test for <model>"' is a highly natural phrase a user would actually say, but coverage is essentially one phrase with limited synonyms or variations, leaving a few natural terms missing.

4 / 5

Distinctiveness Conflict Risk

The 'mflux model' + 'tiny model-saving test' + 'fast twin of the slow save/load test' framing carves a clear niche with a distinct trigger and minimal overlap risk with other skills.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mflux-community/mflux
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.