CtrlK
BlogDocsLog inGet started
Tessl Logo

learning-mechanics

Apply the "learning mechanics" framework — the emerging physics-style theory of deep learning training dynamics — when writing, debugging, scaling, or tuning neural-network code. Use for decisions about learning rate / batch size / width / depth scaling, μP (Maximal Update Parameterization) and hyperparameter transfer, edge-of-stability and progressive sharpening, lazy vs. rich (feature-learning) regimes and initialization scale, neural scaling laws, gradient-flow conservation laws / symmetries, neural collapse, the neural feature ansatz, and for designing scientific experiments on training. Trigger phrases: "learning mechanics", "why is training unstable", "how should I scale the learning rate with width/batch", "muP / mup / hyperparameter transfer", "edge of stability", "lazy vs rich", "feature learning regime", "scaling laws", "tune hyperparameters on a small model". Source: Simon et al., "There Will Be a Scientific Theory of Deep Learning" (2026).

77

Quality

96%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured applied-theory reference: tight decision rules, a sequenced workflow with validation and feedback loops, and clean one-level progressive disclosure into a genuine reference file. The only gap is executability — concrete formulas and thresholds but no runnable code or commands for the instrumentation it prescribes.

Suggestions

Add a minimal runnable snippet for workflow step 2 — e.g. a short PyTorch hook that logs gradient norm, update norm, and update-to-weight ratio per step — so the prescribed instrumentation is copy-paste ready.

Include one concrete μP transfer example (narrow-model LR sweep → wide-model transfer with the `mup` library) to make decision rule #1 executable rather than descriptive.

Show the sharpness-proxy estimation referenced in rule #5 (e.g. a power-iteration snippet for λ_max) so the `λ_max ≳ 2/η` debugging rule can be applied directly.

DimensionReasoningScore

Conciseness

Lean and dense throughout: no explanation of concepts Claude already knows; every section is a decision rule, formula, or checklist ('`2/η` is the max stable sharpness', 'final-layer weights ~`1/width` instead of `1/sqrt(width)`'). The only meta-material (honesty notes, source) is brief and load-bearing for a position-paper-derived skill.

5 / 5

Actionability

Concrete, specific guidance throughout ('If `λ_max ≳ 2/η`, lower `η`', 'scale LR with √(batch size)', 'use a μP-aware library (e.g. `mup`)'), and as an instruction-only skill code absence is not itself penalized — but nothing is copy-paste ready, and the step-2 instrumentation list (gradient norm, update-to-weight ratio, activation RMS…) gives no minimal snippet or command for actually capturing these.

4 / 5

Workflow Clarity

'Workflow (do these in order — don't skip step 2)' provides a clear 6-step sequence with an explicit validation step ('plot observed vs. predicted') and an error-recovery feedback loop ('if it fails, decide whether the proxy, the limit, the independent variable, or the instrumentation was wrong'), plus a reproducibility checklist.

5 / 5

Progressive Disclosure

SKILL.md holds the applied 'what' while the deeper 'why' is split into a real, one-level-deep reference, clearly signaled and gated ('read `references/core-method.md` only when a task needs the *why*'); the reference exists, contains no nested references, and the body's sections are well organized for navigation.

5 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a model of the form: a concrete what-clause, an explicit when-clause, and a literal trigger-phrase list including synonyms. It is long, but every phrase is a specific capability or trigger, not padding.

DimensionReasoningScore

Specificity

Names multiple concrete actions ('writing, debugging, scaling, or tuning neural-network code') and comprehensively enumerates the decision areas ('μP (Maximal Update Parameterization) and hyperparameter transfer', 'edge-of-stability and progressive sharpening', 'neural collapse, the neural feature ansatz') in third person, with no vague filler.

5 / 5

Completeness

Clearly answers both what ('Apply the "learning mechanics" framework … when writing, debugging, scaling, or tuning neural-network code') and when ('Use for decisions about …' plus a literal 'Trigger phrases:' list), matching the top anchor.

5 / 5

Trigger Term Quality

Explicit trigger phrases are natural user utterances with synonyms and variants: 'why is training unstable', 'how should I scale the learning rate with width/batch', 'muP / mup / hyperparameter transfer', 'lazy vs rich', 'tune hyperparameters on a small model' — comprehensive coverage.

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche — physics-style deep-learning training-dynamics theory — with distinct triggers ('edge of stability', 'feature learning regime', 'scaling laws') that are unlikely to fire for unrelated skills.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Ares960826/ares-agent-toolkit
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.