CtrlK
BlogDocsLog inGet started
Tessl Logo

flow-nexus-neural

Train and deploy neural networks in distributed E2B sandboxes with Flow Nexus

65

7.38x
Quality

50%

Does it follow best practices?

Impact

96%

7.38x

Average score across 3 eval scenarios

SecuritybySnyk

—

The risk profile of this skill

Fix and improve this skill with Tessl

tessl review fix ./ai-ml/flow-nexus-neural-mattnigh-skills-collection/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

46%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill is highly actionable — comprehensive, executable MCP tool examples with response shapes — but it is a monolithic ~740-line API reference dumped into SKILL.md with duplicated examples, redundant general-ML explanation, and no progressive disclosure to reference files. Workflows lack validation gates, especially before the destructive cluster-terminate operation.

Suggestions

Split the bulk of the API reference into one-level-deep reference files (e.g., references/training.md, references/clusters.md, references/marketplace.md) and keep SKILL.md as a concise overview with clearly signaled links, per progressive disclosure.

Remove the duplicated LSTM example and the 'Best for:' architecture-primer section that re-explains knowledge Claude already has, and trim full JSON response payloads to the fields that matter.

Add explicit validation checkpoints to the cluster workflow — verify training_status shows completion and the model is saved/published before calling neural_cluster_terminate, with a fix-and-retry loop in Troubleshooting.

DimensionReasoningScore

Conciseness

The ~740-line body inlines an entire API reference with full JSON response payloads, repeats the LSTM architecture example verbatim in both the training section and the Time Series use case, and re-explains knowledge Claude already has ("Best for: Time series, sequences, forecasting" for LSTMs, "Start Small" advice). This matches anchor 2 (noticeably verbose, several padded sections) — more padding than anchor 3's "some unnecessary explanation", but not the continuous beginner-tutorial prose of anchor 1 since most content is still operational API detail.

2 / 5

Actionability

Nearly all guidance is concrete, copy-paste-ready MCP tool invocations with full parameter objects and realistic response shapes (e.g., neural_train with layers/training config, neural_cluster_init with topology/consensus options). It falls short of anchor 5 because of gaps like the GAN example's `generator_layers: [...]` pseudocode placeholders, unexplained `input_dim`, and undefined `user_id` placeholders, matching anchor 4's "mostly executable with minor gaps".

4 / 5

Workflow Clarity

The cluster workflow is sequenced (init → node_deploy → connect → train_distributed → status → terminate) and there is a Troubleshooting section, but there are no validation checkpoints: training_status/cluster_status are shown as calls, not as explicit gates, and the destructive cluster_terminate is never preceded by verifying training completion or saving the published model. Per the guideline capping batch/destructive operations without validation at 3, this cannot exceed anchor 3; it is above anchor 2 because the sequence itself is well defined.

3 / 5

Progressive Disclosure

There is no references/ bundle and no external files at all — the entire ~700-line API reference, architecture patterns, marketplace, and troubleshooting content is inlined in SKILL.md. This matches anchor 2 (content that clearly belongs in separate files is inlined) despite good section headers, because the monolithic single-file structure forces every consumer to load the full reference; it is above anchor 1 only because internal navigation via headers is possible.

2 / 5

Total

11

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description states a clear, specific 'what' anchored to a distinct tool niche, but provides no 'when to use' guidance and has only moderate natural-keyword coverage. Adding an explicit trigger clause (e.g., "Use when the user asks to train, deploy, or run inference on neural networks with Flow Nexus") would substantially improve it.

Suggestions

Add a 'Use when...' clause listing concrete trigger phrases (e.g., "Use when the user wants to train, fine-tune, benchmark, or serve neural networks via Flow Nexus, or mentions E2B sandboxes or distributed ML training") to lift completeness and trigger coverage.

Include natural synonyms users would say — "deep learning", "machine learning model", "distributed training", "federated learning" — to broaden trigger-term coverage.

Mention the additional concrete capabilities documented in the body (inference, template marketplace, cluster/federated training) so the description's action list is more comprehensive.

DimensionReasoningScore

Specificity

"Train and deploy neural networks in distributed E2B sandboxes" names the domain and two concrete actions (train, deploy), but coverage is not comprehensive — inference, the template marketplace, clusters, and model management from the body are absent. It sits at anchor 3 rather than 4 because only two actions are listed with a minor gap; it is above anchor 2 because the actions are concrete rather than generic.

3 / 5

Completeness

The 'what' is clearly stated (train and deploy neural networks in E2B sandboxes via Flow Nexus), but there is no 'when' — no "Use when..." clause or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. Not anchor 4 because 'when' is entirely absent rather than merely imprecise; not anchor 2 because the 'what' is concrete, not vague.

3 / 5

Trigger Term Quality

Terms like "train", "deploy", "neural networks", "E2B sandboxes", and "Flow Nexus" are relevant, but common natural variations users would say — "deep learning", "machine learning model", "distributed training", file/tool synonyms — are missing. This matches anchor 3 (some relevant keywords, missing common variations); it falls short of anchor 4's good coverage and is clearly above anchor 2's generic-only phrasing.

3 / 5

Distinctiveness Conflict Risk

The "Flow Nexus" / "E2B sandboxes" niche is fairly distinct and unlikely to trigger the wrong skill, but the generic opening "Train and deploy neural networks" overlaps with any general ML training skill. This fits anchor 4 (mostly distinct, minor overlap risk) better than anchor 5, which requires fully distinct triggers, and better than anchor 3, since the tool-specific framing does narrow it considerably.

4 / 5

Total

13

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (739 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
majiayu000/claude-skill-registry-data
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.