CtrlK
BlogDocsLog inGet started
Tessl Logo

laya-integration

Add fast, local, typed decisions to any project with Laya, an open-source non-generative decision model (pip install laya). It classifies, routes, scores and answers yes/no questions about a piece of text, returning calibrated probabilities in roughly 20-35 ms on a laptop GPU, with no LLM call and no data leaving the machine. Use this skill whenever the user wants to classify or route text, triage tickets or emails, detect spam, phishing, toxicity or intent, put a guardrail in front of an agent or LLM, score something against a rubric, or replace an LLM-based classifier to cut latency or cost. Also use it when they mention Laya, Jev, a "System 1 model", "typed decisions" or a "decision model", even if they never name Laya.

76

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, unusually specific integration guide: executable code at every step, concrete operational limits and pitfalls, and a genuine evaluate-before-ship feedback loop with a closing checklist. The weaknesses are mild: version/date-pinned details and long calibration/benchmark passages add tokens that could be tightened, and a ~230-line monolith with no reference files leaves detailed material inline that a one-level-deep reference could absorb.

Suggestions

Move the calibration-fitting procedure and the 500-example evaluation narrative into a single one-level-deep reference file (e.g. references/evaluation.md), keeping a two-line pointer plus the headline numbers in SKILL.md — this would relieve both the conciseness and progressive-disclosure pressure from the ~230-line monolith.

Trim or date-stamp the version-pinned caveats: 'As of laya 0.3.4 the detector only recognises...' will silently go stale; consider condensing the supported-language list to the rule (pass lang= whenever the language is known) and moving the enumeration next to the version note.

Tighten the device/latency benchmark sentences ('about 34 ms (English) and 21 ms (multilingual) on an M1 Max GPU, 139 ms and 58 ms on its CPU') to the one decision-relevant fact per platform, cutting roughly a third of that passage's tokens.

DimensionReasoningScore

Conciseness

The body is dense with non-obvious, model-specific knowledge ("Each option is truncated to 48 tokens", "agent.temperature is [choice, score, noul]", "one GPU serves one forward pass at a time") and assumes Claude's competence, but it carries time-sensitive details — "As of laya 0.3.4 the detector only recognises..." and the "September 2026" benchmark narrative — plus the lengthy calibration-math and device-benchmark passages that could be trimmed or split out. This sits between "Lean and efficient; every token earns its place" (5) and the current fit, anchor 4: "Efficient; minor instances of over-explanation that could be trimmed".

4 / 5

Actionability

The code is fully executable and copy-paste ready: a complete `laya.load`/`predict` call with realistic question schemas, `Router` usage, a threshold-gated routing snippet, a working FastAPI sidecar, and a numbered evaluation recipe ending in a concrete fine-tuning pointer ("ships a fine-tuning notebook that runs on free Kaggle GPUs"). This matches anchor 5, "Fully executable; copy-paste ready code or commands; specific examples cover the common cases" — pseudocode and missing key details would put it at 3, which is not the case.

5 / 5

Workflow Clarity

The body sequences a clear install → load-once → warm-up → author-questions → pick-checkpoint → threshold → wire-in → evaluate path, with an explicit validate-and-iterate feedback loop ("Collect 50-200 real examples... Run them through predict and record accuracy per question... Try alternative phrasings and checkpoints... If accuracy is still short, fine-tune") and error-recovery hints ("If a download hangs at 0 bytes, set HF_HUB_DISABLE_XET=1"). It closes with a checklist, matching anchor 5's "explicit validation steps; feedback loops for error recovery; checklists for complex processes".

5 / 5

Progressive Disclosure

The skill is a single well-organized file with clear section headers, a question-type table, and code blocks — good structure with most content appropriately placed inline, since there is no large separable API-reference bulk (anchor 4). It falls short of anchor 5 because at ~230 lines there is no one-level-deep reference split at all: material like the calibration-fitting procedure and the 500-example evaluation narrative could live in a reference file, and the simple-skill exception (under 50 lines) does not apply.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete multi-action capability statement, explicit 'Use when' triggers covering the full range of natural user phrasings, third-person voice, and distinctive brand aliases. Every clause carries information; there is no fluff or over-claiming.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "classifies, routes, scores and answers yes/no questions about a piece of text, returning calibrated probabilities" — plus concrete install and latency specifics ("pip install laya", "roughly 20-35 ms on a laptop GPU"). This matches the anchor "Lists multiple specific concrete actions; comprehensive coverage" exactly; nothing is vague or generic.

5 / 5

Completeness

Both 'what' ("an open-source non-generative decision model... classifies, routes, scores and answers yes/no questions... returning calibrated probabilities") and 'when' ("Use this skill whenever the user wants to classify or route text, triage tickets or emails, detect spam...") are explicit and concrete, matching the anchor "Clearly and explicitly answers both what AND when with concrete trigger phrases".

5 / 5

Trigger Term Quality

It covers the natural phrases a user would actually say — "classify or route text, triage tickets or emails, detect spam, phishing, toxicity or intent, put a guardrail in front of an agent or LLM, score something against a rubric, or replace an LLM-based classifier to cut latency or cost" — plus synonyms and brand aliases ("Laya, Jev, a 'System 1 model', 'typed decisions'"). Coverage is comprehensive including synonyms, matching the top anchor.

5 / 5

Distinctiveness Conflict Risk

The niche is clear — a local, non-generative decision model with no LLM call — and triggers are pinned to distinctive terms ("Laya, Jev, a 'System 1 model', 'typed decisions' or a 'decision model', even if they never name Laya"). While generic triggers like "detect spam" or "guardrail" could superficially overlap with moderation skills, the framing makes conflict risk minimal, matching the anchor "Clear niche with distinct triggers; minimal conflict risk" rather than the 'minor overlap' level 4.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
stevenknowswhy/ProfessionalBuyer
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.