CtrlK
BlogDocsLog inGet started
Tessl Logo

ai-llm-engineering

Operational skill hub for LLM system architecture, evaluation, deployment, and optimization (modern production standards). Links to specialized skills for prompts, RAG, agents, and safety. Integrates recent advances: PEFT/LoRA fine-tuning, hybrid RAG handoff (see dedicated skill), vLLM 24x throughput, multi-layered security (90%+ bypass for single-layer), automated drift detection (18-second response), and CI/CD-aligned evaluation.

46

Quality

49%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills_all/ai-llm-engineering/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

31%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-organized navigation hub, but it carries notable repetition, defers essentially all executable guidance to bundle files that are missing, and its referenced references lead to dead links.

Suggestions

Ship the referenced bundle files (resources/*.md, templates/**, data/sources.json) or remove the broken links, since the skill currently navigates to content that does not exist.

De-duplicate the link listings: keep them once in Resources/Templates and let the Usage and Navigation Summary sections point back with a single pointer instead of re-enumerating every file.

Add at least one small inline executable example (e.g., a minimal RAG or ReAct snippet) in the Quick Reference so the body is actionable even before a user opens a template file.

DimensionReasoningScore

Conciseness

The same resource/template links are repeated across the Resources, Usage, and Navigation Summary sections, and the body carries padded metric claims ('24x throughput', '0.648 accuracy', '18-second response') that add noise without operational value.

2 / 5

Actionability

Aside from a high-level Quick Reference table and a conceptual decision tree, all executable/copy-paste guidance is deferred to template and resource files that are referenced but not actually present, leaving only minimal concrete guidance in the body.

2 / 5

Workflow Clarity

The Usage section gives numbered sequences for new projects, troubleshooting, and ongoing operations, but these are purely navigational ('read this file') steps with no validation checkpoints or feedback loops.

3 / 5

Progressive Disclosure

The body is designed as a clean overview pointing to one-level-deep resources/templates, but none of the referenced bundle files (resources/*.md, templates/**, data/sources.json) exist in the skill bundle, so the disclosure structure is non-functional in practice.

2 / 5

Total

9

/

20

Passed

Description

67%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and capability-rich, clearly conveying what the skill covers, but it omits any explicit 'Use when...' trigger guidance and its broad LLM-engineering scope overlaps with the many specialized sibling skills it links to.

Suggestions

Add an explicit 'Use when...' clause naming concrete user phrasing (e.g., building/deploying RAG or agentic apps, LLM evaluation, production rollout, drift/monitoring, tech-stack selection) to satisfy the completeness 'when' requirement.

Tighten the distinctiveness framing to emphasize the hub/index role versus the delegated depth skills, so the description is less likely to trigger for tasks that belong to a specialized skill.

Trim hard metric claims (e.g., '24x throughput', '90%+ bypass', '18-second response') which read as marketing padding and add little trigger or capability signal.

DimensionReasoningScore

Specificity

Names the LLM-engineering domain and lists multiple concrete capability areas (architecture, evaluation, deployment, optimization, PEFT/LoRA fine-tuning, hybrid RAG, vLLM serving, multi-layered security, drift detection), giving comprehensive coverage rather than just a few actions.

5 / 5

Completeness

Provides a clear 'what' (operational hub for LLM systems and named sub-areas) but contains no 'Use when...' clause or equivalent explicit trigger guidance, which per the rubric caps completeness at 3.

3 / 5

Trigger Term Quality

Includes natural user-facing terms (LLM, RAG, agents, deployment, fine-tuning, evaluation) but leans technical/buzzword-heavy (PEFT/LoRA, vLLM 24x throughput, FP8/FP4), leaving a few common natural phrases and synonyms unspecified.

4 / 5

Distinctiveness Conflict Risk

It carves a hub/navigator niche and explicitly delegates to specialized skills, but the breadth of 'LLM system architecture, evaluation, deployment, and optimization' still overlaps substantially with the several sibling skills it indexes, leaving real trigger conflict risk.

3 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 44 missing, 9 deeper-than-1-level, 18 suspicious

Warning

Total

15

/

16

Passed

Repository
Microck/ordinary-claude-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.