CtrlK
BlogDocsLog inGet started
Tessl Logo

gguf-quantization

GGUF format and llama.cpp quantization for efficient CPU/GPU inference. Use when deploying models on consumer hardware, Apple Silicon, or when needing flexible quantization from 2-8 bit without GPU requirements.

60

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-inference/gguf/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

61%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured skill document dominated by executable commands and complete quantization workflows. Its weaknesses are repetition of the same commands across three sections, and the absence of validation/verification checkpoints in the batch-quantization workflow, which caps workflow clarity.

Suggestions

Add a verification step to the batch workflow (Workflow 3), e.g., check each output exists and has the expected size (`du -h $OUTPUT`) and run a short generation test per quant before declaring success; include an error-recovery branch if a quantization fails.

Deduplicate the repeated convert/quantize/imatrix commands — keep them once in the workflows and have 'Quick start' and 'Common issues' point there, cutting the body significantly.

Move the inlined 'Common issues' section into references/troubleshooting.md (which already covers debugging) and link to it, keeping SKILL.md as the overview.

DimensionReasoningScore

Conciseness

At ~418 lines the body is mostly lean code and commands (little concept-explanation padding), but the same convert/quantize/imatrix commands appear three times (Quick start, Workflow 2, 'Common issues'), and build instructions are repeated in both 'Installation' and 'Hardware optimization'. That is more than the 'minor instances' of anchor 4 — noticeable duplication that could be tightened — though not the pervasive over-explanation of anchor 2.

3 / 5

Actionability

Predominantly copy-paste-ready executable commands: git clone/build, convert_hf_to_gguf.py, llama-quantize, llama-cli, llama-server, and working llama-cpp-python examples covering load, chat, streaming. Minor gaps keep it from 5: a Python snippet sits inside a ```bash block under 'Apple Silicon (Metal)', the bare `make` build flow omits the current cmake steps, and the `--mmap` flag in 'Common issues' is not a valid llama-cli option.

4 / 5

Workflow Clarity

Workflows 1-3 are clearly numbered and sequenced with concrete commands, but validation is implicit or absent: Workflow 1's '4. Test' has no success criteria, and Workflow 3 is a batch loop generating multiple quantizations with no verification checkpoint (e.g., checking output size or running a perplexity/quality check before proceeding). Per the rubric's batch-operation rule, missing validation caps this at 3 despite the good sequencing.

3 / 5

Progressive Disclosure

Two real, one-level-deep references (advanced-usage.md, troubleshooting.md) are clearly signaled in a 'References' section with descriptive labels, and the body has a clean section hierarchy. It falls short of 5 because ~40 lines of 'Common issues' content are inlined in SKILL.md despite troubleshooting.md existing, and the 400+ line body carries integration/hardware detail that could partly live in the reference files — matching anchor 4's 'minor organization gaps'.

4 / 5

Total

14

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit and specific 'Use when...' trigger clause, good natural keywords, and a clearly distinct GGUF/llama.cpp niche. The main weakness is that the 'what' is a noun-phrase fragment rather than enumerated concrete actions, which limits its specificity.

Suggestions

Rewrite the opening as explicit third-person actions, e.g., 'Converts HuggingFace models to GGUF, quantizes them to 2-8 bit K-quants, and runs inference via llama.cpp on CPU or Apple Silicon.'

Add widely-used ecosystem trigger terms users actually say — '.gguf', 'Ollama', 'LM Studio', 'quantize a model', 'local LLM' — to broaden natural keyword coverage.

DimensionReasoningScore

Specificity

The description names the domain concretely — 'GGUF format and llama.cpp quantization' with '2-8 bit' and hardware targets — but contains no explicit action verbs ('convert', 'quantize', 'run'); it reads as a domain-plus-purpose fragment like 'Processes PDF files and extracts content' rather than a list of capabilities. It is above a score of 2 because it names two concrete capability areas (quantization, efficient inference) with a mechanism, but below 4 because no specific actions are enumerated.

3 / 5

Completeness

Both questions are explicitly answered: what — 'GGUF format and llama.cpp quantization for efficient CPU/GPU inference'; when — an explicit 'Use when deploying models on consumer hardware, Apple Silicon, or when needing flexible quantization from 2-8 bit without GPU requirements' clause with concrete trigger conditions. This matches the anchor-5 pattern of a clear 'what' plus an explicit, specific 'when'; anchor 4 ('when' could be more explicit) does not fit since the when-clause enumerates three concrete scenarios.

5 / 5

Trigger Term Quality

Good natural keyword coverage: 'GGUF', 'llama.cpp', 'quantization', 'consumer hardware', 'Apple Silicon', 'CPU/GPU inference', 'deploying models' — terms a user would plausibly say. A few common variations are missing (e.g., '.gguf', 'Ollama', 'LM Studio', 'local LLM', 'quantize a model'), so it falls just short of the comprehensive anchor 5 but clearly above anchor 3's 'some relevant keywords'.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche (GGUF/llama.cpp quantization for CPU/Apple-Silicon/consumer deployment) with distinct triggers like 'GGUF', 'llama.cpp', and 'Apple Silicon'. It even implicitly excludes GPU-calibrated alternatives via 'without GPU requirements', keeping conflict risk with AWQ/GPTQ/TensorRT skills minimal — matching the anchor-5 'clear niche with distinct triggers'.

5 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.