CtrlK
BlogDocsLog inGet started
Tessl Logo

hqq-quantization

Half-Quadratic Quantization for LLMs without calibration data. Use when quantizing models to 4/3/2-bit precision without needing calibration datasets, for fast quantization workflows, or when deploying with vLLM or HuggingFace Transformers.

61

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/hqq/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable reference with executable code for every major use case, explicit workflows with a verification step, and clean one-level-deep references. Its weaknesses are redundancy — the same quantize snippet repeated several times and mixed precision shown twice — and advanced material (PEFT training, vLLM serving) inlined in the main file rather than pushed to the existing references.

Suggestions

Consolidate the four near-identical AutoModelForCausalLM.from_pretrained + HqqConfig snippets into one canonical example and reference it from the workflows.

Move the PEFT/LoRA fine-tuning and vLLM serving sections into references/advanced-usage.md, leaving a one-line pointer each, to cut the main file's length and reduce duplication.

Make the runnable examples self-contained: define 'input_tensor' in the basic example and load the tokenizer in Workflow 2, and either define 'train_dataset'/'data_collator' or mark them explicitly as placeholders.

DimensionReasoningScore

Conciseness

The body avoids explaining concepts Claude already knows, but it repeats the same AutoModelForCausalLM.from_pretrained + HqqConfig snippet nearly verbatim in 'Quick start', 'HuggingFace integration', 'Quantize and save', and 'Workflow 1', and demonstrates mixed precision twice (layer_configs dict and dynamic_config). It fits anchor 3 — mostly efficient but could be tightened by consolidating the duplicate snippets.

3 / 5

Actionability

Nearly all guidance is concrete, executable code covering install, basic quantization, HF, vLLM, PEFT, and troubleshooting, matching the good-example pattern. Not 5 because a few snippets are not copy-paste runnable as-is: 'input_tensor' is never defined in the basic example, Workflow 2 uses 'tokenizer' without loading it, and the Trainer example references undefined 'train_dataset'/'data_collator'.

4 / 5

Workflow Clarity

Both workflows are explicitly numbered, and Workflow 1 includes a real validation checkpoint ('3. Verify quality' with generation test code) before saving; 'Common issues' provides error-recovery fixes (OOM → sequential loading, poor 2-bit quality → smaller group size). Not 5 because the recovery guidance is not framed as validate→fix→retry loops inside the workflows, and Workflow 2 has a gap (tokenizer never loaded before benchmarking).

4 / 5

Progressive Disclosure

The References section clearly signals two real one-level-deep files (references/advanced-usage.md, references/troubleshooting.md), both of which exist and cover genuinely advanced material (sensitivity-based quantization, ONNX export, benchmarking, debugging). Not 5 because the main body still inlines substantial advanced content (full PEFT/Trainer setup, vLLM serving, per-backend tables) that overlaps what belongs in the reference files, making SKILL.md longer than an overview needs to be.

4 / 5

Total

15

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit 'Use when...' trigger clause, concrete precision levels, and named deployment frameworks. Its main gaps are limited action coverage and missing natural trigger synonyms like the skill's own name 'HQQ' and memory-reduction phrasing.

Suggestions

Add 1-2 more concrete capabilities to the 'what' portion, e.g. 'fine-tune quantized models with LoRA/PEFT' and 'select optimized inference backends (Marlin, TorchAO, BitBlas)'.

Include the skill's own name and common synonyms as trigger terms: 'HQQ', 'model compression', 'reduce/shrink model memory'.

Sharpen distinctiveness by signaling when to prefer it over alternatives, e.g. 'use instead of GPTQ/AWQ when no calibration data is available'.

DimensionReasoningScore

Specificity

The description names the domain ('Half-Quadratic Quantization for LLMs') and one concrete action ('quantizing models to 4/3/2-bit precision'), but coverage is not comprehensive — capabilities shown in the body such as backend selection, LoRA/PEFT fine-tuning, and saving/deploying quantized models are absent. It matches anchor 3 (domain plus 1-2 concrete actions); not 4 because several specific actions from the skill are missing, not 2 because the precision levels and calibration-free property are concrete.

3 / 5

Completeness

Both parts are explicit: what it does ('Half-Quadratic Quantization for LLMs without calibration data') and when to use it ('Use when quantizing models to 4/3/2-bit precision without needing calibration datasets, for fast quantization workflows, or when deploying with vLLM or HuggingFace Transformers'). This matches the anchor-5 example structure with a literal 'Use when...' clause and concrete trigger phrases.

5 / 5

Trigger Term Quality

Natural user phrases are present: 'quantizing models', '4/3/2-bit precision', 'calibration datasets', 'fast quantization', 'vLLM', 'HuggingFace Transformers'. Not 5 because common variations a user would say are missing — the skill's own name 'HQQ', 'model compression', 'reduce memory usage', 'shrink model' — and no file/format-style shorthand.

4 / 5

Distinctiveness Conflict Risk

The calibration-free angle ('without needing calibration datasets') is a genuinely distinctive trigger vs. GPTQ/AWQ skills, and precision levels plus vLLM/HF deployment are specific. Not 5 because generic quantization requests ('quantize this model to 4-bit') could plausibly match several quantization skills, since neither 'HQQ' nor an explicit differentiator from calibration-based methods appears as a trigger term.

4 / 5

Total

16

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.