CtrlK
BlogDocsLog inGet started
Tessl Logo

awq-quantization

Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. Use when deploying large models (7B-70B) on limited GPU memory, when you need faster inference than GPTQ with better accuracy preservation, or for instruction-tuned and multimodal models. MLSys 2024 Best Paper Award winner.

61

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/awq/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

68%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Highly actionable reference content with executable code throughout and good decision guidance (when to use AWQ vs GPTQ vs bitsandbytes). The two structural flaws are the orphaned bundle files that are never referenced from the body, and the absence of any validation step for quantized-model output.

Suggestions

Replace the inlined 'Key insight' algorithm explanation and 'Common issues' section with clearly signaled one-level-deep links, e.g. 'Algorithm details: See [advanced-usage.md](references/advanced-usage.md)' and 'Full error/fix catalog: See [troubleshooting.md](references/troubleshooting.md)', so the existing bundle files are actually discoverable.

Add a validation checkpoint after 'model.save_quantized(...)': run a short generation or perplexity check against the FP16 baseline (the accuracy table already shows expected degradation of ~2-3%) so users can confirm quantization quality before deploying.

Trim duplication for token efficiency: the three benchmark tables and the salient-weights explanation (repeated in the body and advanced-usage.md) can be consolidated into the reference file, keeping only the decision-relevant comparison table in SKILL.md.

DimensionReasoningScore

Conciseness

The body is dense and table/code-driven with no padding explaining basics Claude already knows. Minor instances of over-explanation could be trimmed: the 'Key insight' salient-weights explanation is duplicated in references/advanced-usage.md, and three separate benchmark tables could be consolidated.

4 / 5

Actionability

Nearly every section provides complete, copy-paste-ready code: installation commands, loading pre-quantized models, quantization config with commented parameters, kernel backend variants, Transformers/vLLM integration, multi-GPU deployment, and custom calibration. The 'Common issues' section pairs errors with concrete fixes.

5 / 5

Workflow Clarity

The quick-start flow (install → load/configure → quantize → save) is clearly sequenced, but the quantize-your-own-model path has no validation checkpoint — no step to verify the quantized output (e.g., perplexity check or a test generation) before deployment. The 'Common issues' section is reactive rather than an explicit validate-fix-retry loop, matching anchor 3: sequence present but checkpoints missing.

3 / 5

Progressive Disclosure

The bundle contains references/advanced-usage.md and references/troubleshooting.md, but the body never links to either — the 'References' section lists only external URLs. Meanwhile content that belongs in those files is inlined: the 'Key insight' algorithm explanation duplicates advanced-usage.md and the 'Common issues' section duplicates troubleshooting.md. This matches anchor 2: content that clearly belongs in separate files is inlined while the references are buried/unmentioned.

2 / 5

Total

14

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with explicit what/when structure and concrete trigger conditions. Its weaknesses are the single-capability framing (no enumeration of distinct actions like quantizing, loading pre-quantized models, or serving with vLLM) and a second-person 'you' clause that slightly hurts specificity.

DimensionReasoningScore

Specificity

The description names the domain concretely ('Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss') but describes a single capability rather than listing several specific actions. The second-person phrasing 'when you need faster inference than GPTQ' triggers the voice penalty, reducing specificity by 1 from a base of 4.

3 / 5

Completeness

It clearly answers both 'what' (activation-aware weight quantization for 4-bit LLM compression) and 'when' with an explicit 'Use when deploying large models (7B-70B) on limited GPU memory, when you need faster inference than GPTQ with better accuracy preservation, or for instruction-tuned and multimodal models' clause containing concrete trigger phrases, matching the anchor-5 example.

5 / 5

Trigger Term Quality

Natural phrases like 'deploying large models (7B-70B) on limited GPU memory', 'faster inference', 'instruction-tuned and multimodal models', and the GPTQ comparison give good keyword coverage. Common synonyms like 'model compression', 'quantize', or file-format terms are missing, so it falls between anchor 3 and 5 but noticeably above the midpoint.

4 / 5

Distinctiveness Conflict Risk

The description carves a clear niche around a named technique and explicitly positions it against GPTQ, making it mostly distinct. The 'faster inference than GPTQ' phrasing creates minor overlap risk, as a user mentioning GPTQ could trigger this skill, so it fits anchor 4 rather than 5.

4 / 5

Total

16

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.