CtrlK
BlogDocsLog inGet started
Tessl Logo

long-context

Extend context windows of transformer models using RoPE, YaRN, ALiBi, and position interpolation techniques. Use when processing long documents (32k-128k+ tokens), extending pre-trained models beyond original context limits, or implementing efficient positional encodings. Covers rotary embeddings, attention biases, interpolation methods, and extrapolation strategies for LLMs.

60

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/llm-tools/long-context/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

53%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A technically rich body with abundant concrete code, a useful method-comparison table, and a genuinely functional reference bundle. Its weaknesses are redundancy between the inline 'Core Concepts' theory and the reference files, several not-quite-executable snippets (missing imports, undefined variables, a broadcasting bug), and no validation step in the fine-tuning workflow.

Suggestions

Cut the 'Core Concepts' theory/formulas (they duplicate references/rope.md and extension_methods.md) down to one-line summaries that point at the bundle files.

Fix the executable gaps: import math in the ALiBi snippet, import torch.nn.functional as F, expand cos/sin to (1, 1, seq_len, head_dim) before applying rotary embeddings, and replace the undefined `fine_tune(...)` placeholder with the Trainer-based commands already shown.

Add an explicit ordered workflow with a validation checkpoint — e.g. after the 1000-step fine-tune, evaluate perplexity on long documents at the target context length before deploying.

DimensionReasoningScore

Conciseness

The body is mostly dense, concrete code, but the ~80-line 'Core Concepts' section explains well-known techniques Claude already knows ("Encodes absolute position via rotation matrix", "No positional embeddings added to tokens", plus mathematical formulas duplicated in references/rope.md), and 'Best Practices' pads with pseudo-code like `use_method = "ALiBi"` that carries little information. This fits 'mostly efficient but includes some unnecessary explanation or could be tightened' rather than the 'several unnecessary explanations' of 2, since most sections are code-dense.

3 / 5

Actionability

There is substantial concrete code (a full RotaryEmbedding implementation, rope_scaling configs with real model IDs, a Trainer setup, a vLLM deployment snippet), but several flagship snippets are not executable as written: the ALiBi example uses `math` without importing it and references an undefined `attn_scores`, `apply_rotary_pos_emb` broadcasts (batch, heads, seq, dim) tensors against (seq, dim) cos/sin which fails, `F.scaled_dot_product_attention` appears without `import torch.nn.functional as F`, and `fine_tune(model, ...)` is an undefined placeholder. That is 'some concrete guidance but incomplete; missing key details', not the 'minor gaps' of 4.

3 / 5

Workflow Clarity

The implied sequence (choose method → set rope_scaling → fine-tune → deploy) is present across sections and 'Choose the Right Method' plus 'Avoid Common Pitfalls' give decision guidance, but there is no explicit ordered workflow and no validation checkpoint (e.g. evaluate perplexity at the extended length before deploying) for a fine-tuning pipeline that can silently produce a degraded model. This matches 'sequence present but checkpoints missing or implicit'.

3 / 5

Progressive Disclosure

The body is well-sectioned and closes with a 'See Also' list pointing to three real, one-level-deep, self-contained reference files (references/rope.md, extension_methods.md, fine_tuning.md) with clear descriptions of what each contains. The gap keeping it below 5 is that the 'Core Concepts' section inlines theory and formulas that duplicate the reference files, content that clearly belongs in those bundles — 'most content is appropriately placed... minor organization gaps'.

4 / 5

Total

13

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, explicit 'Use when...' triggers, concrete token ranges, and full coverage of the domain's techniques. Its only weakness is that some natural trigger phrasings users might actually type ('long context', 'context length') are absent in favor of technique jargon.

Suggestions

Add natural-language synonyms such as 'long context' or 'extend context length' to the trigger clause so users who don't know the technique names still match.

Mention concrete model families (e.g. LLaMA, Mistral) in the trigger terms, since users typically ask to 'extend LLaMA's context'.

DimensionReasoningScore

Specificity

The description lists multiple specific concrete actions with comprehensive coverage of the domain: "Extend context windows of transformer models using RoPE, YaRN, ALiBi, and position interpolation techniques", plus "processing long documents (32k-128k+ tokens)" and "implementing efficient positional encodings". It names every major technique in the domain rather than stopping at 1-2 actions, which fits the 'comprehensive coverage' anchor and not the 'minor gaps' level 4.

5 / 5

Completeness

It explicitly answers both questions: the 'what' is "Extend context windows of transformer models using RoPE, YaRN, ALiBi, and position interpolation techniques" and the 'when' is a concrete "Use when processing long documents (32k-128k+ tokens), extending pre-trained models beyond original context limits, or implementing efficient positional encodings" — matching the anchor-5 example structure exactly rather than the weaker 'when' at level 4.

5 / 5

Trigger Term Quality

Good keyword coverage including technique names ("RoPE", "YaRN", "ALiBi", "position interpolation"), concrete token ranges ("32k-128k+ tokens"), and phrases like "extending pre-trained models beyond original context limits". A few natural user phrasings are missing — e.g. "long context", "context length", "extend context window" as a literal phrase — so it falls just below the comprehensive-synonyms anchor at 5.

4 / 5

Distinctiveness Conflict Risk

This is a clear niche (positional-encoding-based context extension) with distinct triggers — a user mentioning RoPE, YaRN, ALiBi, or extending context limits would unambiguously land here. The risk of overlapping with generic fine-tuning or LLM-ops skills is minimal, fitting the 'clear niche with distinct triggers' anchor rather than the 'minor overlap risk' level 4.

5 / 5

Total

19

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (546 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

12

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.