CtrlK
BlogDocsLog inGet started
Tessl Logo

chunking

Use when splitting extracted text into chunks for LLM context windows or RAG ingestion. Covers chunk size, overlap, markdown/yaml/semantic chunkers, tokenizer-based sizing, and the standalone `chunk` command.

70

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is chunking in xberg-io/xberg

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, highly executable reference: realistic CLI examples for every mode, a decision table for chunker types, parameter guidance per use case, and valuable non-obvious gotchas. Its main structural weakness is progressive disclosure — the body delegates to `references/configuration.md` and `references/cli-reference.md` that are not present in this bundle, leaving broken navigation.

Suggestions

Ship the referenced files (`references/configuration.md`, `references/cli-reference.md`) in this skill's bundle, or inline the essential `[chunking]` schema and chunk-flag details and drop the dangling references.

State the overlap-vs-size rule once — either in 'Common pitfalls' or in 'Inline during extraction' — instead of repeating it nearly verbatim in both places.

Move the long cross-language field-name alias blockquote (Python/Rust/TOML/Node naming differences) into a reference file (e.g. `references/python-api.md` or a naming section of `references/configuration.md`) so the SKILL.md body stays a lean overview.

DimensionReasoningScore

Conciseness

The body is dense with runnable commands and non-obvious gotchas (capitalized `chunker_type` in JSON output, cross-language field-name aliases) and explains nothing Claude already knows. Not 5 because of minor trimmable redundancy: the overlap-must-be-smaller-than-size rule is stated in full in 'Inline during extraction' and then restated in 'Common pitfalls' ("Overlap ≥ size — rejected on `extract`; clamped to size / 4..."), and the alias blockquote in Programmatic access runs long.

4 / 5

Actionability

Every section carries copy-paste-ready commands with realistic flags and comments — `xberg extract report.pdf --chunk --format json | jq '.chunks | length'`, `xberg chunk --text ... --chunker-type semantic --topic-threshold 0.6`, a runnable async Python snippet, and a TOML config block — plus a documented output schema (`chunks`, `chunk_count`, `config`, `input_size_bytes`). Fully executable with specific examples covering the common cases, matching anchor 5.

5 / 5

Workflow Clarity

This is a single-task reference skill with no multi-step process, and the single action is unambiguous: the choice between inline (`--chunk` on extract) and standalone (`chunk` command) is laid out up front, the chunker-type table plus 'Picking parameters' section gives decision rules per use case (RAG, summarization, topic segmentation), and error behavior is documented (rejected overlap, empty-input bail, missing-feature error). Under the simple-skill exception this qualifies for 5; there are no destructive or batch operations that would require a validation checkpoint.

5 / 5

Progressive Disclosure

The in-body structure is good and references are clearly signaled ("See `references/configuration.md` for the full `[chunking]` schema and `references/cli-reference.md` for every chunk flag"), but this bundle contains no `references/`, `scripts/`, or `assets/` directories — the two referenced files do not exist here, so navigation dead-ends. Dangling references plus some inline material that belongs in those files (the full cross-language alias map, the pitfalls list) fit anchor 3 ('references present but not clearly signaled' in effect: they are signaled but not actually shipped) better than anchor 4's 'references mostly clear'.

3 / 5

Total

17

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: explicit third-person 'Use when' trigger clause paired with a concrete, comprehensive capability list covering chunker types, sizing modes, and the standalone command. The only gaps are a few missing natural synonyms (vector store, embeddings) and mild overlap risk with sibling extraction/RAG skills.

DimensionReasoningScore

Specificity

"Covers chunk size, overlap, markdown/yaml/semantic chunkers, tokenizer-based sizing, and the standalone `chunk` command" lists multiple specific concrete capabilities with comprehensive coverage of the skill's scope, matching the anchor-5 example's breadth. It goes beyond the anchor-4 example ('Extracts text from PDF files, fills forms, converts pages to images') by enumerating the full feature surface of the domain rather than leaving minor gaps.

5 / 5

Completeness

Both questions answered explicitly: what — "chunk size, overlap, markdown/yaml/semantic chunkers, tokenizer-based sizing, and the standalone `chunk` command"; when — "Use when splitting extracted text into chunks for LLM context windows or RAG ingestion". This mirrors the anchor-5 example's structure of a concrete capability list followed by an explicit 'Use when' trigger clause, so it cannot be the 4 ('when' could be more explicit or specific), since the when clause names concrete trigger scenarios.

5 / 5

Trigger Term Quality

Natural terms present: "chunks", "splitting extracted text", "RAG ingestion", "LLM context windows", "chunk size", "overlap", "chunk command" — these are phrases users would actually say. Not anchor 5 because common synonyms like "vector store", "embeddings", or "split text" are absent, leaving a few natural terms missing per the anchor-4 definition.

4 / 5

Distinctiveness Conflict Risk

The chunking/RAG-ingestion niche is distinct with triggers unlikely to fire for unrelated skills, but "splitting extracted text" and "RAG ingestion" could mildly overlap with a sibling extraction skill or a broader RAG-pipeline skill (the body itself notes chunks can be produced inline during extraction). Minor overlap risk with closely related skills matches anchor 4 rather than anchor 5's 'minimal conflict risk'.

4 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 4 missing

Warning

Total

15

/

16

Passed

Repository
xberg-io/xberg
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.