Use when splitting extracted text into chunks for LLM context windows or RAG ingestion. Covers chunk size, overlap, markdown/yaml/semantic chunkers, tokenizer-based sizing, and the standalone `chunk` command.
69
85%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Use this when feeding documents into an LLM context window or a vector
store. Xberg chunks two ways: inline during extraction (chunks land on
each document's chunks field), or standalone via the chunk command for text you
already have. Sizing is character-based by default, or token-based when a
tokenizer model is supplied.
Turn on chunking with --chunk and the chunks appear on the structured
result under chunks:
# 1000-char chunks, 200-char overlap (defaults when --chunk is on)
xberg extract report.pdf --chunk --format json | jq '.chunks | length'
# Explicit size + overlap
xberg extract report.pdf --chunk --chunk-size 1500 --chunk-overlap 300 --format jsonOverlap must be smaller than chunk size — the CLI rejects
--chunk-overlap >= --chunk-size. When you set only --chunk-overlap
against an existing config, an overlap that exceeds the size is clamped to
chunk_size / 4.
chunk commandChunk text you already have, from --text or stdin. Output defaults to
JSON:
# From a flag
xberg chunk --text "long document text ..." --chunk-size 800 --chunk-overlap 100
# From stdin (pipe extracted content straight in)
xberg extract notes.md | xberg chunk --chunk-size 500 --format jsonJSON output carries chunks (array of strings), chunk_count, the
resolved config (max_characters, overlap, chunker_type), and
input_size_bytes. Use --format text for a human-readable dump with
--- chunk N --- separators.
Note: in the JSON output,
chunker_typeis rendered capitalized ("Text","Markdown","Yaml","Semantic") because it is emitted via Rust's Debug formatting, whereas the--chunker-typeinput flag is lowercase (text,markdown,yaml,semantic). Lowercase the value before comparing if you parse it back.
--chunker-type selects the splitting strategy (standalone chunk
command):
| Type | Behavior |
|---|---|
text | Default. Plain character-window splitting with overlap. |
markdown | Markdown-aware — splits on structure (headings, blocks) where possible. |
yaml | YAML-aware splitting for structured config/data documents. |
semantic | Topic-boundary splitting driven by --topic-threshold (0.0–1.0, default 0.75). |
# Markdown-aware chunking keeps headings and blocks intact
xberg chunk --text "$(cat README.md)" --chunker-type markdown
# Semantic chunking — lower threshold = more, smaller topic chunks
xberg chunk --text "$(cat transcript.txt)" --chunker-type semantic --topic-threshold 0.6By default --chunk-size counts characters. To size chunks by tokens for
a specific model, pass --chunking-tokenizer with a HuggingFace tokenizer
id. On the extract command this implicitly enables chunking. Requires the
chunking-tokenizers feature (present in the default CLI build).
# Size chunks by GPT-4o tokens during extraction
xberg extract report.pdf --chunking-tokenizer Xenova/gpt-4o --format json
# Or on the standalone command
xberg chunk --text "$(cat doc.txt)" --chunking-tokenizer Xenova/gpt-4o --chunk-size 512With a tokenizer set, --chunk-size is interpreted in tokens, not
characters.
Field names in config files are snake_case under [chunking]:
[chunking]
max_characters = 1000
overlap = 200
chunker_type = "markdown"xberg extract report.pdf --config xberg.toml --format jsonCLI flags map to config fields as
--chunk-size→max_charactersand--chunk-overlap→overlap. In config files use the snake_case names.
From Python, enable chunking on the config and read the chunks off the
document in the result envelope (result.results[0].chunks):
from xberg import ExtractInput, extract, ExtractionConfig, ChunkingConfig
config = ExtractionConfig(
chunking=ChunkingConfig(max_characters=1000, overlap=200),
)
result = await extract(ExtractInput(uri="report.pdf"), config)
for chunk in result.results[0].chunks or []:
print(len(chunk.content))The public Python
ChunkingConfig(a dataclass) uses constructor kwargsmax_characters/overlap; the Rust core struct fields are alsomax_characters/overlap. TOML/JSON config keys aremax_chars/max_overlap(withmax_characters/overlapaccepted as serde aliases), and dict-form config passed toExtractionConfiglikewise accepts themax_chars/max_overlapaliases; Node'sChunkingConfiginterface usesmaxCharacters/overlap. Seereferences/python-api.mdandreferences/rust-api.mdin the siblingxbergskill.
markdown chunking for docs to keep sections whole.semantic chunker; tune --topic-threshold
down for finer splits, up for coarser ones.extract; clamped to size / 4 when
only overlap is changed against an existing config.--chunking-tokenizer errors if the
CLI was built without chunking-tokenizers. The default build includes it.chunk command bails on empty text;
provide --text or pipe non-empty stdin.See references/configuration.md for the full [chunking] schema and
references/cli-reference.md for every chunk flag.
04336bd
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.