CtrlK
BlogDocsLog inGet started
Tessl Logo

picking-a-format

Use when choosing an output format for extracted documents — plain text, markdown, djot, HTML, JSON, or DocTags. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.

66

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Picking a format

Xberg has two orthogonal format knobs. Get them right up front and the downstream code stays simple.

KnobWhat it controlsValuesDefault
--formatHow the CLI prints the resulttext, json, toontext (extract), json (batch)
--content-formatHow extracted content is rendered inside resultplain, markdown, djot, html, json, doctagsplain
--token-reductionStrip whitespace / boilerplate for LLM contextsoff, light, moderate, aggressive, maximumoff

--format json returns an envelope wrapping the ExtractedDocument — the document lives under .result for extract and under .results[] for batch, with content, metadata, tables, and images as fields of that nested document. --format text prints just content. --content-format is what shows up inside that content field.

Decision tree

Who consumes the output?
├── LLM (Claude, GPT, Gemini, local) — embed/prompt context
│       --format text --content-format markdown
├── Vector store / RAG indexer
│       --format json --content-format markdown
│       (markdown preserves structure for chunking)
├── Downstream parser that expects machine-readable JSON
│       --format json --content-format plain
│       (cleanest text + structured metadata)
├── Human review / archival
│       --format text --content-format markdown
├── HTML re-rendering / web display
│       --format json --content-format html
├── Lossless intermediate for pandoc / academic tooling
│       --format json --content-format djot
└── Token-budget-constrained pipeline
        --format text --content-format plain
        (drops markup; add --token-reduction moderate for further savings)

Examples

Feed a PDF directly into an LLM:

xberg extract paper.pdf --content-format markdown

Index a corpus into a RAG store with tables and headings preserved:

xberg batch docs/*.pdf --format json --content-format markdown \
  | jq -c '.results[] | {content: .content, tables: .tables}'

Strip a file to bare text for a token-tight summarizer:

xberg extract long.pdf \
  --content-format plain \
  --token-reduction moderate

Pull metadata only, ignore content:

xberg extract file.pdf --format json | jq '.result.metadata'

When in doubt

  • Default to markdown as the content format. It is the best compromise across LLMs, RAG, and human review, and Xberg has the most faithful renderer for it.
  • Reach for plain only when downstream cannot tolerate any markup.
  • Reach for djot only if you're already in a djot/pandoc pipeline.
  • Reach for html only when re-rendering for the web.
  • Reach for json for a heading-driven content tree, or doctags for Docling-compatible output.

Token-reduction (orthogonal)

--token-reduction collapses whitespace, strips repeated headers/footers, and trims boilerplate. It composes with any --content-format:

  • off (default), light, moderate, aggressive, maximum.

Use moderate as a safe starting point for LLM context windows. maximum is lossy — verify before relying on it.

See references/cli-reference.md for the full flag set and references/configuration.md for the equivalent output_format and token_reduction keys in xberg.toml.

Repository
xberg-io/xberg
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.