Use when extracting tabular data from PDFs, spreadsheets, or images. Covers layout-aware table detection, table model selection, output formats (markdown / JSON cells), and known limits.
65
78%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Fix and improve this skill with Tessl
tessl review fix ./plugin/.ai-rulez/skills/extracting-tables/SKILL.mdUse this when the user wants structured tabular data — financial statements, scientific tables, invoices, spreadsheet-style PDFs. Xberg detects tables via a layout model (RT-DETR v2) and reconstructs cell structure with a configurable table model.
# Markdown tables embedded in the content stream
xberg extract report.pdf --layout --content-format markdown
# Structured JSON output, tables appear under result.tables
xberg extract report.pdf --layout --format json--layout turns on layout-aware extraction; without it, tables fall back
to plain text reflow and you lose cell boundaries.
Two surfaces, picked via --format (CLI shape) and --content-format
(content rendering):
content — --content-format markdown. Tables
appear inline as | col | col | blocks. Good for LLM ingestion.tables array — --format json. Each entry has
cells[][] (rows × cols), markdown (pre-rendered), page_number,
bounding_box. Use this when downstream code needs exact cell access.
(bounding_box is omitted when no position data is available.)Both are populated at once when --layout is on. The tables array is
always structured; the content stream switches representation.
xberg extract financials.pdf --layout --format json \
| jq '.result.tables[] | {page: .page_number, rows: (.cells | length)}'--layout-table-model picks the reconstruction backend:
| Model | Best for | Notes |
|---|---|---|
tatr | dense complex tables (academic, financial) | Default. Heaviest, highest accuracy. |
slanet_auto | dispatches per-table to wired/wireless | Good when table styles are mixed. |
slanet_wired | tables with visible borders | Faster than tatr. |
slanet_wireless | tables without borders (whitespace-separated) | For invoices, simple grids. |
slanet_plus | hybrid wired / wireless | Lighter than slanet_auto. |
disabled | layout detection only, no table structure | Use to skip table model cost. |
xberg extract bank-statement.pdf \
--layout --layout-table-model tatr --content-format markdownDrop --layout-confidence when the layout model misses tables (default
threshold ~0.5):
xberg extract noisy-scan.pdf --layout --layout-confidence 0.3.xlsx, .ods, .csv, .tsv are extracted by dedicated parsers — no
layout model needed. Each sheet becomes a markdown table (or structured
table) automatically:
xberg extract workbook.xlsx --content-format markdown
xberg extract data.csv --format jsonPass --no-cache=true only when iterating on the same file with different
configs.
# `output_format` in config files equals `--content-format` on the CLI.
output_format = "markdown"
[layout]
confidence_threshold = 0.5
table_model = "tatr"Then:
xberg extract report.pdf --format jsonFrom Python, structured tables live on the document in the result envelope
(result.results[0].tables):
from xberg import ExtractInput, extract, ExtractionConfig, LayoutDetectionConfig
config = ExtractionConfig(
layout=LayoutDetectionConfig(table_model="tatr"),
output_format="markdown",
)
result = await extract(ExtractInput(uri="report.pdf"), config)
for table in result.results[0].tables:
print(table.markdown) # rendered markdown
print(table.cells[0][0]) # cell accessNode.js mirrors this (extract, output.results[0].tables, camelCase fields).
See references/python-api.md and references/nodejs-api.md in the
sibling xberg skill for full type signatures.
--ocr-auto-rotate true for image-based
PDFs before extraction.tables[] entry.
Stitch by matching column headers if needed.tables with --layout on — confidence threshold too high or
table model mismatched. Drop --layout-confidence to 0.3, try
--layout-table-model tatr.--layout-table-model to
slanet_wired for bordered grids or slanet_wireless for invoices.tatr is heavy. Use slanet_auto or
slanet_plus as a default; reach for tatr only when accuracy matters.See references/cli-reference.md for the full layout flag set and
references/advanced-features.md for the layout pipeline internals.
92b4f95
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.