CtrlK
BlogDocsLog inGet started
Tessl Logo

xberg

github.com/xberg-io/xberg

SkillAddedReview
extracting-with-ocr

plugin/.ai-rulez/skills/extracting-with-ocr/SKILL.md

Use when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer. Covers OCR backends, language packs, force-OCR, and performance tuning.

69

extracting-with-ocr

plugin/.cursor-plugin/skills/extracting-with-ocr/SKILL.md

Use when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer. Covers OCR backends, language packs, force-OCR, and performance tuning.

—

—
extracting-with-ocr

plugin/.hermes/package/src/xberg_hermes_plugin/skills/extracting-with-ocr/SKILL.md

Use when extracting text from scanned PDFs, photographed pages, or images that have no embedded text layer. Covers OCR backends, language packs, force-OCR, and performance tuning.

72

extraction-pipeline-patterns

.ai-rulez/skills/extraction-pipeline-patterns/SKILL.md

Change or diagnose Xberg's core extraction orchestration, cache semantics, extractor fallback, post-processing, concurrency defaults, or format-wide quality invariants. Load for pipeline work, not a single parser's syntax.

64

feature-flag-policy

.ai-rulez/skills/feature-flag-policy/SKILL.md

Cargo feature flags for crates/xberg — ORT-incompatible targets (WASM, Android x86_64 emulator), type-only and tract inference companion features, WASM/Android-safe variants, PDF backend, mutually-exclusive ORT variants, platform-conditional deps, aggregate feature sets, and build profiles. Load when adding, wiring, or debugging a Cargo feature, or when reasoning about what compiles on WASM/Android/Windows/macOS-intel targets.

71

format-specific-extraction

.ai-rulez/skills/format-specific-extraction/SKILL.md

Format-specific document extraction workflows

52

mime-detection-routing

.ai-rulez/skills/mime-detection-routing/SKILL.md

MIME type detection and extractor routing in core/mime.rs — the FORMATS registry that EXT_TO_MIME and SUPPORTED_MIME_TYPES are derived from, the path-based and bytes-based detection functions, priority-based registry selection, wildcard MIME families, and the real procedure for adding a format. Load when adding a format, wiring an extractor to a MIME type, or debugging why a file routes to the wrong (or no) extractor.

71

ocr-pipeline-and-quality

.ai-rulez/skills/ocr-pipeline-and-quality/SKILL.md

Change or evaluate Xberg OCR backends, preprocessing, caching, page acceptance, geometry, hOCR structure, table reconstruction, or cross-backend quality. Load for OCR behavior and A/B quality work, not ordinary PDF text extraction.

—

—
pdf-backends

.ai-rulez/skills/pdf-backends/SKILL.md

Change or diagnose Xberg PDF extraction, native/Pdfium backend selection, PDF rendering sessions, encrypted documents, OCR fallback, or backend-specific capability gaps. Load for PDF engine work, not generic image OCR.

72

picking-a-format

plugin/.ai-rulez/skills/picking-a-format/SKILL.md

Use when choosing an output format for extracted documents — plain text, markdown, djot, HTML, JSON, or DocTags. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.

66

picking-a-format

plugin/.hermes/package/src/xberg_hermes_plugin/skills/picking-a-format/SKILL.md

Use when choosing an output format for extracted documents — plain text, markdown, djot, HTML, JSON, or DocTags. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.

72

picking-a-format

plugin/.cursor-plugin/skills/picking-a-format/SKILL.md

Use when choosing an output format for extracted documents — plain text, markdown, djot, HTML, JSON, or DocTags. Maps consumer (LLM, parser, archive) to the right `--format` / `--content-format` pair.

—

—
plugin-architecture-patterns

.ai-rulez/skills/plugin-architecture-patterns/SKILL.md

Design, implement, or diagnose Xberg plugin traits, typed registries, priority collisions, lifecycle, native extractors, and Alef-generated Python plugin bridges. Load for plugin-system work, not ordinary extractor parsing.

66

polyrepo-boundaries

.ai-rulez/skills/polyrepo-boundaries/SKILL.md

Decide which xberg-io repository owns a cross-repository fix or API, and coordinate compatible changes across Xberg, Alef, enterprise, crawler, LLM, and OCR repositories. Load when work spans sibling repos; do not use for a self-contained Xberg edit.

71

release-readiness

.ai-rulez/skills/release-readiness/SKILL.md

Audit Xberg before a push or release by reconciling CI, Publish Release dry-run, Benchmarks, generated freshness, changelog, and remote branch state. Load for release-readiness work, not routine local commits.

75

release-versioning

.ai-rulez/skills/release-versioning/SKILL.md

How xberg versions are synced and released — Cargo.toml is the single source of truth, `task version:sync` propagates it to alef-managed binding manifests AND the integrations under integrations/, which are versioned and published in lockstep with core (including -rc.N). Load before bumping a version, editing the version-sync task, or touching an integration's version/xberg dependency.

73

test-corpus

.ai-rulez/skills/test-corpus/SKILL.md

The test_documents submodule is a bucket-fetched fixture corpus that is not committed. This skill covers read_test_fixture, missing fixtures, valid A/B controls, and submodule push order. Load before running Rust tests on a fresh clone, setting up an A/B control, adding a fixture-backed test, or diagnosing missing-fixture failures.

67

wasm-constraints

.ai-rulez/skills/wasm-constraints/SKILL.md

WASM build constraints for the crates/xberg-wasm crate — the wasm-target feature set, no-tokio sync-only internal APIs, the crate-private SyncExtractor trait, the 2 MB HTML size limit, size-optimized build config (opt-level="z"), and the async-wrapper/sync-internal API pattern. Load when building for wasm32, adding or modifying a WASM-compatible extractor, or debugging WASM build/runtime failures.

70

xberg

plugin/.ai-rulez/skills/xberg/SKILL.md

Extract text, tables, metadata, and images from 107 document formats (PDF, Office, images, HTML, email, archives, academic) using Xberg. Use when writing code that calls Xberg APIs in Python, Node.js/TypeScript, Rust, or CLI. Covers installation, extraction (sync/async), configuration (OCR, chunking, output format), batch processing, error handling, and plugins.

—

—
xberg

plugin/.hermes/package/src/xberg_hermes_plugin/skills/xberg/SKILL.md

Extract text, tables, metadata, and images from 107 document formats (PDF, Office, images, HTML, email, archives, academic) using Xberg. Use when writing code that calls Xberg APIs in Python, Node.js/TypeScript, Rust, or CLI. Covers installation, extraction (sync/async), configuration (OCR, chunking, output format), batch processing, error handling, and plugins.

72