MIME type detection and extractor routing in core/mime.rs — the FORMATS registry that EXT_TO_MIME and SUPPORTED_MIME_TYPES are derived from, the path-based and bytes-based detection functions, priority-based registry selection, wildcard MIME families, and the real procedure for adding a format. Load when adding a format, wiring an extractor to a MIME type, or debugging why a file routes to the wrong (or no) extractor.
69
85%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Policy -> content and/or extension evidence -> validate_mime_type -> registry.get(mime) -> extractor| Function | Location | Behaviour |
|---|---|---|
detect_mime_type(path, check_exists: bool) | core/mime.rs | Path-based only — never reads bytes. Lowercased extension → EXT_TO_MIME, then tree-sitter extension detection (feature tree-sitter), then mime_guess::from_path. check_exists gates a file-existence check, not content inspection. |
detect_mime_type_from_bytes(bytes) | core/mime.rs | Magic-number detection via the infer crate. The only content-sniffing entry point. |
validate_mime_type(mime) | core/mime.rs | Parses the media type, matches its case-insensitive essence against SUPPORTED_MIME_TYPES, and returns the registered MIME spelling. Parameters such as charset do not affect extractor routing. It does not consult the extractor registry. |
FORMATS: &[FormatEntry { extensions, mime_type, aliases }] in core/mime.rs.
EXT_TO_MIME and SUPPORTED_MIME_TYPES are LazyLocks derived from it by iteration —
there is no m.insert call site to add to, and hand-editing either is impossible.
The full registry publishes 106 formats, 140 unique extensions, and 53 aliases, verified by
scripts/sync_supported_counts.py verify. The published count constants describe that static
registry; runtime availability is its intersection with registered extractors. Extension lookup is
case-insensitive (the extension is lowercased before the map hit).
let registry = get_document_extractor_registry(); // plugins/registry/mod.rs
let guard = registry.read()?;
let extractor: Arc<dyn DocumentExtractor> = guard.get(mime_type)?; // Result, not OptionDocumentExtractorRegistry::get (plugins/registry/extractor.rs) returns the
highest-priority() extractor for the MIME type, and returns Err — not None — when none
matches.
An extractor may register a family: "image/*" matches image/png, image/jpeg, and so on
(prefix match on a registered type ending in /*).
FormatEntry to FORMATS in crates/xberg/src/core/mime.rs. EXT_TO_MIME and
SUPPORTED_MIME_TYPES update automatically.scripts/sync_supported_counts.py sync to update published count claims, then run its
verify command.InternalDocumentExtractor (not DocumentExtractor — see
plugin-architecture-patterns) with supported_mime_types() returning the MIME.crates/xberg/src/extractors/mod.rs::register_default_extractors().validate_mime_type() before extraction — but do not treat it as proof an extractor exists.detect_mime_type inspects no content. Extraction defaults to PreferContent, which performs
bounded content inspection and falls back to a supported extension. Use ContentOnly when the
filename must be ignored; use TrustExtension only for trusted sources.application/octet-stream is the exception: it
is a generic placeholder and triggers policy-based detection.EXT_TO_MIME or SUPPORTED_MIME_TYPES — edit FORMATS.04336bd
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.