Xberg v1 Released: A Rust Content Intelligence Framework Supporting 468 Formats
Goldziher · reddit · 2026-08-02
Content intelligence framework Xberg v1 (successor to Kreuzberg) has been officially released. It focuses on high performance and broad format support, handling 101 document formats and 367 code/data formats.
Key Features & Upgrades:
- Engine & Parsing: Replaces pdfium with a pure-Rust PDF backend (pdfoxide), integrates ONNX layout detection for reading order reconstruction, and supports native PaddleOCR alongside Tesseract.
- Multimodal Processing: Features audio/video transcription (Whisper ONNX engine), URL ingestion (JS-rendered), and a pure-Rust Candle stack for OCR/VLM inference.
- Retrieval & Intelligence: Includes SPLADE sparse embeddings, ColBERT retrieval, cross-encoder reranking, NER (GLiNER2), and structured LLM extraction.
- Cross-Platform: Offers 15 language bindings and enables ONNX inference via tract without ONNX Runtime, bringing full support for WASM (in-browser) and mobile (iOS/Android).
Related event: Xberg Launches Local Document Extraction Framework(4 posts)→
More from coding & agent
- Quickly Turn Any Website into an API or MCP Server Using DevTools — dsp_ · 2026-08-04
- Anthropic Engineers Demo Multi-Agent Loops: Full App Built from Scratch in 40 Mins — mathemagie · 2026-08-04
- Running Agents in Prod for a Year: Auditability Trumps Model Capability — KimLikeJ · 2026-08-04
- Open-Source 'The Librarian' MCP Server Manages 3,000 Skills, Saves Tokens and Self-Optimizes — Open-Appeal-9747 · 2026-08-04
- DSPy Launches Experimental Flex Feature for Automated Code and Control Flow Optimization — lateinteraction · 2026-08-04
- OpenAI Hints Next-Gen Models Need More Compute, Codex May Shift to Cloud Agents — haider1 · 2026-08-04