CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — system_design 67 upvotes

Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion

QUESTION — How can heterogeneous enterprise document formats (PDFs, Word, scans) be converted and chunked into retrieval-optimized Markdown while minimizing token costs and processing time?

The authors present Document Retrieval-Aware Chunking (D-RAC), a framework extending enterprise document ingestion by normalizing any input document into PDF to leverage deterministic rendering. A single multimodal LLM pass converts the rendered pages into retrieval-optimized Markdown, transforming tables into self-contained prose statements while preserving heading hierarchy. Parsing then proceeds deterministically into ID-addressable units followed by lightweight LLM-based chunk planning over identifiers. On a 236-document, 795-page PDF benchmark subset, D-RAC processes the corpus error-free, reduces chunking-stage output tokens by 95.7%, and cuts chunking costs by 77.8% to 85.6% compared to agentic chunking.

On the 236-document, 795-page PDF subset of the RAG-Multi-Corpus benchmark, D-RAC converts and chunks the entire corpus in 72 minutes with zero errors, producing 1,748 retrieval-ready chunks.

D-RAC reduces chunking-stage output tokens by 95.7% compared to agentic chunking with frontier LLMs.

It cuts chunking cost by 77.8% (GPT-4.1 pricing) to 85.6% (Gemini 2.5 Pro pricing) and chunking time by 75%.

udayallu · 21 Sept 2026 read the original ↗
↑