How you read a document decides what AI can know
Most AI systems cut documents into arbitrary chunks, and answers inherit the damage. DICER reads documents as they're built, so answers stay grounded.
Ask an AI system a question about a long document, and the quality of its answer is decided long before the question arrives. It’s decided when the document is cut into pieces.
Most retrieval systems split documents into fixed-size chunks: so many characters, so many tokens, with a little overlap. It’s simple, and for plain prose it mostly works. Enterprise documents aren’t plain prose.
What arbitrary chunking breaks
An invoice, a policy, a financial statement or a contract is built from very different kinds of content: tables, key-value fields, headings, paragraphs and footnotes, each with its own structure. Cut that document at arbitrary points and you get:
- Tables split mid-row, so a value ends up in a different chunk from its column header
- Fields separated from their labels, so “₹20,50,840” no longer says what it is the total of
- Clauses cut off from their headings, so an exclusion loses the section that gives it meaning
Every answer built on those chunks inherits the damage. The model may still sound confident, but it’s reasoning over fragments.
Reading documents as they’re built
DICER, Cogniquest’s Document Intelligence & Contextual Extraction Router, takes a different approach. It reads a document as it’s actually built. It recognizes each kind of content, handles each one in the right way, and keeps the layout and context around it.
- Content is kept as what it is. Tables stay tables, fields stay fields, and paragraphs keep their headings.
- Layout is understood. Structure, hierarchy and reading order travel with the content.
- Every unit is complete in its meaning. Each piece carries what it depends on, so it can be retrieved and reasoned over on its own.
Why it matters for decisions
When a decision depends on a document, the details are the decision: the line item, the tolerance, the exclusion, the date. Reliable retrieval and reasoning need complete, meaningful units to work on.
That’s why DICER sits underneath everything Cogniquest does with documents, from APWise reconciling invoices line by line, to FinWise spreading financial tables into a comparable series, to DocWise answering questions with every answer cited to the exact passage it came from.
It’s also why we can say something most AI systems can’t: an answer without support is never returned as an answer.
Explore how DICER fits into the platform on our platform page.