The real problem: the answer is in the document, not in the model

No language model knows what is in your company contract, in your patient report or in the internal policy you published last month. It knows language and it can reason about text you hand it. When someone asks and the model answers without having received the document, it fills the gap with what is plausible — and plausible is not true.

RAG, retrieval-augmented generation, is the design that fixes this: before answering, the system retrieves the relevant passages from your documents and hands them to the model along with the question. The answer becomes grounded in your material, with a citable source. The term comes from a 2020 paper by Lewis and others, and it became the de facto standard for any application that must answer from a private corpus.

Where RAG projects actually fail

Public discussion of RAG revolves around vector databases and model choice. In practice the most common failure sits one step earlier: reading the document. A PDF is not text — it is a set of drawing instructions. Poor extraction means losing structure: the table becomes a continuous line of numbers with no headers, the footnote lands in the middle of a paragraph, the page header repeats on every page, and side-by-side columns interleave.

After that, no model choice saves you. If the retrieved passage says "12 24 36 48" without saying those are months and amounts, the answer will be wrong with full confidence. In Brazilian corporate documents — reports, contracts, exported spreadsheets, numbered policies — the information that matters is almost always in a table or in structure.

The second most common failure is chunking. Cutting the document into fixed-size pieces separates the question from the answer: the condition lands in one chunk and the exception in another. Chunking that respects sections, headings and tables fixes a good share of the errors without changing anything else in the architecture.

What Docling changes

Docling is an open-source project born out of IBM Research and now hosted by LF AI & Data. It handles the step that usually gets underestimated: converting PDF, DOCX, PPTX, spreadsheets and images into a representation that preserves document structure — headings, sections, lists, reading order and, above all, tables as tables.

Three things make it suitable for enterprises. First, it runs locally: the document does not have to leave the client infrastructure to be read, which answers half the legal objections in healthcare, legal and the public sector. Second, it produces an intermediate format with structure preserved, from which you can generate Markdown or JSON for chunking. Third, it treats table recognition and scanned-document reading as first-class problems rather than extras.

The practical effect on a project: the same question, the same model and the same vector database start answering correctly because the retrieved passage finally contains the complete information. It is the best return per hour invested in a RAG project.

Anonymise before sending: Presidio and the LGPD

Brazilian corporate documents are full of personal data: names, tax IDs, addresses, phone numbers, medical records. Sending that to an AI service untreated is a contractual and an LGPD problem. The answer is not to abandon the project, it is to separate what must travel identified from what does not.

Microsoft Presidio, also open source, recognises and replaces personal data in text before it is sent, with recognisers you can tune for Brazilian formats. The design JBKR uses is: Docling reads and preserves structure, Presidio anonymises whatever is identifiable, and only then does the passage go to the model — with the result re-associated locally when needed.

Worth stating the obvious: anonymising does not replace a legal basis or a contract. It does remove the unnecessary risk of sending out what never needed to leave.

How to measure whether it works

RAG without evaluation is a demo, not a product. The minimum evaluation has two parts. Retrieval: given a set of real questions, did the right passage appear among those retrieved? If it did not, the problem is reading or chunking, and swapping models will not help. Answer: is the final answer correct and supported by the cited passage?

Build 30 to 50 real questions with known answers, taken from what people already ask today. That set is the most valuable asset in the project: it lets you compare versions, justify replacing a component and detect regressions when documents change. Log the questions with no good answer too — they tell you what the corpus is missing.

And require source citation in the interface. An answer with a link to the document and page turns the user into a reviewer; without it, errors go unnoticed until they become wrong decisions.

The JBKR clinical case

JBKR delivered a diagnostic-support assistant that answers clinical questions from the institution’s own documents — reports, protocols and literature. The design is the one described here: Docling for structure-preserving reading, Presidio for anonymisation, retrieval with source citation, and mandatory human review before any clinical use.

The lesson applies to any sector: the gain did not come from a better model, it came from reading the documents properly and measuring retrieval with real questions before growing the corpus.

Where to start

Start with a small, heavily consulted set — the policy everyone asks about, the manual that generates tickets, the standard contract. Convert it with Docling, build 30 real questions, measure retrieval, and only then choose the rest of the architecture. A large corpus is the last step, not the first.

At JBKR the initial 30-minute assessment is free and, on this subject, usually ends with an uncomfortable and useful conclusion: the problem is not AI, it is how the documents are stored.

Sources

  1. Lewis et al. (2020) — Retrieval-Augmented Generation
  2. Docling — project repository
  3. Docling Technical Report (IBM Research)
  4. LF AI & Data — Docling project
  5. Microsoft Presidio — personal data anonymisation
  6. LGPD — Brazilian Data Protection Law 13.709/2018
  7. JBKR — case: clinical RAG with Docling and Presidio

See the related service →

Jean C Becker

Jean C Becker

Senior Solutions Architect | AI & Machine Learning Specialist

Founder of JBKR and creator of AironCore and Apicio. Senior solutions architect and AI/ML specialist — RAG, computer vision, agents and LLMOps — on top of 20 years of web, mobile and systems engineering. About JBKR →