In a recent project, I asked an agent a simple question about a maintenance log: what was recorded for the F3 fiber at 1310nm? The correct answer was nothing. The cell was empty. The agent confidently replied, “46.1dB,” borrowing a value from the wrong cell because its extraction had flattened the table’s structure.
If you build with LLMs, you’ve probably run into something similar. Andrej Karpathy put it bluntly: “In my experience there are approx. one thousand different pdf converters that are all equally terrible for anything except the simplest documents.”
Yet the assumption persists that document understanding is solved. Feed the model your files, get structured data back. LLMs are remarkably good at reasoning over document content, but reliable document understanding takes more than a model. It takes a system around the model: deterministic perception to read the page, structure to preserve its meaning, grounding to tie outputs to evidence, and confidence to flag when the system may be wrong.

Why documents are harder than they look
For a machine, reading and understanding a business document is a stack of problems.
Start with the page. Real-world documents have stamps, watermarks, and bleed-through ink. Low-quality scans. Handwriting that borders on illegible. Signatures, curved text, text running in different directions, and tables with no cell borders.
Next comes language. Supporting the world’s documents means supporting hundreds of languages, each with its own characters and conventions, sometimes mixed in the same file. The characters compound it: a 1 and an l, a 0 and an O. Some Russian characters are nearly identical to their English counterparts, so the system has to decide from context which alphabet it’s even looking at.
Locale adds another layer of ambiguity. The same date can mean March 4 or April 3, depending on where it was written. Currency formats shift between markets. The same term carries different meanings in different regions. A system that reads global documents has to read them the way each locale wrote them.
These are document problems, not standalone LLM problems, and they get solved in one unglamorous edge case at a time.
LLMs don’t erase difficulty; they trade one failure mode for another
Then you have to account for reading order. Which line comes “next” on a page with three columns, a sidebar, and a footer? A structured document is text plus layout, and meaning often lives in the layout: which number sits under which column header, which box a label points at.
If you give an LLM only a linearized text version of a page, you throw away spatial relationships that may carry meaning. If you give it the page as an image, the layout survives, but visual question answering is harder and more expensive for LLMs than reasoning over text, and measurably more error-prone.
Foundation model teams are incentivized to meet or exceed benchmarks, not to ensure a 0 didn’t become an O in row 40 of a scanned table. That work falls to whoever builds the document understanding system. And some of the most critical data in documents are not predictable with context or world knowledge, for instance, serial numbers or MRZ code.
On their own, probabilistic outputs break enterprise workflows, and they’re worse for agents. An agent doesn’t simply display the wrong number. It acts on it. An agent paying an invoice off a misread total doesn’t produce a typo. It produces an incorrect wire transfer.
A decade of unglamorous work
Our story begins 10 years before anyone was pointing LLMs at PDFs, when the problem was teaching a machine to read at all.
Lesson one: Perception needs predictable behavior
It started around 2016 with a new OCR engine from Microsoft Research Asia. This was pre-transformer, convolutional, English-only, and industry-leading at the time. Before you can understand a document, you have to see it. You have to be able to find the text lines in a skewed scan, separate characters from stamps and stains, and decide what’s a letter and what’s noise. This layer is deterministic in the operational sense: the same page in, the same reading out, with no sampling.
Lesson two: One language isn’t enough
Our first expansion, Spanish, took the entire team about six months. That pace doesn’t scale, so we invested in the machinery: data collection, labeling pipelines, evaluation, iteration. The unlock was to organize languages into scripts—Latin, Cyrillic, and so on—and build one model per script instead of one per language.
One engineer shipped the last 50 languages in a single month. In 2016, our technology was English-only. Today, we support 309 languages.
Lesson three: Not all figures are made equal
OCR hands you a wall of words. The customer needs specific data. That gap between content extraction and field extraction pushed us into entity recognition and document verticals: receipts, invoices, IDs, passports. The system has to recognize the type of document, understand its purpose, and know which figures matter in it. We built entity-specific benchmarks to test whether it gets the dates, amounts, and identifiers right within a specific document type.
Lesson four: To find the right value, the system has to understand the page’s geometry
With Microsoft Research Asia, we built LayoutLM followed by LayoutXLM in 2021—transformers that take each word’s position as input alongside the word itself. The model reasons in two dimensions, the same way your eye scans a form. We added state-of-the-art table detection and recognition in 2022, which is critical because a small structural mistake in a table can corrupt every downstream answer.
Then LLMs arrived
When retrieval-augmented generation took off, PDFs suddenly mattered a lot more. The knowledge behind every enterprise chatbot lives primarily in documents. But retrieval systems split documents into pieces for indexing, and a naive splitter can cut a table away from the heading that explains it, leaving one chunk full of numbers with no label and another with a label and no numbers. Demand for our layout work exploded.
By then we had mature, purpose-built systems for reading, page structure, and field extraction. But by 2023, when LLMs showed up, everyone in the field was asking the same question about their stack: can LLMs just replace all of this? With a decade of purpose-built neural models on the line, we asked it, too.
How the work changed
We didn’t wait for an answer. In 2024, we started building LLM-based field extraction, which became Azure Content Understanding, one of the first products to combine traditional content extraction with LLMs.
We learned two things:
- LLMs won on language understanding. On unstructured documents, contracts, agreements, anything where meaning lives in the words, the LLM approach beats purpose-built neural models. The specialized models we spent a decade building lost that terrain to a general-purpose one.
- LLMs lost on page layout. We could no longer feed each word’s position into the model. On rigid structured forms like tax forms and mortgage applications, where meaning lives in the geometry, the 10-year-old approach still won.

Today, our perception models convert documents into Markdown, preserving structure such as headings, tables, and reading order before the content is passed to an LLM for field extraction.
We’ve also experimented with sending the rendered page image alongside the Markdown. The results have been mixed. This is an open engineering frontier, and we’re working to close the gap on structured documents.
The LLM couldn’t replace the stack. It improved one layer, on some document types, while still relying on the rest of the existing layers. The deterministic perception models still do the seeing. It turned out that the real question is what you build around an LLM so an enterprise can trust its output.
What we built around the model
The lesson from a decade of work is that an LLM is an important part of a document understanding system, but it is not the whole system. In Azure Content Understanding, we use the LLM as the reasoning engine and build several layers around it to address the failure modes we’ve seen repeatedly in production:

- Deterministic perception. Our OCR and layout models establish what is on the page: text, reading order, tables, and document structure across 309 languages. Given the same page, this layer produces the same representation without relying on LLM sampling. The LLM reasons over that representation rather than having to infer the page from scratch.
- Contextualized reasoning. General-purpose LLMs understand language well, but document extraction often depends on knowledge specific to the document type or customer. We recently introduced Advanced Contextualization, which lets an analyzer use labeled examples and document knowledge to guide the LLM to make extractions. In internal evaluations, it improved average accuracy by up to 3.5% while reducing average LLM token usage by up to 22%. Training data stays in the customer’s own Azure Storage account and is used as a knowledge source rather than copied into the analyzer.
- Grounding. Grounding makes each extracted value traceable to source evidence, giving applications and humans a way to verify what the model produced. In our latest release, merging grounding deeper into the extraction flow cut average inference token usage by up to 28% while improving accuracy, and it preserves traceability without extra post-processing.
- Confidence. Grounding tells you where an answer came from; confidence helps determine how much to trust it. An in-house calibrated confidence model scores every extracted field, so the system signals what it’s unsure about. In production, those scores decide whether a result moves straight through, routes to human review, or gets additional validation. Our latest confidence refresh improved accuracy, measured by AUROC, by up to 14%.

- Schema-bound structured output. The system returns fields in a developer-defined JSON structure instead of free-form prose. Downstream applications can consume the result as an API response while grounding and confidence provide the information needed to decide what to do with it.

These layers reinforce one another. Perception establishes what’s on the page. The LLM reasons about what it means. Grounding connects that reasoning back to evidence. Confidence exposes uncertainty. The schema turns the result into something software can act on.
Just as importantly, this separation lets us apply LLM reasoning selectively, rather than using it for work that more specialized components can already do well. With our most recent contextualization work, the pre-built analyzers in the current preview cut LLM token consumption by up to 99%, and some extractions use no LLM tokens at all. A mature document pipeline uses the model’s reasoning where it adds value, rather than everywhere by default.
What’s next
Closing the structured-document gap is our near-term priority. The goal is for the LLM-based approach to win across every document type rather than only where language dominates layout.
The reasoning layer is expanding as well. Our newest preview adds an agentic mode for the hardest extractions, where evidence is spread across a long document and values must be compared and validated against intermediate results before a final answer. It costs more in latency and tokens, so it’s there for cases where quality matters most.
And all this is relevant beyond PDFs. The same problem—feeding unstructured content in and getting trustworthy structured data out—extends to images, audio, and video, too. And we’re already working on those modalities.
The takeaway
Breakthroughs matter, but trust is built one edge case at a time.
Over a decade of work, the models and hardware were replaced several times, and the failure modes were addressed one by one. We spent years sharpening the system against the thousands of ways real-world documents fail: a missing decimal point, a blurry scan, a malformed table, a signature in the wrong place.
If you’re building with LLMs today, the model matters, but don’t confuse it for the system itself. The system is everything you build around the model so its intelligence holds up on real documents, at real volume, where mistakes have real costs.