Reality of Document AI Comparisons
Benchmarks can be twisted and have severe limitations, always pair them up with a smell test on your own data
Reality of Document AI Comparisons

TL;DR
Benchmarks in document AI can be easily gamed. Anyone can make their system look like the best — especially if the dataset, prompts, or evaluation methods aren’t transparent. I came across one such recent comparison claiming 116/120 wins over Landing AI and Mistral which looks impressive on paper but falls apart under scrutiny: no standard OCR metrics (like CER/WER), no ground truth data, and likely internal reviewers instead of independent evaluators.
Their benchmark mixes real points with marketing spin — cherry-picking test data instances and not showing live playground results or extractions via. API. In contrast, transparent evaluations (like Will It Extract episodes) show how honest, reproducible comparisons should be done.
Bottom line: treat flashy accuracy numbers with skepticism. True credibility comes from transparency, reproducibility, and showing your work — not just winning your own benchmark.
[embed]Don’t trap yourself in the deception of metrics and benchmarks, pair them up with a smell test on the most complex docs you have.
Introduction
The reality of document analysis solutions is that anyone can make it look like their solution is the best. A good and knowledgeable customer/user will never take the benchmark at face value. They will do their own testing. In fact, sometimes it works the opposite way. Whenever someone publishes their doc solution and compares it to GPT or Gemini, you gotta be even more careful, and notice how they have set up their prompts and whether they have made all the prompts and chats public to begin with. Because the outcome of the LLM depends on how you prompt it and how you set up the problem.
Will It Extract Plug! Remember the Will It Extract Episode 1? Have a look here if you haven’t already! I was fully transparent with all the ChatGPT chat links and live comparison videos of me prompting them systematically and showing, step by step, the whole process.
Recap from Comparison Report
According to the comparison report, out of 120 test documents spanning invoices, forms, bank statements, and passports, Docsumo’s OCR was preferred in 116 cases, versus only 4 for Landing AI and 0 for Mistral. These eye-popping results, along with anecdotal examples of competitors’ failures, paint a picture of Docsumo as the clear winner. But how credible are these claims? In this analysis, we closely examine Docsumo’s evaluation methodology and claims, contrast them with known facts, and outline what a “proper eval” should entail. Our goal is to separate marketing spin from a fair, concrete assessment of OCR capabilities.
Proof is in the Pudding — Smell Test using playground
Hey, even before you read further, why don’t we just smell-test Docsumo’s playground and the LandingAI playground? I won’t even talk much. The video will be worth a thousand words.
[embed]ADE Playground is great for smell testing your most complex docs :)
Digging Deeper into the Comparison Report
At first glance, Docsumo’s report appears comprehensive — they even made their sample outputs public via a HuggingFace space for anyone to inspect. However, a deeper look reveals serious credibility issues in how this benchmark was designed and presented. Let’s break down the concerns in detail.
1. Lack of Objective Metrics: The benchmark results rely heavily on subjective human judgments (“preference count”) and qualitative anecdotes, rather than standard OCR accuracy metrics. Docsumo’s own documentation emphasizes that OCR accuracy should be measured by quantifiable metrics like Character Error Rate (CER), Word Error Rate (WER), field-level accuracy, etc., along with processing speed. In fact, Character Error Rate (the percentage of characters misrecognized) and word-level precision/recall are widely used benchmarks in OCR research. Yet, Docsumo’s report did not publish any CER, WER, or precision/recall figures for the three systems. Without these metrics, it’s impossible to verify claims of “higher accuracy” in a rigorous way. The only quantitative result shared was the number of documents (out of 120) where each system’s output was “preferred” by reviewers — a highly subjective criterion unless strict scoring guidelines were in place (which were not described).
2. Opaque Human Evaluation Procedure: Docsumo states that three human reviewers independently rated the outputs, but provides no details on the evaluation criteria or protocol. Were these reviewers employees of Docsumo? (Likely yes, given no external audit is mentioned.) On what basis did they decide one OCR’s output was better — overall visual similarity to the original? number of errors? preservation of formatting? The blog simply presents a tally of preferences with Docsumo winning 97% of the time. Such an extreme skew naturally raises suspicions of confirmation bias. If the evaluators knew which output came from Docsumo (especially since layout differences might make it obvious), their judgments could be unintentionally biased in favor of their home product. A credible benchmark would ideally be double-blind (reviewers don’t know which output is which) and use predefined scoring rubrics or error counts. Docsumo gives no indication that these safeguards were in place.
3. No Public Release of Test Data: While Docsumo did make the outputs available, the actual input documents (the images/PDFs) used in the test are not readily downloadable as a bundle for independent verification. One can manually inspect them on HuggingFace, but Docsumo did not release ground-truth text for each image or any script to reproduce their scoring. Without ground truth and metrics, the community cannot truly verify how many characters or fields each system got wrong. In essence, we are asked to trust Docsumo’s internal evaluation. This runs counter to the spirit of scientific benchmarking, where test sets and metrics are usually shared openly. As one Reddit commenter bluntly put it, “Come on man, don’t use words like objectivity… This is an ad.” The lack of a transparent, reproducible evaluation protocol severely undermines the credibility of the reported results.
- Docsumo’s observations also mix truth with hyperbole — For Landing AI’s Agentic Document Extraction, Docsumo’s observations also mix truth with hyperbole. Generative extraction models can indeed paraphrase or describe content instead of extracting verbatim text — for example, interpreting a logo or stamp in words. Docsumo notes a case where a simple “ABC” logo turned into a 130-word description by Landing’s model. But did they notice the visually grounded image being saved by the Python library? Is it really not a desired outcome for someone’s use case?
Similarly, reading vertical text is a known challenge — a misread like “89000458” becoming “80000456” was reported. These are legitimate errors. Most likely they used the research release Andrew launched in early April, which had some kinks — and note that if someone claims 99% accuracy right from the beginning, it’s likely a fantasy. I mean, come on.
- Speed — If you have read so far, you already know the playground didn’t take more than a few seconds. Docsumo said we took around 1 minute per page, and sometimes over 2 minutes on some pages but forgot to mention that they tried research release version.
Conclusion
So why did I write this article, lol? Because this was a perfect example of why benchmarks, comparisons, and flashy metrics should always be taken carefully and in context. Anyone can claim victory on their own test if they control the setup. What truly matters isn’t just the number, it’s how you got there. Transparency builds trust. When you share prompts, show your setup, and let others reproduce your results, you don’t just prove performance, you prove integrity.
My next article with our Staff MLE, Shankar, on DocVQA validation split is going to be mind-blowing — we got 95%+ accuracy putting us in top 3 — but again, it’s not about the number. It’s about transparency, reproducibility, and earning trust, one evaluation at a time.
Originally published at https://landing.ai on October 17, 2025.
메타데이터
- post_id
- 415fe22c8423
- slug
- reality-of-document-ai-comparisons-415fe22c8423
- url
- https://medium.com/@ankitkhare/reality-of-document-ai-comparisons-415fe22c8423
- canonical_url
- https://medium.com/@ankitkhare/reality-of-document-ai-comparisons-415fe22c8423
- author_url
- https://medium.com/@ankitkhare
- status
- ok
- fetched_at
- 2026-06-24 11:06:28