GenAI PM
tool4 mentions· Updated Aug 18, 2026

ExtractBench

A benchmark for extraction tasks that requires both extracted values and citations to be correct. It evaluates word-level boxes and page-level performance, making it relevant to document AI evaluation.

Key Highlights

  • ExtractBench measures whether systems return both the correct extracted value and the correct citation, making it stronger than answer-only benchmarks.
  • The benchmark is deterministic with zero LLM judges and spans 14 systems, 370 enterprise documents, 4,869 pages, and 67 document types.
  • It is especially useful for exposing silent recall failures on long documents, tables, scans, rotated pages, and handwriting.
  • For AI PMs, ExtractBench is a practical framework for vendor evaluation, production readiness testing, and auditability requirements.

ExtractBench

Overview

ExtractBench is a deterministic benchmark for document extraction systems that tests whether a model or pipeline can return both the correct extracted value and the correct citation/location in the source document. Rather than relying on subjective LLM-as-a-judge scoring, it evaluates extraction quality against ground truth, including word-level bounding boxes and page-level performance. That makes it especially relevant for teams building document AI products where correctness, auditability, and traceability matter.

For AI Product Managers, ExtractBench is useful because it surfaces failure modes that can be hidden by simpler benchmarks or clean-PDF demos. The benchmark covers 14 systems across 370 enterprise documents, 4,869 pages, and 67 document types, and highlights how performance can degrade sharply on long, scanned, rotated, handwritten, or otherwise messy documents. It is particularly valuable for evaluating whether a vendor or internal system can support production workflows that require reliable extraction with verifiable evidence.

Key Developments

  • 2026-08-12: LlamaIndex announced ExtractBench as a deterministic benchmark with zero LLM judges, covering 14 systems across 370 enterprise documents, 4,869 pages, and 67 document types. It reported that beyond 50 pages, commercial VLMs fell below 35% recall while often maintaining high precision, effectively missing large portions of content such as table rows.
  • 2026-08-13: LlamaIndex said it had released ExtractBench the previous day and shared more benchmark findings. Across the longest documents, frontier VLMs scored 8.9% to 35.8% F1 as recall collapsed, while its iterative Agentic Plus tier reached 96.1% F1 on long-list tasks and was the only system reported to maintain performance as document length increased.
  • 2026-08-14: LlamaIndex announced Agentic Plus (Extract Tier) as the only system among 14 tested without a major blind spot in ExtractBench. It scored 95.9%, 93.9%, and 93.8% on rotated, scanned, and handwritten documents, with only a 2-point spread versus swings of 10+ points for other systems.
  • 2026-08-18: LlamaIndex recapped ExtractBench, emphasizing that systems must get both extracted values and citations correct, and that evaluation includes word-level boxes at IoU 0.5. LlamaExtract Agentic Plus reportedly led with 84.9% page-level and 46.4% word-level results, and achieved 87.1% on long documents where other systems scored zero.

Relevance to AI PMs

  • Use it to evaluate real production readiness, not demo quality. ExtractBench helps PMs compare extraction systems on hard enterprise documents, not just clean PDFs, making it a better tool for vendor selection and internal benchmarking.
  • Measure trust and auditability. Because the benchmark requires both the answer and its citation to be correct, it aligns well with workflows in finance, legal, insurance, and operations where users need verifiable evidence, not just plausible outputs.
  • Stress-test long-document and edge-case performance. PMs can use ExtractBench-style criteria to catch silent recall failures on long docs, scans, rotated pages, handwriting, and tables before these issues hit customers in production.

Related

  • llamaindex: Creator and primary promoter of ExtractBench, publishing benchmark results and comparisons across extraction systems.
  • commercial-vlms: A key comparison group in the benchmark; ExtractBench highlighted that these models can maintain precision while suffering major recall drops on long documents.
  • agentic-plus: LlamaIndex's iterative extraction approach, frequently cited in ExtractBench results for strong long-document and edge-case performance.
  • llamaextract-agentic-plus: The specific system variant highlighted as a top performer in later ExtractBench recaps, including page-level and word-level scores.

Newsletter Mentions (4)

2026-08-18
LlamaIndex 🦙 recapped ExtractBench, which requires both extracted values and citations to be correct and evaluates word-level boxes at IoU 0.5.

GenAI PM Daily August 18, 2026 GenAI PM Daily 🎧 Listen to this brief 3 min listen Today's top 20 insights for PM Builders, ranked by relevance from X, YouTube, LinkedIn, and Blogs. Cursor releases Origin, its integrated code hosting platform #1 𝕏 Cursor released Origin, its code hosting platform, with deep Cursor integration and repository syncing from GitHub. Cursor describes Origin as fast and easy to use. Also covered by: @Cursor , @Guillermo Rauch #2 𝕏 Philipp Schmid demonstrated Gemini 3.7 Flash using his Android emulator via ADB for the task “Play 1 round of Wordle.” He said its latency and visual reasoning make it exceptionally good for multimodal agentic use cases such as mobile control and Computer Use. #3 𝕏 LlamaIndex 🦙 recapped ExtractBench, which requires both extracted values and citations to be correct and evaluates word-level boxes at IoU 0.5. LlamaExtract Agentic Plus led with 84.9% page-level and 46.4% word-level results, achieving 87.1% on long documents where other systems scored zero.

2026-08-14
LlamaIndex 🦙 announced Agentic Plus (Extract Tier), which was the only system without a blind spot among 14 tested in ExtractBench, scoring 95.9%, 93.9%, and 93.8% across rotated, scanned, and handwritten documents.

#6 𝕏 LlamaIndex 🦙 announced Agentic Plus (Extract Tier), which was the only system without a blind spot among 14 tested in ExtractBench, scoring 95.9%, 93.9%, and 93.8% across rotated, scanned, and handwritten documents. Its 2-point spread compared with swings of 10+ for other systems highlights the risks of benchmarking extraction tools only on clean PDFs.

2026-08-13
LlamaIndex 🦙 said it released ExtractBench the previous day, benchmarking 14 systems across 370 enterprise documents.

#6 𝕏 LlamaIndex 🦙 said it released ExtractBench the previous day, benchmarking 14 systems across 370 enterprise documents. It reported that frontier VLMs scored 8.9–35.8% F1 on the longest documents as recall collapsed, while its iterative Agentic Plus tier achieved 96.1% F1 on long-list tasks and was the only system to hold performance flat as documents grew.

2026-08-12
"#1 𝕏 LlamaIndex 🦙 announced ExtractBench, a deterministic benchmark with zero LLM judges that tests 14 systems across 370 enterprise docs, 4,869 pages, and 67 doc types."

#1 𝕏 LlamaIndex 🦙 announced ExtractBench, a deterministic benchmark with zero LLM judges that tests 14 systems across 370 enterprise docs, 4,869 pages, and 67 doc types. It found that past 50 pages, commercial VLMs fall below 35% recall while maintaining high precision and silently omitting most table rows.

Stay updated on ExtractBench

Get curated AI PM insights delivered daily — covering this and 1,000+ other sources.

Subscribe Free