Source: Unite.AI
LlamaIndex introduced Extract v2.5 on October 1, 2026, a new generation of its schema-based document extraction agents. In a blog post bylined to Adrian Lyjak and Eli Stewart, the company reported accuracy improvements across all three of its extraction tiers and enhanced grounding for its two highest tiers, Agentic and Agentic Plus, and said the gains come with no increase in per-page pricing.
Tier-By-Tier Benchmark Gains
LlamaIndex reported that on ExtractBench, its open document-extraction benchmark, overall value F1 rose from 87.1 to 93.9 for the Cost Effective tier, from 89.8 to 95.8 for Agentic, and from 95.1 to 96.4 for Agentic Plus, its highest-accuracy tier. According to the company, Cost Effective now outperforms the previous Agentic tier, and Agentic outperforms the prior Agentic Plus.
ExtractBench tests 370 enterprise documents totaling 4,869 pages across 8 business domains and 67 document types, each type with its own schema. LlamaIndex published the benchmark with a public dataset, evaluation harness, and methodology, and has evaluated frontier VLMs, coding agents, and specialized extraction APIs on it.
New Agent Harness and Structural Reasoning
The company said v2.5 runs on a new agent harness purpose-built for document extraction, taking inspiration from the latest coding agents and tuned around the models to handle failure modes across vision, reasoning, and verification. LlamaIndex said the resulting agent can cross-reference context from multiple pages and ground values in exact sources.
A capability the post calls Structural Reasoning works over document representations to tailor extraction based on document type, layout, and information density: the agent spends more effort on complex documents and processes simpler ones with lower latency. For v2.5, the team brought the harness behind Agentic Plus to the Cost Effective and Agentic tiers, adapting its tools and extraction strategies to each tier’s cost constraints after evaluating document representations, tools, and model configurations. The company also said it worked on how agents follow schemas and extraction instructions, including tasks with up to 3,200 fields, and it advises users whose record IDs are essential for matching extracted records to a database to call that out in their prompts.
Long Lists, Cross-Page Records, and Scanned Forms
On long-list tasks, LlamaIndex reported scores rising from 82.0 to 92.6 for Cost Effective, from 86.1 to 95.5 for Agentic, and from 94.5 to 96.3 for Agentic Plus. The company said v2.5 uses intermediate representations to work with the full dataset and validates its output against the user’s schema. In one example, a 17-page fund filing, the previous Cost Effective tier returned 87 of 238 holdings; the company said v2.5 returns all 238, lifting that document’s extraction score from 52.8 to 98.7.
On records spanning pages, reported scores rise from 80.3 to 92.0 for Cost Effective, from 85.5 to 96.5 for Agentic, and from 93.6 to 96.7 for Agentic Plus. In a United Way Worldwide Schedule I grant schedule, the company said Agentic v2.0 stopped at a page break, leaving a grant’s purpose and address incomplete, while v2.5 includes the continuation from the next page in the same record.
On scanned forms, reported scores rise from 89.6 to 93.9 for Cost Effective, from 90.9 to 95.7 for Agentic, and from 94.4 to 96.1 for Agentic Plus. On an annotated W-14 disposal permit, the company said the Agentic tier previously combined a reviewer’s date with the permit number and appended a reviewer note to the freshwater depth, while v2.5 separates the original values from the annotations into the fields requested by the schema.
Advanced Citations and Grounding
The release introduces Advanced Citations for the Agentic and Agentic Plus tiers, which locate bounding boxes of supporting evidence for difficult fields, including values that do not appear word for word in the document. An initial version was previously available in Agentic Plus; LlamaIndex said it improved the feature’s grounding accuracy further and extended it to Agentic. The company reported ExtractBench grounding scores rising from 46.8 to 80.6 for Agentic and from 46.4 to 82.2 for Agentic Plus. Its comparison chart lists grounding scores of 63.2 for Sonnet 5.5, 77.6 for GPT-6 SOL, 77.5 for Opus 5.5, 50.8 for Reducto Deep Extract, 21.8 for Extend Max, and 54.3 for its own Cost Effective tier.
The company said bounding-box citations combine with confidence scores, which help teams decide which extracted values can proceed automatically and which need a closer look, allowing reviewers to check flagged values against the original document in human-in-the-loop workflows.
Spreadsheet Mode and Availability
v2.5 adds native spreadsheet extraction: in spreadsheet mode, agents work directly with workbook cells rather than a flattened representation, and users can enable the mode in the extraction configuration.
Extract v2.5 is available through LlamaCloud, a Go-based command-line tool called llp, an MCP server for agent clients including Claude Code and Codex, and the Python llama-cloud SDK, where extraction jobs select version 2.5 and can enable citesources and confidencescores options. LlamaIndex encouraged users to run v2.5 on documents representative of their workflows using their existing schemas, and said it expects the version to be a significant step up in extraction quality, particularly for documents that previously needed manual cleanup or were not reliable enough to automate.
