[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-extractbench-schema-guided-document-extraction-en":3,"article-related-extractbench-schema-guided-document-extraction-en":29,"series-research-1d615153-f127-4b7a-8043-ba4b2701c1a8":75},{"id":4,"slug":5,"title":6,"content":7,"summary":8,"source":9,"source_url":10,"author":11,"image_url":12,"cover_image":12,"category":13,"language":14,"translated_content":11,"related_article_id":15,"keywords":16,"key_takeaways":22,"views":26,"created_at":27,"published_at":28,"topic_cluster_id":11},"1d615153-f127-4b7a-8043-ba4b2701c1a8","extractbench-schema-guided-document-extraction-en","ExtractBench benchmarks schema-guided document extraction","\u003Cp>Enterprise document extraction often fails in the same boring ways: fields get missed, records get truncated, or the output looks right but cannot be traced back to the source.\u003C\u002Fp>\u003Cp data-speakable=\"summary\">ExtractBench measures schema-guided enterprise document extraction across accuracy, grounding, and cost.\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cstrong>Research org\u003C\u002Fstrong>: Unspecified in arXiv abstract\u003C\u002Fli>\u003Cli>\u003Cstrong>Core data\u003C\u002Fstrong>: 4,869 pages\u003C\u002Fli>\u003Cli>\u003Cstrong>Breakthrough\u003C\u002Fstrong>: First benchmark to score accuracy, completeness, grounding, and cost together\u003C\u002Fli>\u003C\u002Ful>\u003Cp>That matters because “good extraction” is not just about getting a field value into JSON. In enterprise settings, teams also need the record to be complete, the output to follow the user’s schema, and the extracted data to carry source evidence that can survive review, debugging, and compliance checks.\u003C\u002Fp>\u003Ch2>What problem this paper is trying to fix\u003C\u002Fh2>\u003Cp>The paper focuses on \u003Cem>schema-guided extraction\u003C\u002Fem>: a document comes in, a user-defined schema goes out, and the system has to follow that schema faithfully while grounding its answers in the source. That is a more demanding task than plain OCR or generic information extraction, because the model is not just finding text — it is deciding what belongs in a structured record and proving where it came from.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785738793571-kut9.png\" alt=\"ExtractBench benchmarks schema-guided document extraction\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>For developers, the pain point is familiar. A system can look solid on a few easy documents and still fall apart when the document gets longer, the layout changes, or the schema requires many records instead of a single value. ExtractBench is built to make those failure modes visible instead of hiding them behind a single headline score.\u003C\u002Fp>\u003Cp>The \u003Ca href=\"\u002Ftag\u002Fbenchmark\">benchmark\u003C\u002Fa> covers 4,869 pages across 370 enterprise documents, 8 business domains, and 67 document types. It also includes clear tags for different challenge scenarios, which is useful because not all extraction problems are equally hard. Short forms, long record lists, and other document types do not stress the same parts of a system.\u003C\u002Fp>\u003Ch2>How the benchmark is put together\u003C\u002Fh2>\u003Cp>ExtractBench combines three curation strategies to scale up ground truth without pretending every document can be handled the same way. For real documents, it uses independent-system agreement. For synthetic lists, it relies on known values. For forms, it adds human verification.\u003C\u002Fp>\u003Cp>That mix is practical. It acknowledges that enterprise extraction data is expensive to label and that no single annotation method is enough for every document type. By combining methods, the authors can build a larger evaluation set while still keeping the labels credible enough for benchmarking.\u003C\u002Fp>\u003Cp>The paper also says the benchmark is, to its knowledge, the first to score value accuracy, record completeness at scale, grounding, and measured cost together. That is the key design choice here: instead of treating extraction as one metric problem, it treats it as a multi-objective system problem.\u003C\u002Fp>\u003Cul>\u003Cli>Value accuracy is measured with order-insensitive value F1.\u003C\u002Fli>\u003Cli>Grounding is measured with word-level and page-level F1.\u003C\u002Fli>\u003Cli>Cost is measured alongside the quality metrics rather than ignored.\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>What the paper actually shows\u003C\u002Fh2>\u003Cp>The abstract does not give full benchmark tables or exact scores, so there are no detailed numbers to quote beyond the dataset scale and the ranking claims. What it does say is still enough to show the shape of the tradeoff.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785738809196-l62o.png\" alt=\"ExtractBench benchmarks schema-guided document extraction\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>Commercial VLMs perform well on short documents, but they often truncate record lists on long ones. That is an important practical warning: a model that seems strong on single-page or low-complexity inputs may not be reliable once the document demands full-list extraction.\u003C\u002Fp>\u003Cp>Coding agents, by contrast, keep higher accuracy but at much higher cost. That is the classic enterprise tension: better output quality can come with a steep runtime or orchestration bill. ExtractBench makes that tradeoff visible instead of forcing teams to discover it in production.\u003C\u002Fp>\u003Cp>The paper reports that LlamaExtract Agentic Plus ranks first on all three metrics, and that its accuracy is comparable to coding agents at a fraction of the cost. The abstract does not provide the exact metric values, so that is as far as the source lets us go. Still, the result suggests that agentic extraction systems may be able to close much of the quality gap without paying the full coding-\u003Ca href=\"\u002Ftag\u002Fagent\">agent\u003C\u002Fa> price.\u003C\u002Fp>\u003Ch2>Why developers should care\u003C\u002Fh2>\u003Cp>If you are building document pipelines, vendor intake flows, compliance tooling, or any workflow that turns messy PDFs into structured records, this benchmark is useful because it tests the stuff that breaks real systems. A model that only scores well on value extraction is not enough if it drops records or cannot show its work.\u003C\u002Fp>\u003Cp>Grounding is especially relevant for enterprise use. When an extracted field needs to be audited, reviewed, or reconciled, the system has to point back to the source text or page. ExtractBench’s word- and page-level grounding metrics give teams a way to compare systems on traceability, not just on final output shape.\u003C\u002Fp>\u003Cp>The cost dimension matters too. Teams often compare models on quality alone and then discover that the “best” system is too expensive to run at scale. By putting cost into the evaluation, the benchmark pushes the conversation toward deployable systems rather than demo-friendly ones.\u003C\u002Fp>\u003Ch2>Limits and open questions\u003C\u002Fh2>\u003Cp>The abstract leaves several things unspecified. It does not provide the full metric tables, exact scores, or detailed per-category breakdowns in the summary we have here. It also does not tell us how the benchmark behaves across every challenge tag, so you should not assume the reported ranking generalizes uniformly across all document types.\u003C\u002Fp>\u003Cp>Another limitation is that benchmark design can only approximate production reality. Even with 370 enterprise documents and multiple domains, real deployments may face different schemas, noisier scans, or more adversarial formatting. The benchmark is a strong step toward better evaluation, but it is still an evaluation set, not the whole problem.\u003C\u002Fp>\u003Cp>Still, ExtractBench points in a direction that matters: extraction systems should be judged on whether they are accurate, complete, grounded, and affordable at the same time. That is a much better fit for enterprise workflows than a single accuracy number.\u003C\u002Fp>\u003Cp>For teams choosing between commercial VLMs, coding agents, or more agentic extraction systems, the paper’s message is straightforward: short-document success is not enough, and cost cannot be an afterthought. A benchmark like this helps you ask the right question before you ship.\u003C\u002Fp>\u003Cp>In that sense, ExtractBench is less about one model winning and more about changing what “good” means for document extraction. For developers, that is the useful part.\u003C\u002Fp>","ExtractBench measures schema-guided enterprise document extraction across accuracy, grounding, and cost.","arxiv.org","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.29677",null,"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785738793571-kut9.png","research","en","2dabbf39-7875-4653-91da-0b7cd9db4185",[17,18,19,20,21],"document extraction","schema-guided extraction","benchmark","grounding","enterprise AI",[23,24,25],"ExtractBench evaluates extraction on accuracy, completeness, grounding, and cost together.","Commercial VLMs struggle more on long documents and record lists.","Coding agents are accurate but expensive; LlamaExtract Agentic Plus is reported to match them more cheaply.",0,"2026-08-03T06:32:48.963876+00:00","2026-08-03T06:32:48.954+00:00",{"tags":30,"relatedLang":34,"relatedPosts":38},[31,32],{"name":19,"slug":19},{"name":21,"slug":33},"enterprise-ai",{"id":15,"slug":35,"title":36,"language":37},"extractbench-schema-guided-document-extraction-zh","ExtractBench 盯住企業文件抽取","zh",[39,45,51,57,63,69],{"id":40,"slug":41,"title":42,"cover_image":43,"image_url":43,"created_at":44,"category":13},"0832d329-8eb2-4a7b-9dca-8e52ba1f2d04","private-mode-finding-regression-clustering-en","Private mode finding for regression and clustering","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785740577783-kwop.png","2026-08-03T07:02:30.144339+00:00",{"id":46,"slug":47,"title":48,"cover_image":49,"image_url":49,"created_at":50,"category":13},"ec3d8c98-79d8-4be6-a08b-6f34972e3125","toktier-stateful-tokenization-agentic-llm-serving-en","TokTier cuts tokenization overhead for agentic LLMs","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785736974393-o8su.png","2026-08-03T06:02:30.190251+00:00",{"id":52,"slug":53,"title":54,"cover_image":55,"image_url":55,"created_at":56,"category":13},"8117c0b0-2d1a-41eb-bf11-a66f1b28c6db","systema-turns-aivc-scores-into-a-harder-test-en","Systema turns AIVC scores into a harder test","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785632613702-lswn.png","2026-08-02T01:03:09.600967+00:00",{"id":58,"slug":59,"title":60,"cover_image":61,"image_url":61,"created_at":62,"category":13},"f97ca9f3-d2db-477c-8c89-e1af7450e9e4","stablecoin-remittances-hit-9-percent-bank-italy-test-en","Stablecoin remittances hit 9% in Bank of Italy test","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785591177728-mx7u.png","2026-08-01T13:32:33.046591+00:00",{"id":64,"slug":65,"title":66,"cover_image":67,"image_url":67,"created_at":68,"category":13},"3183c9c7-4249-450c-9a5b-8938785357fe","stablecoins-hit-308b-as-svbs-shock-echoes-en","Stablecoins Hit $308B as SVB’s Shock Still Echoes","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785585774053-zcn7.png","2026-08-01T12:02:28.552018+00:00",{"id":70,"slug":71,"title":72,"cover_image":73,"image_url":73,"created_at":74,"category":13},"fddc60df-68a9-44b4-8e0e-79f3985d2b49","rust-compiler-speed-wins-july-2026-en","Rust compiler speed wins from July 2026","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785569563875-w1xs.png","2026-08-01T07:32:21.551218+00:00",[76,81,86,91,96,101,106,111,116,121],{"id":77,"slug":78,"title":79,"created_at":80},"a2715e72-1fe8-41b3-abb1-d0cf1f710189","ai-predictions-2026-big-changes-en","AI Predictions for 2026: Brace for Big Changes","2026-03-26T01:25:07.788356+00:00",{"id":82,"slug":83,"title":84,"created_at":85},"8404bd7b-4c2f-4109-9ec4-baf29d88af2b","ml-papers-of-the-week-github-research-desk-en","ML Papers of the Week Turns GitHub Into a Research Desk","2026-03-27T01:11:39.480259+00:00",{"id":87,"slug":88,"title":89,"created_at":90},"87897a94-8065-4464-a016-1f23e89e17cc","ai-ml-conferences-to-watch-in-2026-en","AI\u002FML Conferences to Watch in 2026","2026-03-27T01:51:54.184108+00:00",{"id":92,"slug":93,"title":94,"created_at":95},"6f1987cf-25f3-47a4-b3e6-db0997695be8","openclaw-agents-manipulated-self-sabotage-en","OpenClaw Agents Can Be Manipulated Into Failure","2026-03-28T03:03:18.899465+00:00",{"id":97,"slug":98,"title":99,"created_at":100},"a53571ad-735a-4178-9f93-cb09b699d99c","vega-driving-language-instructions-en","Vega: Driving with Natural Language Instructions","2026-03-28T14:54:04.698882+00:00",{"id":102,"slug":103,"title":104,"created_at":105},"a34581d6-f36e-46da-88bb-582fb3e7425c","personalizing-autonomous-driving-styles-en","Drive My Way: Personalizing Autonomous Driving Styles","2026-03-28T14:54:26.148181+00:00",{"id":107,"slug":108,"title":109,"created_at":110},"2bc1ad7f-26ce-4f02-9885-803b35fd229d","training-knowledge-bases-writeback-rag-en","Training Knowledge Bases with WriteBack-RAG","2026-03-28T14:54:45.643433+00:00",{"id":112,"slug":113,"title":114,"created_at":115},"71adc507-3c54-4605-bbe2-c966acd6187e","packforcing-long-video-generation-en","PackForcing: Efficient Long-Video Generation Method","2026-03-28T14:55:02.646943+00:00",{"id":117,"slug":118,"title":119,"created_at":120},"675942ef-b9ec-4c5f-a997-381250b6eacb","pixelsmile-facial-expression-editing-en","PixelSmile Framework Enhances Facial Expression Editing","2026-03-28T14:55:20.633463+00:00",{"id":122,"slug":123,"title":124,"created_at":125},"6954fa2b-8b66-4839-884b-e46f89fa1bc3","adaptive-block-scaled-data-types-en","IF4: Smarter 4-Bit Quantization That Adapts to Your Data","2026-03-31T06:00:36.65963+00:00"]