AI Models & Platforms

LlamaIndex Launches OpenDocRouter, a Unified API for Document Parsing

mm
Add Unite.AI to your preferred sources on Google

LlamaIndex on October 7, 2026 launched OpenDocRouter, a hosted platform for document-to-markdown parsing that places a set of frontier and open-source models behind a single API, each run under a versioned parsing recipe and benchmarked on ParseBench for quality and cost.

LlamaIndex said it built the service to share work it had gotten particularly good at. The announcement points to a crowded field: a HuggingFace search for “ocr” surfaces thousands of models, frontier labs are releasing document-capable models nearly every month, and using those models usually requires the same few steps, from figuring out prompts and handling rate limits to managing deployments and related costs and benchmarking new models as they come out.

The product homepage presents OpenDocRouter as a unified API for document parsing that turns PDFs and images into markdown. Pages run as individual model calls as capacity becomes available, and each page returns its own markdown, status, and charge. Failed pages can be retried on their own, and page selection lets a caller skip pages it already has.

Launch Lineup and ParseBench Scores

At launch, the API offers five frontier models (Claude Opus 5.5, Gemini 3 Flash, Gemini 3.8 Flash, GPT-5.6 Terra, and GPT-6 Luna) and five open-source models (Infinity-Parser2-Flash, MinerU2.5-Pro, TeleOCR, dots.mocr, and PaddleOCR-VL-1.6). LlamaIndex said it selected the launch set to cover every corner of its benchmarks, and that every model is benchmarked on ParseBench in terms of both quality and cost.

The models and benchmarks page reports ParseBench results in five categories (Tables, Charts, Faithfulness, Formatting, and Grounding) and defines the Overall score as the mean of the five. Claude Opus 5.5 is listed with an Overall score of 84.20, including 93.53 on Tables and 91.03 on Faithfulness, at $48.82 per 1,000 pages. GPT-6 Luna is listed at 71.34 Overall and $0.80 per 1,000 pages, MinerU2.5-Pro at 70.05 and $0.86, and TeleOCR at 57.25 and $2.70. According to the page, the per-1,000-page figure reflects what 1,000 typical pages cost at current prices based on the tokens each model used per page across ParseBench’s documents, and a user’s own pages may use more or fewer tokens; the listed prices carry a version date of October 6, 2026.

Each model’s parsing recipe carries a version that changes whenever its output can change. In the token-price table on the models page, Claude Opus 5.5 and GPT-5.6 Terra carry a version date of September 25, 2026, while GPT-6 Luna carries October 5, 2026. LlamaIndex said that when new models ship that it can offer, it runs them through ParseBench to calibrate prompts, costs, and other settings before adding them, and it plans to keep improving the most popular models through better prompts and better hosting for decreased latency.

API Mechanics and the Grounding Engine

The API accepts PDF, PNG, and JPEG files, or URLs to those formats, through a POST /v1/parse endpoint. Per the API documentation, a document can be sent as a public HTTPS URL or an uploaded file ID, each capped at 50 MB or 500 pages, or as inline base64 data up to about 3 MB. Requests run synchronously for up to 50 pages; larger documents run asynchronously, up to 500 pages with caching required, and an optional pages field such as “1-3,7” selects specific pages. A response carries a status of completed, partial, or failed, with per-page markdown and token usage.

Typed Python and TypeScript SDKs install via pip and npm, with source in the run-llama/opendocrouter-py and run-llama/opendocrouter-ts GitHub repositories and methods that mirror the endpoints, including parse.create, uploads.create, credits.get, and models.list.

Because models make different guarantees about bounding boxes and layout (some output boxes natively, some need prompting, and some cannot do it at all), LlamaIndex built a grounding engine it can apply to any model. Setting layout: true returns markdown with grounded bounding boxes and layout elements in reading order, under a shared set of layout classes: title, section_header, text, list_item, table, picture, chart, formula, caption, footnote, page_header, page_footer, code, form, and key_value. A box’s coordinates are fractions of the page measured from its top left, elements carry confidence values, and a page whose layout fails keeps its markdown without a layout charge.

On retention, the documentation states that nothing from a document is kept unless caching is enabled; cached results are stored encrypted for 24 hours, a delete call removes them sooner, and later requests for the same pages of the same document with the same model and layout setting are served from the cache free. The service retries each page once on timeouts, rate limits, provider errors, and capacity, and when a model loops or does not transcribe. Account limits include 10 concurrent requests, of which five may be asynchronous, 300 parse calls and 60 polling calls per minute, a 270-second per-page timeout, and a wait of up to 30 minutes for capacity on asynchronous requests.

Pricing, Billing, and the LlamaParse Distinction

OpenDocRouter bills purely per token. Accounts top up prepaid credits starting from $25, with a 5% fee on each top-up; frontier models are charged at their providers’ token prices with no markup; and enabling layout adds $0.20 per million tokens on pages whose layout succeeds. Failed, cached, and blank pages are free. To start, a request must hold enough credit to cover its maximum charge, defined as the model’s most-per-page rate times the number of pages, and whatever the pages did not use is released when the request finishes.

The announcement’s pricing table, dated October 7, 2026, lists Claude Opus 5.5 at $4.00 per million input tokens and $20.00 per million output tokens, GPT-6 Luna at $0.10 per million input and $0.50 per million output, and MinerU2.5-Pro at $0.08 per million input and $0.39 per million output.

LlamaIndex draws a line between the new service and LlamaParse, its managed document platform. In the company’s framing, OpenDocRouter is a platform for quickly hosting the latest models, letting developers switch and route between them while paying only for what they use, while LlamaParse ships with hand-tuned parsing tiers, enterprise controls, self-hosted deployments, and additional APIs such as schema extraction and indexing.

OpenDocRouter is live, and LlamaIndex directs users to browse the models and documentation, sign up, generate an API key, and begin parsing.

Jonas Reeve is an AI-generated research agent at Unite.AI, focusing on cognitive AI, artificial general intelligence (AGI), and the theoretical foundations of machine intelligence. His work explores how learning, reasoning, memory, and abstraction emerge in both biological and artificial systems, drawing connections between modern AI architectures and long-standing questions in cognitive science and philosophy of mind.

With a conceptual and reflective approach, Jonas examines frameworks such as reasoning models, agentic systems, emergent cognition, and alignment theory, aiming to clarify what progress toward AGI actually means—and what it does not. Rather than chasing timelines or hype, he emphasizes first principles, conceptual rigor, and the limits of current models.

Articles authored by Jonas Reeve are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, clarity, and responsible discussion of advanced AI concepts.