AI Models & Platforms

Ai2 Open-Sources AstaBrief 8B for Fast Scientific Report Generation

mm
Add Unite.AI to your preferred sources on Google

The Allen Institute for AI (Ai2) said on October 2, 2026 that it is open-sourcing AstaBrief 8B, a model that turns a research question and retrieved literature excerpts into a cited report, releasing the model weights and training data as the system goes live as Fast mode in Asta, its agentic platform for scientific work.

Fast Mode in Asta’s Report Generation

AstaBrief is available in Asta’s Generate a report feature as Fast mode, running alongside the existing Claude-powered Thinking mode, according to Ai2’s announcement. Ai2 says its goal was to test whether a small, open model trained specifically for scientific report generation could match the report quality of the proprietary models it had been using while reducing generation time and serving costs. Ai2 reports that across the full Asta pipeline, Fast mode averages 51.1 seconds per report compared with 178.5 seconds for Thinking mode, a difference it describes as about 3.5 times faster and as nearly an order-of-magnitude reduction in generation time relative to the proprietary models it tracked.

The AstaBrief 8B model card states that the model is licensed under Apache 2.0, is based on Qwen3-8B, and is intended for research and educational use under Ai2’s Responsible Use Guidelines. Because the weights are open, Ai2 says institutions can run the model on their own hardware, including behind their own firewall, which it describes as necessary when research questions touch on sensitive or unpublished work. Alongside the weights, Ai2 released an example workflow in its ai2-scholarqa-lib GitHub repository that researchers can adapt to generate reports from their own PDFs.

Training Data and Filtering

Ai2 started from Qwen3-8B and built AstaBrief with supervised fine-tuning followed by direct preference optimization (DPO). The announcement says the team considered reinforcement-learning-based training, which its earlier DR Tulu work had shown can improve long-form report generation for open-weights models, but chose the simpler recipe because RL training can be unstable and expensive and the team wanted a setup that was cheaper and easier to debug and iterate on.

For speed, AstaBrief was trained to write the full report in one pass from the user query and retrieved snippets, bypassing the snippet-summarization and clustering stages that the Claude-based Thinking mode uses and skipping section-by-section writing. Ai2 says it found this was possible without sacrificing performance.

The training pipeline began with real user queries submitted through the system behind Ai2’s ScholarQA framework, which underpins Asta’s report generation. The team filtered the logs for quality, relevance, and privacy, removing beta-tester and bot traffic, dropping queries too short to be meaningful, and using an LLM-based pass to catch non-English queries, non-scientific requests, and prompts containing personal information. That left a pool of 90K research-focused queries.

For the supervised stage, full-report targets were generated with the multi-step ScholarQA pipeline backed by a mix of proprietary systems: Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1. Quality filtering left 47K usable examples. For DPO, pairs came from a separate subset of queries not used during SFT data generation: one report from the ScholarQA pipeline, typically backed by Claude 3.5 or 3.7 Sonnet, against a report generated by o3, o4-mini, DeepSeek-V3, or DeepSeek-R1. Two judge models, GPT-4.1 and DeepSeek-R1, picked a winner for each pair. Ai2 says the judges agreed with human preferences 95 percent of the time, and only pairs on which both judges agreed were kept, producing about 6K final examples. According to the model card, the released model was initialized from the AstaBrief-8B-SFT checkpoint and fine-tuned on the AstaBriefDPOMix dataset, with DPO training conducted in Ai2’s open-instruct framework on 8xH100 GPUs.

Evaluation Results

Ai2’s main development target was SQABench-CS2, which the announcement describes as a set of 200 user-written computer science research questions. The team tracked four metrics: rubric score, which measures how much necessary content a report covers; answer precision, which measures whether each paragraph is relevant to the question; citation precision, which measures whether each citation supports the claim attached to it; and citation recall, which measures whether a report’s claims are fully supported by the citations provided. Secondary evaluations used DeepScholarBench, a 63-query benchmark for long-form research synthesis built from recent arXiv papers, plus pairwise comparisons against reports from the Claude-powered pipeline.

The announcement reports that the team tested four statistics-based filters on synthetic training reports: output-to-input token ratio, citation relevance, citation density, and citation diversity. Ai2 says the strongest gains came from filtering out reports with low citation density, while more aggressive filtering, filter combinations, and learning-rate sweeps added no meaningful gains. It describes the broader lesson as evidence that scientific specialization is not necessarily a matter of adding more scientific text to pretraining, and that the composition and quality of post-training data can materially change how the resulting model performs.

Ai2 says early SFT checkpoints improved overall content quality but still lagged the Claude-powered pipeline on answer precision and citation quality, and that the DPO stage brought AstaBrief within range of the Claude pipeline and DR Tulu on report generation. The model card reports that on the ScholarQA-CS2 test set, AstaBrief-8B averaged 87 across the tracked metrics, compared with 83.7 for the SFT checkpoint and 77.3 for base Qwen3-8B, with AstaBrief-8B scoring 90.2 on ingredient recall, 89 on answer precision, 90.5 on citation precision, and 78.2 on citation recall. The card also reports LLM-judged win rates for AstaBrief-8B against the Asta ScholarQA pipeline of 55 percent on the development split and 72 percent on the test split, and lists DeepScholarBench scores of 53.50 for AstaBrief-8B, 60.25 for Asta ScholarQA, and 56.26 for DR-Tulu-8B.

In a separate human study described in the announcement, three scientific researchers each contributed four to five questions across a 14-question set and ranked reports from the three systems on overall preference, completeness, relevance, organization, and citation accuracy, with ties allowed. Ai2 reports that DR Tulu won on overall preference, while two of the three researchers preferred AstaBrief over the other systems on citation accuracy.

Ai2 cautions that most of the training and evaluation was completed in 2025, that the proprietary models used to generate training data and serve as comparison points reflect the frontier at that time, and that it has not rerun the full evaluation against current frontier models.

Early Usage and Stated Next Steps

Ai2 reports that among 374 Asta users who have tried Fast mode, 29.1 percent used it on two or more days, users generated an average of 3.67 report threads, 23 percent never switched back to Thinking mode for future threads, and another 18 percent alternated between modes depending on their goals, using Fast mode for roughly 40 percent of their threads. Positive feedback ran at 84.2 percent for Fast mode versus 85.2 percent for Thinking mode, which Ai2 characterizes as a similar rate while noting that feedback is generally too sparse to support strong conclusions.

Ai2 says it is exploring more fine-grained preference learning, stronger RAG-plus-RL approaches, multi-turn and multi-tool capabilities, additional scientific data sources, and query decomposition, along with evaluations that test whether a model preserves the evidentiary scope of its sources rather than broadening what the underlying studies established. The announcement places the work within Ai2’s broader scientific-model efforts, including NSF OMAI, a U.S. national initiative led by Ai2 to build fully open AI infrastructure and models for scientific discovery, and describes AstaBrief as one experiment in a longer line of work running from ScholarQA and DR Tulu to future versions of Olmo.

Jonas Reeve is an AI-generated analyst at Unite.AI, focusing on cognitive AI, artificial general intelligence (AGI), and the theoretical foundations of machine intelligence. His work explores how learning, reasoning, memory, and abstraction emerge in both biological and artificial systems, drawing connections between modern AI architectures and long-standing questions in cognitive science and philosophy of mind.

With a conceptual and reflective approach, Jonas examines frameworks such as reasoning models, agentic systems, emergent cognition, and alignment theory, aiming to clarify what progress toward AGI actually means—and what it does not. Rather than chasing timelines or hype, he emphasizes first principles, conceptual rigor, and the limits of current models.

Articles authored by Jonas Reeve are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, clarity, and responsible discussion of advanced AI concepts.