· via Hacker News – Front Page (hnrss.org)
Ai2 open-sources AstaBrief 8B, a fast evidence-cited report model for science
Ai2 has released the weights, training data and pipeline behind AstaBrief 8B, a Qwen3-based model that writes cited science reports about 3.5 times faster than the Claude-backed mode in its Asta platform.

What Ai2 released
The Allen Institute for AI (Ai2) has open-sourced AstaBrief 8B, the model behind Fast mode in Asta, its agentic platform for scientific work. Given a research question plus excerpts retrieved from the literature, the model produces a report with citations attached. Alongside the weights, Ai2 says it is publishing the training data and a sample workflow that researchers can adapt to build reports from their own PDFs, making local report generation straightforward.
According to Ai2's blog post, which surfaced on the front page of Hacker News, AstaBrief also remains available inside Asta's report-generation feature, where it now sits next to a slower Thinking mode powered by Claude.
A training recipe built on real queries
AstaBrief starts from Qwen3-8B, with most of the effort going into post-training data, evaluation and the surrounding generation scaffolding. Ai2 considered reinforcement learning, noting that its own DR Tulu work showed RL with judge models in the loop can improve long-form generation for open-weight models. It ultimately chose supervised fine-tuning plus direct preference optimization instead, arguing that RL is unstable and costly while an SFT and DPO pipeline is cheaper, easier to debug and simpler to iterate on. That decision shifted the burden onto data quality.
The supervision came from genuine user queries collected through ScholarQA, the retrieval-augmented framework that underpins Asta's report feature. After removing bot and beta-tester traffic, discarding prompts too short to be meaningful, and running an LLM-based pass to catch non-English, non-scientific and privacy-sensitive submissions, Ai2 was left with about 90,000 research-oriented queries. Target reports were synthesized with the multi-step ScholarQA pipeline using Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini and GPT-4.1, yielding roughly 47,000 usable examples after quality filtering. DPO pairs came from a separate pool of queries, each paired with two competing reports and a preference label for one of them.
Speed from a one-pass pipeline
The larger speed gain came from restructuring how reports are written. AstaBrief drafts the finished report in a single pass from the query and retrieved snippets, skipping the snippet summarization and clustering stages that the Claude-backed mode depends on and dropping section-by-section composition. Ai2 reports that quality held up under this design. Across the full Asta pipeline, Fast mode averages 51.1 seconds per report versus 178.5 seconds for Thinking mode, roughly 3.5 times faster, and Ai2 measured generation time falling close to tenfold relative to the proprietary systems it tracked.
One caveat Ai2 itself raises: most training and evaluation was finished in 2025, so the proprietary models used for data generation and as baselines reflect that period's frontier, and the evaluation has not been rerun against newer systems.
Why it matters
The release is a solid data point for the argument that small, openly licensed models trained on carefully curated, domain-specific data can stand in for large commercial APIs inside narrow professional workflows. For research institutions, open weights carry a practical benefit beyond cost: AstaBrief can be self-hosted when a question would reveal sensitive or unpublished work that should not leave institutional infrastructure. The release also fits Ai2's wider push, through the NSF-backed OMAI initiative, to build fully open models shaped by how scientists actually work, and the accompanying training data gives other labs a reproducible foundation for similar domain adaptation.
- #open-source
- #ai2
- #language-models
- #scientific-research
- #qwen