· via Hacker News – Front Page (hnrss.org)
Multiverse Computing claims Europe's top AI benchmark score with new 438B reasoning model
Multiverse Computing's first large model, Quasar 438B, scores 43 on the Artificial Analysis Intelligence Index — the highest European result, per the company — and answers in 15.3 seconds.

Multiverse Computing has released Quasar 438B, a reasoning model it presents as its flagship offering for enterprise agents and coding. According to the company's announcement, which surfaced on the Hacker News front page, the model scores 43 on the Artificial Analysis Intelligence Index — the highest result of any European model, the company says.
A first large model from an efficiency-focused lab
Quasar 438B is the first large model Multiverse Computing has shipped. It works in English and Spanish and is designed for multi-step tasks that involve planning, tool use, code execution and large context windows. Rather than asking customers to provision infrastructure, the company serves the model through its CompactifAI API, so teams can begin testing by signing up and sending requests. Multiverse has built its position on making AI more efficient and deployable, and Quasar carries that approach into the 400-billion-plus parameter class while keeping response times lower than is typical for models of that size.
Benchmark position
The Artificial Analysis Intelligence Index v4.1.1 aggregates nine evaluations, among them GPQA Diamond, Humanity's Last Exam, SciCode, Terminal-Bench v2.1 and AA-LCR. Quasar's composite score of 43 puts it ahead of Mistral Medium 3.5 at 30 and Inkling at 42, as well as NVIDIA's Nemotron 3 Ultra. The announcement is inconsistent on that last comparison, listing Nemotron at 38 in one section and 36 in another, which makes the gap four to five points. The field overall is led by Claude Opus 5 at 63, leaving Quasar clearly behind the frontier group even as it tops the European field.
Speed as the practical differentiator
Quasar returns a 500-token response, including the time it spends reasoning, in 15.3 seconds. Only three models in the company's comparison are faster — Nemotron 3.5 Lightning at 9.4 seconds, Gemini 3.5 Flash-Lite at 10.8 seconds and Gemini 3.7 Flash at 11.5 seconds — and only that last one also scores higher, at 56. The gap over closer competitors is wider: Inkling, one point behind on the index, needs 48.3 seconds, more than three times Quasar's time, while Mistral Medium 3.5 takes 18.8 seconds for a score of 30. Among the models that outscore Quasar, only two respond in under 25 seconds; the rest take between 38 and 156 seconds.
Long-context and terminal results
On AA-LCR, which measures how well a model can pull together and reason over information distributed across long documents, Quasar scores 75.0 — level with Grok 4.6 (high) and within a point of Claude Opus 5 at 75.7 and Qwen3.8 2.4T A95B at 75.3. On Terminal-Bench v2.1, which evaluates agents operating in live terminal environments, it scores 69.3, leading Mistral Medium 3.5 by 18.7 points and Nemotron 3 Ultra by 15.4 while trailing a frontier group led by Claude Opus 5 at 89.1. The company acknowledges that terminal work is where the most headroom remains, calls it the focus of its next round of improvements, and says further updates are coming.
Why it matters
European labs have struggled to field models that compete with US frontier systems, and a European model outscoring Mistral's current mid-tier offering on a third-party composite index is a notable data point for the region's ecosystem. The latency-per-capability trade-off is arguably the more interesting claim: if the figures hold up under independent testing, a 438B-parameter model that answers in 15 seconds while scoring 43 would fit naturally into interactive agent products rather than batch workloads. The caveats are real, though. Every number here comes from the company's own announcement, drawn from Artificial Analysis charts it selected; the deficit to Claude Opus 5 stands at 20 points; and the agentic benchmark where Quasar trails most is precisely the one enterprise buyers care about. Treat this as a strong vendor claim awaiting outside verification.
- #llm
- #benchmarks
- #european-ai
- #reasoning-models
- #enterprise-ai