deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

MiMo-v2.6-Pro benchmarked on intelligence, speed and price by Artificial Analysis

Artificial Analysis has published its evaluation of MiMo-v2.6-Pro, combining intelligence index scores with throughput, latency and per-task cost data, and the results page reached Hacker News' front page.

MiMo-v2.6-Pro benchmarked on intelligence, speed and price by Artificial Analysis

What happened

A new model, MiMo-v2.6-Pro, has been put through Artificial Analysis' independent benchmarking process, and the results page — covering intelligence, performance and price — drew enough attention to reach the Hacker News front page on September 22, 2026. Artificial Analysis measures models against a standardized set of evaluations and publishes the scores alongside throughput, latency and pricing data, and the MiMo-v2.6-Pro page follows that template.

What the intelligence score is built from

The headline figure is the Artificial Analysis Intelligence Index, currently at version v4.3.2. It aggregates ten evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1.

Two of those are worth explaining because Artificial Analysis developed them in-house. AA-Briefcase v1.1 tests agentic knowledge work and reports a combined Elo score that blends a rubric pass rate with analytical-quality and presentation ratings. AA-Omniscience measures knowledge reliability: correct answers add points, hallucinated answers subtract them, and declining to answer carries no penalty. Scores on that measure run from -100 to 100, with zero meaning as many wrong answers as right ones.

The remaining evaluations cover agentic coding and terminal use, scientific code and demanding reasoning, among other areas. Beyond the single index, the page breaks out capability indexes for agentic knowledge work, agentic real-world tasks, coding and terminal use, professional document reasoning, and medical long-context reasoning. Artificial Analysis cautions that while intelligence tends to transfer across use cases, individual evaluations can matter more depending on the application.

The page also carries an Openness Index, a 0-to-100 normalized measure of how open a model is, and flags when a model's license restricts or prohibits commercial use of its weights.

Performance and cost metrics

Performance is reported through several lenses. Output speed counts tokens per second while the model is generating, based on its first-party API or the median across providers where no first-party API exists. Latency is measured as time to the first answer token, which for reasoning models includes the thinking phase before the visible answer begins. There is also an end-to-end figure for a 500-token response and a weighted average decode time per Intelligence Index task. Context window size, which matters for retrieval-augmented workflows, is listed alongside.

On the cost side, the central number is the weighted average price of running one Intelligence Index task. It is computed from input, cache-hit, cache-write, reasoning and answer token prices, divided by task count and weighted by each benchmark's share of the index. The analysis also totals what it would cost to run every evaluation in the index and shows per-million-token pricing, including the discounted rate for cached prompts.

Why it matters

Choosing a model usually means weighing capability claims from different vendors against invoices measured in tokens. What Artificial Analysis provides is a common yardstick: the same tasks, run the same way, priced the same way, so a model's intelligence score can be compared directly with its cost per completed task. That price-per-unit-of-capability view is often more decision-relevant than raw benchmark wins, since a cheaper model with a slightly lower score can still deliver more capability per dollar.

For teams building agentic applications, the performance columns matter as much as the scores. Time to first token and output speed determine how a model feels in interactive tools, and cost per finished task rather than per token is what a budget actually scales with.

The usual caveats apply: an index is an average across ten evaluations, and a model that scores well overall can still be the wrong pick for a specific workload, so the per-capability breakdowns deserve scrutiny before committing. The full results for MiMo-v2.6-Pro, including its scores and provider pricing, are on the Artificial Analysis model page.

  • #llm-benchmarks
  • #artificial-analysis
  • #model-evaluation
  • #mimo