deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (hnrss.org)

StepFun's Step 5 Preview reaches the Pareto frontier on Artificial Analysis benchmarks

StepFun's Step 5 Preview appears on Artificial Analysis's Pareto frontier, meaning no tracked model currently beats it on both measured intelligence and cost per benchmark task.

StepFun's Step 5 Preview reaches the Pareto frontier on Artificial Analysis benchmarks

What happened

A Hacker News submission flags Step 5 Preview, a large language model from Shanghai-based AI lab StepFun, as having reached the Pareto frontier on Artificial Analysis, the independent benchmarking service that compares model quality, speed and price across providers. The post links to Artificial Analysis's model page for Step 5, which carries the "On AA Pareto frontier" designation.

In Artificial Analysis's framing, that label means no other model it currently tracks offers both a higher Intelligence Index score and a lower cost per benchmark task. Anyone wanting more measured capability has to pay for it somewhere else in the chart; anyone wanting a cheaper task has to accept a lower score.

How the ranking is built

Artificial Analysis scores models with its Intelligence Index, currently at version 4.3, which aggregates ten evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1.

The mix is deliberately broad. According to the methodology text on the page, AA-Briefcase is an agentic knowledge-work benchmark whose Elo score combines analytical quality, presentation quality and rubric pass rates. Terminal-Bench 4.0 exercises agentic coding and command-line use, while separate capability indexes cover professional document reasoning and medical long-context work. AA-Omniscience probes knowledge reliability, rewarding correct answers and penalising hallucinations on a scale from -100 to 100, where zero means as many right answers as wrong ones.

Cost is computed per Intelligence Index task using a model's input, cached-prompt, cache-write, reasoning and output token prices, weighted by each evaluation's share of the index. The service also tracks output speed, time to first token and context window, but the frontier designation itself concerns the intelligence-versus-cost trade-off.

One detail makes the claim notable: Artificial Analysis charts open-weight and proprietary models together, marking whether weights are available and flagging licences that restrict commercial use. A frontier position for an open-weight model is therefore a comparison against the closed frontier labs' flagship APIs, not just against other open releases.

What remains unclear

The available source material is a benchmark listing and an aggregator headline rather than a StepFun announcement, so several practical questions are open:

  • The model is labelled a preview, and preview scores on third-party benchmarks routinely move as models are updated or evaluation runs are repeated.
  • The captured page text does not specify Step 5's licensing terms, parameter count, context window or hosting options, so the practical implications of the weights being available depend on details not present here.
  • The frontier is defined relative to the current field. Competitor price changes, or a new release from any lab, can knock a model off that frontier without the model itself changing.
  • Intelligence Index v4.3 is itself a recent revision of Artificial Analysis's methodology, and the service periodically recomputes scores when it updates its eval set.

Why it matters

A competitive open-weight model at the price-performance frontier is significant for two reasons. First, it applies pricing pressure on proprietary API providers: if a downloadable or cheaply hosted model matches flagship quality per dollar, the room for premium pricing narrows. Second, it gives developers a credible self-hosting or multi-provider option, which matters for cost control, data residency and avoiding lock-in.

It also continues a broader pattern in which Chinese AI labs have repeatedly shipped open-weight models that cluster near the top of independent capability-per-dollar rankings. A single benchmark listing is not proof, and a preview is not a product. But if Step 5 Preview's position holds under independent scrutiny, it would mark StepFun as a serious contender in the segment where open weights compete directly with the largest closed models.

  • #llm
  • #open-weights
  • #benchmarks
  • #artificial-analysis
  • #stepfun

Related posts