deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Mistral Large 4 hits public preview: 1T-parameter open-weight model with weights due this month

Mistral has opened a public preview of Mistral Large 4, a 1.05T-parameter multimodal open-weight model with strong cybersecurity and coding scores, and promises to release the weights by the end of October.

Mistral Large 4 hits public preview: 1T-parameter open-weight model with weights due this month

Mistral AI has put its next flagship model, Mistral Large 4, into public preview. The model — unofficially nicknamed "le Chonk" by the company — is a roughly 1-trillion-parameter open-weight system, and Mistral says it will publish the downloadable weights by the end of the month. A preview API is available now through Mistral Studio, according to the company's October 6 announcement.

A trillion-parameter mixture-of-experts model

Mistral describes ML4 as its largest and most capable model to date: a natively multimodal system with 49 billion active parameters. Mistral's documentation puts the exact figure at 1.05 trillion total parameters arranged in what the company calls a granular Mixture-of-Experts architecture, alongside a 1.6-billion-parameter vision encoder and a 1 million-token context window.

The model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs inside Mistral's own European datacenters, and the preview is served from that same infrastructure. Mistral says a significant share of the training data was multilingual, covering more than 160 languages, including every official EU language. The company also plans a deployment it operates end-to-end in Europe under European law — part of a sovereignty pitch aimed at organisations that want to self-host advanced AI rather than depend on US providers.

Benchmark results, led by cybersecurity

Mistral's headline claims centre on security work. On the Artificial Analysis Cyber Index, an independent evaluation of how well models find and fix flaws in real software, the company says ML4 ranks among the top five models globally and leads open-weight models developed outside China. On one index task — reproducing a real vulnerability in open-source software and then patching it — ML4 scored 82%, which Mistral calls the highest of any model. The company notes that several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on that test because they refuse the task, while ML4 also solved 93% of the Cybench challenge set.

On agentic coding, Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4.0, for a combined Coding Agent Index score of 49.8% — ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. A caveat: those three harness scores were evaluated privately by Artificial Analysis before the harness's public launch. In a blind human evaluation run with Surge AI, professional annotators rated ML4 Preview second of five models at 3.74 out of 5, behind Claude Opus 5 (4.22) and ahead of Kimi K3, GLM-5.3 and GLM-5.2.

For general agents, Mistral cites 59.9% on AutomationBench, a suite of 657 business workflows across tools like Gmail, Slack and Salesforce, and 1,393 Elo on AA-Briefcase for long-horizon knowledge work. In multimodal tests, the company says ML4 beats GPT-6 Astra on the Dense 200 visual grounding benchmark, 42% to 41%. All of these figures come from Mistral itself unless otherwise noted, so independent verification will largely wait on the weight release.

Availability, pricing and what comes next

Until the weights ship, Mistral says it is red-teaming the model in real-world settings with cybersecurity leaders, vetted partners and state authorities, who get access to the same model with reduced moderation and expanded cyber capabilities. The company argues this matters for defenders, since provider-level refusals in closed models can block legitimate vulnerability research and incident response.

Mistral's documentation lists pricing between $0.68 and $1.36 per million input tokens (with cached input at $0.07–$0.14) and $2.09 to $4.18 per million output tokens. Supported features include structured outputs, function calling, document question answering, batching and built-in tools for agents. Mistral says further architecture details, additional benchmarks and its post-training methodology will be published alongside the weights, and that ML4 will serve as the foundation for a new generation of specialised and optimised models.

Why it matters

If the weights land on schedule, ML4 would put a trillion-parameter frontier-class model into the open-weight ecosystem at a moment when the strongest open models increasingly come from China. Its cybersecurity positioning is the sharpest edge of the story: Mistral is explicitly marketing a model that will do offensive-adjacent security work that leading closed models refuse, which is a genuine capability for defenders and incident responders but also intensifies the dual-use debate around open weights. For European enterprises and governments, an end-to-end EU-operated deployment of a top-tier model is a concrete sovereignty option rather than a slogan. The open questions are whether the benchmark claims hold up once the weights and the new Artificial Analysis harnesses are public, and how the reduced-moderation red-teaming phase shapes the final release.

  • #mistral
  • #open-weights
  • #large-language-models
  • #cybersecurity
  • #multimodal

Related posts