deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Meta releases Llama 4 Scout with native 1M-token context and Mixture-of-Depths routing

Meta's FAIR team has reportedly released Llama 4 Scout, an open-weights model with 105B total parameters, 24B active per token, and a native 1M-token context window.

Meta releases Llama 4 Scout with native 1M-token context and Mixture-of-Depths routing

Open-weights frontier model goes public

Meta's Fundamental AI Research team has released Llama 4 Scout, according to a post on dev.to, which describes it as the first frontier-class open-weights foundation model built natively on Mixture-of-Depths routing. The model carries 105 billion total parameters, of which 24 billion are active for any given token, and is reported to trade blows with leading closed reasoning APIs on coding, mathematics and long-context tasks while running on enterprise-class hardware rather than a data-centre cluster.

How Mixture-of-Depths changes the compute budget

The post lays out the motivation: a conventional dense transformer applies identical depth and compute to every token, whether it is a comma or a knotty line of code. Llama 4 Scout instead places a learned top-k routing gate inside each transformer block. At inference time the router judges each token's difficulty; trivial tokens are passed through residual shortcut paths that skip the heavy self-attention and feed-forward layers, while information-dense tokens are routed through the full reasoning stack. The claimed result is a 58% reduction in generation latency compared with dense 70-billion-parameter models, even as total capacity expands to 105 billion parameters.

A million tokens, natively

Rather than stretching context after the fact with RoPE interpolation techniques such as YaRN — which the post notes tends to trigger perplexity spikes and lost-in-the-middle failures beyond 64K tokens — Llama 4 Scout was reportedly trained from the first pre-training step on an eight-stage progressive curriculum that extends context to 1,048,576 tokens. Attention combines RingAttention with sliding-window multi-head latent attention (MLA). On the Needle In A Haystack benchmark across the full million-token span, the post reports 99.8% retrieval accuracy over 2,500 distinct document depths. In practical terms, the model is said to take in entire code repositories, multi-volume legal document sets or large scientific collections within a single prompt.

Reported benchmark numbers

The dev.to post, which says the evaluations were conducted by independent third-party auditing teams, lists the following results against two well-known commercial models:

  • MMLU-Pro: 78.4%, against 78.0% for Claude 3.5 Sonnet and 77.2% for GPT-4o
  • HumanEval: 92.4%, against 93.7% and 90.2% for the two competitors respectively
  • SWE-bench Verified: 48.6% of real GitHub issues resolved autonomously, versus 49.0% for Claude 3.5 Sonnet and 38.8% for GPT-4o
  • MATH 500: 81.2%, versus 78.3% and 74.6%

The pattern is near-parity on knowledge and coding, a slight edge on the agentic software-engineering and maths suites, and a clear win over GPT-4o on SWE-bench.

Availability and deployment

Base, instruction-tuned and quantized checkpoints — including official FP8 and INT4 AWQ variants — are said to be downloadable from the Hugging Face Hub and Meta's own model distribution portal. In 4-bit mode the model is reported to require 24GB of VRAM, putting it within reach of a single NVIDIA RTX 4090 or 5090. Runtime support has reportedly already merged into vLLM, Ollama, LM Studio and Hugging Face TGI, so self-hosted deployments across Kubernetes clusters should work without waiting on ecosystem catch-up. The post also credits the model with multimodal comprehension and frames the release as handing startups and data-sensitive organisations full data sovereignty with no per-token API charges.

Why it matters

If the reported specifications hold up, Llama 4 Scout would mark an open-weights milestone on two axes at once: a frontier-class model that is genuinely self-hostable on modest hardware, and native million-token context without the degradation that afflicts interpolated long-context models. That combination — no API bill, no data leaving the premises, whole-repository-scale prompts — is precisely what enterprises have said they want from open models. The caveats deserve equal weight, however. The details come from a single community post rather than Meta's official announcement channel, and the headline benchmark tables are near-ties rather than decisive wins, so the frontier-parity claim warrants independent replication before anyone rebuilds their stack around it.

  • #open-source
  • #meta
  • #llama
  • #large-language-models
  • #long-context

Related posts