· via dev.to (home feed)
VIDRAFT's Darwin-180B-RSI takes first place on five Hugging Face leaderboards
Korean startup VIDRAFT has released Darwin-180B-RSI, a 180B-parameter mixture-of-experts reasoning model trained via a recursive self-improvement loop, with self-reported first places on five Hugging Face leaderboards.

What VIDRAFT released
Korean AI startup VIDRAFT has released Darwin-180B-RSI, an open reasoning model with 180 billion parameters built on a mixture-of-experts (MoE) architecture. According to a write-up on dev.to that cites AI Market Watch, the model currently occupies first place on five Hugging Face leaderboards spanning mathematics, science, general knowledge and multimodal reasoning. Every one of those scores is self-reported, and the write-up itself notes that no independent validation has taken place.
The dev.to article, published on 29 September 2026, lists the core specifications. Darwin-180B-RSI is not a from-scratch pretrain: it is an adaptation of Qwen3.8-Flash-Next produced through selective model merging, a technique that combines strengths from existing checkpoints rather than training new capabilities from zero. The merged model carries its 180 billion total parameters across 512 experts, with only 10 experts activated per inference request, and offers a context window of roughly 260,000 tokens. VIDRAFT describes it as an open model, with the reported results hosted on Hugging Face.
The sparse routing design matters for deployment: per-token compute stays well below what a dense 180B model would cost, although the hardware still needs enough memory to hold every expert's weights.
How the self-improvement loop works
The RSI in the name stands for recursive self-improvement, the training method VIDRAFT says it applied. As the dev.to write-up describes it, the model generates candidate solutions to reasoning problems, an automatic check compares those solutions against verifiable ground-truth answers, and only the confirmed-correct outputs survive. The retained outputs then retrain the model, and the generate, filter and retrain cycle repeats, in principle letting the model bootstrap better reasoning from its own verified work.
The article places this in the same family as rejection sampling fine-tuning and outcome-reward methods such as STaR and GRPO, which train on a model's own outputs instead of human-labelled data. Two significant unknowns remain. VIDRAFT has not disclosed how many iterations the loop ran, what the filtering criteria were, or whether any reward modelling was involved. And because the approach depends on answers that a machine can verify, it fits mathematics and formal science far more naturally than subjective domains.
The reported scores
VIDRAFT's reported results, as relayed by dev.to, are:
- AIME 2026: 100%
- HMMT 2026: 100%
- GPQA Diamond: 94.44%
- MMLU-Pro: 88.12%
- MMMU-Pro: 79.48%
Perfect scores on two competition mathematics benchmarks are the eye-catching part of the announcement, and they are also the part that most deserves scrutiny. The dev.to write-up explicitly warns that leaderboard rankings measure performance on fixed datasets and may not transfer to production workloads, novel problem distributions or domain-specific tasks. Anyone considering the model should treat these numbers as a prompt to run their own held-out evaluations rather than as settled fact.
Availability and self-hosting
The write-up describes Darwin-180B-RSI as an open model on Hugging Face but could not confirm the exact repository path, and it mentions no GitHub repository, API endpoint or OpenAI-compatible serving URL. On the hardware side, the MoE design keeps active-parameter compute modest, but serving still requires memory for all 180B parameters; the article suggests quantisation or expert offloading through tools such as vLLM or llama.cpp may be necessary depending on available hardware.
Why it matters
If VIDRAFT's approach holds up under replication, it is a useful datapoint showing that measurable reasoning gains can come from a merge-and-retrain pipeline on an existing base model, without a full pretraining run and without human annotation. That would lower the cost of building strong open reasoning models considerably. If it does not hold up, it is a reminder of how quickly self-reported leaderboard claims, particularly perfect scores, should be met with independent testing. Either way, the practical advice is the same: evaluate the model on your own tasks before betting on it.
- #mixture-of-experts
- #open-weights
- #llm
- #benchmarks
- #model-merging