· via dev.to (home feed)
Open-weight Darwin-180B-RSI tops Swiss legal exam benchmarks with zero legal training data
VIDRAFT's open-weight Darwin-180B-RSI leads the LEXam and LEXam-hard Swiss legal-reasoning benchmarks despite training only on self-generated math and science solutions, scoring 68.94 on the multiple-choice set.

Darwin-180B-RSI, an open-weight 180-billion-parameter model from the Korean startup VIDRAFT, has taken first place on LEXam and LEXam-hard, the two legal-reasoning leaderboards that Hugging Face lists as official benchmarks. According to a dev.to post describing the work, the model was never trained on legal data — its only practice material was math and science problems, chosen because their answers can be checked automatically.
A benchmark built from real Swiss exams
LEXam was assembled by researchers at ETH Zurich, the University of Zurich and the Max Planck Institute from 340 genuine law-school exams administered at Swiss universities, in German and English, with answers verified by legal experts. As the dev.to post explains, it does not test whether a model has memorized statutes; it asks the model to apply law to a set of facts and reason to a conclusion. It is hard: the multiple-choice section has four options, so guessing earns 25 points, yet the strongest entries had clustered near 50. LEXam-hard is a subset of the open-ended questions on which leading open models score worst.
The reported scores
On LEXam's multiple-choice set of 1,655 questions, Darwin-180B-RSI scored 68.94, well clear of the previous leader's 52.41 (DeepSeek-R1). On LEXam-hard's 518 open-ended questions it scored 45.72 against a prior best of 40.82 (Inkling). The post also cites figures the LEXam authors published for frontier models — GPT-5 at 62.65, Claude-4.5-Sonnet at 58.01, Gemini-2.5-Pro at 55.72 — and notes Darwin's headline 68.94 came from a majority vote over four samples. Even with a single sample it scored 60.54, above the reported Claude and Gemini numbers, though those baselines come from the benchmark authors' own paper rather than a re-run under identical conditions.
The evaluation protocol, per the post: four samples per question with majority voting and a 32,000-token thinking budget on LEXam, and a single sample on LEXam-hard, where the 60 answers that hit the token limit were regenerated at 120,000 tokens. Open-ended answers were graded by DeepSeek-R1-0528, the judge named in the benchmark's official eval.yaml — meaning an LLM, not a human examiner, scored the free-text portion.
Learning by checking its own work
The training method is the part researchers will scrutinize. The team used model-level recursive self-improvement: the model solves practice problems, each answer is verified automatically, and the model trains only on its own correct solutions; the improved model then becomes the solver for the next round. The problems were math and science only, no human-written solutions or reasoning traces were used, and no legal text appeared in the training set.
Why would that transfer to law? The team's reading, stated in the post, is that self-improvement on verifiable problems trains a habit of reasoning — break the problem down, check each step, commit to a conclusion — that carries into a domain the model never studied. They concede the causation is unproven, because the parent model was never measured on LEXam under the same protocol. One effect they did measure: the model reaches its parent's accuracy with roughly 11% shorter reasoning.
Model-level, not harness-level
The post contrasts the approach with Google's recently released RRSI, which keeps the model frozen and improves the prompts, tools and workflow around it — harness-level improvement. Darwin is model-level: the weights change, so the gain ships inside the model file and works for anyone who downloads it, with no harness required. The two are described as complementary.
Under the hood, Darwin evolves models rather than pretraining from scratch: it diagnoses a parent model layer by layer and expert by expert, transplants the strongest parts from several models, and fuses them weighted by diagnostic trust. The family spans more than 50 official models and 400-plus community derivatives (arXiv 2605.14386). The RSI step changed only about 0.02% of the parent's parameters — attention paths and shared experts — leaving all 512 routed experts, the router and the vision encoder untouched. A component called ZTC reads the model's internal state once, before generation, to estimate the probability an answer will be correct.
With the two law results, the model holds first place on seven Hugging Face official leaderboards, which the post says is the most of any organization on the Hub; Zhipu AI has four, and DeepSeek, Xiaomi and Moonshot AI three each. The other wins: AIME 2026 (100), HMMT Feb 2026 (100), GPQA Diamond (94.44), MMLU-Pro (88.12) and MMMU-Pro (79.48). Weights and evaluation settings are published on Hugging Face under FINAL-Bench/Darwin-180B-RSI.
Why it matters
If the scores hold up under independent evaluation, the result suggests something more interesting than a leaderboard shuffle: reasoning ability honed on automatically checkable problems can transfer to a domain the model has never encountered. That matters for anyone building specialist models, because it hints you may not need a domain corpus to get domain competence. It is also open-weight, so the capability travels with the file rather than a proprietary harness. The caveats are real — the results are self-reported in a single post, the open-ended grading was done by an LLM judge, and the team itself says it has not isolated why the transfer works. Treat it as a striking research signal awaiting replication.
- #ai
- #llm
- #benchmarks
- #open-weights
- #legal-tech