· via dev.to (home feed)
Alibaba says Qwen3.8-Max ran 33 automated training rounds and targets 10-trillion-parameter models
Alibaba says Qwen3.8-Max completed 33 automated training and tuning rounds without human input, lifting its Artificial Analysis score while keeping the same model ID, and plans models with up to 10 trillion parameters.

What Alibaba claims
At its Apsara Conference in Hangzhou on 22 September, Alibaba said its flagship model Qwen3.8-Max had completed 33 back-to-back cycles of automated training and post-training tuning spanning more than a month, with no person stepping in. The claim comes from Alibaba Cloud's own press release, summarised in a dev.to write-up by Nokka, which says the automated runs covered pipeline design, data verification, repeated experimentation and diagnosing failures.
The reported result is a move from 40 to 45 on Artificial Analysis's Intelligence Index — while the model identifier stayed the same. Anyone calling the endpoint by its usual name would have been served a model reworked by a process nobody supervised. A related OrcaRouter post, also cited in the write-up, puts the model at 2.4 trillion parameters with 95 billion active.
Plainly, these are vendor numbers. No outside party has verified how the 33 rounds actually worked.
The parts outsiders can check
The weightier evidence sits in two case studies relayed in the dev.to piece, with figures attributed to coverage by The Satyajit.
The first is oh-my-cli, a command-line tool the model built starting from an empty repository. Left alone, the model assembled its own workflow: a state machine that moved tasks through stages, checks for builds, unit tests and end-to-end tests, and a feedback path that routed failures back to the originating task for another attempt. By 30 July, roughly 16 days in, the repository held 265 commits, 127 pull requests and 151 issues — with no human review or merge at any point. Because the code sits in a public GitHub repository, the history can be inspected rather than taken on faith.
The second case handed the model a single research paper and a set of GPUs, with no starter code or scaffolding. Over roughly 125 hours — around five days of unattended work — it wrote about 7,600 lines of code, performed more than 1,100 operations and ran 33 training cycles on the GPUs. It reportedly rebuilt the paper's pipeline from scratch and reproduced all six of its findings within the first 37 hours, then spent the remaining 88 hours generating and testing 18 ideas of its own. The final approach is said to beat the paper's baseline, with a cited figure of +7.7% on AIME24.
A road to 10 trillion parameters
Alibaba also laid out a scaling plan. Reuters reported, per the dev.to write-up, that the company intends to build Qwen 4.5 and Qwen 5 at 5–10 trillion parameters, with Qwen 4 already in training. The same event brought the launch of the Zhenwu V900 AI chip, and Alibaba Cloud's announcement also describes a 60-hour automated chip-design experiment — more than 10,000 calls to EDA tools and a 42% reduction in chip area — alongside a 20GW data-centre capacity target for 2032.
How to read the numbers
The dev.to piece flags three cautions. First, "self-evolution" here means automated rounds of training and tuning, not a model rewriting its own architecture; the term in the press release is narrower than the headline reading suggests. Second, a five-point gain on one benchmark index is a positive signal, not proof of across-the-board superiority over rivals with similar scores. Third, because the model identifier never changed, last month's evaluation may not describe what the endpoint serves today; teams testing models for production use should log dates and model IDs with every run.
Why it matters
Two stories collided in the final week of September. Reuters and CNBC reported that Anthropic's IPO filing warns advanced models could show self-preserving behaviour — resisting shutdown or concealing information — and that models realising they are being evaluated limits safety assessment. The same week, Alibaba was presenting month-long autonomous training loops as progress. Even read sceptically, the direction is clear: the cycle from training to tuning to testing is closing with fewer people inside it. That makes every benchmark a snapshot of a moving target, and it sharpens the question of who verifies systems that improve themselves.
- #qwen
- #alibaba
- #llm
- #automated-training
- #ai-chips