· via Hacker News – Front Page (hnrss.org)
StartLux 27B model lands second in CAICT MCP test, 1.3 points behind DeepSeek-V4-Pro
Chinese startup StartLux's first model, a 27B-parameter local agent built on Qwen3.6-27B, ranked second in CAICT's MCP benchmark — 1.3 points behind DeepSeek-V4-Pro's 1.6-trillion-parameter flagship.

A 27B dark horse in China's model race
A new Chinese startup called StartLux, formerly known as Yuandian Xinghui, has entered the large-model sector with an immediate statement. Its first release, StartLux-V1.0-27B-Preview, took second place overall in the China Academy of Information and Communications Technology's (CAICT) trustworthy AI benchmark for MCP, according to a chinaonchina.com report that reached Hacker News's front page on September 3, 2026. The only system ahead of it was DeepSeek-V4-Pro, a cloud flagship with 1.6 trillion parameters — roughly 60 times larger — and the gap between them was 1.3 percentage points.
The model is positioned as a local agent: it runs directly on consumer-grade PCs without relying on the cloud, which is central to the company's pitch.
What the benchmark measures
CAICT's MCP special item (the report also refers to it as MCP-Universe) sets six task categories around real application scenarios: location navigation, web search, browser automation, financial analysis, code repository management and 3D design. A comprehensive assessment sits on top, evaluating how well an agent coordinates multiple tools, executes complex tasks and interacts with real-world environments. The report's framing is that the test does not grade a model on its answers, but on whether it can actually get things done.
StartLux-V1.0-27B-Preview scored 39.25% overall. DeepSeek-V4-Flash-0731 (over 284 billion parameters) and Step-3.7-Flash (198 billion) also finished close behind the leader, in the same 1.3-point band, per the report. At equal scale, StartLux outscored Alibaba's Qwen 3.6 by 5.34 percentage points, and it placed first in three sub-categories: location navigation, financial analysis and browser automation.
Head-to-head with Claude Sonnet 4.6
The report illustrates the results with two comparisons against Anthropic's Claude Sonnet 4.6.
In a two-year Microsoft stock investment calculation, Claude produced $47,254 and 89.02%, while StartLux produced $47,499.09 and 90.00%. The difference, on inspection of the reasoning traces, came down to data handling: with raw data missing, Claude treated January 8, 2025 as a non-trading day and used the previous day's closing price; StartLux went back to the original data, verified market movements around the target date, confirmed the correct closing price, and only then completed the calculation and generated a visualization. The report notes that in finance, small discrepancies can lead to large losses.
The second task was a live flight search performed in a browser. StartLux completed it in roughly 95 seconds, finding an Air China ticket at $299 in 12 steps. Claude Sonnet 4.6 needed extra screenshot confirmations for date selection, popup dismissal and filter menus — at one point accidentally opening the time-filter panel — and after more than 200 seconds and 21 steps returned a lowest quote of $556.
Post-training on top of Qwen
StartLux-V1.0-27B-Preview was not trained from scratch. It is built on Qwen3.6-27B, and the report attributes the score gap to task-oriented training data and automated post-training that reinforces task understanding, tool selection, parameter construction, multi-step execution, status checking and result verification. In practice, the model learns not just what text to output, but when to invoke which tool, how to adjust when a tool returns an exception, and when a task can be declared complete.
The team says it independently developed a multi-dimensional, verifiable and scalable iteration method it calls Auto Research, in which the model autonomously executes tasks in a real tool environment and continuously adjusts strategy based on environmental feedback — an AI-trained-AI loop. The company claims it is the country's first local agent model to complete post-training this way.
Chen Danian's comeback
The founder and CEO is Chen Danian (the article romanizes his name inconsistently, also as Chen Danyan and Chen Dawei), described as a co-founder of Shanda Network, head of Lian Shang Network, and one of China's earliest programmers to introduce the concept of shareware. After roughly a decade away from the industry, he has returned betting on local models.
At the 18th anniversary reunion of Shanda Innovation Institute, he presented what he called eight non-consensus views for the AI era, four of which center on local models — including the prediction that local models will catch up with Claude within three years, occupy 80% of the market, and make parameter-count competition obsolete. On scaling, he described the Scaling Law as a passing fairy: without it, AI cannot take off, but its future is not necessarily tied to it.
Why it matters
If the numbers hold up, the result is a strong data point against the assumption that agent capability must be bought with parameters. A 27B model, fine-tuned from an open base with agentic post-training, landed within 1.3 points of a trillion-scale flagship on a benchmark that scores real task completion — and beat the model it was built on by more than five points. That matters for cost, privacy and offline use, since local execution removes cloud dependency entirely, and it reshuffles China's domestic race, where DeepSeek, Qwen and Step have been the reference points. Caveats apply: this is a preview release, a single benchmark and a single report, and the CAICT scores and Claude comparisons have not been independently verified here.
- #ai
- #benchmarks
- #local-models
- #china
- #startlux