· via dev.to (home feed)
DeepSeek 4.1 ships open weights: 670B-parameter MoE, 37B active, $0.55 per million tokens
DeepSeek released open weights and an enterprise endpoint for its 4.1 model: a 670B-parameter MoE activating 37B per token, with MTP-4x multi-token decoding and DualPipe 2.0 scheduling, at a claimed $0.55 per million output tokens.

DeepSeek 4.1 ships open weights and an enterprise endpoint
According to an engineering write-up published on dev.to on 3 October 2026, DeepSeek has released the open weights of DeepSeek 4.1 — also referred to as the DeepSeek-V4.1 Foundation series — alongside a corporate API endpoint, making the model globally available. The model is a sparse mixture-of-experts (MoE) system with 670 billion parameters in total, of which only 37 billion are active for each generated token. The post claims it matches closed frontier competitors, naming Anthropic's Claude Mythos 5.1 and OpenAI's GPT-6 Astra, while charging a small fraction of their API rates.
256 routed experts and one shared expert
The dev.to piece describes a router that scores each token against 256 routed experts and activates the top eight, alongside a single shared expert of 3.2 billion parameters that processes every token unconditionally. The shared expert absorbs general-purpose work — syntax, JSON and Markdown formatting, common grammar — so the routed experts can specialize in narrower domains such as mathematics, compiler engineering, contract law and distributed systems design.
Load balancing is handled without auxiliary losses. The post explains that older MoE designs used heavy auxiliary-loss terms to stop routers from overloading a handful of popular experts, at the cost of final model quality. DeepSeek 4.1 instead adjusts dynamic bias terms on the router's affinity scores during training and inference, keeping expert traffic balanced without a loss penalty. The net effect, according to the write-up, is a model with the memory footprint of a roughly 700-billion-parameter network and the per-token compute cost of a mid-sized one.
Speed claims: MTP-4x decoding and DualPipe 2.0
Two systems drive the throughput numbers. MTP-4x adds four hierarchical prediction heads during pre-training; at inference time the model can emit up to four probable tokens in a single forward pass and validate them in parallel — speculative decoding built into the weights rather than bolted on afterward. The post cites an acceptance rate above 88 percent on coding workloads and a sequential decoding rate above 240 tokens per second per node, which it frames as a fourfold improvement in decode speed.
DualPipe 2.0 targets the pipeline bubble that appears when MoE routing forces all-to-all communication between GPUs across nodes, leaving processors idle while tensors travel over the interconnect. While a node computes dense attention and the shared expert for the current batch, its network engines transmit and receive sparse-expert tensors for the next batch, and KV-cache circular buffers are aligned with the all-to-all communication channels. The write-up claims average streaming-multiprocessor utilization of 98.2 percent, close to the hardware ceiling.
Pricing, and what is unverified
The post puts output pricing at $0.55 per million generated tokens and characterizes that as roughly one-twentieth of what closed competitors charge. It also promises a practical deployment guide with asynchronous production code for teams that want to self-host.
Every hard number in this story — throughput, acceptance rate, utilization, pricing and the frontier-parity comparisons — comes from a single dev.to post. No independent benchmark results or official documentation were available to corroborate them, so the performance claims should be treated as the author's figures until third-party evaluations appear.
Why it matters
If even part of the reported efficiency holds up, the release changes the arithmetic of frontier-scale AI. Extreme sparsity plus native multi-token prediction attacks inference cost — the dominant line item in most AI deployments — and open weights let organizations with suitable GPU capacity skip per-token billing entirely. A managed endpoint alongside the weights gives developers a middle path: start on the API, then move to self-hosting as volume grows. For closed-model providers, an open alternative claiming comparable quality at a twentieth of the price adds the kind of pressure that has historically pushed API prices down.
- #deepseek
- #open-weights
- #mixture-of-experts
- #inference
- #llm