· via dev.to (home feed)
VIDRAFT's POCKET-Darwin-180B brings a 180B sparse MoE model to an 8GB-VRAM gaming laptop
VIDRAFT says POCKET-Darwin-180B, a 180B-parameter sparse MoE model, can run on a gaming laptop with 8GB VRAM by streaming expert weights from SSD through llama.cpp.

VIDRAFT has published details of POCKET-Darwin-180B, a 180-billion-parameter language model that it says runs on an ordinary consumer gaming laptop with 8GB of VRAM and 32GB of system memory. According to a post on dev.to, the model combines a sparse mixture-of-experts architecture, SSD streaming and the llama.cpp runtime to fit a network of that size onto hardware many people already own.
The post notes that a Chinese AI outlet covered the release under a headline the developer translates as "a 180B model moves into a gaming laptop," and says the resulting framing matches the project's own goal: the hardware needed to run very large models locally has dropped sharply. The stated target is not the data center but the offline, isolated machine.
How a 180B model fits into 8GB
Three mechanisms do the work, according to the write-up:
- Sparse Mixture of Experts (MoE): the model contains 180B parameters in total, but only a small share of them — a few billion, by the post's account — activate for any given token. Compute is spent on the active path rather than the whole network.
- SSD streaming: expert weights are kept on disk rather than loaded in full. The runtime streams in only the pieces a request needs, so system RAM and VRAM hold a working set instead of the entire model.
- llama.cpp runtime: a portable, widely used inference engine built for modest hardware and quantized weights.
Together, the developer argues, these techniques turn a model that would normally require server-class hardware into one that runs on a laptop the buyer may already have.
Slower, and that is accepted
The project is explicit that speed is not the selling point. In the post's FAQ, local inference on a laptop is described as slower than a tuned cloud endpoint. The compromise gives up throughput in exchange for reach: the value lies in being able to run at all in places no external API can serve.
Where it is meant to run
The intended users are organizations that cannot send data to an outside service in the first place. Defense, healthcare and legal workloads often run on isolated networks because of law or policy, the post explains. For those buyers, a model that runs fully offline on existing local hardware is less a convenience than a requirement, and the developer positions POCKET-Darwin-180B as serving a market distinct from the cloud API race.
Why it matters
If the claims hold, the release pushes the boundary of what consumer hardware can run locally: a 180B-parameter model within reach of an 8GB gaming laptop, at the cost of per-token speed. It also frames offline inference as a real deployment target rather than a demonstration. Capable local inference is no longer confined to cloud data centers; it now extends to disconnected hardware in secure rooms, and the question for buyers there is not which service is fastest but which model runs at all on site.
The details come from a single developer-published post, which reports the hardware setup but includes no benchmark figures, so the performance claims remain unverified. Even so, the recipe on display — sparse expert activation, disk-backed weights and a lightweight runtime — matches how other oversized models have been squeezed onto modest machines, and it signals where local inference is heading next.
- #local-inference
- #mixture-of-experts
- #llama-cpp
- #llm
- #on-device-ai