· via dev.to (home feed)
VIDRAFT ships 4-bit GGUF build of its 180B MoE model that runs on an 8 GB-VRAM laptop
VIDRAFT has released POCKET-Darwin-180B-GGUF, a 4-bit quantized version of its 180B-parameter Mixture-of-Experts model that runs on a gaming laptop with 8 GB VRAM via llama.cpp SSD streaming.

What VIDRAFT released
VIDRAFT has published POCKET-Darwin-180B-GGUF, a compressed release of its Darwin-180B-RSI foundation model designed to run on consumer hardware. According to a write-up on dev.to, which credits AI타임스 with the original report, the full model carries 180 billion parameters on a base VIDRAFT calls Qwen3.8-Flash-Next and occupies about 360 GB at BF16 precision — territory that normally requires a multi-GPU server.
The POCKET edition applies 4-bit quantization across the entire model, shrinking it to roughly 111 GB spread over four GGUF files. It went up on Hugging Face and ModelScope at the same time, and it targets on-premises, air-gapped and otherwise local deployments where there is no room for a cloud account or a dedicated inference server. The naming reflects VIDRAFT's product split: RSI (Recursive Self-Improvement) covers how the model was trained, while POCKET covers how it ships.
How an 180B model fits on a laptop
Three techniques stack up here.
The first is the Mixture-of-Experts architecture. Darwin-180B contains 512 expert sub-networks, but a learned router activates only 10 of them per token, so the arithmetic for each token corresponds to about 3 billion active parameters even though all 180 billion are stored. As the dev.to post explains, that gap between total and active parameters is what makes aggressive compression far less destructive here than it would be for a dense model.
The second is the 4-bit conversion itself. Weights move from BF16 to 4-bit integers, cutting storage from 360 GB to 111 GB, while the routing structure survives intact: the model still selects 10 of 512 experts per token, just with quantized weights.
The third is llama.cpp's ability to stream weights from an SSD on demand instead of holding the whole 111 GB in memory. Only the expert weights needed for the current token have to be resident at any moment, which is what makes an 8 GB VRAM plus 32 GB RAM laptop a viable host — the SSD effectively becomes one more tier of the memory hierarchy.
Reported performance
VIDRAFT's own testing, as relayed by dev.to, produced the following figures:
- Server CPU, one processor with 16 threads: 18.4 to 21.0 tokens per second, with 78.8 GB of memory in use.
- A mini-PC with 128 GB RAM kept the full model in memory for GPU-free, CPU-only inference.
- A gaming laptop with 8 GB VRAM and 32 GB RAM reached 4.17 tokens per second in SSD-streaming mode.
On accuracy, VIDRAFT reported an MMLU-Pro score of 87.65% both before and after quantization on a matched set of 2,000 questions. The dev.to post labels this a self-reported result, encourages independent replication, and cautions that quantization effects can vary by task type and prompt style.
How to run it
The GGUF files are public on Hugging Face and ModelScope. Prerequisites are a recent llama.cpp build and about 111 GB of free disk. After pulling the files with the Hugging Face CLI, the model runs through llama-cli: setting --n-gpu-layers to 0 forces pure CPU inference, while a positive value offloads layers to a GPU. The dev.to guide defers to the model card for exact file names, recommended context length, chat template and the mmap settings that matter for SSD streaming on low-RAM machines.
Why it matters
This is a real marker for local AI. Until recently, 180B parameters meant server hardware; 4.17 tokens per second on a laptop is not fast, but it demonstrates that very large MoE models — whose sparse activation already trims the compute bill — can be combined with 4-bit quantization and on-demand weight streaming from disk to run in settings where a 7B dense model was previously the practical ceiling.
For enterprises, government bodies and regulated industries that cannot let data leave the network, a fully local 180B-class model with no self-reported benchmark loss is a genuinely notable option. The caveats are real too: the numbers come from VIDRAFT rather than independent testers, real-world throughput will hinge on SSD speed, and a 111 GB download is still a serious commitment. Even so, the gap between frontier-scale capability and consumer hardware keeps closing.
- #local-ai
- #gguf
- #llama-cpp
- #mixture-of-experts
- #quantization