deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Redis creator antirez ships ds4, a lean C engine for running DeepSeek and Qwen locally

Redis creator antirez has released ds4, a compact C inference engine that runs DeepSeek V4, GLM and Qwen models locally on high-memory Macs and CUDA or ROCm hardware using asymmetric 2-bit quantization.

Redis creator antirez ships ds4, a lean C engine for running DeepSeek and Qwen locally

What ds4 is

Salvatore Sanfilippo, the developer known as antirez and the original author of Redis, has released ds4 (DwarfStar 4), a compact inference engine written in C for running large open-weight models on local hardware. According to the project page at dwarfstar.sh, which surfaced on the Hacker News front page, the engine targets high-memory Macs along with CUDA and ROCm machines, ships under the MIT license, and deliberately supports only a handful of model families: DeepSeek V4 and V4.1 Flash, GLM 5.x, and Qwen3.8 Flash Next, in both text and vision variants.

The narrowness is the point. Rather than acting as a universal GGUF runner, ds4 validates each supported model layout end to end and expects users to fetch its own prepared GGUF files, built with asymmetric 2-bit quantization plus an imatrix. Generic GGUF downloads from elsewhere are explicitly not the target.

How a 284-billion-parameter model fits in local memory

The headline workload is DeepSeek V4 Flash, which the project describes as a 284-billion-parameter mixture-of-experts model that would ordinarily be served from remote infrastructure. ds4's approach is asymmetric quantization: the routed experts, which make up the bulk of an MoE model's parameters, are compressed to 2 bits, while shared and otherwise critical paths through the network are kept at higher precision. That combination, the project says, is what lets the supported builds fit the memory of the machines it targets.

A second design choice is a KV cache that persists to SSD. Long prompt prefixes are written to disk and can be resumed by prompt hash, so a restart or a revisited conversation does not force a full re-prefill. The architecture page also lists tensor parallelism, session batching, speculative decoding and vision input among the engine's capabilities.

One engine, three front ends

All three interfaces share the same model state and cache: ./ds4 is an interactive chat CLI, ./ds4-server exposes local OpenAI- and Anthropic-style APIs, and ./ds4-agent is a native coding agent intended for persistent sessions. Because the server speaks both API dialects, existing tools such as Codex, Claude Code and OpenCode can point their base URL at a local ds4 instance instead of a cloud endpoint.

Setup is deliberately short: clone the antirez/ds4 repository on GitHub, run the bundled download script to fetch a prepared model such as ds4f-q2, build with make (plus optional targets such as cuda-spark), then launch the CLI or start the server with a chosen context length.

Reported performance and hardware fit

The project's benchmark table lists reference figures for DeepSeek V4 Flash in Q2. On an M5 Max with 128 GB of memory, it reports 790.2 tokens per second of prefill and 39.4 tokens per second of generation at a 2,048-token context, falling to 398.5 and 27.6 respectively at 65,536 tokens. On a DGX Spark with 128 GB, prefill stays nearly flat — 825.8 tokens per second at 2,048 tokens and 823.0 at 65,536 — while generation runs at 18.1 and 13.8 tokens per second. These are self-reported numbers without third-party verification.

For hardware sizing, the project treats V4 Flash Q2 as the baseline configuration; at 128 GB of memory, GLM 5.3 in Q2 and Qwen in Q4 are also said to fit, while V4.1 Q2 streams its weights from SSD. The site additionally advertises running Qwen on a 64 GB machine.

Why it matters

A developer with antirez's track record in infrastructure software betting on local inference is a signal in itself. ds4 takes a contrarian position within the local-LLM ecosystem: instead of a general-purpose runtime that loads anything, it pairs a tiny C engine with a fixed, validated set of frontier open-weight models and pushes most of the compression burden onto the routed experts of those MoE architectures. If the quality holds up at 2 bits on the expert layers, the practical effect is that very large open models become usable on hardware people already own, and local coding agents gain a credible path off cloud APIs.

The caveats are real: the supported model list is short, generic GGUFs are out of scope, and the performance claims come from the project's own tables. But as a demonstration that frontier-scale weights can be squeezed onto a workstation with a deliberately opinionated tool, ds4 is one of the more interesting releases in the local inference space.

  • #local-inference
  • #llm
  • #open-source
  • #quantization
  • #deepseek

Related posts