· via Hacker News – Front Page (hnrss.org)
Strata runs 125B-parameter Qwen3.8 Flash Next on consumer GPUs with 12GB VRAM
Open-source engine Strata serves the 125B-parameter Qwen3.8-Flash-Next model from a gaming PC, with local OpenAI- and Anthropic-compatible APIs and support for 12GB graphics cards.
Strata, a free and open-source inference engine published by GitHub user Niko1221, has reached the Hacker News front page with a bold pitch: run Qwen3.8-Flash-Next, a 125-billion-parameter model that the project says normally needs a server, on an ordinary gaming PC. The tool ships one-click installers for Windows and Linux, a built-in chat app, and OpenAI- and Anthropic-compatible APIs served entirely from localhost.
The performance claim, in context
The Hacker News post title advertises "100T/s" on an RTX 4090, which in local-LLM shorthand reads as 100 tokens per second. That specific card does not appear in the project's own published numbers. According to the repository's README, the measured benchmarks cover an NVIDIA RTX 5070 (12GB) paired with a Ryzen 5 7600 and 64GB of RAM, and an AMD RX 9070 XT (16GB) with a Ryzen 9 3900X and 47GB of RAM.
On the RTX 5070, Strata writes answers at 94 tokens per second with the Q2_0 quantization, 79 at IQ2_XS, 62 at IQ3_XXS and 53 at IQ3_S, while ingesting a 32,000-token prompt at up to 2,650 tokens per second. The RX 9070 XT lands at 60 tokens per second for Q2_0. The README offers the rule of thumb that a token is roughly three-quarters of a word and that 60 tokens per second already outpaces reading speed. It also notes that cards with more VRAM run faster, estimating an RTX 3090 (24GB) at roughly 100-140 tokens per second — which makes the 4090 headline plausible even though the repository does not document it.
What Strata ships
The project's emphasis is packaging as much as raw speed. The installer detects the graphics card, asks which model size and context length to use, downloads roughly 70GB, and opens a local web app at 127.0.0.1:8080 with chat, a live system monitor and settings. For tooling, Strata exposes an OpenAI-compatible endpoint at /v1, Anthropic's API shape at /v1/messages for apps such as Claude Code, and the OpenAI Responses API at /v1/responses for Codex CLI — with any API key and any model name accepted. An MCP server lets AI coding assistants install, start and stop the engine themselves, and image input is an optional setup choice, though AMD cards on Windows cannot read pictures yet. By default Strata answers one request at a time; a "parallel": 2 setting enables concurrent responses at the cost of slower individual replies.
Hardware requirements and trade-offs
Strata asks for an NVIDIA RTX 20, 30, 40 or 50 series card or a range of AMD Radeons, with 12GB of VRAM minimum, at least 32GB of system RAM (64GB fits every model size), and about 80GB of free disk, preferably on an SSD. Startup loads 35-55GB into RAM and can leave the PC slow or unresponsive for one to three minutes, which the README flags as normal.
The speeds that fit small cards come from heavy compression. The fastest options in the benchmark table are the smallest quants, Q2_0 and IQ2_XS. A "Coder" build removes half of the model's experts and reportedly reaches 91% of the full model's SWE-bench Verified score, as measured by its authors, in order to fit 32GB machines — but it is weaker outside code, including Chinese and other CJK text. A "Swift 1.5" fine-tune shortens thinking time before answers. Larger quants such as Unsloth's UD-Q4_K_XL read mostly from the SSD during generation and fall to 7-8.5 tokens per second on a 64GB PC. Older GPUs, from Tesla P40 and V100 to GTX 10-series cards, work experimentally through community-written guides.
Why it matters
A model in the 100B-plus parameter class was, until recently, a server-class workload. Serving one interactively on a single consumer GPU — even at 2-bit quality — changes the economics of local AI: coding agents like Claude Code or Codex can point at a localhost endpoint instead of a metered cloud API, and sensitive workloads never leave the device. Strata's broader contribution may be the plumbing: hardware auto-detection, resumable downloads, guided setup and MCP control remove much of the friction that has kept local LLMs a hobbyist pursuit. The caveat is equally clear — the headline speeds on 12GB cards depend on aggressive quantization, so smaller machines carry a real quality trade-off, and the strongest numbers in the README are estimates rather than measurements.
- #local-llm
- #open-source
- #qwen
- #inference
- #consumer-hardware