· via Hacker News – Front Page (native)
OpenTPU: AI agents design an open-source accelerator that runs real language models
AI agents reportedly built openTPU, an open-source accelerator with RTL, ISA, simulator and compiler in one repo. It runs ten models on a Kintex-7 FPGA card, matching its simulator bit for bit.
AI agents built the whole accelerator stack
A project named openTPU, hosted on GitHub under the account FeSens and highlighted on the Hacker News front page, makes an unusual authorship claim: its entire accelerator stack was developed by AI agents. According to the repository, that stack spans the chip design itself in SystemVerilog, the instruction set, a bit-exact Python simulator, a kernel language with its compiler, and the host software that drives a real PCIe card. The project frames itself as an extension of an earlier effort called auto-arch-tournament, and asks two questions: how far agents can go at hardware design, and whether they can build the chip that runs their own inference.
The repository doubles as a teaching artifact. Everything lives in one monorepo meant to be read end to end, from a matrix multiply in Python down to the wires, and the host side ships chat, monitoring and profiling tools named otpu-chat, otpu-smi and otpu-lens.
Verified results on an FPGA card
The design is not confined to simulation. According to the project's page, ten recent models run with their real weights on an Inspur YPCB-00338 card carrying a Xilinx Kintex-7 xc7k480t FPGA with two DDR3 channels, and the card emits the same tokens as the simulator, bit for bit, in every configuration tested.
Throughput is modest but consistent. In int8, LFM2.5-230M decodes at 59.0 tokens per second, Qwen3-0.6B at 21.6 and Qwen3.5-0.8B at 17.6; larger models drop to single digits, with Phi-4-mini (3.8B) at 3.99 and Gemma 4 E4B at 3.78. Four-bit weights lift decode throughput by roughly 40 to 45 percent, pushing LFM2.5-230M to 85.8 tokens per second. While decoding, the card sustains 82 to 94 percent of its 17.1 GB/s DDR3 peak, which the project attributes to decode being memory-bandwidth-bound. Measurements dated 2026-09-29 to 2026-10-01 used a 133.33 MHz image with a single bitstream covering all models, hosted in a machine built around an Intel Core i7-4790.
A deliberately simple architecture
The machine avoids complexity on purpose. A sequencer issues one instruction per cycle, each instruction being eight 32-bit words, to a handful of units: a DMA engine, a four-column systolic matrix unit multiplying int8 weights streamed from DRAM, an fp32 vector unit, and a quantizer that converts results back to int8. There is no cache and no hidden scheduling; every data movement is an explicit instruction, so a trace shows exactly where the cycles go. The memory controllers are LiteDRAM instances calibrated at startup by a small CPU inside the memory core, in twelve seconds, without host involvement.
Quantization and expert offload extend the reach
Four-bit weights use FP4 values with two-level block scales, working out to 4.25 bits per weight, with the language-model head kept in int8 for accuracy. The project reports this cuts bytes per token by about a third, at a measurable perplexity cost documented per model.
Mixture-of-experts models larger than the card's 4 GiB of memory run with experts streamed from host storage. LFM2.5-8B-A1B (8.5B parameters, 1.7B active) decodes at 10.6 tokens per second, with 98.5 percent of expert uses hitting on-card slots; Qwen3.5-35B-A3B reaches 3.95 tokens per second, streaming 153 MB per token over PCIe at 1.41 GB/s. Both configurations match the simulator bit for bit. The host is otherwise nearly out of the loop: for several models the card runs one precompiled decode program, with the host adding only 0.17 to 0.30 milliseconds per token.
Why it matters
Hardware design has resisted automation more stubbornly than software, and this project claims agent authorship across the full vertical stack: not just RTL generation, but an ISA, compiler, simulator and host tooling that agree bit for bit between simulation and physical silicon. If the claim holds up under scrutiny, it is evidence that AI agents can handle the cross-layer constraint satisfaction that accelerator design demands.
The practical angle matters too. A decade-old CPU paired with an off-the-shelf FPGA card runs models up to 35B parameters, and the entire design is open and readable, making it a rare end-to-end reference for anyone learning how inference hardware actually works.
The caveats are real: the AI-authored claim rests on the project's own description, decode speeds sit far below commercial accelerators, and the results have not been independently verified. Even so, as a demonstration of agents operating across the hardware-software boundary, openTPU marks a notable milestone.
- #ai-agents
- #fpga
- #open-source-hardware
- #llm-inference
- #chip-design