deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (hnrss.org)

Seven ESP32-S3 boards run a 1.58-bit BitNet LLM over an SPI daisy-chain

An open-source project runs a roughly half-billion-parameter BitNet model across seven ESP32-S3 microcontrollers, splitting its transformer layers over an SPI daisy-chain.

Seven ESP32-S3 boards run a 1.58-bit BitNet LLM over an SPI daisy-chain

A developer has published an open-source project that runs a BitNet-style large language model on a cluster of seven ESP32-S3 microcontrollers, with six of the boards each handling a slice of the network's transformer layers. According to the project's GitHub repository, which surfaced on the Hacker News front page, the design is a pipeline: the boards pass a single hidden-state vector down a chain, each one transforming it with its own four transformer blocks before handing it on.

How the cluster is wired

One ESP32-S3 acts as the master. It runs the BPE tokenizer and the token-embedding lookup, then transmits the resulting FP32 hidden-state vector over one SPI channel to the first compute node. Each of the six compute nodes executes four transformer blocks — node one covers layers 0 through 3, and the final node covers layers 20 through 23, for 24 blocks in total — and forwards the updated state to the next board. The last node sends the state back to the master, which applies the final RMS normalization, a language-model head tied to the embedding matrix, and greedy sampling to emit the next token.

The README describes the interconnect as a high-speed SPI daisy-chain, and the hidden state is the only data that crosses between boards.

The model size is stated slightly differently in the two places this story appeared: the Hacker News summary calls it a 0.4B-parameter model, while the README says a 0.5B model is sliced across the cluster.

Where the 1.58 bits go

The project relies on BitNet-style ternary quantization, which stores each weight as one of three values, working out to roughly 1.58 bits per weight. According to the repository, the attention projections (Q, K, V and output) and the MLP projections (gate, up and down) in every block use this ternary format, with rotary position embeddings and a KV cache held in PSRAM. Normalization layers run in FP16 scaled to FP32.

Not everything is ternary. Embeddings are quantized to INT4 and occupy about 14MB of flash, the final-norm weights sit in a dedicated 64KB partition, and the LM head reuses the INT4 embedding weights. The node firmware, built on ESP-IDF, includes a hand-written assembly multiply-accumulate kernel for the ternary linear layers plus lookup tables for additional speed. The attention runtime lives in a file named qwen_attention.cpp, which points to a Qwen-family model as the likely base.

Preparing the model

A suite of PC-side Python scripts does the heavy lifting before anything is flashed. One prunes the tokenizer vocabulary down to 32,000 tokens; another slices the embedding matrix to match; a third runs quantization-aware fine-tuning so the model adapts to ternary weights rather than being quantized after the fact. Further scripts pack the weights and tokenizer into flash-aligned binaries, and batch scripts flash multiple boards in parallel. The project is MIT-licensed and credits the BitNet concept along with two earlier ESP32 LLM efforts — a single-board quantized deployment and a distributed multi-node design — as inspiration.

Why it matters

An ESP32-S3 is microcontroller-class silicon: no GPU, limited on-chip memory, and inexpensive enough to embed in everyday hardware. Running any LLM on such chips is only plausible because ternary weights replace most of the multiply-accumulate work in transformer layers with simple additions, which is exactly the trade BitNet was designed to make. The cluster approach adds a second trick: by spreading 24 transformer layers across seven boards, the design sidesteps the memory ceiling of any single microcontroller while moving only a small hidden-state vector between them.

The caveats are real — the repository does not publish throughput figures, sampling is greedy only, and a model in the 0.4B-to-0.5B range has limited capabilities. Even so, the project is a working demonstration that fully local, offline text inference is drifting down toward commodity embedded hardware, with implications for devices where connectivity, power budgets or privacy rules out cloud-based models.

  • #edge-ai
  • #esp32
  • #bitnet
  • #llm
  • #quantization

Related posts