deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (hnrss.org)

Rust-built PSSA model with plastic weights outlearns a matched transformer

A language model written from scratch in Rust pairs a recurrent state-space core with episodic memory and self-updating weights, and reports lower loss and about 12x faster CPU generation than a matched transformer.

Rust-built PSSA model with plastic weights outlearns a matched transformer

No attention, no ML framework

PSSA, short for plastic state-space architecture, is a small language model that abandons both the transformer design and the conventional deep-learning stack. According to the project's README on GitHub, which reached the Hacker News front page, the implementation is written in Rust with no PyTorch, TensorFlow or any other framework underneath: the linear algebra is hand-written, training can use a CUDA device, and every gradient is checked against a scalar CPU reference, with a reported maximum discrepancy of 2.98e-8.

The architecture has four notable parts. A recurrent state-space core processes text one token at a time, carrying a fixed-size state forward through learned continuous matrices instead of attending over a context window. An episodic memory bank of 512 slots, using Poincare-style hyperbolic retrieval with a bounded top-4 search, is written to and read from while the model runs. Plastic weights are updated on the fly: fast changes reinforce what works, novelty drives growth, and a refractory gate rate-limits overwrites so repeated contradictory input does less damage. Finally, a closed-form ridge-regression step periodically folds those fast updates back into the base transition matrix, an analogy the README draws to sleep consolidating learning.

Because the state is fixed in size, the cost per token does not grow with the length of prior text, unlike attention's quadratic scaling over the context.

The head-to-head numbers

The comparison pits PSSA against a standard transformer of roughly equal size (the baseline is listed at 1,541,120 parameters), trained on the same cleaned WikiText-103 pass with the same tokenizer, optimizer schedule and seed. Both runs were organized as 64 links of 200,000 encoded tokens, each resuming from the previous checkpoint so the cosine schedule and optimizer state continue across the whole run.

Over 12.7M training tokens, PSSA finished at 3.98 cross-entropy against the transformer's 4.43, a gap of 0.45 nats, or perplexity of 53.7 versus 83.7. The transformer spent its entire token budget reaching a loss PSSA had already passed at around the 2M-token mark.

On a held-out slice of 198,939 tokens neither run touched, the final checkpoints scored 3.997 versus 4.429 cross-entropy and 24.1% versus 18.0% next-token accuracy. Across every checkpoint of both runs scored on bounded windows of that slice, the transformer's curve never overtook PSSA's, and the held-out gap of 0.43 nats essentially matches the training gap, which the author reads as better generalization rather than harder memorization. One wrinkle is disclosed: the baseline's first session was cut off at link 43 by a notebook time limit and its loss CSV was lost, so those links were reconstructed from the session's run log before the chain resumed and finished.

About twelve times faster on CPU

Generating 200 tokens on the same CPU, prompt and sampler took 226 ms for PSSA versus 2,735 ms for the transformer. The README stresses that this CPU-to-CPU comparison is the fair one, while the training throughput figures are not hardware-matched: PSSA trained on a Kaggle T4 at roughly 900 tokens per second while the baseline ran CPU-only, so those numbers should not be read as an architectural result.

Stated limits

The author is blunt about scale. These are 1.5M-parameter models trained on 12.7M tokens, a research prototype rather than a rival to any production system. Text quality is poor for both at this size; the README samples 'a barget of the Prian Academy' from PSSA and 'a material circulation of the United States' from the transformer. Two experiments remain unmeasured: whether earlier skills survive a switch of corpus, and whether removing the memory bank changes the loss.

Running it and what the project needs

The project builds with a Rust toolchain supporting Edition 2024 via cargo build --release, with two direct runtime dependencies (ureq for dataset downloads and tokenizers for byte-level BPE); CUDA is optional and the CPU path always works. The main request is compute: everything so far ran on a free hosted notebook with an entry-level GPU in sessions cut after a few hours, and questions such as whether the gap holds at 10x or 100x the parameters, or whether the memory bank matters at scale, need real VRAM and allocations measured in days. The most wanted code contributions are kernel performance, a modern recurrent baseline to compare against, and evaluation beyond next-token loss.

Why it matters

Dominant language models are transformers, and their cost grows with context length. PSSA is a compact, framework-free test of a different bet: fixed-size recurrent state plus retrieval memory plus weights that keep changing after training. At toy scale it beats a matched transformer on both learning efficiency and inference cost, inside an auditable codebase with unusually candid caveats. Whether the advantages survive larger scale, modern recurrent baselines and memory ablations is precisely what the project cannot yet test, which is why it is asking for GPUs.

  • #rust
  • #language-models
  • #state-space-models
  • #open-source
  • #machine-learning

Related posts