deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

TurboGPT trains a 22KiB byte-level GPT in 13 seconds using CUDA C++

An MIT-licensed project on GitHub trains a tiny byte-level GPT directly in CUDA C++, packaging the full training loop with resumable checkpoints and TensorBoard-compatible logs.

TurboGPT trains a 22KiB byte-level GPT in 13 seconds using CUDA C++

A project called TurboGPT reached Hacker News's front page on 29 September with a striking claim in its title: a 22KiB transformer trained in 13 seconds. The GitHub repository behind the post, published by the developer lostmsu under an MIT licence, describes itself more modestly as a way to "train a tiny GPT in under a minute (CUDA only)".

A GPT written in CUDA C++

The unusual part is the implementation language. Most public GPT training code lives in Python on top of frameworks like PyTorch; TurboGPT is written directly in CUDA C++, NVIDIA's toolkit for general-purpose GPU programming. The model works at the byte level, reading raw bytes of text rather than a tokenised vocabulary, which removes the tokenizer from the pipeline entirely.

The build requirements are narrow: Windows, Visual Studio 2022's C++ tools and CUDA 13.4. The README also notes that a build setting called CudaArch must match the compute capability of your GPU, as listed in NVIDIA's own CUDA GPU documentation.

Training, checkpoints and logs

The tool runs from the command line:

.\build\turbogpt.exe --data hn1g.txt --log-to runs/ctx4

According to the repository, each run writes a checkpoint to runs/ctx4/ctx4.pt containing the model, optimizer, scheduler and trainer state together, and a --load flag pointing at that file resumes training from it. Despite the .pt extension, which PyTorch also uses, the README makes no explicit claim of PyTorch compatibility.

Logging is designed to slot into familiar tooling. A report. file is derived from the log directory, and the logs themselves are TensorBoard-compatible, with one report per batch, capped at 8Mi reports and flushed alongside periodic or final checkpoints.

The reported numbers

The README quotes one headline result: 2.52435 BPB on the hn1g dataset after 1.5G training tokens. Bits per byte is the standard loss metric for byte-level language models, so the figure measures how well the model predicts the next byte of text, with lower being better.

Beyond that, the published material offers few benchmark details. Neither the Hacker News title nor the README excerpt specifies which GPU produced the 13-second figure or the "under a minute" tagline, so the timings should be read as hardware-dependent claims from the author rather than verified results.

Why it matters

Most engineers encounter GPTs through several layers of abstraction: a Python API, an autodiff framework, pre-built kernels. A complete, working GPT trainer in raw CUDA C++ is rare, and a 22KiB model is small enough to read, understand and modify in a single sitting. For anyone learning how attention layers, optimizers and schedulers map onto actual GPU code rather than framework calls, it is a compact reference implementation.

The project also handles the unglamorous half of training infrastructure: resumable checkpoints that bundle optimizer and scheduler state, periodic flushing, and TensorBoard-compatible reporting. Those details separate a demo from something you can run repeatable experiments with, and their presence suggests the author prioritised reproducibility alongside speed.

Finally, byte-level modelling keeps the pipeline simple and makes metrics like the BPB figure directly comparable across datasets without tokenizer differences muddying the picture. As a small, fast, permissively licensed testbed for GPU training work, TurboGPT is less a product than a piece of instructive reading material, and that is precisely its value for AI engineers.

  • #cuda
  • #gpu-computing
  • #open-source
  • #language-models
  • #machine-learning

Related posts