deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (hnrss.org)

Tiny Audio trains a 1.8% WER speech-to-text model for about $25

An open-source project called Tiny Audio claims a speech-to-text system trainable for around $25, reaching 1.8% WER on LibriSpeech test-clean by linking a frozen speech encoder to a Qwen LLM.

Tiny Audio trains a 1.8% WER speech-to-text model for about $25

A $25 speech-to-text stack

An open-source project called Tiny Audio, which surfaced on the Hacker News front page, offers a speech-to-text system its author says you can train for about $25. According to the project's GitHub repository, the published model reaches a 1.84% word error rate (WER) on LibriSpeech test-clean and 7.42% pooled across 12 benchmarks covering 11,822 samples — while training only about 80 million parameters.

The pitch is as much educational as practical. The codebase is described as small enough to read in an afternoon, and a real training loop runs on a laptop in roughly five minutes before any GPU hardware is rented.

Frozen encoder, small projector, LLM decoder

The architecture follows the pattern of attaching a language model to a perception frontend. Audio at 16 kHz passes through a pretrained speech encoder to produce frame embeddings. A small MLP projector stacks neighbouring frames and maps them into the LLM's embedding space — it is the only component trained from scratch. The LLM then reads those projected frames as if they were tokens and writes out the transcript.

Two recipes ship with the repo, and the repository notes that encoder, projector and decoder are each swappable from config:

  • The published model uses a Granite Speech 470M encoder and a frozen Qwen3.5 decoder with LoRA, training the projector plus LoRA adapters (~80M parameters).
  • The default course recipe uses a GLM-ASR-Nano 635M encoder and a fine-tuned Qwen3-0.6B decoder.

Because the decoder is a language model, output arrives punctuated, capitalized and with numbers formatted, with no post-processing step. Weights are bf16, requiring roughly 6 GB of GPU or Apple Silicon memory, and the model runs through the standard Hugging Face automatic-speech-recognition pipeline (published as mazesmazes/tiny-audio, with trust_remote_code).

The benchmark picture

The repository reports per-dataset numbers measured with its own eval tooling against its HTTP API on an RTX 4090. Clean read speech scores well: SPGISpeech at 2.29% and TED-LIUM at 3.71% follow LibriSpeech test-clean. Harder, noisier material pushes the numbers up — LibriSpeech test-other at 6.38%, Common Voice at 7.18%, GigaSpeech at 9.06%, Earnings22 at 10.58%, People's Speech at 17.59% and AMI's single-distant-microphone setup at 23.53%. The mean across the 12 sets is 8.71%. Two datasets (LoquaciousSet and Earnings22) were held out from training entirely.

The same eval harness can push identical samples through commercial APIs — AssemblyAI, Deepgram, ElevenLabs and Apple Speech — so users can verify the comparison themselves rather than trusting a vendor-supplied table.

More than plain text

The pipeline supports word-level timestamps via forced alignment and speaker diarization, implemented as a zero-shot port of NeMo's masked_asr recipe: each speaker found by Nemotron-3-Diarization is transcribed separately on a copy of the audio where everyone else is silenced. The repository is candid about limits — overlapping speech is only partly handled, with another speaker's words inside a turn remaining in that speaker's stream. Token-by-token streaming is also available.

A bundled ta serve command puts the model behind a batched HTTP server, where concurrent requests share GPU batches; the repo reports roughly 460x real time at 128 concurrent requests on an RTX 4090, with 45 minutes of audio transcribed in about 10 seconds (21 seconds with speaker labels). It runs on RunPod, CUDA, Apple Silicon or CPU.

Training tiers and a free course

Training is tiered. A free smoke test trains on 73 LibriSpeech clips on a laptop. The course run — the basis of the $25 claim — uses roughly 250 hours of audio on a rented GPU for a few hours. A production recipe over about 3 million clips across ten corpora needs an 80 GB GPU for a day or more and costs hundreds of dollars. A free 3.5-hour course walks through the full loop, though it trains the smaller recipe, so student results will trail the published table.

Why it matters

High-quality ASR has been dominated by large commercial APIs and models too expensive for individuals to reproduce. Tiny Audio demonstrates that a frozen-encoder-plus-LLM recipe with ~80M trainable parameters can hit competitive accuracy at a training cost almost anyone can afford, with a codebase small enough to actually understand. For teams, the swappable components and cheap training loop make domain-specific ASR realistic; for learners, it demystifies how modern speech systems are built. The caveats are real — all figures are self-reported, and the hardest benchmarks sit far above the headline 1.8% — but the accessibility argument stands on its own.

  • #speech-to-text
  • #open-source
  • #machine-learning
  • #asr
  • #llm

Related posts