· via Hacker News – Front Page (native)
Samsung Labs' LittleBit shrinks LLM weights to 0.1 bits per weight
Samsung Labs has open-sourced LittleBit and LittleBit-2, which compress LLMs to between 0.1 and 1.0 bits per weight by factorizing weight matrices into binarized low-rank factors.
A sub-1-bit approach from Samsung Labs
Samsung Labs has published the reference implementation for LittleBit, a compression scheme for large language models that operates below one bit per weight. The GitHub repository, which surfaced on the Hacker News front page, covers two papers: the original LittleBit, a NeurIPS 2025 paper by Banseok Lee, Dongkyu Kim, Youngcheon You and Youngmin Kim, and the follow-up LittleBit-2, an ICML 2026 paper from Lee and Youngmin Kim.
Factorize first, then binarize
Most quantization work rounds each weight to a small number of levels in place. LittleBit instead decomposes every dense weight matrix into low-rank latent factors and binarizes those factors, so each stored entry is effectively a sign. Small learned scales then reintroduce the magnitude information that binarization discards, and according to the repository this is what lets the method reach effective precisions from 1.0 down to 0.1 bits per weight. The project also states that the original model architecture is preserved at inference time, meaning the compressed layers act as drop-in replacements rather than a new architecture.
LittleBit-2 fixes the initialization
The sequel targets a subtlety in how training starts. Factors produced by singular value decomposition do not naturally line up with the corners of the binary hypercube where binarized values live, and the repository describes this latent geometry misalignment in the initialization stage as a weakness of the original recipe. LittleBit-2 applies Internal Latent Rotation with Joint Iterative Quantization (Joint-ITQ), rotating the SVD-derived factors into alignment with that hypercube before quantization-aware training begins.
The change is opt-in through a --use_itq flag and touches only initialization: the deployed factorized layer stays the same, so the improvement adds no inference overhead.
Running it
The codebase is training-based rather than a post-training shortcut. Quantization-aware training uses a SmoothSign quantization function with optional residual factorization, and the example commands run five epochs on the repo's c4_wiki dataset, with DeepSpeed offered for multi-GPU setups.
Model coverage is broad: OPT, Llama and Llama 2/3, Phi-4, Qwen2.5 and QwQ, Gemma 2 and Gemma 3, and Qwen3. The recommended environment is Python 3.12 with the CUDA 12.4.1 toolkit and PyTorch 2.8.0, and the authors advise pinning transformers to the 4.51.x line when reproducing paper results, since newer releases may change model internals or evaluation behavior. A separate evaluation script reports perplexity on WikiText-2 and C4 and runs zero-shot benchmarks including BoolQ, PIQA, HellaSwag, WinoGrande, ARC and OpenBookQA, against either a local checkpoint or a model hosted on the Hugging Face Hub.
One licensing caveat: the project ships under CC BY-NC 4.0, which rules out commercial use without a separate arrangement.
Why it matters
Memory capacity and bandwidth, more than raw compute, decide which models can run on laptops, phones and edge devices, and quantization is the main lever for shrinking that footprint. A scheme that holds quality at 0.1 bits per weight changes the arithmetic considerably: a seven-billion-parameter model would in principle need well under a gigabyte for its weights at that rate, before accounting for the learned scales, activations and KV cache.
The trade-offs are real. This is a quantization-aware training pipeline, so it needs the compute to retrain a model rather than a quick post-training pass, and the non-commercial license limits direct product use. The repository itself publishes no headline benchmark numbers, so the honest test of the quality claim at these extreme bit rates is running the provided evaluation suite, or reading the NeurIPS and ICML papers, before betting a constrained-hardware deployment on it.
- #llm
- #quantization
- #compression
- #machine-learning
- #open-source