deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Saluki 27B: 2-bit, 7.89 GB, outscores its 54 GB parent on one bench, slips on math

Conway Research's 2-bit build of Qwen3.8-27B shrinks 54 GB to 7.89 GB, scores 88 vs 84 on the team's own tool-calling bench and runs on stock llama.cpp — but retains only 82-85% on competitive math.

Saluki 27B: 2-bit, 7.89 GB, outscores its 54 GB parent on one bench, slips on math

A 7.89 GB file that edged its 54 GB parent

Conway Research's Underdog team has released Saluki 27B, a 2-bit quantized build of Qwen3.8-27B that compresses the roughly 54 GB dense model into a single 7.89 GB GGUF file — 6.84 times smaller. According to a dev.to writeup, Saluki scored 88 of 120 on the team's own tool-calling benchmark while the full model scored 84.

The headline number is real but narrow. The same writeup collects three caveats the developers publish themselves, plus measurable regressions in competitive math, and those details determine what the file is actually good for.

What it is and how to run it

Per the dev.to article, the base model is Qwen3.8-27B: a dense model combining gated-delta-net linear attention with full attention, a 262,144-token context window and an Apache 2.0 license. Saluki is not a retrain. The team started from ISTA-DASLab's earlier 2-bit GSQ/RCO GGUF release and applied its own mixed IQ2 quantization, with calibration data selected to preserve tool-calling behavior.

The selling point is deployment: the file loads in stock llama.cpp, with no fork or separate runtime. The documented invocation is llama-server with --jinja, -ngl 99, -fa on and -c 32768, exposing an OpenAI-compatible chat API. Recommended sampling is temperature 0.6, top_p 0.95 and top_k 20 for general use, or temperature 0 when tool-calling speed matters. Thinking mode is on by default and can be disabled per request. The main file is text-only; vision add-ons of 629–928 MB are separate downloads.

Some coverage claims the file also runs in Ollama, LM Studio and vLLM, but as the dev.to author notes, the official page only emphasizes standard llama.cpp, and the article declines to settle the discrepancy.

The benchmark ladder

Underdog Bench is 120 tasks drawn from version 4 of the Berkeley Function Calling Leaderboard, split into five categories of 24 and frozen on September 27, 2026, before any model was tested. The reported scores:

  • Underdog main build, 7.9 GB — 91
  • Saluki, 7.89 GB — 88
  • Mia 2-bit, 9.1 GB — 85
  • Qwen3.8-27B full, about 54 GB — 84
  • ISTA 2-bit, 8.4 GB — 76
  • Ternary Bonsai 2, 5.95 GB — 70

On single-tool selection, Saluki scored 47 of 48 across the two simplest categories against Mia's 31 of 48, and the team says it and the main build beat the full model on 76–77 of the 84 tasks the full model solved.

The caveats the team publishes itself

First, the team's main build scores higher — 91 versus 88 — and Saluki was released anyway because the main build answers real web questions worse: 48 of 150 (32%) against Saluki's 60 of 150 (40%). Second, in parallel tool calling under strict checking, Saluki manages 42 of 100 while Mia leads clearly at 69 of 100; a lenient checker that forgives formatting moves Saluki to 55 and Mia only to 70. Third, the team states outright that it does not claim Saluki is smarter than the original, calls 120 tasks a modest test, attributes part of the gap to run-to-run variance, and notes that the base model's public scores came from a different harness.

Where quantization bites

Competitive math takes the biggest hit: AIME 2025 falls from 96.7 to 79.2 and AIME 2026 from 94.6 to 80.0, roughly 82–85% retained. SWE-bench Verified drops from 33 to 30 solved out of 50, MBPP+ from 83.9 to 78.0, and MuSR from 79.6 to 67.5. IFEval is the one inversion, at 93.5 versus 91.5. Letter-manipulation puzzles are the weakest task family, and the strict parallel-call checker could not parse 23 of 100 answers, mostly for formatting reasons.

Compatibility versus size

The dev.to article argues the real differentiator is the runtime, not the score. Mia ships in ExLlamaV3's EXL3 format and cannot load in llama.cpp at all. Ternary Bonsai 2 requires a PrismML fork of llama.cpp; stock llama.cpp rejects its PQ2_0 and PTQ1_0 files as unknown types, and a Q2_0 load silently produces garbage because the Hadamard rotation runtime is missing. Bonsai is 1.33 times smaller and claims 98.2% capability retention on its own 14-benchmark suite, but the two claim sets come from different tests and harnesses and cannot be compared directly.

As of October 11, Saluki showed 59,191 downloads within 30 days, against 4.46 million for Bonsai and 1.48 million for the ISTA original. AlphaSignal had reported only about 15,000 Saluki downloads a day earlier.

Why it matters

Saluki is a data point that 2-bit quantization, with calibration aimed at one behavior, can hold a commercially important capability — tool calling — at roughly one-seventh of the original footprint, on the runtime local users already run. Just as useful is the disclosure pattern: the team publishes where it loses, which makes the 88-versus-84 result credible rather than promotional. For anyone running local agents with about 8 GB to spare, this is a practical option; for math-heavy work, the full model or a higher-bit quant remains the safer choice.

  • #quantization
  • #local-llm
  • #llama-cpp
  • #tool-calling
  • #open-source

Related posts