deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Magnitude open-sources self-tuning inference engine claiming up to 2x faster than llama.cpp

YC S25 startup Magnitude has open-sourced an inference engine that compiles and tunes kernels on the user's own hardware, claiming up to 2x faster open-model agent inference than llama.cpp.

Magnitude open-sources self-tuning inference engine claiming up to 2x faster than llama.cpp

Magnitude, a startup from Y Combinator's S25 batch, has released its self-optimizing inference engine for AI agents as open source. According to the company's launch post on Hacker News, the engine compiles and tunes its compute kernels on the user's own hardware before a model runs, and the team claims this lets open-weight models run up to twice as fast as llama.cpp. The project is published under an Apache 2.0 licence and ships as a desktop app for macOS, Windows and Linux.

How the claimed speedup works

The company's core argument is that established local engines — it names llama.cpp, Ollama and LM Studio — distribute kernels compiled in advance for broad classes of hardware, which leaves performance unused on any specific chip. Magnitude instead compiles and tunes its kernels on the actual device before a model executes, so the generated code fits the exact hardware in the machine.

The headline "up to 2x" claim breaks down unevenly across platforms, though. According to the launch post, decode is 92% faster on Apple's Metal graphics API, which is where the near-doubling figure comes from, but only 19% faster on NVIDIA's CUDA. The company also says it writes hand-optimized kernels for the most popular open-weight model families, which it credits for beating generalist engines, and points to benchmark details and a supported-model list on its website. All of these figures are the company's own measurements; no independent results accompanied the announcement.

Built around agent workloads

Rather than positioning itself as a chat client, Magnitude treats local inference as infrastructure for coding agents and similar tools. The launch post claims 27% less memory per agent, with that memory released once an agent stops, and says concurrent sessions share prefix caches so that running several agents at the same time does not cause slowdowns.

The desktop app bundles the magnitude CLI, so no separate installation is needed. Getting started involves downloading the app, choosing a recommended model from a Discover tab, and connecting an agent from a Connections screen. One-click integrations are offered for Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi and Cline, and any other tool can connect through an OpenAI-compatible API.

Hardware and privacy

Magnitude runs on Apple Silicon, NVIDIA GPUs, AMD GPUs, or machines with nothing but a CPU. The company says there is no fixed minimum specification: smaller machines run smaller models, and more memory simply allows larger ones. On privacy, the launch post states that prompts, files and models stay on the machine, that there are no token costs, and that no internet connection is required once a model has been downloaded.

Why it matters

Agent workloads are unusually sensitive to inference speed, because agents make many model calls in sequence and per-token throughput directly shapes how usable a local agent feels. Magnitude's bet is that per-device kernel tuning, rather than shipping precompiled binaries, is a better way to squeeze performance out of the varied hardware people run local models on — and that open weights plus local execution will keep growing as agent usage does.

The Apache 2.0 licence and the OpenAI-compatible API keep switching costs low for developers who want to test that bet. The caveats are just as real: the performance numbers are self-reported, the "up to 2x" figure reflects the best case on Metal, and the 19% gain on CUDA is far more modest. llama.cpp, Ollama and LM Studio are entrenched with large communities. Whether on-device self-optimization is a genuine architectural advantage or a marginal optimization will only become clear when independent benchmarks appear.

  • #open-source
  • #llm-inference
  • #ai-agents
  • #local-ai
  • #developer-tools

Related posts