deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (hnrss.org)

MacStories: M5 Ultra Mac Studio makes local AI agents fast enough for daily use

A MacStories hands-on review finds the M5 Ultra Mac Studio with 256GB of unified memory runs local AI agents at genuinely usable speeds, thanks to 50% more memory bandwidth and a much faster GPU.

MacStories: M5 Ultra Mac Studio makes local AI agents fast enough for daily use

A hands-on review published by MacStories and currently on the Hacker News front page makes a striking claim: the new M5 Ultra Mac Studio, configured with 256GB of unified memory, is good enough to run local AI agents as everyday personal assistants — quickly, quietly, and with no recurring cloud bill.

The reviewer, who describes himself as a longtime tinkerer with local models rather than a professional AI developer, spent four days testing the machine against its predecessor, an M3 Ultra Mac Studio with 512GB of RAM, and a large gaming PC built around an NVIDIA RTX 5090.

Where the gains come from

According to MacStories, the M5 Ultra is the first Apple chip to use UltraFusion to join two dual-die M5 Max chips into a quad-die package. For local model inference, the review homes in on two specifications: GPU compute and memory bandwidth. The 80-core GPU pairs each core with a Neural Accelerator, which Apple credits with up to 4.5x the peak GPU compute for AI compared with the M3 Ultra, while memory bandwidth climbs from 819 GB/s to 1.2 TB/s — a 50% jump. Unified memory still maxes out at 512GB, but that configuration is not due to ship until late October.

In practice, the review says the machine spends far less time ingesting a prompt before generation begins, and then streams text noticeably faster. In side-by-side tests using the Open Minis app on iOS, the M5 Ultra produced responses roughly 70% faster than the M3 Ultra on average, and the Qwen3.8-Flash-Next model cleared 100 tokens per second.

Agents that hold up over long sessions

Those numbers matter less than the workflow they unlock. The reviewer calls the M5 Ultra a major performance leap for local models powered by MLX, and has made Qwen3.8-Flash-Next his default in both Open Minis for iOS and Hermes Agent — assistants he says he now reaches for more than Siri AI. He also runs local models inside the Codex app on his Mac, either as primary threads or as subagents orchestrated by GPT-6 Astra. Because the GPU and memory bandwidth are higher, agents start replying sooner, stay responsive at larger context windows, and sustain long multi-turn loops without bogging down as a session grows — exactly the usage pattern that defines agentic AI.

A 99-day workload for zero dollars

The review's backdrop is a real project. Over the summer, the author built an internal app called Desk to organize 310 documents — notes, PDFs, clipped webpages and review chapters — for MacStories' iOS and iPadOS 27 coverage. A team of agents built on DeepSeek V4 Flash, with olmOCR handling PDFs, ran 24/7 for 99 days: transcribing WWDC sessions, extracting features from multiple sources, cross-referencing features to chapters, pulling details and bugs from screenshots, and syncing everything through the Notion API. Priced against the OpenAI or Anthropic APIs, that always-on workload would have been prohibitively expensive, the review argues; run locally on a Mac Studio, it cost nothing. The published review was written the traditional way, but its research stack was assembled by agents.

The caveats

The piece is frank about limits. An RTX 5090 still holds an edge thanks to higher memory bandwidth, and cloud frontier models remain both more capable and often faster than local ones. Local setups are fiddly and firmly experimental, and the hardware is expensive enough that a buyer could fund years of a premium AI subscription for the same money — something the reviewer says he would happily recommend to anyone who just wants a polished cloud product. His preference for the Mac Studio rests as much on its compact, quiet chassis and the macOS app ecosystem as on raw benchmarks.

Why it matters

Apple's unified memory architecture lets a single quiet desktop hold models that would not fit on consumer GPUs, and the M5 Ultra's bandwidth gains turn that capacity into interactive speed. If always-on personal agents can run at usable latencies for the cost of electricity, the trade-offs around privacy, vendor lock-in and per-token pricing start to look different. The M5 Ultra does not beat frontier cloud models — but according to MacStories it narrows the gap enough that local agents stop being a hobbyist curiosity and become a practical daily option.

  • #apple-silicon
  • #local-ai
  • #ai-agents
  • #mac-studio
  • #mlx

Related posts