· via The Verge
The Verge tests local AI: an open-source agent and a 125-billion-parameter Qwen model
The Verge's Antonio G. Di Benedetto shares early lessons from running an open-source agent with a 125-billion-parameter Qwen model locally, motivated by privacy concerns over cloud AI services.

The Verge has published the first installment of a testing diary in which reviewer Antonio G. Di Benedetto tries to make local AI an everyday tool rather than a novelty. The motivation is privacy: he writes that handing personal data to cloud services has been the main reason he avoided regular AI use, and running models on his own hardware offers a way around that.
The setup
Di Benedetto installed Hermes Agent, an open-source, self-hosted AI agent desktop app that runs on macOS, Windows and Linux and is free to use as long as it is paired with local models. The host machine is an M5 Ultra Mac Studio with 256GB of unified memory — enough, he notes, to run just about any model available. He chose Qwen 3.8 Flash Next, a 125-billion-parameter model roughly 105GB in size, reasoning that token costs disappear when a model runs locally, so there is little incentive to start small. He also plans to test smaller Qwen models on an M6 Mac Mini, an M5 MacBook Air, an Asus TUF Gaming A14 with an AMD Strix Halo processor, and forthcoming RTX Spark machines.
The hardware context matters here. According to The Verge, Apple is pitching its new Mac desktops partly on their local AI capability, and a line of RTX Spark Windows machines with up to 128GB of RAM is landing soon, aimed at agentic AI. Apple also demonstrated a cluster of four Mac Studios — nearly $50,000 of compute — wired up for local AI at a briefing Di Benedetto attended.
What the agent actually did
The first tasks were deliberately small. A daily 7:30am briefing, scheduled as a cron job and delivered through a Telegram bot, scans his email and calendar and appends a short weather report. It failed repeatedly at first, until he realized the Mac could not be asleep when the job fired. He concedes the briefing is not especially useful and does not require a $12,000 computer, but it now runs like clockwork.
A more substantial job was reorganizing a Steam library of more than 400 games. After registering a Steam web API key and granting it to Hermes — a permission he revoked once the work was done — the agent sorted the collection by genre within minutes while preserving his own categories such as favorites, co-op titles and party games.
Data analysis proved the privacy point most clearly. He had Hermes crunch personal financial records and build a spec-comparison spreadsheet for a laptop, the latter involving information under embargo. He writes that keeping that data on his own machine was the deciding factor in whether he would use AI for these tasks at all.
The biggest ongoing project is automating laptop benchmark testing, where a dozen or so tests are run three times each to produce averages. Even walking Hermes through the required procedures and coaxing usable Python scripts out of it has been an endeavor, and the work remains unfinished.
The rough edges
The diary's headline conclusion — exciting, overwhelming and frustrating — holds up in the details. The sheer number of available models, many with specialized uses, was the first source of overwhelm. More telling, even a 125-billion-parameter model does not add up to a magical assistant: after being fixed, the daily briefing broke again, multiple times. Di Benedetto treats Hermes and local AI strictly as a tool rather than a digital companion, and says he intends to stay vigilant and cautious as testing continues.
Why it matters
This is an early field report on a shift hardware makers are actively betting on: capable models that run entirely on a user's own machine, connected to an agent that can act on files, calendars and third-party APIs. For people whose barrier to AI has been data privacy — financial records, embargoed material, personal correspondence — local execution removes the central objection, and the Steam library and data analysis tasks show genuinely useful results. But the trade-offs are equally clear: significant spending on memory-rich hardware, a confusing model landscape, permission management for every integration, and reliability that still demands troubleshooting. Local AI is becoming practical for tinkerers willing to put in the effort, but it is not yet a hands-off substitute for cloud assistants.
- #local-ai
- #on-device-ai
- #llm
- #qwen
- #privacy
- #apple-silicon