· via Hacker News – Front Page (hnrss.org)
Open-weight inference keeps older NVIDIA GPUs earning years longer, Ornn Data analysis argues
An Ornn Data analysis of GPU rental markets argues that open-weight inference demand keeps older NVIDIA hardware economically productive, with the A100 undercutting the H100 on sparse workloads.
A rental-market study from Ornn Data, titled "The Economics of Open-Weight Inference" and published on 7 September 2026, has surfaced on Hacker News's front page. Its central claim runs against standard data-center accounting: demand for open-weight model inference is extending the economic life of older NVIDIA GPUs rather than stranding them each time a new architecture ships.
The core argument
GPUs are commonly depreciated on the assumption that every new NVIDIA generation renders its predecessor obsolete. Ornn Data pushes back on that thesis. Closed models are reachable only through provider-controlled endpoints and subscription allowances, so a posted per-token rate card effectively sets the marginal price of additional usage. Open-weight models, by contrast, can be deployed by anyone on compatible hardware, which lets price-sensitive customers route work to whichever GPU serves a given model most cheaply.
Across eleven open-weight and eight closed models on the Artificial Analysis Intelligence Index, the paper finds that the cheapest qualifying open-weight model, once cost is adjusted for intelligence score, completes a task at roughly one fifth the price of a comparable closed model. The advantage is not universal: closed models remain cheaper at some score thresholds, and the closed frontier still exceeds the open sample at the top of the range.
The cost numbers
According to the paper, self-hosting on rented hardware brings compute-only costs down to $0.12 to $0.35 per million output tokens at full utilization, based on published throughput figures and Ornn's 1 September 2026 spot rents.
Which GPU wins depends heavily on the workload. For dense Llama-2-70B serving, newer silicon is cheapest: the B200 lands at $0.39 per million output tokens in the base case, versus $0.68 for the H100 and an estimated $0.75 for the A100. For the sparse gpt-oss-120b — 5.1 billion active parameters, MXFP4 expert weights, small enough to have been served on a single 80 GB A100 — the ranking flips. The A100's base-case cost is $0.29 per million output tokens, less than half the H100's $0.64, and the A100 stays cheaper at spot and at three- and five-year term prices. Ornn cautions that the A100 and H100 sparse figures come from different third-party serving setups, and that the dense A100 row is estimated.
Rental-market signals
Ornn presents its own rental data as market evidence. The A100's five-year term price retains 80.2 percent of its one-month mark — for a contract that would run until the Ampere family is more than eleven years old — while Hopper parts retain 43.7 to 59.8 percent and Blackwell parts 53.8 percent. Separately, A100 occupancy rose from 74 to 90 percent between March and September 2026 even as listed rental capacity grew 13 percent and the spot index climbed 20 percent, so the family's strength cannot be explained by shrinking rental supply alone.
Caveats
The paper is candid about its limits. It does not establish that open-weight demand caused the A100's occupancy or price behavior; its forward term marks are analyst-produced indicators rather than executable quotes; and portability is imperfect. September 2026 OpenRouter snapshots listed eighteen to twenty-two providers for widely served open models, with output-price ratios up to 5.6 between highest and lowest, but quantization, capacity, reliability and migration costs all constrain substitution. The workloads best suited to older hardware — long-running agents, batch evaluation and parts of reinforcement learning — tolerate latency and are hardware-agnostic, unlike interactive chat. Electricity at nameplate power, the paper adds, runs a few percent of spot rent and does not change the rankings. The text also references NVIDIA's acquisition of Hugging Face, dated 3 September 2026, in connection with its argument about demand routing.
Why it matters
If the thesis holds, the consequences reach both balance sheets and procurement plans. Operators and lessors that depreciate Ampere-era fleets aggressively may be writing off hardware capable of earning for years on latency-tolerant inference, and buyers shopping for open-weight serving can treat older GPUs as a genuinely competitive tier rather than salvage. The analysis also implies that the cost floor for usable AI keeps falling: when a roughly five-year-old accelerator undercuts a current-generation flagship on a capable sparse model, the marginal price of intelligence becomes a question of workload fit, not hardware age. That pressure lands squarely on closed-model rate cards — and on any forecast that assumes new silicon retires the old.
- #nvidia
- #gpu-cloud
- #inference
- #open-weight-models
- #cloud-economics