deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Gemini 3.8 Flash price doubles in January as Meta's Muse Spark 1.3 sells cheap tokens for training data

Google shipped Gemini 3.8 Flash with an intro price that doubles on January 1, 2027, while Meta's Muse Spark 1.3 offers a tier roughly 20x cheaper — paid for with training rights on your data.

Gemini 3.8 Flash price doubles in January as Meta's Muse Spark 1.3 sells cheap tokens for training data

On September 2, 2026 Google released Gemini 3.8 Flash, its third Flash model in six weeks, alongside Gemini 3.8 Flash Cyber, a security-tuned variant that cannot be bought through normal channels. Four hours later Meta published Muse Spark 1.3, which includes an endpoint up to about twenty times cheaper than its standard tier on one condition: Meta may train on those sessions. According to a dev.to roundup of both launches, the real story is not the benchmarks but the pricing fine print on each side.

What Google shipped

Google calls Gemini 3.8 Flash its best reasoning and coding model yet, at the speed and cost of the 3.7 version released three weeks earlier. Its claims include 54.9 percent on HLE-Verified plus wins on finance and legal agent benchmarks. Google's Logan Kilpatrick posted a 73.7 percent DeepSWE score, while the independent Datacurve leaderboard lists 74 percent — a small gap, but a reminder to check whose numbers you are reading. The Hacker News thread reached 1,157 points; Simon Willison's quick test, a request to "make me a cool thing in html", returned a particle simulation in 13 seconds for 1.8 cents, complete with a "60 FPS" counter that another user found was hard-coded into the page.

The price that doubles in January

Flash lists at $0.75 per million input tokens and $3.75 per million output tokens — as an introductory rate. A footnote in the launch post sets the expiry at December 31, 2026, after which the price doubles to $1.50 and $7.50. A Hacker News commenter quoted by dev.to put the problem plainly: users are being signed up to pay twice as much for a model that will no longer be frontier. Google says 3.7 Flash remains fully supported for efficiency-first workloads, which gives cost-sensitive buyers a fallback.

Cheap per token, expensive per task

On the independent DeepSWE v1.1 board (113 tasks, the same agent harness for every model), Flash ties Claude Opus 5 at 74 percent while costing $2.36 per task against $11.84 — arguably the strongest number in the launch. The trade-off is verbosity: Flash averages 166 steps and 143,000 output tokens per task, versus 61 steps and 60,000 tokens for GPT-5.6 Sol. Google's own explanation is that "3.8 Flash works harder", spending extra reasoning steps and iterative tool calls to maximize performance. Artificial Analysis found the same pattern from another angle: running its index on the model consumed 140 million output tokens against a median of 79 million, at a grading cost of $1,077.95. Output is quick at 280.8 tokens per second, but time to first token averages 12.74 seconds against a 3.33-second median — tolerable for background agents, painful for chat interfaces.

Flash Cyber sits behind a gate

Gemini 3.8 Flash Cyber is the same model with safety limits loosened for what Google calls trusted defenders. Sundar Pichai posted an 86.2 percent score on CyberGym, a benchmark for autonomous vulnerability discovery; Google adds 47.2 percent on CWE-Bench patching, just under the 47.8 percent it attributes to a leading frontier model, and says Wiz measured 7.5 to 9.7 percent better recall at 2.3 to 5.2 times lower cost. Access runs only through the new Fairwind Program: more than 650 vetted partners, restricted to internal security, incident-response and pentest teams, with multi-factor authentication and Google's CodeMender harness. The dev.to author reads the gating as the tell — a model that only patches vulnerabilities could be sold to anyone, so restricting it suggests Google believes it can also walk through the holes it finds.

Meta prices your transcript

Muse Spark 1.3's private endpoint costs $1.25 per million input and $4.25 per million output tokens — more than Flash's introductory rate. The contributor endpoint costs $0.10 and $0.20, roughly 12.5 times cheaper on input and 21 times on output, and the only stated difference is that Meta may use the data to improve its products. Mark Zuckerberg called it frontier performance almost too cheap to meter. Meta's own table puts Spark 1.3 at 75.4 on DeepSWE v1.1, ahead of Opus 5 at 74.0 — though that is Meta's figure, not the independent board — and claims about 20 percent fewer tool calls and 25 percent fewer tokens than Spark 1.2, the opposite design bet from Google's. Hacker News commenters noted this may be the first quantified price on training with user tokens, and asked whether secrets like cloud keys could be extracted from a model trained on customer input.

Why it matters

Per-token pricing is now the least useful number in a model announcement. Flash is cheap per token but spends more than twice the output tokens of rivals per task, and its rate doubles in four months; Spark's cheap tier is paid in data rather than dollars. Teams should benchmark on cost per finished task — logging tokens, steps and wall time — budget for the January 2027 prices now, and pin 3.7 Flash for workloads where it was already good enough. The price war between Google and Meta is escalating, but it is being fought with footnotes, verbosity and data rights rather than sticker prices.

  • #google-gemini
  • #meta
  • #llm-pricing
  • #ai-models
  • #security

Related posts