· via dev.to (home feed)
ElderAI shares lessons from $100 coding-model fine-tuning; no run passed its quality gate
On dev.to, ElderAI recounts fine-tuning its coding model on about $100 of rented GPU time: pre-registered quality gates, a tool-use skill tradeoff, and watchdogs that keep each run to a few dollars.

ElderAI, a small team building ATLAS Code — a coding model aimed at agent tools such as Cline, Aider, Continue and Cursor — has published a candid writeup on dev.to about its recent fine-tuning work. The headline outcome is a negative result: none of the team's fine-tunes has yet cleared its own quality gate, so the invite-only preview still runs the starting checkpoint they are trying to beat. All recent attempts combined — several pilots, two full runs stopped midway and the latest gated run — cost about $15 of GPU time from a prepaid budget of roughly $100.
Quality gates written before the results
The team's first discipline is procedural. On a small budget it is tempting to pick whichever metric improved and call the run a success, so ElderAI now writes a small gate file before any metric is computed. The file is hashed and the launcher refuses to start if it changes. For the latest run, the gate required the fine-tune to produce more byte-exact correct files than the starting checkpoint on an edit test set (a tie fails), lose at most one problem on a standard Python coding benchmark, and keep at least 97% of tool calls parseable — all measured against the baseline in the same job with the same harness.
The gate has already prevented self-deception. In the latest run, the fine-tune was ahead on the edit metric at the 40% checkpoint but missed the tool-call parse bar by about one call in a hundred, so the run stopped. The team notes the starting checkpoint sat right at that bar in the same job, suggesting the threshold is very tight — but the rules forbid loosening it after seeing results.
Tool-use training can quietly erode coding skill
The team's initial instinct was to pour agent-style trajectories into training. Tool-call format improved, but general coding slipped just enough to fail the benchmark rule on several runs. Their countermeasures: mixing identical plain code-generation examples into each run as rehearsal, sticking to small LoRA adapters at low learning rates (a larger update touching MLP layers moved the edit metric more but is riskier), and evaluating a merged checkpoint at 40% of training, killing the run early if general coding has already dropped. They describe the tradeoff as the main thing to manage, not a bug to patch.
When the test set is the problem
An early edit metric demanded byte-exact matches after a tool call, and nearly everything failed it — including the baseline. Manual review showed why: instructions were real commit messages such as asking to increase spacing for quadrature encoders, where the actual commit changed a value from 3 to 6, which no model could infer. Rather than tune the scorer until numbers looked better, ElderAI changed what it measures: a mechanical check that each edit finds a unique match and changes the file, a whitespace-normalized exact match that still counts indentation, a split where instructions spell out the change precisely, and credit for files corrected within three tool calls with real tool errors fed back. Byte-exact remains the official number, with the other metrics reported for diagnosis.
Watchdogs and dollar caps
Every run sits under an on-box watchdog with a hard per-run dollar stop, a projected-cost stop, wall-clock and nightly caps, automatic stops for stalled logs, idle GPUs or NaN loss, and verification that the rented machine is actually deleted afterward. The projected-cost rule proved too aggressive — early ETA jitter killed two otherwise healthy jobs, a lesson priced at about $1.18. The biggest saving comes from the 40% midcheck: a failing run costs about $1.40 instead of about $3.
Next up is one more gated run using the precise-instruction edit data. If it passes, ATLAS Code gets the fine-tune and the team will publish what changed; if not, they say they will write that up too. The service remains in invite-only preview behind an OpenAI-compatible API with capped plans, no overage charges, and, per the team, no training on customer prompts or code.
Why it matters
Most fine-tuning discussion assumes budgets far beyond a small team's reach. This writeup is valuable precisely because it is a reproducible template for constrained training: pre-registered gates that survive contact with results, honest reporting of failures, the recognition that agent-format training can tax general skill, and cost watchdogs that make experimentation cheap enough to iterate. It is also a useful reminder that a harsh metric often says more about the test set than the model.
- #fine-tuning
- #llms
- #coding-agents
- #machine-learning
- #evaluation