deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

AI model Jev scored below a coin flip predicting Bitcoin in dev.to test, losing to gradient boosting

A two-part dev.to test gave the AI model Jev the same indicators as a gradient-boosting baseline. Jev scored below a coin flip on next-candle BTC direction and leaned the opposite way from the market.

AI model Jev scored below a coin flip predicting Bitcoin in dev.to test, losing to gradient boosting

Can an AI model read a market? A two-part hands-on test published on dev.to answers no for one specific case: the model Jev, asked to predict the direction of Bitcoin's next five-minute candle, scored slightly below a coin flip while a plain gradient-boosting model fed the same numbers did measurably better. The more consequential finding is that Jev's probability estimates leaned the opposite way from both the trained model and the market itself.

How the test was run

The author, writing on dev.to, narrowed the question to a single binary call: after a BTCUSDT five-minute candle closes, will the next one finish above its open? A gradient-boosting model was trained on roughly 314,000 candles covering 2022 through 2024 and scored once on about 181,000 held-out candles from January 2025 to late September 2026. Fifty-one technical indicators were reduced to the twelve that mattered most on a validation window.

Jev, version 1.13.0, was tested on 1,500 randomly drawn candles from the test period, in two formats: raw numbers with definitions and a plain-English description. The data was blind — no symbol, dates or price levels — so the model could not recognise a famous day in Bitcoin's history. Each request also carried control questions with known answers, such as whether RSI sat above 50, to separate parsing failures from prediction failures. Jev answered 99.8% of those correctly. The whole exercise cost about $0.19 across 4,560 requests.

Below chance, and pointing the wrong way

On the same 1,500 candles, boosting reached an AUC of 0.522, against 0.486 for Jev with numeric input and 0.494 in plain text, where 0.500 is a coin flip. The author is careful about error bars: at 1,500 samples every confidence interval overlaps 0.500, boosting's included. What settles the interpretation is the direction of the signal. Boosting, fitted on three years of data, learned short-term reversal — price stretched far from its short moving average tends to snap back. Jev moved the other way, assigning higher up-probabilities exactly when price was stretched. The two models' predictions correlate at −0.71, close to opposites.

The market sided with the statistical model. Far below its nine-period EMA, 52.8% of next candles closed up; far above it, only 46.9% did. Jev's average prediction ran from about 41.6% to 54.4% across that same range, in the reverse direction.

The filter problem

Earlier parts of the series examined QuantDinger, a trading bot that lets Jev veto entry orders — a gate the author notes lets orders proceed when the model fails and cannot be backtested as shipped. Tested as a filter, Jev made boosting worse or left it unchanged. The starkest number: on the 300 candles where boosting was most confident, calls that went its way 58% of the time, Jev objected to 294 of them in numeric format. A gate built on those answers would have blocked nearly every trade the model liked best.

None of this is tradable anyway, the author stresses: even the boosting model's most confident decile earns 0.74 basis points per trade before costs, while a Binance spot round trip costs 15 to 20.

Where the model does work

The series is not a dismissal. On tasks outside trading, Jev scored 92.4% on a 77-label banking intent classification set, close to a fine-tuned BERT at 93.66%, and caught 74 of 74 mislabelled entries when reviewing a trading journal. The pattern the author draws: it performs well when the answer is already in the input, and poorly when the answer sits in the future and depends on how one market on one timeframe actually behaves. The recommendation is to use it for the paperwork around trading, not in the order path.

Why it matters

The test is a cheap template for evaluating any AI-in-trading claim. Of 19 projects reviewed in the first part of the series, only four compared the model against something simpler, and it won none of those comparisons. A confident probability is not evidence, and a model trained on text will encode textbook technical analysis even where a specific market disagrees with the textbook. The practical advice from the author: before letting any model gate live orders, log the state for every historical signal and backtest the gate with and without it — at Jev's pricing that costs cents, far cheaper than finding out live.

Limits

One model version, one market and one timeframe were tested, on price data alone; the author flags that tasks involving news, filings or order flow were not covered. The code is public, and the whole test can be rerun for about $0.19 whenever a new version ships.

  • #ai
  • #machine-learning
  • #bitcoin
  • #trading
  • #benchmark

Related posts