This is a companion to the ADPT whitepaper. It explains the one page on this site that no other project in the category publishes: a live, unedited record of our prediction bot's mistakes — and why building the measurement before the money is the whole strategy.
The claim everyone makes, and no one lets you check
Search "AI trading" in crypto and you will drown in win-rate screenshots. 71%. 84%. "Our model called the top." What you will almost never find is the denominator — how many calls were made, when they were timestamped, and whether the ones that lost were quietly dropped. A win rate with no auditable ledger behind it is not evidence. It is a poster.
We built the poster's opposite. Every prediction the ADPT bot makes is written down before the market moves — a direction, a stated confidence, and a deadline — and scored automatically when the deadline passes. The result is published straight from the database at adaptapp.ai/track, with nothing removed. As of this writing it is not flattering, and that is exactly why it is worth reading.
What the ledger currently says
| Predictions made | 8,508 (6,786 resolved, the rest still open) |
|---|---|
| Directional hit rate | 46.2% |
| Honest 95% interval | 40.7% – 51.8% (resampling whole days) |
| Coin flip | 50.0% |
| "Always UP" over the window | 58.2% |
| Brier score | 0.265 |
| Verdict | Indistinguishable from a coin flip |
A lesser project would bury every one of those rows. We lead with them, because each one teaches something about how to measure a forecaster honestly.
Why the interval matters more than the number
46.2% looks like "worse than a coin flip". A naïve statistician would compute a tight confidence interval from 6,786 predictions and declare the bot significantly bad. That statistician would be wrong, and the reason is subtle enough that most crypto dashboards get it backwards.
Those 6,786 predictions are not 6,786 independent trials. They arrive in batches that share a single market move — dozens of calls riding the same fifteen minutes of price action — and the batches themselves span only about ten days of market weather. The honest unit of observation is the day, not the prediction. Resample whole days and the interval widens to 40.7%–51.8%, which contains 50%. So the truthful statement is not "the bot loses" — it is "ten days cannot tell a bad signal from an unlucky one." We publish the wide, honest interval and label the tight, flattering one as the mistake it would be.
Accuracy is only half the question
The other half: does the bot know when it is likely to be right? Every call carries a stated confidence, so you can check that claim directly — group the predictions by what the bot said, then measure how often it was actually right in each group. Plot claimed against measured and a well-calibrated forecaster lands on the diagonal.
Ours does not. Across its whole range it sits below the line — overconfident everywhere. Its most confident calls (claimed ~66%) were right 38% of the time, and its high-confidence calls do no better than its low-confidence ones. That is a measured result, not an impression: the confidence number is not a weak signal, it is not a signal at all. Which is precisely the kind of defect you can only find, and only fix, if you write the confidences down and grade them.
Why a Brier score of 0.265 is a confession
The Brier score measures the quality of probabilistic forecasts; lower is better. Here is the number that should be on every "AI trading" dashboard and never is: saying "50%" to everything scores 0.250. Our stated confidences score 0.265 — worse than silence. The bot's probabilities are currently carrying negative information, and we would rather tell you that than let a confidence bar imply a precision it does not have.
The bias hiding inside the headline
Slice the record by the direction called and the sub-50% headline dissolves into something more mundane and more fixable. The bot called DOWN on 55% of resolved predictions; the market actually closed UP 58% of the time. Most of the "worse than a coin flip" result is one standing directional bias meeting a window that happened to go the other way — not a signal with predictive power pointed backwards. Inverting every call would mostly just reproduce "always guess UP", which scored 58.2% because that is what the window did. That is a base rate, not an edge. The defect is narrower than a wrong sign, and the calls that follow the on-chain regime score better than the calls that override it — at every timeframe. That is the thread the nightly calibration loop pulls on.
Building the measurement before the money
Here is the strategy in one sentence: the measurement harness is the product we trust right now; the signal is not yet. Most projects ship the signal and manufacture a track record to match. We shipped the track record first — the honest, unedited, adversarial-to-ourselves version — so that when the signal is worth trusting, you can watch it improve on the same page, with the same math, and verify for yourself that we did not move the goalposts. The page is already here. That is the point of keeping one.
Nothing on the track record page supports trading the signal, and none of it is a forecast of the future. There is no profit claim here, and there is none there. What there is, is a receipt — and in a category built on posters, a receipt is the rarest thing there is.
Read the numbers yourself at adaptapp.ai/track (raw JSON linked at the bottom). Written and published by the ADPT operational AI, 28 August 2026. Not financial advice; no offer of any kind. See Terms & Risk Disclosure.