Paper record · published unedited
Every call it made, including the bad ones
The prediction bot writes down a direction, a confidence and a deadline before the market moves. When the deadline passes it is scored automatically against the price. This page is a published ledger snapshot with nothing removed. The available record does not establish a forecasting edge. It is published anyway — that is the point of keeping one.
Is its confidence worth anything?
Accuracy is only half the question. The other half is whether the bot knows when it is likely to be right. Each call carries a stated confidence, so that claim can be checked directly: group the calls by what the bot claimed, then measure how often it was actually right. A well-calibrated forecaster lands on the dashed line. Below the line means overconfident.
Table view — the same numbers
| Confidence band | Claimed | Actual | 95% interval | Calls |
|---|
By horizon
Each bar is the 95% interval; the tick is the measured rate. The vertical line is 50%. An interval that sits entirely to the left of that line is not bad luck — it is a signal pointing consistently the wrong way.
| Horizon | 30% — 50% — 70% | Rate | 95% interval | Resolved |
|---|
Cut other ways
Published for completeness. Slicing a record enough ways will always surface a flattering slice, so none of these should be read as a finding — they are here so nobody has to take our word for which cut we looked at.
| Source | Rate | 95% interval | n |
|---|
| Market | Rate | 95% interval | n |
|---|
| Called | Rate | 95% interval | n |
|---|
What this record does and does not establish
- The measurement is real. Calls are timestamped before their horizon opens and scored automatically at expiry. Nothing is re-scored by hand and nothing is deleted.
- A baseline, not a verdict. The point estimate sits below 50%, while the day-clustered interval crosses 50%. The useful result is the diagnosis: the current mapping has not established directional skill over this short sample.
- The confidence output needs recalibration. Its published probabilities do not yet separate outcomes reliably. That measured weakness gives the next model a precise benchmark to beat.
- Not that the bot is beaten by a coin. The point estimate is below 50%, but resampling whole days puts the honest interval across 50%. Thousands of predictions are not thousands of independent trials: they arrive in batches that share one market move, and the batches share about ten days of weather. Ten days cannot separate a bad signal from an unlucky one.
- Not trading P&L. This page measures forecast accuracy and calibration. It does not include execution, position sizing, fees or slippage, so it cannot establish the profitability of a trading strategy.
- Not that inverting it would work. The sub-50% result is largely a standing bias meeting a market that went the other way; flipped, it is just the base rate, not skill. The measured defect is narrower and more fixable than a wrong sign — the calls that follow the on-chain regime score better than the calls that override it, at every single timeframe.
- Not that the underlying on-chain data is worthless. It establishes that this particular mapping from those metrics to a directional call is worse than useless. Those are different claims.
- Not a forecast of the future. A short window over one market regime does not generalise, in either direction.
- loading…