ACCURACY · TRACK RECORD · WE PUBLISH OUR MISSES

We grade our own work —
and keep the misses.

Every Fahali detection is recorded and later resolved against what the market actually did. Hits and misses live in the same ledger, on the same scale. The honesty is the point — and it's why we publish the record one stratum at a time, with its misses, instead of one platform-wide percentage.

Why there's no single big percentage on this page. A platform-wide accuracy number is not a thing we will ever quote: pooled across engines, horizons and asset classes it mixes populations with different base rates, and the result is a figure nobody can act on. What we publish instead is the record per stratum — one engine, one horizon tier, one instrument class — with its recall and its matched base rate beside it, and its measurement date attached. A stratum that fails the publishability gate is withheld with a reason rather than rounded up. You can read those on the front page and pull every row from the endpoint yourself.

How a call becomes a verified outcome

1 · Emit

A detection fires and is written to the signal-to-outcome ledger with a unique ID, the engine(s) involved, the symbol, a timestamp, the predicted direction where applicable, and a confidence score.

2 · Resolve

After the forecast horizon (1h, 4h, 24h, up to 48h+ depending on the engine), the call is resolved against realized price action and marked correct or incorrect — and that result is stored permanently.

3 · Keep the misses

Incorrect calls stay in the same ledger as correct ones. There is no separate scorekeeping, no filtered dataset, no sanitized retrospective. The record compounds with every market session and cannot be back-filled.

The four axes we score

Detections are not all the same kind of claim, so they are scored on the axis that fits.

Direction

For directional engines: was the predicted up/down correct against the net return over the horizon?

Magnitude

Did a meaningful move actually occur (e.g. a move beyond a small threshold), regardless of direction?

Volatility

For regime/expansion signals: did volatility expand as flagged? Directionless by design.

Crash

For crash-class signals: did a rapid decline materialize within the warning window?

The ledger, right now.

These numbers are pulled live from the signal-to-outcome ledger. They are interim while the ledger grows across a mixed market regime, and they include misses.

SCORECARD · accumulating data

Per-engine scorecard pending

The fleet-wide per-engine scorecard is still accumulating across a mixed market regime, and is withheld until it earns publication. That is a narrower statement than it used to be: individual strata that already clear the gate are published — see the judged record on the front page and the full row set at /api/track-record/lead-time. What is withheld here is the fleet roll-up, not the record.

Does agreement between engines make the call better?

Here is one number we can stand behind today — because it is relative lift measured inside the same market window, so it does not depend on which way the market happened to go. We group the ledger's already-resolved signals into moments where several independent engines fired on the same instrument at once, then ask a simple question: did those higher-agreement moments resolve better than the average single signal over the same period?

Reading the ledger…

MULTI-ENGINE CONSENSUS · live from the ledger

Loading…

Pulling consensus tiers from the live signal-to-outcome ledger.

Read the lift, not the raw rate. Lift = a tier's hit-rate minus the average single signal's hit-rate over the same window. Because both are measured in the same regime, lift isolates the value of agreement from market drift — the metric we said we would report. A negative lift is a real result and stays on the page: it says agreement did not help over this window, which is a finding about our own engines, not a rendering fault. This is a relative measure, not a per-engine accuracy figure, and it is not advice — observation only.

Why base-rate lift, not raw percentage

The number that will matter is lift, not the raw percentage. In a falling market, a naive "always bearish" strategy can score ~100% on direction while adding zero information — it just rode the market. Base-rate lift measures how much better a signal performs than that naive baseline (lift = precision − base rate). When we publish the scorecard, we'll show both the raw figure and the lift, so you can separate genuine skill from market drift. Reporting a raw "accuracy %" alone would be a vanity metric, and we won't do it.

What this page is not

It is not a marketing scorecard of impressive-looking numbers. Anyone can publish a percentage; few publish the misses behind it and fewer still withhold the number until it's statistically honest. That restraint — observable here — is itself the credibility signal. When the ledger qualifies, this page becomes the live per-engine scorecard with direction / magnitude / volatility / crash breakdowns and base-rate lift for each engine.

See the live read in the meantime.

The daily plain-language market read is free, no signup wall.