DarwinResearchera 2hourly calls

A forecasting agent that has to earn every change to itself.

Darwin is a research agent in its second era. Every hour it makes direction calls on BTC and ETH, records them before the outcome exists, and scores them as they resolve. At most once a day it tries a refit, which replaces the model in use only when it does better on data it was not fitted on. Every call is hash-committed to a public ledger before its outcome exists and revealed after, so the record can be checked. We report how often it is right, with the interval, and nothing about money.

Recordsince generation 10Read 29 Sep 2026, 09:37 UTC
Calls resolved
1,988
Correct
51.2%
95% Wilson interval
49% to 53.4%

The interval includes 50%. This record does not yet show an advantage over chance. Live from Darwin's statistics. Accuracy only; no profit and loss.

gen 10
current generation, era 2
hourly
calls made and scored
daily
refit tried at most once a day; last attempt 28 Sep 2026, 19:00 UTC
51.2%
of 1,988 calls since generation 10 correct, 95% interval 49–53.4%

What it does. Calls first, scores later, changes only on evidence.

Calls first, scores later

Every hour Darwin records a 1-hour call per market, with 4-hour calls every fourth hour and daily calls at midnight UTC, all before the outcomes exist. Each is scored against what the market did once it resolves.

A logistic model per timeframe

From generation 10 onward, each forecasting timeframe (1h, 4h, 1d) has its own logistic model. A refit on resolved calls is tried at most once a day, while the recent record is under 60%.

A change has to earn its place

A refit is only a candidate. It replaces the model in use when held-out accuracy improves by at least half a point and neither the Brier score nor the log loss gets worse. If any check fails, the candidate is discarded.

Challengers are declared first

Rival models and selection rules are written down, with the test they must pass, before the outcomes they are scored on exist. Hindsight picks become challengers, not findings.

One number, with its interval

The record is the share of calls correct since generation 10, with a 95% Wilson interval, the cohort, and the time it was read. No profit and loss, no dollar figures.

Every call is committed before it resolves

Each forecast is hash-committed to a public ledger before the outcome exists and revealed once it resolves, so no call can be edited after the fact. Accepted model changes are kept in Darwin's evolution history.

The loop. Forecast, score, refit and test.

  1. Forecast

    A 1-hour call per market every hour, 4-hour calls every fourth hour, daily calls at midnight UTC, each recorded before the outcome exists.

  2. Score

    Calls resolve against the market and join the record: resolved count, share correct, interval.

  3. Refit and test

    At most once a day, while the recent record is under 60%, a refit is tried. It is accepted only if it beats the model in use on held-out calls by at least half a point, without a worse Brier score or log loss. Anything else is discarded.

Method. How a model earns its place.

The acceptance rule

Darwin calls every hour, but it tries a refit at most once a day, and only while its last 200 resolved calls are under 60% correct. A candidate is a regularised refit on the older 70% of the record, pulled toward a year-long backtest; a small, bounded adjustment suggested by a language model sometimes joins it. Every candidate faces the same test on the newest 30%, which it was not fitted on.

Accuracy alone is a weak test: a refit can get one more call right on the validation data while becoming more confidently wrong on the rest. So a candidate has to pass three checks at once. Held-out accuracy rises by at least half a percentage point (and reaches at least 50%), the Brier score (the mean squared gap between the forecast probability and what happened) is no worse, and the log loss (which punishes confident mistakes hardest) is no worse.

Passing means a change is not worse in probability. It does not mean the model is calibrated. The rule is conservative on purpose, and most attempts are rejected.

at most once a day, while the last 200 resolved calls
are under 60% correct:

candidate = refit(older 70% of resolved calls)

on the newest 30% (held out, never fitted on):
  accuracy(candidate)  >= max(50%, accuracy(current) + 0.5 pts)
  brier(candidate)     <= brier(current)
  log_loss(candidate)  <= log_loss(current)

all three pass → accept, record in history
any one fails  → discard, keep current

Pre-registered challengers

A challenger is a rival model or selection rule that is written down, with the test it must pass, before the outcomes it will be scored on exist. It then runs beside the current policy on new time blocks.

Choosing the better of two models after seeing their results flatters whichever one got lucky. A subgroup that looks strong in hindsight is treated the same way: it becomes a challenger and is tested on new outcomes. It is not reported as a finding.

How accuracy is reported

One number for the whole cohort since generation 10: the share of resolved calls that were correct, with a 95% Wilson interval. Wilson behaves well near 50% and at small samples, where the normal approximation can run past 0% or 100%.

The interval assumes calls are independent. Calls made in consecutive hours, or on markets that move together, are not fully independent, so the real uncertainty is wider than the interval shows.

Read 29 Sep 2026, 09:37 UTC: 1,988 calls, 51.2% correct, interval 49% to 53.4%.

z = 1.96
p̂ = correct / resolved

centre     = (p̂ + z²/2n) / (1 + z²/n)
half-width = z · √( p̂(1 − p̂)/n + z²/4n² ) / (1 + z²/n)

What we do not report

No profit and loss, and no dollar figures: Darwin's record is a statement about forecasts. No subgroup chosen after the fact: the whole cohort since generation 10 is the only number we publish. No figure without its cohort and the time it was read.

Read the method. Then check the record against it.