How Darwin decides which model to keep
Darwin is a forecasting agent that rewrites its own model. This note sets out the rule it uses to accept a change, and how we report what it gets right.
Contents
Summary
Darwin makes forecasts, waits for them to resolve, and refits its model. A refit is only a candidate. It replaces the model in use when it does better on data it was not fitted on, by count and by probability, and the change is recorded in Darwin's evolution history.
We report one number for the whole record since generation 10, with its interval, and nothing about money.
- calls resolved
- 1,988
- correct
- 51.2%
- 95% Wilson interval
- 49–53.4%
The interval includes 50%. This record does not yet show an advantage over chance.
Accuracy only. No profit and loss, and no dollar figures.
What Darwin is
Darwin is a research agent in its second era. Every hour it makes direction calls on two crypto markets, with 4-hour calls every fourth hour and daily calls at midnight UTC. Each call is hash-committed to a public ledger repository on the Twigpine network before its outcome exists, and revealed and scored once it resolves.
From generation 10 onward, each forecasting timeframe has its own logistic model. Era 2 rebuilt how calls are settled and validated; it was released on 5 Sep 2026, and the acceptance rule below was adopted with it.
The acceptance rule
Accuracy alone is a weak test. A refit can get one more call right on the validation data while becoming more confidently wrong on the rest. So a candidate has to pass three checks on held-out data,11 Held-out dataOutcomes the candidate was not fitted on. Scoring on them is what keeps a refit from grading its own homework. and all three at once:
- 1Held-out accuracy improves.By at least half a percentage point over the model in use, and to at least 50%.
- 2The Brier score is no worse.The mean squared gap between the forecast probability and what happened. Lower is better.
- 3The log loss is no worse.Punishes confident mistakes hardest. Lower is better.
A refit is tried at most once a day, and only while the last 200 resolved calls are under 60% correct. The candidate is fitted on the older 70% of the record and scored on the newest 30%; a small, bounded adjustment suggested by a language model sometimes joins it and faces the same test. If any check fails, the candidate is discarded and the current model stays. If all three pass, the candidate becomes the model in use, and the change is recorded in Darwin's evolution history.
The rule is conservative on purpose.22 What the rule does not provePassing the checks means a change is not worse in probability. It does not mean the model is calibrated. It rejects the change that looks better by count but worse in probability, which is the kind of change an accuracy-only rule would let through.
Pre-registered challengers
A challenger is a rival model or selection rule that is written down, with the test it must pass, before the outcomes it will be scored on exist. It then runs beside the current policy on new time blocks.
Choosing the better of two models after seeing their results flatters whichever one got lucky. Declaring the comparison first does not. A subgroup that looks strong in hindsight is treated the same way: it becomes a challenger and is tested on new outcomes. It is not reported as a finding.
How accuracy is reported
We report the share of resolved calls that were correct, for the whole cohort since generation 10, with a 95% Wilson interval.33 Why WilsonIt behaves well near 50% and at small sample sizes, where the simpler normal approximation can run past 0% or 100%. As of 26 Sep 2026 that was 1,796 calls, 51.1% correct, interval 48.8% to 53.4%. The record box above reads the same figures live.
p̂ = correct / resolved centre = (p̂ + z²/2n) / (1 + z²/n) half-width = z · √( p̂(1 − p̂)/n + z²/4n² ) / (1 + z²/n) n = 1,796 correct = 918 → 48.8% to 53.4% (as of 26 Sep)
The interval assumes calls are independent.44 Whole countsThe statistics publish the rate rounded to a tenth of a point, so the correct count is the whole number nearest resolved × rate. The rounding moves the interval by less than its last digit. Calls made in consecutive hours, or on markets that move together, are not fully independent, so the real uncertainty is wider than the interval shows. We state the interval anyway, because a bare percentage hides how little 1,796 calls can tell you.
Fig. 1 makes the same point from the other side. If the rate held at 51.1%, the lower edge of the interval would clear 50% at about 7,700 resolved calls. That is arithmetic about the width of the interval, not a prediction that the rate will hold.
What we do not report
- No profit and loss, and no dollar figures. Darwin's record is a statement about forecasts, and we keep it that way.
- No subgroup chosen after the fact. The whole cohort since generation 10 is the only number we publish.
- No figure without its cohort and date. Every number on this page says which calls it covers and when it was read.