How Darwin decides which model to keep

Darwin is a forecasting agent that rewrites its own model. This note sets out the rule it uses to accept a change, and how we report what it gets right.

Published
Method adopted 5 Sep 2026
Figures as of
The record box reads live
Reading time
6 min
Subject
Darwin, era 2, generation 10 onward
Type
Research, method note
Written by
Twigpine
Where this entry sits among 28 dated releases, models, notes and specs, 14 May to 26 Sep 2026. One twig per day and kind; releases grow up, research and specs grow down.
Contents
  1. Summary
  2. What Darwin is
  3. The acceptance rule
  4. Pre-registered challengers
  5. How accuracy is reported
  6. What we do not report
  7. Sources
  8. Revision history

Summary

Darwin makes forecasts, waits for them to resolve, and refits its model. A refit is only a candidate. It replaces the model in use when it does better on data it was not fitted on, by count and by probability, and the change is recorded in Darwin's evolution history.

We report one number for the whole record since generation 10, with its interval, and nothing about money.

The record since generation 10Live, read 29 Sep 2026, 09:37 UTC
calls resolved
1,988
correct
51.2%
95% Wilson interval
49–53.4%

The interval includes 50%. This record does not yet show an advantage over chance.

Accuracy only. No profit and loss, and no dollar figures.

Fig. 1The 95% Wilson interval at a 51.1% rate, by number of resolved calls
Arithmetic, not a forecast
Fig. 1. The shaded band is the 95% Wilson interval for a 51.1% rate at each sample size, on a log scale; the dashed line is 50%, chance. The bar is where Darwin's record stood on 26 Sep 2026. The band narrows as calls resolve; this shows how wide the uncertainty is, not where Darwin's accuracy will go.

What Darwin is

Darwin is a research agent in its second era. Every hour it makes direction calls on two crypto markets, with 4-hour calls every fourth hour and daily calls at midnight UTC. Each call is hash-committed to a public ledger repository on the Twigpine network before its outcome exists, and revealed and scored once it resolves.

From generation 10 onward, each forecasting timeframe has its own logistic model. Era 2 rebuilt how calls are settled and validated; it was released on 5 Sep 2026, and the acceptance rule below was adopted with it.

The acceptance rule

Accuracy alone is a weak test. A refit can get one more call right on the validation data while becoming more confidently wrong on the rest. So a candidate has to pass three checks on held-out data,11 Held-out dataOutcomes the candidate was not fitted on. Scoring on them is what keeps a refit from grading its own homework. and all three at once:

  1. 1
    Held-out accuracy improves.By at least half a percentage point over the model in use, and to at least 50%.
  2. 2
    The Brier score is no worse.The mean squared gap between the forecast probability and what happened. Lower is better.
  3. 3
    The log loss is no worse.Punishes confident mistakes hardest. Lower is better.

A refit is tried at most once a day, and only while the last 200 resolved calls are under 60% correct. The candidate is fitted on the older 70% of the record and scored on the newest 30%; a small, bounded adjustment suggested by a language model sometimes joins it and faces the same test. If any check fails, the candidate is discarded and the current model stays. If all three pass, the candidate becomes the model in use, and the change is recorded in Darwin's evolution history.

The rule is conservative on purpose.22 What the rule does not provePassing the checks means a change is not worse in probability. It does not mean the model is calibrated. It rejects the change that looks better by count but worse in probability, which is the kind of change an accuracy-only rule would let through.

Fig. 2The acceptance rule, drawn as a branch and a merge
Schematic
Fig. 2. The rule in the same drawing system as the network figure on the homepage. Candidates branch from the model in use; a discarded one is pruned, an accepted one merges back and replaces it. Schematic: the number and order of refits are illustrative, not data.

Pre-registered challengers

A challenger is a rival model or selection rule that is written down, with the test it must pass, before the outcomes it will be scored on exist. It then runs beside the current policy on new time blocks.

Choosing the better of two models after seeing their results flatters whichever one got lucky. Declaring the comparison first does not. A subgroup that looks strong in hindsight is treated the same way: it becomes a challenger and is tested on new outcomes. It is not reported as a finding.

How accuracy is reported

We report the share of resolved calls that were correct, for the whole cohort since generation 10, with a 95% Wilson interval.33 Why WilsonIt behaves well near 50% and at small sample sizes, where the simpler normal approximation can run past 0% or 100%. As of 26 Sep 2026 that was 1,796 calls, 51.1% correct, interval 48.8% to 53.4%. The record box above reads the same figures live.

The 95% Wilson intervalz = 1.96
p̂          = correct / resolved
centre      = (p̂ + z²/2n) / (1 + z²/n)
half-width  = z · √( p̂(1 − p̂)/n + z²/4n² ) / (1 + z²/n)
n = 1,796   correct = 918   →   48.8% to 53.4%   (as of 26 Sep)

The interval assumes calls are independent.44 Whole countsThe statistics publish the rate rounded to a tenth of a point, so the correct count is the whole number nearest resolved × rate. The rounding moves the interval by less than its last digit. Calls made in consecutive hours, or on markets that move together, are not fully independent, so the real uncertainty is wider than the interval shows. We state the interval anyway, because a bare percentage hides how little 1,796 calls can tell you.

Fig. 1 makes the same point from the other side. If the rate held at 51.1%, the lower edge of the interval would clear 50% at about 7,700 resolved calls. That is arithmetic about the width of the interval, not a prediction that the rate will hold.

What we do not report

  • No profit and loss, and no dollar figures. Darwin's record is a statement about forecasts, and we keep it that way.
  • No subgroup chosen after the fact. The whole cohort since generation 10 is the only number we publish.
  • No figure without its cohort and date. Every number on this page says which calls it covers and when it was read.

Notes

  1. 1
    Held-out data. Outcomes the candidate was not fitted on. Scoring on them is what keeps a refit from grading its own homework. ↩
  2. 2
    What the rule does not prove. Passing the checks means a change is not worse in probability. It does not mean the model is calibrated. ↩
  3. 3
    Why Wilson. It behaves well near 50% and at small sample sizes, where the simpler normal approximation can run past 0% or 100%. ↩
  4. 4
    Whole counts. The statistics publish the rate rounded to a tenth of a point, so the correct count is the whole number nearest resolved × rate. The rounding moves the interval by less than its last digit. ↩

Sources

  • Method
    Darwin's README and the acceptance rule as adopted with era 2 on 5 Sep 2026, as the Darwin page states it.
  • Figures
    Darwin's live statistics, generation 10 cohort. The text and figures are fixed as of 26 Sep 2026; the record box reads the same statistics live. The interval is computed from the cohort's resolved and correct counts.
  • Release

Revision history

  • First published. The method itself was adopted on 5 Sep 2026 with era 2.

How to cite. Twigpine, “How Darwin decides which model to keep”, 26 Sep 2026. https://twigpine.com/research/darwin-method