Skip to content

Research infrastructure · Intermediate

Running backtests at scale

A backtest can be trusted only if it can be reproduced and counted, and both have to be built into the pipeline before the first sweep runs.

36 min read15 referencesIntroduction to the series

Educational material only. Not investment advice.

Contents
  1. 01What a result must carry
  2. 02The pipeline
  3. 03Leases, retries and idempotence
  4. 04Count every trial
  5. 05An interactive laboratory
  6. 06A real sweep, deflated
  7. 07Register the prediction first
  8. 08A checklist
  9. —References

Abstract

Run thousands of backtests and two things go wrong: results that cannot be reproduced, and winners chosen from searches nobody counted. We identify each run by a hash of its code, parameters, data and engine, run the jobs through a leased queue that tells transient failures from permanent ones, and write each attempt to a ledger. The deflated Sharpe ratio then judges the best result against the size of the search. On a real sweep of 156 variants of the 200-day rule, the best variant’s edge over a matched market exposure has a probabilistic Sharpe ratio of 99.8% judged alone and a deflated one of 93%, short of the 95% bar we set.

Key takeaways

  • Identify each run by a hash of its code, parameters, data and engine versions; cache and reproduce by that key.
  • Lease jobs, classify each failure as transient or permanent, retry transient ones with jittered exponential backoff, and retire a job that keeps killing its workers.
  • Keep a ledger of all trials, including abandoned sweeps and failures; the trial count is part of every result.
  • Judge the best result with the deflated Sharpe ratio, and the choice among variants with the probability of backtest overfitting; neither is a p-value.
  • Register a prediction and its refutation condition before a sweep runs, and score it afterwards.
  • Summarise a parameter surface by its level and its neighbourhoods, never by its best cell.

Before you start

  • Market anomalies, for data snooping and the noise ceiling
  • QuantConnect and LEAN, or any engine that runs a backtest from code, for what a run depends on
  • Basic Python, SQL or pandas, and the Sharpe ratio

A single backtest is easy to trust and easy to get wrong. A research programme runs thousands, and its errors are quiet ones. A result computed on last month’s code sits beside one computed on today’s, and the strongest variant of a sweep is reported without the hundred weaker ones tried beside it. Neither can be repaired from the results afterwards, so the fixes have to be in place before the first sweep runs.

What a result must carry

A backtest result is a function of four things: the code, the parameters, the data and the engine that ran them. If any of the four can change without the result changing its name, the result cannot be reproduced, and in a long research programme one of them always changes. Code is edited, a data vendor corrects a history, a library upgrade alters a default. Recording exactly how each result was produced is the first rule of reproducible computation.1 We make the record the result’s name: each run is identified by a cryptographic hash of all four inputs, so the same specification always has the same key and a different specification can never borrow an old one.

InputWhat we recordWhat goes wrong without it
CodeA hash of the bytes of each file that computes the stored resultA result survives a bug fix it never saw
ParametersCanonical JSON: sorted keys and sets, numbers as floats, seeds included; the caller writes out every defaultOne idea gets two names, or a parameter left out lets two different runs share one
DataA hash of the raw files, which fixes the vintageA corrected history silently changes old results
EngineLanguage and library versions, and the backtest engine’s versionDefaults change underneath the run
site_research/fieldnotes/pipeline.py
def digest(*chunks: bytes) -> str:
    h = hashlib.sha256()
    for c in chunks:
        h.update(hashlib.sha256(c).digest())
    return h.hexdigest()


def canonical(x):
    """Parameters in one form, the same in every process: numbers become floats (so 1 and 1.0
    are the same parameter) except integers beyond 2**53, which a float would merge with their
    neighbours; sets are sorted, since their order changes with Python's hash seed; nested
    dictionaries and lists are canonicalised too. Keys become strings and are sorted when the
    JSON is written. Defaults are the caller's to write out."""
    if isinstance(x, np.bool_):
        return bool(x)
    if isinstance(x, bool) or x is None or isinstance(x, str):
        return x
    if isinstance(x, (int, np.integer)) and abs(int(x)) > 2 ** 53:
        return int(x)
    if isinstance(x, (int, float, np.integer, np.floating)):
        return float(x)
    if isinstance(x, dict):
        return {str(k): canonical(v) for k, v in x.items()}
    if isinstance(x, (set, frozenset)):
        return sorted((canonical(v) for v in x), key=lambda v: json.dumps(v, sort_keys=True))
    return [canonical(v) for v in x]


def fingerprint(code_files: list[Path], params: dict, data_digest: str) -> tuple[str, dict]:
    """The key of one run: a hash of the code's bytes, the parameters in canonical form, the
    data's digest and the engine's versions. Change any of them and the key changes. A random
    seed, where a run has one, is a parameter: two seeds are two trials."""
    spec = {
        "code": digest(*(Path(p).read_bytes() for p in sorted(map(str, code_files)))),
        "params": canonical(params),
        "data": data_digest,
        "engine": {"python": platform.python_version(), "numpy": np.__version__, "pandas": pd.__version__,
                   "scipy": scipy.__version__},
    }
    canon = json.dumps(spec, sort_keys=True, separators=(",", ":"))
    return hashlib.sha256(canon.encode()).hexdigest()[:16], spec       # 64 bits: ample below ~10^8 runs
Change a comment in any hashed file, a single parameter, one byte of the data or a recorded library version, and the key changes. Hashing source files is deliberately strict: a harmless edit also creates a new run. A random seed is a parameter, so two seeds are two trials.

A result keyed this way can be cached, so a sweep that is interrupted and restarted, or a colleague’s identical request, costs nothing the second time. It can be reproduced, because the specification is stored beside it. And a result that disagrees with its own rerun is a finding in itself: the engine is not deterministic, or an input escaped the fingerprint. The public data in this series are corrected and extended from time to time, which is why each note states the range it used and the build hashes the files it read.2 On a hosted engine whose data cannot be hashed, the practical substitute is a canary: one fixed backtest, rerun on a schedule, whose result must not change. The sources of drift inside an engine, from fill models to corporate actions, are described in QuantConnect and LEAN.

The pipeline

The pipeline that turns ideas into counted, reproducible results has the same shape at any scale. A hypothesis, with a prediction written down before anything runs, becomes a set of specifications; each specification is fingerprinted and queued; workers lease jobs from the queue, run them and write results; the ledger records every attempt, successful or not; and analysis reads the ledger.

Figure 1A research pipeline
Pipeline: a hypothesis and its prediction become fingerprinted specifications, which are queued; workers lease jobs and write results keyed by fingerprint; transient failures and expired leases return jobs to the queue, permanent failures retire them; every attempt goes to the trial ledger, which analysis reads before the next sweep.Hypothesisand predictionSpecificationsfingerprintedQueueleases, backoffWorkersrun and reportResult storeone per fingerprintRetiredpermanent failureTrial ledgerevery attemptAnalysissurface, DSR, PBO, verdicttransient: retry after a jittered backofflease expiredresultsfailuresreadnext sweep
Solid arrows carry work forward. A worker holds a job only for the length of its lease. A transient failure, or a lease that expired because the worker died, sends the job back to the queue; a permanent failure, or a job that has used up its attempts, retires it. All outcomes go to the ledger, and the deflated Sharpe ratio is computed from the ledger rather than from the winner alone.

The arrow from the ledger to analysis is the one that is easiest to leave out. A results store keeps what worked; a ledger keeps what was tried. The machine-learning literature describes the cost of the untraced alternative: systems that accumulate configurations and data dependencies nobody can trace, and results nobody can reproduce, pay for it later as technical debt.3 The whole state fits in four tables:

site_research/fieldnotes/pipeline.py
SCHEMA = """
CREATE TABLE IF NOT EXISTS jobs (
    fingerprint TEXT PRIMARY KEY,
    sweep       TEXT NOT NULL,
    spec        TEXT NOT NULL,            -- canonical JSON: code, params, data, engine
    state       TEXT NOT NULL DEFAULT 'queued',   -- queued | leased | done | dead
    attempts    INTEGER NOT NULL DEFAULT 0,
    lease_owner TEXT,
    lease_until REAL,
    not_before  REAL NOT NULL DEFAULT 0,  -- earliest time a retry may start
    error       TEXT
);
CREATE TABLE IF NOT EXISTS results (
    fingerprint TEXT PRIMARY KEY,         -- one result per specification, ever
    metrics     TEXT NOT NULL,
    worker      TEXT NOT NULL,
    finished    REAL NOT NULL
);
CREATE TABLE IF NOT EXISTS ledger (       -- every attempt, including the failures
    id          INTEGER PRIMARY KEY AUTOINCREMENT,
    sweep       TEXT NOT NULL,
    fingerprint TEXT NOT NULL,
    outcome     TEXT NOT NULL,            -- result | transient | permanent | lease_expired | stale
    sharpe      REAL, n_obs INTEGER, skew REAL, kurt REAL,
    note        TEXT,
    at          REAL NOT NULL
);
CREATE TABLE IF NOT EXISTS predictions (  -- written BEFORE the sweep runs
    sweep       TEXT PRIMARY KEY,
    claim       TEXT NOT NULL,
    refuted_if  TEXT NOT NULL,
    verdict     TEXT
);
"""

Leases, retries and idempotence

At scale, workers die, services time out and some jobs are simply wrong, and a queue has to treat each case differently.

Leases

A worker that takes a job holds it on a lease: a claim that expires after a fixed time unless the worker renews it or reports the job finished. If the worker crashes, loses its network or is shut down, nobody has to notice; the lease runs out and the job becomes available again. Gray and Cheriton introduced leases for distributed file caches for this reason: a lock held by a dead machine is held forever, a lease is not.4 The lease should be comfortably longer than the slowest normal run, and a worker on a long run renews it. A job whose lease expires on every attempt is probably killing its workers, so after 4 attempts it is retired like any other permanent failure.

site_research/fieldnotes/pipeline.py
def lease(db, worker: str, now: float, ttl: float = 900.0, max_attempts: int = MAX_ATTEMPTS):
    """Claim one runnable job: queued and past its backoff, or leased by a worker whose lease
    has expired (that worker is presumed dead). A job whose lease expires on its last attempt is
    retired instead, like any job that has used its attempts: when leases keep expiring, the job, not the worker, is the likely culprit."""
    db.execute("BEGIN IMMEDIATE")
    try:
        while True:
            row = db.execute(
                "SELECT fingerprint, sweep, spec, state, attempts FROM jobs WHERE (state = 'queued' AND not_before <= ?) "
                "OR (state = 'leased' AND lease_until < ?) ORDER BY not_before, fingerprint LIMIT 1", (now, now)).fetchone()
            if row is None:
                db.execute("COMMIT")
                return None
            fp, sweep, spec, state, attempts = row
            if state == "leased":
                db.execute("INSERT INTO ledger (sweep, fingerprint, outcome, note, at) VALUES (?, ?, 'lease_expired', ?, ?)",
                           (sweep, fp, "worker lost its lease", now))
                if attempts >= max_attempts:
                    db.execute("UPDATE jobs SET state = 'dead', lease_owner = NULL, error = ? WHERE fingerprint = ?",
                               ("lease expired on its last attempt", fp))
                    db.execute("INSERT INTO ledger (sweep, fingerprint, outcome, note, at) VALUES (?, ?, 'permanent', ?, ?)",
                               (sweep, fp, "lease expired on its last attempt", now))
                    continue
            db.execute("UPDATE jobs SET state = 'leased', lease_owner = ?, lease_until = ?, attempts = attempts + 1 "
                       "WHERE fingerprint = ?", (worker, now + ttl, fp))
            db.execute("COMMIT")
            return fp, sweep, json.loads(spec)
    except BaseException:
        db.execute("ROLLBACK")
        raise


def renew(db, fp: str, worker: str, now: float, ttl: float = 900.0) -> bool:
    """Extend a lease the worker still holds. False means another worker has it now: the heartbeat
    stops renewing, and whatever the run reports later is harmless (a result is kept only if it is
    the first for its fingerprint, and a failure from a worker without the lease changes nothing)."""
    return db.execute("UPDATE jobs SET lease_until = ? WHERE fingerprint = ? AND state = 'leased' AND lease_owner = ?",
                      (now + ttl, fp, worker)).rowcount == 1
One transaction claims one runnable job: queued and past its backoff, or leased to a worker whose lease has expired. An expired lease is written to the ledger when the job is reclaimed.

Transient and permanent failures

A timeout from a data server is likely to succeed on a second attempt; a parameter outside its valid range will fail every time. Retrying the second kind wastes capacity and hides the bug; not retrying the first kind loses work. So each failure is classified before anything else happens. Transient failures go back to the queue after a wait that grows exponentially with each attempt and is drawn at random up to that bound (“full jitter”), so that workers that failed together do not all retry in the same second.5 After 4 attempts they are retired too. Permanent failures are retired at once, with the error kept for whoever fixes them. Only the worker that still holds the lease may re-queue or retire the job, so a worker that was presumed dead and wakes up cannot undo the work of its replacement. A late result is still accepted, because it is identical to the one the replacement will produce.

site_research/fieldnotes/pipeline.py
def backoff(attempt: int, rng: random.Random, base: float = 30.0, cap: float = 3600.0) -> float:
    """Exponential backoff with full jitter: a uniform wait up to base * 2^attempt, capped, so
    workers that failed together do not retry together."""
    return rng.uniform(0, min(cap, base * 2 ** attempt))


def fail(db, fp: str, worker: str, exc: BaseException, now: float, rng: random.Random,
         max_attempts: int = MAX_ATTEMPTS) -> str:
    """Transient failures go back to the queue after a backoff, until the attempts run out;
    anything else is permanent at once. Only the worker that still holds the lease may change the
    job: a report from any other is a stale one, kept in the ledger as such and changing nothing."""
    transient = isinstance(exc, (TransientError, TimeoutError, ConnectionError))
    db.execute("BEGIN IMMEDIATE")
    try:
        sweep, attempts = db.execute("SELECT sweep, attempts FROM jobs WHERE fingerprint = ?", (fp,)).fetchone()
        retry = transient and attempts < max_attempts
        owned = db.execute(
            "UPDATE jobs SET state = ?, lease_owner = NULL, not_before = ?, error = ? "
            "WHERE fingerprint = ? AND state = 'leased' AND lease_owner = ?",
            ("queued" if retry else "dead", now + backoff(attempts, rng) if retry else 0.0,
             f"{type(exc).__name__}: {exc}", fp, worker)).rowcount == 1
        outcome = ("transient" if transient else "permanent") if owned else "stale"
        db.execute("INSERT INTO ledger (sweep, fingerprint, outcome, note, at) VALUES (?, ?, ?, ?, ?)",
                   (sweep, fp, outcome, f"{type(exc).__name__}: {exc}", now))
        db.execute("COMMIT")
    except BaseException:
        db.execute("ROLLBACK")
        raise
    return "stale" if not owned else "retry" if retry else "dead"

Idempotent results

A worker can lose its lease and still finish. The job has meanwhile been leased to someone else, and two results for the same fingerprint arrive. Because a fingerprint fixes the code, parameters, data and engine, the two results should be identical, so the store accepts the first and ignores the second. Writing results keyed by fingerprint makes recording them idempotent: doing it twice, or again after a crash halfway through, leaves the same state as doing it once.

site_research/fieldnotes/pipeline.py
def complete(db, fp: str, worker: str, metrics: dict, now: float) -> bool:
    """Record a result. A worker that lost its lease may still finish: the result is the same
    (same code, parameters and data), so the first write wins and a second changes nothing. It is
    accepted from any worker, even for a job its replacement holds or one already retired, because
    it is the result the replacement would produce."""
    db.execute("BEGIN IMMEDIATE")
    try:
        sweep = db.execute("SELECT sweep FROM jobs WHERE fingerprint = ?", (fp,)).fetchone()[0]
        new = db.execute("INSERT OR IGNORE INTO results VALUES (?, ?, ?, ?)",
                         (fp, json.dumps(metrics), worker, now)).rowcount == 1
        if new:
            db.execute("INSERT INTO ledger (sweep, fingerprint, outcome, sharpe, n_obs, skew, kurt, at) "
                       "VALUES (?, ?, 'result', ?, ?, ?, ?, ?)",
                       (sweep, fp, metrics["sharpe"], metrics["n_obs"], metrics["skew"], metrics["kurt"], now))
        db.execute("UPDATE jobs SET state = 'done', lease_owner = NULL, error = NULL WHERE fingerprint = ?", (fp,))
        db.execute("COMMIT")
        return new
    except BaseException:
        db.execute("ROLLBACK")
        raise


def worker_loop(path: str, name: str, run, ttl: float = 900.0) -> int:
    """A real worker: lease, run with a heartbeat that renews the lease, record the outcome, and
    repeat until nothing is runnable now (jobs waiting out a backoff are left for the next pass).
    Any number of these, in threads or processes, can share one database file."""
    db, rng, done = connect(path), random.Random(name), 0
    while (job := lease(db, name, time.time(), ttl)) is not None:
        fp, _, spec = job
        stop = threading.Event()

        def heartbeat():
            hb = connect(path)
            while not stop.wait(ttl / 3) and renew(hb, fp, name, time.time(), ttl):
                pass
            hb.close()

        threading.Thread(target=heartbeat, daemon=True).start()
        try:
            complete(db, fp, name, run(spec["params"]), time.time())
            done += 1
        except Exception as exc:                          # noqa: BLE001 -- classified in fail()
            fail(db, fp, name, exc, time.time(), rng)
        finally:
            stop.set()
    return done
A worker is a loop: lease, run under a heartbeat that renews the lease, record the result or classify the failure. Any number of them, in threads or processes, can share one database file.
EventClassWhat the queue doesWhat the ledger records
Worker crashes mid-runLease expiryThe job returns to the queue when the lease runs out; retired once it has used 4 attemptslease_expired, then permanent
Data server timeoutTransientRetry after a jittered, growing wait; retire after 4 attemptstransient, with the error
Parameter out of range, code errorPermanentRetire at once, keep the errorpermanent, with the error
A stale worker reports lateHarmlessThe first result for a fingerprint is kept and a second is ignored; a late failure cannot change the jobstale, for a late failure; nothing, for a second result

The lab runs this queue, ported line for line, on six jobs and three workers, with a clock that moves in thirty-second ticks. The faults are the reader’s to choose.

Figure 2Break the queue
Worker w2 dies on its first job
The data server times out on the 250-day job
The 2,000-day typo
Later, dead w2 wakes up and

At 5 min: 8 attempts have started and 5 of 6 jobs have a result; 1 transient failure is in the ledger, 1 lease has expired and 1 job has been retired. A late report from w2, which no longer held the lease, is in the ledger as stale and changed nothing. The queue is empty after 5 min, and every fingerprint has at most one result.

Each job’s attempts, from lease to outcome (worker named on the bar)0m1m2m3m4m5m100 days150 days200 days250 days300 days2,000 daysw1✓w2 died; lease expiredw3✓w1timeoutw3✓w1retired: bad parameterw1✓w3✓w2 wakes: late timeout ignored: not the lease holder
Results
5
one per fingerprint
Transient
1
in the ledger; a job is retried until 4 attempts
Leases expired
1
a lease lasts 2 min 30 s unless renewed
Retired
1
permanent failures, kept with their error
The trial ledger up to 5 min
Ledger, at min:sJobOutcome
1:30100 daysresult
1:30200 daysresult
2:30250 daystransient failure
3:002,000 dayspermanent failure
3:00300 daysresult
3:00150 dayslease expired
3:30150 dayslate report, not the lease holder: ignored
4:30150 daysresult
5:00250 daysresult

Worker w2 vanishes the moment it leases its first job, as a machine that loses power would; the 250-day job times out after one minute on the attempts chosen above. A late report from w2 arrives one tick after another worker has taken its job over.

A simulated clock; the queue’s functions are those of pipeline.py, and the site’s build checks that the page writes the same ledger as the Python for the same faults. Each row is a job and each bar an attempt, from its lease to its outcome. Suggested experiments: let the worker die and watch its job sit untouched until the lease runs out; make the data server time out every time and count the attempts before the job is retired; wake the dead worker with a timeout while its replacement is running, and see that the job is not re-queued under it; then wake it with a result and find the replacement’s identical result ignored.

From one machine to many

SQLite with an immediate transaction is enough for the workers of one machine. Across machines the same design carries over unchanged in its logic: a PostgreSQL table claimed with SELECT … FOR UPDATE SKIP LOCKED, or a managed queue whose visibility timeout is a lease under another name.6,7 The worker’s run function is where an engine plugs in. For a cloud backtest it submits the project, polls until the run finishes and returns the statistics; the platform’s rate limits and timeouts are raised as transient errors, and a compile error as a permanent one.

Throughput and cost

The arithmetic of a sweep is plain: the number of runs, times the minutes each takes, divided by the number of workers, plus retries. On a shared engine much of a backtest’s wall-clock time can go on waiting for capacity and loading data. Loading data once per worker and caching by fingerprint remove most of that. A coarse sweep to find the plateau, then a finer one around it, spends compute where the decision is. Each extra run also raises the bar the best one must clear.

Count every trial

In Market anomalies we showed how good the best of many useless strategies looks: with a hundred independent tries on ten years of data, the best is expected to show an annual Sharpe ratio of about 0.80 with no edge at all. The number of trials is therefore part of every result. A Sharpe ratio of 1.0 reported after three attempts is a different claim from the same figure reported after three hundred, and only the ledger records which was made. White’s reality check and Hansen’s test for superior predictive ability test the best of many models against that selection, and, as Market anomalies noted, Harvey, Liu and Zhu put the bar for a new factor at a t-statistic of about three.8,9,10

What counts as a trial

A trial is any configuration whose result anyone saw on the same question, in any sweep. That includes sweeps abandoned because they looked bad, and reruns after a code change, since a changed fingerprint is a new trial. The count N is the number of distinct fingerprints with a result, and the spread below is taken over all of them. A separate question gets a separate ledger. Failures that never produced a result did not inform the choice, but they are recorded too: a sweep whose failures cluster in one region of the parameter space has not explored that region, and its apparent optimum may simply be the edge of what ran. What should not inflate the count is the same specification run twice, which the fingerprint makes easy to tell apart.

The deflated Sharpe ratio

Bailey and López de Prado’s probabilistic Sharpe ratio gives the probability that a strategy’s true Sharpe ratio exceeds a benchmark SR*, given an estimate SR from n observations with skewness γ3 and kurtosis γ4:11

PSR(SR*)  =  Φ( (SR − SR*) √(n − 1) / √(1 − γ3 SR + (γ4 − 1)/4 · SR2) )
probabilistic Sharpe ratio; Sharpe ratios per period, not annualised; γ₄ = 3 for a normal distribution

The deflated Sharpe ratio sets the benchmark to the best Sharpe ratio that N trials with no edge would be expected to produce, given how widely the trials’ estimates were spread:12

SR0  =  √V[SRn] · ( (1 − γE) Φ−1(1 − 1/N) + γE Φ−1(1 − 1/(Ne)) ),    DSR = PSR(SR0)
expected best of N under no edge; e is Euler’s number, γE ≈ 0.5772 the Euler–Mascheroni constant, Φ⁻¹ the normal quantile

Correlation between trials enters twice. Near-copies of a configuration agree with each other, so the spread, and with it the penalty, is small; but the count N still treats each near-copy as a separate try. When trials fall into a few clusters, that makes the correction conservative, which is why López de Prado and Lewis estimate the effective number of trials by clustering the trials’ returns.13 The deflated Sharpe ratio is also not a p-value. Its benchmark is only the typical best of a search with no edge, but a 95% reading asks the winner to beat that typical best by a further 1.65 standard errors of a single estimate, far more than the best of many varies from one search to the next. With no edge it passes much less often than one time in twenty, and the lab below shows how much less.

site_research/fieldnotes/pipeline.py
def probabilistic_sharpe(sr: float, sr0: float, n_obs: int, skw: float = 0.0, kurt: float = 3.0) -> float:
    """P(true Sharpe > sr0) given an estimate `sr` from n_obs returns, allowing for skewness and
    fat tails (Bailey and Lopez de Prado 2012). All Sharpe ratios per period, not annualised."""
    return float(norm.cdf((sr - sr0) * np.sqrt(n_obs - 1) / np.sqrt(1 - skw * sr + (kurt - 1) / 4 * sr ** 2)))


def deflated_sharpe(sr: float, trial_srs, n_obs: int, skw: float = 0.0, kurt: float = 3.0) -> dict:
    """The deflated Sharpe ratio (Bailey and Lopez de Prado 2014): the probabilistic Sharpe ratio
    of the selected strategy, measured against the best Sharpe ratio that `len(trial_srs)` trials
    with no edge would be expected to produce, given how spread out the trials were."""
    trial_srs = np.asarray(trial_srs, float)
    n = len(trial_srs)
    sr0 = float(np.std(trial_srs, ddof=1) * expected_max_sharpe(n, 1.0)) if n > 1 else 0.0
    return {"trials": n, "sr0": sr0, "psr": probabilistic_sharpe(sr, 0.0, n_obs, skw, kurt),
            "dsr": probabilistic_sharpe(sr, sr0, n_obs, skw, kurt)}
Both functions take per-period Sharpe ratios. expected_max_sharpe(n, 1.0) is the helper that draws the noise ceiling in Market anomalies, evaluated with a standard error of one.

Any backtest can be deflated the same way, given its Sharpe ratio, its length, the size of the search behind it and the shape of its returns.

Figure 3Deflate your own backtest

Judged alone, the chance that the true Sharpe ratio is above zero is above 99.9%. After 200 trials, the best result that no edge at all would be expected to produce is an annual Sharpe ratio of 0.59, and the deflated Sharpe ratio is 95.5%. To reach 95% after this search, the backtest would have needed an annual Sharpe ratio of 1.18. It stops passing at about 263 trials.

50%75%95%1101001,00010,000Configurations tried (log scale)stops passing at 263
Deflated Sharpe ratioJudged alone (PSR)
Daily data, 252 sessions a year. “Spread of the trials’ Sharpe ratios” sets the spread of the trials’ Sharpe estimates as a share of one estimate’s standard error: one for independent trials, lower for near-copies of each other. The curve is the deflated Sharpe ratio for each trial count from one to ten thousand. Suggested experiments: find the trial count at which a Sharpe ratio of 1.2 on eight years stops passing; halve the spread and see how much searching near-copies costs; add fat tails and a negative skew.

An interactive laboratory

The lab makes the correction visible on a synthetic grid of 576 configurations, the kind of two-parameter surface a sweep produces. Each configuration has a true annual Sharpe ratio, chosen by the reader: zero everywhere, a broad plateau, or a narrow peak of the same height. Its estimate is the true value plus noise with the standard error of a Sharpe ratio measured over the chosen number of years, and by default neighbouring configurations share much of their noise, as neighbouring parameters share most of their trades. The configurations are tried in a random order, and the reader decides how many.

Figure 4The best of many, deflated
True edge
Years of data
Neighbours share noise

Of 50 configurations tried on 10 years of data, the best shows an annual Sharpe ratio of 0.54; its true value is 0.00. Judged alone, the chance its true Sharpe ratio is above zero looks like 96%. Charged for all 50 trials, the deflated figure is 37%. The median of its five-by-five neighbourhood is 0.52.

Estimated Sharpe ratio, by configurationBest annual Sharpe ratio so far−0.50.00.51.01.5Judged alone vs deflated0%50%95%110100576Configurations tried (log scale)
Best estimate so farExpected best with no edgeJudged alone: chance the true Sharpe ratio is above zero (PSR)Deflated: chance it beats the best of no edge (DSR)Grid: estimate above zerobelow zero▢best cell
Synthetic data, drawn by a seeded generator; the page runs the same code as the Python build. The grid shows the estimated Sharpe ratio of each configuration tried (aqua above zero, red below, darker further from zero; empty cells not yet tried), with the best outlined and its five-by-five neighbourhood dashed. The chart shows, as more are tried, the best estimate climbing with the expected best under no edge, and the probabilistic Sharpe ratio of the winner, judged alone, climbing towards certainty while the deflated Sharpe ratio does not. Suggested experiments: with no edge, slide to the full grid and watch the undeflated probability pass 95%; switch to the plateau and compare the best cell with its neighbourhood median; then the narrow peak, where the two part company.

Averaged over 300 surfaces with no edge and independent noise, on 10 years of data, the best of one configuration shows an annual Sharpe ratio of about 0.02; the best of 20, about 0.59; the best of all 576, about 0.97. Judged alone, the winner’s probabilistic Sharpe ratio exceeds 95% in 65% of surfaces once 20 are tried, and in 100% of them at the full grid. The deflated Sharpe ratio of the same winner averages 49% at the full grid, and at no single count from five upwards did it exceed 95% in any of the 300 surfaces. When neighbours share their noise the count overstates the number of independent trials, and the average deflated figure drifts lower still, to 36%.

The plateau shows the cost of that strictness. A real edge is there, yet after the whole grid the winner’s true Sharpe ratio averages only 0.27 against an estimate of 1.02, its deflated Sharpe ratio averages 49%, lower than after ten trials (64%), and it cleared 95% in at most 1.0% of surfaces at any single count from twenty upwards. Searching harder for a modest edge makes it harder to demonstrate, because each extra configuration raises the bar. Twenty years of data raise the full-grid average only to 55%: more data shrinks the noise, but not the spread of the true Sharpe ratios across the grid, which the deflation also charges for. A deflated Sharpe ratio below 95% is only weak evidence against an idea.

A real sweep, deflated

To show the pipeline on real data we ran 156 variants of the 200-day moving-average rule from Indicators and the 200-day moving average, on the daily US market from 5 May 1927 to 31 August 2026: lookbacks from 50 to 300 days in steps of ten, a hysteresis band of 0, 1% or 2%, trading at the close that gives the signal (a market-on-close order) or at the next close, and 10 bp a switch. The first 252 sessions of the data warm up the averages, so the five longest lookbacks begin the scored sample in cash, for up to 48 sessions. Each variant was fingerprinted and queued, with one deliberate typo (a lookback of 2,000 days) and two kinds of injected fault on first attempts: some workers vanished mid-run, and some jobs met a simulated data-server timeout.

Jobs queued
157
156 variants and one typo
Attempts
182
13 transient failures retried, 12 leases expired
Results
156
1 job retired as a permanent failure
Trials in the ledger
156
the N the deflation uses
Step by stepFrom a sweep of winners to a refuted prediction
Annual Sharpe ratio over the matched market exposure, each variantSignal’s close, no bandSignal’s close, 1% bandSignal’s close, 2% bandNext close, no bandNext close, 1% bandNext close, 2% bandHolding the market, 0.46Median 0.20240 days, 0.29Best of 156 with no edge, 0.14Probability the best variant’s true edge beats a benchmark50%75%100%95% bar, registered in advanceJudged alone 99.8%Deflated 93%: refutedOverfitting probability (PBO) 59%, see below0.00.10.20.30.40.50.60.7Annual Sharpe ratio (bottom: of the benchmark)
  1. Step 1 of 4

    Against cash

    Each dot is one of the 156 variants, placed at its annual Sharpe ratio in excess of cash. They run from 0.48 to 0.71, and all of them beat holding the market (0.46). That says little about the rule: anything that holds stocks for most of a century earns the equity premium.

  2. Step 2 of 4

    Against the matched market

    Measured over a constant market exposure with the same volatility, the dots slide left, to between 0.02 and 0.29, with a median of 0.20. Trading at the signal’s close keeps more of the edge (a median of 0.24) than trading a close later (0.18).

  3. Step 3 of 4

    The best cell, judged alone

    The best is a 240-day lookback with no band, trading at the signal’s close, at 0.29. Judged as if it were the only variant tried, the probability that its true edge is above zero, its probabilistic Sharpe ratio, is 99.8%.

  4. Step 4 of 4

    Charged for the search

    It was the best of 156. Trials spread as widely as these would produce a best of about 0.14 with no edge at all, and the probability of beating that benchmark is 93%: the curve clears the 95% bar only for benchmarks up to 0.13. The prediction registered before the sweep is refuted.

Kenneth R. French Data Library daily data, 5 May 1927 to 31 August 2026, 10 bp a switch. Top: one dot per variant, one row per trade timing and band, lookbacks from 50 days (top of each row) to 300 (bottom). Bottom: the probabilistic Sharpe ratio of the best variant against a benchmark Sharpe ratio, computed with the formula above from its returns’ skewness and kurtosis. Without script, or with reduced motion, the figure shows its final state.

The story above follows the sweep from its easiest test to its hardest. The harder test, in steps 2 to 4, measures each variant against a constant holding of the market scaled to the same volatility as the rule, a yardstick Volatility targeting develops further. Precisely, we take the Sharpe ratio of the difference between the rule’s excess return and the market’s, scaled to the rule’s volatility over the whole period; it is positive exactly when the rule’s Sharpe ratio beats the market’s. (The cost of 10 bp a switch and the next-close variants put the range against cash below the one in Indicators and the 200-day moving average.)

Figure 5The real sweep’s surface
Sharpe ratio over the matched exposure:0.000.100.200.29signal’s close, band 0%lookback 50 days, signal’s close, band 0%: 0.26lookback 60 days, signal’s close, band 0%: 0.28lookback 70 days, signal’s close, band 0%: 0.25lookback 80 days, signal’s close, band 0%: 0.26lookback 90 days, signal’s close, band 0%: 0.25lookback 100 days, signal’s close, band 0%: 0.25lookback 110 days, signal’s close, band 0%: 0.25lookback 120 days, signal’s close, band 0%: 0.24lookback 130 days, signal’s close, band 0%: 0.23lookback 140 days, signal’s close, band 0%: 0.23lookback 150 days, signal’s close, band 0%: 0.26lookback 160 days, signal’s close, band 0%: 0.27lookback 170 days, signal’s close, band 0%: 0.24lookback 180 days, signal’s close, band 0%: 0.23lookback 190 days, signal’s close, band 0%: 0.26lookback 200 days, signal’s close, band 0%: 0.27lookback 210 days, signal’s close, band 0%: 0.27lookback 220 days, signal’s close, band 0%: 0.27lookback 230 days, signal’s close, band 0%: 0.29lookback 240 days, signal’s close, band 0%: 0.29lookback 250 days, signal’s close, band 0%: 0.26lookback 260 days, signal’s close, band 0%: 0.28lookback 270 days, signal’s close, band 0%: 0.26lookback 280 days, signal’s close, band 0%: 0.24lookback 290 days, signal’s close, band 0%: 0.24lookback 300 days, signal’s close, band 0%: 0.24signal’s close, band 1%lookback 50 days, signal’s close, band 1%: 0.17lookback 60 days, signal’s close, band 1%: 0.20lookback 70 days, signal’s close, band 1%: 0.20lookback 80 days, signal’s close, band 1%: 0.22lookback 90 days, signal’s close, band 1%: 0.20lookback 100 days, signal’s close, band 1%: 0.19lookback 110 days, signal’s close, band 1%: 0.23lookback 120 days, signal’s close, band 1%: 0.25lookback 130 days, signal’s close, band 1%: 0.24lookback 140 days, signal’s close, band 1%: 0.24lookback 150 days, signal’s close, band 1%: 0.25lookback 160 days, signal’s close, band 1%: 0.22lookback 170 days, signal’s close, band 1%: 0.21lookback 180 days, signal’s close, band 1%: 0.23lookback 190 days, signal’s close, band 1%: 0.24lookback 200 days, signal’s close, band 1%: 0.25lookback 210 days, signal’s close, band 1%: 0.25lookback 220 days, signal’s close, band 1%: 0.25lookback 230 days, signal’s close, band 1%: 0.27lookback 240 days, signal’s close, band 1%: 0.28lookback 250 days, signal’s close, band 1%: 0.27lookback 260 days, signal’s close, band 1%: 0.24lookback 270 days, signal’s close, band 1%: 0.25lookback 280 days, signal’s close, band 1%: 0.23lookback 290 days, signal’s close, band 1%: 0.21lookback 300 days, signal’s close, band 1%: 0.19signal’s close, band 2%lookback 50 days, signal’s close, band 2%: 0.06lookback 60 days, signal’s close, band 2%: 0.11lookback 70 days, signal’s close, band 2%: 0.12lookback 80 days, signal’s close, band 2%: 0.14lookback 90 days, signal’s close, band 2%: 0.10lookback 100 days, signal’s close, band 2%: 0.15lookback 110 days, signal’s close, band 2%: 0.17lookback 120 days, signal’s close, band 2%: 0.22lookback 130 days, signal’s close, band 2%: 0.23lookback 140 days, signal’s close, band 2%: 0.23lookback 150 days, signal’s close, band 2%: 0.23lookback 160 days, signal’s close, band 2%: 0.24lookback 170 days, signal’s close, band 2%: 0.25lookback 180 days, signal’s close, band 2%: 0.25lookback 190 days, signal’s close, band 2%: 0.28lookback 200 days, signal’s close, band 2%: 0.27lookback 210 days, signal’s close, band 2%: 0.22lookback 220 days, signal’s close, band 2%: 0.22lookback 230 days, signal’s close, band 2%: 0.22lookback 240 days, signal’s close, band 2%: 0.22lookback 250 days, signal’s close, band 2%: 0.22lookback 260 days, signal’s close, band 2%: 0.19lookback 270 days, signal’s close, band 2%: 0.17lookback 280 days, signal’s close, band 2%: 0.19lookback 290 days, signal’s close, band 2%: 0.19lookback 300 days, signal’s close, band 2%: 0.22next close, band 0%lookback 50 days, next close, band 0%: 0.02lookback 60 days, next close, band 0%: 0.06lookback 70 days, next close, band 0%: 0.14lookback 80 days, next close, band 0%: 0.14lookback 90 days, next close, band 0%: 0.14lookback 100 days, next close, band 0%: 0.17lookback 110 days, next close, band 0%: 0.14lookback 120 days, next close, band 0%: 0.17lookback 130 days, next close, band 0%: 0.18lookback 140 days, next close, band 0%: 0.21lookback 150 days, next close, band 0%: 0.20lookback 160 days, next close, band 0%: 0.19lookback 170 days, next close, band 0%: 0.16lookback 180 days, next close, band 0%: 0.14lookback 190 days, next close, band 0%: 0.18lookback 200 days, next close, band 0%: 0.18lookback 210 days, next close, band 0%: 0.20lookback 220 days, next close, band 0%: 0.20lookback 230 days, next close, band 0%: 0.20lookback 240 days, next close, band 0%: 0.21lookback 250 days, next close, band 0%: 0.19lookback 260 days, next close, band 0%: 0.18lookback 270 days, next close, band 0%: 0.17lookback 280 days, next close, band 0%: 0.15lookback 290 days, next close, band 0%: 0.16lookback 300 days, next close, band 0%: 0.16next close, band 1%lookback 50 days, next close, band 1%: 0.08lookback 60 days, next close, band 1%: 0.11lookback 70 days, next close, band 1%: 0.12lookback 80 days, next close, band 1%: 0.15lookback 90 days, next close, band 1%: 0.13lookback 100 days, next close, band 1%: 0.15lookback 110 days, next close, band 1%: 0.18lookback 120 days, next close, band 1%: 0.21lookback 130 days, next close, band 1%: 0.21lookback 140 days, next close, band 1%: 0.22lookback 150 days, next close, band 1%: 0.20lookback 160 days, next close, band 1%: 0.20lookback 170 days, next close, band 1%: 0.19lookback 180 days, next close, band 1%: 0.15lookback 190 days, next close, band 1%: 0.18lookback 200 days, next close, band 1%: 0.19lookback 210 days, next close, band 1%: 0.21lookback 220 days, next close, band 1%: 0.20lookback 230 days, next close, band 1%: 0.22lookback 240 days, next close, band 1%: 0.23lookback 250 days, next close, band 1%: 0.23lookback 260 days, next close, band 1%: 0.23lookback 270 days, next close, band 1%: 0.20lookback 280 days, next close, band 1%: 0.20lookback 290 days, next close, band 1%: 0.20lookback 300 days, next close, band 1%: 0.17next close, band 2%lookback 50 days, next close, band 2%: 0.04lookback 60 days, next close, band 2%: 0.09lookback 70 days, next close, band 2%: 0.10lookback 80 days, next close, band 2%: 0.11lookback 90 days, next close, band 2%: 0.09lookback 100 days, next close, band 2%: 0.09lookback 110 days, next close, band 2%: 0.13lookback 120 days, next close, band 2%: 0.15lookback 130 days, next close, band 2%: 0.18lookback 140 days, next close, band 2%: 0.18lookback 150 days, next close, band 2%: 0.18lookback 160 days, next close, band 2%: 0.18lookback 170 days, next close, band 2%: 0.18lookback 180 days, next close, band 2%: 0.19lookback 190 days, next close, band 2%: 0.23lookback 200 days, next close, band 2%: 0.21lookback 210 days, next close, band 2%: 0.17lookback 220 days, next close, band 2%: 0.18lookback 230 days, next close, band 2%: 0.17lookback 240 days, next close, band 2%: 0.19lookback 250 days, next close, band 2%: 0.18lookback 260 days, next close, band 2%: 0.17lookback 270 days, next close, band 2%: 0.17lookback 280 days, next close, band 2%: 0.18lookback 290 days, next close, band 2%: 0.18lookback 300 days, next close, band 2%: 0.210.2950100150200250300Lookback, trading days
Annual Sharpe ratio of each variant over a constant market exposure of the same volatility, 5 May 1927 to 31 August 2026, Kenneth R. French Data Library daily data, 10 bp a switch. Rows: when the trade happens, and the width of the hysteresis band. Trading at the signal’s close with no band, every lookback from 50 to 300 days sits on a plateau (0.23 to 0.29); a band pulls down the short lookbacks, the 2% band most; trading a close later lowers the whole surface (median 0.18 against 0.24).

The best variant’s deflated Sharpe ratio already allows for the daily returns’ kurtosis of about 35; at a daily Sharpe ratio this small the adjustment hardly moves it. Most of the edge was earned early: 0.42 from 5 May 1927 to 31 December 1975, and 0.16 since. Trading at the signal’s close also assumes a market-on-close order sent on an estimate of the closing signal, which is feasible for an index but not free.

The probability of backtest overfitting

A second test asks whether the choice among variants carries information. Bailey, Borwein, López de Prado and Zhu’s combinatorially symmetric cross-validation splits the history into 16 blocks, and for each of the 12,870 ways of choosing half of them as the in-sample period, it picks the best variant in sample and finds its rank in the other half. The probability of backtest overfitting is the share of splits in which the in-sample winner lands in the bottom half out of sample.14 Here it is 59%. The in-sample winner traded at the signal’s close in 98% of splits, so the timing choice was learned reliably; the choice of lookback and band within it was not, and the winner landed in the bottom half out of sample slightly more often than not. The honest summary of this sweep is the level of its plateau, the median of 0.20, and the median of the best cell’s neighbourhood (itself and the cells one step away in lookback and band), 0.28, rather than 0.29.

site_research/fieldnotes/pipeline.py
def pbo(returns: np.ndarray, blocks: int = 16) -> dict:
    """Probability of backtest overfitting by combinatorially symmetric cross-validation (Bailey,
    Borwein, Lopez de Prado and Zhu 2017). `returns` is observations x configurations. Split the
    rows into `blocks` pieces; for every way of choosing half of them as the in-sample set, pick
    the best configuration in sample and find its rank out of sample. PBO is the share of splits
    in which the in-sample winner lands in the bottom half out of sample. Also returns how often
    each configuration was the in-sample winner."""
    t, n = returns.shape
    edges = np.linspace(0, t, blocks + 1).astype(int)
    s1 = np.array([returns[a:b].sum(0) for a, b in zip(edges[:-1], edges[1:])])        # blocks x n
    s2 = np.array([(returns[a:b] ** 2).sum(0) for a, b in zip(edges[:-1], edges[1:])])
    cnt = np.diff(edges)
    logits, winners = [], []
    for ins in itertools.combinations(range(blocks), blocks // 2):
        m = np.zeros(blocks, bool)
        m[list(ins)] = True
        sr = []
        for part in (m, ~m):
            k = cnt[part].sum()
            mu = s1[part].sum(0) / k
            sr.append(mu / np.sqrt((s2[part].sum(0) - k * mu ** 2) / (k - 1)))
        best = int(np.argmax(sr[0]))
        winners.append(best)
        rank = (sr[1] < sr[1][best]).sum() + 0.5 * ((sr[1] == sr[1][best]).sum() - 1)   # 0 .. n-1
        w = (rank + 1) / (n + 1)
        logits.append(np.log(w / (1 - w)))
    logits = np.array(logits)
    return {"splits": len(logits), "pbo": float((logits <= 0).mean()), "median_logit": float(np.median(logits)),
            "winners": np.bincount(winners, minlength=n)}
Block sums and sums of squares are computed once, so each of the 12,870 splits costs a few vector operations.

Register the prediction first

Analyses chosen after seeing the data drift towards whatever the data happened to show, and the sciences that depend on statistics increasingly ask for the prediction and the analysis to be registered first.15 The build writes a prediction into the ledger before it queues the sweep, with the condition that would refute it. Ours was written after Indicators and the 200-day moving average had explored the same rule, and its 95% threshold was fixed before the sweep ran. The prediction is part of the code, so its history is in version control. We wrote:

The first half was never at much risk: Indicators and the 200-day moving average had already shown it on the same data, and a prediction made with the answer in view tests nothing. Our count of trials is a floor for the same reason, since the looks taken in that note are not in this ledger. The second half was open, and it failed. Because it was written down with its threshold, a near miss cannot become “close enough”. Scored over many sweeps, predictions also measure the researcher: someone whose predictions usually fail is choosing ideas worse than they believe, and no single backtest can show that.

A checklist

What a stored result should carry, and what an analysis should report:

  1. The fingerprint, and the full specification it was computed from: code digest, canonical parameters with seeds, data digest and engine versions.
  2. The sweep it belongs to, the prediction registered for that sweep, and the verdict.
  3. The per-period Sharpe ratio, the number of observations, skewness and kurtosis, so that it can be deflated later.
  4. The number of trials in the ledger when it was selected, and the deflated Sharpe ratio at that count.
  5. The surface around it: the median of the sweep and of the winner’s neighbourhood, never the best cell alone.
  6. Every failure, with its class, so that unexplored regions of the parameter space are visible.

Sizing comes next on the path: Volatility targeting takes a rule that has earned a place and decides how much of it to hold.

References

  1. Sandve, G. K., Nekrutenko, A., Taylor, J., & Hovig, E. (2013). Ten simple rules for reproducible computational research. PLoS Computational Biology, 9(10), e1003285. https://doi.org/10.1371/journal.pcbi.1003285
  2. French, K. R. (n.d.). Data Library: Fama/French factors (daily) and momentum factor (daily). Tuck School of Business, Dartmouth College. mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html
  3. Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J.-F., & Dennison, D. (2015). Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems 28 (NIPS 2015). papers.nips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html
  4. Gray, C. G., & Cheriton, D. R. (1989). Leases: An efficient fault-tolerant mechanism for distributed file cache consistency. Proceedings of the Twelfth ACM Symposium on Operating Systems Principles, 202–210. https://doi.org/10.1145/74850.74870
  5. Brooker, M. (2015). Exponential backoff and jitter. AWS Architecture Blog, Amazon Web Services. aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/
  6. PostgreSQL Global Development Group (n.d.). SELECT: the locking clause (FOR UPDATE, SKIP LOCKED). PostgreSQL documentation. www.postgresql.org/docs/current/sql-select.html
  7. Amazon Web Services (n.d.). Amazon SQS visibility timeout. Amazon Simple Queue Service Developer Guide. docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-visibility-timeout.html
  8. White, H. (2000). A reality check for data snooping. Econometrica, 68(5), 1097–1126. https://doi.org/10.1111/1468-0262.00152
  9. Hansen, P. R. (2005). A test for superior predictive ability. Journal of Business & Economic Statistics, 23(4), 365–380. https://doi.org/10.1198/073500105000000063
  10. Harvey, C. R., Liu, Y., & Zhu, H. (2016). …and the cross-section of expected returns. Review of Financial Studies, 29(1), 5–68. https://doi.org/10.1093/rfs/hhv059
  11. Bailey, D. H., & López de Prado, M. (2012). The Sharpe ratio efficient frontier. Journal of Risk, 15(2), 3–44. https://doi.org/10.21314/JOR.2012.255
  12. Bailey, D. H., & López de Prado, M. (2014). The deflated Sharpe ratio: Correcting for selection bias, backtest overfitting, and non-normality. Journal of Portfolio Management, 40(5), 94–107. https://doi.org/10.3905/jpm.2014.40.5.094
  13. López de Prado, M., & Lewis, M. J. (2019). Detection of false investment strategies using unsupervised learning methods. Quantitative Finance, 19(9), 1555–1565. https://doi.org/10.1080/14697688.2019.1622311
  14. Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2017). The probability of backtest overfitting. Journal of Computational Finance, 20(4), 39–69. https://doi.org/10.21314/JCF.2016.322
  15. Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606. https://doi.org/10.1073/pnas.1708274114