Research infrastructure · Intermediate
Running backtests at scale
A backtest can be trusted only if it can be reproduced and counted, and both have to be built into the pipeline before the first sweep runs.
Educational material only. Not investment advice.
Contents
Abstract
Run thousands of backtests and two things go wrong: results that cannot be reproduced, and winners chosen from searches nobody counted. We identify each run by a hash of its code, parameters, data and engine, run the jobs through a leased queue that tells transient failures from permanent ones, and write each attempt to a ledger. The deflated Sharpe ratio then judges the best result against the size of the search. On a real sweep of 156 variants of the 200-day rule, the best variant’s edge over a matched market exposure has a probabilistic Sharpe ratio of 99.8% judged alone and a deflated one of 93%, short of the 95% bar we set.
Key takeaways
- Identify each run by a hash of its code, parameters, data and engine versions; cache and reproduce by that key.
- Lease jobs, classify each failure as transient or permanent, retry transient ones with jittered exponential backoff, and retire a job that keeps killing its workers.
- Keep a ledger of all trials, including abandoned sweeps and failures; the trial count is part of every result.
- Judge the best result with the deflated Sharpe ratio, and the choice among variants with the probability of backtest overfitting; neither is a p-value.
- Register a prediction and its refutation condition before a sweep runs, and score it afterwards.
- Summarise a parameter surface by its level and its neighbourhoods, never by its best cell.
Before you start
- Market anomalies, for data snooping and the noise ceiling
- QuantConnect and LEAN, or any engine that runs a backtest from code, for what a run depends on
- Basic Python, SQL or pandas, and the Sharpe ratio
A single backtest is easy to trust and easy to get wrong. A research programme runs thousands, and its errors are quiet ones. A result computed on last month’s code sits beside one computed on today’s, and the strongest variant of a sweep is reported without the hundred weaker ones tried beside it. Neither can be repaired from the results afterwards, so the fixes have to be in place before the first sweep runs.
What a result must carry
A backtest result is a function of four things: the code, the parameters, the data and the engine that ran them. If any of the four can change without the result changing its name, the result cannot be reproduced, and in a long research programme one of them always changes. Code is edited, a data vendor corrects a history, a library upgrade alters a default. Recording exactly how each result was produced is the first rule of reproducible computation.1 We make the record the result’s name: each run is identified by a cryptographic hash of all four inputs, so the same specification always has the same key and a different specification can never borrow an old one.
| Input | What we record | What goes wrong without it |
|---|---|---|
| Code | A hash of the bytes of each file that computes the stored result | A result survives a bug fix it never saw |
| Parameters | Canonical JSON: sorted keys and sets, numbers as floats, seeds included; the caller writes out every default | One idea gets two names, or a parameter left out lets two different runs share one |
| Data | A hash of the raw files, which fixes the vintage | A corrected history silently changes old results |
| Engine | Language and library versions, and the backtest engine’s version | Defaults change underneath the run |
def digest(*chunks: bytes) -> str:
h = hashlib.sha256()
for c in chunks:
h.update(hashlib.sha256(c).digest())
return h.hexdigest()
def canonical(x):
"""Parameters in one form, the same in every process: numbers become floats (so 1 and 1.0
are the same parameter) except integers beyond 2**53, which a float would merge with their
neighbours; sets are sorted, since their order changes with Python's hash seed; nested
dictionaries and lists are canonicalised too. Keys become strings and are sorted when the
JSON is written. Defaults are the caller's to write out."""
if isinstance(x, np.bool_):
return bool(x)
if isinstance(x, bool) or x is None or isinstance(x, str):
return x
if isinstance(x, (int, np.integer)) and abs(int(x)) > 2 ** 53:
return int(x)
if isinstance(x, (int, float, np.integer, np.floating)):
return float(x)
if isinstance(x, dict):
return {str(k): canonical(v) for k, v in x.items()}
if isinstance(x, (set, frozenset)):
return sorted((canonical(v) for v in x), key=lambda v: json.dumps(v, sort_keys=True))
return [canonical(v) for v in x]
def fingerprint(code_files: list[Path], params: dict, data_digest: str) -> tuple[str, dict]:
"""The key of one run: a hash of the code's bytes, the parameters in canonical form, the
data's digest and the engine's versions. Change any of them and the key changes. A random
seed, where a run has one, is a parameter: two seeds are two trials."""
spec = {
"code": digest(*(Path(p).read_bytes() for p in sorted(map(str, code_files)))),
"params": canonical(params),
"data": data_digest,
"engine": {"python": platform.python_version(), "numpy": np.__version__, "pandas": pd.__version__,
"scipy": scipy.__version__},
}
canon = json.dumps(spec, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(canon.encode()).hexdigest()[:16], spec # 64 bits: ample below ~10^8 runsA result keyed this way can be cached, so a sweep that is interrupted and restarted, or a colleague’s identical request, costs nothing the second time. It can be reproduced, because the specification is stored beside it. And a result that disagrees with its own rerun is a finding in itself: the engine is not deterministic, or an input escaped the fingerprint. The public data in this series are corrected and extended from time to time, which is why each note states the range it used and the build hashes the files it read.2 On a hosted engine whose data cannot be hashed, the practical substitute is a canary: one fixed backtest, rerun on a schedule, whose result must not change. The sources of drift inside an engine, from fill models to corporate actions, are described in QuantConnect and LEAN.
The pipeline
The pipeline that turns ideas into counted, reproducible results has the same shape at any scale. A hypothesis, with a prediction written down before anything runs, becomes a set of specifications; each specification is fingerprinted and queued; workers lease jobs from the queue, run them and write results; the ledger records every attempt, successful or not; and analysis reads the ledger.
The arrow from the ledger to analysis is the one that is easiest to leave out. A results store keeps what worked; a ledger keeps what was tried. The machine-learning literature describes the cost of the untraced alternative: systems that accumulate configurations and data dependencies nobody can trace, and results nobody can reproduce, pay for it later as technical debt.3 The whole state fits in four tables:
SCHEMA = """
CREATE TABLE IF NOT EXISTS jobs (
fingerprint TEXT PRIMARY KEY,
sweep TEXT NOT NULL,
spec TEXT NOT NULL, -- canonical JSON: code, params, data, engine
state TEXT NOT NULL DEFAULT 'queued', -- queued | leased | done | dead
attempts INTEGER NOT NULL DEFAULT 0,
lease_owner TEXT,
lease_until REAL,
not_before REAL NOT NULL DEFAULT 0, -- earliest time a retry may start
error TEXT
);
CREATE TABLE IF NOT EXISTS results (
fingerprint TEXT PRIMARY KEY, -- one result per specification, ever
metrics TEXT NOT NULL,
worker TEXT NOT NULL,
finished REAL NOT NULL
);
CREATE TABLE IF NOT EXISTS ledger ( -- every attempt, including the failures
id INTEGER PRIMARY KEY AUTOINCREMENT,
sweep TEXT NOT NULL,
fingerprint TEXT NOT NULL,
outcome TEXT NOT NULL, -- result | transient | permanent | lease_expired | stale
sharpe REAL, n_obs INTEGER, skew REAL, kurt REAL,
note TEXT,
at REAL NOT NULL
);
CREATE TABLE IF NOT EXISTS predictions ( -- written BEFORE the sweep runs
sweep TEXT PRIMARY KEY,
claim TEXT NOT NULL,
refuted_if TEXT NOT NULL,
verdict TEXT
);
"""Leases, retries and idempotence
At scale, workers die, services time out and some jobs are simply wrong, and a queue has to treat each case differently.
Leases
A worker that takes a job holds it on a lease: a claim that expires after a fixed time unless the worker renews it or reports the job finished. If the worker crashes, loses its network or is shut down, nobody has to notice; the lease runs out and the job becomes available again. Gray and Cheriton introduced leases for distributed file caches for this reason: a lock held by a dead machine is held forever, a lease is not.4 The lease should be comfortably longer than the slowest normal run, and a worker on a long run renews it. A job whose lease expires on every attempt is probably killing its workers, so after 4 attempts it is retired like any other permanent failure.
def lease(db, worker: str, now: float, ttl: float = 900.0, max_attempts: int = MAX_ATTEMPTS):
"""Claim one runnable job: queued and past its backoff, or leased by a worker whose lease
has expired (that worker is presumed dead). A job whose lease expires on its last attempt is
retired instead, like any job that has used its attempts: when leases keep expiring, the job, not the worker, is the likely culprit."""
db.execute("BEGIN IMMEDIATE")
try:
while True:
row = db.execute(
"SELECT fingerprint, sweep, spec, state, attempts FROM jobs WHERE (state = 'queued' AND not_before <= ?) "
"OR (state = 'leased' AND lease_until < ?) ORDER BY not_before, fingerprint LIMIT 1", (now, now)).fetchone()
if row is None:
db.execute("COMMIT")
return None
fp, sweep, spec, state, attempts = row
if state == "leased":
db.execute("INSERT INTO ledger (sweep, fingerprint, outcome, note, at) VALUES (?, ?, 'lease_expired', ?, ?)",
(sweep, fp, "worker lost its lease", now))
if attempts >= max_attempts:
db.execute("UPDATE jobs SET state = 'dead', lease_owner = NULL, error = ? WHERE fingerprint = ?",
("lease expired on its last attempt", fp))
db.execute("INSERT INTO ledger (sweep, fingerprint, outcome, note, at) VALUES (?, ?, 'permanent', ?, ?)",
(sweep, fp, "lease expired on its last attempt", now))
continue
db.execute("UPDATE jobs SET state = 'leased', lease_owner = ?, lease_until = ?, attempts = attempts + 1 "
"WHERE fingerprint = ?", (worker, now + ttl, fp))
db.execute("COMMIT")
return fp, sweep, json.loads(spec)
except BaseException:
db.execute("ROLLBACK")
raise
def renew(db, fp: str, worker: str, now: float, ttl: float = 900.0) -> bool:
"""Extend a lease the worker still holds. False means another worker has it now: the heartbeat
stops renewing, and whatever the run reports later is harmless (a result is kept only if it is
the first for its fingerprint, and a failure from a worker without the lease changes nothing)."""
return db.execute("UPDATE jobs SET lease_until = ? WHERE fingerprint = ? AND state = 'leased' AND lease_owner = ?",
(now + ttl, fp, worker)).rowcount == 1Transient and permanent failures
A timeout from a data server is likely to succeed on a second attempt; a parameter outside its valid range will fail every time. Retrying the second kind wastes capacity and hides the bug; not retrying the first kind loses work. So each failure is classified before anything else happens. Transient failures go back to the queue after a wait that grows exponentially with each attempt and is drawn at random up to that bound (“full jitter”), so that workers that failed together do not all retry in the same second.5 After 4 attempts they are retired too. Permanent failures are retired at once, with the error kept for whoever fixes them. Only the worker that still holds the lease may re-queue or retire the job, so a worker that was presumed dead and wakes up cannot undo the work of its replacement. A late result is still accepted, because it is identical to the one the replacement will produce.
def backoff(attempt: int, rng: random.Random, base: float = 30.0, cap: float = 3600.0) -> float:
"""Exponential backoff with full jitter: a uniform wait up to base * 2^attempt, capped, so
workers that failed together do not retry together."""
return rng.uniform(0, min(cap, base * 2 ** attempt))
def fail(db, fp: str, worker: str, exc: BaseException, now: float, rng: random.Random,
max_attempts: int = MAX_ATTEMPTS) -> str:
"""Transient failures go back to the queue after a backoff, until the attempts run out;
anything else is permanent at once. Only the worker that still holds the lease may change the
job: a report from any other is a stale one, kept in the ledger as such and changing nothing."""
transient = isinstance(exc, (TransientError, TimeoutError, ConnectionError))
db.execute("BEGIN IMMEDIATE")
try:
sweep, attempts = db.execute("SELECT sweep, attempts FROM jobs WHERE fingerprint = ?", (fp,)).fetchone()
retry = transient and attempts < max_attempts
owned = db.execute(
"UPDATE jobs SET state = ?, lease_owner = NULL, not_before = ?, error = ? "
"WHERE fingerprint = ? AND state = 'leased' AND lease_owner = ?",
("queued" if retry else "dead", now + backoff(attempts, rng) if retry else 0.0,
f"{type(exc).__name__}: {exc}", fp, worker)).rowcount == 1
outcome = ("transient" if transient else "permanent") if owned else "stale"
db.execute("INSERT INTO ledger (sweep, fingerprint, outcome, note, at) VALUES (?, ?, ?, ?, ?)",
(sweep, fp, outcome, f"{type(exc).__name__}: {exc}", now))
db.execute("COMMIT")
except BaseException:
db.execute("ROLLBACK")
raise
return "stale" if not owned else "retry" if retry else "dead"Idempotent results
A worker can lose its lease and still finish. The job has meanwhile been leased to someone else, and two results for the same fingerprint arrive. Because a fingerprint fixes the code, parameters, data and engine, the two results should be identical, so the store accepts the first and ignores the second. Writing results keyed by fingerprint makes recording them idempotent: doing it twice, or again after a crash halfway through, leaves the same state as doing it once.
def complete(db, fp: str, worker: str, metrics: dict, now: float) -> bool:
"""Record a result. A worker that lost its lease may still finish: the result is the same
(same code, parameters and data), so the first write wins and a second changes nothing. It is
accepted from any worker, even for a job its replacement holds or one already retired, because
it is the result the replacement would produce."""
db.execute("BEGIN IMMEDIATE")
try:
sweep = db.execute("SELECT sweep FROM jobs WHERE fingerprint = ?", (fp,)).fetchone()[0]
new = db.execute("INSERT OR IGNORE INTO results VALUES (?, ?, ?, ?)",
(fp, json.dumps(metrics), worker, now)).rowcount == 1
if new:
db.execute("INSERT INTO ledger (sweep, fingerprint, outcome, sharpe, n_obs, skew, kurt, at) "
"VALUES (?, ?, 'result', ?, ?, ?, ?, ?)",
(sweep, fp, metrics["sharpe"], metrics["n_obs"], metrics["skew"], metrics["kurt"], now))
db.execute("UPDATE jobs SET state = 'done', lease_owner = NULL, error = NULL WHERE fingerprint = ?", (fp,))
db.execute("COMMIT")
return new
except BaseException:
db.execute("ROLLBACK")
raise
def worker_loop(path: str, name: str, run, ttl: float = 900.0) -> int:
"""A real worker: lease, run with a heartbeat that renews the lease, record the outcome, and
repeat until nothing is runnable now (jobs waiting out a backoff are left for the next pass).
Any number of these, in threads or processes, can share one database file."""
db, rng, done = connect(path), random.Random(name), 0
while (job := lease(db, name, time.time(), ttl)) is not None:
fp, _, spec = job
stop = threading.Event()
def heartbeat():
hb = connect(path)
while not stop.wait(ttl / 3) and renew(hb, fp, name, time.time(), ttl):
pass
hb.close()
threading.Thread(target=heartbeat, daemon=True).start()
try:
complete(db, fp, name, run(spec["params"]), time.time())
done += 1
except Exception as exc: # noqa: BLE001 -- classified in fail()
fail(db, fp, name, exc, time.time(), rng)
finally:
stop.set()
return done| Event | Class | What the queue does | What the ledger records |
|---|---|---|---|
| Worker crashes mid-run | Lease expiry | The job returns to the queue when the lease runs out; retired once it has used 4 attempts | lease_expired, then permanent |
| Data server timeout | Transient | Retry after a jittered, growing wait; retire after 4 attempts | transient, with the error |
| Parameter out of range, code error | Permanent | Retire at once, keep the error | permanent, with the error |
| A stale worker reports late | Harmless | The first result for a fingerprint is kept and a second is ignored; a late failure cannot change the job | stale, for a late failure; nothing, for a second result |
The lab runs this queue, ported line for line, on six jobs and three workers, with a clock that moves in thirty-second ticks. The faults are the reader’s to choose.
At 5 min: 8 attempts have started and 5 of 6 jobs have a result; 1 transient failure is in the ledger, 1 lease has expired and 1 job has been retired. A late report from w2, which no longer held the lease, is in the ledger as stale and changed nothing. The queue is empty after 5 min, and every fingerprint has at most one result.
| Ledger, at min:s | Job | Outcome |
|---|---|---|
| 1:30 | 100 days | result |
| 1:30 | 200 days | result |
| 2:30 | 250 days | transient failure |
| 3:00 | 2,000 days | permanent failure |
| 3:00 | 300 days | result |
| 3:00 | 150 days | lease expired |
| 3:30 | 150 days | late report, not the lease holder: ignored |
| 4:30 | 150 days | result |
| 5:00 | 250 days | result |
Worker w2 vanishes the moment it leases its first job, as a machine that loses power would; the 250-day job times out after one minute on the attempts chosen above. A late report from w2 arrives one tick after another worker has taken its job over.
pipeline.py, and the site’s build checks that the page writes the same ledger as the Python for the same faults. Each row is a job and each bar an attempt, from its lease to its outcome. Suggested experiments: let the worker die and watch its job sit untouched until the lease runs out; make the data server time out every time and count the attempts before the job is retired; wake the dead worker with a timeout while its replacement is running, and see that the job is not re-queued under it; then wake it with a result and find the replacement’s identical result ignored.From one machine to many
SQLite with an immediate transaction is enough for the workers of one machine. Across machines the same design carries over unchanged in its logic: a PostgreSQL table claimed with SELECT … FOR UPDATE SKIP LOCKED, or a managed queue whose visibility timeout is a lease under another name.6,7 The worker’s run function is where an engine plugs in. For a cloud backtest it submits the project, polls until the run finishes and returns the statistics; the platform’s rate limits and timeouts are raised as transient errors, and a compile error as a permanent one.
Throughput and cost
The arithmetic of a sweep is plain: the number of runs, times the minutes each takes, divided by the number of workers, plus retries. On a shared engine much of a backtest’s wall-clock time can go on waiting for capacity and loading data. Loading data once per worker and caching by fingerprint remove most of that. A coarse sweep to find the plateau, then a finer one around it, spends compute where the decision is. Each extra run also raises the bar the best one must clear.
Count every trial
In Market anomalies we showed how good the best of many useless strategies looks: with a hundred independent tries on ten years of data, the best is expected to show an annual Sharpe ratio of about 0.80 with no edge at all. The number of trials is therefore part of every result. A Sharpe ratio of 1.0 reported after three attempts is a different claim from the same figure reported after three hundred, and only the ledger records which was made. White’s reality check and Hansen’s test for superior predictive ability test the best of many models against that selection, and, as Market anomalies noted, Harvey, Liu and Zhu put the bar for a new factor at a t-statistic of about three.8,9,10
What counts as a trial
A trial is any configuration whose result anyone saw on the same question, in any sweep. That includes sweeps abandoned because they looked bad, and reruns after a code change, since a changed fingerprint is a new trial. The count N is the number of distinct fingerprints with a result, and the spread below is taken over all of them. A separate question gets a separate ledger. Failures that never produced a result did not inform the choice, but they are recorded too: a sweep whose failures cluster in one region of the parameter space has not explored that region, and its apparent optimum may simply be the edge of what ran. What should not inflate the count is the same specification run twice, which the fingerprint makes easy to tell apart.
The deflated Sharpe ratio
Bailey and López de Prado’s probabilistic Sharpe ratio gives the probability that a strategy’s true Sharpe ratio exceeds a benchmark SR*, given an estimate SR from n observations with skewness γ3 and kurtosis γ4:11
The deflated Sharpe ratio sets the benchmark to the best Sharpe ratio that N trials with no edge would be expected to produce, given how widely the trials’ estimates were spread:12
Correlation between trials enters twice. Near-copies of a configuration agree with each other, so the spread, and with it the penalty, is small; but the count N still treats each near-copy as a separate try. When trials fall into a few clusters, that makes the correction conservative, which is why López de Prado and Lewis estimate the effective number of trials by clustering the trials’ returns.13 The deflated Sharpe ratio is also not a p-value. Its benchmark is only the typical best of a search with no edge, but a 95% reading asks the winner to beat that typical best by a further 1.65 standard errors of a single estimate, far more than the best of many varies from one search to the next. With no edge it passes much less often than one time in twenty, and the lab below shows how much less.
def probabilistic_sharpe(sr: float, sr0: float, n_obs: int, skw: float = 0.0, kurt: float = 3.0) -> float:
"""P(true Sharpe > sr0) given an estimate `sr` from n_obs returns, allowing for skewness and
fat tails (Bailey and Lopez de Prado 2012). All Sharpe ratios per period, not annualised."""
return float(norm.cdf((sr - sr0) * np.sqrt(n_obs - 1) / np.sqrt(1 - skw * sr + (kurt - 1) / 4 * sr ** 2)))
def deflated_sharpe(sr: float, trial_srs, n_obs: int, skw: float = 0.0, kurt: float = 3.0) -> dict:
"""The deflated Sharpe ratio (Bailey and Lopez de Prado 2014): the probabilistic Sharpe ratio
of the selected strategy, measured against the best Sharpe ratio that `len(trial_srs)` trials
with no edge would be expected to produce, given how spread out the trials were."""
trial_srs = np.asarray(trial_srs, float)
n = len(trial_srs)
sr0 = float(np.std(trial_srs, ddof=1) * expected_max_sharpe(n, 1.0)) if n > 1 else 0.0
return {"trials": n, "sr0": sr0, "psr": probabilistic_sharpe(sr, 0.0, n_obs, skw, kurt),
"dsr": probabilistic_sharpe(sr, sr0, n_obs, skw, kurt)}Any backtest can be deflated the same way, given its Sharpe ratio, its length, the size of the search behind it and the shape of its returns.
Judged alone, the chance that the true Sharpe ratio is above zero is above 99.9%. After 200 trials, the best result that no edge at all would be expected to produce is an annual Sharpe ratio of 0.59, and the deflated Sharpe ratio is 95.5%. To reach 95% after this search, the backtest would have needed an annual Sharpe ratio of 1.18. It stops passing at about 263 trials.
An interactive laboratory
The lab makes the correction visible on a synthetic grid of 576 configurations, the kind of two-parameter surface a sweep produces. Each configuration has a true annual Sharpe ratio, chosen by the reader: zero everywhere, a broad plateau, or a narrow peak of the same height. Its estimate is the true value plus noise with the standard error of a Sharpe ratio measured over the chosen number of years, and by default neighbouring configurations share much of their noise, as neighbouring parameters share most of their trades. The configurations are tried in a random order, and the reader decides how many.
Of 50 configurations tried on 10 years of data, the best shows an annual Sharpe ratio of 0.54; its true value is 0.00. Judged alone, the chance its true Sharpe ratio is above zero looks like 96%. Charged for all 50 trials, the deflated figure is 37%. The median of its five-by-five neighbourhood is 0.52.
Averaged over 300 surfaces with no edge and independent noise, on 10 years of data, the best of one configuration shows an annual Sharpe ratio of about 0.02; the best of 20, about 0.59; the best of all 576, about 0.97. Judged alone, the winner’s probabilistic Sharpe ratio exceeds 95% in 65% of surfaces once 20 are tried, and in 100% of them at the full grid. The deflated Sharpe ratio of the same winner averages 49% at the full grid, and at no single count from five upwards did it exceed 95% in any of the 300 surfaces. When neighbours share their noise the count overstates the number of independent trials, and the average deflated figure drifts lower still, to 36%.
The plateau shows the cost of that strictness. A real edge is there, yet after the whole grid the winner’s true Sharpe ratio averages only 0.27 against an estimate of 1.02, its deflated Sharpe ratio averages 49%, lower than after ten trials (64%), and it cleared 95% in at most 1.0% of surfaces at any single count from twenty upwards. Searching harder for a modest edge makes it harder to demonstrate, because each extra configuration raises the bar. Twenty years of data raise the full-grid average only to 55%: more data shrinks the noise, but not the spread of the true Sharpe ratios across the grid, which the deflation also charges for. A deflated Sharpe ratio below 95% is only weak evidence against an idea.
A real sweep, deflated
To show the pipeline on real data we ran 156 variants of the 200-day moving-average rule from Indicators and the 200-day moving average, on the daily US market from 5 May 1927 to 31 August 2026: lookbacks from 50 to 300 days in steps of ten, a hysteresis band of 0, 1% or 2%, trading at the close that gives the signal (a market-on-close order) or at the next close, and 10 bp a switch. The first 252 sessions of the data warm up the averages, so the five longest lookbacks begin the scored sample in cash, for up to 48 sessions. Each variant was fingerprinted and queued, with one deliberate typo (a lookback of 2,000 days) and two kinds of injected fault on first attempts: some workers vanished mid-run, and some jobs met a simulated data-server timeout.
- Step 1 of 4
Against cash
Each dot is one of the 156 variants, placed at its annual Sharpe ratio in excess of cash. They run from 0.48 to 0.71, and all of them beat holding the market (0.46). That says little about the rule: anything that holds stocks for most of a century earns the equity premium.
- Step 2 of 4
Against the matched market
Measured over a constant market exposure with the same volatility, the dots slide left, to between 0.02 and 0.29, with a median of 0.20. Trading at the signal’s close keeps more of the edge (a median of 0.24) than trading a close later (0.18).
- Step 3 of 4
The best cell, judged alone
The best is a 240-day lookback with no band, trading at the signal’s close, at 0.29. Judged as if it were the only variant tried, the probability that its true edge is above zero, its probabilistic Sharpe ratio, is 99.8%.
- Step 4 of 4
Charged for the search
It was the best of 156. Trials spread as widely as these would produce a best of about 0.14 with no edge at all, and the probability of beating that benchmark is 93%: the curve clears the 95% bar only for benchmarks up to 0.13. The prediction registered before the sweep is refuted.
Kenneth R. French Data Library daily data, 5 May 1927 to 31 August 2026, 10 bp a switch. Top: one dot per variant, one row per trade timing and band, lookbacks from 50 days (top of each row) to 300 (bottom). Bottom: the probabilistic Sharpe ratio of the best variant against a benchmark Sharpe ratio, computed with the formula above from its returns’ skewness and kurtosis. Without script, or with reduced motion, the figure shows its final state.
The story above follows the sweep from its easiest test to its hardest. The harder test, in steps 2 to 4, measures each variant against a constant holding of the market scaled to the same volatility as the rule, a yardstick Volatility targeting develops further. Precisely, we take the Sharpe ratio of the difference between the rule’s excess return and the market’s, scaled to the rule’s volatility over the whole period; it is positive exactly when the rule’s Sharpe ratio beats the market’s. (The cost of 10 bp a switch and the next-close variants put the range against cash below the one in Indicators and the 200-day moving average.)
The best variant’s deflated Sharpe ratio already allows for the daily returns’ kurtosis of about 35; at a daily Sharpe ratio this small the adjustment hardly moves it. Most of the edge was earned early: 0.42 from 5 May 1927 to 31 December 1975, and 0.16 since. Trading at the signal’s close also assumes a market-on-close order sent on an estimate of the closing signal, which is feasible for an index but not free.
The probability of backtest overfitting
A second test asks whether the choice among variants carries information. Bailey, Borwein, López de Prado and Zhu’s combinatorially symmetric cross-validation splits the history into 16 blocks, and for each of the 12,870 ways of choosing half of them as the in-sample period, it picks the best variant in sample and finds its rank in the other half. The probability of backtest overfitting is the share of splits in which the in-sample winner lands in the bottom half out of sample.14 Here it is 59%. The in-sample winner traded at the signal’s close in 98% of splits, so the timing choice was learned reliably; the choice of lookback and band within it was not, and the winner landed in the bottom half out of sample slightly more often than not. The honest summary of this sweep is the level of its plateau, the median of 0.20, and the median of the best cell’s neighbourhood (itself and the cells one step away in lookback and band), 0.28, rather than 0.29.
def pbo(returns: np.ndarray, blocks: int = 16) -> dict:
"""Probability of backtest overfitting by combinatorially symmetric cross-validation (Bailey,
Borwein, Lopez de Prado and Zhu 2017). `returns` is observations x configurations. Split the
rows into `blocks` pieces; for every way of choosing half of them as the in-sample set, pick
the best configuration in sample and find its rank out of sample. PBO is the share of splits
in which the in-sample winner lands in the bottom half out of sample. Also returns how often
each configuration was the in-sample winner."""
t, n = returns.shape
edges = np.linspace(0, t, blocks + 1).astype(int)
s1 = np.array([returns[a:b].sum(0) for a, b in zip(edges[:-1], edges[1:])]) # blocks x n
s2 = np.array([(returns[a:b] ** 2).sum(0) for a, b in zip(edges[:-1], edges[1:])])
cnt = np.diff(edges)
logits, winners = [], []
for ins in itertools.combinations(range(blocks), blocks // 2):
m = np.zeros(blocks, bool)
m[list(ins)] = True
sr = []
for part in (m, ~m):
k = cnt[part].sum()
mu = s1[part].sum(0) / k
sr.append(mu / np.sqrt((s2[part].sum(0) - k * mu ** 2) / (k - 1)))
best = int(np.argmax(sr[0]))
winners.append(best)
rank = (sr[1] < sr[1][best]).sum() + 0.5 * ((sr[1] == sr[1][best]).sum() - 1) # 0 .. n-1
w = (rank + 1) / (n + 1)
logits.append(np.log(w / (1 - w)))
logits = np.array(logits)
return {"splits": len(logits), "pbo": float((logits <= 0).mean()), "median_logit": float(np.median(logits)),
"winners": np.bincount(winners, minlength=n)}Register the prediction first
Analyses chosen after seeing the data drift towards whatever the data happened to show, and the sciences that depend on statistics increasingly ask for the prediction and the analysis to be registered first.15 The build writes a prediction into the ledger before it queues the sweep, with the condition that would refute it. Ours was written after Indicators and the 200-day moving average had explored the same rule, and its 95% threshold was fixed before the sweep ran. The prediction is part of the code, so its history is in version control. We wrote:
The first half was never at much risk: Indicators and the 200-day moving average had already shown it on the same data, and a prediction made with the answer in view tests nothing. Our count of trials is a floor for the same reason, since the looks taken in that note are not in this ledger. The second half was open, and it failed. Because it was written down with its threshold, a near miss cannot become “close enough”. Scored over many sweeps, predictions also measure the researcher: someone whose predictions usually fail is choosing ideas worse than they believe, and no single backtest can show that.
A checklist
What a stored result should carry, and what an analysis should report:
- The fingerprint, and the full specification it was computed from: code digest, canonical parameters with seeds, data digest and engine versions.
- The sweep it belongs to, the prediction registered for that sweep, and the verdict.
- The per-period Sharpe ratio, the number of observations, skewness and kurtosis, so that it can be deflated later.
- The number of trials in the ledger when it was selected, and the deflated Sharpe ratio at that count.
- The surface around it: the median of the sweep and of the winner’s neighbourhood, never the best cell alone.
- Every failure, with its class, so that unexplored regions of the parameter space are visible.
Sizing comes next on the path: Volatility targeting takes a rule that has earned a place and decides how much of it to hold.
References
- Sandve, G. K., Nekrutenko, A., Taylor, J., & Hovig, E. (2013). Ten simple rules for reproducible computational research. PLoS Computational Biology, 9(10), e1003285. https://doi.org/10.1371/journal.pcbi.1003285
- French, K. R. (n.d.). Data Library: Fama/French factors (daily) and momentum factor (daily). Tuck School of Business, Dartmouth College. mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html
- Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J.-F., & Dennison, D. (2015). Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems 28 (NIPS 2015). papers.nips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html
- Gray, C. G., & Cheriton, D. R. (1989). Leases: An efficient fault-tolerant mechanism for distributed file cache consistency. Proceedings of the Twelfth ACM Symposium on Operating Systems Principles, 202–210. https://doi.org/10.1145/74850.74870
- Brooker, M. (2015). Exponential backoff and jitter. AWS Architecture Blog, Amazon Web Services. aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/
- PostgreSQL Global Development Group (n.d.). SELECT: the locking clause (FOR UPDATE, SKIP LOCKED). PostgreSQL documentation. www.postgresql.org/docs/current/sql-select.html
- Amazon Web Services (n.d.). Amazon SQS visibility timeout. Amazon Simple Queue Service Developer Guide. docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-visibility-timeout.html
- White, H. (2000). A reality check for data snooping. Econometrica, 68(5), 1097–1126. https://doi.org/10.1111/1468-0262.00152
- Hansen, P. R. (2005). A test for superior predictive ability. Journal of Business & Economic Statistics, 23(4), 365–380. https://doi.org/10.1198/073500105000000063
- Harvey, C. R., Liu, Y., & Zhu, H. (2016). …and the cross-section of expected returns. Review of Financial Studies, 29(1), 5–68. https://doi.org/10.1093/rfs/hhv059
- Bailey, D. H., & López de Prado, M. (2012). The Sharpe ratio efficient frontier. Journal of Risk, 15(2), 3–44. https://doi.org/10.21314/JOR.2012.255
- Bailey, D. H., & López de Prado, M. (2014). The deflated Sharpe ratio: Correcting for selection bias, backtest overfitting, and non-normality. Journal of Portfolio Management, 40(5), 94–107. https://doi.org/10.3905/jpm.2014.40.5.094
- López de Prado, M., & Lewis, M. J. (2019). Detection of false investment strategies using unsupervised learning methods. Quantitative Finance, 19(9), 1555–1565. https://doi.org/10.1080/14697688.2019.1622311
- Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2017). The probability of backtest overfitting. Journal of Computational Finance, 20(4), 39–69. https://doi.org/10.21314/JCF.2016.322
- Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606. https://doi.org/10.1073/pnas.1708274114