Production operations · Intermediate
Live versus backtest reconciliation
Every gap between a live book and its backtest has a cause, and a daily comparison of holdings, decisions and returns names it.
Educational material only. Not investment advice.
Contents
Abstract
We rerun the frozen release from the live account’s starting state and compare, session by session, what the broker holds with what the backtest holds, and the live decisions with what the release makes of the live inputs. The return gap is split among its causes by switching them on one at a time and by averaging over every order of switching them on. A band that widens with the square root of time catches excess noise, and a t-test on the daily difference catches drift the band cannot see.
Key takeaways
- Run the frozen release as a backtest from the live account’s starting state over the same sessions, and use broker statements as the live reference.
- Compare holdings with holdings at every close, keyed by permanent identifier, because plans show only intent.
- Rerun the release on the live inputs: if it reproduces the live decisions, data or state explain them; if not, the code drifted.
- Attribute the return gap by switching causes on one at a time, and average over every ordering of the causes when the split must not depend on one.
- Test the gap twice: a band for excess noise and a t-test for drift; measure costs fill by fill.
- Classify every close as matched, explained, unexplained or not compared, by the next day, and feed what is learned back into research.
Before you start
- Going live, for research parity and the frozen release
- Monitoring a live book, for broker snapshots and three-valued checks
- TWAP, slicing and jitter, for implementation shortfall
A live book always earns something different from its backtest, and the difference has causes that can be named: a feed, a late fill, a cost the model missed, a parameter that drifted. Reconciliation is the daily work of naming them, and of noticing the day one cannot be named.
Why live diverges
Six things make a live book diverge from its simulation. It decides on different data. Its timing differs: orders go out before the close, fill late, or wait for an auction. Its fills differ from the fill model, and its costs (fees, financing, borrow) from the cost model. An order can be missed altogether. And its code or configuration can differ from what was tested. The first four are expected and can be measured; a missed order is a defect the holdings comparison catches on the day; the last should never happen, and research parity in Going live exists to prevent it.
The published evidence is sobering. Across 888 algorithms on one research platform, Wiecki, Campbell, Lent and Stauth found that the backtest Sharpe ratio explained almost none of the variation in out-of-sample Sharpe ratios, and that the more a strategy had been backtested, the larger the shortfall.1 Suhonen, Lennkh and Perez found the same pattern in commercial alternative risk premia strategies after launch;2 Bailey, Borwein, López de Prado and Zhu explain why overfitting makes it likely.3 Reconciliation cannot fix an overfitted strategy, but it can separate that problem from the ones an engineer can fix.
Setting up the comparison
A comparison is only as good as its setup. The backtest must run the same release (the same fingerprint), start from the same holdings and cash as the live account on a declared go-live date, and cover exactly the sessions the live book traded. A live book compared with a backtest that began years earlier inherits that backtest’s luck. For the live side, the broker’s statements are the reference: positions, cash, fills and fees as the broker records them, not as the strategy believes them. Where a statement has not arrived, the strategy’s own records stand in, labelled as such, until the statement replaces them. Platforms help: QuantConnect runs an out-of-sample backtest alongside each live deployment and overlays the two equity curves, with a catalogue of the usual causes of deviation.4 Starting that backtest from the account’s actual holdings, as well as its date and equity, is what makes the curves comparable.
Holdings against holdings
The first comparison is holdings against holdings: what the broker says the account holds at the close, instrument by instrument, against what the backtest holds at the same close. Plans are the wrong thing to compare, because targets and orders say what the code intended, and the point of the exercise is to find where intention and reality part. Positions are keyed by the full identifier (a permanent id, not a ticker, as Data for systematic trading explains), and a position that exists on only one side is itself a finding.
Every close gets one of four results. Matched: the broker holds what the backtest holds and what the live records say. Explained: positions differ from the backtest, but the broker holds exactly what the live records say (targets less orders still working) and the tested release, rerun on the live book’s own inputs, makes the same decisions. Unexplained: the broker disagrees with the live records, or the live decisions are not the release’s. Not compared: the statement did not arrive, and the session stays open until it does.
def reconcile(bt: dict, live: dict, rerun: dict) -> list:
"""Holdings against holdings at every close, keyed by the instrument's full identifier.
`rerun` is the tested release run on the live book's own inputs. A close is:
matched -- the broker holds what the backtest holds, and what the live records say;
explained -- positions differ, but the broker holds exactly what the live records say
(targets less working orders) and the rerun reproduces the live decisions;
unexplained -- the broker disagrees with the live records, or the live decisions are not
the tested release's (code or configuration drift);
not compared -- the broker's statement never arrived. No day is skipped."""
out = []
for t in range(DAYS):
if t == NO_STATEMENT_DAY:
out.append({"day": t, "status": "not compared", "why": "no statement"})
continue
held = live["positions"][t]
records_ok = bool(np.all(held == live["targets"][t] - live["working"][t]))
decisions_ok = bool(np.all(live["targets"][t] == rerun["targets"][t]))
same = bool(np.all(held == bt["positions"][t]))
if not records_ok:
status, why = "unexplained", "the broker disagrees with the live records"
elif not decisions_ok:
status, why = "unexplained", "the tested release would not have made these decisions"
elif same:
status, why = "matched", ""
else:
status, why = "explained", "a decision on different data, or an order still working"
out.append({"day": t, "status": status, "why": why})
return out120 closes: 9 matched, 109 explained, 1 unexplained and 1 not compared. The live book decided differently from the backtest on 42 sessions. Click any session.
| Session 67: unexplained (the broker disagrees with the live records) | W | X | Y | Z |
|---|---|---|---|---|
| Backtest holds | 2,975 | 3,666 | 905 | 2,830 |
| Broker holds (live) | 2,975 | 3,664 | 2,836 | 2,831 |
| The release rerun on live inputs | 2,975 | 3,664 | 906 | 2,831 |
| Live targets, less working orders | 2,975 | 3,664 | 906 | 2,831 |
Explained: the broker holds exactly what the live book’s records say (targets less orders still working), and the tested release, rerun on the live inputs, makes the same decisions. Unexplained: the broker disagrees with the records, or the decisions are not the release’s. Session 91 has no statement.
In the synthetic book the feed difference alone changes 42 decisions and leaves 110 closes explained: positions differ from the backtest, and each difference is accounted for. The lost order produces 1 unexplained close, the only one: the book believes its order filled, and the broker disagrees. Because the book trades the difference between target and the broker’s positions every session, the next session’s trade makes up the shortfall; a book that trusted its own records would carry the error until someone noticed.
Decision by decision
The second comparison is decision by decision: on each session, did the live book change the same targets, by the same amounts, as the backtest? A different decision has a small set of possible sources: different inputs (the feed, a corporate action, a clock), different state (a late fill, a restart, a warm-up that read less history), nondeterminism, or different code. Rerunning the frozen release on the live book’s recorded inputs and state separates them. If the rerun reproduces the live decisions, inputs or state explain them; if it does not, the code or its configuration drifted, and that is a release problem.
In the synthetic book, a single parameter nudged from its tested value (the large weight from 30% to 31%) changes 34 decisions, and 106 closes become unexplained, because the rerun of the tested release disagrees with them. Its effect on returns, +0.20% of the backtest’s final equity over the period, is small enough that no return test would find it. Counting differing decisions is a better daily statistic than the P&L gap: it is exact, and it does not need months to become significant.
Attributing the gap
The third comparison is the return gap, and the aim is to divide it among its causes. Perold’s implementation shortfall measures a real portfolio against a paper one traded at decision prices;5 Kissell separates the shortfall of an order into delay, trading and opportunity components, with fees alongside.6 With a simulation in hand, the attribution can be done by replacement: start from the backtest, switch on one live difference at a time, and credit each with what it adds.
The parts add up to the gap exactly, but the split depends on the order in which causes are switched on, because causes interact: a late fill matters more on a day the data changed the decision. Averaging each cause’s contribution over every ordering, its Shapley value, removes the arbitrariness at the cost of more runs. Both books start with 1,000,000 in cash, and the backtest ends at 1,056,938; the shares below are of that final figure. With all six causes on, the live book ends 1.14% below the backtest. Slippage costs 0.94% on the average over orderings and the data difference 0.67%; late fills are credited with +0.27% in the fixed order but +0.07% on average, which is what a cause whose sign is luck looks like.
def attribute(on) -> dict:
"""The gap in final equity, split among the causes switched on. `parts` switches them on one
at a time in a fixed order (the parts add up to the gap exactly, but the split depends on the
order); `average` is each cause's contribution averaged over every order (its Shapley value),
which also adds up to the gap and depends on no order."""
on = [c for c in CAUSES if c in on]
final = {}
def value(subset):
key = frozenset(subset)
if key not in final:
final[key] = float(book(**live_kwargs(subset))["equity"][-1])
return final[key]
parts, prev = [], value(())
for k, c in enumerate(on):
cur = value(on[:k + 1])
parts.append((c, cur - prev))
prev = cur
avg = {c: 0.0 for c in on}
orders = list(permutations(on))
for order in orders:
for k, c in enumerate(order):
avg[c] += (value(order[:k + 1]) - value(order[:k])) / len(orders)
return {"backtest": value(()), "live": value(on), "gap": value(on) - value(()), "parts": parts,
"average": [(c, avg[c]) for c in on]}Tolerances: noise and drift
Tolerances differ by quantity. Holdings should match exactly, up to lot rounding: any other difference has a cause that can be named. Cash should match within the day’s fees and interest. Returns are noisy, and need two tests. The first asks whether the gap is larger than noise would make it: a band that widens with the square root of time,
with σ the daily tracking difference expected from all sources of noise, measured in a shadow period or the first live stage (paper fills are the model’s own and carry none of the fill noise), and z about 2. In the synthetic book the feed difference and the late fills together give 17.5 bp. The √t rests on the daily differences being uncorrelated; if they are not (late fills move P&L from one day to the next), estimate σ from the spread of multi-day sums, or with a Newey–West correction, rather than from daily ones.
Checked every day, even a correct band is crossed by pure noise often: in 100,000 simulated paths, a random walk leaves ±2σ√t at least once in 120 sessions 36% of the time. We flag only a gap that stays outside for five sessions in a row, which brings that to 17%; a strict sequential test would do better. The second test asks whether there is a steady drift: a t-statistic on the mean daily difference. Recomputed and checked every session, a threshold of |t| > 3 fires on pure noise 3.6% of the time over 120 sessions, about 11 times the rate of a single test at the end; the warm-up of twenty sessions and the high threshold are what keep it that low. The band is blind to drift for a long time. Slippage four times the model’s, alone, never leaves the band in 120 sessions; the drift test flags it on session 21, the first session its warm-up allows, at −0.77 bp a day. Mixed with the noise of late fills, the same slippage is never flagged by the drift test and only on session 57 by the noise test, which is why costs are measured fill by fill against the cost model, where each order is evidence. Frazzini, Israel and Moskowitz, with a large manager’s own executions, found real costs an order of magnitude smaller than models built on public data suggested:7 the gap can run either way, and only the fills say which.
- Step 1 of 4
The gap
Take a live book whose feed differs from the official close by 5 bp and whose fills cost 8 bp where the model assumed 2 bp. Both books start with 1,000,000 in cash; after 120 sessions the live book ends 18,388 below the backtest’s 1,056,938, −1.74%. The line is the cumulative return gap, session by session.
- Step 2 of 4
Split by cause
Switched on in order, the feed difference accounts for −0.80% of the backtest’s final equity and the slippage for −0.94%; averaged over both orderings, −0.81% and −0.93%. The feed’s part comes from 42 decisions made on slightly different numbers. The slippage’s is a cost paid on each fill.
- Step 3 of 4
The noise band
With σ at 17.5 bp, the band ±2σ√t is wide. The gap stays inside it on all 120 sessions, never using more than 45% of the room the band allows. The noise test sees nothing wrong.
- Step 4 of 4
The drift test
The t-statistic on the mean daily difference passes −3 on session 46: a steady −1.47 bp a day, too small for the band and too regular for noise. Alone, the same slippage is flagged on session 21; the feed’s noise delays the verdict by 25 sessions. The attribution names the cause, and comparing each fill with the cost model confirms it.
Synthetic: the book of the laboratory below with two causes switched on, the feed difference and fills at 8 bp against the model’s 2 bp; σ is the noise the book should expect. Computed by the TypeScript port of reconcile.py, checked against the Python in all combinations of causes. Without script, or with reduced motion, the figure shows its final state.
Both books start with 1,000,000 in cash. After 120 sessions the live book ends 10,110 below the backtest’s 1,056,938. The noise test flags it on session 57, the fifth in a row outside ±2σ√t: more than σ allows for, so something to explain. The drift test does not flag it (t = −0.6 at the end).
Fixed order: each cause is switched on in the order data, timing, slippage, missed fill, fees, config, and credited with what it adds. Average over orderings: each cause’s contribution averaged over every ordering of the causes (its Shapley value). Both add up to the gap exactly; only the second does not depend on an arbitrary order.
| Cause, alone | Holdings | Decisions that differ | Noise test | Drift test |
|---|---|---|---|---|
| data | 110 explained | 42 | never | never |
| timing | 22 explained | 0 | never | never |
| slippage | all matched | 0 | never | session 21 |
| missed fill | 1 unexplained | 0 | never | never |
| fees | all matched | 0 | never | session 21 |
| config | 106 unexplained | 34 | never | never |
Computed by reconcile.py, one cause at a time at the sizes in the lab. Holdings and decisions find their causes on the day they happen; the return tests find costs only when nothing noisier is going on.
def band(bt: dict, live: dict, sigma_bp: float, z: float = 2.0, persist: int = 5) -> dict:
"""The cumulative return gap against +/- z sigma sqrt(t), flagged only when it stays outside
for `persist` sessions in a row: a noise test. It is blind to slow drift."""
gap = live["equity"] / bt["equity"] - 1.0
t = np.arange(1, DAYS + 1)
limit = z * sigma_bp / 1e4 * np.sqrt(t)
out = np.abs(gap) > limit
run, first = 0, None
for i, o in enumerate(out):
run = run + 1 if o else 0
if run >= persist and first is None:
first = i
return {"gap": gap, "limit": limit, "flagged_on": first, "days_outside": int(out.sum())}
def drift_test(bt: dict, live: dict, threshold: float = 3.0, warmup: int = 20) -> dict:
"""A t-statistic on the mean daily difference, computed each session on everything so far:
a drift test. It flags a small, steady cost the band cannot see."""
d = daily_differences(bt, live)
first, tstat = None, []
for n in range(2, len(d) + 1):
x = d[:n]
s = np.std(x, ddof=1)
tv = float(np.mean(x) / (s / np.sqrt(n))) if s > 0 else 0.0
tstat.append(tv)
if first is None and n >= warmup and abs(tv) > threshold:
first = n
return {"t": tstat, "flagged_on": first, "t_final": tstat[-1], "mean_bp": float(np.mean(d) * 1e4)}A daily routine
Each morning the previous session is compared: holdings, decisions, the return gap and both tests. We promise a result for every session by the next day, and a backlog of sessions not yet compared is itself an alert, routed as Alerting that people act on describes, next to the evening checks of Monitoring a live book.
| Finding | What we do |
|---|---|
| Holdings differ, records and rerun agree | Explained: name the decision or timing difference and log it |
| The broker disagrees with the live records | Unexplained: stop trading the instrument until the broker, the journal and the fills agree |
| The rerun on live inputs does not reproduce a decision | Code or configuration drift: a release problem; roll back if it cannot be fixed before the next session |
| Decisions differ, and the rerun reproduces them | Data or state: find which input differed, and whether the backtest’s source or the live feed is right |
| The drift test flags, holdings and decisions agree | Costs or fills: compare every fill with the cost model; recalibrate the model or the execution |
| The noise test flags, and nothing else does | Re-measure σ; if it holds, look for a cause the checks do not cover |
| No statement for a session | Not compared: chase the statement |
References
- Wiecki, T., Campbell, A., Lent, J., & Stauth, J. (2016). All that glitters is not gold: Comparing backtest and out-of-sample performance on a large cohort of trading algorithms. Journal of Investing, 25(3), 69–80. https://doi.org/10.3905/joi.2016.25.3.069
- Suhonen, A., Lennkh, M., & Perez, F. (2017). Quantifying backtest overfitting in alternative beta strategies. Journal of Portfolio Management, 43(2), 90–104. https://doi.org/10.3905/jpm.2017.43.2.090
- Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2014). Pseudo-mathematics and financial charlatanism: The effects of backtest overfitting on out-of-sample performance. Notices of the American Mathematical Society, 61(5), 458–471. https://doi.org/10.1090/noti1105
- QuantConnect (n.d.). Reconciliation. QuantConnect documentation. www.quantconnect.com/docs/v2/cloud-platform/live-trading/reconciliation
- Perold, A. F. (1988). The implementation shortfall: Paper versus reality. Journal of Portfolio Management, 14(3), 4–9. https://doi.org/10.3905/jpm.1988.409150
- Kissell, R. (2006). The expanded implementation shortfall: Understanding transaction cost components. Journal of Trading, 1(3), 6–16. https://doi.org/10.3905/jot.2006.644083
- Frazzini, A., Israel, R., & Moskowitz, T. J. (2018). Trading costs. SSRN working paper. https://doi.org/10.2139/ssrn.3229719