Skip to content

Production operations · Intermediate

Live versus backtest reconciliation

Every gap between a live book and its backtest has a cause, and a daily comparison of holdings, decisions and returns names it.

18 min read7 referencesIntroduction to the series

Educational material only. Not investment advice.

Contents
  1. 01Why live diverges
  2. 02Setting up the comparison
  3. 03Holdings against holdings
  4. 04Decision by decision
  5. 05Attributing the gap
  6. 06Tolerances: noise and drift
  7. 07A daily routine
  8. —References

Abstract

We rerun the frozen release from the live account’s starting state and compare, session by session, what the broker holds with what the backtest holds, and the live decisions with what the release makes of the live inputs. The return gap is split among its causes by switching them on one at a time and by averaging over every order of switching them on. A band that widens with the square root of time catches excess noise, and a t-test on the daily difference catches drift the band cannot see.

Key takeaways

  • Run the frozen release as a backtest from the live account’s starting state over the same sessions, and use broker statements as the live reference.
  • Compare holdings with holdings at every close, keyed by permanent identifier, because plans show only intent.
  • Rerun the release on the live inputs: if it reproduces the live decisions, data or state explain them; if not, the code drifted.
  • Attribute the return gap by switching causes on one at a time, and average over every ordering of the causes when the split must not depend on one.
  • Test the gap twice: a band for excess noise and a t-test for drift; measure costs fill by fill.
  • Classify every close as matched, explained, unexplained or not compared, by the next day, and feed what is learned back into research.

Before you start

  • Going live, for research parity and the frozen release
  • Monitoring a live book, for broker snapshots and three-valued checks
  • TWAP, slicing and jitter, for implementation shortfall

A live book always earns something different from its backtest, and the difference has causes that can be named: a feed, a late fill, a cost the model missed, a parameter that drifted. Reconciliation is the daily work of naming them, and of noticing the day one cannot be named.

Why live diverges

Six things make a live book diverge from its simulation. It decides on different data. Its timing differs: orders go out before the close, fill late, or wait for an auction. Its fills differ from the fill model, and its costs (fees, financing, borrow) from the cost model. An order can be missed altogether. And its code or configuration can differ from what was tested. The first four are expected and can be measured; a missed order is a defect the holdings comparison catches on the day; the last should never happen, and research parity in Going live exists to prevent it.

The published evidence is sobering. Across 888 algorithms on one research platform, Wiecki, Campbell, Lent and Stauth found that the backtest Sharpe ratio explained almost none of the variation in out-of-sample Sharpe ratios, and that the more a strategy had been backtested, the larger the shortfall.1 Suhonen, Lennkh and Perez found the same pattern in commercial alternative risk premia strategies after launch;2 Bailey, Borwein, López de Prado and Zhu explain why overfitting makes it likely.3 Reconciliation cannot fix an overfitted strategy, but it can separate that problem from the ones an engineer can fix.

Figure 1Two books from one release
Reconciliation: the frozen release runs live, as a backtest from the same start, and as a rerun on live inputs; holdings, decisions and returns are compared daily; the gap is attributed and tested; residuals are investigatedFrozen releaseone fingerprintLive bookbroker statementsRerun on live inputssame code, live dataBacktestsame start and sessionsCompare dailyholdings, decisions, returnsAttributeby causeTestnoise band, driftInvestigatewhat is left
The frozen release is run twice over the same sessions: live, against the market and the broker, and as a backtest from the same starting state on the official data. A third run, the release on the live book’s own inputs, separates data from code. Holdings, decisions and returns are compared every day; the gap is attributed and tested, and whatever is left is investigated.

Setting up the comparison

A comparison is only as good as its setup. The backtest must run the same release (the same fingerprint), start from the same holdings and cash as the live account on a declared go-live date, and cover exactly the sessions the live book traded. A live book compared with a backtest that began years earlier inherits that backtest’s luck. For the live side, the broker’s statements are the reference: positions, cash, fills and fees as the broker records them, not as the strategy believes them. Where a statement has not arrived, the strategy’s own records stand in, labelled as such, until the statement replaces them. Platforms help: QuantConnect runs an out-of-sample backtest alongside each live deployment and overlays the two equity curves, with a catalogue of the usual causes of deviation.4 Starting that backtest from the account’s actual holdings, as well as its date and equity, is what makes the curves comparable.

Holdings against holdings

The first comparison is holdings against holdings: what the broker says the account holds at the close, instrument by instrument, against what the backtest holds at the same close. Plans are the wrong thing to compare, because targets and orders say what the code intended, and the point of the exercise is to find where intention and reality part. Positions are keyed by the full identifier (a permanent id, not a ticker, as Data for systematic trading explains), and a position that exists on only one side is itself a finding.

Every close gets one of four results. Matched: the broker holds what the backtest holds and what the live records say. Explained: positions differ from the backtest, but the broker holds exactly what the live records say (targets less orders still working) and the tested release, rerun on the live book’s own inputs, makes the same decisions. Unexplained: the broker disagrees with the live records, or the live decisions are not the release’s. Not compared: the statement did not arrive, and the session stays open until it does.

site_research/fieldnotes/reconcile.py
def reconcile(bt: dict, live: dict, rerun: dict) -> list:
    """Holdings against holdings at every close, keyed by the instrument's full identifier.

    `rerun` is the tested release run on the live book's own inputs. A close is:
      matched     -- the broker holds what the backtest holds, and what the live records say;
      explained   -- positions differ, but the broker holds exactly what the live records say
                     (targets less working orders) and the rerun reproduces the live decisions;
      unexplained -- the broker disagrees with the live records, or the live decisions are not
                     the tested release's (code or configuration drift);
      not compared -- the broker's statement never arrived. No day is skipped."""
    out = []
    for t in range(DAYS):
        if t == NO_STATEMENT_DAY:
            out.append({"day": t, "status": "not compared", "why": "no statement"})
            continue
        held = live["positions"][t]
        records_ok = bool(np.all(held == live["targets"][t] - live["working"][t]))
        decisions_ok = bool(np.all(live["targets"][t] == rerun["targets"][t]))
        same = bool(np.all(held == bt["positions"][t]))
        if not records_ok:
            status, why = "unexplained", "the broker disagrees with the live records"
        elif not decisions_ok:
            status, why = "unexplained", "the tested release would not have made these decisions"
        elif same:
            status, why = "matched", ""
        else:
            status, why = "explained", "a decision on different data, or an order still working"
        out.append({"day": t, "status": status, "why": why})
    return out
Figure 2Every close, holdings against holdings
Switch on a cause of divergence (any combination)

120 closes: 9 matched, 109 explained, 1 unexplained and 1 not compared. The live book decided differently from the backtest on 42 sessions. Click any session.

Every close, holdings against holdingsSession 1: matchedSession 2: matchedSession 3: matchedSession 4: matchedSession 5: matchedSession 6: matchedSession 7: matchedSession 8: matchedSession 9: matchedSession 10: explainedSession 11: explainedSession 12: explainedSession 13: explainedSession 14: explainedSession 15: explainedSession 16: explainedSession 17: explainedSession 18: explainedSession 19: explainedSession 20: explainedSession 21: explainedSession 22: explainedSession 23: explainedSession 24: explainedSession 25: explainedSession 26: explainedSession 27: explainedSession 28: explainedSession 29: explainedSession 30: explainedSession 31: explainedSession 32: explainedSession 33: explainedSession 34: explainedSession 35: explainedSession 36: explainedSession 37: explainedSession 38: explainedSession 39: explainedSession 40: explainedSession 41: explainedSession 42: explainedSession 43: explainedSession 44: explainedSession 45: explainedSession 46: explainedSession 47: explainedSession 48: explainedSession 49: explainedSession 50: explainedSession 51: explainedSession 52: explainedSession 53: explainedSession 54: explainedSession 55: explainedSession 56: explainedSession 57: explainedSession 58: explainedSession 59: explainedSession 60: explainedSession 61: explainedSession 62: explainedSession 63: explainedSession 64: explainedSession 65: explainedSession 66: explainedSession 67: unexplainedSession 68: explainedSession 69: explainedSession 70: explainedSession 71: explainedSession 72: explainedSession 73: explainedSession 74: explainedSession 75: explainedSession 76: explainedSession 77: explainedSession 78: explainedSession 79: explainedSession 80: explainedSession 81: explainedSession 82: explainedSession 83: explainedSession 84: explainedSession 85: explainedSession 86: explainedSession 87: explainedSession 88: explainedSession 89: explainedSession 90: explainedSession 91: not comparedSession 92: explainedSession 93: explainedSession 94: explainedSession 95: explainedSession 96: explainedSession 97: explainedSession 98: explainedSession 99: explainedSession 100: explainedSession 101: explainedSession 102: explainedSession 103: explainedSession 104: explainedSession 105: explainedSession 106: explainedSession 107: explainedSession 108: explainedSession 109: explainedSession 110: explainedSession 111: explainedSession 112: explainedSession 113: explainedSession 114: explainedSession 115: explainedSession 116: explainedSession 117: explainedSession 118: explainedSession 119: explainedSession 120: explained
■Matched■Explained■Unexplained■Not compared
Session 67: unexplained (the broker disagrees with the live records)WXYZ
Backtest holds2,9753,6669052,830
Broker holds (live)2,9753,6642,8362,831
The release rerun on live inputs2,9753,6649062,831
Live targets, less working orders2,9753,6649062,831

Explained: the broker holds exactly what the live book’s records say (targets less orders still working), and the tested release, rerun on the live inputs, makes the same decisions. Unexplained: the broker disagrees with the records, or the decisions are not the release’s. Session 91 has no statement.

Synthetic: four instruments over 120 sessions, a signal-flip strategy on a fixed notional, its backtest, a live twin and the release rerun on the live inputs. The lab opens with a feed that differs slightly from the official close and one lost order. Open the red session; switch the feed difference off and the grey sessions go; switch on the drifted parameter and the grid turns red.

In the synthetic book the feed difference alone changes 42 decisions and leaves 110 closes explained: positions differ from the backtest, and each difference is accounted for. The lost order produces 1 unexplained close, the only one: the book believes its order filled, and the broker disagrees. Because the book trades the difference between target and the broker’s positions every session, the next session’s trade makes up the shortfall; a book that trusted its own records would carry the error until someone noticed.

Decision by decision

The second comparison is decision by decision: on each session, did the live book change the same targets, by the same amounts, as the backtest? A different decision has a small set of possible sources: different inputs (the feed, a corporate action, a clock), different state (a late fill, a restart, a warm-up that read less history), nondeterminism, or different code. Rerunning the frozen release on the live book’s recorded inputs and state separates them. If the rerun reproduces the live decisions, inputs or state explain them; if it does not, the code or its configuration drifted, and that is a release problem.

In the synthetic book, a single parameter nudged from its tested value (the large weight from 30% to 31%) changes 34 decisions, and 106 closes become unexplained, because the rerun of the tested release disagrees with them. Its effect on returns, +0.20% of the backtest’s final equity over the period, is small enough that no return test would find it. Counting differing decisions is a better daily statistic than the P&L gap: it is exact, and it does not need months to become significant.

Attributing the gap

The third comparison is the return gap, and the aim is to divide it among its causes. Perold’s implementation shortfall measures a real portfolio against a paper one traded at decision prices;5 Kissell separates the shortfall of an order into delay, trading and opportunity components, with fees alongside.6 With a simulation in hand, the attribution can be done by replacement: start from the backtest, switch on one live difference at a time, and credit each with what it adds.

Elive − Ebacktest  =  Σk ( Ek − Ek − 1 )
the gap as a telescoping sum over causes switched on in order

The parts add up to the gap exactly, but the split depends on the order in which causes are switched on, because causes interact: a late fill matters more on a day the data changed the decision. Averaging each cause’s contribution over every ordering, its Shapley value, removes the arbitrariness at the cost of more runs. Both books start with 1,000,000 in cash, and the backtest ends at 1,056,938; the shares below are of that final figure. With all six causes on, the live book ends 1.14% below the backtest. Slippage costs 0.94% on the average over orderings and the data difference 0.67%; late fills are credited with +0.27% in the fixed order but +0.07% on average, which is what a cause whose sign is luck looks like.

site_research/fieldnotes/reconcile.py
def attribute(on) -> dict:
    """The gap in final equity, split among the causes switched on. `parts` switches them on one
    at a time in a fixed order (the parts add up to the gap exactly, but the split depends on the
    order); `average` is each cause's contribution averaged over every order (its Shapley value),
    which also adds up to the gap and depends on no order."""
    on = [c for c in CAUSES if c in on]
    final = {}

    def value(subset):
        key = frozenset(subset)
        if key not in final:
            final[key] = float(book(**live_kwargs(subset))["equity"][-1])
        return final[key]

    parts, prev = [], value(())
    for k, c in enumerate(on):
        cur = value(on[:k + 1])
        parts.append((c, cur - prev))
        prev = cur
    avg = {c: 0.0 for c in on}
    orders = list(permutations(on))
    for order in orders:
        for k, c in enumerate(order):
            avg[c] += (value(order[:k + 1]) - value(order[:k])) / len(orders)
    return {"backtest": value(()), "live": value(on), "gap": value(on) - value(()), "parts": parts,
            "average": [(c, avg[c]) for c in on]}

Tolerances: noise and drift

Tolerances differ by quantity. Holdings should match exactly, up to lot rounding: any other difference has a cause that can be named. Cash should match within the day’s fees and interest. Returns are noisy, and need two tests. The first asks whether the gap is larger than noise would make it: a band that widens with the square root of time,

| gapt |  ≤  z · σ · √t
the noise band for the cumulative return gap after t sessions

with σ the daily tracking difference expected from all sources of noise, measured in a shadow period or the first live stage (paper fills are the model’s own and carry none of the fill noise), and z about 2. In the synthetic book the feed difference and the late fills together give 17.5 bp. The √t rests on the daily differences being uncorrelated; if they are not (late fills move P&L from one day to the next), estimate σ from the spread of multi-day sums, or with a Newey–West correction, rather than from daily ones.

Checked every day, even a correct band is crossed by pure noise often: in 100,000 simulated paths, a random walk leaves ±2σ√t at least once in 120 sessions 36% of the time. We flag only a gap that stays outside for five sessions in a row, which brings that to 17%; a strict sequential test would do better. The second test asks whether there is a steady drift: a t-statistic on the mean daily difference. Recomputed and checked every session, a threshold of |t| > 3 fires on pure noise 3.6% of the time over 120 sessions, about 11 times the rate of a single test at the end; the warm-up of twenty sessions and the high threshold are what keep it that low. The band is blind to drift for a long time. Slippage four times the model’s, alone, never leaves the band in 120 sessions; the drift test flags it on session 21, the first session its warm-up allows, at −0.77 bp a day. Mixed with the noise of late fills, the same slippage is never flagged by the drift test and only on session 57 by the noise test, which is why costs are measured fill by fill against the cost model, where each order is evidence. Frazzini, Israel and Moskowitz, with a large manager’s own executions, found real costs an order of magnitude smaller than models built on public data suggested:7 the gap can run either way, and only the fills say which.

Step by stepOne gap, split by cause and tested twice
The gap in final equity, as a share of the backtest’s, by causeFeed differs by 5 bp−0.80%Fills cost 8 bp, not 2−0.94%Total gap−1.74%Cumulative return gap, live minus backtest−4%−2%0%2%4%±2σ√t, σ = 17.5 bp0 sessions outside−1.7%Drift test: t-statistic on the mean daily difference−303flagged, session 46session 1306090120
  1. Step 1 of 4

    The gap

    Take a live book whose feed differs from the official close by 5 bp and whose fills cost 8 bp where the model assumed 2 bp. Both books start with 1,000,000 in cash; after 120 sessions the live book ends 18,388 below the backtest’s 1,056,938, −1.74%. The line is the cumulative return gap, session by session.

  2. Step 2 of 4

    Split by cause

    Switched on in order, the feed difference accounts for −0.80% of the backtest’s final equity and the slippage for −0.94%; averaged over both orderings, −0.81% and −0.93%. The feed’s part comes from 42 decisions made on slightly different numbers. The slippage’s is a cost paid on each fill.

  3. Step 3 of 4

    The noise band

    With σ at 17.5 bp, the band ±2σ√t is wide. The gap stays inside it on all 120 sessions, never using more than 45% of the room the band allows. The noise test sees nothing wrong.

  4. Step 4 of 4

    The drift test

    The t-statistic on the mean daily difference passes −3 on session 46: a steady −1.47 bp a day, too small for the band and too regular for noise. Alone, the same slippage is flagged on session 21; the feed’s noise delays the verdict by 25 sessions. The attribution names the cause, and comparing each fill with the cost model confirms it.

Synthetic: the book of the laboratory below with two causes switched on, the feed difference and fills at 8 bp against the model’s 2 bp; σ is the noise the book should expect. Computed by the TypeScript port of reconcile.py, checked against the Python in all combinations of causes. Without script, or with reduced motion, the figure shows its final state.

Figure 3Attributing the gap, and testing it
Switch on a cause of divergence (any combination)
Attribution

Both books start with 1,000,000 in cash. After 120 sessions the live book ends 10,110 below the backtest’s 1,056,938. The noise test flags it on session 57, the fifth in a row outside ±2σ√t: more than σ allows for, so something to explain. The drift test does not flag it (t = −0.6 at the end).

Gap in final equity
−10,110
−0.96% of the backtest’s
Largest single cause
slippage
from the attribution below
Noise test
session 57
9 sessions outside ±2σ√t
Drift test
quiet
t on the mean daily difference: −0.6
Cumulative return gap, live against backtest, with the tolerance band ±2σ√t−4.0%−2.0%0.0%2.0%4.0%noise test flagssession 1session 20session 40session 60session 80session 100session 120
Where the gap comes from, in final equity (starting cash 1,000,000)timing−396slippage−9,714total gap−10,110

Fixed order: each cause is switched on in the order data, timing, slippage, missed fill, fees, config, and credited with what it adds. Average over orderings: each cause’s contribution averaged over every ordering of the causes (its Shapley value). Both add up to the gap exactly; only the second does not depend on an arbitrary order.

Synthetic: the same book. Switch causes on and off; the lab opens with late fills and costly slippage, and σ at the noise the book should expect. Switch timing off to leave slippage alone, then switch everything on and compare the fixed-order attribution with the average over orderings.
Cause, aloneHoldingsDecisions that differNoise testDrift test
data110 explained42nevernever
timing22 explained0nevernever
slippageall matched0neversession 21
missed fill1 unexplained0nevernever
feesall matched0neversession 21
config106 unexplained34nevernever

Computed by reconcile.py, one cause at a time at the sizes in the lab. Holdings and decisions find their causes on the day they happen; the return tests find costs only when nothing noisier is going on.

site_research/fieldnotes/reconcile.py
def band(bt: dict, live: dict, sigma_bp: float, z: float = 2.0, persist: int = 5) -> dict:
    """The cumulative return gap against +/- z sigma sqrt(t), flagged only when it stays outside
    for `persist` sessions in a row: a noise test. It is blind to slow drift."""
    gap = live["equity"] / bt["equity"] - 1.0
    t = np.arange(1, DAYS + 1)
    limit = z * sigma_bp / 1e4 * np.sqrt(t)
    out = np.abs(gap) > limit
    run, first = 0, None
    for i, o in enumerate(out):
        run = run + 1 if o else 0
        if run >= persist and first is None:
            first = i
    return {"gap": gap, "limit": limit, "flagged_on": first, "days_outside": int(out.sum())}


def drift_test(bt: dict, live: dict, threshold: float = 3.0, warmup: int = 20) -> dict:
    """A t-statistic on the mean daily difference, computed each session on everything so far:
    a drift test. It flags a small, steady cost the band cannot see."""
    d = daily_differences(bt, live)
    first, tstat = None, []
    for n in range(2, len(d) + 1):
        x = d[:n]
        s = np.std(x, ddof=1)
        tv = float(np.mean(x) / (s / np.sqrt(n))) if s > 0 else 0.0
        tstat.append(tv)
        if first is None and n >= warmup and abs(tv) > threshold:
            first = n
    return {"t": tstat, "flagged_on": first, "t_final": tstat[-1], "mean_bp": float(np.mean(d) * 1e4)}

A daily routine

Each morning the previous session is compared: holdings, decisions, the return gap and both tests. We promise a result for every session by the next day, and a backlog of sessions not yet compared is itself an alert, routed as Alerting that people act on describes, next to the evening checks of Monitoring a live book.

FindingWhat we do
Holdings differ, records and rerun agreeExplained: name the decision or timing difference and log it
The broker disagrees with the live recordsUnexplained: stop trading the instrument until the broker, the journal and the fills agree
The rerun on live inputs does not reproduce a decisionCode or configuration drift: a release problem; roll back if it cannot be fixed before the next session
Decisions differ, and the rerun reproduces themData or state: find which input differed, and whether the backtest’s source or the live feed is right
The drift test flags, holdings and decisions agreeCosts or fills: compare every fill with the cost model; recalibrate the model or the execution
The noise test flags, and nothing else doesRe-measure σ; if it holds, look for a cause the checks do not cover
No statement for a sessionNot compared: chase the statement

References

  1. Wiecki, T., Campbell, A., Lent, J., & Stauth, J. (2016). All that glitters is not gold: Comparing backtest and out-of-sample performance on a large cohort of trading algorithms. Journal of Investing, 25(3), 69–80. https://doi.org/10.3905/joi.2016.25.3.069
  2. Suhonen, A., Lennkh, M., & Perez, F. (2017). Quantifying backtest overfitting in alternative beta strategies. Journal of Portfolio Management, 43(2), 90–104. https://doi.org/10.3905/jpm.2017.43.2.090
  3. Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2014). Pseudo-mathematics and financial charlatanism: The effects of backtest overfitting on out-of-sample performance. Notices of the American Mathematical Society, 61(5), 458–471. https://doi.org/10.1090/noti1105
  4. QuantConnect (n.d.). Reconciliation. QuantConnect documentation. www.quantconnect.com/docs/v2/cloud-platform/live-trading/reconciliation
  5. Perold, A. F. (1988). The implementation shortfall: Paper versus reality. Journal of Portfolio Management, 14(3), 4–9. https://doi.org/10.3905/jpm.1988.409150
  6. Kissell, R. (2006). The expanded implementation shortfall: Understanding transaction cost components. Journal of Trading, 1(3), 6–16. https://doi.org/10.3905/jot.2006.644083
  7. Frazzini, A., Israel, R., & Moskowitz, T. J. (2018). Trading costs. SSRN working paper. https://doi.org/10.2139/ssrn.3229719