Skip to content

Production operations · Intermediate

Going live

Going live tests the machinery around a strategy, and most of what can go wrong there can be measured before money moves.

19 min read10 referencesIntroduction to the series

Educational material only. Not investment advice.

Contents
  1. 01Two paths
  2. 02Paper trading and its limits
  3. 03Freezing a release
  4. 04Where live and backtest part ways
  5. 05A staged start
  6. 06Restart and recovery
  7. 07Kill switches, deployment and rollback
  8. —References

Abstract

We freeze each release under one fingerprint and rerun it over the recent past until its decisions match research exactly. Some live differences are designed in, and the gap between a backtest and its live twin splits into decisions, fills and costs. A power calculation shows that early live P&L cannot detect a realistic shortfall for months, so capital stages are gated on mechanics that can be checked each session. A journal of orders, state rebuilt from the broker, a kill switch that stays tripped and a one-step rollback let the system fail without doubling anything.

Key takeaways

  • Paper trading exercises the plumbing; its fills are the backtest’s own assumptions, so it says little about performance.
  • Freeze code, parameters, engine, library versions and data sources under one fingerprint, and rerun the release over the recent past until its decisions match research exactly.
  • Backtest the decision time you can actually meet live. What remains of the decision gap is noise; costs are a certain loss.
  • Early live P&L cannot detect a realistic shortfall for months, so capital stages are gated on mechanics checked every session.
  • Read working orders, then positions, from the broker at each start; journal an order before sending it, and look it up, once the broker has reported, before resending; size any resend from what the broker holds, never from the journal alone.
  • Keep a kill switch that trips on the first breach and needs a person to reset it, and a rollback one redeploy away, which holds only if the new release wrote no state the old one cannot read.

Before you start

  • QuantConnect and LEAN, for one engine in backtest and live
  • Running backtests at scale, for fingerprinted runs
  • TWAP, slicing and jitter, for costs and fills
  • Data for systematic trading, for timestamps and the close
  • Hypothesis tests and statistical power

A strategy that works in a backtest has passed a test of an idea. Going live tests the machinery around it, and most of what goes wrong there is engineering. The habit that pays most is to ask of each difference between live and backtest whether it was designed in or discovered.

Two paths

A live system has two paths. The release path takes a tested strategy to production and back out again. The live path runs every session: data in, decisions, orders out, fills back, and the account’s true state read from the broker. The two meet in one place, the build that is running, and that build must be the one that was tested.

Figure 1The release path and the live path
Release path: build with fingerprint, parity rerun, paper trading, staged capital, with rollback to the previous build. Live path: market data, strategy, order manager with pre-trade checks and kill switch, broker, exchange; the broker’s positions and working orders rebuild the state the strategy reads, and feed monitoring.Release pathLive path, every sessionon failureFrozen buildone fingerprintResearch paritysame decisions?PaperplumbingStaged capitalgates on mechanicsRollbackprevious fingerprintMarket datafeed, calendarStrategysignal → targetsOrder managerchecks, journal, killBrokerorders, fills, positionsExchangeStaterebuilt from brokerMonitoringchecks, alerts
A generic design, simplified. Top lane: a build is fingerprinted, rerun over the recent past, paper traded, then given capital in stages; rollback redeploys the previous fingerprint. Bottom lane: each session, the running build reads data, decides, passes its orders through pre-trade checks and the kill switch, and rebuilds its state from the broker’s records. Monitoring can also trip the kill switch (not drawn).

Paper trading and its limits

Paper trading runs the strategy on live data with simulated fills, and it comes in two kinds. A platform’s simulator fills orders inside the platform: QuantConnect’s paper brokerage, for example, part of the engine described in QuantConnect and LEAN, fills market orders at once and in full, at the bid or ask when quotes are available, with no slippage, and models fees, buying power and margin, but it never talks to a broker.1 A broker’s own paper account goes through the broker’s interface, order types and rejections, with fills that are still simulated. Use the first to test the schedule, the handling of holidays and half-days, and a restart in the middle of a session; use the second to test the connection, the order formats and the broker’s rules. Exercise each of them before capital is at risk.

Neither tells you much about performance. A paper fill has no queue position and no market impact, and in a platform simulator it is the backtest’s own fill model. A paper account that matches the backtest has confirmed the backtest’s assumptions, not the market’s. The first real fills, measured against the cost model order by order, are the earliest evidence about performance there is.

Freezing a release

Before a strategy trades, we freeze it: the code, the parameters, the engine version, the versions of the libraries it imports and the data sources become one release with one identity. Humble and Farley make the general case for building every release once, from version control, and promoting the same artefact through each stage;2 Google’s release engineers add that builds must be hermetic, so the same inputs always give the same result.3 A content hash over everything that decides behaviour makes that identity checkable. It is the fingerprint that Running backtests at scale uses to name a research run, with the data named by its source rather than digested, since a live feed has no file to hash.

site_research/fieldnotes/golive.py
def fingerprint(sources: dict, params: dict, engine: str, data: dict) -> str:
    """One hash for everything that decides behaviour: source files, parameters, the engine's
    version and the data sources (vendor, dataset, normalisation). Put the dependency lockfile
    among the sources, so a library upgrade changes the fingerprint too. Two runs with the same
    fingerprint must make the same decisions on the same data."""
    h = hashlib.sha256()
    for name in sorted(sources):
        h.update(name.encode() + b"\0" + sources[name].encode() + b"\0")
    for part in (params, {"engine": engine}, data):
        h.update(json.dumps(part, sort_keys=True).encode() + b"\0")
    return h.hexdigest()[:16]
The release-level fingerprint. It hashes the same things as the research fingerprint, but takes the data as a description of its source (vendor, dataset, normalisation) where the research version takes a digest of the files, and it keeps the source files by name.

Research parity

The frozen release is then rerun as a backtest over the most recent weeks, on the data the live system would have seen, and its decisions are compared with those of the research version, session by session and order by order. The accepted number of differences is zero, because nothing about the market can explain one.

Cause of a differenceHow it shows up
A changed default in the engine or a libraryEvery session differs a little from one date on
A data vintage adjusted after the factDifferences cluster around dividends, splits or rolls
A parameter read from the wrong fileOne rule or one instrument behaves differently throughout
Randomness or thread order (seeds, parallel sums)Two reruns of the same release disagree with each other

Each of these would also make the live system disagree with its own backtest, and they are far cheaper to find before the money moves. Once the strategy is live, the same rerun, over the session just traded, becomes the simulation that Live versus backtest reconciliation compares with the account.

Where live and backtest part ways

Even a perfectly frozen strategy trades differently live than in its backtest, for reasons that are designed in and should be known in advance.

DifferenceBacktestLiveWhat to do
Decision timeDecides on the day’s closeDecides minutes before it, on the last tradeBacktest the decision time you can actually meet
Data resolutionA daily barA stream of trades and quotes, or minute barsBuild the live signal from the same bars the backtest used
The close itselfThe official closing priceYour feed’s last print, which can differKnow which close each source reports
FillsAt the close, in fullQueue position, partial fills, the auction’s cut-offModel the order type you will send; measure live fills against it
CostsA modelled spread and feeSpread, impact, fees, borrow, financingMeasure each fill against the cost model
StateStarts clean, never crashesRestarts mid-session, inherits positionsRebuild state from the broker at each start
Calendar and clockAssumed sessionsHolidays, half-days, daylight savingSchedule in exchange time from an exchange calendar

Anything else found live was discovered, not designed, and is a defect until proved otherwise. Which close a source reports, and how it is timestamped, is covered in Data for systematic trading.

A market-on-close order has to reach the exchange before the auction’s cut-off, so the decision must come earlier still: LEAN, for instance, by default refuses a market-on-close order sent less than 15.5 minutes before the close (the buffer is configurable).4 A simulator makes the designed-in part of the gap concrete. A 50-day trend rule is backtested on closes and run live on a feed, and the gap between the two is built up one difference at a time.

Step by stepA backtest and its live twin, one difference at a time
Growth of 1, 1,000 synthetic sessions1.01.21.40250500750100014 decisions differ (red marks below)The gap in annual return, percentage points−1.0−0.50+0.5+1.0Decisions−0.56Fills0.00Costs−0.63Whole gap−1.20±1.21 across 40 paths
  1. Step 1 of 5

    The backtest

    The backtest decides on each session’s official close and trades at it, paying the modelled 2 bp per unit traded. Over 1,000 sessions, about 4.0 years, the rule switches 78 times and returns 7.6% a year.

  2. Step 2 of 5

    Decisions before the close

    The live twin decides 20 minutes before the close, in time for a market-on-close order, on a feed whose price differs from the official one by 2 bp. The two versions decide differently on 14 sessions of the 951 on which both have a signal, marked in red. Those sessions change the live book’s annual return by −0.56 percentage points.

  3. Step 3 of 5

    Fills

    A market-on-close order fills at the official close, the price the backtest assumed, so the fills change the return by 0.00 percentage points. Filling at the price seen at the decision instead would have changed it by −0.49 percentage points on this path; the lab below has the switch.

  4. Step 4 of 5

    Costs

    Live trading pays 5 bp per unit traded against the backtest’s 2. That takes −0.63 percentage points a year, and the whole gap is −1.20 percentage points: the live book returns 6.4% a year against the backtest’s 7.6%.

  5. Step 5 of 5

    Luck and certainty

    The sign of the decision gap is luck. Over 40 synthetic paths it averaged 0.00 percentage points a year with a standard deviation of 1.21 percentage points, and filling at the decision price averaged −0.07 percentage points with a standard deviation of 0.49. Costs are the opposite: each unit traded live pays the extra 3 bp, so that step is small, certain and against you on any path.

Synthetic: one trend rule on a random walk of 1,000 sessions, the simulator of the laboratory below at its default settings; the spread across paths is the same simulator on 40 other random walks. Without script, or with reduced motion, the figure shows its final state.

The lab runs the same simulator with the settings in your hands: the decision time, the gap between the feed and the official close, the fill and the cost.

Figure 2The gap between a backtest and its live twin
Live fill

The live process decides differently from the backtest on 14 of 951 sessions, all of them days when the price was close to its 50-day average. Backtest 7.6% a year, live 6.4%: decisions −0.56, fills 0.00 and costs −0.63 percentage points a year. Day to day the two differ by 10.5 bp, so a true shortfall of 2 bp a day would take about 216 sessions to detect.

Decisions that differ
14 of 951
the backtest itself switches 78 times
Backtest, a year
7.6%
decides and fills at the close
Live, a year
6.4%
market-on-close fills
Daily tracking difference
10.5 bp
216 sessions to detect 2 bp a day
Growth of 1: the backtest and its live twin0.901.101.301.50session 1session 250session 500session 750session 1000
Price against its average: close (x), live view (y)−4%−4%−2%−2%0%0%2%2%4%4%at the close, what the backtest usedwhat the live process sawWhere the gap comes fromBacktest7.64%Decisions−0.56Fills0.00Costs−0.63Live6.44%return a year; steps in percentage points
BacktestLive●A session where both decide alike●A session where live and backtest decide differently

The rule holds the asset when its price is above the 50-day average, otherwise cash earning nothing; the intraday price is a Brownian bridge between closes. The steps add up exactly: each adds one live difference to the one before.

Synthetic: one trend rule on a random walk of 1,000 sessions, backtested on closes and run live with the settings above. Left: each session’s distance from the average at the close against what the live process saw; red points sit in opposite quadrants and are the decisions that differ. Right: the gap between the backtest’s and the live book’s annual return, split into three steps. Move the decision to an hour before the close, then switch the fill to the decision price.

Deciding on the close itself leaves one difference, from the feed alone, which still moves the annual return by −0.43 percentage points: the gap depends on which sessions flip, not only on how many. Deciding an hour early raises the count to 30 sessions. Whatever the settings, backtest the decision time you can actually meet, and treat what remains of the decision gap as noise and the cost gap as a certain loss.

A staged start

Whether a strategy works and whether live matches its backtest are separate questions. The first was answered imperfectly by research, and live returns will usually be worse than the backtest’s for reasons Bailey, Borwein, López de Prado and Zhu describe.5 The second can be tested, and the arithmetic of detection shows how slowly. If the daily differences between live and backtest returns have a standard deviation σ (daily here, unlike the series notation), telling a true mean shortfall δ from zero needs about ((1.96 + 0.84)·σ/δ)² sessions.

Daily difference, σShortfall 1 bp a day2 bp a day5 bp a day
5 bp196 sessions49 sessions8 sessions
10 bp784 sessions196 sessions32 sessions
20 bp3,136 sessions784 sessions126 sessions

Sessions needed for a two-sided 5% test with 80% power on paired daily differences. A 2 bp daily shortfall is about 5% a year. The formula assumes independent, roughly normal differences. When a few days with different decisions dominate them, as in the simulator, the test rejects a true mean of zero far more often than 5%: 13% of the time at 216 sessions (the lab’s count for 2 bp a day) and 30% at 35 (its count for 5 bp a day), drawing from the simulator’s own daily differences. Treat the short entries as lower bounds.

The lab’s default has a daily difference of 10.5 bp (close to the table’s 10 bp row), so a 2 bp daily shortfall, about 5% a year, would take 216 sessions, most of a year, to see in P&L. Costs are better measured fill by fill, against the cost model, where each order is a data point. Gates between capital stages rest on mechanics that can be checked each session: decisions identical to the research parity rerun, all orders acknowledged and then filled or cancelled for a known reason, fills within the cost model’s band, positions at the broker equal to the targets, and all alerts explained. We move from paper to a small allocation, then to larger ones, only after a run of clean sessions at each stage, and each stage is sized so that its worst plausible loss is a price we are willing to pay for the information.

Restart and recovery

Processes die: a host restarts, a network drops, a deploy goes wrong mid-session. A live strategy must be able to start at any moment and reach the right positions without doubling anything. Three rules do the work. Read the broker’s working orders first, with their remaining quantities, and its positions second, so a fill landing between the two reads can only make you trade too little. Write each order to a durable journal, under an id derived from the release, the session and the instrument, before sending it. And on restart, look each journalled order up at the broker before sending anything, among all of the session’s orders and executions, filled and cancelled ones included (a list of open orders drops an order that filled while the process was down), and size anything resent from the positions and working orders just read, never from the journal alone. Some brokers reject a repeated client order id, others do not or do so only within a session, and the order manager should not depend on which. The lookup is safe only once the broker has caught up: an order sent in the instant before the crash may not be listed yet, so the restart first waits for the broker’s order and execution reports, or for a quiet interval (a heuristic: a report that arrives after it still doubles the order, unless the resend reuses the journalled id and the broker rejects repeated ids). In a simulation that skips the wait, the restart sends that order again and ends with 60 of C against a target of 30. Kleppmann’s treatment of idempotence and exactly-once processing is the best general reference.6

site_research/fieldnotes/golive.py
def rebalance(broker: FakeBroker, journal: dict, targets: dict, session: str, release: str,
              crash_after: int = None, wait_for_reports: bool = True) -> None:
    """Move the account to `targets`, safely after any crash.

    0. Wait until the broker has reported every order sent before the crash (its order and
       execution reports, or a quiet interval). Until then a journalled order missing from its
       list may still be on its way, and resending it would double it.
    1. Working orders are read FIRST (their REMAINING quantity), then positions, so a fill
       landing between the two reads can only make us trade too little, never too much.
    2. Every intent is written to a durable journal before it is sent, under an id made of the
       release, the session, the symbol and a sequence number. On restart, an intent already
       at the broker is never resent: we look it up by id among all the session's orders and
       executions (a list of open orders drops one that filled while we were down) instead of
       trusting the broker to reject duplicates.
    3. Only the difference between target and (position + working) is traded, and that
       includes a journalled order resent after a crash: its size comes from what was just
       read, never from the journal alone."""
    if wait_for_reports:
        broker.catch_up()
    working = {}
    for symbol, qty, filled, status in broker.orders.values():
        if status == "working":
            working[symbol] = working.get(symbol, 0) + qty - filled
    held = dict(broker.positions)
    sent = 0
    for symbol, target in sorted(targets.items()):
        mine = [cid for cid, j in journal.items() if j["symbol"] == symbol and j["session"] == session]
        unsent = [cid for cid in mine if cid not in broker.orders]
        delta = target - held.get(symbol, 0) - working.get(symbol, 0)
        if delta == 0:
            continue
        if unsent:                                           # journalled, crashed before sending
            cid = unsent[0]
            journal[cid]["qty"] = delta                      # re-sized from positions + working
        else:
            cid = f"{release}-{session}-{symbol}-{len(mine) + 1}"
            journal[cid] = {"symbol": symbol, "session": session, "qty": delta}
        if crash_after is not None and sent == crash_after:
            raise RuntimeError("process killed after journaling, before sending")
        broker.submit(cid, symbol, journal[cid]["qty"])
        sent += 1
Figure 3Kill the process, restart it
The process dies
While it is down, the first order fills

Start: A 20, B 0, C 0, D 80. Target: A 100, B −50, C 30, D 80. The process is killed, restarted, and asked to finish the rebalance. 1 of 3 designs ends on target.

Memory
✗ off target
Trusts the positions it remembers from the start of the day; a fresh id for every order.
Orders sentSide
before the crash · Abuy 80
before the crash · Bsell 50
after restart · Abuy 80
after restart · Bsell 50
after restart · Cbuy 30
FinalABCD
held180−1003080
Broker positions
✗ off target
Reads positions from the broker, but not the orders still working there.
Orders sentSide
before the crash · Abuy 80
before the crash · Bsell 50
after restart · Abuy 40
after restart · Bsell 50
after restart · Cbuy 30
FinalABCD
held140−1003080
Broker state and a journal
✓ on target
Reads working orders, then positions; journals each order and looks it up before resending.
Orders sentSide
before the crash · Abuy 80
before the crash · Bsell 50
after restart · Cbuy 30
FinalABCD
held100−503080

The fake broker accepts a repeated client id, as some real ones do, so nothing but the order manager’s own design prevents a double. Try “after all three”: even a routine restart doubles every order in the first design.

Synthetic: four instruments, two already held, and a fake broker that accepts repeated client ids. With the defaults the process dies after two orders (having journalled a third), half of the first order fills while it is down, and the restart finishes the rebalance: memory ends at A 180, B −100, C 30, D 80, the journalled design at A 100, B −50, C 30, D 80. Kill it after one order and let that order fill completely, and the second design recovers; kill it before any order and all three do.

The same discipline covers the rest of a strategy’s state. Anything it needs to decide, such as indicator histories, is either recomputed from data at start-up (the warm-up) or stored with the fingerprint of the release that wrote it, so a new release never reads an old release’s state by accident. Building strategies from components shows how to test which state a component can re-derive. A second release running beside the first, as a canary, needs its own account or its own share of one, or the two will net each other’s orders.

Kill switches, deployment and rollback

From 27 July 2012 Knight Capital deployed new order-routing code to its eight servers in stages, and a technician missed one. The new code reused a flag that had once switched on a retired function; when trading opened on 1 August the eighth server woke that function up, and in about 45 minutes it sent millions of orders and lost more than 460 million dollars. Removing the new code from the other seven servers during the incident made it worse, because the flag was still set. The SEC’s order found no written deployment procedure, no second person reviewing the deployment, and no automated control that stopped orders once losses or positions passed a threshold.7 Two lessons are specific: delete dead code, and roll back configuration together with the code that reads it.

Regulators have since written the missing controls down. The SEC’s market access rule requires brokers to apply pre-trade risk checks to every order.8 The European rules for algorithmic trading require firms to test algorithms before deployment, to deploy them with limits on what they may trade at first, to check each order before it leaves, and to keep a kill function that cancels all unexecuted orders at once.9 A strategy of any size should hold itself to the same standard.

site_research/fieldnotes/golive.py
class KillSwitch:
    """Trips on the first breach, calls `on_trip` (which cancels working orders and blocks new
    ones), and stays tripped until a person resets it: it never re-arms itself. The order
    manager asks `allow()` before every order. Call `check` only while the market is open (by
    the exchange calendar), or the data-age limit trips on every overnight gap. Run a second
    copy outside the strategy's own process, so a hung strategy cannot keep trading."""

    def __init__(self, limits: dict, on_trip):
        self.limits, self.on_trip = limits, on_trip
        self.tripped, self.reason = False, None

    def check(self, session_pnl: float, orders_last_min: int, gross: float, data_age_s: float) -> None:
        L = self.limits
        for hit, why in ((session_pnl < -L["max_loss"], "session loss"),
                         (orders_last_min > L["max_orders_per_min"], "order rate"),
                         (gross > L["max_gross"], "gross exposure"),
                         (data_age_s > L["max_data_age_s"], "stale data")):
            if hit and not self.tripped:
                self.tripped, self.reason = True, why
                self.on_trip(why)

    def allow(self) -> bool:
        return not self.tripped

    def reset(self, who: str) -> None:
        self.tripped, self.reason = False, f"reset by {who}"
A kill switch that trips on the first breach, cancels working orders through its callback, blocks new ones and does not re-arm itself. The order manager asks it before every order. Call it only while the market is open, by the exchange calendar, or the data-age limit trips on every overnight gap.

Deployment follows the same logic. We deploy outside market hours, one change at a time, with the previous fingerprint kept ready. After a deploy the system proves which build it is running before it trades, and the first session is compared with the parity rerun. Rollback means redeploying the previous fingerprint with its own configuration, which works only if the new release wrote no state the old one cannot read. Google’s practice of canarying, releasing to a small share first and comparing it with the rest, carries over directly: a new release can run on paper or on a small allocation beside the old one before it replaces it.10

Before the first live sessionChecked by
The running build’s fingerprint equals the tested oneThe system, at start-up
Research parity: zero differences over the recent pastA rerun of the frozen release
Paper sessions through a restart, a holiday and a half-dayPaper trading, both kinds
Pre-trade limits on order size, price and positionThe order manager, on every order
Kill switch tested by tripping it on paperA person, with a record
Stages, their length and their exit criteria written downA person, before the first session
Rollback rehearsed: previous fingerprint and its configurationA deploy to paper
Decisions that differ, across paths
17.4
of about 950 sessions on average, over 40 synthetic paths
Decision gap across paths
±1.2
percentage points a year; mean 0.00
Sessions to see 2 bp a day
216
at the lab’s 10.5 bp daily difference
Orders doubled after the restart
none
with broker state, a journal, and lookup by id once the broker has reported

Once the strategy is live, the work becomes watching it: Monitoring a live book.

References

  1. QuantConnect (n.d.). QuantConnect paper trading. QuantConnect documentation. www.quantconnect.com/docs/v2/cloud-platform/live-trading/brokerages/quantconnect-paper-trading
  2. Humble, J., & Farley, D. (2010). Continuous delivery: Reliable software releases through build, test, and deployment automation. Addison-Wesley.
  3. McNutt, D. (2016). Release engineering. In B. Beyer, C. Jones, J. Petoff, & N. R. Murphy (Eds.), Site reliability engineering: How Google runs production systems (ch. 8). O’Reilly Media. sre.google/sre-book/release-engineering/
  4. QuantConnect (n.d.). LEAN source: MarketOnCloseOrder (DefaultSubmissionTimeBuffer). GitHub. github.com/QuantConnect/Lean/blob/master/Common/Orders/MarketOnCloseOrder.cs
  5. Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2014). Pseudo-mathematics and financial charlatanism: The effects of backtest overfitting on out-of-sample performance. Notices of the American Mathematical Society, 61(5), 458–471. https://doi.org/10.1090/noti1105
  6. Kleppmann, M. (2017). Designing data-intensive applications: The big ideas behind reliable, scalable, and maintainable systems. O’Reilly Media.
  7. U.S. Securities and Exchange Commission (2013). In the matter of Knight Capital Americas LLC (Exchange Act Release 34-70694). Administrative proceeding. www.sec.gov/files/litigation/admin/2013/34-70694.pdf
  8. U.S. Securities and Exchange Commission (2010). Risk management controls for brokers or dealers with market access (Exchange Act Release 34-63241). Final rule. www.sec.gov/files/rules/final/2010/34-63241.pdf
  9. European Commission (2017). Commission Delegated Regulation (EU) 2017/589 of 19 July 2016 on the organisational requirements of investment firms engaged in algorithmic trading (Articles 5–8, 12 and 15). Official Journal of the European Union. eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32017R0589
  10. Warner, A., & Davidovič, Š. (2018). Canarying releases. In B. Beyer, N. R. Murphy, D. K. Rensin, K. Kawahara, & S. Thorne (Eds.), The site reliability workbook: Practical ways to implement SRE (ch. 16). O’Reilly Media. sre.google/workbook/canarying-releases/