Production operations · Intermediate
Going live
Going live tests the machinery around a strategy, and most of what can go wrong there can be measured before money moves.
Educational material only. Not investment advice.
Contents
Abstract
We freeze each release under one fingerprint and rerun it over the recent past until its decisions match research exactly. Some live differences are designed in, and the gap between a backtest and its live twin splits into decisions, fills and costs. A power calculation shows that early live P&L cannot detect a realistic shortfall for months, so capital stages are gated on mechanics that can be checked each session. A journal of orders, state rebuilt from the broker, a kill switch that stays tripped and a one-step rollback let the system fail without doubling anything.
Key takeaways
- Paper trading exercises the plumbing; its fills are the backtest’s own assumptions, so it says little about performance.
- Freeze code, parameters, engine, library versions and data sources under one fingerprint, and rerun the release over the recent past until its decisions match research exactly.
- Backtest the decision time you can actually meet live. What remains of the decision gap is noise; costs are a certain loss.
- Early live P&L cannot detect a realistic shortfall for months, so capital stages are gated on mechanics checked every session.
- Read working orders, then positions, from the broker at each start; journal an order before sending it, and look it up, once the broker has reported, before resending; size any resend from what the broker holds, never from the journal alone.
- Keep a kill switch that trips on the first breach and needs a person to reset it, and a rollback one redeploy away, which holds only if the new release wrote no state the old one cannot read.
Before you start
- QuantConnect and LEAN, for one engine in backtest and live
- Running backtests at scale, for fingerprinted runs
- TWAP, slicing and jitter, for costs and fills
- Data for systematic trading, for timestamps and the close
- Hypothesis tests and statistical power
A strategy that works in a backtest has passed a test of an idea. Going live tests the machinery around it, and most of what goes wrong there is engineering. The habit that pays most is to ask of each difference between live and backtest whether it was designed in or discovered.
Two paths
A live system has two paths. The release path takes a tested strategy to production and back out again. The live path runs every session: data in, decisions, orders out, fills back, and the account’s true state read from the broker. The two meet in one place, the build that is running, and that build must be the one that was tested.
Paper trading and its limits
Paper trading runs the strategy on live data with simulated fills, and it comes in two kinds. A platform’s simulator fills orders inside the platform: QuantConnect’s paper brokerage, for example, part of the engine described in QuantConnect and LEAN, fills market orders at once and in full, at the bid or ask when quotes are available, with no slippage, and models fees, buying power and margin, but it never talks to a broker.1 A broker’s own paper account goes through the broker’s interface, order types and rejections, with fills that are still simulated. Use the first to test the schedule, the handling of holidays and half-days, and a restart in the middle of a session; use the second to test the connection, the order formats and the broker’s rules. Exercise each of them before capital is at risk.
Neither tells you much about performance. A paper fill has no queue position and no market impact, and in a platform simulator it is the backtest’s own fill model. A paper account that matches the backtest has confirmed the backtest’s assumptions, not the market’s. The first real fills, measured against the cost model order by order, are the earliest evidence about performance there is.
Freezing a release
Before a strategy trades, we freeze it: the code, the parameters, the engine version, the versions of the libraries it imports and the data sources become one release with one identity. Humble and Farley make the general case for building every release once, from version control, and promoting the same artefact through each stage;2 Google’s release engineers add that builds must be hermetic, so the same inputs always give the same result.3 A content hash over everything that decides behaviour makes that identity checkable. It is the fingerprint that Running backtests at scale uses to name a research run, with the data named by its source rather than digested, since a live feed has no file to hash.
def fingerprint(sources: dict, params: dict, engine: str, data: dict) -> str:
"""One hash for everything that decides behaviour: source files, parameters, the engine's
version and the data sources (vendor, dataset, normalisation). Put the dependency lockfile
among the sources, so a library upgrade changes the fingerprint too. Two runs with the same
fingerprint must make the same decisions on the same data."""
h = hashlib.sha256()
for name in sorted(sources):
h.update(name.encode() + b"\0" + sources[name].encode() + b"\0")
for part in (params, {"engine": engine}, data):
h.update(json.dumps(part, sort_keys=True).encode() + b"\0")
return h.hexdigest()[:16]Research parity
The frozen release is then rerun as a backtest over the most recent weeks, on the data the live system would have seen, and its decisions are compared with those of the research version, session by session and order by order. The accepted number of differences is zero, because nothing about the market can explain one.
| Cause of a difference | How it shows up |
|---|---|
| A changed default in the engine or a library | Every session differs a little from one date on |
| A data vintage adjusted after the fact | Differences cluster around dividends, splits or rolls |
| A parameter read from the wrong file | One rule or one instrument behaves differently throughout |
| Randomness or thread order (seeds, parallel sums) | Two reruns of the same release disagree with each other |
Each of these would also make the live system disagree with its own backtest, and they are far cheaper to find before the money moves. Once the strategy is live, the same rerun, over the session just traded, becomes the simulation that Live versus backtest reconciliation compares with the account.
Where live and backtest part ways
Even a perfectly frozen strategy trades differently live than in its backtest, for reasons that are designed in and should be known in advance.
| Difference | Backtest | Live | What to do |
|---|---|---|---|
| Decision time | Decides on the day’s close | Decides minutes before it, on the last trade | Backtest the decision time you can actually meet |
| Data resolution | A daily bar | A stream of trades and quotes, or minute bars | Build the live signal from the same bars the backtest used |
| The close itself | The official closing price | Your feed’s last print, which can differ | Know which close each source reports |
| Fills | At the close, in full | Queue position, partial fills, the auction’s cut-off | Model the order type you will send; measure live fills against it |
| Costs | A modelled spread and fee | Spread, impact, fees, borrow, financing | Measure each fill against the cost model |
| State | Starts clean, never crashes | Restarts mid-session, inherits positions | Rebuild state from the broker at each start |
| Calendar and clock | Assumed sessions | Holidays, half-days, daylight saving | Schedule in exchange time from an exchange calendar |
Anything else found live was discovered, not designed, and is a defect until proved otherwise. Which close a source reports, and how it is timestamped, is covered in Data for systematic trading.
A market-on-close order has to reach the exchange before the auction’s cut-off, so the decision must come earlier still: LEAN, for instance, by default refuses a market-on-close order sent less than 15.5 minutes before the close (the buffer is configurable).4 A simulator makes the designed-in part of the gap concrete. A 50-day trend rule is backtested on closes and run live on a feed, and the gap between the two is built up one difference at a time.
- Step 1 of 5
The backtest
The backtest decides on each session’s official close and trades at it, paying the modelled 2 bp per unit traded. Over 1,000 sessions, about 4.0 years, the rule switches 78 times and returns 7.6% a year.
- Step 2 of 5
Decisions before the close
The live twin decides 20 minutes before the close, in time for a market-on-close order, on a feed whose price differs from the official one by 2 bp. The two versions decide differently on 14 sessions of the 951 on which both have a signal, marked in red. Those sessions change the live book’s annual return by −0.56 percentage points.
- Step 3 of 5
Fills
A market-on-close order fills at the official close, the price the backtest assumed, so the fills change the return by 0.00 percentage points. Filling at the price seen at the decision instead would have changed it by −0.49 percentage points on this path; the lab below has the switch.
- Step 4 of 5
Costs
Live trading pays 5 bp per unit traded against the backtest’s 2. That takes −0.63 percentage points a year, and the whole gap is −1.20 percentage points: the live book returns 6.4% a year against the backtest’s 7.6%.
- Step 5 of 5
Luck and certainty
The sign of the decision gap is luck. Over 40 synthetic paths it averaged 0.00 percentage points a year with a standard deviation of 1.21 percentage points, and filling at the decision price averaged −0.07 percentage points with a standard deviation of 0.49. Costs are the opposite: each unit traded live pays the extra 3 bp, so that step is small, certain and against you on any path.
Synthetic: one trend rule on a random walk of 1,000 sessions, the simulator of the laboratory below at its default settings; the spread across paths is the same simulator on 40 other random walks. Without script, or with reduced motion, the figure shows its final state.
The lab runs the same simulator with the settings in your hands: the decision time, the gap between the feed and the official close, the fill and the cost.
The live process decides differently from the backtest on 14 of 951 sessions, all of them days when the price was close to its 50-day average. Backtest 7.6% a year, live 6.4%: decisions −0.56, fills 0.00 and costs −0.63 percentage points a year. Day to day the two differ by 10.5 bp, so a true shortfall of 2 bp a day would take about 216 sessions to detect.
The rule holds the asset when its price is above the 50-day average, otherwise cash earning nothing; the intraday price is a Brownian bridge between closes. The steps add up exactly: each adds one live difference to the one before.
Deciding on the close itself leaves one difference, from the feed alone, which still moves the annual return by −0.43 percentage points: the gap depends on which sessions flip, not only on how many. Deciding an hour early raises the count to 30 sessions. Whatever the settings, backtest the decision time you can actually meet, and treat what remains of the decision gap as noise and the cost gap as a certain loss.
A staged start
Whether a strategy works and whether live matches its backtest are separate questions. The first was answered imperfectly by research, and live returns will usually be worse than the backtest’s for reasons Bailey, Borwein, López de Prado and Zhu describe.5 The second can be tested, and the arithmetic of detection shows how slowly. If the daily differences between live and backtest returns have a standard deviation σ (daily here, unlike the series notation), telling a true mean shortfall δ from zero needs about ((1.96 + 0.84)·σ/δ)² sessions.
| Daily difference, σ | Shortfall 1 bp a day | 2 bp a day | 5 bp a day |
|---|---|---|---|
| 5 bp | 196 sessions | 49 sessions | 8 sessions |
| 10 bp | 784 sessions | 196 sessions | 32 sessions |
| 20 bp | 3,136 sessions | 784 sessions | 126 sessions |
Sessions needed for a two-sided 5% test with 80% power on paired daily differences. A 2 bp daily shortfall is about 5% a year. The formula assumes independent, roughly normal differences. When a few days with different decisions dominate them, as in the simulator, the test rejects a true mean of zero far more often than 5%: 13% of the time at 216 sessions (the lab’s count for 2 bp a day) and 30% at 35 (its count for 5 bp a day), drawing from the simulator’s own daily differences. Treat the short entries as lower bounds.
The lab’s default has a daily difference of 10.5 bp (close to the table’s 10 bp row), so a 2 bp daily shortfall, about 5% a year, would take 216 sessions, most of a year, to see in P&L. Costs are better measured fill by fill, against the cost model, where each order is a data point. Gates between capital stages rest on mechanics that can be checked each session: decisions identical to the research parity rerun, all orders acknowledged and then filled or cancelled for a known reason, fills within the cost model’s band, positions at the broker equal to the targets, and all alerts explained. We move from paper to a small allocation, then to larger ones, only after a run of clean sessions at each stage, and each stage is sized so that its worst plausible loss is a price we are willing to pay for the information.
Restart and recovery
Processes die: a host restarts, a network drops, a deploy goes wrong mid-session. A live strategy must be able to start at any moment and reach the right positions without doubling anything. Three rules do the work. Read the broker’s working orders first, with their remaining quantities, and its positions second, so a fill landing between the two reads can only make you trade too little. Write each order to a durable journal, under an id derived from the release, the session and the instrument, before sending it. And on restart, look each journalled order up at the broker before sending anything, among all of the session’s orders and executions, filled and cancelled ones included (a list of open orders drops an order that filled while the process was down), and size anything resent from the positions and working orders just read, never from the journal alone. Some brokers reject a repeated client order id, others do not or do so only within a session, and the order manager should not depend on which. The lookup is safe only once the broker has caught up: an order sent in the instant before the crash may not be listed yet, so the restart first waits for the broker’s order and execution reports, or for a quiet interval (a heuristic: a report that arrives after it still doubles the order, unless the resend reuses the journalled id and the broker rejects repeated ids). In a simulation that skips the wait, the restart sends that order again and ends with 60 of C against a target of 30. Kleppmann’s treatment of idempotence and exactly-once processing is the best general reference.6
def rebalance(broker: FakeBroker, journal: dict, targets: dict, session: str, release: str,
crash_after: int = None, wait_for_reports: bool = True) -> None:
"""Move the account to `targets`, safely after any crash.
0. Wait until the broker has reported every order sent before the crash (its order and
execution reports, or a quiet interval). Until then a journalled order missing from its
list may still be on its way, and resending it would double it.
1. Working orders are read FIRST (their REMAINING quantity), then positions, so a fill
landing between the two reads can only make us trade too little, never too much.
2. Every intent is written to a durable journal before it is sent, under an id made of the
release, the session, the symbol and a sequence number. On restart, an intent already
at the broker is never resent: we look it up by id among all the session's orders and
executions (a list of open orders drops one that filled while we were down) instead of
trusting the broker to reject duplicates.
3. Only the difference between target and (position + working) is traded, and that
includes a journalled order resent after a crash: its size comes from what was just
read, never from the journal alone."""
if wait_for_reports:
broker.catch_up()
working = {}
for symbol, qty, filled, status in broker.orders.values():
if status == "working":
working[symbol] = working.get(symbol, 0) + qty - filled
held = dict(broker.positions)
sent = 0
for symbol, target in sorted(targets.items()):
mine = [cid for cid, j in journal.items() if j["symbol"] == symbol and j["session"] == session]
unsent = [cid for cid in mine if cid not in broker.orders]
delta = target - held.get(symbol, 0) - working.get(symbol, 0)
if delta == 0:
continue
if unsent: # journalled, crashed before sending
cid = unsent[0]
journal[cid]["qty"] = delta # re-sized from positions + working
else:
cid = f"{release}-{session}-{symbol}-{len(mine) + 1}"
journal[cid] = {"symbol": symbol, "session": session, "qty": delta}
if crash_after is not None and sent == crash_after:
raise RuntimeError("process killed after journaling, before sending")
broker.submit(cid, symbol, journal[cid]["qty"])
sent += 1Start: A 20, B 0, C 0, D 80. Target: A 100, B −50, C 30, D 80. The process is killed, restarted, and asked to finish the rebalance. 1 of 3 designs ends on target.
| Orders sent | Side |
|---|---|
| before the crash · A | buy 80 |
| before the crash · B | sell 50 |
| after restart · A | buy 80 |
| after restart · B | sell 50 |
| after restart · C | buy 30 |
| Final | A | B | C | D |
|---|---|---|---|---|
| held | 180 | −100 | 30 | 80 |
| Orders sent | Side |
|---|---|
| before the crash · A | buy 80 |
| before the crash · B | sell 50 |
| after restart · A | buy 40 |
| after restart · B | sell 50 |
| after restart · C | buy 30 |
| Final | A | B | C | D |
|---|---|---|---|---|
| held | 140 | −100 | 30 | 80 |
| Orders sent | Side |
|---|---|
| before the crash · A | buy 80 |
| before the crash · B | sell 50 |
| after restart · C | buy 30 |
| Final | A | B | C | D |
|---|---|---|---|---|
| held | 100 | −50 | 30 | 80 |
The fake broker accepts a repeated client id, as some real ones do, so nothing but the order manager’s own design prevents a double. Try “after all three”: even a routine restart doubles every order in the first design.
The same discipline covers the rest of a strategy’s state. Anything it needs to decide, such as indicator histories, is either recomputed from data at start-up (the warm-up) or stored with the fingerprint of the release that wrote it, so a new release never reads an old release’s state by accident. Building strategies from components shows how to test which state a component can re-derive. A second release running beside the first, as a canary, needs its own account or its own share of one, or the two will net each other’s orders.
Kill switches, deployment and rollback
From 27 July 2012 Knight Capital deployed new order-routing code to its eight servers in stages, and a technician missed one. The new code reused a flag that had once switched on a retired function; when trading opened on 1 August the eighth server woke that function up, and in about 45 minutes it sent millions of orders and lost more than 460 million dollars. Removing the new code from the other seven servers during the incident made it worse, because the flag was still set. The SEC’s order found no written deployment procedure, no second person reviewing the deployment, and no automated control that stopped orders once losses or positions passed a threshold.7 Two lessons are specific: delete dead code, and roll back configuration together with the code that reads it.
Regulators have since written the missing controls down. The SEC’s market access rule requires brokers to apply pre-trade risk checks to every order.8 The European rules for algorithmic trading require firms to test algorithms before deployment, to deploy them with limits on what they may trade at first, to check each order before it leaves, and to keep a kill function that cancels all unexecuted orders at once.9 A strategy of any size should hold itself to the same standard.
class KillSwitch:
"""Trips on the first breach, calls `on_trip` (which cancels working orders and blocks new
ones), and stays tripped until a person resets it: it never re-arms itself. The order
manager asks `allow()` before every order. Call `check` only while the market is open (by
the exchange calendar), or the data-age limit trips on every overnight gap. Run a second
copy outside the strategy's own process, so a hung strategy cannot keep trading."""
def __init__(self, limits: dict, on_trip):
self.limits, self.on_trip = limits, on_trip
self.tripped, self.reason = False, None
def check(self, session_pnl: float, orders_last_min: int, gross: float, data_age_s: float) -> None:
L = self.limits
for hit, why in ((session_pnl < -L["max_loss"], "session loss"),
(orders_last_min > L["max_orders_per_min"], "order rate"),
(gross > L["max_gross"], "gross exposure"),
(data_age_s > L["max_data_age_s"], "stale data")):
if hit and not self.tripped:
self.tripped, self.reason = True, why
self.on_trip(why)
def allow(self) -> bool:
return not self.tripped
def reset(self, who: str) -> None:
self.tripped, self.reason = False, f"reset by {who}"Deployment follows the same logic. We deploy outside market hours, one change at a time, with the previous fingerprint kept ready. After a deploy the system proves which build it is running before it trades, and the first session is compared with the parity rerun. Rollback means redeploying the previous fingerprint with its own configuration, which works only if the new release wrote no state the old one cannot read. Google’s practice of canarying, releasing to a small share first and comparing it with the rest, carries over directly: a new release can run on paper or on a small allocation beside the old one before it replaces it.10
| Before the first live session | Checked by |
|---|---|
| The running build’s fingerprint equals the tested one | The system, at start-up |
| Research parity: zero differences over the recent past | A rerun of the frozen release |
| Paper sessions through a restart, a holiday and a half-day | Paper trading, both kinds |
| Pre-trade limits on order size, price and position | The order manager, on every order |
| Kill switch tested by tripping it on paper | A person, with a record |
| Stages, their length and their exit criteria written down | A person, before the first session |
| Rollback rehearsed: previous fingerprint and its configuration | A deploy to paper |
Once the strategy is live, the work becomes watching it: Monitoring a live book.
References
- QuantConnect (n.d.). QuantConnect paper trading. QuantConnect documentation. www.quantconnect.com/docs/v2/cloud-platform/live-trading/brokerages/quantconnect-paper-trading
- Humble, J., & Farley, D. (2010). Continuous delivery: Reliable software releases through build, test, and deployment automation. Addison-Wesley.
- McNutt, D. (2016). Release engineering. In B. Beyer, C. Jones, J. Petoff, & N. R. Murphy (Eds.), Site reliability engineering: How Google runs production systems (ch. 8). O’Reilly Media. sre.google/sre-book/release-engineering/
- QuantConnect (n.d.). LEAN source: MarketOnCloseOrder (DefaultSubmissionTimeBuffer). GitHub. github.com/QuantConnect/Lean/blob/master/Common/Orders/MarketOnCloseOrder.cs
- Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2014). Pseudo-mathematics and financial charlatanism: The effects of backtest overfitting on out-of-sample performance. Notices of the American Mathematical Society, 61(5), 458–471. https://doi.org/10.1090/noti1105
- Kleppmann, M. (2017). Designing data-intensive applications: The big ideas behind reliable, scalable, and maintainable systems. O’Reilly Media.
- U.S. Securities and Exchange Commission (2013). In the matter of Knight Capital Americas LLC (Exchange Act Release 34-70694). Administrative proceeding. www.sec.gov/files/litigation/admin/2013/34-70694.pdf
- U.S. Securities and Exchange Commission (2010). Risk management controls for brokers or dealers with market access (Exchange Act Release 34-63241). Final rule. www.sec.gov/files/rules/final/2010/34-63241.pdf
- European Commission (2017). Commission Delegated Regulation (EU) 2017/589 of 19 July 2016 on the organisational requirements of investment firms engaged in algorithmic trading (Articles 5–8, 12 and 15). Official Journal of the European Union. eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32017R0589
- Warner, A., & Davidovič, Š. (2018). Canarying releases. In B. Beyer, N. R. Murphy, D. K. Rensin, K. Kawahara, & S. Thorne (Eds.), The site reliability workbook: Practical ways to implement SRE (ch. 16). O’Reilly Media. sre.google/workbook/canarying-releases/