Production operations · Intermediate
Monitoring a live book
End-of-day checks that show their evidence, and that say so when they could not look.
Educational material only. Not investment advice.
Contents
Abstract
Each evening we check that the broker holds what the strategy intended, that each instrument filled its intended change, that exposure is within limits, that prices are from today and that the day’s P&L is explained. Every check returns pass, fail or unavailable, and the report ranks unavailable above pass. Each result carries a witness, the dated evidence it read in this pass, and each check is trusted only after it has caught the fault it was written for.
Key takeaways
- Compare every position the broker holds with what the strategy intended, and check exposure separately, because a wrong target passes the holdings check.
- Every check returns pass, fail or unavailable, and the headline is fail, else unavailable, else pass.
- Each result carries its witness, the dated evidence read in this pass; nothing is stored as passed and read back.
- A check is trusted only after a planted fault has made it fail; a matrix of faults against checks shows where one check stands alone.
- Explain the day’s P&L from holdings, price moves, fills and fees; the unexplained remainder points at its cause.
- Dashboards help people explore; checks run whether or not anyone looks; only pre-trade controls prevent a bad order.
Before you start
- Going live, for the live process being monitored
- Data for systematic trading, for timestamps and calendars
- TWAP, slicing and jitter, for orders and fills
Each evening a live book raises one question: did it do what the strategy intended? A monitor answers it with evidence or admits that it cannot. The most dangerous monitor is the one that cannot tell those two apart, because on the evening the broker stops answering it reports that all is well.
The daily loop
We run the same loop after every session. The strategy stamps its targets when it produces them. When the market closes, the monitor takes fresh snapshots from the broker (positions, fills, the statement of P&L) and from the price feed. Each check reads those snapshots, recomputes what it needs, and reports one of three results with the evidence it used. The report lists all the checks, each day, and only the results that need a person become alerts. The monitor also sends a heartbeat on each pass to a watcher that runs elsewhere, so a monitor that has stopped is noticed.
What to measure each session
Each check compares two things that should agree and were produced independently. A position is only wrong relative to something, and the choice of that something decides what the check can catch.
| Check | Compares | Catches | Cannot see |
|---|---|---|---|
| Holdings | Every position at the broker with the strategy’s targets, including positions no strategy owns, and the targets’ timestamp with the session | Missed, doubled and reversed orders; unknown positions; a strategy that did not run | A target that was itself wrong |
| Orders and fills | Our order log with the broker’s fills: repeated order ids, and each instrument’s net fill with its intended change | Duplicates, partial and missed fills; slicing an order is fine | A target that was itself wrong: the intended change is computed from it |
| Signs | Each position’s side with its target’s | A buy sent as a sell | Anything short of a reversal |
| Exposure | Gross exposure of the broker’s positions, over the account’s equity, against its limit | A strategy asking for too much | Risk inside the limit |
| Freshness | Each price’s timestamp, date included, against the close | Yesterday’s prices, frozen feeds, time zones read wrongly | A fresh price that is wrong |
| P&L | The broker’s P&L with holdings × price moves + fills − fees | Wrong prices, unexpected fees and charges | A doubled order: the broker’s own fills explain it |
The six checks in the listings and the labs. A production suite adds more that this synthetic session does not model: continuity (yesterday’s positions plus today’s fills equal today’s), cash and margin, corporate actions, orders left working at the close, and the strategy’s own invariants.
Holdings are read from the broker, never from our own records, because our records are what a bug corrupts. The last column matters as much as the third. Holdings can equal targets even when the target itself is wrong: in the synthetic session, a strategy that asks for forty times its intended position passes the holdings check and is caught only by the exposure check.
Pass, fail and unavailable
A check has three possible results. It passed: it looked, and what it saw was right. It failed: it looked, and something was wrong. Or it was unavailable: it could not look, because the broker did not answer, returned an empty list, or the check itself crashed. The report’s headline is fail if anything failed, else unavailable if anything could not look, and pass only when every check looked and passed.
def run(d: Day) -> list:
"""Every check, each isolated: a check that raises is UNAVAILABLE, never a crash of the report.
On a day the exchange calendar has no session there is nothing to check, and each check says so."""
if d.close_utc is None:
return [Result(name, PASS, "no session today by the exchange calendar: nothing to check", "calendar: exchange closed")
for name in CHECKS]
out = []
for name, fn in CHECKS.items():
try:
out.append(Result(name, *fn(d)))
except Exception as e: # noqa: BLE001
out.append(Result(name, UNAVAILABLE, f"the check itself failed ({type(e).__name__})", "no result"))
return out
def overall(results: list, unavailable_is_pass: bool = False) -> str:
"""The report's headline. Correctly: any FAIL, else any UNAVAILABLE, else PASS. The naive
version, which many dashboards amount to, reads UNAVAILABLE as PASS."""
s = {r.status for r in results}
if FAIL in s:
return FAIL
if UNAVAILABLE in s and not unavailable_is_pass:
return UNAVAILABLE
return PASSThe rule matters on the evening the broker’s interface is down. On many dashboards an empty panel, a metric that stopped updating and a green light look the same at a glance.
- Step 1 of 4
A clean evening
The strategy stamps its targets at 19:40 UTC, the market closes at 20:00, and at 20:05 the monitor reads the broker. All six checks look, pass and say what they read: “6 broker positions read at 20:05, 6 targets from 19:40”. The report says pass.
- Step 2 of 4
An order sent twice
Now one order goes out twice. The holdings and orders checks fail and say where: “C: held 1,000, target 500”. The P&L check passes, because the broker’s P&L follows the broker’s own fills. One failure is enough, and the report says fail.
- Step 3 of 4
The broker stops answering
The same evening, the broker’s interface is down. Five of the six checks cannot look, and each says why; only the freshness check, which reads the price feed, still passes. No check can see the doubled order now.
- Step 4 of 4
Unavailable is not a pass
The three-valued report says unavailable: nobody knows what the book holds, and someone must find out before the next session. A board that reads missing data as fine says pass, on an evening when the broker holds 1,000 of C against a target of 500. Only the first reading sends someone to look.
Synthetic: the session of the laboratory below, through the six checks of monitor.py (the figure runs the TypeScript port, checked word for word against the Python). Without script, or with reduced motion, the figure shows its final state.
| Check | Result | What it found | Witness: what it read in this pass |
|---|---|---|---|
| Holdings equal targets | Unavailable | broker positions could not be read | no position snapshot this pass |
| Every order filled once | Unavailable | the broker’s fills could not be read | 5 orders in our log, no fills |
| Positions on the right side | Unavailable | no broker positions to compare | no position snapshot this pass |
| Gross exposure within limit | Unavailable | no broker positions to value | no position snapshot this pass |
| Prices fresh at the close | Pass | every price within 15 minutes of the close | 6 prices, the oldest stamped 20:00 |
| P&L explained | Unavailable | the broker’s P&L could not be read | no statement this pass |
Each check recomputes what it needs from this pass’s snapshots; nothing is read back from an earlier run.
| Result | Before the next session |
|---|---|
| Fail | Follow the check’s runbook; the book does not trade again until the failure is explained |
| Unavailable | Get the evidence another way (the broker’s website, a statement, a second feed) and rerun the check; if it still cannot look, treat it as a failure |
| Pass | Nothing, except reading the witnesses once a week to be sure they still say what they should |
Witnesses and planted faults
Each result carries a witness: the evidence the check examined in this pass, with its date. “6 broker positions read at 20:05, 6 targets from 19:40” is a witness; “OK” is not. Witnesses make silent failures visible. A check that read nothing (an empty snapshot, a missing file) is unavailable, not passed, and says so. A price from the day before is named as such: “stale: B stamped the day before at 20:00”. We never store “the check passed” and read it back; each pass recomputes from snapshots taken in that pass.
A check is trusted only after it has been seen to fail. For each check we plant the fault it exists for and confirm that it fires, a practice borrowed from mutation testing in software, where a test suite is judged by the deliberate defects it catches.1 Run over all the planted faults, the synthetic suite gives the matrix below. Each planted fault is caught by at least one check, and 5 of the 11 rest on a single check: those are the places where a second, independent check would earn its keep.
def session_close_utc(day: date):
"""The session's close in minutes from midnight UTC of `day`, from the exchange calendar,
or None when the exchange does not open that day."""
if day.weekday() >= 5 or day in HOLIDAYS:
return None
utc = datetime.combine(day, EARLY_CLOSES.get(day, REGULAR_CLOSE), EXCHANGE_TZ).astimezone(ZoneInfo("UTC"))
return (utc.date() - day).days * 1440 + utc.hour * 60 + utc.minute
def check_freshness(d: Day) -> tuple:
if not d.price_utc:
return UNAVAILABLE, "no prices were read", "empty price snapshot"
age = {s: d.close_utc - t for s, t in d.price_utc.items()}
stale = {s: a for s, a in age.items() if abs(a) > FRESH_MIN}
w = f"{len(age)} prices, the oldest stamped {when(d.close_utc - max(age.values()))}"
if not stale:
return PASS, "every price within 15 minutes of the close", w
offsets = set(stale.values())
off = next(iter(offsets))
if len(stale) == len(age) and len(offsets) == 1 and off % 30 == 0 and abs(off) <= 14 * 60:
return FAIL, (f"every price is off by exactly {abs(off) / 60:g} h: check the time zone before "
"blaming the feed"), w
return FAIL, "stale: " + ", ".join(f"{s} stamped {when(d.close_utc - a)}" for s, a in sorted(stale.items())), wThe close the check measures against is the session’s own, from the exchange calendar, never a fixed clock time. The regular close in New York falls at a different hour in UTC in summer and in winter, and earlier on a half-day, so a fixed time would report every price off by a whole hour for half the year and blame the time zone for what is really the calendar. On a holiday there is no session, and nothing to be fresh: the report says so for every check (run, in the first listing).
| Date | The day | Close, UTC |
|---|---|---|
| 2026-06-15 | the synthetic session (summer time) | 20:00 |
| 2026-01-15 | a winter session | 21:00 |
| 2026-11-27 | a half-day | 18:00 |
| 2026-11-26 | a holiday | no session |
From session_close_utc in monitor.py, which converts the exchange’s local close through the tz database. The calendar in the listing holds only the few dates this example needs; a production monitor reads the exchange’s published calendar.
Explaining the day’s P&L
The day’s P&L should be explainable from three things we already have: yesterday’s holdings times the day’s price moves, each fill marked to the close, and fees. Whatever the broker reports beyond that is unexplained, and its size and sign say where to look. A gap equal to a position times a price difference points at a price; a round number points at a fee; a gap that recurs every day points at something we do not model, such as financing or borrow.
Positions, fills and our prices explain 2,058; the broker reports 2,458. The remaining 400 is unexplained, and the size and sign point to the cause: a price, a fee, or a fill we do not know about.
When an order is doubled, this check still passes: the broker’s P&L follows the broker’s fills, and those fills explain it exactly. The P&L is right; the book is wrong, and the holdings and orders checks are the ones that say so. A P&L check run against the intended book instead would also catch it. The checks are designed as a set. Live versus backtest reconciliation compares the same holdings with the backtest’s, where a doubled order shows up at once.
Dashboards and checks
A dashboard answers the questions a person thinks to ask; a check answers one question every day whether anyone asks or not. Google’s site reliability engineers list dashboards and ad hoc analysis among the uses of monitoring, and keep alerts for symptoms that need a person now.2 Few’s rules for dashboard design, one screen with the exceptions made obvious, help the person looking;3 they cannot help on an evening when nobody looks.
| Dashboard | Check | |
|---|---|---|
| Question | Whatever the viewer asks | One question, fixed in advance |
| Runs | When someone looks | Every session, on a schedule |
| Missing data | An empty panel, easily missed | Unavailable, reported and alerted |
| Proof it works | None | A planted fault it must catch |
| Output | Pictures for a person | Pass, fail or unavailable, with a witness |
The SEC’s order on Knight Capital describes the cost of relying on the first kind. Knight’s position monitoring tool relied entirely on people watching it, generated no automated alerts, and did not display the limits it was meant to police, and it could not stop the entry of orders, so the orders kept flowing while positions grew.4 End-of-day checks find problems after the fact; preventing them takes pre-trade controls in the order path, which the SEC’s market access rule requires of brokers and which Going live describes.5
Which of these results should wake someone is the subject of Alerting that people act on.
References
- DeMillo, R. A., Lipton, R. J., & Sayward, F. G. (1978). Hints on test data selection: Help for the practicing programmer. Computer, 11(4), 34–41. https://doi.org/10.1109/c-m.1978.218136
- Ewaschuk, R. (2016). Monitoring distributed systems. In B. Beyer, C. Jones, J. Petoff, & N. R. Murphy (Eds.), Site reliability engineering: How Google runs production systems (ch. 6). O’Reilly Media. sre.google/sre-book/monitoring-distributed-systems/
- Few, S. (2006). Information dashboard design: The effective visual communication of data. O’Reilly Media.
- U.S. Securities and Exchange Commission (2013). In the matter of Knight Capital Americas LLC (Exchange Act Release 34-70694). Administrative proceeding. www.sec.gov/files/litigation/admin/2013/34-70694.pdf
- U.S. Securities and Exchange Commission (2010). Risk management controls for brokers or dealers with market access (Exchange Act Release 34-63241). Final rule. www.sec.gov/files/rules/final/2010/34-63241.pdf