Foundations · Intermediate
Data for systematic trading
The data errors that make a backtest look better than it is, and how to catch each one in code before a strategy reads it.
Educational material only. Not investment advice.
Contents
Abstract
Many errors in systematic research are data errors that look like results. On synthetic data where the true answer is known, we measure four of them: survivorship, adjusted prices read on the wrong date, futures contracts joined the wrong way, and fundamentals used before they were published. A survivors-only universe of identical stocks gains about five percentage points a year from selection alone. Each problem has a remedy that fits in a few lines of code, and the checks that enforce them are tested against planted faults.
Key takeaways
- Store every fact with the time it became known and read it at that time; a naive join on the period it describes leaks the future.
- A universe built from today’s list is survivors only; keep dead securities, their delisting returns and permanent identifiers.
- Use total-return adjusted prices (or raw prices plus cash flows) for returns and signals, and raw prices of the day for orders, filters and anything in price levels.
- Ratio-adjusted futures series reproduce a rolled position’s returns and difference-adjusted series one contract’s P&L; the unadjusted splice reproduces neither.
- Keep timestamps in UTC, key daily data by session, and use the IANA database rather than fixed offsets.
- Test every check against a planted fault, compare a suspect move with a reference before calling it an error, and reconcile vendors on overlapping history before switching.
Before you start
- Market anomalies, for why a backtest flatters itself
- pandas: joins, time series and time zones
- What a futures contract and its expiry are
A backtest cannot tell a data error from a result. A universe that holds only the companies that survived, a price that falls 74% on the day a stock splits, an earnings figure used weeks before it was published: each produces returns that look like an edge. The remedies are old and unglamorous, and all of them can be written as code and tested.
From delivery to strategy
We treat market data as a pipeline with a record of what arrived, when, and in what form. Raw deliveries are kept exactly as received and never edited. Everything a strategy reads is derived from them by code: adjusted prices, continuous futures, calendars, and the checks that decide whether a day’s data may be used at all. When a number looks wrong, the raw record shows whether the vendor sent it or we made it.
A raw record is useful only if it can answer, months later, what the system knew and when. Ours holds, for every delivery:
| Field | Why |
|---|---|
| Vendor, dataset and file | Which source to ask when a number is disputed |
| Received at (UTC) | What was knowable at any moment; the basis of every point-in-time read |
| The date or time the vendor says the data describe | Which session a delivery belongs to, separately from when it arrived |
| A content hash | Proof that a later read saw the same bytes; duplicates detected for free |
| The delivery it supersedes, if any | Corrections and restatements kept as vintages, never overwritten |
The same reader serves research, the backtest and the live system. When they read different copies of the data, the backtest and the live book disagree for reasons that have nothing to do with the strategy; QuantConnect and LEAN describes how one engine keeps that path shared.
Point-in-time data
A fact about a period has at least two timestamps: the period it describes and the moment it became known. Quarterly earnings for a quarter that ends in March are published weeks later and are sometimes restated months after that. Macroeconomic series are routinely re-estimated after publication. Croushore and Stark built a real-time data set for US macroeconomic variables because results computed from the latest figures can differ materially from what could have been computed at the time;1 the Federal Reserve Bank of St. Louis publishes every vintage of its series free of charge in ALFRED.2 Index membership, sector classifications and credit ratings change too, and a database that stores only the latest value rewrites history each time.
A point-in-time record keeps every vintage with its publication time, and a backtest asks it one question: what was known when the decision was made? Fama and French handled this for accounting data by matching fiscal years ending in calendar year t − 1 to returns from July of year t, a gap of at least six months that makes sure the figures were public.3 A lag rule of that kind is a safe default when publication dates are missing. Where they exist, use them, with a time of day: a figure released after the close cannot inform that day’s closing trade. Our listing assumes a date-only publication arrives after the close.
def known_on(facts: pd.DataFrame, day: pd.Timestamp) -> pd.DataFrame:
"""The record a strategy could use at the close of `day`: each period in its latest vintage.
Publication dates carry no time of day, so a fact is assumed to arrive after the close and
is first usable at the next session's close (strictly earlier date). With timestamps,
compare them with the decision time instead."""
seen = facts[facts["published"] < day].sort_values("published")
return seen.groupby("period_end", as_index=False).last()
def point_in_time(days: pd.DatetimeIndex, facts: pd.DataFrame, period: str = None) -> pd.Series:
"""What a backtest may use at each day's close: the latest period's value as known then,
or, with `period`, that period's value in the vintage known then."""
out = []
for d in days:
k = known_on(facts, d).sort_values("period_end")
if period is not None:
k = k[k["period_end"] == pd.Timestamp(period)]
out.append(k["value"].iloc[-1] if len(k) else float("nan"))
return pd.Series(out, index=days)
def naive(days: pd.DatetimeIndex, facts: pd.DataFrame) -> pd.Series:
"""The look-ahead mistake: today's final vintage, joined on the period it describes."""
final = facts.sort_values("published").groupby("period_end", as_index=False).last()
left = pd.DataFrame({"day": days.astype(final["period_end"].dtype)})
out = pd.merge_asof(left, final.sort_values("period_end"), left_on="day", right_on="period_end")
return out.set_index("day")["value"]At the close of 16 October 2023 a backtest may use 1.25, the quarter ending 30 June 2023 as published on 3 August 2023. The naive join reads 1.18 for the quarter ending 30 September 2023, a figure nobody could see until 15 March 2024.
| Quarter to | Published | EPS | At this close | Used by |
|---|---|---|---|---|
| 2023-03-31 | 2023-05-04 | 1.10 | known | — |
| 2023-06-30 | 2023-08-03 | 1.25 | known | point in time |
| 2023-09-30 | 2023-11-02 | 1.32 | not yet published | — |
| 2023-12-31 | 2024-02-22 | 1.41 | not yet published | — |
| 2023-09-30 | 2024-03-15 | 1.18 | not yet published | naive join |
| 2024-03-31 | 2024-05-02 | 1.05 | not yet published | — |
Synthetic, the note’s record. A date-only release is assumed to come after the close, so it is first usable on the next weekday.
Banz and Breen showed how much this matters: treating accounting figures as known at the fiscal year end, and using only the firms still on the current file, gave significantly different answers to standard questions from those of a point-in-time record.4 Kothari, Shanken and Sloan argued that part of the book-to-market effect in early studies came from the way firms entered the accounting database with their histories already attached;5 Chan, Jegadeesh and Lakonishok found that bias small in their sample.6 Either way, the question has to be asked of every dataset.
Survivorship and delisting returns
A list of stocks that trade today is a list of winners. Companies that went bankrupt, were delisted for low prices or were taken over have left it, and a backtest that starts from today’s list and looks backwards never holds them. Brown, Goetzmann, Ibbotson and Ross showed that truncating a sample on survival can make performance look persistent and predictable when it is neither,7 and Elton, Gruber and Blake measured the effect on mutual fund returns directly.8
Keeping dead companies is not the end of it. When a stock is delisted, its holders receive something: a takeover price, or the last trade on an over-the-counter market after a fall. Shumway found that the Center for Research in Security Prices (CRSP) file was missing most delisting returns for stocks delisted for poor performance, that for NYSE and AMEX stocks the missing returns averaged about −30%, and suggested using that figure when the true one is unknown;9 for Nasdaq stocks, Shumway and Warther put it nearer −55%.10 CRSP’s files carry delisting returns, with codes for why one is missing.11
- Step 1 of 4
Identical stocks
Our universe has 400 synthetic stocks. Each month every one of them earns the market’s return plus noise with a mean of zero, so none is built to beat another. An equal-weighted portfolio of every stock listed at the time, which is what an investor could have held, compounds at 6.16% a year over 20 years.
- Step 2 of 4
The ones that left
A stock is delisted when its price ends a month below 20% of where it started, and its holders receive −30% on the way out. 167 of the 400 go, the first in month 7 and 75 within five years. The point-in-time portfolio held each of them until it left, losses included.
- Step 3 of 4
Today’s list
A list of the stocks trading today holds the 233 survivors, none of which ever ended a month below the floor. Take their histories back 20 years, weight them equally, and the portfolio compounds at 10.93% a year.
- Step 4 of 4
The gap
The gap, 4.77 percentage points a year, comes from selection alone. Every stock had the same expected return; the survivors’ history leaves out the ones that fell through the floor, and their delisting losses with them.
Synthetic: 400 stocks over 20 years, each earning the market factor plus zero-mean noise, delisted when the price ends a month below 20% of its start, with a delisting return of −30%. Both lines are equal-weighted and rebalanced monthly. The laboratory below varies the floor, the delisting return and the length of history. Without script, or with reduced motion, the figure shows its final state.
Of 400 stocks, 233 are still listed after 20 years. A history built from today’s list compounds at 10.93% a year; the point-in-time universe, which is what an investor could have held, at 6.16%. The difference, 4.77 percentage points a year, is pure selection: every stock was built with the same expected return.
If the database drops delisting returns instead of recording them at −30%, the point-in-time universe compounds at 6.87%: 0.71 percentage points a year better than the truth.
The bias is worst for strategies that buy weakness. Buying each month the fifth of stocks with the worst past year earns 12.8% a year on the survivors and 3.0% on the point-in-time universe. Among survivors, no stock fell far enough to be delisted, so buying the fallen ones never meant buying a stock on its way to zero.
def listed(r: np.ndarray, floor: float, delist_return: float = None):
"""Returns as a point-in-time database records them, and who was listed when.
A stock whose price ends a month below `floor` is delisted at that month's end. Following
CRSP's convention, that month's recorded return compounds the month's trading return with
the delisting return (what the shares fetch after leaving the exchange). None means the
database has no delisting return: the month shows only the trading return, and the loss on
exit is missing (Shumway's problem with early CRSP data).
Returns (rec, alive, survivor): `alive[t, i]` is True if stock i is listed at the start of
month t; `rec[t, i]` is its return that month (NaN once it is gone); `survivor[i]` is True
if the stock is still listed at the end.
"""
months, n = r.shape
price = np.ones(n)
alive = np.zeros((months, n), dtype=bool)
rec = np.full((months, n), np.nan)
live = np.ones(n, dtype=bool)
for t in range(months):
alive[t] = live
rec[t, live] = r[t, live]
price[live] = price[live] * (1.0 + r[t, live])
out = live & (price < floor)
if delist_return is not None:
rec[t, out] = (1.0 + r[t, out]) * (1.0 + delist_return) - 1.0
live = live & ~out
return rec, alive, liveIdentifiers
A related error is quieter. Tickers change when companies rename or move exchange, and a ticker freed by a delisting is reassigned to another company. A history joined on ticker can splice two companies into one series or lose a company halfway. Key every series by a permanent identifier (CRSP’s PERMNO, or a vendor’s own), map tickers to it by date, and take the universe as it stood on each rebalancing date from a source that keeps dead securities and their delisting returns. The same logic applies to futures markets that were discontinued, funds that closed and indices whose rules changed.
Corporate actions and adjusted prices
Prices on the tape are quoted in the shares of the day. When a company splits its stock four for one, the price falls to a quarter and the holder owns four times as many shares. When it pays a dividend, the price drops by about the dividend on the ex-date, which fixes who is entitled; the cash arrives on the payment date, often weeks later. A raw price series treats both as losses. Adjusted series correct them by scaling every earlier price, and data vendors publish them in two common forms: split-adjusted, and total-return adjusted, which adjusts for dividends as well.
For a split of ratio k on day e, prices before e are divided by k. For a cash dividend D going ex on day e, prices before e are multiplied by
where Pe − 1 is the raw close before the ex-date. The cumulative factor for any day is the product over every event after it.
An adjusted history is therefore a function of the date it is read. Each new dividend or split lowers every price before it. In the synthetic stock below, the first close printed at 180.00; the total-return series read a year later shows 176.99, and read at the end shows 41.17. Returns computed from the series do not change; price levels do. A rule written in price levels (“buy below 100”, “stocks above 5 dollars”) tested on today’s adjusted history uses information from the future.
The adjusted return on an ex-date is Pe / (Pe − 1 − D) − 1, while the holder’s return is (Pe + D) / Pe − 1 − 1. They differ by a second-order term, about the dividend yield times the day’s return: in the synthetic stock by 0.6 bp on an average ex-date and at most 2.1 bp. Over five years the total-return series compounds to 57.90% against an exact 57.94%. For precise work, compute returns from raw prices and the cash flows, as CRSP does for its holding-period return. Both forms assume the gross dividend is reinvested; an investor who suffers withholding tax, or holds the stock in another currency, earns something else, and the backtest should say which.
Read on session 1,260, after 20 dividends and the split, the first close is 41.17 on the total-return series and 45.00 split-adjusted; it printed at 180.00. A rule that buys below 100 would buy on the first day of this history, although on the day itself the price was nowhere near 100.
def factors(raw, div, split, asof: int):
"""The cumulative adjustment factors as a vendor computes them on day `asof`: for each day,
the product over events AFTER it (and no later than `asof`) of 1/ratio for a split and
1 - D / previous close for a dividend. Events after `asof` have not happened yet."""
n = len(raw)
fs = np.ones(n) # splits only
fd = np.ones(n) # dividends only
for e in range(asof, 0, -1): # walk back from the reading date
fs[e - 1] = fs[e] / split[e]
fd[e - 1] = fd[e] * (1.0 - div[e] / raw[e - 1]) if div[e] > 0 else fd[e]
return fs, fdOn raw prices the split shows up as a one-day return of −74.4%; the holder actually made 2.3%. A 50-day moving-average rule run on raw and on adjusted prices disagrees on 64 sessions, 50 of them in the 60 sessions after the split, when the raw series sits far below its own average. Over the whole sample the raw price “returned” −63.9%, the split-adjusted price 44.5% and the holder 57.9%.
| Use | Series | Why |
|---|---|---|
| Returns, volatility, correlations, backtest profit and loss (P&L) | Total-return adjusted, or raw prices plus cash flows | The holder’s actual experience, with no jumps at corporate actions |
| Signals on price ratios (moving averages, momentum) | Total-return adjusted | Ratios survive a constant factor; they do not survive a split in the raw series |
| Order prices, limit and stop levels, share quantities | Raw, on the day | The exchange trades the shares of the day |
| Price filters, tick sizes, lot sizes, option strikes | Raw, on the day | Rules are set in the day’s prices |
| Liquidity (dollar volume, participation) | Raw price times raw volume | Split adjustment scales price and volume in opposite directions; keep them consistent |
| Price levels in any rule | Raw, or adjusted only for actions known at the time | Today’s adjusted history encodes later dividends and splits |
Split-adjusted series without dividends suit charts of price, not measures of return.
Backtesting engines make this choice for you unless you make it yourself. LEAN, for example, feeds split- and dividend-adjusted prices by default (its Adjusted mode) and lets each subscription choose another, including Raw and SplitAdjusted; a history request can also use ScaledRaw, which adjusts prices only for the events known at the algorithm’s current time. Its TotalReturn mode applies splits and then adds the accumulated dividends to the price instead of scaling earlier prices for them, so it is a different series from the one called total-return adjusted here.12,13 Know which one your engine uses before you read a fill price.
Continuous futures and the roll
A futures contract expires, so a history of “the” futures price has to be assembled from a chain of contracts. A trader holding the front contract sells it some days before expiry and buys the next: the roll. On the roll date the two contracts trade at different prices, and the difference is carry: financing and storage, less dividends or convenience yield, depending on the market. When later contracts are dearer the curve is in contango; when they are cheaper, in backwardation. Three constructions are common, and they disagree:
| Series | How it is built | Reproduces | Breaks |
|---|---|---|---|
| Unadjusted (spliced) | The held contract’s price; switches at each roll | The prices that actually traded | Returns and P&L: every roll is a jump that never traded |
| Difference-adjusted (“Panama”) | Each roll’s gap (new − old) added to all earlier prices | The P&L of holding one contract throughout, day by day | Percentage returns; old prices can go negative |
| Ratio-adjusted | Earlier prices multiplied by each roll’s ratio (new / old) | The return of a position whose notional is kept whole at each roll | P&L in points; old price levels |
LEAN calls these Raw, BackwardsPanamaCanal and BackwardsRatio; a forward-adjusted variant keeps the first contract’s prices instead of the last. Both identities hold only when the gap is measured at the prices actually traded on the roll, with no spread cost.
Over 15 years and 60 rolls, spot moved 5.5% while a fully collateralised position rolled through the contracts returned −49.9% before interest on its collateral: each contract drifts towards spot as it nears expiry, and in contango that drift is down. The ratio-adjusted series reproduces that return exactly and the difference-adjusted series the P&L of one contract; the unadjusted series reproduces neither, because every roll adds a jump that never traded.
| Series | P&L, points | Return | Low |
|---|---|---|---|
| The position itself (the truth) | −59.5 | −49.9% | — |
| Unadjusted (spliced) | 5.7 ✗ | 5.5% ✗ | 46.4 |
| Difference-adjusted | −59.5 ✓ | −35.5% ✗ | 83.1 |
| Ratio-adjusted | −107.6 ✗ | −49.9% ✓ | 82.6 |
The return excludes interest on the collateral, which in practice earns roughly the carry back. A tick means the series reproduces the position exactly, on every day, not just at the end.
Nothing is lost at the roll itself: one contract is exchanged for another at market prices. At the laboratory’s default carry of +5% a year, the loss in the readout above accrues day by day as each contract, priced above spot, drifts down to it at expiry. When the carry is the financing rate, as here, interest on the collateral earns back roughly what the roll-down costs: the collateralised position tracks spot plus interest minus carry, which here is about spot, while the continuous contract on its own, an excess return, tracks spot minus carry. The ratio-adjusted series gets the excess return right to the last digit. The difference-adjusted series, read as a percentage, says −35.5%, because in contango each gap is added to every earlier price, the early prices are inflated, and every percentage computed from them is too small. In backwardation the same construction goes the other way: at −5% the difference-adjusted history bottoms at −13.9, a negative price for an asset that never traded below 45.
def continuous(carry: float = 0.05, offset: int = 5, years: int = YEARS) -> dict:
"""The unadjusted, difference-adjusted and ratio-adjusted series, and the truth they are
judged against: a one-contract position's P&L, and the growth of a fully collateralised
position (its excess return: no interest on the collateral).
On a roll day the series still shows the old contract; the new one takes over the next day.
Each roll's gap (new minus old, or new over old, at the roll's close) is applied to that
day and every day before it."""
spot, today, prev, rolls = chain(carry, offset, years)
n = len(spot)
gap_add, gap_mul = np.zeros(n), np.ones(n)
for d, old, new in rolls:
gap_add[d], gap_mul[d] = new - old, new / old
add = np.cumsum(gap_add[::-1])[::-1] # the sum of the gaps on or after each day
mul = np.cumprod(gap_mul[::-1])[::-1]
pnl = np.r_[0.0, np.cumsum(today[1:] - prev[1:])]
growth = np.r_[1.0, np.cumprod(today[1:] / prev[1:])]
return {"spot": spot, "unadjusted": today, "difference": today + add, "ratio": today * mul,
"pnl": pnl, "growth": growth, "roll_days": [d for d, _, _ in rolls]}We use ratio-adjusted series for signals and returns, difference-adjusted series for P&L in points, and the raw contract’s price to trade and size. The roll rule belongs in the backtest: the date (a fixed number of sessions before expiry, before first notice day for physically delivered contracts, or when open interest moves to the next contract), the price at which both legs trade, and the cost of the spread trade. The daily “close” of a future is usually the exchange’s settlement price, set by its own procedure and published after the session, which can differ from the last trade. A continuous series built with one roll rule and a strategy traded with another will not reconcile.
Time zones and session calendars
Exchanges set their hours in local time, and local time moves against UTC when daylight saving starts and ends. The United States moves its clocks on the second Sunday of March and the first Sunday of November;14 the European Union, and the United Kingdom under its own law, on the last Sundays of March and October.15,16 For a few weeks a year New York is four hours behind London instead of five. In 2025, the London and New York sessions overlapped for two hours on most weekdays and for three hours from 10 to 28 March and from 27 to 31 October.
Store every timestamp in UTC with its source’s time zone recorded, and convert to local time only to apply an exchange’s rules. Use the IANA database through the language’s own library, never a fixed offset.17 Key daily data by the session it belongs to, not the calendar date of its timestamp: many futures sessions open on the evening before their trade date, and a daily bar stamped at midnight UTC can belong to either day. Keep an exchange calendar with holidays and shortened sessions, and treat a bar on a day the exchange was closed as an error.
def utc_hours(day: date, zone: str, local: time) -> float:
"""The UTC clock time, in hours, of a local time on `day`."""
t = datetime.combine(day, local, tzinfo=ZoneInfo(zone)).astimezone(UTC)
return t.hour + t.minute / 60Signals that combine markets need the same care. A close in Tokyo is known hours before the close in New York; a signal that uses both “closes of day t” is fine for a New York trade and a look-ahead for a Tokyo one. Write down, for every input, the UTC time at which it becomes available, and join on that.
Quality checks
Every delivery goes through the same checks before anything reads it. Each check is a small function, tested against the fault it is meant to catch. Two standard references on tick data, by Brownlees and Gallo and by Barndorff-Nielsen and colleagues, describe cleaning rules that follow the same logic at a finer scale: remove what is impossible, then compare each print with its neighbours.18,19
| Check | Catches | Blind spots and false alarms |
|---|---|---|
| Calendar | Missing sessions, bars on closed days, duplicated timestamps | Misses a bar with the right date and the wrong prices |
| OHLC consistency | High below the close, low above the open, zero or negative prices | Misses a bar that is consistent but wrong as a whole |
| Stale prints | A frozen feed: unchanged closes or no volume for several sessions | False alarm on an illiquid instrument that really did not trade |
| Spikes against a reference | Bad prints that reverse, unadjusted splits, unexplained jumps | Misses errors smaller than the instrument’s own noise; blind for its first weeks of history |
Every check has a blind spot, and the next check in the suite is chosen to cover it.
The spike check is the one that needs thought. A 20% fall is suspect in a quiet stock on a quiet day and unremarkable on a day the whole market fell 20%. So the check compares each return with a reference, here the market with a beta of one, and scales the residual by the median absolute deviation of recent residuals, which a single bad print cannot inflate. A flagged move that reverses the next session is a bad print. One that stays, at a split ratio relative to the market, is an unadjusted split if the corporate-actions file has one that day and a move for review if it does not. On the day a bar arrives its reversal cannot yet be seen, so a live check quarantines the bar and classifies it a session later; an earnings calendar removes most of the false alarms on single stocks.
def check_spikes(bars: pd.DataFrame, ref: pd.Series, actions: dict = None, z: float = 8.0, window: int = 63) -> pd.DataFrame:
"""Returns that the market does not explain.
The residual is the bar's return minus the reference's over the same interval (a beta of
one; use a rolling beta or a sector reference for stocks far from it). Its scale is robust:
1.4826 times the median absolute deviation of recent residuals from their rolling median,
so one bad print cannot widen it. A residual beyond `z` scales is flagged and classified:
- it reverses the next session (in log terms, by more than half): a bad print;
- it stays, and the move relative to the reference is a split ratio: an unadjusted split
if the corporate-actions file (`actions`, date -> ratio) has one that day, otherwise a
real move at a suspicious size, for review;
- anything else: a jump for a person to look at.
On the latest bar the reversal cannot be seen yet, so the flag says so and the bar stays in
quarantine until the next session. The first `window` sessions or so have no scale and are
not checked: a new listing needs a cross-sectional scale instead."""
actions = actions or {}
close = bars["close"].where(bars["close"] > 0) # a non-positive price is ohlc's to report
r = close.pct_change(fill_method=None)
rr = ref.reindex(close.index).pct_change(fill_method=None)
e = r - rr
m = window // 2
dev = (e - e.rolling(window, min_periods=m).median()).abs()
scale = 1.4826 * dev.rolling(window, min_periods=m).median().shift(1)
lr = np.log1p(r)
rows, skip = [], -1
for i in np.flatnonzero((e.abs() > z * scale).to_numpy()):
if i == skip: # the day a bad print reverted
continue
d = close.index[i]
rel = (1 + r.iloc[i]) / (1 + rr.iloc[i]) # the move the market does not explain
at_ratio = min(abs(np.log(rel / q)) for q in SPLIT_RATIOS) < 3 * scale.iloc[i]
if i + 1 >= len(e) or not np.isfinite(lr.iloc[i + 1]):
kind = "unconfirmed until the next session: quarantine"
elif np.sign(lr.iloc[i + 1]) == -np.sign(lr.iloc[i]) and abs(lr.iloc[i + 1]) > 0.5 * abs(lr.iloc[i]):
kind, skip = "spike that reverses: a bad print", i + 1
elif at_ratio and d in actions:
kind = "unadjusted split: one is on the corporate-actions file"
elif at_ratio:
kind = "move at a split ratio, none on file: review"
else:
kind = "unexplained jump: review"
rows.append((d, "spike", kind))
return pd.DataFrame(rows, columns=["date", "check", "detail"])
def run_checks(bars: pd.DataFrame, ref: pd.Series, sessions: pd.DatetimeIndex, actions: dict = None) -> pd.DataFrame:
"""The suite. Duplicates are reported by the calendar check and then dropped, keeping the
later delivery (a resent bar is usually a correction), so the price checks see one bar per
timestamp."""
once = bars[~bars.index.duplicated(keep="last")].sort_index()
out = [check_calendar(bars, sessions), check_ohlc(once), check_stale(once), check_spikes(once, ref, actions)]
return pd.concat(out, ignore_index=True).sort_values(["date", "check"], ignore_index=True)| Planted in the synthetic history | Session | Flagged by | Reason given |
|---|---|---|---|
| A session is missing | 2021-06-22 | calendar | missing session |
| A bar is duplicated | 2021-10-13 | calendar | duplicate timestamp |
| The feed freezes for four sessions | 2022-07-27 | stale | 4 sessions unchanged |
| One bar is ten times too high | 2022-12-14 | spike | spike that reverses: a bad print |
| A high below the close | 2023-03-10 | ohlc | inconsistent or non-positive |
| A zero close | 2023-05-19 | ohlc | inconsistent or non-positive |
| An unadjusted two-for-one split | 2023-07-18 | spike | unadjusted split: one is on the corporate-actions file |
| A Friday bar stamped Saturday | 2023-09-29 | calendar | bar on a non-session day; missing session |
| Control: a genuine 20% market crash | 2022-03-04 | nothing (correct) | — |
Synthetic: 767 sessions of one stock and a market reference, with one fault of each kind planted and one genuine 20% market crash as a control. On the clean history the suite raises no flags. Run without the market reference, the same spike check flags the crash as an unexplained jump.
Missing data in pandas
| Habit | What goes wrong |
|---|---|
| Forward-filling gaps | A stale price looks like a quiet day: volatility falls and correlations rise |
| Back-filling gaps | Look-ahead: a missing day takes tomorrow’s value |
| Resampling with default labels | A bar can be stamped with a time before the data in it existed; state the label and closed side |
| Aligning two series by an inner join | Days one market traded and the other did not disappear silently |
When vendors disagree
Two vendors rarely agree on every number, and the differences have causes that can be listed. Ince and Porter compared Datastream’s US equity returns with CRSP’s and found enough discrepancies, in coverage, in rounding and in the treatment of dead and foreign securities, to change the results of standard studies unless the data were screened first.20
| Question | Why vendors differ |
|---|---|
| What is the close? | Last trade anywhere, the primary exchange’s closing auction, a consolidated official price, or for futures the settlement price |
| What time is the bar? | Bars stamped at their start or their end; exchange time, the vendor’s time or UTC |
| Which sessions? | Regular hours only, or extended hours folded in; holidays and half-days handled differently |
| How is it adjusted? | Split only or total return; multiplicative or subtractive dividends; the ex-date or the pay date |
| Who is it? | Tickers change and are reused; vendors map them to permanent identifiers differently |
| What is missing? | Dead securities, delisting returns, halted days, and fields back-filled after the fact |
| When was it known? | Latest vintage only, or every revision with its publication time |
Historical bars change after the fact as well: late prints are added, cancelled trades removed, closes corrected. The history a backtest reads next month is not what the live system saw today. Keep what the live feed actually delivered in the raw store, and treat the vendor’s later history as a second source to reconcile against.
Before switching vendors, or adding one, reconcile the two on the overlapping history: returns day by day, corporate actions event by event, and the universe on each date. Explain every difference larger than rounding. The same exercise, run between a backtest and a live account, is the subject of Live versus backtest reconciliation.
With the data in order, the next step is to turn an idea into a rule and test it: Indicators and the 200-day moving average.
References
- Croushore, D., & Stark, T. (2001). A real-time data set for macroeconomists. Journal of Econometrics, 105(1), 111–130. https://doi.org/10.1016/s0304-4076(01)00072-0
- Federal Reserve Bank of St. Louis (n.d.). ALFRED: Archival Federal Reserve Economic Data. Federal Reserve Bank of St. Louis. alfred.stlouisfed.org/
- Fama, E. F., & French, K. R. (1992). The cross-section of expected stock returns. Journal of Finance, 47(2), 427–465. https://doi.org/10.1111/j.1540-6261.1992.tb04398.x
- Banz, R. W., & Breen, W. J. (1986). Sample-dependent results using accounting and market data: Some evidence. Journal of Finance, 41(4), 779–793. https://doi.org/10.1111/j.1540-6261.1986.tb04548.x
- Kothari, S. P., Shanken, J., & Sloan, R. G. (1995). Another look at the cross-section of expected stock returns. Journal of Finance, 50(1), 185–224. https://doi.org/10.1111/j.1540-6261.1995.tb05171.x
- Chan, L. K. C., Jegadeesh, N., & Lakonishok, J. (1995). Evaluating the performance of value versus glamour stocks: The impact of selection bias. Journal of Financial Economics, 38(3), 269–296. https://doi.org/10.1016/0304-405x(94)00818-l
- Brown, S. J., Goetzmann, W. N., Ibbotson, R. G., & Ross, S. A. (1992). Survivorship bias in performance studies. Review of Financial Studies, 5(4), 553–580. https://doi.org/10.1093/rfs/5.4.553
- Elton, E. J., Gruber, M. J., & Blake, C. R. (1996). Survivor bias and mutual fund performance. Review of Financial Studies, 9(4), 1097–1120. https://doi.org/10.1093/rfs/9.4.1097
- Shumway, T. (1997). The delisting bias in CRSP data. Journal of Finance, 52(1), 327–340. https://doi.org/10.1111/j.1540-6261.1997.tb03818.x
- Shumway, T., & Warther, V. A. (1999). The delisting bias in CRSP’s Nasdaq data and its implications for the size effect. Journal of Finance, 54(6), 2361–2379. https://doi.org/10.1111/0022-1082.00192
- Center for Research in Security Prices (2012). Data descriptions guide: CRSP US Stock and US Index Databases (delisting returns and cumulative adjustment factors). University of Chicago Booth School of Business. clouddc.chass.utoronto.ca/ds/crsp/en/manuals/data_descriptions_guide.pdf
- QuantConnect (n.d.). US equity: requesting data (data normalization modes). QuantConnect documentation. www.quantconnect.com/docs/v2/writing-algorithms/securities/asset-classes/us-equity/requesting-data
- QuantConnect (n.d.). LEAN source: DataNormalizationMode (Common/Global.cs). GitHub. github.com/QuantConnect/Lean/blob/master/Common/Global.cs
- National Institute of Standards and Technology (n.d.). Daylight saving time rules. NIST Time and Frequency Division. www.nist.gov/pml/time-and-frequency-division/popular-links/daylight-saving-time-dst
- European Parliament and Council of the European Union (2001). Directive 2000/84/EC of 19 January 2001 on summer-time arrangements. Official Journal of the European Communities. eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32000L0084
- United Kingdom (2002). The Summer Time Order 2002 (SI 2002/262). legislation.gov.uk. www.legislation.gov.uk/uksi/2002/262/made
- Internet Assigned Numbers Authority (n.d.). Time zone database. IANA. www.iana.org/time-zones
- Brownlees, C. T., & Gallo, G. M. (2006). Financial econometric analysis at ultra-high frequency: Data handling concerns. Computational Statistics & Data Analysis, 51(4), 2232–2245. https://doi.org/10.1016/j.csda.2006.09.030
- Barndorff-Nielsen, O. E., Hansen, P. R., Lunde, A., & Shephard, N. (2009). Realized kernels in practice: Trades and quotes. Econometrics Journal, 12(3), C1–C32. https://doi.org/10.1111/j.1368-423x.2008.00275.x
- Ince, O. S., & Porter, R. B. (2006). Individual equity return data from Thomson Datastream: Handle with care! Journal of Financial Research, 29(4), 463–479. https://doi.org/10.1111/j.1475-6803.2006.00189.x