Skip to content

Foundations · Intermediate

Data for systematic trading

The data errors that make a backtest look better than it is, and how to catch each one in code before a strategy reads it.

28 min read20 referencesIntroduction to the series

Educational material only. Not investment advice.

Contents
  1. 01From delivery to strategy
  2. 02Point-in-time data
  3. 03Survivorship and delisting returns
  4. 04Corporate actions and adjusted prices
  5. 05Continuous futures and the roll
  6. 06Time zones and session calendars
  7. 07Quality checks
  8. 08When vendors disagree
  9. —References

Abstract

Many errors in systematic research are data errors that look like results. On synthetic data where the true answer is known, we measure four of them: survivorship, adjusted prices read on the wrong date, futures contracts joined the wrong way, and fundamentals used before they were published. A survivors-only universe of identical stocks gains about five percentage points a year from selection alone. Each problem has a remedy that fits in a few lines of code, and the checks that enforce them are tested against planted faults.

Key takeaways

  • Store every fact with the time it became known and read it at that time; a naive join on the period it describes leaks the future.
  • A universe built from today’s list is survivors only; keep dead securities, their delisting returns and permanent identifiers.
  • Use total-return adjusted prices (or raw prices plus cash flows) for returns and signals, and raw prices of the day for orders, filters and anything in price levels.
  • Ratio-adjusted futures series reproduce a rolled position’s returns and difference-adjusted series one contract’s P&L; the unadjusted splice reproduces neither.
  • Keep timestamps in UTC, key daily data by session, and use the IANA database rather than fixed offsets.
  • Test every check against a planted fault, compare a suspect move with a reference before calling it an error, and reconcile vendors on overlapping history before switching.

Before you start

  • Market anomalies, for why a backtest flatters itself
  • pandas: joins, time series and time zones
  • What a futures contract and its expiry are

A backtest cannot tell a data error from a result. A universe that holds only the companies that survived, a price that falls 74% on the day a stock splits, an earnings figure used weeks before it was published: each produces returns that look like an edge. The remedies are old and unglamorous, and all of them can be written as code and tested.

From delivery to strategy

We treat market data as a pipeline with a record of what arrived, when, and in what form. Raw deliveries are kept exactly as received and never edited. Everything a strategy reads is derived from them by code: adjusted prices, continuous futures, calendars, and the checks that decide whether a day’s data may be used at all. When a number looks wrong, the raw record shows whether the vendor sent it or we made it.

Figure 1From vendor delivery to what a strategy reads
Data pipeline: vendor deliveries into an immutable raw store, quality checks, derivation with reference data, a point-in-time store, and one reader for research, backtest and live; failed checks go to quarantinepassfailVendorsprices, actions, factsRaw storeas received, stampedCheckscalendar, OHLC, stale, spikesQuarantinereviewed by handReference dataactions, contracts, calendarsDeriveadjust, roll, key by sessionPoint-in-time storevalue, period, publishedReadersresearch, backtest, live
Raw deliveries are immutable and stamped with their arrival time. They are checked, then derived into point-in-time series with the reference data (corporate actions, contract specifications, calendars); a day that fails a check goes to quarantine until someone reviews it.

A raw record is useful only if it can answer, months later, what the system knew and when. Ours holds, for every delivery:

FieldWhy
Vendor, dataset and fileWhich source to ask when a number is disputed
Received at (UTC)What was knowable at any moment; the basis of every point-in-time read
The date or time the vendor says the data describeWhich session a delivery belongs to, separately from when it arrived
A content hashProof that a later read saw the same bytes; duplicates detected for free
The delivery it supersedes, if anyCorrections and restatements kept as vintages, never overwritten

The same reader serves research, the backtest and the live system. When they read different copies of the data, the backtest and the live book disagree for reasons that have nothing to do with the strategy; QuantConnect and LEAN describes how one engine keeps that path shared.

Point-in-time data

A fact about a period has at least two timestamps: the period it describes and the moment it became known. Quarterly earnings for a quarter that ends in March are published weeks later and are sometimes restated months after that. Macroeconomic series are routinely re-estimated after publication. Croushore and Stark built a real-time data set for US macroeconomic variables because results computed from the latest figures can differ materially from what could have been computed at the time;1 the Federal Reserve Bank of St. Louis publishes every vintage of its series free of charge in ALFRED.2 Index membership, sector classifications and credit ratings change too, and a database that stores only the latest value rewrites history each time.

A point-in-time record keeps every vintage with its publication time, and a backtest asks it one question: what was known when the decision was made? Fama and French handled this for accounting data by matching fiscal years ending in calendar year t − 1 to returns from July of year t, a gap of at least six months that makes sure the figures were public.3 A lag rule of that kind is a safe default when publication dates are missing. Where they exist, use them, with a time of day: a figure released after the close cannot inform that day’s closing trade. Our listing assumes a date-only publication arrives after the close.

site_research/fieldnotes/pointintime.py
def known_on(facts: pd.DataFrame, day: pd.Timestamp) -> pd.DataFrame:
    """The record a strategy could use at the close of `day`: each period in its latest vintage.
    Publication dates carry no time of day, so a fact is assumed to arrive after the close and
    is first usable at the next session's close (strictly earlier date). With timestamps,
    compare them with the decision time instead."""
    seen = facts[facts["published"] < day].sort_values("published")
    return seen.groupby("period_end", as_index=False).last()


def point_in_time(days: pd.DatetimeIndex, facts: pd.DataFrame, period: str = None) -> pd.Series:
    """What a backtest may use at each day's close: the latest period's value as known then,
    or, with `period`, that period's value in the vintage known then."""
    out = []
    for d in days:
        k = known_on(facts, d).sort_values("period_end")
        if period is not None:
            k = k[k["period_end"] == pd.Timestamp(period)]
        out.append(k["value"].iloc[-1] if len(k) else float("nan"))
    return pd.Series(out, index=days)


def naive(days: pd.DatetimeIndex, facts: pd.DataFrame) -> pd.Series:
    """The look-ahead mistake: today's final vintage, joined on the period it describes."""
    final = facts.sort_values("published").groupby("period_end", as_index=False).last()
    left = pd.DataFrame({"day": days.astype(final["period_end"].dtype)})
    out = pd.merge_asof(left, final.sort_values("period_end"), left_on="day", right_on="period_end")
    return out.set_index("day")["value"]
A synthetic record: one quarter is restated after the next quarter’s release. Over 325 weekdays, the naive join (today’s final figures, keyed on the period) gives a different value from the point-in-time one on 177 of them, 54% of the sample: before each release it already knows the number, and after the restatement it has always known the restated one.
Figure 2What a backtest knew on each day

At the close of 16 October 2023 a backtest may use 1.25, the quarter ending 30 June 2023 as published on 3 August 2023. The naive join reads 1.18 for the quarter ending 30 September 2023, a figure nobody could see until 15 March 2024.

Point in time
1.25
quarter to Jun 2023, published 3 August 2023
Naive join
1.18
quarter to Sep 2023, published 15 March 2024
Weekdays that differ
177 of 325
54% of the sample
Earnings per share each weekday uses, two ways1.001.101.201.301.40this closeJul 2023Oct 2023Jan 2024Apr 2024
Point in time (as known at the close)Naive join (final figures, keyed on the period)A release (first weekday it can be used)Weekdays the two differ
The record as known at the close of 16 October 2023
Quarter toPublishedEPSAt this closeUsed by
2023-03-312023-05-041.10known—
2023-06-302023-08-031.25knownpoint in time
2023-09-302023-11-021.32not yet published—
2023-12-312024-02-221.41not yet published—
2023-09-302024-03-151.18not yet publishednaive join
2024-03-312024-05-021.05not yet published—

Synthetic, the note’s record. A date-only release is assumed to come after the close, so it is first usable on the next weekday.

Synthetic: the note’s earnings record, joined to every weekday from 3 April 2023 to 28 June 2024 by the two functions in the listing; the table under the chart lists it. Move the day to just before a release, then to any day from 3 November to 29 December 2023: the naive join already shows the quarter ending 30 September 2023 as restated (1.18), months before the restatement of 15 March 2024, when a backtest could only have known 1.32.

Banz and Breen showed how much this matters: treating accounting figures as known at the fiscal year end, and using only the firms still on the current file, gave significantly different answers to standard questions from those of a point-in-time record.4 Kothari, Shanken and Sloan argued that part of the book-to-market effect in early studies came from the way firms entered the accounting database with their histories already attached;5 Chan, Jegadeesh and Lakonishok found that bias small in their sample.6 Either way, the question has to be asked of every dataset.

Survivorship and delisting returns

A list of stocks that trade today is a list of winners. Companies that went bankrupt, were delisted for low prices or were taken over have left it, and a backtest that starts from today’s list and looks backwards never holds them. Brown, Goetzmann, Ibbotson and Ross showed that truncating a sample on survival can make performance look persistent and predictable when it is neither,7 and Elton, Gruber and Blake measured the effect on mutual fund returns directly.8

Keeping dead companies is not the end of it. When a stock is delisted, its holders receive something: a takeover price, or the last trade on an over-the-counter market after a fall. Shumway found that the Center for Research in Security Prices (CRSP) file was missing most delisting returns for stocks delisted for poor performance, that for NYSE and AMEX stocks the missing returns averaged about −30%, and suggested using that figure when the true one is unknown;9 for Nasdaq stocks, Shumway and Warther put it nearer −55%.10 CRSP’s files carry delisting returns, with codes for why one is missing.11

Step by stepHow a list of today’s stocks rewrites 20 years
233 of 400 still listedGrowth of 1, equal-weighted (log scale)125051015year 206.16%167 delisted10.93%+4.77percentagepoints a year
  1. Step 1 of 4

    Identical stocks

    Our universe has 400 synthetic stocks. Each month every one of them earns the market’s return plus noise with a mean of zero, so none is built to beat another. An equal-weighted portfolio of every stock listed at the time, which is what an investor could have held, compounds at 6.16% a year over 20 years.

  2. Step 2 of 4

    The ones that left

    A stock is delisted when its price ends a month below 20% of where it started, and its holders receive −30% on the way out. 167 of the 400 go, the first in month 7 and 75 within five years. The point-in-time portfolio held each of them until it left, losses included.

  3. Step 3 of 4

    Today’s list

    A list of the stocks trading today holds the 233 survivors, none of which ever ended a month below the floor. Take their histories back 20 years, weight them equally, and the portfolio compounds at 10.93% a year.

  4. Step 4 of 4

    The gap

    The gap, 4.77 percentage points a year, comes from selection alone. Every stock had the same expected return; the survivors’ history leaves out the ones that fell through the floor, and their delisting losses with them.

Synthetic: 400 stocks over 20 years, each earning the market factor plus zero-mean noise, delisted when the price ends a month below 20% of its start, with a delisting return of −30%. Both lines are equal-weighted and rebalanced monthly. The laboratory below varies the floor, the delisting return and the length of history. Without script, or with reduced motion, the figure shows its final state.

Figure 3Survivorship bias on a synthetic universe
Delisting return

Of 400 stocks, 233 are still listed after 20 years. A history built from today’s list compounds at 10.93% a year; the point-in-time universe, which is what an investor could have held, at 6.16%. The difference, 4.77 percentage points a year, is pure selection: every stock was built with the same expected return.

Still listed
233 of 400
167 delisted along the way
Point-in-time universe
6.16%
market factor 6.8%; each exit costs holders a further 30%
Survivors only
10.93%
the same stocks, chosen after the fact
Buying last year’s losers
12.8% vs 3.0%
survivors only, against point in time
Growth of 1, equal-weighted and rebalanced monthly (log scale)1.02.05.0year 0year 4year 8year 12year 16year 20
Survivors only (today’s constituents, back-filled)Point in time (every stock listed at the time, with delisting returns)
Synthetic: 400 stocks, each with the same expected return (the market factor plus zero-mean noise), delisted when the price ends a month below a fraction of where it started. The two lines are equal-weighted portfolios of the same stocks. Set the floor to 50% and compare the point-in-time line under −30% and “not recorded”; then shorten the history to five years.

If the database drops delisting returns instead of recording them at −30%, the point-in-time universe compounds at 6.87%: 0.71 percentage points a year better than the truth.

The bias is worst for strategies that buy weakness. Buying each month the fifth of stocks with the worst past year earns 12.8% a year on the survivors and 3.0% on the point-in-time universe. Among survivors, no stock fell far enough to be delisted, so buying the fallen ones never meant buying a stock on its way to zero.

site_research/fieldnotes/survivorship.py
def listed(r: np.ndarray, floor: float, delist_return: float = None):
    """Returns as a point-in-time database records them, and who was listed when.

    A stock whose price ends a month below `floor` is delisted at that month's end. Following
    CRSP's convention, that month's recorded return compounds the month's trading return with
    the delisting return (what the shares fetch after leaving the exchange). None means the
    database has no delisting return: the month shows only the trading return, and the loss on
    exit is missing (Shumway's problem with early CRSP data).

    Returns (rec, alive, survivor): `alive[t, i]` is True if stock i is listed at the start of
    month t; `rec[t, i]` is its return that month (NaN once it is gone); `survivor[i]` is True
    if the stock is still listed at the end.
    """
    months, n = r.shape
    price = np.ones(n)
    alive = np.zeros((months, n), dtype=bool)
    rec = np.full((months, n), np.nan)
    live = np.ones(n, dtype=bool)
    for t in range(months):
        alive[t] = live
        rec[t, live] = r[t, live]
        price[live] = price[live] * (1.0 + r[t, live])
        out = live & (price < floor)
        if delist_return is not None:
            rec[t, out] = (1.0 + r[t, out]) * (1.0 + delist_return) - 1.0
        live = live & ~out
    return rec, alive, live
The point-in-time record: who was listed each month, and what holders received when a stock left. The survivors-only history keeps only the stocks still listed at the end.

Identifiers

A related error is quieter. Tickers change when companies rename or move exchange, and a ticker freed by a delisting is reassigned to another company. A history joined on ticker can splice two companies into one series or lose a company halfway. Key every series by a permanent identifier (CRSP’s PERMNO, or a vendor’s own), map tickers to it by date, and take the universe as it stood on each rebalancing date from a source that keeps dead securities and their delisting returns. The same logic applies to futures markets that were discontinued, funds that closed and indices whose rules changed.

Corporate actions and adjusted prices

Prices on the tape are quoted in the shares of the day. When a company splits its stock four for one, the price falls to a quarter and the holder owns four times as many shares. When it pays a dividend, the price drops by about the dividend on the ex-date, which fixes who is entitled; the cash arrives on the payment date, often weeks later. A raw price series treats both as losses. Adjusted series correct them by scaling every earlier price, and data vendors publish them in two common forms: split-adjusted, and total-return adjusted, which adjusts for dividends as well.

For a split of ratio k on day e, prices before e are divided by k. For a cash dividend D going ex on day e, prices before e are multiplied by

fe  =  1 − D / Pe − 1
the multiplicative dividend factor of vendors’ adjusted-close series

where Pe − 1 is the raw close before the ex-date. The cumulative factor for any day is the product over every event after it.

An adjusted history is therefore a function of the date it is read. Each new dividend or split lowers every price before it. In the synthetic stock below, the first close printed at 180.00; the total-return series read a year later shows 176.99, and read at the end shows 41.17. Returns computed from the series do not change; price levels do. A rule written in price levels (“buy below 100”, “stocks above 5 dollars”) tested on today’s adjusted history uses information from the future.

The adjusted return on an ex-date is Pe / (Pe − 1 − D) − 1, while the holder’s return is (Pe + D) / Pe − 1 − 1. They differ by a second-order term, about the dividend yield times the day’s return: in the synthetic stock by 0.6 bp on an average ex-date and at most 2.1 bp. Over five years the total-return series compounds to 57.90% against an exact 57.94%. For precise work, compute returns from raw prices and the cash flows, as CRSP does for its holding-period return. Both forms assume the gross dividend is reinvested; an investor who suffers withholding tax, or holds the stock in another currency, earns something else, and the backtest should say which.

Figure 4Raw, split-adjusted and total-return prices

Read on session 1,260, after 20 dividends and the split, the first close is 41.17 on the total-return series and 45.00 split-adjusted; it printed at 180.00. A rule that buys below 100 would buy on the first day of this history, although on the day itself the price was nowhere near 100.

Split day, raw prices
−74.4%
4-for-1: the price falls, the holder does not
Split day, total-return adjusted
2.3%
the holder’s actual return that day
First close, read on session 1,260
41.17
printed at 180.00
50-day rule disagrees
64 sessions
raw against total-return prices
One stock, three price histories, as read on the chosen session0501001502002504-for-1 splitread hereyear 0year 1year 2year 3year 4year 5
Raw close, in the shares of the daySplit-adjustedTotal-return adjusted (splits and dividends)
Synthetic: five years of a stock that pays a quarterly dividend and splits four for one. Move the reading date back before the split: the whole adjusted history rises about fourfold, while returns between any two days stay the same. Then read the history on session 700, the day before the split.
site_research/fieldnotes/adjust.py
def factors(raw, div, split, asof: int):
    """The cumulative adjustment factors as a vendor computes them on day `asof`: for each day,
    the product over events AFTER it (and no later than `asof`) of 1/ratio for a split and
    1 - D / previous close for a dividend. Events after `asof` have not happened yet."""
    n = len(raw)
    fs = np.ones(n)            # splits only
    fd = np.ones(n)            # dividends only
    for e in range(asof, 0, -1):          # walk back from the reading date
        fs[e - 1] = fs[e] / split[e]
        fd[e - 1] = fd[e] * (1.0 - div[e] / raw[e - 1]) if div[e] > 0 else fd[e]
    return fs, fd

On raw prices the split shows up as a one-day return of −74.4%; the holder actually made 2.3%. A 50-day moving-average rule run on raw and on adjusted prices disagrees on 64 sessions, 50 of them in the 60 sessions after the split, when the raw series sits far below its own average. Over the whole sample the raw price “returned” −63.9%, the split-adjusted price 44.5% and the holder 57.9%.

UseSeriesWhy
Returns, volatility, correlations, backtest profit and loss (P&L)Total-return adjusted, or raw prices plus cash flowsThe holder’s actual experience, with no jumps at corporate actions
Signals on price ratios (moving averages, momentum)Total-return adjustedRatios survive a constant factor; they do not survive a split in the raw series
Order prices, limit and stop levels, share quantitiesRaw, on the dayThe exchange trades the shares of the day
Price filters, tick sizes, lot sizes, option strikesRaw, on the dayRules are set in the day’s prices
Liquidity (dollar volume, participation)Raw price times raw volumeSplit adjustment scales price and volume in opposite directions; keep them consistent
Price levels in any ruleRaw, or adjusted only for actions known at the timeToday’s adjusted history encodes later dividends and splits

Split-adjusted series without dividends suit charts of price, not measures of return.

Backtesting engines make this choice for you unless you make it yourself. LEAN, for example, feeds split- and dividend-adjusted prices by default (its Adjusted mode) and lets each subscription choose another, including Raw and SplitAdjusted; a history request can also use ScaledRaw, which adjusts prices only for the events known at the algorithm’s current time. Its TotalReturn mode applies splits and then adds the accumulated dividends to the price instead of scaling earlier prices for them, so it is a different series from the one called total-return adjusted here.12,13 Know which one your engine uses before you read a fill price.

Continuous futures and the roll

A futures contract expires, so a history of “the” futures price has to be assembled from a chain of contracts. A trader holding the front contract sells it some days before expiry and buys the next: the roll. On the roll date the two contracts trade at different prices, and the difference is carry: financing and storage, less dividends or convenience yield, depending on the market. When later contracts are dearer the curve is in contango; when they are cheaper, in backwardation. Three constructions are common, and they disagree:

SeriesHow it is builtReproducesBreaks
Unadjusted (spliced)The held contract’s price; switches at each rollThe prices that actually tradedReturns and P&L: every roll is a jump that never traded
Difference-adjusted (“Panama”)Each roll’s gap (new − old) added to all earlier pricesThe P&L of holding one contract throughout, day by dayPercentage returns; old prices can go negative
Ratio-adjustedEarlier prices multiplied by each roll’s ratio (new / old)The return of a position whose notional is kept whole at each rollP&L in points; old price levels

LEAN calls these Raw, BackwardsPanamaCanal and BackwardsRatio; a forward-adjusted variant keeps the first contract’s prices instead of the last. Both identities hold only when the gap is measured at the prices actually traded on the roll, with no spread cost.

Figure 5Three continuous series from one contract chain

Over 15 years and 60 rolls, spot moved 5.5% while a fully collateralised position rolled through the contracts returned −49.9% before interest on its collateral: each contract drifts towards spot as it nears expiry, and in contango that drift is down. The ratio-adjusted series reproduces that return exactly and the difference-adjusted series the P&L of one contract; the unadjusted series reproduces neither, because every roll adds a jump that never traded.

Spot and three ways to join the same contracts050100150200year 0year 3year 6year 9year 12year 15
Spot, close to the unadjusted lineUnadjusted (spliced)Difference-adjustedRatio-adjusted
SeriesP&L, pointsReturnLow
The position itself (the truth)−59.5−49.9%—
Unadjusted (spliced)5.7 ✗5.5% ✗46.4
Difference-adjusted−59.5 ✓−35.5% ✗83.1
Ratio-adjusted−107.6 ✗−49.9% ✓82.6

The return excludes interest on the collateral, which in practice earns roughly the carry back. A tick means the series reproduces the position exactly, on every day, not just at the end.

Synthetic: a driftless spot price and quarterly contracts priced by cost of carry, rolled a set number of sessions before expiry. Set the carry to −5% and the difference-adjusted series crosses zero; set it to 0 and all three coincide with spot.

Nothing is lost at the roll itself: one contract is exchanged for another at market prices. At the laboratory’s default carry of +5% a year, the loss in the readout above accrues day by day as each contract, priced above spot, drifts down to it at expiry. When the carry is the financing rate, as here, interest on the collateral earns back roughly what the roll-down costs: the collateralised position tracks spot plus interest minus carry, which here is about spot, while the continuous contract on its own, an excess return, tracks spot minus carry. The ratio-adjusted series gets the excess return right to the last digit. The difference-adjusted series, read as a percentage, says −35.5%, because in contango each gap is added to every earlier price, the early prices are inflated, and every percentage computed from them is too small. In backwardation the same construction goes the other way: at −5% the difference-adjusted history bottoms at −13.9, a negative price for an asset that never traded below 45.

site_research/fieldnotes/rolls.py
def continuous(carry: float = 0.05, offset: int = 5, years: int = YEARS) -> dict:
    """The unadjusted, difference-adjusted and ratio-adjusted series, and the truth they are
    judged against: a one-contract position's P&L, and the growth of a fully collateralised
    position (its excess return: no interest on the collateral).

    On a roll day the series still shows the old contract; the new one takes over the next day.
    Each roll's gap (new minus old, or new over old, at the roll's close) is applied to that
    day and every day before it."""
    spot, today, prev, rolls = chain(carry, offset, years)
    n = len(spot)
    gap_add, gap_mul = np.zeros(n), np.ones(n)
    for d, old, new in rolls:
        gap_add[d], gap_mul[d] = new - old, new / old
    add = np.cumsum(gap_add[::-1])[::-1]          # the sum of the gaps on or after each day
    mul = np.cumprod(gap_mul[::-1])[::-1]
    pnl = np.r_[0.0, np.cumsum(today[1:] - prev[1:])]
    growth = np.r_[1.0, np.cumprod(today[1:] / prev[1:])]
    return {"spot": spot, "unadjusted": today, "difference": today + add, "ratio": today * mul,
            "pnl": pnl, "growth": growth, "roll_days": [d for d, _, _ in rolls]}
The continuous series and the two truths they are judged against. `chain` in the same file prices the contracts and records each roll.

We use ratio-adjusted series for signals and returns, difference-adjusted series for P&L in points, and the raw contract’s price to trade and size. The roll rule belongs in the backtest: the date (a fixed number of sessions before expiry, before first notice day for physically delivered contracts, or when open interest moves to the next contract), the price at which both legs trade, and the cost of the spread trade. The daily “close” of a future is usually the exchange’s settlement price, set by its own procedure and published after the session, which can differ from the last trade. A continuous series built with one roll rule and a strategy traded with another will not reconcile.

Time zones and session calendars

Exchanges set their hours in local time, and local time moves against UTC when daylight saving starts and ends. The United States moves its clocks on the second Sunday of March and the first Sunday of November;14 the European Union, and the United Kingdom under its own law, on the last Sundays of March and October.15,16 For a few weeks a year New York is four hours behind London instead of five. In 2025, the London and New York sessions overlapped for two hours on most weekdays and for three hours from 10 to 28 March and from 27 to 31 October.

Figure 6Three exchanges in UTC through one year
Exchange sessions in UTC through the year2025, weekdaysJanAprJulOct00:00 UTC04:0008:0012:0016:0020:0024:00Both on standard timeTokyoLondonNew Yorkoverlap 2 hNew York on summer time, London notTokyoLondonNew Yorkoverlap 3 hBoth on summer timeTokyoLondonNew Yorkoverlap 2 h
Regular sessions of the Tokyo (09:00–15:30), London (08:00–16:30) and New York (09:30–16:00) stock exchanges in local time, converted to UTC for 2025 with the IANA time zone database. Tokyo does not observe daylight saving, and its midday break is not drawn. The strip colours each weekday by its regime; holidays are not removed.

Store every timestamp in UTC with its source’s time zone recorded, and convert to local time only to apply an exchange’s rules. Use the IANA database through the language’s own library, never a fixed offset.17 Key daily data by the session it belongs to, not the calendar date of its timestamp: many futures sessions open on the evening before their trade date, and a daily bar stamped at midnight UTC can belong to either day. Keep an exchange calendar with holidays and shortened sessions, and treat a bar on a day the exchange was closed as an error.

site_research/fieldnotes/sessions.py
def utc_hours(day: date, zone: str, local: time) -> float:
    """The UTC clock time, in hours, of a local time on `day`."""
    t = datetime.combine(day, local, tzinfo=ZoneInfo(zone)).astimezone(UTC)
    return t.hour + t.minute / 60
A local exchange time on a given date, in UTC. It returns the clock time only, which is enough for sessions that do not cross midnight UTC.

Signals that combine markets need the same care. A close in Tokyo is known hours before the close in New York; a signal that uses both “closes of day t” is fine for a New York trade and a look-ahead for a Tokyo one. Write down, for every input, the UTC time at which it becomes available, and join on that.

Quality checks

Every delivery goes through the same checks before anything reads it. Each check is a small function, tested against the fault it is meant to catch. Two standard references on tick data, by Brownlees and Gallo and by Barndorff-Nielsen and colleagues, describe cleaning rules that follow the same logic at a finer scale: remove what is impossible, then compare each print with its neighbours.18,19

CheckCatchesBlind spots and false alarms
CalendarMissing sessions, bars on closed days, duplicated timestampsMisses a bar with the right date and the wrong prices
OHLC consistencyHigh below the close, low above the open, zero or negative pricesMisses a bar that is consistent but wrong as a whole
Stale printsA frozen feed: unchanged closes or no volume for several sessionsFalse alarm on an illiquid instrument that really did not trade
Spikes against a referenceBad prints that reverse, unadjusted splits, unexplained jumpsMisses errors smaller than the instrument’s own noise; blind for its first weeks of history

Every check has a blind spot, and the next check in the suite is chosen to cover it.

The spike check is the one that needs thought. A 20% fall is suspect in a quiet stock on a quiet day and unremarkable on a day the whole market fell 20%. So the check compares each return with a reference, here the market with a beta of one, and scales the residual by the median absolute deviation of recent residuals, which a single bad print cannot inflate. A flagged move that reverses the next session is a bad print. One that stays, at a split ratio relative to the market, is an unadjusted split if the corporate-actions file has one that day and a move for review if it does not. On the day a bar arrives its reversal cannot yet be seen, so a live check quarantines the bar and classifies it a session later; an earnings calendar removes most of the false alarms on single stocks.

site_research/fieldnotes/quality.py
def check_spikes(bars: pd.DataFrame, ref: pd.Series, actions: dict = None, z: float = 8.0, window: int = 63) -> pd.DataFrame:
    """Returns that the market does not explain.

    The residual is the bar's return minus the reference's over the same interval (a beta of
    one; use a rolling beta or a sector reference for stocks far from it). Its scale is robust:
    1.4826 times the median absolute deviation of recent residuals from their rolling median,
    so one bad print cannot widen it. A residual beyond `z` scales is flagged and classified:
      - it reverses the next session (in log terms, by more than half): a bad print;
      - it stays, and the move relative to the reference is a split ratio: an unadjusted split
        if the corporate-actions file (`actions`, date -> ratio) has one that day, otherwise a
        real move at a suspicious size, for review;
      - anything else: a jump for a person to look at.
    On the latest bar the reversal cannot be seen yet, so the flag says so and the bar stays in
    quarantine until the next session. The first `window` sessions or so have no scale and are
    not checked: a new listing needs a cross-sectional scale instead."""
    actions = actions or {}
    close = bars["close"].where(bars["close"] > 0)        # a non-positive price is ohlc's to report
    r = close.pct_change(fill_method=None)
    rr = ref.reindex(close.index).pct_change(fill_method=None)
    e = r - rr
    m = window // 2
    dev = (e - e.rolling(window, min_periods=m).median()).abs()
    scale = 1.4826 * dev.rolling(window, min_periods=m).median().shift(1)
    lr = np.log1p(r)
    rows, skip = [], -1
    for i in np.flatnonzero((e.abs() > z * scale).to_numpy()):
        if i == skip:                                      # the day a bad print reverted
            continue
        d = close.index[i]
        rel = (1 + r.iloc[i]) / (1 + rr.iloc[i])           # the move the market does not explain
        at_ratio = min(abs(np.log(rel / q)) for q in SPLIT_RATIOS) < 3 * scale.iloc[i]
        if i + 1 >= len(e) or not np.isfinite(lr.iloc[i + 1]):
            kind = "unconfirmed until the next session: quarantine"
        elif np.sign(lr.iloc[i + 1]) == -np.sign(lr.iloc[i]) and abs(lr.iloc[i + 1]) > 0.5 * abs(lr.iloc[i]):
            kind, skip = "spike that reverses: a bad print", i + 1
        elif at_ratio and d in actions:
            kind = "unadjusted split: one is on the corporate-actions file"
        elif at_ratio:
            kind = "move at a split ratio, none on file: review"
        else:
            kind = "unexplained jump: review"
        rows.append((d, "spike", kind))
    return pd.DataFrame(rows, columns=["date", "check", "detail"])


def run_checks(bars: pd.DataFrame, ref: pd.Series, sessions: pd.DatetimeIndex, actions: dict = None) -> pd.DataFrame:
    """The suite. Duplicates are reported by the calendar check and then dropped, keeping the
    later delivery (a resent bar is usually a correction), so the price checks see one bar per
    timestamp."""
    once = bars[~bars.index.duplicated(keep="last")].sort_index()
    out = [check_calendar(bars, sessions), check_ohlc(once), check_stale(once), check_spikes(once, ref, actions)]
    return pd.concat(out, ignore_index=True).sort_values(["date", "check"], ignore_index=True)
The spike check and the suite. SPLIT_RATIOS and the other three checks are in the same file.
Planted in the synthetic historySessionFlagged byReason given
A session is missing2021-06-22calendarmissing session
A bar is duplicated2021-10-13calendarduplicate timestamp
The feed freezes for four sessions2022-07-27stale4 sessions unchanged
One bar is ten times too high2022-12-14spikespike that reverses: a bad print
A high below the close2023-03-10ohlcinconsistent or non-positive
A zero close2023-05-19ohlcinconsistent or non-positive
An unadjusted two-for-one split2023-07-18spikeunadjusted split: one is on the corporate-actions file
A Friday bar stamped Saturday2023-09-29calendarbar on a non-session day; missing session
Control: a genuine 20% market crash2022-03-04nothing (correct)—

Synthetic: 767 sessions of one stock and a market reference, with one fault of each kind planted and one genuine 20% market crash as a control. On the clean history the suite raises no flags. Run without the market reference, the same spike check flags the crash as an unexplained jump.

Missing data in pandas

HabitWhat goes wrong
Forward-filling gapsA stale price looks like a quiet day: volatility falls and correlations rise
Back-filling gapsLook-ahead: a missing day takes tomorrow’s value
Resampling with default labelsA bar can be stamped with a time before the data in it existed; state the label and closed side
Aligning two series by an inner joinDays one market traded and the other did not disappear silently

When vendors disagree

Two vendors rarely agree on every number, and the differences have causes that can be listed. Ince and Porter compared Datastream’s US equity returns with CRSP’s and found enough discrepancies, in coverage, in rounding and in the treatment of dead and foreign securities, to change the results of standard studies unless the data were screened first.20

QuestionWhy vendors differ
What is the close?Last trade anywhere, the primary exchange’s closing auction, a consolidated official price, or for futures the settlement price
What time is the bar?Bars stamped at their start or their end; exchange time, the vendor’s time or UTC
Which sessions?Regular hours only, or extended hours folded in; holidays and half-days handled differently
How is it adjusted?Split only or total return; multiplicative or subtractive dividends; the ex-date or the pay date
Who is it?Tickers change and are reused; vendors map them to permanent identifiers differently
What is missing?Dead securities, delisting returns, halted days, and fields back-filled after the fact
When was it known?Latest vintage only, or every revision with its publication time

Historical bars change after the fact as well: late prints are added, cancelled trades removed, closes corrected. The history a backtest reads next month is not what the live system saw today. Keep what the live feed actually delivered in the raw store, and treat the vendor’s later history as a second source to reconcile against.

Before switching vendors, or adding one, reconcile the two on the overlapping history: returns day by day, corporate actions event by event, and the universe on each date. Explain every difference larger than rounding. The same exercise, run between a backtest and a live account, is the subject of Live versus backtest reconciliation.

Survivorship, synthetic
+4.77
percentage points a year, from selection alone
Split on raw prices
−74.4%
the holder made 2.3%
Contango, synthetic
−49.9%
rolled position before interest; spot 5.5%
Naive join, synthetic
54%
of weekdays use a figure not yet published

With the data in order, the next step is to turn an idea into a rule and test it: Indicators and the 200-day moving average.

References

  1. Croushore, D., & Stark, T. (2001). A real-time data set for macroeconomists. Journal of Econometrics, 105(1), 111–130. https://doi.org/10.1016/s0304-4076(01)00072-0
  2. Federal Reserve Bank of St. Louis (n.d.). ALFRED: Archival Federal Reserve Economic Data. Federal Reserve Bank of St. Louis. alfred.stlouisfed.org/
  3. Fama, E. F., & French, K. R. (1992). The cross-section of expected stock returns. Journal of Finance, 47(2), 427–465. https://doi.org/10.1111/j.1540-6261.1992.tb04398.x
  4. Banz, R. W., & Breen, W. J. (1986). Sample-dependent results using accounting and market data: Some evidence. Journal of Finance, 41(4), 779–793. https://doi.org/10.1111/j.1540-6261.1986.tb04548.x
  5. Kothari, S. P., Shanken, J., & Sloan, R. G. (1995). Another look at the cross-section of expected stock returns. Journal of Finance, 50(1), 185–224. https://doi.org/10.1111/j.1540-6261.1995.tb05171.x
  6. Chan, L. K. C., Jegadeesh, N., & Lakonishok, J. (1995). Evaluating the performance of value versus glamour stocks: The impact of selection bias. Journal of Financial Economics, 38(3), 269–296. https://doi.org/10.1016/0304-405x(94)00818-l
  7. Brown, S. J., Goetzmann, W. N., Ibbotson, R. G., & Ross, S. A. (1992). Survivorship bias in performance studies. Review of Financial Studies, 5(4), 553–580. https://doi.org/10.1093/rfs/5.4.553
  8. Elton, E. J., Gruber, M. J., & Blake, C. R. (1996). Survivor bias and mutual fund performance. Review of Financial Studies, 9(4), 1097–1120. https://doi.org/10.1093/rfs/9.4.1097
  9. Shumway, T. (1997). The delisting bias in CRSP data. Journal of Finance, 52(1), 327–340. https://doi.org/10.1111/j.1540-6261.1997.tb03818.x
  10. Shumway, T., & Warther, V. A. (1999). The delisting bias in CRSP’s Nasdaq data and its implications for the size effect. Journal of Finance, 54(6), 2361–2379. https://doi.org/10.1111/0022-1082.00192
  11. Center for Research in Security Prices (2012). Data descriptions guide: CRSP US Stock and US Index Databases (delisting returns and cumulative adjustment factors). University of Chicago Booth School of Business. clouddc.chass.utoronto.ca/ds/crsp/en/manuals/data_descriptions_guide.pdf
  12. QuantConnect (n.d.). US equity: requesting data (data normalization modes). QuantConnect documentation. www.quantconnect.com/docs/v2/writing-algorithms/securities/asset-classes/us-equity/requesting-data
  13. QuantConnect (n.d.). LEAN source: DataNormalizationMode (Common/Global.cs). GitHub. github.com/QuantConnect/Lean/blob/master/Common/Global.cs
  14. National Institute of Standards and Technology (n.d.). Daylight saving time rules. NIST Time and Frequency Division. www.nist.gov/pml/time-and-frequency-division/popular-links/daylight-saving-time-dst
  15. European Parliament and Council of the European Union (2001). Directive 2000/84/EC of 19 January 2001 on summer-time arrangements. Official Journal of the European Communities. eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32000L0084
  16. United Kingdom (2002). The Summer Time Order 2002 (SI 2002/262). legislation.gov.uk. www.legislation.gov.uk/uksi/2002/262/made
  17. Internet Assigned Numbers Authority (n.d.). Time zone database. IANA. www.iana.org/time-zones
  18. Brownlees, C. T., & Gallo, G. M. (2006). Financial econometric analysis at ultra-high frequency: Data handling concerns. Computational Statistics & Data Analysis, 51(4), 2232–2245. https://doi.org/10.1016/j.csda.2006.09.030
  19. Barndorff-Nielsen, O. E., Hansen, P. R., Lunde, A., & Shephard, N. (2009). Realized kernels in practice: Trades and quotes. Econometrics Journal, 12(3), C1–C32. https://doi.org/10.1111/j.1368-423x.2008.00275.x
  20. Ince, O. S., & Porter, R. B. (2006). Individual equity return data from Thomson Datastream: Handle with care! Journal of Financial Research, 29(4), 463–479. https://doi.org/10.1111/j.1475-6803.2006.00189.x