Field note No. 01 · Foundations
Market anomalies
What they are, why some of them persist, how they have held up, and how not to fool yourself when looking for new ones.
Educational material only. Not investment advice.
Contents
Abstract
A working map of the best-documented return anomalies (momentum, time-series momentum, value, profitability, low volatility, carry, seasonality, post-earnings drift and short-term reversal); the risk-based, behavioural and limits-to-arbitrage explanations for them; the evidence on decay after publication and on replication; and the multiple-testing problem that makes many newly discovered anomalies illusory.
Key takeaways
- An anomaly is an alpha relative to a risk model, so its explanation (risk, behaviour or limits to arbitrage) indicates how it is likely to fail.
- A small number of families (momentum, trend, value, profitability, low volatility, carry, seasonality, post-earnings drift and reversal) account for most of the credible evidence.
- Published effects shrink: in one large study, by about a quarter out of sample and by more than half after publication. Costs take a further share.
- Testing many ideas all but guarantees an impressive-looking best result. Record every trial and correct for the number.
- A small operator's advantages are patience, capacity and discipline, not novelty.
Before you start
- Means, standard deviations and the t-statistic
- Long-short portfolios, factor regressions and backtests
- The Sharpe ratio (excess return per unit of volatility)
Every systematic strategy starts from a claim that some pattern in prices will keep repeating. The academic name for such a pattern is an anomaly: an average return that a standard model of risk does not explain. This note is a working map of the best-documented anomalies, of the arguments about why they exist, of the evidence on how they have held up, and of the statistical traps that await anyone who goes looking for new ones.
What an anomaly is, and is not
In an efficient market, in the sense Fama set out in 1970, prices already reflect the available information, so no rule built from that information should earn more than the compensation for the risk it carries.1 An anomaly is therefore always defined relative to a model. In practice it is measured as the intercept, or alpha, of a regression of a strategy’s returns on the factors of a risk model, and judged by that intercept’s t-statistic. A strategy that beats the market by taking more market risk is not an anomaly; one whose alpha is large and statistically reliable is.
That framing has a consequence that is easy to miss: every anomaly is a joint test. When a pattern survives, either markets are not fully efficient or the risk model is incomplete, and the data alone cannot say which. The distinction matters to a practitioner because it indicates how the pattern is likely to fail. A premium paid for bearing a real risk will occasionally produce large losses at the worst possible moment. A premium created by other investors’ errors can disappear once enough of them stop making the errors.
It also helps to separate two kinds of claim. A cross-sectional anomaly says that, on a given date, some assets will outperform others: buy the top of a ranking, sell the bottom, and the spread is positive on average. A time-series anomaly says that an asset’s own history predicts its next move. The first is naturally long-short and close to market-neutral; the second is naturally directional. They behave differently in a crisis and require different engineering.
Why a pattern can persist
If a pattern is known and profitable, why does competition not remove it? There are three families of explanation, and most documented anomalies are attributed to a mixture of them.
Risk: the return is compensation
The risk-based view holds that the extra return compensates for exposure to something investors genuinely dislike, typically losses that arrive when they are already poorer, more constrained or forced to sell. Fama and French argued that the returns of small and cheap stocks behave like exposures to common risk factors, and built their three-factor model on that interpretation.2,3 If this view is right, the premium should be durable, and the strategy should lose money in precisely the states of the world in which its holders can least afford it.
Behaviour: the return is someone else’s error
The behavioural view holds that prices deviate from fundamental value because investors process information imperfectly: they update too slowly on news, or they extrapolate a run of good results too far. Barberis, Shleifer and Vishny built a model in which the same investors underreact to individual pieces of news and overreact to long streaks, producing short-term continuation and long-term reversal.4 Lakonishok, Shleifer and Vishny interpreted the value premium in this way: investors overpay for firms with a history of growth and underpay for those with a history of disappointment.5
Limits to arbitrage: the error is expensive to correct
A mispricing disappears only if someone with capital trades against it. Shleifer and Vishny observed that real arbitrageurs are specialised, manage other people’s money and are judged on interim results; when a position moves against them before it converges, their investors may withdraw capital at exactly the wrong moment.6 Add trading costs, short-sale constraints, benchmark tracking and career risk, and a pattern can persist for decades after it is published. Baker, Bradley and Wurgler argue that benchmark-relative mandates are a central reason professional managers do not arbitrage away the low-volatility anomaly.7
The principal families
Hundreds of anomalies have been published, but most are variations on a small number of ideas. The families below have the longest record, the widest replication and the clearest economic rationale.
Cross-sectional momentum
Jegadeesh and Titman showed that US stocks with the highest returns over the previous three to twelve months continued to outperform those with the lowest returns over the following months, and that the effect was not explained by market risk.8 Fama and French found that their three-factor model could not account for it either,3 and Carhart added a momentum factor in his study of mutual-fund performance.9 Momentum appears across asset classes and countries, and Asness, Moskowitz and Pedersen found it to be negatively correlated with value, which makes the two complementary.10 Its weakness is well documented: Daniel and Moskowitz show that momentum suffers infrequent but severe crashes, concentrated in market rebounds that follow a decline, while volatility is high and the short side (the previous year’s losers) rallies hardest.11 Scaling momentum by its own recent volatility removes much of that crash risk,12 as Field note 02 (Volatility targeting) shows on the same factor. Implementations usually skip the most recent month because of short-term reversal, discussed below.
Time-series momentum (trend)
Moskowitz, Ooi and Pedersen studied several dozen liquid futures and forward contracts on equity indices, currencies, commodities and government bonds, and found that an instrument’s own excess return over the past twelve months predicted its return over the next month, with partial reversal at longer lags.13 Hurst, Ooi and Pedersen extended the evidence back more than a century.14 Trend is the core of most managed-futures programmes. It has tended to perform well in prolonged market declines, which is much of its appeal, and poorly in range-bound markets where signals reverse before they pay (Field note 03 (Indicators, and the 200-day moving average) examines the simplest trend rule on a century of daily data); its returns were notably weaker in the decade after the 2008 crisis than in the decades before.14
Value
Stocks that are cheap relative to fundamentals, measured by book-to-market, earnings or cash-flow yield, have historically outperformed expensive ones. Fama and French documented the effect in the cross-section of US returns and made it one of their factors.15,2 Its explanation is the classic contest between the risk view (cheap firms are distressed and fail together) and the behavioural view (investors extrapolate growth).5 Value can underperform for many years in succession, as it did through much of the 2010s; that is the single most important fact to know before holding it.
Profitability and quality
Novy-Marx showed that firms with high gross profits relative to assets earn higher average returns, and that profitability is negatively correlated with value, so the two combine well.16 Fama and French added profitability and investment factors to their model in 2015.17 Profitability is now treated as a family in its own right, often under the broader label of quality.
Low volatility and low beta
The capital asset pricing model predicts that higher beta should mean higher expected return. Empirically the relationship is much flatter than the theory predicts, a finding that goes back to Black, Jensen and Scholes.18 Frazzini and Pedersen explain the flat line with leverage constraints: investors who cannot borrow buy high-beta assets to obtain more risk, bidding up their prices.19 A related but distinct result concerns idiosyncratic volatility: Ang, Hodrick, Xing and Zhang found that stocks with high idiosyncratic volatility had strikingly low average returns.20 Low-risk strategies are attractive because they can be implemented long-only, and vulnerable because low-volatility stocks can become crowded and expensive.
Carry
Koijen, Moskowitz, Pedersen and Vrugt defined carry as the return an asset would earn if its price did not change, such as the interest-rate differential in currencies or the roll yield in commodity futures, and found that carry predicted returns across equities, bonds, currencies and commodities.21 For a trader of futures and currencies it is as central as trend. Its characteristic risk is a sharp loss when crowded positions unwind together.
Calendar and seasonal effects
Rozeff and Kinney documented that US stock returns were unusually high in January.22 Bouman and Jacobsen found that returns from November to April exceeded those from May to October in almost every market they studied.23 A subtler version appears in the cross-section: Heston and Sadka showed that stocks which performed well in a given calendar month tended to do so in the same month in subsequent years.24 Calendar effects are cheap to test and particularly easy to overfit, because the number of possible calendar rules is very large and each is tested on few independent observations. They warrant more scepticism than any other family here.
Post-earnings-announcement drift
Ball and Brown showed in 1968 that abnormal returns continued in the direction of an earnings surprise after the announcement month.25 Bernard and Thomas examined this drift in detail and concluded that it was better explained as a delayed reaction to the news than as compensation for risk.26 It is among the most replicated patterns in accounting research. It requires clean, point-in-time earnings data, and much of the effect is concentrated in smaller, less liquid stocks where trading costs are highest.
Short-term reversal
Over periods of a week to a month the pattern inverts: recent losers tend to recover and recent winners tend to give back part of their gains. Jegadeesh documented negative serial correlation in monthly returns, and Lehmann found large reversal profits at a weekly frequency.27,28 Nagel showed that reversal returns behave like compensation for supplying liquidity, rising when liquidity is scarce.29 That interpretation also explains the effect’s sensitivity to costs: the strategy trades constantly, in exactly the stocks whose prices have just moved.
| Family | Signal lookback | Leading explanation | Characteristic risk |
|---|---|---|---|
| Cross-sectional momentum | 3 to 12 months | Underreaction to news | Severe crashes in market rebounds |
| Time-series momentum | 1 to 12 months | Underreaction; slow diffusion of information | Losses in range-bound markets |
| Value | Years | Distress risk, or extrapolation | Long periods of underperformance |
| Profitability | Annual accounts | Mispricing of profitable firms, or risk | Overlap with other factors |
| Low volatility / low beta | Months to years | Leverage constraints; benchmarking | Crowding and valuation |
| Carry | Current yield or roll | Compensation for crash and liquidity risk | Abrupt unwinds |
| Calendar and seasonality | Days to months | Contested | Overfitting; few independent samples |
| Post-earnings drift | Weeks | Delayed reaction to news | Costs in small, illiquid stocks |
| Short-term reversal | Days to a month | Compensation for supplying liquidity | Turnover and trading costs |
A summary of the literature discussed above, not a ranking. Each row compresses a debate that remains open.
Decay, replication and costs
McLean and Pontiff studied what happened to 97 published return predictors after their papers appeared.30 They separated two effects. Outside the original sample period but before publication, returns were about a quarter lower, which they attribute to statistical bias in the original studies. After publication they were more than half lower; the additional decline is consistent with investors learning of the pattern and trading on it, and it was larger where arbitrage was cheaper. The size effect is the textbook example: first documented by Banz in 1981,31 it was much weaker in the decades that followed.
The public factor data are consistent with that pattern, but they also show how weak a single factor’s evidence is. Figure 1 compounds each of three long-short factors from the Kenneth R. French Data Library (1927-01 to 2026-08)32 and splits it at the month its best-known paper appeared. All three averaged less afterwards: size 0.30% a month before Banz and 0.02% after, value 0.44% and 0.18%, momentum 0.73% and 0.38%. Yet none of the three declines is itself statistically significant: the t-statistics of the differences are 1.5, 1.3 and 1.2. The split month also matters. Most of size’s pre-publication return in this series comes from 1976-01 to 1981-02, between the end of Banz’s sample and his paper (1.19% a month, t = 3.5); over his own sample window the library’s size factor earned 0.16% a month (t = 1.2). Three factors are an illustration, not a test: the library’s factors are not the papers’ own portfolios, each series is one path, the t-statistics treat months as independent, and the drawdowns in the chart are as informative as the averages. The large cross-section studies above are the evidence.
Momentum (UMD) averaged 0.73% a month before Jegadeesh and Titman (1993) (t = 4.42) and 0.38% a month after it (t = 1.60).
| Factor, published | Mean / month, before | t, before | Mean / month, after | t, after | Sharpe before → after | t of the drop |
|---|---|---|---|---|---|---|
| Size (SMB) Banz (1981) | 0.30% | 2.29 | 0.02% | 0.17 | 0.31 → 0.03 | 1.52 |
| Value (HML) Fama and French (1992) | 0.44% | 3.35 | 0.18% | 1.11 | 0.41 → 0.19 | 1.26 |
| Momentum (UMD) Jegadeesh and Titman (1993) | 0.73% | 4.42 | 0.38% | 1.60 | 0.54 → 0.28 | 1.20 |
Kenneth R. French Data Library: monthly Fama-French factors and momentum factor. 1927-01 to 2026-08. Long-short factors built by the library, not the papers’ own portfolios; before costs. The Sharpe ratio needs no risk-free rate because each factor is self-financing. “t of the drop” is Welch’s t for the difference of the two means, treating months as independent; none clears 2.
def split_at(factor: pd.Series, published: str) -> pd.DataFrame:
"""Mean monthly return, its t-statistic and the annualised Sharpe ratio, before and after a
publication date. A long-short factor is self-financing, so its Sharpe needs no risk-free."""
rows = {}
for name, part in (("before", factor[factor.index < published]),
("after", factor[factor.index >= published])):
mu, sd, n = part.mean(), part.std(ddof=1), len(part)
rows[name] = {"months": n, "mean_pct_month": 100 * mu,
"t_stat": mu / (sd / np.sqrt(n)), "sharpe": mu / sd * np.sqrt(12)}
out = pd.DataFrame(rows).T
# IS THE DROP ITSELF SIGNIFICANT? Welch's t for the difference of the two means. Months are
# treated as independent, which overstates precision where returns are autocorrelated.
a, b = factor[factor.index < published], factor[factor.index >= published]
out["t_diff"] = (a.mean() - b.mean()) / np.sqrt(a.var(ddof=1) / len(a) + b.var(ddof=1) / len(b))
return outdef french_csv(zip_bytes: bytes) -> pd.DataFrame:
"""The first table in a French Data Library CSV zip file, as decimal returns by date."""
with zipfile.ZipFile(io.BytesIO(zip_bytes)) as zf:
text = zf.read(zf.namelist()[0]).decode("latin-1")
lines = text.splitlines()
start = next(i for i, l in enumerate(lines) if l.strip().startswith(",") or l.lower().startswith(",mom"))
rows = []
for l in lines[start + 1:]:
parts = [p.strip() for p in l.split(",")]
if not parts[0].isdigit() or len(parts[0]) not in (6, 8):
break
rows.append(parts)
cols = ["date"] + [c.strip() for c in lines[start].split(",")[1:]]
df = pd.DataFrame(rows, columns=cols)
fmt = "%Y%m%d" if len(df["date"].iloc[0]) == 8 else "%Y%m"
df.index = pd.to_datetime(df.pop("date"), format=fmt)
return df.astype(float) / 100.0import urllib.request
FRENCH = "https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/ftp/"
raw = urllib.request.urlopen(FRENCH + "F-F_Momentum_Factor_CSV.zip").read()
mom = french_csv(raw).iloc[:, 0] # monthly momentum factor, decimal returns
print(split_at(mom, "1993-03-01"))A second line of work asks whether the original results replicate at all, and the answers depend on method. Hou, Xue and Zhang re-tested several hundred anomalies with a common procedure that limits the influence of microcap stocks, and found that a majority lost statistical significance.33 Using different protocols, Chen and Zimmermann reproduced the large majority of the published predictors they examined,34 and Jensen, Kelly and Pedersen found that most factors replicate and survive when evaluated jointly with a Bayesian model.35 The disagreement is largely about small, illiquid stocks and about how much weight a single t-statistic should carry.
Costs are the final filter (Field note 05 (Execution: TWAP, slicing and jitter) covers how they arise). Novy-Marx and Velikov show that trading costs eliminate most of the profit from high-turnover anomalies, that low-turnover strategies such as value survive largely intact, and that mid-turnover strategies such as momentum survive much better with simple rules that trade only when the expected benefit exceeds the cost.36
The practical reading: assume a published effect is smaller than the paper reports, possibly much smaller; assume costs will take a meaningful share of what remains; and prefer families whose economic rationale does not depend on other investors remaining uninformed.
Data snooping and multiple testing
A more fundamental problem than changing markets is selection. Researchers, professional and amateur, test far more ideas than they report, and the ideas that are reported are those that happened to look good. Lo and MacKinlay quantified how severely this kind of data snooping can distort inference in asset pricing.37
The arithmetic is unforgiving. If a signal has no predictive power, a test at the 5% level still passes one time in twenty. Test twenty independent signals with no predictive power, and the probability that at least one passes is 64%. The same logic applies to Sharpe ratios. For serially uncorrelated returns and a true Sharpe ratio of zero, an annualised Sharpe ratio estimated from T years of data has a standard error of approximately 1/√T.38 The best of N such independent estimates then has an expected value of roughly
where Φ−1 is the standard normal quantile function and γ ≈ 0.577 is the Euler–Mascheroni constant; the bracket approximates the expected maximum of N standard normal draws.39 With ten years of data, the best of a hundred strategies with no edge is expected to show an annual Sharpe ratio of about 0.80; with five years and twenty attempts, about 0.85. Either would pass a casual screen as a genuine strategy. Figure 2 shows how quickly this noise ceiling rises with the number of trials.
Test 100 independent ideas with no edge on 10 years of data and the best is expected to show an annual Sharpe ratio of about 0.80. The chance that at least one passes a test at the 5% level is 99%.
The same arithmetic, in code, so the curve can be checked or extended:
def expected_max_sharpe(trials: int, years: float) -> float:
"""Expected best annual Sharpe among `trials` independent strategies with no edge
(Bailey and Lopez de Prado 2014), with the standard error 1/sqrt(years) under the null."""
if trials < 2:
return 0.0
z = (1 - EULER_GAMMA) * norm.ppf(1 - 1 / trials) + EULER_GAMMA * norm.ppf(1 - 1 / (trials * np.e))
return z / np.sqrt(years)The profession’s response has been to raise the bar. Harvey, Liu and Zhu catalogued hundreds of published factors and argued that, given how many have been tried, a new one should clear a t-statistic of about 3.0 rather than the conventional 2.0.40 White’s reality check tests whether the best of a set of strategies outperforms by more than chance, given the whole set.41 Bailey and López de Prado’s deflated Sharpe ratio adjusts an observed Sharpe ratio for the number of trials, the length of the sample and the shape of the return distribution.39 All three depend on a complete record of what was tried: the number of trials for the t-statistic hurdle, the number and dispersion of trial Sharpe ratios for the deflated Sharpe ratio, and the full set of trial returns for White’s test.
Implications for a small operator
A small operator cannot out-research a large quantitative firm, and should not try to find exotic new anomalies. It has different advantages: it is small enough to trade capacity-constrained ideas, it answers to no benchmark, and it can hold a position through a poor year without a committee. Those advantages suggest a way of working, sketched in Figure 3 and set out below.
- Start from families with an economic rationale and a long out-of-sample record. Momentum, trend, value, profitability, carry and low volatility are not secret, which is why they are a sensible foundation. Most now have decades of post-publication data, and all have at least several years, which is genuine evidence of what has survived.
- Write the hypothesis down before the backtest. State the mechanism, the universe, the holding period and the result that would lead you to abandon the idea. A hypothesis written afterwards is a description of the data.
- Record every trial. Keep a ledger of every configuration tested, including the failures, and use it when judging the best result. The deflated Sharpe ratio is a sensible default.
- Prefer broad, simple, stable rules. A signal that works across many instruments and a wide range of parameters is more credible than one that works well at a single setting. A result that depends on a lookback of exactly 47 days is probably noise.
- Model costs conservatively and hold data back. Charge realistic spreads, commissions and market impact, and reserve a period of history, and ideally a set of markets, that is not examined until the design is fixed.
- Combine families rather than perfecting one. Anomalies with different explanations tend to fail at different times. A modest allocation to several is usually more robust than a finely tuned version of one.
- Expect decay, and decide in advance how to recognise it. Specify what live underperformance, over what period, would count as evidence that an edge has gone, and allow for how long a genuine edge can underperform by chance: a strategy with a true Sharpe ratio of 0.5 has about a 19% chance of a negative three-year record.
None of this guarantees a profitable strategy. It makes it much more likely that the strategy eventually traded is the one the research actually found, rather than the luckiest of the ones that were tried.
References
- Fama, E. F. (1970). Efficient capital markets: A review of theory and empirical work. Journal of Finance, 25(2), 383-417.
- Fama, E. F., & French, K. R. (1993). Common risk factors in the returns on stocks and bonds. Journal of Financial Economics, 33(1), 3-56.
- Fama, E. F., & French, K. R. (1996). Multifactor explanations of asset pricing anomalies. Journal of Finance, 51(1), 55-84.
- Barberis, N., Shleifer, A., & Vishny, R. W. (1998). A model of investor sentiment. Journal of Financial Economics, 49(3), 307-343.
- Lakonishok, J., Shleifer, A., & Vishny, R. W. (1994). Contrarian investment, extrapolation, and risk. Journal of Finance, 49(5), 1541-1578.
- Shleifer, A., & Vishny, R. W. (1997). The limits of arbitrage. Journal of Finance, 52(1), 35-55.
- Baker, M., Bradley, B., & Wurgler, J. (2011). Benchmarks as limits to arbitrage: Understanding the low-volatility anomaly. Financial Analysts Journal, 67(1), 40-54.
- Jegadeesh, N., & Titman, S. (1993). Returns to buying winners and selling losers: Implications for stock market efficiency. Journal of Finance, 48(1), 65-91.
- Carhart, M. M. (1997). On persistence in mutual fund performance. Journal of Finance, 52(1), 57-82.
- Asness, C. S., Moskowitz, T. J., & Pedersen, L. H. (2013). Value and momentum everywhere. Journal of Finance, 68(3), 929-985.
- Daniel, K., & Moskowitz, T. J. (2016). Momentum crashes. Journal of Financial Economics, 122(2), 221-247.
- Barroso, P., & Santa-Clara, P. (2015). Momentum has its moments. Journal of Financial Economics, 116(1), 111-120.
- Moskowitz, T. J., Ooi, Y. H., & Pedersen, L. H. (2012). Time series momentum. Journal of Financial Economics, 104(2), 228-250.
- Hurst, B., Ooi, Y. H., & Pedersen, L. H. (2017). A century of evidence on trend-following investing. Journal of Portfolio Management, 44(1), 15-29.
- Fama, E. F., & French, K. R. (1992). The cross-section of expected stock returns. Journal of Finance, 47(2), 427-465.
- Novy-Marx, R. (2013). The other side of value: The gross profitability premium. Journal of Financial Economics, 108(1), 1-28.
- Fama, E. F., & French, K. R. (2015). A five-factor asset pricing model. Journal of Financial Economics, 116(1), 1-22.
- Black, F., Jensen, M. C., & Scholes, M. (1972). The capital asset pricing model: Some empirical tests. In M. C. Jensen (Ed.), Studies in the Theory of Capital Markets (pp. 79-121). Praeger.
- Frazzini, A., & Pedersen, L. H. (2014). Betting against beta. Journal of Financial Economics, 111(1), 1-25.
- Ang, A., Hodrick, R. J., Xing, Y., & Zhang, X. (2006). The cross-section of volatility and expected returns. Journal of Finance, 61(1), 259-299.
- Koijen, R. S. J., Moskowitz, T. J., Pedersen, L. H., & Vrugt, E. B. (2018). Carry. Journal of Financial Economics, 127(2), 197-225.
- Rozeff, M. S., & Kinney, W. R. (1976). Capital market seasonality: The case of stock returns. Journal of Financial Economics, 3(4), 379-402.
- Bouman, S., & Jacobsen, B. (2002). The Halloween indicator, “Sell in May and go away”: Another puzzle. American Economic Review, 92(5), 1618-1635.
- Heston, S. L., & Sadka, R. (2008). Seasonality in the cross-section of stock returns. Journal of Financial Economics, 87(2), 418-445.
- Ball, R., & Brown, P. (1968). An empirical evaluation of accounting income numbers. Journal of Accounting Research, 6(2), 159-178.
- Bernard, V. L., & Thomas, J. K. (1989). Post-earnings-announcement drift: Delayed price response or risk premium? Journal of Accounting Research, 27 (Supplement), 1-36.
- Jegadeesh, N. (1990). Evidence of predictable behavior of security returns. Journal of Finance, 45(3), 881-898.
- Lehmann, B. N. (1990). Fads, martingales, and market efficiency. Quarterly Journal of Economics, 105(1), 1-28.
- Nagel, S. (2012). Evaporating liquidity. Review of Financial Studies, 25(7), 2005-2039.
- McLean, R. D., & Pontiff, J. (2016). Does academic research destroy stock return predictability? Journal of Finance, 71(1), 5-32.
- Banz, R. W. (1981). The relationship between return and market value of common stocks. Journal of Financial Economics, 9(1), 3-18.
- French, K. R. (2026). Data Library: Fama/French 3 factors and momentum factor (monthly). Tuck School of Business, Dartmouth College. https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html (retrieved September 2026; data through August 2026).
- Hou, K., Xue, C., & Zhang, L. (2020). Replicating anomalies. Review of Financial Studies, 33(5), 2019-2133.
- Chen, A. Y., & Zimmermann, T. (2022). Open source cross-sectional asset pricing. Critical Finance Review, 11(2), 207-264.
- Jensen, T. I., Kelly, B., & Pedersen, L. H. (2023). Is there a replication crisis in finance? Journal of Finance, 78(5), 2465-2518.
- Novy-Marx, R., & Velikov, M. (2016). A taxonomy of anomalies and their trading costs. Review of Financial Studies, 29(1), 104-147.
- Lo, A. W., & MacKinlay, A. C. (1990). Data-snooping biases in tests of financial asset pricing models. Review of Financial Studies, 3(3), 431-467.
- Lo, A. W. (2002). The statistics of Sharpe ratios. Financial Analysts Journal, 58(4), 36-52.
- Bailey, D. H., & López de Prado, M. (2014). The deflated Sharpe ratio: Correcting for selection bias, backtest overfitting, and non-normality. Journal of Portfolio Management, 40(5), 94-107.
- Harvey, C. R., Liu, Y., & Zhu, H. (2016). …and the cross-section of expected returns. Review of Financial Studies, 29(1), 5-68.
- White, H. (2000). A reality check for data snooping. Econometrica, 68(5), 1097-1126.