Statistical Arbitrage: How It Works and Why It Stops Working
Statistical arbitrage is one bet made thousands of times at once. Two related prices have separated, and the gap between them closes back up. Each bet is worth a fraction of a percent, and the money comes from the count, and from trading cheaply enough to run that many.
The record is good and it is not flawless. Aurum, which publishes composite returns across the hedge fund industry, puts the stat arb composite at 13.4% in 2024 and 9.8% in 2025. The second year is the more useful one: quant as a whole finished seventh of the eight strategy groups Aurum tracks, returning 8.8% against an industry composite of 12.6%.
Good years attract imitators, and plenty of things now get sold as statistical arbitrage that hold positions for months, run forty positions, and rest on a backtest fitted to five years of daily price data. A family office pays for that in fees, for years. Someone building it at home finds out in the first live week.
If you check one thing, check how long it holds a position. Real stat arb turns over in days or less. Most of what gets sold alongside it runs for months.
- How does statistical arbitrage work?
- How is stat arb different from pairs trading?
- What it takes to build a stat arb strategy
- Why has stat arb made money lately?
- What happens when a stat arb fund gets too big?
- Why do stat arb signals stop working?
- How can you tell real stat arb from a slow factor strategy?
- How rare is a program that passes all six checks?
How does statistical arbitrage work?
Statistical arbitrage buys one security and sells a related one whenever the price gap between them has stretched wider than usual, then closes both when the gap narrows again. Each trade earns a fraction of a percent. A portfolio runs hundreds or thousands of them at once, held for days.
The same idea comes in four common shapes. Market-neutral sets one security against another so the pair does not care which way the market moves. Cross-market trades a security against its own listing on a second exchange. Cross-asset trades an instrument against whatever it is meant to track, a convertible bond against the shares it converts into. ETF arbitrage trades a fund against the basket of shares inside it.
Picking the pairs is where the work in statistical arbitrage sits. Two securities qualify when their prices wander freely but the gap between them keeps returning to a stable average, a relationship called cointegration. A regression on their price histories then sets how much of the second security it takes to hedge the first.
Two standard deviations wide, you open. Back near the average, you close. The hard part is deciding which pairs deserve that treatment, and pricing the round trip before you take it.
A thousand positions cannot be run by hand, and one costing mistake repeats a thousand times, which is why data and execution infrastructure decides more of the result here than in any slower strategy. Gatev, Goetzmann and Rouwenhorst record that the Morgan Stanley desk run by Nunzio Tartaglia reportedly made $50 million from pairs trading in 1987, then disbanded in 1989 after two bad years.
How is stat arb different from pairs trading?
A pairs trade is one spread between two securities. Statistical arbitrage is the portfolio version of the same idea, run across hundreds or thousands of spreads at once, where no single pair matters and the money comes from the count. Pairs trading is a part of stat arb, not another name for it.
Take one pair on its own. Two regional utilities have tracked each other for years, and the regression says every $100 in one is hedged by $80 of the other. One drops 6% on a credit downgrade while the other barely moves, and the gap opens to two standard deviations.
Buy $100 of the one that fell, short $80 of the one that did not, wait. If the relationship was real, the gap closes and the pair earns a fraction of a percent on the money tied up.
If the downgrade was the first sign of something worse, the gap keeps widening instead. The loss on the short side has no ceiling, because a share you have borrowed and sold can rise without limit.
That is exactly why the count matters. One pair like that is a coin flip with a long, ugly tail on one side. Nine hundred of them, uncorrelated enough, is a business. Most tutorials teach the single pair and describe it as stat arb, which is how so many home-built versions end up running forty positions and behaving nothing like the funds they get compared to.
What it takes to build a stat arb strategy
Four decisions carry most of a statistical arbitrage strategy’s result, and the signal formula is not one of them: which securities you scan, how you test the relationship, what you assume it costs to trade, and how you validate. Tutorials spend their length on the formula because it is the part that fits inside a code block.
Universe selection sets the ceiling before a single test runs. Every combination in a 3,000-name universe is roughly 4.5 million pair tests, and at the 5% threshold most people use, around 225,000 of those pairs will look statistically sound while being nothing of the kind. Narrowing to economically related securities first cuts the search to a size a test result can survive.
Cointegration testing is the step most builds get wrong. Two price series can track each other closely for a year and share no long-run relationship at all, and that is precisely the pair that falls apart in month thirteen.
The Engle-Granger and Johansen tests exist to separate the two cases. Both mean something only when the relationship still holds outside the window it was estimated on, which is what out-of-sample testing is for.
Then the arithmetic that kills most home-built versions: transaction costs. A book at a two-day holding period turns its capital over roughly 125 times a year, so one basis point of unmodelled slippage per round trip, one hundredth of a percentage point, costs 125 basis points a year. Run it the other way: if a realistic round trip costs 10 basis points, the gaps have to pay 12.5% a year before the strategy keeps anything.

What a two-day holding period costs before the strategy makes anything. Algotrader.ch, 2026.
Walk-forward validation comes last and gets skipped most. Settings are re-estimated on a rolling window and tested only on the stretch of history after it, so the lookback length, the entry threshold and the pair list are never chosen with knowledge of the returns they are scored on.
On a statistical arbitrage book the setting that does the most damage is the lookback used to estimate the hedge ratio, and it is also the one most often picked after the returns have been seen. A backtest built in that order is measuring that choice and reporting it as the strategy.
Why has stat arb made money lately?
Three things outside any manager’s control did most of the work. Shares in the same sector have been moving in different directions instead of together, which widens the gaps the strategy trades. Crowded positions came apart. And the dealers who normally close small gaps have been leaving more of them open.
The middle one is worth spelling out. By late 2021 a great many funds were holding the same handful of share characteristics, cheapness or momentum or company size, in the same direction. That pile-up came apart over the following years, and prices moved in ways the models could trade.
None of the three is published as a number you can check against a particular fund. A manager who credits their own research for a run of years when all three were present is claiming something the public figures cannot separate.
The category average hides the rest. Aurum publishes returns by sub-strategy and never by fund, so you can see that stat arb made 9.8% in 2025 while quant as a whole made 8.8%, and nothing at all about the distance between the best and worst fund inside either number.
That calm is partly the averaging. A standard deviation of 3.8% describes the category once hundreds of funds are pooled together, and no single fund inside it is anywhere near that steady.
What happens when a stat arb fund gets too big?
Its returns shrink, and nothing the manager does prevents it. The price gaps stat arb lives on are small and short-lived, and more money does not create more of them. A larger fund chases the same gaps with bigger orders, and the price has moved by the time those orders fill.
Which makes “what is your capacity” the wrong question on its own. Ask what that number assumes about how often the book rebalances, how much of it is in use today, and what happens to expected return as the fund approaches it. The same three questions work on your own strategy’s capacity.
Two numbers also get quietly swapped. The capacity a research team calculates for a signal at full strength is not the one that reaches a marketing deck, and a fund that quotes only the second has handed you a sales figure.
Renaissance closed Medallion to new outside money in 1993 and bought out its last outside investor in 2005. Across 2010 to 2018 the fund’s assets stayed near $10 billion while it compounded at more than 29% a year. Assets hold flat against returns like that only if the profits leave each year, which is the practice Renaissance is reported to follow.
Turning capital away is what that discipline costs, and few firms pay it.
None of this reaches a small account. A limit that forces a multi-billion-dollar fund to close does not register on $200,000, because an order that size is too small a share of the day’s trading to move the price against itself. Small size also opens the thinly traded end of the market that a large fund has to skip.
What limits a small account instead is cost per trade, and whether the shares the short side needs can be borrowed at all.
Why do stat arb signals stop working?
Because other people find them. Once enough money is chasing the same small price gap, the gap closes faster and pays less, until nothing is left in it. The industry calls this signal decay, it happens to every statistical arbitrage signal eventually, and it is the market doing its job.
The funds worth backing measure how fast a signal is fading and retire it on a schedule, before the returns make the decision for them.
Data leakage is the quieter version of the same problem. A test that lets tomorrow’s information touch today’s decision produces a return that never existed, and stat arb is unusually exposed, because the gaps are small enough that a tiny leak explains most of what shows on screen.
In your own code it comes from one of three places. A rebalance that fires on the same day’s close. A fundamentals field carrying the restated figure instead of the one published at the time. A universe built from today’s index membership, which quietly deletes every company that got taken over or went to zero.
Each has a dull fix: lag the signal by one day, buy point-in-time data that records what was known on the day, and build the universe from a price file that still contains the dead companies. None is expensive. All three get skipped.
Search enough pairs and thresholds against one history and some will fit it beautifully. That is overfitting, a result of how hard you searched rather than anything about the strategy. If a fund or a repository cannot say how look-ahead bias was ruled out, the safe reading is that nobody checked. Turnover makes both problems expensive: at 125 round trips a year, a leak worth two basis points a round trip is two and a half percentage points of yearly return that was never there.
- Dozens of quant equity funds, many of them marketed as statistical arbitrage, had independently ended up holding near-identical long and short positions
- Khandani and Lo attribute the losses to a coordinated deleveraging of similarly constructed portfolios, not to any single signal breaking
- Their transaction data puts the sharpest damage inside two intraday windows, on 1 August 2007 between 10:45 and 11:30 New York time and on 6 August 2007 from the open until 1pm. Months of return went in hours
- A backtest run on that period’s own data priced the crowding as ordinary noise, because the risk sat outside the window
In January 2025 the SEC found that Two Sigma had left recognised weaknesses in models managing client assets unaddressed until August 2023, and had failed to supervise an employee who made unauthorised changes to more than a dozen models. The firm repaid $165 million and paid a $90 million penalty. Decay and leakage usually arrive that way, as an open item nobody closed.
How can you tell real stat arb from a slow factor strategy?
Six questions separate statistical arbitrage from what imitates it, and the first two settle most cases. Holding period and position count are facts a fund either states or avoids, and if those two do not line up the other four rarely rescue it. They work on a pitch deck and on your own code.
- What is the holding period? Intraday to a few days. Never weeks.
- How many positions are open at once? Hundreds or thousands. Forty is a different strategy.
- What does the monthly return pattern look like? Steady, with no single month explaining the year.
- How is capacity described? By signal and rebalancing frequency, with a number attached and a share of it in use today.
- What is the data and execution stack made of? Specific vendors, specific feeds, and who keeps them running. “Proprietary technology” is not an answer.
- How was look-ahead bias ruled out? With a specific, testable answer.
The hardest case to catch is a slower factor strategy with a hedge stapled on, sold as statistical arbitrage. Different return shape, different capacity ceiling, and a failure mode that shows up years after the money went in.
Harvey, Liu and Zhu catalogued 316 published equity factors up to 2016 and showed that so many had been tried that a good number were always going to look convincing by luck alone, which is why they argue a new one should clear a much stricter statistical bar than the conventional one. If the factor underneath fails that bar, no amount of hedging turns it into a strategy.
| What to check | Real stat arb | Slow factor book sold as stat arb |
|---|---|---|
| Position count | Hundreds to thousands | Dozens, often concentrated |
| Holding period | Intraday to a few days | Weeks to months |
| Monthly return pattern | Low, steady volatility, no month explaining the year | Lumpy, a few standout months carrying the return |
| Infrastructure claim | Vendors, feeds and who runs them, stated | “Proprietary technology”, rarely specified |
Holding period is the fastest tell, because trading at that frequency needs infrastructure most funds do not have and cannot fake. A fund holding for weeks is carrying factor exposure that a pair hedge does not remove over that length of time, and it should be sized against that exposure rather than against a spread.
How rare is a program that passes all six checks?
Rare, and not because the bar is exotic. None of the six questions needs access a private investor lacks, which is what makes the answers so telling: a fund that will not answer them has told you what it is able to evidence.
That standard is what the strategies we feature in The Review get checked against, and the research-discipline criteria behind our scoring are published in full so you can run them yourself.
They are scored as five separate dimensions with no overall grade, so a weak answer on capacity cannot be averaged away by a strong one somewhere else.