Skip to content

Statistical Arbitrage: How It Works and Why It Stops Working

Statistical arbitrage is one bet made thousands of times at once. Two related prices have separated, and the gap between them closes back up. Each bet is worth a fraction of a percent, and the money comes from the count, and from trading cheaply enough to run that many.

The record is good and it is not flawless. Aurum, which publishes composite returns across the hedge fund industry, puts the stat arb composite at 13.4% in 2024 and 9.8% in 2025. The second year is the more useful one: quant as a whole finished seventh of the eight strategy groups Aurum tracks, returning 8.8% against an industry composite of 12.6%.

Good years attract imitators, and plenty of things now get sold as statistical arbitrage that hold positions for months, run forty positions, and rest on a backtest fitted to five years of daily price data. A family office pays for that in fees, for years. Someone building it at home finds out in the first live week.

If you check one thing, check how long it holds a position. Real stat arb turns over in days or less. Most of what gets sold alongside it runs for months.

How does statistical arbitrage work?

Statistical arbitrage buys one security and sells a related one whenever the price gap between them has stretched wider than usual, then closes both when the gap narrows again. Each trade earns a fraction of a percent. A portfolio runs hundreds or thousands of them at once, held for days.

The same idea comes in four common shapes. Market-neutral sets one security against another so the pair does not care which way the market moves. Cross-market trades a security against its own listing on a second exchange. Cross-asset trades an instrument against whatever it is meant to track, a convertible bond against the shares it converts into. ETF arbitrage trades a fund against the basket of shares inside it.

Picking the pairs is where the work in statistical arbitrage sits. Two securities qualify when their prices wander freely but the gap between them keeps returning to a stable average, a relationship called cointegration. A regression on their price histories then sets how much of the second security it takes to hedge the first.

Two standard deviations wide, you open. Back near the average, you close. The hard part is deciding which pairs deserve that treatment, and pricing the round trip before you take it.

A thousand positions cannot be run by hand, and one costing mistake repeats a thousand times, which is why data and execution infrastructure decides more of the result here than in any slower strategy. Gatev, Goetzmann and Rouwenhorst record that the Morgan Stanley desk run by Nunzio Tartaglia reportedly made $50 million from pairs trading in 1987, then disbanded in 1989 after two bad years.

2024 was strong and still had two losing months
+13.4%
Aurum’s stat arb composite, built from fund-reported returns, with two losing months and a one-year standard deviation of 3.8% (Aurum Hedge Fund Industry Deep Dive, 2024 review).
Q2 2025 was the biggest inflow quarter since 2014
$24.8bn
Net hedge fund inflow in that quarter alone. Relative value arbitrage, the wider family stat arb sits inside, took $7.7bn. Separately, firms managing more than $5bn took $22.9bn of the same total (HFR, 2025).
Quant trailed the industry in 2025
7th of 8
Where quant ranked by return among the eight strategy groups Aurum tracks in 2025, at 8.8%, against an industry composite of 12.6%. Stat arb inside it made 9.8% (Aurum, January 2026).

How is stat arb different from pairs trading?

A pairs trade is one spread between two securities. Statistical arbitrage is the portfolio version of the same idea, run across hundreds or thousands of spreads at once, where no single pair matters and the money comes from the count. Pairs trading is a part of stat arb, not another name for it.

Take one pair on its own. Two regional utilities have tracked each other for years, and the regression says every $100 in one is hedged by $80 of the other. One drops 6% on a credit downgrade while the other barely moves, and the gap opens to two standard deviations.

Buy $100 of the one that fell, short $80 of the one that did not, wait. If the relationship was real, the gap closes and the pair earns a fraction of a percent on the money tied up.

If the downgrade was the first sign of something worse, the gap keeps widening instead. The loss on the short side has no ceiling, because a share you have borrowed and sold can rise without limit.

That is exactly why the count matters. One pair like that is a coin flip with a long, ugly tail on one side. Nine hundred of them, uncorrelated enough, is a business. Most tutorials teach the single pair and describe it as stat arb, which is how so many home-built versions end up running forty positions and behaving nothing like the funds they get compared to.

What it takes to build a stat arb strategy

Four decisions carry most of a statistical arbitrage strategy’s result, and the signal formula is not one of them: which securities you scan, how you test the relationship, what you assume it costs to trade, and how you validate. Tutorials spend their length on the formula because it is the part that fits inside a code block.

Universe selection sets the ceiling before a single test runs. Every combination in a 3,000-name universe is roughly 4.5 million pair tests, and at the 5% threshold most people use, around 225,000 of those pairs will look statistically sound while being nothing of the kind. Narrowing to economically related securities first cuts the search to a size a test result can survive.

Cointegration testing is the step most builds get wrong. Two price series can track each other closely for a year and share no long-run relationship at all, and that is precisely the pair that falls apart in month thirteen.

The Engle-Granger and Johansen tests exist to separate the two cases. Both mean something only when the relationship still holds outside the window it was estimated on, which is what out-of-sample testing is for.

Then the arithmetic that kills most home-built versions: transaction costs. A book at a two-day holding period turns its capital over roughly 125 times a year, so one basis point of unmodelled slippage per round trip, one hundredth of a percentage point, costs 125 basis points a year. Run it the other way: if a realistic round trip costs 10 basis points, the gaps have to pay 12.5% a year before the strategy keeps anything.

Statistical arbitrage cost arithmetic: a two-day holding period means about 125 round trips a year, so a 10 basis point all-in round trip requires 12.5% gross spread capture before the book keeps anything

What a two-day holding period costs before the strategy makes anything. Algotrader.ch, 2026.

Walk-forward validation comes last and gets skipped most. Settings are re-estimated on a rolling window and tested only on the stretch of history after it, so the lookback length, the entry threshold and the pair list are never chosen with knowledge of the returns they are scored on.

On a statistical arbitrage book the setting that does the most damage is the lookback used to estimate the hedge ratio, and it is also the one most often picked after the returns have been seen. A backtest built in that order is measuring that choice and reporting it as the strategy.

Why has stat arb made money lately?

Three things outside any manager’s control did most of the work. Shares in the same sector have been moving in different directions instead of together, which widens the gaps the strategy trades. Crowded positions came apart. And the dealers who normally close small gaps have been leaving more of them open.

The middle one is worth spelling out. By late 2021 a great many funds were holding the same handful of share characteristics, cheapness or momentum or company size, in the same direction. That pile-up came apart over the following years, and prices moved in ways the models could trade.

None of the three is published as a number you can check against a particular fund. A manager who credits their own research for a run of years when all three were present is claiming something the public figures cannot separate.

The category average hides the rest. Aurum publishes returns by sub-strategy and never by fund, so you can see that stat arb made 9.8% in 2025 while quant as a whole made 8.8%, and nothing at all about the distance between the best and worst fund inside either number.

That calm is partly the averaging. A standard deviation of 3.8% describes the category once hundreds of funds are pooled together, and no single fund inside it is anywhere near that steady.

What happens when a stat arb fund gets too big?

Its returns shrink, and nothing the manager does prevents it. The price gaps stat arb lives on are small and short-lived, and more money does not create more of them. A larger fund chases the same gaps with bigger orders, and the price has moved by the time those orders fill.

Which makes “what is your capacity” the wrong question on its own. Ask what that number assumes about how often the book rebalances, how much of it is in use today, and what happens to expected return as the fund approaches it. The same three questions work on your own strategy’s capacity.

Two numbers also get quietly swapped. The capacity a research team calculates for a signal at full strength is not the one that reaches a marketing deck, and a fund that quotes only the second has handed you a sales figure.

What refusing capital looks like from outside

Renaissance closed Medallion to new outside money in 1993 and bought out its last outside investor in 2005. Across 2010 to 2018 the fund’s assets stayed near $10 billion while it compounded at more than 29% a year. Assets hold flat against returns like that only if the profits leave each year, which is the practice Renaissance is reported to follow.

Turning capital away is what that discipline costs, and few firms pay it.

None of this reaches a small account. A limit that forces a multi-billion-dollar fund to close does not register on $200,000, because an order that size is too small a share of the day’s trading to move the price against itself. Small size also opens the thinly traded end of the market that a large fund has to skip.

What limits a small account instead is cost per trade, and whether the shares the short side needs can be borrowed at all.

Why do stat arb signals stop working?

Because other people find them. Once enough money is chasing the same small price gap, the gap closes faster and pays less, until nothing is left in it. The industry calls this signal decay, it happens to every statistical arbitrage signal eventually, and it is the market doing its job.

The funds worth backing measure how fast a signal is fading and retire it on a schedule, before the returns make the decision for them.

Data leakage is the quieter version of the same problem. A test that lets tomorrow’s information touch today’s decision produces a return that never existed, and stat arb is unusually exposed, because the gaps are small enough that a tiny leak explains most of what shows on screen.

In your own code it comes from one of three places. A rebalance that fires on the same day’s close. A fundamentals field carrying the restated figure instead of the one published at the time. A universe built from today’s index membership, which quietly deletes every company that got taken over or went to zero.

Each has a dull fix: lag the signal by one day, buy point-in-time data that records what was known on the day, and build the universe from a price file that still contains the dead companies. None is expensive. All three get skipped.

Search enough pairs and thresholds against one history and some will fit it beautifully. That is overfitting, a result of how hard you searched rather than anything about the strategy. If a fund or a repository cannot say how look-ahead bias was ruled out, the safe reading is that nobody checked. Turnover makes both problems expensive: at 125 round trips a year, a leak worth two basis points a round trip is two and a half percentage points of yearly return that was never there.

What happened · August 2007
  • Dozens of quant equity funds, many of them marketed as statistical arbitrage, had independently ended up holding near-identical long and short positions
  • Khandani and Lo attribute the losses to a coordinated deleveraging of similarly constructed portfolios, not to any single signal breaking
  • Their transaction data puts the sharpest damage inside two intraday windows, on 1 August 2007 between 10:45 and 11:30 New York time and on 6 August 2007 from the open until 1pm. Months of return went in hours
  • A backtest run on that period’s own data priced the crowding as ordinary noise, because the risk sat outside the window

In January 2025 the SEC found that Two Sigma had left recognised weaknesses in models managing client assets unaddressed until August 2023, and had failed to supervise an employee who made unauthorised changes to more than a dozen models. The firm repaid $165 million and paid a $90 million penalty. Decay and leakage usually arrive that way, as an open item nobody closed.

A stat arb fund that has never lost money to crowding has not been running long. The ones worth backing know which of their positions everyone else is holding too, before the day it matters.

Algotrader.ch Editorial Team

How can you tell real stat arb from a slow factor strategy?

Six questions separate statistical arbitrage from what imitates it, and the first two settle most cases. Holding period and position count are facts a fund either states or avoids, and if those two do not line up the other four rarely rescue it. They work on a pitch deck and on your own code.

  1. What is the holding period? Intraday to a few days. Never weeks.
  2. How many positions are open at once? Hundreds or thousands. Forty is a different strategy.
  3. What does the monthly return pattern look like? Steady, with no single month explaining the year.
  4. How is capacity described? By signal and rebalancing frequency, with a number attached and a share of it in use today.
  5. What is the data and execution stack made of? Specific vendors, specific feeds, and who keeps them running. “Proprietary technology” is not an answer.
  6. How was look-ahead bias ruled out? With a specific, testable answer.

The hardest case to catch is a slower factor strategy with a hedge stapled on, sold as statistical arbitrage. Different return shape, different capacity ceiling, and a failure mode that shows up years after the money went in.

Harvey, Liu and Zhu catalogued 316 published equity factors up to 2016 and showed that so many had been tried that a good number were always going to look convincing by luck alone, which is why they argue a new one should clear a much stricter statistical bar than the conventional one. If the factor underneath fails that bar, no amount of hedging turns it into a strategy.

What to checkReal stat arbSlow factor book sold as stat arb
Position countHundreds to thousandsDozens, often concentrated
Holding periodIntraday to a few daysWeeks to months
Monthly return patternLow, steady volatility, no month explaining the yearLumpy, a few standout months carrying the return
Infrastructure claimVendors, feeds and who runs them, stated“Proprietary technology”, rarely specified

Holding period is the fastest tell, because trading at that frequency needs infrastructure most funds do not have and cannot fake. A fund holding for weeks is carrying factor exposure that a pair hedge does not remove over that length of time, and it should be sized against that exposure rather than against a spread.

Very few programs clear this

How rare is a program that passes all six checks?

Rare, and not because the bar is exotic. None of the six questions needs access a private investor lacks, which is what makes the answers so telling: a fund that will not answer them has told you what it is able to evidence.

That standard is what the strategies we feature in The Review get checked against, and the research-discipline criteria behind our scoring are published in full so you can run them yourself.

They are scored as five separate dimensions with no overall grade, so a weak answer on capacity cannot be averaged away by a strong one somewhere else.

Have a stat arb program checked against the six questions
Get in touch →

Common questions about stat arb

Is statistical arbitrage the same as classical arbitrage?
No, and the shared word causes real confusion. Classical arbitrage is close to risk-free: buying and selling the same asset at two prices at the same moment. Statistical arbitrage is a bet that a price relationship reverts, with no such guarantee, and the short side can lose without limit if it does not. Anyone leaning on the word arbitrage to imply safety is stretching it.
How risky is statistical arbitrage?
Less risky than a directional bet on any one share, and more risky than the word arbitrage suggests. Because the positions are hedged and spread across hundreds of pairs, a normal month is quiet. The danger is concentrated: several pairs can break for the same underlying reason at the same time, and every fund holding similar positions then sells into the same market. August 2007 is the standing example, and it cost quant funds months of return in hours.
Is statistical arbitrage legal?
Yes. It trades on published prices and public data, and it is run openly by regulated funds and by private individuals. What is regulated is the conduct around it: short selling rules differ by market, borrowing shares requires a broker who can locate them, and using information that is not public turns any strategy into insider dealing regardless of the mathematics attached to it.
Can retail investors access statistical arbitrage?
Two routes, worth keeping apart. Buying into an institutional program is hard: high minimums, closed share classes, and capacity limits that keep those funds deliberately small. Running a version yourself is workable at small size, because the capacity ceiling that constrains a multi-billion-dollar fund does not reach down that far. What limits a small account is cost per trade, whether the shares the short side needs can be borrowed, and data quality.
Do you need a coding background to run stat arb?
Yes, at least enough to handle data and run tests without trusting a black box. Cointegration tests are standard in Python and R and in most commercial research platforms, so the modelling library is rarely what separates a workable setup from a fragile one. The data is. You need split-and-dividend-adjusted price history that still contains delisted companies, point-in-time fundamentals if the signal touches them, and borrow data for the short side. A universe built only from today’s survivors tests a strategy that never had to hold a loser to the end.