Pairs trading takes two instruments whose prices have historically moved in a stable relationship, sells the one that has become expensive relative to the other, buys the other, and waits for the gap to close. The appeal is obvious: if both legs are equally exposed to the market, the direction of the NIFTY stops mattering and only the relationship does. The danger is equally obvious once stated plainly — the entire position is a bet that a gap which has widened will stop widening, and nothing in the arithmetic tells you when it will not.
What is pairs trading, in one paragraph?
You hold two positions at once: long one instrument, short another, sized so that the market exposure of the two roughly cancels. What is left after that cancellation is called the spread — a single synthetic number that goes up when the long leg outperforms and down when it underperforms.
The strategy has nothing to say about whether either company is good, cheap or growing. It is a claim about one number: that the spread has a level it keeps returning to, and that it is currently far from it.
Statistical arbitrage is the same idea run across many such relationships at once rather than one at a time. It is not arbitrage in the risk-free sense — nothing here locks in anything. The word describes the statistical framing, not a guaranteed outcome.
Note — This page is education, not advice. It deliberately names no pair of real NSE stocks. Naming two live tickers as a tradeable relationship would be a recommendation published under a research analyst registration, which is not what a learning page is for.
Why is correlation not cointegration?
This is the sentence that decides whether a retail pairs strategy has any chance at all, and it is the one almost every tutorial gets wrong.
Correlation measures whether two series tend to move in the same direction over some window. Cointegration is a much stronger and much rarer claim: that some fixed combination of the two prices is stationary — that the gap between them has a mean, and keeps coming back to it.
Two stocks in the same sector, both riding the same broad market, will show a high correlation number almost mechanically. That number is entirely compatible with a gap that widens every single month and never closes. A pairs trade placed on correlation alone is a bet on a relationship nobody ever tested for.
- Correlation
- A measure of co-movement, computed over a chosen window. It is unstable — the same two names can show a high number over 60 candles and a low one over 250. It says nothing at all about whether a gap closes.
- Cointegration
- A property of the combination, not of the returns. If A and B are cointegrated, then A minus h×B is stationary: it has a mean, a spread around that mean, and it revisits it. This is the property a pairs trade actually requires.
- Stationary
- A series whose statistical behaviour does not drift — its mean and variance stay roughly constant. A trending price series is not stationary. A well-behaved spread is. Testing for stationarity is the whole test.
Watch out — A high correlation coefficient is not a screening result. It is the first thing that goes wrong. If your candidate list was built by sorting a correlation matrix, you have selected for co-movement and tested for nothing.
What exactly is the spread?
The spread is a single constructed series — one number per candle — that you then treat exactly as you would treat any other price series you were mean-reverting.
The hedge ratio is what makes the construction meaningful. It is the number of units of B you hold against one unit of A so that the combined position has as little directional exposure as possible. Estimating it is a modelling decision, not a detail: use the wrong ratio and the series you are calling a spread still carries a large slice of plain market direction, so you are running a directional trade while telling yourself you are market-neutral.
Spread(t) = Price A(t) − h × Price B(t)- Price A, Price B — the two instrument prices at the same timestamp, adjusted identically for splits, bonuses and dividends. An unadjusted price series will manufacture a spread break out of a corporate action that changed nothing economically.
- h — the hedge ratio: how many units of B are held against one of A. Usually estimated by regressing A on B over a lookback window.
- The lookback window used to estimate h is a choice you must record. Re-estimating it every candle and re-estimating it once a quarter produce different strategies, not different settings of one strategy.
How does the z-score framing work?
A raw spread in rupees is not comparable across pairs, so the spread is usually restated in standard deviations from its own recent mean. That restated number is the z-score, and the entry and exit rules are written against it.
The convention you will see everywhere is: enter when the z-score is beyond some band, close when it returns to zero, and abandon the position at some wider band. What matters far more than the numbers is that every one of them is a choice with consequences, and that the last one — the abandon level — is the only thing standing between you and an unbounded loss.
What are the parameters, and what does each one actually decide?
Six choices, and they interact. Changing the lookback changes what counts as a two-standard-deviation move, which changes the trade count, which changes how much any single event matters. This is why pairs strategies are unusually easy to over-fit: there are enough dials that some combination will always look tidy on the history you tested it on.
| Choice | What it really controls | The failure it creates |
|---|---|---|
| Lookback N for the mean and standard deviation | How much history defines 'normal' for this spread | Short N makes any new level look normal within weeks — the z-score quietly adapts to a broken relationship instead of flagging it. |
| Entry band | How stretched the spread must be before you act | A wide band gives you very few trades, so the result rests on a handful of events. A narrow band buries you in costs. |
| Exit level | What counts as the relationship having normalised | Exiting exactly at zero assumes the mean is knowable. It is an estimate that moves. |
| Stop / abandon band | The point at which you accept the relationship is gone | Leaving this out is the single decision that turns a bounded idea into an unbounded one. |
| Maximum holding period | How long you are willing to be wrong before exiting anyway | Without it, a dead pair sits in the book indefinitely consuming margin and attention. |
| Re-estimation frequency for h | Whether the hedge ratio tracks the relationship or is frozen | Re-estimating too often chases noise; freezing it lets a genuine structural change go unnoticed. |
Why does screening a large universe find so many pairs that are not real?
Take 200 instruments and you can form nearly 20,000 candidate pairs. Test every one of them for a statistical property at a conventional significance level and a large number will pass by chance alone. That is not a flaw in the test; it is what the significance level means.
The result is a shortlist that looks impressively rigorous and is mostly noise. Worse, the pairs that pass are the ones whose historical spread happened to look tidiest, which is precisely the selection you would make if you were trying to over-fit deliberately.
This is why the economic reason has to come first. A pair you can justify before testing narrows the search to a handful of candidates, so a passing test carries some information. A pair discovered by sweeping 20,000 combinations carries almost none, no matter how good the number looks.
Watch out — If your process is 'test everything, keep what passes', then the number of things you tested is part of the result and must be recorded alongside it. Nine tests with one survivor is a different finding from one test with one survivor, and the research log is where that distinction survives contact with your memory six months later.
Why do costs hit this strategy harder than most?
A round trip in a pairs trade is four executions, not two: open both legs, close both legs. Each one carries brokerage, exchange transaction charges, SEBI turnover fees, GST on the charges, stamp duty on the buy side, and STT where applicable to the segment and side. Then add slippage on all four, financing or borrow cost on the short leg, and rollover costs on every expiry if the short leg is a futures position.
Set against that, the per-trade edge a spread strategy targets is small by design — it is trying to capture a normalisation, not a trend. Small edge, high trade count and four executions per round trip is the exact profile that costs erase first, which is why costs have to be itemised per trade in the research rather than estimated as an average at the end.
Why do pairs break?
Because the relationship was a description of the past, not a law. Two things end it.
The first is that it was never structural. Two names moved together because they shared a sector, a commodity input or a rate cycle, and for a stretch of time that was enough. Nothing bound them. When the shared driver stopped operating, the co-movement stopped, and no test run on the earlier data could have known.
The second is that the relationship was real and then something ended it — a demerger, a merger, a large equity issuance, a regulatory change that lands on one leg and not the other, or a slow divergence in what the two businesses actually do. In each case the historical spread stops being a description of anything current.
Why is 'the spread always reverts' the assumption that produces an unbounded loss?
Consider what the rule does as the spread widens. At two standard deviations you enter. At three, the position is losing and the model says the signal is stronger. At four, stronger still. If the rule includes any form of adding to the position as the z-score grows — and many published versions do — the position gets larger precisely as the evidence that the relationship is dead accumulates.
Now add the short leg. A long position's worst case is that the instrument goes to zero: the loss is large, but bounded and knowable in advance. A short position has no such ceiling. There is no price above which an instrument cannot go. If the leg you are short re-rates upward on a merger, an order win or an index inclusion, the loss on that leg is limited by nothing except when you close it.
So the honest statement is: pairs trading takes a series of small, frequent, comfortable gains and pays for them with rare, open-ended losses. That trade can be perfectly reasonable — but only if you have decided in advance, in writing, at what point you stop believing the relationship exists.
Watch out — Averaging into a widening spread and 'the spread has to come back eventually' are the same sentence. The first is a position-sizing rule, the second is a belief, and neither has any mechanism behind it once the relationship that produced the reversion has ended.
Why is the short leg the hard part in the Indian market?
Everything above is generic. This is where an Indian pairs strategy meets reality, and it is a constraint on the strategy itself, not an execution detail to be handled later.
In the cash segment on NSE and BSE, a short position taken without borrowed stock must be squared off within the same session. It cannot be carried overnight. A spread that takes several candles to normalise — which is the entire premise — therefore cannot be held in the cash market at all on the short side.
That leaves two routes. Single-stock futures let you carry a short, but only for names in the derivatives-eligible universe, which is a small fraction of the listed market and changes as SEBI and the exchanges revise eligibility. Or you borrow the stock through the securities lending and borrowing mechanism, where availability, cost and depth vary by name and by day, and where a recall is a risk you do not control.
What does that constraint do to a textbook pairs strategy?
- The universe collapses first. If both legs must be shortable overnight, your candidate list is effectively the derivatives-eligible names — and both legs must be in it, not one.
- Contract mechanics intrude. Futures have expiries, lot sizes and a basis that moves for reasons unrelated to your spread. A lot size fixes the minimum position, which may be far larger than the size your risk rule allows.
- The hedge ratio becomes lumpy. A ratio of 1.37 is trivial in cash and awkward in whole lots. Rounding to the nearest lot leaves residual directional exposure you must acknowledge rather than ignore.
- Rollover is a recurring, forced execution event. Every expiry you must close and reopen both legs, paying costs and slippage on a schedule the strategy did not choose.
- Margin is charged on both legs and can be revised intraday. A widening spread raises the margin requirement at exactly the moment the position is losing.
- Borrow can disappear. In the lending route, a recall forces you out of the short leg on someone else's timetable, leaving you holding a naked long you never intended.
- Corporate actions must be handled on both legs, in the price history and in the live position. A bonus issue on one leg alone will fabricate a spread break in an unadjusted series.
Note — Contract specifications, eligibility criteria, margin rules and lending mechanics are set by SEBI and the exchanges and are revised periodically. Verify the current rules on the NSE and SEBI websites before designing around any of them — do not take a number from an article, including this one.
What would a serious research process for a pair look like?
- 1
Start with a reason, not a screen
Write the economic mechanism first: why should these two prices be tied together? Two firms exposed to the same input cost, the same regulator, the same demand cycle. If you cannot write that sentence, the statistics will only tell you what happened to co-move.
- 2
Check the relationship is even tradeable here
Before any maths — can both legs be shorted and carried for the intended horizon? If not, the research is finished. Doing this last is how people build a strategy they cannot execute.
- 3
Build the spread on properly adjusted data
Both series adjusted identically for splits, bonuses and dividends, aligned on timestamps, with holidays and suspensions handled explicitly rather than forward-filled by accident.
- 4
Test the spread for stationarity, out of sample
Estimate the hedge ratio on one window and test the spread's behaviour on a later window you did not look at. A spread that is stationary only on the data used to fit it has told you nothing.
- 5
Write the abandon rule before the entry rule
The level, or the holding period, or the event, at which you stop believing the relationship exists. This is the rule that makes the loss bounded, and it must be written when you are calm rather than when the position is losing.
- 6
Subtract the real costs of both legs
Brokerage twice, STT, exchange and SEBI charges, stamp duty, GST, financing or borrow cost on the short, rollover on every expiry, and slippage on four executions per round trip. Spread strategies are high-frequency in trade count and low in per-trade edge, which is exactly the profile costs destroy first.
- 7
Size it as one position, not two
A pairs trade is a single risk with two legs. Sizing each leg independently understates what you are actually exposed to when the relationship breaks and both legs move against you at once.
What invalidates a pairs trade?
Every method has to state what would prove it wrong. For a spread trade, these are the conditions that end it — some of them before the stop level is ever reached.
- The spread passing your written abandon band. Not a re-examination, an exit. The band exists because the judgement made at that moment would be worse than the one made in advance.
- A corporate action on either leg — demerger, merger, large issuance, scheme of arrangement. The instrument on one side is no longer the one the relationship was measured on.
- A regulatory or tax change that lands on one leg and not the other. The shared driver has stopped being shared.
- The hedge ratio moving materially on re-estimation. If h has drifted a long way, the historical spread was not measuring the thing you thought it was measuring.
- The stationarity test failing on recent data. The property the whole trade rests on is testable, and it can stop holding.
- The maximum holding period expiring. A spread that has not normalised in the horizon you allowed it is evidence, not bad luck.
- Borrow being recalled, or the name leaving the derivatives-eligible list. The trade is no longer executable as designed, whatever the statistics say.
- You explaining to yourself why this particular divergence is different. That sentence is the reliable early warning that the abandon rule is about to be overridden.
Should a retail systematic trader in India run pairs at all?
It is a legitimate question and the honest answer is: only after several other things are in place.
The strategy is unusually demanding. It needs two clean, corporate-action-adjusted price series; a short leg that can be carried; costs on four executions per round trip; margin on both legs; and the discipline to close a position that the model says is getting more attractive. It is also unusually easy to over-fit, because the number of dials is high and the number of independent events in any test is low.
What it is genuinely good for, even if you never trade it, is teaching the difference between a relationship you have described and a relationship that exists. That distinction transfers to every other systematic idea you will ever test.
Note — Nothing on this page is a recommendation to buy, sell or short any security, or to run any strategy. Market-neutral does not mean risk-neutral: a spread position can lose on both legs at once, and the short leg's loss has no upper bound.
Key points
Pro tip — Before you write a single line of the entry rule, write the abandon rule and the maximum holding period, and write down the economic reason the two prices should be tied together in the first place. If you cannot state that reason in one sentence without using the word 'correlated', you do not have a pair — you have two charts that happened to move together on the data you looked at.
Frequently asked questions
What is the difference between correlation and cointegration in pairs trading?
Correlation says two series tended to move in the same direction over a chosen window. Cointegration says a specific combination of the two prices is stationary — that the gap between them has a mean it keeps returning to. Two stocks in the same sector will usually show high correlation simply because both follow the broad market, and that is entirely compatible with a gap that widens forever. Only the cointegration property is what a pairs trade needs, and it must be tested for directly rather than inferred from a correlation number.
How is the hedge ratio in a pairs trade calculated?
The common approach is to regress one price series on the other over a chosen lookback window and use the slope as the number of units of the second instrument held against one unit of the first. The lookback length and how often it is re-estimated are modelling decisions that materially change the strategy, so both should be recorded as part of the rule rather than treated as settings. A poorly estimated ratio leaves directional market exposure in a position you believe is neutral.
Can you do pairs trading in the Indian stock market?
The mechanics work, but the short leg is the binding constraint. A short position in the cash segment on NSE or BSE has to be squared off within the same session unless the stock is borrowed, so a multi-candle spread trade generally has to be run through single-stock futures — which limits the universe to the derivatives-eligible list — or through the securities lending and borrowing route, where availability and cost vary by name. Verify current eligibility, margin and lending rules on the NSE and SEBI websites, since they are revised periodically.
What is the z-score used for in a spread trade?
It restates the spread in standard deviations from its own recent mean, so that a gap can be compared across pairs and across time instead of being read in rupees. Entry, exit and abandon rules are then written against that number. The bands people quote — commonly two standard deviations to enter and zero to exit — are conventions used to illustrate the mechanic, not recommended parameters, and they depend entirely on the lookback window used to compute the mean.
Why do pairs trading strategies fail?
Usually because the relationship was never structural — it was co-movement produced by a shared sector or a shared rate cycle, with nothing binding the two prices together. The rest fail because something real ended the relationship: a demerger, a merger, a large equity issuance, a regulatory change hitting one leg, or a gradual divergence in the two businesses. In both cases the spread leaves its historical range and does not come back, while a rule written around 'enter when stretched' reads the widening as an increasingly strong signal.
Is pairs trading market neutral and therefore low risk?
Market neutral is not the same as low risk. It means the position aims to remove broad market direction, not that it removes loss. The specific risk is that both legs move against you at once when the relationship ends, and because one leg is short, that loss has no upper bound — there is no price above which an instrument cannot go. A written abandon level and a maximum holding period are what make the exposure bounded; the market-neutral construction does not.
What is statistical arbitrage, and is it actually arbitrage?
Statistical arbitrage is the same mean-reversion idea applied across many relationships at once rather than a single pair, usually with position sizes set by a model across the whole book. It is not arbitrage in the risk-free sense: nothing is locked in, no price discrepancy is guaranteed to close, and the outcome depends on relationships continuing to hold. The word describes the statistical framing, not the certainty of the result.