Every systematic edge has a shelf life, and none of them announces the expiry date. The code keeps running, the orders keep going to the exchange, the logs stay clean — and the results quietly get worse. The hard problem is not accepting that this happens. It is telling the difference between a bad run your own validated history already predicted and a real deterioration that will not come back, using a rule you wrote before you had any money at stake. Get that rule wrong in one direction and you abandon a working system at its worst moment. Get it wrong in the other and you keep feeding a dead one.
What does strategy decay actually mean?
Decay is a persistent, structural fall in what a strategy produces, relative to what its own validated history said to expect. The two halves of that sentence both matter.
'Persistent and structural' rules out a run of losses. Losing runs are not evidence of anything on their own — they are the ordinary texture of a probabilistic process, and a system that has never had one has almost certainly not been run long enough.
'Relative to its own validated history' is the part people skip. Without a stated expectation, every drawdown feels like decay, because a drawdown always feels worse from inside it than it looked in a table. The expectation has to exist on paper before the drawdown arrives, or you will be judging a system with the one instrument that is guaranteed to be biased: your mood while it is losing.
Why do systematic edges erode?
Four mechanisms, and they are unrelated to each other. Applying the response that suits one of them to a situation caused by another is the most common expensive mistake in this part of the work.
- Crowding
- An edge is a payment for doing something other participants are not doing. As more capital finds the same trade, the payment shrinks — and the first place it shows up is execution, not results. Your fills sit further from the signal price, partial fills become common, and the gap between the assumed cost in your study and the cost you actually pay widens. Watch the execution log before the equity, because execution degrades first.
- Regime change
- The conditions the method needs are simply absent. A breakout system in a long directionless range is not broken; it is in the environment that hurts it, and that environment will end. Nothing about the rule has changed and nothing about it needs to. This is the case where doing nothing is the correct, and hardest, action.
- Market-structure change
- Something in the plumbing was revised. A derivative contract's lot size or expiry day changes, the settlement cycle moves, session timings shift, a tick size or a price band is revised, or SEBI's framework for how retail algorithms reach the exchange is amended. These arrive as a dated break rather than a slow fade, which is what makes them diagnosable — you can go and look at what changed that week.
- It was never there
- The fourth possibility, and the one nobody checks first: there was no edge, you fitted noise, and live trading is the first honest sample the idea has ever met. Nothing decayed, because nothing existed. This is more common than the other three combined, and it is why the diagnosis has to start with the study rather than with the market.
Watch out — If live results never resembled the study — not 'worse than', but different in character from the very first block of trades — treat the fourth explanation as the leading one until you have ruled it out. Genuine decay usually looks like a system that worked and then stopped. A fitted artefact usually looks like a system that never started.
Why is the diagnosis harder than the decision?
Once you know which of the four you are facing, what to do is nearly obvious. Getting to that knowledge is the whole difficulty, because the evidence available in the moment — a sequence of disappointing outcomes — is consistent with all four.
So the diagnosis cannot be made from the recent result. It has to be made against something fixed. That something is a band: the range of behaviour your validation work already said to expect, extracted and written down before a rupee was committed. Inside the band, a bad stretch is information you already had. Outside it, and with enough decisions behind it to mean anything, you have something new.
What has to come out of the validation work before you go live?
The band is not something you can construct after the drawdown starts, because by then you know the answer you want. It is an output of the walk-forward and out-of-sample work covered earlier in this module, recorded at the time.
- The deepest peak-to-trough decline the strategy produced across every validation window — and the second deepest, since one exceptional stretch is a weak reference point.
- The longest run of consecutive losing decisions, and the longest stretch between new highs measured in decisions rather than in calendar time.
- How wide the spread of outcomes was across walk-forward windows. A method whose windows disagreed sharply has a wide band, and a wide band demands a larger sample before you may conclude anything.
- The market conditions each window contained, in plain words — so you can later ask whether the conditions the method needs were even present.
- The cost assumptions used, so a live shortfall can be attributed to worse fills rather than to a vanished edge.
- The number of variants tried before this one was kept. A survivor of two hundred attempts needs a far larger live sample before you take any live result seriously.
Note — Time gaps here are counted in decisions, not days. A method that trades twice a month and one that trades twice a session reach the same sample at very different points on the calendar, and the calendar is what panic reads from.
How much evidence before you are allowed to have an opinion?
Set a minimum sample, in advance, below which no verdict may be reached at all — not 'watch closely', not 'reduce a little', nothing. This is the single most useful sentence in the whole procedure, because it removes the first six weeks of live trading from the argument entirely.
There is no universal number. What you can do is make the number a consequence of your own study rather than a preference: if your validation windows disagreed a lot, you need more decisions before live results separate from noise; if they agreed closely, fewer.
Minimum decisions before judging = longest validated losing run × safety multiple- Longest validated losing run — the worst consecutive-loss stretch your walk-forward work actually produced.
- Safety multiple — a number greater than one that you fix in writing before going live. It is a statement about how much noise you are prepared to sit through, not a tuned parameter.
- The result is a count of decisions, never a number of weeks. Convert to calendar time only to set a review date in your diary.
Example — Illustrative only. If the worst losing run across your validation windows was eleven consecutive decisions and you fix the multiple at three, you may not form a view before thirty-three live decisions. That number is now a rule, and the point of the rule is that it was set while nothing was going wrong.
What does a normal drawdown look like next to genuine decay?
Read the last row first. Most of the arguments people have with themselves about decay are settled by noticing they have not collected enough decisions to be having the argument.
| A drawdown your history predicted | Genuine decay | |
|---|---|---|
| Depth relative to the band | Uncomfortable, but inside the range the validation windows produced | Persistently outside it, and not returning |
| Shape | Sharp down, then a recovery attempt; the pattern repeats | A grinding drift, or a dated step-change with a cause you can name |
| Execution quality | Fills roughly where the study assumed | Fills consistently further away; more partials, more slippage than modelled |
| Conditions it needs | Absent — the environment simply does not suit the method right now | Present, and it is still not working. This is the damning combination |
| Behaviour of related instruments | Similar names show the same difficulty | Only your implementation is struggling — which points at execution or a bug |
| Sample behind the judgement | Below your stated minimum, so no verdict is permitted | Above it, in the same direction, across more than one review block |
What does crowding look like before it reaches the results?
Crowding is the one cause that gives you warning, and the warning arrives in the execution log rather than in the equity. That is worth knowing, because the execution log is the part nobody reads.
The signature is a widening gap between the price your study assumed and the price you actually got. Fills land further from the signal price. Partial fills become routine on sizes that used to complete in one go. The queue position you used to get at the open stops being available. On less liquid NSE names, the depth simply is not there at the moment your rule wants it, and your own order becomes a visible part of the move.
None of this shows up as a losing trade at first. It shows up as a slightly worse average, distributed across every decision, which is exactly the shape that hides inside ordinary variation until it has been accumulating for months. If you record the assumed price and the achieved price on every order — one extra column — the drift becomes visible far earlier than any drawdown threshold would catch it.
Pro tip — Log realised slippage per order against the price your study assumed, and review the rolling average of that one number monthly. It is the cheapest early-warning instrument in systematic trading, and it distinguishes 'the edge is being competed away' from 'the environment does not suit the method' without waiting for the results to decide.
Why is retuning the parameters the wrong response?
The instinct when live disappoints is to adjust — widen the stop, lengthen the lookback, add a filter that would have skipped the worst trades. It feels like maintenance. It is refitting.
What you are doing is selecting parameters using data that includes the recent losses. That is exactly the procedure that produces an overfitted study, except now it is being run on a smaller sample, under emotional pressure, without a held-out set, and without any of the discipline you applied during research. Every one of the protections built earlier in this module has been suspended precisely when the temptation to suspend them is strongest.
There is a further problem: once you retune, you no longer know what you are running. The live record before the change and after it are records of two different strategies, so your sample resets to zero at the worst possible moment. A system that gets retuned every time it underperforms has never been allowed to be wrong for long enough to be evaluated at all.
Watch out — Retuning because live results disappointed is overfitting wearing maintenance clothing. If a change is genuinely warranted, it belongs in the research process — new hypothesis, fresh validation, held-out data — not in a live configuration file edited on a Friday evening.
Is this the same thing as an AI model drifting?
No, and the distinction is worth holding clearly because the two get merged constantly.
The AI module's treatment of model drift is about a dependency you rent: a hosted model is changed by its provider, so the same prompt meets a different system and the output quietly shifts. You did not change anything; the tool moved underneath you.
This topic is the other case entirely. Your model has not moved at all. The parameters are exactly the ones you fitted, the code is byte-identical, and the market it was fitted to is the thing that changed. Nothing in your stack is faulty. The diagnostic question is different too: for a hosted-model change you re-run a fixed regression prompt against a frozen input and compare; for strategy decay you compare live behaviour against a band your own validation produced. Same symptom of 'it got worse without erroring', two unrelated causes, two unrelated tests.
What should the response ladder look like?
Write the ladder before you need it. Each rung pairs an observed condition with a response you have committed to, so that the decision under pressure is a lookup rather than a judgement.
How do you switch off without abandoning at the worst moment?
- 1
Reduce before you stop
The first response to a band breach with sufficient sample is halving the allocation, not switching off. It is reversible, it caps the damage while the diagnosis is still open, and it keeps the strategy generating the decisions you need in order to reach a verdict at all. A system you switched off produces no further evidence.
- 2
Separate the four causes deliberately
Check execution quality against the study's cost assumptions. Check whether the conditions the method needs were present. Check whether anything in market structure changed on a specific date. Only after those three does 'the edge has gone' become the leading explanation.
- 3
Check the plumbing before the premise
A stale feed, a corporate action not adjusted for, a renamed field, an order type your broker now handles differently — these produce exactly the same disappointing results as decay, and they are far more common. Confirm the system is doing what you think before you conclude the market changed.
- 4
Switch off as a state, not a verdict
Off does not mean the idea was worthless. It means the allocation is zero while the question is open. Keep the strategy running in paper mode against live prices so it continues to accumulate a record you can read later.
- 5
Send it back to research, not to the config file
If the idea is worth rescuing, it re-enters the research workflow at the hypothesis stage with the live period as fresh out-of-sample data. It does not get a quick parameter change and a fresh allocation.
- 6
Write the post-mortem while it stings
One page: what you expected, what happened, which of the four causes the evidence supports, what you would need to see before it is switched back on, and what you would do differently. Written six months later, this is fiction. Written the same week, it is the most valuable document in your research log.
How often should you look at any of this?
Monitoring fails for boring reasons. It gets proposed as a daily discipline, survives a fortnight, and then quietly stops — which is the same failure mode as the thing it was meant to catch.
So separate the two tiers. The scheduled tier runs whether or not anything seems wrong, because nothing will seem wrong until it is late. The triggered tier runs on events known to precede a break. Neither replaces the other: scheduled checks catch slow erosion, triggered checks catch discrete breaks.
Anchor the scheduled review to something that already happens rather than to a date you have to remember — the first weekend of the month, or the session after expiry week ends. And keep a log entry for the reviews where nothing changed, because a run of unremarkable entries is precisely what makes an eventual change legible, and lets you bracket when it began.
| Tier | When it runs | What it asks |
|---|---|---|
| Daily, automated only | Every session, by the monitoring you already built | Did the system do what it was told — orders sent, fills received, no unhandled errors? |
| Monthly, by hand | On a fixed anchor, whether or not anything feels wrong | Rolling slippage against the study's assumption; decisions accumulated; position against the band |
| Quarterly, by hand | After the quarter closes | Were the conditions this method needs actually present? This is the regime question, and no software answers it |
| Triggered | On a market-structure change, a broker change, or a single error found by accident | Has anything in the plumbing or the contract specification changed since the study was built? |
Why is the psychology the part that breaks?
Both failure modes are emotional, and they are opposite.
Holding on too long is sunk cost dressed as conviction. You built it, you validated it, you told someone about it. Switching it off feels like admitting the months were wasted, so you find reasons the current period is unrepresentative — and every drawdown supplies those reasons for free.
Switching off too early is capitulation timed by discomfort. The pain of a drawdown peaks near its worst point, which is exactly where the pre-committed ladder is hardest to follow and exactly where abandoning is most expensive. If the decision is being made because the loss feels intolerable, the real problem is position size, not the strategy — and that problem does not get solved by switching anything off.
The defence against both is the same, and it is structural rather than emotional: the threshold, the minimum sample and the response were all fixed while nothing was at stake. Under pressure you are not deciding. You are executing a decision your calmer self already made.
Pro tip — Ask one question when you feel the urge to intervene: has anything changed except the recent outcomes? If the answer is no, you are responding to a sequence of results, which is the one input your rule was specifically written to ignore.
What invalidates this framework?
This procedure has limits, and they should be stated as plainly as the procedure itself.
- A band built from a study that was itself overfitted describes nothing. If the validation was weak, everything downstream of it — including the thresholds you are so carefully obeying — is decoration.
- A strategy that trades rarely may never reach a meaningful sample within a horizon you care about. The honest response is a smaller allocation from the start, not a faster verdict.
- A dated market-structure change breaks the comparison outright. Behaviour before and after are not the same series, and a band spanning both is measuring two different things.
- Costs that drifted, rather than an edge that vanished, produce an identical picture. If brokerage, impact or taxes changed, the study's assumptions must be updated before any conclusion is drawn.
- A bug is indistinguishable from decay in the results and completely different in cause. The plumbing check is not optional.
- The framework assumes you actually recorded the validated expectations. If they were never written down, you do not have a band — you have a memory, and memory drifts toward whatever you currently believe.
The pre-commitment checklist
Write these down before the strategy takes its first live order
- The minimum number of decisions below which no verdict may be reached.
- The band: the worst validated drawdown, the longest validated losing run, and the spread across walk-forward windows.
- The threshold at which the allocation is halved, stated as a condition and not as a feeling.
- The threshold at which the strategy is switched off, and the further block of decisions required before it is applied.
- The conditions the method needs, in plain words, so you can later check whether they were present.
- The execution quality the study assumed, so a shortfall can be attributed correctly.
- The explicit statement that parameters will not be changed live, under any drawdown, for any reason.
- The review date in your diary, anchored to something that already happens every month.
Key points
Pro tip — Put the switch-off rule in the same file as the strategy code, dated, in a comment block at the top, and read it before every review. A threshold kept in your head is not a threshold — it is a preference that will renegotiate itself the moment the drawdown makes it inconvenient. This page is education, not advice, and nothing here is a recommendation to run, stop or fund any strategy.
Frequently asked questions
How do I know if my trading strategy has stopped working?
You cannot know it from recent results alone, because a losing run is consistent with a healthy system. Compare live behaviour against the range your own validation work produced — the worst drawdown, the longest losing run, the spread across walk-forward windows — and require a minimum number of decisions before any verdict is permitted. Persistently outside that range, with enough decisions behind it, is evidence. Anything less is noise you already knew to expect.
What is the difference between a drawdown and strategy decay?
A drawdown is a decline your validated history already predicted; decay is a persistent fall relative to that history. The practical tests are depth against the band, whether the conditions the method needs were even present, whether execution quality degraded, and whether enough decisions have accumulated to judge. A method in the environment that hurts it is out of season, not broken.
Should I change my strategy parameters when live results disappoint?
No. Selecting parameters using data that includes the recent losses is refitting, run on a smaller sample under emotional pressure with none of the safeguards you applied in research. It also resets your live sample to zero, because the record before and after the change describes two different strategies. A genuine change belongs in the research process with fresh validation, not in a live configuration file.
Why do algorithmic trading edges disappear over time?
Four unrelated reasons. Crowding, as more capital finds the same trade and execution degrades first. Regime change, where the rule is unchanged but the conditions it needs are absent. Market-structure change, such as a revised lot size, expiry, settlement cycle or session timing on NSE and BSE, which arrives as a dated break. And the fourth, that the edge was never present and the study fitted noise.
When should I switch off an algo strategy completely?
When it has stayed outside its expected band across more than one review block, with sufficient sample, after you have ruled out a data or execution fault and confirmed the conditions the method needs were actually present. Reduce the allocation first rather than stopping outright — reduction is reversible and keeps generating the decisions you need to reach a verdict.
Is strategy decay the same as AI model drift?
No. Model drift in a hosted AI tool means the provider changed the model behind your prompt, so the same input meets a different system. Strategy decay means your model is unchanged and the market it was fitted to has moved. The symptom — quietly worse output with no error — is shared, but the causes and the diagnostic tests are unrelated.
How long should I run a strategy before judging it?
Count decisions rather than weeks. Set the minimum in advance as a multiple of the longest losing run your validation work actually produced, and treat that count as a rule. A system trading twice a month and one trading twice a session reach the same sample at very different points on the calendar, and the calendar is what impatience reads from.