Advanced8-12 min readTopic 19 of 20

    Strategy Decay & When to Switch Off

    Rohit Singh

    Mr. Chartist · SEBI RA

    Module Progress
    0/20
    Module

    Every systematic edge has a shelf life, and none of them announces the expiry date. The code keeps running, the orders keep going to the exchange, the logs stay clean — and the results quietly get worse. The hard problem is not accepting that this happens. It is telling the difference between a bad run your own validated history already predicted and a real deterioration that will not come back, using a rule you wrote before you had any money at stake. Get that rule wrong in one direction and you abandon a working system at its worst moment. Get it wrong in the other and you keep feeding a dead one.

    Four different reasons an edge stops paying. They need different responses, which is why guessing between them is how a sound rule gets retuned to death.Four panels. Crowding, where enough participants find the same trade and fills worsen before results do. Regime change, where the rule is unchanged but the conditions it needs are absent. Market-structure change, such as a lot size, settlement cycle, session timing or tick size revision, which arrives as a dated break rather than a slow fade. And the fourth, that the edge was never present and the study fitted noise.Why an edge stops payingthe fourth is the one nobody checks first1CrowdingEnough participants find the same tradeFills worsen before results do2Regime changeThe conditions it needs are absentRule is fine; the environment left3Market structureLot size, settlement, session, tickA dated break, not a slow fade4It was never thereYou fitted noise and called it edgeLive never resembled the studyThe four need different responses. Guessing which one you have is how a good rule gets retuned to death.
    Four different reasons an edge stops paying. They need different responses, which is why guessing between them is how a sound rule gets retuned to death.

    What does strategy decay actually mean?

    Decay is a persistent, structural fall in what a strategy produces, relative to what its own validated history said to expect. The two halves of that sentence both matter.

    'Persistent and structural' rules out a run of losses. Losing runs are not evidence of anything on their own — they are the ordinary texture of a probabilistic process, and a system that has never had one has almost certainly not been run long enough.

    'Relative to its own validated history' is the part people skip. Without a stated expectation, every drawdown feels like decay, because a drawdown always feels worse from inside it than it looked in a table. The expectation has to exist on paper before the drawdown arrives, or you will be judging a system with the one instrument that is guaranteed to be biased: your mood while it is losing.

    Why do systematic edges erode?

    Four mechanisms, and they are unrelated to each other. Applying the response that suits one of them to a situation caused by another is the most common expensive mistake in this part of the work.

    Crowding
    An edge is a payment for doing something other participants are not doing. As more capital finds the same trade, the payment shrinks — and the first place it shows up is execution, not results. Your fills sit further from the signal price, partial fills become common, and the gap between the assumed cost in your study and the cost you actually pay widens. Watch the execution log before the equity, because execution degrades first.
    Regime change
    The conditions the method needs are simply absent. A breakout system in a long directionless range is not broken; it is in the environment that hurts it, and that environment will end. Nothing about the rule has changed and nothing about it needs to. This is the case where doing nothing is the correct, and hardest, action.
    Market-structure change
    Something in the plumbing was revised. A derivative contract's lot size or expiry day changes, the settlement cycle moves, session timings shift, a tick size or a price band is revised, or SEBI's framework for how retail algorithms reach the exchange is amended. These arrive as a dated break rather than a slow fade, which is what makes them diagnosable — you can go and look at what changed that week.
    It was never there
    The fourth possibility, and the one nobody checks first: there was no edge, you fitted noise, and live trading is the first honest sample the idea has ever met. Nothing decayed, because nothing existed. This is more common than the other three combined, and it is why the diagnosis has to start with the study rather than with the market.

    Watch out — If live results never resembled the study — not 'worse than', but different in character from the very first block of trades — treat the fourth explanation as the leading one until you have ruled it out. Genuine decay usually looks like a system that worked and then stopped. A fitted artefact usually looks like a system that never started.

    Why is the diagnosis harder than the decision?

    Once you know which of the four you are facing, what to do is nearly obvious. Getting to that knowledge is the whole difficulty, because the evidence available in the moment — a sequence of disappointing outcomes — is consistent with all four.

    So the diagnosis cannot be made from the recent result. It has to be made against something fixed. That something is a band: the range of behaviour your validation work already said to expect, extracted and written down before a rupee was committed. Inside the band, a bad stretch is information you already had. Outside it, and with enough decisions behind it to mean anything, you have something new.

    Schematic only — no scale, and not an equity curve. One track dips hard but stays inside the range the validated history already predicted; the other leaves it. Only the second one is evidence.A schematic band, drawn with no numeric scale, represents the range of behaviour the strategy's own validated history already predicts. One live track wanders inside the band throughout, including a deep dip, and is labelled hold. A second track drifts steadily lower and leaves the band, and the crossing point is marked as the first evidence of decay. A vertical line early on marks the minimum number of decisions below which no verdict may be reached at all.Is this the drawdown you already expected?rolling quality - schematic, no scaledecisions taken, not days elapsedwhat the validated history already predictsno verdict before this many decisionsleaves the band - first evidenceINSIDE THE BAND: HOLDA dip is only information once it falls outside the range the strategy's own history had already priced in.
    Schematic only — no scale, and not an equity curve. One track dips hard but stays inside the range the validated history already predicted; the other leaves it. Only the second one is evidence.

    What has to come out of the validation work before you go live?

    The band is not something you can construct after the drawdown starts, because by then you know the answer you want. It is an output of the walk-forward and out-of-sample work covered earlier in this module, recorded at the time.

    • The deepest peak-to-trough decline the strategy produced across every validation window — and the second deepest, since one exceptional stretch is a weak reference point.
    • The longest run of consecutive losing decisions, and the longest stretch between new highs measured in decisions rather than in calendar time.
    • How wide the spread of outcomes was across walk-forward windows. A method whose windows disagreed sharply has a wide band, and a wide band demands a larger sample before you may conclude anything.
    • The market conditions each window contained, in plain words — so you can later ask whether the conditions the method needs were even present.
    • The cost assumptions used, so a live shortfall can be attributed to worse fills rather than to a vanished edge.
    • The number of variants tried before this one was kept. A survivor of two hundred attempts needs a far larger live sample before you take any live result seriously.

    Note — Time gaps here are counted in decisions, not days. A method that trades twice a month and one that trades twice a session reach the same sample at very different points on the calendar, and the calendar is what panic reads from.

    How much evidence before you are allowed to have an opinion?

    Set a minimum sample, in advance, below which no verdict may be reached at all — not 'watch closely', not 'reduce a little', nothing. This is the single most useful sentence in the whole procedure, because it removes the first six weeks of live trading from the argument entirely.

    There is no universal number. What you can do is make the number a consequence of your own study rather than a preference: if your validation windows disagreed a lot, you need more decisions before live results separate from noise; if they agreed closely, fewer.

    Minimum decisions before judging = longest validated losing run × safety multiple
    • Longest validated losing run — the worst consecutive-loss stretch your walk-forward work actually produced.
    • Safety multiple — a number greater than one that you fix in writing before going live. It is a statement about how much noise you are prepared to sit through, not a tuned parameter.
    • The result is a count of decisions, never a number of weeks. Convert to calendar time only to set a review date in your diary.

    Example — Illustrative only. If the worst losing run across your validation windows was eleven consecutive decisions and you fix the multiple at three, you may not form a view before thirty-three live decisions. That number is now a rule, and the point of the rule is that it was set while nothing was going wrong.

    What does a normal drawdown look like next to genuine decay?

    Read the last row first. Most of the arguments people have with themselves about decay are settled by noticing they have not collected enough decisions to be having the argument.

    A drawdown your history predictedGenuine decay
    Depth relative to the bandUncomfortable, but inside the range the validation windows producedPersistently outside it, and not returning
    ShapeSharp down, then a recovery attempt; the pattern repeatsA grinding drift, or a dated step-change with a cause you can name
    Execution qualityFills roughly where the study assumedFills consistently further away; more partials, more slippage than modelled
    Conditions it needsAbsent — the environment simply does not suit the method right nowPresent, and it is still not working. This is the damning combination
    Behaviour of related instrumentsSimilar names show the same difficultyOnly your implementation is struggling — which points at execution or a bug
    Sample behind the judgementBelow your stated minimum, so no verdict is permittedAbove it, in the same direction, across more than one review block

    What does crowding look like before it reaches the results?

    Crowding is the one cause that gives you warning, and the warning arrives in the execution log rather than in the equity. That is worth knowing, because the execution log is the part nobody reads.

    The signature is a widening gap between the price your study assumed and the price you actually got. Fills land further from the signal price. Partial fills become routine on sizes that used to complete in one go. The queue position you used to get at the open stops being available. On less liquid NSE names, the depth simply is not there at the moment your rule wants it, and your own order becomes a visible part of the move.

    None of this shows up as a losing trade at first. It shows up as a slightly worse average, distributed across every decision, which is exactly the shape that hides inside ordinary variation until it has been accumulating for months. If you record the assumed price and the achieved price on every order — one extra column — the drift becomes visible far earlier than any drawdown threshold would catch it.

    Pro tip — Log realised slippage per order against the price your study assumed, and review the rolling average of that one number monthly. It is the cheapest early-warning instrument in systematic trading, and it distinguishes 'the edge is being competed away' from 'the environment does not suit the method' without waiting for the results to decide.

    Why is retuning the parameters the wrong response?

    The instinct when live disappoints is to adjust — widen the stop, lengthen the lookback, add a filter that would have skipped the worst trades. It feels like maintenance. It is refitting.

    What you are doing is selecting parameters using data that includes the recent losses. That is exactly the procedure that produces an overfitted study, except now it is being run on a smaller sample, under emotional pressure, without a held-out set, and without any of the discipline you applied during research. Every one of the protections built earlier in this module has been suspended precisely when the temptation to suspend them is strongest.

    There is a further problem: once you retune, you no longer know what you are running. The live record before the change and after it are records of two different strategies, so your sample resets to zero at the worst possible moment. A system that gets retuned every time it underperforms has never been allowed to be wrong for long enough to be evaluated at all.

    Watch out — Retuning because live results disappointed is overfitting wearing maintenance clothing. If a change is genuinely warranted, it belongs in the research process — new hypothesis, fresh validation, held-out data — not in a live configuration file edited on a Friday evening.

    Is this the same thing as an AI model drifting?

    No, and the distinction is worth holding clearly because the two get merged constantly.

    The AI module's treatment of model drift is about a dependency you rent: a hosted model is changed by its provider, so the same prompt meets a different system and the output quietly shifts. You did not change anything; the tool moved underneath you.

    This topic is the other case entirely. Your model has not moved at all. The parameters are exactly the ones you fitted, the code is byte-identical, and the market it was fitted to is the thing that changed. Nothing in your stack is faulty. The diagnostic question is different too: for a hosted-model change you re-run a fixed regression prompt against a frozen input and compare; for strategy decay you compare live behaviour against a band your own validation produced. Same symptom of 'it got worse without erroring', two unrelated causes, two unrelated tests.

    What should the response ladder look like?

    Write the ladder before you need it. Each rung pairs an observed condition with a response you have committed to, so that the decision under pressure is a lookup rather than a judgement.

    Condition on the left, pre-committed response on the right — and the one response that is banned at any rung. Written while nothing is going wrong, this is a rule; written during a drawdown, it is a mood.Four rungs, each pairing an observed condition with the response committed to in advance: below the minimum sample, change nothing; inside the expectation band, change nothing but log it; outside the band with sufficient sample, halve the allocation; still outside after a further block of decisions, switch off and return the idea to research. Below them, drawn with a dashed caution border, sits the one response that is banned — retuning parameters because live results disappointed.Decide this while nothing is going wrong1Below the minimum sampleThere is no verdict to reach yetChange nothing at all2Inside the band, drawdown deepeningThis is the case the study anticipatedChange nothing; log it3Outside the band, sample sufficientReversible, and it buys reading timeHalve the allocation4Still outside after a further blockOff is a state, not a verdictSwitch off; return to researchNot on the ladderRetuning the parameters because live disappointed. That is fitting to the newest data, wearing maintenance clothing.condition on the left, committed response on the right
    Condition on the left, pre-committed response on the right — and the one response that is banned at any rung. Written while nothing is going wrong, this is a rule; written during a drawdown, it is a mood.

    How do you switch off without abandoning at the worst moment?

    1. 1

      Reduce before you stop

      The first response to a band breach with sufficient sample is halving the allocation, not switching off. It is reversible, it caps the damage while the diagnosis is still open, and it keeps the strategy generating the decisions you need in order to reach a verdict at all. A system you switched off produces no further evidence.

    2. 2

      Separate the four causes deliberately

      Check execution quality against the study's cost assumptions. Check whether the conditions the method needs were present. Check whether anything in market structure changed on a specific date. Only after those three does 'the edge has gone' become the leading explanation.

    3. 3

      Check the plumbing before the premise

      A stale feed, a corporate action not adjusted for, a renamed field, an order type your broker now handles differently — these produce exactly the same disappointing results as decay, and they are far more common. Confirm the system is doing what you think before you conclude the market changed.

    4. 4

      Switch off as a state, not a verdict

      Off does not mean the idea was worthless. It means the allocation is zero while the question is open. Keep the strategy running in paper mode against live prices so it continues to accumulate a record you can read later.

    5. 5

      Send it back to research, not to the config file

      If the idea is worth rescuing, it re-enters the research workflow at the hypothesis stage with the live period as fresh out-of-sample data. It does not get a quick parameter change and a fresh allocation.

    6. 6

      Write the post-mortem while it stings

      One page: what you expected, what happened, which of the four causes the evidence supports, what you would need to see before it is switched back on, and what you would do differently. Written six months later, this is fiction. Written the same week, it is the most valuable document in your research log.

    How often should you look at any of this?

    Monitoring fails for boring reasons. It gets proposed as a daily discipline, survives a fortnight, and then quietly stops — which is the same failure mode as the thing it was meant to catch.

    So separate the two tiers. The scheduled tier runs whether or not anything seems wrong, because nothing will seem wrong until it is late. The triggered tier runs on events known to precede a break. Neither replaces the other: scheduled checks catch slow erosion, triggered checks catch discrete breaks.

    Anchor the scheduled review to something that already happens rather than to a date you have to remember — the first weekend of the month, or the session after expiry week ends. And keep a log entry for the reviews where nothing changed, because a run of unremarkable entries is precisely what makes an eventual change legible, and lets you bracket when it began.

    TierWhen it runsWhat it asks
    Daily, automated onlyEvery session, by the monitoring you already builtDid the system do what it was told — orders sent, fills received, no unhandled errors?
    Monthly, by handOn a fixed anchor, whether or not anything feels wrongRolling slippage against the study's assumption; decisions accumulated; position against the band
    Quarterly, by handAfter the quarter closesWere the conditions this method needs actually present? This is the regime question, and no software answers it
    TriggeredOn a market-structure change, a broker change, or a single error found by accidentHas anything in the plumbing or the contract specification changed since the study was built?

    Why is the psychology the part that breaks?

    Both failure modes are emotional, and they are opposite.

    Holding on too long is sunk cost dressed as conviction. You built it, you validated it, you told someone about it. Switching it off feels like admitting the months were wasted, so you find reasons the current period is unrepresentative — and every drawdown supplies those reasons for free.

    Switching off too early is capitulation timed by discomfort. The pain of a drawdown peaks near its worst point, which is exactly where the pre-committed ladder is hardest to follow and exactly where abandoning is most expensive. If the decision is being made because the loss feels intolerable, the real problem is position size, not the strategy — and that problem does not get solved by switching anything off.

    The defence against both is the same, and it is structural rather than emotional: the threshold, the minimum sample and the response were all fixed while nothing was at stake. Under pressure you are not deciding. You are executing a decision your calmer self already made.

    Pro tip — Ask one question when you feel the urge to intervene: has anything changed except the recent outcomes? If the answer is no, you are responding to a sequence of results, which is the one input your rule was specifically written to ignore.

    What invalidates this framework?

    This procedure has limits, and they should be stated as plainly as the procedure itself.

    • A band built from a study that was itself overfitted describes nothing. If the validation was weak, everything downstream of it — including the thresholds you are so carefully obeying — is decoration.
    • A strategy that trades rarely may never reach a meaningful sample within a horizon you care about. The honest response is a smaller allocation from the start, not a faster verdict.
    • A dated market-structure change breaks the comparison outright. Behaviour before and after are not the same series, and a band spanning both is measuring two different things.
    • Costs that drifted, rather than an edge that vanished, produce an identical picture. If brokerage, impact or taxes changed, the study's assumptions must be updated before any conclusion is drawn.
    • A bug is indistinguishable from decay in the results and completely different in cause. The plumbing check is not optional.
    • The framework assumes you actually recorded the validated expectations. If they were never written down, you do not have a band — you have a memory, and memory drifts toward whatever you currently believe.

    The pre-commitment checklist

    Write these down before the strategy takes its first live order

    • The minimum number of decisions below which no verdict may be reached.
    • The band: the worst validated drawdown, the longest validated losing run, and the spread across walk-forward windows.
    • The threshold at which the allocation is halved, stated as a condition and not as a feeling.
    • The threshold at which the strategy is switched off, and the further block of decisions required before it is applied.
    • The conditions the method needs, in plain words, so you can later check whether they were present.
    • The execution quality the study assumed, so a shortfall can be attributed correctly.
    • The explicit statement that parameters will not be changed live, under any drawdown, for any reason.
    • The review date in your diary, anchored to something that already happens every month.

    Key points

    Decay is a persistent fall relative to a strategy's own validated history — not a run of losses.
    The four causes are crowding, regime change, market-structure change, and an edge that was never there.
    'It was never there' is the most common of the four and the one checked last.
    The expectation band must be extracted from validation and written down before any money is committed.
    Set a minimum number of decisions below which no verdict of any kind may be reached.
    Count decisions, never days — trade frequency decides how fast evidence accumulates.
    Crowding shows up in execution quality before it shows up anywhere else.
    Retuning parameters because live disappointed is overfitting in maintenance clothing, and it resets your sample to zero.
    Halve the allocation before switching off: reduction is reversible and keeps producing evidence.
    Off is a state, not a verdict. The idea returns to research at the hypothesis stage, not to the config file.
    Check the plumbing — feed, corporate actions, order handling — before concluding the market changed.
    Both failure modes are emotional and opposite; the defence is a rule written while nothing was at stake.

    Pro tip — Put the switch-off rule in the same file as the strategy code, dated, in a comment block at the top, and read it before every review. A threshold kept in your head is not a threshold — it is a preference that will renegotiate itself the moment the drawdown makes it inconvenient. This page is education, not advice, and nothing here is a recommendation to run, stop or fund any strategy.

    Frequently asked questions

    How do I know if my trading strategy has stopped working?

    You cannot know it from recent results alone, because a losing run is consistent with a healthy system. Compare live behaviour against the range your own validation work produced — the worst drawdown, the longest losing run, the spread across walk-forward windows — and require a minimum number of decisions before any verdict is permitted. Persistently outside that range, with enough decisions behind it, is evidence. Anything less is noise you already knew to expect.

    What is the difference between a drawdown and strategy decay?

    A drawdown is a decline your validated history already predicted; decay is a persistent fall relative to that history. The practical tests are depth against the band, whether the conditions the method needs were even present, whether execution quality degraded, and whether enough decisions have accumulated to judge. A method in the environment that hurts it is out of season, not broken.

    Should I change my strategy parameters when live results disappoint?

    No. Selecting parameters using data that includes the recent losses is refitting, run on a smaller sample under emotional pressure with none of the safeguards you applied in research. It also resets your live sample to zero, because the record before and after the change describes two different strategies. A genuine change belongs in the research process with fresh validation, not in a live configuration file.

    Why do algorithmic trading edges disappear over time?

    Four unrelated reasons. Crowding, as more capital finds the same trade and execution degrades first. Regime change, where the rule is unchanged but the conditions it needs are absent. Market-structure change, such as a revised lot size, expiry, settlement cycle or session timing on NSE and BSE, which arrives as a dated break. And the fourth, that the edge was never present and the study fitted noise.

    When should I switch off an algo strategy completely?

    When it has stayed outside its expected band across more than one review block, with sufficient sample, after you have ruled out a data or execution fault and confirmed the conditions the method needs were actually present. Reduce the allocation first rather than stopping outright — reduction is reversible and keeps generating the decisions you need to reach a verdict.

    Is strategy decay the same as AI model drift?

    No. Model drift in a hosted AI tool means the provider changed the model behind your prompt, so the same input meets a different system. Strategy decay means your model is unchanged and the market it was fitted to has moved. The symptom — quietly worse output with no error — is shared, but the causes and the diagnostic tests are unrelated.

    How long should I run a strategy before judging it?

    Count decisions rather than weeks. Set the minimum in advance as a multiple of the longest losing run your validation work actually produced, and treat that count as a rule. A system trading twice a month and one trading twice a session reach the same sample at very different points on the calendar, and the calendar is what impatience reads from.