Intermediate5-8 min readTopic 12 of 20

    Reading a Backtest Report

    Rohit Singh

    Mr. Chartist · SEBI RA

    Module Progress
    0/20
    Module

    A backtest report is a set of claims dressed as a set of facts. Every headline statistic on it — compound growth, maximum drawdown, Sharpe, trade count, exposure — is a summary that throws away far more than it keeps, and each one can be technically correct while being deeply misleading. This page is about the second half of that sentence: what each number actually measures, what it deliberately hides, and the specific question that makes a weak report admit what it is.

    The value column is deliberately blank. Every headline statistic is a claim, and each one has a question attached that decides whether it means anything.A backtest summary shown as a list of statistic names with no values attached. Beside each name is the question that decides whether the statistic means anything: the period behind a compound growth rate, the duration rather than depth of a drawdown, the sampling frequency and distribution assumption behind a Sharpe ratio, whether the trade count is large enough to separate skill from luck, what share of the period capital was deployed, what survives the removal of the largest few contributors, and whether Indian transaction costs were modelled at all.A summary page is a set of claims, not a set of factsWHAT THE REPORT PRINTSWHAT YOU ASK BEFORE BELIEVING ITCAGR?Over which years? Which regime did it live in?Max drawdown?How long was it under water, not just how deep?Sharpe ratio?Daily or monthly? Which risk-free rate? Gaussian?Number of trades?Enough to distinguish an edge from a run of luck?Exposure?What share of the period was capital actually at work?Top contributors?What remains once the best few trades are removed?Costs?Brokerage, STT, stamp, GST, slippage — modelled or wished away?The value column is deliberately blank. This page teaches interrogation, not benchmarks — no figure here is a result, ours or anyone's.
    The value column is deliberately blank. Every headline statistic is a claim, and each one has a question attached that decides whether it means anything.

    Why can a backtest report be accurate and still useless?

    Because every statistic on it is a compression. A whole decade of decisions, thousands of candles and hundreds of positions get squeezed into eight numbers, and compression by definition discards information. The question is never whether the number is arithmetically right — it usually is. The question is what it discarded on the way.

    This matters most when someone else hands you the report. A vendor, a course seller, a forum post, an AI-generated strategy write-up: all of them produce reports that are individually defensible and collectively unreadable unless you know what to ask.

    The habit worth building is simple. For every number on the page, ask two things: **over what period was this measured, and what would make it look different?** Almost every weak result falls apart under those two questions alone.

    Note — This page is education, not advice. It deliberately quotes no benchmark values for any statistic — not ours, not anyone's — because a 'good' number depends entirely on the strategy type, the instrument, the period and the costs charged, and publishing a threshold would invite exactly the lazy comparison this page argues against.

    What does CAGR actually measure, and what does it hide?

    Compound annual growth rate collapses a whole test into one smoothed annual figure. Two things immediately follow.

    First, **it is a property of the period, not of the strategy**. A long-only system measured across a period that contained a sustained NSE advance is being credited for the market's direction, not for its own rules. Change the start and end dates by a year in either direction and the number often changes more than any parameter change ever did.

    Second, **it says nothing about the path**. The same CAGR can describe a steady climb or a violent sequence you would have abandoned in month four. The path is what you actually have to live through, and CAGR is specifically the statistic that erases it.

    • **Ask for the exact start and end dates**, then ask what the same test shows starting one year earlier and one year later. A result that only exists on one window is a window, not a result.
    • **Ask what the underlying index did over the identical window.** A long-only equity system needs to be read against the market it was long in.
    • **Ask whether the figure is before or after costs and taxes.** For anything intraday on Indian markets this single question routinely decides whether the strategy exists.
    • **Ask whether it is compounded on a fixed capital base or on a base that grew.** The convention chosen changes the headline substantially.

    Why is drawdown duration more important than drawdown depth?

    Maximum drawdown is the deepest peak-to-trough fall in the account curve. It is the second number everyone quotes, and it is the wrong half of the picture.

    Depth is vertical. Duration is horizontal — the time from the old high to the day the old high is regained. Depth is what the report prints; duration is what you have to sit through, watching the account do nothing while the market makes new highs without you.

    Two systems can report the identical maximum drawdown while one recovers within weeks and the other stays under water for years. They are not comparable objects, and only the second one gets switched off in frustration two months before it would have recovered.

    Schematic underwater profile with no vertical scale. The vertical measure is the depth a report quotes; the horizontal measure — old high to old high regained — is the one that decides whether a human keeps running the system.A schematic underwater profile: the line sits at zero when the account is at a new high and drops below whenever it is not. Two measurements are marked on the same episode. Depth is the vertical distance to the lowest point, the single number a report usually quotes. Duration is the horizontal distance from the old high to the day the old high is regained, which is the part a trader actually experiences. The vertical axis carries no numbers and no returns are shown.The quoted number is vertical. The experience is horizontal.at a highdepth — the one line in the reportduration — old high to old high regainedTwo systems can print the identical maximum drawdown while one recovers within a few weeksand the other stays under water for years. Only the second one gets switched off in anger.Schematic shape with no vertical scale. Not an equity curve and not a result.
    Schematic underwater profile with no vertical scale. The vertical measure is the depth a report quotes; the horizontal measure — old high to old high regained — is the one that decides whether a human keeps running the system.

    How is maximum drawdown misread?

    What the report saysWhat it actually meansThe question that exposes it
    Maximum drawdownThe single worst episode in this specific history. It is a sample maximum — the worst thing that happened, not the worst thing that can happen.How many separate drawdowns exceeded half of it? One outlier reads very differently from a recurring pattern.
    Measured on closing valuesIntra-period pain is invisible. Positions that swung hard inside the measurement interval show up smoothed.Is this end-of-day, or does it capture intraday excursion?
    Recovery not statedThe report is silent on the part that breaks discipline.What was the longest time under water, and how many months in total were spent below a previous high?
    Drawdown on a growing baseA percentage fall late in a compounding test is a much larger rupee figure than the same percentage early on.What was this in rupees, at the capital actually deployed at the time?
    No leverage disclosureA modest drawdown at three times leverage is not a modest drawdown.Was any leverage or margin assumed, and were margin requirements modelled at all?

    What assumptions are buried inside the Sharpe ratio?

    Sharpe expresses return in excess of a risk-free rate per unit of volatility. It is genuinely useful and it is routinely quoted without any of the four things it depends on.

    The sampling frequency
    Sharpe computed on daily returns and then annualised is a different number from one computed on monthly returns, and the annualisation itself assumes returns are independent from period to period. Trading returns frequently are not.
    The risk-free rate used
    In an Indian context this should be a domestic short-term government rate over the same period. Reports built on foreign templates often silently use a different rate, or zero.
    The symmetry assumption
    Volatility punishes upside moves exactly as hard as downside ones. A strategy with occasional large gains is penalised for them, and a strategy that grinds out small gains and then suffers one catastrophic loss can carry a flattering Sharpe right up until the loss.
    The distribution assumption
    Sharpe is most meaningful when returns are roughly normal. Option-selling and other short-volatility profiles are the classic case where they are not — long stretches of small positive returns give low measured volatility, which is precisely the thing being mismeasured.

    Watch out — Ranking two strategies by Sharpe alone is safe only when their return distributions have a similar shape. Comparing a trend system that makes its money in a few large moves against a premium-collection system on Sharpe is comparing two different measurements that happen to share a name.

    How many trades do you need before a result means anything?

    This is the least glamorous question on the page and it invalidates more reports than any other.

    A backtest with a small number of trades is not a small measurement — it is not a measurement. Any statistic computed over it inherits enormous uncertainty, and the report will show none of that uncertainty because summary tables have no column for it.

    Worse, the trade count interacts with topic 11's problem. Every parameter you tuned had to be paid for out of the same limited set of trades. Forty trades and six parameters is not a validated system; it is a description of forty events.

    Independent evidence ≈ trades that were genuinely separate decisions
    • Two hundred trades taken across the same three-day event on sixty correlated NSE stocks is closer to a handful of independent observations than to two hundred.
    • Overlapping positions in the same sector, or a basket entered on one signal, count once — not once per leg.
    • The right question is not 'how many rows in the trade log' but 'how many times did the market independently agree with these rules'.

    Example — Illustrative, with invented round numbers and no real instrument: a report showing 400 trades looks statistically comfortable until you notice 340 of them were entered on eleven distinct days across a correlated basket. The effective sample is closer to eleven observations than to 400. This is arithmetic about structure, not a claim about any strategy's outcome.

    What does exposure tell you that the return does not?

    Exposure — time in market — is the share of the test period during which capital was actually deployed. It is frequently omitted, and it changes the reading of everything above it.

    A result produced while deployed a small fraction of the time carried far less risk than the same headline suggests, and the idle capital had an opportunity cost the report almost never charges against it. A result produced while fully invested throughout carried every gap, every policy day and every circuit in the period.

    Neither is better. But a headline return with no exposure figure beside it is an unanswerable claim, because you cannot tell how much risk bought it.

    Schematic timeline with no return figure attached. Shaded segments are periods with a position open; plain segments are time flat in cash. Two reports with identical headlines and different shading are describing different things.A single timeline broken into alternating segments. The shaded segments are the periods a position was open; the rest of the bar is time spent flat in cash. The point made alongside is that a headline return earned while deployed for a small fraction of the period is a very different object from the same return earned while fully invested throughout, both in risk taken and in what the idle capital could have been doing.A return earned in a quarter of the period is not the same objectstart of testend of testshaded = position open · plain = flat in cashWhy it changes the readingLow exposure means the capital sat idle for most of the period, so the risk actually carriedwas smaller than a headline return suggests — and the idle rupees had an opportunity costthe report almost never charges against them. High exposure means the opposite: you wereexposed to every gap, every policy day and every circuit for the whole test.Schematic timeline. No return figure is shown or implied.
    Schematic timeline with no return figure attached. Shaded segments are periods with a position open; plain segments are time flat in cash. Two reports with identical headlines and different shading are describing different things.

    What happens when you delete the best few trades?

    This is the single most revealing test you can run on someone else's report, and it takes one line of code on your own.

    Sort every trade by its contribution. Delete the top two or three. Now re-read every headline statistic.

    If the result survives, you are looking at a system that made its money across many decisions. If it collapses, the system was carried by a handful of episodes. That is not automatically a flaw — trend-following is built to work exactly that way and makes no secret of it — but it completely changes what the report is claiming. It is no longer a claim about many repeatable decisions; it is a claim about two or three rare conditions recurring, which is a far harder thing to bank on and a far longer wait.

    Schematic sorted contribution ladder with no scale and no rupee or per cent figure attached. The shape is the lesson: a few episodes usually dominate, and the tail has to carry the system once they are removed.A schematic ladder of per-trade contributions sorted from largest to smallest. The first three bars tower over the rest, and the remainder form a long flat tail. The test described alongside is to delete the largest two or three contributions and re-read every headline statistic: if the result collapses, the system was carried by a handful of episodes rather than by a repeatable behaviour. The vertical axis is unnumbered and no rupee or percentage figure is shown.Delete your best three trades, then re-read the reportREMOVE THESEthe long tail that has to carry the system on its ownA handful of episodes usually dominates a backtest.That is not automatically a flaw — trend systems are built that way —but it changes what the report is telling you. It is no longer a claimabout many repeatable decisions; it is a claim about two or three raremarket conditions recurring, which is a far harder thing to bank on.Sorted contribution per trade, largest first — schematic bars, no scale, no rupee or per centfigure attached to any of them. The shape is the lesson, not the heights.Illustrative only. Nothing here describes any actual strategy's outcome.
    Schematic sorted contribution ladder with no scale and no rupee or per cent figure attached. The shape is the lesson: a few episodes usually dominate, and the tail has to carry the system once they are removed.

    What is missing from almost every backtest report you will be shown?

    • **The full cost model.** Brokerage, STT, exchange transaction charges, SEBI turnover fees, stamp duty, GST and a slippage assumption — stated as numbers, not as 'costs included'. Costs are covered in full in Costs, Slippage & Taxation.
    • **Slippage realism at the traded size.** A backtest fills at the printed price. Real orders in mid-cap NSE names move the book, and the assumed size decides whether the strategy is transferable to your capital at all.
    • **The trade log itself.** Summary statistics without an inspectable trade list cannot be audited. Ask for it. The refusal is the answer.
    • **The instrument universe and how it was chosen.** Testing on today's index constituents bakes the answer in. Survivorship and corporate actions are covered in Market Data, Corporate Actions & Bias.
    • **Dividends, splits, bonuses and symbol changes.** Whether the price series was adjusted, and how, silently changes every result on an equity strategy.
    • **How many other variants were tested before this one.** The single most informative fact about any report, and the one you will never be shown unless you ask.
    • **The rolling picture.** A year-by-year or segment-by-segment breakdown, which is what turns 'it worked' into 'it worked here and not there'.
    • **A statement of what would falsify it.** A report that cannot say what would make the author abandon the strategy is marketing.

    How much do Indian costs change every number on the page?

    Costs do not reduce a result proportionally. They act per trade, which means their effect scales with turnover — and turnover is exactly what the headline statistics never show.

    A position-held-for-months system pays its costs a handful of times a year and barely notices them. A system turning over the book several times a week pays them on every leg, and the same rupee cost that was a rounding error for the first strategy is the entire difference between a workable and an unworkable result for the second. Two reports can quote comparable headline figures while one of them has quietly been handed a subsidy the other has not.

    On Indian markets the components are specific and each one has to be modelled, not gestured at: brokerage, Securities Transaction Tax, exchange transaction charges, SEBI turnover fees, stamp duty, GST on the charges, and slippage — which is not a fee at all but is usually larger than all of them combined.

    Example — Illustrative structure with invented round numbers and no real instrument: if a round trip costs ₹120 all-in and the average gross gain per trade is ₹400, then a system with 50 trades a year loses ₹6,000 to costs while one with 500 trades loses ₹60,000 — from an identical rule set. This is arithmetic about turnover, not a claim about any strategy's returns.

    What does a report worth reading look like?

    Turn the whole page around. Instead of what to distrust, here is the structure of a report that can actually be evaluated — which is also the structure you should be producing for yourself.

    SectionWhat it has to containWhy it is there
    SpecificationEntry, exit, sizing, universe and the exact date range, written before any result.Fixes what was tested, so nobody can quietly redefine it after the fact.
    Cost modelEvery Indian charge itemised, plus the slippage assumption in rupees or basis points.The single most common place a result is manufactured.
    Sample descriptionTrade count, exposure, and an honest note on how many trades were independent.Says how much evidence exists before saying what the evidence shows.
    Risk sectionDrawdown depth, longest time under water, and the rupee figure at deployed capital.Turns an abstract percentage into something a human can decide about.
    Robustness sectionParameter neighbours, the result with the top contributors removed, a different universe.Shows the result was attacked, not just produced.
    Research historyHow many variants were tested, and what was discarded.The context without which no statistic on the page can be interpreted.
    FalsificationWhat would make the author stop running it, written in advance.Separates a research document from a sales document.

    How do you interrogate a report someone else shows you?

    1. 1

      Read the period before the number

      Cover the statistics with your hand. Find the start date, the end date and the instrument universe first. Half the reports you are shown will already be explained by those three facts.

    2. 2

      Ask for exposure and trade count

      Neither is impressive, which is why neither is on the headline slide. Together they tell you how much risk was taken and how much evidence exists.

    3. 3

      Ask for drawdown duration, not just depth

      Specifically: the longest time under water. Watch whether the answer arrives immediately or has to be computed. That tells you whether it was ever looked at.

    4. 4

      Ask what the costs were, in rupees per trade

      A precise answer means a real model. A vague one means an assumption, and on Indian intraday strategies an optimistic cost assumption is usually the entire result.

    5. 5

      Ask how many variants were tested

      This is topic 11's question, and it is the one that decides whether the number in front of you is a measurement or the winner of a sweep.

    6. 6

      Ask to see the losing periods and the abandoned versions

      A research process that has never discarded anything has never tested anything. The discarded work is the evidence that the surviving work was actually examined.

    7. 7

      Ask what would make them switch it off

      An honest systematic trader has a pre-written answer — a drawdown limit, a duration limit, a behavioural change in the market. No answer means no plan.

    What makes a backtest report worthless outright?

    Any one of these is enough to discard the report entirely rather than adjust for it.

    • Costs are excluded, or described as 'nominal' without figures.
    • No trade log is available for inspection, and none can be produced.
    • The instrument universe was chosen with hindsight, or the price series was not corporate-action adjusted.
    • The period is short enough to contain one market condition and nothing else.
    • The trade count cannot support the number of parameters that were tuned.
    • The report is presented alongside an offer to sell you the strategy, a subscription or a course — the incentive to present the best of many runs is structural, and it does not require anyone to be dishonest.
    • A guaranteed, assured or risk-free outcome is implied anywhere. Under SEBI rules that framing is not permitted for a registered research analyst, and its presence tells you what kind of document you are holding.
    • The author cannot state what would falsify the strategy or when they would stop running it.

    Watch out — A backtest — including a perfectly honest one — is a statement about the past under a set of assumptions. It is never a projection, and no statistic on this page becomes a forecast by being computed carefully. Nothing here is a recommendation to build, buy, subscribe to or run any strategy.

    The checklist to run against any report

    Before you believe a backtest report

    • Start date, end date, instrument universe — written down before you read any statistic.
    • Costs stated as rupees per trade, with the slippage assumption named.
    • Exposure figure present, and read alongside the headline return.
    • Trade count, and an honest estimate of how many of those were independent decisions.
    • Drawdown depth AND longest time under water, in months.
    • Result re-read with the largest two or three contributions deleted.
    • Year-by-year or segment-by-segment breakdown, not just the aggregate.
    • The number of variants tested before this one was chosen.
    • The trade log, inspectable.
    • A written statement of what would falsify the strategy.

    Key points

    Every headline statistic is a compression — the useful question is what it discarded.
    CAGR is a property of the measurement period as much as of the strategy; always ask for exact dates.
    Drawdown depth is what is printed; drawdown duration is what gets a system switched off.
    Maximum drawdown is a sample maximum — the worst that happened, not the worst that can happen.
    Sharpe hides four assumptions: sampling frequency, risk-free rate, symmetry and distribution shape.
    Trade count means little until you ask how many of those trades were independent decisions.
    Without an exposure figure, a return is uninterpretable — you cannot see how much risk bought it.
    Delete the best two or three trades and re-read everything. That is the fastest structural test.
    The trade log, the cost model in rupees and the number of variants tested are the three things to demand.
    A backtest is a statement about the past under assumptions. It never becomes a projection.

    Pro tip — When someone shows you a backtest, cover the numbers with your hand and read only the period, the universe and the cost assumption first. Then ask one question before looking at anything else: 'how many versions did you test before this one?' The pause before the answer is more informative than the entire report.

    Frequently asked questions

    What is the most important number in a backtest report?

    None of them in isolation. The most informative single fact is usually not a statistic at all — it is the number of variants that were tested before the one being shown was chosen, because that decides whether the figures are a measurement or the top of a sweep. After that, the trade count and the exposure figure do more to tell you what a report means than the headline return does.

    What is the difference between maximum drawdown and drawdown duration?

    Maximum drawdown is the depth of the deepest peak-to-trough fall — a vertical measure. Drawdown duration is how long the account stayed below its previous high before regaining it — a horizontal one. Reports almost always quote the first and almost never the second, yet duration is the part a trader actually experiences and the reason most systems are abandoned. Two strategies with identical maximum drawdowns can be completely different things to live with.

    Is a high Sharpe ratio always better?

    No, because Sharpe assumes a return distribution that many trading strategies do not have. It treats upside volatility as risk, so a system that makes its money in a few large moves is penalised for exactly the behaviour that defines it. Strategies that produce long runs of small gains punctuated by rare large losses can carry a flattering Sharpe until the loss arrives. Comparing two strategies on Sharpe is only meaningful when their return shapes are similar, and the sampling frequency and risk-free rate used must be disclosed.

    How many trades does a backtest need to be credible?

    There is no universal number, and any source quoting one has skipped the part that matters: how many of those trades were independent decisions. Two hundred positions entered across a correlated basket on eleven distinct days is closer to eleven observations than two hundred. The count also has to be read against how many parameters were tuned — every parameter is paid for out of the same evidence.

    Why does exposure matter when reading a backtest?

    Exposure is the share of the test period during which capital was actually deployed. A headline figure earned while in the market a small fraction of the time carried much less risk than the same figure earned while fully invested, and the idle capital had an opportunity cost the report rarely charges. Without an exposure figure you cannot tell how much risk produced the result, which makes the result uninterpretable rather than merely incomplete.

    What happens to a backtest if you remove the best few trades?

    It is the fastest way to find out what the report is really claiming. Sort trades by contribution, delete the top two or three, and re-read every statistic. If the result survives, the system earned across many decisions. If it collapses, a handful of episodes carried it — which is normal for trend-following and does not automatically condemn the system, but it means you are betting on rare conditions recurring rather than on a frequently repeatable behaviour.

    Should I trust a backtest report from a strategy seller?

    Treat it as a marketing document until it is audited, regardless of the seller's sincerity. The incentive to present the best of many runs is structural and does not require anyone to lie. Ask for the full trade log, the exact period, the cost model in rupees per trade, the exposure figure and the number of variants tested. In India, note also that any assured, guaranteed or risk-free framing around returns is not permitted for a SEBI-registered research analyst, and its presence tells you what kind of document you are reading.