A backtest report is a set of claims dressed as a set of facts. Every headline statistic on it — compound growth, maximum drawdown, Sharpe, trade count, exposure — is a summary that throws away far more than it keeps, and each one can be technically correct while being deeply misleading. This page is about the second half of that sentence: what each number actually measures, what it deliberately hides, and the specific question that makes a weak report admit what it is.
Why can a backtest report be accurate and still useless?
Because every statistic on it is a compression. A whole decade of decisions, thousands of candles and hundreds of positions get squeezed into eight numbers, and compression by definition discards information. The question is never whether the number is arithmetically right — it usually is. The question is what it discarded on the way.
This matters most when someone else hands you the report. A vendor, a course seller, a forum post, an AI-generated strategy write-up: all of them produce reports that are individually defensible and collectively unreadable unless you know what to ask.
The habit worth building is simple. For every number on the page, ask two things: **over what period was this measured, and what would make it look different?** Almost every weak result falls apart under those two questions alone.
Note — This page is education, not advice. It deliberately quotes no benchmark values for any statistic — not ours, not anyone's — because a 'good' number depends entirely on the strategy type, the instrument, the period and the costs charged, and publishing a threshold would invite exactly the lazy comparison this page argues against.
What does CAGR actually measure, and what does it hide?
Compound annual growth rate collapses a whole test into one smoothed annual figure. Two things immediately follow.
First, **it is a property of the period, not of the strategy**. A long-only system measured across a period that contained a sustained NSE advance is being credited for the market's direction, not for its own rules. Change the start and end dates by a year in either direction and the number often changes more than any parameter change ever did.
Second, **it says nothing about the path**. The same CAGR can describe a steady climb or a violent sequence you would have abandoned in month four. The path is what you actually have to live through, and CAGR is specifically the statistic that erases it.
- **Ask for the exact start and end dates**, then ask what the same test shows starting one year earlier and one year later. A result that only exists on one window is a window, not a result.
- **Ask what the underlying index did over the identical window.** A long-only equity system needs to be read against the market it was long in.
- **Ask whether the figure is before or after costs and taxes.** For anything intraday on Indian markets this single question routinely decides whether the strategy exists.
- **Ask whether it is compounded on a fixed capital base or on a base that grew.** The convention chosen changes the headline substantially.
Why is drawdown duration more important than drawdown depth?
Maximum drawdown is the deepest peak-to-trough fall in the account curve. It is the second number everyone quotes, and it is the wrong half of the picture.
Depth is vertical. Duration is horizontal — the time from the old high to the day the old high is regained. Depth is what the report prints; duration is what you have to sit through, watching the account do nothing while the market makes new highs without you.
Two systems can report the identical maximum drawdown while one recovers within weeks and the other stays under water for years. They are not comparable objects, and only the second one gets switched off in frustration two months before it would have recovered.
How is maximum drawdown misread?
| What the report says | What it actually means | The question that exposes it |
|---|---|---|
| Maximum drawdown | The single worst episode in this specific history. It is a sample maximum — the worst thing that happened, not the worst thing that can happen. | How many separate drawdowns exceeded half of it? One outlier reads very differently from a recurring pattern. |
| Measured on closing values | Intra-period pain is invisible. Positions that swung hard inside the measurement interval show up smoothed. | Is this end-of-day, or does it capture intraday excursion? |
| Recovery not stated | The report is silent on the part that breaks discipline. | What was the longest time under water, and how many months in total were spent below a previous high? |
| Drawdown on a growing base | A percentage fall late in a compounding test is a much larger rupee figure than the same percentage early on. | What was this in rupees, at the capital actually deployed at the time? |
| No leverage disclosure | A modest drawdown at three times leverage is not a modest drawdown. | Was any leverage or margin assumed, and were margin requirements modelled at all? |
What assumptions are buried inside the Sharpe ratio?
Sharpe expresses return in excess of a risk-free rate per unit of volatility. It is genuinely useful and it is routinely quoted without any of the four things it depends on.
- The sampling frequency
- Sharpe computed on daily returns and then annualised is a different number from one computed on monthly returns, and the annualisation itself assumes returns are independent from period to period. Trading returns frequently are not.
- The risk-free rate used
- In an Indian context this should be a domestic short-term government rate over the same period. Reports built on foreign templates often silently use a different rate, or zero.
- The symmetry assumption
- Volatility punishes upside moves exactly as hard as downside ones. A strategy with occasional large gains is penalised for them, and a strategy that grinds out small gains and then suffers one catastrophic loss can carry a flattering Sharpe right up until the loss.
- The distribution assumption
- Sharpe is most meaningful when returns are roughly normal. Option-selling and other short-volatility profiles are the classic case where they are not — long stretches of small positive returns give low measured volatility, which is precisely the thing being mismeasured.
Watch out — Ranking two strategies by Sharpe alone is safe only when their return distributions have a similar shape. Comparing a trend system that makes its money in a few large moves against a premium-collection system on Sharpe is comparing two different measurements that happen to share a name.
How many trades do you need before a result means anything?
This is the least glamorous question on the page and it invalidates more reports than any other.
A backtest with a small number of trades is not a small measurement — it is not a measurement. Any statistic computed over it inherits enormous uncertainty, and the report will show none of that uncertainty because summary tables have no column for it.
Worse, the trade count interacts with topic 11's problem. Every parameter you tuned had to be paid for out of the same limited set of trades. Forty trades and six parameters is not a validated system; it is a description of forty events.
Independent evidence ≈ trades that were genuinely separate decisions- Two hundred trades taken across the same three-day event on sixty correlated NSE stocks is closer to a handful of independent observations than to two hundred.
- Overlapping positions in the same sector, or a basket entered on one signal, count once — not once per leg.
- The right question is not 'how many rows in the trade log' but 'how many times did the market independently agree with these rules'.
Example — Illustrative, with invented round numbers and no real instrument: a report showing 400 trades looks statistically comfortable until you notice 340 of them were entered on eleven distinct days across a correlated basket. The effective sample is closer to eleven observations than to 400. This is arithmetic about structure, not a claim about any strategy's outcome.
What does exposure tell you that the return does not?
Exposure — time in market — is the share of the test period during which capital was actually deployed. It is frequently omitted, and it changes the reading of everything above it.
A result produced while deployed a small fraction of the time carried far less risk than the same headline suggests, and the idle capital had an opportunity cost the report almost never charges against it. A result produced while fully invested throughout carried every gap, every policy day and every circuit in the period.
Neither is better. But a headline return with no exposure figure beside it is an unanswerable claim, because you cannot tell how much risk bought it.
What happens when you delete the best few trades?
This is the single most revealing test you can run on someone else's report, and it takes one line of code on your own.
Sort every trade by its contribution. Delete the top two or three. Now re-read every headline statistic.
If the result survives, you are looking at a system that made its money across many decisions. If it collapses, the system was carried by a handful of episodes. That is not automatically a flaw — trend-following is built to work exactly that way and makes no secret of it — but it completely changes what the report is claiming. It is no longer a claim about many repeatable decisions; it is a claim about two or three rare conditions recurring, which is a far harder thing to bank on and a far longer wait.
What is missing from almost every backtest report you will be shown?
- **The full cost model.** Brokerage, STT, exchange transaction charges, SEBI turnover fees, stamp duty, GST and a slippage assumption — stated as numbers, not as 'costs included'. Costs are covered in full in Costs, Slippage & Taxation.
- **Slippage realism at the traded size.** A backtest fills at the printed price. Real orders in mid-cap NSE names move the book, and the assumed size decides whether the strategy is transferable to your capital at all.
- **The trade log itself.** Summary statistics without an inspectable trade list cannot be audited. Ask for it. The refusal is the answer.
- **The instrument universe and how it was chosen.** Testing on today's index constituents bakes the answer in. Survivorship and corporate actions are covered in Market Data, Corporate Actions & Bias.
- **Dividends, splits, bonuses and symbol changes.** Whether the price series was adjusted, and how, silently changes every result on an equity strategy.
- **How many other variants were tested before this one.** The single most informative fact about any report, and the one you will never be shown unless you ask.
- **The rolling picture.** A year-by-year or segment-by-segment breakdown, which is what turns 'it worked' into 'it worked here and not there'.
- **A statement of what would falsify it.** A report that cannot say what would make the author abandon the strategy is marketing.
How much do Indian costs change every number on the page?
Costs do not reduce a result proportionally. They act per trade, which means their effect scales with turnover — and turnover is exactly what the headline statistics never show.
A position-held-for-months system pays its costs a handful of times a year and barely notices them. A system turning over the book several times a week pays them on every leg, and the same rupee cost that was a rounding error for the first strategy is the entire difference between a workable and an unworkable result for the second. Two reports can quote comparable headline figures while one of them has quietly been handed a subsidy the other has not.
On Indian markets the components are specific and each one has to be modelled, not gestured at: brokerage, Securities Transaction Tax, exchange transaction charges, SEBI turnover fees, stamp duty, GST on the charges, and slippage — which is not a fee at all but is usually larger than all of them combined.
Example — Illustrative structure with invented round numbers and no real instrument: if a round trip costs ₹120 all-in and the average gross gain per trade is ₹400, then a system with 50 trades a year loses ₹6,000 to costs while one with 500 trades loses ₹60,000 — from an identical rule set. This is arithmetic about turnover, not a claim about any strategy's returns.
What does a report worth reading look like?
Turn the whole page around. Instead of what to distrust, here is the structure of a report that can actually be evaluated — which is also the structure you should be producing for yourself.
| Section | What it has to contain | Why it is there |
|---|---|---|
| Specification | Entry, exit, sizing, universe and the exact date range, written before any result. | Fixes what was tested, so nobody can quietly redefine it after the fact. |
| Cost model | Every Indian charge itemised, plus the slippage assumption in rupees or basis points. | The single most common place a result is manufactured. |
| Sample description | Trade count, exposure, and an honest note on how many trades were independent. | Says how much evidence exists before saying what the evidence shows. |
| Risk section | Drawdown depth, longest time under water, and the rupee figure at deployed capital. | Turns an abstract percentage into something a human can decide about. |
| Robustness section | Parameter neighbours, the result with the top contributors removed, a different universe. | Shows the result was attacked, not just produced. |
| Research history | How many variants were tested, and what was discarded. | The context without which no statistic on the page can be interpreted. |
| Falsification | What would make the author stop running it, written in advance. | Separates a research document from a sales document. |
How do you interrogate a report someone else shows you?
- 1
Read the period before the number
Cover the statistics with your hand. Find the start date, the end date and the instrument universe first. Half the reports you are shown will already be explained by those three facts.
- 2
Ask for exposure and trade count
Neither is impressive, which is why neither is on the headline slide. Together they tell you how much risk was taken and how much evidence exists.
- 3
Ask for drawdown duration, not just depth
Specifically: the longest time under water. Watch whether the answer arrives immediately or has to be computed. That tells you whether it was ever looked at.
- 4
Ask what the costs were, in rupees per trade
A precise answer means a real model. A vague one means an assumption, and on Indian intraday strategies an optimistic cost assumption is usually the entire result.
- 5
Ask how many variants were tested
This is topic 11's question, and it is the one that decides whether the number in front of you is a measurement or the winner of a sweep.
- 6
Ask to see the losing periods and the abandoned versions
A research process that has never discarded anything has never tested anything. The discarded work is the evidence that the surviving work was actually examined.
- 7
Ask what would make them switch it off
An honest systematic trader has a pre-written answer — a drawdown limit, a duration limit, a behavioural change in the market. No answer means no plan.
What makes a backtest report worthless outright?
Any one of these is enough to discard the report entirely rather than adjust for it.
- Costs are excluded, or described as 'nominal' without figures.
- No trade log is available for inspection, and none can be produced.
- The instrument universe was chosen with hindsight, or the price series was not corporate-action adjusted.
- The period is short enough to contain one market condition and nothing else.
- The trade count cannot support the number of parameters that were tuned.
- The report is presented alongside an offer to sell you the strategy, a subscription or a course — the incentive to present the best of many runs is structural, and it does not require anyone to be dishonest.
- A guaranteed, assured or risk-free outcome is implied anywhere. Under SEBI rules that framing is not permitted for a registered research analyst, and its presence tells you what kind of document you are holding.
- The author cannot state what would falsify the strategy or when they would stop running it.
Watch out — A backtest — including a perfectly honest one — is a statement about the past under a set of assumptions. It is never a projection, and no statistic on this page becomes a forecast by being computed carefully. Nothing here is a recommendation to build, buy, subscribe to or run any strategy.
The checklist to run against any report
Before you believe a backtest report
- Start date, end date, instrument universe — written down before you read any statistic.
- Costs stated as rupees per trade, with the slippage assumption named.
- Exposure figure present, and read alongside the headline return.
- Trade count, and an honest estimate of how many of those were independent decisions.
- Drawdown depth AND longest time under water, in months.
- Result re-read with the largest two or three contributions deleted.
- Year-by-year or segment-by-segment breakdown, not just the aggregate.
- The number of variants tested before this one was chosen.
- The trade log, inspectable.
- A written statement of what would falsify the strategy.
Key points
Pro tip — When someone shows you a backtest, cover the numbers with your hand and read only the period, the universe and the cost assumption first. Then ask one question before looking at anything else: 'how many versions did you test before this one?' The pause before the answer is more informative than the entire report.
Frequently asked questions
What is the most important number in a backtest report?
None of them in isolation. The most informative single fact is usually not a statistic at all — it is the number of variants that were tested before the one being shown was chosen, because that decides whether the figures are a measurement or the top of a sweep. After that, the trade count and the exposure figure do more to tell you what a report means than the headline return does.
What is the difference between maximum drawdown and drawdown duration?
Maximum drawdown is the depth of the deepest peak-to-trough fall — a vertical measure. Drawdown duration is how long the account stayed below its previous high before regaining it — a horizontal one. Reports almost always quote the first and almost never the second, yet duration is the part a trader actually experiences and the reason most systems are abandoned. Two strategies with identical maximum drawdowns can be completely different things to live with.
Is a high Sharpe ratio always better?
No, because Sharpe assumes a return distribution that many trading strategies do not have. It treats upside volatility as risk, so a system that makes its money in a few large moves is penalised for exactly the behaviour that defines it. Strategies that produce long runs of small gains punctuated by rare large losses can carry a flattering Sharpe until the loss arrives. Comparing two strategies on Sharpe is only meaningful when their return shapes are similar, and the sampling frequency and risk-free rate used must be disclosed.
How many trades does a backtest need to be credible?
There is no universal number, and any source quoting one has skipped the part that matters: how many of those trades were independent decisions. Two hundred positions entered across a correlated basket on eleven distinct days is closer to eleven observations than two hundred. The count also has to be read against how many parameters were tuned — every parameter is paid for out of the same evidence.
Why does exposure matter when reading a backtest?
Exposure is the share of the test period during which capital was actually deployed. A headline figure earned while in the market a small fraction of the time carried much less risk than the same figure earned while fully invested, and the idle capital had an opportunity cost the report rarely charges. Without an exposure figure you cannot tell how much risk produced the result, which makes the result uninterpretable rather than merely incomplete.
What happens to a backtest if you remove the best few trades?
It is the fastest way to find out what the report is really claiming. Sort trades by contribution, delete the top two or three, and re-read every statistic. If the result survives, the system earned across many decisions. If it collapses, a handful of episodes carried it — which is normal for trend-following and does not automatically condemn the system, but it means you are betting on rare conditions recurring rather than on a frequently repeatable behaviour.
Should I trust a backtest report from a strategy seller?
Treat it as a marketing document until it is audited, regardless of the seller's sincerity. The incentive to present the best of many runs is structural and does not require anyone to lie. Ask for the full trade log, the exact period, the cost model in rupees per trade, the exposure figure and the number of variants tested. In India, note also that any assured, guaranteed or risk-free framing around returns is not permitted for a SEBI-registered research analyst, and its presence tells you what kind of document you are reading.