Intermediate5-8 min readTopic 4 of 20

    The Systematic Research Workflow

    Rohit Singh

    Mr. Chartist · SEBI RA

    Module Progress
    0/20
    Module

    A system and a hunch can look identical on a chart. What separates them is not the idea — it is the process that produced it. A hunch is an observation you went looking for evidence to support. A system is a hypothesis you wrote down first, tested once, and were willing to throw away. The difference shows up nowhere in the rules themselves and everywhere in the record behind them, which is why the research log matters more than any single result it contains.

    Seven stages, each with a kill gate. Failing a stage ends the idea; it does not send it back for another parameter. The only loop that is allowed runs back to the hypothesis.Seven narrowing stages run down the page, from observation to a small live allocation. Every stage has an arrow leading right into a shared bin marked killed, because failing a stage ends the idea rather than sending it forward. A separate arrow runs down the left margin from the forward-test stage back up to the hypothesis stage.One idea, seven gates1Observationsomething looks repeatable2Written hypothesiswith a named cause3Crude testone slice of data4Full in-sample testwith costs subtracted5Out-of-sample checkdata you never looked at6Paper / forward runlive prices, no money7Small live, then reviewsmallest size that is realKILLEDlogged, notforgottenback to the hypothesis, not to the parametersA stage you cannot fail is not a gate. Write the kill condition before you run the test.
    Seven stages, each with a kill gate. Failing a stage ends the idea; it does not send it back for another parameter. The only loop that is allowed runs back to the hypothesis.

    Why does the order of operations decide whether research is honest?

    There are only two possible sequences, and they produce results that look the same on paper.

    Sequence one: you state what you expect and why, then you look. If the data disagrees, the idea is dead and you have learned something cheaply.

    Sequence two: you look, find something that stands out, and then explain why it must be true. This feels identical from the inside. It produces a story that fits the data perfectly, because it was built out of that data. And it tells you nothing about any session that has not happened yet.

    Given enough freedom in what you look for, something in a price series will always stand out. Markets contain enormous amounts of noise, and noise has structure in it by definition. A process that lets you search until you are satisfied is guaranteed to find something, whether or not anything is there.

    Watch out — The tell is simple. If you cannot state, right now and in writing, what result would have made you abandon the idea, you did not test it. You searched until it looked good, and searching until it looks good is not evidence — it is the absence of evidence, dressed as a finding.

    What does a research hypothesis look like?

    A hypothesis is not 'let me see if this works'. It has four parts, and it is written before any data is touched.

    The claim
    A statement precise enough to be false. 'A daily close above a 30-candle range tends to be followed by continuation over the next 10 candles' is a claim. 'Breakouts work' is not, because nothing could contradict it.
    The candidate cause
    Why this would be true about how buyers and sellers behave. Supply that was overhead has been absorbed; participants who were trapped have been relieved; a position had to be adjusted for reasons unrelated to price. Without a cause you have a coincidence you have not disproved yet.
    The kill condition
    The result that ends the idea, written before you look. This is the part that turns a search into a test, and it is the part almost everyone omits.
    The scope
    Which instruments, which timeframe, which window of data, and which costs are subtracted. An idea that is only stated for NIFTY 500 daily bars cannot later be claimed for intraday Bank NIFTY without a fresh test.

    Note — The candidate cause is not decoration. It is the thing that makes the difference between an effect that might persist and a pattern that existed in one window because something had to. A finding with no plausible mechanism should be held to a far higher bar, not a lower one.

    What are the seven stages of the research workflow?

    Each stage narrows the funnel, and each has a gate the idea can fail. A stage you cannot fail is not a gate — it is a formality.

    1. 1

      Observation

      Something on a chart looks repeatable. This is where every idea legitimately starts. It is not evidence of anything; it is a reason to write a hypothesis.

    2. 2

      Written hypothesis

      The claim, the cause, the kill condition and the scope — dated, and filed before any test runs. Once it is written it cannot be edited to match a result.

    3. 3

      Crude test

      The cheapest possible check on one slice of data. The purpose is not to be right; it is to kill obviously dead ideas before you spend a week on them.

    4. 4

      Full in-sample test

      The proper test on your chosen window, with realistic costs subtracted. Anything that only survives with costs ignored has already failed.

    5. 5

      Out-of-sample check

      Data you set aside at the start and have never looked at. You get one look. The moment you re-tune after seeing it, that data has become in-sample and you no longer have an out-of-sample set.

    6. 6

      Paper or forward run

      Live prices, no money, decisions logged in real time. This catches the thing no historical test can: signals that were obvious in hindsight and ambiguous as they formed.

    7. 7

      Small live, then review

      The smallest size that is still psychologically real, run for a stated period, reviewed against the written expectation. Not scaled up because it started well.

    Why is the only permitted loop the one back to the hypothesis?

    When an idea fails a gate, there are two things you can do. One is legitimate.

    The legitimate one is to go back to the hypothesis. The failure told you something about your candidate cause. Maybe the mechanism you proposed does not exist, or exists only under a condition you had not specified. That is new information, and it earns a new, separately dated hypothesis — which then has to run the whole funnel again from the top.

    The illegitimate one is to go back to the parameters. You keep the idea, change the number, and re-run. Nothing about your understanding changed; you simply took another draw. Do that enough times and something passes, and the thing that passed is the one that fit the noise best.

    On the figure above, that is why the return arrow runs to stage two and not to stage four. It is a single line on a diagram, and it is the entire difference between research and self-deception.

    Example — Illustrative arithmetic, no real instrument involved. Suppose you test 3 range lengths × 3 volume thresholds × 3 holding periods. That is 27 combinations from a single afternoon's work. Even on data with nothing in it, some of those 27 will look better than the rest — that is what variation means. Reporting the best of 27 as though it were the result of one test overstates it enormously, and it is the most common mistake in retail strategy research.

    What is p-hacking, and what does it look like in trading?

    P-hacking is the practice — usually unintentional — of searching a data set with enough freedom that a finding is inevitable, then reporting the finding as though it came from a single planned test. In academic research it destroyed the credibility of entire literatures. In trading it is the default behaviour of anyone with a backtesting tool and an afternoon.

    It rarely feels like cheating. It feels like diligence.

    • Running dozens of parameter combinations and reporting only the best one, without stating how many you ran.
    • Adding a filter after seeing which trades lost. Every such filter is a description of the past wearing the costume of a rule.
    • Sliding the test window until the result improves — trimming off the year that spoiled it, usually with a reason that sounds sensible.
    • Changing the exit rule after seeing the entry results, then re-testing on the same data and treating the second result as independent.
    • Trying the idea across many instruments and keeping the ones where it worked, without noting how many it was tried on.
    • Stopping the moment a result looks acceptable — the stopping rule itself is a choice, and stopping on success biases everything.
    • Re-testing an idea you killed six months ago because a recent move reminded you of it, without knowing you had already killed it.

    Watch out — The last one is why the log exists. A result you cannot remember producing is a result you will produce again, count as fresh confirmation, and act on twice.

    How do you defend against searching until something passes?

    None of these makes a result true. Together they make a false result far less likely to survive, which is the most any process can offer.

    The defenceWhat it preventsWhat it costs you
    Write the hypothesis and kill condition first, datedRewriting the expectation to match the outcomeA few minutes, and the comfort of vagueness
    Move one variable at a timeNot knowing which change caused the changeMore runs to cover the same ground
    Record how many combinations you triedReporting the best of many as the result of oneNothing — but it makes a weak finding look weak
    Hold out data and look at it onceTuning on the data you claim validated youLess data for the main test
    Set the stopping rule before startingStopping on success, which biases every resultIdeas you must abandon while still curious
    Log the failures with their reasonsRe-running a dead idea and calling it new evidenceAn honest record that is uncomfortable to read
    Prefer fewer parametersA rule flexible enough to fit any historyElegance you have to give up on

    Why does moving one variable at a time matter so much?

    If you change the range length and the volume filter and the holding period together, and the outcome changes, you have learned nothing about any of them. You cannot attribute the change, so you cannot form a next hypothesis — you can only take another draw.

    There is a second, subtler reason. A rule whose behaviour changes sharply when you nudge one parameter is fragile, and you only discover that by nudging one parameter. Testing a whole grid at once and picking the peak hides exactly the information you needed: whether the neighbours of that peak behave anything like it. A result standing alone in a field of failures is not a discovery. It is the most likely thing to be an accident.

    Note — This is also why the specimen research record on this page ends in a KILL despite one parameter surviving. A lone survivor out of nine is evidence against the idea, not for it — nine draws will usually produce one that looks good even when nothing is there.

    What goes into a research log?

    The log is not admin. It is the only thing that keeps a result readable a year later, and the only defence against re-running an idea you have already killed. One record per test, written at the time, never edited afterwards.

    One research-log record, field by field. The three annotated fields — candidate cause, parameters tried, and the verdict with its reason — are the ones that do the actual work.One record from a research log, listing eight fields: identifier, date, hypothesis, candidate cause, data window, parameters tried, result, and verdict with its reason. Three annotations point at the candidate cause, the parameter count and the verdict, explaining why each field earns its place.One research-log recordID2026-041DATE2026-08-14HYPOTHESISrange breakout holds 10 candlesCANDIDATE CAUSEsupply above the range is exhaustedDATA WINDOWNIFTY 500, daily, 2015-2021PARAMS TRIEDrange 20/30/40 x vol 1.2/1.5/2.0 (9)RESULTsurvives costs on 30 onlyVERDICT + WHYKILL - lone survivor of 9No named cause means thepattern is a coincidenceuntil something proves it.Record how many you tried.Nine runs, one survivor -the survivor is worth less.Write the reason it died.In six months this is whatstops you re-running it.The log is not admin. It is the only thing that keeps a result readable a year later.
    One research-log record, field by field. The three annotated fields — candidate cause, parameters tried, and the verdict with its reason — are the ones that do the actual work.

    Which fields does every record need?

    1. ID and date — so tests can be referenced, and so the sequence of your own thinking is reconstructable.
    2. Hypothesis — copied verbatim from what you wrote before the test, not paraphrased afterwards.
    3. Candidate cause — the mechanism you claim is behind it. If this field is empty, the record is of a coincidence.
    4. Kill condition — the result that was going to end the idea, recorded before the run.
    5. Data window and universe — instruments, timeframe, exact date range, and which data you deliberately held out.
    6. Costs assumed — brokerage, exchange charges, taxes, and the slippage figure you used. State the number even if it is a guess, so a later you can challenge it.
    7. Parameters tried — how many combinations, not just the one you kept. This single field defuses most self-deception.
    8. Result — qualitative is fine and often better. What happened, where it broke, what surprised you.
    9. Verdict and why — kill, park, or advance, with the reason. In six months this is the field that stops you re-running a dead idea.

    Note — A spreadsheet is enough. The medium has never been the problem — writing the record before you know the outcome is. A log filled in after the fact is a memoir, and memoirs are kind to their authors.

    Why must the failures be logged, not just the survivors?

    A log of successes is not a record, it is a highlight reel — and it is exactly the thing that makes a published track record worthless. The same logic applies inward, to your own research.

    There are three concrete reasons the failures earn their space. First, without them you cannot know how many ideas produced your survivors, and that ratio is what tells you how much the survivors are worth. Second, a killed idea has a reason attached, and that reason is often more transferable than any passing result — it tells you something about the market that the winner does not. Third, ideas come back. Something in a chart six months from now will remind you of an idea you have already tested and buried, and only a written verdict will stop you from spending another week on it.

    Watch out — Count the entries. If your research log contains almost no kills, that is not a sign of good idea generation. It is a sign that you stop testing at the point where a result becomes usable, which means every entry in the log was produced by a process designed to find something.

    What are the honest limits of this workflow?

    • It cannot make a small sample large. A rule that triggers rarely produces few decisions to judge, and no amount of process turns a handful of observations into evidence.
    • Historical data does not contain the fill you would have got. Prices you can see are not prices you could have traded, particularly away from the most liquid names.
    • Indian corporate actions rewrite the price series. Splits, bonuses and demergers, and the way your data source adjusts for them, can create or destroy an apparent effect entirely — that is a topic of its own later in this module.
    • Survivorship in the universe biases everything. Testing today's index constituents over the last decade quietly excludes everything that fell out of the index along the way.
    • Costs are assumptions. Brokerage, STT, exchange transaction charges, stamp duty, GST and slippage all have to be modelled, and a wrong slippage assumption can flip a conclusion on its own.
    • Out-of-sample data is finite. Every idea you test consumes some of the only genuinely unseen data you will ever have, which is a real argument for testing fewer ideas more carefully.
    • The process guarantees nothing about the future. It only ensures that when something stops working, you can tell — because you wrote down what working was supposed to look like.

    What would tell you your research process itself is broken?

    Every honest method has to state what invalidates it. These are the signals that the process, not the strategy, has failed — and none of them is a losing result.

    • You cannot produce the hypothesis you wrote before a test, because you did not write one.
    • Your log has no kills, or the kills stop appearing after the first few weeks.
    • You cannot say how many parameter combinations produced the result you are relying on.
    • You have looked at your held-out data more than once, or re-tuned after looking at it. There is no out-of-sample set any more; there is only a second in-sample set.
    • The rule now has more conditions than it started with, and each one was added after seeing a specific loss.
    • A test cannot be reproduced from the record — the data window, universe or cost assumption was never written down.
    • You are testing an idea you already killed, and you only found out afterwards. The log exists and you are not reading it.
    • The stopping decision was made by the result. If you would have kept going had it looked worse, the result is not a result.

    Note — This page is education, not investment advice, and nothing here is a recommendation to buy or sell any security. It describes how to evaluate an idea honestly — it makes no claim that any particular idea, tested this way or otherwise, will work.

    How do you run your next research idea?

    Before, during and after a single test

    • Write the claim, the candidate cause, the kill condition and the scope. Date it. Do not open the data first.
    • Decide the held-out window now, and put it out of reach until stage five.
    • State your cost assumptions in numbers before you know whether they matter to the answer.
    • Change one variable per run, and write down every combination you try — including the ones you abandon halfway.
    • Decide in advance how many runs you will do, and stop there whether or not you have found anything.
    • Write the verdict — kill, park or advance — with its reason, on the day, before the result becomes a story.
    • If it advances, forward-log it against live prices before any money is involved.
    • Re-read the log before starting any new idea. Half of what feels new has already been tested by you.

    Key points

    A system and a hunch differ in process, not in appearance. The record behind the rule is the only place the difference shows.
    Write the hypothesis, the candidate cause, the kill condition and the scope before touching the data.
    A hypothesis you cannot state a failing result for was never tested — it was searched.
    Seven stages, each with a gate the idea can fail. A stage you cannot fail is a formality.
    When an idea fails, go back to the hypothesis. Going back to the parameters is just another draw.
    Move one variable at a time, or you cannot attribute anything you observe.
    Record how many combinations you tried. The best of twenty-seven is not the result of one test.
    Look at held-out data once. Re-tuning after the look destroys the only unseen data you had.
    Log the failures with their reasons — a log with no kills describes a process designed to find something.
    Set the stopping rule before you start. Stopping on success biases every result you keep.

    Pro tip — Before your next test, write one sentence and date it: 'If ___ happens, I will drop this idea.' Then file it where you cannot quietly edit it. Almost every research failure downstream of that sentence is a failure to have written it — and the sentence costs nothing except the option to change your mind about what you were looking for.

    Frequently asked questions

    What is p-hacking in trading?

    P-hacking is searching a data set with enough freedom that a good-looking result is inevitable, then reporting it as though it came from a single planned test. In trading it usually appears as running many parameter combinations and keeping the best, adding filters after seeing which trades lost, or trimming the test window until the outcome improves. The defence is to write the hypothesis and the kill condition before the test, record how many combinations you ran, and set the stopping rule in advance.

    How many parameters should a trading strategy have?

    As few as express the idea. Every additional parameter is another dimension you can search, and searching is what produces false findings. There is no magic count, but there is a useful check: the more parameters a rule has, the more of its apparent quality you should attribute to fitting rather than to any real effect, and the more independent evidence it needs before you rely on it.

    What is the difference between in-sample and out-of-sample testing?

    In-sample data is what you develop and tune on. Out-of-sample data is set aside at the very start and looked at once, at the end, to check whether the rule survives on data that had no influence on it. The rule that makes it meaningful is the single look: if you re-tune after seeing the out-of-sample result, that data has become in-sample and you no longer have a held-out set.

    Do I need to keep a research log if I am the only one using it?

    Especially then. The log's main reader is you in six months, when an idea you already killed looks new again and the reasoning that killed it has faded. It is also the only way to know how many ideas produced your survivors, which is what tells you how much those survivors are worth. A spreadsheet is enough; writing the record before you know the outcome is the part that matters.

    Why should I record failed backtests?

    Because the ratio of failures to survivors is what makes a survivor interpretable. One idea passing out of two is a very different claim from one passing out of ninety, and without the failures you cannot tell those apart. Killed ideas also carry reasons, and a reason is often more transferable than a result — it tells you something about the market that the winner does not.

    How much data do I need to test a trading idea?

    Enough that the number of trades, not the number of years, is meaningful — and enough to span more than one kind of market. There is no single threshold, and any answer that ignores how often the rule triggers is not an answer. A rule producing a handful of signals over a decade cannot be validated by lengthening the window; it can only be recognised as untestable at that frequency.

    Is forward testing better than backtesting?

    They answer different questions. A backtest asks whether the idea would have made sense historically and can be run in an afternoon. A forward or paper run asks whether you can identify the signal as it forms, in real time, without hindsight — which is the one thing no historical test can check. Neither replaces the other, and both come after the hypothesis is written.