Somewhere on the landing page there is a big number. An accuracy percentage, usually, sitting in a large font with a short label under it. It is the first thing your eye lands on and the last thing anybody interrogates. But an accuracy figure with no sample size, no time period and no cost assumption is not evidence of anything. It is a design element. You cannot tell from it whether the tool was tested on twenty trades or twenty thousand, over one trending quarter or five mixed years, with brokerage and slippage deducted or with none at all.
This article is not about spotting scams, although some of it applies. Most AI trading products are sold by people who genuinely believe their thing works. The problem is that believing something works and having demonstrated it are very different states, and the marketing language for both is identical. So the job is not to detect dishonesty. It is to find out which state the vendor is in — and that is done with questions, not instincts.
What follows is the buyer’s question list, the difference between a claim that carries information and one that carries none, the structural red flags that should end a conversation regardless of how good the product looks, and the registration check that comes before all of it in India. Apply every bit of it to this site as well. The author here is a SEBI Registered Research Analyst, and that is a reason to check the registration, not a reason to skip the check.
The one thing to remember
A claim you cannot attach a sample size, a time period and a cost assumption to is marketing — and the vendor’s willingness to answer those three questions in one line each tells you more than the number ever will.
What a Headline Accuracy Number Actually Hides
Take the number off the landing page and try to reconstruct what it means. You immediately find that you cannot, because at least four things have been left out, and each of them can move the figure enormously. How many trades is it computed over. What period those trades span. What counted as a win. And whether costs were deducted before the win was counted.
Sample size comes first because small samples produce impressive numbers by accident. A handful of outcomes can look extraordinary for reasons that have nothing to do with skill, in the same way a short run of coin flips can come up heads repeatedly. The larger the sample, the harder it becomes for luck to carry the result — which is exactly why a vendor with a genuinely large sample tends to state it prominently, and a vendor without one tends to state a percentage instead.
Period matters because market conditions are not uniform. A method tested only across a stretch when the broad market trended upward has not been tested against the conditions that hurt it. Ask which years, and ask specifically whether the test period includes a sharp drawdown and a long directionless range. A method that has only met one kind of market has not been tested; it has been introduced.
Then costs, which is where more claims die than anywhere else. Indian equity and derivative trading carries brokerage, exchange charges, statutory levies and — the one most often ignored — slippage, the gap between the price you expected and the price you got. A short-holding-period method touches those costs constantly. A result computed before costs and a result computed after costs are not the same result with a small adjustment; for high-turnover approaches they can be different conclusions entirely.
A percentage without a sample, a period and a cost assumption is not a weak claim. It is not a claim at all.
Autopsy of a claim
The test is falsifiability, not confidence. Cut the sentence apart and ask of each piece: could any observation prove this wrong?
The Six Questions a Genuine Vendor Answers in One Line
The useful property of the questions below is that they are cheap to answer if you have done the work and impossible to answer if you have not. Someone who has genuinely tested a method knows the sample size the way you know your own phone number. Someone who has not will produce a paragraph about their proprietary approach, and the length of the paragraph is inversely related to the strength of the evidence.
Send them in writing, together, and read what comes back as a whole. You are not looking for perfect answers — a small sample honestly disclosed is far more encouraging than a large one vaguely implied. You are looking for whether the vendor treats measurement as something they did or as something they are being interrogated about.
One question deserves special attention because it is the one almost nobody thinks to ask: how many variants were tried before this one was chosen. If a hundred parameter combinations were tested and the best one is being sold to you, the reported performance is partly a measure of how many attempts were made, not how good the method is. Testing many things and reporting the winner is the most common way an honest person produces a misleading result.
The follow-up that separates the serious from the rest: is this backtested, forward-tested, or live? A backtest is a simulation on history the developer could see while building. A forward test runs on data that arrived after the rules were fixed. Live means real money, real fills, real costs. These sit in ascending order of credibility, and the gap between the first and the last is where almost all disappointment lives.
Below is the marketing copy for a trading product. Do not evaluate whether the product is good. I only want the claims audited. For every performance or capability claim in the text, produce a row with: 1. The claim, quoted verbatim 2. What would have to be measured for this claim to be verifiable 3. Which of these is missing: sample size, time period, universe tested, cost and slippage assumption, backtest vs forward-test vs live, maximum drawdown 4. A single specific question I should send the vendor about this claim Then list separately: - Any language that promises, implies or hints at a return, profit or outcome - Any claim stated as a number with no source - Anything presented as evidence that is actually a testimonial or a screenshot Do not soften anything and do not add a verdict. If a claim is fully specified, say so plainly. MARKETING COPY: [paste it here]
When to use — Paste a landing page, a brochure PDF or a sales message before you reply to it, so the questions you send back are specific rather than general.
A good answer — Verbatim quotes rather than paraphrases, a genuinely specific missing-information column, and a question list you could send unedited. Treat a summary that says “the claims look reasonable” as a failed run — that is the model being agreeable, not analytical.
The six questions — send them in writing, before you pay
A vendor who has done the work answers each of these in a sentence. Watch for which ones get a paragraph instead of an answer, and for which ones get skipped entirely.
- Over what period, and across how many individual trades, is the reported performance computed?
- On which universe of stocks — and does that universe include names that were later delisted or suspended, or only names that still exist today?
- Are brokerage, exchange charges, statutory levies and slippage deducted from the reported result, and at what assumed slippage per trade?
- Is this backtested on history, forward-tested on data that arrived after the rules were fixed, or live with real money?
- What was the maximum drawdown — the deepest peak-to-trough fall in account value — and how long did it take to recover from it?
- How many strategy or parameter variants were tested before this one was selected, and were the discarded ones measured the same way?
A Claim That Means Something vs a Claim That Means Nothing
The difference is not tone and it is not confidence. Plenty of meaningless claims are stated modestly and plenty of meaningful ones are stated with pride. The difference is falsifiability: can you, in principle, check it and find it wrong? A claim you could never disprove is a claim that was never at risk of being false, which is another way of saying it told you nothing.
“Our AI analyses millions of data points” cannot be wrong, because analysing is undefined and millions is unverifiable. “We processed every quarterly filing for the NIFTY 50 constituents from this date to that date, and here is the extraction schema” can be wrong — you could check a filing and find the extraction missing. The second sentence is less exciting and vastly more informative.
This test also works on capability claims, not just performance ones. “Understands market context” means nothing. “Reads a PDF you upload and returns a fixed set of fields with a page reference beside each one” means something, because you can upload a PDF and see whether the page references are correct. Ask for the checkable version of any claim, and notice how often the vendor cannot produce one.
Apply the table below to whatever is in front of you. Most marketing copy sits entirely in the right-hand column, and the exercise of trying to move a claim from right to left is usually enough to end the evaluation without any further work.
| Topic | A claim that means something | A claim that means nothing |
|---|---|---|
| Performance | Result stated with trade count, dates, universe and deducted costs | A large accuracy percentage with no context beside it |
| Testing | Forward-tested on data that arrived after the rules were frozen | “Extensively tested by our team over many years” |
| Risk | Deepest peak-to-trough fall stated, with recovery time | “Advanced risk management built in” |
| Data | Named sources, stated coverage period, stated update frequency | “Analyses millions of data points in real time” |
| Capability | Returns a fixed field set with a page reference you can check | “Understands market context and sentiment” |
| Selection | Number of variants tested disclosed, discarded ones measured too | Only the winning configuration is ever mentioned |
| Evidence | A method description you could reproduce yourself | Screenshots of profitable trades and customer testimonials |
Pro tip — Ask for one worked example on a name you choose, not one from their demo list. A tool that performs on the vendor’s three favourite stocks and stumbles on the fourth you pick has been tuned to a demo rather than built for a job.
The Structural Red Flags
Some signals are not about weak evidence. They are about the shape of the offer itself, and they should end the conversation regardless of how impressive the technology looks. The first is any language that promises or implies a return. Guaranteed, assured, risk-free, sure-shot, fixed monthly income from trading — none of these can honestly be said about a market-linked activity by anyone, and a party willing to say them has already told you how they handle inconvenient truths.
The second is testimonials as evidence. A screenshot of a profitable trade proves that one profitable trade existed. It says nothing about the trades not screenshotted, and the selection is being done by the person trying to sell you something. Reviews and profit screenshots are the weakest form of evidence in this domain and are treated as the strongest by most buyers, which is precisely why they are used.
The third is manufactured urgency. Closing in six hours, three seats left, price doubles tomorrow. Urgency exists to prevent the thing you are doing right now — thinking about it. A product with real evidence does not need you to decide before you have read the evidence, and a genuine capacity limit can be stated once without a countdown timer attached.
The fourth is subtler and specific to India. A “signal” service that tells you which security to buy or sell, at what price, with a target and a stop, is not a neutral piece of software however it is packaged. Telling people what to buy or sell is a regulated activity here, and the label on the product does not change what the product does. If a person or entity is issuing recommendations to you, the registration question below is not optional.
How to run an evaluation without being managed through it
Do
- Send your questions in writing and read the whole reply as one piece — the pattern of what gets dodged is the real answer.
- Ask for one worked example on a name you choose, on a date you choose, and check it against the primary source yourself.
- Keep your own log for the trial period: what you asked, what it returned, what you verified, what was wrong.
- Judge the tool on whether it saves you defensible time, not on whether its outputs happened to be right during a favourable stretch.
- Check registration status on SEBI’s own site before anything else, if the offering involves recommendations.
Don't
- Do not accept profit screenshots, testimonials or member counts as evidence of a method working.
- Do not let a countdown, a closing cohort or a limited-seats message compress a decision you have not finished making.
- Do not evaluate a tool during only one kind of market and conclude anything durable from it.
- Do not pay for a year up front to save money on a product you have used for a week.
- Do not assume that impressive technology implies tested performance — those are unrelated claims and they are almost always bundled.
Watch out — Any offer that promises or implies assured returns, guaranteed profits or a fixed monthly income from market activity should end your evaluation immediately, whatever the technology behind it looks like. Nobody can honestly make that promise about a market-linked outcome.
In India, the Registration Check Comes First
Before the accuracy questions, before the trial, before anything: if the offering involves someone telling you what to buy or sell, find out whether that person or entity is registered with SEBI. Research analysts and investment advisers operate under registration, and registration is what makes a real person accountable for what they publish — with disclosures about their own positions and conflicts, and a defined channel when something goes wrong. An unregistered party offering recommendations has none of that structure behind them, and neither do you.
The check itself is quick and you should do it yourself rather than accept a claim of registration. SEBI publishes lists of registered intermediaries on sebi.gov.in. Look up the registration number, confirm the name attached to it matches the party you are dealing with, and confirm the category is the one relevant to what they are offering. A registration number printed on a website is a claim like any other until you have seen it on the regulator’s own site.
Be clear about what registration does and does not tell you. It tells you there is an accountable, identifiable party operating under a regulatory framework with disclosure obligations. It does not tell you the person is skilled, that their calls will work, or that you will make money — SEBI itself is explicit that registration guarantees neither performance nor returns. It is a floor, not a recommendation, and treating it as an endorsement is its own error.
This applies to this site too. Rohit Singh is a SEBI Registered Research Analyst, registration number INH000015297, trade name INVESTOLOGY, and you should verify that on sebi.gov.in rather than take this sentence for it. Every question in this article — what is the sample, what is the period, were costs included, what is the evidence — applies to content published here exactly as it applies anywhere else. A writer who asks you to be sceptical of everyone except themselves is asking for the wrong thing.
Registration tells you someone accountable stands behind the words. It does not tell you the words are right — you still have to check.
Running a Trial That Actually Tells You Something
Suppose the vendor answered well and the registration checks out. You still know nothing about whether the tool helps you, because that depends on your process, not on their evidence. So the trial is a real experiment and needs to be run like one, which mostly means deciding what would count as success before you start rather than deciding afterwards.
Pick one task you already do — reading quarterly filings for a set of names, say, or the first pass over your watchlist — and run the tool against work you have already done or can verify. This is important. If you test it on unfamiliar material, you cannot tell a good output from a confident wrong one. Testing on your own ground is the only way to see the failure modes that matter to you.
Keep a plain log for the trial: date, what you asked, what it returned, what you checked, what was wrong. Two weeks of that log is worth more than any amount of reading about the product, because it measures the only thing that matters — whether the output was defensible on your material. And the log has a second use: it stops you from remembering the trial through the lens of whichever outcome happened last.
At the end, ask three questions. Did it save time that was genuinely being wasted, or did it move the work from producing to checking? Did the errors it made cluster somewhere predictable, so you know where not to trust it? And would you still want it if the market had been flat throughout? If the honest answer to the third is no, you were evaluating conditions rather than the tool.
I am ending a trial of a research tool. Below is my log of every task I gave it and what I found on verification. Analyse the log and return: 1. Error rate by task type — where did verification most often find something wrong? 2. Error type breakdown: invented figures, wrong attribution, missed material items, correct but useless 3. Which tasks it handled reliably enough that I could keep using it with a light check 4. Which tasks it should never be given again, based on this log 5. Where my log is too thin to conclude anything — say so rather than filling the gap Rules: - Use only what is in the log. Do not infer performance for task types I did not test. - Do not tell me whether to buy the product. Just characterise what the log shows. - If the sample is too small for a conclusion, say that first and plainly. LOG: [paste your dated log here]
When to use — At the end of a trial period, before the renewal decision, when memory of the last few outputs is about to stand in for the whole trial.
A good answer — Error rates tied to specific task types, an explicit statement about where your sample is too thin, and no purchase recommendation. If it produces confident conclusions from a dozen entries, it is pattern-matching to your hopes rather than reading the log.
Common questions
Not on its own. Accuracy without a trade count, a date range, a tested universe and a cost assumption cannot be interpreted, because each of those can move the figure dramatically. Ask for all four in writing. A vendor who has genuinely measured performance answers in one line; one who has not will send a paragraph about their methodology instead.
Knowledge Check
A vendor states a strong accuracy figure. Which single missing detail most often makes such a claim misleading?
Written By
Rohit Singh
Mr. Chartist
With 14+ years of experience in Indian financial markets, Rohit Singh (Mr. Chartist) is a SEBI Registered Research Analyst, Amazon #1 bestselling author, and the founder of Investology — a premium trading ecosystem trusted by a 1.5 Lakh+ strong community across India.
Keep reading
Guardrails, SEBI and the Human in the Loop — Where AI Ends and Advice Begins
Advice is a regulated activity with a named, accountable person behind it. A model has none of that, and cannot acquire it.
Quant & Machine LearningBacktesting an AI-Generated Strategy — The Honest Test
The model wrote it in thirty seconds. Proving it is not curve-fitted noise takes considerably longer, and skipping that is the whole risk.
Quant & Machine LearningFeatures, Labels and Overfitting — Why Most ML Backtests Are Fiction
Look-ahead bias, survivorship bias and a model that memorised the past. Three failure modes that produce beautiful, worthless equity curves.
