intermediate10 min read12 of 24

    AI-Assisted Screening — Casting a Wider Net, Not Picking the Winner

    Numbers still belong in a real scanner. What AI adds is the qualitative layer no numeric filter can reach.

    Rohit Singh

    Mr. Chartist · SEBI RA INH000015297

    Module

    There is a version of this article that a lot of people would rather read. It would explain how to describe an investment idea in plain English, hand that description to a language model, and receive back a list of NSE companies that match. It would feel like magic, and it would be wrong in a way that is very hard to detect, because the list would look exactly like a list a real screener produced.

    A language model does not have the NSE and BSE universe in front of it. It has a compressed impression of text it read at some point in the past, and no ability to sort three thousand companies by a ratio. Ask it for "companies with debt-to-equity under 0.5 and improving margins" and it will produce names — recognisable, plausible, sector-appropriate names — with numbers attached that were generated rather than read. Every part of that answer is the kind of thing that gets a retail investor into trouble quietly.

    So the boundary in this article is hard and it is not negotiable. Numbers belong in a real scanner — a tool connected to exchange data that sorts and filters deterministically. What AI adds sits entirely on the other side of that line: the qualitative layer. Language. Tone. What management said this quarter that they did not say last quarter. Which complaints keep recurring across an entire sector’s calls. That layer is genuinely hard to screen for with a numeric filter, and it is where AI earns its place in the process.

    And even there, the output is not a decision. It is a shortlist of things to go and check. The net gets wider; the verification does not get lighter.

    The one thing to remember

    AI is not a screener. A scanner filters numbers deterministically; AI reads language and hands you candidates. Anything a model produces that looks like a filtered list of stocks with figures attached was written, not computed.

    The Line: Numbers Go To a Scanner, Language Goes To a Model

    A screener is a database query. You specify conditions — market cap above a threshold, promoter holding above a level, five-year sales growth above a rate — and the tool walks the whole listed universe and returns exactly the rows that satisfy them. The output is reproducible. Run it again on the same data and you get the same set. If a company is missing, either it failed a condition or the data is wrong, and both are checkable.

    A language model does none of that. It has no table to walk, no sort, no comparison operator. When you ask it for companies meeting numeric conditions, it produces text that statistically resembles an answer to that question — which means well-known names in the right sector with figures in the plausible range. It cannot tell you it did not compute anything, because from its own point of view nothing unusual happened.

    The failure is dangerous specifically because the output format is identical. A real screener output and an invented one both arrive as a tidy list with a company name, a ratio and a two-line reason. There is no visual tell. The only defence is a rule you apply before you look at the output: if the question required sorting or filtering the universe by a number, the answer did not come from a model.

    This site ships a real scanner for exactly that reason — numeric conditions against live NSE and BSE data, run deterministically. The division of labour is clean. The scanner narrows three thousand companies to a workable set using conditions you can defend. AI then reads the language around that set, or around a sector, and surfaces things a ratio cannot express.

    If a question can only be answered by sorting the universe by a number, a language model is the wrong tool — and its answer will look exactly like the right one.

    The questionWhyWhere it belongs
    Companies with ROCE above 18% for five straight yearsA sort and filter across the full universeScanner — never a model
    Stocks trading above their 200-day average with rising volumeDeterministic price and volume computationScanner or charting platform
    Debt reduced in each of the last four annual reportsA numeric comparison across dated statementsScanner, then verify in the filings
    Which of these 40 concalls changed their guidance languageReading tone and wording across documentsAI, on transcripts you supply
    What cost complaint keeps recurring across this sectorPattern in language, not in a data fieldAI, on transcripts you supply
    Which filings this month describe a genuine capacity changeClassifying prose by what it actually saysAI, on filings you supply
    Which tool for which question. The middle column is the reason, not a preference — it is about whether the operation is a computation or a reading.

    Watch out — Never accept a stock list with numbers attached from a language model, even when it names its source. The citation is generated by the same process as the number. Treat any figure that did not come from a scanner or a filing you opened as fiction until proven otherwise.

    What a Numeric Filter Structurally Cannot See

    A screener sees fields. Ratios, growth rates, holding patterns, price and volume. It is very good at that and it is the right place to start. But an enormous amount of what actually moves a business lives in prose that never becomes a field.

    Consider a management team that spent four consecutive quarters saying they were "confident of maintaining margins" and this quarter says they are "watching the input cost situation closely". No ratio has moved yet. The reported numbers are identical to what a screener would have shown last quarter. Something material has nevertheless happened, and it happened entirely in language.

    Or consider a cost complaint appearing in eleven of thirty concalls across an industry within the same fortnight — freight, a particular raw material, a wage revision. Individually each mention is unremarkable. As a cluster it is a sector-level observation that no single company’s financial fields would ever surface, because the pattern exists across documents rather than within any one of them.

    Or the shape of a filing calendar. A company that files three related announcements in eight days — a board meeting notice, a fund-raising enabling resolution and a plant-related intimation — is telling you something about its own tempo that no ratio encodes. AI can read all three and describe what they collectively refer to. Reading them is the value; concluding anything from them is still your job.

    Guidance language shifts before the reported numbers do, and no numeric field captures the shift.
    Cross-company patterns — a complaint recurring across a sector’s calls — exist between documents, not inside any one of them.
    Filing clusters and their tempo carry information that never becomes a screenable field.
    Risk-factor wording that changes year over year in an annual report is a language signal, not a data one.
    None of these are predictive. They are reasons to go and look, which is a different and more modest claim.

    The Funnel — Where AI Sits, and Where It Does Not

    The right mental model is a funnel with four stages, and AI occupies exactly one of them. Getting this arrangement right is most of the article.

    Stage one is the universe: every listed company on NSE and BSE. Stage two is the numeric filter, run in a real scanner, which cuts that to a set small enough to read — perhaps thirty to eighty names, depending on how strict you were. Stage three is the qualitative pass, and this is the AI stage: you feed it the documents belonging to those names and ask it to read across them. Stage four is verification, which is you, opening the source and confirming every claim before it goes anywhere near a decision.

    Notice what the AI stage does to the funnel shape. It does not narrow further in the way a filter does. It re-sorts — it tells you which of the thirty are worth reading first and why, and it occasionally widens by surfacing a name adjacent to your set that shares the language pattern. That is the "wider net" in the title. The net gets wider at stage three and then the verification stage does the actual narrowing.

    The most common way this goes wrong is stage-skipping: running the AI pass on the whole universe rather than on a filtered set, which is the exact moment you stop supplying documents and start asking the model to recall. If you cannot paste the sources for the names you are asking about, the stage is being run on memory, and memory is where invented figures live.

    Numbers in the scanner, language in the model

    The AI stage widens the net. It does not narrow it to a winner.

    Stage 1 — a real scannernumeric filters on price andfinancial data, computed exactlywhole universepasses the filtersNever ask a language model to do this arithmetic.Stage 2 — the AI passreads language only, no mathsGuidance tone across concallsRecurring cost complaintsClustered filings in one sectorthe net gets wider hereCandidatesfor verificationStock AcheckStock BcheckStock CcheckStock DcheckPlaceholder labels. Not names, not picks.Nothing below this line existsThere is no final stage. The screen ends at a shortlist you take to the chart and thefilings yourself. A screen that hands you one name has stopped being a screen — and nooutput on this page is a recommendation to buy or sell anything.
    The four-stage funnel. The numeric filter runs in a scanner against exchange data; AI reads language across the documents belonging to the survivors; verification against the source is the only stage that produces anything you can act on.
    1. 1

      Start with the full listed universe

      Every NSE and BSE company. Nothing about this stage involves AI — it is simply the set you are filtering down from.

    2. 2

      Apply numeric conditions in a real scanner

      Balance-sheet, growth, holding and price conditions, run deterministically against exchange data. Write the conditions down; you will need them when you review what the filter excluded.

    3. 3

      Collect the documents for the survivors

      Concall transcripts, exchange filings, annual report sections. This is manual and it is the step people skip. Without the documents, stage four is recall, not reading.

    4. 4

      Run the qualitative pass across those documents

      One fixed prompt, applied to the pasted text, asking for language patterns and quoted evidence. The output is an ordering and a set of quotes, never a ranking of attractiveness.

    5. 5

      Verify every quote against the source

      Open the transcript, find the sentence, confirm the wording. A quote that cannot be located in the document is deleted along with whatever conclusion rested on it.

    6. 6

      Do your own work on what survives

      The shortlist is an input to your research process, not a substitute for it. Nothing in stages one to five has told you anything about whether a business is worth owning.

    The Qualitative Pass — Reading Thirty Calls In One Sitting

    Here is the honest version of what AI does at stage three. You have thirty concall transcripts. Reading all of them carefully is perhaps two full days of work, which is why most people read four and guess about the rest. A model can pass over all thirty in an afternoon and tell you, with quotes, which four contain the language you specified. You then read those four properly, and you read them knowing what you are looking for.

    That is a real gain and it is worth being precise about what kind of gain it is. It is triage. It changes the order in which you read, and it means you are unlikely to leave a document unopened purely because it was twenty-eighth in the pile. It has not evaluated any company and it has not compared any two businesses on merit.

    The prompt for this pass has the same architecture as every other extraction prompt in this module: pasted source, fixed output structure, mandatory quotes, and explicit permission to come back with nothing. That last part matters more here than anywhere else. Ask a model to find guidance softening in thirty calls and it will find guidance softening in thirty calls, because producing an empty result reads to it as failure. You have to say, in writing, that "no such language in this transcript" is the expected answer for most of them.

    Run it one document at a time, in a fresh thread, and collect the outputs yourself. Batching all thirty into one enormous prompt sounds efficient and produces the worst of both worlds: the earlier transcripts get less attention as the context fills, and the framing from company one leaks into the reading of company nine.

    Qualitative screening pass — one transcript at a time, identical wording every time
    You are reading ONE earnings call transcript as part of a screening pass
    across many companies. Work only from the text between the markers. Do not use
    anything you know about this company from any other source.
    
    COMPANY: [full registered name]
    PERIOD: [e.g. Q2 FY26]
    
    --- BEGIN TRANSCRIPT ---
    [paste the full transcript here]
    --- END TRANSCRIPT ---
    
    Report on exactly these five things, in this order:
    
    1. GUIDANCE LANGUAGE — every sentence in which management describes future
       performance. Quote each one verbatim with the speaker.
    2. HEDGING AND QUALIFIERS — any wording that softens, conditions or withdraws
       a commitment ("subject to", "we will watch", "assuming", "barring").
       Quote verbatim.
    3. COST PRESSURE MENTIONED — each input, freight, power or wage cost the
       management or an analyst raised, with the sentence it appeared in.
    4. CAPACITY, CAPEX OR EXPANSION LANGUAGE — as stated, with amounts and
       timelines exactly as given, in the units used in the call.
    5. QUESTIONS DEFLECTED — analyst questions that were asked but not actually
       answered. Quote the question and the reply.
    
    Rules, all of which override the desire to be helpful:
    - If a heading has nothing in this transcript, write "NOTHING IN THIS
      TRANSCRIPT". Most transcripts should return that for at least one heading.
    - Every item must be a verbatim quote. Never paraphrase and never summarise.
    - Do not compare this company to any other company.
    - Do not say whether any of this is positive or negative, and do not mention
      valuation or the share price.

    When to use — Once per transcript, in a fresh thread, across every name that survived your numeric filter. The wording never changes between companies — that is what makes thirty outputs comparable.

    A good answer — Verbatim quotes with speakers attached, at least one "NOTHING IN THIS TRANSCRIPT" across the five headings, no cross-company commentary, and no view on whether any of it is good news.

    Pro tip — Keep the outputs in one file, one company per block, with the transcript date at the top of each. The cross-company pattern only becomes visible when you can scroll through thirty identically-shaped blocks — and that reading is yours to do, not the model’s.

    Cross-Document Patterns — The Thing Only This Layer Finds

    Once you have thirty identically-structured outputs, a second kind of question opens up: what is common across them? This is where the qualitative layer stops being a faster way to read and starts producing something a numeric screen genuinely could not.

    Paste your own collected outputs — not the raw transcripts, your extracted and structured blocks — and ask for recurrence. Which cost line was mentioned by the most companies. Which phrase appears in several unrelated managements’ guidance. Where two companies in the same sector describe the same demand environment in opposite terms, which is often the most interesting single observation available, because at most one of them can be right about a shared market.

    The output of this pass is a set of observations with counts and citations back to your own blocks. It is descriptive. "Eleven of thirty calls mentioned freight cost; here are the eleven quotes" is a fact about a set of documents. It is not a forecast, it does not imply that any of those eleven companies is attractive, and translating it into a view is entirely your job — done with the documents open.

    It is also worth running deliberately in the opposite direction. Ask what appears in only one of the thirty. Sector-wide language tells you about the environment; a single company saying something nobody else is saying tells you either that they see something others do not, or that they are describing their own situation differently from their peers. Both are worth a careful read of that one transcript.

    A pattern that exists between thirty documents rather than inside any one of them is invisible to every numeric screen ever built. That gap is the whole reason this layer exists.

    Cross-company recurrence pass — run on your own collected outputs
    Below are my structured extraction blocks from [N] earnings calls in the
    same sector, one block per company. Every quote in them was taken verbatim from
    the transcript and checked by me.
    
    --- BEGIN MY BLOCKS ---
    [paste your collected per-company outputs here]
    --- END MY BLOCKS ---
    
    Working only from these blocks, report:
    
    1. RECURRING THEMES — every subject raised by three or more companies. Give
       the count and quote one line from each company that raised it.
    2. UNIQUE MENTIONS — subjects raised by exactly one company, with the quote.
    3. DIRECT CONTRADICTIONS — where two companies describe the same market,
       input cost or demand condition in incompatible terms. Quote both.
    4. HEDGING CONCENTRATION — which companies’ guidance carried the most
       qualifying language, listed with their quotes. Do not interpret this.
    5. WHAT I APPEAR TO HAVE MISSED — headings that came back empty across an
       implausible number of companies, which may mean my extraction was faulty
       rather than the language being absent.
    
    Rules:
    - Every claim must cite the company block it came from.
    - Do not rank these companies, score them, or suggest which are attractive.
    - Do not introduce any company, figure or fact that is not in the blocks above.
    - Say "no contradictions found" rather than manufacturing one.

    When to use — After you have run the per-transcript pass across a whole sector and verified the quotes. This is the step that finds what no ratio filter could.

    A good answer — Counts with citations back to your own blocks, a populated contradictions section or an honest "none found", and no ranking of any kind. Section 5 is the one to read first — it usually reveals a flaw in your extraction prompt.

    The Output Is a Shortlist of Questions, Not a List of Picks

    Everything above produces one thing: a small set of names you are going to look at more carefully, each attached to a specific reason and a specific quote. That is the deliverable. It is deliberately unexciting.

    Write the shortlist down in a form that keeps the reason attached to the name, because the reason is the part that decays. "Company X — guidance wording moved from confident to conditional between the Q1 and Q2 calls; quotes attached; source Q2 FY26 transcript, management opening remarks" is a research task. "Company X — looks interesting" is a memory you will misremember as a conclusion within a fortnight.

    Then do the actual work. Read the full transcript rather than the extract. Check the corresponding numbers in the filed results, from the filing itself. Look at what the company said in the previous two calls, not just the one you screened. The AI pass got you to the right four documents faster; it did not read them for you, and nothing in it constitutes a view on whether any business should be owned.

    And keep a note of what your numeric filter excluded before stage three ever ran. A screening process is defined as much by what it rejects as by what it surfaces, and the exclusions are where a filter’s hidden assumptions live. A debt condition that removes every company mid-way through a capex cycle is a decision about what kinds of businesses you will never see, made silently by a threshold you set once.

    Using the shortlist

    Do

    • Keep the reason and the verbatim quote attached to every name on the list.
    • Treat each entry as a question to investigate, phrased as a question in your notes.
    • Read the full source document for every name that survives, not the extraction.
    • Record the numeric conditions your scanner ran, so you know what was excluded and why.
    • Delete any entry whose supporting quote you could not locate in the source.

    Don't

    • Ask the model which of the shortlisted names is the most attractive.
    • Ask it to rank, score or assign conviction levels to companies.
    • Let a name onto the list without a quote you have personally verified.
    • Treat a recurring sector theme as a reason to act rather than a reason to read.
    • Skip the scanner and let a model produce the numeric filter, in any form, ever.

    Before a name goes on your shortlist

    Run this on every entry. Any single failure means the entry is not ready, regardless of how compelling the reason sounded.

    • Did the numeric filtering happen in a real scanner against exchange data, not in a model?
    • Can you state the exact conditions the scanner ran, and what they excluded?
    • Was the source document pasted into the prompt rather than named and recalled?
    • Is every claim on the entry supported by a verbatim quote you located yourself in the source?
    • Does the entry name the document and the location inside it — call, quarter, section?
    • Does the entry record a question to investigate, rather than a conclusion?
    • Is the entry free of any figure that came from the model rather than from a filing or the scanner?
    • Have you separately noted what your filter would have thrown away, so you know your blind spot?

    The Four Ways This Process Quietly Breaks

    The failures here are not dramatic. Nothing errors out. The process keeps producing tidy shortlists while the ground under it has gone.

    The first is the numeric slide. You start with the discipline intact, then one day you ask the model a small quantitative question — "roughly what is this company’s debt level?" — because opening the filing is inconvenient. It answers. The answer is plausible. From then on the line between computed and generated figures inside your own notes is gone, and you will not be able to reconstruct which is which.

    The second is confirmation. If you tell the model which four names you already like before asking it to read thirty transcripts, you will get a reading that supports those four. Screening prompts must be blind to your existing preferences — no mention of what you own, what you are watching, or what you hope to find.

    The third is the illusion of coverage. Thirty transcripts processed feels like thirty companies understood. It is not. It is thirty documents skimmed by a process that reports what you asked it to look for and is silent about everything else. The set of things your prompt does not ask about is enormous, and it is invisible in the output by construction.

    The fourth is drift in the prompt. Edit the screening prompt halfway through a sector and the last fifteen outputs are no longer comparable with the first fifteen — but they still look identical, and the cross-document pass will happily find patterns in a set that was produced two different ways. Version and date the prompt, freeze it for the duration of a sector run, and note the version at the top of the output file.

    Numeric questions creep back in gradually; the defence is a rule, not vigilance.
    A screening prompt that knows what you already like will find reasons to like it.
    Processing a document is not understanding it — the prompt defines a narrow window and hides everything outside it.
    A prompt edited mid-run silently destroys comparability across the set.
    Every one of these failures produces output that looks exactly like success.

    Watch out — A shortlist is not a recommendation, and this article is not one either. Every company or sector referenced anywhere in this module appears purely to illustrate a method. Nothing here constitutes investment advice, and no process described here — however carefully run — can tell you what a business is worth or what a share price will do.

    Common questions

    No. A language model has no access to the NSE or BSE universe and cannot sort or filter it. When asked for companies meeting numeric conditions it generates plausible names with plausible-looking figures, and the output is indistinguishable from a real screener’s. Numeric filtering must run in a scanner connected to exchange data. AI’s role begins after that filter, on documents you supply.

    Knowledge Check

    Question 1 of 3Score: 0

    Why must numeric filtering happen in a scanner rather than in a language model?

    Rohit Singh — Mr. Chartist

    Written By

    Rohit Singh

    Mr. Chartist

    With 14+ years of experience in Indian financial markets, Rohit Singh (Mr. Chartist) is a SEBI Registered Research Analyst, Amazon #1 bestselling author, and the founder of Investology — a premium trading ecosystem trusted by a 1.5 Lakh+ strong community across India.

    INH000015297Full Bio

    Keep reading