beginner11 min read6 of 24

    Prompting for Market Research — Getting a Useful Answer, Not a Confident One

    Scope, source, structure and permission to say “not stated”. The four-part prompt that separates research from fiction.

    Rohit Singh

    Mr. Chartist · SEBI RA INH000015297

    Module

    Type "analyse Reliance" into any AI tool and you will get something back within seconds. It will be well organised, it will use the right vocabulary, and it will read like research. Look closely and you will find that almost nothing in it is checkable. There is no source, no period, no page reference, and no way to tell which sentences came from a document and which came from the model’s general sense of what a large Indian conglomerate is usually like.

    That output is not a failure of the tool. It is exactly what the question asked for. A vague question has no wrong answer, so the model produces the safest thing available — a fluent average of everything it has seen written about companies of that type. The vagueness was yours, and the model politely filled it in.

    A research prompt has four parts, and each one closes a specific gap where invention creeps in. Exact scope. The source, pasted in rather than recalled. A demanded output structure. And explicit permission to say "not stated in the source". Get those four right and the same tool that produced marketing copy starts producing extraction you can defend.

    The one thing to remember

    A model treats "I do not know" as a failure it should avoid, so unless you explicitly permit that answer, every gap in your source gets filled with something plausible instead of being reported as a gap.

    Why "Analyse Reliance" Returns Nothing You Can Use

    A model answers by producing what usually follows a question like yours. When the question is broad, what usually follows is a broad answer: a paragraph on the business, a paragraph on the sector, a nod to growth drivers, a nod to risks. It is a shape, filled with the most typical content that fits the shape. Nothing is being retrieved from a document, because you did not give it one.

    The trouble is that this output is indistinguishable from real extraction. It has the same confident register, the same tidy structure, sometimes the same specific-sounding numbers. There is no visual difference between a figure read off page 118 of a filing and a figure the model produced because a number of that magnitude is what usually appears in that sentence position.

    The deeper problem is that a vague prompt cannot be graded. If you ask for "an analysis" and receive an analysis, on what basis do you reject it? You have no criterion. Precision in the question is not politeness towards the machine — it is what gives you the standing to say the answer is wrong.

    So the first move in every research prompt is to narrow until the question has exactly one correct answer, or a small set of them. Which company. Which document. Which period. Which metric. Once a question is that narrow, a wrong answer becomes visible, and visible wrong answers are the only kind you can defend against.

    A question with no wrong answer cannot produce a right one. Narrow it until being wrong is possible.

    How to phrase the ask

    Do

    • Name the document, the financial year and the exact metric: "from the pasted segment note for FY24, list revenue by segment with the note reference".
    • State the unit you want and the period you want it for, so a figure cannot arrive stripped of either.
    • Ask one question per prompt, even when three related questions are on your mind.
    • Tell the model what to do when the answer is absent, in the same sentence that asks the question.
    • Use the company’s full registered name once at the top, so a similarly-named peer cannot be substituted.

    Don't

    • Ask for "an analysis", "your view", "key highlights" or "anything important" — none of these can be graded.
    • Ask whether a stock looks attractive, or what a figure implies for the share price.
    • Stack four questions into one message and accept a single blended paragraph in reply.
    • Use a ticker or a short nickname alone as the subject of the prompt.
    • Ask the model to "be accurate" or "do not hallucinate" — these read as tone instructions and change nothing structural.

    The Four-Part Research Prompt

    Every prompt that reliably produces usable research has the same four components. They are not stylistic preferences. Each one closes a specific route by which invention gets into the output, and removing any one of them reopens that route immediately.

    Scope tells the model exactly what it is being asked about, so the answer is about one thing rather than an average of many. Source supplies the text the answer must come from, so it is reading rather than recalling. Structure fixes the shape of the reply, which forces specificity and, more usefully, makes gaps visible. Permission tells the model that reporting an absence is a correct answer rather than a failure.

    The order matters as much as the presence. Scope and structure go above the pasted source. The source sits in the middle, clearly fenced. Then the permission rule and a one-line restatement of the question go underneath, so the last thing in view before the model starts writing is what you actually asked.

    Once you have written this once, you never write it again. It becomes a template with four or five blanks in it. The work in a research session shifts from composing prompts to selecting which section of a document to paste, which is where the work belongs.

    Anatomy of a research prompt

    Four working parts around a fenced source. Each one closes a specific route by which invention gets into the answer.

    A prompt broken into five stacked blocks. Scope and the demanded output structure sit above a fenced block of pasted source text. Below the source sit the permission clause allowing the answer "not stated in the source" and a one-line restatement of the question, so the ask is the last thing in view. Each block is annotated with the failure it prevents.1 · SCOPEabove the sourceCOMPANY: [full registered name]DOCUMENT: [annual report / concall]PERIOD + SECTION: [FY], pages [x–y]Closes: a blended answer aboutthe wrong company or period.2 · STRUCTUREabove the sourceReturn a table with these columns:Item | Figure | Period | Page ref | QuoteCloses: prose that smooths agap over with a transition.3 · SOURCEfenced, in the middle--- SOURCE START ---[paste one section here]--- SOURCE END ---Closes: recall. Fencing turnsthe task into reading.4 · PERMISSIONbelow the sourceIf it is not in the source, write:"not stated in the source".Do not estimate. Do not use memory.Closes: the pull towards aplausible answer over a gap.5 · RESTATED ASKbelow the sourceRestating the task: fill the table from the SOURCE text only.Closes: an instruction buriedfar above the reply.Written once, reused as a template. The work moves from composing to choosing what to paste.
    The four parts of a research prompt and where each one sits. Scope and structure frame the source from above; the permission rule and a restated question sit below it, so the request is the last thing in view.
    Scope: name the company, the document, the period and the metric — never a ticker on its own.
    Source: paste the relevant section rather than relying on what the model remembers about the company.
    Structure: demand a table or a fixed set of headings, so an answer cannot hide behind fluent prose.
    Permission: say in writing that "not stated in the source" is an acceptable and expected answer.
    Placement: instructions above the source, permission rule and restated question below it.
    Reuse: save the finished prompt as a template with blanks; composing it fresh each time invites drift.

    Part One — Scope Tight Enough to Fail

    Scope has four dimensions and dropping any of them lets the answer wander. The entity: the full registered name, not a ticker or a nickname, because several listed Indian companies share a first word and a model will happily blend two of them. The document: an annual report, a quarterly result, a concall transcript, an exchange filing. The period: the financial year or the specific quarter. And the metric: the exact line item you want, in the words the document itself uses.

    That last point is worth dwelling on. Financial statements use precise terms — revenue from operations, other income, finance costs, profit before exceptional items. Asking about "profit" when the statement reports four different profit lines forces the model to pick one for you, silently. Ask for the line by the name it carries in the document and the ambiguity disappears.

    Scope also decides what a good answer looks like before you see one. If you asked for segment revenue for FY24 from the pasted note, you know in advance that a correct reply is a short list of segments with figures and a note reference. Anything longer than that is padding, and anything without the note reference is unsourced. You are grading against a standard you set, not reacting to whatever arrived.

    The test for whether your scope is tight enough is simple. Read your prompt back and ask: could two different but reasonable answers both satisfy this? If yes, narrow it further before you press enter.

    What you askedWhat the model doesWhat comes back
    Analyse this companyProduces the typical shape of a company write-upFluent prose, no source, nothing checkable
    What are the key risks?Recalls risks common to the sectorGeneric risks that fit any peer equally well
    Summarise the FY24 annual reportCompresses whatever it holds, or recallsA summary with no page references anywhere
    From the pasted MD&A, list the risks management statesReads the supplied text and extractsManagement’s own stated risks, quotable
    From the pasted segment note FY24, give revenue by segmentReads one note and reports the rowsFigures with units, period and note reference
    The same underlying curiosity, asked four ways. Only the fourth produces something you could quote in your own notes with a reference beside it.

    Part Two — Paste the Source, Do Not Trust the Memory

    A model’s training has a cutoff date, after which it has seen nothing. It also does not hold a clean, indexed copy of the documents it trained on. What it has is a compressed statistical impression of enormous quantities of text. Asking it to recall a specific figure from a specific Indian company’s filing is asking it to reconstruct a detail from an impression, and reconstruction produces plausible detail rather than correct detail.

    Pasting the source changes the task from recall to reading. This is the single largest quality difference available to you, and it is bigger than any refinement of wording. A mediocre prompt over a pasted document beats an elegant prompt over the model’s memory, every time and by a wide margin.

    Fence the source clearly. Put a marker line before it and after it, and tell the model explicitly that the answer must come from between those markers and nowhere else. Without the fence, the model blends what you pasted with what it half-remembers, and the join is invisible in the output.

    Where the document is genuinely too long to paste, split it and paste one section per prompt rather than accepting a summary built on recall. A partial answer from a real source is worth more than a complete answer from an imagined one, because the partial answer tells you honestly what it does not cover.

    A mediocre prompt over a pasted document beats an elegant prompt over the model’s memory. Nothing else you change matters as much.

    Pro tip — Where a tool offers web search or document upload, it may genuinely be reading rather than recalling — but check the citation it gives you actually exists and actually says what the answer claims. A link is not a verification; opening the link is.

    Watch out — A model asked for a figure it does not have will often produce one anyway, correctly formatted and roughly the right magnitude. There is no visual difference between that and a figure read from a document. Never accept a number that did not come from text you supplied.

    Part Three — Demand a Structure, Because Prose Hides Gaps

    Ask for a paragraph and you will get a paragraph, and a paragraph can absorb an enormous amount of vagueness without looking vague. Transitions cover missing links. Qualifiers cover missing figures. The prose reads smoothly precisely because the gaps have been smoothed over.

    Ask for a table with fixed columns and the gaps become holes you can see. A column headed "figure" that is empty is obviously empty. A column headed "page reference" with nothing in it tells you instantly that the claim beside it is unsourced. Structure does not make the model more accurate — it makes the model’s inaccuracy visible to you, which is the part you actually need.

    Fixed headings do the same work for qualitative material. If you always ask for the same five headings from every concall transcript, then a heading that comes back thin is a signal about the transcript rather than about your prompt. Consistency of structure is what turns a series of one-off answers into something you can compare across quarters and across companies.

    Add one more heading than you think you need: a section for what is missing. Ask explicitly for the things you would normally expect in a document of this type that are absent from this one. An empty "missing" section is either good news or a lazy answer, and you will learn quickly which one you are looking at.

    Reusable filing-summary template
    You are extracting facts from ONE section of a company filing. Answer only from the text between the SOURCE markers. Do not use any knowledge of this company from outside that text.
    
    COMPANY: [full registered name]
    DOCUMENT: [annual report / quarterly results / exchange filing]
    PERIOD: [FY or quarter]
    SECTION: [section name], pages [x-y]
    
    --- SOURCE START ---
    [paste the section here]
    --- SOURCE END ---
    
    Return a table with exactly these columns:
    | Item | Figure (with unit) | Period | Page or note reference | Exact quote |
    
    Fill it for each of these rows, in this order. Use one row per item even when the answer is an absence:
    1. Revenue from operations
    2. Operating margin, or the two line items needed to compute it
    3. Total borrowings
    4. Promoter shareholding
    5. Related-party transactions — total value
    6. Auditor's opinion type, and any emphasis of matter
    7. The single largest change versus the comparative period shown in this text
    
    Then, below the table, three short lists:
    MANAGEMENT'S OWN STATED RISKS — quoted, not paraphrased.
    LANGUAGE WORTH NOTING — hedged, conditional or unusually specific wording, quoted exactly.
    NOT PRESENT IN THIS SECTION — items above that this text does not cover.
    
    Rules:
    - If an item is not in the source, write "not stated in the source" in the Figure column. This is a correct and expected answer. Do not estimate, and do not supply it from memory.
    - Do not compute any ratio unless both inputs appear in the text above. Where you compute one, show the two numbers used.
    - Every Exact quote cell must be a verbatim string from the source. If you cannot quote it, the row does not belong in the table.
    
    Restating the task: fill the table and the three lists from the SOURCE text only.

    When to use — Your default first pass on any filing section. Change only the four header lines and the pasted text — keeping the rules identical is what makes outputs comparable across companies and quarters.

    A good answer — A complete table where several rows honestly say "not stated in the source", every populated row carries a page reference and a verbatim quote, and the NOT PRESENT list is not empty. A table with no gaps at all from a single section is a warning sign, not a good result.

    Part Four — Permission to Say "Not Stated in the Source"

    This is the highest-leverage sentence in any research prompt, and it is the one people leave out. Written into the prompt, it says: if the answer is not in the text I gave you, reply "not stated in the source" rather than producing something.

    The reason it matters so much sits in how these systems are built. A model is optimised to produce a helpful, complete-looking response. Across the enormous body of text it learned from, a confident answer is overwhelmingly what follows a question, and an admission of ignorance is rare. Later training stages reward responses people rate as useful, and a blank is rarely rated useful. The result is a strong default pull towards answering — "I do not know" behaves like a failure state to be avoided rather than a legitimate output.

    You are not fixing that tendency. You are overriding it locally, for this prompt, by redefining what counts as success. When the prompt states that reporting an absence is a correct answer, the absence stops being a failure and becomes one of the things you asked for. That is a change in the task, not a change in the model.

    Say it in the imperative and be specific about the format. "If it is not in the source, write: not stated in the source." Then add the two clauses that close the side doors: do not estimate it, and do not supply it from memory. Without those, you sometimes get a helpful approximation instead of an honest gap, which is the exact failure you were trying to prevent.

    Unless you say otherwise, silence is the one answer the model has learned not to give you.

    Write the permission as an instruction, not a hope: give the exact phrase you want back.
    Add "do not estimate" and "do not supply it from memory" — permission alone leaves both doors open.
    Treat honest gaps as a quality signal: an extraction with no gaps from one short section deserves suspicion.
    Never write "do not hallucinate" instead — it reads as tone and changes nothing about the task.
    Repeat the rule once at the bottom of the prompt, after the pasted source, where it is last in view.

    Pro tip — Test any new template by asking it for something you know is absent from the source you pasted. If it invents an answer, the permission clause is not doing its job and the template needs fixing before you rely on it.

    Ask for the Quote Behind Every Figure

    Add one column to any extraction and the whole exercise changes character: the exact sentence the claim came from. Not a paraphrase, not a summary — the verbatim string, plus the page or note reference where it sits.

    This does two jobs at once. It gives you something to verify against, so checking a claim means finding one quoted sentence in the document rather than re-reading a section. And it acts as a filter at the moment of generation. A claim that has no source sentence has nowhere to put its quote, which makes an unsupported row conspicuous instead of invisible.

    It is not a guarantee. A quote can itself be fabricated, and a fabricated quote from a document you have in front of you is trivially caught — you search for the string and it is not there. That thirty-second search is the whole verification, and it is only possible because you asked for the quote.

    Concall transcripts are where this pays off most. The prepared remarks are written to be reassuring and are usually the least informative part of the call. The analyst question-and-answer section is where the specifics live: someone asks about a margin decline directly, and management either answers it or visibly does not. Quoting the exchange preserves that distinction, which a summary would erase.

    Concall Q&A extraction prompt
    Below is the transcript of one earnings call. Work only from the text between the markers.
    
    COMPANY: [full registered name]
    QUARTER: [e.g. Q2 FY25]
    CALL DATE: [date printed on the transcript]
    
    --- TRANSCRIPT START ---
    [paste the transcript, or the Q&A section only]
    --- TRANSCRIPT END ---
    
    Ignore the prepared remarks. Work only on the analyst question-and-answer section.
    
    For every analyst question asked, produce one entry with these five fields:
    
    QUESTION — what was asked, in one line, in the analyst's own terms.
    ASKED BY — name and firm, if the transcript states them; otherwise "not stated in the source".
    MANAGEMENT'S ANSWER — quoted verbatim. If long, quote the two most specific sentences.
    ANSWERED OR DEFLECTED — pick one: DIRECT (a specific figure or commitment given), PARTIAL (some specifics, key part unaddressed), or DEFLECTED (no specifics given). State in one line which part of the question went unanswered.
    NUMBERS GIVEN — any figure stated in the answer, with its unit and period. If none, write "none".
    
    After all entries, add two closing lists:
    
    REPEATED THEMES — subjects more than one analyst asked about, with the count.
    PRESSED BUT NOT ANSWERED — questions marked DEFLECTED or PARTIAL, listed in one line each.
    
    Rules:
    - Quote verbatim in MANAGEMENT'S ANSWER. Never paraphrase into that field.
    - Do not infer intent, tone or confidence. Report what was said, not what it suggests.
    - If the transcript does not contain an analyst Q&A section, say exactly that and stop.
    
    Restating the task: one entry per analyst question, from the transcript text only.

    When to use — Your second pass on any quarterly call, after the prepared-remarks summary. Run it on the same transcript each quarter so the DEFLECTED list can be compared over time.

    A good answer — Verbatim answer quotes you can find with a text search of the transcript, an honest mix of DIRECT and DEFLECTED labels, and a PRESSED BUT NOT ANSWERED list with something in it. If every question is marked DIRECT, the model is being agreeable rather than reading.

    Iterate In the Thread, Do Not Start Over

    When an answer is not quite right, the instinct is to rewrite the whole prompt and try again from scratch. That is usually the wrong move. The source is already in the thread, the structure is already established, and a follow-up question inherits both. Rewriting throws away context you paid for.

    Narrow instead. If the extraction is too broad, ask for one row of it in more detail. If a figure lacks its comparative, ask for the comparative from the same source. If a claim has no quote, ask for the quote and watch what happens — a request for a quote that cannot be met is how an invented claim announces itself.

    Reserve the fresh start for two situations. When you are switching to a different document, because a new source in an old thread competes with the one already there. And when a wrong answer has gone uncorrected for several turns, because by then it has settled into the transcript as established context and follow-ups will build on it rather than question it.

    The most valuable follow-up in market research is the comparison. One quarter’s extraction tells you the state of things. Two quarters in the same format tell you the direction, and direction is what a single snapshot can never give you. Because both extractions used an identical structure, the comparison is mechanical rather than interpretive — which is exactly the kind of work to hand to a machine.

    Quarter-on-quarter comparison prompt
    Below are two extractions I produced earlier, using the same template, from two different quarters of the same company. Compare them.
    
    COMPANY: [full registered name]
    EARLIER PERIOD: [e.g. Q1 FY25]
    LATER PERIOD: [e.g. Q2 FY25]
    
    --- EARLIER EXTRACTION START ---
    [paste the earlier quarter's extraction, including its page references]
    --- EARLIER EXTRACTION END ---
    
    --- LATER EXTRACTION START ---
    [paste the later quarter's extraction, including its page references]
    --- LATER EXTRACTION END ---
    
    Produce a table with these columns:
    | Item | Earlier | Later | Change | Both figures present? | References |
    
    Then three short sections:
    
    MATERIAL MOVES — items where the change is large relative to the item's own size. State the direction and the size. Do not explain why; the reason is not in this text.
    LANGUAGE CHANGES — wording that appears in one period and not the other, or that was hedged in one and firm in the other. Quote both versions side by side.
    DISAPPEARED OR APPEARED — anything disclosed in one extraction and absent from the other, in either direction.
    
    Rules:
    - Compare only items present in both extractions. Where an item exists in one only, put it in DISAPPEARED OR APPEARED and leave it out of the table.
    - Compute a change only where both figures are present in the text above and carry the same unit and basis. Otherwise write "not comparable" and say why in four words.
    - Do not speculate about causes, outlook or what any change means for the share price. Report the movement only.
    - Carry both page references into the References column for every row.
    
    Restating the task: compare the two pasted extractions and report movements, not explanations.

    When to use — Once you have two quarters extracted in the same format. It is the point at which the template stops being a summary tool and starts being a monitoring one.

    A good answer — A table where references from both quarters survive into every row, at least one honest "not comparable", and a LANGUAGE CHANGES section quoting both versions. Any sentence explaining why something moved, or what it means for the stock, is the model exceeding the task — delete it.

    Before you press enter

    Six checks, roughly twenty seconds. Run them until they stop being a list and start being a habit.

    • Have I named the company in full, the document, the period and the exact metric?
    • Is the source pasted and fenced, with an instruction to answer only from between the markers?
    • Have I demanded a specific output structure — a table or fixed headings — rather than a summary?
    • Have I written the permission clause, plus "do not estimate" and "do not supply from memory"?
    • Have I asked for a page or note reference and a verbatim quote beside every figure?
    • Is my exact question restated on the last line, below the source, so it is the final thing in view?

    Common questions

    There is no single best prompt, but every good one has the same four parts: exact scope, the source pasted in, a demanded output structure, and explicit permission to answer "not stated in the source". A plain prompt with all four beats a clever one missing any of them.

    Knowledge Check

    Question 1 of 3Score: 0

    Which instruction most reduces invented figures in a research prompt?

    Rohit Singh — Mr. Chartist

    Written By

    Rohit Singh

    Mr. Chartist

    With 14+ years of experience in Indian financial markets, Rohit Singh (Mr. Chartist) is a SEBI Registered Research Analyst, Amazon #1 bestselling author, and the founder of Investology — a premium trading ecosystem trusted by a 1.5 Lakh+ strong community across India.

    INH000015297Full Bio

    Keep reading