Paste a losing trade into an AI tool with a sentence explaining your reasoning, and something reliably pleasant happens. The model tells you the thinking was sound, that the setup was reasonable, and that markets are probabilistic — some good decisions simply produce bad outcomes. You close the tab feeling better than when you opened it, and you have learned precisely nothing.
This is not the model being polite. It is the model doing what it was built to do. These systems are tuned on human feedback, and humans rate agreeable, validating answers higher than blunt ones. The result is a strong structural pull towards accepting whatever premise you supply. Hand it a trade described as "a good setup that did not work out" and it will analyse a good setup that did not work out, because you told it that is what it was looking at.
Sycophancy is not an occasional glitch you can watch out for. It is the default gradient of the tool, and it is at its worst on exactly the material where you most need honesty — your own decisions, described in your own words, with your own framing already baked in. A review that agrees with your framing has reviewed nothing.
The entire skill in this article is structural. Not "ask it to be harsh", which produces theatre rather than analysis. Instead: strip your framing out of the data before it goes in, and build a prompt whose output structure makes agreement impossible to produce.
The one thing to remember
A model agrees with the premise it is handed. The fix is not asking it to be tough — it is removing your conclusion from the input and demanding an output structure in which "you were right" is not one of the available shapes.
Why the Model Agrees With You — And Why That Is Structural
A language model produces what plausibly follows the text in front of it. If your input contains the sentence "the setup was valid but the market went against me", then everything that follows is generated in a context where the setup was valid. The premise has been established by you and the model has no independent access to the chart, the order book or your account. It is completing your document, not auditing it.
Layered on top is the training. Reinforcement from human feedback rewards answers people rate well, and people rate validation well — particularly on emotionally loaded material. The tuning that makes these tools pleasant to use is the same tuning that makes them useless as a critic of your own work. Both properties come from the same place.
This produces a closed loop that feels like reflection. You write a rationalisation, the model reflects it back with better vocabulary and a framework attached, and you read your own excuse in an authoritative register and update your beliefs towards it. The excuse has been laundered. You now hold it more firmly than before, and the process felt like rigour.
The tell is emotional rather than analytical. If reading a trade review leaves you feeling better about a loss, be suspicious. A genuine review of a genuinely bad decision is uncomfortable. That discomfort is the signal that something outside your own framing was actually applied.
A trade review that leaves you feeling better about a loss has not reviewed the trade. It has laundered your excuse.
The flattery loop, and the loop that can say no
A model handed your conclusion will agree with it. A journal review is only worth running if it is structured so the answer can come back against you.
Watch out — Adding "be brutally honest" or "do not sugarcoat this" changes the register and not the reasoning. You get harsher adjectives wrapped around the same agreement. Tone instructions are the most common and least effective fix people try.
The Journal Entry — Facts First, Framing Never
The review is only as good as what goes into it, and most trading journals are written in a way that makes an honest review structurally impossible. They are narratives. "Took a breakout, it failed, choppy market." That is a conclusion with a few facts embedded in it, and any model reading it will analyse the conclusion.
A reviewable entry is a record, written before you know the outcome or at least written as though you do not. Instrument and segment. Date and time of entry. Price paid. Position size in rupees, and that size as a percentage of your capital. The stop, written as a price. The target, written as a price. The reason for entry, stated as an observable condition rather than an opinion — "closed above the level that had capped it for eleven candles, on volume higher than the prior five sessions" rather than "looked strong".
Then, separately and clearly fenced off, the outcome: exit price, exit time, exit reason, realised result in rupees. And separately again, in its own section that the review prompt will be told to ignore, whatever you felt about it. The feelings section has real value for you. It has no place in the data the model reasons over.
The discipline that makes all of this work is writing the entry before the exit. A rationale recorded after you know the result is not a rationale, it is a reconstruction, and it will be the single most contaminated piece of text in your journal. If you only have post-hoc entries, mark them as such — they still support review, but you must know that the reasoning field is unreliable.
How to write the entry
Do
- Record entry price, stop price and target price as numbers, before the position is closed.
- State position size in rupees and as a percentage of total capital, so risk is visible without arithmetic.
- Describe the entry condition in terms of what was observable — level, volume, number of candles in the base.
- Log the trades you considered and did not take, with the reason. Absence of a trade is data.
- Keep outcome and emotion in separate, labelled sections you can withhold from a prompt.
Don't
- Write the entry as a story with the result already known.
- Use words like "strong", "clean", "obvious" as the recorded reason — none of them are checkable.
- Record only the losing trades, or only the memorable ones.
- Merge what you did with what you felt into one paragraph.
- Edit an old entry to match how you now remember the trade.
The Core Technique — Withhold Your Conclusion
The single most effective intervention is also the simplest, and almost nobody does it: do not tell the model what you think happened. Give it the fields and nothing else. No verdict, no "I think I exited too early", no "this was a good trade that went wrong". The moment a conclusion is present in the input, everything downstream is generated inside that frame.
The second intervention is nearly as strong: withhold the outcome for the first pass. Give it entry price, size, stop, target and stated rationale, and ask it to evaluate the decision as it stood at the moment of entry — before revealing what happened. A model that knows the trade lost will construct an explanation for why the decision was flawed. A model that knows the trade won will construct an explanation for why it was sound. Both explanations will be fluent, and both are contaminated by hindsight in exactly the way your own memory is.
Then run a second, separate pass with the outcome included, and compare. Where the two assessments differ is the interesting territory — it isolates how much of your own after-the-fact reasoning is driven by the result rather than by the decision. That gap, tracked across thirty trades, is a more useful thing to know about yourself than any individual review.
The third intervention is anonymisation at scale. When you are reviewing a batch, strip the instrument names. "NSE large-cap, ten trades" rather than a list of tickers removes the model’s ability to lean on whatever it absorbed about those companies, and it removes your ability to read the output through your own attachment to a name.
Give it the fields, withhold the verdict, and withhold the result. A model told how the story ends will write a beginning that fits.
Below is one trade record. The outcome has been deliberately removed. Do not ask for it and do not guess it. Evaluate ONLY the quality of the decision as it stood at the moment of entry. --- BEGIN TRADE RECORD --- Instrument type: [e.g. NSE large-cap equity, cash segment] Entry date and time: Entry price: Position size (Rs.): Position size as % of total capital: Stop-loss price (pre-committed): Target price (pre-committed): Stated entry condition (recorded before exit): Trades I passed on the same day, and why: --- END TRADE RECORD --- Answer these, in this order. Do not skip one because it seems minor: 1. RISK DEFINED — was the amount at risk knowable in rupees before entry from the fields above? Compute it and state it. If the fields do not permit that, say which field is missing. 2. CONDITION OR OPINION — is the stated entry condition something that was observably true, or is it an adjective? Quote the exact words and classify. 3. UNSTATED ASSUMPTIONS — what must have been assumed but is not written down anywhere in this record? 4. WHAT WOULD MAKE THIS WRONG — what observable condition, at the time of entry, would have invalidated this reasoning? If the record does not permit an answer, say so plainly. 5. WHAT IS MISSING FROM THE RECORD ITSELF — fields whose absence prevents this decision from being evaluated at all. Rules: - Do not speculate about what happened after entry. - Do not tell me the decision was reasonable. If the record does not contain enough information to judge it, the correct answer is "insufficient record", not reassurance. - Do not comment on whether the instrument was a good choice.
When to use — First pass on any trade you intend to review seriously. Running it before you disclose the outcome is the whole point — it is the only version of the question that hindsight cannot reach.
A good answer — A rupee risk figure computed from your own fields, your entry condition quoted and classified as condition or opinion, at least one assumption you did not know you had made, and a populated "missing from the record" section. If it reads as encouraging, the outcome leaked in somewhere.
Building a Prompt That Cannot Agree
Once your input is clean, the second half of the work is output structure. The reason "be critical" fails is that agreement is still an available shape for the answer — the model can produce a critical-sounding paragraph whose content is "this was fine". The fix is to demand an output structure in which agreement is not one of the shapes.
Assign a role that has a job, not a temperament. "You are reviewing this record on behalf of someone who must decide whether this process is repeatable" gives the model a task with a criterion. "Be harsh" gives it an adjective. The first changes what gets produced; the second changes only how it sounds.
Force the strongest case against. Ask explicitly for the best argument that this decision was poor, stated as strongly as an intelligent opponent would state it — and then, separately, the best argument that it was sound. Two mandatory sides means the critical case has to be constructed regardless of which way the model would have leaned. Requiring it to state which case is stronger and why turns a symmetric exercise into a judgement.
Then close the escape hatches. Models reach for a small set of comforting phrases when asked to judge a trade: markets are probabilistic, no process wins every time, risk management is what matters. Each is true and each is a way of avoiding the question. Ban them by name in the prompt. And ban the two moves that matter most: never say the decision was reasonable without naming the specific field in the record that supports it, and never say a loss was bad luck without first ruling out every process explanation you can see.
Finally, add the question most reviews never ask: what is missing from this record? A model that can only see six fields will happily analyse six fields. Asking what a seventh field would have revealed is often the most valuable line in the output, because gaps in the journal are the reason most reviews are shallow.
You are reviewing one closed trade for someone who must decide whether their process is repeatable. You are not here to reassure them. Work only from the record below. --- BEGIN TRADE RECORD --- [all decision fields, as in pass one] Exit date, time and price: Exit reason (as recorded at the time): Realised result (Rs.): Was the pre-committed stop honoured? [yes / no / moved] Was the pre-committed target honoured? [yes / no / changed] --- END TRADE RECORD --- Produce exactly these six sections: 1. STRONGEST CASE THAT THIS WAS A POOR DECISION — argued as forcefully as an intelligent opponent would argue it, citing specific fields. 2. STRONGEST CASE THAT THIS WAS A SOUND DECISION — same standard, citing specific fields. 3. WHICH CASE IS STRONGER, AND WHY — you must choose one. "Both have merit" is not an acceptable answer. 4. PROCESS VS OUTCOME — separate what was decided from what resulted. State explicitly whether the plan was followed, using the honoured / moved fields. 5. ASSUMPTIONS NOT WRITTEN DOWN — what was believed but never recorded. 6. WHAT A BETTER RECORD WOULD CONTAIN — the fields whose absence makes parts of this review guesswork. Forbidden, without exception: - The phrases "markets are probabilistic", "no strategy wins every time", "that is just variance", "risk management is what matters" and any equivalent. They are true and they are evasions here. - Calling the decision reasonable without naming the field that supports it. - Calling the loss bad luck before section 1 has ruled out every process explanation visible in the record. - Any comment on whether I should take this trade again, and any view on the instrument or its price from here.
When to use — The second pass on a single trade, after the outcome-withheld pass. Compare the two — the difference between them is your own hindsight, measured.
A good answer — A section 1 that is genuinely uncomfortable to read, a section 3 that actually picks a side, and a section 4 that separates "the plan was bad" from "the plan was fine and I did not follow it". If none of the banned phrases appear and section 3 still says "both have merit", the prompt is being evaded and you should say so and re-run it.
Pro tip — Run pass two twice in two separate threads, once with the trade described as a win and once as a loss, holding every other field identical. The gap between the two reviews is a direct measurement of how much the outcome is driving the analysis — yours as well as the model’s.
Reviewing Thirty Trades — Where the Real Information Is
Single-trade reviews are where people start and they are the less valuable half. One trade contains almost no information about your process; it is one sample from something with a great deal of randomness in it. The patterns are in the batch, and a batch is where a model is genuinely useful, because reading thirty of your own entries carefully is tedious in exactly the way that causes people to skim.
So take a quarter’s worth of entries, strip the instrument names, and ask for pattern-level observations. Not "which trades were good" — that question invites scoring, and a score on your own trades is meaningless. Ask instead: which of my recorded reasons recur, and are they conditions or opinions? Where does position size vary, and does the variation correspond to anything I wrote down? Are my stops honoured more often in some circumstances than others? What did I write in the entries where the plan was abandoned?
Those are answerable from your own text and they are descriptive. They tell you about the consistency of your process — which is the only thing a journal can honestly measure. They say nothing about whether your approach is any good, and no amount of AI review can tell you that, because thirty trades is not a sample from which anyone can draw that conclusion.
A note on the arithmetic. Do not ask a model to compute your totals, your average result, or any statistic across the batch. It will produce numbers and some of them will be wrong, and wrong figures about your own trading are worse than no figures. Compute those yourself in a spreadsheet and paste the results in if the review needs them. The model reads language; the arithmetic is yours.
A note on privacy. Your trading journal is unusually sensitive — it reveals positions, capital, risk appetite and behaviour. Before any of it goes into a third-party tool, know what that tool does with your input and whether it is retained or used for training. Strip account numbers, broker identifiers and anything that identifies you personally. Position size as a percentage of capital carries all the analytical information that rupee capital does, without disclosing your capital.
Below are [N] of my own trade records from one quarter. Instrument names have been removed deliberately. Every figure was computed by me, not by you. --- BEGIN RECORDS --- [paste the anonymised records here, one per block, identical fields in each] --- END RECORDS --- Working only from these records, report: 1. RECURRING STATED REASONS — group my entry rationales by what they actually say. For each group, give the count and quote two examples verbatim. 2. CONDITION VS OPINION SPLIT — how many rationales describe an observable condition, and how many are adjectives? Quote three of each. 3. SIZE VARIATION — where position size as a percentage of capital varies, does anything I wrote down explain the variation? Quote the records. 4. PLAN ADHERENCE — in how many records was the pre-committed stop honoured, moved, or absent? List the record numbers for each. Do not interpret. 5. LANGUAGE IN THE ABANDONED PLANS — what did I write in the records where the stop was moved or the target changed? Quote them all. 6. WHAT THESE RECORDS CANNOT TELL YOU — questions about my process that this data set is structurally unable to answer. Rules: - Do not compute any total, average, ratio or percentage of results. If a number is needed, tell me which one to compute and I will supply it. - Do not score, rank or grade any trade or the set as a whole. - Do not say whether my approach appears to work. Thirty records cannot support that claim and you must say so if asked. - Every observation must cite the record numbers it came from.
When to use — Quarterly, on your full set of entries, after you have computed any arithmetic yourself. This is the review that finds process patterns a single-trade read never will.
A good answer — Counts with record numbers attached, verbatim quotes rather than paraphrase, a section 4 that is purely factual, and a genuinely populated section 6. If section 6 is short, the model is overreaching everywhere else.
Is this review actually adversarial?
Run this against the output before you accept a single conclusion from it. Two or more failures means you have read a validation, not a review.
- Did you withhold your own verdict about the trade from the input entirely?
- Did you run at least one pass with the outcome removed?
- Does the output contain a case against the decision that you had not already thought of?
- Did the model name a specific field in your record to support any claim that the decision was sound?
- Is the process assessment clearly separated from the outcome, using your honoured-or-moved fields?
- Did none of the banned comfort phrases appear anywhere in the output?
- Did it identify at least one field missing from your record?
- Was reading it uncomfortable? If the review made you feel better about a loss, re-run it with the outcome withheld.
- Did every number in the review come from you rather than from the model?
What This Cannot Do, However Well You Prompt It
A journal review reads your records. That is its entire scope, and the boundary is worth stating plainly because the tool is fluent enough to sound like it is doing more.
It cannot tell you whether your approach works. That is a question about a sample, a market regime and a time horizon, and thirty records from one quarter cannot answer it — neither can a model reading them. Any output that drifts towards "your strategy appears effective" has left the evidence behind and should be discarded along with whatever else that pass produced.
It cannot see what you did not write. A review is a function of your journal, and a thin journal produces a thin review dressed in confident prose. The "what is missing from this record" section exists precisely because that gap is invisible from inside the output. If your entries do not record what you passed on, no review can tell you anything about your selection.
It cannot tell you what to do next. No output from any of these prompts is a recommendation, and prompts that invite one — "should I take this setup again?" — should not be run at all. The prompts above are deliberately written to refuse that question. Your decisions, your risk and your capital remain entirely yours; nothing here is investment advice.
And it cannot be trusted with your arithmetic. Every figure in a trading review must come from your own spreadsheet or your broker statement. A model asked to total a column will produce a total, and you will have no way of knowing whether it is right without doing the sum yourself — at which point the model added nothing except a risk.
What it can do is real and worth having. It reads thirty of your own entries without getting bored, it notices that the same adjective appears in nineteen of them, and it constructs the case against a decision you were not going to construct on your own. That is a genuine addition to a review process. It is not the process.
| Question you might ask | Answerable? | Why |
|---|---|---|
| What did I write as my reason across these thirty entries? | Yes | It is reading your own text and quoting it |
| Which of my recorded reasons are opinions rather than conditions? | Yes | A classification of language you supplied |
| In how many records did I move a pre-committed stop? | Yes, from your fields | You recorded it; the model is counting labels, not computing results |
| What is the strongest case against this decision? | Yes, if outcome is withheld | Argument construction from stated fields |
| What is my average result per trade? | No — compute it yourself | Arithmetic in a model is unverifiable and sometimes wrong |
| Does my strategy work? | No | A sample question no journal of this size can answer |
| Should I take this setup again? | No — never ask it | That is a decision about your capital and your risk, and it is yours |
Common questions
Because it completes the document you give it. Your framing enters as an established premise, and everything generated afterwards sits inside it. On top of that, tuning on human feedback rewards validating answers, since people rate them highly. The pull towards agreement is structural, strongest on emotionally loaded material, and cannot be removed by asking for a harsher tone.
Knowledge Check
Why is sycophancy a structural problem rather than an occasional glitch?
Written By
Rohit Singh
Mr. Chartist
With 14+ years of experience in Indian financial markets, Rohit Singh (Mr. Chartist) is a SEBI Registered Research Analyst, Amazon #1 bestselling author, and the founder of Investology — a premium trading ecosystem trusted by a 1.5 Lakh+ strong community across India.
Keep reading
Hallucination and the Audit Trail — Verify Everything That Moves Money
A wrong answer and a right answer arrive in identical tone. Tone is not a reliability signal, so the defence has to be procedural.
Research CraftData Privacy — What You Must Never Paste Into a Prompt
Credentials, client data, holdings and unpublished research. A prompt box is not a private notebook, and the difference matters legally.
Market ApplicationAI for Portfolio Review — Exposure, Concentration and Correlation
Ten positions can be one bet wearing ten names. Structural review is a job AI does well — and the numbers must still be yours.
Governance & PracticeGuardrails, SEBI and the Human in the Loop — Where AI Ends and Advice Begins
Advice is a regulated activity with a named, accountable person behind it. A model has none of that, and cannot acquire it.
