intermediate10 min read11 of 24

    AI for Charts and Pattern Recognition — What a Vision Model Misses

    It matches pixels to descriptions it has seen. It does not measure price. That distinction decides every safe use of it.

    Rohit Singh

    Mr. Chartist · SEBI RA INH000015297

    Module

    Upload a candlestick chart to a modern AI model and it will say something remarkably plausible. It will name a structure, describe a base, mention where price seems to be finding support, and do all of it in the confident register of somebody who has been reading charts for years. The output is genuinely impressive. It is also produced by a process that has almost nothing in common with what you do when you read the same chart.

    A vision model is matching the pixels in your screenshot against the enormous number of chart images and chart descriptions it saw during training. It is answering the question “what does this picture most resemble?” It is not answering “what did price do here?” Those two questions have the same answer often enough to be dangerous and different enough to lose you money.

    This article is about the gap between them. Three structural blind spots that no amount of prompting fixes, one use where a vision model genuinely saves you time on a large watchlist, and one use where it should never appear. The dividing line is simple and it holds up: it can help you decide what to look at. It cannot decide what you saw.

    The one thing to remember

    A vision model tells you what a chart resembles, never what price actually did — so it is a fast way to decide which five names deserve your eyes, and never a reason to take a trade.

    What a Vision Model Is Actually Doing With Your Screenshot

    A vision model — a model that accepts images as well as text — converts your screenshot into a numerical representation and then generates a description of it, one word at a time, based on statistical associations learned from a very large collection of images and the text that accompanied them.

    That training set contained a lot of charts. Textbook diagrams, blog posts, screenshots from social media, annotated illustrations from trading courses. Alongside each of them sat text describing what the picture showed. The model learned the association between certain visual arrangements and certain words. When your chart arrives, it produces the words that arrangement is associated with.

    This is why the output sounds so knowledgeable. It is reproducing the vocabulary of chart commentary, correctly matched to the visual shape it detected. What it is not doing is measuring anything. It did not calculate the distance from the swing low to the breakout level. It did not count the candles in the consolidation. It recognised a shape and reached for the language that shape usually travels with.

    Hold on to that distinction, because everything else follows from it. A scanner running on price data compares numbers. A trained human reading a chart measures levels, counts candles and checks volume against the move. A vision model describes an appearance. All three can output the word “breakout”, and only two of them arrived at it by looking at price.

    The model is answering “what does this picture resemble?” You were asking “what did price do?” The answers overlap often enough to be genuinely dangerous.

    Three Blind Spots That No Prompt Can Fix

    The limitations here are not bugs waiting for a better model. They are consequences of the input being a picture. A screenshot simply does not contain some of the information you need, so nothing downstream can recover it.

    The first is scale. The image contains an axis with numbers on it, but the model’s reading of small axis text is unreliable, and even when it reads correctly it has no sense of what those numbers mean in proportion. A move that looks large on a compressed axis and a move that looks large on an expanded one produce the same visual impression.

    The second is everything outside the frame. Your screenshot starts somewhere. Whatever price did in the candles just before that edge — the ones that created the level you are now watching — does not exist as far as the model is concerned. A base that looks clean in the frame may sit directly under supply built up a few dozen candles earlier, entirely off-screen.

    The third is volume. Unless the volume panel is actually in the image and legible, the model has no volume information at all. And a breakout without volume confirmation is exactly the kind of breakout you were trying to filter out. This is the blind spot with the highest cost, because it removes the single most useful piece of confirming evidence on the chart.

    It reads the picture, not the price

    A vision model matches your screenshot to the language that shape usually travels with.

    Outside the frameWhat the screenshot containssupply builthere — invisible₹ ---₹ ---₹ ---volume panel not in the image — so it does not existpixelsWhat comes back“consolidation, then abreakout above the range”Correct vocabulary forthat shape. Nothing wasmeasured, counted orconfirmed against volume.Three things the image never carried1. Scale — the axis is decoration, so a large move and a trivial one look the same2. The left edge — the candles that built the level are simply not there3. Volume — absent unless the panel is visible and legible inside the frame
    What the screenshot carries and what it silently drops: the price scale becomes decoration, the candles before the left edge cease to exist, and volume is absent unless the panel is visible and legible in the image itself.
    The model has no reliable sense of the price scale, so it cannot judge whether a move is large or trivial in rupee terms.
    Anything outside the screenshot’s frame does not exist for it, including the structure that created the level you are watching.
    Volume information is present only if the volume panel is visibly and legibly included in the image.
    Nothing in the output tells you which of these it was missing, because the description reads identically either way.
    These are properties of using a picture as the input, so a better model does not remove them.

    It Does Not Know What the Prices Are

    Take a chart where price has moved from a consolidation into a fresh advance. On screen the move looks decisive: a long candle, clean separation from the range. Now consider that the vertical axis might span a very wide range or a very narrow one. The visual impression is nearly identical in both cases. The significance is not remotely identical.

    A human reads the axis and immediately converts the picture into rupees. That conversion is what tells you whether the breakout is worth acting on, where a sensible invalidation level sits, and what the position size should be. The vision model does not perform that conversion. It saw a long candle clearing a horizontal boundary and produced the words that pattern is associated with.

    You can help it by stating the levels in your prompt. Telling the model the current price, the range boundaries and the recent swing points gives it real numbers to reason over instead of an impression. This genuinely improves the output — but note what has happened. You supplied the measurement. The model is now working with your reading of the chart, which means it can no longer independently confirm anything about it.

    That is the honest ceiling. Either the model works from the picture, in which case it has no scale, or you give it the numbers, in which case its answer is downstream of your own analysis. There is no third arrangement where it independently verifies a level you have not already established yourself.

    Pro tip — When you do state levels in the prompt, state them as plain numbers with the timeframe attached, and ask the model to repeat them back before it reasons. If it repeats a level you did not give it, you have caught a fabrication before it entered your thinking.

    The Candles Just Off the Edge of the Screenshot

    Every chart screenshot is a window, and windows have edges. Whatever price did before the left edge is gone. The model cannot know it, cannot ask for it, and will not mention its absence.

    This matters most for supply and demand left behind by earlier structure. Suppose the visible portion shows a tidy consolidation with price pressing against the top of it. That reads as constructive. Now suppose that thirty or forty candles before your left edge, price had broken down through that exact zone on heavy volume. The area you are watching is not clean overhead space. It is the scene of a previous failure, and the people trapped there are still trapped.

    A human dealing with this simply zooms out. That is such an automatic reflex it barely registers as a step. The model has no equivalent — the image it received is the entire universe of what it knows about this instrument, and it will describe that universe with complete assurance.

    The practical fix is to include more history in every screenshot you send, and to send a higher-timeframe view alongside the one you are working on. That reduces the problem. It does not remove it, because the new image also has edges, and the model still cannot tell you what it might be missing beyond them.

    Watch out — The model never says “I may be missing context outside this frame”. It describes what it can see with exactly the same confidence whether the frame contains the whole story or a misleading fragment of it. Absence of a caveat is not evidence that the frame was sufficient.

    Volume Is the First Thing the Image Loses

    Most chart screenshots people share are price-only. The volume panel is switched off, or cropped away, or compressed into a strip too small for the bars to be distinguishable. In every one of those cases the model is describing price behaviour with no idea what participation looked like underneath it.

    That removes the most important confirming evidence you have. A move through a level on expanding volume and the same move on thin volume look nearly identical in the price panel alone. They are not the same event, and the difference between them is a large part of what separates a continuation from a failure that pulls straight back into the range.

    Even when the volume panel is included, treat its reading with caution. The model must resolve relative bar heights in a small strip of pixels to say anything useful, and it does not compute an average to compare against. “Volume appears elevated relative to the preceding bars” is roughly the ceiling of what an image supports, and that is a much softer claim than the one you actually need.

    So the rule follows directly. Always include the volume panel, at a size where the bars are genuinely distinguishable. Then read the volume yourself, from your own chart, at full size. Whatever the model says about it is a prompt to go and look, never a substitute for looking.

    Sending a chart to a vision model

    Do

    • Include the volume panel at a size where individual bars are clearly distinguishable.
    • Include substantially more history than the section you are focused on, so the relevant prior structure is inside the frame.
    • Send a higher-timeframe view alongside the working timeframe, as two images or two passes.
    • State the instrument, the timeframe and the candle interval in the prompt text.
    • Ask for a description of what is visible, and require it to say when something is not determinable from the image.

    Don't

    • Send a tightly cropped screenshot of only the section that looks interesting.
    • Ask it to confirm whether a breakout is genuine, which is a question the image cannot answer.
    • Accept a named pattern without going to your own chart and checking the structure yourself.
    • Ask for entry, stop-loss or target levels from an image, since it has no reliable price scale.
    • Assume the absence of a caveat means the frame contained everything that mattered.

    Where It Genuinely Helps — Triage Across a Large Watchlist

    There is one job a vision model does well, and it is worth doing. If you track a long watchlist and cannot give every name proper attention every week, the model can take a first pass and tell you which handful look structurally interesting enough to deserve your eyes.

    The economics of this are what make it work. A first pass that is roughly right is useful when the alternative is not looking at forty names at all. You are not asking it to be correct. You are asking it to be a slightly better filter than random, and to hand you a shortlist that is shorter than the watchlist.

    Crucially, the cost of an error is small in this direction. If it flags a name that turns out to be nothing, you have lost ninety seconds looking at a chart. If it misses something, you were not going to look at that chart this week anyway. Neither outcome puts money at risk, because nothing downstream of the triage happens without your own full read.

    Phrase the triage prompt as a sorting question rather than a judgement. “Which of these look like they are in a defined range, and which look like they are trending?” is answerable from a picture. “Which of these is about to break out?” is not, and asking it that way invites exactly the confident fabrication you are trying to avoid.

    Watchlist triage prompt — sorting, not judging
    I am attaching chart images for several names from my watchlist. Each image
    includes the volume panel and roughly the same amount of price history. Treat
    each one independently.
    
    For each chart, in this fixed format:
    
    NAME/LABEL: [as I have labelled it]
    STRUCTURE: one of — trending up, trending down, defined range, expanding
      volatility, unclear.
    CONSOLIDATION LENGTH: approximate number of candles the sideways phase covers,
      or “not applicable”.
    POSITION IN RANGE: near the upper boundary, near the lower boundary, mid-range,
      or not in a range.
    VOLUME PANEL: legible or not legible in this image. If legible, whether recent
      bars look larger or smaller than the preceding ones.
    NOT DETERMINABLE FROM THIS IMAGE: list what you cannot assess here — for
      example the price scale, history before the left edge, or volume.
    DESERVES A MANUAL LOOK: yes or no, with one sentence of reasoning.
    
    Rules:
    - Describe only what is visible in each image.
    - Do not state entry, stop-loss or target levels.
    - Do not say whether anything is a buy or a sell.
    - Do not tell me a breakout is confirmed. Confirmation is not available from a
      screenshot and I will do it myself.

    When to use — Weekly, on a watchlist too long to review chart by chart, purely to decide which names get your attention first.

    A good answer — The same fields filled for every chart, candle counts rather than dates, an honest “not legible” whenever the volume panel is too small, and a populated “not determinable” line on every single one. If that line is ever empty, the model is overreaching.

    Where It Must Never Be Used — Confirming a Trade

    Asking a vision model whether a breakout is real, and treating its answer as the reason to enter, is the failure mode this entire article exists to prevent. Every blind spot compounds at exactly that moment.

    Consider what confirmation actually requires. You need the level established from prior structure, which needs history the frame may not contain. You need the move measured against the price scale, which the model does not reliably read. You need volume expansion on the move, which is absent unless the panel is legible. You need to know how price behaved on the retest, which is about candles the model cannot count reliably. The image supports none of these, and the answer will still arrive fluent and specific.

    There is a second problem underneath the first: you asked a leading question. Ask “is this a valid breakout?” and you have already supplied the frame. The model will tend to work within it, listing features consistent with the premise you handed over. That is not confirmation. That is a well-written restatement of your own hope.

    The house discipline here is unchanged by the technology. Confirmation comes from your own read of price and volume on your own chart at full size — where the level came from, whether the move carried participation, how many candles the consolidation ran, and what happened when price came back to test the level. That work is yours. A screenshot description does not do it and cannot be made to.

    It can help you choose what to look at. It cannot tell you what you saw, and it must never be the reason you took the trade.

    What the Manual Read Still Has to Cover

    If the triage step worked, you are now sitting in front of five charts instead of forty. Everything that decides anything happens here, on your own screen, at full size, with your own eyes.

    Establish where the level came from before you assess whether it broke. Zoom out far enough that the structure which created it is on screen. Ask how many candles price spent building the range, and how many candles have passed since it left. Counting candles rather than dates keeps the reading tied to what actually happened on the chart instead of to the calendar.

    Then check participation. Compare the volume on the move against the volume through the consolidation. Then look at behaviour rather than position: whether price held above the level on the retest, whether it closed back inside the range, how the candles behaved when it came back to the boundary. Location tells you where price is. Behaviour tells you what it did when it got there, and that is the part that carries information.

    You can still use the model at this stage — just not as the judge. Describing your own read to it in words and asking what evidence would contradict you is a useful exercise, because it is working with your measurements rather than a picture. Use it to argue against your reading, never to bless it.

    Arguing against your own read — text, not images
    Here is my own reading of a chart. I measured all of this myself from a
    full-size chart; you are not being shown an image.
    
    - Instrument type: [large-cap NSE equity / index / etc.]
    - Timeframe: [daily / weekly] candles
    - Range: price consolidated for roughly [N] candles between [lower] and [upper]
    - Move: price closed above [upper] [M] candles ago
    - Volume: the breakout candle’s volume was [larger / similar / smaller] than the
      average of the consolidation candles
    - Retest: price returned to [level] after [K] candles and [held / closed back inside]
    
    Do this and nothing else:
    1. List the evidence in what I have written that argues AGAINST my reading.
    2. List what I have not told you that would materially change the picture.
    3. Name the specific observation that would tell me I am wrong.
    
    Do not agree with me. Do not suggest an entry, a stop or a target. Do not tell
    me whether to take this trade.

    When to use — After your own manual read, before you act on it — as a structured way to hunt for what you might have skipped over.

    A good answer — A genuine list of gaps rather than agreement: the higher-timeframe structure you did not mention, the volume through the retest you did not describe, and a specific, observable condition that would invalidate your reading.

    Common questions

    It can describe what a chart resembles, which is not the same as reading it. A vision model matches your screenshot against chart images it saw in training and produces the associated language. It does not measure the price scale, cannot see candles outside the frame, and has no volume data unless the volume panel is visibly included and legible.

    Knowledge Check

    Question 1 of 3Score: 0

    A vision model looking at your chart screenshot is fundamentally doing what?

    Rohit Singh — Mr. Chartist

    Written By

    Rohit Singh

    Mr. Chartist

    With 14+ years of experience in Indian financial markets, Rohit Singh (Mr. Chartist) is a SEBI Registered Research Analyst, Amazon #1 bestselling author, and the founder of Investology — a premium trading ecosystem trusted by a 1.5 Lakh+ strong community across India.

    INH000015297Full Bio

    Keep reading