How AI Models Forecast Stock Trends — and How Much to Trust Them
An ML price model is a pattern-matcher over history, not a crystal ball. How ours actually works, how backtests get inflated, and why an honest 55% is worth more than a claimed 90%.
PSX Expert Editorial
Market research desk
Published 24 July 2026
Updated 20 August 2026
8 min read
We run an AI model on this site. Every stock page carries its signal. That makes us either the best or the worst people to explain how these models work — best because we build one, worst because we have something to sell. Judge for yourself by the end. The short version: machine learning applied to stock prices is real, modestly useful, and dramatically oversold, and the entire difference between a useful model and a marketing prop is how honestly its accuracy is measured.
A model is a pattern-matcher, not a forecaster
Strip away the branding and an ML price model does one thing. It takes each stock on each historical day and reduces it to a row of numbers — engineered features such as the return over the last five, twenty and sixty sessions, momentum relative to the market, recent volatility, volume against its own average, and fundamentals like earnings trend and valuation multiples. Then it looks at what happened over the following days, across hundreds of thousands of such rows, and learns which combinations of numbers were statistically followed by rises and which by falls.
That is the whole trick. The model never sees "OGDC" as a company with gas fields and a government shareholder. It sees a numeric row that resembles other numeric rows from the past. "Forecast" is a generous word for what it produces; "resemblance score" is closer. When today's pattern looks like patterns that historically preceded gains, the score goes up.
This is not a criticism. Pattern persistence in prices is real — momentum is one of the most stubbornly documented effects in finance. But it defines the ceiling. A model can only find patterns that repeat, and markets are adversarial: patterns that made money attract money until they stop working. Whatever edge exists is small, unstable, and fights for its life.
Out-of-sample is the only honest number
Every serious model splits its history three ways: a training set the model learns from, a validation set used to tune it, and a test set touched only once, at the end. In markets the split must be chronological — train on the past, test on the later period — because shuffling the years lets the model peek at the future.
Why the ritual? Because accuracy on training data is meaningless. A sufficiently flexible model can effectively memorise its training set and score near-perfectly on it while knowing nothing usable. Validation accuracy is better but degrades too: every time you tune settings to improve the validation score, you leak information from that data into the model. After enough rounds of tuning, the validation set has quietly become training data.
Only the untouched test period — data the model never saw and you never tuned against — produces a number you can believe. And even that number ages. The genuinely honest record is live, out-of-sample performance: predictions timestamped before the outcome, published whether or not they worked. Ours is on the model performance page, losing stretches included. If a signal service will not show you that, you have learnt what you needed to know.
How backtests get inflated
Whenever you see a spectacular backtest — and PSX Facebook groups are full of them — one of three mechanisms is usually doing the work.
- Lookahead bias. The strategy uses information that was not available at the time. The crude version tests a rule discovered in 2024 on data from 2018. The subtle versions are everywhere: computing a signal from today's close and pretending you bought at that same close; using financial statements as of their period-end date rather than the date they were actually published; using today's index membership for historical trades.
- Survivorship bias. The backtest runs only on companies still listed today. Every delisted, defaulted and merged-away company — precisely the ones that destroyed capital — has been silently removed from the exam. This is the same trick that makes tipsters look good, which we dissected in why most stock tips lose money, performed by a spreadsheet instead of a screenshot.
- Parameter mining. Test two hundred combinations of indicator settings, publish the best one. The winner of two hundred coin-flipping contests looks like a genius on the flips you selected them by. This is the core disease of technical-strategy backtesting, and it is why technical indicators cannot tell you most of what their sellers claim.
None of these requires dishonesty. They happen by default, to careful people, unless deliberately engineered out. Which is exactly why "we backtested it" means nothing on its own.
The accuracy numbers that matter
Suppose the inflation is dealt with and you get a clean number. Which number should it be? Not the one usually advertised.
| The claim you will see | The question that deflates it |
|---|---|
| "90% accurate" | Measured on what data? On training data, 90% measures memorisation, not skill. |
| "62% hit rate" | Direction only? If the average loss is twice the average win, 62% still loses money. |
| "300% backtested returns" | After spreads and slippage — and could those orders actually have been filled in thin stocks? |
| "AI-powered signals" | Where is the timestamped out-of-sample record, including the losing months? |
Hit rate — how often the direction call is right — is the headline number, and alone it is close to useless. What matters is hit rate joined to magnitude: being right often on small moves and wrong occasionally on catastrophic ones is a losing system with an impressive accuracy figure.
The second number almost nobody advertises is calibration: when the model expresses 70% confidence, is it right about 70% of the time? A calibrated model that honestly says "55% up" is useful, because it is telling you how much it does not know. An overconfident one is dangerous in proportion to how convincing it sounds.
And 55% is roughly what honest looks like. A genuine 55% directional edge, applied with sensible position sizing across many independent decisions, compounds meaningfully — that is casino economics, and casinos are profitable. A claimed 90% on stock direction is fraud or measurement error, full stop. If it were real, the owner would be quietly trading it with leverage, not selling it to you for a monthly fee.
What no model can see coming
Our model does not read the news. It does not know a CEO resigned, a plant flooded, a gas tariff notification landed, or the FBR changed a tax rule. Fundamentals reach it late, as numbers, after publication. Everything it knows is already inside historical prices, volumes and filed statements.
That blindness matters most at exactly the moments that matter most. Consider mid-2023: the policy rate was on its way to a peak of 22%, the market was priced for something close to default, and then the IMF staff-level agreement in late June flipped the regime. What followed was one of the strongest rallies in KSE-100 history. No price-history model positioned for that in advance, ours included — the patterns preceding it had previously preceded nothing but grind. Statisticians call this a regime break: the future stops resembling the training data, and a pattern-matcher is, by construction, wrong until the new regime has run long enough to become the pattern.
So a model is at its best in boring markets and at its worst at turning points — precisely where you most want help.
Frontier-market data makes everything harder
Models trained on PSX data inherit problems that S&P 500 models do not have, and it is worth knowing them because they shape how far any signal here can be trusted.
Thin volume is the first. A large share of listed companies trade so lightly that the daily close is set by a handful of small orders. Features computed from those prices are mostly noise, and a model trained on them learns noise. Signals mean more on deep, liquid names like OGDC, HBL or LUCK than on a stock that traded eleven thousand shares yesterday — check the volume on the stock page before you weight the signal at all.
Circuit breakers are the second. When a stock locks at its daily limit, the close is not a market-clearing price; it is a queue of unfilled orders wearing a price's clothes. A model reads it as a price anyway, and a backtest cheerfully "trades" at prices no human could have obtained.
History is the third. The PSX in its current form dates only from the 2016 merger of the Karachi, Lahore and Islamabad exchanges, and clean, survivorship-free digital data is shorter still. A model here has seen perhaps two full macro cycles. That is a thin book to have studied before an exam this hard.
The correct dose
Here is how we use our own model, and how we suggest you use it. It ranks. It does not decide. Run the screener to shortlist companies, do the fundamental reading in the order set out in how to read a stock page, and let the signal act as a tiebreaker and a timing nudge between candidates you already understand — one input, weighted by the live record on the performance page, never an instruction.
And notice what the model conspicuously does not output: how much to buy. A 55% model is an edge only if you survive the 45%, and survival is a position-sizing decision, not a forecasting one. The forecast is the part that looks clever. The sizing is the part that keeps you solvent — and no model on this site, or anywhere else, will make it for you.
Written by PSX Expert Editorial, Market research desk at PSX Intelligence — the desk that builds and publishes the models behind this site. More about who writes this.
This is education, not advice
Nothing here is a recommendation to buy or sell any security. We are not licensed investment advisers. Everything on this site is general information; your circumstances are not. See our full disclaimer.