Editorial Note: The opening scenarios are fictional composites based on recurring patterns in quantitative trading practice. Institutional examples and research findings are drawn from publicly available sources. This article examines how AI is being used in quantitative trading and where current evidence remains limited; it is not investment advice.
How AI Is Actually Used in Quantitative Trading — and Where It Still Breaks
Two stories from the trenches: a backtest that worked, and a factor factory that didn't
Marcus, 32, ex-SaaS engineer. He left a stable job two years ago to trade on his own. Sixty thousand dollars in the account, Python on the laptop, a retail broker's API, no team, no risk colleague. Days are spent shipping code; nights are spent running backtests. When he gets stuck, he asks an LLM.
Last month he generated a mean-reversion strategy with one: go long when price sits more than one standard deviation below its 20-day mean, close when it reverts. The backtest printed a Sharpe of 2.1, a controlled drawdown, and an equity curve that looked like a textbook figure. Three weeks live, the account was down 6%.
He went back and read the generated code line by line. The bug lived in two lines:
signal = (close - close.rolling(20).mean()) / close.rolling(20).std()
position = np.where(signal < -1, 1, 0) # decision uses this bar's close
returns = position * close.pct_change() # fill also uses this bar's close
The signal was computed from the close of the current bar, and the fill was booked at that same close. The backtest therefore assumed he knew the closing price before it printed. Look-ahead bias like this usually gets caught in a code review when a human writes it. In LLM-generated code it slides through, because the syntax is clean, the logic is self-consistent, and the comments are helpful. The second problem is that the snippet contains no slippage and no commissions — and a 20-day mean-reversion system is not a low-turnover strategy.
Priya, 26, researcher at a mid-sized systematic fund. Fourteen months in, a six-person team, responsible for alternative data and factor research, data science background. Last month the team switched on an internal LLM assistant, and her assignment was to surface testable hypotheses from published research.
One afternoon produced forty factor candidates, each shipped with a smooth piece of financial reasoning: inventory turns leading earnings revisions, hiring efficiency predicting subsequent revenue, capex cadence signalling a demand inflection. She handed the list to risk. Thirty-eight were cut for duplication, unstable out-of-sample behaviour, or an economic story that wouldn't hold. The two survivors turned out to be variants of factors that were already crowded.
The failure modes differ — one is engineering, one is process — but they point at the same thing: the cost of producing a hypothesis has fallen sharply, while the capacity to falsify one has not moved with it.
Which makes the question concrete: along the quant pipeline, where does AI actually hold up, and where does it only look like it does?
Where AI sits in the quant pipeline: from research to order placement
A systematic investment shop's operations can be sketched as a pipeline: turn raw data into something computable, find signals in it, decide how much to hold and how to trade it, then execute and control risk. AI occupies very different positions at each stage, and the position tells you more than the model name does.
The most solid ground is at the top of the pipeline. Filings, press releases, news, earnings-call transcripts — material that used to be read one document at a time — now gets converted into computable fields in batches. Data plumbing and backtest scaffolding have also moved substantially. Man Group is deliberately modest about its own setup: an off-the-shelf LLM is roughly intern-grade, good at writing code and summarising research. That analogy is useful — it saves time, and it doesn't sign anything.
One stage down is signal generation, the factor-mining work covered below. This is where some of the most visible experimentation is taking place, and where published results diverge the most.
Further down is portfolio construction: weights, risk budgets, rebalancing cadence. AI has made progress here, but the public evidence is thinner than at the signal layer, for a straightforward reason — there is no single metric to optimise against. The constraints stack up, and "better" has to be defined before it can be measured.
The last two stages are execution and risk control. Machine learning has been in execution for well over a decade, but mostly as gradient-boosted trees and deep networks on order-book features; large language models are largely absent from the order path. Risk control keeps human gates in place across the industry.
The pattern in the public examples is fairly consistent: the closer a system sits to text and research, the more mature the documented LLM deployment; the closer it sits to order placement and millisecond-level execution, the less visible large-model deployment becomes.
Three non-technical questions do a lot of work when you're judging whether a given stage is genuinely production-grade. Can a human review the output? Is the cost of a single error bounded? Is it on the critical path to an order? A stage is easier to trust when its output can be reviewed, the cost of failure is bounded, and it is kept off the critical order path. If it sits directly on that path while review and failure controls remain vague, that's worth a follow-up.
LLM alpha factor generation: how large language models mine trading signals
The generator–implementer–evaluator loop
The most instructive public institutional case is Man Group's AlphaGPT, described as a three-person digital research team that never sleeps. Each role owns one segment of the research process:
- An ideator proposes testable hypotheses in bulk from prices, filings, news and published research, outputting strategy ideas in natural language — time window, how the signal is computed, expected direction.
- An implementer translates a hypothesis into production code against internal databases and tooling.
- An evaluator tests it against the same standards applied to human research: statistical significance, risk analysis, and whether the economic logic holds.
Two constraints come straight from Man Group's own material: the workflow does not run unsupervised, and any AI-generated signal must carry a clear economic explanation and clear the same bar as human research before it is considered for live capital.
Academic implementations share the structure with the roles replaced by components: an LLM proposes a factor expression, deterministic code implements it, an independent evaluator backtests and scores it, and the score feeds back for another pass. The value of iterating is moving a candidate from "computable" to "computable and explicable."
The detail that matters most is where authority sits. In the documented workflows discussed here, the LLM is confined largely to hypothesis generation and code assistance; judgement stays with the deterministic evaluator and with people. In practice, this points toward a recurring design pattern: let the model act as the researcher, while fixed code and explicit controls govern what reaches the trading system.
The story that ships with a factor is a bug, not a feature
A framework paper on LLM factor mining (Hubble, roughly 500 US equities, 104 surviving candidate factors) is candid in its limitations section about two cognitive traps, and it reads as some of the clearest writing on this topic.
The first is temporal leakage in the model. Pre-training corpora can contain financial literature, research and market events dated after the backtest's cutoff. Isolating the data temporally does not necessarily isolate the model temporally — the system may be recalling rather than inferring. Recent research on information leakage in LLM-based financial agents finds substantial degradation when evaluation moves beyond the model's knowledge window, highlighting the difficulty of separating genuine inference from memorised information.
The second is narrative bias. When an LLM produces a factor, it almost always attaches a fluent mechanism. The trouble is that a spurious signal wrapped in a plausible story is harder to reject than an obviously flawed one — reviewers get persuaded by fluency, and fluency carries no information about validity.
The engineering implication is blunt: treat a generated hypothesis as a prompt for scrutiny, not as evidence for it. The economic rationale should be constructed independently and attacked, rather than accepted from the model's own account.
What the numbers actually say
Once evaluation windows are lengthened and bias controls are added, published results get noticeably flatter.
FINSABER (Edinburgh, UCLA, Oxford; arXiv 2505.07078) extends evaluation to 2004–2024 across more than a hundred instruments and reports that previously observed advantages of LLM-based timing strategies deteriorate substantially under broader cross-sectional and longer-horizon testing. In the paper's reported comparisons, conventional benchmarks remain competitive across much of the evaluation, with some individual assets producing different results. The broader lesson is less about whether one model wins a particular asset and more about how quickly apparent advantages can weaken when the evaluation window and universe are expanded.
StockBench (arXiv 2510.02209) takes the other route, building a contamination-free multi-month environment where agents make daily buy/sell/hold decisions from prices, fundamentals and news. Its evaluation finds that most tested LLM agents struggle to outperform a simple buy-and-hold baseline consistently, although several models show positive returns and better risk characteristics. The paper's broader conclusion is that strong performance on static financial-knowledge tests does not automatically transfer into trading ability.
Hubble's out-of-sample result is similarly cautious: a fixed top-5 factor set shows mixed persistence across the 2025–2026 holdout period, with the weakest in-sample trend factor decaying materially. The authors also note that they did not run cost-sensitive backtests, making turnover an important unresolved question rather than a settled result.
Taken together, the defensible role for LLMs in factor mining is hypothesis engine plus research assistant: they widen the hypothesis space, compress coding time, and sweep corners that human attention doesn't reach. What they don't supply is falsification.
AI portfolio optimization: the quieter layer, and the harder one to verify
Portfolio construction attracts visibly less public discussion than factor mining, and the reason is structural. There is no information coefficient to chase here. The objective carries risk budgets, turnover limits, capacity constraints, impact costs and exposures to multiple risk factors at once, so "better" has to be defined across dimensions before it can be measured.
Three technical threads dominate. Reinforcement learning treats rebalancing as a sequential decision problem, with market features as state and weight adjustments as actions. Generative and diffusion models are used to improve priors and covariance estimates, supplementing or replacing frameworks like Black-Litterman that depend on subjective view inputs. Multi-objective agents allocate across sub-portfolios with different risk preferences, targeting drawdown reduction rather than return enhancement.
The difficulties cluster in three places. First, the gap between simulator and market: reinforcement learning performance depends heavily on simulator fidelity, and impact, liquidity drying up, and queue behaviour are exactly where simulators are weakest. Second, more degrees of freedom to overfit: factor mining searches expressions, while portfolio optimisation also tunes constraints, penalties and rebalancing rules — an order of magnitude more search space, and correspondingly harder out-of-sample validation. Third, regime coverage: an allocation policy trained through a trending market can fail wholesale when style rotation accelerates.
The practical status is closer to "usable in research, cautious in production." There is also an institutional reason: portfolio and risk decisions are where human governance is most entrenched — position limits, kill switches, investment committee sign-off. AI tends to enter these as decision support rather than as decision maker.
AI in high-frequency trading: latency, execution, and where LLMs don't go
One common misunderstanding needs clearing first. Machine learning has been central to this track for well over a decade, but the workhorses are gradient-boosted trees and deep networks on order-book features, plus reinforcement learning for execution optimisation. Large language models are largely absent from the critical order path, and the reason is not capability — it is latency, determinism, reproducibility and auditability. The same prompt returning different conclusions on two runs is not workable when tick-to-trade is measured in microseconds.
The economics of the track are readable from public filings. XTX Markets reported combined net revenue of £3.93 billion across its three UK entities in 2025, up 43% year on year, with net profit of £1.71 billion, up 33%, on roughly 250 employees, average daily volumes around $250 billion, and a data centre build-out in Finland. Output per head at that scale points to competition driven by algorithms, infrastructure and talent density rather than by wiring up a frontier model.
A recurring engineering pattern in the public examples is to keep large models offline or on an asynchronous side channel, where they can assist with research, parameter analysis, anomaly attribution or risk review, while the live order path remains deterministic and tightly constrained, with low-latency infrastructure — FPGA included where appropriate — carrying execution.
Counter-evidence is equally concrete. Between 18 October and 4 November 2025, Nof1's "Alpha Arena" gave six frontier models $10,000 each of real capital to trade crypto autonomously for 17 days. Two open-weight models finished positive (+22.32% and +4.89%); all four closed-source frontier models finished negative, the worst at −62.66% (a loss of $6,266). Win rates ranged from 24% to 30%, and the highest fee line item was $1,331. Reading filings, writing analysis and explaining a mechanism clearly is a different skill from managing position size, controlling turnover and absorbing slippage.
A second development worth logging: on 27 May 2026, Robinhood launched "Agentic Trading," allowing customers to connect AI agents to trade through the platform, with built-in safety controls and a real-time activity feed. The launch is a notable example of agentic trading reaching the retail brokerage layer as a product. Its use of account controls, monitoring and revocation mechanisms also illustrates an engineering response to the risks of delegating trading actions to AI agents.
Factor crowding and alpha decay: what 2026 taught the industry
The mechanism is simple. AI lowers the marginal cost of research, similar hypotheses get generated in bulk, factors converge, positions converge, and drawdowns converge with them. The important part is that this compounds: more firms using similar tools to mine similar signals accelerates homogenisation faster than a linear addition would suggest.
Crowding is directly observable. Goldman Sachs' Prime Book data in mid-2026 showed long-crowding and short-crowding factor exposures both approaching five-year extremes, with medium-term momentum exposure around the 98th percentile. July then brought violent deleveraging: the largest two-day unwind of global technology longs since January 2021, the largest recorded position clearing in memory-chip names, and net exposure in the "Magnificent Seven" falling to roughly the 3rd percentile of its history.
To be accurate, though, "AI causes crowding" remains an open risk rather than an established finding. The NBER study on AI in asset management, issued in May 2026, finds lower return comovement among AI-driven funds than among non-AI funds, rather than evidence of greater herding. It also finds that early outperformance of AI-driven hedge funds diminished over time and was statistically indistinguishable from zero in later years. In comparisons with conventional sibling funds managed by the same adviser, however, AI-driven funds continued to show an average return advantage of roughly 35 basis points per month. The evidence therefore points to performance decay over time without establishing that AI-driven strategies are becoming more homogeneous.
Jane Street recorded a reported monthly loss of roughly $15 billion in July, its first negative trading month since 2016, with losses linked to its exposure to Situational Awareness and to positions in Asian and technology stocks. Situational Awareness, an AI-focused investment fund, fell roughly 67% in July and sold most of its public-equity portfolio to Citadel after margin calls. These events are relevant here not because they demonstrate that AI trading fails, but because they show how concentrated positions, leverage and liquidity can overwhelm a strategy regardless of how sophisticated its research process appears.
The variables driving that sequence are concentration, leverage and evaporating liquidity. They do not establish that AI trading fails, and treating them as such evidence is a shortcut. What they do show is that the boundary between market making, proprietary trading and directional betting is quite blurred at the top of the industry, and that regulators have started looking seriously at non-bank intermediaries' funding chains — both the Federal Reserve and the Bank of England have tightened scrutiny of banks' exposures to firms like these.
One side observation: events of this kind damage the narrative around AI trading more than the technology. Public discussion routinely merges the two, even though their evidentiary bases are entirely different.
Model risk and governance: what the Two Sigma case changed
In January 2025 the SEC penalised Two Sigma. Per the regulator's description, employees identified a flaw in an investment model around March 2019, the firm did not address it until August 2023, and it ultimately repaid investors $165 million voluntarily and paid a $90 million civil penalty.
The case is useful because it converts "model risk" from an academic topic into an operational one: a known defect persisted in production for more than four years while continuing to inform investment decisions. The corresponding governance requirements are concrete — full lifecycle model management, change logging, independent review, and post-incident analysis.
Practice is uneven. Human-in-the-loop controls and kill switches are common at systematic firms, and several large institutions describe the same arrangement publicly: machines generate decisions, humans own risk and data, and no fully unsupervised automated trading runs. Written AI usage policies, explicit model admission standards and periodic audits remain in the minority.
Real deployment or slide deck: how to read an AI trading claim
The table below is not an argument against AI. It is a way to separate engineering from marketing. The right-hand column lists probabilistic flags — any single one may be innocent, combinations deserve attention.
| Common claim | What to ask | What a credible answer looks like | Flags worth noting |
|---|---|---|---|
| "Our AI strategy returns X% annualised" | Out-of-sample and live periods, regime coverage, cost assumptions | Rolling out-of-sample and live track records shown separately, with fee and slippage assumptions stated | Single year only, backtest only, costs undisclosed |
| "The LLM produced N new factors" | Independent economic rationale, multiple-testing treatment, turnover and capacity | Prior hypothesis stated, FDR or White's Reality Check–style correction, turnover reported | Factor count as the headline, no multiple-testing correction, high turnover unaddressed |
| "AI participates directly in order placement" | Whether it is on the critical path, how determinism and latency are guaranteed, degradation behaviour on failure | Model on a side channel, deterministic code live, hard limits and a switch | No clear failure-degradation path |
| "Reinforcement learning for allocation" | Simulator-to-market gap, how hyperparameters were chosen, robustness across regimes | Parameter sensitivity and cross-period results reported | Works only in a single trending window |
| "The model is interpretable" | Whether the explanation drives decisions or is generated afterwards | Explanations feed a risk veto process and are used to kill signals | Explanations used only for external presentation |
| "We have AI governance" | Written policy, change logs, independent review, incident post-mortems | Policy documents and audit records can be produced | Verbal commitments only |
| "AI helps us beat peers" | Return correlation with peers, factor crowding, strategy capacity | Crowding and capacity monitoring reported | Crowding and capacity never mentioned |
What changed is the workflow, not the direction of the market
Across the three threads the common ground is clearer than the differences. In factor mining, AI compresses the time cost of generating hypotheses and writing code, while the burden of falsification sits entirely with people and deterministic programs. In portfolio optimisation, AI enters as decision support, and position limits, kill switches and committee sign-off have not moved. In execution, large models are parked on a side channel and live systems run fixed code.
The change is concentrated in process rather than in market direction. When "sounds plausible" becomes cheap, the scarce resource moves from ideas to verification: how multiple testing is handled, what cost assumptions are used, how out-of-sample periods are cut, whether an economic rationale can be independently overturned. The quality of those steps determines whether an institution ends up with capacity or with noise.
There is not yet an agreed standard for telling those two apart. Firms define out-of-sample differently, treat costs differently, and apply different admission thresholds to AI-generated signals, and disclosure is not standardised. That is worth holding in mind whenever this topic comes up.
Read More of Intelligenr
-
Does AI Make You Worse at Your Job? The Feedback Problem Nobody Measures
-
The Evolutionary Tree of AI: Why Transformers Dominated & What's Next
References
-
Man Group (2025) What AI Can (and Can't Yet) Do for Alpha https://www.man.com/insights/what-ai-can-do-for-alpha
-
Shi, R., Yan, S., Cai, Y., & Lv, C. (2026) Hubble: An LLM-Driven Agentic Framework for Safe, Diverse, and Reproducible Alpha Factor Discovery https://arxiv.org/abs/2604.09601
-
Li, W. W., Kim, H., Cucuringu, M., & Ma, T. (2025) Can LLM-based Financial Investing Strategies Outperform the Market in Long Run? https://arxiv.org/abs/2505.07078
-
Li, X., Zeng, Y., Xing, X., Xu, J., & Xu, X. (2025) Profit Mirage: Revisiting Information Leakage in LLM-based Financial Agents https://arxiv.org/abs/2510.07920
-
Chen, S., Sialm, C., & Xu, D. X. (2026) The Growth and Performance of Artificial Intelligence in Asset Management https://www.nber.org/papers/w35273
Author Note: This article focuses on the research and verification layer of AI-assisted trading. The aim is not to predict whether AI will outperform, but to examine what changes when generating trading ideas becomes cheaper than validating them.