Editorial Note: This essay explores AI architectures through the lens of evolution: not as a race toward a single “best” model, but as a continuous process of adaptation, trade-offs, and recombination. The goal is not to predict which architecture will replace another, but to provide a framework for understanding why certain approaches succeed, where their limitations emerge, and how to read the signals of architectural change.
The Evolutionary Tree of AI: Why Transformers Dominated — and What Comes Next. Why AI Types So Slow, Hallucinates So Confidently, and Where the Field Is Headed.
Two Moments You’ve Definitely Had With AI
Moment one: watching the cursor blink.
You ask an AI to draft a market analysis. The cursor pulses, spitting out words one at a time. You go make coffee; when you return, it’s still on paragraph three.
You upload a 100-page contract as a PDF. It replies, politely, that the file exceeds its context limit.
You toggle on “deep thinking” mode and pose a math problem. The screen tells you it spent 47 seconds reasoning. You can’t see what it was doing, but you know those 47 seconds burned through six thousand tokens and your bill is climbing. Then it outputs three lines of answer.
Three ordinary afternoons, three mild frustrations. They feel like engineering problems — wait a bit, upgrade, next version will be better. But look closer and each scene encodes an uncomfortable question: Why must it speak one word at a time? Why does reading more cost it exponentially more, while for humans longer texts just mean more fatigue? Why does its “thinking” require it to talk more?
Moment two: confidently wrong.
Switch scenes. You ask, “What happens if I put a ceramic mug in a paper bag and lift it?” The AI gives you a well-structured, courteous answer about paper tensile strength and center of gravity — fluent, complete, and physically absurd.
Or you ask for book recommendations in an obscure field. It names two titles that don’t exist, with plausible authors and publishers, everything except reality.
A picture forms in your mind: it has read ten thousand books about swimming, but it has never been in the water.
The answer to both moments is the same: these are not bugs. They are expressions of the current AI species’ genetic code.
What follows is an anatomy of that code. One disclaimer up front: this article does not cover the open-vs-closed debate. That is a different war. Here we examine only the technical genome.
Many of the AI systems people interact with daily — including ChatGPT, Claude, and Gemini — belong to the Transformer lineage. In 2017, a paper with a deliberately provocative title, Attention Is All You Need, established its dominance. Nine years later, it still sits on the throne.
But on the evolutionary tree, Homo sapiens was never the only hominin. Neanderthals, Denisovans, Homo erectus all coexisted with our ancestors. Some left genes behind before vanishing; others were absorbed into our bodies. Today’s AI stands at the edge of its own multi-species era. The reigning king carries unfixable flaws in its DNA. Extinct branches donated genes to the winner. Several challengers have reached the gate.
This article does five things: explains why this species won; identifies the diseases baked into its genes; inventories what the dead branches left behind; introduces the challengers at the door; and hands you a ruler you can keep — so the next time you encounter an AI headline, you can judge for yourself what is signal and what is noise.
For now, hold four words in mind: sequence, memory, representation, reasoning. They are the coordinate system for everything that follows.
How the Transformer Works: From Telephone Game to Round Table
The telephone game, replaced by a round table
Before 2017, AI’s language world ran on a telephone game.
In recurrent neural networks (RNNs), information flowed serially: the first word passed to the second, the second to the third, and by the end of a long sentence the message had degraded beyond recognition. Worse, it was inherently sequential — no one could speak until the person before them finished. This serial nature made parallel computation impossible. Scale up the model or the data, and training speed collapsed.
The Transformer changed the game into a round-table meeting. Every token takes a seat simultaneously. Each token can directly attend to every other token, scoring relevance on the spot. In the sentence “The cat couldn’t fit in the basket because it was too big,” the pronoun “it” had to be resolved through layers of degraded transmission in the telephone game. At the round table, “it” and “cat” lock eyes directly, in a single step.
The essence of this shift deserves its own sentence, because it is among the most important in this entire article: The Transformer’s breakthrough was not simply that it understood language better; it was that it converted a traditionally sequential problem into a highly parallelizable computation problem. Sequential problems must queue; matrix computations parallelize completely. AI’s fate turned on this pivot.
Three components
The round table runs on three parts.
The first is self-attention. Each token assigns a relevance score to every other token, then aggregates information weighted by those scores. The cost hides in “every other token”: N tokens require N×N pairwise computations. Double the text length, quadruple the compute. Hold that number; Part 2 will turn it into the most expensive line item in the entire stack.
The second is positional encoding. A round table has no inherent order — “cat chases dog” and “dog chases cat” look identical from above. So each token must be stamped with an artificial seat number, injecting sequence information after the fact. This is a patch: RNNs had order natively; Transformers bolt it on. Patches carry hidden liabilities, and this one seeded the long-context extrapolation problems we’ll revisit shortly.
The third is a fork in the road: why the “next-word” camp beat the “fill-in-the-blank” camp. Two training philosophies competed. BERT-style masked modeling excelled at understanding; GPT-style autoregressive generation excelled at production. History chose the latter, for three reasons. Its training objective is perfectly unified — every scrap of text on the internet serves as curriculum, requiring zero human annotation. Its scaling efficiency is highest — a simple, brute-force objective converts compute into capability along the cleanest curve. And nearly every emergent ability — including chain-of-thought, the trick that taught AI to narrate its own reasoning — appeared exclusively on this branch.
Why it won: four factors multiplied
Transformer’s victory is often attributed to a great paper. The real answer is a multiplication formula: hardware resonance × infinite data × predictable scaling × ecosystem lock-in. Remove any factor and the product collapses to zero.
Hardware resonance. RNN-based approaches often struggled to fully exploit modern GPU parallelism, making large-scale training inefficient. Transformers, by contrast, aligned naturally with GPU architecture, pushing utilization to levels that made massive scaling economically viable. The essential framing remains: GPUs chose Transformers as much as Transformers chose GPUs. Nvidia’s rise and this architecture are mutually causal.
Infinite data. Autoregressive training needs no labels. Trillions of tokens on the open web serve as ready-made curriculum. When data acquisition cost approaches zero, “scale” stops being an option and becomes a strategy.
Predictable scaling. In 2020, OpenAI published the Scaling Laws paper: model capability follows a power-law relationship with compute, data, and parameters. For the first time, you could calculate how much intelligence a given investment would buy. This curve turned AI from alchemy into civil engineering. Capital doesn’t bet on miracles, but it will pour into a predictable engineering curve. Without this line of research, the economic case for scaling frontier models would have been far less compelling.
Ecosystem lock-in. Nine years of accumulated infrastructure: operator libraries, inference engines, chip-level optimizations, skill stacks held by millions of engineers. Technical merit is the entry ticket; ecosystem lock-in is the moat. Windows was never the best operating system. QWERTY was never the most efficient keyboard layout. Both won anyway. Transformers are tracing the same path.
That is the complete answer to “why it won.” But the best way to understand a reigning species is never through its strengths. It is through the flaws it cannot fix. And the four success factors above contain, precisely, the structural constraints that define its limits. The reasons it won are the sources of its limitations.
Four Flaws Baked Into the DNA
The four words previewed at the start now take their seats: sequence, memory, representation, reasoning. These are the Transformer’s four structural constraints — and the four cracks through which every challenger enters.
Flaw one: the sequence constraint — one word at a time, always
Autoregressive generation means producing the 1,000th token requires all 999 preceding tokens to be complete. This is a physical property of the architecture, not an engineering shortfall. No amount of GPU stacking eliminates it. Generation speed caps at single-token latency times token count. There is no shortcut.
The forty seconds you spend watching the cursor in Moment One are the direct expression of this constraint. Note the asymmetry: parallelization accelerates reading — attention processes all inputs in one pass — but speaking remains serial. An AI that can ingest a hundred pages in parallel still outputs like a telegraph operator. The peculiar waiting sensation in human-AI interaction originates here.
Flaw two: the memory constraint — the quadratic curse of reading more
The N×N bill comes due here. The metaphor is handshakes at a company retreat: 100 people shaking everyone else’s hand yields 4,950 pairs; 1,000 people yields nearly half a million. Increase text length tenfold, increase attention compute a hundredfold.
On the inference side, a hidden ledger called KV Cache compounds the cost. During generation, every produced token leaves a note in GPU memory for future tokens to consult. In long-context inference, the notes often consume more memory than the content itself. The dominant line item on many inference bills is memory bandwidth, not raw compute. The concrete figure: In certain long-context settings, standard attention can consume substantially more memory than linear-time alternatives.
“File too long” in Moment One, and the industry-wide pain point of prohibitively expensive long documents, both originate here.
Flaw three: the representation constraint — word co-occurrence is not causation
Transformers learn which words tend to appear together. They do not learn how the world works. If the training corpus contains enough sci-fi dialogue, the model will sincerely believe pigs can fly alongside airplanes.
The deeper explanation for hallucination lives here, and it is the most sobering insight in this article: incorrect answers and correct answers are the same operation statistically — both are “generate the most probable continuation given the context.” There is no internal switch between “I know” and “I’m making this up.” Hallucination therefore may not disappear simply through better training; it requires additional mechanisms such as verification, grounding, or external knowledge. It is a genetic condition, not a cold. The mug-and-paper-bag moment is its symptom.
Flaw four: the reasoning constraint — thinking equals talking
Chain-of-thought is the Transformer branch’s own brilliant adaptation: forcing the model to write out intermediate steps dramatically improved reasoning performance. Every product’s “deep thinking” mode today is the commercialized form of this patch.
But the patch carries a price tag: thinking consumes tokens. Deeper thinking means longer thinking means slower and more expensive. The six thousand invisible tokens in Moment One are this patch’s invoice.
Dig deeper and you find an assumption that has gone almost entirely unquestioned: why must AI’s thinking become speech? Humans rotating a 3D object mentally do not narrate the rotation. Mental arithmetic proceeds without language. But because Transformers can only speak, we taught them to think by speaking — and then normalized the equation.
Four flaws, one table
| Constraint | What it burns | Who attacks it | Weapon |
|---|---|---|---|
| Sequence (autoregressive generation) | Latency, interaction quality | Diffusion LLMs | Parallel generation |
| Memory (quadratic cost + KV cache) | VRAM, long-context expense | SSM/Mamba, hybrid architectures | Linear complexity |
| Representation (statistics ≠ causation) | Commonsense, physical understanding | World models / JEPA | Physical state prediction |
| Reasoning (thinking = talking) | Token cost, inference latency | Recurrent depth | Latent-space iteration |
Every challenger that follows exists as a mirror image of this table. None arrived from nowhere.
Evolution tells the same story. Homo erectus lacked speech. Neanderthals had larger brains than sapiens. Every species carries a genetically determined ceiling. The ceiling doesn’t prevent dominance; it determines the shape of that dominance — and who stands ready to fill the niche when conditions shift.
Extinct Branches and the Genes They Left Behind
Before meeting the challengers, visit the graveyard.
There are no true dead ends on the evolutionary tree, only dormancy and donated genes. Every human alive today outside sub-Saharan Africa carries roughly 2% Neanderthal DNA. AI’s genome carries four extinct branches, equally alive.
RNN/LSTM: the original telephone game
Before its dethronement, it was king. From 2014 to 2017, machine translation, speech recognition, and text generation all ran on this memory-equipped telephone game. LSTM variants partially mitigated information decay via dedicated long-term memory cells. Its fall came not from a smarter opponent but from hardware mismatch: serial computation starved GPUs. When Transformers reframed the same tasks as matrix operations, the training-speed gap became insurmountable. The 2017 paper’s title was itself a declaration: Attention Is All You Need — implying, unmistakably, that everything RNNs offered was unnecessary. Within three years, it abdicated.
But its core idea — recurrence — did not die. In 2026, it is reincarnating in a different form. Hold that thread; it returns in Part 4.
BERT: the copy editor who lost to the improviser
Fill-in-the-blank versus next-word prediction: the route war of 2018–2020. BERT acted as a copy editor, reading entire passages bidirectionally, excelling at classification, retrieval, and question answering. GPT acted as an improviser, continuing from prior context, naturally suited to generation. The improviser won, for reasons already enumerated: unified objective, superior scaling efficiency, emergent abilities concentrated on its branch.
The copy editor left two genes. The first reincarnated as embedding models: virtually all semantic search, recommendation systems, and RAG retrieval layers today descend from the BERT lineage. The second is its mask-and-denoise mechanism, which later became the core component of diffusion language models. That second gene hasn’t fully expressed yet. See Part 4.
Symbolic AI: the stubborn scholar who rebranded and won enterprise
This is the oldest branch, predating every other in this article — traceable to the 1956 Dartmouth Conference. The approach: humans write rules, machines execute logic. Expert systems, knowledge graphs, if-then rule bases. In the 1980s, scholars spending a year encoding ten thousand rules watched neural networks derive equivalent capabilities from data in a week. Its defeat came from expansion velocity, not conceptual error. Hand-written rules could never scale at the pace of automatic learning. The underlying philosophy was never wrong.
It returned under a new name: RAG, retrieval-augmented generation. Attach a verifiable knowledge base to a large model. Essentially, give the improviser a copy editor with a library. Nearly every serious enterprise AI deployment today includes this component. A branch that lost the theoretical war won engineering practice — perhaps the most successful resurrection in AI history.
Data quality: reading textbooks while everyone else scrolls short videos
The industry’s default faith holds that more data is always better. Microsoft’s Phi series ran the opposite experiment: train small models on carefully curated, synthetic “textbook-quality” data. On specific benchmarks, they matched models dozens of times larger. The metaphor writes itself: short-video scrolling versus textbook study.
It didn’t win the throne — on general capability, the brute-force combination of data volume and parameter count still prevails — but it became the throne’s foundation. Every frontier model’s training mix today incorporates this lineage’s lessons.
They all died of economics
Place the four branches side by side and the pattern snaps into focus.
None died because their ideas were wrong. All died of economics — couldn’t feed GPUs, couldn’t scale past data costs, couldn’t justify the expense. And none truly perished. Their genes were absorbed into the victor’s body.
Remember this pattern. Now meet the challengers. To judge any challenger’s fate, skip the elegance of its paper. Look at its economics.
Four Challengers at the Gate
Four challengers, attacking four cracks. Ordered by certainty, highest to lowest: hybrid architectures (already happening), MoE (now standard), diffusion LLMs (closing in), recurrent depth (just emerging).
Each receives the same four-part examination: which crack it attacks, how it works, current battle status, why it hasn’t won yet.
Mamba: a self-erasing whiteboard (attacking the memory constraint)
The mechanism fits in a stationery metaphor. Transformers spread the entire source record across a desk — lossless, but the desk fills up. Mamba and its state-space model (SSM) family use a fixed-size whiteboard: each new token intelligently revises the existing notes. Reading 100 tokens or 1 million tokens, the whiteboard stays the same size. The trade-off is lossy compression; the payoff is linear complexity and near-zero inference-time caching overhead.
Battle status has reached an inflection point. Born at CMU and Princeton, the architecture has iterated to Mamba-3. Recent Mamba-family models have demonstrated competitive or superior efficiency on selected benchmarks, especially for long-context tasks. More importantly, consensus has formed around the landing configuration: hybrid architectures. Most network layers run Mamba; a few critical layers retain attention. The corporate analogy: most processes follow SOPs; a handful of key decisions require face-to-face meetings.
The economic significance of this memory gap deserves explicit framing. It means “fit an entire codebase or full book into context” shifts from prohibitively expensive to routinely feasible.
Why it hasn’t won: hybrid architecture is gradual integration, not regime change. Transformers aren’t overthrown; they’re partially replaced. Pure Mamba still trails attention on precise retrieval tasks — whiteboards are lossy by design.
MoE: a hospital with 100 departments (attacking the cost constraint)
Dense models are general practitioners handling every case personally. Mixture-of-experts models are hospitals with 100 departments — each patient routes to only two or three relevant specialists. Total parameter count is enormous (broad knowledge); active parameters per forward pass are small (low cost).
MoE’s story turns on its comeback. Proposed early, ignored for years, then suddenly ubiquitous. The inflection point was a phase transition in the industry’s competitive axis. During the training race, the metric was who could train the largest model; dense architectures were simpler, and MoE looked ornamental. As the industry shifted to the inference race — who can serve hundreds of millions of users at lowest cost — MoE’s economic advantage flipped overnight from irrelevant to existential. The landmark case is DeepSeek-V3: trained at a fraction of peers’ cost while matching frontier performance, sending shockwaves through Western tech circles in early 2025. Since then, MoE appears on virtually every frontier model’s architecture sheet.
The intellectual payload of this subsection: the same technology can have radically different valuations across different industry phases. The technology didn’t change. The race did. Before evaluating any technique, ask: which race is the industry running right now?
Strictly speaking, MoE leaves the Transformer core intact — attention remains in place. It is cost reform within the throne room, which also makes it the fastest-deployed of the four challengers.
Diffusion LLMs: typewriter versus photograph development (attacking the sequence constraint)
Autoregressive generation resembles a typewriter, striking keys left to right. Diffusion generation resembles developing a photograph — the entire negative emerges simultaneously, sharpening from blur to clarity, all tokens materializing at once.
This lineage is distinguished. In image generation, diffusion models displaced GANs entirely within roughly three years to become the universal substrate — many leading text-to-image systems today are built on diffusion-based architectures. Its core mechanism, mask-and-denoise, is precisely BERT’s fill-in-the-blank resurrected. Like Denisovan gene fragments lying dormant in the human genome for tens of thousands of years until high-altitude hypoxia triggered expression, genes donated by extinct branches achieve secondary expression here — this time as the primary attacker.
Battle status is the most aggressive of the four. Some emerging diffusion language models have demonstrated dramatically faster generation speeds than traditional autoregressive approaches. The forty-second typing wait from Moment One becomes a four-second photo-development wait. Google’s open-source DiffusionGemma (Apache 2.0 license, high developer-community traction) has brought long-context diffusion to single consumer-grade GPUs.
Honest accounting on limitations: quality still carries a discount, trailing autoregressive models by several percentage points on certain benchmarks. Current engineering compromise pairs diffusion for rapid drafting with a small autoregressive model for proofreading — hiring a copy editor for the fast typist.
Observation signal for readers: the day a flagship-class pure diffusion model ships without quality compromise is the day interaction paradigms shift. Extrapolating from the three-year displacement curve in image generation, that day may arrive sooner than consensus expects.
Recurrent depth: thinking without speaking (attacking the reasoning constraint)
Thread pickup: the RNN’s recurrence gene has returned, but rotated ninety degrees. No longer cycling along the sequence axis (horizontal), it cycles along the depth axis (vertical): the same neural network layer applies repeatedly, iterating dozens of times within latent space, producing zero intermediate tokens.
Chain-of-thought resembles spreading scratch paper across a desk and working step by step. Recurrent depth resembles closing your eyes and reasoning internally. Mental multiplication of 27×34 proceeds through spatial operations, entirely without language. Experiments demonstrate small recurrent models iterating 32 times can match dense models orders of magnitude larger on specific reasoning tasks. Indeed, “deep thinking” products already partially acknowledge this direction — their reasoning traces are invisible to users, unfolding entirely in latent space.
This may be the deepest thread in the article: language’s limit is where thought begins. Chain-of-thought’s enormous success led the field to normalize thinking-as-speech. But this may be historical accident: because Transformers happen to speak, we taught them to think by speaking. If thought can decouple from language, the foundational assumption of the entire interpretability field — auditing AI cognition by reading tokens — requires reconstruction.
Why it hasn’t won: the answer is singular. This is the only challenge line where the primary bottleneck is human, not technical. Chain-of-thought became the default in AI safety research precisely because it provides the sole observation window — humans can audit what the model “thought” token by token. Recurrent depth closes that window. Internal iterations are opaque. Technically feasible, it is blocked by regulatory frameworks and trust infrastructure. Human society hasn’t prepared itself to accept an AI that thinks silently. Technology waits on humans.
The incumbent is a gene absorber
Placing the four challengers back onto the constraint table reveals an easily overlooked fact: Transformers themselves are absorbing every challenger’s genes. Subsequent versions already incorporate sparse activation (from MoE), grouped attention (inspired by linear-architecture thinking), speculative decoding (borrowing the draft-and-proofread division of labor). Incumbents are never stationary targets. They are voracious gene absorbers.
Every battle described above unfolds within the territory of language. But the swimming expert from the opening reminds us: the true next front may lie beyond language altogether.
When AI Needs to Enter the Water
From metaphor to proposition
“Read ten thousand books about swimming but never entered the water” began as metaphor. It now upgrades to technical proposition: current LLM intelligence is linguistic intelligence. What it models is how humans talk about the world, not the world itself.
A thought experiment: delete every instance of “cups break when dropped” from the internet, and the LLM no longer “knows” this fact. A two-year-old child knows it without any textual exposure. The source of knowledge determines the nature of knowledge.
Predicting the next state, not the next token
The Transformer’s objective function predicts the next token. World models predict the next state: given a current scene and an action, what happens next? The prediction target shifts from words to physics.
The representative approach is JEPA, championed by Yann LeCun. Its key design choice: prediction occurs not at pixel level but in abstract representational space. It ignores reflections on the cup, angles of shadow; it cares only about causal elements like “it will break.” The human analog: watching someone carry a full glass, your brain doesn’t render individual finger pixels, but you possess an intuitive model predicting spillage. LeCun’s widely cited claim — current LLMs are less intelligent than a cat — carries precise technical meaning. Linguistic intelligence and physical intelligence are distinct currencies. High scores in one do not automatically convert to the other.
2026: theory, tools, and products converging
Theory line: multiple studies demonstrate systematic linear correspondence between large-model internal representations and physical variables. Models acquire rudimentary physical intuition as a byproduct of language learning. Good news for the linguistic camp; also a hint that the two routes may converge in middle ground.
Tools line: Google’s Genie family generates interactive virtual worlds directly from text and images. For the first time, AI has a pool to enter — it can attempt actions, make mistakes, observe consequences inside generated environments rather than merely reading about consequences in text.
Products line: Fei-Fei Li’s World Labs launched Marble, a commercially available 3D world-generation product. “Spatial intelligence” moved from paper titles to keynote themes. Industry chatter about “world model year” has begun circulating.
Déjà vu economics
Per protocol: enumerate the obstacles, then interpret through Part 3’s pattern.
Three obstacles. Data economics: physical interaction data costs orders of magnitude more than text. Internet text is free; robots breaking ten thousand cups is not. Evaluation benchmark deficit: language models have mature benchmarks; world models lack even industry consensus on what “good” means. Hardware-dependent commercial loops: world models’ largest buyers — humanoid robotics and autonomous driving — operate on hardware adoption cycles measured in decades.
Now the point: these obstacles are identical to those that killed RNNs against Transformers. Ideas didn’t lose. Economics lost. World models are conceptually unassailable; even the most committed linguistic-school researchers concede that commonsense cannot be fully solved within language alone. What it awaits is its own GPU moment, its own internet-text moment.
One observation signal for readers: watch for the day linguistic and physical intelligence fuse within a single model — not an LLM with a bolted-on physics engine, but a single forward pass that understands both text and gravity. When that day arrives, Moment Two’s problem will have found its real answer.
Homo sapiens left Africa not through larger brains but through fire and tools — extending species capability beyond biology via external organs. World models are AI’s fire.
A Ruler You Can Keep
Three signals
The article reaches its final station. The most valuable delivery is not conclusions but tools. To evaluate whether any AI technical branch can break through, ignore citation counts and keynote volume. Check three objective signals:
One: does the pain point it solves burn money every day? Real pain means someone pays continuously to alleviate it. Pseudo-pain’s telltale sign is that only researchers care.
Two: does its data/hardware economics work? Part 3’s entire historical induction compresses to one sentence: every branch died of economics, none of ideas.
Three: has it developed its own Scaling Law? Predictable progress curves enable capital commitment and ecosystem formation. Technologies without power laws cannot assemble the resources required for regime change.
Reverse validation: this ruler was forged from history
Apply the three signals to Transformer’s ascent: GPU resonance plus free internet text (signal two), the 2020 power-law paper (signal three), inference-cost pain exploding with user scale (signal one). All three lit. Four-factor product detonated.
The ruler wasn’t reverse-engineered post hoc. It was induced from history, then deployed to forecast the future. This is the “new thinking skill” this article aims to deliver: next time an AI architecture headline appears, skip “is it powerful” and ask these three questions first.
Forward projection: four trajectories ranked by certainty
Hybrid architectures: highest certainty, already underway. Pain burns daily ✓, economics hold ✓, progress curve predictable ✓. Observation signal: Mamba/SSM layer share rising year-over-year in frontier model technical reports. Progress proceeds via gradual integration, like sapiens absorbing Neanderthal genes. Completion will arrive without ceremony.
MoE: settled, no prediction needed. Remaining observation point: how far sparsity can compress.
Diffusion LLMs: 2–4 year horizon, medium-high certainty. All three signals lit — latency pain burns daily, GPU-native friendliness, quality-speed curve already predictable. Observation signal: first flagship-class pure diffusion model shipping without quality discount.
Recurrent depth: 3–5 year horizon, medium certainty, one special variable. Technology and economics both hold. The sole uncertain signal is human acceptance — how long regulatory frameworks and safety-research paradigms need to digest an AI that thinks silently. Observation signal: first latent-space reasoning flagship passing mainstream safety audit.
World models: 5–10 year horizon, highest long-term weight. Signal one (embodied AI’s burn rate) brightening. Signal two (physical data cost) unresolved. Signal three (own power-law curve) not yet emerged. Observation signal: native fusion model integrating linguistic and physical intelligence.
Convergence endpoint: a fused species
One final trend, and the evolutionary metaphor’s closure: these branches are not competing routes. They are fusing into a new species. Hybrid architecture as skeleton. Diffusion head for rapid response. Recurrent depth for deliberation. RAG for verification. World models for commonsense.
Neanderthal and Denisovan fragments persist in the sapiens genome to this day. The eventual dominant AI architecture will likely be a similar genetic fusion — and attention will be the longest-surviving gene within it.
Next Time You Watch the Cursor Blink
Return to the two opening scenes.
Next time you stare at the cursor, waiting for words to emerge one by one, you won’t see “AI is slow.” You’ll see a species expressing its genes: the sequence constraint, and the diffusion challenger closing in. Perhaps within a few years, what you wait for will shift from typing to developing.
Next time it confidently misstates a piece of physical commonsense, you won’t see “AI is dumb.” You’ll see the representation constraint, and the long economics of the “enter the water” route — awaiting its GPU moment, just as RNNs once awaited theirs, just as Transformers once awaited theirs.
What this article ultimately aims to leave behind is not predictions about which branch wins. No one predicts accurately. What it leaves is a perspective shift: AI is not a finished product. It is an ongoing evolutionary process. Those who understand the process and those who treat AI purely as a tool will accumulate fundamentally different judgment over the coming decade. When the next architecture headline breaks, others will see spectacle. You will see four constraints shifting.
The evolutionary tree has never contained a “best species.” Only the species best suited to this moment. Right now, Transformers are best suited. And “right now” is passing.
Read More of Intelligenr
-
The Cognitive Trap Zone: Redesigning AI Tools Around Human Attention
-
Copilot vs. Cognitive Exoskeleton: How AI Is Reshaping Human Memory & Thinking
References
-
Vaswani, A., Shazeer, N., Parmar, N., et al. (2017).
Attention Is All You Need.
Advances in Neural Information Processing Systems (NeurIPS).
https://arxiv.org/abs/1706.03762 -
Kaplan, J., McCandlish, S., Henighan, T., et al. (2020).
Scaling Laws for Neural Language Models.
https://arxiv.org/abs/2001.08361 -
Gu, A., & Dao, T. (2023).
Mamba: Linear-Time Sequence Modeling with Selective State Spaces.
https://arxiv.org/abs/2312.00752 -
LeCun, Y. (2022).
A Path Towards Autonomous Machine Intelligence.
OpenReview.
https://openreview.net/forum?id=BZ5a1r-kVsf
Author Note: This article reflects an independent analysis of AI architectures, combining research literature with broader observations about how technology evolves through constraints, economics, and engineering trade-offs.