Editorial Note: AI performance is rarely just about compute. Memory capacity, bandwidth, context length, and power constraints often determine how quickly a model can actually respond. This article looks beneath the interface to explain the hardware mechanics behind AI latency—and why understanding the bottleneck matters more than chasing a single performance number.
Why AI Gets Slower the More You Use It: Memory, Bandwidth, and the Physics Behind the Lag
Two moments almost every AI user has lived through—and the hardware reality that explains both.
Two Moments You've Probably Lived Through
Moment A. You open ChatGPT or Claude, start a fresh conversation, and ask something complex. The answer streams out faster than you can read. You keep going, follow-up after follow-up. Forty minutes later, the thread has tens of thousands of words of context—and now you ask something simple. You watch the screen as the model dribbles out one word at a time, like toothpaste from a nearly empty tube. You refresh. You re-ask. Still slow.
Moment B. You install LM Studio or Ollama, pull down a 7B open-source model, and it runs surprisingly well—dozens of tokens per second, fans barely spinning. Emboldened, you download a 70B model. You hit load. The progress bar completes, memory usage slams into the red line, and the app either crashes or crawls along at one or two words per second. Forty seconds pass before it finishes a single sentence.
These two moments look unrelated. One is a cloud service stalling mid-conversation; the other is a local model that never had a chance. But set them side by side and the same physics is doing the work in both—an accounting problem involving memory, bandwidth, compute, and power.
Once you can read that account, a lot changes. AI's behavior stops feeling arbitrary. You start seeing which of your actions load the system, where the boundary between cloud and local capability actually sits, and which numbers on a spec sheet matter for AI workloads—versus which ones are marketing. That understanding is becoming a baseline skill for working with AI, not a specialist's hobby.
So let's start at the bottom: how the model actually "speaks."
1. AI Isn't Writing—It's Playing Word Chain
Why LLMs generate text one token at a time
The intuition many people carry is that a large model "thinks through a paragraph, then outputs it." The real process is closer to word chain: the model predicts the next token (roughly, a word or word fragment), appends it to the input, and predicts again. Token #100 cannot begin until tokens 1–99 exist. That dependency chain has no shortcut, and it is the root of everything you experience as "slow."
But there's a detail that gets missed: generation happens in two phases with completely different characters.
Reading is fast. Writing is slow.
Phase one is called prefill. The model ingests your entire prompt at once. Because the prompt is known and complete, all tokens can be processed in parallel—thousands of GPU cores working simultaneously. This phase is often compute-heavy and highly parallel, although the exact bottleneck depends on model, prompt length, hardware, and implementation. This is why the interface feels "ready" almost the instant you hit enter.
Phase two is called decode. Now the model generates new tokens, one at a time. For a dense model running a single-stream decode, each generated token typically requires streaming most or all of the model's weights through the memory hierarchy again, in addition to accessing the KV cache. This cannot be fully parallelized—you don't know token 20 until token 19 has been generated. This phase is often memory-bandwidth-bound.
That difference produces a result most people don't expect: what determines AI generation speed is often not how much compute a machine has, but how much memory bandwidth.
Run the numbers. A 70B model at 4-bit quantization has roughly 35GB of raw weight data before metadata and runtime overhead. As a back-of-the-envelope illustration, 40GB × 30 tokens/sec = 1.2TB/sec of weight traffic. That is not a measured bandwidth requirement—it is an approximation of the memory-traffic pressure created by repeatedly streaming the weights during decode. Only the very top of consumer GPUs—around 1–1.8 TB/s of bandwidth—can approach that theoretical weight-traffic rate. In a memory-bound decode regime, much of the accelerator's theoretical compute capacity can remain underutilized.
Longer context makes every step more expensive
The second mechanism is attention. For each new token, the model compares its query against the cached keys and values from previous tokens. With 1,000 tokens of history, that means roughly 1,000 positions to consider; at 32,000 tokens, that per-token attention work is roughly thirty-two times larger. Full attention has quadratic cost during prefill, but with a KV cache, the incremental attention work during decode grows roughly linearly with context length.
To avoid recomputing all that history at every step, engineering introduces the KV cache: previously computed context is stored and reused. The catch is that the cache itself occupies memory, and it grows linearly with conversation length.
Which explains a counterintuitive phenomenon: in a long conversation, the thing eating memory may not be the model—it's the chat history. A 7B model quantized down to 4–5GB can carry several more GB of KV cache at 32K context. Stretch the context window further, and the bill grows accordingly.
Now back to Moment A. Early in the session, context is a few hundred tokens, the KV cache is negligible, every step is cheap. Tens of thousands of words later, each generated token has a larger history to work with—and the provider is also dealing with a more expensive request. How that request is scheduled depends on the provider's infrastructure and policies, but long-context sessions can consume substantially more memory and compute than short ones. The two effects can compound into the toothpaste effect on your screen.
The "stutter" you see is the backend's true speed
One more thing routinely gets folded into "slow": streaming output.
Early AI apps waited for the full response, leaving you staring at a spinner for fifteen seconds. Modern apps push each token the moment it's generated, so you see progress immediately. That's a genuine experience improvement—with a side effect: any fluctuation in backend generation speed now maps onto your screen with zero buffering.
Cloud interfaces with optimization layers absorb network jitter and queueing delays, so it just feels "a bit slow." Local clients typically stream raw, and every millisecond your GPU gets squeezed shows up as a visible hitch. This is why the same model can feel completely different across two clients.
2. Cloud and Local: Same Symptom, Different Disease
Local LLM vs cloud AI—and why your local model is so slow
Both moments are "slow." But lay the two situations side by side and they're bottlenecked in different places, with different levers available.
When you use a web app or API, the model runs in the provider's datacenter—machines with tens of gigabytes of high-bandwidth memory, purpose-built inference engines, cluster-level load balancing. In principle, the speed ceiling there dwarfs anything a personal device can reach. So why does it still feel slow?
Several factors stack up. Peak-hour requests get queued into batches, with unpredictable waits. The KV cache problem from Section 1 exists in the cloud too—and providers have an economic incentive to deprioritize extremely long sessions, since those requests cost far more resources than average. Network distance and connection setup can add latency before generation begins, while network jitter can affect how smoothly streamed output reaches the client. And one factor that rarely gets named: some of the slowness can come from the client itself. Long conversations can increase rendering, layout, and JavaScript work in the browser, and that client-side jank can be perceptually indistinguishable from model latency.
So in the cloud scenario, the variable you can actually manage is context discipline: start new conversations at sensible points, keep prompts tight, feed long documents as summaries rather than full text. These moves work because they directly reduce the provider-side computation and memory footprint.
Move the model onto your own machine and the picture changes completely. All pressure lands on one box—no load balancing, no provider-side inference optimization, no buffer layer.
Now the hardest constraint is memory capacity. Model weights must live in GPU memory (VRAM) for normal-speed inference—for the reason from the accounting above: every token requires a full pass over the weights. If those weights aren't in VRAM, they must be hauled over the PCIe bus from system RAM. And the bandwidth gap between those two paths is one to two orders of magnitude: a high-end consumer GPU offers 1,000–1,800 GB/s of memory bandwidth; PCIe 5.0 x16 offers about 64 GB/s; dual-channel DDR5 system RAM, roughly 80–100 GB/s. The moment a model "spills" into system RAM (the infamous offloading), generation speed collapses from 30–40 tokens/sec to 1–3 tokens/sec—a drop of over 90%. The struggling 70B model in Moment B is almost certainly stuck exactly here.
How do you budget VRAM? A useful first-pass estimate is: VRAM needed ≈ weight memory + runtime overhead + KV cache. Weight memory is roughly parameters × effective bytes per parameter, but the real footprint varies with quantization format, model architecture, runtime, context length, and batch size. FP16 uses 2 bytes per parameter in the raw representation, while 8-bit and 4-bit quantization use roughly 1 and 0.5 bytes per parameter before accounting for metadata and implementation overhead. As a rough guide, a 7B model at 4-bit may fit in around 4–5GB of weight memory, but the total VRAM requirement depends on context and runtime overhead. A 70B model at 4-bit still needs roughly 35–40GB+ once the practical memory footprint is included—more than the VRAM of mainstream single-GPU consumer cards. No software trick removes that ceiling.
The recurring "4-bit" refers to quantization—representing model weights with roughly four bits per weight rather than 16-bit floating-point values. Different 4-bit schemes use different representations and metadata, so the effective memory footprint is usually somewhat above the theoretical four-bit minimum. This is the enabling technology for local AI; without it, a 7B model would need 14GB and a 70B model would need 140GB. In practice, well-designed 4-bit quantization can preserve much of a model's capability while substantially reducing memory use, although the quality trade-off varies by model, quantization method, and workload. Push down to 3-bit or 2-bit and the quality trade-offs can become more pronounced.
One easy-to-miss wrinkle: context length eats back what quantization saved. Compress the model to 4-bit, free up several GB—then open a 64K context window and the KV cache takes those GB right back. The VRAM budget is shared between model weights and conversation history. Both sides of the ledger count.
One architecture deserves its own paragraph, because it changes the accounting fundamentally: Apple Silicon's unified memory. On a traditional PC, the CPU uses system RAM and the GPU uses its own VRAM, with the PCIe bus between them—data must be copied across. Under unified memory, CPU, GPU, and other accelerators share a common physical memory pool, avoiding the separate CPU-RAM/VRAM capacity boundary and the PCIe transfer path between them. For the first time, a personal device can locally run 70B-class or larger models—something no single consumer GPU can do. The trade-off: unified memory bandwidth typically sits in the hundreds of GB/s, clearly below a high-end discrete GPU's 1TB+, so on small models the discrete card generates faster. There's also the ecosystem question—the mainstream AI toolchain grew up around CUDA. Other ecosystems have improved fast, but you'll still hit cases needing manual adaptation or tools that don't yet support the platform.
There's no "which is better" answer here. Need to run big models? Capacity wins. Need raw speed? Bandwidth and ecosystem win. The two camps serve barely overlapping audiences—knowing which side you're on matters more than any individual spec.
Laid out together, the two routes look like this:
| Dimension | Cloud API | Local inference |
|---|---|---|
| Capability ceiling | Frontier models, no hardware barrier | Hard-limited by local VRAM |
| Privacy & offline | Data leaves the machine; network required | Data never leaves the machine; works offline |
| Speed predictability | Depends on provider load & scheduling | Depends on your hardware; stable and predictable |
| Cost structure | Per-token billing or subscription | One-time hardware + electricity |
| Primary bottleneck | Network, context length, provider scheduling | VRAM capacity and bandwidth |
3. How AI Is Redefining Hardware: VRAM and Bandwidth Became the Hard Currency
Why AI tools need GPUs—and how much VRAM you actually need
A CPU has a few very complex general-purpose cores—good at branching, scheduling, serial logic. A GPU has thousands of simple cores—good at doing enormous amounts of the same arithmetic simultaneously. The bulk of LLM inference is matrix multiplication: one operation repeated billions of times. That sits precisely at the center of what a GPU does well. An NPU is the same idea pushed further: it sacrifices generality for matrix-math efficiency at lower power, which is why it lives in phones and thin laptops. This also explains why GPU inference is usually much faster than CPU inference for large LLMs, although the gap varies dramatically with model size, quantization, memory bandwidth, and implementation. CPU inference remains perfectly viable for smaller models and memory-rich systems.
But revisit the bandwidth account from earlier and something counterintuitive surfaces: for text generation, the speed limiter is usually memory bandwidth, not compute.
During decode, every token requires a full read of the weights—that's pure memory traffic. Stronger compute doesn't help if the data can't be fed. Plug different GPUs' bandwidth against model size and you'll see it: two cards with similar peak FLOPS will generate text at different speeds, and the one with higher bandwidth wins. Spec sheets and reviews put FLOPS in the biggest font while bandwidth hides in a small print row—but for this workload, the priority is reversed.
That said, this ordering holds only for text generation. Switch workloads and the picture changes entirely:
| AI workload | Compute character | Primary bottleneck |
|---|---|---|
| Text generation (LLM inference) | Token-by-token, serial; repeated weight reads | Memory bandwidth + VRAM capacity |
| Image / video / 3D generation | Multi-step iteration or sustained parallel compute | Peak compute + VRAM + sustained thermals |
| Fine-tuning / training | Weights + gradients + optimizer states in memory | Much higher memory requirements than inference; exact requirements depend heavily on optimizer, precision, batch size, and training method |
Image generation (diffusion) runs full-image parallel computation each denoising step, so peak FLOPS matter again. Video and 3D generation sustain heavy parallel loads for long stretches, stressing both compute and thermal endurance. Fine-tuning and training keep weights, gradients, and optimizer states resident simultaneously—roughly 3–4× the VRAM of inference. The same GPU might be bandwidth-bound with idle compute during text generation, then compute-saturated with just-enough memory during image generation. "AI performance" is never a single number. It depends entirely on what you're running.
On-device AI in phones is a different constraint set altogether. A phone's sustained power budget is roughly 3–6W; a high-performance laptop GPU sustains 60–115W; a desktop flagship runs 300–450W. Generation speed is strongly constrained by that power budget—thermodynamics is honest at this layer. So on-device models are generally much smaller than frontier cloud models, often using aggressive quantization and architecture-specific optimizations. Model sizes vary widely depending on the device and workload, so parameter count alone is a poor proxy for the amount of AI a phone can run. Phones also use unified memory, which removes the separate CPU-RAM/VRAM capacity boundary—but the model shares that pool with the OS and background apps, and the system itself already consumes several GB. For phones intended to run larger general-purpose models locally, 12–16GB of system memory provides considerably more headroom, but there is no universal RAM floor for "on-device AI." Actual requirements depend on model size, quantization, accelerator design, OS memory pressure, and which AI features are being run.
And one variable almost never printed on spec sheets: sustained thermals. Phones and thin laptops throttle under continuous load. That "20 tokens/sec" figure is often the first 30 seconds on a cold device; ten minutes later it might be 8. When evaluating a mobile device's AI capability, sustained performance tells you far more than peak performance.
4. From Mechanisms to Trade-offs: How These Constraints Shape Your Choices
Do you need a powerful GPU for AI? It depends which of these you're doing.
The last three sections established three constraints: context-length cost inflation, VRAM capacity and bandwidth as the local-inference gatekeeper, and power/thermals as the mobile endurance limit. Each carries a different weight depending on how you actually use AI—which means "what hardware do I need for AI" is really several different questions.
If you live in cloud APIs and web apps, constraint #1 is the only variable you actively manage. Network quality, client rendering efficiency, and your own context hygiene drive the experience. The GPU is nearly irrelevant here; whether the CPU and RAM can hold up under dozens of tabs and long documents, screen quality, and portability matter more.
If you run models locally, constraint #2 is a hard gate. VRAM capacity decides whether a model runs; bandwidth decides how it feels. The priority order becomes VRAM > bandwidth > system RAM > CPU—with unified memory as a special case on the capacity side.
If your work is image generation, video, or 3D assets, the bottleneck shifts again. Peak compute and software ecosystem compatibility suddenly outweigh what mattered for pure text inference. Ecosystem fit often beats paper specs in daily experience—a slightly weaker platform with mature tooling can be far easier to live with than a stronger one you're constantly hand-configuring.
If on-device phone AI is your world, constraint #3 takes the lead: SoC generation (which determines NPU performance), memory capacity, and thermal design. 12–16GB of RAM, plus how well the vendor has integrated AI at the system level and how consistently they update it, predicts daily experience better than any single peak metric.
Side by side:
| Usage pattern | Real bottleneck | What hardware attention goes to |
|---|---|---|
| Cloud-API-first | Network, client rendering, context management | CPU/RAM, screen, portability; GPU barely matters |
| Local-inference | VRAM capacity, then bandwidth | VRAM > bandwidth > RAM > CPU |
| AI creative work | Compute + ecosystem compatibility | Ecosystem fit first, then compute and VRAM |
| Mobile-first | SoC/NPU, memory, thermals | RAM ≥ 12–16GB; depth of system-level AI integration |
Notice that the gaps between these categories are much wider than the gaps within any of them. The same budget spent for a cloud-first user versus a heavy local-inference user produces diametrically opposite configurations.
One more trend worth watching is the growing use of automatic local/cloud task routing. Light work—live captions, notification summaries, quick rewrites—runs on the local NPU at zero latency, never leaving the device. Heavy work—complex reasoning, long documents, image generation—shifts seamlessly to the cloud. When that dispatch is transparent enough, the "local camp vs. cloud camp" choice dissolves. It's worth checking whether a device has this hybrid capability, and whether switching requires manual intervention. The maturity of that dispatch logic is becoming as important as the hardware specs themselves.
One last plain observation: AI hardware's performance-per-dollar improves every year, and quantization advances keep expanding what fits in a given amount of memory. Which means "wait for the next generation" and "buy now" both tend to look reasonable in hindsight. In practice, what keeps shaping the experience is understanding which category of workload you actually run—and that doesn't expire when the hardware does.
5. Constraints Will Shift. The Physics Won't Disappear.
Where local AI is headed—and whether it will get faster
Back to the two opening moments. Within a few years, they'll likely be rarer.
Several technical paths are already underway. Speculative decoding uses a smaller model or another prediction mechanism to draft multiple tokens while the larger model verifies them in batches, potentially increasing generation speed without changing the target model's output distribution when implemented correctly. Linear-attention and sliding-window variants are softening the quadratic context penalty. Finer quantization schemes keep fitting more capability into the same memory. Distillation keeps pushing frontier-level ability down into smaller parameter counts. On the hardware side, memory bandwidth per dollar keeps climbing.
All of it is real. But it relieves today's constraints—it doesn't abolish constraints.
Some things won't change: the autoregressive dependency chain means token N still waits for token N−1. Weights still have to move through a memory hierarchy, and that movement costs bandwidth, time, and energy. Cooling capacity still sets an important ceiling on sustained performance.
And there's a treadmill effect: every efficiency gain gets partly consumed by bigger models, longer contexts, and higher quality expectations. Today 16GB comfortably runs a 7B model; when efficiency doubles, the expectation becomes running 70B in 16GB—and you hit a new wall, freshly built.
That's the point of understanding the mechanisms. Specific product names and specs go stale, but "generation is token-by-token," "bandwidth sets decode speed," "VRAM capacity is the hard gate," and "power budget determines endurance" stay valid for a long time. With those in hand, any new tool or device can be assessed in minutes: where is its bottleneck, and is my usage pattern loading it?
In that sense, the depth of your understanding of the machine's constraints shapes the ceiling of your collaboration with it. It lands in very concrete behaviors: knowing long sessions grow heavier, you segment conversations and start fresh on purpose; knowing local VRAM is a hard gate, you stop expecting a thin laptop to run big models and route heavy work to the cloud; knowing on-device models have a capability ceiling, you judge them by the right standard—and decide deliberately which tasks live in your pocket versus on your desk.
Those judgments outlast any "best configuration" recommendation.
One final question to sit with.
Devices are beginning to decide automatically what runs locally and what goes to the cloud. "My computer" and "some card in some datacenter" relay seamlessly within a single conversation. So where does the device end?
We're used to thinking of a tool as a physical object: a laptop, a phone, a GPU. But computation in the AI era is becoming a fluid resource, assembled on demand. The box in your hands looks less and less like a machine, and more and more like a doorway.
That shift has consequences for what "I own this," "my data is here," and "this is my compute" even mean. Worth thinking through before the next device decision—not after.
Read More of Intelligenr
-
The Evolutionary Tree of AI: Why Transformers Dominated & What's Next
-
The Cognitive Trap Zone: Redesigning AI Tools Around Human Attention
References
-
Hugging Face. Cache strategies.
-
Leviathan, Y., Kalman, M., & Matias, Y. Fast Inference from Transformers via Speculative Decoding. Proceedings of Machine Learning Research, 2023.
Author Note: AI can feel like software, but underneath the interface it is still a physical system. Understanding where the time goes—from memory and bandwidth to power and thermals—makes the behavior of AI a little less mysterious.