The alert came through at 2 a.m. An ML engineer running a RAG-based Q&A service watched GPU memory warnings stack up in the ops channel. Their system prepended a five-thousand-word knowledge base summary to every request as a system prompt. The first user question went through fine. By the third or fourth turn in the same conversation, the inference service had slowed to a crawl.
They were running vLLM, one of the most battle-tested serving engines in the industry. So why was it choking?
The answer lives at a specific boundary in how vLLM handles KV Cache. Understanding that boundary is what separates a frustrated on-call rotation from a system that actually holds up under load.
—
Where the Bottleneck Actually Hides
The Transformer attention mechanism has a simple requirement: when generating any new token, it needs to access the Key and Value vectors of every token that came before it. Without caching, that means recomputing the same history on every step, an O(n²) operation that gets worse with every token added.
KV Cache is the standard fix. Store those Key-Value pairs once, read them back on subsequent steps. Within a single generation pass, this drops decode-phase computation from quadratic to something close to linear. The gains are real and well-established.
The catch is memory. A 70B parameter model processing an 8K-token context can generate a KV Cache exceeding 10 GB. Stack a dozen concurrent requests on the same GPU, and you hit the ceiling fast.
vLLM’s answer to this was PagedAttention, and it worked extremely well within the problem it was designed to solve.
—
PagedAttention: Managing GPU Memory Like an OS
The conceptual debt PagedAttention draws on is virtual memory paging from operating systems. Traditional KV Cache allocation required contiguous GPU memory blocks reserved upfront per request. Sequence lengths are hard to predict in advance, so you either over-allocate and waste memory, or under-allocate and pay for frequent resizing.
PagedAttention breaks KV data into fixed-size blocks (“pages”) that can be non-contiguous in physical memory. A block manager handles allocation and deallocation centrally. Requests sharing a common prefix, say, the same system prompt, can share the same physical blocks rather than each holding a private copy. Memory fragmentation drops significantly. Concurrent throughput goes up.
This is a real engineering achievement. For most general-purpose LLM serving, it makes the memory problem tractable.
But there is a hard edge to what it covers. PagedAttention manages KV Cache *within a single request’s lifetime*. When a request finishes, its blocks get marked for reclamation. The next request, even if it carries the exact same system prompt and the exact same document context, runs through the full prefill phase again from scratch.
That is not a flaw in vLLM. It reflects the scope of what vLLM was designed to do: manage memory efficiently inside a single inference pass. Cross-request KV reuse is a different problem entirely, and vLLM does not claim to solve it.
For stateless API calls, that is fine. The real trouble starts when your workload looks like long multi-turn conversations, repeated large system prompts, or RAG pipelines that repeatedly encode the same document corpus across many requests. In those cases, you keep paying the prefill cost on every single request, and it adds up.
—
What LMCache Adds to the Picture
LMCache (GitHub: LMCache/LMCache, roughly 12,000 stars at the time of writing) positions itself as a layer above vLLM, not a replacement for it. The framing matters: it does not compete with vLLM’s core inference engine; it extends the caching scope beyond individual request boundaries.
The core idea is straightforward. After vLLM computes KV data for a given input sequence, LMCache serializes that KV data and writes it to a configurable storage backend. On the next request that shares a matching prefix, LMCache intercepts the prefill phase, restores the cached KV data directly into GPU memory, and hands the remaining uncached portion off to vLLM for normal processing. The overlap between what was cached and what the new request needs drives the speedup.
The storage options form a tiered hierarchy:
CPU memory sits closest to the GPU and offers the lowest restore latency. Space is still limited compared to disk, but for frequently accessed prompts it is the right first stop.
Local NVMe SSD trades some latency for much larger capacity. Modern NVMe drives are fast enough to make this practical for production workloads, particularly when cache hit rates are high.
Remote storage (Redis or distributed object stores) enables sharing cached KV data across multiple inference instances. For a fleet of nodes all serving the same document corpus, this avoids each node independently recomputing the same prefill work.
LMCache’s interface is designed to be transparent to the application layer. It presents the same API as vLLM, so swapping it in does not require changes upstream.
—
The Relationship Between the Two
A common confusion worth clearing up: people sometimes frame this as a choice between LMCache and vLLM. It is not. LMCache requires vLLM (or a compatible engine) underneath it. There is no version of using LMCache that does not involve vLLM doing the actual inference.
The analogy that fits best is a query cache in front of a database engine. The database does the real computation. The cache intercepts queries it can answer from stored results, reducing how often the full computation runs. You do not replace the database to install a cache; you add the cache in front of it.
The table below maps the main differences:
| Aspect | vLLM (PagedAttention) | LMCache |
|---|---|---|
| Scope | KV memory within a single request | KV reuse across multiple requests |
| Cache lifetime | Ends when the request finishes | Persists across requests and restarts |
| Storage location | GPU memory only | CPU memory / NVMe SSD / remote store |
| Dependency | Standalone inference engine | Requires vLLM or compatible engine |
| Best fit | General-purpose LLM serving | Long prompts, multi-turn chat, RAG |
| Deployment overhead | Low | Medium (cache backend configuration needed) |
| Extra memory cost | None | Yes, CPU memory or disk for cached KV data |
Neither row makes the other irrelevant. They solve different things at different scopes.
—
Where the Speedup Shows Up in Practice
Multi-Turn Conversations
As a conversation stretches to ten or fifteen turns, the accumulated history can easily reach several thousand tokens. Every new user message has to carry that full history into the prefill phase.
With vLLM alone, each request is stateless. The conversation history’s KV data does not survive between calls. LMCache holds that history in its cache. When the next turn arrives, only the incremental new tokens need fresh computation. The longer the conversation runs, the larger the fraction of total prefill work that gets skipped, and the more noticeable the drop in time-to-first-token (TTFT).
Fixed or Near-Fixed System Prompts
The RAG engineer’s situation from the opening is the clearest version of this problem. If your system prompt is a large, fixed document, a knowledge base or a reference corpus, and that same prompt goes out with every single request, you are paying full prefill cost on every request for content that never changes.
LMCache caches that prompt’s KV data once. Subsequent requests share the cache hit. TTFT for requests with long, stable system prompts can fall sharply, not because the model is running faster, but because most of the work simply does not run at all.
Multi-Instance RAG Services
When you are running multiple inference nodes and all of them repeatedly encode the same document set, the redundancy adds up fast. Remote shared cache means a document encoded once by any node can be reused by all nodes. GPU cycles that used to be spent on re-encoding get redirected to actual generation work.
This also helps when GPU memory is tight. Offloading less-recently-accessed KV data to CPU memory or SSD effectively expands the working cache pool without adding more GPUs.
—
The Engineering Costs
No acceleration tool comes without trade-offs, and LMCache is no exception.
Storage footprint. KV Cache data is large. Serializing it to CPU memory or disk means planning for that size from the start. Long-running services need a cache eviction policy, otherwise the store grows without bound.
Serialization overhead. Moving KV tensors from GPU to CPU and writing them out takes time. For workloads with low cache hit rates, this round-trip can hurt more than it helps. The math only works in LMCache’s favor when the same prefix shows up repeatedly.
Cache invalidation on model updates. If you deploy a fine-tuned model update, cached KV data from the previous checkpoint is no longer valid. You need to track which cache entries belong to which model version and flush accordingly. This is manageable but adds operational complexity.
External dependencies. Adding Redis or a distributed object store introduces another component that can fail. The inference pipeline now has a cache backend in its critical path, which means more failure modes to monitor.
For single-node deployments with moderate concurrency and low prompt similarity, these costs may outweigh the gains. vLLM’s native prefix caching covers a lot of ground without any of this overhead.
—
When vLLM’s Built-In Caching Is Enough
vLLM ships with prefix caching that handles a fair amount of what LMCache targets. If your system prompt is a few hundred tokens or less, if concurrent load fits comfortably in GPU memory, if you are running a single inference node, and if most requests are one-shot queries rather than multi-turn sessions, then vLLM’s native mechanisms are probably sufficient. Adding LMCache in that configuration introduces complexity without a proportional return.
The signal to look for LMCache is a specific kind of repeated work: large prompts appearing across many requests, long conversations where history accumulates, or multi-node setups where the same documents get encoded over and over. Those patterns are where the cross-request caching actually pays off.
A rough set of signals that suggest LMCache is worth evaluating:
- System prompt exceeds roughly 2,000 tokens and is shared across a significant portion of traffic
- Multi-turn sessions commonly run longer than 10 turns before the user drops off
- You are running multiple inference instances serving the same document corpus
- GPU memory pressure is causing request queuing or OOM errors under normal load
If none of those apply, the simpler setup is usually the right choice.
—
Back to That Engineer
After the 2 a.m. incident, the team spent a week evaluating options. They eventually layered LMCache on top of their existing vLLM deployment and moved the fixed knowledge base summary’s KV data into CPU memory cache.
The results were measurable. TTFT on requests that hit the cache dropped by a factor they would not have gotten from hardware upgrades alone. GPU memory pressure stabilized. The operations channel went quiet.
The cost was real too. A few hundred gigabytes of CPU memory reserved for the cache, a cache invalidation process tied to their document update workflow, and an extra Redis instance in the monitoring rotation. For their workload, that trade made sense.
What the story illustrates is not that LMCache is better than vLLM. It is that the two tools address different parts of the same problem. vLLM makes inference efficient at the request level. LMCache makes it efficient across requests. In workloads where the same content shows up repeatedly, that second dimension matters quite a bit.
The question worth asking is not “which one should I use?” It is “where is my system paying full price for work it has already done?” If the answer involves large repeated prompts or long conversation histories, there is a specific kind of caching that was built for exactly that.
[//]: # (SEO)
[//]: # (title: LMCache vs vLLM: How Much Faster Is KV Cache Acceleration for LLM Inference?)
[//]: # (desc: A deep dive comparing LMCache and vLLM KV Cache mechanisms, analyzing acceleration in multi-turn chat, RAG, and long system prompt scenarios.)
[//]: # (kw: LMCache,vLLM,KV Cache,LLM inference acceleration,KV cache optimization)
[//]: # (category: comparisons)
Related reading
- LLM Inference Cost Is Collapsing: What Happens When AI Is Nearly Free in 2026
- On-Device AI Is Eating Cloud Inference From the Inside Out
- Modal vs Replicate vs Baseten vs RunPod: Which Serverless GPU Platform Should You Use for AI Inference in 2026?
- Claude Code Silently Runs Git Reset Every 10 Minutes: How Much Should You Trust AI Coding Tools?
- Forethought Is Too Expensive: 6 AI Customer Service Alternatives That Cost Less and Deploy Faster



