Akamai Developer Sheilah Kirui Demonstrates Speculative Decoding

On October 7, 2026, Akamai developer advocate Sheilah Kirui demonstrated live on the AI Engineer podcast that speculative decoding cuts inference latency 1.6x on structured coding tasks using a single NVIDIA Blackwell GPU, while failing to accelerate creative workloads where wide-open output spaces cause draft token rejection, BigGo Finance reported.

Hardware Constraints and the Dual-Model Memory Tax

Running speculative decoding is not a free performance tier. As Kirui outlined during her live session on stage, the technique forces inference servers to host two separate models simultaneously on the same silicon. Both the baseline target model and the smaller draft model require their own distinct KV cache allocations, which expand rapidly as context lengths grow.

During the demonstration, hardware limitations dictated the architecture. Kirui deployed a setup split across two vLLM servers running on a single Blackwell GPU. The configuration relied on a target model carrying approximately 16 GB of weights alongside a draft model scaled down to roughly 2.5 GB. This specific sizing left enough VRAM headroom to accommodate the necessary working memory for both KV caches.

In high-concurrency environments where GPUs are already fully saturated processing active requests, introducing a secondary draft model creates fatal memory pressure without delivering any latency reduction.

Structured Output Wins While Creative Tasks Stagnate

The operational value of speculative decoding is strictly bounded by the nature of the workload. Parallel tabs run during the demonstration exposed a stark performance divide between deterministic programming tasks and open-ended generative writing.

When processing structured outputs such as writing JSON, generating SQL prompts, or writing code, the auxiliary draft model successfully anticipated the target model’s choices. This high acceptance rate—defined as tokens accepted by the target model divided by total tokens generated by the draft model—delivered a 1.6x speedup and boosted overall token throughput.

Conversely, running creative workloads such as poetry or brainstorming collapsed the acceptance rate. High temperature settings combined with an expansive output space meant the draft model’s autoregressive guesses frequently missed. Every rejected token forced the larger target model to step in and recompute the sequence, negating any speed advantages.

Component Memory Footprint Operational Notes
Baseline Target Model ~16 GB of weights Leaves substantial room for KV caching on a single Blackwell GPU.
Draft Model ~2.5 GB of weights Small enough to co-host, sized 10x to 50x smaller than the target.

Why Prompt Length Limits Speculative Speedups

Engineering teams adopting speculative decoding must account for where the optimization actually applies. The technique accelerates token generation during the decode phase, but it leaves the initial prefill phase completely untouched.

As Kirui emphasized in her presentation, when a workload demands massive amounts of input context—such as retrieval-augmented generation or extensive document analysis—the model spends the vast majority of its processing cycle reading the prompt and building the initial KV cache. Because generated tokens remain far fewer than input tokens in context-heavy pipelines, the absolute latency savings shrink relative to total request duration.

Reading model weights from High Bandwidth Memory to produce a single token dominates execution time. Speculative decoding bypasses this by letting a smaller model handle rapid speculative passes, but this mechanism provides zero relief during prefill operations.

Using the vLLM Proposer Menu

Beyond traditional draft models, serving architectures have evolved to offer multiple paths for sequence acceleration. As documented by blog.teliaz.com, vLLM now ships a menu of proposers, each with a different profile designed to balance latency and throughput.

  • Draft Model: The classic approach utilizing a smaller secondary checkpoint from the same model family to share tokenizers.
  • EAGLE: Attaches a lightweight prediction head directly to the target model’s hidden states, avoiding the memory footprint of a separate draft model.
  • Multi-Token Prediction (MTP): Relies on native MTP modules trained directly alongside the target model weights where available.
  • N-gram & Suffix Decoding: Uses zero-model approaches that match patterns in the prompt and generated text, proving exceptionally effective for repetitive code editing and structured RAG tasks without adding VRAM pressure.

The unresolved tension across enterprise deployments remains clear. Low-concurrency environments have the spare VRAM headroom required to host secondary draft models, but they benefit least from absolute latency cuts at scale. High-concurrency environments need token acceleration the most, yet their saturated GPUs leave no room for auxiliary model weights.

Photo of author

Sophie Lin - Technology Editor

Sophie is a tech innovator and acclaimed tech writer recognized by the Online News Association. She translates the fast-paced world of technology, AI, and digital trends into compelling stories for readers of all backgrounds.

EU Customs Charges: €31.95 Hidden Fees for Irish Buyers