What You Are Actually Measuring When You Benchmark Speculative Decoding Under FP8: A measurement case study on vLLM 0.29.0 and NVIDIA L4
Description
An independent measurement study of speculative decoding and FP8 quantization on vLLM 0.29.0 / NVIDIA L4 (SM89), released with the harness, the raw records, and the full history of its own corrections.
We set out to measure how FP8 quantization interacts with speculative decoding and could not, four times, until the measurement apparatus itself was fixed. Four confounds are identified and quantified:
- Attention-backend selection is conditioned on KV-cache dtype. At BF16 KV vLLM selects FlashAttention-2; at fp8_e4m3 it selects FlashInfer, so a precision A/B under default settings is also a kernel A/B. Target and draft models select independently, so pinning one does not pin the other.
- Speculative throughput is unreproducible across boots of an identical configuration while acceptance is not: coefficient of variation 13.92% over 12 fixed-prompt boots, collapsing to 1.44% under --enforce-eager. The dispersion is associated with the CUDA-graph path; the mechanism is unidentified, and boot and host effects were not separated. It is specific to draft-model speculation (n-gram gives 2.67% at the same n).
- Acceptance length has a boot-to-boot coefficient of variation of 0.62% under repeated identical boots, which we could not find reported. We do not convert this into a detection threshold.
- "Inter-token latency" is inter-chunk latency when computed from stream-arrival timestamps. A speculative step emits every accepted token in one chunk, so at concurrency 64 a p95-ITL objective scores the highest-throughput configuration at zero compliant goodput, at every threshold from 20 to 75 ms. This one does not add error, it inverts the decision.
The first three add error; the fourth reverses the ranking.
Contents of this archive: the LaTeX source and PDF of the paper; 34 numbered findings, of which 8 are retractions, each carrying a mandatory "what this does NOT establish" section; the raw append-only measurement records; the analysis scripts that derive every table in the paper from those records; and the audit scripts that verify every numeric claim traces to a stored record and that no finding rests on terminal output.
Reproducibility: every result is re-derivable by a third party on rented hardware. Requires a CUDA GPU of compute capability 8.9 or above and roughly 22 GiB of VRAM. The harness never imports the engine; it launches vLLM as a separate server process and speaks only HTTP. Two disclosed provenance defects (F009, F014) rest on probe output predating per-measurement recording; neither is cited in the paper and both are disclosed in its threats section.
Licensing: the harness and analysis code are MIT (see LICENSE). The measurements, findings and the paper are CC BY 4.0 (see LICENSE-DATA). The split is deliberate: the harness should be reusable with minimal friction, and the scientific record should carry attribution.
This is a preprint. It has not been peer reviewed.
Notes
Files
akshathtiwari/spec-fp8-study-v1.3.0.zip
Files
(9.0 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:fe385c79417410b8cef3e38d595f31ad
|
9.0 MB | Preview Download |
Additional details
Related works
- Is supplement to
- Software: https://github.com/akshathtiwari/spec-fp8-study/tree/v1.3.0 (URL)
Software
- Repository URL
- https://github.com/akshathtiwari/spec-fp8-study