Google Achieves 3x LLM Inference Speedup on TPUs with DFlash — Diffusion-Style Speculative Decoding Breaks the Sequential Bottleneck
Summary Google and UCSD researchers have announced DFlash, a novel diffusion-style speculative decoding framework that achieves a 3.13x average speedup in LLM inference on Google’s TPU v5p hardware, with peak…
