Block-Diffusion LLMs Need Different Hardware Than Autoregressive Models. The Numbers Now Exist.
Hardware teams targeting edge AI silicon with autoregressive workload assumptions are designing the wrong memory hierarchy: block-diffusion LLMs have immutable prefix blocks that autoregressive KV-caches cannot exploit, and the co-design penalty is now quantified at 3.8x energy.