Optimizing Token Throughput with Speculative Decoding & FP8 Quantization
Serving frontier 70B+ parameter models at low latency is one of the toughest challenges in production AI infrastructure. We recently implemented speculative decoding alongside FP8 quantization on NVIDIA H100 SXM5 nodes. Speculative Decoding