Optimizing Token Throughput with Speculative Decoding & FP8 Quantization
Serving frontier 70B+ parameter models at low latency is one of the toughest challenges in production AI infrastructure. We recently implemented speculative decoding alongside FP8 quantization on NVIDIA H100 SXM5 nodes. Speculative Decoding
72
6