FlashSpec: Adaptive LLM Inference with Speculative Decoding — Six Hard-Won Lessons
FlashSpec implements speculative decoding using GPU-native Triton kernel verification and online bandit-based draft model selection, sharing six practical lessons on specification-first development, hidden temperature bugs, cross-platform packaging pitfalls, kernel performance trade-offs, property-based testing value, and adaptive algorithm prerequisites.
