Tagged articles

Triton kernel

1 articles · Page 1 of 1
DeepHub IMBA
DeepHub IMBA
Sep 3, 2026 · Artificial Intelligence

FlashSpec: Adaptive LLM Inference with Speculative Decoding — Six Hard-Won Lessons

FlashSpec implements speculative decoding using GPU-native Triton kernel verification and online bandit-based draft model selection, sharing six practical lessons on specification-first development, hidden temperature bugs, cross-platform packaging pitfalls, kernel performance trade-offs, property-based testing value, and adaptive algorithm prerequisites.

CI/CDLLM inferenceThompson sampling
0 likes · 15 min read
FlashSpec: Adaptive LLM Inference with Speculative Decoding — Six Hard-Won Lessons