[AINews] Megakernels are so dead and so back
NVIDIA's Rubin GPU architecture kills megakernels, ending a major inference optimization research direction
“the GPU is designed in such a way that it kills mega kernels. So it seems like that entire research field won't be continued.”
Practitioners on the Latent Space Inference Engineering Masterclass argue that megakernels—hand-fused GPU kernels meant to reduce launch overhead—are dying in production, with modular kernels from TensorRT-LLM already outperforming them. NVIDIA's Rubin GPU introduces dependency triggers and straggler-CTA handling that eliminates the core bottlenecks megakernels were designed to solve. This effectively closes off a notable systems research direction and has implications for inference startups that bet on kernel fusion as a competitive moat.