The Hallway Track
Research Findings

Weight Folding, CUDA Streams, and the Bug That Made My Model Speak Backwards — Filip Makraduli

AI Engineer · Sep 19, 2026 · Research Findings

Two lines of algebra make transformer RMS norm cheaper and faster via weight folding and deferred division.

“the GPUs are not slow or bad at math, but they're bad at everything else around the actual math”

A co-authored arXiv paper proposes algebraic tricks to fuse normalization into matrix multiplication, apply weight folding, and defer division, reducing the wall-clock overhead of RMS norm layers that fire many times per decode step. It matters as a low-cost efficiency improvement to the transformer architecture, though it is an incremental optimization rather than a major industry signal.

transformers rms-norm gpu-optimization inference weight-folding

Watch / read the original source →