A startup is betting on diffusion-based LLMs because they are inherently more parallel at inference time than autoregressive transformers.
“the bitter lesson is that the more parallel solution is the one that is eventually going to win”
8 tracked signals on transformers.
A startup is betting on diffusion-based LLMs because they are inherently more parallel at inference time than autoregressive transformers.
“the bitter lesson is that the more parallel solution is the one that is eventually going to win”
Transformer architecture standardization will enable extreme chip specialization and massive inference performance gains
“I think we're gonna have crazy levels of specialization because there's gonna be so much inference demand. And that's gonna unlock just like super exciting levels of performance.”
Two lines of algebra make transformer RMS norm cheaper and faster via weight folding and deferred division.
“the GPUs are not slow or bad at math, but they're bad at everything else around the actual math”
DiScoFormer introduces a single transformer that models both density and score across multiple distributions.
Hugging Face analyzes which token types hybrid attention-SSM models predict better than transformers.
GPU design shifted from compute efficiency to memory bandwidth after transformers emerged
“attention is equal to n-squared”
NVIDIA NeMo AutoModel accelerates fine-tuning of Hugging Face Transformers models.
Optimizing transformers for low-precision training cuts GPU hours and speeds up experimentation and model scaling.
“Accelerating transformers is therefore not just a performance optimization, but directly affects how quickly teams can experiment and how large a model they can afford to train.”