Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth
Unsloth fixed a gradient accumulation bug improving training accuracy by 1-3% across the entire stack.
“fixed a gradient accumulation bug fix um which increased accuracy by 1 to 3% um across the entire training stack”
Daniel Han of Unsloth (one of Hugging Face's top model distributors with 300M+ downloads) outlines the org's broad contributions to the open-source AI stack, including async gradient checkpointing, flex attention, and a gradient accumulation bug fix that boosted training accuracy 1-3%. The talk is a workshop intro covering kernels, RL, and reward hacking in agents. Content is cut off before the substantive technical sections begin, limiting assessable signal.