How Cursor Ships a 1TB Model Across the World Mid-Training
Cursor ships RL model updates as compressed weight deltas ~20x smaller than the full 1TB model.
“My delta might be like 20 times smaller than was shipping the full model with and this makes it practical”
Cursor engineers describe a system for shipping a 1TB model between distributed training clusters by exploiting the fact that RL changes only a regular subset of weights per step, letting them transmit lossless deltas roughly 20x smaller than the full model. This matters because it makes cross-cluster RL training practical and fast while controlling weight staleness, revealing the kind of database-style storage and reconciliation infrastructure required to train frontier coding models at scale.