Another DeepSeek Moment Has Arrived
DeepSeek's updated flash model beats its larger pro version via post-training improvements alone
“Just the post-training step changed?”
DeepSeek's updated flash model achieves dramatic benchmark gains—some results more than doubling, one improving 7x—without changing the underlying model architecture or size, solely through improved post-training. The smaller flash model now outperforms the pro model that is five times larger, demonstrating that post-training is a critical and underappreciated lever for capability gains. The model is open-weight and significantly cheaper than frontier alternatives, continuing the pattern of Chinese labs delivering high-efficiency disruptions to Western AI incumbents.