Reinforcement Learning without Verifiable Rewards — Will Brown, Prime Intellect
Prime Intellect is extending RL to real-world tasks lacking verifiable reward signals.
Will Brown of Prime Intellect presented on the challenge of applying reinforcement learning to messy, real-world tasks where clear verifiable rewards don't exist — a gap left by RLVR approaches that dominated the past 18 months. The talk synthesizes internal research and broader literature on methods to extend RL beyond clean, checkable benchmarks. This matters because most production agentic use-cases lack neat verifiers, making reward design a key bottleneck for real-world RL deployment.