The Hallway Track

reinforcement-learning

38 tracked signals on reinforcement-learning.

Frontier Tuning: Microsoft Build 2026

Microsoft Developer (Build) · Jun 03, 2026

Microsoft launches Frontier Tuning, letting enterprises reinforcement-fine-tune AI models on their own M365 data and workflows.

“With Frontier Tuning, we're making it possible for you to create your own enterprise AI.”
Taking Reinforcement Learning Cross Datacenter — Nan Jiang, Modal

AI Engineer · Aug 10, 2026

Modal proposes decoupling RL rollout workers from trainer clusters to use distributed GPU capacity across datacenters.

“IO wants all four of these at the same time. Enough GPU, same region, fast fabric, and available now. Any of these like is manageable, but all four of them that are pretty hard to get at the same time.”
Hugging Face Journal Club: Kimi K3

Hugging Face · Jul 29, 2026

Kimi K3 is trained to be comparable to Claude Opus 4.8 using novel 3x3 domain expert distillation.

“they've trained a model that is probably comparable to Opus 4.8”
Frontier Tuning: Teaching AI to work the way you do

Microsoft 365 Dev Blog · Jun 02, 2026

Microsoft launches Frontier Tuning: RL-based AI customization within enterprise compliance boundaries

“a new approach to making AI work the way your business does by applying reinforcement learning inside your compliance boundary with your own data, processes, and conventions”
Cursor | The Hidden Bug in Every Large-Scale RL Run

Sequoia Capital · Jun 02, 2026

Numerical mismatch in inference log-probabilities is a hidden bug plaguing large-scale asynchronous RL runs.

“Hopefully next composer versions are going to be our own model instead of basing it off an open source base.”
How Cursor Ships a 1TB Model Across the World Mid-Training

Sequoia Capital · Jun 01, 2026

Cursor ships RL model updates as compressed weight deltas ~20x smaller than the full 1TB model.

“My delta might be like 20 times smaller than was shipping the full model with and this makes it practical”
The Base Model Is Dead — Varun Singh, Arcee AI

AI Engineer · Jul 31, 2026

The traditional base-model paradigm of web-scale pre-training is being displaced by post-training

“RL was mostly just a cherry on top, shaping the flavor of the interactions more than conferring extra knowledge or quality onto the base model itself.”
Using RL Agent to Detect and Remediate ETL Pipeline Failures - Anna Marie Benzon

AI Engineer · Jun 29, 2026

An RL agent autonomously diagnoses and remediates ETL pipeline failures with safety bounds, cutting manual recovery from ~2.5 days.

“The central question is simply whether an agent can act, but whether it can act usefully, explainably, and within the boundaries that an operation would actually trust.”
Cursor |Why Online RL Is Just the Cherry on Top

Sequoia Capital · Jun 03, 2026

Cursor uses online (real-time) RL only to polish already-shipped models, not build them from scratch.

“That's kind of the paradox of online RL or how we like to call it real time is that, you know, we can't use this to really create the model from scratch because users need to be using the model.”
Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI

AWS Machine Learning Blog · Oct 02, 2026

Multi-turn RL on SageMaker lets small models match frontier reliability for search agents

“Fine-tuning offers a third path: you teach a small model your tools and environment directly. The result is a small model's speed and cost with the reliability that would otherwise require a frontier model.”
From RL to IRL — Gaurav Mishra, Amazon AGI Lab

AI Engineer · Aug 14, 2026

Amazon AGI Lab researcher details failure modes of RL-trained agents in real-world deployment

“That's why we've been able to train really compelling coding agents using RL.”
NVIDIA's AI Learns Why Copying Humans Isn't Enough

Two Minute Papers · Aug 02, 2026

NVIDIA combines imitation learning and goal-based RL to train parkour AI on just 30 seconds of data

“Systems that copy us are beautiful, but brittle, repetitive. Systems that chase goals adapt, but often stop moving like us.”
5 Papers That Show Where AI Research Is Heading Right Now

Y Combinator · Jun 12, 2026

AlphaZero-style self-play, unbiased by human data, is the likely path to far more intelligent systems and possibly AGI.

“alpha zero unbiased by um humans meandering is uh the way we'll get to much more intelligent systems, maybe even dare say agi”
Training Agents 3: Reinforcement Learning

Hugging Face · Jul 28, 2026

Hugging Face is teaching GRPO reinforcement learning for agent training in a live stream series

“it's pretty straightforward to learn from the available options and you can apply it on most use cases”
We got @unsloth a DGX Station!

NVIDIA GTC · Sep 23, 2026

NVIDIA gave Unsloth a DGX Station to accelerate its model quantization and RL research.

“We're going to be use utilizing this to do all of the models. And just provide more and more for the community.”