DatologyAI generated 12 trillion synthetic tokens for pre-training across web, math, and code domains.
“we have hit a "data wall." You need to spend exponentially more computing power and data to get models that only get linearly better.”
9 tracked signals on synthetic-data.
DatologyAI generated 12 trillion synthetic tokens for pre-training across web, math, and code domains.
“we have hit a "data wall." You need to spend exponentially more computing power and data to get models that only get linearly better.”
Simile AI raises $2B to build behavioral foundation models simulating human decision-making at scale
“What if we can just recreate the world that we live in?”
Six open reproductions of the closed 'Jev' decision model appeared within two days of its viral launch.
“Of course, not enough people are talking about the data side, which is acknowledged to be 100% synthetic”
Poolside releases open-weight Laguna M and Laguna XS models, shifting from enterprise-only distribution.
“at least at Pulsar we don't see it as a way to replace organic data”
Bespoke Labs says AI post-training has shifted from knowledge data to RL environments for agentic tasks
“we have moved on from knowing to doing right so that's the idea of agents”
NVIDIA releases synthetic 3D medical imaging pipeline to unblock radiology AI model training
ONESTRUCTION built a BIM foundation model using synthetic data and RLVR to overcome niche data scarcity.
NVIDIA proposes GPU-native medical physics simulation to solve healthcare robotics data scarcity
“Unlike autonomous driving or industrial robotics, healthcare robotics can't rely on internet-scale data collection or unlimited real-world experimentation.”
No content was provided to extract signal from.