10% worse, 100x cheaper, 10000x faster: Why Simulation is taking over
Since 2022, each year has seen one more component of the ML pipeline transition from human-made to model-made, moving from synthetic reward signals to synthetic environments and human subjects The progression follows seven stages: reward signal (2022), training data (2023), teacher (2023), curriculum (2024), researcher (2026), environment (2026), and human subject (2025) Each flipped component follows the same pattern: ~10% worse than human-made, but 100x cheaper and 10,000x faster, with a "pati
Analysis
TL;DR
- Since 2022, each year has seen one more component of the ML pipeline transition from human-made to model-made, moving from synthetic reward signals to synthetic environments and human subjects
- The progression follows seven stages: reward signal (2022), training data (2023), teacher (2023), curriculum (2024), researcher (2026), environment (2026), and human subject (2025)
- Each flipped component follows the same pattern: ~10% worse than human-made, but 100x cheaper and 10,000x faster, with a "patient zero" paper or product marking the frontier adoption
- Key milestones include InstructGPT/Constitutional AI for reward modeling, Phi/Nemotron for synthetic training data, Alpaca/DeepSeek-R1 for distillation, Meta's Self-Rewarding models for curriculum, Karpathy's autoresearch for automated experimentation, and Z.ai/GLM-5.3 for synthetic RL environments
- The loop is now closing on itself: models generate their own data, teach themselves, design their own curricula, run their own experiments, and simulate their own training environments and human subjects
Why It Matters
This framework provides AI practitioners with a strategic map of where the field is heading — the entire ML pipeline is becoming self-sustaining, which fundamentally changes how organizations should think about R&D investment, talent, and competitive moats. The implication is that human expertise shifts from being the primary producer of training components to being the architect of systems that produce those components autonomously, creating both opportunity and disruption across the industry.
Technical Details
- Stage 1 (Reward Signal, 2022): InstructGPT introduced reward modeling from human preferences; Constitutional AI (RLAIF) had models critique themselves against principles; LLM-as-judge became standard via MT-Bench and AlpacaEval, replacing human evaluators entirely.
- Stage 2 (Training Data, 2023): Microsoft's Phi series demonstrated textbook-quality synthetic data outperformed scale; Apple's WRAP rephrased the entire web for 3x pretraining efficiency; NVIDIA's Nemotron-4 340B shipped a permissively licensed synthetic data pipeline; reasoning-trace corpora became standard by 2025.
- Stage 3 (Teacher, 2023): Stanford's Alpaca showed a $600 fine-tune on GPT instructions could clone frontier behavior; Vicuna and Orca advanced distillation; on-policy generalized knowledge distillation resolved train/inference mismatch; DeepSeek-R1 made distilled model families the default release pattern.
- Stage 4 (Curriculum, 2024): Meta's Self-Rewarding Language Models and SPIN demonstrated models generating their own tasks, judging their own outputs, and improving beyond human preference data ceilings — curriculum design became self-directed.
- Stage 5 (Researcher, 2026): DeepMind's AlphaEvolve evolved new algorithms in 2025; Sakana's AI Scientist (published in Nature) automated paper writing; Karpathy's autoresearch used a minimal ratchet loop (700 experiments, 20 kept improvements) to cut time-to-GPT-2 from 2.02 to 1.80 hours.
- Stage 6 (Environment, 2026): Z.ai/GLM-5.3 synthesized environments end-to-end with research agents mining real work patterns, judge agents confirming solvability, and verifiers stress-tested without seeing reference solutions; Ornith-1.5 claimed end-to-end self-improvement with models proposing tasks and generating RL rollouts.
- Stage 7 (Human Subject, 2025): Simile replaces human subjects in the loop; lineage from Generative Agents (Smallville, 2023) to 1,000-person simulations achieving 85% accuracy in reproducing human survey and behavioral responses.
Industry Insight
- Organizations should invest in synthetic pipeline infrastructure now — the models that can most efficiently produce synthetic reward signals, data, environments, and researchers will have compounding advantages as each stage feeds the next.
- The "10% worse, 100x cheaper, 10,000x faster" tradeoff means near-term synthetic components will have quality ceilings; the strategic play is to build systems where synthetic and human components coexist during transition periods, with clear upgrade paths as synthetic quality improves.
- Talent strategy must shift: the highest-value roles move from hands-on data labeling, reward engineering, and curriculum design to system architecture, verification, and oversight of autonomous synthetic pipelines — the human becomes the verifier, not the producer.
Disclaimer: The above content is generated by AI and is for reference only.