Our chart this week shows how frontier labs split their compute between three jobs. In early 2024, pre-training took 67% and post-training with reinforcement learning (RL) took 4%. By the third quarter of 2026 those figures had almost reversed, to 10% and 53%. Inference stayed at about a third throughout.
These are shares, not totals. Lab compute has grown many times over since 2024, so labs are probably still doing more pre-training in absolute terms than before. What has changed is how fast RL has grown. In RL, a model attempts a task, its output is graded, and it learns from the result, millions of times over. Most of the gains in coding, maths and agent work over the past 18 months have come from this process.
RL also needs different inputs. Pre-training draws on text that already exists. RL needs practice environments, such as software sandboxes, simulated websites and copies of business tools. Each one comes with a test that can tell good work from bad, and building them is now an industry in its own right.
