.@​nathanbenaich explains why AI labs are shifting compute from pretraining to RL: "The last 4-5 years have been trying to figure out what the recipe is for pretraining, what is the best data mixture, what are the best ingredients to this whole magical soup." "Then over the years we figured out how to do that better. So it strikes me as normal that one would end up spending less money on exploratory work because the solution is more in the plain eye." "The whole pitch with RL is this idea that what we have as solutions today are really local maxima that were discovered through human ingenuity and sharing, through reading and whatnot. But that doesn't necessarily mean it's the global maxima." ▻