Researchers train a sparse autoencoder on a language model's activations so that each input is explained by a few active latents, which tend to be more distinct than raw neurons. Latents can be tested causally by steering, meaning adding or subtracting their direction during generation. OpenAI used SAE latents to study metagaming in an o3 reinforcement learning run.