The term covers open-source stacks that bundle the four parts of an RL post-training loop: a rollout engine that generates completions, a trainer that computes advantages and updates weights, an orchestrator that schedules both across nodes and survives failures, and an environment layer that defines and scores the task. Projects in this category include Miles from RadixArk, SkyRL, Prime Intellect's stack, and OpenRLHF; the framework layer is increasingly free while environments and GPU hours remain the scarce inputs.