An LLM post-training framework for RL scaling, designed to improve training efficiency and performance in reinforcement learning scenarios
1.4k upvotes
Visit GitHub →
On mobile tap Share for WeChat / RED (Xiaohongshu) / X; on desktop use Copy text and paste into the app.