AI HUM SU Sulbha Jain Why Agentic RL Breaks (and How rStar2-Agent Fixes It) — Paper Review If you’ve ever watched an LLM use tools during reasoning, you’ve seen the magic: