כתבה
arXiv cs.CL ·
OpenForgeRL: Train Harness-native Agents in Any Environment
תקציר מקורי באנגליתarXiv:2607.21557v2 Announce Type: replace-cross Abstract: Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training harness-based agents end-to-end in diverse environments. OpenForgeRL achieves this with a lightweight proxy that serves the harness's model calls while recording them as training data for a standard RL codebase (e.g., veRL), and a Kubernetes orchestrator that runs each rollout in its own remote con
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית