כתבה
arXiv cs.LG ·
Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
תקציר מקורי באנגליתarXiv:2509.03059v2 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have shown that their reasoning capabilities can be significantly improved through Reinforcement Learning with Verifiable Reward (RLVR), particularly in domains like mathematics and programming, where ground-truth correctness can be automatically evaluated. However, extending this success to other reasoning-intensive domains remains challenging due to the scarcity of high-quality, verifiable datasets and the high cost of human supervision. In this work, we introduce the Loong Project: an open-source framework for scalable synthetic data generation and verification across a diverse range of reasoning-intensive domains. The framework consists of two key components: (1) LoongBench, a curated se
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית