יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

ORCA-bench: האם מוכנים מודלי שפה לניתוח שורשי?

ORCA-bench: How Ready Are Language Model Agents for Oncall?
ORCA-bench הוא בנץ'מרק שבודק את יכולתם של מודלי שפה לניתוח שורשי. הוא משתמש במערכת מיקרו-שירותים עם נתונים אמיתיים ומודלים כמו Claude. התוצאות מראות שיש עוד הרבה עבודה לפני שמודלים יוכלו להיות מהימנים בייצור.
תקציר מקורי באנגליתarXiv:2607.28545v3 Announce Type: replace Abstract: Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began. We introduce ORCA-bench, a benchmark that puts general-purpose coding agents in a production-fidelity oncall setting. ORCA-bench pairs 1,079 RCA tasks with six days of metrics, logs, and traces collected from an OpenTelemetry-instrumented microservice system under continuous simulated user load. Agents investigate this recorded history through real observability interfaces---Prometheus, Jaeger, and OpenSearch via Grafana---with full access to application source code. Tasks sys
קרא במקור המקורי