כתבה
arXiv cs.AI ·
Evaluating and Benchmarking the System One Model Jev
תקציר מקורי באנגליתarXiv:2609.37647v1 Announce Type: cross Abstract: Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed options, a position on a rubric, or the probability that a statement is true, with probabilities the vendor describes as calibrated. Such models target small decisions in information access pipelines, such as routing queries, checking grounding, moderating content, or rating against a rubric. We evaluate Jev (jev-1.13.0) zero-shot on 37 datasets spanning classification, routing, natural language inference, reading comprehension, commonsense reasoning, moderation, legal clause analysis and rubric scoring, with one frozen template per dataset and full evaluation splits: 346,009 requests for under
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית