יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

TypedBench: בנצ'מרק לכיול ועלות במודלים

TypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision Models
TypedBench הוא בנצ'מרק למודלים מוכווני טיפוס. הוא בודק כיול, רגישות למילים ועלות. הבנצ'מרק כולל 7 יוצרים מתויגים ו-9 חבילות הערכה. הוא מדווח על דיוק ושגיאת כיול.
תקציר מקורי באנגליתarXiv:2610.11392v1 Announce Type: new Abstract: System One models output calibrated probabilities over typed answers such as categorical choices, ordinal levels, or binary outcomes, via a non-generative interface. Software can act on these probabilities through thresholds, cost-weighted choices, and escalation rules. Consequently, if these probabilities are miscalibrated or wording-sensitive, the software ma take unintended actions leaving human operators with no textual rationale to inspect. Current evaluations largely report accuracy and calibration on public classification datasets without a clear reference. We present TypedBench, a benchmark for typed decision models like Jev built from seven policy-labelled generators and nine evaluation suites. We report accuracy as median and range
קרא במקור המקורי