יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

בדיקת חישובי הסתברות ל-AI גנרטיבי: תשתית לעבודות עסקי

Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads
בדיקת חישובי הסתברות ל-AI גנרטיבי, המודל GPT-5, והתשתית לעבודות עסקי
תקציר מקורי באנגליתarXiv:2610.07094v1 Announce Type: cross Abstract: LLM deployment is shifting from single-turn completion to agentic trajectories in which a model plans, calls tools, reads results and reasons at test time before acting. This inverts the economics of inference hardware: chat serving amortises weight reads across large batches, whereas agent trajectories are sequentially dependent, run at effective batch one, and make per-token decode latency (TPOT) the dominant term in task completion time. Using a roofline analysis and a closed-form episode-latency model, we show why this regime favours accelerators that keep weights in on-die SRAM (Cerebras WSE-3/3T, Groq/NVIDIA LPU) or compiler-managed tiered memory (SambaNova SN40L/SN50), and why three vendor ecosystems converged in 2026 on disaggregate
קרא במקור המקורי