יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

SchemaFill: כלי קריאה יעיל ל-LLM דרך דקודינג ספקולטיבי-סלוט-פאראלל

SchemaFill: Efficient LLM Tool Calling via Slot-Parallel Speculative Decoding
SchemaFill מציע כלי קריאה יעיל ל-LLM דרך דקודינג ספקולטיבי-סלוט-פאראלל. הכלי משפר את הערכת התוצאות של קריאת כלים על ידי דקודינג סלוט-פאראלל, ומציע תמיכה למודלי LLM כמו Claude ו-GPT-5.
תקציר מקורי באנגליתarXiv:2610.07086v1 Announce Type: cross Abstract: LLM agents interact with external systems by generating structured tool calls. Given a user request, conversational context, and a catalog of tool schemas, a tool-calling model must select tools and generate their arguments, potentially producing multiple calls in a single response. Standard autoregressive decoding generates these calls token by token, incurring substantial latency for requests involving multiple calls or many argument fields. The explicit argument structure offers opportunities for parallel generation, but later argument values may depend on preceding fields and calls, so independently generated values can differ from the target model's output. We present SchemaFill, a framework for efficient LLM tool calling through slot-
קרא במקור המקורי