כתבה
arXiv cs.AI ·
PetQA: בנצ'מרק לידע וטרינרי
PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning
PetQA הוא בנצ'מרק לידע וטרינרי וניהול קליני. הוא מכיל 10,076 זוגות QA טקסט בלבד ו-8,751 זוגות QA רב-מודאליים. הבנצ'מרק מודד 18 מודלים, כולל LLM.
תקציר מקורי באנגליתarXiv:2609.04598v1 Announce Type: cross Abstract: We introduce PetQA, a Korean long-form question-answering (QA) benchmark for evaluating veterinary knowledge and clinical reasoning in large language models (LLMs) and large vision-language models (LVLMs). PetQA contains 10,076 text-only and 8,751 multimodal QA pairs derived from real-world questions about dogs and cats, paired with answers from expert veterinarians. Its test split, PetQA-Bench, further includes annotations for question types and clinical conditions. We evaluate eighteen models using ROUGE, BERTScore, and LLM-as-a-judge metrics for factuality and helpfulness under three settings: zero-shot inference, retrieval-augmented generation (RAG), and supervised fine-tuning (SFT). The benchmarking results provide an overview of the s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית