כתבה
arXiv cs.AI ·
Statistically Supported LLM Ingredient and Recipe Data Collection in Computational Nutrition
תקציר מקורי באנגליתarXiv:2607.23273v1 Announce Type: cross Abstract: Computational nutrition needs precise ingredient data, but current databases are incomplete, inconsistent, and built for human reference rather than automated reasoning. LLMs could help fill these gaps, but single-pass outputs are unreliable and can introduce silent errors into downstream computation. We present a quality-controlled LLM pipeline for ingredient data acquisition that combines robust statistical estimation, domain-specific invariant checks, and a web-fetch fallback. An illustrative Heap's Law fit to 233 recipes suggests that unique-ingredient growth is sub-linear and front-loaded: the projected ratio of unique ingredients to recipes falls from 1.74 at 100 recipes to 0.19 at 5,000. For each ingredient attribute, repeated LLM qu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית