יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

בניית מערכת הכרה שפה יוונית-אנגלית ליישום

Building a Production Greek-English Speech Recognizer
במאמר זה, צוות חקר תיאר את תהליך בניית מערכת הכרה שפה יוונית-אנגלית ליישום. הם תיארו פיתוח של שישה שלבים בפייפל הנתונים, כולל עריכת חומרת אודיו-כמותי. הם גם תיארו רשימת שבעה ניסיונות שנבדקו אך לא נשלחו.
תקציר מקורי באנגליתarXiv:2609.13498v1 Announce Type: cross Abstract: We report a multi-month engineering program to build Sophea, a production bilingual Greek-English automatic speech recognition system. We evaluate the system against nine production gates covering Greek and English word error rate, language identification, and hallucinations on non-speech audio. Across twenty-three training iterations and two model architectures, no training-data composition passed all nine gates simultaneously. Meeting the Greek noisy-environment target required about 1,500 steps of dense domain exposure, while preserving English language identification tolerated only about 250 steps, or about 1,250 with a rebalanced mix that reduced Greek accuracy. We describe a six-stage data pipeline in which calibrating an audio-qualit
קרא במקור המקורי