?
Экспертная оценка ответов больших генеративных моделей на запросы врачей в клинической практике
Large language models (LLMs) in healthcare currently attract significant interest. This study compared GigaChat and YandexGPT by evaluating 1,200 LLM-generated responses to physician queries from outpatient and inpatient settings across 7 topics. Thirteen expert physicians assessed the responses using a 5-point scale across five criteria: relevance, correctness, safety, completeness, and language literacy. We performed paired t-tests and calculated effect size. Heterogeneous differences between LLM scores were observed depending on assessment criteria and topic. Safety and correctness were unsatisfactory on several topics, including “Medications” for GigaChat (3.1±1.5 correctness, 3.3±1.6 safety) and “Diagnostic parameters, reference ranges, and their interpretation” for YandexGPT (3.3±1.5 for both). Overall, our findings suggest that LLMs are not yet ready to function as physician assistants in clinical practice.