Readability and Appropriateness of Responses Generated by ChatGPT 3.5, ChatGPT 4.0, Gemini, and Microsoft Copilot for FAQs in Refractive Surgery

Fahri Onur Aydın; Burakhan Kürşat Aksoy; Ali Ceylan; Yusuf Berk Akbaş; Serhat Ermiş; Burçin Kepez Yıldız; Yusuf Yıldırım

doi:10.4274/tjo.galenos.2024.28234

Readability and Appropriateness of Responses Generated by ChatGPT 3.5, ChatGPT 4.0, Gemini, and Microsoft Copilot for FAQs in Refractive Surgery

Turk J Ophthalmol. 2024 Dec 31;54(6):313-317. doi: 10.4274/tjo.galenos.2024.28234.

Authors

Fahri Onur Aydın¹, Burakhan Kürşat Aksoy¹, Ali Ceylan¹, Yusuf Berk Akbaş¹, Serhat Ermiş¹, Burçin Kepez Yıldız¹, Yusuf Yıldırım¹

Affiliation

¹ University of Health Sciences Türkiye, Başakşehir Çam and Sakura City Hospital, Clinic of Ophthalmology, İstanbul, Türkiye.

PMID: 39743925
DOI: 10.4274/tjo.galenos.2024.28234

Abstract

Objectives: To assess the appropriateness and readability of large language model (LLM) chatbots' answers to frequently asked questions about refractive surgery.

Materials and methods: Four commonly used LLM chatbots were asked 40 questions frequently asked by patients about refractive surgery. The appropriateness of the answers was evaluated by 2 experienced refractive surgeons. Readability was evaluated with 5 different indexes.

Results: Based on the responses generated by the LLM chatbots, 45% (n=18) of the answers given by ChatGPT 3.5 were correct, while this rate was 52.5% (n=21) for ChatGPT 4.0, 87.5% (n=35) for Gemini, and 60% (n=24) for Copilot. In terms of readability, it was observed that all LLM chatbots were very difficult to read and required a university degree.

Conclusion: These LLM chatbots, which are finding a place in our daily lives, can occasionally provide inappropriate answers. Although all were difficult to read, Gemini was the most successful LLM chatbot in terms of generating appropriate answers and was relatively better in terms of readability.

Keywords: Artificial intelligence; ChatGPT; Copilot; Gemini; chatbots; refractive surgery FAQs.

MeSH terms

Comprehension*
Humans
Language
Patient Education as Topic
Reading
Refractive Surgical Procedures*
Surveys and Questionnaires