Mapping Gaps and Improvement Targets in Large Language Model-Generated Melanoma Patient Education in a Non-English Setting
DOI:
https://doi.org/10.58600/eurjther3197Keywords:
health literacy, large language models, melanoma, patient education, public health, therapeutic communicationAbstract
Objective: Large language models (LLMs) are increasingly being used to develop medical education materials; however, it remains unclear how reliable, readable, or guideline-compliant the content generated by these models is for non-English-speaking patient groups. We evaluated the quality of Turkish melanoma patient education texts generated by seven frontier LLMs.
Methods: A standardized 22-item Turkish prompt, built from international melanoma guidelines, was put to seven models in zero-shot sessions: ChatGPT 4.0 Turbo, Gemini 2.0 Flash, Claude 3.7 Sonnet, Grok 3, Qwen 2.5 Plus, DeepSeek R1, and Mistral Large 2. Each output was rated for readability (Ateşman Index), understandability, how clearly medical terminology was explained, scientific reliability (DISCERN instrument), empathy, and adherence to a 31-item guideline-based checklist. Model comparisons were summarized descriptively, using model-level absolute scores, score ranges, and rankings.
Results: Model performance differed across readability, understandability, reliability, empathy, and guideline-adherence domains. DeepSeek R1 led on both readability (81.6) and understandability (23.5/25). Guideline adherence was strongest for Grok 3 and DeepSeek R1, at 96.8% and 93.5%, respectively, and Grok 3, DeepSeek R1, and Gemini 2.0 Flash each scored above 90% on the normalized total DISCERN measure. DeepSeek R1 also recorded the highest empathy score (90%). Gemini 2.0 Flash had the lowest readability score and produced the longest output (Ateşman 65.8). None of the models provided citations or verifiable sources, so every model received the lowest possible DISCERN Source Reliability score; Mistral Large 2 showed the weakest overall performance.
Conclusion: How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next. A few produced text that was clear, empathetic, and reasonably guideline-concordant, but the lack of verifiable citations and uneven guideline coverage remain genuine limitations. These findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.
References
[1] Siegel RL, Giaquinto AN, Jemal A (2024) Cancer statistics, 2024. CA Cancer J Clin. 74(1):12–49. https://doi.org/10.3322/caac.21820
[2] Whiteman DC, Green AC, Olsen CM (2016) The growing burden of invasive melanoma: projections of incidence rates and numbers of new cases in six susceptible populations through 2031. J Invest Dermatol. 136(6):1161–1171. https://doi.org/10.1016/j.jid.2016.01.035
[3] Gershenwald JE, Scolyer RA, Hess KR, Sondak VK, Long GV, Ross MI, Lazar AJ, Faries MB, Kirkwood JM, McArthur GA, Haydu LE, Eggermont AMM, Flaherty KT, Balch CM, Thompson JF (2017) Melanoma staging: evidence-based changes in the American Joint Committee on Cancer eighth edition cancer staging manual. CA Cancer J Clin. 67(6):472–492. https://doi.org/10.3322/caac.21409
[4] Luke JJ, Flaherty KT, Ribas A, Long GV (2017) Targeted agents and immunotherapies: optimizing outcomes in melanoma. Nat Rev Clin Oncol. 14(8):463–482. https://doi.org/10.1038/nrclinonc.2017.43
[5] Swetter SM, Johnson D, Albertini MR, Barker CA, Bateni S, Baumgartner J, Bhatia S, Bichakjian C, Boland G, Chandra S, Chmielowski B, DiMaio D, Dronca R, Fields RC, Fleming MD, Galan A, Guild S, Hyngstrom J, Karakousis G, Kendra K, Kiuru M, Lange JR, Lanning R, Logan T, Olson D, Olszanski AJ, Ott PA, Ross MI, Rothermel L, Salama AK, Sharma R, Skitzki J, Smith E, Tsai K, Wuthrick E, Xing Y, McMillian N, Espinosa S (2024) NCCN Guidelines® Insights: Melanoma: Cutaneous, Version 2.2024. J Natl Compr Canc Netw. 22(5):290–298. https://doi.org/10.6004/jnccn.2024.0036
[6] Manne S, Heckman CJ, Frederick S, Schaefer AA, Studts CR, Khavjou O, Honeycutt A, Berger A, Liu H (2024) A digital intervention to improve skin self-examination among survivors of melanoma: protocol for a type-1 hybrid effectiveness-implementation randomized trial. JMIR Res Protoc. 13:e52689. https://doi.org/10.2196/52689
[7] Choi J, Cho Y, Woo H (2018) mHealth approaches in managing skin cancer: systematic review of evidence-based research using integrative mapping. JMIR Mhealth Uhealth. 6(8):e164. https://doi.org/10.2196/mhealth.8554
[8] Heckman CJ, Mitarotondo A, Lin Y, Khavjou O, Riley M, Manne SL, Yaroch AL, Niu Z, Glanz K (2024) Digital interventions to modify skin cancer risk behaviors in a national sample of young adults: randomized controlled trial. J Med Internet Res. 26:e55831. https://doi.org/10.2196/55831
[9] Mumtaz H, Riaz MH, Wajid H, Saqib M, Zeeshan MH, Khan SE, Chauhan YR, Sohail H, Vohra LI (2023) Current challenges and potential solutions to the use of digital health technologies in evidence generation: a narrative review. Front Digit Health. 5:1203945. https://doi.org/10.3389/fdgth.2023.1203945
[10] Busch F, Hoffmann L, Rueger C, van Dijk EHC, Kader R, Ortiz-Prado E, Makowski MR, Saba L, Hadamitzky M, Kather JN, Truhn D, Cuocolo R, Adams LC, Bressem KK (2025) Current applications and challenges in large language models for patient care: a systematic review. Commun Med (Lond). 5(1):26. https://doi.org/10.1038/s43856-024-00717-2
[11] Hager P, Jungmann F, Holland R, Bhagat K, Hubrecht I, Knauer M, Vielhauer J, Makowski M, Braren R, Kaissis G, Rueckert D (2024) Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat Med. 30(9):2613–2622. https://doi.org/10.1038/s41591-024-03097-1
[12] Schlicht IB, Zhao Z, Sayin B, Flek L, Rosso P (2025) Do LLMs provide consistent answers to health-related questions across languages? In: Hauff C, Macdonald C, Jannach D, Kazai G, Nardini FM, Pinelli F, Silvestri F, Tonellotto N (eds) Advances in Information Retrieval. ECIR 2025. Lecture Notes in Computer Science, vol 15574. Springer, Cham, pp 314–322. https://doi.org/10.1007/978-3-031-88714-7_30
[13] Sallam M, Al-Mahzoum K, Alshuaib O, Alhajri H, Alotaibi F, Alkhurainej D, Al-Balwah M, Barakat M, Egger J (2024) Language discrepancies in the performance of generative artificial intelligence models: an examination of infectious disease queries in English and Arabic. BMC Infect Dis. 24(1):799. https://doi.org/10.1186/s12879-024-09725-y
[14] Gallegos IO, Rossi RA, Barrow J, Tanjim MM, Kim S, Dernoncourt F, Yu T, Zhang R, Ahmed NK (2024) Bias and fairness in large language models: a survey. Comput Linguist Assoc Comput Linguist. 50:1097–1179. https://doi.org/10.1162/coli_a_00524
[15] Stall M, Germann JN, Orta M Jr, Winick N, Kaye EC (2024) Equitable communication for pediatric cancer patients and families who speak languages other than English. Pediatr Blood Cancer. 71(3):e30828. https://doi.org/10.1002/pbc.30828
[16] Ateşman E (1997) Türkçede okunabilirliğin ölçülmesi. Dil Dergisi. 58:71–74.
[17] Charnock D, Shepperd S, Needham G, Gann R (1999) DISCERN: an instrument for judging the quality of written consumer health information on treatment choices. J Epidemiol Community Health. 53(2):105–111. https://doi.org/10.1136/jech.53.2.105
[18] Okuhara T, Furukawa E, Okada H, Yokota R, Kiuchi T (2025) Readability of written information for patients across 30 years: a systematic review of systematic reviews. Patient Educ Couns 135:108656. https://doi.org/10.1016/j.pec.2025.108656
[19] Solak İ, Kozanhan B, Ay E (2021) Readability of Turkish websites containing COVID-19 information. Anatol JFM. 4:57–62. https://doi.org/10.5505/anatoljfm.2020.21939
[20] Alper Şahin A, Boz M, Keçeci T, Ünal A, Çıraklı A (2022) Readability and quality levels of websites that contain written information about anterior cruciate ligament injury: a survey of Turkish websites. Acta Orthop Traumatol Turc. 56:88–93. https://doi.org/10.5152/j.aott.2022.21142
[21] Armache M, Assi S, Wu R, Jain A, Lu J, Gordon L, Jacobs LM, Fundakowski CE, Rising KL, Leader AE, Fakhry C, Mady LJ (2024) Readability of patient education materials in head and neck cancer: a systematic review. JAMA Otolaryngol Head Neck Surg. 150(8):713–724. https://doi.org/10.1001/jamaoto.2024.1569
[22] Mirza FN, Wu E, Abdulrazeq HF, Connolly ID, Tang OY, Zogg CK, Williamson T, Galamaga PF, Roye GD, Sampath P, Telfeian AE, Qureshi AA, Groff MW, Shin JH, Asaad WF, Libby TJ, Gokaslan ZL, Kohane IS, Zou J, Ali R (2024) The literacy barrier in clinical trial consents: a retrospective analysis. EClinicalMedicine 75:102814. https://doi.org/10.1016/j.eclinm.2024.102814
[23] Goorman E, Mittal S, Choi JN (2026) Assessing readability of skin cancer screening resources: a comparison of online websites and ChatGPT responses. J Cancer Educ. 41(3):483–487. https://doi.org/10.1007/s13187-025-02683-2
[24] Bernal J, Mazo C (2022) Transparency of artificial intelligence in healthcare: insights from professionals in computing and healthcare worldwide. Appl Sci (Basel). 12(20):10228. https://doi.org/10.3390/app122010228
[25] Sanders JJ, Dubey M, Hall JA, Catzen HZ, Blanch-Hartigan D, Schwartz R (2021) What is empathy? Oncology patient perspectives on empathic clinician behaviors. Cancer. 127(22):4258–4265. https://doi.org/10.1002/cncr.33834
Downloads
Published
How to Cite
License
Copyright (c) 2026 Niyazi Çetin, Ahmet Uğur Atılan

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
The content of this journal is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.









