Extending Epuskesmas With A Domain-Specific Large Language Model For Orthopaedic Documentation of Osteoarthritis And Osteoporosis: A Proof-Of-Concept Study Toward Indonesian Primary Health Care Strengthening

Authors

  • Ahmad Azmul A. Irfan Orthopaedic Department, Faculty of Medicine UIN Syarif Hidayatullah Jakarta, Indonesia
  • Nur Ahmad Khatim Division of Information Science, NAIST (Nara Institute of Science and Technology), Japan
  • Achmad Zaki Orthopaedic Department, Faculty of Medicine UIN Syarif Hidayatullah Jakarta, Indonesia
  • Bisatyo Mardjikoen Orthopaedic Department, Faculty of Medicine UIN Syarif Hidayatullah Jakarta, Indonesia
  • Mansur M. Arief Industrial and Systems Engineering, King Fahd University of Petroleum and Minerals, Saudi Arabia
  • Amril Nur Ismail General Practiotioner, uskesmas Banggae II, Indonesia
  • Ahmad Farhap Undergraduate Student, Faculty of Medicine UIN Syarif Hidayatullah Jakarta, Indonesia

DOI:

https://doi.org/10.51601/ijhp.v6i3.696

Abstract

Background. Clinical documentation is a major driver of workload in primary care, and Indonesia's mandated transition to electronic medical records has increased the recording burden on community health centres (Puskesmas). A recent proof-of-concept study showed that a browser-based pipeline combining automatic speech recognition (ASR) with large language model (LLM) summarisation can convert Bahasa Indonesia doctor-patient conversations into ePuskesmas fields for general primary care. Musculoskeletal complaints, particularly knee osteoarthritis and osteoporosis, are common, disabling and largely manageable at primary level, yet they require documentation elements that a general-purpose template does not explicitly capture. Objective. To develop a domain-specific orthopaedic extension of the ePuskesmas LLM documentation framework and to evaluate the clinical adequacy of its generated documentation through independent clinician review. Methods. Six scripted roleplay consultations covering knee osteoarthritis and osteoporosis-related conditions were recorded in Bahasa Indonesia, transcribed with a Whisper model, and summarised with an LLM using an orthopaedic-specific prompt mapped to ePuskesmas fields. Transcripts and structured outputs were assessed independently by three practising clinicians (two general practitioners working in Puskesmas and one orthopaedic specialist) against a seven-domain rubric scored 1-5, yielding 126 ratings. Inter-rater reliability was quantified using the intraclass correlation coefficient (ICC). Results. Across 126 ratings the overall mean was 3.29 of 5 (range 2-5), corresponding to the rubric anchor acceptable (usable after moderate editing). Agreement between reviewers was high: ICC(2,k) = 0.96, with all three reviewers assigning an identical score on 73.8% of items and agreeing to within one scale point on every item. Domain means were highest for appropriateness of referral recommendation (4.00) and appropriateness of Puskesmas-level management (3.83), and intermediate for red-flag identification (3.44). Scenario means ranged from 4.19 (suspected fragility fracture) to 2.57 (moderate knee osteoarthritis with obesity and gastritis). Qualitative review identified one materially unsafe analgesic recommendation, systematic loss of pertinent negative findings, incomplete transfer of examination detail, over-triage of one chronic case and conflation of fracture risk with established diagnosis. Conclusion. A domain-specific ePuskesmas LLM is technically feasible and produces documentation drafts of acceptable but not clinically final quality. Performance was adequate for management and referral recommendations, whereas fidelity to the source consultation was the limiting factor.

Downloads

Download data is not yet available.

References

[1] Khatim NA, Irfan AAA, Arief MM. Using Large Language Models for Real-time Transcription and Summarization of Doctor-patient Interactions into ePuskesmas in Indonesia: a proof-of-concept study. arXiv preprint arXiv:2409.17054. 2024. https://arxiv.org/abs/2409.17054

[2] Tierney AA, Gayre G, Hoberman B, Mattern B, Ballesca M, Kipnis P, et al. Ambient Artificial Intelligence Scribes to Alleviate the Burden of Clinical Documentation. NEJM Catal Innov Care Deliv. 2024;5(3):CAT.23.0404. https://doi.org/10.1056/CAT.23.0404

[3] Shah SJ, Devon-Sand A, Ma SP, Jeong Y, Crowell T, Smith M, et al. Ambient Artificial Intelligence Scribes: Physician Burnout and Perspectives on Usability and Documentation Burden. J Am Med Inform Assoc. 2025;32(2):375–380. https://doi.org/10.1093/jamia/ocae295

[4] Duggan MJ, Gervase J, Schoenbaum A, Hanson W, Howell JT, Sheinberg M, et al. Clinician Experiences with Ambient Scribe Technology to Assist with Documentation Burden and Efficiency. JAMA Netw Open. 2025;8(2):e2460637. https://doi.org/10.1001/jamanetworkopen.2024.60637

[5] Van Veen D, Van Uden C, Blankemeier L, Delbrouck JB, Aali A, Bluethgen C, et al. Adapted Large Language Models can Outperform Medical Experts in Clinical Text Summarization. Nat Med. 2024;30(4):1134–1142. https://doi.org/10.1038/s41591-024-02855-5

[6] Yang X, Chen A, PourNejatian N, Shin HC, Smith KE, Parisien C, et al. A Large Language Model for Electronic Health Records. NPJ Digit Med. 2022;5(1):194. https://doi.org/10.1038/s41746-022-00742-2

[7] Wornow M, Xu Y, Thapa R, Patel B, Steinberg E, Fleming S, et al. The Shaky Foundations of Large Language Models and Foundation Models for Electronic Health Records. NPJ Digit Med. 2023;6(1):135. https://doi.org/10.1038/s41746-023-00879-8

[8] Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language Models Encode Clinical Knowledge. Nature. 2023;620(7972):172–180. https://doi.org/10.1038/s41586-023-06291-2

[9] Ji Z, Lee N, Frieske R, Yu T, Su D, Xu Y, et al. Survey of Hallucination in Natural Language Generation. ACM Comput Surv. 2023;55(12):1–38. https://doi.org/10.1145/3571730

[10] Asgari E, Montaña-Brown N, Dubois M, Khalil S, Balloch J, Au Yeung J, et al. A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation. NPJ Digit Med. 2025;8(1):274. https://doi.org/10.1038/s41746-025-01670-7

[11] Koenecke A, Choi ASG, Mei KX, Schellmann H, Sloane M. Careless Whisper: Speech-to-text Hallucination Harms. In: Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT '24). New York: ACM; 2024. p. 1672–1681. https://doi.org/10.1145/3630106.3658996

[12] Radford A, Kim JW, Xu T, Brockman G, McLeavey C, Sutskever I. Robust Speech Recognition Via Large-scale Weak Supervision. arXiv preprint arXiv:2212.04356. 2022. https://arxiv.org/abs/2212.04356

[13] Ciecierski-Holmes T, Singh R, Axt M, Brenner S, Barteit S. Artificial Intelligence for Strengthening Healthcare Systems in low- and middle-income Countries: a Systematic Scoping Review. NPJ Digit Med. 2022;5(1):162. https://doi.org/10.1038/s41746-022-00700-y

[14] GBD 2021 Osteoarthritis Collaborators. Global, Regional, and National Burden of Osteoarthritis, 1990-2020 and projections to 2050: a systematic analysis for the Global Burden of Disease Study 2021. Lancet Rheumatol. 2023;5(9):e508–e522. https://doi.org/10.1016/S2665-9913(23)00163-7

[15] Kloppenburg M, Namane M, Cicuttini F. Osteoarthritis. Lancet. 2025;405(10472):71–85. https://doi.org/10.1016/S0140-6736(24)02322-5

[16] National Institute for Health and Care Excellence. Osteoarthritis in Over 16s: Diagnosis and Management. NICE guideline NG226. London: NICE; 2022. https://www.nice.org.uk/guidance/ng226

[17] Gregson CL, Armstrong DJ, Bowden J, Cooper C, Edwards J, Gittoes NJL, et al. UK Clinical Guideline for the Prevention and Treatment of Osteoporosis. Arch Osteoporos. 2022;17(1):80. https://doi.org/10.1007/s11657-022-01115-8

[18] Gregson CL, Armstrong DJ, Avgerinou C, Bowden J, Cooper C, Edwards J, et al. The 2024 UK Clinical Guideline for the Prevention and Treatment of Osteoporosis. Arch Osteoporos. 2025;20(1):119. https://doi.org/10.1007/s11657-025-01588-3

[19] Morin SN, Feldman S, Funnell L, Giangregorio L, Kim S, McDonald-Blumer H, et al. Clinical Practice Guideline for Management of Osteoporosis and Facture Prevention in Canada: 2023 update. CMAJ. 2023 ;195 (39) :E1333–E1348.https://doi.org/10.1503/cmaj.221647

[20] World Health Organization. Package of interventions for rehabilitation. Geneva: World Health Organization; 2023. https://www.who.int/teams/noncommunicable-diseases/sensory-functions-disability-and-rehabilitation/package-of-interventions-for-rehabilitation

[21] Kementerian Kesehatan Republik Indonesia. Peraturan Menteri Kesehatan Nomor 24 Tahun 2022 tentang Rekam Medis. Jakarta: Kemenkes RI; 2022. https://peraturan.bpk.go.id/Details/234978/permenkes-no-24-tahun-2022

[22] Aisyah DN, Setiawan AH, Lokopessy AF, Faradiba N, Setiaji S, Manikam L, Kozlakidis Z. The Information and Communication Technology Maturity Assessment at Primary Health Care Services Across 9 provinces in Indonesia: evaluation study. JMIR Med Inform. 2024;12:e55959. https://doi.org/10.2196/55959

[23] US Preventive Services Task Force. Screening for osteoporosis to prevent fractures: US Preventive Services Task Force Recommendation Statement. JAMA. 2025;333(6):498-508. https://doi.org/10.1001/jama.2024.27154

[24] Dinas Kesehatan Kabupaten Rembang. Monitoring morbiditas pasien Kabupaten Rembang berdasarkan data Platform SATUSEHAT Januari–Mei 2025. 2025 Jun 3. https://dinkes.rembangkab.go.id/monitoring-morbiditas-pasien-kabupaten-rembang-berdasarkan-data-platform-satusehat-januari-mei-2025/

[25] Tazakka RVM, Lestari D, Purwarianti A, Tanaya D, Azizah K, Sakti S. Indonesian-English code-switching Speech Recognition Using the Machine Speech Chain Based Semi-supervised Learning. In: Proceedings of the 3rd Annual Meeting of the Special Interest Group on Under-resourced Languages at LREC-COLING 2024. 2024. p. 143-148. https://aclanthology.org/2024.sigul-1.18/

[26] Adila A, Lestari D, Purwarianti A, Tanaya D, Azizah K, Sakti S. Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities. In: Proceedings of the 27th Conference of the Oriental COCOSDA. IEEE; 2024. https://doi.org/10.1109/O-COCOSDA64382.2024.10800336

[27] Afonja T, Olatunji T, Ogun S, Etori NA, Owodunni A, Yekini M. Performant ASR Models for Medical Entities in Accented Speech. In: Proceedings of Interspeech 2024. 2024. p. 2315–2319. https://doi.org/10.21437/Interspeech.2024-2261

[28] Mulyanto J, Wibowo Y, Kringos DS. Exploring General Practitioners' Perceptions about the Primary Care Gatekeeper Role in Indonesia. BMC Fam Pract. 2021;22:5. https://doi.org/10.1186/s12875-020-01365-w

[29] Yanthi B, Hendratini J, Sulistyo DH. Determinan Rujukan non Spesialistik dengan Kriteria TACC di FKTP Kabupaten Batang Hari tahun 2022. J Jaminan Kesehatan Nasional. 2023;3(1). https://doi.org/10.53756/jjkn.v3i1.63

Downloads

Published

2026-08-29

How to Cite

Ahmad Azmul A. Irfan, Nur Ahmad Khatim, Achmad Zaki, Bisatyo Mardjikoen, Mansur M. Arief, Amril Nur Ismail, & Ahmad Farhap. (2026). Extending Epuskesmas With A Domain-Specific Large Language Model For Orthopaedic Documentation of Osteoarthritis And Osteoporosis: A Proof-Of-Concept Study Toward Indonesian Primary Health Care Strengthening. International Journal of Health and Pharmaceutical (IJHP), 6(3), 838–851. https://doi.org/10.51601/ijhp.v6i3.696

Issue

Section

Articles