Full Papers
The impact of colchicine treatment on long-term complications of familial Mediterranean fever since its introduction in 1972: an evaluation of artificial intelligence performance
E. Ben-Chetrit1, P.D. Levin2, E. Ben-Chetrit3
- Infectious Diseases Unit, Shaare Zedek Medical Center, Jerusalem, Israel. elibc@szmc.org.il
- Department of Intensive Care, Shaare Zedek Medical Center, Jerusalem, Israel.
- Rheumatology Unit, Department of Medicine, Hadassah Hebrew University Medical Center, Jerusalem, Israel.
CER20050
Full Papers
Free to view
(click on article PDF icon to read the article)
PMID: 42814599 [PubMed]
Received: 17/04/2026
Accepted : 23/06/2026
In Press: 22/09/2026
Abstract
OBJECTIVES:
Familial Mediterranean fever (FMF) is a hereditary autoinflammatory disease characterised by recurrent febrile serositis and complicated in severe cases by AA amyloidosis, renal failure, infertility and adverse pregnancy outcomes. Since its introduction in 1972, colchicine has dramatically improved the prognosis of FMF.
METHODS:
In this study, we try to explore the ability of three artificial intelligence (AI)-based medical information platforms: Open Evidence, ChatGPT-4, and ChatGPT-5.2, to summarise the impact of colchicine treatment on FMF-related complications. Each platform was presented with two identical clinically focused questions addressing AA amyloidosis complications, fertility, and pregnancy outcomes. Only the initial responses were analysed. Outputs were compared descriptively according to response length, structure, clinical content, contextualisation, citation characteristics, and overall presentation. References were manually verified, and textual similarity was assessed using TF-IDF cosine similarity analysis.
RESULTS:
All three systems correctly conveyed the major clinical benefits of colchicine in FMF. However, substantial differences were observed in depth, contextualisation and citation quality. Open Evidence generated concise and focused responses with directly accessible references but limited contextual discussion. ChatGPT-4 produced more detailed and historically contextualised summaries, whereas ChatGPT-5.2 generated the most comprehensive responses, including discussion beyond the scope of the questions. Most references corresponded to real publications, although occasional bibliographic inaccuracies were identified, particularly among ChatGPT-generated citations. Cosine similarity analysis demonstrated overlap in core medical concepts despite marked differences in phrasing and organisation..
CONCLUSIONS:
AI-based medical information systems may provide rapid, clinically useful summaries of FMF and colchicine-associated outcomes. Nevertheless, human verification and expert interpretation remain essential because inaccuracies, hallucinations and citation inconsistencies persist. These tools should currently be regarded as complementary resources rather than replacements for conventional reviews or expert clinical judgment.



