Labos: Dr. Google is always in, but too often wrong
Dr. Google is doing his best. As a health care resource, he tries very hard. He answers millions of queries in any language at any hour of the night or day. He is quick, efficient, never complains and will diagnose you if you type your symptoms into the chat box. The only downside to Dr. Google’s medical advice is that he’s often wrong.
Lest you think I am a Luddite, I have no inherent objections to technology. I doubt that the computers will one day become self-aware and wipe out humanity, although the fact that our current reality is increasingly starting to mirror the Terminator movies does give me pause on occasion. But I do believe that artificial intelligence has the potential to be greatly beneficial in medicine.
I’ve seen many physicians use it as a dictation tool to auto-generate medical notes. Admittedly, most physicians are not great note takers. The old joke about doctors’ handwriting being illegible is firmly grounded in reality. The promise of artificial intelligence is that it can draft a comprehensive, typed note based on the medical encounter.
But ChatGPT-4 didn’t do a great job when humans checked its work. One study fed transcripts of 14 simulated patient encounters into the AI program and asked it to generate a medical note. ChatGPT-4 made on average 23.6 errors per clinical case.
Although I make no claim to perfection, this seems like a lot to me. Mostly these were errors of omission, but 3.2 per cent were facts recorded incorrectly and 10.5 per cent were addition errors — that is, it hallucinated facts not in the interview transcripts, like lab results that were never performed. Say what you will about your physician, they probably don’t randomly make up stuff about your medical history.
But by far, the most common use of AI for medical purposes comes from people Googling their symptoms. In an era where accessing a doctor is difficult at the best of times, the temptation to pop your symptoms into a search bar and see what comes out is understandable. In mere seconds, the AI algorithm will give you a diagnosis and medical advice, which is an obvious plus if your doctor makes you wait in the waiting room for hours before being seen. Except here, too, AI performs surprisingly badly.
A recent study tested four large language models like ChatGPT and Gemini to see how well they could answer questions about menopause. Researchers drafted 35 questions (20 patient-level questions and 15 doctor-level questions) and fed them into the programs. For the patient-level questions, ChatGPT 3.5 (the free version) amazingly and counterintuitively did better than ChatGPT 4.0 (the paid version). They were accurate 70 and 60 per cent of the time, respectively. Gemini, Google’s AI, was accurate 30 per cent of the time.
For the doctor-level questions, ChatGPT 4.0 had the highest accuracy at 67 per cent, followed by ChatGPT 3.5 and OpenEvidence at 60 per cent each, with Gemini coming in last at 47 per cent. To add insult to injury, the answers were generally scored as “difficult” or “very difficult” to read on the Flesch Reading Ease Score. Which begs the question, if the answers generated by AI are inaccurate but also incomprehensible, do these two shortcomings cancel each other out?
I take no smug satisfaction in seeing how badly AI does when asked to perform in the medical arena. I don’t use AI to write my medical notes for me because I find I waste more time fixing its mistakes than typing the notes myself. I also wouldn’t rely on Dr. Google to make a diagnosis. He tries his best, but he does tend to make stuff up. I try not to be too critical of my colleagues. But Dr. Google tends to hallucinate too much for my liking.
Related Stories
AI News
Calgary uses AI to determine the 50 highest
43 minutes ago
AI News
Globus Medical Acquires Higgs Boson Health To Add AI
48 minutes ago
AI News
IITM Pravartak Announces Batch 03 of Advanced Certificate in Applied Artificial Intelligence & Deep Learning
48 minutes ago
AI News
Google rolls out Gemini AI feature to help readers analyse their e
1 hour ago
AI News
How AI and Smart Manufacturing Are Transforming China's Optical Fiber Giant
2 hours ago
AI News
Sunny Hostin of ‘The View’ bets on AI clones as Hollywood fights over its soul
2 hours ago
AI News
130,000 faces scanned in a week: The technology changing Australian policing
3 hours ago
AI News
Mom Who Travels For Work Builds AI Clone of Herself to Keep Her Teenage Son Company
3 hours ago