News: 1610440384

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Study: AI designed to detect diabetic eye disease blinks in the real world, makes more work for doctors

(2021/01/12)


A number of AI programs trained to detect diabetic eye damage struggle to perform consistently in the real world despite apparently excelling in clinical tests, say scientists in the US.

Academics led by the University of Washington School of Medicine tested seven algorithms from five companies: Eyenuk and Retina-AI Health in America, Airdoc in China, Retmaker of Portugal, and OphtAI in France. All of the models have gone through clinical studies, and are used – or can be used –to diagnose diabetic retinopathy, a complication of diabetes that damages blood vessels in the eye, leading to impaired vision or blindness.

[1]

The research team said it found at least some of the software packages wanting during its own testing, and this month [2]published its findings in the Diabetes Care journal.

“It’s alarming that some of these algorithms are not performing consistently since they are being used somewhere in the world," [3]said lead researcher Aaron Lee, assistant professor of ophthalmology at the university.

Top doctors slam Google for not backing up incredible claims of super-human cancer-spotting AI [4]READ MORE

The team tested the code by showing it a dataset of 311,604 photos from 23,724 patients at hospitals in Seattle and Atlanta from 2006 to 2018, and found some of the software's diagnosis of these patients was sub-par. When the algorithms' decisions were compared to a real physician, the team said three performed reasonably well, only one of them was as good as a human expert, and the rest were worse.

The AI models tended to over-predict if a patient had the disease or not, Lee told The Register . Although it’s better to be safe than sorry, it meant the systems would more often than not flag up patients for examinations by professional eye doctors. Instead of reducing the workload for ophthalmologists, by filtering out those without the disease, the software would increase it.

“The study design prevents us from disclosing which company supplied which algorithm unfortunately," Lee added. "It is my understanding that all of these algorithms are in clinical use somewhere in the world however."

The programs did better with imagery from Atlanta, we're told, a sign that performance depends heavily on the quality of the data. “We believe one of the reasons for the discrepancy in performance was that Atlanta has a more stringent protocol for image quality at the time of screening," Lee told us. "This suggests that AI models may be more sensitive to image quality issues than human beings."

The academics suggested medical algorithms should be evaluated on larger real-world datasets before being validated for public use. “AI algorithms are not all created equal and they can, but not always, recapitulate biases in datasets,” Lee warned.

So, are they safe for use?

Airdoc declined to comment on the study, and Retmaker did not respond to El Reg 's questions.

Stephen Odaibo, CEO and founder of Retina-AI Health, told us in a statement he thought the researchers' conclusions were not supported by the study.

"First, the study was a retrospective study based on heterogeneous unstructured data from the [veteran patients]," he said. "The data included pictures that were not of the retina, for example, pictures of people's faces, eyelids, or even their driver's licenses. The algorithms were made to first sort through to identify which of the images were of retinas, and then which were of right eyes or left eyes; after which it was then to determine the disease stage."

"Furthermore, it was not known from what types of cameras the images had been taken over the years. This is a completely different scenario from the use case for which these AI algorithms were developed and subsequently clinically validated in prospective clinical trials for FDA approval," Odaibo continued, referring to America's medical watchdog, the Food and Drug Administration.

"The indications of use are in primary care settings and with a trained camera operator who selects two specific images per eye from a known camera device on which the algorithm has been specifically validated prospectively. The above discrepancy is the source of the big logical gap between the study and its claims. To make an evidence-based recommendation to the FDA one would need to design a study that reflects the indications of use and intended use of the medical device."

Frank Cheng, president and chief customer office of the other US-based company in the study, Eyenuk, agreed that the experiments carried out by the academics didn't quite mirror how a system would be tested after FDA approval: "Our view is that systems that have FDA clearance are already going through more rigorous prospective clinical trial validation than the University of Washington study, additional testing is not necessary, so long as photographers and imaging protocol training takes place ... In real world clinical use, FDA-cleared systems such as Eyenuk's are integrated with the camera, and photographers are trained on the imaging protocol to be used."

Cheng said he believed "Eyenuk's EyeArt AI system is very much ready for prime time and is available for clinical use," and said he thought the "study analysis was well conducted in general."

OphtAI's CTO Bruno Lay told The Register that the research group's conclusions were fair. Lay claimed OphtAI's algorithms were ranked as the best and second-best of the seven algorithms tested, and that the technology from three out of the five companies trialed probably isn't yet good enough to be used in the real-world.

"The experiments were very challenging," he said. "We had no idea of the quality of the images used in the test. We were able to process the whole dataset in just three days, and our system is already available for use in hospitals in France."

[5]

Diabetic retinopathy is a widely studied area in medical AI research. Several Alphabet subsidiaries, including Google, Verily, and DeepMind have demonstrated how machine-learning software can automatically analyze retinal scans. ®

Get our [6]Tech Resources



[1] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_emergent_tech/artificial_intelligence&sz=300x250%7C300x252%7C300x600&tile=3&c=33X-2BRQao00Ep7tVW8hET0gAAAUo&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dtop%26test%3D0

[2] https://care.diabetesjournals.org/content/early/2021/01/01/dc20-1877

[3] https://newsroom.uw.edu/print/10591

[4] https://www.theregister.com/2020/10/16/google_ai_research/

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_emergent_tech/artificial_intelligence&sz=300x100%7C300x250%7C300x251&tile=4&c=44X-2BRQao00Ep7tVW8hET0gAAAUo&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[6] https://whitepapers.theregister.com/

The AIs have it (or not)

Chris G

What I get from the makers responses, is that for their systems to function well, depends on humans trained to use the correct type of camera, to also be trained well enough to submit the right kind of photo so the AI is not really achieving the aim of simplifying, speeding up and making the process more efficient.

I would argue that new AI diagnostic aids should have comparable testing and approval as that applied to drugs

Re: The AIs have it (or not)

Anonymous Coward

Spot on - clinical tools with false positives have a massive detrimental effect on patients.

30 years ago my PhD investigated the effects of noise / distortions in 'AI', or pattern recognition as it was known then. The reason being lab systems had a very good accuracy (as they'd been tuned to high heaven on the consistent data sets) which wasn't matched in the real world.

A practical issue is labs / dev / research tends to go for the best sensors, cameras, mic's whatnot, real-world implementations go for the cheapest with differing response curves and SNR.

quartzz

"It's better to be safe than sorry" - this phrase is being used to say "false positives" are ok. no they aren't. false positives can be just as negative as missing the real thing.

it's like email security*. keeping the right people out is just as harmful as letting the wrong people in. no.....false positives aren't good enough. great for capitalism, but not for that diagnosis.

*I can't let this one slide without mentioning instagram. who are currently/last few weeks locking loads of peoples accounts under the guise of "security", when it's actually under the guise of "your aren't uploading enough profit for us, so we're locking your account out". you know. pandemic. more people online. facebook wants more profit. (this is why your insta stats have gone down - cos other ppl have got locked out)

Whitter

Anything that would increase the rate at which diabetic disease is flagged up from an eye exam would inevitably increase the workload of human reviewers. Is this result really a bad one? Difficult to tell from the article, which implies poor performance but doesn't show the data. e.g. How many true detections were made that would have been (or were) missed by eye docs? Maybe that is in the original research, just not the El Reg snippet.

I guess it ultimately depends on just how many false detections were being made: alarm fatigue is a well-known problem in medical institutions.

Image Quality

Anonymous Coward

I've had annual retinal screening checks for about 10 years now, arranged via the local GP practice, but performed by a contractor. Admittedly I am a difficult subject as invariably I blink when the flash fires and the operator will need to repeat the process several times, but sometimes only a single attempt.

I always ask to see the images before leaving.

Once during a regular eye test, a retinal screening was done. The optician took many images until she was satisfied of the quality of the images - many more than I have experienced in the past in the retinal screening service. When I saw the final images, it was a revelation - there was a big difference compared to what I had seen before - the blood vessels were clearly much better resolved and sharper. Back on the regular screening, the image quality was never as good,and in some instances, only a single image on each eye was performed.

So, it is not at all surprising to hear “We believe one of the reasons for the discrepancy in performance was that Atlanta has a more stringent protocol for image quality at the time of screening,"

I no longer attend the community retinal screening, as complications (not picked up by that service) meant I am now under the care of the hospital directly.

Re: Image Quality

Loyal Commenter

This is actually a pretty good example of the creeping privatisation of the NHS, and the negative effects of doing so. The motivations of private companies (in this case, the "contractor") and medical professionals are quite different.

The contractor will almost certainly have been chosen either as the lowest quote, or as the only quote (in which case, the outsourcing has probably been put in place specifically to give work to that organisation, not because of clinical need).

Private companies are motivated by profit, so the person doing the scans will be the cheapest available, and thus will have been trained to the minimum standard required to do the job. Spending any more money on someone more qualified, or on more training will be seen as wasted cash.

Medical professionals, however, are motivated by a desire to try and help people. They will want to do the best job possible, to save them having to repeat the job later, and to make sure they don't miss anything and potentially get sued. The private individual doing the retinography doesn't have this worry, as it is their company that gets sued, not them personally.

In theory, outsourcing such tasks gains an economy of scale (e.g. multiple NHS trusts using the same outsourcer). In practice, it doesn't save costs, but instead ends up with corners being cut to maximise the bottom line. In the end, this make the people who put that outsourcing in place very rich, at the expense of everyone else.

AI trial sucessful

Anonymous Coward

AI software successfully interchanges shortage of trained ophthalmologists with shortage of trained camera and AI software operators.

Things past redress and now with me past care.
-- William Shakespeare, "Richard II"