Prime Time or Prototype? A Clinician’s Guide to Which Healthcare AI Actually Works
By Kevin L Watson Jr, MD, FAAP, Pediatric Gastroenterologist, Assistant Director of Clinical Informatics, Specialty Systems, Akron Children’s Hospital
Artificial Intelligence (AI) is everywhere now, even when we are not aware of its presence. Healthcare is no different. In the past few years, clinicians have been inundated with AI and claims of its amazing features. Healthcare organizations have seen an exponential expansion of AI algorithms in the marketplace with promises of supporting clinical decisions, improving diagnostic accuracy and allowing clinicians to focus more on patients, while improving time management. Some of these tools have delivered on this promise, while others may need more time to develop.
The current focus is not solely on how these AI models perform in validation studies, but how they perform in clinical workflows. Clinicians will often put a particular AI model in two categories, either it works persistently well, or it impedes workflow and slows things down. The real question is which AI is ready for patient care, and which still belongs in a pilot sandbox. Differentiating these categories of tools requires a different framework that clinicians may not have typically used.
When assessing AI tools, assigning a “Level” is a beneficial way to categorize the AI model. This framework provides 4 levels that give an idea of risk and potential for benefit.
Ultimately, a successful AI tool will not replace clinicians but allow them to spend time thinking again.
Level 1 – Observational
Level 1 tools are those that give some sort of observational insight. These can be thought of as tools that “run in the background”. Some examples may be those assessing risk for readmission, dashboards that show population analytics, or models that predict risk of patient no shows. These tools may be beneficial to administrators, but not something that may be directly used by clinicians in day-to-day life and rarely affect real-time decisions.
Level 2 – Advisory
Level 2 tools are advisory tools. These provide suggestions but don’t act without the clinician acting upon them. Many clinical decision support AI tools and best practice advisories fall into this category. AI models that provide assistance, such as differential diagnoses suggestions or order recommendations, are examples. These AI models can provide valuable recommendations to clinicians but are interruptive and can be ignored even if correct and beneficial.
Level 3 – Assistive
Level 3 tools are assistive and are usually embedded in the workflow. These tools can save time immediately and can reduce cognitive burden on clinicians. These tools are still voluntary and require the clinician to initially interact in order to be used. Real time examples of these tools are ambient documentation and chart summarization tools. These AI models are the first iteration of true generative AI clinical adoption.
Level 4 – Autonomous
Level 4 tools are autonomous in that they act independently of human interaction. These AI models present the greatest risk in that there is no human directly overseeing individual interactions. Trust should be present before tools such as these are implemented. Examples would be an AI model that reads radiology images without review or automated clinical triaging.
Determining the level of an AI tool serves the purpose of assessing risk and categorization of the tool. It is always important to know what type of human intervention is needed and having a level associated helps to give a clear identification of this. Governance of these AI models is crucial to have in place and having explicit level identification for clinicians utilizing these tools can help with understanding and adoption.
Healthcare AI relies on clinician adoption and there are various reasons why adoption may not occur or why it is not continued. First, it needs to be easy. Clinicians already have large workloads throughout the day; anything that causes additional verification may create work and decrease adoption. AI models also need to be seamless; interruptions create obstacles and if the AI is impeding workflow, adoption will be low. Clinicians, and patients, demand accuracy. If an AI model gives inaccurate outputs, trust is diminished not only in that model, but also in potential subsequent ones. Real-time benefits are also expected with clinical AI models. Timeliness is a large factor, if the perfect output occurs after a decision is made, there is little clinical benefit.
So, what AI tools are ready for prime time? There are three categories of tools that have demonstrated overall success in improving clinical workflows in real time.
1.) Ambient AI Scribes: These generative AI tools record clinical conversations and can generate an encounter note in real time. There is typically a very low learning curve with little to no changes in behavior for the clinician during the encounter. They allow the clinician to decrease typing but not the thought process. Human review is still imperative as errors can still occur, even with overall high accuracy.
2.) Chart Summarization: Many EHR vendors have such embedded tools available that will summarize previous encounters for a patient in inpatient and outpatient settings. This can greatly reduce clinician time with chart navigation and often comes with transparency and references to the patient’s chart. Integrity of the underlying data is very important with summarization tools, as these AI models rely on information that is input into the EHR. If data is incorrect, this can be further magnified in summaries that are taken as fact. This presents a risk for these tools and, again, clinician review is imperative.
3.) Communication AI: Tools that are in this category include those that generate patient instructions, patient communication portal responses and care coordination drafts. Many of the same underlying AI technologies power these models. They can help with adjusting language to fit specific reading levels, translate into different languages and reword communications to display greater empathy. These tools are lower risk as clinical reasoning is not included.
As stated previously, some tools, while encouraging, may not be ready for prime time. With tools that may not be ready, the major concerns are accuracy and risk. Tools that diagnose, not just create differentials, are promising but still in prototype phases. There is a question of liability without a clinician in the loop and oversight. AI prognostic and clinical prediction tools may also not quite be ready for full clinical deployment. It is still not consistent where some of these tools fit into the clinical workflow and if they can autonomously compare with seasoned clinician interpretations. There is much nuance in medicine and encompassing all clinical and socioeconomic factors may not yet be possible. Any AI tool that is fully automated should be approached with healthy skepticism and caution.
Considerations should be made by any healthcare organization or clinician regarding a new AI model. Does it remove or reduce work? Would the clinician notice if it disappeared? What education is needed surrounding it? What impact does it have on patient care? Adoption of these tools should be recommended but not necessarily mandated. If adoption requires a mandate, then the tool may not be quite ready to incorporate universally. If an AI tool works well, clinicians will use it. Education will be needed, not only regarding the tool itself, but also about slight adjustments in workflow that may make the tool work even better. At its basic form, clinical AI tools should make the clinician’s job easier. If it is functioning, then there will not be an interruption. Ultimately, a successful AI tool will not replace clinicians but allow them to spend time thinking again.

