When evaluating an artificial intelligence system in healthcare, attention often focuses on model performance. Accuracy, sensitivity, specificity, and calibration are fundamental measures for understanding how the system performs with respect to a given task.

But between the prediction produced by the model and the clinical decision, there is a space that metrics alone do not describe.

The professional receives an output, interprets it together with the other available information, decides how much to trust it, and ultimately makes a decision. AI therefore does not operate in isolation: it enters an existing workflow and becomes a component of a broader decision-making process. [1,2]

For this reason, an important question is not only how well the model performs, but how its output is used in the real-world context. [1]

From the Model to the Decision

A decision-support system can be represented, in simplified form, through the following sequence: patient data โ†’ AI system โ†’ output โ†’ professional โ†’ decision โ†’ action

Technical validation mainly concerns the first part of the chain: the relationship between data, prediction, and observed outcome. However, when the system enters practice, the subsequent steps also become relevant. [1]

  • Who receives the output?
  • At what point in time?
  • What other information is available?
  • Is the output presented as a probability, risk category, recommendation, or alert?
  • Does the professional correctly understand what the system is communicating?

These questions do not replace technical evaluation, but describe a different level of the same problem. A model may produce a correct prediction, but that prediction may arrive too late to influence the decision. It may be displayed to a professional other than the one intended, or interpreted as a diagnosis when it was designed simply as decision support.

The model’s performance may therefore remain unchanged while the operational meaning of its output changes.

The Same Model Can Become a Different System

Imagine that two hospitals use the same model to estimate a patient’s risk of deterioration.

  • Hospital A: The score appears in the electronic health record as additional information. The physician can consult it together with test results, symptoms, and their own clinical assessment.
  • Hospital B: The same score automatically generates a visible alert and activates an escalation pathway.

The model may be identical. Even the value produced for the same patient may be identical. But the use is not. In the first case, AI provides one piece of information among many; in the second, the output directly enters an organizational process and may change priorities, timing, and responsibilities.

This means that the unit of analysis cannot always be limited to the model: the system must also be considered in its concrete use. [1,2]

The question therefore becomes more specific: not simply โ€œhas this model been validated?โ€, but:

โ€œDoes the available evidence support the way its output is being used here?โ€

Humanโ€“AI Interaction Is Not Only a Matter of Interface

Talking about humanโ€“AI interaction may immediately bring usability to mind: colors, buttons, screen organization, or ease of use. These are important aspects, but the issue is broader. Interaction also concerns the weight that the output carries in the decision.

A professional may rely excessively on a recommendation because it is presented with a level of authority that the system should not have. In other cases, the opposite may happen: frequent or irrelevant alerts may gradually be ignored. [1,2]

There is also the problem of disagreement. What happens when the professional believes that the AI output is not appropriate for that patient? The technical possibility of ignoring or modifying a recommendation does not, by itself, answer the question. It is necessary to understand what role the system plays in the decision-making process and whether professional autonomy can genuinely still be exercised. [2]

Human oversight, therefore, does not simply mean having a person in front of the screen. A person may be formally present while, at the same time, having very little real opportunity to challenge the system’s output.

An Alert Without Responsibility Is Not a Control

The interaction between AI and the professional also raises an organizational question: who is expected to act?

Suppose a system correctly identifies a high-risk patient and generates an alert. From the model’s perspective, the result may be correct. But if it is unclear who should receive the alert, how quickly it should be assessed, or what action is expected, the correctness of the prediction does not guarantee the effectiveness of the process.

The problem is no longer only algorithmic. It is a workflow problem.

The same applies when several professionals interact with the system:

  • Who is responsible for assessing the output?
  • Who can ignore it?
  • Who verifies that an alert has been acknowledged?

Technology can produce information, but the workflow determines how that information becomesโ€”or fails to becomeโ€”action. [1,2]

Failure Is Also Part of the Environment of Use

Another way to understand the role of context is to ask what happens when the system does not work as expected. The service may be temporarily unavailable, some data may be missing, an integration with the information system may fail, or the output may arrive after the decision has already been made.

These situations do not necessarily demonstrate a defect in the model, but they show that the model depends on an execution environment. For this reason, it is useful to consider fallback conditions as well: how does the process continue when the AI is unavailable?

The answer cannot be universal, because it depends on the type of system, the decision being supported, and the clinical context. The important point is that if the safe functioning of a process depends on the availability of an AI system, then managing its unavailability is also part of the real conditions of its use. [1,2]

The Workflow Changes the Meaning of the Evidence

This leads to an important consequence for validation. Evidence produced about a model does not automatically describe every possible way in which that model can be used.

High sensitivity demonstrates something about the model’s predictive behavior in the population and under the conditions analyzed. By itself, it does not demonstrate that a particular alert is presented at the right time, that it is interpreted correctly, or that the workflow clearly assigns responsibility for action.

These are different questions and require different evidence. This does not mean that every organizational difference makes previous validation unusable. It means that we should avoid extending a conclusion beyond what the evidence can actually support.

The same model can be incorporated into different workflows, used by different professionals, and connected to decisions with different consequences. The validity of its use therefore cannot be inferred exclusively from the technical validity of the model. [1]

From Model Performance to the System in Use

Clinical AI is often described as if the sequence were simple: input โ†’ model โ†’ output

In practice, the sequence is longer: input โ†’ model โ†’ output โ†’ interpretation โ†’ decision โ†’ action

It is precisely in the second part that the professional, the workflow, and the organization come into play. For this reason, the analysis of humanโ€“AI interaction should not be considered an additional check performed after technical validation, but a necessary part of understanding what the system is actually doing in the context in which it is used. [1,2]

The question is not whether the human is formally โ€œin the loop,โ€ but what role the human actually plays in the process:

  • Do they see the output before or after forming their own judgment?
  • Do they understand the limitations of the information they receive?
  • Can they deviate from it?
  • Is disagreement managed?
  • Is responsibility for the subsequent action clearly assigned?
  • Is there an operational procedure for when the system is unavailable?

These questions are simple to formulate, but the answers can completely change the meaning of the same output.

Conclusion

A good model does not automatically result in a good decision-making process.

When an AI system enters a clinical environment, its performance remains fundamental, but it becomes one component of a broader system consisting of data, technology, professionals, procedures, and responsibilities. [1,2]

For this reason, the question โ€œdoes the model work?โ€ should always be accompanied by a second question:

โ€œDoes its output support the intended decision, for the intended professional, and under the conditions in which it is actually used?โ€

It is at this point that humanโ€“AI interaction becomes a question of validity rather than merely a question of usability.

Evaluating AI within the workflow therefore means making visible what happens between prediction and action: who interprets the output, how much weight they give it, how disagreement is managed, who assumes responsibility, and what happens when the expected conditions are not present.

Because, in the real-world context, it is not the model alone that makes a decision. It is the system in use that produces consequences.

References

  1. Vasey, B., Nagendran, M., Campbell, B., et al. (2022). “Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI.” Nature Medicine, 28, 924-933. DOI: 10.1038/s41591-022-01772-9.
  2. World Health Organization (2021). “Ethics and governance of artificial intelligence for health: WHO guidance.” Geneva: World Health Organization. ISBN 978-92-4-002920-0.


Pubblicato

in

da

Commenti

Lascia un commento

Il tuo indirizzo email non sarร  pubblicato. I campi obbligatori sono contrassegnati *

Share with