Stressing Wisely
Check out other digital tools ↗

Learn · Part 3. Choosing · 5 min read

Reading the ratings

What Appropriate, May Be Appropriate and Rarely Appropriate mean, and what the tool's other labels and confidence grades tell you.

60-second take

  • Appropriate, May Be Appropriate and Rarely Appropriate come from the AUC panels. They describe a scenario, they are not rankings and they are not mandates.
  • Not applicable, No data and Not specified by the AUCs are different things. None of them is a recommendation against a test.
  • Verbatim, High, Moderate and Low confidence describe how directly the evidence supports a cell, not how appropriate the test is.
  • Open the tile's explanation to see the justification, the source passages and the patient factors.

Where the three categories come from

The scenario builder shows each test’s rating for the scenario you describe. Those ratings come from the appropriate use criteria (AUC). In an AUC, a writing group builds clinical scenarios, and an independent rating panel scores each test in each scenario in two rounds with a discussion in between, using a modified Delphi process. Each test ends up as Appropriate, May Be Appropriate or Rarely Appropriate . The panel works from numerical scores, but this tool does not show them: it shows the categories only.

Label in the toolWhat it means
AppropriateA reasonable option for this population because benefits generally outweigh risks. It is an effective option for individual care plans, though not always necessary, depending on physician judgment and patient preferences .
May Be AppropriateAt times an appropriate option, because evidence or agreement about the benefit-risk balance varies, or the population varies . A scenario where panelists disagreed is placed here, whatever the median score was .
Rarely AppropriateRarely an appropriate option because there is no clear benefit-to-risk advantage. Exceptions should have the clinical reasons for proceeding documented .

The authors also stress that the three groups are a simplification: appropriateness is most accurately viewed as a continuum, depending on the benefits and risks for the individual patient .

What ratings are not

They are not rankings. The 2023 AUC says its ratings are explicitly not competitive rankings , and the 2024 pre-operative AUC adds that within a category the numerical score is not a rank order, and that local expertise may favor one modality over another . They are not mandates either: the AUC is meant to support, not replace, clinician judgment . A rating applies to the scenario as described, and it presumes the patient has no contraindications .

The tool’s other labels

Some cells do not carry one of the three categories. The tool uses three other labels, and each means something different:

  • Not applicable: the source document deliberately did not rate that test for that scenario, most often because it is the test that was just performed. The AUC tables mark these cells as grayed out .
  • No data: no in-scope document rates the combination and no defensible inference exists. It is not a recommendation against the test.
  • Not specified by the AUCs: the existing AUCs do not cover this combination. The main example is a patient with known coronary disease who has already been tested during the current evaluation. The decision is individualized with the patient’s cardiologist, and reference ratings are shown for context only.

Example from the tool

Normal ET

After a normal exercise ECG, the exercise ECG tile has no rating to show. Open the scenario to see why.

  • Exercise ECGNot applicable
  • Exercise SPECTMay Be Appropriate
  • Pharmacologic SPECTMay Be Appropriate
  • Pharmacologic PETMay Be Appropriate
  • Exercise echoMay Be Appropriate
  • Dobutamine echoMay Be Appropriate
  • Stress CMRMay Be Appropriate
  • CAC scoreMay Be Appropriate
  • CCTAMay Be Appropriate
  • Invasive angiographyRarely Appropriate
  • No testMay Be Appropriate

See the reasons for each rating in the tool →

Verbatim, High, Moderate and Low confidence

Not every cell is printed in an AUC table. Where the documents rate a neighbouring scenario or a related test, the tool derives the cell and marks it Inferred, shows the rule it used and attaches a confidence grade. The grade describes how direct the evidence for that cell is. It is the tool’s own scale, not an AUC concept:

  • Verbatim: the document rates this exact scenario and modality.
  • High: one input differs from a verbatim scenario and the document’s neighbouring ratings are consistent on both sides, or the document itself states the inference (for example, SPECT and PET rated together).
  • Moderate: rests on a single analogous documented scenario, or combines an AUC rating with a protocol or guideline statement.
  • Low: crosses indication families or borrows from a similar modality, or rests mainly on clinical reasoning.
  • No data: no defensible analogue exists. It is not a recommendation against the test.

A grade says nothing about how appropriate the test is. A Low-confidence “Appropriate” and a Verbatim “Appropriate” are the same category, but the second is printed in a source table. One example of the documents’ own reasoning about grouping: the 2023 writing group judged that splitting a modality into its subtypes would not make a substantial difference to the ratings (for example, both SPECT and PET could be appropriate for recurrent symptoms after PCI), so the subtypes are not rated separately .

How to use the explanation on each tile

Open any result tile, then work down the “Why this rating” panel:

  1. Check the rating, the confidence grade and whether it is marked Inferred.
  2. Read the justification and the rules applied. They explain why the cell has this rating.
  3. Read the source passages, which show the document, table and page.
  4. Look for a document discordance note. It appears where two source documents rated the same scenario differently.
  5. Check the patient factors affecting this test. These are contraindications and cautions shown beside the rating. They do not change it.

These examples show how the other labels appear on the result cards.

Example from the tool

Incomplete revascularization — after exercise ECG — normal

Known coronary disease with incomplete revascularization, after a normal exercise ECG: every cell reads “Not specified by the AUCs”.

  • Exercise ECGNot specified by the AUCs
  • Exercise SPECTNot specified by the AUCs
  • Pharmacologic SPECTNot specified by the AUCs
  • Pharmacologic PETNot specified by the AUCs
  • Exercise echoNot specified by the AUCs
  • Dobutamine echoNot specified by the AUCs
  • Stress CMRNot specified by the AUCs
  • CAC scoreNot specified by the AUCs
  • CCTANot specified by the AUCs
  • Invasive angiographyNot specified by the AUCs
  • No testNot specified by the AUCs

See the reasons for each rating in the tool →

Example from the tool

Patient undergoing high-risk nonvascular surgery — No New or Worsening Symptoms AND a Functional Status <4 METs

Before high-risk nonvascular surgery, no known heart disease, functional capacity under 4 METs: the 2024 AUC does not rate deferral, so the No test cell reads “No data”.

  • Exercise ECGMay Be Appropriate
  • Exercise SPECTMay Be Appropriate
  • Pharmacologic SPECTMay Be Appropriate
  • Pharmacologic PETMay Be Appropriate
  • Exercise echoMay Be Appropriate
  • Dobutamine echoMay Be Appropriate
  • Stress CMRMay Be Appropriate
  • CAC scoreRarely Appropriate
  • CCTAMay Be Appropriate
  • Invasive angiographyRarely Appropriate
  • No testNo data

See the reasons for each rating in the tool →

Open the scenario builder → and click through the tiles for a scenario you know. The reasons behind each rating are shown there, not on this page.

Try it in the tool

Sources cited on this page

Click a tag for the full reference. All references