RADPEER scoring: six step QA checklist for radiology teams

RADPEER scoring for US radiologists and QA leads: six post discrepancy steps, key blind spots, and when to add teleradiology to protect your peer-learning

Published 31 August 2026

RADPEER Scoring: Six Step QA Checklist for US Radiology Teams

Hands marking flagged radiology films

RADPEER scoring is the American College of Radiology's standardized peer-review system for diagnostic imaging. It runs on a 3-point scale (1, 2, 3), with optional a/b modifiers that flag clinical significance, and it happens during routine interpretation rather than as a separate review step. Radiology groups use it to support Ongoing Professional Practice Evaluation (OPPE), meet accreditation requirements, and drive departmental peer learning.


TL;DR:

  • Most RADPEER discrepant scores involve cases flagged for arbitration, which should trigger internal review before submission to ensure data accuracy.
  • A low reported discrepancy rate may result from underreporting due to fear of punitive use, not necessarily indicating high-quality interpretation.
  • Shifting from merely collecting scores to actively integrating flagged cases into peer learning and conference discussions enhances practice improvement.
  • Internal arbitration of concerning scores and embedding case reviews into regular workflows reduces data noise and promotes meaningful feedback.
  • Effective RADPEER programs depend on a supportive culture that prioritizes learning over blame, with clear processes for case review and arbitration.

Table of Contents

What Radpeer Scoring Means: The Lexicon and Its Modifiers

A score of 1 means the current reader agrees with the prior interpretation. A score of 2 means there's a discrepancy, but one a reasonable radiologist might make given the same information. A score of 3 means the discrepancy shouldn't happen most of the time, given the imaging findings. Each score can carry an a or b modifier, where "a" indicates the discrepancy is not likely clinically significant and "b" indicates it is.

Before 2016, RADPEER used a 4-point scale that split discrepancies into two tiers of severity. The ACR RADPEER Committee's scoring white paper collapsed scores 3 and 4 into a single category, arguing the finer split invited endless committee debate without improving data quality. Archived 4-point data maps forward into the current 3-point framework for longitudinal reporting.

Score Meaning Modifier
1 Concurrence with prior read None
2 Discrepancy, understandable miss 2a (not significant) / 2b (significant)
3 Discrepancy, should be caught most of the time 3a (not significant) / 3b (significant)

How Radpeer Fits Into Daily Reads and OPPE

RADPEER doesn't require a separate review session. A peer-review event is triggered automatically whenever a radiologist reads a new study and has access to a relevant prior exam. At that moment, the reviewing radiologist forms an opinion about the earlier interpretation and, per the ACR's RADPEER instructions, may assign a score without leaving the normal interpretation workflow.

Scores get logged through eRADPEER, the secure web-based submission tool that adds minimal friction to the reporting process. In practice:

  • The reviewing radiologist assigns a score while dictating the current study.
  • Scores route through eRADPEER for the group's records.
  • Practice chairs and medical directors pull aggregated, practice-level reports to track trends over time.
  • The ACR reports that RADPEER has processed more than 30 million records, with participation from over 1,100 groups and roughly 18,000 physicians as of early 2017.

That scale is exactly why RADPEER doubles as an OPPE data source: it generates a continuous, low-effort stream of peer input tied directly to real cases rather than artificial audits.

The commercial interest belongs before the checklist, because it runs through the rest of this piece: AstraRad is a teleradiology group that operates its own peer-review program, so the failure modes below are ones we manage on our own reads. One in 20 signed reports is independently double-read blind by a second subspecialist, in a program modeled on ACR standards, with a major discrepancy rate under 0.3% reviewed at a monthly discrepancy meeting. Each client receives that as a monthly quality report, which is the part most peer-review programs never produce.

Internal Arbitration and Why It Matters for Accreditation

Not every discrepant score should go straight into the RADPEER database. Scores of 2b, 3a, or 3b need a second look before submission, since these carry clinical-significance flags that deserve scrutiny from someone with authority over the case, typically the department chair, medical director, or a designated QA committee member.

A workable internal workflow looks like this:

  1. The reviewing radiologist flags a 2b, 3a, or 3b score at the point of interpretation.
  2. The case routes to the chair, medical director, or QA committee for arbitration before submission.
  3. The committee documents its decision, including whether the score stands, changes, or gets reclassified.
  4. Only the arbitrated, confirmed score is submitted through eRADPEER.

Skipping this step inflates noise in the data and makes practice-level reports less useful for spotting real patterns. RADPEER participation itself supports both accreditation requirements and Maintenance of Certification (MOC) activities, but it's worth being clear about what the ACR does and doesn't require: the organization does not set numeric benchmarks for acceptable discrepancy rates. There's no target percentage a group must hit.

Pro Tip: Build arbitration into your PACS worklist rather than treating it as a separate administrative task. Groups that route flagged cases automatically see far fewer scores stall in limbo.

The Blind Spots: Subjectivity, Hindsight Bias, and Underreporting

RADPEER data looks more objective than it is. Independent analysis in Clinical Radiology documents real interobserver disagreement when the same cases get rescored by different reviewers, along with clear evidence of hindsight bias, where knowing the eventual diagnosis colors how harshly an earlier read gets judged.

Submission bias compounds the problem. Radiologists tend to under-report the discrepancies that matter most, often out of concern for how the score will be used against a colleague.

A few adjustments protect the value of the data:

  • Anonymize scores where the QA culture allows it, especially for lower-stakes discrepancies.
  • Sample cases representatively rather than only reviewing the ones that stand out.
  • Capture free-text feedback alongside the numeric score, since the number alone rarely explains the miss.
  • Never tie individual RADPEER scores to compensation or discipline decisions.

Pro Tip: If your group's serious-discrepancy rate looks suspiciously low, that's not necessarily good news. It may mean radiologists have learned the score gets used punitively, so they've stopped submitting the ones that count.

Radpeer QI and the Shift Toward Peer Learning

RADPEER QI represents the ACR's next iteration of the program, and it's built around a different premise: a numeric score by itself teaches nothing unless someone acts on it. The updated platform adds peer learning tracking, tools for building case conferences directly from flagged studies, and a modernized interface that replaces the older static scoring forms.

The practical shift for departments is to stop treating RADPEER as a data collection exercise and start treating it as a case pipeline. A 3b score isn't a filing task, it's raw material for a subspecialty learning conference. Groups that harvest these cases and track whether the resulting recommendations changed practice see measurable improvement in error patterns over time, something a spreadsheet of scores alone never delivers.

A Practical QA Checklist After a Discrepant Score

Turning a single flagged case into departmental learning takes a consistent sequence, not ad hoc handling. Use this as a starting template:

  1. Triage immediately for clinical urgency. If the discrepancy affects patient management now, notify the treating clinical team before anything else happens.
  2. Route 2b, 3a, and 3b scores to internal arbitration and document the committee's final determination.
  3. Confirm the arbitrated score before it's submitted through eRADPEER.
  4. Pull the case into a teaching file if it illustrates a recognizable pattern or a subtle finding worth revisiting.
  5. Schedule the case for the next peer-learning conference rather than letting it sit in the database.
  6. Track whether the conference discussion changed downstream reads, and revisit the pattern in six to twelve months.

Departments that follow all six steps consistently tend to see their peer-learning conferences drive actual practice change, not just attendance credit.

Why Does RADPEER Only Work If the Culture Supports It?

Why RADPEER only works if the culture supports it , overview diagram

The scoring mechanics are the easy part. What separates a RADPEER program that actually improves care from one that just generates a compliance report is whether radiologists trust that a flagged case leads to learning instead of blame. I've seen departments where a single change, moving arbitration off email and into a standing weekly huddle, doubled the number of cases that made it to a peer-learning conference instead of disappearing into a spreadsheet.

None of that shows up in the numeric score. It shows up in whether people submit the miss that actually matters. Look at how your own QA committee handles a 3b score today, and ask honestly whether the process invites learning or just invites silence.

Rafael

Where a Teleradiology Partner Fits Into Your QA Picture

RADPEER only works if radiologists have the bandwidth to review priors carefully and route flagged cases through arbitration without rushing. When backlog pressure squeezes that time, peer review is usually the first thing that suffers. AstraRad gives radiology groups and hospitals a way to offload volume, STAT, urgent, or routine, to US board-certified subspecialists without adding another portal, since every read integrates directly into your existing PACS.

AstraRad

AstraRad's own compliance track record, including a 99.4% SLA rate across STAT reads, reflects the same peer-review discipline this article describes: every study still passes through structured quality checks before it's signed. That matters most for groups managing overflow volume or backlog, where internal QA teams need capacity freed up to focus on arbitration and case-based learning rather than chasing report turnaround. If your department is evaluating how a reading partner might fit around your existing peer-review structure, see how our workflow integrates with your PACS and check current per-report pricing to see whether it fits your volume.

This article is general information, not a substitute for advice from a qualified doctor. Consult a qualified healthcare professional about your own circumstances before acting on anything here.

Sources

Put a radiologist's name on your next read.

Tell us your modalities and monthly volume. A complete per-report rate card, with turnaround tiers and SLA terms in writing, lands in your inbox within one business day.