Skip to main content

ACR Digital Mammography Phantom QC

By Jiali Wang, PhD, DABR
July 23, 2024 16 min read

The ACR Digital Mammography Phantom test is a routine quality-control check in which simulated fibers, speck groups, and masses are imaged and scored to confirm the system can display the small, low-contrast structures that matter for cancer detection. It is scored by technologists on a routine schedule and by the medical physicist at least annually, alongside contrast-to-noise ratio, signal-to-noise ratio, and dose. Knowing how the test works — and where it is subjective — is what makes it useful.12

Introduction

Mammography lives or dies on the ability to see fine detail: microcalcifications a few hundred micrometers across, subtle masses, and thin spiculations. A digital mammography system that drifts — a failing detector element map, a processing change, a dose creep — can lose that detail before anyone notices on clinical images. The phantom image test exists to catch that drift with a known, repeatable object set.1

The American College of Radiology's Digital Mammography (DM) Quality Control program centers on the ACR DM Phantom, a test object containing simulated fibers, speck groups, and masses. Technologists image and score it on the facility's routine QC schedule, and the medical physicist evaluates it during the annual survey together with quantitative metrics. The test is deceptively simple to run and surprisingly easy to run badly.12

This article explains the ACR DM Phantom, how objects are scored, the pass criteria, why those criteria differ from the older screen-film accreditation phantom, the quantitative metrics that accompany the visual score, and the observer variability that makes consistent technique essential. DRPS delivers this testing through mammography physics and MQSA surveys and accreditation support across Florida, Maryland, Virginia, Washington DC, California, and Nevada.

Topic Explanation

What the phantom simulates

The ACR DM Phantom is designed to mimic the radiographic appearance of a compressed breast and to embed objects that represent the three findings mammography must detect: fibers (fibrous structures), speck groups (microcalcifications), and masses (tumor-like densities). The reader scores how many of each object type are visible, starting from the largest and most conspicuous and working toward the smallest and faintest.23

The three object classes map onto real diagnostic tasks:

  • Fibers simulate fibrous and spiculated structures.
  • Speck groups simulate clustered microcalcifications — often the earliest sign of malignancy.
  • Masses simulate the low-contrast, rounded densities of tumors.

Because these objects are manufactured to graded sizes and contrasts, the point at which a reader can no longer see them is a sensitive indicator of system performance.

Two phantoms, two eras

There is real confusion between the legacy phantom and the current one, and getting them mixed up leads to citing the wrong pass criteria.

The legacy screen-film ACR mammographic accreditation phantom (models such as the RMI 156 and Gammex 156D) simulates a 4.2 cm compressed breast of roughly 50% adipose and 50% glandular composition. It contains 6 fibers, 5 speck groups, and 5 masses, and the historic accreditation criterion was to visualize at least 4 fibers, 3 speck groups, and 3 masses — the familiar "4-3-3."34

The newer ACR DM Phantom, used in the ACR Digital Mammography QC program and its 2018 manual, contains 6 fibers, 6 speck groups, and 6 masses, and its pass criterion is 2 fibers, 3 speck groups, and 2 masses with no clinically significant artifacts.15 The lower counts do not mean a lower standard — the objects, their contrasts, and the scoring rules are different, so the two thresholds are not interchangeable.

Feature Legacy screen-film ACR accreditation phantom ACR DM Phantom
Typical models RMI 156 / Gammex 156D ACR DM Phantom
Objects present 6 fibers, 5 speck groups, 5 masses 6 fibers, 6 speck groups, 6 masses
Passing score ≥4 fibers, ≥3 speck groups, ≥3 masses ≥2 fibers, ≥3 speck groups, ≥2 masses
Primary metric alongside score Optical density (film) CNR, SNR, average glandular dose
Program MQSA accreditation (screen-film era) ACR Digital Mammography QC Manual (2018)

Using the legacy 4-3-3 threshold to judge an ACR DM Phantom image — or vice versa — is one of the most common phantom-scoring errors. Always match the criterion to the phantom in the tunnel.13

Key Technical Principles

How objects are scored

Each object is scored on a partial-credit scale — commonly 0.0, 0.5, or 1.0 — reflecting whether it is fully seen, partially seen, or not seen. Fibers are judged by the length of the visible segment, speck groups by how many individual specks in a group are visible, and masses by the completeness of their circular border and density.6 Scoring proceeds from the largest object of each type down to the smallest, and the count stops at the last clearly visible object.

The subtlety is that the smallest fibers and faintest masses sit right at the threshold of visibility. Small changes in windowing, ambient light, monitor calibration, or reader diligence can move an object across the seen/not-seen line, which is exactly why the test is standardized and why the same reader and viewing conditions should be used over time.

The quantitative companions: SNR and CNR

The visual score is necessary but coarse. To catch drift earlier, the medical physicist measures signal-to-noise ratio (SNR) and contrast-to-noise ratio (CNR) from defined regions in the phantom image. SNR characterizes the detector's raw signal relative to noise; CNR characterizes how well a low-contrast object stands out from background.

CNR is defined from mean pixel values and background noise:

where is the mean pixel value in a region containing a test object (or a contrast disk), is the mean in an adjacent background region, and is the standard deviation of the background.

As an illustration, suppose a contrast region reads a mean of 1,050 and the adjacent background reads a mean of 1,000 with a standard deviation of 10:

The absolute number depends on the phantom, technique, and detector, so each system is tracked against its own established baseline and the action limits in the ACR DM QC Manual rather than a universal target. A CNR trending downward over successive surveys flags a real change — a detector, technique, or processing problem — often before any phantom objects are actually lost.1 SNR and CNR turn a pass/fail visual test into a trend you can act on early.

Why average glandular dose belongs in the same test

The physicist evaluates the phantom image quality together with average glandular dose (AGD) because image quality and dose are a trade-off. A system can "pass" the phantom by simply using more dose, so the phantom score is only meaningful when read next to the dose required to produce it. The goal is adequate object visibility at an appropriately low glandular dose, consistent with MQSA dose limits and ALARA. For a deeper treatment of the dose side, see our discussion of mean glandular dose in mammography.

Clinical Impact

The test protects the detection task

Every object class in the phantom stands in for a real clinical finding. When speck-group visibility drops, the system's ability to show early microcalcifications is at risk. When masses or fibers fade, subtle soft-tissue findings are threatened. A phantom failure is therefore not a bureaucratic event — it is a warning that the modality's core purpose, finding small cancers, may be degraded.2

Observer variability is real and measurable

Phantom scoring is a human perception task, and humans disagree, especially near the visibility threshold. Studies comparing multiple physicists scoring the same ACR DM Phantom images report meaningful inter-observer variability, with the finest objects — particularly fibers below roughly 0.6 mm — scored least consistently.56 One multi-vendor study of ACR DM Phantom images across many observers quantified this spread and compared it against automated software.7

This variability has practical consequences. If the same borderline image is a "pass" for one reader and a "fail" for another, the facility's QC record depends partly on who happened to score it. That is why consistent viewing conditions, a fixed reader when possible, and clear scoring rules matter — and why automated, software-based scoring is an active area of development.

Automated scoring and its promise

A growing body of work applies template matching, model observers, and convolutional neural networks to score mammography phantom images automatically. Systematic review of these methods finds that computerized scoring is generally more consistent than inter-observer human scoring, particularly for the higher-contrast objects, and that template matching is reliable for well-defined structures.8 Neural-network approaches have reported high agreement with expert scoring, with the largest disagreements at the lowest doses and smallest masses — precisely the borderline objects where humans also struggle.910 Automation does not replace the physicist's judgment, but it can reduce the reader-to-reader noise in the QC record.

Practical Optimization Tips

Running the phantom test well is mostly about controlling everything except the system under test.

1. Standardize the setup

Image the phantom in the same position, with the same technique (or the clinically relevant automatic-exposure-control mode), the same compression, and any required attenuator, every time. A change in setup masquerades as a change in system performance.

2. Fix the viewing conditions

Score on a calibrated mammography-grade display, in consistent low ambient light, at a consistent window/level and magnification. The single biggest source of spurious score changes is inconsistent viewing.

3. Score from largest to smallest and record partial credit

Follow the manual's scoring order and rules, count the last clearly visible object, and record the partial-credit values rather than a bare integer. The detail supports trending and dispute resolution.

4. Read the artifacts, not just the objects

The pass criterion requires no clinically significant artifacts. Grid lines, detector nonuniformities, ghosting, and processing artifacts can fail an otherwise well-scored image and, more importantly, can hide or mimic pathology. Evaluate the whole field, not only the object columns. For the broader modality QC context, see mammography quality control and MQSA.

5. Trend SNR and CNR, do not just pass/fail

Keep the quantitative metrics over time. A phantom that still "passes" but whose CNR has fallen 20% from baseline is telling you something is changing. Trending catches it early.

Common pitfalls to avoid

  • Applying the wrong pass criteria. The legacy 4-3-3 and the ACR DM Phantom 2-3-2 are not interchangeable.
  • Changing readers or viewing conditions between surveys. This injects variability that looks like system drift.
  • Ignoring the dose side. A rising AGD can "buy" a passing score while masking a real problem.
  • Scoring only the objects. Artifacts can fail an image and threaten clinical reads.
  • Not baselining. Without an established SNR/CNR baseline, you cannot recognize meaningful drift.

Regulatory Considerations

Mammography quality control in the United States is governed by the Mammography Quality Standards Act (MQSA) and its implementing regulation, 21 CFR 900.12, with image-quality QC performed under the facility's chosen program — commonly the ACR Digital Mammography QC Manual or the manufacturer's QC manual.111

Key frameworks:

  • MQSA — 21 CFR 900.12. Establishes quality standards for mammography facilities, including required quality-control testing and the annual medical physicist survey. Phantom image quality is part of the required QC.11
  • FDA MQSA Final Rule (2023). The amendments to the MQSA regulations, with enforcement beginning September 10, 2024, added breast-density reporting to patients and providers, communication timelines, and medical-outcomes-audit metrics. They did not remove phantom image-quality QC; facilities should continue to follow their current QC manual.12
  • ACR Digital Mammography QC Manual (2018). Defines the ACR DM Phantom test, scoring, pass criteria, and the physicist's evaluation of CNR, SNR, and average glandular dose as an accepted QC program under MQSA.1
  • ACR–AAPM–SIIM practice parameter on digital mammography image quality. Provides the professional framework for determinants of image quality in digital mammography.13

A facility must operate under an accepted QC program, keep phantom QC records, and correct failures before continued clinical use. The medical physicist's annual survey documents phantom score, CNR, SNR, AGD, and artifact evaluation. DRPS provides these surveys and helps facilities maintain accreditation through mammography physics and MQSA and accreditation support.

Frequently Asked Questions (FAQs)

What is the ACR Digital Mammography Phantom test?

It is a quality-control test in which a phantom containing simulated fibers, speck groups, and masses is imaged and scored to confirm that a digital mammography system can display small, low-contrast structures. Technologists perform it routinely and the medical physicist evaluates it at least annually along with contrast-to-noise ratio and dose.

What are the pass criteria for the ACR DM Phantom?

Per the ACR Digital Mammography QC Manual, the ACR DM Phantom passes when at least 2 fibers, 3 speck groups, and 2 masses are visible and there are no clinically significant artifacts. Objects are scored from largest to smallest using a partial-credit scale.

How is the ACR DM Phantom different from the old accreditation phantom?

The legacy screen-film ACR mammographic accreditation phantom contains 6 fibers, 5 speck groups, and 5 masses and used a 4-3-3 accreditation criterion. The newer ACR DM Phantom contains 6 fibers, 6 speck groups, and 6 masses and uses the 2-3-2 pass criterion defined in the ACR Digital Mammography QC Manual.

Who scores the phantom, the technologist or the physicist?

Both. The technologist images and scores the phantom on the facility's routine QC schedule to catch drift, and the medical physicist evaluates the phantom image, contrast-to-noise ratio, signal-to-noise ratio, and average glandular dose during the annual survey.

Why do two people score the same phantom differently?

Phantom scoring is a human detection task, so the smallest, lowest-contrast objects sit near the threshold of visibility and different readers count them differently. Studies show meaningful inter-observer variability, especially for the finest fibers, which is why consistent technique and, increasingly, automated scoring are valuable.

What does the CNR in a phantom image tell you?

Contrast-to-noise ratio measures how well a low-contrast object stands out from background noise. A falling CNR signals a real change in detector performance, technique, or processing before it becomes visible as lost phantom objects, so it is a sensitive early-warning metric the physicist tracks over time.

Does the MQSA 2023 Final Rule change phantom QC?

The 2023 MQSA amendments, enforced beginning September 10, 2024, focus on breast-density reporting, communication timelines, and audit metrics. Phantom image quality remains a required quality-control test; facilities should follow the current ACR DM QC Manual or the manufacturer's QC program.

Key Takeaways

  • The phantom protects the detection task. Fibers, speck groups, and masses stand in for spiculations, microcalcifications, and tumors.
  • Match the criteria to the phantom. The ACR DM Phantom (6/6/6 objects) passes at 2 fibers, 3 speck groups, 2 masses; the legacy screen-film phantom used 4-3-3.13
  • Score with partial credit, from largest to smallest, on a 0.0/0.5/1.0 scale, and record the detail.6
  • Trend SNR, CNR, and AGD. Quantitative metrics catch drift before objects are visibly lost.
  • Observer variability is real. The finest objects are scored least consistently; standardized viewing and automated scoring reduce the noise.58
  • Read the artifacts. No clinically significant artifacts is part of passing.

Conclusion

The ACR Digital Mammography Phantom test looks like a simple counting exercise, but it is one of the most direct checks that a mammography system can still do its job. Its value comes from discipline: the right phantom, the right pass criteria, standardized technique and viewing, careful object and artifact scoring, and quantitative SNR/CNR/AGD trending that catches problems early.

The medical physicist's role is to make the test reproducible and to interpret it in context — separating true system drift from reader-to-reader noise, and connecting image quality to dose. Facilities that treat phantom QC as a trend to manage, not a box to check, protect both image quality and the patients who depend on it.

How DRPS Can Help

Diagnostic Radiation Physics Services performs mammography physicist surveys and quality-control support, including ACR DM Phantom evaluation, CNR/SNR analysis, average glandular dose measurement, artifact assessment, and QC program setup, plus help preparing for and maintaining ACR accreditation. This is delivered through mammography physics and MQSA, accreditation support, and medical physics consulting.

DRPS supports facilities across our service locations, including Florida, Maryland, Virginia, Washington DC, California, Nevada, New York, Pennsylvania, New Jersey, and Delaware.

Related Resources

References

  1. American College of Radiology. Digital Mammography Quality Control Manual. 2nd ed. Reston, VA: ACR; 2018. acr.org
  2. U.S. Food and Drug Administration. Mammography Quality Standards Act (MQSA) and Program. fda.gov
  3. Brooks KW, Trueblood JH, Kearfott KJ, Lawton DT. Automated analysis of the American College of Radiology mammographic accreditation phantom images. Med Phys. 1997;24(5):709-723. doi:10.1118/1.597992. PubMed
  4. American College of Radiology. Mammography Accreditation Program Requirements. Reston, VA: ACR. accreditationsupport.acr.org
  5. Alawaji Z, Taba ST, Cartwright L, Rae W. Automated quality control analysis for American College of Radiology (ACR) digital mammography (DM) phantom images. J Appl Clin Med Phys. 2024;25(12):e14548. doi:10.1002/acm2.14548. PubMed
  6. Yun H, Noh S, Cho H, Ko EY, Yang Z, Woo OH. AI-driven quality assurance in mammography: enhancing quality control efficiency through automated phantom image evaluation in South Korea. PLoS One. 2025;20(9):e0330091. doi:10.1371/journal.pone.0330091. PubMed
  7. Alawaji Z, Tavakoli Taba S, Alshabibi AS, Cartwright L, Rae W. Quality control of mammographic systems using ACR digital mammography (DM) phantom images: a comparative study of automated and human observer scoring. Acta Radiol. 2026;67(6):528-539. doi:10.1177/02841851261425207. PubMed
  8. Alawaji Z, Tavakoli Taba S, Rae W. Automated image quality assessment of mammography phantoms: a systematic review. Acta Radiol. 2022;64(3):971-986. doi:10.1177/02841851221112856. PubMed
  9. Sundell VM, Mäkelä T, Vitikainen AM, Kaasalainen T. Convolutional neural network-based phantom image scoring for mammography quality control. BMC Med Imaging. 2022;22(1):216. doi:10.1186/s12880-022-00944-w. PubMed
  10. Gennaro G, Contento G, Ballaminut A, Caumo F. Inter-phantom variability in digital mammography: implications for quality control. Eur Radiol Exp. 2025;9(1):42. doi:10.1186/s41747-025-00583-0. PubMed
  11. U.S. Food and Drug Administration. 21 CFR 900.12: Quality standards for mammography. ecfr.gov
  12. U.S. Food and Drug Administration. Final Rule to Amend the Mammography Quality Standards Act (MQSA). 2023; enforcement effective September 10, 2024. fda.gov
  13. American College of Radiology, American Association of Physicists in Medicine, Society for Imaging Informatics in Medicine. ACR–AAPM–SIIM Practice Parameter for Determinants of Image Quality in Digital Mammography. Revised 2020. acr.org