Skip to main content

FDG PET/CT in Lymphoma: Deauville Response

By Di Zhang, PhD, DABR, DABSNM
November 29, 2023 • 16 min read

FDG PET/CT is central to staging and response assessment in FDG-avid lymphomas, and the standard language for reporting that response is the Deauville five-point scale — a visual comparison of residual tumor uptake to two internal reference regions, the mediastinal blood pool and the liver.12 The scale was formalized at the First International Workshop on Interim-PET in Deauville, France, in 2009 and has since been embedded in the Lugano classification that governs modern lymphoma staging and response criteria.134

Behind that deceptively simple 1-to-5 score is a real physics problem. The score is only as reliable as the quantitative comparability of the images it grades, both within a single scan and between scans acquired weeks or months apart. Making Deauville scoring and its quantitative companion, ΔSUVmax, reproducible depends on standardized acquisition, accurate standardized uptake value (SUV) quantification, and cross-scanner harmonization.256

Introduction

The genius of the Deauville scale is that it grades disease against internal, patient-specific references rather than an absolute number, which makes it far more robust to the acquisition variables that plague absolute SUV. Injected activity, uptake time, blood glucose, patient size, and scanner calibration all shift absolute uptake values; anchoring the score to the patient's own mediastinum and liver cancels much of that variability.12

That robustness is not unlimited. The reference organs themselves have uptake that depends on acquisition, and quantitative endpoints such as ΔSUVmax require that two scans be directly comparable. In diffuse large B-cell lymphoma, an SUVmax reduction of about 66% after two cycles of chemotherapy improved outcome prediction over visual analysis by reducing false positives, and a reduction near 73% was reported after four cycles — figures that are only meaningful when the baseline and interim scans are quantitatively matched.5

This guide explains what the Deauville scale is, the reference-region concept, the key quantitative principles of SUV and ΔSUVmax, a worked calculation, the clinical impact on response-adapted therapy, practical tips for making scores reproducible, the accreditation and regulatory context, frequently asked questions, and the quality-assurance chain a medical physicist maintains to keep the numbers trustworthy.

Topic Explanation

What is the Deauville five-point scale?

The Deauville five-point scale scores the most intense residual FDG uptake in a site of lymphoma against two internal references — the mediastinal blood pool and the liver — on a scale from 1 to 5.12 It was designed to be simple enough for routine reporting yet reproducible enough for multicenter trials, and it can be adapted to both interim (mid-treatment) and end-of-treatment assessment with good interobserver agreement.3

The five scores are:

  • Score 1 — no uptake above background.
  • Score 2 — uptake at or below the mediastinal blood pool.
  • Score 3 — uptake above the mediastinum but at or below the liver.
  • Score 4 — uptake moderately increased above the liver at any site.
  • Score 5 — uptake markedly increased above the liver, or new FDG-avid disease consistent with lymphoma.23

An additional designation, X, marks new uptake unlikely to be related to lymphoma, prompting further evaluation.2

Why internal reference regions?

Absolute FDG uptake in a lesion depends on many factors that vary between patients and between visits. Rather than fight all of them with a single universal SUV threshold, the Deauville scale uses the patient's own mediastinal blood pool and liver as reference points on the same image, so many acquisition-related shifts affect the lesion and the reference together and largely cancel in the comparison.12

This internal-reference design is why the scale travels well across institutions and scanner platforms — but it also means the reference organs must be measured consistently. The liver reference in particular is sensitive to acquisition and to hepatic physiology, so guidelines specify how it should be sampled.26

How does Deauville relate to Lugano response categories?

The Deauville score is the metabolic input to the Lugano classification, the consensus framework that combines PET-based metabolic assessment with CT-based anatomic measurement and clinical findings into formal response categories.4 For FDG-avid lymphomas, complete metabolic response, partial metabolic response, no metabolic response, and progressive metabolic disease are defined largely by the Deauville score and its change from baseline.34 Reporting a Deauville number without the Lugano context — interim versus end-of-treatment, and the threshold in use — invites misinterpretation.34

Key Technical Principles

The standardized uptake value

The standardized uptake value normalizes measured tissue radioactivity concentration to the injected activity and the patient's body mass, giving a semi-quantitative index of FDG accumulation:

where is the decay-corrected activity concentration in the region of interest (in MBq per mL), is the injected activity (in MBq), and is the patient body mass (in g, so that SUV is dimensionless in units of g/mL). SUVmax is the single hottest voxel in the region; SUVpeak is a small fixed-volume average around it, designed to be less noise-sensitive.26

SUV depends explicitly on accurate knowledge of the injected activity, the uptake (decay) time, the patient mass, and the scanner's cross-calibration to the dose calibrator. An error in any of these propagates directly into SUV, which is why quantitative response endpoints demand a maintained calibration chain.6

The reference-region comparison

Deauville scoring is fundamentally a comparison of the lesion's uptake to the reference organs' uptake. A useful quantitative surrogate for the visual comparison is the ratio to the liver:

A ratio at or below 1 corresponds broadly to a visual score of 3 or lower (lesion at or below liver), while a ratio meaningfully above 1 corresponds to a score of 4 or 5. Approaches such as qPET formalize this liver-normalized ratio to reduce interobserver variability at the clinically critical score-3/score-4 boundary.23

The change in maximum SUV

The quantitative companion to the visual score is the fractional reduction in SUVmax between baseline (PET0) and a later scan (PETn):

In diffuse large B-cell lymphoma, a ΔSUVmax of about 66% after two cycles of chemotherapy predicted event-free survival better than visual analysis by reducing false-positive interim reads, and an optimal cutoff near 73% was reported after four cycles.5 ΔSUVmax is only interpretable when PET0 and PETn are acquired and reconstructed comparably; otherwise the subtraction mixes a real biological change with a technical one.56

Worked example

Consider an interim PET after two cycles of chemotherapy for diffuse large B-cell lymphoma:

  • Baseline nodal SUVmax (PET0): 18.0.
  • Interim nodal SUVmax (PET2): 5.4.
  • Liver mean SUV on the interim study: 2.5.
  • Mediastinal blood-pool SUV on the interim study: 1.8.

The change in maximum SUV is:

A 70% reduction exceeds the roughly 66% two-cycle threshold associated with a favorable outcome, a quantitatively encouraging interim result.5 For the visual Deauville score, the residual lesion SUV of 5.4 is well above both the mediastinum (1.8) and the liver (2.5), so the liver ratio is:

A ratio near 2.2 places the lesion clearly above the liver — a Deauville score of 4. This illustrates a common and clinically important scenario: a favorable quantitative trend (large ΔSUVmax) coexisting with a still-positive visual score, which is exactly why interim response-adapted trials specify in advance which metric and which threshold govern the treatment decision.35

The Deauville five-point scale at a glance

Score Uptake relative to references Typical metabolic interpretation
1 No uptake above background Complete metabolic response
2 At or below mediastinal blood pool Complete metabolic response
3 Above mediastinum, at or below liver Complete metabolic response (in most end-of-treatment settings)
4 Moderately increased above liver Partial, no response, or progression by context
5 Markedly increased above liver, or new disease Partial, no response, or progression by context
X New uptake unlikely related to lymphoma Investigate separately

Whether score 3 is treated as "negative" (complete metabolic response) or grouped with positive scores depends on the clinical question: end-of-treatment assessment generally counts 1–3 as complete metabolic response, whereas some interim response-adapted trials use a more stringent threshold to intensify or de-escalate therapy.34

Clinical Impact

Deauville scoring directly drives treatment in modern lymphoma care. Response-adapted trials in Hodgkin lymphoma and diffuse large B-cell lymphoma use the interim PET score to decide whether to escalate, de-escalate, or maintain therapy — for example, omitting radiotherapy or reducing chemotherapy cycles in patients with a negative interim scan.34 The reliability of the score therefore translates into real decisions about exposing patients to more or less toxic therapy.

Because those decisions turn on the score-3/score-4 boundary and on ΔSUVmax thresholds, small technical inconsistencies matter clinically. A liver reference measured on a noisy, non-harmonized reconstruction can nudge a borderline lesion across the boundary; an interim scan acquired at a different uptake time or on an uncalibrated scanner can distort ΔSUVmax. The quantitative infrastructure is not a back-office concern — it is part of the treatment pathway.56

End-of-treatment PET also guides surveillance. A complete metabolic response (Deauville 1–3) supports a favorable prognosis and argues against routine surveillance scanning, whereas a residual score of 4–5 prompts biopsy or closer follow-up, sparing patients both unnecessary radiation and unnecessary intervention when the assessment is done well.34

Practical Optimization Tips

Standardize the acquisition

Fix the uptake time (commonly about 60 minutes), verify blood glucose is within protocol limits, and use consistent injected-activity-per-body-weight so that SUV and the reference organs are comparable within and between studies. Document these on every report so a downstream reader can judge comparability.26

Sample the reference organs consistently

Measure the liver reference with a standardized volume of interest in normal right-lobe parenchyma, avoiding vessels and lesions, and sample the mediastinal blood pool in the descending thoracic aorta. Consistency here is what preserves the internal-reference advantage of the Deauville scale.2

Keep the calibration chain tight

  • Cross-calibrate the scanner to the dose calibrator on the required schedule.
  • Synchronize clocks between the dose calibrator, injection area, and scanner so decay correction is accurate.
  • Verify SUV accuracy with a uniform phantom and, for multicenter or longitudinal work, participate in a harmonization program.26

Report the metric and the threshold

Always state whether a score is interim or end-of-treatment, which reference defines the positive threshold, and whether ΔSUVmax was used. A bare "Deauville 3" without context can be read as either a complete response or a positive interim result depending on the trial framework.34

Common pitfalls

  1. Comparing scans across scanners without harmonization, letting technical drift masquerade as response.6
  2. Inconsistent uptake time or unmanaged hyperglycemia, distorting both lesion and reference uptake.2
  3. Measuring the liver reference on a lesion, vessel, or noisy region, shifting the score-3/score-4 boundary.2
  4. Reporting ΔSUVmax from non-comparable reconstructions, invalidating the subtraction.5
  5. Omitting the interim-versus-end-of-treatment context, inviting misinterpretation of score 3.34

Regulatory Considerations

FDG PET/CT for lymphoma sits within both the materials-licensing framework for the radiopharmaceutical and the accreditation framework for quantitative imaging quality. F-18 FDG is byproduct material whose medical use falls under 10 CFR Part 35 (or the equivalent Agreement State program), with occupational and public dose limits under 10 CFR Part 20; in Florida the state administers these requirements as an NRC Agreement State under Florida Administrative Code Chapter 64E-5. DRPS also serves Maryland, Virginia, Washington DC, California, Nevada, Pennsylvania, New York, New Jersey, and Delaware, where the NRC or the Agreement-State authority imposes parallel requirements.

For quantitative reliability, the operative standards are professional. The ACR–ACNM–SNMMI–SPR practice parameter for FDG-PET/CT and the SNMMI/EANM procedure guidance define acquisition and quality-control expectations, and the EANM procedure guidelines for tumour imaging version 2.0 specifically target SUV harmonization so that quantitation is comparable across systems and multicenter settings.6 ACR PET accreditation requires phantom-based image-quality and SUV-accuracy verification by a qualified medical physicist. Documented calibration, harmonization, and SUV verification are what make a Deauville score or a ΔSUVmax measurement defensible for clinical decisions. For related quantitative topics, see our guides to PET SUV quantification and EARL PET SUV harmonization.

Frequently Asked Questions (FAQs)

Is a Deauville score the same as an SUV threshold?

No. A Deauville score is a visual comparison of residual uptake to the patient's own mediastinum and liver, which makes it more robust than a single absolute SUV cutoff. Quantitative ratios and ΔSUVmax complement the visual score but do not replace it.12

Does score 3 mean the treatment worked?

Usually, at end of treatment — scores 1–3 are generally read as a complete metabolic response. In some interim response-adapted trials a more stringent threshold is used, so the answer depends on whether the scan is interim or end-of-treatment and on the governing protocol.34

Why is ΔSUVmax sometimes added to the visual score?

Because a large fractional drop in SUVmax can reclassify a visually positive interim scan as a favorable responder, reducing false positives. Reductions on the order of 66% at two cycles improved outcome prediction in diffuse large B-cell lymphoma.5

Can I compare a PET done here to one done at another center?

Only if both were acquired and reconstructed to a common harmonized standard. Without harmonization, SUV and reference comparisons can drift and mimic real change.6

Who is responsible for the quantitative accuracy of the scan?

A qualified or board-certified medical physicist maintains the calibration and harmonization chain and verifies SUV accuracy, working with the nuclear medicine physician and technologists.6

Key Takeaways

  • The Deauville five-point scale grades residual lymphoma uptake against the mediastinal blood pool and the liver.12
  • Internal references make the score robust to acquisition variability that would corrupt an absolute SUV threshold.12
  • Scores 1–3 generally indicate complete metabolic response at end of treatment; scores 4–5 indicate residual disease, interpreted by context.34
  • ΔSUVmax complements the visual score; a roughly 66% reduction at two cycles improved outcome prediction in diffuse large B-cell lymphoma.5
  • Longitudinal and multicenter comparison requires PET harmonization so SUV and reference values are comparable.6
  • A maintained calibration chain and physicist-verified SUV accuracy are what make the score trustworthy.26

How DRPS Can Help

Diagnostic Radiation Physics Services (DRPS) supports PET/CT and nuclear medicine programs across Florida, Maryland, Virginia, Washington DC, California, Nevada, Pennsylvania, New York, New Jersey, and Delaware with PET/CT and nuclear medicine physics, SUV and cross-calibration verification, harmonization support, and ACR accreditation support delivered by board-certified medical physicists.

A reliable lymphoma response program is not just a reader assigning a number. It is a documented quantitative pipeline — calibrated, harmonized, and verified — that lets a Deauville score or a ΔSUVmax measurement carry the clinical weight of a treatment decision.

Conclusion

The Deauville five-point scale turned FDG PET/CT into a practical, reproducible response tool for lymphoma by grading disease against the patient's own liver and mediastinum, and the Lugano classification built that score into formal staging and response criteria. But the reliability of a Deauville number, and of the ΔSUVmax that complements it, rests on a quantitative foundation: standardized acquisition, accurate SUV, consistent reference sampling, and cross-scanner harmonization. When the medical physics infrastructure is maintained and verified, the score can safely guide response-adapted therapy; when it is neglected, technical drift can move a patient across a treatment-defining boundary.1356

Related Resources

References

  1. Meignan M, Gallamini A, Haioun C. Report on the First International Workshop on Interim-PET-Scan in Lymphoma. Leuk Lymphoma. 2009;50(8):1257-1260. doi:10.1080/10428190903040048. doi.org
  2. Barrington SF, Mikhaeel NG, Kostakoglu L, et al. Role of imaging in the staging and response assessment of lymphoma: consensus of the International Conference on Malignant Lymphomas Imaging Working Group. J Clin Oncol. 2014;32(27):3048-3058. doi:10.1200/JCO.2013.53.5229. doi.org
  3. Meignan M, Barrington S, Itti E, Gallamini A, Haioun C, Polliack A. Report on the 4th International Workshop on Positron Emission Tomography in Lymphoma held in Menton, France, 3-5 October 2012. Leuk Lymphoma. 2014;55(1):31-37. doi:10.3109/10428194.2013.802784. doi.org
  4. Cheson BD, Fisher RI, Barrington SF, et al. Recommendations for initial evaluation, staging, and response assessment of Hodgkin and non-Hodgkin lymphoma: the Lugano classification. J Clin Oncol. 2014;32(27):3059-3068. doi:10.1200/JCO.2013.54.8800. doi.org
  5. Itti E, Lin C, Dupuis J, et al. Prognostic value of interim 18F-FDG PET in patients with diffuse large B-cell lymphoma: SUV-based assessment at 4 cycles of chemotherapy. J Nucl Med. 2009;50(4):527-533. doi:10.2967/jnumed.108.057703. doi.org
  6. Boellaard R, Delgado-Bolton R, Oyen WJG, et al. FDG PET/CT: EANM procedure guidelines for tumour imaging: version 2.0. Eur J Nucl Med Mol Imaging. 2015;42(2):328-354. doi:10.1007/s00259-014-2961-x. doi.org
  7. American College of Radiology. ACR–ACNM–SNMMI–SPR Practice Parameter for the Performance of FDG-PET/CT in Oncology. Reston, VA: ACR. acr.org
  8. American College of Radiology. Nuclear Medicine and PET Accreditation Program Requirements. Reston, VA: ACR. accreditationsupport.acr.org