Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

3+3 Design vs. Continual Reassessment Method (CRM) in Phase 1 Oncology Dose-Escalation

A statistical comparison of the 3+3 rule-based dose-escalation design and the model-based Continual Reassessment Method (CRM) used to find the maximum tolerated dose in Phase 1 oncology trials.

Phase 1 oncology trials exist to answer one narrow but consequential question: how much of an investigational drug can be given before toxicity outweighs benefit? The design chosen to answer that question determines how many patients are exposed to ineffective or unsafe doses along the way, and how reliable the resulting maximum tolerated dose (MTD) or recommended dose actually is. Two design families dominate the literature and practice: the algorithmic 3+3 rule-based design and model-based designs, of which the Continual Reassessment Method (CRM) is the original and most widely cited example. This guide compares the two on their mechanics, statistical basis, sample-size and safety tradeoffs, and where each still fits in current oncology drug development.

For the regulatory and policy context around dose selection in oncology specifically — including FDA’s shift away from MTD-only dose selection — see our companion guide on FDA’s Project Optimus. This page focuses on the underlying statistical design question: rule-based vs. model-based dose escalation.

What the 3+3 design is

The 3+3 design is an algorithmic, rule-based dose-escalation method that requires no statistical model and no real-time data analysis. Patients are enrolled in successive cohorts of three at a pre-specified sequence of dose levels. After each cohort completes its dose-limiting toxicity (DLT) observation window, a fixed decision rule determines the next step:

  • 0 of 3 patients experience a DLT: escalate to the next dose level.
  • 1 of 3 patients experiences a DLT: expand the cohort to 6 at the same dose. If no additional DLTs occur (1/6 total), escalate; if a second DLT occurs (2/6 or more), that dose is considered too toxic.
  • 2 or more of 3 (or 6) patients experience a DLT: the dose is declared too toxic; the trial de-escalates, and the MTD is typically defined as the highest dose level at which fewer than 2 of 6 patients experienced a DLT (sometimes with a small additional confirmation cohort).

The appeal of 3+3 is operational: it requires no statistician to run in real time, no software, and no pre-specified dose-toxicity model, which is why it became the de facto default for first-in-human oncology dose-finding for decades and remains in wide use today.

What the Continual Reassessment Method (CRM) is

The CRM, first proposed by O’Quigley, Pepe, and Fisher (1990), is a model-based design. Before the trial starts, investigators specify a working dose-toxicity model (commonly a one-parameter model such as a power or logistic function) and a target DLT probability (for example, 25-33%). After each patient or small cohort is evaluated, the model is re-fit using all accumulated toxicity data across all dose levels tested so far, and the next patient is assigned to whichever dose the updated model estimates is closest to the target toxicity probability. The process continues, “reassessing” the dose-toxicity curve continually, until a pre-specified sample size is reached or a stopping rule is met.

Because CRM uses information from every patient at every dose level — not just the current cohort — it is more statistically efficient than a design that only looks at the most recent cohort’s outcomes. Variants developed since the original proposal address practical concerns with the first formulation, including two-stage CRM (an initial rule-based escalation phase followed by model-based reassessment), Bayesian CRM with different prior specifications, and escalation-with-overdose-control (EWOC), which explicitly constrains the model to limit the probability of assigning a patient to an excessively toxic dose.

3+3 vs. CRM: side-by-side comparison

Dimension 3+3 (rule-based) CRM (model-based)
Statistical basis None — fixed algorithmic rule applied to the current cohort only Pre-specified dose-toxicity model, updated (Bayesian or likelihood-based) after each patient/cohort
Use of accumulated data Only the current (and sometimes immediately prior) cohort’s outcomes drive the next decision All patients at all dose levels tested so far inform the next dose assignment
Target toxicity rate Implicit and imprecise — the design does not pre-specify a numeric target Explicit, pre-specified target DLT probability (e.g., 25-30%)
Accuracy at identifying the true MTD Comparative and simulation studies have repeatedly found 3+3 identifies the true MTD less consistently, especially as the number of dose levels increases Simulation and comparison studies have generally found CRM-type designs identify the true MTD more often across a range of true dose-toxicity scenarios
Patient allocation efficiency Tends to treat more patients at doses well below the eventual MTD because escalation only proceeds one small cohort at a time Converges toward the target dose faster once enough data has accrued, generally allocating fewer patients to markedly sub-therapeutic doses
Operational/statistical burden Low — no biostatistician or model-fitting software required to run in real time Higher — requires a pre-specified model, simulation-based design validation before the trial opens, and a statistician available to re-fit the model as each cohort’s data comes in
Flexibility on cohort/dose spacing Rigid, fixed dose sequence and cohort size defined in advance More flexible — can in principle assign the next patient to any dose the model currently recommends, subject to safety/overdose-control constraints
Regulatory/methodological standing Long track record and broad familiarity, but increasingly criticized in the biostatistics literature as statistically inefficient for modern targeted and immuno-oncology agents Endorsed in the methodological literature as more efficient and accurate; increasingly encouraged (though not mandated) alongside broader dose-optimization expectations such as FDA’s Project Optimus initiative

Safety tradeoffs

Neither design is simply “safer.” Each manages risk differently:

  • 3+3’s conservatism cuts both ways. Because it escalates cautiously and stops at the first sign of excess toxicity, 3+3 rarely assigns a large cohort to a dose far above the true MTD. But that same caution means it frequently under-shoots: trials often stop escalating — and declare an MTD — below the dose level that would in fact be tolerated, meaning some patients are treated at doses unlikely to be therapeutically optimal, purely as an artifact of the algorithm’s coarse decision rule rather than a considered assessment of the dose-response relationship.
  • CRM requires explicit overdose safeguards. Because CRM’s model can, in principle, recommend a large dose jump if early data are misleadingly favorable, well-designed CRM implementations build in constraints: maximum dose-escalation increments between successive cohorts, a requirement that a dose be tested in at least one prior cohort before jumping past it, and/or explicit overdose-control terms (as in EWOC) that cap the model-estimated probability of overdosing the next patient. A CRM design without these safeguards is a legitimate methodological concern; a properly specified one is not inherently riskier than 3+3, and simulation evidence generally shows comparable or better safety performance alongside better accuracy.
  • Sample size is comparable at low dose-level counts, diverges as levels increase. Published comparisons have found that with a small number of pre-specified dose levels, 3+3 and CRM tend to require similar numbers of patients to reach a recommended dose. As the number of dose levels tested grows — common with modern targeted agents and combination regimens, which may explore finer dose gradations — CRM tends to reach a recommended dose using fewer patients than 3+3, and with better precision around the estimated toxicity rate at that dose (Iasonos & O’Quigley, 2014; direct comparison literature such as Iasonos et al., 2008, PubMed 18827039).

Why 3+3 is still used despite the methodological critique

The biostatistics literature has been critical of 3+3 for years — a 2020 Annals of Oncology piece went so far as to title itself a “requiem” for the design — yet 3+3 remains common in practice. The reasons are largely operational rather than statistical: 3+3 requires no dedicated biostatistical support to execute, no simulation-based design validation before the trial opens, and no software to run in real time, which matters for smaller sponsors and investigator-initiated trials without in-house quantitative oncology-trial design expertise. CRM and its variants require that infrastructure up front, which is a real cost even though it typically pays off in a more accurate, more efficient trial.

Where model-based interval designs fit in

Since CRM was introduced, a family of “model-assisted” designs has emerged that tries to combine CRM’s statistical efficiency with 3+3’s operational simplicity. Designs such as the Bayesian Optimal Interval design (BOIN) and the modified toxicity probability interval design (mTPI-2) pre-calculate dose-escalation/de-escalation decision boundaries using an underlying statistical model, so that — like 3+3 — the clinical team can simply look up the next decision from a table, but — like CRM — the boundaries themselves are statistically calibrated against a target toxicity rate and use accumulated trial data. These designs are not the direct subject of this comparison, but are worth knowing as the practical middle ground many newer oncology Phase 1 protocols now choose between the two poles described above.

Frequently asked questions

Is CRM always better than 3+3?

Statistically, simulation and comparison studies have consistently favored CRM-type designs on accuracy in identifying the true MTD and on patient allocation efficiency, particularly as the number of dose levels tested increases. “Better,” however, depends on what a specific trial needs: if the number of dose levels is small and a sponsor lacks biostatistical infrastructure to specify and validate a model-based design before opening the trial, 3+3’s simplicity remains a genuine practical advantage even though it is the statistically weaker design.

Does the FDA require CRM or another model-based design?

No single design is mandated. FDA’s Project Optimus initiative and its 2024 final guidance on dosage optimization for oncology drugs push sponsors toward evaluating more than one dose and supporting dose selection with pharmacokinetic, pharmacodynamic, safety, and efficacy data rather than relying on MTD alone — a goal that model-based and model-assisted designs are generally better suited to support than a fixed 3+3 algorithm, but the guidance addresses dose optimization broadly rather than prescribing a specific escalation algorithm. See our Project Optimus guide for that regulatory context.

What is a dose-limiting toxicity (DLT)?

A DLT is a protocol-defined adverse event — typically a specific grade of toxicity per a standard toxicity grading scale, occurring within a pre-specified observation window after the first dose — that triggers the design’s decision rule (in 3+3) or is entered into the model (in CRM) as an indicator that the dose was not tolerated. The precise DLT definition is protocol-specific and pre-specified before the trial opens, regardless of which escalation design is used.

Can 3+3 and CRM be combined in the same trial?

Yes — this is the basis of the “two-stage” CRM design, which typically opens with a rule-based or otherwise conservative initial escalation phase (functioning similarly to 3+3) while too little data exists to fit a reliable model, then switches to model-based reassessment once enough toxicity data has accumulated to estimate the dose-toxicity curve with reasonable stability.

How does dose escalation relate to the expansion cohort that typically follows it?

Dose escalation (via 3+3, CRM, or a model-assisted design) determines the MTD or a recommended dose range. That dose is then typically carried into one or more expansion cohorts — larger patient groups treated at the selected dose to further characterize safety, tolerability, and preliminary efficacy signals before later-phase development. See our Expansion Cohort dictionary entry and Phase 1 Trial entry for how these pieces fit into the overall first-in-human trial structure.

Key takeaways

  • 3+3 is a fixed, algorithmic rule applied cohort-by-cohort with no underlying statistical model; CRM is a model-based design that re-fits a dose-toxicity model using all accumulated data after each patient or cohort.
  • CRM is generally more accurate at identifying the true MTD and more efficient in patient allocation, particularly as the number of dose levels increases, but requires biostatistical support and simulation-based design validation before the trial opens that 3+3 does not.
  • 3+3’s simplicity is why it persists despite a substantial methodological literature critical of its efficiency; well-specified CRM (with overdose-control safeguards) is not inherently less safe, just operationally heavier to set up.
  • Model-assisted interval designs (BOIN, mTPI-2) exist specifically to bridge the two, offering CRM-like statistical calibration with 3+3-like table-lookup simplicity.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →