X-ray crystallography is the technique that has produced the majority of experimentally determined atomic-resolution structures in the Protein Data Bank, and it remains the workhorse method for small-molecule and many macromolecular structures despite the rise of cryo-electron microscopy. This guide covers the physical principle, the practical workflow from crystal to deposited coordinates, the data-quality metrics that actually tell you whether a structure is trustworthy, and the deposition and integrity obligations that make crystallography one of the more self-correcting corners of experimental science.
The core idea: diffraction, Bragg’s law, and electron density
A crystal is a three-dimensional, periodically repeating arrangement of molecules. When a beam of X-rays passes through that ordered lattice, it scatters off the electrons in the structure and produces a discrete pattern of diffraction spots rather than a continuous smear, because constructive interference only occurs at specific angles set by the spacing of the lattice planes. That relationship is described by Bragg’s law, nλ = 2d sinθ, where λ is the X-ray wavelength, d is the spacing between a set of lattice planes, θ is the angle of incidence, and n is an integer order of reflection.
Each spot in a diffraction pattern is a “reflection” with a measurable intensity and, in principle, a phase. Intensities alone can be recorded directly by the detector. Mathematically, the electron density inside the unit cell — which is what you actually want, because it lets you build an atomic model — is obtained by a Fourier transform of the complete set of structure-factor amplitudes and phases. Because both amplitude and phase are needed and only amplitude is measured, crystallography has a problem to solve before any model can be built at all.
The phase problem
Diffraction detectors record intensity (proportional to the square of the structure-factor amplitude), not phase. This is the “phase problem”: without phase information, the same set of measured intensities is consistent with infinitely many different electron-density maps, and there is no direct experimental way to measure phase at X-ray wavelengths. Several complementary strategies are used to recover it:
- Molecular replacement (MR) — the dominant method today for macromolecules. A previously solved structure of a homologous protein or domain is used as a search model; its known phases are combined with the new crystal’s measured amplitudes to generate an initial map, which is then rebuilt and refined against the new data.
- Isomorphous replacement (SIR/MIR) — heavy-atom derivatives (e.g., mercury or platinum compounds soaked into the crystal) are compared against the native crystal; the small intensity differences they cause let you locate the heavy atoms and derive initial phases.
- Anomalous dispersion (SAD/MAD) — exploits the fact that atoms with absorption edges near the X-ray wavelength used (commonly selenium, incorporated via selenomethionine-substituted protein) scatter with a measurable anomalous signal. Single-wavelength anomalous dispersion (SAD) or multi-wavelength anomalous dispersion (MAD) data, collected at a tunable synchrotron beamline, provide enough information to solve phases experimentally without a search model.
- Direct methods — for small molecules with well-ordered, high-resolution data (typically better than about 1.2 Å), statistical relationships between reflection phases can be exploited computationally to solve the structure directly, without heavy atoms or a search model. This is the standard approach in small-molecule and organic/inorganic crystallography, where structures deposited to the Cambridge Structural Database are routinely solved this way.
Once an initial, approximate set of phases produces an interpretable map, the model is built and improved iteratively — each cycle of model building generates better phases, which produce a better map, which supports further model building.
The workflow, end to end
Purification and crystallisation
For macromolecules, the starting point is a highly pure, homogeneous, monodisperse sample — heterogeneity in folding state, oligomeric state, or post-translational modification is one of the most common reasons a construct simply will not crystallise. Crystallisation itself is typically approached by vapour diffusion (hanging-drop or sitting-drop), mixing protein with a reservoir solution containing a precipitant, buffer, and additives, then allowing the drop to slowly equilibrate toward supersaturation. Because there is no reliable way to predict which chemical conditions will produce diffraction-quality crystals for a given macromolecule, initial broad sparse-matrix screens (typically covering hundreds of condition combinations) are followed by targeted optimisation around any promising hits — fine-tuning pH, precipitant concentration, temperature, and additives to improve crystal size and order.
Cryoprotection and mounting
Most macromolecular data collection today is done at cryogenic temperature (around 100 K, using a stream of cold nitrogen gas) to reduce radiation damage during exposure. Crystals are equilibrated in a cryoprotectant (commonly glycerol, ethylene glycol, or a higher concentration of the precipitant itself) to prevent ice-crystal formation from disrupting the lattice, then looped in a small fibre or nylon loop and flash-cooled directly in liquid nitrogen or in the cold gas stream at the beamline.
Data collection: home source vs. synchrotron
Data can be collected on a laboratory (“home source”) rotating-anode or microfocus sealed-tube X-ray generator, which is adequate for well-diffracting crystals and routine structures. For weakly diffracting crystals, small crystals, anomalous-signal phasing experiments, or high-throughput structural biology, synchrotron beamlines are standard: they provide several orders of magnitude higher X-ray flux, tunable wavelength (essential for SAD/MAD phasing), and finely focused beams down to micron scale, all of which let structures be solved from crystals too small or too weakly ordered for a home source to handle.
Indexing, integration, and scaling
Raw diffraction images are processed computationally: indexing assigns Miller indices to each reflection based on the crystal’s unit cell and orientation, integration measures each reflection’s intensity, and scaling merges and puts multiple images (and multiple crystals, if needed) onto a common scale while correcting for absorption, radiation damage, and detector effects. The output is a merged, scaled set of structure-factor amplitudes ready for phasing.
Phasing, model building, and refinement
Once phases are obtained (by one of the methods above) and an initial map calculated, the atomic model is built into the electron density, initially often with substantial automation, then refined iteratively by adjusting atomic coordinates, occupancies, and atomic displacement (B-)factors to maximise agreement between the model’s calculated diffraction pattern and the observed data, subject to stereochemical restraints (bond lengths, angles, planarity) that keep the model chemically sensible.
Validation
Before a structure is considered finished, it goes through geometric and statistical validation — checking the model against the experimental data and against expected chemistry — which is exactly what wwPDB’s mandatory validation report covers at deposition (see below).
Quality metrics, and what they actually mean
A resolution number alone tells you almost nothing about whether a structure is reliable. These are the metrics that matter, and why:
- Resolution — the smallest d-spacing (in Å) at which reflections are still meaningfully measured above background. Lower numbers mean finer detail is resolved (side-chain rotamers and ordered water become visible below about 2 Å; below about 1.2 Å individual atoms can be resolved). Resolution describes the data, not automatically the model — a poorly refined model can still be built into high-resolution data.
- R-work and R-free — R-work is the standard crystallographic residual measuring the disagreement between observed and model-calculated structure-factor amplitudes for the data used in refinement; lower is better, but R-work alone can be driven arbitrarily low by overfitting (adding parameters that fit noise rather than signal). R-free, introduced by Brünger in 1992, guards against exactly that: a small fraction of reflections (commonly around 5-10%) is set aside at the start of refinement and never used to fit the model, only to check it. R-free should track R-work reasonably closely (typically within a few percentage points at a given resolution); a large gap between them is a specific, well-recognised sign of overfitting. This is precisely why a model can report a flattering R-work and still be a poor structure — R-free, not R-work, is the number to check first.
- Completeness — the percentage of theoretically possible unique reflections that were actually measured, in the relevant resolution range. Incomplete data (from radiation damage, crystal geometry, or detector gaps) can bias the resulting map.
- Redundancy (multiplicity) — how many times, on average, each unique reflection was measured. Higher redundancy improves the accuracy of the merged intensities and the reliability of downstream statistics.
- I/σI — the average intensity of a reflection divided by its estimated error, used to judge at what resolution the data stop being statistically meaningful (data are conventionally still considered usable down to an I/σI of roughly 1-2, though modern practice increasingly favours CC½ over a hard I/σI cutoff).
- CC½ — the correlation coefficient between two randomly split halves of the same dataset, proposed by Karplus and Diederichs as a more statistically grounded way to define the useful resolution limit than traditional I/σI or completeness cutoffs, because it is on the same scale as the correlation between the model and the true underlying data.
- Ramachandran outliers — the percentage of protein backbone dihedral angles (φ/ψ) that fall outside statistically favoured and allowed regions of the Ramachandran plot. A well-refined model at reasonable resolution should have few or no outliers; a high outlier count signals either genuine and rare local strain or, far more often, a poorly built or over-fitted model.
- Clashscore — a count of unfavourable atomic overlaps (steric clashes) per 1,000 atoms, as computed by tools such as MolProbity. A low clashscore indicates the model is not just fitting the data but is also physically and chemically plausible.
The reason to read several of these together rather than any one in isolation: a model can achieve a low, flattering R-free and still have poor stereochemistry — bad geometry, Ramachandran outliers, high clashscore — if it was refined carelessly or if the underlying data were weaker than the reported resolution number suggests. Reviewers, structural genomics consortia, and the wwPDB validation pipeline all treat geometry and clashscore as independent checks precisely because R-free alone does not catch every failure mode.
Deposition is mandatory, not optional
Structural biology is one of the few fields where public data deposition is not a norm under debate but a near-universal, enforced requirement. The Protein Data Bank (PDB), operated by the worldwide Protein Data Bank (wwPDB) consortium, is the single global archive for macromolecular structures determined by crystallography, NMR, and cryo-EM. Small-molecule and organic/inorganic crystal structures instead go to the Cambridge Structural Database (CSD), maintained by the Cambridge Crystallographic Data Centre; CASRAI also covers the Crystallography Open Database (COD) as an open-access alternative for small-molecule structures.
Practically, this means:
- Journals require a PDB accession code before or at publication. Virtually every major structural biology journal (and most general biology and biochemistry journals publishing a structure) requires deposition of coordinates, and increasingly the raw or processed structure-factor data, as a condition of publication, not merely as a recommendation.
- wwPDB runs a mandatory validation report at deposition. Every deposited structure is automatically checked against the metrics above (geometry, clashscore, R-free behaviour, data-processing statistics) and against comparable structures in the archive, and this validation report is made available to depositors, reviewers, and — after release — the public alongside the entry itself.
- Structure factors should be deposited alongside coordinates. Depositing only the final coordinate model, without the underlying structure-factor amplitudes, makes independent re-refinement or re-examination of the model against the experimental data impossible. wwPDB member archives require structure-factor deposition for X-ray structures as part of standard practice.
- Embargo/hold options exist but are time-limited. Depositors can request that a deposited entry be held back from public release, typically until publication or for a defined maximum period, rather than being released immediately — but the norm is release, not indefinite non-disclosure, and most journals require the hold to be lifted at publication.
This deposition requirement connects directly to the broader norms CASRAI covers around data availability statements and the FAIR data principles — a PDB or CSD accession code is, in effect, a structured, machine-checkable data availability statement that predates the general open-data movement by decades.
The integrity angle: why deposited structure factors make the field self-correcting
Because a deposited PDB entry ideally includes not just the final coordinates but the experimental structure factors and the model’s refinement statistics, an incorrect structure can, in principle, be caught by anyone re-examining the deposited data — not only by the original authors. This is not a hypothetical safeguard. A well-documented case from the mid-2000s illustrates it directly: a laboratory studying ABC-transporter membrane proteins discovered that a sign error in an in-house data-processing script had inverted the phase information used to build several published structures, producing models with an incorrect hand and topology despite otherwise reasonable-looking crystallographic statistics. The error came to light when an independently solved, related structure showed a different and more chemically sensible fold, prompting re-examination of the original data — and led to the retraction of multiple papers in 2006-2007 (Science published the formal retraction). The episode is frequently cited in crystallography training precisely because the underlying data, not just the published conclusion, were what ultimately exposed the error.
This is the structural argument for why deposition mandates matter for research integrity beyond simple transparency: a claim built on data nobody else can check is much harder to independently verify or refute than one built on data anyone can re-process. It is also why reproducibility in crystallography has a genuinely different character than in many other experimental fields — re-refining a deposited structure against its own deposited structure factors is a direct, data-level check, not merely a repeat of the experiment. It is the same logic behind retraction mechanisms generally: the ability to catch an error after publication depends entirely on whether the underlying data were ever made available to check.
X-ray crystallography vs. cryo-EM vs. NMR
These three techniques are complementary, not strictly competing, and the right choice depends on the sample:
- X-ray crystallography requires a well-ordered crystal but, once one is obtained, routinely delivers the highest resolution of the three methods (well below 2 Å is common, sub-1 Å achievable for very well-ordered small proteins and small molecules), making it the method of choice when atomic-level detail — precise bond geometry, water networks, ligand-binding poses for drug design — is the goal.
- Cryo-electron microscopy (cryo-EM) does not require a crystal at all — samples are vitrified directly and imaged — which is exactly why it has displaced crystallography for large, flexible, or heterogeneous assemblies (ribosomes, viral capsids, membrane-protein complexes) that are difficult or impossible to crystallise. Since the “resolution revolution” driven by direct electron detectors in the early-to-mid 2010s, cryo-EM has closed much of the resolution gap with crystallography for large complexes, though very small proteins remain harder for cryo-EM than for crystallography. Raw cryo-EM image data is archived separately, in EMPIAR, alongside PDB deposition of the resulting model.
- NMR spectroscopy works in solution rather than requiring a crystal or vitrified grid, and is uniquely able to report on dynamics and conformational flexibility directly, but is practically limited to smaller proteins and complexes (roughly under 40-50 kDa for routine structure determination) by signal overlap and relaxation effects at larger molecular weight.
In practice, many structural biology groups use more than one of these techniques on the same system, choosing crystallography when a well-diffracting crystal is achievable and atomic detail matters most, cryo-EM when the assembly is too large or heterogeneous to crystallise, and NMR when solution dynamics are the actual question.
A note on AlphaFold and predicted models
AlphaFold and related deep-learning structure-prediction methods produce computationally predicted models, not experimentally determined structures — no diffraction data, no electron density, and no experimental phase or refinement statistics underlie a prediction. Predicted models should never be deposited to the PDB as though they were experimental structures; the wwPDB archive is explicitly for experimentally determined structures, and confusing the two undermines exactly the kind of verifiability this guide describes. Predicted models are appropriately shared through dedicated resources built for that purpose (such as the AlphaFold Protein Structure Database), clearly labelled as predictions, and are commonly used to provide molecular-replacement search models or to guide interpretation of a low-resolution cryo-EM map — a genuinely useful role, but a distinct one from experimental structure determination.
Frequently asked questions
Why isn’t a low R-free enough to trust a crystal structure?
Because R-free measures agreement with the experimental data, not chemical or geometric plausibility. A model can fit the diffraction data reasonably well and still have implausible bond geometry, backbone torsion angles outside allowed Ramachandran regions, or steric clashes between atoms — all signs of a poorly built or over-interpreted model. Reliable validation checks R-free alongside geometry and clashscore together, which is exactly what the wwPDB validation report does automatically at deposition.
Do I have to deposit structure factors, or just the final coordinates?
wwPDB member archives require structure-factor deposition alongside coordinates for X-ray structures as standard practice, and most journals require it as a condition of publication. Coordinates alone cannot be independently re-refined or checked against the original experimental data.
Can I deposit an AlphaFold-predicted model to the PDB?
No. The PDB is an archive of experimentally determined structures. Predicted models belong in dedicated prediction databases, clearly labelled as computational predictions rather than experimental results.
Why has cryo-EM become so much more common for large complexes?
Cryo-EM does not require growing a crystal — samples are vitrified and imaged directly — which removes the single biggest bottleneck for large, flexible, or compositionally heterogeneous assemblies. Since the resolution improvements driven by direct electron detectors in the 2010s, cryo-EM can now reach resolutions competitive with crystallography for many such complexes, while crystallography still generally leads for small, well-ordered targets and for the highest attainable atomic resolution.
What is the phase problem, in one sentence?
Diffraction detectors record the intensity, but not the phase, of each reflection, and both are mathematically required to reconstruct electron density — so phases have to be recovered indirectly, by molecular replacement, isomorphous replacement, anomalous dispersion, or (for small molecules) direct methods.







