A common misreading of a phylogenetic tree happens before the reader gets to the numbers: assuming that the order the tips (the species, sequences or samples at the ends of the branches) are listed in, top to bottom, tells you something about how closely related they are. It doesn’t. A tree’s topology — which lineages share which ancestors — is fixed, but any internal node can be rotated on its branch without changing what the tree asserts, so two published trees of the identical dataset can list the same tips in opposite orders and be exactly the same tree. What actually carries information is which tips share a common node, not where they sit vertically.
This guide covers what a phylogenetic tree’s parts mean, what branch length does and does not represent, and — the part most explainers skip — how to read the numbers printed on or near the branches: bootstrap values and Bayesian posterior probabilities. Both are support values, but they are calculated differently, are not on a comparable scale, and license different claims about how confident you should be in a given branching pattern.
The anatomy of a tree: tips, nodes, branches and root
Four parts recur on every phylogenetic tree you’ll encounter:
- Tips (leaves, OTUs) — the ends of the branches. Each tip is an operational taxonomic unit: a species, a strain, a gene sequence, a sample. Tips are observed data; everything else on the tree is inferred.
- Internal nodes — branch points. Each internal node represents a hypothesized common ancestor: the point at which, according to the tree, one lineage split into two (or more, on a polytomy — a node with more than two children, usually meaning the data couldn’t resolve the order of splitting).
- Branches (edges) — the lines connecting nodes. A branch represents a lineage persisting through time between one splitting event and the next.
- Root — the single node representing the most recent common ancestor of everything on the tree. Not every tree is rooted; an unrooted tree shows relationships (which lineages group together) without asserting a direction of descent, and rooting typically requires an outgroup — a lineage known, from evidence outside the tree, to have diverged before everything else on it.
The grouping that does the analytical work is the clade (or monophyletic group): a node plus every tip descended from it. When a paper claims two species are “sister taxa,” it means they share a node with each other before either shares one with anything else — not that they appear next to each other in the figure.
Why tip order carries no information
Because any internal node is free to rotate — swap the left and right branches below it — without changing which tips share which ancestors, the vertical or left-right order of tips in a published tree is a drawing choice, not a result. Software typically orders tips alphabetically, by ladder pattern, or to minimize crossing lines; none of those choices are biologically meaningful. If you find yourself inferring a relationship because two tips happen to be adjacent in the figure, check whether they actually share a node before either shares one with anything else — adjacency in the drawing is not evidence of relatedness.
Branch length: what it does and does not mean
This is the second most common misreading, after tip order, and it matters because the same tree shape can be drawn two different ways that mean very different things:
- Cladogram — branch lengths carry no meaning at all; they’re drawn a convenient length purely to make the topology legible. All the information is in the branching pattern (topology), not the line lengths. If a figure doesn’t label its branch-length scale, assume it’s a cladogram and don’t read anything into how long or short a branch looks.
- Phylogram — branch lengths are scaled to the amount of evolutionary change (typically substitutions per site) inferred to have occurred along that branch. A longer branch means more inferred change, not necessarily more elapsed time — a lineage under strong selection or a faster mutation rate accumulates more change per unit time than a slowly evolving one. Phylograms carry a scale bar (e.g., “0.01 substitutions/site”) for exactly this reason; without it, you can’t tell a phylogram from a cladogram at a glance.
- Chronogram (time tree) — branch lengths are scaled to elapsed time, usually calibrated against fossil or molecular-clock evidence, with a time axis instead of (or alongside) a substitution scale.
The practical check before you interpret any branch length: find the scale bar or axis label. If there isn’t one, the tree is very likely a cladogram and branch length is not telling you anything.
Reading the numbers on the branches: bootstrap values and posterior probabilities
The numbers printed at or near internal nodes are support values — a measure of how confident the analysis is that a given node (that specific split) is correct, as opposed to an artifact of the particular dataset or method. The two you’ll see most often are calculated in unrelated ways and are not directly comparable to each other:
- Bootstrap values (also called bootstrap support or bootstrap proportion) come from resampling the underlying sequence alignment with replacement, rebuilding the tree from each resampled version many times (typically 1,000 or more replicates), and recording what percentage of those replicate trees recovered that same node. A node that appears in 950 of 1,000 replicate trees gets a bootstrap value of 95.
- Bayesian posterior probabilities come from a different process entirely — typically Markov chain Monte Carlo sampling of tree space under an explicit evolutionary model — and represent the estimated probability that the clade is correct, given the model and priors. They’re usually reported as a proportion (0 to 1) rather than a percentage.
Because they’re estimated differently, posterior probabilities and bootstrap values are not on a directly comparable scale, and posterior probabilities tend to run higher than bootstrap values for the same node in the same dataset — a well-documented pattern in the phylogenetics literature, not a sign that one method is simply more powerful. Treat a bootstrap value of 95 and a posterior probability of 0.95 as similar in spirit (both say “high support”) but do not treat them as numerically equivalent measurements of the same thing.
Worked example
The table below is an illustrative composite — a small, made-up set of support values in the style of what a real maximum-likelihood-plus-bootstrap and Bayesian analysis of the same alignment might report side by side — not a result from any actual study, included to show how the two numbers get written up together.
| Node (clade) | Bootstrap value | Posterior probability |
|---|---|---|
| Species A + Species B | 98 | 1.00 |
| (Species A + B) + Species C | 72 | 0.94 |
| Species D + Species E | 45 | 0.71 |
The sentence a researcher would write about that pattern in a methods or results section: “The clade uniting Species A and B was strongly supported (bootstrap = 98, posterior probability = 1.00). Support for the placement of Species C as sister to this clade was moderate (bootstrap = 72, posterior probability = 0.94). The relationship between Species D and E was weakly supported (bootstrap = 45, posterior probability = 0.71) and should be treated as unresolved rather than reported as a confirmed finding.” Note what that sentence does not do: it doesn’t claim the D–E clade is wrong, only that the data don’t support it strongly enough to treat it as settled — a genuine polytomy or a different topology are both still live possibilities at that node.
What a support value licenses you to claim — and what it doesn’t
A widely cited rule of thumb, from Hillis and Bull’s 1993 simulation study, is that a bootstrap value of 70% or higher tends to correspond to a real probability of 95% or greater that the clade is correctly resolved. That figure gets repeated often enough in textbooks and lab meetings that it’s worth being precise about its limits: Hillis and Bull derived it under specific simulated conditions (symmetric trees, equal rates of change, and a modest amount of change between splits), and later work has shown it doesn’t hold uniformly across real datasets — under some conditions bootstrap support is conservative (understating real confidence), under others it isn’t. Treat “70%” as a widely used convention worth knowing, not a statistical law that converts a bootstrap value into a fixed confidence level.
What a high support value does license: confidence that, given this alignment, this model and this method, the same node would very likely reappear if you resampled the data again. What it does not license: proof that the clade reflects the true evolutionary history, immunity from systematic error (a support value can be high and the tree still wrong if the underlying model is misspecified, if there’s long-branch attraction, or if the alignment itself has errors), or a claim that a low-support node is definitely incorrect rather than simply unresolved by this dataset.
Assumptions and when not to take a tree at face value
- The alignment has to be right first. A tree is only as good as the sequence alignment it was built from; misaligned or poorly conserved regions propagate directly into the topology and into inflated or deflated support values.
- Long-branch attraction can pull two unrelated, fast-evolving lineages together in the tree regardless of true relationship, and can do so with artificially high bootstrap support, because the same systematic bias reappears in each bootstrap replicate.
- Model misspecification (using an evolutionary model that doesn’t fit the real substitution process) biases the whole tree, not just support values, and bootstrapping does not detect or correct for it.
- Gene trees are not always species trees. A tree built from a single gene reflects that gene’s history, which can differ from the species’ actual history through incomplete lineage sorting, hybridization or horizontal transfer — high support at a node only means the gene-tree topology is well supported, not that it necessarily matches the species tree.
- A polytomy is a result, not a failure. When a node splits into three or more branches with no further resolution, that’s frequently the data honestly reporting it can’t distinguish the order of splitting — don’t read it as a drawing shortcut or an error.
Frequently asked questions
How do you read a phylogenetic tree with numbers on it?
The numbers at internal nodes are support values, most commonly bootstrap percentages (0–100) or Bayesian posterior probabilities (0–1), showing how confident the analysis is in that specific branching point. Numbers directly on the branches themselves, rather than at nodes, are usually branch lengths on a phylogram or chronogram — check for a scale bar to confirm which.
What does a bootstrap value of 100 mean?
That the node in question was recovered in every one of the resampled replicate trees (typically 1,000 replicates) built during the bootstrap analysis — the maximum possible value on that scale, indicating strong, consistent support for that split within this dataset and method. It is not a guarantee the split reflects true evolutionary history, only that it’s highly robust to resampling this particular data.
Do longer branches always mean more time has passed?
No. On a phylogram, branch length represents inferred evolutionary change (commonly substitutions per site), not elapsed time; a branch can be long because the lineage evolved quickly over a short period, not just because a long period passed. Only on a chronogram, calibrated to a time axis, does branch length directly represent time.
Can two differently drawn trees represent the same relationships?
Yes. Because internal nodes can rotate freely without changing which tips share which ancestors, the same topology can be drawn with tips in a different top-to-bottom order, as a triangle instead of a rectangle, or rooted on a different side of the page, and still assert exactly the same relationships.
Is a low bootstrap value proof that a relationship is wrong?
No. A low support value means the dataset and method don’t strongly support that specific resolution — it should be reported as unresolved rather than confidently wrong. The true relationship could still be what’s drawn; the data simply aren’t decisive enough at that node to say so with confidence.
Related reading
Bootstrap resampling and support values sit alongside other measures researchers use to express confidence in a result rather than just report a point estimate — see how that logic plays out with a confidence interval, a p-value, and effect size. If the phylogenetic analysis in question used Bayesian methods, Markov Chain Monte Carlo (MCMC): What It Is and How to Read the Diagnostics covers how the posterior probabilities on the tree were actually generated and how to check the sampler converged. For assembling and pooling evidence across multiple studies rather than a single tree, see systematic review and the PRISMA 2020 reporting standard.







