Written and maintained by CASRAI Editorial Board
Last updated
Systems biology is the study of living things as integrated systems rather than as lists of separate parts. Instead of asking what a single gene or protein does in isolation, it asks how many genes, proteins, metabolites and cells interact to produce behaviour that none of them shows alone — a cell dividing, a signalling pathway switching on, a tissue responding to a drug. It does this by pairing large-scale measurement with mathematical and computational models that can be tested against experiments. This guide explains what the field studies, how its methods work, how it differs from the neighbouring field of computational biology, where its standards and funding come from, and why it matters for research data management and reproducibility.
What Is Systems Biology?
A widely cited short statement of the idea comes from Hiroaki Kitano’s 2002 overview in Science, “Systems Biology: A Brief Overview”: to understand biology at the system level, one has to examine the structure and dynamics of cellular and organismal function, rather than the characteristics of isolated parts of a cell or organism. In practice, three commitments define the field:
- Networks and interactions are the unit of study. Gene regulatory networks, signalling cascades, metabolic networks and protein-interaction networks are analysed as connected structures, because behaviours such as oscillation, switching, robustness and feedback emerge from the wiring, not from any single node.
- Quantitative, often dynamic, measurement. Systems biology relies on data that can constrain a model: concentrations, rates, time courses and single-cell measurements, increasingly gathered with high-throughput technologies.
- Models that make predictions. A model is not just a diagram; it is a formal description that can be simulated, compared with data, and revised when it fails. The cycle of model, prediction, experiment and refinement is the field’s signature workflow.
Because it spans measurement, theory and computation, systems biology is usually done by teams: experimental biologists who generate the data, and mathematicians, physicists, engineers and computer scientists who build and analyse the models. Many research groups aim to have both skills in the same person or the same lab.
Systems Biology vs. Computational Biology
The two names are often used as if they were synonyms, and in day-to-day research they overlap heavily. They are nevertheless defined differently, and the difference is useful to know when reading funding calls or choosing a degree.
- Computational biology is defined by its methods. The NIH Biomedical Information Science and Technology Initiative (BISTI), established in 2000, offered a working definition of computational biology as the development and application of data-analytical and theoretical methods, mathematical modeling and computational simulation techniques to the study of biological, behavioral and social systems. It is a toolbox that can be applied to any biological question, including ones that have nothing to do with networks.
- Systems biology is defined by its scientific stance and question: understanding how components interact to give rise to the behaviour of a whole system. It uses computational biology heavily, but it also requires a specific experimental commitment — perturbing and measuring the system quantitatively — and it is not limited to computation. A purely experimental systems-level study, such as a multi-layer omics profile of a disease, can be systems biology without any simulation.
- Bioinformatics, in the same BISTI document, emphasises tools for acquiring, storing, organising, archiving, analysing or visualising biological, medical, behavioural or health data. See What Is Bioinformatics? for that field in detail.
A simple way to hold the distinction: computational biology and bioinformatics describe how the work is done (computational and mathematical methods; data tools), while systems biology describes what is being asked (how the parts of a living system work together). A systems biologist routinely uses both, and a computational biologist may never work on a whole-system question. Funders and journals do not apply these boundaries consistently, so check the scope statement of any program or journal rather than relying on the label alone.
Core Methods in Systems Biology
Mathematical and computational modelling
Models range from qualitative to fully quantitative, and the choice depends on how much is known about the system.
- Ordinary differential equation (ODE) models describe how concentrations change over time given reaction rates, and are the standard for signalling and metabolic pathways with well-characterised kinetics.
- Stochastic models account for random fluctuations that matter when molecule numbers are small, as in single-cell gene expression.
- Boolean and logic-based network models represent each component as on or off, useful when quantitative parameters are unknown.
- Constraint-based models, such as flux balance analysis of genome-scale metabolic reconstructions, predict metabolic flows from the stoichiometry of a network without needing detailed kinetics.
- Agent-based and multiscale models link molecular events to cells, tissues or organs.
Omics measurement and data integration
Models need data, and the omics technologies supply it at scale: genomics, transcriptomics, proteomics and metabolomics. A distinguishing practice in systems biology is integration — combining several layers measured on the same system to infer regulation that no single layer would show. Examples of the underlying analysis steps are covered in CASRAI’s guides on RNA-seq experimental design and analysis and differential gene expression analysis, and on spatial proteomics.
Network inference and analysis
Researchers reconstruct networks from data and from curated knowledge. Public knowledge resources commonly used as starting points include KEGG for pathways and STRING for protein-protein interaction networks. Graph-theoretic measures, clustering and motif analysis then help identify hubs, modules and regulatory circuits.
Parameter estimation, sensitivity and uncertainty
A model with many unknown parameters must be fitted to data, and fitted parameters are rarely uniquely determined. Sensitivity analysis shows which parameters drive the output, and identifiability and uncertainty analysis shows what the data can and cannot constrain. Honest reporting of this uncertainty is a hallmark of good modelling practice.
Model and data standards
Because models must be exchanged, reproduced and reused, the field has built shared formats, many of them coordinated by the COMBINE initiative:
- SBML (Systems Biology Markup Language) — a free, open, XML-based format for representing models of biological processes. The SBML team began coming together around 1999, and the foundational paper by Hucka and colleagues, summarising SBML Level 1, appeared in Bioinformatics in March 2003.
- SED-ML (Simulation Experiment Description Markup Language) — encodes which models to use, which simulation tasks to run and which results to produce, so a simulation can be re-run.
- SBGN (Systems Biology Graphical Notation) — a set of standardised graphical languages (Process Description, Entity Relationship and Activity Flow) for drawing biological knowledge unambiguously.
COMBINE also covers other formats including CellML, BioPAX, NeuroML and SBOL. The group publishes periodic overviews of the status of these specifications in the Journal of Integrative Bioinformatics.
Workflows and reproducible computation
Modern systems biology analyses chain many steps, from raw measurements to fitted models. Workflow managers and environment tools make those chains re-runnable; see CASRAI’s comparisons of Snakemake vs Nextflow and Mamba vs Conda.
A Short History
The ideas are older than the name. Mihajlo Mesarovic organised a “Systems Theory and Biology” symposium at Case Institute of Technology and edited a volume of that title published in 1968, arguing for applying systems and control theory to living organisms. Through the following decades, mathematical biology, biochemical kinetics and metabolic control analysis kept the quantitative, network-oriented tradition alive, while molecular biology concentrated on isolating and characterising individual components.
Genome sequencing and high-throughput technology changed the position around the year 2000. Complete parts lists made the question “how do these parts work together?” both urgent and tractable. The Institute for Systems Biology in Seattle was co-founded in 2000 by Leroy Hood, Alan Aderem and Ruedi Aebersold, and describes itself as the first institute dedicated to systems biology. In the United States, the National Institute of General Medical Sciences (NIGMS) began funding centres focused on systems-level biomedical research in the early 2000s, and its National Centers for Systems Biology program was set up to promote multidisciplinary research, training and outreach on systems-level studies of biomedical phenomena. Kitano’s 2002 Science overview helped establish the framing, SBML Level 1 was published in 2003, and the journal Molecular Systems Biology launched in 2005 as one of the first open access journals in the area.
Since then the field has absorbed single-cell technologies, spatial measurements and machine learning, and has produced a closely related applied branch, systems medicine, which applies the same network-level thinking to disease mechanisms and treatment.
Where Systems Biology Is Applied
- Cell signalling and regulation — explaining how pathways process information, make switch-like decisions or oscillate.
- Metabolism and metabolic engineering — genome-scale reconstructions to predict growth and guide strain design.
- Development and cell fate — modelling gene regulatory networks that stabilise cell types.
- Disease and drug action — network-level views of cancer, immune responses and drug combinations, overlapping with pharmacometrics and model-informed drug development.
- Microbial communities and ecology — modelling interactions among species in a community.
- Synthetic biology — designing circuits with predictable behaviour, using the same modelling and standards toolkit.
Research Funding and Programs
There is rarely a funding line named exactly “systems biology”; the work is usually funded through broader programs, and names change. Treat the following as an orientation and confirm the current scope with each funder before citing a program or deadline.
- NIH. NIGMS is the institute most identified with foundational, non-disease-specific systems-level and computational biology, including its National Centers for Systems Biology. Disease-focused institutes, such as the National Cancer Institute, fund systems approaches within their own missions, and the National Human Genome Research Institute funds genomic data science.
- National Science Foundation (NSF). The Directorate for Biological Sciences and the mathematical and computer-science directorates fund quantitative and computational work on living systems.
- Department of Energy (DOE). The Office of Science supports systems and computational biology for bioenergy and environmental questions through its biological and environmental research programs.
- UK and Europe. In the UK, the BBSRC is a principal funder of bioscience, including quantitative and systems approaches. European funding runs through European Union framework programmes and national agencies.
- Private foundations. Several philanthropic funders support open-source scientific software, reference datasets and quantitative biology, though the specific programs vary over time.
Proposals commonly straddle categories, so read the stated scope, review criteria and any data-sharing requirements closely. The NIH policy on sharing genomic data, for example, is described in the NIH Genomic Data Sharing policy entry, and many systems biology projects that use human data fall under it.
Journals, Societies and Conferences
- Journals. Molecular Systems Biology (EMBO Press, launched 2005), Cell Systems (Cell Press), npj Systems Biology and Applications (Nature Portfolio, in partnership with The Systems Biology Institute, Tokyo), PLOS Computational Biology, Bioinformatics and the Journal of Integrative Bioinformatics, which hosts the COMBINE standards updates.
- Societies and meetings. The International Society for Systems Biology organises the International Conference on Systems Biology (ICSB). The International Society for Computational Biology (ISCB) hosts the ISMB conference and is the main professional society for computational biology and bioinformatics.
- Standards community. COMBINE coordinates the development of the modelling standards described above.
Training and Career Paths
Entry routes are varied because the field is, by nature, cross-disciplinary. A typical person arrives from biology, biochemistry, physics, applied mathematics, engineering or computer science, then adds the missing side: biologists learn programming, differential equations and statistics; quantitative scientists learn enough laboratory biology to judge what is measurable and what is biologically meaningful. Related foundations are described in What Is Biology?, What Is Molecular Biology?, What Is Biochemistry?, What Is Biostatistics? and What Is Mathematics?.
Dedicated master’s and PhD programs in systems biology, quantitative biology and computational biology exist at many universities and research institutes, usually combining molecular biology, dynamical systems, statistics, and programming with a thesis project that closes the loop between a model and an experiment. Careers run through academic research and core facilities, biotechnology and pharmaceutical companies (target identification, mechanistic modelling and biomarker work) and research software engineering. Because specific programs and job markets change, check current program pages and postings directly.
Why It Matters for Research Administration
Systems biology is a field where data and models are the product, which makes research data management and reproducibility concerns central rather than peripheral.
- Data management planning. Projects generating multi-omic and time-course data need a plan for formats, metadata, storage volumes and repositories from the start. See the data management plan (DMP) entry, and note that repositories such as MetaboLights exist for specific data types.
- Model sharing. A model that exists only as a figure or as code in one lab cannot be checked. Encoding it in SBML, with the simulation described in SED-ML, lets others run it in different software. Curating models into public repositories is an established practice in the community.
- Reproducibility. Reproducing a result requires the data, the model, the parameter values, the software versions and the exact simulation settings. Workflow tools and containerised environments help, and a methods section that names them makes review and replication far easier. See the reproducibility entry for the vocabulary.
- Human-derived data. Studies using patient samples must reconcile open sharing with consent and controlled-access requirements, which is why the GA4GH standards for responsible genomic data sharing are relevant to many projects.
- Team structure and credit. Cross-disciplinary teams include modellers, experimentalists and software developers; recording each contribution transparently supports fair authorship and credit.
- Compute and software costs. High-performance or cloud computing, software maintenance and data storage should be budgeted explicitly rather than treated as overhead.
Frequently Asked Questions
What is systems biology in simple terms?
It is the study of how the many parts of a living system — genes, proteins, metabolites, cells — work together as a connected network to produce the behaviour of the whole, using quantitative measurements and mathematical or computational models that can be tested.
What is the difference between systems biology and computational biology?
Computational biology is defined by methods: computational, mathematical and statistical techniques applied to biological questions. Systems biology is defined by its question — how components interact to produce system-level behaviour — and it uses computational biology as one of its tools while also requiring quantitative experiments. The two overlap heavily and are often used loosely as synonyms.
Is systems biology the same as bioinformatics?
No. Bioinformatics focuses on tools and databases for storing, analysing and visualising biological data. Systems biology draws on bioinformatics to handle its data but is concerned with understanding system behaviour, often through dynamic models. See What Is Bioinformatics?.
What is SBML and why does it matter?
SBML is a free, open XML-based format for representing computational models of biological processes. Because a model saved in SBML can be opened in many different software tools, it makes models shareable and reproducible rather than locked to one program.
Do I need a lot of mathematics to work in systems biology?
Some, but the depth varies by role. Experimental systems biologists need enough quantitative training to design measurements that constrain a model and to read modelling results critically; modellers need enough biology to know what is realistic. Differential equations, statistics and programming are the common foundations.
What does a systems biology project need from a data-management perspective?
A data management plan covering formats and metadata, repository choice, model and code sharing in standard formats, documented software versions, and attention to consent when human data are involved.
Related Guides
This guide is part of CASRAI’s Branches of Science series. For neighbouring fields see bioinformatics, genomics, proteomics and genetics.








