Written and maintained by CASRAI Editorial Board
Last updated
Cheminformatics (also spelled chemoinformatics or chemical informatics) is the application of computing and information science to chemical problems: how molecules are represented in software, stored and searched in databases, described with numerical properties, and used to build models that predict behavior. If analytical chemistry is about measuring what a sample contains, cheminformatics is about what a research community can do with the structures, properties and results once they exist as data. This guide defines the field, explains the molecular representations everything else depends on, surveys the main open chemical databases, and walks through the methods (similarity searching, QSAR, virtual screening) that sit on top of them. It also covers history, training paths, journals and societies, and how cheminformatics connects to research data management and the FAIR principles.
What Cheminformatics Is
There is no single agreed definition, and the field has gone by several names. A frequently cited definition of chemoinformatics in the context of drug discovery came from F. K. Brown in 1998: the mixing of information resources to transform data into information, and information into knowledge, for the purpose of making better decisions faster in drug lead identification and optimization. That definition is tied to pharmaceutical research, but the working scope is wider. Cheminformatics is used in agrochemicals, materials, environmental chemistry, toxicology, natural products research and chemical safety, wherever large collections of molecules and their measured properties need to be organized, searched and modeled.
The activity itself is older than the label. Computer-based handling of chemical structures and chemical documentation goes back to the 1960s and 1970s in academic departments and in pharmaceutical R&D, and both “cheminformatics” and “chemoinformatics” are in current use. The spelling is mostly regional and institutional habit, not a difference in meaning; this guide uses the first form.
Cheminformatics, Computational Chemistry, and Bioinformatics
The boundaries overlap, and people draw them differently. A useful rule of thumb:
- Computational chemistry usually means calculating the physics of molecules from first principles or force fields: electronic structure, energies, geometries, molecular dynamics. It asks what one molecule or system will do.
- Cheminformatics usually treats molecules as data objects (graphs, fingerprints, descriptor vectors) and works across many molecules at once: storing, searching, comparing, classifying and predicting. It asks what a collection of molecules tells you.
- Bioinformatics does the equivalent for biological sequences, structures and omics data; see What Is Bioinformatics? for that parallel field. The two meet in drug discovery, where a compound’s chemistry has to be related to a protein target’s biology.
In practice the same person often does all three, and several of the methods below (docking, for instance) sit on the border.
Molecular Representations
Every cheminformatics workflow begins with a decision about how to write a molecule down in a form a computer can store, compare and compute on. The choice of representation limits what can be done afterward, and it is also a research-data-quality issue, because a structure that cannot be unambiguously reproduced cannot be reliably reused.
Connection Tables and Molfiles
The most basic representation is the molecular graph: atoms as nodes, bonds as edges, usually with 2D or 3D coordinates. Chemical drawing programs and databases exchange these as connection-table formats such as the MDL molfile and SDF (a molfile container for many records). They are explicit and widely supported, but they are verbose and, depending on the software that wrote them, may encode the same molecule in different ways.
SMILES
SMILES (Simplified Molecular Input Line Entry System) writes a molecular graph as a short line of text. David Weininger initiated it in the 1980s at the US EPA laboratory in Duluth. Atoms are written as their element symbols, branches are enclosed in parentheses, ring closures are marked with matching digits, and aromaticity can be written with lowercase letters. Two concrete examples that anyone can verify:
- Ethanol is written
CCO. - Benzene is written
c1ccccc1(aromatic form) orC1=CC=CC=C1(Kekulé form).
The second example shows the main caveat: the same molecule can be written as many different valid SMILES strings, depending on atom order and the notation chosen. Software therefore generates a canonical SMILES using a canonicalization algorithm, so that one molecule maps to one string within that toolkit. Canonical SMILES from different toolkits are not guaranteed to match each other, so a SMILES string alone is not a portable identifier unless the generating software and version are recorded. SMILES has also been extended (for example to encode stereochemistry, isotopes and reactions), and an open community specification, OpenSMILES, was written to reduce ambiguity between implementations.
InChI and InChIKey
The IUPAC International Chemical Identifier (InChI) was developed under IUPAC auspices and released as a standard in 2006. Unlike SMILES, it is designed from the start as an identifier: a given structure should generate the same InChI regardless of how it was drawn. The string is built in layers (formula, connectivity, hydrogens, charge, stereochemistry and so on), and the Standard InChI settings are fixed so that results from different databases line up. For ethanol the standard InChI is InChI=1S/C2H6O/c1-2-3/h3H,2H2,1H3, which can be checked against any major chemistry database.
Long InChI strings are awkward for web searches, so there is also a fixed-length hashed form, the InChIKey: 27 characters, split by hyphens into a 14-character block derived from the core connectivity layer, an 8-character block derived from the remaining layers, and a few flag and check characters. Because it is a hash it cannot be reversed to a structure, but it is well suited to exact-match lookup and to linking records across databases.
SMILES versus InChI
| Property | SMILES | Standard InChI / InChIKey |
|---|---|---|
| Main design goal | Compact, human-readable structure line notation | Canonical, software-independent identifier |
| One molecule, one string? | Only after canonicalization, and only within one toolkit | Yes, by design (for the same standard settings) |
| Readability | Often readable by a trained chemist | Not designed to be read by hand |
| Reversible to a structure? | Yes | InChI yes; InChIKey no (hash) |
| Best use | Input to models, data exchange, quick entry | Deduplication, cross-database linking, citation of a specific structure |
The practical advice for authors and data curators: record both. Keep a SMILES for reuse in software and an InChI or InChIKey for unambiguous lookup, and note which toolkit produced them. Because both are plain text, both can be checked by anyone who parses the string back into a structure, which is what makes them verifiable in a way a drawn picture in a PDF is not.
Fingerprints and Descriptors
For modeling, a structure is usually converted to numbers. Molecular descriptors are calculated properties such as molecular weight, calculated lipophilicity (see LogP vs LogD), polar surface area and counts of hydrogen-bond donors and acceptors. Fingerprints are long bit vectors that record the presence or absence of structural features, such as circular fingerprints built from each atom’s neighborhood. The Tanimoto coefficient, a ratio of shared to total set bits, is the most widely used way to turn two fingerprints into a similarity score. More recently, learned representations from graph neural networks and language-model-style approaches applied to SMILES have been added to this toolbox.
Chemical Databases and Open Chemistry Data
A large share of cheminformatics is made possible by public databases that curate structures and measurements. Three points matter for a researcher choosing one: what the database actually contains (structures only, or structures with measured activity), how records are curated, and under what terms the data can be reused.
PubChem
PubChem is a public repository of information on small molecules and their biological activities. It is maintained by the National Library of Medicine, part of the US National Institutes of Health, and was launched in 2004 as a component of the NIH Molecular Libraries Roadmap Initiatives. It aggregates records contributed by many depositors, links compounds to bioassay results, and offers programmatic access, so it is often the first stop for looking up a structure, synonyms, identifiers and screening data. See CASRAI’s guide to PubChem as a compound and bioassay data repository for details of how it is organized and how to cite and reuse it.
ChEMBL
ChEMBL is a bioactivity database developed and maintained at the European Bioinformatics Institute (EMBL-EBI), part of the European Molecular Biology Laboratory. Its core activity data are manually extracted by curators from the full text of peer-reviewed medicinal chemistry literature, in journals such as the Journal of Medicinal Chemistry, and recorded as standardized measurements (for example IC50, Ki and EC50) against annotated targets. That curation makes it a standard training set for target-prediction and QSAR work. It cross-links to PubChem and other resources. The dictionary entry for ChEMBL gives the short definition.
Other Resources Worth Knowing
- The Protein Data Bank holds experimentally determined macromolecular structures, including many protein–ligand complexes that structure-based design depends on.
- The Crystallography Open Database provides open-access small-molecule and mineral crystal structures. For the experimental side, see What Is Crystallography?
- The Powder Diffraction File is a curated reference database for phase identification.
- Commercial and subscription databases (reaction and substance databases) also exist; their licenses usually restrict redistribution, which matters if your analysis needs to be reproducible by others.
Open Source Toolkits
Much cheminformatics is done with open-source libraries, among them RDKit, Open Babel and the Chemistry Development Kit (CDK). They read and write the formats above, generate canonical SMILES and fingerprints, calculate descriptors, and do substructure searching. Using an open toolkit, and recording its version, is one of the simplest things a lab can do to make a computed result reproducible.
Core Methods
Substructure and Similarity Searching
Substructure search finds every molecule in a collection that contains a specified pattern (a functional group or scaffold). Similarity search ranks a collection by fingerprint similarity to a query molecule. Both rest on the guiding hypothesis that similar structures tend to have similar properties, which is useful and also imperfect: small structural changes sometimes cause large changes in activity, often called activity cliffs.
QSAR and QSPR
Quantitative structure–activity relationship (QSAR) modeling relates numerical descriptions of molecular structure to a measured biological activity, and quantitative structure–property relationship (QSPR) modeling does the same for physical or chemical properties such as solubility or boiling point. The classical approach is usually dated to the work of Corwin Hansch and Toshio Fujita in the early 1960s, whose 1964 paper introduced a hydrophobicity substituent constant for correlating structure with activity. Modern QSAR uses the same logic with far larger descriptor sets and machine-learning algorithms.
The methodological issues are the same as in any statistical modeling, and they are where published QSAR models most often fail:
- Applicability domain. A model is only trustworthy for molecules that resemble its training set. Predictions outside that domain should be flagged.
- Validation. Reporting only fit to the training data overstates performance. External test sets and honest cross-validation (split by scaffold or time, not just randomly) are expected.
- Data quality. Activity values merged from different assays, units or conditions add noise. Curated sources and clear provenance matter. See the Cheng-Prusoff equation guide for one example of why IC50 values from different assays are not directly interchangeable with Ki values.
- Interpretation. A good statistical fit does not by itself show a causal mechanism.
In regulatory toxicology, QSAR predictions are one input among several in some assessment frameworks, which is one reason the documentation of model version, training data and applicability domain is treated as part of the record. See What Is Toxicology? for the neighboring field.
Virtual Screening
Virtual screening uses computation to rank a large library of molecules, real or enumerable, so that experimental testing can focus on the most promising ones. There are two broad families:
- Ligand-based methods start from known active compounds and search for others like them, using similarity, pharmacophore models or QSAR models.
- Structure-based methods use the three-dimensional structure of the target, usually from the Protein Data Bank or a predicted model, and dock candidate molecules into the binding site, scoring the predicted fit.
Virtual screening is a prioritization tool, not a substitute for an assay: hit lists contain false positives, scoring functions are approximate, and a computational hit means very little until it is confirmed experimentally. Pre-filters such as Lipinski’s “rule of five” and related drug-likeness rules are often applied to remove molecules unlikely to have acceptable oral properties, though they are heuristics and have many exceptions.
Machine Learning and AI
Machine learning has been part of cheminformatics for decades in the form of random forests, support vector machines and neural networks used for QSAR. Deep learning on molecular graphs and sequence-style representations has widened the use of the same data for property prediction, generative design and reaction outcome prediction. The same cautions apply, and more strongly: benchmark leakage, narrow training data and weak external validation are the common problems. For the general field, see What Is Artificial Intelligence?
FAIR Data and Reproducibility in Chemistry
Chemistry has a particular reproducibility challenge: a molecule is a structure, but papers and supplements often present it as a drawing in a PDF, a name that several software packages interpret differently, or a table of spectra without the structure behind it. The FAIR data principles (findable, accessible, interoperable, reusable), published in 2016, translate into concrete practices for cheminformatics:
- Findable. Deposit structures and assay results in a repository that assigns a persistent identifier, and include the InChIKey in the metadata so records can be matched.
- Accessible. Use repositories with standard retrieval protocols and clear access conditions, rather than only a supplementary PDF.
- Interoperable. Provide structures in machine-readable formats (SDF, molfile, SMILES, InChI) and use controlled vocabularies for assay and target descriptions, so datasets from different sources can be combined.
- Reusable. State the license, the provenance of each value, and the processing steps. Record software names and versions, including the toolkit that generated canonical SMILES and descriptors, to support computational reproducibility.
Electronic lab notebooks that capture structures natively make these practices much easier, because the structure is entered once and exported in a standard form; see Electronic Lab Notebooks for Chemistry. Preprint servers help early sharing of chemistry results; see What Is ChemRxiv? Datasets and models also belong in a data management plan; see the dictionary entries for data management plan (DMP) and open data. Open data is not the same as unrestricted data: license terms such as those discussed under the Open Database License determine what reuse is allowed.
Applications
- Drug discovery. Library design, hit identification, lead optimization, target prediction and early ADME property prediction. See What Is Pharmacology? for the downstream discipline.
- Toxicology and chemical safety. Predicting hazard endpoints from structure and grouping similar chemicals.
- Analytical chemistry. Matching measured spectra to candidate structures, for example in mass spectrometry; see What Is Analytical Chemistry? and the NIST library match-factor guide.
- Materials and synthesis. Property prediction for new materials and computer-assisted planning of synthetic routes; see What Is Organic Chemistry? and What Is Chemistry?
- Natural products and chemical biology. Organizing and mining collections of compounds and bioassay results, such as those aggregated in PubChem.
A Brief History
- 1960s–1970s. Early computer-based structure handling and chemical documentation. Hansch and Fujita’s work in the early 1960s lays the groundwork for QSAR. The American Chemical Society journal that is now the Journal of Chemical Information and Modeling began in 1961 as the Journal of Chemical Documentation.
- 1980s. SMILES is initiated at the US EPA laboratory in Duluth.
- 1990s. Combinatorial chemistry and high-throughput screening create large datasets; the term chemoinformatics appears in its drug-discovery sense, as in Brown’s 1998 definition.
- 2000s. The IUPAC InChI standard (2006) and PubChem (2004) give the field open identifiers and an open, large-scale structure and assay resource. The Journal of Cheminformatics becomes open access in 2009.
- 2010s onward. Open-source toolkits, curated bioactivity databases such as ChEMBL and data-sharing expectations from funders move cheminformatics toward open, reusable data; deep learning enters the toolbox.
Training Paths and Careers
There is no single route. Cheminformaticians typically come from one of three backgrounds: chemistry or biochemistry with added programming and statistics, computer science or data science with added chemistry, or computational chemistry and molecular modeling. Some universities offer dedicated master’s programs or specialization tracks in chemoinformatics, but many practitioners are trained through a chemistry or bioinformatics PhD plus self-directed or workshop-based learning. Skills that matter in practice:
- Organic and medicinal chemistry fundamentals, enough to judge whether a structure or a prediction is chemically sensible.
- Programming (Python is common in the open-source toolkits) and database querying.
- Statistics and machine-learning validation practice.
- Data curation: standardizing structures, handling salts, tautomers and stereochemistry, and tracking provenance.
Employers include pharmaceutical and biotechnology companies, agrochemical and chemical-industry R&D groups, academic research groups, public databases and research infrastructure, and software and data vendors. Titles vary (cheminformatician, computational chemist, molecular modeler, scientific data scientist), so search by skills rather than by one job title.
Journals, Societies and Venues
- Journal of Cheminformatics, an open-access journal published by BioMed Central, covers chemical information systems, software and databases, molecular modeling, structure representation, similarity searching, computer-aided molecular design, QSAR and data mining.
- Journal of Chemical Information and Modeling, published by the American Chemical Society, covers computational chemistry and chemical informatics.
- ACS Division of Chemical Information (CINF) brings together the chemistry-librarian and chemical-informatics research communities and organizes sessions at ACS meetings.
- Chemistry preprints can be posted on ChemRxiv; see the guide linked above and ChemRxiv vs. bioRxiv for how it compares with other servers.
Check each journal’s current scope statement and author instructions before submitting, since scope and data-deposition requirements change.
Funding and the Link to Research Administration
Cheminformatics projects are usually funded as part of larger programs in drug discovery, chemical biology, toxicology, materials or data infrastructure, rather than under a stand-alone “cheminformatics” heading. Check the specific funder’s program descriptions for the closest fit. For research administrators, a few features distinguish these projects:
- Data and software outputs. Models, curated datasets and code are deliverables. Funder data-sharing policies and the budget for curation, storage and repository fees should be addressed in the data management plan.
- Licensing and IP. Compound libraries, commercial databases and proprietary software come with license terms that can restrict publication of derived data, and models trained on third-party data may carry obligations. Sort these out in the agreement stage, not at the time of publication.
- Compute and software costs. High-performance computing, cloud usage and software licenses are budgeted items; whether they are direct or indirect costs depends on the sponsor and the institution’s policies.
- Reproducibility expectations. Reviewers and funders increasingly look for shared code, versioned software and deposited datasets, which connects to broader research data management practice.
- Collaboration with wet labs. Computational predictions are tested experimentally, which usually means sub-agreements, material transfer or compound-sharing arrangements, and shared authorship decisions.
Frequently Asked Questions
What is cheminformatics in simple terms?
It is the use of computers and information science to store, search, describe and model chemical structures and their properties, so that researchers can make decisions from large collections of chemical data.
Is cheminformatics the same as chemoinformatics?
Yes. The two spellings refer to the same field; usage varies by region and institution.
What is the difference between cheminformatics and computational chemistry?
Computational chemistry mostly calculates the physics of individual molecules or systems. Cheminformatics mostly manages and models data about many molecules. The two overlap, and many projects use both.
What is the difference between SMILES and InChI?
SMILES is a compact line notation for writing a structure, convenient as input to software but not unique across toolkits unless canonicalized. InChI is an IUPAC-standardized identifier designed so that one structure gives one identifier, with the fixed-length InChIKey hash intended for lookup and linking.
What are PubChem and ChEMBL, and how do they differ?
PubChem, run by the NIH National Library of Medicine, is a broad repository of small molecules and bioactivity data from many depositors. ChEMBL, run by EMBL-EBI, is a manually curated bioactivity database extracted largely from the medicinal chemistry literature. They cross-link and are often used together.
What is QSAR used for?
It builds models linking molecular structure to activity or properties, in order to predict values for untested compounds, prioritize synthesis and testing, and screen for hazards. Its predictions are only as reliable as the training data, the validation and the model’s applicability domain.
Does virtual screening replace laboratory screening?
No. It ranks candidates so experimental effort can be focused. Hits must be confirmed in assays.
How does FAIR apply to chemistry data?
Share structures in machine-readable formats, include standard identifiers such as InChIKey, deposit data in repositories with persistent identifiers, state licenses and provenance, and record software versions.
Do I need to be able to program to work in cheminformatics?
For research roles, almost always. Python and the open-source toolkits are the common entry point, though some graphical workflow tools reduce the amount of code required.
Where should I publish cheminformatics research?
Specialist venues include the Journal of Cheminformatics and the Journal of Chemical Information and Modeling; application-focused work often appears in medicinal chemistry, toxicology or materials journals. Match the venue to the audience, and check its data and code sharing requirements first.








