Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Materials Genome Initiative (MGI): The US Push for FAIR Materials-Science Data

The Materials Genome Initiative (MGI) is a US federal multi-agency effort, launched in 2011 and coordinated via NIST and the NSTC, to accelerate materials discovery through integrated computation, experiment, and FAIR data infrastructure.

The Materials Genome Initiative (MGI) is a U.S. federal multi-agency initiative, launched in 2011, aimed at cutting the time and cost of discovering, developing, manufacturing, and deploying advanced materials by integrating computational tools, experimental methods, and shared digital data infrastructure. It is coordinated through a subcommittee of the National Science and Technology Council (NSTC) and currently spans roughly 16 federal agencies and bodies — including the Department of Energy, NSF, NIST, DARPA, NASA, NIH, FDA, and the military service research organizations — with the National Institute of Standards and Technology (NIST) playing a leading role given its existing mandate for curating and provisioning critically evaluated materials data and models.

For research-data-management practitioners, MGI matters less as a funding program to track and more as the policy lineage behind why materials science now has a distinct, federally reinforced expectation that research data be shared, machine-actionable, and reusable — a expectation that predates, and helped shape the vocabulary for, the general FAIR (Findable, Accessible, Interoperable, Reusable) data movement across other disciplines.

Origins and strategic direction

MGI began with a June 2011 White House white paper, Materials Genome Initiative for Global Competitiveness, which set the goal of halving the time and cost to bring new materials from discovery to market. The initiative’s founding premise was that materials innovation had become a bottleneck for U.S. manufacturing and clean-energy competitiveness, and that closing the gap required building out three interlocking pillars: computational tools (simulation and modeling), experimental tools (high-throughput synthesis and characterization), and digital data infrastructure connecting the two.

The NSTC subcommittee overseeing MGI has published two formal strategic plans since the initial white paper — one in 2014 and an update in 2021 — along with periodic progress reports (including a 5th-anniversary report in 2016) and PI-meeting proceedings. The 2021 plan sharpened the initiative’s focus on four goals: developing and disseminating advanced materials data infrastructure and standards, integrating experimental and computational research more tightly, promoting a trained workforce fluent in both materials science and data/computational methods, and improving access to materials data and models. That data-infrastructure goal is where FAIR principles enter explicitly: the strategic planning documents describe building infrastructure, tools, and standards for materials data that are findable, accessible, interoperable, and reusable, treating this as necessary groundwork for the initiative’s broader acceleration goal.

MGI and FAIR data for materials science

Materials science has data characteristics that make FAIR compliance harder than in many other fields: results come from a heterogeneous mix of simulation codes (density functional theory, molecular dynamics, phase-field models), physical characterization instruments, and synthesis protocols, each with its own file formats, metadata conventions, and provenance requirements. Before MGI-era investment, most of this data lived in lab notebooks, unpublished spreadsheets, or supplementary files with little shared structure — making it neither easily findable nor reusable outside the group that generated it.

MGI’s approach has been to fund the infrastructure and standards layer rather than mandate a single repository or format. That has translated into support for:

  • High-throughput computational databases built from large volumes of density-functional-theory calculations — efforts such as the Materials Project (Lawrence Berkeley National Laboratory, DOE-funded), AFLOW, and OQMD emerged in the MGI era as queryable, API-accessible repositories of computed crystal structures, thermodynamic, and electronic-structure data.
  • Experimental and multi-modal data infrastructure — the Materials Data Facility, developed at Argonne National Laboratory and the University of Chicago, was funded as part of the MGI-era push for shared experimental databases, providing publication and discovery services for materials datasets that don’t fit neatly into a computational-only repository.
  • Standards and interoperability work through NIST itself, which develops reference data, evaluated property databases, and data-format standards intended to let materials data move between simulation codes, instruments, and repositories without manual reformatting.

It’s worth being precise about one relationship that is easy to overstate: the NOMAD Laboratory (Novel Materials Discovery), which launched its repository and archive at the end of 2014 and is widely cited as the first FAIR-native storage infrastructure for computational materials-science data, is a German/European effort (originally Max Planck Society and Humboldt University Berlin-led, now the NOMAD CoE under EU funding) — not an MGI program. It emerged in parallel with, and is commonly discussed alongside, the MGI-funded U.S. databases because the two ecosystems interoperate and share the same underlying FAIR-for-materials goals, but NOMAD itself is not a NIST or MGI deliverable.

What this means in practice for materials-science researchers

If you generate materials data under U.S. federal funding, MGI’s legacy shows up less as a direct compliance requirement and more as the reason your funding agency’s data-sharing expectations and the repository landscape look the way they do. Practical implications:

  • Data management plans should name a domain-appropriate repository, not a generic one. For computational results (DFT outputs, simulation trajectories), the Materials Project, NOMAD, and similar databases accept structured submissions with standardized metadata. For experimental and mixed-modality data, the Materials Data Facility and other domain repositories are typically a better fit than a generalist repository, which won’t enforce materials-specific metadata schemas.
  • Metadata and provenance matter more than in many fields because reuse of materials data (e.g., feeding it into a machine-learning model trained across labs) depends on knowing exactly which code, version, basis set, or instrument settings produced a result — MGI-funded infrastructure projects have generally prioritized capturing this provenance over simply hosting files.
  • Interoperability is an explicit design goal, not an afterthought. Several of these repositories expose APIs specifically so that data can be queried programmatically and combined across sources (e.g., cross-referencing Materials Project and NOMAD entries), which is directly downstream of MGI’s original computational/experimental integration goal.
  • Check your funder’s specific expectations. DOE, NSF, and NIST each fund and reference MGI-era infrastructure somewhat differently in their own data-management-plan guidance; MGI itself does not operate a single centralized data policy or repository mandate, so researchers should confirm which agency-specific DMP requirements apply to their award rather than assuming a uniform MGI data policy exists.

Frequently asked questions

Is the Materials Genome Initiative still active?

Yes. MGI remains an active NSTC-coordinated interagency initiative as of its most recent published strategic plan (2021) and subsequent PI meetings and challenge documents; it has not been formally sunset.

Does MGI operate its own data repository?

No. MGI is a coordinating policy and funding initiative, not a single repository. It has funded and helped catalyze specific infrastructure — including the Materials Data Facility and high-throughput computational databases like the Materials Project — but researchers should identify the specific repository appropriate to their data type rather than looking for a single “MGI repository.”

Is NOMAD part of the Materials Genome Initiative?

No. NOMAD is a European (originally German-led, now EU Centre of Excellence) infrastructure effort. It is not funded or administered under MGI, though it shares MGI’s FAIR-data goals for computational materials science and the two ecosystems are often discussed together and interoperate.

How does MGI relate to the general FAIR Principles?

MGI’s data-infrastructure goal explicitly adopts FAIR data principles — findable, accessible, interoperable, reusable — as the target standard for materials data infrastructure, applying them to a domain (materials science) with unusually heterogeneous computational and experimental data types.

What agency leads MGI?

MGI is coordinated by a subcommittee of the National Science and Technology Council, with NIST playing a leading operational role due to its existing mission around evaluated reference data and measurement standards; funding and participation is distributed across roughly 16 federal agencies rather than concentrated in one.

Related CASRAI resources

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →