Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Environmental Data Initiative (EDI): The LTER Network’s Data Repository

EDI is the repository that publishes LTER Network and broader environmental/ecological data as EML-described data packages — distinct from the EML standard itself and from NEON’s data-generating role.

The Environmental Data Initiative (EDI) is a domain repository for ecological and environmental data, operated jointly by the University of New Mexico (UNM) and the University of Wisconsin–Madison and funded primarily by the US National Science Foundation’s Division of Environmental Biology. EDI is the primary data-publication venue for the US Long Term Ecological Research (LTER) Network, but it is not the same thing as LTER, and it is not the same thing as the metadata standard its data packages are built on. This guide explains what EDI actually is, how it differs from the standards and networks researchers most often confuse it with, and how the publication workflow works in practice.

What is the Environmental Data Initiative?

EDI is a repository — a hosted service where researchers deposit, describe, and publish environmental and ecological datasets so that other researchers can find, evaluate, and reuse them. It grew out of informatics work at the LTER Network Office and formally launched in 2016 as a joint project of UNM and UW–Madison, with an initial National Science Foundation award of roughly $3.1 million. EDI’s founding mission was to extend the LTER Network’s decades of data-management practice to a broader set of environmental science communities that lacked a domain-appropriate archive of their own, including NSF’s Long-Term Research in Environmental Biology (LTREB) program, the Organization of Biological Field Stations (OBFS), and the Museum of Southwestern Biology’s Arctic/environmental collections program, alongside the LTER sites themselves.

Today EDI accepts submissions from any environmental or ecological researcher, not only NSF-funded programs, provided the data fit its scope and quality requirements. Datasets published through EDI receive a persistent DOI, are indexed by the DataONE federation of earth- and environmental-science repositories, and are discoverable through Google Dataset Search — giving depositors the citable, fundable-data-management-plan-satisfying publication record that funders such as NSF increasingly expect.

EDI is a repository, not the EML standard

EDI’s data packages are described using Ecological Metadata Language (EML), the XML-based metadata specification that originated from the same LTER/NCEAS informatics lineage EDI itself grew out of. This is where the two are easiest to conflate, so it is worth being precise about the distinction:

  • EML is a metadata standard — a schema that defines how to structure the description of a dataset (its creators, methods, coverage, and the attributes of each data table). EML is maintained as an open, community-governed specification and is used by multiple independent repositories, not just EDI — also by the Knowledge Network for Biocomplexity (KNB) and other DataONE member repositories.
  • EDI is a repository — an organization that hosts data, runs a submission workflow, performs quality checks, mints DOIs, and provides long-term access and preservation. EDI happens to use EML as the metadata layer for every package it hosts, but EDI is the service, not the standard.

The relationship is analogous to the one between Darwin Core and GBIF: Darwin Core is a vocabulary for describing occurrence records, and GBIF is one of several repositories that aggregates data published using it. A researcher choosing where to submit data needs to evaluate EDI on repository-level criteria — scope, curation support, cost, retention commitments — not on EML’s technical properties, which are the same regardless of which EML-based repository ultimately hosts the package.

EDI vs. NEON: submission repository vs. data-generating observatory

EDI and the National Ecological Observatory Network (NEON) are both frequently cited environmental-data infrastructure, and both serve the same broad research community, but they occupy fundamentally different roles:

  • NEON is a data-generating facility. It operates its own standardized sensors, airborne surveys, and field-sampling protocols at 81 sites across the US to produce a single, internally consistent long-term dataset, and it publishes that data itself through its own data portal.
  • EDI is a submission repository. It does not generate its own measurements; it accepts data packages from many independent researchers and projects — LTER sites, individual PIs, field stations, and others, including NEON-affiliated researchers publishing derived products — each with its own methods, instruments, and sampling design.

In practice this means a researcher does not choose “EDI or NEON” as competing options for the same data the way they might choose between two general-purpose repositories. NEON is a place to obtain a specific, standardized continental-scale dataset; EDI is a place to deposit a dataset a researcher or project has independently produced, whether or not it has any connection to NEON or the LTER Network.

History and governance

EDI’s technical foundation is PASTA (Provenance Aware Synthesis Tracking Architecture), a repository software platform originally developed at the LTER Network Office beginning around 2009, with a production release in January 2013, under the technical leadership of Mark Servilla at UNM. In 2016, that PASTA-based infrastructure and its operating team joined with information-management researchers at UW–Madison — including Corinna Gries, who has led much of EDI’s community outreach and researcher-support work — to form EDI as a formally NSF-funded, cross-institutional initiative. EDI has subsequently received continued NSF support to sustain operations. The repository’s software (PASTA+ and related tools) is developed openly, with documentation and source maintained by the EDI organization (EDIorg) on GitHub.

The data package model

EDI organizes everything it hosts into data packages: a thematically coherent bundle of one or more data entities (tables, spatial files, or other resources) plus a single EML metadata document describing the package as a whole — its creators, methods, temporal and geographic coverage, and the structure of each constituent file down to individual column names, units, and missing-value codes. This mirrors how EML is used generally (see the EML dictionary entry for the standard itself), but EDI adds repository-specific behavior on top of the standard:

  • Quality checking — submitted packages are run through automated congruency checks that verify the EML metadata actually matches the submitted data files (correct column counts, consistent units, valid coverage information) before publication, catching a common source of unusable “published” datasets.
  • DOI assignment — each published data package receives a persistent DOI, and revisions to a package are versioned, with the DOI/citation record tracking which version is being cited.
  • REST API and provenance tracking — PASTA exposes a REST API for programmatic search, download, and workflow automation, and tracks provenance relationships between packages, including which derived products were built from which source packages.
  • Discovery — published packages are searchable directly through the EDI data portal and are also harvested into DataONE, extending discoverability beyond EDI’s own site.

Who publishes through EDI

EDI’s core constituency remains the LTER Network’s member sites across the US, but its scope is broader in practice:

  • LTER Network sites, for both core long-term monitoring data and site-based synthesis products.
  • NSF’s Long-Term Research in Environmental Biology (LTREB) awardees.
  • Organization of Biological Field Stations (OBFS) member stations.
  • Independent environmental and ecological researchers — including those affiliated with NEON, USGS, and other agencies — who need a domain-appropriate archive and do not have one through their home institution or funder.

This “open archive of last resort” role is one of EDI’s stated founding purposes: it exists partly to serve environmental scientists whose institutions or projects have no other viable place to deposit and preserve data long-term, which is a distinct value proposition from generalist repositories (e.g., Dryad or a university’s institutional repository) that don’t apply ecology-specific curation, or from a discipline-general metadata schema comparison — see how to choose a metadata schema for a dataset for that broader decision.

Frequently asked questions

Is EDI the same organization as the LTER Network?

No. The LTER Network is a set of NSF-funded long-term research sites; EDI is a separate, though closely affiliated, repository organization that grew out of LTER’s data-management infrastructure and now serves as the LTER Network’s primary data-publication venue, alongside serving other environmental science communities.

Does EDI generate its own environmental data, like NEON does?

No. EDI does not operate sensors, field surveys, or sampling programs. It is a submission repository: researchers and projects that generate their own data deposit it with EDI for curation, DOI assignment, and long-term preservation. NEON, by contrast, is itself a data-generating observatory.

What metadata standard does EDI use?

Ecological Metadata Language (EML). Every data package published through EDI is described with an EML document, but EML is an independent, community-governed standard also used by other repositories (such as KNB) — it is not exclusive to EDI.

Can researchers outside the LTER Network publish data through EDI?

Yes. While LTER sites are EDI’s founding and largest constituency, EDI also serves LTREB and OBFS-affiliated researchers and accepts submissions from other environmental scientists whose data fit its scope and quality requirements, particularly where no other domain-appropriate repository is available.

Do EDI data packages get a DOI?

Yes. Published data packages receive a persistent DOI, and revised versions of a package are tracked so that citations can point to a specific version.

Related reading

References

  • Environmental Data Initiative, edirepository.org — organizational overview, data package model, and services.
  • PASTA+ documentation (pastaplus-core.readthedocs.io) — PASTA/EDI repository software history and architecture.
  • University of New Mexico Newsroom and UNM Center for Advanced Research Computing — EDI founding, NSF funding, and institutional roles.
  • “The Environmental Data Initiative: Connecting the past to the future through data reuse,” PMC (PMC9817195).

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →