CyVerse is a National Science Foundation (NSF)-funded, open-source cyberinfrastructure platform that provides data storage, computational resources, and bioinformatics tools for plant biology and life-sciences research. It began in 2008 as The iPlant Collaborative, an NSF award (Award #0735191) focused specifically on plant sciences and genomics, before expanding its scope to all of the life sciences in 2013 and rebranding as CyVerse in 2016 (NSF Award #1743442). For research data management (RDM) purposes, CyVerse functions as both a computing platform and a discipline-specific data repository, making it relevant to researchers and research administrators handling large, computationally intensive biological datasets — genomic sequences, phenotyping images, environmental sensor data, and similar.
What CyVerse Is, and Who Runs It
CyVerse is led by the University of Arizona, with the Texas Advanced Computing Center (TACC) providing a significant share of the underlying computing infrastructure, and is sustained through ongoing NSF funding rather than a single one-time grant. As of a 2024 NSF award reported by the University of Arizona, CyVerse continues as part of a larger, multi-institution NSF-supported framework for data-driven bioscience discovery, with the University of Arizona named the lead institution for the CyVerse component (University of Arizona News). The platform is free to use for U.S.-based academic research and is built on open-source software, which distinguishes it from commercial cloud-computing providers a lab might otherwise turn to for similar computational needs.
From iPlant Collaborative to CyVerse: Why the History Matters
Researchers and administrators who worked with plant genomics data before 2016 may still know this infrastructure as “iPlant” or “the iPlant Collaborative” — older grant reports, data management plans, and publications may cite tools or the platform itself under that name. The rebrand to CyVerse in 2016 reflected a genuine scope change, not just a name change: the underlying cyberinfrastructure broadened from a plant-sciences-specific mandate to a general life-sciences one, though plant biology, genomics, and agricultural research remain a core and historically foundational part of its user base. When verifying citations or legacy data management plans that reference “iPlant,” it is the same organizational lineage as current CyVerse infrastructure, per CyVerse’s own About page.
Core Services CyVerse Provides
Data Store
The Data Store is CyVerse’s cloud storage system for research data — a place to upload, organize, and share large datasets (CyVerse’s tools are built around workloads well beyond what a typical institutional email attachment or shared drive can handle) with lab members or collaborators, with access controls researchers manage themselves.
Discovery Environment (DE)
The Discovery Environment is CyVerse’s web-based analysis workbench: a graphical interface for running bioinformatics applications and workflows — sequence alignment, genome assembly, phenotyping image analysis, and similar pipelines — without requiring a researcher to independently install and configure that software or provision their own computing hardware. This is one of the platform’s most direct value propositions for research administrators evaluating cyberinfrastructure needs during a grant proposal: it lowers the computational and technical barrier to data-intensive plant-biology and life-sciences analysis for labs that don’t have dedicated bioinformatics or IT support.
CyVerse Data Commons and Curated Data
CyVerse’s Data Commons is where the platform’s role as a data repository becomes most concrete. Datasets published through CyVerse Curated Data receive a permanent identifier — a DOI or an ARK (Archival Resource Key) — and are expected to remain stable and citable long-term; publication requires meeting CyVerse’s own data-organization guidelines (a README, an inventory, and documented metadata) and an open data license. CyVerse also offers a lighter-weight “Community Released Data” option for datasets that are actively evolving or not yet ready for long-term preservation commitments, which can later be promoted to Curated Data (and issued a permanent identifier) once the underlying dataset is stable, per CyVerse’s own Data Commons publishing documentation.
Computing, APIs, and Newer AI/ML Capabilities
Beyond the Data Store and Discovery Environment, CyVerse provides programmatic Science APIs for integrating CyVerse storage and compute into external pipelines, a cloud automation and orchestration layer (CACAO) for deploying services, and — reflecting the platform’s more recent direction — GPU-backed infrastructure supporting AI/ML workloads, including a 2024 NSF award specifically supporting CyVerse’s role in providing cyberinfrastructure and training for a new NSF Artificial Intelligence Institute (per CyVerse’s own announcement). CyVerse also documents enhanced security configurations for regulated data types, including ITAR- and HIPAA-relevant considerations, for labs working with export-controlled or health-related data.
CyVerse and FAIR Data Compliance
For research administrators tracking FAIR data principles compliance — Findable, Accessible, Interoperable, Reusable — CyVerse’s Data Commons is built to support several of these directly rather than as an afterthought:
- Findable: Curated Data receives a persistent DOI or ARK and a landing page, satisfying the core FAIR requirement that data be assigned a globally unique, persistent identifier.
- Accessible: the identifier and landing page remain resolvable independent of whether the underlying files move, and access conditions are documented.
- Interoperable: CyVerse requires structured metadata for Curated Data submissions, including support for the DataCite metadata schema via a DOI-request template built into the Discovery Environment.
- Reusable: publication requires an open data license, a README, and a documented file inventory — the minimum context another researcher needs to reuse a dataset correctly.
This makes CyVerse a plausible repository choice to name in a data management plan (DMP) for plant-biology or life-sciences projects where a discipline-specific repository is preferred over a generalist one — see CASRAI’s guide on how to choose an open data repository for the broader decision framework, and how to make your dataset FAIR: a step-by-step checklist for the underlying principles CyVerse’s Curated Data requirements are built around. CASRAI has not independently verified CyVerse’s certification status against a third-party framework such as CoreTrustSeal; researchers with a funder mandate that specifically requires certified-repository status should confirm current standing directly with CyVerse or the funder before naming it in a DMP.
Who Uses CyVerse
CyVerse’s user base spans plant biology, genomics, agricultural science, and increasingly the broader life sciences — its own materials describe supporting research from single-lab projects up through large, multi-institution consortia. Because it is free for U.S. academic use and does not require institutional subscription, it is also a practical option for labs at institutions without their own high-performance computing or bioinformatics core facility. For research administrators, this makes CyVerse worth knowing about specifically in the context of genome-project data management and other large-scale biological research data workflows, where local infrastructure and storage limits are a common bottleneck in a data management plan.
How CyVerse Compares to Generalist Repositories
CyVerse is a discipline-specific repository paired with an active computing platform, which distinguishes it from generalist repositories such as Dryad or Figshare that primarily provide deposit, DOI-minting, and preservation without domain-specific analysis tools built in. A researcher publishing a plant-genomics dataset might choose CyVerse specifically because the same platform can host the analysis pipeline that produced the data, not just the resulting files — while a researcher whose funder or journal requires deposit in a generalist, discipline-agnostic repository regardless of subject area may still need Dryad, Figshare, or an institutional repository alongside or instead of CyVerse. Checking the specific repository requirement in a funder’s data-sharing policy or a journal’s data-availability policy before choosing is worth doing either way.
Frequently Asked Questions
What is CyVerse used for?
CyVerse is used to store, analyze, and share large biological research datasets — most heavily in plant biology, genomics, and agricultural science, though its scope now covers the life sciences generally. Its Discovery Environment provides bioinformatics analysis pipelines, and its Data Store and Data Commons provide storage and, for published datasets, a citable, permanent identifier.
Is CyVerse the same thing as iPlant Collaborative?
Yes — CyVerse is the continuation of the same NSF-funded infrastructure originally launched in 2008 as The iPlant Collaborative. It was renamed CyVerse in 2016 when its scope expanded from plant sciences specifically to the life sciences more broadly.
Is CyVerse free to use?
CyVerse is free to use for U.S.-based academic research, funded through NSF awards rather than user subscriptions. Researchers should confirm current terms directly with CyVerse for their specific use case, particularly for non-academic or international use.
Does CyVerse issue DOIs for datasets?
Yes. Datasets published through CyVerse Curated Data receive a permanent identifier — a DOI or an ARK — along with a landing page, provided the submission meets CyVerse’s data-organization and metadata requirements and includes an open data license.
Is CyVerse only useful for plant science research?
No, not any longer. CyVerse began as plant-science-specific infrastructure (as The iPlant Collaborative) but broadened to support the life sciences generally starting in 2013, ahead of its 2016 rebrand. Plant biology and agricultural genomics remain a core part of its user base, but it is not restricted to that domain.







