Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

Data Stewardship and the Data Curator Role in Research

What data stewards and data curators actually do day to day, how the role differs from a data manager or data librarian, and where it typically sits in a research institution.

“Data steward,” “data curator,” “data manager,” and “data librarian” all show up in job postings at research institutions, sometimes for functionally identical positions and sometimes for genuinely different ones. There is no single, universally agreed job description — a 2024 conference poster surveying research data management (RDM) job titles concluded that titles are often “invented from scratch” and that responsibilities vary substantially even when the title is the same (Wilbrandt, IDCC24). That ambiguity is a real operational problem for a research office trying to decide whether to create this role, what to call it, and where it should sit. This guide covers what the role actually does day to day, how it differs from adjacent titles, and where it typically fits in an institution’s organizational chart.

What a data steward or data curator actually does, day to day

Strip away the title variation and the core function is consistent across the literature: someone has to sit between the researcher generating data and the infrastructure and communities that will eventually reuse it. The Digital Curation Centre (DCC) defines digital curation as “the management and preservation of digital data/information over the long-term,” spanning a full curation lifecycle from initial data creation through appraisal, ingest, preservation action, and reuse — not a single point-in-time task. In practice, that breaks down into four recurring areas of day-to-day work.

Metadata quality

Writing, checking, and correcting the metadata attached to a dataset is the single most consistently cited task across role descriptions. This is also formally recognized in the CRediT taxonomy: the Data Curation role is officially defined as “management activities to annotate (produce metadata), scrub data and maintain research data … for initial use and later re-use.” A steward or curator selects an appropriate metadata schema for the discipline (Dublin Core, DataCite, or a domain-specific schema), fills in required fields, checks the values against controlled vocabularies, and fixes gaps that would otherwise make a dataset hard to find or hard to reuse correctly — units not recorded, ambiguous variable names, missing provenance about how the data was collected or processed.

FAIR compliance checking

Someone has to actually check whether a dataset meets the FAIR principles (Findable, Accessible, Interoperable, Reusable) before or at deposit — that’s not automatic, and it’s routinely a named task in the role. This overlaps closely with metadata quality (Findability and Interoperability are largely metadata problems) but also covers Accessibility (is there a clear access/reuse license attached, not just a vague “available on request”) and Reusability (is the documentation sufficient for someone outside the original team to actually use the data correctly).

Repository liaison

Data doesn’t deposit itself. A steward or curator typically helps a researcher choose an appropriate repository, prepares the deposit package, and — where the institution runs its own repository — is often the person doing the work to keep that repository meeting CoreTrustSeal certification requirements or equivalent trusted digital repository standards: documented preservation policies, defined workflows, and demonstrable procedures for handling deposits from ingest through long-term access. For institutions depositing to an external generalist or discipline-specific repository instead of running their own, this liaison work is lighter but still real — matching a dataset to the right home is covered in more depth in CASRAI’s guide to choosing an open data repository.

Researcher support

This is the least metadata-focused and most people-facing part of the role: reviewing draft Data Management Plans (DMPs) before submission, running training sessions, answering ad hoc questions about a funder’s data-sharing requirements, and — in institutions with a central stewardship function — sitting on institutional working groups that shape RDM policy. The Turing Way’s description of the role lists exactly this mix: reviewing DMPs, supporting grant proposals, and guiding researchers through the practical steps of data sharing, alongside the more technical curation work above.

Data steward vs. data curator vs. data manager vs. data librarian

These four titles are used inconsistently enough across institutions and countries that no single distinction holds everywhere — this is a genuinely unresolved area of the profession, not a gap in this guide. But three real, recurring distinctions are worth knowing:

  • Curation vs. management, by scope. Where the two terms are used deliberately rather than interchangeably, “curation” tends to be used for the archiving/long-term-preservation end of the lifecycle, while “(data) management” covers the full process from collection through analysis to eventual archiving or reuse — curation as a subset of management, not a synonym for it.
  • “Steward” as the deliberately generic term. Several European RDM programs adopted “data steward” specifically because it is less domain-specific than “librarian,” “curator,” or “scientist” — useful where one role has to work across disciplines with very different data practices (a genomics dataset and a survey dataset have almost nothing in common structurally). CASRAI’s own Dictionary reflects this pattern: the Data Steward Role (in DMP) term defines it as “a named individual or role accountable for the day-to-day execution of a DMP’s commitments,” explicitly distinct from the principal investigator and from a repository (a repository is infrastructure the steward liaises with, not a steward itself).
  • “Librarian” implies a specific professional background, not just a task set. Where “data librarian” is used, it usually signals the person also carries traditional librarianship responsibilities — reference support, information literacy instruction, outreach — alongside data-specific curation work, and is typically embedded in the library rather than a standalone research-support unit.

The honest summary for a research-administration audience: don’t assume a candidate’s or a peer institution’s job title tells you what they actually do. Write (or read) the position description against the four task areas above, not against the title.

Where the role sits organizationally

Placement varies by institution and there is no single dominant model. A 2024 comparative study of data stewardship programs across North American and European universities (Rousi, Boehm & Wang, Journal of Documentation, 2024) found that policies, funding structures, and organizational placement all differed meaningfully across the institutions studied, even where the underlying goals were the same. The three placements that recur most often in the wider RDM literature are:

  • Centralized within the library. Academic libraries were early adopters of RDM support (data management is frequently cited as the top new responsibility taken on by liaison/subject librarians in recent years), so a centralized library-based data curation team is a common model, particularly where the institution already has strong institutional-repository infrastructure to plug into.
  • A dedicated central research-support office. Larger or research-intensive institutions sometimes stand up a stewardship function inside the research office (alongside grants administration and compliance staff) rather than the library, on the logic that DMP review and funder-compliance work sits closer to the grants lifecycle than to collections work.
  • Distributed / embedded within faculties or departments. Some institutions place stewards within individual faculties or research groups rather than centrally, trading consistency for closer domain fit — a steward embedded in a genomics department can develop real expertise in that discipline’s metadata standards in a way a fully central, generalist role often cannot.

None of these is “correct” in the abstract; each trades off consistency, domain expertise, and cost differently. A small institution with one or two people in this function will usually end up centralized by necessity; a large, research-intensive institution has a genuine choice to make and should make it deliberately rather than by default.

Deciding whether to create or fill this role

For a research-administration audience actually weighing this decision, three practical signals are worth checking before writing a job description:

  • Is DMP review currently happening at all, and by whom? If DMPs are being submitted to funders without anyone at the institution checking they’re realistic and consistent with institutional policy, that’s the clearest signal the function is missing, not just under-resourced.
  • Does your institution run its own repository, or rely on external ones? Running your own repository (especially if pursuing CoreTrustSeal certification) creates ongoing curation and metadata-quality work that doesn’t stop once the repository is built — that’s an argument for a dedicated, ongoing role rather than a one-off project hire.
  • Are researchers actually asking for help, and where does that request currently land? If DMP and data-sharing questions currently get routed inconsistently — sometimes IT, sometimes the library, sometimes nobody — that’s a sign the function exists in practice already, informally, and formalizing it (with a real title, reporting line, and time allocation) is more a recognition than a new cost.

The RDA Professionalising Data Stewardship Interest Group has published a landscape report and career-paths survey specifically aimed at this maturing-profession problem — a useful reference for job-description language and training-needs planning if standing up the role for the first time.

How this connects to CASRAI’s own vocabulary

Because CASRAI both originated and co-stewards CRediT (with NISO, as ANSI/NISO Z39.104-2022) and maintains the Dictionary’s machine-actionable-DMP vocabulary, the data steward/curator role connects to several pieces of CASRAI infrastructure directly rather than needing a separate explainer:

  • The Data Curation CRediT role gives a steward’s curation work a standard, citable way to be credited in a publication’s author-contribution statement — distinct from authorship itself.
  • The RDA DMP Common Standard and machine-actionable DMP (maDMP) both support recording a named data steward as a DMP “contributor” entity, identified by ORCID, so the role is machine-readable rather than only appearing in narrative text.
  • Metadata-quality work connects directly to the DataCite metadata schema and, downstream, to whether a dataset actually surfaces in Google Dataset Search — good curation work has a concrete discoverability payoff, not just a compliance one.

Frequently asked questions

What is data curation, in one sentence?

Data curation is the ongoing work of organizing, describing (via metadata), validating, and preserving research data so it remains findable and usable — by the original researcher and by others — well beyond the point it was first collected. It’s a process across the data lifecycle, not a one-time cleanup step.

What does a data curator actually do?

Day to day: writing and checking metadata, verifying FAIR compliance before deposit, liaising with a repository (internal or external) on ingest requirements, and supporting researchers with DMP review and data-sharing questions. See the four task areas above for detail.

Is a database administrator the same as a data curator?

No — a database administrator (DBA) manages the technical infrastructure a database runs on (performance, backups, access control) for operational systems generally. A research data curator works with the content and description of specific research datasets across their lifecycle, usually with no responsibility for the underlying database software or servers at all. The similar-sounding titles are a genuine, common source of confusion, particularly in job-search contexts.

Is “data steward” the same job as “data curator”?

Often close enough in practice to be interchangeable, but not guaranteed — some institutions use “steward” for a broader, more generalist accountability role and reserve “curator” for the more hands-on metadata/preservation work. Read the actual position description rather than assuming from the title (see the distinctions section above).

Does a research institution need a dedicated data curator, or can existing staff absorb the work?

It depends on scale. An institution running its own certified repository, or with a high volume of externally-funded DMPs requiring review, generally needs a dedicated role — the work is ongoing, not a one-off project. Smaller programs sometimes absorb it into an existing subject-librarian or research-support position; the decision criteria in the section above are a reasonable starting checklist either way.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →