Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

When Does Data Curation or Stewardship Work Become Co-Authorship?

A judgment framework, built on ICMJE’s authorship criteria and the CRediT taxonomy, for deciding when curating, cleaning, or maintaining a shared or reused dataset rises to co-authorship versus acknowledgment.

A shared or reused dataset almost always has someone behind it who cleaned it, documented it, checked it for errors, or kept it usable over time — a data curator, data steward, repository manager, or lab data manager. When that dataset ends up underpinning a publication, a genuine question follows: does that curation work amount to authorship, or is it acknowledgment-worthy technical support?

This is a distinct question from who counts as an author on a data paper itself — that guide covers the formal publication describing a dataset. Here the dataset is being reused: a researcher pulls a shared, previously curated dataset (a lab’s internal collection, a repository holding, a consortium resource) into a new analysis, and the person who maintained that dataset was not necessarily involved in writing the paper at all. No publisher, standards body, or funder sets a single numeric threshold for this. What exists instead is a converging judgment framework, built from ICMJE’s authorship criteria and the CRediT taxonomy, that research teams can apply consistently. This guide sets out that framework rather than a fixed rule.

Two frameworks that answer different questions

Two standards get invoked in this conversation, and conflating them is the single most common source of confusion:

  • CRediT (Contributor Roles Taxonomy), formalized as ANSI/NISO Z39.104-2022, defines a “Data Curation” role: “Management activities to annotate (produce metadata), scrub data and maintain research data (including software code, where it is necessary for interpreting the data itself) for initial use and later re-use.” CRediT describes what someone did. It was built to let a contribution statement say “this person did data curation” separately from “this person did formal analysis” or “this person wrote the manuscript” — see CRediT’s Data Curation role page for the full definition and worked examples.
  • ICMJE’s authorship criteria answer a different question entirely: who qualifies to be named as an author. ICMJE requires all four of the following, for every named author: substantial contribution to the conception, design, acquisition, analysis, or interpretation of the work; drafting the work or revising it critically for important intellectual content; final approval of the version to be published; and agreement to be accountable for all aspects of the work.

The critical point: being assigned a CRediT “Data Curation” credit does not, by itself, make someone an author — and the reverse is also true, someone can meet the ICMJE bar for authorship through curation work even though “Data Curation” is not typically thought of as an authorship-track role. CRediT is a contribution-description system; it does not adjudicate the authorship question. That adjudication runs entirely through ICMJE’s four-part test (or an equivalent institutional/journal policy built on it).

Applying the ICMJE test to curation work specifically

Curation and stewardship work spans a wide range, from mechanical file handling to genuinely expert judgment calls. Walking through each ICMJE criterion against that range is what makes the threshold visible:

1. Substantial contribution to conception, design, acquisition, analysis, or interpretation

This is where the real variation lives. Contribution counts toward this criterion when curation involves domain-expert judgment that shapes what the data means for the study — for example, developing the quality-control logic that determines which records are usable, resolving conflicting values across merged sources in a way that requires subject-matter expertise, deriving a harmonized or de-duplicated dataset from disparate inputs, or making documented decisions about how ambiguous or missing values are handled that materially affect downstream results. It does not typically count when the work is applying an existing, previously specified protocol mechanically — running a standard cleaning script, reformatting files to a required submission schema, or backing up and hosting data without exercising interpretive judgment about its content.

2. Drafting or critically revising the work for important intellectual content

This is frequently the criterion that separates curation contributors who meet the full authorship bar from those who don’t, independent of how substantial their data work was. A curator who never sees the manuscript, or who reviews only a methods paragraph describing the dataset, is unlikely to satisfy this criterion. One who is materially involved in interpreting results, reviewing the discussion, or shaping how the dataset’s limitations are represented in the paper is closer to meeting it.

3. Final approval of the version to be published

A procedural but real requirement — the person has actually reviewed and approved the final manuscript, not just an earlier draft or a data-description paragraph sent for factual sign-off.

4. Agreement to be accountable for all aspects of the work

Authors are accountable for the integrity of the whole paper, not just their own section. A data steward willing to be named as accountable for the dataset’s provenance and handling, but not for the study’s conclusions generally, is signaling they don’t intend to meet this criterion — which is a legitimate position, and one ICMJE’s own guidance anticipates by directing such contributors to acknowledgment instead.

All four criteria must be met — meeting only the data-related first criterion, however substantial, is not sufficient on its own. This is the same all-four-or-none logic covered in more general terms in CASRAI’s ICMJE Authorship Criteria guide.

Signals the work is crossing into authorship territory

  • The curator made non-obvious methodological choices (inclusion/exclusion logic, harmonization rules, handling of conflicting or missing data) that a reader would need to understand to properly interpret the results, and that required domain expertise rather than a generic technical skill.
  • The curator’s decisions materially shaped the analytic sample or the variables available for analysis, not just the dataset’s format or accessibility.
  • The curator is engaged with the manuscript itself — reviewing drafts, discussing interpretation, commenting on how findings should be framed — not only responding to isolated factual queries about the data.
  • The curator is willing to stand behind the published paper’s integrity as a whole, not only the dataset’s provenance.
  • The work was original and non-routine enough that it could plausibly have been described, on its own, as a methods contribution to this specific study.

Signals the work remains acknowledgment-worthy technical support

  • The curation followed an existing, previously documented protocol, checklist, or repository ingestion standard without requiring new judgment calls specific to this study.
  • The dataset was curated and made available independently of, and prior to, this specific research question — a shared resource maintained for general reuse rather than built or adapted for this paper.
  • The curator’s involvement ended once the dataset was delivered; they did not review interpretation, drafts, or conclusions.
  • The contribution is better described as infrastructure or service provision (hosting, format conversion, routine metadata tagging, backup and preservation) than as a substantive intellectual contribution to this study’s design or interpretation.

This is the same underlying distinction CASRAI’s guide on core facility staff and co-authorship works through for instrument and service work — routine, protocol-following service work sits on the acknowledgment side of the line, while service work that involves genuine intellectual judgment specific to the study can cross into authorship. Data curation and stewardship follow the identical logic; the type of work differs, the test does not.

The specific case of a shared or reused dataset

The scenario this guide is scoped to — a curator maintaining a dataset that other researchers reuse, rather than a curator working on a single study’s own data paper — adds a wrinkle: the curator’s work often predates, and was never intended for, the specific paper now citing it. A repository data manager who maintains a public genomic database, a consortium data coordinator, or a lab’s long-serving data steward is typically doing ongoing, general-purpose curation for many downstream users, not for any one study.

In that situation, the default expectation in most institutional and publisher practice is that reuse is credited through formal data citation and, where the paper uses a CRediT statement, a “Data Curation” contributor credit for the specific individuals involved — not co-authorship on every paper that cites the dataset. Authorship becomes the right call only when the specific engagement for this paper independently meets the four ICMJE criteria above: for instance, the steward was consulted specifically for this analysis, made judgment calls that shaped it, and engaged with the resulting manuscript. Ongoing stewardship of a shared resource, on its own, is what the data curator role is for — see CASRAI’s guide on data stewardship and the data curator role for what that day-to-day work typically covers. It is a real, credit-worthy contribution; it is just not, by itself, authorship.

A practical decision framework

  1. Identify what the curator actually did for this specific paper, distinct from their general maintenance of the dataset. General stewardship of a shared resource does not carry authorship weight on every downstream paper; project-specific judgment calls might.
  2. Check each of the four ICMJE criteria against that specific work — not against the curator’s overall expertise or the dataset’s overall importance to the field.
  3. Raise the question early, ideally when the dataset is first incorporated into the study, not at manuscript submission. ICMJE’s own recommendations note that anyone meeting authorship criteria should be given the opportunity to be listed as an author, which requires the conversation to happen while there’s still time to involve them in drafting and revision.
  4. If authorship isn’t the right call, credit the contribution properly rather than omitting it — a CRediT “Data Curation” statement, a named acknowledgment, and/or a formal data citation to the dataset’s own DOI are the standard mechanisms, and using more than one is common and appropriate.
  5. Document the decision, particularly for long-running or multi-institution projects with shared data infrastructure, so the same reasoning applies consistently across the project’s outputs rather than being re-litigated paper by paper.

Where the assessment is genuinely contested, most journals and institutions follow COPE’s authorship-dispute guidance rather than leaving it to informal negotiation between the parties.

Frequently asked questions

Does a “Data Curation” CRediT credit ever, by itself, mean someone should be a co-author?

No. CRediT credits describe what a contributor did; they do not determine who qualifies as an author. A person can hold a Data Curation credit as a non-author contributor, and, separately, a person doing curation work can meet the full ICMJE authorship bar — the two questions are evaluated independently.

Is there an industry-standard percentage or hours threshold for curation work to become authorship?

No verified standards body publishes a numeric threshold. ICMJE’s criteria are qualitative and apply to all disciplines and contribution types identically — the assessment is a judgment call against the four criteria, not a quantitative cutoff.

What if the data steward maintains a public repository and doesn’t know their dataset was used in a particular paper?

That is the clearest case for acknowledgment via data citation rather than authorship. ICMJE requires that authors give final approval of the manuscript and agree to be accountable for it — someone unaware their data was used in a specific study cannot meet either criterion for that study, whatever the quality of their original curation work.

Should this be decided by the research team alone, or does the journal or institution have a say?

Both. ICMJE sets the baseline criteria that most journals and institutions adopt directly or adapt into their own authorship policy, but many institutions also have internal authorship guidelines research teams are expected to follow, and journals may ask for a contribution statement (CRediT or narrative) at submission that makes the underlying reasoning visible to editors and readers.

Related CASRAI resources

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →