A data paper (Scientific Data calls its version a Data Descriptor) is reviewed on the quality and reusability of the dataset it documents, not on the novelty of a scientific finding. That difference in review focus creates a genuine authorship question that the standard research-article model does not answer cleanly: generating, curating, and describing a dataset involves people whose work looks very different from designing an experiment or drafting a discussion section — sample collectors, instrument operators, repository curators, data managers, statisticians who cleaned the file. Some of that work clears the bar for authorship. A meaningful amount of it, done well, still does not. This guide sets out how to tell the two apart, using Scientific Data’s own published policy as the primary reference point, since it is the most widely used dedicated data-paper venue and its authorship policy is representative of how data journals generally handle the question.
The baseline: Scientific Data follows the same four-part test as any other journal
Scientific Data (published by Nature Portfolio / Springer Nature) does not run a separate, looser authorship standard for data papers. Its editorial policies state that the journal follows the Nature Portfolio authorship policy, which is built on the ICMJE authorship criteria. Under that framework, authorship requires satisfying all four of the following, not just one or two:
- Substantial contribution to the conception or design of the work, or the acquisition, analysis, or interpretation of data for the work;
- Drafting the work or reviewing/revising it critically for important intellectual content;
- Final approval of the version to be published; and
- Agreement to be accountable for all aspects of the work, including ensuring that questions about its accuracy or integrity are appropriately investigated and resolved.
ICMJE is explicit that these are joined by AND, not OR: meeting only one or two of the four does not qualify someone for authorship, however important that single contribution was. See ICMJE Authorship Criteria: The 4 Requirements for the full test applied outside the data-paper context, and the ICMJE dictionary entry for background on the committee itself.
What makes data papers distinct is not a different authorship test — it is that the underlying activities the test gets applied to are different. Scientific Data’s own guidance notes that acceptance is not based on the perceived novelty or impact of scientific conclusions drawn from the dataset, and Data Descriptors are explicitly not expected to contain in-depth analysis or new scientific findings. That reframes what “substantial contribution to acquisition, analysis, or interpretation of data” and “drafting the work” mean in practice for this content type.
Who typically qualifies as an author on a data paper
Applying the four-part test to the actual work of producing a dataset and its descriptor, the people who most often clear all four criteria are those who:
- Designed the data-generation protocol — determined what would be measured, sampled, or collected, and how, in a way that shapes the resulting dataset’s structure and scope (this is “conception or design of the work” for a data paper).
- Substantially performed or supervised the data collection, processing, or quality-control pipeline that produced the dataset described — not routine execution of someone else’s fully specified protocol with no analytical input, but the acquisition/processing work with real judgment involved.
- Curated and validated the dataset for reuse — checked completeness, resolved quality issues, structured the metadata so the dataset is actually usable by someone else. This overlaps closely with the CRediT Data Curation role, though contributorship credit and authorship are separate determinations (see below).
- Drafted or substantively revised the Data Descriptor text — the methods narrative, technical validation section, and usage notes that make the dataset interpretable, not just a light copy-edit pass.
- Reviewed and approved the final submitted version and are willing to be accountable for the dataset’s accuracy and integrity if a problem surfaces after publication.
In a typical multi-person data-generation project, this usually converges on a smaller group than everyone who ever touched the data: the person(s) who designed the collection protocol, whoever ran and quality-controlled the pipeline with real technical judgment, and whoever wrote and stands behind the descriptor.
Who is typically a data contributor rather than an author
The following roles are common on data-generation projects and are real, valuable contributions — but on their own, without also meeting the other three ICMJE criteria, they do not clear the authorship bar:
- Providing access to a facility, instrument, or existing sample/specimen collection without participating in the design, curation, or write-up of the resulting dataset. Granting access is a form of Resources under CRediT, not on its own conception/design or drafting.
- Technical staff who executed a fully specified protocol with no discretionary input into design, analysis, or curation — running the assay, operating the sequencer, or performing the fieldwork exactly as instructed.
- Repository or infrastructure staff who deposited, hosted, or assigned identifiers to the dataset as part of their institutional role, without contributing to its scientific design or intellectual content.
- Funders and funding bodies, credited in a funding statement rather than the author list, per standard ICMJE guidance on acknowledgments.
- Statisticians, data managers, or writers who performed a bounded technical task (cleaning a file, formatting for submission) without drafting or substantively revising the descriptor and without accountability for the dataset as a whole.
ICMJE’s own recommendation for this situation is explicit: contributors who do not meet all four authorship criteria should not be listed as authors, but they should be acknowledged. Scientific Data supports exactly this distinction — a contributor who does not qualify for authorship can and should still be named, either in an acknowledgments section or, where the dataset itself credits specific roles, in the paper’s contributions statement. Omitting a real contributor rather than acknowledging them correctly is itself a policy problem, not a safer default. See Acknowledgments Section and the comparison Acknowledgments vs. Authorship: Who Should Be Listed and Why for how that distinction is documented in practice.
Why CRediT contribution statements do not replace this determination
Many data papers, including many published in Scientific Data, now carry a CRediT contribution statement alongside the byline, listing which of the 14 CRediT contributor roles (formalized as ANSI/NISO Z39.104-2022) each named person held — commonly Data Curation, Investigation, Methodology, and Writing – Original Draft for a typical data paper. A CRediT statement is useful precisely because it records granular contribution detail that a flat author list cannot — but it answers a different question than authorship does. CRediT records what a listed person did; it does not itself decide whether that person should be listed as an author in the first place. It is possible, and normal, for someone to appear in a data paper’s CRediT statement (for example, credited with the Data Curation role) while not independently meeting all four ICMJE authorship criteria, if their contribution stopped short of drafting/revising the text and taking accountability for it. The authorship determination and the CRediT statement are made separately, even though they are usually published together.
A practical decision checklist
For each person proposed for the author list on a data paper, confirm all four are true, not just the ones that feel most obviously satisfied:
- Did they substantially shape the dataset’s design/scope, or substantially contribute to acquiring, processing, or interpreting the data — beyond executing someone else’s fully specified instructions?
- Did they draft part of the Data Descriptor, or critically revise it for important intellectual content (methods accuracy, technical validation, usage notes) — not just proofread it?
- Have they reviewed and approved the final submitted version?
- Are they willing to be accountable if a data-quality or integrity question about this dataset comes up after publication?
Anyone who fails even one of the four belongs in the acknowledgments section, a named data-contributors note, or the CRediT statement without being listed as an author — not omitted, but also not credited as an author for a contribution that does not meet the full test. This is the same logic CASRAI’s broader authorship guidance applies outside the data-paper context; see Types of Authorship in Research for how it plays out for standard manuscripts, and ICMJE Authorship Criteria: The 4 Requirements for the underlying test in full.
Data papers vs. the data underlying a separate research article
This authorship question is specific to data papers as their own citable, peer-reviewed output — not to the (more common) situation where a dataset simply underlies a conventional research article’s findings. When a dataset is deposited alongside a standard article rather than described in its own data paper, the people who generated that data are usually already covered by the article’s own authorship determination, and a separate dataset-specific author list does not apply. For the broader distinction between publishing a dataset as a citable data paper versus depositing it as a repository record with no accompanying peer-reviewed narrative, see Data Papers vs. Dataset Records: Data Journals vs. Repository Deposits and the overview guide Data Publication Practices: Repositories, Data Journals, and Data Papers.
Large or consortium data-generation projects
Multi-site data-generation efforts (biobanks, large sequencing consortia, long-running monitoring networks) raise this question at scale: dozens or hundreds of people may have touched the dataset in some capacity. The same four-criteria test applies regardless of project size — it does not loosen because the contributor pool is large. Two mechanisms commonly used to keep the author list accurate without either inflating it or erasing real contributions are a named group/consortium listing (crediting a group name on the byline while identifying which named individuals within it meet full authorship criteria) and a separate, explicit data-contributors list distinct from the author list, naming individuals or sites that supplied data or samples without meeting the full test. Both approaches keep the distinction between “contributed data” and “authored the paper describing it” visible rather than collapsing the two.
Frequently asked questions
Does providing the raw dataset automatically make someone an author on the data paper?
No. Supplying data, samples, or instrument access is a real contribution (typically the CRediT Resources role), but on its own it does not satisfy the design/acquisition-and-analysis, drafting, approval, and accountability criteria together. It should be acknowledged, not silently omitted.
Does Scientific Data use a different authorship standard than a standard research journal?
No. Scientific Data applies the same Nature Portfolio authorship policy, built on the ICMJE four-criteria test, that applies to its other journals. What differs is the nature of the underlying work being evaluated against that test, since a Data Descriptor is not expected to contain novel scientific analysis.
Can someone be in the CRediT statement but not listed as an author?
Yes, and it is common. A CRediT role credits a specific type of contribution; it does not by itself establish that the person meets all four ICMJE authorship criteria. The two determinations are related but independent.
Where should a data contributor who does not qualify for authorship be listed?
In the acknowledgments section, or, on projects that use one, a named data-contributors note distinct from the byline. See Acknowledgments Section.
Who should be the corresponding author on a data paper?
Journal practice generally follows the same expectations as any article: someone accountable for the dataset and the descriptor, typically someone meeting all four authorship criteria who is available to respond to post-publication queries about the data. See Corresponding Author.







