Skip to main content
v2026.11,858 entries · CC-BY 4.0

C2PA Content Credentials: How Provenance Actually Works

Content Credentials are often described as a nutrition label for digital media. The metaphor oversells them. C2PA is a signed-assertion format with a cryptographic binding to specific bytes: it can tell you that a particular signer said a particular thing about a particular file and that the file has not changed since. It cannot tell you that what the signer said is true, and the chain breaks on the ordinary act of uploading a picture to a website. This guide walks the actual data model, the trust machinery that decides whose signature counts, the documented failure modes, and where the EU AI Act’s Article 50 marking duty does and does not meet the standard.

Written and maintained by CASRAI Editorial Board

Last updated

Last verified against primary sources: 25 September 2026. “Content Credentials” is the user-facing name for provenance data built to the specifications of the Coalition for Content Provenance and Authenticity (C2PA), a project of the Joint Development Foundation, an affiliate of the Linux Foundation. The technology is routinely described in press coverage as a nutrition label for digital media, and that metaphor is the single biggest source of confusion about what it does. A nutrition label implies an audited statement of fact. A Content Credential is something narrower and, once you understand it, more useful: a set of statements signed by an identifiable party, cryptographically bound to a specific sequence of bytes.

That distinction is not a quibble invented here. It is the position of the specification itself. The C2PA technical specification states that C2PA specifications “SHOULD NOT provide value judgments about whether a given set of provenance data is ‘good’ or ‘bad,’ merely whether the assertions included within can be validated as associated with the underlying asset, correctly formed, and free from tampering.” The companion explainer is blunter still: “provenance information alone cannot tell you whether the digital content is true, accurate or factual.”

This guide walks the data model as the specification defines it, explains where the cryptography actually reaches and where it stops, sets out the trust machinery that decides whose signature counts, summarises the documented failure modes from recent security analysis, and locates the standard against the EU AI Act’s Article 50 marking duty. It closes with the one place a research-administration office is likely to encounter this in practice, because a US federal research-integrity body has already published guidance pointing at it.

The Data Model: Assertion, Claim, Signature, Manifest

Almost every misunderstanding of C2PA dissolves once the four-layer structure is clear. The specification defines them as follows (section numbers are from technical specification 2.4, released April 2026; the definitions are materially unchanged from 2.2, released May 2025):

Layer Specification definition What it means in practice
Assertion “A data structure which represents a statement either made by the signer or gathered at claim generation-time, concerning the asset” (2.3.1) One discrete statement: this was captured by this camera model, this was generated by an AI system, these editing actions were applied, this creator does not permit AI training use
Claim “A digitally signed and tamper-evident data structure that references a set of assertions” (2.3.2) The bundle. It gathers the assertions and the content-binding information as they stood at one moment
Claim signature “The digital signature on the claim created using the private key owned by a signer” (2.3.3) The cryptographic commitment. Since specification 2.0, only X.509 certificates may be used for signing
C2PA Manifest The combination of one or more assertions, a single claim, and a claim signature (2.3.4) One verifiable unit of provenance, corresponding to one point in the asset’s history
Manifest Store “A collection of C2PA Manifests that can either be embedded into an asset or be external” (2.3.5) The accumulated history. The store, not any single manifest, is the provenance record
Active Manifest “The last manifest in the list…which is the one with the set of content bindings that are able to be validated” (2.3.7) The manifest that describes the file as it exists right now. Earlier manifests describe earlier states

The component that produces all of this is the claim generator: the hardware or software that assembles the assertions, constructs the claim, and obtains the signature. A camera body, an image editor, or a generative model’s serving infrastructure can each be a claim generator.

Read the layers in order and the honest reading of a Content Credential falls out. A validator can establish that signer S committed to assertion set A over byte-range B. Everything a reader then wants to conclude — that the photograph depicts what it appears to depict, that the AI-generation assertion is complete, that the edit history is the whole edit history — rests on what you already believe about signer S. The cryptography carries the binding. It does not carry the truth.

Hard Bindings and Soft Bindings: Where the Cryptography Reaches

The link between the signed statements and the actual file is called a content binding, and the specification defines two kinds that behave very differently.

A hard binding is defined as “one or more cryptographic hashes that uniquely identifies either the entire asset or a portion thereof” (2.3.12). The specification supports byte-range hashing for arbitrary formats, box-based hashing for non-BMFF formats such as JPEG, PNG and GIF, BMFF-based hashing for ISO Base Media File Format files, and collection hashing for multi-asset workflows. A manifest “shall not contain more than one assertion defining a hard binding but may contain zero or more assertions defining soft bindings.”

A hard binding is the strong part of the system and also its most brittle part. It is a hash. Any re-encoding, recompression, resizing or format conversion changes the bytes and therefore breaks the hash, even when the image on screen is visually identical. That is not a defect in the hashing; it is what hashing means. But it collides badly with how media actually moves across the internet, where recompression in a content-delivery pipeline is the normal case rather than the exception.

A soft binding is the specification’s answer to that collision: “a content identifier that is either (a) not statistically unique, such as a fingerprint, or (b) embedded as an invisible watermark” (2.3.13). Soft bindings survive transformations that destroy a hash, which is precisely why they are weaker. They are not unique by construction, and the specification itself acknowledges their exposure to collision-based attack.

Soft bindings underpin what C2PA calls durable Content Credentials: manifests discoverable in an external repository by fingerprint or watermark even after the embedded metadata has been stripped from the file. The C2PA FAQ describes this as support for rediscovering an associated Content Credential “even if it’s removed from the file.” This is a genuine mitigation and it is worth understanding correctly: durability comes from a lookup, not from the signature surviving. The strength of the guarantee drops from these exact bytes were signed to something statistically similar to this was signed once.

What Happens When the File Is Edited

C2PA is designed around the assumption that media is edited, not around the assumption that it is pristine. When a conforming tool modifies an asset, three things happen: the previous manifests are preserved in the manifest store, an ingredient assertion records the source material that was used, and a new manifest is appended describing the new state. As the specification puts it, each time an asset is changed the existing provenance is preserved, with each new change added to it.

The practical consequence is that a Content Credential chain is a record of declared edits by conforming tools, not a record of all edits. A file that passes through a non-conforming editor simply arrives at the next conforming tool as an ingredient with no prior provenance. Specification 2.4 added a digitalSourceType capability for ingredients that carry no C2PA manifest of their own, which is an improvement in expressiveness — it lets a tool say “this input came from somewhere I cannot vouch for” — but it does not reconstruct the missing history.

Validation: Well-Formed, Valid, Trusted

The specification separates three questions that ordinary usage collapses into one word, “verified”:

State The question it answers What it does not answer
Well-formed Is the manifest structurally correct according to the format rules? Whether the signature checks out
Valid Is it well-formed, with verified signatures, and do the content bindings match the asset? Whether anyone should believe the signer
Trusted Is the signer on a trust list the validator recognises? Whether the assertions are factually accurate

The third row is where the governance lives, and it is the row most public discussion skips. “Trusted” is not a property of the file. It is a property of a policy decision made by whichever validator you happen to be using. Trust is rooted in the identity of the signer associated with the signing key, and validators may use the default C2PA Trust List of X.509 trust anchors or a trust list configured locally. Two validators with different trust lists can look at the same valid manifest and reach opposite conclusions, entirely correctly.

The Trust List and the Conformance Program

Deciding whose certificates belong on the list is an institutional problem, not a cryptographic one, and C2PA built institutions for it comparatively late. The C2PA Conformance Program and the official C2PA Trust List both launched in mid-2025. C2PA describes the program as a risk-based, transparent and unbiased governance process intended to ensure that products comply with the Content Credentials specification and meet the security requirements set out in the separate C2PA Generator Product Security Requirements document.

Before that, the ecosystem ran on an Interim Trust List (ITL), a temporary measure for early implementations. Its retirement matters for anyone validating older assets. C2PA’s conformance page records that the ITL remained operational and accepted new certificates through 31 December 2025, and that on 1 January 2026 the ITL was frozen: “No new entries will be added, and no updates will be made.” The program is oriented around the 2.x specification series, with the eventual sunset of 1.x-based certificates as a stated goal. Current entries in the Conforming Products List, the C2PA Trust List and the timestamp-authority trust list are published through the C2PA Conformance Explorer.

So the practical state of affairs in 2026 is: an official trust list exists, it is governed by a program that opened in mid-2025 and is therefore still filling, and the interim arrangement that carried the ecosystem through its first years is frozen but not deleted. A validator’s behaviour on a three-year-old asset depends heavily on which of these it consults.

Standardisation Status: ISO/DIS 22144

C2PA’s version 2 architecture has been put through the ISO fast-track route as ISO/DIS 22144, Authenticity of information — Content credentials. The draft describes the technical aspects of the Content Credentials architecture, including how to create and process a C2PA Manifest and its components and the use of digital signature technology for tamper-evidence and trust establishment.

The important qualifier is the letters DIS: Draft International Standard. As of this writing the project is in progress and has not been published as a final International Standard, and the fast-track ballot drew substantive comments — the Association for Information Science and Technology, for instance, voted against with comments seeking revision of the normative-references chapter. Anyone drafting a procurement clause or a policy that cites “the ISO standard for content credentials” should check the current stage at ISO before relying on the citation.

Where Regulation Actually Meets This

The most common claim about C2PA in compliance writing is that it is “the standard behind the EU AI Act.” That claim does not survive contact with the texts.

Article 50 of the EU AI Act does impose a machine-readable marking duty. The European Commission’s own summary of the transparency rules states that providers must apply “a machine-readable mark to synthetic content generated or manipulated by AI,” and that deployers must mark deepfakes and text published on matters of public interest with “clear and perceivable labels.” The Commission records that these transparency rules apply from 2 August 2026, with a grace period for the marking obligation until December 2026 for generative AI systems placed on the market before 2 August 2026. Systems placed on the market on or after that date get no transitional window.

What the Commission does not do is name a technology. Its transparency fact page references Commission Guidelines on Article 50, a Code of Practice on Transparency of AI-generated Content, and three optional labelling icons — and specifies no particular technical standard. The Code of Practice itself, developed through a process that opened with a kick-off plenary on 5 November 2025, produced a first draft on 17 December 2025 and a second on 3 March 2026 before final publication on 10 June 2026, likewise refers to technical solutions generically without naming C2PA, watermarking or fingerprinting. Adherence to the Code is voluntary even though the Article 50 obligations it addresses are legal ones; the Commission has confirmed it as an adequate voluntary tool for demonstrating compliance.

The honest formulation is therefore: Article 50 creates a marking obligation that C2PA is a plausible way to discharge, and no EU instrument obliges anyone to use it. Other jurisdictions have drafted closer to the mechanism — California’s AI Transparency Act (SB 942/AB 853), which does specify provenance metadata alongside latent disclosures, and China’s 2025 labelling rules, which distinguish explicit from implicit labels, are both more prescriptive about the artefact than the AI Act is. For the wider picture of what disclosure law actually compels versus what publisher policy asks for, see our guide to AI disclosure laws and how they differ from journal policy.

The Documented Failure Modes

Three distinct bodies of evidence bear on how well this works in deployment, and they should be read together rather than traded off against each other.

The stripping problem is mundane and pervasive. Writing in The Scholarly Kitchen on 13 March 2025, Todd A Carpenter reported that his own test image lost its credentials simply by being processed through WordPress, and noted that it is easy to strip the metadata “by screen captures or images of images.” This is not an exotic attack. It is what happens when media passes through ordinary publishing infrastructure. Carpenter’s second observation is arguably more damaging: in a widely discussed case, images did carry C2PA data and nobody examined it. A provenance signal that no one checks provides no assurance regardless of how sound its cryptography is.

Government cyber agencies have treated it as promising but requiring care. On 29 January 2025, the NSA, together with the Australian Signals Directorate’s Australian Cyber Security Centre, the Canadian Centre for Cyber Security and the UK National Cyber Security Centre, issued a Cybersecurity Information Sheet on Content Credentials, Strengthening Multimedia Integrity in the Generative AI Era. It sets out questions organisations should work through before implementation and recommends practices to ensure the preservation of unaltered metadata throughout the media lifecycle — framing the technology as an evolving standard that can significantly increase transparency of media provenance, rather than a finished control.

Security analysis of the specifications has been sharply critical. In Verifying Provenance of Digital Media: Why the C2PA Specifications Fall Short (arXiv:2604.24890, 23 April 2026), Enis Golaszewski and ten co-authors, with affiliations including UMBC, Hacker Factor and NSA, argue that the specifications as written do not achieve their claimed security goals. Their documented categories include:

Failure category The finding
Unprotected timestamps “Nothing in the signed data references the timestamp, allowing removal and replacement without detection”
Revocation gaps Revocation checking is optional, so validators may accept revoked or compromised credentials
Validator inconsistency Different tools return contradictory validity assessments of identical content
Partial file protection Exclusion ranges permit undetected modification of excluded metadata, GPS location among the examples given
Certificate expiry Signed media can become unverifiable within months, which conflicts with legal and archival retention needs
Weak certification Conformance rests on self-reported compliance without source-code examination

Their recommendations run to mandating strict revocation checking, making timestamps tamper-evident, requiring consistent validation behaviour across tools, protecting whole files rather than selected portions, introducing independent security audits of certified products, and clarifying C2PA’s actual limitations in public communications. That last item is a communications finding, not a cryptographic one, and it is the one this guide is most directly a response to.

None of this makes Content Credentials worthless. It makes them a transparency mechanism with a known and bounded assurance level, which is a different thing from a proof of authenticity. It is worth contrasting the failure mode with that of post-hoc AI-detection classifiers: a detector that is wrong produces a confident false accusation, whereas a provenance chain that is broken produces an absence of information. Absence of a credential is not evidence of manipulation, and any policy built on this technology has to state that explicitly or it will generate exactly the false accusations it was meant to prevent.

The Research-Administration Angle: ORI Has Already Pointed At This

There is one place where research administrators are likely to meet Content Credentials as an operational question rather than an abstract policy one, and it comes with published federal guidance.

The US Department of Health and Human Services Office of Research Integrity (ORI) maintains a page on content provenance that identifies C2PA Content Credentials as, in its words, “one available content provenance tool” that provides a verifiable record by adding metadata. ORI is careful about the status of that mention, and the careful wording is itself instructive. The page states that identification or discussion of specific content provenance tools or implementing software is “for information only. It is not intended to imply recommendation or endorsement by the Office of Research Integrity or any U.S. Government agency.”

ORI’s practical guidance to institutions and researchers is closer to workflow hygiene than to technology adoption:

  • Enable content provenance features in the software actually used to prepare figures
  • Verify provenance data using a free verification tool before submission
  • Submit original, full-resolution files directly to journals, rather than versions that have been through intermediate processing
  • Ask journals what provenance verification they are capable of performing
  • Treat provenance solutions as complementary to expert scientific review, not a replacement for it

Item three is the operationally important one, and it is a direct consequence of the hard-binding mechanics described above. Every intermediate step that recompresses an image — pasting into a slide deck, exporting from a shared drive with transformation enabled, routing through a submission portal that resizes — breaks the hash. An institution that wants figure provenance to survive to the publisher has to look at its whole manuscript-preparation path, not just at whether Photoshop has the feature switched on.

The coordination burden is the honest obstacle. Carpenter’s assessment is that implementation demands significant coordination and collaboration across the entire research community, from device manufacturers through the software that processes data and produces summary images to article-submission vendors and publishers. That is a long chain, and a provenance chain is only as good as its weakest conforming link. Research-data teams will recognise the underlying discipline from a related context: fixity checking for archived research data solves a narrower version of the same problem, establishing that stored bytes have not changed, without attempting to say anything about who made them or what they depict.

Reading a Content Credential Correctly: A Short Checklist

If you are assessing a file, or writing a policy that asks someone else to:

  • Ask which validation state you were given. Well-formed, valid and trusted are three different results. A tool that reports a green tick without saying which one it means is not giving you enough to act on.
  • Ask whose trust list produced “trusted.” The answer differs between the official C2PA Trust List, a locally configured list, and the frozen Interim Trust List still relevant to older assets.
  • Distinguish hard from soft binding. A hard-binding match is a statement about exact bytes. A soft-binding match is a statement about similarity, with the collision exposure the specification acknowledges.
  • Read the assertions, not the badge. The badge means a signature verified. The assertions are the actual content, and they are claims by the signer, not findings by an auditor.
  • Never treat absence as evidence. Most media has no credential, and ordinary publishing infrastructure removes credentials from media that had them.
  • Do not write “C2PA-compliant” into a contract without a version and a conformance reference. The specification moved from 2.2 in May 2025 to 2.4 in April 2026, the conformance program opened only in mid-2025, and the ISO draft is not yet a published standard.

Sources

  • C2PA, Content Credentials: C2PA Technical Specification, versions 2.2 (May 2025) and 2.4 (April 2026) — spec.c2pa.org
  • C2PA, C2PA and Content Credentials Explainer, version 2.2 — spec.c2pa.org
  • C2PA, Conformance Program and Conformance Explorer — c2pa.org/conformance
  • C2PA, Frequently Asked Questions — c2pa.org/faqs
  • European Commission, Quick Facts: Transparency rules for AI systems and Code of Practice on Transparency of AI-generated Content — digital-strategy.ec.europa.eu
  • US HHS Office of Research Integrity, Content Provenance — ori.hhs.gov/content-provenance
  • E. Golaszewski et al., Verifying Provenance of Digital Media: Why the C2PA Specifications Fall Short, arXiv:2604.24890, 23 April 2026
  • NSA, ACSC, CCCS and NCSC, Content Credentials: Strengthening Multimedia Integrity in the Generative AI Era, Cybersecurity Information Sheet, 29 January 2025
  • T. A. Carpenter, Research Integrity, Image Manipulation, Content Provenance and the C2PA, The Scholarly Kitchen, 13 March 2025
  • ISO/DIS 22144, Authenticity of information — Content credentials (draft, in progress)

This page describes a technical specification and the regulatory instruments that touch it. It is not legal advice, and it is not a conformance assessment of any product. Specification versions, trust-list contents and regulatory deadlines all move; confirm against the linked primary sources before relying on any of it for a compliance decision.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about C2PA Content Credentials: How Provenance Actually Works

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

An AI assistant specialized in research administration. It cites the sources behind every answer, labels web answers and says when it can't answer.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

Research-admin question? Get an answer that links its sources.

An AI assistant specialized in research administration. Every answer links its sources to check before you act. 2 questions free, no account. $29/month after.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.