Forensic Tools for Content Provenance
Why Content Provenance Matters for Research
Content Provenance
The rise of automated and AI-enabled image manipulation tools jeopardizes research integrity and creates challenges for scientific progress and public trust. Manipulation tools are becoming more powerful and accessible, and altered images harder to detect. New software can create completely AI-fabricated images, which can be indistinguishable from authentic data.
When manipulations are discovered, they may not only damage the reputations of the individuals involved but also cast doubt over the scientific enterprise, undermining the trust that the scientific community and the public rely on.
Content provenance is the ability to create a verifiable record of where a piece of digital content originated, any modifications made, and where it has been published.1
How Content Provenance Protects Data
Adopting a content provenance solution offers several benefits:
- Asserting Authenticity and Integrity: Content provenance provides a verifiable record that data are authentic. A report from the International Association of Scientific, Technical, and Medical Publishers highlights that content provenance allows users to demonstrate that an image originated from a specific instrument at a specific time, providing a defense against accusations of misconduct.
- Enhancing Trust with Publishers and Funders: Submitting data with a clear provenance trail demonstrates a commitment to transparency.
- Combatting Broader Research Misconduct: Widespread adoption creates a deterrent against bad actors. When publishers, funders, or researchers can quickly verify the provenance of submitted images, it becomes significantly more difficult for fabricated content to enter the scholarly record.
- Future-Proofing Research: As AI-driven misinformation becomes more common, provenance solutions may also become more common, and the absence of provenance may then raise concern. Adopting these tools helps future-proof data and position research as a source of accurate information.
Content Provenance Solution: C2PA Content Credentials
One available content provenance tool is the Content Credential, offered by the Coalition for Content Provenance and Authenticity (C2PA). Content Credentials, explained here by C2PA, provide a verifiable record by adding metadata. When a user creates a file in a C2PA-enabled software, the Content Credential is created and attached to the file. Then it is updated every time the file is edited. The Content Credential can answer questions like:
Who created or modified this file?
When was it created and last edited?
What software or device was used to create or modify it?
What modifications were made (e.g., color adjustments, cropping)?
This information is cryptographically signed to ensure that it cannot be altered without detection. It acts as a tamper-evident, verifiable trust signal embedded in data files.
Content Provenance Solution: Google SynthID
Another available digital provenance tool is Google DeepMind’s SynthID. SynthID provides a durable identification layer by embedding an invisible digital watermark directly into the pixels of AI-generated images and the waveforms of AI-generated audio, rather than appending it to the file's metadata. Further information is available from Google here.
Users can verify the presence of a SynthID watermark by analyzing the image through Google's proprietary verification systems, such as the SynthID Detector portal available here. This embedded pixel data helps answer the question: “Was this image generated by a supported AI model?” However, absence of a SynthID watermark does not guarantee that the file was not created by another, unsupported AI model.
Because this system modifies the image at the pixel level, the watermark is robust to most downstream edits, including cropping, resizing, heavy JPEG compression, color adjustments, and even taking a screenshot of the image (a form of image-to-image reproduction).
We highlight Google’s SynthID because it has been adopted by multiple generative AI companies, including OpenAI, NVIDIA, and ElevenLabs. Other generative AI companies also have watermarks, and these watermarks tend to work similarly, but can typically only be verified by the company that produced the content. Information about watermarks including the SynthID watermark is provided for informational purposes only and should not be construed as an endorsement.
Understanding the Current Limitations
While content provenance solutions present a significant step forward, they are not a silver bullet. This table explains some current challenges.
Challenge | Description | Potential Mitigation |
Metadata Stripping | Some online platforms automatically strip metadata from files to save bandwidth or for privacy reasons. This action would render content unverifiable. | Always submit original, full-resolution files directly to journals or repositories. Use institutional repositories designed to preserve metadata. |
Incomplete Adoption | The value of a content provenance solution relies on "end-to-end" adoption. If an instrument does not create a provenance record, the provenance chain is incomplete. Likewise, if publishers do not have workflows to read the credentials, their value is lost at the final step. | Inquire with journals about provenance verification in their submission systems. |
Image-to-Image Reproduction | Content Credentials attach to the digital file, not the visual information itself. Other content provenance solutions, such as Google’s SynthID attach to the image. When provenance information is tied to a file instead of an image, bad actors can take pictures of images and pass these reproductions off as original work. | Foster an environment where the absence of a verifiable provenance credential on a submitted piece of data is viewed with skepticism and triggers further scrutiny. Consider combining provenance solutions. |
Context vs. Content | Content Credentials can demonstrate that an edit was made, but not why. It cannot distinguish between a legitimate adjustment to improve clarity and a questionable manipulation. | View content provenance solutions as tools that enable and accelerate expert review, rather than replacing it. The provenance trail provides reviewers with the data they need to assess whether edits are scientifically appropriate. |
How to Get Started with Content Provenance
Enabling Google SynthID Watermarks
SynthID watermarks are automatically embedded in content produced by participating models and do not need to be turned on manually.
Enabling C2PA Content Credentials
Content Credentials are not automatically enabled. Steps for enabling them in the Adobe suite are reproduced below, and more information is available from C2PA.
For Adobe Photoshop® (Desktop Application)
This option ensures Content Credentials are automatically applied to data and files in Adobe Photoshop®.
- Open the Photoshop application.
- Access Preferences
- Windows: Go to Edit > Preferences
- Mac: Go to Photoshop > Preferences
- Navigate to the History & Content Credentials submenu
- Enable the feature
- Select “Enable for saved documents with Content Credentials” to enable Content Credentials to update when documents that already have them are edited
- Select “Enable for new and saved documents with Content Credentials” to enable Content Credentials on all documents
- Click "OK" to save
For Adobe Creative Cloud
Users with Adobe Content Authenticity as part of Adobe Creative Cloud, can adopt this option to apply Content Credentials to photos.
- Sign in to Adobe at contentauthenticity.adobe.com
- Select apply at the top
- Upload up to 20 photos at a time
- Select apply
Verifying Data
The free, open-source C2PA Viewer tool allows users to upload a file to display its origin, creator, and edit history. This tool can be used to confirm data are signed before submission or to inspect the provenance of data from other sources.
The C2PA Viewer tool checks for Content Credentials but does not check for SynthID or other watermarks. To determine whether an image or other file contains a watermark, users must generally upload the image to proprietary tools such as the SynthID Detector portal.
The presence of Content Credentials and/or watermarks are positive indicators about the provenance trail of an uploaded file. The absence of either indicator, however, does not prove that a file is authentic. For example, an image generated by a non-participating software would not show a SynthID watermark but may still be synthetic. Users should take care to understand the scope of the provenance claims that any one verification tool can provide.
Adopting a content provenance solution helps protect research and usher in a new standard of transparency and trust for publicly funded science.
Disclaimer
Identification or discussion of specific content provenance tools or implementing software is for information only. It is not intended to imply recommendation or endorsement by the Office of Research Integrity or any U.S. Government agency, nor is it intended to imply that the tools or software are necessarily the best available for content provenance.
1. See a detailed explanation of provenance features and use cases from the UK and Canadian governments at https://www.ncsc.gov.uk/collection/public-content-provenance-for-organisations and a survey of content provenance benefits and risks from the Center for Democracy and Technology at https://cdt.org/insights/the-promise-and-risk-of-digital-content-provenance/.
