Understanding AI Marks: Watermarking, Provenance, and What They Mean for Information Management
Intelligent Information Management (IIM) | Artificial Intelligence (AI)
Social media has been abuzz this week about Anthropic's announcement about its approach to marking AI-generated content, and we wanted to follow up with what this may mean for information professionals.
What are watermarks or marks?
AI marks (or AI watermarks) are invisible, machine-readable signals embedded directly into text, images, or audio generated by artificial intelligence.
Why is watermarking useful?
AIIM has advocated for source citations and watermarking and was pleased that the administration is committed to further research into authenticating, labeling, and detecting synthetic data (i.e., generative AI output). In our past policy statements and comments, we have supported watermarking of AI-generated output so consumers can use software to detect AI output.
AIIM's support of watermarking comes down to a few connected reasons:
- Consumer certainty and trust. Watermarking gives "users of AI-generated information more certainty and awareness" so they can judge whether output is accurate and credible. If people can tell what's AI-generated, they can calibrate how much to trust it.
- Shared responsibility between developers and consumers. AIIM's position is that accuracy and trust in AI output isn't just on the user to figure out. Developers have tools available to make that job easier, and watermarking is one valid method for accomplishing that.
- Combating disinformation without touching free speech. Watermarking is a labeling mechanism, not a content restriction. It protects free speech while still giving the public a way to protect themselves.
What are the different types of AI marks?
Marks can take different forms. The EU AI Act Recital 133 notes that marks can include "watermarks, metadata identifications, cryptographic methods for proving provenance and authenticity of content, logging methods, fingerprints or other techniques."
There are a few common approaches for marking, and these can be used in combination.
- Technical markers (embedded watermarks) are signals baked into the content itself, at the level of the words, pixels, or audio waveform. Technical watermarking asks: can software detect that this content was generated or altered? It does this by hiding a signal inside the content itself.
- Embedded pixel-level watermarks encode a signal directly into the pixel values of an image, rather than attaching it as separate metadata.
- Provenance markers (signed metadata) ask a different question: can I trace this content's history? Rather than altering the content, they attach a record alongside it, a kind of certificate that says where the file came from and whether it's been tampered with. This is more akin to a chain-of-custody log.
For provenance markers, Anthropic and the EU AI Act specifically reference the C2PA standard. C2PA (the Coalition for Content Provenance and Authenticity) is an open technical standard that attaches signed provenance data to digital content, backed by companies including Adobe, Microsoft, and Google.
Are AI marks required?
The EU AI Act's Article 50(2) requires that "[providers] of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated." This is the regulation Anthropic is specifically complying with, as they signed the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generated Content.
However, the EU AI Act isn't the only regulation requiring AI marks. China's Generative AI Measures and California SB 942 (the California AI Transparency Act) also contain requirements for AI marks.
How do I detect AI marks?
The EU's Code of Practice on Transparency of AI-Generated Content, which Anthropic and other major providers have signed, requires providers to make detection mechanisms available so users and relevant stakeholders can verify whether content has been generated or manipulated by AI.
Do AI marks devalue data?
There are many debates on LinkedIn right now about whether writers and CEOs should avoid generative AI tools that mark their work as "AI processed." The question of whether to use AI in communications, and whether to disclose that usage, is a broader debate.
The role of an information professional isn't to judge the value of data based on how it was developed. We value content (and the need to preserve it) based on its purpose, use, and the decisions it generates.
What's important to an information professional is documenting the use of AI. When I spoke with AIIM Chairman Jason Cassidy about AI marks, he used his experience playing guitar to explain his thinking. "The guitar strings still produce sound, but the type of guitar you use will change the vibrating sound of the strings," Cassidy explained. "I don't care if it was processed by AI, I care if it helped me solve the problem."
Are AI marks and detection ready for enterprise operations?
AI marks are a step in the right direction for transparency, but AI marks and mark detection are still immature. At an enterprise level, I'm not convinced we're prepared to use AI marks or mark detection mechanisms effectively.
AI marks are provider-specific and lack universal standardization. C2PA is a shared, cross-vendor standard for provenance metadata, but it is optional to use. An enterprise that uses Claude, ChatGPT, Gemini, and various image and video generators would need to either integrate multiple provider-specific detection tools, or rely on third-party aggregator platforms if and when those mature enough to check across vendors.
Back in 2023, AIIM recognized this in our advocacy work: watermarking wasn't mature enough to implement at scale. The Article 50 mandate ensures that providers will build or make available detection tools, but it doesn't address how an organization will feasibly detect AI and verify content. That's a major gap.
Information professionals should become acquainted with the available regulations and standards and begin to lead these conversations about AI marks within their own organizations.
-
When is it important for your organization to know when content was processed by AI?
-
What are the processes for identifying AI content?
What questions would you add to the list? Join the conversation in our member-only online forum. Not a member yet? Learn more about membership.
AI Disclosure
Oh, and in the interest of transparency: I originally wrote this blog post in OneNote like a complete weirdo, and then used Claude to refine (aka correct) my horrendous grammar and spelling. I also used Claude to verify the accuracy of my research.
About Tori Miller Liu, CIP
Tori Miller Liu, MBA, FASAE, CAE, CIP is the President & CEO of the Association for Intelligent Information Management. She is an experienced association executive, technology leader, speaker, and facilitator. Previously, she served as the Chief Information Officer of the American Speech-Language-Hearing Association (ASHA) and been working in association management since 2006. Tori is a current member of the ASAE Executive Management Advisory Council and Association Coalition for AI. She is a former member of the ASAE Technology Professional Advisory Council and a former Board Member of Association Women Technology Champions. She was named a 2020 Association Trends Young & Aspiring Professional and 2021 Association Forum Forty under 40 award recipient. She is also an alumna of the ASAE NextGen program. She is a Certified Association Executive and holds an MBA from George Washington University. In 2023, Tori was named as a Fellow of the American Society of Association Executives (ASAE).