Order Kate’s book today and discover What Matters Next!
On August 11, Anthropic published an article in its help center explaining how Claude now marks what it produces: an imperceptible statistical watermark woven into generated text, and signed C2PA provenance metadata attached to supported generated files. The marking is applied at the model level, so it moves with the output across the API, the apps, the coding tools, and the cloud platforms — in short, wherever Claude runs. Models launched on or after August 2, 2026 carry it from launch. Older ones are being retrofitted. Although its catalyst was the EU AI Act (more on that below), it applies worldwide, regardless of whether the user is in Europe.
Within about a day, the loudest part of the conversation had almost nothing to do with the technology. Writers, students, lawyers, researchers, and developers were arguing about exposure: who would be caught, by whom, and on what evidence. Some canceled their subscriptions over it. Others pointed out the irony of a model trained on other people’s work now signing its own.
That reaction is the real story here, and before we argue about detection accuracy it’s worth examining that reaction more closely.
The technical announcement itself was fairly modest, and unusually candid about its own limitations. But it landed in an institutional vacuum. A great many schools, newsrooms, employers, publishers, and platforms now have access to a signal they have no agreed way to read. Detection capability has outrun interpretation policy, which is a familiar sequence and rarely a comfortable one. Detection and evasion have always moved together, as regulation and capability so often do, and the paraphrasing tools that wash a watermark out will improve at roughly the same rate the watermarks do. When a capability arrives before the norms that govern its use, that capability tends to get used badly for a while. And questions this contested don’t get settled once; they get relitigated every time the technology moves, which it will.
None of the underlying limits are new information, and they were not discovered in the last week. I’ve participated in discussions convened around the UN’s High-level Advisory Body on AI, where watermarking came up regularly as part of the provenance conversation, and I don’t recall anyone in those rooms describing it as a detector. It was understood as one layer among several: useful for making synthetic content traceable at population scale, close to useless for adjudicating any individual case. The people building these systems have generally been clear about that. The gap has always been on the receiving end, in what institutions do with a signal once they have one.
Provenance is a record of what a piece of content passed through on its way to you. That is a narrower claim than this whole thing sounds like, and the narrowness is the point.
Two mechanisms in play work differently and fail differently, which matters more than most coverage suggests. A text watermark is statistical. It biases word choices according to a key the provider holds, which means it needs enough text to carry a reliable signal, travels through copy and paste, and degrades under paraphrase, translation, and heavy editing. Short passages may fall below the detection threshold entirely. C2PA metadata, by contrast, is a cryptographically signed manifest attached to a file. It is strong evidence when present and fragile in completely ordinary ways: format conversion, re-saving, resizing through a content pipeline, or a screenshot will strip it without leaving a trace. Most organizations have been destroying provenance metadata in their own upload workflows for years without noticing.
I ran into the inverse of this in 2019, writing about the 10 Year Challenge meme. The main popular objection to that piece was that the data already existed, since people had been uploading photos to these platforms for a decade already. What that objection missed is that the ambient corpus had been through a decade of ordinary handling. Re-uploads, edits, format changes, and platform processing strip EXIF, so the timestamps and provenance on all that abundant content were unreliable or gone. The meme produced something the abundance of content could not: a clean set, fixed interval, labeled by the people in the photos. Volume was never a scarce resource. Reliable provenance was.
The same asymmetry applies here, and it has an uncomfortable consequence. Text pasted straight out of the model and published as-is will carry a detectable mark. The same text after a round of revision, a translation, or a pass through a system that rewrites the file may not. Detection tracks how much handling a piece of content received on its way to publication, which is a fact about someone’s workflow rather than a fact about their honesty. The people who show up as AI-involved will skew toward those with the shortest pipelines, and some of the heaviest users will be structurally invisible. Any institution planning to treat detection as evidence should be sure to sit with that before it writes its first policy.
So a detected mark supports one narrow inference: this content probably passed through this system. Everything a leader actually wants to know sits above that line.
It cannot establish authorship. Someone may have used the model to proofread, translate, summarize, or convert a file whose ideas, structure, reporting, and argument are entirely their own. The mark records processing, not origination, and it flattens a genuinely layered thing into a binary one.
It cannot establish legitimacy. Knowing that a tool was involved says nothing about whether its use was appropriate under the rules that apply. Copyediting a memo and fabricating a citation both leave the same kind of trace. The interesting distinctions live entirely in policy, not in the signal.
It cannot establish accuracy. A watermark says nothing about whether the content is correct, sourced, or responsible. Marked content can be excellent. Unmarked content can be nonsense.
And absence establishes nothing at all. Content may be unmarked because it came from an older model, a different provider, a platform that rewrote the file, a passage too short to carry a signal, or a draft that was edited enough to wash the signal out. Treating a clean result as a clean bill of health is a mistake with real consequences, and it is the mistake I’d expect institutions to make first.
That last failure mode has a name I’ve been using for a while (and will be talking about more in my forthcoming book): certainty theater. It’s what happens when we perform confidence we haven’t actually earned, usually because the alternative feels like weakness. A detection score is very good raw material for certainty theater. It arrives with a number attached, it looks like evidence, and it lets a busy administrator skip the harder conversation about what the number means.
Having said all that, I want to be clear that I think this is a good development, and I’d rather have the signal than not.
An information environment with no machine-readable history at all is worse than one with partial, contestable history. Provenance marks make synthetic content traceable at the scale where scale is the actual problem: coordinated inauthentic campaigns, automated persuasion, bulk-generated pages, media that circulates faster than anyone can verify it by hand. They give platforms and regulators a common technical layer to build policy on. They create a reason for the disclosure conversation to happen at all, in organizations that would otherwise keep deferring it. And they push the industry away from total opacity, which has been the default and has served almost no one well.
The honest framing is that watermarking is weak evidence about any individual piece of content and useful infrastructure across a whole information system. Those two statements sit together comfortably. Most of the current argument comes from collapsing them into one.
This did not come from nowhere, and it isn’t one company’s choice. The catalyst is Article 50 of the EU AI Act, which took effect on August 2, 2026 and requires providers of generative systems to mark synthetic output in a machine-readable format and make it detectable. The Commission finalized its implementation guidelines in July, alongside a Code of Practice on Transparency of AI-Generated Content that providers can sign to demonstrate compliance. Penalties run to fifteen million euros or three percent of worldwide turnover. Systems already on the market got a short runway into December for the marking obligation specifically.
The part worth watching is that Anthropic applied it globally rather than only in the EU, which is how a regional rule becomes a default operating norm everywhere. Regulation is turning into product architecture, and product architecture turns into public expectation of what responsible output looks like.
Meanwhile the international layer has been consolidating in parallel. The UN General Assembly established an Independent International Scientific Panel on AI and a Global Dialogue on AI Governance in 2025, and the first Dialogue convened in Geneva in July, with information integrity among the themes running through it. C2PA continues to develop the file-level standard. None of these bodies is going to tell a school district what to do with a detection result on a term paper. That interpretive work stays local. It’s ordinary planning under uncertainty, and waiting for it to arrive from somewhere else is not a strategy.
Disclose by role, not by presence. The binary question, was AI involved, produces no useful information and invites suspicion. Say what the tool actually did. Try something like: “AI disclosure: Claude was used to summarize four source reports and to copyedit the final draft. The argument, the structure, the reporting, and all quoted material are mine. No AI-generated passages appear in the published text.” That takes twenty seconds to write, it survives scrutiny, and it moves the conversation from accusation to specifics.
Write your interpretation rule before you need it. Decide now what a detected mark triggers in your organization. In most cases the right answer is a conversation, a request for the working materials, and a documented decision. Name who reviews it, what the person can show in response, and what the appeal looks like. An institution that starts drafting this policy after its first contested case will draft it badly, under pressure, in public.
Keep your own provenance trail. Outlines, notes, source files, interview recordings, prompts where they materially shaped the output, and revision history. This isn’t defensiveness. It’s the difference between being able to explain your work and having to insist on it. If you are producing anything with professional stakes attached, you want the record to exist before anyone asks.
Test what your own pipeline strips. Run a Claude-generated image through your CMS, your export process, and your image optimizer, then check whether the C2PA manifest survived. Most teams assume provenance travels with the file. It often doesn’t, and finding that out during a dispute is expensive.
The copy-able version, for sharing in your organization:
There’s a version of the next two years where watermarking becomes a genuine layer of trust infrastructure that helps people decide, communicate, and act responsibly under uncertainty. There’s another version where it becomes a blunt instrument, used mostly to accuse people with less institutional power than their accusers, on evidence nobody involved fully understands.
The variable isn’t the technology. It’s whether the institutions holding the detector are willing to do the slower work of interpretation. That work is unglamorous and mostly consists of writing things down before you need them.
It’s also worth noticing what these marks don’t touch. Provenance on outputs says nothing about inputs, about how the data that trained these systems became something other than what it was or whether the people whose work made the model possible were ever compensated. That question isn’t answered by a signature on a file. Marking outputs is the easier half of the accountability problem, and we should be honest that we’ve started with the easier half.
A provenance signal is an invitation to ask a better question, and the organizations that get value from it will be the ones that already know how to ask. That capacity is built the same way it always has been, by treating these systems as thinking partners rather than oracles and staying curious about what you actually know versus what you’ve merely been told with confidence.
So here’s the test I’d suggest, and it doesn’t involve a detector. Take the last significant thing your organization published. Can you describe how it was made? Not defend it. Describe it: who wrote what, what tools touched it, what was verified and by whom. If the answer requires an investigation, the mark was never going to help you. If the answer is already written down somewhere, you were never going to need one.
Six years after its pandemic-era launch, The Tech Humanist Show has grown into a seasonal platform for thoughtful conversations about technology, humanity, and what...
Before we dive into implications and applications, let’s establish a clear foundation in the basics of agentic AI. This emerging technology represents a fundamental...
Search is changing. With more AI-generated overviews, your traditional SEO playbook is obsolete. Here’s what to focus on instead. Years ago, I owned an...
When developing an AI digital transformation strategy, many leaders turn to historical precedents for guidance. The comparison between AI adoption and global electrification has...
"Language shapes reality. In my years guiding digital transformation initiatives, I've noticed something fascinating: the way different groups within organizations talk about their work...
Today let’s take a journey into the future of workspaces and workplaces of the future. How will they look? How will they function? And...
Let’s talk about AI. Not the AI of science fiction, not the AI that promises to take over the world or replace humans. No,...
Have you found yourself worried about the role of independent human thought in a world increasingly influenced by artificial intelligence (AI)? Many of our...
As we step into 2024 and look even further into the future, it’s evident that digital transformation continues to reshape the world of business...
Digital Transformation isn’t a one-time project, but a continuous journey that evolves with time and technology. It’s not just one magical digital transformation project...
Many leaders have digital transformation on their minds, and everyone is talking about AI. So with the rapid pace of technological change, it’s natural...
Lately I’ve been asked by several prospective clients whether I speak about AI, and — can I be honest? — that question surprised me...
By using this form you agree with our Privacy Policy and the storage and handling of your data by this website.
info@koinsights.com
27 W 60th St #20560
New York, NY 10023-9991
By using this form you agree with our Privacy Policy and the storage and handling of your data by this website.
© 2026 KO Insights. All rights reserved. Privacy Policy. Cookie Policy.
Made with love by Unicorn Designer Studio.
We use cookies to improve your experience and understand what’s working, so we can keep making this site more useful for you. Read our Privacy Policy.