Most AI systems aren't ready. Check yours in 15 min →

Anthropic добавит водяные знаки в Claude после подписания акта ЕС

AuthorAndrew
Published on:
Published in:AI

Watermarking AI output sounds like one of those “finally, an adult is in the room” ideas. And it probably is. But it’s also the kind of policy move that looks clean on a slide and gets messy the second it hits real life.

Here’s the basic fact, from what’s been shared publicly: Anthropic says its Claude models will mark texts and files with watermarks. They’re doing it after signing the EU’s AI regulation act. And there’s a timing detail that matters: only Claude models launched after August 2, 2026 will watermark right away. Older versions are supposed to get the feature after a transition period.

I actually agree with the direction. Not because I think a watermark magically “solves” deepfakes, cheating, scams, or misinformation. It won’t. I agree because right now we’re living in a weird in-between world where AI content is everywhere, and the default is plausible deniability. People can generate something in two minutes, post it, and if it causes damage they can shrug: “Who knows where it came from?” That’s not a healthy norm.

A watermark, if it’s done well, changes the default from “we can’t tell” to “maybe we can check.” That’s already a big shift. It makes lying slightly harder. It makes accountability slightly more possible. And in systems like this, “slightly” matters, because the bad stuff scales. One person cheating on an essay is annoying; a thousand people automating fake customer support chats, fake reviews, fake job applications, fake medical advice—now you’ve got a trust problem that spreads into everything.

But I don’t buy the comforting version of this story where watermarking becomes a simple truth machine. The people who need this to work most—the teachers, the small business owners, the overwhelmed moderators, the normal person trying to figure out if a document is real—are not in a position to run forensic checks all day. If the watermark is invisible and requires special tools to verify, most people will never verify. If it’s visible, it will be cropped, paraphrased, retyped, screenshot, or just regenerated by a different model that doesn’t cooperate.

And that’s where the stakes get real: if watermarking becomes a “good actor badge,” it could punish the wrong people. Imagine you’re a student who uses Claude to help clean up grammar in a second language. Your final text gets watermarked, and suddenly you’re “the AI kid,” even if the ideas and structure are yours. Meanwhile someone else uses an older model, or a different tool, or just rewrites the output by hand, and sails through with no mark. That’s not fairness. That’s just selecting for people who know how to hide it.

There’s also a business angle that’s going to make people uncomfortable. If only models after August 2, 2026 are watermarked by default, you’ve got a window where older versions—and whatever else is out there—can become the go-to option for anyone who wants content with no trace. That could create a perverse little market: “Need something clean? Use the unmarked model.” When you create two lanes—marked and unmarked—you shouldn’t be surprised when the worst behavior concentrates in the unmarked lane.

Still, I’d rather have imperfect tracing than none. The right way to think about this is not “watermarks will stop abuse.” It’s “watermarks give society a handle.” A handle for disputes. A handle for platforms deciding what to label. A handle for companies trying to protect their own documents from getting mixed up with AI-generated ones. A handle for journalists who receive a “leaked memo” and need at least one more signal before they hit publish. Even if bad actors can bypass it, a lot of everyday harm comes from low-effort misuse. Raising the effort level can reduce the volume.

The uncomfortable part is what happens next. Once watermarking exists, people will start demanding it everywhere. Schools will demand it. Employers will demand it. Courts will demand it. And then we’re not just watermarking “AI text.” We’re building a culture where you may have to prove you didn’t use a tool. That flips the burden onto ordinary people. It’s easy to imagine a future where a job applicant gets rejected because their cover letter “looks like AI” and there’s no watermark to check, no clear standard, just vibes and risk avoidance.

And I can’t ignore the obvious counterpoint: a watermark can become a privacy and control lever. If a company can mark outputs, it can also potentially track patterns of usage, or at least create pressure for identity checks and logging “for safety.” Maybe that won’t happen here, but the incentive is sitting right there. Regulators want traceability. Companies want compliance. Users want convenience. Those goals don’t always line up with anonymity or freedom to experiment.

So yes, I’m glad Claude is moving toward watermarking, and I think more AI systems will end up doing it, whether they want to or not. But I’m also wary of the story we’ll tell ourselves—that this is a clean fix, that we can relax now, that truth will be easy again. It won’t. It’s one tool, and it will create new edge cases, new unfairness, and new games.

If watermarking becomes normal, should unwatermarked AI output be treated as suspicious by default, or should we treat watermarking as optional context and keep the burden of proof on the accuser?

Frequently asked questions

What is AI agent governance?

AI agent governance is the set of policies, controls, and monitoring systems that ensure autonomous AI agents behave safely, comply with regulations, and remain auditable. It covers decision logging, policy enforcement, access controls, and incident response for AI systems that act on behalf of a business.

Does the EU AI Act apply to my company?

The EU AI Act applies to any organisation that develops, deploys, or uses AI systems in the EU, regardless of where the company is headquartered. High-risk AI systems face strict obligations starting 2 August 2026, including risk management, data governance, transparency, human oversight, and conformity assessments.

How do I test an AI agent for security vulnerabilities?

AI agent security testing evaluates agents for prompt injection, data exfiltration, policy bypass, jailbreaks, and compliance violations. Talan.tech's Talantir platform runs 500+ automated test scenarios across 11 categories and produces a certified security score with remediation guidance.

Where should I start with AI governance?

Start with a free AI Readiness Assessment to benchmark your current maturity across 10 dimensions (strategy, data, security, compliance, operations, and more). The assessment takes about 15 minutes and produces a prioritised roadmap you can act on immediately.

Ready to secure and govern your AI agents?

Start with a free AI Readiness Assessment to benchmark your maturity across 10 dimensions, or dive into the product that solves your specific problem.