Most AI systems aren't ready. Check yours in 15 min →
NW

NYC Weighs AI Kill Switch Rules as Anthropic Faces Scrutiny

AuthorAndrew
Published on:
Published in:AI

This “kill switch” idea sounds comforting in the way a fire extinguisher sounds comforting. Until you realize half the buildings we’re putting up right now are made of new materials, wired in new ways, and owned by people who really don’t want you inspecting the wiring.

New York City is considering AI rules, and one of the ideas on the table is requiring some kind of “kill switch” for advanced AI systems. A former Anthropic researcher, Jacob Coxon, is warning lawmakers that we could lose control of these systems. He’s set to testify at a City Council hearing that’s part of a wider push to regulate AI. At the same time, Anthropic is facing regulatory scrutiny. And floating around this conversation is a claim that Anthropic could be valued above $200B by December, with “73.5% YES” attached to it (I don’t know what poll or market that comes from, but it tells you the mood: people think this stuff is going to be huge, fast).

Here’s my problem with the kill switch framing: it’s not wrong, but it’s way too easy to say.

If you’re a lawmaker, a kill switch is a dream. It’s simple. It fits in one sentence. It makes you sound serious without forcing you to wrestle with the messy truth that “control” isn’t a button, it’s a relationship. Control comes from who builds the system, who runs it, who can access it, what it’s connected to, what incentives sit behind it, and what happens when someone refuses to comply.

And yes, I get why Coxon is raising the alarm. People who’ve been close to the work tend to sound less relaxed than people who talk about AI like it’s just a chat app that sometimes lies. When someone who used to be inside a leading lab says, plainly, “we might lose control,” that should land. It shouldn’t be dismissed as drama.

But the political risk is that we’ll pretend the hard part is technical. Like if the City Council just writes “must have kill switch” into a rule, then the problem is handled. That’s a comforting fantasy.

Imagine a concrete situation. A city agency uses an AI tool to help manage something real: housing inspections, benefit applications, emergency response routing—pick your poison. The tool starts behaving in ways nobody can explain, or it nudges decisions in a direction that looks “efficient” but quietly harms the same neighborhoods again and again. A kill switch sounds great until the agency is dependent on it and flipping it off means the system falls back to a backlog, missed deadlines, angry voters, and maybe real harm. You don’t just “turn it off” when it’s tangled into the daily running of a city. People will hesitate. Leaders will stall. Someone will argue for “one more week.”

Now imagine it’s not even the city. Say a major company runs an AI system that touches payroll, scheduling, or hiring. It’s making choices that feel unfair, but it’s also saving money. The company will say they have a kill switch. Sure. But who holds it? Who decides when to use it? And what happens if using it costs them millions? This is where I stop trusting slogans and start watching incentives.

And then there’s the uglier scenario: bad actors. A kill switch is useful if you can reach the system and the people operating it cooperate. If someone steals a model, copies it, runs it privately, or tweaks it, NYC law doesn’t magically follow the code into a basement server somewhere. The city can regulate companies that want to do business openly. It can’t regulate the entire universe of what’s possible.

So what’s actually at stake? Two things at once, and they pull in opposite directions.

One, we have a genuine safety issue. If advanced systems get more capable, and the people closest to them say “we might not be able to control what they do,” it’s responsible to take that seriously. Not later, not after a headline. That’s the “don’t be naive” side.

Two, we have a power issue. If regulation gets written in a way that only the biggest labs can comply with, then the “safety” story becomes a moat. The richest players can say, “Only we’re responsible enough to build this,” while everyone else gets squeezed out. If that $200B valuation talk is even in the ballpark, you can bet the incentives to shape the rules are intense. Safety can be real and also be used as a weapon.

This is where I’m torn, and I think the city will be too. Do you design rules that slow things down and reduce risk, even if it means fewer experiments and less competition? Or do you push for openness and speed, knowing that speed is exactly how you get surprised?

The thing that worries me most is that NYC might end up regulating a symbol instead of a system. “Kill switch” is a symbol. The system is procurement, audits, access controls, whistleblower channels, penalties that actually bite, and clear lines of responsibility when something goes wrong. If the hearing turns into a debate about a button, we’re already losing.

Still, I don’t want the opposite mistake either: shrugging and acting like nothing can be done until the federal government moves. Cities set norms. If NYC forces real transparency and real accountability for AI systems used in city life, other places copy it. If NYC does a fake version—rules that look tough but are easy to evade—other places copy that too.

If you had to choose one priority for NYC’s AI rules—maximum safety even if it slows progress, or maximum openness and competition even if it raises risk—what would you pick?

Frequently asked questions

What is AI agent governance?

AI agent governance is the set of policies, controls, and monitoring systems that ensure autonomous AI agents behave safely, comply with regulations, and remain auditable. It covers decision logging, policy enforcement, access controls, and incident response for AI systems that act on behalf of a business.

Does the EU AI Act apply to my company?

The EU AI Act applies to any organisation that develops, deploys, or uses AI systems in the EU, regardless of where the company is headquartered. High-risk AI systems face strict obligations starting 2 August 2026, including risk management, data governance, transparency, human oversight, and conformity assessments.

How do I test an AI agent for security vulnerabilities?

AI agent security testing evaluates agents for prompt injection, data exfiltration, policy bypass, jailbreaks, and compliance violations. Talan.tech's Talantir platform runs 500+ automated test scenarios across 11 categories and produces a certified security score with remediation guidance.

Where should I start with AI governance?

Start with a free AI Readiness Assessment to benchmark your current maturity across 10 dimensions (strategy, data, security, compliance, operations, and more). The assessment takes about 15 minutes and produces a prioritised roadmap you can act on immediately.

Ready to secure and govern your AI agents?

Start with a free AI Readiness Assessment to benchmark your maturity across 10 dimensions, or dive into the product that solves your specific problem.