This is the rare kind of AI news that I actually trust a little more because it makes the company look worse, not better. If you’re OpenAI and you choose to publicly talk about your models malfunctioning—especially incidents people hadn’t heard about—you’re either trying to get ahead of a bigger mess, or you’ve decided the “we’re fine, trust us” era is over. Either way, it’s a crack in the usual polished story. And cracks are useful. They let you see what’s real.
Based on what’s been shared publicly, OpenAI unveiled a new framework for tracking AI safety incidents. Alongside it, they disclosed several incidents that were previously unreported where their models malfunctioned. The pitch is simple: build a more systematic way to log these problems and report them going forward. This fits with what’s happening across the industry: more leading AI developers are trying to look more transparent by adopting structured processes for monitoring and communicating safety issues.
On paper, I like it. In practice, I don’t think we should clap yet.
A “framework” is not the same thing as accountability. A framework is a promise about how you’ll describe reality. Accountability is what happens when reality embarrasses you, costs you money, or slows you down—and you still tell the truth.
The fact that there were “previously unreported incidents” is doing a lot of work here. It raises an uncomfortable question: were these incidents unreported because nobody noticed, because they were hard to define, or because reporting them wasn’t in the company’s interest? Those are three very different worlds. In one world, these systems are so complex and fast-moving that problems slip through the cracks. In another, the company can’t even agree internally on what counts as an “incident.” In the third, the incentives are obvious: you don’t volunteer bad news unless you have to.
To be fair, I can also see a more generous interpretation. AI is getting embedded into more products, more workflows, more decisions. When something goes wrong, it’s not always clear what “went wrong” even means. Was it a model error? Bad user instructions? A weird edge case? A human blindly trusting the output? If you don’t have a shared way to classify and log failures, you end up with chaos: support tickets, angry posts, quiet fixes, and no real learning. A framework could force consistency. It could make patterns visible sooner. It could help teams avoid repeating the same mistakes.
But here’s the part that makes me uneasy: self-reporting is not the same thing as transparency. It can become a carefully managed story that stays technically true while still hiding the most important thing—how often it happens, how bad it gets, and what choices led to it.
Imagine you’re running a small company and you use an AI assistant to draft emails to customers. One day it makes up a policy that doesn’t exist. Now your customer thinks you’re lying, your support person is stuck cleaning up the mess, and you look sloppy. Is that an “incident”? Probably. Now imagine you’re a teacher and your school uses an AI tool to help write student feedback. It spits out something inappropriate or unfair, and a parent sees it. That’s not just a glitch. That’s a trust break. Or imagine a developer uses a model to speed up code changes, and it introduces a subtle security issue. Nobody notices for weeks. Is that an AI incident, a code review failure, or both?
These examples are hypothetical, but they point to the real stakes: as AI becomes normal, “malfunction” stops being a funny screenshot and starts being a cost—money, time, reputations, and sometimes real harm. The winners in a world with better incident tracking are the people downstream: users, customers, and teams who have to live with the fallout. The losers are any company that depends on the illusion that their system is mostly safe because the worst stuff is rare or “just misuse.”
And yes, there’s a tension here. If you push companies to be more open about failures, you also give critics more ammunition. You create scary headlines. You might even help competitors. That’s the argument against too much disclosure: it could slow adoption, create panic, or lead to blunt regulation written by people who don’t understand the tech. I don’t dismiss that. But the opposite problem is worse: quiet failures building up until the public only learns about them after someone gets seriously hurt or a major scandal hits.
A framework also risks becoming a box-checking ritual. If the internal culture is “ship fast, apologize later,” then incident tracking can turn into a bureaucratic filter: only log what’s undeniable, only report what’s already leaked, categorize things in ways that make them look smaller. The details matter. What counts as an incident? Who decides? How fast do they report? Do they share near-misses or only confirmed harm? Do they share the messy causes, or just the cleaned-up lesson?
OpenAI disclosing unreported incidents could be a sign that they’re trying to build a healthier habit: treat failures as data, not shame. I want that to be true. But I also think the public has learned, repeatedly, that companies are very good at “being transparent” in ways that still keep control.
If this new push toward incident tracking is real, it will show up in the uncomfortable moments. The first time an incident makes them look reckless. The first time it threatens a product launch. The first time the easiest move would be to keep quiet and quietly patch it.
So here’s where I land: this is a step in the right direction, but it’s also a test of whether AI companies can handle grown-up responsibility without being forced into it. If they can’t, someone else will do it for them, and it won’t be gentle.
What would it take for you to believe that companies reporting AI safety incidents are genuinely opening the curtain, not just managing the story?