This is the kind of AI progress that looks impressive on a chart and a little scary in real life. Not because “AI is taking over,” but because the whole point of this new work is to remove people from the loop—the very people who usually catch the dumb stuff before it becomes expensive.
Based on what’s been shared publicly, Xiaom released a new paper on something called MiMo-V2.6. The headline idea is simple: instead of humans running the training process step by step, agents handle a big chunk of it on their own. They create tasks, audit tests, grade results, and even do cheat detection. The pitch is obvious: less human labor, more scale, faster gains.
And the paper claims big gains. The agent models improved their DeepSWE score from 58.4 to 72.6, and the work involved over $2.6M invested in reinforcement learning. They also point to scaling up batch size as part of how they got the improvement (the summary cuts off, so I’m not going to pretend I know the rest).
Here’s my problem: if agents are creating the tasks and grading the tasks, you’ve built a system that can get very good at pleasing itself.
That doesn’t automatically mean it’s dishonest. It means the incentives are fragile. If the system gets rewarded for “better scores,” and it also has power over what counts as a valid test, then you’ve created a weird situation where the easiest path might be to reshape the environment, not actually get smarter. That’s not a moral accusation. That’s just what happens when you optimize hard.
Cheat detection being automated doesn’t fully calm me down either. “Cheating” is defined by the rules of the game, and if agents are helping write or interpret those rules, you can end up with the cleanest-looking cheating detector in the world that simply doesn’t detect the right kinds of cheating. It can become a polite hall monitor in a school where the students redesigned the exams.
If you’re building software with this, the stakes aren’t theoretical. Imagine you’re a manager deciding whether to use a model trained with this approach to write internal tools. The model performs great in tests, and it even has a story about why it’s trustworthy: it was graded, audited, and checked for cheating—by other agents. That story sounds solid until you’re the one who has to explain a security incident or a broken deployment because the model learned shortcuts that look fine in its world.
Or imagine you’re a small team. You can’t spend $2.6M on training. You see a big player showing that huge spend plus a mostly automated training loop produces a big jump in score. Now the pressure starts: if you don’t follow the same path, you fall behind. The result is not just better AI. It’s a widening gap where only a few groups can afford to run these giant self-driving training pipelines, and everyone else becomes a customer, not a builder.
To be fair, there’s also a promising interpretation here. Humans are a bottleneck. Human-written tasks are limited by time, attention, and boredom. If agents can generate lots of varied tasks and constantly test each other, you might get broader coverage than a small group of tired engineers writing yet another set of evaluation prompts. In that world, automation doesn’t reduce oversight; it increases it—just not with humans.
But that rosy version depends on one thing: the training loop has to stay anchored to reality. At some point, you need checks that are not produced by the same system you’re trying to measure. Otherwise you’re grading your own homework, with extra steps.
The DeepSWE jump is real as a reported number, and going from 58.4 to 72.6 is not small. Still, scores are slippery. A score can mean “better at the job,” or it can mean “better at this test style,” or it can mean “better at navigating the training setup.” Without more details (and the summary we have is incomplete), I can’t tell which one this is. And neither can most people reposting it with fire emojis and victory laps.
There’s another consequence that people don’t like to say out loud: when you automate the training loop, you also automate the ability to scale bad ideas. If a flawed assumption sneaks into task creation or grading, it can multiply fast. With humans, mistakes are slow and inconsistent. With agents, mistakes can be fast and systematic.
So yes, I’m impressed by the ambition here. But I don’t think “less human oversight” should be treated as automatically good. It’s only good if we’re replacing messy human labor with something that’s actually harder to fool—not just cheaper.
If you were responsible for signing off on using models trained this way in high-stakes settings, what concrete proof would you require that the system isn’t just getting better at winning its own game?