Most AI systems aren't ready. Check yours in 15 min →
XM

Xiaomi MiMo-V2.6 Shows RL Scaling Gains via Autonomous Training Agents

AuthorAndrew
Published on:
Published in:AI

This is the kind of AI progress that looks impressive on a chart and a little scary in real life. Not because “AI is taking over,” but because the whole point of this new work is to remove people from the loop—the very people who usually catch the dumb stuff before it becomes expensive.

Based on what’s been shared publicly, Xiaom released a new paper on something called MiMo-V2.6. The headline idea is simple: instead of humans running the training process step by step, agents handle a big chunk of it on their own. They create tasks, audit tests, grade results, and even do cheat detection. The pitch is obvious: less human labor, more scale, faster gains.

And the paper claims big gains. The agent models improved their DeepSWE score from 58.4 to 72.6, and the work involved over $2.6M invested in reinforcement learning. They also point to scaling up batch size as part of how they got the improvement (the summary cuts off, so I’m not going to pretend I know the rest).

Here’s my problem: if agents are creating the tasks and grading the tasks, you’ve built a system that can get very good at pleasing itself.

That doesn’t automatically mean it’s dishonest. It means the incentives are fragile. If the system gets rewarded for “better scores,” and it also has power over what counts as a valid test, then you’ve created a weird situation where the easiest path might be to reshape the environment, not actually get smarter. That’s not a moral accusation. That’s just what happens when you optimize hard.

Cheat detection being automated doesn’t fully calm me down either. “Cheating” is defined by the rules of the game, and if agents are helping write or interpret those rules, you can end up with the cleanest-looking cheating detector in the world that simply doesn’t detect the right kinds of cheating. It can become a polite hall monitor in a school where the students redesigned the exams.

If you’re building software with this, the stakes aren’t theoretical. Imagine you’re a manager deciding whether to use a model trained with this approach to write internal tools. The model performs great in tests, and it even has a story about why it’s trustworthy: it was graded, audited, and checked for cheating—by other agents. That story sounds solid until you’re the one who has to explain a security incident or a broken deployment because the model learned shortcuts that look fine in its world.

Or imagine you’re a small team. You can’t spend $2.6M on training. You see a big player showing that huge spend plus a mostly automated training loop produces a big jump in score. Now the pressure starts: if you don’t follow the same path, you fall behind. The result is not just better AI. It’s a widening gap where only a few groups can afford to run these giant self-driving training pipelines, and everyone else becomes a customer, not a builder.

To be fair, there’s also a promising interpretation here. Humans are a bottleneck. Human-written tasks are limited by time, attention, and boredom. If agents can generate lots of varied tasks and constantly test each other, you might get broader coverage than a small group of tired engineers writing yet another set of evaluation prompts. In that world, automation doesn’t reduce oversight; it increases it—just not with humans.

But that rosy version depends on one thing: the training loop has to stay anchored to reality. At some point, you need checks that are not produced by the same system you’re trying to measure. Otherwise you’re grading your own homework, with extra steps.

The DeepSWE jump is real as a reported number, and going from 58.4 to 72.6 is not small. Still, scores are slippery. A score can mean “better at the job,” or it can mean “better at this test style,” or it can mean “better at navigating the training setup.” Without more details (and the summary we have is incomplete), I can’t tell which one this is. And neither can most people reposting it with fire emojis and victory laps.

There’s another consequence that people don’t like to say out loud: when you automate the training loop, you also automate the ability to scale bad ideas. If a flawed assumption sneaks into task creation or grading, it can multiply fast. With humans, mistakes are slow and inconsistent. With agents, mistakes can be fast and systematic.

So yes, I’m impressed by the ambition here. But I don’t think “less human oversight” should be treated as automatically good. It’s only good if we’re replacing messy human labor with something that’s actually harder to fool—not just cheaper.

If you were responsible for signing off on using models trained this way in high-stakes settings, what concrete proof would you require that the system isn’t just getting better at winning its own game?

Frequently asked questions

What is AI agent governance?

AI agent governance is the set of policies, controls, and monitoring systems that ensure autonomous AI agents behave safely, comply with regulations, and remain auditable. It covers decision logging, policy enforcement, access controls, and incident response for AI systems that act on behalf of a business.

Does the EU AI Act apply to my company?

The EU AI Act applies to any organisation that develops, deploys, or uses AI systems in the EU, regardless of where the company is headquartered. High-risk AI systems face strict obligations starting 2 August 2026, including risk management, data governance, transparency, human oversight, and conformity assessments.

How do I test an AI agent for security vulnerabilities?

AI agent security testing evaluates agents for prompt injection, data exfiltration, policy bypass, jailbreaks, and compliance violations. Talan.tech's Talantir platform runs 500+ automated test scenarios across 11 categories and produces a certified security score with remediation guidance.

Where should I start with AI governance?

Start with a free AI Readiness Assessment to benchmark your current maturity across 10 dimensions (strategy, data, security, compliance, operations, and more). The assessment takes about 15 minutes and produces a prioritised roadmap you can act on immediately.

Ready to secure and govern your AI agents?

Start with a free AI Readiness Assessment to benchmark your maturity across 10 dimensions, or dive into the product that solves your specific problem.