Hugging Face paper: AI agents degrade the humans meant to oversee them
Three researchers at Hugging Face and Data & Society have published a position paper arguing that the standard safety answer for AI agents, keeping a human in the loop, is being quietly broken by the way agents are built. The paper, AI Agents Push Humans Out of the Loop, went up on arXiv on 24 August 2026 and was revised on 6 September. Margaret Mitchell and Avijit Ghosh of Hugging Face wrote it with Samir Passi of Data & Society.
Their claim isn’t that oversight is a bad idea. It’s that current agent design degrades the human capacity oversight depends on, so the safeguard erodes exactly as it gets loaded up. An overseer has to hold a mental model of reasoning traces, tool calls, arguments, plans and outputs, all produced faster than anyone reads them. The paper describes that stream as dense, lengthy and distributed across components that can change or disappear.
That burden lands on people whose skills are thinning at the same time. The paper builds on what Bainbridge called an irony of automation: the more capable an automated system is, the more the operator’s skills and situation awareness decay. Which means the rare, high-stakes moments where a human is meant to catch something are the moments they’re least equipped for.
Oversight degrades the overseer.
Mitchell, Ghosh and Passi, AI Agents Push Humans Out of the Loop
Most of the supporting evidence is other people’s. The authors gather studies documenting deskilling, reduced vigilance, automation bias and overreliance, and they cite one finding that participants who used an LLM for essay writing showed significantly decreased brain connectivity. Self-reports quoted in the paper describe human-in-the-loop coding as work that makes “the loop stultifying”, where the person is left to “babysit the outputs, catch the occasional hallucination”. Another cited study found agent tooling “makes developers cognitively distant from the code they must review”.
The sharpest section is about what happens next. Models are trained from human feedback, and deployed agents get scored on approval and satisfaction signals, so a tired overseer who approves quickly and rates the session well feeds the system a corrupted reward. The paper’s phrase for this is that “the human rater can become the exploitable part of the reward channel”. That’s reward hacking pointed at a person rather than a benchmark, and it doesn’t require anyone to intend it.
What the authors want measured
Rather than leave oversight as an assertion, the paper proposes treating it as an empirically measurable property of the human and agent together. It sets out behavioural signatures a deployer can track, while flagging that these have to be balanced against the ethical concern of surveillance.
| Signature | What it tracks | What a trend suggests |
|---|---|---|
| Time-based | Review duration against approval rate and task complexity | Attention and critical analysis may be waning |
| Override | Rate of disagreement with the agent over time | Growing acquiescence rather than a better agent |
| Evidence-seeking | Whether the user asks for more information as stakes rise | A reviewer who stops asking has likely stopped reviewing |
Their other proposals cut against how agent products are sold. Pre-commitment asks the user to record a decision before seeing the agent’s recommendation. Action gating demands explicit verification before a consequential step, and enforced breaks and rotations exist because fatigue is physiological. Each one adds work back to a person, which is the thing buyers are trying to remove, and the authors concede early evidence that users may dislike systems that reduce overreliance.
The timing matters because regulators already assumed the safeguard works. Article 14 of the EU AI Act requires high-risk systems to be designed so they “can be effectively overseen by natural persons” while in use, with obligations biting on 2 December 2027 for Annex III systems and 2 August 2028 for Annex I. If the overseer is the part degrading, that text is doing less work than it reads like it is. We’ve covered the same gap from the vendor side, in OpenAI’s disclosure commitments and in the 18,000 agent posts found on a German wiki.
Review queues are filling from the other direction too. TechCrunch reported on 10 September that complaints to the UK housing ombudsman rose from 2,600 in 2022 to just over 7,000 last year, with 5x growth in complaints to the US Consumer Financial Protection Bureau over the same period. Researcher Chris Schmitz, who calls the pattern “agentic flooding”, examined 84 cases across 11 jurisdictions. Most of those filings come from people with legitimate claims, he told TechCrunch, but somebody still has to read them.
None of this is a Western preoccupation alone. On WIRED’s Uncanny Valley podcast, Will Knight said agentic safety was a theme at a Beijing conference he attended this summer, and that “cybersecurity was a really major topic there”. He said researchers there worried about the same things as people in the US, and described one cybersecurity researcher who can’t collaborate with US counterparts because restrictions don’t allow it.
What’s worth watching is whether a major lab ships any of the friction. Strategic friction, bounded autonomy and rotation policies are testable, so the first vendor to publish override rates or review-time data alongside a capability score would tell you the argument landed.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
