ICML position paper: evaluate human-AI teams, not autonomous AI
A position paper accepted to the ICML 2026 Position Paper Track argues that the way the field measures AI is steering development the wrong way. Jan Kulveit, Gavin Leech, Tomáš Gavenčiak and Raymond Douglas posted AI Evaluation Should Work With Humans to arXiv on July 6. Their claim is that benchmarks chasing superhuman autonomous performance implicitly target the goal of replacing people.
The proposed fix is a pivot, and it’s a simple one to state. Score the performance of human-AI teams rather than the model working alone. The authors write that the dominant paradigm “is guiding AI development in the wrong direction”, and they argue a collaborative shift would produce systems that act as true complements to human capabilities. The abstract opens by calling itself a position paper, so it’s an argument about direction rather than a set of measurements.
The AI community should pivot to evaluating the performance of human-AI teams.
Kulveit, Leech, Gavenčiak and Douglas, via arXiv
A separate group has already tried running evaluation that way. In Research-Oriented Human-Centric Evaluation for Foundation Models, revised on August 14, Yijin Guo and seven co-authors report 604 human evaluation sessions across various disciplines. Their framework captures user perceptions across three dimensions: problem-solving ability, information quality, and interaction experience.
Those sessions weren’t quiz-style, which is the whole point of the design. The project’s GitHub repository says participants choose a task based on their major and interests, then interact freely with a foundation model for 20 minutes before completing a questionnaire. The team used Pearson correlation analysis to check the validity of the evaluation dimensions. The repository puts the motivation plainly: quiz performance “is hard to reflect human experience”.
The most awkward result arrived when they swapped the humans out. The authors ran an LLM-as-a-judge experiment and found that even sophisticated models “struggle to accurately replicate human subjective judgment”. They read that as evidence for what they call the irreplaceable value of first-person human assessment. That matters because the same substitute, a model standing in for a human grader, shows up again in the papers below.
| Paper | Method | Finding |
|---|---|---|
| AI Evaluation Should Work With Humans | ICML 2026 position argument, no experiment | Evaluate human-AI teams, not autonomous performance |
| Research-Oriented Human-Centric Evaluation | 604 human sessions, 20 minutes of free interaction each | Models struggle to replicate human subjective judgment |
| Agreement Is Not Alignment | 500-item ETHICS-derived benchmark, five moral domains | Matching labels can hide divergent reasoning |
| Evaluating Agentic Learning Harness Capabilities Without Labels | Teacher-model scoring in place of labelled benchmarks | Peer LLM judging yields no usable signal |
Two other papers push at the same seam. Agreement Is Not Alignment, presented at the AI Transparency Conference in Nuremberg in June, tests a curated 500-item ETHICS-derived benchmark spanning five domains of moral judgment. Agreement with human annotator majority labels is often high, yet the rationales behind those labels diverge systematically, so the authors conclude that label-based evaluation “can therefore be misleadingly reassuring unless complemented by analysis of the reasons, principles, and moral priorities expressed in model judgments”. A separate evaluation paper reports that LLM-as-a-judge between similarly powered models “yields no usable signal”.
None of this is fresh pressure on the scoreboard, though the direction is getting harder to ignore. We covered the Stanford AI Index finding that benchmarks now saturate within months, and our report on agent behavioral consistency made a similar case for measuring process instead of outcome. The catch is cost, because 604 human sessions don’t scale the way an automated benchmark does.
So the open question is whether an ICML badge moves anything in practice. A team score doesn’t compress into a single leaderboard row, and the cheap substitute, a model grading a model, is what these papers keep finding fault with. The signal worth watching is the first model card that carries a human-AI team result next to the solo benchmark number.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
