A research team documented how humans and AI agents worked together to build a new AI model. The findings challenge some expectations about how independently agents can work.
When AI agents help build new AI models, who makes the decisions? A team involving researchers from China’s Fudan University studied its own project to find out. It analyzed more than 700 task logs from 56 participants, along with logs from the agents they used.
The project centered on developing an agentic language model called Atria Dawn Preview, built on a mixture-of-experts architecture with 744 billion parameters and designed for research and engineering tasks.
The model was trained through a pipeline that ties each task to a real execution environment. It calls tools, generates intermediate results, and gets checked against external signals like tests, metrics, or source evidence. The team says it leads on five of 16 benchmarks, including web search and cybersecurity, though it doesn’t hold an overall edge over competitors.

A third of completed AI-assisted tasks wouldn’t have been attempted without AI
AI was used in 96.5 percent of the tasks reviewed. Over the course of the project, participants handed off more and more to agents. The median ratio of agent actions to human inputs rose from 11 to 28.5 over four weeks. The team cautions against reading this as growing autonomy. Each human decision led to more agent steps, which didn’t mean the agents were making more decisions themselves.

Participants were also asked whether they could have completed their share of a task without AI, at the same scope and quality. Of 455 completed AI-assisted tasks, 151 were rated infeasible without AI, roughly a third. These tasks were spread across 27 of the 56 participants, so they didn’t come from just a handful of power users. AI didn’t speed up existing work in these cases. It made work possible that would never have been started otherwise.

AI proposes, humans choose
For methods and parameters, the most common pattern was “AI proposes, human selects” at 55.4 percent. Overall, humans made 85.5 percent of decisions about methods and parameters, while AI made just 9.2 percent. Humans made the final decision on goals and scope in 93.4 percent of cases.
AI’s share of proposals ranged from 17 to 55 percent depending on the decision type. Its share of final decisions stayed in the single digits. Who proposed the options varied widely, but humans consistently made most of the final choices. Even among the 151 tasks rated infeasible without AI, humans chose the goal 95.4 percent of the time.

Humans supply context, not manual labor
The same pattern shows up when things go wrong. Of 588 tasks with a recorded difficulty, 76 percent moved forward through human intervention, and in 23 percent the agent solved the problem on its own. Human help almost always came in the form of information, either by adding context or clarifying requirements (35.2 percent) or by diagnosing issues and switching methods (34.7 percent). Humans rarely did the work themselves. Partial edits accounted for 3.2 percent of cases, and full takeovers just 0.7 percent.

When AI outputs needed revision, the AI handled the changes itself 75.4 percent of the time after receiving human feedback. Human judgment, rather than execution, was the bottleneck.
The team describes three phases in AI’s role, from a subject of research to a tool for individual tasks and now a project partner. In that current role, AI drafts and adjusts plans within goals set by humans. A speculative fourth phase would involve recursive self-improvement, with stronger models producing stronger successors.

The authors say a model can improve at its training tasks without getting better at developing its successor. How AI could propose varied research directions and assess their value before results are available remains an open question.
The rubber-stamp risk
When every decision rests on a longer chain of agent work than any human can review, oversight gets hard. In the worst case, humans become reviewers who can only rubber-stamp what they see, the team writes. Many participants also ran agents in autonomous modes to avoid interrupting long runs with constant approvals. That boundary was drawn out of convenience, not from any deliberate choice about how much authority AI should have.
The paper lands in the middle of a debate about recursive self-improvement. Anthropic considers an AI that develops its own successor possible sooner than expected, and CEO Dario Amodei is calling for a speed limit for the industry as a result. According to Anthropic, humans now make only a single-digit percentage of decisions about research direction at the company. OpenAI uses GPT-5.6 Sol across its entire development cycle, while Google and DeepMind let AI agents explore alternative strategies through recorded search trajectories with Dream-RSI, though they only improve the search strategy, not the model itself.
Over a thousand employees at leading AI companies recently warned that their organizations may be on the verge of automating AI research. A separate study from Princeton and the UK AI Security Institute reached a conclusion more in line with the Atria team’s findings, showing that frontier models can handle research engineering but fail at the judgment calls that actually matter.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive “AI Radar” frontier report six times a year, full archive access, and access to our comment section.
Read on for the full picture.
Subscribe for hype-free coverage.
- Full access to every article on THE DECODER
- No ads
- Join the comments and community discussions
- A weekly AI news recap via mail
- 6x/year: “AI Radar” — deep dives on the AI topics that matter most
- Daily AI news, always up to date
- Our full ten-year archive
- Covered by a team with 10+ years in AI








