Bitcoin

Bitcoin

$77,213.55

BTC 0.28%

Ethereum

Ethereum

$2,106.63

ETH 0.42%

  • Login
  • Register
Metaverse Media Group
  • Home
  • Crypto
  • NFTs
  • Artificial Intelligence
  • More
    • Technology
    • Business
    • Newsletter
No Result
View All Result
  • Home
  • Crypto
  • NFTs
  • Artificial Intelligence
  • More
    • Technology
    • Business
    • Newsletter
No Result
View All Result
Metaverse Media Group

AI agents do more of the work in model development, but humans still make the decisions

AI agents do more of the work in model development, but humans still make the decisions

The Decoderby The Decoder
27 September 2026
A research team analyzed 769 task logs from building its own AI model. AI agents supplied up to 55 percent of method proposals, but humans made more than 85 percent of final decisions. A third of the tasks wouldn’t have been attempted without AI. The authors warn that more agent activity doesn’t mean more autonomy. The article AI agents do more of the work in model development, but humans still make the decisions appeared first on The Decoder….


Jonathan Kemper


Sep 27, 2026

Image description

Nano Banana Pro prompted by THE DECODER

A research team documented how humans and AI agents worked together to build a new AI model. The findings challenge some expectations about how independently agents can work.

When AI agents help build new AI models, who makes the decisions? A team involving researchers from China’s Fudan University studied its own project to find out. It analyzed more than 700 task logs from 56 participants, along with logs from the agents they used.

The project centered on developing an agentic language model called Atria Dawn Preview, built on a mixture-of-experts architecture with 744 billion parameters and designed for research and engineering tasks.

The model was trained through a pipeline that ties each task to a real execution environment. It calls tools, generates intermediate results, and gets checked against external signals like tests, metrics, or source evidence. The team says it leads on five of 16 benchmarks, including web search and cybersecurity, though it doesn’t hold an overall edge over competitors.

Eight bar charts comparing Atria Dawn Preview with DeepSeek V4 Pro, Kimi K3, Qwen 3.8 Max, GLM 5.3, GPT 5.6 Sol, and Claude Opus 5 across AutomationBench, SkillsBench, Workspace-Bench-Lite, GDPval, CyberGym, MLE-Bench Lite, DeepResearch Bench II, and SWE-Bench Pro.
Atria Dawn Preview leads in AutomationBench, CyberGym, and MLE-Bench Lite but trails the field in GDPval and SWE-Bench Pro. | Image: Atria Team

A third of completed AI-assisted tasks wouldn’t have been attempted without AI

AI was used in 96.5 percent of the tasks reviewed. Over the course of the project, participants handed off more and more to agents. The median ratio of agent actions to human inputs rose from 11 to 28.5 over four weeks. The team cautions against reading this as growing autonomy. Each human decision led to more agent steps, which didn’t mean the agents were making more decisions themselves.

Line chart showing the daily median of agent actions per human prompt from August 7 to September 4, 2026, rising from 11.0 to 28.5 with a gray shaded interquartile range.
Over four weeks, the median number of agent actions per human input rose from 11 to 28.5. | Image: Atria Team

Participants were also asked whether they could have completed their share of a task without AI, at the same scope and quality. Of 455 completed AI-assisted tasks, 151 were rated infeasible without AI, roughly a third. These tasks were spread across 27 of the 56 participants, so they didn’t come from just a handful of power users. AI didn’t speed up existing work in these cases. It made work possible that would never have been started otherwise.

Heatmap table showing estimated time effort without AI by task category, ranging from under half an hour to infeasible, with the infeasible share at 33.2 percent overall.
Participants said about a third of AI-assisted tasks couldn’t have been completed without AI. | Image: Atria Team

AI proposes, humans choose

For methods and parameters, the most common pattern was “AI proposes, human selects” at 55.4 percent. Overall, humans made 85.5 percent of decisions about methods and parameters, while AI made just 9.2 percent. Humans made the final decision on goals and scope in 93.4 percent of cases.

AI’s share of proposals ranged from 17 to 55 percent depending on the decision type. Its share of final decisions stayed in the single digits. Who proposed the options varied widely, but humans consistently made most of the final choices. Even among the 151 tasks rated infeasible without AI, humans chose the goal 95.4 percent of the time.

Stacked bar chart showing who proposes and who decides on goals, methods, and acceptance criteria, with AI proposing methods in 55.4 percent of cases and humans making the final selection.
Agents often supply the method proposals, but humans make the final call in over 80 percent of cases. | Image: Atria Team

Humans supply context, not manual labor

The same pattern shows up when things go wrong. Of 588 tasks with a recorded difficulty, 76 percent moved forward through human intervention, and in 23 percent the agent solved the problem on its own. Human help almost always came in the form of information, either by adding context or clarifying requirements (35.2 percent) or by diagnosing issues and switching methods (34.7 percent). Humans rarely did the work themselves. Partial edits accounted for 3.2 percent of cases, and full takeovers just 0.7 percent.

Horizontal bar chart showing how tasks progressed after a problem, with 35.2 percent resolved by adding context, 34.7 percent by human diagnosis, and 23.0 percent by the AI recovering on its own.
In three quarters of problem cases, human intervention moved work forward, mostly through context or diagnosis rather than taking over. | Image: Atria Team

When AI outputs needed revision, the AI handled the changes itself 75.4 percent of the time after receiving human feedback. Human judgment, rather than execution, was the bottleneck.

The team describes three phases in AI’s role, from a subject of research to a tool for individual tasks and now a project partner. In that current role, AI drafts and adjusts plans within goals set by humans. A speculative fourth phase would involve recursive self-improvement, with stronger models producing stronger successors.

Four-panel diagram showing a human and a robot in the roles of research object, task-level runner, project-level coworker, and a question mark for the next stage.
The Atria team describes AI’s evolution from research object to project partner. The next stage, recursive self-improvement, remains an open question. | Image: Atria Team

The authors say a model can improve at its training tasks without getting better at developing its successor. How AI could propose varied research directions and assess their value before results are available remains an open question.

The rubber-stamp risk

When every decision rests on a longer chain of agent work than any human can review, oversight gets hard. In the worst case, humans become reviewers who can only rubber-stamp what they see, the team writes. Many participants also ran agents in autonomous modes to avoid interrupting long runs with constant approvals. That boundary was drawn out of convenience, not from any deliberate choice about how much authority AI should have.

The paper lands in the middle of a debate about recursive self-improvement. Anthropic considers an AI that develops its own successor possible sooner than expected, and CEO Dario Amodei is calling for a speed limit for the industry as a result. According to Anthropic, humans now make only a single-digit percentage of decisions about research direction at the company. OpenAI uses GPT-5.6 Sol across its entire development cycle, while Google and DeepMind let AI agents explore alternative strategies through recorded search trajectories with Dream-RSI, though they only improve the search strategy, not the model itself.

Over a thousand employees at leading AI companies recently warned that their organizations may be on the verge of automating AI research. A separate study from Princeton and the UK AI Security Institute reached a conclusion more in line with the Atria team’s findings, showing that frontier models can handle research engineering but fail at the judgment calls that actually matter.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive “AI Radar” frontier report six times a year, full archive access, and access to our comment section.


Subscribe now

Read on for the full picture.
Subscribe for hype-free coverage.

  • Full access to every article on THE DECODER
  • No ads
  • Join the comments and community discussions
  • A weekly AI news recap via mail
  • 6x/year: “AI Radar” — deep dives on the AI topics that matter most
  • Daily AI news, always up to date
  • Our full ten-year archive
  • Covered by a team with 10+ years in AI


Subscribe to The Decoder

Read the full article on The-Decoder.com
in AI
Reading Time: 7 mins read
0
0
24
VIEWS
Share on TwitterShare on Facebook

Subscribe to our newsletter

For the latest news & monthly prize giveaways
Join Now

Subscribe to our newsletter

For the latest news & monthly prize giveaways
Join Now
ADVERTISEMENT

Related Posts

Researchers plug GPT-6 Astra directly into a robot and let it clean up an unfamiliar kitchen
AI

Researchers plug GPT-6 Astra directly into a robot and let it clean up an unfamiliar kitchen

12 hours ago
23
Tens of thousands of security probes show OpenAI’s Hugging Face incident was just the beginning
AI

Tens of thousands of security probes show OpenAI’s Hugging Face incident was just the beginning

14 hours ago
24
xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6
AI

Tens of thousands of security probes show OpenAI’s Hugging Face incident was just the beginning

14 hours ago
26

Comments

Please login to join discussion
ADVERTISEMENT

Latest News

  • All
  • Crypto
  • NFTs
  • Technology
  • Business
Crypto Weekly: BCH and NEAR Surge 34% as Altcoins Outpace Bitcoin
Crypto

Crypto Weekly: BCH and NEAR Surge 34% as Altcoins Outpace Bitcoin

Bitcoin.com News
by Bitcoin.com News
21 minutes ago
19
We Asked 8 AI Chatbots to Predict Bitcoin’s Price, Here’s the Verdict
Crypto

We Asked 8 AI Chatbots to Predict Bitcoin’s Price, Here’s the Verdict

Bitcoin.com News
by Bitcoin.com News
1 hour ago
24
Buterin Says Ethereum’s Last ‘Normal’ Fork Is Planned for 2027
Crypto

Buterin Says Ethereum’s Last ‘Normal’ Fork Is Planned for 2027

Bitcoin.com News
by Bitcoin.com News
2 hours ago
24
Thorchain Faces Heat as Bitget Hack Revives Bybit Controversy
Crypto

Thorchain Faces Heat as Bitget Hack Revives Bybit Controversy

Bitcoin.com News
by Bitcoin.com News
5 hours ago
22
Same Name, Different Company: Blockstream Addresses $6.7M River Lawsuit
Crypto

Same Name, Different Company: Blockstream Addresses $6.7M River Lawsuit

Bitcoin.com News
by Bitcoin.com News
6 hours ago
23
Peter Brandt Says XRP Charts Alone Give Him Reason to Make a Bet
Crypto

Peter Brandt Says XRP Charts Alone Give Him Reason to Make a Bet

Bitcoin.com News
by Bitcoin.com News
6 hours ago
23
Load More
Next Post
Peter Brandt Predicts $8,600 ETH, Coinbase Adds IPO Share Trading, And More – Week In Review

Peter Brandt Predicts $8,600 ETH, Coinbase Adds IPO Share Trading, And More – Week In Review

ADVERTISEMENT

Follow Us

Categories

  • Crypto
  • NFTs
  • AI
  • Technology
  • Business
  • Crypto
  • NFTs
  • AI
  • Technology
  • Business
Subscribe to our Newsletter

© 2022 Metaverse Media Group – The Metaverse Mecca

Privacy and Cookie Policy | Sitemap

Welcome Back!

Sign In with Google
OR

Login to your account below

Forgotten Password? Sign Up

Create New Account!

Sign Up with Google
OR

Fill the forms below to register

*By registering into our website, you agree to the Terms & Conditions and Privacy Policy.
All fields are required. Log In

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Crypto
  • NFTs
  • Artificial Intelligence
  • More
    • Technology
    • Business
    • Newsletter
Bitcoin

Bitcoin

$77,213.55

BTC 0.28%

Ethereum

Ethereum

$2,106.63

ETH 0.42%

  • Login
  • Sign Up
This website uses cookies. By continuing to use this website you are giving consent to cookies being used. Visit our Privacy and Cookie Policy.

Subscribe to our newsletter

Get the latest news & win monthly prizes

Subscribe to our newsletter

For the Latest News and Monthly Prize Giveaways

Join Now
Join Now