Humans Still Call the Shots in AI‑Assisted Model Development, Study Finds
A Fudan University team’s analysis of 700+ task logs shows AI agents handle most routine steps, but humans make the overwhelming majority of strategic decisions.

What happened, who, and when
On September 27 2026, researchers from China’s Fudan University published a study examining how human participants and AI agents collaborated to build a new agentic language model, Atria Dawn Preview. The investigation covered more than 700 task logs generated by 56 participants working with the model’s development pipeline.
Key findings from the study
- AI involvement is near‑ubiquitous: AI agents were used in 96.5 % of the reviewed tasks.
- Increasing reliance on agents: Over a four‑week period the median ratio of agent actions to human inputs rose from 11 to 28.5, indicating participants handed off more work to the agents as the project progressed.
- Tasks enabled by AI: Of 455 completed AI‑assisted tasks, 151 (about one‑third) were judged infeasible without AI, a sentiment shared across 27 of the 56 participants.
- Decision‑making split: For choices about methods and parameters, “AI proposes, human selects” accounted for 55.4 % of interactions. Humans made 85.5 % of decisions on methods/parameters and 93.4 % of final decisions on goals and scope, while AI contributed only 9.2 % of method/parameter decisions and remained in single‑digit percentages for final choices.
- Human intervention in problem cases: In 588 tasks with recorded difficulty, humans intervened in 76 % of cases, primarily by adding context (35.2 %) or diagnosing issues and switching methods (34.7 %). Agents solved 23 % of difficult tasks on their own. Full human takeovers occurred in just 0.7 % of cases; partial edits were 3.2 %.
- Feedback loop: After receiving human feedback, AI revised its own outputs 75.4 % of the time, suggesting execution, not judgment, is the bottleneck.
- Evolution of AI role: The authors outline three phases—research subject, task‑level tool, and project partner—where AI now drafts and adjusts plans within human‑set goals. A speculative fourth phase would involve recursive self‑improvement, though the study notes current models can improve at training tasks without becoming better at designing successors.
- Rubber‑stamp risk: The paper warns that as agents handle longer chains of work, human oversight may degrade to mere “rubber‑stamping,” especially when participants run agents autonomously for convenience rather than by design.
Industry context and related developments
The findings arrive amid a heated debate over recursive self‑improvement. Anthropic’s CEO Dario Amodei has called for a speed limit on AI research, noting that at Anthropic humans now make only single‑digit percentages of research‑direction decisions. OpenAI reportedly employs GPT‑5.6 Sol throughout its development cycle, while Google and DeepMind use agents like Dream‑RSI to explore alternative search strategies, though not to redesign models themselves. Over a thousand AI‑industry employees have recently warned that their companies may be on the brink of automating AI research entirely. A separate study by Princeton and the UK AI Security Institute aligns with the Fudan results, showing frontier models excel at research engineering but falter on the high‑level judgment calls that ultimately matter.

