Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems
Patrick Emami, Sameera Horawalavithana, Truc Nguyen, Gihan Panapitiya, Bruno Jacob, Siddhisanket Raskar, Saumya Sinha, Jared D. Willard, Andrew Glaws, Nithin Somasekharan, Ling Yue, Brian Lu, Shaowu Pan, Jason Eisner
COLM 2nd Workshop on Languag
cs.AI, cs.HC
2026-08-02
A position paper from four US labs argues AI Scientists should be studied as human-agent systems: most current systems stop at humans setting goals and grading artifacts, AI-augmented research triples paper output while cutting topic diversity 5%, and GPT-5-mini asks for help once across ten tasks.
Papers on AI Scientist systems keep arriving with a consistent direction: toward higher autonomy, with less human involvement celebrated as progress. This position paper (National Laboratory of the Rockies, Pacific Northwest National Laboratory, RPI, Johns Hopkins) argues that direction skips an entire layer: science is social, teams are 6.3 times more likely than solo authors to produce a highly cited paper, and the collaboration dynamics of inserting agents into scientific teams are exactly what the autonomy narrative does not cover.
The evidence is already accumulating. A 2026 study found AI-augmented research produces three times as many papers at a 5% reduction in the diversity of explored topics. A post-hoc review found 51 accepted NeurIPS 2025 papers containing 100 confirmed hallucinated citations, and ACL 2026 similarly flagged over 100 accepted papers citing non-existent literature. These are the direct consequences of deploying agents without accounting for human-agent dynamics.
The position: shift the unit of analysis from "can the agent solve it autonomously" to the human-agent pair, a direction both underexplored and undervalued.
As a position paper, the argument runs in four steps.
Step one, a literature review classifying existing systems by the primary channel through which human feedback enters the discovery loop. Artifact-level feedback (humans only set goals and grade final outputs) is the largest group, holding AI Scientist v2, Kosmos, CycleResearcher, VirSci, and AgentRxiv, all evaluative, coarse-grained, and post-task. Discovery-phase-level systems (InternAgent, Co-Scientist, AgentLaboratory, Denario) gate input at phase boundaries. Research-plan-level systems (MAPPS, El Agente Q, Organa, SciSciGPT) follow plan-critique-execute with humans shaping course early. Asynchronous steering and interruption (FreePhDLabor, EvoScientist, OmniScientist, ScienceClaw) permits mid-task intervention at arbitrary points, the finest control grain and the rarest.
Step two adds a small experiment. On the AstaBench end-to-end discovery benchmark, a ReAct GPT-5-mini agent gets an ask-user tool across ten tasks. The bare harness asks zero questions; adding the tool plus taxonomy guidance yields one; enforcing a question before every action yields 92. Frontier models do not know when to ask, preferring to assume their way past missing critical resources.
Step three argues risks along three lines: reliability and bias (hallucinated citations, paradigm homogenization, homogenized writing), institutions and skills (NeurIPS and ICLR submissions up 10.4x over a decade without matching reviewer growth, reduced conceptual understanding from AI coding assistants, higher cognitive offloading among younger users), and accountability (36% of LLM-generated papers containing noticeable plagiarism, NIST naming CBRN risk amplification).
Step four offers a utility model: team utility equals collaboration advantage minus collaboration disadvantage. CA measures team performance against its best member alone (complementarity); CD captures communication and coordination costs, tokens spent and human cognitive load. The research program is finding environmental factors and team dimensions that raise CA or lower CD, with three operational proxy metrics: review time for cognitive load, intervention density as edit distance between agent artifacts and the human's final version, and simulated oracle synergy using a stronger LLM as a proxy human to measure how well an agent ingests critical feedback.
For a position paper, the "results" are qualitative unpackings of two case studies.
Case one: a complexity theorist and Gemini 3 Pro inside Google Antigravity turning a long-shelved theorem into a paper. How AI augmented the human: within eight prompt turns, a roadmap and LaTeX draft, iteratively expanded proofs, pushing an informal result toward publishable. How the human augmented the AI: the expert early identified missing technical context (relevant prior work, a search-versus-decision framing) and corrected an erroneous assumption in a lemma. The authors' judgment: without that expert diagnosis, the same workflow would likely produce a misleading paper. The pattern resembles expert plus junior collaborator.
Case two: GPT-5 helping a physicist build a reduced model of thermonuclear burn propagation in four linked activities, a PDE model for intuition, numerical implementation, experiments exposing sensitivities, and a theory explaining the simulations. The expert reported the strongest outcome was not time saved but the fourth: a theory-level closed-form predictor verifying the numerically elicited trends. The human's role sat in numerical tuning (hot-spot and cold-fuel conditions, conductivities, alpha-particle stopping assumptions) and epistemic checking (recognizing when outputs were still noise, rejecting silent substitution of full solves with coarse approximations). The shared risk across both cases: premature confidence and smoothing over thorny issues, requiring expert oversight.
For teams building research agents, the feedback-taxonomy table is a ready-made design coordinate system: locate your system, see what separates it from the asynchronous-interruption tier. The utility model turns "is human-AI synergy worth it" from talk into something measurable, and the three proxy metrics drop into evaluations directly. The ask-user data point is practical: without enforcement, models do not ask, which means prompting will not fix help-seeking, interaction design or training must.
For research institutions, the risk section's citation list (3x output with 5% diversity loss, reviewer load, skill erosion) is direct ammunition for AI-use policy.
A position paper's evidence is inherently weak: the literature classification rests on the authors' framework choices; the ten-task experiment covers deep but not broad territory (each task is a full idea-to-report workflow); the case studies come from published optimistic accounts with no correction for publication bias. The utility model is toy-grade, two terms with no accounting for measurement cost itself (conceded in the alternative-views section), and CA and CD are hard to separate empirically. The taxonomy assigns systems to a primary feedback channel while table footnotes admit most systems support several, so boundary calls are subjective. Both case studies pair experts with frontier models; the early-career-researcher-with-weaker-model combination, which the risk section predicts goes worse, has no corresponding data. No controlled human-agent comparison experiment appears anywhere, which is precisely what the paper calls on others to run, so every claim stays at "worth doing" rather than "done and effective".