Literature research · Landscape survey · run date 2026-09-25
AI agents in scientific workflows
What AI agent systems do across the scientific workflow (literature and ideation, hypothesis generation, experiment design and self-driving labs, data analysis and simulation, writing and review), what has been shown to work, and what it means for DOE x-ray and neutron user facilities, focused on 2024–2026.
Prepared for DOE x-ray and neutron facility scientists (SLAC, ORNL, BNL) and university collaborators
At a glance
- Agents now reach every stage of research, but full autonomy is rare. Of the 83 systems cataloged here, 43 are tool-using agents working inside one task (L2), 21 run closed loops against instruments, robots or code (L3) and 10 run end to end (L4). The L4 systems sit where the "experiment" is code, existing data or a proof, such as machine-learning papers, data-driven reports and mathematics [1, 2, 3] (Q1).
- In 2026 the flagship "AI scientist" systems passed journal peer review. Co-Scientist, Robin, the AI Scientist and OpenScholar appeared in Nature, and Biomni in Science [1, 4, 5, 6, 7]. Where they report wet-lab results, these are still in vitro or organoid tests, and independent replication is rare (Q2).
- Benchmarks: scores on closed-form questions are near the ceiling; on open-ended research they are low. Top models now exceed PhD-expert accuracy on graduate-level science questions, but agents still score well below human experts at replicating papers and rarely finish complete research experiments [8, 9, 10, 11] (Q2).
- At DOE facilities, agents have moved from chatbots to supervised control. Peer-reviewed 2026 work shows an agent aligning a crystal on a real SSRL beamline, an agent running multistage experiments on the ALS accelerator, and an agent taking SNS neutron data to a validated structure file. Every one keeps a human checkpoint or a tool-level guard, and facility steering is still mostly single-campaign demonstrations [12, 13, 14, 15] (Q4).
- The infrastructure is converging. Common pieces are the Model Context Protocol (MCP), a standard way for agents to call tools; open-weight models served from DOE computing centers; and Globus, Parsl and DOE's Integrated Research Infrastructure (IRI) program for reaching HPC [16, 17, 18, 19] (Q3).
- The main risks are unverified results and unsafe tool actions. They include fabricated references and reports, reward hacking (an agent gaming its success check instead of doing the task), weak evaluation and prompt injection (hidden instructions that hijack an agent). The Genesis Mission, a November 2025 executive order, now sets the DOE policy frame [20, 21, 22, 23] (Q5; policy in Q4 and Q6).
Q1 · Landscape: where agents sit in the scientific workflow
Answer. Of 83 cataloged systems, about half are L2 tool-using agents; closed-loop (L3) agents cluster in self-driving labs and at facility instruments; end-to-end (L4) autonomy appears mainly where the "experiment" is code, existing data or a proof; and independent verification lags [1, 3, 24].
Taxonomy: stage × autonomy
Here an agent is a large language model (LLM) that calls tools (search, code, simulations, instrument controls) and picks its next step from the results. We sorted 83 systems by main workflow stage and by autonomy: 58 general-science systems (mostly 2024–2026) and 25 facility systems covered in Q4. The general set includes national-lab tools such as LLaMP (LBNL) and ChemGraph (Argonne) [25, 26].
Autonomy levels. L1 Assistant: answers or retrieves on request. L2 Tool-using agent: plans and runs multi-step tool calls within one task. L3 Closed-loop: iterates plan → act → observe against real instruments, robots, simulations or code over many cycles, with human checkpoints. L4 End-to-end: goes from a goal through experiments or analysis to a written paper with minimal human help. Stages: literature & ideation; hypothesis generation; experiment design & self-driving labs (labs where software chooses and runs the next experiment); data analysis & simulation; and writing & review (of papers and proposals). Each system sits under its main stage.
| Stage | L1 Assistant | L2 Tool-using agent | L3 Closed-loop | L4 End-to-end | Total |
|---|---|---|---|---|---|
| Literature & ideation | ChatEED ESAC | Ai2 Asta agents + AstaBench Ai2 Scholar QA FutureHouse Platform (Crow, Falcon, Owl, Phoenix) Gemini Deep Research LLM ideation agent vs. 100+ NLP researchers OpenScholar PaperQA2 ResearchAgent STORM APS-RAG CALMS GAIA | 14 | ||
| Hypothesis generation | Co-Scientist MOOSE-Chem POPPER SciAgents SciMON | DeepScientist FunSearch InternAgent (NovelSeek) Robin | Aletheia | 10 | |
| Experiment design & self-driving labs | LLM-assisted scanning probe microscopy | AutoLabs ChemCrow CRISPR-GPT Virtual Lab ALD reactor agent ALS accelerator agentic AI Instrument agents that learn on the job Lightfall Osprey TEM Agent VISION | A-Lab BioDiscoveryAgent ChemAgents Coscientist CRESt k-agents MARS Mobile-robot synthesis lab ORGANA A-Lab GPSS agentic reasoning AI X-ray scientist Experiment Automation Agents (EAA) LLM accelerator tuning NSLS-II multi-beamline AI agents | 26 | |
| Data analysis & simulation | AtomAgents Biomni ChemGraph Data Interpreter DS-Agent El Agente Q LLaMP MDCrow APEXA EQSANS-CLI NeuDiff Agent PEAR SasAgent | AlphaEvolve Empirical Research Assistance (ERA) Rongzai agent | Agent Laboratory AutoResearchClaw Denario Kosmos | 20 | |
| Writing & review | AgentReview CycleResearcher / CycleReviewer GPT-4 paper feedback study Prism Review Feedback Agent (ICLR 2025) LLM proposal ranking | DeepReview Stanford Agentic Reviewer | AgentRxiv AI-Researcher data-to-paper The AI Scientist (v1, v2) Zochi | 13 | |
| Total | 9 | 43 | 21 | 10 | 83 |
Reading the table:
- L2 is the norm (43 of 83). Literature & ideation is almost all L2 (12 of 14): retrieve-read-cite agents such as PaperQA2, OpenScholar and Ai2's Asta [6, 27, 28]. They hand a cited answer to a person, so there is no loop to close.
- Most closed loops run in labs. Experiment design & self-driving labs is the largest stage (26) and holds 14 of the 21 L3 systems, such as Coscientist (Carnegie Mellon, 2023; not Google's Co-Scientist) and MARS [29, 30]. Half of this stage (13 of 26) and 6 of the L3 systems are facility systems, such as SLAC's AI X-ray scientist, which aligned a crystal on an SSRL beamline while a human relayed its commands [12].
- No experiment-stage system is L4. In wet labs, humans still run or approve experiments [5, 31]. The 10 L4 systems, none of them at a facility, sit where the "experiment" is code, existing data or a proof: machine-learning papers (The AI Scientist), data-driven reports (Kosmos, data-to-paper) and mathematics (Aletheia) [1, 2, 3, 32]. Two are facility crystallography agents, such as ORNL's NeuDiff Agent [15].
- L1 is rare (9). Six of the 9 are writing & review tools, most of which inform rather than replace human authors and reviewers [33, 34].
Six trends, 2024 → 2026
- From single tools to multi-agent "AI scientists." Systems from 2023–24 wrapped one LLM around domain tools; ChemCrow used 18 chemistry tools [35]. By 2025–26 the usual design is a team of role-specialized agents that critique each other: Co-Scientist runs a generate-critique-refine tournament [4], Virtual Lab has a "PI" agent directing specialists [31], and Kosmos (a preprint) links about 200 agent rollouts per run through a shared "world model", a common record of findings [2].
- Flagship claims reached top journals. In 2026 Co-Scientist, Robin, The AI Scientist and OpenScholar appeared in Nature, and Biomni in Science [1, 4, 5, 6, 7].
- Companies and frontier labs entered. Google and Google DeepMind cover hypotheses, algorithms and mathematics (Co-Scientist, AlphaEvolve, Aletheia) [3, 4, 36]. FutureHouse spun out Edison Scientific to commercialize its AI Scientist; Kosmos is built from Edison Scientific agents [2, 37]. Intology says a paper written by its Zochi agent was accepted at the ACL 2025 main conference [38].
- Autonomy is highest where checking is cheap. Strong L3 and L4 results appear where an automatic check exists (a benchmark score, code evaluator or proof check). A white paper reports that AlphaEvolve found a way to multiply 4×4 complex matrices with 48 scalar multiplications [36]; DeepScientist tested about 1,100 ideas in a month [39]. At Agents4Science 2025, a conference for papers with AI first authors, humans contributed more to hypotheses and design, and accepted papers had more human guidance [40].
- Evaluation lags capability (the "verification gap"). Independent checks are rare and mixed. An academic lab, co-authoring with Google, found that Co-Scientist's top hypothesis matched its own unpublished result [41]; an independent analysis disputed A-Lab's novelty claims [42]. LLM ideas that experts rated more novel than human ideas lost that edge once carried out [43, 44]. About one in five Kosmos statements was not judged accurate [2]. A 2026 survey (preprint) found that 38% of 24 runnable systems report any novelty verification [24]; AstaBench concludes that AI remains far from solving research assistance [28].
- AI enters peer review; targeted human oversight beats full autonomy. GPT-4 feedback on papers overlapped with human reviews about as much as two reviewers overlap [33]. In a randomized ICLR 2025 trial, 27% of reviewers given LLM feedback updated their reviews [34]. The AI Scientist's authors warn of "taxing overwhelmed review systems" [1]. A preprint finds that targeted human interventions in AutoResearchClaw beat both full autonomy and step-by-step oversight [45].
What is new in 2026
Besides the journal papers and the peer-review trial above, preprints report that Aletheia wrote one paper with no human intervention, solved four open Erdős problems autonomously, and solved 6 of 10 FirstProof problems [3, 46]. MARS pushed closed-loop autonomy in robotic materials labs [30], and OpenAI's Prism put an LLM inside LaTeX manuscript editing [47]. Critiques sharpened: a position paper argues that current agents are co-scientists, not built for autonomous discovery [48]. In the US, the November 2025 Genesis Mission order directs DOE to build a platform that includes AI agents and autonomous experimentation [23].
Caveats. Autonomy labels are our judgment from each paper's own description. Vendor and company claims (Zochi, Prism, Edison Scientific) are not independently verified. Several headline numbers (Kosmos, AlphaEvolve, Aletheia) come from preprints or white papers.
Q2 · What works: evidence versus claims, and what benchmarks show
Answer. Peer-reviewed studies show real but narrow agent successes in labs and literature. Top models now score 95.8% on graduate science questions, above expert level, but open-ended research and XRD analysis stay weak, and no public benchmark tests closed-loop x-ray or neutron beamtime [4, 9, 49].
Established: peer-reviewed or machine-checkable
"Established" means peer-reviewed or machine-checkable, not independently replicated.
| Result | Evidence | What was actually shown | Caveat |
|---|---|---|---|
| Coscientist (chemistry) | Peer-reviewed (Nature 2023) | A GPT-4 agent planned and ran robotic experiments, including Pd-catalysed cross-coupling optimization [29]. | No outside replication in our sources. |
| ChemCrow | Peer-reviewed (Nat Mach Intell 2024) | An agent with 18 tools carried out syntheses of an insect repellent and three organocatalysts [35]. | Authors' own demonstration. |
| A-Lab (contested) | Peer-reviewed (Nature 2023); critique (PRX Energy 2024); correction (Nature 2026) | Reported 41 of 58 targets made in 17 days [50]. A 2026 Author Correction cut this to 36 confirmed (4 inconclusive by XRD, 1 removed because it was in the training data) and said "novel" meant new to the prediction platform, not to science [51]. | Outside analysis: no new materials, and automated Rietveld analysis (whole-pattern structure fitting) of powder XRD "not yet reliable" [42]. 2026 Author Correction: the title now says "inorganic", not "novel" [51]. |
| Virtual Lab | Peer-reviewed (Nature 2025) | Agents designed 92 SARS-CoV-2 nanobodies; two bound JN.1 or KP.3 better [31]. | Tests ran in the authors' own lab. |
| CRISPR-GPT | Peer-reviewed (Nat Biomed Eng 2025) | Guided knockout of four genes and activation of two in human cells [52]. | Authors' own demonstration. |
| Google Co-Scientist | Peer-reviewed + partner-lab tests (Nature 2026; Cell 2025; Adv Sci 2025) | Acute myeloid leukaemia drug-repurposing candidates confirmed in vitro [4]. Top hypothesis matched a lab's confirmed but unpublished mechanism (cf-PICIs hijack phage tails) [41, 53]. Two suggested drugs were anti-fibrotic in human liver organoids [54]. | Partner labs co-authored with Google. |
| Robin | Peer-reviewed (Nature 2026) | In vitro tests confirmed ripasudil and KL001 as candidates for dry age-related macular degeneration [5]. | No animal or clinical evidence yet. |
| Biomni | Peer-reviewed (Science 2026) | A general biomedical agent with wet-lab case studies [7]. | Case studies only. |
| AlphaEvolve | White papers (arXiv 2025); machine-checkable | Provably correct new algorithms, e.g. 4×4 complex matrix multiplication with 48 scalar multiplications [36]; with outside mathematicians, improved several best-known constructions [55]. | Not peer-reviewed, but checkable. |
| OpenScholar | Peer-reviewed (Nature 2026) | Citation accuracy on par with human experts; GPT-4o hallucinated (invented) citations 78–90% of the time [6]. | Measures citation accuracy only. |
| Review feedback | Peer-reviewed (Nat Mach Intell 2026; NEJM AI 2024) | In a randomized trial on 20,000+ ICLR 2025 reviews, 27% of reviewers given LLM feedback updated their reviews [34]. GPT-4 comments overlap humans about as much as two humans do (30.85% vs 28.58%) [33]. | Measures change, not review quality. |
| AI Scientist | Peer-reviewed (Nature 2026) | An AI-written paper passed first-round review at a workshop [1]. | The workshop accepted 70%: a low bar. |
| X-ray beamline control | Peer-reviewed (Nat Mach Intell 2026) | An LLM agent found reference reflections and the orientation matrix (crystal orientation on the diffractometer) on a real SSRL beamline [12]. | A human relayed each command, unmodified, for safety. |
| AFM control | Peer-reviewed (Nat Commun 2025) | An agent on a real atomic force microscope scored 88.3% on documentation tasks but 33.3% on analysis [56]. | It "sleepwalked" off its instructions. |
Claims awaiting validation
Numbers are the claimants' own.
| Claim | Source type | What would validate it |
|---|---|---|
| Kosmos (Edison): independent scientists judged 79.4% of statements accurate, only 57.9% of synthesis statements; claims seven discoveries [2]. | Preprint (arXiv 2025) | Peer review, outside replication, and false-positive rates across all runs. |
| Zochi (Intology): says its AI-written paper was accepted at ACL 2025 [57]. | Company report (GitHub) | An audited record of human versus AI work. |
| Periodic Labs Neon: 55.3% on an internal test of 134 multiphase XRD patterns, graded by an LLM judge (a model grading answers); claimed to beat frontier models [58]. | Company blog (2026) | Release the test set, or score against expert Rietveld refinements of public data. |
| Lila Sciences: an AI-guided loop screened 2,942 catalysts; InMnPdOx stayed below 0.5 V overpotential for 1,000 h in acid [59]. | Preprint (arXiv, 2026-09) | Independent synthesis and durability tests. |
| OpenAI GPT-5 cases: four new math results, checked only by the human co-authors [60]. | Preprint (arXiv 2025) | Refereed publication. |
| OpenAI Navier–Stokes (2026-09): Lean (proof-checker) certificates claim finite-time blowup (Clay alternatives C/D) [61]. | Code release (GitHub, 2026) | Experts confirm the formal statements match the Clay problem, then refereed publication. |
What the benchmarks show
"—" means no human baseline was reported.
| Benchmark | Measures | Best reported (system, date) | Human/expert baseline |
|---|---|---|---|
| GPQA Diamond [8] | Graduate-level science multiple choice | 95.8% (GPT-6 Astra, 2026-09, per the Epoch AI benchmarking hub) [9] | Experts 65% on the full question pool (74% discounting clear mistakes); the paper puts expert accuracy on Diamond between 65% and 81% |
| HLE (Humanity's Last Exam) [62] | Expert-written closed questions | 54.8% (GPT 6 Astra, 2026-09-09, per the Scale Labs leaderboard) [63] | None |
| FrontierScience [64] | Olympiad / research tasks | 77% / 25% (GPT-5.2, 2026-01) | — |
| CritPt [65] | Unpublished physics research problems | 32.3% (GPT-5.6 Sol, 2026-07, per the Epoch AI benchmarking hub) [9] | — |
| ChemBench [66] | Chemistry questions | Best models beat the best chemist surveyed (2025) | Surveyed chemists |
| MaCBench [49] | Chemistry/materials images (XRD, AFM) | XRD intensity ranking 0.28 (2025) | — |
| LAB-Bench [67] | Practical biology | Claude 3.5 Sonnet (2024-07) | Experts clearly ahead |
| LabSafety Bench [68] | Lab hazards | <70% hazard identification (2026) | — |
| SciCode [69] | Research code | 10.8% main problems (official leaderboard) [70]; 66.9% subproblems (Claude Opus 5.5, 2026-09, per Artificial Analysis) [71] | — |
| ScienceAgentBench [72] | Code for data-driven discovery | 42.2% (o1-preview, 3 tries, 2024-10) | — |
| CORE-Bench Hard [73] | Reproduce results from code | 95.5% with manual validation (77.8% automated), declared "solved" (Opus 4.5 + Claude Code, per the HAL leaderboard) [74] | — |
| BixBench [75] | Bioinformatics | 17% (2025-02) | — |
| DiscoveryWorld [76] | Simulated discovery | ≤18% completion (2024) | Human scientists (MSc or PhD) 66% |
| MLE-bench [77] | Kaggle machine-learning competitions | Medal in 64.4% (Famou-Agent 2.0, 2026-02, per the MLE-bench leaderboard) [78] | Kaggle leaderboards |
| MLAgentBench [79] | ML experiments | 37.5% (Claude 3 Opus, 2024) | — |
| RE-Bench [80] | ML research and development | Agents 4× the expert score with a 2 h budget (2024-11) | Experts 2× the agent score with 32 h |
| PaperBench [10] | Replicate ICML papers | 21.0% (Claude 3.5 Sonnet, 2025-04) | ML PhDs 41.4% (subset) |
| EXP-Bench [11] | Full AI experiments | 0.5% (2025-05) | — |
| SciReplicate-Bench [81] | Implement algorithms from papers | 39% (2025) | — |
| ReplicationBench [82] | Astrophysics paper replication | ~20% (Claude Sonnet 4.5, 2025-10) | — |
| AstaBench [28] | Research assistance | 53.0% (Asta v0, 2025-08) [83] | — |
| CURIE [84] | Long-context science | 32% (2025-03) | — |
| FIRE-Bench [85] | Rediscover ML findings | <50 F1 (2026) | — |
| Collider-Bench [86] | Reproduce LHC analyses | None reliably beats the baseline (2026-05) | Physicist-in-the-loop |
| AFMBench [56] | Real AFM hardware | 33.3% analysis (GPT-4o, 2025) | — |
| APEXA-Bench [22] | Synchrotron data reduction (58 tasks) | Not yet scored (2026-09) | — |
Reading the benchmarks
- Closed science exams no longer separate models. Experts score about 65–81% on GPQA Diamond [8], against a best reported 95.8% [9]. HLE's answer key also has problems: by FutureHouse's estimate about 29% of its text-only chemistry and biology answers conflict with the literature, and HLE's own expert re-review found about 18% of a subset problematic. New HLE scores also carry contamination flags (the questions may be in training data) [63, 87].
- Well-specified computing tasks are nearly solved (CORE-Bench Hard, MLE-bench) [74, 78]; humans still lead RE-Bench at long time budgets [80].
- Open-ended research is weak: 25% on FrontierScience-Research, about 20% on ReplicationBench, 0.5% on EXP-Bench [11, 64, 82].
- Lab-facing skills lag: experts lead LAB-Bench, BixBench tops out at 17%, and no model exceeds 70% on hazard identification [67, 68, 75].
- X-ray and neutron relevance is thin. On XRD, models find the highest peak (0.74 accuracy) but rank intensities at only 0.28; shown crystal structures, they assign space groups at only 0.45 [49]. AFMBench is a rare real-instrument benchmark [56]. In APEXA, a frontier model fabricated a calibration report for commands that never ran [22]. No public benchmark covers closed-loop x-ray or neutron beamtime (diffraction, spectroscopy, imaging); the SSRL work is a demonstration, not a benchmark [12].
- Read scores with care. Many rest on LLM judges or private test sets, and saturated (near-ceiling) benchmarks still have construct-validity problems: they may not measure the skill they name [10, 58, 88].
Q3 · Infrastructure: models, frameworks, protocols and the path to HPC
Answer. Science agents are built from interchangeable parts: a vendor or center-hosted model, a general agent framework, the Model Context Protocol (MCP) for tools, and Globus Compute or Parsl for HPC. A facility's main job is therefore to expose the services it already runs (Bluesky/EPICS, Tiled, IRI APIs) as guarded, logged tools [17, 19, 22, 89].
Models
Frontier models such as GPT-5 [90], called through vendor APIs, are still the default engine; Co-Scientist, for example, is built on Gemini, and Biomni's setup expects Claude [4, 91]. But frameworks can usually swap them: ChemGraph accepts OpenAI, Anthropic, Google, Argonne's Argo gateway or a local Ollama server [89].
Open-weight models, whose weights can be downloaded and run locally, make self-hosting practical. gpt-oss-120b is Apache-2.0 licensed and tuned for tool use [92]. OLCF's inference service offers it [18]. Science-trained models also compete: ether0, a 24B chemistry reasoning model trained by reinforcement learning on 640,730 problems, beat frontier models and experts on molecular design [93].
Where the model runs is often a data-handling decision. ALCF built FIRST for private, secure inference; it lets researchers generate "billions of tokens daily on-premises" without commercial cloud [17]. DOE's National Nuclear Security Administration (NNSA) runs frontier AI models on the classified network of its Venado supercomputer at Los Alamos [94].
Agent frameworks and science toolkits
Frameworks typically build on ReAct, the loop in which the model reasons, calls a tool (acts), reads the result (observes) and repeats [95]. General orchestrators include LangGraph [96]; AutoGen is now in maintenance mode, with new users pointed to Microsoft Agent Framework [97]. Vendor SDKs package the same loop; the Claude Agent SDK, for example, adds hooks, subagents and permission controls [98].
A science toolkit is usually a thin layer over one of these orchestrators plus a registry of domain tools. ChemGraph combines LangGraph with ASE, RDKit and MCP, and runs jobs through Parsl or Globus Compute on Aurora and Polaris [89]. Academy runs stateful, asynchronous agents across HPC, experimental facilities and data repositories [99]. The value is in the tools, not the loop.
Tool and agent protocols
The Model Context Protocol (MCP) is an open standard that lets an agent call tools and read data through one common interface. A2A (Agent2Agent) links agents to each other; its maintainers call MCP the "vertical" agent-to-tool layer and A2A the "horizontal" agent-to-agent layer [100]. Both now sit in the Linux Foundation's Agentic AI Foundation: MCP since 9 Dec 2025, A2A (v1.0) since Aug 2026 [16, 100].
MCP is still changing. Its 2026-07-28 revision made the protocol stateless, moved long-running "tasks" into an official extension, and deprecated Roots, Sampling and Logging [101]. A lighter option is a "skill": a folder with a SKILL.md instruction file that the agent loads only when needed [102]. ALS staff write skills for the Lightfall beamline-control platform, and at the SNS one skill document lets a Slack bot run EQSANS reductions [103, 104].
The tool layer is also an attack surface. In tool poisoning, hidden instructions in a tool's description redirect the agent. MCPTox tested 45 live MCP servers: o1-mini's attack success rate was 72.8%, and no model refused more than 3% of attacks [105]. On HPC systems, a "hijacked authorized agent" acts with a user's valid credentials [106].
Connecting to workflows, HPC and instruments
One request passes through six layers, mostly existing DOE software:
- Model. A vendor API or a center-hosted open model. ALCF (FIRST, with Globus Auth) and OLCF (S3M tokens) serve models with vLLM, an open-source serving engine, behind OpenAI-compatible endpoints. These accept OpenAI-style requests, so code switches by changing an address and key [17, 18].
- Agent framework. An agent built with LangGraph, a vendor SDK or Academy reasons and picks the next tool [95, 99].
- MCP server or skill. The call reaches a thin wrapper around an existing service, such as Globus Labs' MCP servers for Globus Transfer, Compute and Search and for facility status APIs [19].
- Execution fabric. Globus Compute (formerly funcX) or Parsl runs the work on HPC [107, 108]. On Aurora, a gpt-oss-120b agent swarm had a planner split work among executors that shared one Parsl-backed MCP server [109]. The Integrated Research Infrastructure (IRI) Facility API standardizes status, account, compute/jobs, filesystem, storage and task endpoints, with instances at NERSC, ALCF and ESnet [110]; SLAC hosts a fork [111].
- Instrument and data. Bluesky and its Ophyd device layer, EPICS, and the Tiled data service are the plug points [112, 113, 114]. At SSRL, an agent with MCP tools wrote SPEC commands, relayed unmodified by a human, to align a single crystal [12]. At NSLS-II, VISION ran a voice-controlled scattering experiment by sending generated code to Bluesky [115]. ORNL agents steered an additive-manufacturing workflow linking OLCF with a simulated version of its Manufacturing Demonstration Facility [116].
- Provenance. Prompts, responses and decisions are logged with the workflow record, as PROV-AGENT does by extending W3C PROV [117].
Sandboxing and permissions. Biomni runs LLM-generated code with full system privileges by default [91]. NERSC recommends sandboxes and workspace-write mode (the agent may change files only inside its working directory) for coding agents, and holds users responsible for agent actions [118]. Facility projects put the limits in code, not prompts. APEXA's deterministic guard allowed 0/200 adversarial motor violations, against 15/200 for a safety prompt, and caught a fabricated calibration report [22]. NeuDiff, at the SNS TOPAZ instrument, uses allowlisted tools and fail-closed gates [15]. EnvTrace checks agent-written control code against a beamline digital twin before it runs [119].
Stack at a glance
| Layer | Common choices | What facilities should watch |
|---|---|---|
| Models | GPT-5 and other vendor APIs; self-hosted gpt-oss, DeepSeek-R1; science-tuned ether0 [92, 93, 120] | Center-hosted OpenAI-compatible endpoints [17, 18]; data-sensitivity rules [94] |
| Frameworks | LangGraph, vendor SDKs, Academy, ChemGraph [96, 98, 99] | AutoGen churn [97]; code-execution privileges [91] |
| Protocols | MCP, A2A, skills [16, 100, 102] | Stateless MCP and tasks extension [101]; tool poisoning [105] |
| Workflow and HPC | Globus Compute, Parsl, IRI Facility API [107, 108, 110] | MCP-wrapped facility APIs [19]; agents acting under user credentials [106] |
| Instruments and data | Bluesky/Ophyd, EPICS, SPEC, Tiled [12, 112, 113, 114] | Guards in code, not prompts [15, 22]; provenance [117] |
Implications for facility deployments
- Wrap, do not rebuild. Put MCP servers or skill/CLI layers over existing services such as the IRI API, Bluesky and Tiled, as Globus Labs, ALS and SNS teams did for their own services [19, 103, 104].
- Plan for model portability. Write agents against OpenAI-compatible interfaces; open-weight models behind such endpoints already run at ALCF and OLCF [17, 18].
- Enforce safety in the tool layer, not the prompt. Use allowlists, execution-integrity guards and fail-closed gates [15, 22].
- Treat tool metadata, logs and shared files as untrusted input. Agents acting under user credentials are a new HPC threat class [105, 106].
- Record agent provenance from day one [117].
Q4 · Facilities: agents at DOE x-ray and neutron user facilities and national labs
Answer. LLM agents at DOE facilities have moved from documentation assistants and first instrument-control prototypes (2023–24) to supervised control of real beamlines, accelerators and microscopes, peer-reviewed in 2025–26, but steering is still mostly single-campaign demonstrations; deployed systems are mainly knowledge assistants, data-reduction agents and one accelerator control-room framework (Osprey, at the ALS) [12, 14, 22, 115, 121].
Beamline, instrument and accelerator steering
This subsection covers agents that steer beamlines, instruments or accelerators, grouped by lab. Most agents that reduce or analyze data, or answer staff and user questions, are listed under Facility data analysis and knowledge assistants; that is where the SNS agents, and APEXA, APS-RAG and PEAR at the APS, appear.
SLAC
- SSRL. An LLM agent aligned a single crystal on the BL17-2 six-circle diffractometer. MCP tools wrap SPEC commands; the agent found reference reflections and the orientation matrix. It was developed on a virtual diffractometer, where Claude Sonnet 4 and Gemini 2.5 Flash were benchmarked; Claude Opus 4 ran on the real beamline. For safety, a human relayed each command without modification [12].
- LCLS. This survey found no public LLM-agent paper on LCLS experiment control. The LCLS items we found are physics-guided Bayesian optimization (not an LLM) that aligned a 12-parameter split-and-delay optic in wave-optics simulations checked against measured data [122], and a computational cognitive model of how operators, analysts and managers interact during LCLS instrument operations, which stresses that LCLS beam time is "extremely valuable and limited" [123].
- Accelerator operations. ChatEED answers operator questions from electronic logbooks and wikis [124].
BNL (NSLS-II, CFN)
- VISION turns voice or text into Bluesky code at the 11-BM CMS beamline. It ran the first voice-controlled x-ray scattering experiment. Beamline models run locally; cloud GPT-4o serves other functions [115].
- EnvTrace scores control code from more than 30 LLMs against a beamline digital twin, so code can be checked before a live experiment [119].
Argonne (APS, CNM)
- CALMS combines retrieval over facility documentation with instrument tool calls [125].
- Human-in-the-loop agents operate an x-ray nanoprobe beamline and a robotic materials station, and improve with feedback [126].
- Vision-language agents (EAA) automate zone-plate focusing and feature search at an APS imaging beamline [127].
- On an MCP-compatible atomic layer deposition reactor, only recent models (o1, o3, GPT-5, Claude Opus 4) performed well on process-discovery tasks [128].
LBNL (ALS, Molecular Foundry, A-Lab)
- An agent ran multistage machine-physics experiments on the ALS accelerator. It cut preparation time about 100× within operator safety constraints [13].
- Osprey, the framework behind it, requires human review of a complete plan before any hardware action. It runs in production across hundreds of thousands of ALS control channels [14].
- Semantic channel finding maps natural-language intent to control signals; proof-of-concepts ran at four facilities [129].
- Lightfall, a beamline control platform with a built-in agent, is in testing at COSMIC-Scattering (preprint) [103].
- TEM Agent lets a commercial LLM control microscope subsystems, data management and HPC through MCP [130].
- In the air-free A-Lab, LLM agents proposed 263 of 352 syntheses; researchers designed the rest (preprint) [131].
ORNL (CNMS, SNS)
- At CNMS, GPT-4 wrote scanning-probe microscope code but struggled with in-depth experiment design [132].
- INTERSECT provides federated microservices for multi-facility autonomous experiments [133].
Other accelerators and international facilities
- Retrieval-augmented e-log search was demonstrated at Fermilab, JLab, LBNL and SLAC [134].
- At DESY, an LLM tuned a simulated accelerator subsystem (an ARES linac beam-tuning task) from an operator prompt, as a proof of principle [135].
- CSNS (China) runs a Rietveld refinement agent for external users [136]. We found no comparable deployed system from ESRF, Diamond, European XFEL or PSI. European XFEL has reported two prototypes of an agent that retrieves facility knowledge and writes analysis code in its HPC environment, evaluated with facility experts [137].
Pre-LLM autonomous-experiment precursors
These closed loops use Bayesian or machine-learning methods, not language models. Most predate the agents above.
- Gaussian-process autonomous x-ray scattering, the gpCAM lineage [138].
- CAMEO: closed-loop Bayesian active learning at SSRL found a new Ge-Sb-Te phase-change material [139].
- ANDiE: autonomous neutron diffraction determined Néel temperatures 5-fold more efficiently [140].
- Argonne FAST: autonomous scanning microscopy needed under 25% of the sample [141].
- Bluesky-hosted agents coordinated multi-beamline measurements at NSLS-II [142, 143].
- LCLS physics-guided Bayesian optimization of a 12-parameter split-and-delay optic, tested in simulation [122]; ORNL edge-to-exascale steering at SNS, still a proof of concept [144].
Facility data analysis and knowledge assistants
- NeuDiff Agent (SNS TOPAZ) goes from raw data to a validated CIF (crystallographic information file). Its safeguards are allowlisted tools, fail-closed gates and full provenance. Processing time fell from 435 min (manual) to 86.5–94.4 min, with no level A or B checkCIF alerts (the most serious validation warnings) [15].
- ORNL neutron small-angle and diffraction tools. SasAgent drives SasView [145]. EQSANS-CLI lets an external agent run reductions from Slack (preprint) [104]. VULCAN's reduction pipeline was written with an AI coding agent [146].
- APEXA (APS) is a deployed calibration and integration agent with 61 tools (preprint). A frontier model fabricated a calibration report for commands that never ran. Against a simulated IOC, motor-control violations were 0/200 with its deterministic guard and 15/200 with a safety prompt alone [22].
- Other analysis agents. PEAR uses multiple agents for ptychography [147]. Rongzai (CSNS) runs GSAS-II from a natural-language task to a report [136].
- Knowledge and user-program assistants. ESAC is a chatbot for visiting EQ-SANS users [148]. APS-RAG, deployed for APS staff, reached 70.3% strict recall versus 63.8% for a BM25 baseline [121]. For SNS proposals, LLM rankings correlated with human rankings at ρ≈0.2–0.8, at over 100× lower cost [149].
DOE programs and policy
- Genesis Mission. The executive order of Nov 24, 2025 creates a platform that explicitly includes "AI agents" and autonomous experimentation. It sets milestones at 60, 90, 120, 240 and 270 days; the 240-day step reviews robotic laboratories. It requires classification, cybersecurity and export-control compliance [23].
- Genesis follow-up. DOE named 26 challenges, including "Enhancing Particle Accelerators for Discovery" and "Achieving AI-Driven Autonomous Laboratories" [150]. The autonomous-labs challenge names DOE user facilities and national laboratories as its "nucleus" [151]. DOE signed agreements with 24 organizations, among them Anthropic, AWS, Google, NVIDIA and OpenAI [152]. In July 2026 the White House announced more than $5B in federal commitments from more than 15 agencies, and DOE selected 278 projects [153]. DOE also announced more than $800M from partners [154]. An Office of Science advisory subcommittee's high-performance-magnets milestones include connecting autonomous labs to user facilities [155].
- American Science Cloud (AmSC). Public Law 119-21, Sec. 50404 ("Transformational artificial intelligence models") defines the American Science Cloud and appropriates $150M, through Sept 30, 2026, for curating DOE data and seeding science AI models to be shared through it [156]. An ASCR lab call builds on IRI and includes API access to user facilities and instruments [157]. AmSC aims to unify access to the 17 DOE labs and 28 user facilities that ESnet already links, adding a common identity layer and shared service interfaces [158]. The multi-lab SYNAPS-I project (LBNL-led, with Argonne, BNL, ORNL and SLAC) targets real-time analysis at light and neutron sources [158, 159].
- Funding calls. A DOE package of more than $320M includes 14 robotics and autonomous-experiment projects [160]. A robotics and automation testbed call is about $30M [161].
- FASST. DOE released a roadmap in 2024 for FASST (Frontiers in Artificial Intelligence for Science, Security and Technology) [162]. Our inference, not a DOE statement: Genesis now plays FASST's umbrella role.
- 1,000 Scientist AI Jam (Feb 28, 2025). More than 1,400 scientists at nine labs took part; SLAC was not one of them [163].
- Lab–company partnerships. OpenAI models run on LANL's Venado supercomputer [94, 164]. Anthropic's DOE partnership plans MCP servers that connect Claude to scientific instruments [165]. DOE announced NVIDIA-based Solstice and Equinox at Argonne [166] and AMD-based Lux and Discovery at ORNL [167].
- Integrated Research Infrastructure (IRI). A 2023 blueprint lays out how to link experimental facilities, computing and data [168].
Status, integration points and facility constraints
Facility agent systems
| System | Facility/lab | Technique | What the agent controls or analyzes | Status | Evidence |
|---|---|---|---|---|---|
| AI X-ray scientist | SLAC SSRL | Crystal diffraction | Diffractometer motors, detector | Demonstrated | Peer-reviewed [12] |
| VISION | BNL NSLS-II | X-ray scattering | Bluesky motors, detector | Demonstrated | Peer-reviewed [115] |
| Learn-on-the-job agents | Argonne APS/CNM | Nanoprobe, robotics | Multi-task workflows | Demonstrated | Peer-reviewed [126] |
| EAA | Argonne APS | X-ray imaging | Focusing, feature search | Demonstrated | Peer-reviewed [127] |
| ALS machine-physics agent | LBNL ALS | Accelerator | Plans multistage machine-physics experiments; archive retrieval, control-channel resolution | Demonstrated | Peer-reviewed [13] |
| Osprey framework | LBNL ALS | Accelerator control room | Plan-first orchestration of control channels, with human review before hardware actions | Deployed | Peer-reviewed [14] |
| Lightfall | LBNL ALS | Coherent scattering | Devices, GP scans | In testing | Preprint [103] |
| TEM Agent | LBNL Molecular Foundry | TEM | Microscope, HPC | Demonstrated | Peer-reviewed [130] |
| NeuDiff Agent | ORNL SNS | Neutron crystallography | Reduction to CIF | Demonstrated | Peer-reviewed [15] |
| SasAgent / EQSANS-CLI | ORNL; SNS EQ-SANS | SANS | Fitting, reduction | Demonstrated | Peer-reviewed; preprint [104, 145] |
| APEXA | Argonne APS | Diffraction | Calibration, integration | Deployed | Preprint [22] |
| APS-RAG | Argonne APS | Operations | Logbooks, control data | Deployed | Preprint [121] |
| ChatEED | SLAC | Accelerator operations | E-log retrieval | Demonstrated | Peer-reviewed workshop [124] |
| Rongzai | CSNS | Neutron powder diffraction | Rietveld refinement | Deployed | Preprint [136] |
| Genesis platform | DOE | All | Autonomous experimentation | In development | Executive order; White House release [23, 153] |
Integration points
- Instrument control. Agents reach hardware through SPEC [12], Bluesky [115, 169] and control-system channels [14].
- Data and workflows. A Tiled, Prefect and gpCAM workflow ran at both ALS and PETRA III with minimal facility-specific changes [170].
- Tool layer. MCP is a common way to expose instruments and data to agents [12, 121, 127, 130].
- Computing. Agents can be served on leadership-class HPC [171]; a planned AmSC API, built on the IRI framework, is to give code-level access to user facilities and instruments [157].
Facility-specific constraints
- Beamtime. LCLS beam time is "extremely valuable and limited" [123]. The SSRL team developed the workflow in simulation and deployed it on the beamline with only interface-level changes [12].
- Safety. Current safeguards sit outside the model: human relay [12], plan review [14], allowlisted tools [15], deterministic guards [22], or a digital-twin check before execution [119].
- User turnover. Visiting users need guidance [148], and agents must work zero-shot, without long training periods [172].
- Security and data policy. Genesis mandates classification, cybersecurity and export-control compliance [23]. VISION runs its beamline-control models on a local GPU server and sends only chatbot and refinement tasks to cloud GPT-4o [115].
- Verification. Checking results, not generating ideas, is now the bottleneck [173].
Q5 · Failure modes and risks
Answer. Documented failures include irreproducible runs, fabricated references and results, benchmarks that overstate real performance, unsafe or hijacked tool actions, and unclear credit; the mitigations in use are full logging, independent checks, sandboxed tools and disclosure rules [22, 174, 175, 176].
Reproducibility
- Agent evaluations often ignore cost and lack holdout sets, and non-standard evaluation practices lead to "a pervasive lack of reproducibility" [174].
- Reproducing papers from their own code and data was hard for agents: when CORE-Bench was released, the best agent scored 21% on its hardest tier (270 tasks, 90 papers) [73]. Newer agents now score far higher (Q2).
- "Deterministic" settings still gave accuracy swings of up to 15% across runs [177]. GPT-4 accuracy at identifying primes fell from 84% to 51% between its March and June 2023 versions [178].
- One standardized re-evaluation took 21,730 rollouts and about $40,000. Its logs showed agents looking up benchmark answers on HuggingFace [179].
- A-Lab: an independent critique examined all 43 synthetic products discussed in the A-Lab paper (which reported 41 successful targets) and concluded that no new materials had been discovered. About two-thirds were likely known disordered phases, and automated Rietveld analysis of XRD "is not yet reliable" [42].
- In an independent test of Sakana's AI Scientist, 42% of experiments failed from coding errors [180].
For facilities: Pin model versions. Log prompts and tool calls with the data provenance; the NeuDiff Agent at SNS keeps full provenance from raw data to a validated CIF [15]. Have an expert check any structure refinement made by an agent.
Fabricated citations or results
- 55% of GPT-3.5 citations and 18% of GPT-4 citations were fabricated [181].
- About 1 in 20 NeurIPS and USENIX Security 2025 papers has at least 2 likely hallucinated references. An automated audit costs about $0.04 per paper [20].
- ICLR 2026 treated hallucinated references as an ethics violation and desk-rejected papers containing them; area chairs found the cases, with LLM-detection tools used for triage [182]. At Agents4Science, only about 44% of submissions had no flagged reference [40].
- AI Scientist papers contained hallucinated numbers [180]. Data leakage, metric misuse and post-hoc selection are easier to spot in trace logs and code than in the final paper [183].
- Research agents reward-hacked without being asked in 30.5% of open-ended research-pipeline tasks [21].
- Agents given different personas reached opposite conclusions from the same data, yet 86% of their reports passed AI review [184].
For facilities: Check references automatically. Require the data, code and agent traces behind any agent-assisted result. At the APS, a frontier model fabricated a calibration report for commands that never ran, so an agent's report must be checked against what actually executed [22].
Evaluation gaps
- Contamination: accuracy dropped by up to 8% on fresh GSM1k questions compared with GSM8k [185].
- Flawed task or grader design can mis-estimate agent performance by up to 100% (relative). A checklist cut overestimation on CVE-Bench by 33% [186].
- On 102 ScienceAgentBench tasks drawn from 44 real papers, the best of the main agents solved 32.4%; OpenAI's o1-preview, evaluated in addition with self-debugging, reached 42.2% [72].
- Multi-agent systems show 14 failure modes, grouped into design, inter-agent misalignment and verification [187].
- LLM research ideas were rated more novel than expert ideas (p<0.05) [43]. After 43 experts carried the ideas out, the scores of the LLM ideas fell significantly more [44].
- LLM judges show position, length and self-preference biases [188]. At Agents4Science, mean reviewer scores ranged from 2.30 to 4.23 depending on the model [40]. An LLM review panel missed 6.5% of confirmed reward hacks [21].
- Beyond scores: AI tools can foster "illusions of understanding", where scientists believe they understand more than they do, and can narrow the range of methods a field uses [189].
For facilities: Test agents on held-out, in-house beamtime data graded by experts. Do not rely on a single LLM judge.
Safety and dual use
- No model tested scored above 70% on lab hazard identification [68].
- Instructions hidden in retrieved data can hijack LLM applications [190]. The safest agent tested still failed 23.9% of high-stakes tool cases [175]. The AI Scientist tried to edit its own code to extend its time limit [191].
- In 2025, 18 arXiv manuscripts hid prompts such as "GIVE A POSITIVE REVIEW ONLY" [192]. Agents4Science caught 2 papers trying to manipulate its reviewers [40].
- A drug-design model generated 40,000 candidate toxic molecules, including VX, in under 6 hours [193]. ChemCrow added checks for controlled chemicals and explosives [35].
- A safety assessment of one self-driving lab found 16 hazards. Robot–human collision and chemical exposure were the most critical, and collaborative robots may exceed force limits [194]. A DOE workshop hosted by ORNL called for safety protocols and human oversight [195].
For facilities: Sandbox agents and allow only listed tools; NeuDiff at SNS uses allowlisted tools and fail-closed gates [15]. Keep interlocks independent of the agent, and treat user files as untrusted. At SSRL a human relayed each agent command unmodified [12]; at the ALS, Osprey requires human review of a complete plan before any hardware action [14]. At the APS, a deterministic guard allowed 0 of 200 motor-control violations against a simulated IOC, versus 15 of 200 with a safety prompt alone [22].
Credit and authorship
- Nature: an LLM cannot be an author, and its use must be documented [196]. ICMJE (International Committee of Medical Journal Editors): AI cannot be an author, AI use must be disclosed, and humans remain responsible [176].
- Conferences: NeurIPS 2025 allows only human authors [197]. ICLR 2026 requires disclosure of all LLM use and counts hidden prompts as collusion [198].
- ICML 2026 hid instructions in submission PDFs to catch reviewers who used LLMs against its policy. It removed 795 reviews (~1%) and desk-rejected 497 papers written by the offending reviewers [199].
- 6.5–16.9% of peer-review text at four AI conferences may have been substantially LLM-modified [200]. A detector vendor estimates that 21% of ICLR 2026 reviews were fully AI-generated [201].
- Agents4Science 2025 required an AI first author [202]. It accepted 48 of 315 submissions [40].
- Funders: NIH bans generative AI in peer review [203]. It does not count applications substantially developed by AI as original, and caps each PI at 6 applications a year [204]. NSF bars reviewers from uploading proposals to non-approved AI tools [205]. This search found no DOE-wide equivalent.
For facilities: Require AI-use disclosure in beamtime proposals, and do not allow external AI tools in proposal review panels.
Risk register
| Risk | Example | Signal from evidence | Mitigation seen in the literature |
|---|---|---|---|
| Irreproducible results | Model version drift | Up to 15% swing between runs [177]; 84% to 51% across versions [178] | Pin versions; shared logs of every run [179] |
| Overclaimed discovery | A-Lab materials | About two-thirds likely known phases [42] | Expert crystallography review [42] |
| Fake references | NeurIPS 2025 papers | About 1 in 20 papers [20] | Automated reference checks [20]; desk rejection [182] |
| Fabricated execution report | Beamline calibration agent | Report for commands that never ran [22] | Execution-integrity enforcement [22] |
| Reward hacking or selective reporting | Research-pipeline tasks | 30.5% unprompted [21] | Review traces and code; multiverse checks [183, 184] |
| Benchmark–lab mismatch | Science data tasks | Best agents solve 32.4–42.2% [72] | Benchmark checklist; in-house tasks [186] |
| Biased LLM judges | AI reviewers | Mean scores 2.30–4.23 by model [40] | Final human review [40] |
| Prompt injection | Hidden review prompts | 18 manuscripts [192] | Treat inputs as untrusted [190]; misconduct rules [198] |
| Unsafe tool actions | High-stakes tools; motor moves | 23.9% failures [175]; 15 of 200 violations with a safety prompt alone [22] | Sandboxing; deterministic guards; human plan review [14, 22] |
| Physical lab hazards | Robot–human collision | 16 hazards in one lab [194]; no model above 70% on hazard identification [68] | ISO-based risk assessment [194]; human oversight [195] |
| Dual use | Toxin design | 40,000 molecules in under 6 h [193] | Controlled-substance checks [35] |
| Undisclosed AI in review | LLM-written reviews | 6.5–16.9% of text [200] | Disclosure; canary prompts; bans [199, 203] |
Q6 · Discussion: open questions for the session
Answer. The open questions for facilities are no longer whether agents can touch instruments (they already align crystals, run accelerator experiments and reduce neutron data) but where humans must stay in the loop, how agent results are verified and credited, and what shared infrastructure DOE facilities should build [12, 13, 15, 24].
Each prompt below states the evidence it rests on and links to the section where that evidence is discussed. The prompts are ordered from beamline-level to program-level questions.
Suggested running order for a 60–90 minute session. Three blocks of about 20 minutes each: at the beamline (prompts 1, 2 and 4), shared infrastructure and data (prompts 3, 5 and 6), and programs and people (prompts 7, 8 and 9). If time is short, use prompts 1, 2, 4, 6 and 8. Each block works best if it ends with one concrete output, for example a list of beamline actions that always need human approval, one candidate joint benchmark, or a first joint deliverable.
- Where should the human stay in the loop at a beamline? In its first real-beamline trial, the SSRL agent had a person relay every command [12], and the ALS accelerator framework requires human review of a complete plan before any hardware action [14] (Q4 steering). Outside facilities, targeted human interventions beat both full autonomy and step-by-step oversight [45], and accepted Agents4Science papers had more human guidance [40] (Q1 trends). For discussion: which actions (moving motors, changing sample environment, discarding data, ending a run) need explicit approval, and could facilities agree on a shared autonomy scale for beamtime, like the L1–L4 scale used here?
- Should safety be enforced in the tool layer rather than in the prompt? At APS, a deterministic execution guard allowed 0 of 200 adversarial motor-control violations, against 15 of 200 with a safety prompt alone [22]; NeuDiff at SNS uses allowlisted tools and fail-closed gates [15] (Q3 workflow and HPC). Tool poisoning succeeds against many models on real MCP servers [105] (Q3 protocols), and no model tested scored above 70% on lab hazard identification [68] (Q5 safety). For discussion: should DOE facilities define a common "agent-safe" instrument interface, with allowlisted tools, simulation-first execution and interlocks that do not depend on the agent?
- What evaluation would convince facility scientists? Closed-form science benchmarks are saturated [8, 9], while models rank XRD peak intensities poorly [49] and no public benchmark covers closed-loop x-ray or neutron beamtime (Q2 benchmarks). A beamline digital twin already scores LLM-written control code [119] (Q4 steering), and benchmark design flaws can badly mis-estimate agent performance [186] (Q5 evaluation gaps). For discussion: could SLAC, ORNL, BNL and university partners build a shared, held-out benchmark from archived beamtime (diffraction, SAXS/SANS, spectroscopy, imaging) with expert grading, plus digital twins for control tasks?
- How are agent-produced analyses verified before they reach a paper? An outside analysis disputed the A-Lab's claims of new materials and called automated Rietveld analysis of powder XRD not yet reliable [42] (Q2 established results). A frontier model fabricated a calibration report for commands that never ran [22], and research agents reward-hacked without being asked in open-ended tasks [21] (Q5 fabrication). A 2026 survey calls verification the field's central bottleneck [24]. For discussion: what provenance should be required for agent-assisted results from facility data (prompts, tool calls, code, model version, raw-data hashes), and who checks it: the user, the beamline scientist, or the journal? Provenance tooling for agents already exists [117].
- Which models may touch user data, and where do they run? Center-hosted, OpenAI-compatible inference on open-weight models already runs at ALCF and OLCF [17, 18] (Q3 models). VISION keeps its beamline models local [115], NERSC tells users not to place credentials in external AI services [118], and the Genesis Mission order requires security and export-control compliance [23] (Q3 workflow and HPC; Q4 constraints). For discussion: what is the policy for proprietary or sensitive user data: commercial APIs, center-hosted open-weight models, or both? Who pays for inference during beamtime?
- What should facilities standardize on: MCP servers, skills, or the APIs they already have? MCP moved to a Linux Foundation body in December 2025 [16], and its July 2026 revision changed core behavior [101] (Q3 protocols). MCP servers already wrap Globus and facility services [19], the IRI Facility API standardizes job and data endpoints [110], and Bluesky is shared across facilities [169] (Q3 workflow and HPC; Q4 status). Anthropic's DOE partnership describes MCP servers for instruments [165]. For discussion: should facilities jointly publish and maintain agent interfaces for Bluesky, EPICS, Tiled and IRI, rather than each beamline building its own? Who owns them as the protocols change?
- How should agent use be disclosed and credited in proposals, papers and review? Journals and ICMJE bar AI authorship and require disclosure [176, 196]. ICML 2026 found 795 reviews that broke its no-LLM policy and desk-rejected 497 papers from the reviewers involved [199], and NIH does not count applications substantially developed by AI as original [204] (Q5 credit). At SNS, LLM proposal rankings correlated with human rankings at ρ≈0.2–0.8, at over 100× lower cost [149] (Q4 analysis). For discussion: should beamtime proposals and facility-data papers declare agent use? Should LLMs help triage proposals, and if so, with what disclosure to proposers?
- Where will Genesis Mission funding meet facility operations? DOE's autonomous-laboratories challenge names user facilities as its nucleus [151]. The July 2026 announcement lists more than $5 billion in federal commitments, and DOE selected 278 projects [153]; the American Science Cloud aims to unify access to the 17 DOE labs and 28 user facilities that ESnet already links [158] (Q4 programs). Most facility agent systems are still single-campaign demonstrations (Q4 status). For discussion: what is a realistic first joint deliverable, for example one agent-steered experiment that spans a light or neutron source and an HPC center, and how do teams avoid one-off demonstrations that are never maintained?
- How do users and staff keep their expertise as agents take over routine steps? A facility chatbot already serves visiting EQ-SANS users [148] (Q4 analysis), and AI can create "illusions of understanding" and narrow the range of methods scientists use [189] (Q5 evaluation gaps). LLM-generated research ideas lost their apparent edge once experts carried them out [44] (Q1 trends). For discussion: which skills (alignment, reduction, refinement) must users still show before an agent acts for them, and how should university partners train students for agent-assisted beamtime?
Appendix
System catalog
All 83 systems placed in the Q1 taxonomy, sorted by stage and then year. Autonomy uses the L1–L4 scale in the glossary. Evidence type describes the cited source: peer-reviewed, peer-reviewed with independent validation, preprint, workshop, product or vendor report, or lab or project page. The one-line description under each name summarizes what the cited source reports. Facility and national-lab systems (Q4) are included alongside general-science systems.
| System | Organization | Year | Stage | Autonomy | Evidence type | Reference |
|---|---|---|---|---|---|---|
| CALMSRetrieval- and tool-augmented LLM answering facility-documentation questions and conversationally operating instruments. | Argonne (APS, CNM, ALCF) | 2023 | Literature & ideationspans: Experiment design & self-driving labs | L2 Tool-using agent | Peer-reviewed | [125] |
| ESACRetrieval-augmented chatbot serving as an interactive instrument reference for visiting EQ-SANS users. | ORNL (SNS EQ-SANS) | 2024 | Literature & ideationspans: Experiment design & self-driving labs | L1 Assistant | Peer-reviewed | [148] |
| GAIAReAct-style LLM assistant combining logbook and document retrieval with control-script generation for accelerator operations. | DESY | 2024 | Literature & ideationspans: Experiment design & self-driving labs | L2 Tool-using agent | Preprint | [206] |
| Gemini Deep ResearchConsumer deep-research agent that browses the web on the user's behalf and returns a cited, exportable report; launched December 2024. | 2024 | Literature & ideation | L2 Tool-using agent | Product/vendor report | [207] | |
| LLM ideation agent vs. 100+ NLP researchersBlinded human study: agent ideas judged more novel (p<0.05) than experts' ideas but slightly weaker on feasibility. | Stanford University | 2024 | Literature & ideationspans: Hypothesis generation | L2 Tool-using agent | Peer-reviewed | [43] |
| OpenScholarRetrieval-augmented LM over 45M open-access papers; citation accuracy on par with experts; experts preferred OpenScholar-8B answers over expert-written ones 51% of the time. | University of Washington / Ai2 | 2024 | Literature & ideation | L2 Tool-using agent | Peer-reviewed | [6] |
| PaperQA2Literature-search agent that matched or exceeded subject experts on three literature tasks; found 2.34 contradictions per biology paper, 70% expert-validated. | FutureHouse | 2024 | Literature & ideation | L2 Tool-using agent | Preprint | [27] |
| ResearchAgentProposes problems, methods and experiment designs from a core paper plus academic graph, iteratively refined by LLM reviewing agents. | KAIST / Microsoft Research | 2024 | Literature & ideationspans: Hypothesis generation | L2 Tool-using agent | Peer-reviewed | [208] |
| STORMResearches a topic via multi-perspective question asking, then writes cited Wikipedia-like articles; 25% absolute gain in judged organization over a retrieval baseline. | Stanford University | 2024 | Literature & ideationspans: Writing & review | L2 Tool-using agent | Peer-reviewed | [209] |
| Ai2 Asta agents + AstaBenchScience-agent ecosystem with a 2,400+ problem benchmark; evaluation of 57 agents concludes AI remains far from solving science research assistance. | Ai2 | 2025 | Literature & ideationspans: Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [28] |
| Ai2 Scholar QAFree, open-source literature-synthesis app producing organized reports with attribution; outperformed competing systems on a recent scientific QA benchmark. | Ai2 | 2025 | Literature & ideation | L2 Tool-using agent | Peer-reviewed | [210] |
| ChatEEDAgentic retrieval assistant over electronic logbooks and wikis for accelerator operators. | SLAC (accelerator operations) | 2025 | Literature & ideation | L1 Assistant | Peer-reviewed | [124] |
| FutureHouse Platform (Crow, Falcon, Owl, Phoenix)Web platform of literature agents plus a ChemCrow-based chemistry planner; vendor claims better precision than PhD-level researchers on literature search. | FutureHouse | 2025 | Literature & ideationspans: Experiment design & self-driving labs | L2 Tool-using agent | Product/vendor report | [211] |
| APS-RAGDeployed agentic GraphRAG over logbooks, documents, wikis and control data for APS staff; 70.3% vs 63.8% strict recall. | Argonne (APS) | 2026 | Literature & ideation | L2 Tool-using agent | Preprint | [121] |
| FunSearchLLM paired with an automated evaluator in an evolutionary program search; found cap-set constructions beyond the best known and better bin-packing heuristics. | Google DeepMind | 2023 | Hypothesis generationspans: Data analysis & simulation | L3 Closed-loop | Peer-reviewed | [212] |
| SciMONRetrieves literature inspirations and iteratively revises generated research ideas to increase novelty relative to prior work. | UIUC / Ai2 | 2023 | Hypothesis generationspans: Literature & ideation | L2 Tool-using agent | Peer-reviewed | [213] |
| MOOSE-ChemInspiration-retrieval, composition and ranking agents rediscovered many hypotheses from 51 high-impact 2024 chemistry papers using a pre-2024 LLM. | Nanyang Technological University / Shanghai AI Laboratory | 2024 | Hypothesis generationspans: Literature & ideation | L2 Tool-using agent | Peer-reviewed | [214] |
| SciAgentsMulti-agent LLMs reason over ontological knowledge graphs to generate, critique and refine research hypotheses for bio-inspired materials. | MIT | 2024 | Hypothesis generationspans: Literature & ideation | L2 Tool-using agent | Peer-reviewed | [215] |
| Co-ScientistGemini multi-agent generate-critique-refine tournament; AML drug-repurposing candidates validated in vitro; separately matched an external lab's unpublished gene-transfer mechanism. | 2025 | Hypothesis generationspans: Literature & ideation | L2 Tool-using agent | Peer-reviewed + independent validation | [4] | |
| DeepScientistMonth-long hypothesize-verify-analyze loop; ~5,000 ideas, ~1,100 tested, beat human state of the art on three AI tasks (183.7%, 1.9%, 7.9%). | Westlake University | 2025 | Hypothesis generationspans: Data analysis & simulation | L3 Closed-loop | Peer-reviewed | [39] |
| InternAgent (NovelSeek)Closed-loop hypothesis-to-verification framework across 12 tasks; e.g., reaction-yield prediction improved from 27.6% to 35.4% in 12 hours. | Shanghai AI Laboratory | 2025 | Hypothesis generationspans: Data analysis & simulation | L3 Closed-loop | Preprint | [216] |
| POPPERAgents design and run falsification tests with Type-I error control; matched human scientists validating biological hypotheses in one-tenth the time. | Stanford University | 2025 | Hypothesis generationspans: Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [217] |
| RobinLab-in-the-loop multi-agent system; proposed ripasudil for dry macular degeneration and analyzed follow-up RNA-seq; humans executed the wet-lab experiments. | FutureHouse | 2025 | Hypothesis generationspans: Literature & ideation; Experiment design & self-driving labs; Data analysis & simulation | L3 Closed-loop | Peer-reviewed | [5] |
| AletheiaGemini Deep Think math agent that generates, verifies and revises proofs; one paper produced without human intervention; four open Erdős problems solved autonomously. | Google DeepMind | 2026 | Hypothesis generationspans: Writing & review | L4 End-to-end | Preprint | [3] |
| A-LabRobotic solid-state synthesis with computation, literature-trained recipe models and active learning; reported 41 of 58 targets in 17 days; 2026 correction confirms 36. | Lawrence Berkeley National Laboratory | 2023 | Experiment design & self-driving labsspans: Data analysis & simulation | L3 Closed-loop | Peer-reviewed | [50] |
| ChemCrowGPT-4 agent with 18 expert-designed chemistry tools; autonomously planned and executed syntheses of an insect repellent and three organocatalysts. | EPFL / University of Rochester / IBM Research | 2023 | Experiment design & self-driving labsspans: Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [35] |
| CoscientistGPT-4 agent with search, code and lab-automation tools that designed, planned and ran experiments, including optimizing palladium-catalysed cross-couplings. | Carnegie Mellon University | 2023 | Experiment design & self-driving labsspans: Literature & ideation | L3 Closed-loop | Peer-reviewed | [29] |
| BioDiscoveryAgentIterative design of genetic perturbation screens; 21% better hit prediction than Bayesian-optimization baselines across six datasets. | Stanford University / Arc Institute | 2024 | Experiment design & self-driving labsspans: Hypothesis generation | L3 Closed-loop | Peer-reviewed | [218] |
| ChemAgentsOn-board Llama-3.1-70B hierarchical agents (literature, design, computation, robot) ran seven robotic chemistry tasks with minimal human intervention. | University of Science and Technology of China | 2024 | Experiment design & self-driving labsspans: Literature & ideation; Data analysis & simulation | L3 Closed-loop | Peer-reviewed | [219] |
| CRISPR-GPTGene-editing design and analysis agent; guided knockout of four genes (Cas12a) and activation of two genes (dCas9) in human cell lines. | Stanford University / Princeton University | 2024 | Experiment design & self-driving labsspans: Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [52] |
| k-agentsLLM agents encoding lab knowledge ran a superconducting quantum processor for hours, producing entangled states at the level of human scientists. | University of Oxford / University of Toronto | 2024 | Experiment design & self-driving labsspans: Data analysis & simulation | L3 Closed-loop | Peer-reviewed | [220] |
| LLM accelerator tuningLLM tunes an accelerator subsystem from an operator's natural-language prompt, benchmarked against Bayesian optimization and RL-trained optimizers. | DESY | 2024 | Experiment design & self-driving labs | L3 Closed-loop | Peer-reviewed | [135] |
| LLM-assisted scanning probe microscopyGPT-4 with instrument APIs converts experimental workflow ideas into microscope code; limited for in-depth experimental design. | ORNL (CNMS) | 2024 | Experiment design & self-driving labsspans: Data analysis & simulation | L1 Assistant | Peer-reviewed | [132] |
| Mobile-robot synthesis labMobile robots operate shared synthesis, UPLC-MS and benchtop NMR equipment; a heuristic decision-maker picks and re-checks hits in exploratory chemistry. | University of Liverpool | 2024 | Experiment design & self-driving labsspans: Data analysis & simulation | L3 Closed-loop | Peer-reviewed | [221] |
| ORGANALLM-interfaced robotic assistant for solubility, pH, recrystallization and electrochemistry; user study reported 80.3% average time saved. | University of Toronto | 2024 | Experiment design & self-driving labs | L3 Closed-loop | Peer-reviewed | [222] |
| Virtual LabLLM 'PI' agent leads specialist agents to build a nanobody design pipeline; 92 designed, two with improved binding to recent SARS-CoV-2 variants. | Stanford University / Chan Zuckerberg Biohub | 2024 | Experiment design & self-driving labsspans: Hypothesis generation; Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [31] |
| VISIONModular LLM assistant turning voice/text into Bluesky code at NSLS-II 11-BM; first voice-controlled X-ray scattering experiment. | BNL (CFN, NSLS-II) | 2024 | Experiment design & self-driving labsspans: Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [115] |
| AI X-ray scientistLLM agent with MCP tools and SPEC commands aligned a single crystal on SSRL BL17-2; developed in a virtual diffractometer; human relayed commands. | SLAC (SSRL) | 2025 | Experiment design & self-driving labsspans: Data analysis & simulation | L3 Closed-loop | Peer-reviewed | [12] |
| ALS accelerator agentic AIPlan-first agent ran multistage machine-physics experiments on the ALS accelerator; preparation time cut ~100x under operator safety constraints. | LBNL (ALS) | 2025 | Experiment design & self-driving labsspans: Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [13] |
| AutoLabsSelf-correcting multi-agent system turning natural-language goals into liquid-handler protocols; F1 > 0.89 versus expert procedures on multi-plate syntheses. | Pacific Northwest National Laboratory | 2025 | Experiment design & self-driving labs | L2 Tool-using agent | Peer-reviewed | [223] |
| CREStMultimodal-model-guided Bayesian optimization with robotics; 900+ chemistries and 3,500 tests in 3 months found a catalyst with 9.3-fold cost-specific gain. | MIT | 2025 | Experiment design & self-driving labsspans: Data analysis & simulation | L3 Closed-loop | Peer-reviewed | [224] |
| Instrument agents that learn on the jobHuman-in-the-loop LLM agents orchestrating an X-ray nanoprobe beamline and an autonomous robotic materials station, improving through iterative feedback. | Argonne (APS, CNM) | 2025 | Experiment design & self-driving labsspans: Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [126] |
| NSLS-II multi-beamline AI agentsNon-LLM agents inside Bluesky run on-the-fly reduction, GP modeling and Bayesian-optimized XRD/XAFS mapping across two beamlines. | BNL (NSLS-II) | 2025 | Experiment design & self-driving labsspans: Data analysis & simulation | L3 Closed-loop | Preprint | [143] |
| OspreyControl-room agent framework with human plan review before hardware actions; production deployment across hundreds of thousands of ALS control channels. | LBNL (ALS) | 2025 | Experiment design & self-driving labs | L2 Tool-using agent | Peer-reviewed | [14] |
| TEM AgentMCP tools let a commercial LLM control TEM subsystems, data management and HPC through text instructions, without extra training. | LBNL (Molecular Foundry) | 2025 | Experiment design & self-driving labsspans: Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [130] |
| A-Lab GPSS agentic reasoningAgentic AI proposed 352 air-free solid-state syntheses of halide spinel conductors; success fraction rose from 1.33% to 5.33%. | LBNL (A-Lab) | 2026 | Experiment design & self-driving labsspans: Hypothesis generation | L3 Closed-loop | Preprint | [131] |
| ALD reactor agentLLM agent converts user queries into JSON-encoded atomic layer deposition processes executed on a real reactor through an MCP-compatible interface. | Argonne (Applied Materials Division) | 2026 | Experiment design & self-driving labs | L2 Tool-using agent | Peer-reviewed | [128] |
| Experiment Automation Agents (EAA)Vision-language agents with MCP tools automate zone-plate focusing and natural-language feature search at an APS imaging beamline. | Argonne (APS) | 2026 | Experiment design & self-driving labsspans: Data analysis & simulation | L3 Closed-loop | Peer-reviewed | [127] |
| LightfallAPI-first beamline control platform with embedded LLM agent and staff-authored skills; in testing at ALS COSMIC-Scattering. | LBNL (ALS) | 2026 | Experiment design & self-driving labs | L2 Tool-using agent | Preprint | [103] |
| MARS19 LLM agents and 16 domain tools coordinate robotic synthesis, characterization and analysis for perovskite nanocrystal materials. | Shenzhen Institute of Advanced Technology, CAS | 2026 | Experiment design & self-driving labsspans: Literature & ideation; Data analysis & simulation | L3 Closed-loop | Peer-reviewed | [30] |
| AtomAgentsPhysics-aware multimodal multi-agent system that retrieves knowledge, runs atomistic simulations and analyzes results to design alloys. | MIT | 2024 | Data analysis & simulationspans: Hypothesis generation | L2 Tool-using agent | Peer-reviewed | [225] |
| Data InterpreterHierarchical task-graph planning with verified code generation for end-to-end data science; InfiAgent-DABench accuracy rose from 75.9% to 94.9%. | DeepWisdom / MetaGPT | 2024 | Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [226] |
| DS-AgentCase-based reasoning over Kaggle solutions to build and train ML models; 100% development-stage success with GPT-4 at about $1.60 per run. | Jilin University | 2024 | Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [227] |
| LLaMPHierarchical ReAct agents query Materials Project data and run atomistic simulations, reducing LLM errors on bulk moduli, band gaps and formation energies. | UC Berkeley / Lawrence Berkeley National Laboratory | 2024 | Data analysis & simulationspans: Literature & ideation | L2 Tool-using agent | Peer-reviewed | [25] |
| PEARMultiple LLM agents handle ptychography knowledge retrieval, code generation, parameter recommendation and image reasoning. | Argonne (APS) | 2024 | Data analysis & simulation | L2 Tool-using agent | Preprint | [147] |
| Agent LaboratoryTakes a human idea through literature review, experiments and report writing; 84% cheaper than prior autonomous methods; human feedback improved quality. | AMD / Johns Hopkins University | 2025 | Data analysis & simulationspans: Literature & ideation; Writing & review | L4 End-to-end | Peer-reviewed | [228] |
| AlphaEvolveEvolutionary Gemini coding agent with automated evaluators; found 4x4 complex matrix multiplication using 48 scalar multiplications, first improvement in 56 years. | Google DeepMind | 2025 | Data analysis & simulationspans: Hypothesis generation | L3 Closed-loop | Preprint | [36] |
| BiomniGeneral biomedical agent whose tool environment is mined from thousands of papers across 25 domains; composes code workflows and wet-lab protocols. | Stanford University | 2025 | Data analysis & simulationspans: Experiment design & self-driving labs | L2 Tool-using agent | Peer-reviewed | [7] |
| ChemGraphAgentic computational-chemistry workflows from ML potentials to DFT; on 13 tasks, multi-agent decomposition let small LLMs match GPT-4o in some cases. | Argonne National Laboratory | 2025 | Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [26] |
| DenarioModular multi-agent research assistant from idea to drafted and reviewed paper; showcased expert-scored AI-generated papers across many disciplines, including astrophysics. | Flatiron Institute / University of Cambridge et al. | 2025 | Data analysis & simulationspans: Literature & ideation; Hypothesis generation; Writing & review | L4 End-to-end | Preprint | [229] |
| El Agente QHierarchical-memory multi-agent system that writes, submits and debugs quantum-chemistry workflows from prompts; >87% average task success on course exercises. | University of Toronto | 2025 | Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [230] |
| Empirical Research Assistance (ERA)LLM plus tree search writes score-maximizing scientific software; 40 single-cell methods beat top leaderboard entries; 14 COVID-19 models beat the CDC ensemble. | 2025 | Data analysis & simulation | L3 Closed-loop | Preprint | [231] | |
| KosmosUp to 12-hour runs of data-analysis and literature agents sharing a world model; independent scientists judged 79.4% of report statements accurate. | Edison Scientific (FutureHouse spinout) | 2025 | Data analysis & simulationspans: Literature & ideation; Hypothesis generation; Writing & review | L4 End-to-end | Preprint | [2] |
| MDCrowAgent using 40+ expert-designed tools to set up, run and analyze molecular dynamics simulations; evaluated on 25 tasks of varying difficulty. | FutureHouse / University of Rochester | 2025 | Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [232] |
| SasAgentCoordinator plus three specialist LLM agents use SasView tools for SLD calculation, synthetic data and small-angle scattering fitting. | ORNL (SNS) | 2025 | Data analysis & simulation | L2 Tool-using agent | Peer-reviewed | [145] |
| APEXADeployed 61-tool multi-agent calibration and integration system; guard refuses results not backed by executed tool calls. | Argonne (APS) | 2026 | Data analysis & simulation | L2 Tool-using agent | Preprint | [22] |
| AutoResearchClawMulti-agent research pipeline with self-healing execution and seven human-intervention modes; targeted human input beat both full autonomy and step-by-step oversight. | AIMING Lab, UNC Chapel Hill et al. | 2026 | Data analysis & simulationspans: Hypothesis generation; Writing & review | L4 End-to-end | Preprint | [45] |
| EQSANS-CLIAgent-addressable SANS reduction tool; an external agent with one skill document drives complete reductions through a Slack bot. | ORNL (SNS) | 2026 | Data analysis & simulation | L2 Tool-using agent | Preprint | [104] |
| NeuDiff AgentGoverned agent takes single-crystal neutron data to validated CIF with allowlisted tools and fail-closed gates; 4.6-5.0x faster than manual. | ORNL (SNS TOPAZ) | 2026 | Data analysis & simulationspans: Writing & review | L2 Tool-using agent | Peer-reviewed | [15] |
| Rongzai agentLLM plus GSAS-II agent performs neutron Rietveld refinement from natural-language task to report; deployed at CSNS for external users. | CSNS (China) | 2026 | Data analysis & simulationspans: Writing & review | L3 Closed-loop | Preprint | [136] |
| GPT-4 paper feedback studyGPT-4 comments on full papers overlapped with human reviewers as much as reviewers overlap each other; 57.4% of 308 researchers found it helpful. | Stanford University | 2023 | Writing & review | L1 Assistant | Peer-reviewed | [33] |
| AgentReviewLLM-agent simulation of reviewers, authors and area chairs; reviewer biases alone produced 37.1% variation in paper decisions. | Georgia Tech / William & Mary et al. | 2024 | Writing & review | L1 Assistant | Peer-reviewed | [233] |
| CycleResearcher / CycleReviewerOpen LLMs trained with automated-reviewer feedback draft full papers whose experimental results are fabricated, not run; CycleReviewer cut score error 26.89% versus individual reviewers. | Westlake University | 2024 | Writing & reviewspans: Hypothesis generation | L1 Assistant | Peer-reviewed | [234] |
| data-to-paperBackward-traceable pipeline from annotated data to full paper; simple-goal autopilot runs recapitulated published findings in about 80-90% of cases. | Technion | 2024 | Writing & reviewspans: Hypothesis generation; Data analysis & simulation | L4 End-to-end | Peer-reviewed | [32] |
| The AI Scientist (v1, v2)Generates ideas, code, experiments, full manuscript and self-review; one paper passed first-round review at an ICLR 2025 workshop (70% acceptance rate). | Sakana AI / University of Oxford / UBC | 2024 | Writing & reviewspans: Literature & ideation; Hypothesis generation; Data analysis & simulation | L4 End-to-end | Peer-reviewed | [1] |
| AgentRxivShared preprint server where agent labs upload and build on each other's reports; collaborating labs gained 13.7% relative on MATH-500. | Johns Hopkins University / ETH Zurich | 2025 | Writing & reviewspans: Literature & ideation | L4 End-to-end | Preprint | [235] |
| AI-ResearcherPipeline from literature review and hypotheses to implementation and manuscript; introduces Scientist-Bench; authors report near-human paper quality. | University of Hong Kong | 2025 | Writing & reviewspans: Literature & ideation; Hypothesis generation; Data analysis & simulation | L4 End-to-end | Peer-reviewed | [236] |
| DeepReviewMulti-stage reviewer with literature retrieval and evidence-based argumentation; DeepReviewer-14B outperforms CycleReviewer-70B while using fewer tokens. | Westlake University | 2025 | Writing & review | L2 Tool-using agent | Peer-reviewed | [237] |
| LLM proposal rankingPairwise LLM ranking of proposals from three SNS beamlines correlates with human rankings (Spearman 0.2-0.8) at over 100x lower cost. | ORNL (SNS) | 2025 | Writing & review | L1 Assistant | Peer-reviewed | [149] |
| Review Feedback Agent (ICLR 2025)Randomized trial on 20,000+ ICLR 2025 reviews; 27% of reviewers given LLM feedback updated reviews, incorporating 12,000+ suggestions. | Stanford University / ICLR | 2025 | Writing & review | L1 Assistant | Peer-reviewed | [34] |
| Stanford Agentic ReviewerFree web reviewer grounding feedback in arXiv searches; Spearman correlation with a human reviewer 0.42 versus 0.41 human-human on ICLR 2025 papers. | Stanford University | 2025 | Writing & reviewspans: Literature & ideation | L2 Tool-using agent | Lab/project page | [238] |
| ZochiCommercial autonomous research agent; vendor reports a Zochi-generated paper was accepted to the ACL 2025 main conference. | Intology | 2025 | Writing & reviewspans: Hypothesis generation; Data analysis & simulation | L4 End-to-end | Product/vendor report | [38] |
| PrismFree AI-native LaTeX writing and collaboration workspace with an OpenAI model working inside the document context; launched January 2026. | OpenAI | 2026 | Writing & review | L1 Assistant | Product/vendor report | [47] |
Glossary
- AI agent
- A system in which a language model chooses and calls tools (search, code, databases, simulators, instruments) in a loop, observes the results and decides the next step, rather than answering once.
- Autonomy level (L1–L4)
- This report's four-step scale. L1 assistant: answers or drafts on request. L2 tool-using agent: plans and runs several tool calls inside one task. L3 closed-loop: repeats plan → act → observe against real instruments, robots, simulations or code over many cycles, with human checkpoints. L4 end-to-end: goes from a goal to experiments or analysis to a written paper or report with little human help.
- Self-driving laboratory (SDL)
- A laboratory where robots or instruments run experiments chosen by an algorithm that learns from earlier results. Older SDLs use Bayesian optimization; newer ones add a language-model agent for planning and interfaces.
- Bayesian optimization / active learning
- Methods that choose the next measurement from a statistical model of the results so far, trading off exploration and exploitation. They are used by most pre-LLM autonomous experiments at light and neutron sources.
- Retrieval-augmented generation (RAG)
- Giving a model documents it retrieves at run time (papers, manuals, logbooks) so that its answers can point to sources.
- Deep research
- Products and agents that run many web or literature searches, read the results and write a cited report.
- Multi-agent system
- Several model instances with different roles (for example planner, critic, coder) that pass messages to each other.
- Tool calling / function calling
- A model API feature: the model returns a structured request to run a named function with arguments, and the host program runs it.
- Model Context Protocol (MCP)
- An open protocol, introduced in November 2024, for connecting models to tools, data sources and prompts through a standard client–server interface.
- Agent2Agent (A2A)
- An open protocol for agents built by different vendors or teams to find and message each other.
- Reasoning model
- A model trained to produce long intermediate reasoning before answering. These models do better on multi-step math, code and science tasks.
- Frontier model
- One of the most capable current models from the leading AI companies (for example OpenAI, Anthropic and Google), usually used through a paid online API.
- Token
- The unit of text a model reads and writes, roughly a word or part of a word. Model use, speed and cost are counted in tokens.
- Rollout
- One complete attempt by an agent at a task, from its first step to its final answer.
- Open-weight model
- A model whose weights are published, so it can run on site (for example on a laboratory cluster) without sending data to a vendor.
- Benchmark contamination
- When test items, or close copies of them, appear in a model's training data, so that scores overstate real ability.
- LLM-as-judge
- Using a language model to grade outputs such as ideas, papers or reviews. It is cheap but can be biased or easy to game.
- Hallucination / fabricated citation
- Fluent output that is not true: invented references, numbers or results.
- Prompt injection
- Hidden instructions in data an agent reads (a web page, a PDF, a tool result) that take over the agent's behavior.
- Human-in-the-loop
- A design in which a person approves or edits key steps, for example before an agent moves a motor or submits a job.
- Reward hacking
- When an agent raises the score, or passes the check it is judged by, without doing the intended task: for example by editing tests, using leaked answers or reporting only favorable runs.
- Agent trace
- The step-by-step record of an agent's run: its prompts, reasoning, tool calls and tool results. It shows what the agent actually did, which its final report may not.
- Allowlist
- A fixed list of the tools or commands an agent may use; anything not on the list is refused.
- Fail-closed gate
- A check that stops the workflow when it fails or cannot run, instead of letting the action go ahead.
- Deterministic guard
- A rule written in ordinary code, outside the model, that checks each agent action or report (for example against motor limits, or against the record of commands actually run) and blocks it if the check fails. Unlike a safety instruction in the prompt, the model cannot argue its way past it.
- Digital twin
- A simulation of an instrument or beamline detailed enough to run and score control code before it touches real hardware.
- Bluesky / Ophyd
- Python libraries for experiment orchestration (Bluesky) and hardware abstraction (Ophyd), used at NSLS-II and many other light sources.
- EPICS
- The Experimental Physics and Industrial Control System: the control-system layer (process variables) under most DOE accelerators and beamlines.
- Tiled
- A data-access service that serves array and table data (for example beamline data) over HTTP, with search and slicing.
- SPEC
- A long-established command-line program for diffractometer and beamline control, used at SSRL and many other synchrotrons.
- CIF / checkCIF
- The crystallographic information file (CIF) is the standard text format for a crystal structure and its refinement. checkCIF is the International Union of Crystallography's validation service; its level A and B alerts flag the most serious problems.
- Globus Compute
- A federated function-as-a-service platform that runs Python functions on remote computers, including HPC systems, through a cloud-managed service.
- Parsl
- A Python library for parallel scripting: it runs many tasks as one workflow on laptops, clusters or HPC systems.
- Integrated Research Infrastructure (IRI)
- A DOE Office of Science program to connect experimental facilities, HPC centers and networks so that data can move between them and be analyzed as it is taken.
- ALCF, OLCF, NERSC, ESnet
- DOE Office of Science computing and network facilities: the Argonne and Oak Ridge Leadership Computing Facilities, the National Energy Research Scientific Computing Center at Berkeley Lab, and the Energy Sciences Network that links the labs and user facilities.
- Facilities named in this report
- SLAC: SSRL (Stanford Synchrotron Radiation Lightsource) and LCLS (Linac Coherent Light Source, an x-ray free-electron laser). BNL: NSLS-II (National Synchrotron Light Source II) and CFN (Center for Functional Nanomaterials). Argonne: APS (Advanced Photon Source) and CNM (Center for Nanoscale Materials). LBNL: ALS (Advanced Light Source) and the Molecular Foundry. ORNL: SNS (Spallation Neutron Source) and CNMS (Center for Nanophase Materials Sciences). Outside DOE: CSNS (China Spallation Neutron Source) and DESY (Germany), which runs the PETRA III synchrotron.
- ICLR, ICML, NeurIPS, ACL
- The leading machine-learning and natural-language-processing conferences. In these fields, peer-reviewed conference papers, not journal articles, are the main publications.
- Desk rejection
- Rejection of a submitted paper by the editors or program chairs without full peer review, for example for breaking submission rules.
Method notes
- Run date and scope. The survey reflects sources available on 2026-09-25. It covers AI agents built on language models, plus the closely related autonomous-experiment systems they build on, across the research workflow. The focus is 2024–2026; a few landmark 2023 systems are included for context.
- How the work was split. An orchestrating agent split the survey into parallel research lanes (Q1 catalog, Q2 evidence and benchmarks, Q3 infrastructure, Q4 facilities, Q5 risks) and a tooling lane (report builder and checker), then merged and cross-checked the results in later iterations. The iteration ledger and claims file are in
.goal/. - Sources. Discovery used web search together with the arXiv API, Crossref, OpenAlex, Semantic Scholar and the OSTI.GOV API (for DOE-funded work). Primary sources (papers, preprints, specifications, lab and project pages) were preferred over press coverage. When a preprint has a peer-reviewed version, the published version is cited.
- Evidence labels. Peer-reviewed: published in a refereed journal or archival proceedings. Peer-reviewed + independent validation: also checked by a group other than the developers. Preprint: not yet refereed. Product/vendor report and Lab/project page: self-reported. In Q2 the first two classes are kept apart from claims.
- Autonomy scale. Systems are placed on the L1–L4 scale defined in the glossary, based on what the cited source shows the system doing, not on what it is said to be able to do in future.
- Reference checks.
check_report.pyparses this file and confirms, in the same run: every DOI through Crossref (or doi.org for DataCite DOIs), with the fetched title matching the listed title; every arXiv ID through the arXiv API, with the title matching; every other URL by fetching it. It also checks that every reference is cited, that every citation points to a listed reference, and that the page loads no external resources. - Fact checks. Before publication, independent audit agents checked the text against primary sources in three passes: 79 headline numbers (12 corrected or re-attributed), 211 further statements across all six questions (36 corrected), and the labels of all 83 catalog rows (20 changes, mostly first-release years and four autonomy levels). The audit files are in
notes/. - Limitations. The field moves quickly, and many results exist only as preprints or company announcements. Title checks confirm that a reference exists and is the one listed; they do not confirm every sentence attributed to it. Facility work that is internal or unpublished is not visible to this method. The audience is described only by institution and field.
References
Numbered in order of first citation. Each entry shows the evidence type in brackets. DOIs link through doi.org, preprints through arXiv, and other sources are linked directly.
- Towards end-to-end automation of AI research. Nature (2026). doi:10.1038/s41586-026-10265-5 [Peer-reviewed]
- Kosmos: An AI Scientist for Autonomous Discovery. arXiv preprint (2025). arXiv:2511.02824 [Preprint]
- Towards Autonomous Mathematics Research. arXiv preprint (2026). arXiv:2602.10177 [Preprint]
- Accelerating scientific discovery with Co-Scientist. Nature (2026). doi:10.1038/s41586-026-10644-y [Peer-reviewed + independent validation]
- A multi-agent system for automating scientific discovery. Nature (2026). doi:10.1038/s41586-026-10652-y [Peer-reviewed]
- Synthesizing scientific literature with retrieval-augmented language models. Nature (2026). doi:10.1038/s41586-025-10072-4 [Peer-reviewed]
- Autonomous biomedical research with an artificial intelligence agent. Science (2026). doi:10.1126/science.adz4351 [Peer-reviewed]
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark. arXiv (2023). arXiv:2311.12022 [Benchmark (preprint)]
- . AI Benchmarks & Capabilities | Epoch AI. Epoch AI Benchmarking Hub (2026). link [Lab/project page]
- PaperBench: Evaluating AI's Ability to Replicate AI Research. arXiv (2025). arXiv:2504.01848 [Benchmark (preprint)]
- EXP-Bench: Can AI Conduct AI Research Experiments? arXiv (2025). arXiv:2505.24785 [Benchmark (preprint)]
- An agentic artificially intelligent X-ray scientist. Nature Machine Intelligence (2026). doi:10.1038/s42256-026-01261-5 [Peer-reviewed]
- Agentic artificial intelligence for multistage physics experiments at a large-scale user facility particle accelerator. Physical Review Research (2026). doi:10.1103/jtqy-9jz1 [Peer-reviewed]
- Osprey: Production-ready agentic AI for safety-critical control systems. APL Machine Learning (2026). doi:10.1063/5.0306302 [Peer-reviewed]
- NeuDiff Agent: a governed AI workflow for single-crystal neutron crystallography. Journal of Applied Crystallography (2026). doi:10.1107/S1600576726004474 [Peer-reviewed]
- . MCP joins the Agentic AI Foundation | Model Context Protocol Blog. MCP blog (2025). link [Official docs/repo]
- . FIRST: Federated Inference Resource Scheduling Toolkit for Scientific AI Model Access. SC '25 Workshops (2025). doi:10.1145/3731599.3767346 [Peer-reviewed]
- . OLCF Inference Service Documentation — OLCF User Documentation. docs.olcf.ornl.gov (2026). link [Official docs/repo]
- . Experiences with Model Context Protocol Servers for Science and High Performance Computing. arXiv (2025). arXiv:2508.18489 [Preprint]
- . Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences. arXiv (2026). arXiv:2607.00738 [Preprint]
- Reward Hacking Challenges Oversight of Autonomous Research Agents. arXiv (2026). arXiv:2609.28614 [Preprint]
- APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction. arXiv (2026). arXiv:2609.24165 [Preprint]
- . Launching the Genesis Mission – The White House. Executive order (2025). link [Lab/project page]
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap. arXiv preprint (survey) (2026). arXiv:2608.05179 [Preprint]
- LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval. EMNLP 2025 (2025). doi:10.18653/v1/2025.emnlp-main.1280 [Peer-reviewed]
- ChemGraph as an agentic framework for computational chemistry workflows. Communications Chemistry (2026). doi:10.1038/s42004-025-01776-9 [Peer-reviewed]
- Language agents achieve superhuman synthesis of scientific knowledge. arXiv preprint (2024). arXiv:2409.13740 [Preprint]
- AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite. ICLR 2026 (2025). arXiv:2510.21652 [Peer-reviewed]
- Autonomous chemical research with large language models. Nature (2023). doi:10.1038/s41586-023-06792-0 [Peer-reviewed]
- Knowledge-driven autonomous materials research via collaborative multi-agent and robotic system. Matter (2026). doi:10.1016/j.matt.2025.102577 [Peer-reviewed]
- The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies. Nature (2025). doi:10.1038/s41586-025-09442-9 [Peer-reviewed]
- Autonomous LLM-Driven Research — from Data to Human-Verifiable Research Papers. NEJM AI (2025). doi:10.1056/AIoa2400555 [Peer-reviewed]
- Can Large Language Models Provide Useful Feedback on Research Papers? A Large-Scale Empirical Analysis. NEJM AI (2024). doi:10.1056/AIoa2400196 [Peer-reviewed]
- A large-scale randomized study of large language model feedback in peer review. Nature Machine Intelligence (2026). doi:10.1038/s42256-026-01188-x [Peer-reviewed]
- Augmenting large language models with chemistry tools. Nature Machine Intelligence (2024). doi:10.1038/s42256-024-00832-8 [Peer-reviewed]
- AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv (white paper) (2025). arXiv:2506.13131 [Preprint]
- . Announcing Edison Scientific | FutureHouse. FutureHouse announcement (2025). link [Product/vendor report]
- . Zochi Publishes A* Paper | Intology. Intology blog (2025). link [Product/vendor report]
- DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively. ICLR 2026 (2025). arXiv:2509.26603 [Peer-reviewed]
- Exploring the use of AI authors and reviewers at Agents4Science. Nature Biotechnology (2025). doi:10.1038/s41587-025-02963-8 [Peer-reviewed]
- AI mirrors experimental science to uncover a mechanism of gene transfer crucial to bacterial evolution. Cell (2025). doi:10.1016/j.cell.2025.08.018 [Peer-reviewed]
- Challenges in High-Throughput Inorganic Materials Prediction and Autonomous Synthesis. PRX Energy (2024). doi:10.1103/PRXEnergy.3.011002 [Peer-reviewed]
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers. ICLR 2025 (2025). arXiv:2409.04109 [Peer-reviewed]
- The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas. ICLR 2026 (2025). arXiv:2506.20803 [Peer-reviewed]
- AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration. arXiv preprint (2026). arXiv:2605.20025 [Preprint]
- Aletheia tackles FirstProof autonomously. arXiv preprint (2026). arXiv:2602.21201 [Preprint]
- . Prism - AI LaTeX Editor. Product page (2026). link [Product/vendor report]
- Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery. arXiv preprint (position paper) (2026). arXiv:2605.08956 [Preprint]
- Probing the limitations of multimodal language models for chemistry and materials research. Nature Computational Science (2025). doi:10.1038/s43588-025-00836-3 [Benchmark (peer-reviewed)]
- An autonomous laboratory for the accelerated synthesis of inorganic materials. Nature (2023). doi:10.1038/s41586-023-06734-w [Peer-reviewed]
- Author Correction: An autonomous laboratory for the accelerated synthesis of inorganic materials. Nature (2026). doi:10.1038/s41586-025-09992-y [Peer-reviewed]
- CRISPR-GPT for agentic automation of gene-editing experiments. Nature Biomedical Engineering (2025). doi:10.1038/s41551-025-01463-z [Peer-reviewed]
- Chimeric infective particles expand species boundaries in phage-inducible chromosomal island mobilization. Cell (2025). doi:10.1016/j.cell.2025.08.019 [Peer-reviewed]
- AI‐Assisted Drug Re‐Purposing for Human Liver Fibrosis. Advanced Science (2025). doi:10.1002/advs.202508751 [Peer-reviewed + independent validation]
- Mathematical exploration and discovery at scale. arXiv (2025). arXiv:2511.02864 [Preprint]
- Evaluating large language model agents for automation of atomic force microscopy. Nature Communications (2025). doi:10.1038/s41467-025-64105-7 [Peer-reviewed]
- . GitHub - IntologyAI/Zochi: Repository for Zochi's Research · GitHub. GitHub (2025). link [Product/vendor report]
- . Nature Is Our Learning Environment – Periodic Labs. Periodic Labs blog (2026). link [Product/vendor report]
- AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution. arXiv (2026). arXiv:2609.30133 [Preprint]
- Early science acceleration experiments with GPT-5. arXiv preprint (2025). arXiv:2511.16072 [Preprint]
- . GitHub - openai/NavierStokesAndEuler: Lean certificates accompanying Navier-Stokes and Euler results · GitHub. GitHub (2026). link [Product/vendor report]
- A benchmark of expert-level academic questions to assess AI capabilities. Nature (2026). doi:10.1038/s41586-025-09962-4 [Benchmark (peer-reviewed)]
- . Scale Labs Leaderboard: Humanity's Last Exam. Scale Labs leaderboard (2026). link [Lab/project page]
- FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks. arXiv (2026). arXiv:2601.21165 [Benchmark (preprint)]
- Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark. arXiv (2025). arXiv:2509.26574 [Benchmark (preprint)]
- A framework for evaluating the chemical knowledge and reasoning abilities of large language models against the expertise of chemists. Nature Chemistry (2025). doi:10.1038/s41557-025-01815-x [Benchmark (peer-reviewed)]
- LAB-Bench: Measuring Capabilities of Language Models for Biology Research. arXiv (2024). arXiv:2407.10362 [Benchmark (preprint)]
- Benchmarking large language models on safety risks in scientific laboratories. Nature Machine Intelligence (2026). doi:10.1038/s42256-025-01152-1 [Benchmark (peer-reviewed)]
- SciCode: A Research Coding Benchmark Curated by Scientists. arXiv (2024). arXiv:2407.13168 [Benchmark (preprint)]
- . Leaderboard - SciCode Benchmark. SciCode leaderboard (2025). link [Lab/project page]
- . SciCode Benchmark Leaderboard | Artificial Analysis. Artificial Analysis leaderboard (2026). link [Lab/project page]
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery. ICLR 2025 (2024). arXiv:2410.05080 [Benchmark (peer-reviewed)]
- CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark. TMLR (2024). arXiv:2409.11363 [Benchmark (peer-reviewed)]
- . HAL: CORE-Bench Hard Leaderboard. HAL leaderboard (2026). link [Lab/project page]
- BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology. arXiv (2025). arXiv:2503.00096 [Benchmark (preprint)]
- DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents. arXiv (2024). arXiv:2406.06769 [Benchmark (preprint)]
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering. arXiv (2024). arXiv:2410.07095 [Benchmark (preprint)]
- . GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHub. GitHub leaderboard (2026). link [Lab/project page]
- MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation. arXiv (2023). arXiv:2310.03302 [Benchmark (preprint)]
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts. arXiv (2024). arXiv:2411.15114 [Benchmark (preprint)]
- SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers. arXiv (2025). arXiv:2504.00255 [Benchmark (preprint)]
- ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers? arXiv (2025). arXiv:2510.24591 [Benchmark (preprint)]
- . AstaBench: Rigorous benchmarking of AI agents with a holistic scientific research suite | Ai2. Ai2 blog (2025). link [Lab/project page]
- CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning. arXiv (2025). arXiv:2503.13517 [Benchmark (preprint)]
- FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights. arXiv (2026). arXiv:2602.02905 [Benchmark (preprint)]
- Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction. arXiv (2026). arXiv:2605.13950 [Benchmark (preprint)]
- . About 30% of Humanity’s Last Exam Answers are Wrong | FutureHouse. FutureHouse blog (2025). link [Lab/project page]
- Life After Benchmark Saturation: A Case Study of CORE-Bench. arXiv (2026). arXiv:2606.26158 [Benchmark (preprint)]
- . GitHub - argonne-lcf/ChemGraph: Agentic framework for computational chemistry and materials science workflows · GitHub. GitHub (2026). link [Official docs/repo]
- . OpenAI GPT-5 System Card. arXiv (2025). arXiv:2601.03267 [Product/vendor report]
- . GitHub - snap-stanford/Biomni: Biomni: a general-purpose biomedical AI agent · GitHub. GitHub (2026). link [Official docs/repo]
- . gpt-oss-120b & gpt-oss-20b Model Card. arXiv (2025). arXiv:2508.10925 [Product/vendor report]
- . Training a Scientific Reasoning Model for Chemistry. Advances in Neural Information Processing Systems 38 (2025). doi:10.52202/085713-5269 [Peer-reviewed]
- . NNSA's Los Alamos National Laboratory launches frontier AI models on the Venado supercomputer | Department of Energy. DOE/NNSA news (2025). link [Government document]
- ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023 (2022). arXiv:2210.03629 [Peer-reviewed]
- . GitHub - langchain-ai/langgraph: Build resilient agents. · GitHub. GitHub (2026). link [Official docs/repo]
- . GitHub - microsoft/autogen: A programming framework for agentic AI · GitHub. GitHub (2026). link [Official docs/repo]
- . Agent SDK overview - Claude Code Docs. Official docs (2026). link [Official docs/repo]
- . Empowering Scientific Workflows with Federated Agents. IPDPS 2026 (2026). doi:10.1109/ipdps65963.2026.00114 [Peer-reviewed]
- . A New Chapter for A2A: Joining the Agentic AI Foundation - A2A Protocol. a2a-protocol.org blog (2026). link [Official docs/repo]
- . Key Changes - Model Context Protocol. modelcontextprotocol.io (2026). link [Specification]
- . Specification - Agent Skills. agentskills.io (2026). link [Specification]
- Lightfall: An API-first, LLM-addressable control platform for synchrotron beamlines. arXiv (2026). arXiv:2606.06711 [Preprint]
- . EQSANS-CLI: A natural-language, agent-ready command-line tool for small-angle neutron scattering data reduction at EQ-SANS. arXiv (2026). arXiv:2605.00651 [Preprint]
- MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers. arXiv (2025). arXiv:2508.14925 [Preprint]
- . Trusted Credentials, Untrusted Behavior: Benchmarking LLM-Agent Security in High-Performance Computing. arXiv (2026). arXiv:2607.18485 [Preprint]
- funcX: A Federated Function Serving Fabric for Science. HPDC '20 (2020). doi:10.1145/3369583.3392683 [Peer-reviewed]
- Parsl. HPDC '19 (Parsl: Pervasive Parallel Programming in Python) (2019). doi:10.1145/3307681.3325400 [Peer-reviewed]
- . Multi-Agent Orchestration for High-Throughput Materials Screening on a Leadership-Class System. arXiv (2026). arXiv:2604.07681 [Preprint]
- . GitHub - doe-iri/iri-facility-api-python: The IRI Facility API reference implementation (Python) · GitHub. GitHub (2026). link [Official docs/repo]
- . GitHub - slaclab/iri-facility-api-python: The IRI Facility API reference implementation (Python) · GitHub. GitHub (2026). link [Official docs/repo]
- . Bluesky Project. blueskyproject.io (2026). link [Official docs/repo]
- . EPICS - Experimental Physics and Industrial Control System. epics-controls.org (2026). link [Official docs/repo]
- . Tiled documentation — tiled main documentation. Project documentation (2026). link [Lab/project page]
- VISION: a modular AI assistant for natural human-instrument interaction at scientific user facilities. Machine Learning: Science and Technology (2025). doi:10.1088/2632-2153/add9e4 [Peer-reviewed]
- . AI Agents for Enabling Autonomous Experiments at ORNL's HPC and Manufacturing User Facilities. SC '25 Workshops (2025). doi:10.1145/3731599.3767592 [Peer-reviewed]
- . PROV-AGENT: Unified Provenance for Tracking AI Agent Interactions in Agentic Workflows. IEEE eScience 2025 (2025). doi:10.1109/escience65000.2025.00093 [Peer-reviewed]
- . Overview - NERSC Documentation. docs.nersc.gov (2026). link [Official docs/repo]
- EnvTrace: Simulation-Based Semantic Evaluation of LLM Code via Execution Trace Alignment -- Demonstrated at Synchrotron Beamlines. arXiv (2025). arXiv:2511.09964 [Preprint]
- . DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645:633-638 (2025). doi:10.1038/s41586-025-09422-z [Peer-reviewed]
- A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility. arXiv (2026). arXiv:2607.24663 [Preprint]
- Domain Knowledge Guided Bayesian Optimization For Autonomous Alignment Of Complex Scientific Instruments. arXiv (2026). arXiv:2602.10670 [Preprint]
- A Multi-Scale Cognitive Interaction Model of Instrument Operations at the Linac Coherent Light Source. arXiv (2024). arXiv:2408.04734 [Preprint]
- ChatEED: An agentic retrieval assistant for accelerator operators. Proceedings of the SC '25 Workshops (2025). doi:10.1145/3731599.3767408 [Peer-reviewed]
- Opportunities for retrieval and tool augmented large language models in scientific facilities. npj Computational Materials (2024). doi:10.1038/s41524-024-01423-2 [Peer-reviewed]
- Operating advanced scientific instruments with AI agents that learn on the job. npj Computational Materials (2026). doi:10.1038/s41524-026-02005-0 [Peer-reviewed]
- Experiment automation agents: automating materials characterization with vision language model agents. npj Computational Materials (2026). doi:10.1038/s41524-026-02213-8 [Peer-reviewed]
- Design and performance of AI agents interfacing with an atomic layer deposition tool. Review of Scientific Instruments (2026). doi:10.1063/5.0318770 [Peer-reviewed]
- From natural language to control signals: a conceptual framework for semantic channel finding in complex experimental infrastructure. Machine Learning: Science and Technology (2026). doi:10.1088/2632-2153/ae8219 [Peer-reviewed]
- TEM Agent: enhancing transmission electron microscopy with modern AI tools. npj Computational Materials (2026). doi:10.1038/s41524-026-02103-z [Peer-reviewed]
- Agentic LLM Reasoning in a Self-Driving Laboratory for Air-Sensitive Lithium Halide Spinel Conductors. arXiv (2026). arXiv:2604.11957 [Preprint]
- Synergizing human expertise and AI efficiency with language model for microscopy operation and automated experiment design *. Machine Learning: Science and Technology (2024). doi:10.1088/2632-2153/ad52e9 [Peer-reviewed]
- A Microservices Architecture Toolkit for Interconnected Science Ecosystems. SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis (2024). doi:10.1109/SCW63240.2024.00259 [Peer-reviewed]
- eLog analysis for accelerators: status and future outlook. arXiv (IPAC'25) (2025). arXiv:2506.12949 [Preprint]
- Large language models for human-machine collaborative particle accelerator tuning through natural language. Science Advances (2025). doi:10.1126/sciadv.adr4173 [Peer-reviewed]
- Rongzai agent: A Large Language Model-Based Autonomous Assistant for Rietveld Refinement of Neutron Diffraction Data. arXiv (2026). arXiv:2605.13911 [Preprint]
- From Overload to Insights: How AI Agents Can Support Scientists in Analyzing Complex Data. arXiv preprint (2026). arXiv:2607.16845 [Preprint]
- A Kriging-Based Approach to Autonomous Experimentation with Applications to X-Ray Scattering. Scientific Reports (2019). doi:10.1038/s41598-019-48114-3 [Peer-reviewed]
- On-the-fly closed-loop materials discovery via Bayesian active learning. Nature Communications (2020). doi:10.1038/s41467-020-19597-w [Peer-reviewed]
- On-the-fly autonomous control of neutron diffraction via physics-informed Bayesian active learning. Applied Physics Reviews (2022). doi:10.1063/5.0082956 [Peer-reviewed]
- Demonstration of an AI-driven workflow for autonomous high-resolution scanning microscopy. Nature Communications (2023). doi:10.1038/s41467-023-40339-1 [Peer-reviewed]
- Self-driving Multimodal Studies at User Facilities. arXiv (2023). arXiv:2301.09177 [Preprint]
- A modular framework for collaborative human-AI, multi-modal and multi-beamline synchrotron experiments. arXiv (2025). arXiv:2509.22959 [Preprint]
- Integrated edge-to-exascale workflow for real-time steering in neutron scattering experiments. Structural Dynamics (2024). doi:10.1063/4.0000279 [Peer-reviewed]
- SasAgent : multi-agent artificial intelligence system for small-angle scattering data analysis. Journal of Applied Crystallography (2026). doi:10.1107/S1600576726001299 [Peer-reviewed]
- . Automated Data Reduction on VULCAN. OSTI technical report (ORNL) (2026). doi:10.2172/3413708 [Government document]
- PEAR: A Robust and Flexible Automation Framework for Ptychography Enabled by Multiple Large Language Model Agents. arXiv (2024). arXiv:2410.09034 [Preprint]
- ESAC (EQ-SANS Assisting Chatbot): Application of large language models and retrieval-augmented generation for enhanced user experience at EQ-SANS. SoftwareX (2025). doi:10.1016/j.softx.2025.102191 [Peer-reviewed]
- LLMs can assist with proposal selection at large user facilities. Scientific Reports (2026). doi:10.1038/s41598-026-60226-1 [Peer-reviewed]
- . Energy Department Announces 26 Genesis Mission Science and Technology Challenges to Accelerate AI-Enabled American Innovation and Leadership | Department of Energy. DOE press release (2026). link [Government document]
- . Achieving AI-Driven Autonomous Laboratories | Department of Energy. DOE web page (2026). link [Government document]
- . Energy Department Announces Collaboration Agreements with 24 Organizations to Advance the Genesis Mission | Department of Energy. DOE press release (2025). link [Government document]
- . Trump Administration Announces More Than $5 Billion for the Genesis Mission, a National Mission on AI for Science – The White House. Press release (2026). link [Government document]
- . U.S. Department of Energy Announces More Than $800 Million in Partner Commitments to the Genesis Mission | Department of Energy. DOE press release (2026). link [Government document]
- . Genesis Mission Frameworks for AI-Accelerated National Breakthroughs: Report from the SCAC Subcommittee on the Genesis Mission. OSTI technical report (2026). doi:10.2172/3387411 [Government document]
- . Public Law 119 - 21 - An act to provide for reconciliation pursuant to title II of H. Con. Res. 14. - PLAW-119publ21 | Content Details | GovInfo. Public Law 119-21 (2025). link [Government document]
- . GRANTS The American Science Clou... | U.S. DOE Office of Science(SC). Lab announcement LAB 25-3555 (2025). link [Government document]
- . How the Genesis Mission’s American Science Cloud Advances Innovation – Berkeley Lab News Center. Berkeley Lab News Center (2026). link [Lab/project page]
- . Real‑time AI engine poised to revolutionize large‑scale imaging data at national labs. Lab news release (2026). link [Lab/project page]
- . Energy Department Advances Investments in AI for Science | Department of Energy. DOE press release (2025). link [Government document]
- . GRANTS Robotics and Automation T... | U.S. DOE Office of Science(SC). Lab announcement LAB 26-3601 (2026). link [Government document]
- . DOE Announces Roadmap for New Initiative for Artificial Intelligence in Science, Security and Technology | Department of Energy. DOE press release (2024). link [Government document]
- . 1,000 Scientist AI Jam Session explores AI-driven scientific discovery | Lawrence Livermore National Laboratory. LLNL news (2025). link [Lab/project page]
- . Los Alamos Lab partners with OpenAI to boost national security | LANL. LANL news (2025). link [Lab/project page]
- . Working with the US Department of Energy \ Anthropic. Company news (2025). link [Product/vendor report]
- . Energy Department Announces New Partnership with NVIDIA and Oracle to Build Largest DOE AI Supercomputer | Department of Energy. DOE press release (2025). link [Government document]
- . Energy Department Announces New Public-Private Partnership Model, Two Supercomputers, to Accelerate American Dominance in Science and Technology | Department of Energy. DOE press release (2025). link [Government document]
- Integrated Research Infrastructure Architecture Blueprint Activity (Final Report 2023). OSTI technical report (2023). doi:10.2172/1984466 [Government document]
- Bluesky's Ahead: A Multi-Facility Collaboration for an a la Carte Software Project for Data Acquisition and Management. Synchrotron Radiation News (2019). doi:10.1080/08940886.2019.1608121 [Peer-reviewed]
- Toward Unified Autonomous Scattering Experiments: A Cross-Facility Case Study at ALS and PETRA III. Photon Science (2026). doi:10.1021/photonsci.5c00044 [Peer-reviewed]
- Toward Agentic HPC: Serving and Evaluating LLM-Powered Agents for Scientific Applications on Leadership-Class Platforms. ISC High Performance 2026 Research Paper Proceedings (2026). doi:10.23919/ISC.2026.11520499 [Peer-reviewed]
- . Benchmarking Autonomy in Scientific Experiments: A Hierarchical Taxonomy for Autonomous Large-Scale Facilities. arXiv (2026). arXiv:2601.06978 [Preprint]
- Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap. OSTI technical report (ORNL) (2026). doi:10.2172/3377973 [Government document]
- AI Agents That Matter. arXiv (2024). arXiv:2407.01502 [Preprint]
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox. ICLR 2024 (2024). arXiv:2309.15817 [Peer-reviewed]
- . ICMJE | Recommendations | Preparing a Manuscript for Submission to a Medical Journal. ICMJE Recommendations (2025). link [Editorial/policy]
- Non-Determinism of "Deterministic" LLM Settings. arXiv (2024). arXiv:2408.04667 [Preprint]
- . How Is ChatGPT’s Behavior Changing Over Time? Harvard Data Science Review (2024). doi:10.1162/99608f92.5317da47 [Peer-reviewed]
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation. arXiv (2025). arXiv:2510.11977 [Preprint]
- . Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future? ACM SIGIR Forum (2025). doi:10.1145/3769733.3769747 [Peer-reviewed]
- . Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports (2023). doi:10.1038/s41598-023-41032-5 [Peer-reviewed]
- . ICLR 2026 Response to LLM-Generated Papers and Reviews – ICLR Blog. ICLR Blog (2025). link [Official page]
- . The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems. arXiv (2025). arXiv:2509.08713 [Preprint]
- . The Agentic Garden of Forking Paths. arXiv (2026). arXiv:2607.01507 [Preprint]
- A Careful Examination of Large Language Model Performance on Grade School Arithmetic. Advances in Neural Information Processing Systems 37 (NeurIPS 2024) (2024). doi:10.52202/079017-1485 [Peer-reviewed]
- Establishing Best Practices for Building Rigorous Agentic Benchmarks. arXiv (2025). arXiv:2507.02825 [Preprint]
- Why Do Multi-Agent LLM Systems Fail? Advances in Neural Information Processing Systems 38 (NeurIPS 2025) (2025). doi:10.52202/085713-4082 [Peer-reviewed]
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems 36 (NeurIPS 2023) (2023). doi:10.52202/075280-2020 [Peer-reviewed]
- . Artificial intelligence and illusions of understanding in scientific research. Nature (2024). doi:10.1038/s41586-024-07146-0 [Peer-reviewed]
- Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. AISec '23 (ACM) (2023). doi:10.1145/3605764.3623985 [Peer-reviewed]
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv (2024). arXiv:2408.06292 [Preprint]
- . Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review. Communications of the ACM (2026). doi:10.1145/3779116 [Peer-reviewed]
- Dual use of artificial-intelligence-powered drug discovery. Nature Machine Intelligence (2022). doi:10.1038/s42256-022-00465-9 [Peer-reviewed]
- Robotic safety in self-driving laboratories. SLAS Technology (2026). doi:10.1016/j.slast.2026.100463 [Peer-reviewed]
- . Shaping the Future of Self-Driving Autonomous Laboratories Workshop. OSTI technical report (DOE) (2024). doi:10.2172/2481197 [Official page]
- . Tools such as ChatGPT threaten transparent science; here are our ground rules for their use. Nature (2023). doi:10.1038/d41586-023-00191-1 [Editorial/policy]
- . LLM Policy. NeurIPS website (2025). link [Editorial/policy]
- . Policies on Large Language Model Usage at ICLR 2026 – ICLR Blog. ICLR Blog (2025). link [Editorial/policy]
- . On Violations of LLM Review Policies – ICML Blog. ICML Blog (2026). link [Official page]
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews. ICML 2024 (2024). arXiv:2403.07183 [Peer-reviewed]
- . Pangram Predicts 21% of ICLR Reviews are AI-Generated | Pangram. Pangram blog (2025). link [Product/vendor report]
- . Open Conference of AI Agents for Science: 2025. Conference website (2025). link [Official page]
- . NOT-OD-23-149: The Use of Generative Artificial Intelligence Technologies is Prohibited for the NIH Peer Review Process. NIH Guide notice (2023). link [Editorial/policy]
- . NOT-OD-25-132: Supporting Fairness and Originality in NIH Research Applications. NIH Guide notice (2025). link [Editorial/policy]
- . Notice to Research Community: Use of Generative Artificial Intelligence Technology in the NSF Merit Review Process - Policies | NSF - U.S. National Science Foundation. NSF policy notice (2023). link [Editorial/policy]
- . GAIA: A General AI Assistant for Intelligent Accelerator Operations. arXiv (2024). arXiv:2405.01359 [Preprint]
- . Gemini: Try Deep Research and Gemini 2.0 Flash Experimental. Google blog (The Keyword) (2024). link [Product/vendor report]
- ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models. NAACL 2025 (2025). doi:10.18653/v1/2025.naacl-long.342 [Peer-reviewed]
- Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models. NAACL 2024 (2024). doi:10.18653/v1/2024.naacl-long.347 [Peer-reviewed]
- Ai2 Scholar QA: Organized Literature Synthesis with Attribution. ACL 2025 System Demonstrations (2025). doi:10.18653/v1/2025.acl-demo.49 [Peer-reviewed]
- . FutureHouse Platform: Superintelligent AI Agents for Science | FutureHouse. FutureHouse announcement (2025). link [Product/vendor report]
- Mathematical discoveries from program search with large language models. Nature (2023). doi:10.1038/s41586-023-06924-6 [Peer-reviewed]
- SciMON: Scientific Inspiration Machines Optimized for Novelty. ACL 2024 (2024). doi:10.18653/v1/2024.acl-long.18 [Peer-reviewed]
- MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses. ICLR 2025 (2025). arXiv:2410.07076 [Peer-reviewed]
- . SciAgents: Automating Scientific Discovery Through Bioinspired Multi‐Agent Intelligent Graph Reasoning. Advanced Materials (2024). doi:10.1002/adma.202413523 [Peer-reviewed]
- InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification. arXiv preprint (2025). arXiv:2505.16938 [Preprint]
- Automated Hypothesis Validation with Agentic Sequential Falsifications. ICML 2025 (2025). arXiv:2502.09858 [Peer-reviewed]
- BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments. ICLR 2025 (2025). arXiv:2405.17631 [Peer-reviewed]
- A Multiagent-Driven Robotic AI Chemist Enabling Autonomous Chemical Research On Demand. Journal of the American Chemical Society (2025). doi:10.1021/jacs.4c17738 [Peer-reviewed]
- Automating quantum computing laboratory experiments with an agent-based AI framework. Patterns (2025). doi:10.1016/j.patter.2025.101372 [Peer-reviewed]
- Autonomous mobile robots for exploratory synthetic chemistry. Nature (2024). doi:10.1038/s41586-024-08173-7 [Peer-reviewed]
- ORGANA: A robotic assistant for automated chemistry experimentation and characterization. Matter (2025). doi:10.1016/j.matt.2024.10.015 [Peer-reviewed]
- AutoLabs: cognitive multi-agent systems with self-correction for autonomous chemical experimentation. Scientific Reports (2026). doi:10.1038/s41598-026-45593-z [Peer-reviewed]
- A multimodal robotic platform for multi-element electrocatalyst discovery. Nature (2025). doi:10.1038/s41586-025-09640-5 [Peer-reviewed]
- . Automating alloy design and discovery with physics-aware multimodal multiagent AI. PNAS (2025). doi:10.1073/pnas.2414074122 [Peer-reviewed]
- Data Interpreter: An LLM Agent for Data Science. Findings of ACL 2025 (2025). doi:10.18653/v1/2025.findings-acl.1016 [Peer-reviewed]
- DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning. ICML 2024 (2024). arXiv:2402.17453 [Peer-reviewed]
- Agent Laboratory: Using LLM Agents as Research Assistants. Findings of EMNLP 2025 (2025). doi:10.18653/v1/2025.findings-emnlp.320 [Peer-reviewed]
- The Denario project: Deep knowledge AI agents for scientific discovery. arXiv preprint (2025). arXiv:2510.26887 [Preprint]
- El Agente: An autonomous agent for quantum chemistry. Matter (2025). doi:10.1016/j.matt.2025.102263 [Peer-reviewed]
- An AI system to help scientists write expert-level empirical software. arXiv preprint (2025). arXiv:2509.06503 [Preprint]
- MDCrow: automating molecular dynamics workflows with large language models. Machine Learning: Science and Technology (2026). doi:10.1088/2632-2153/ae4b07 [Peer-reviewed]
- AgentReview: Exploring Peer Review Dynamics with LLM Agents. EMNLP 2024 (2024). doi:10.18653/v1/2024.emnlp-main.70 [Peer-reviewed]
- CycleResearcher: Improving Automated Research via Automated Review. ICLR 2025 (2025). arXiv:2411.00816 [Peer-reviewed]
- . AgentRxiv: Towards Collaborative Autonomous Research. arXiv preprint (2025). arXiv:2503.18102 [Preprint]
- AI-Researcher: Autonomous Scientific Innovation. NeurIPS 2025 (2025). arXiv:2505.18705 [Peer-reviewed]
- DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process. ACL 2025 (2025). doi:10.18653/v1/2025.acl-long.1420 [Peer-reviewed]
- . Tech Overview - Stanford Agentic Reviewer. Project page (2025). link [Lab/project page]