At a glance

Q1 · Landscape: where agents sit in the scientific workflow

Answer. Of 83 cataloged systems, about half are L2 tool-using agents; closed-loop (L3) agents cluster in self-driving labs and at facility instruments; end-to-end (L4) autonomy appears mainly where the "experiment" is code, existing data or a proof; and independent verification lags [1, 3, 24].

Taxonomy: stage × autonomy

Here an agent is a large language model (LLM) that calls tools (search, code, simulations, instrument controls) and picks its next step from the results. We sorted 83 systems by main workflow stage and by autonomy: 58 general-science systems (mostly 2024–2026) and 25 facility systems covered in Q4. The general set includes national-lab tools such as LLaMP (LBNL) and ChemGraph (Argonne) [25, 26].

Autonomy levels. L1 Assistant: answers or retrieves on request. L2 Tool-using agent: plans and runs multi-step tool calls within one task. L3 Closed-loop: iterates plan → act → observe against real instruments, robots, simulations or code over many cycles, with human checkpoints. L4 End-to-end: goes from a goal through experiments or analysis to a written paper with minimal human help. Stages: literature & ideation; hypothesis generation; experiment design & self-driving labs (labs where software chooses and runs the next experiment); data analysis & simulation; and writing & review (of papers and proposals). Each system sits under its main stage.

Stage L1 AssistantL2 Tool-using agentL3 Closed-loopL4 End-to-end Total
Literature & ideationChatEED ESACAi2 Asta agents + AstaBench Ai2 Scholar QA FutureHouse Platform (Crow, Falcon, Owl, Phoenix) Gemini Deep Research LLM ideation agent vs. 100+ NLP researchers OpenScholar PaperQA2 ResearchAgent STORM APS-RAG CALMS GAIA14
Hypothesis generationCo-Scientist MOOSE-Chem POPPER SciAgents SciMONDeepScientist FunSearch InternAgent (NovelSeek) RobinAletheia10
Experiment design & self-driving labsLLM-assisted scanning probe microscopyAutoLabs ChemCrow CRISPR-GPT Virtual Lab ALD reactor agent ALS accelerator agentic AI Instrument agents that learn on the job Lightfall Osprey TEM Agent VISIONA-Lab BioDiscoveryAgent ChemAgents Coscientist CRESt k-agents MARS Mobile-robot synthesis lab ORGANA A-Lab GPSS agentic reasoning AI X-ray scientist Experiment Automation Agents (EAA) LLM accelerator tuning NSLS-II multi-beamline AI agents26
Data analysis & simulationAtomAgents Biomni ChemGraph Data Interpreter DS-Agent El Agente Q LLaMP MDCrow APEXA EQSANS-CLI NeuDiff Agent PEAR SasAgentAlphaEvolve Empirical Research Assistance (ERA) Rongzai agentAgent Laboratory AutoResearchClaw Denario Kosmos20
Writing & reviewAgentReview CycleResearcher / CycleReviewer GPT-4 paper feedback study Prism Review Feedback Agent (ICLR 2025) LLM proposal rankingDeepReview Stanford Agentic ReviewerAgentRxiv AI-Researcher data-to-paper The AI Scientist (v1, v2) Zochi13
Total943211083

Reading the table:

  1. From single tools to multi-agent "AI scientists." Systems from 2023–24 wrapped one LLM around domain tools; ChemCrow used 18 chemistry tools [35]. By 2025–26 the usual design is a team of role-specialized agents that critique each other: Co-Scientist runs a generate-critique-refine tournament [4], Virtual Lab has a "PI" agent directing specialists [31], and Kosmos (a preprint) links about 200 agent rollouts per run through a shared "world model", a common record of findings [2].
  2. Flagship claims reached top journals. In 2026 Co-Scientist, Robin, The AI Scientist and OpenScholar appeared in Nature, and Biomni in Science [1, 4, 5, 6, 7].
  3. Companies and frontier labs entered. Google and Google DeepMind cover hypotheses, algorithms and mathematics (Co-Scientist, AlphaEvolve, Aletheia) [3, 4, 36]. FutureHouse spun out Edison Scientific to commercialize its AI Scientist; Kosmos is built from Edison Scientific agents [2, 37]. Intology says a paper written by its Zochi agent was accepted at the ACL 2025 main conference [38].
  4. Autonomy is highest where checking is cheap. Strong L3 and L4 results appear where an automatic check exists (a benchmark score, code evaluator or proof check). A white paper reports that AlphaEvolve found a way to multiply 4×4 complex matrices with 48 scalar multiplications [36]; DeepScientist tested about 1,100 ideas in a month [39]. At Agents4Science 2025, a conference for papers with AI first authors, humans contributed more to hypotheses and design, and accepted papers had more human guidance [40].
  5. Evaluation lags capability (the "verification gap"). Independent checks are rare and mixed. An academic lab, co-authoring with Google, found that Co-Scientist's top hypothesis matched its own unpublished result [41]; an independent analysis disputed A-Lab's novelty claims [42]. LLM ideas that experts rated more novel than human ideas lost that edge once carried out [43, 44]. About one in five Kosmos statements was not judged accurate [2]. A 2026 survey (preprint) found that 38% of 24 runnable systems report any novelty verification [24]; AstaBench concludes that AI remains far from solving research assistance [28].
  6. AI enters peer review; targeted human oversight beats full autonomy. GPT-4 feedback on papers overlapped with human reviews about as much as two reviewers overlap [33]. In a randomized ICLR 2025 trial, 27% of reviewers given LLM feedback updated their reviews [34]. The AI Scientist's authors warn of "taxing overwhelmed review systems" [1]. A preprint finds that targeted human interventions in AutoResearchClaw beat both full autonomy and step-by-step oversight [45].

What is new in 2026

Besides the journal papers and the peer-review trial above, preprints report that Aletheia wrote one paper with no human intervention, solved four open Erdős problems autonomously, and solved 6 of 10 FirstProof problems [3, 46]. MARS pushed closed-loop autonomy in robotic materials labs [30], and OpenAI's Prism put an LLM inside LaTeX manuscript editing [47]. Critiques sharpened: a position paper argues that current agents are co-scientists, not built for autonomous discovery [48]. In the US, the November 2025 Genesis Mission order directs DOE to build a platform that includes AI agents and autonomous experimentation [23].

Caveats. Autonomy labels are our judgment from each paper's own description. Vendor and company claims (Zochi, Prism, Edison Scientific) are not independently verified. Several headline numbers (Kosmos, AlphaEvolve, Aletheia) come from preprints or white papers.

Q2 · What works: evidence versus claims, and what benchmarks show

Answer. Peer-reviewed studies show real but narrow agent successes in labs and literature. Top models now score 95.8% on graduate science questions, above expert level, but open-ended research and XRD analysis stay weak, and no public benchmark tests closed-loop x-ray or neutron beamtime [4, 9, 49].

Established: peer-reviewed or machine-checkable

"Established" means peer-reviewed or machine-checkable, not independently replicated.

ResultEvidenceWhat was actually shownCaveat
Coscientist (chemistry)Peer-reviewed (Nature 2023)A GPT-4 agent planned and ran robotic experiments, including Pd-catalysed cross-coupling optimization [29].No outside replication in our sources.
ChemCrowPeer-reviewed (Nat Mach Intell 2024)An agent with 18 tools carried out syntheses of an insect repellent and three organocatalysts [35].Authors' own demonstration.
A-Lab (contested)Peer-reviewed (Nature 2023); critique (PRX Energy 2024); correction (Nature 2026)Reported 41 of 58 targets made in 17 days [50]. A 2026 Author Correction cut this to 36 confirmed (4 inconclusive by XRD, 1 removed because it was in the training data) and said "novel" meant new to the prediction platform, not to science [51].Outside analysis: no new materials, and automated Rietveld analysis (whole-pattern structure fitting) of powder XRD "not yet reliable" [42]. 2026 Author Correction: the title now says "inorganic", not "novel" [51].
Virtual LabPeer-reviewed (Nature 2025)Agents designed 92 SARS-CoV-2 nanobodies; two bound JN.1 or KP.3 better [31].Tests ran in the authors' own lab.
CRISPR-GPTPeer-reviewed (Nat Biomed Eng 2025)Guided knockout of four genes and activation of two in human cells [52].Authors' own demonstration.
Google Co-ScientistPeer-reviewed + partner-lab tests (Nature 2026; Cell 2025; Adv Sci 2025)Acute myeloid leukaemia drug-repurposing candidates confirmed in vitro [4]. Top hypothesis matched a lab's confirmed but unpublished mechanism (cf-PICIs hijack phage tails) [41, 53]. Two suggested drugs were anti-fibrotic in human liver organoids [54].Partner labs co-authored with Google.
RobinPeer-reviewed (Nature 2026)In vitro tests confirmed ripasudil and KL001 as candidates for dry age-related macular degeneration [5].No animal or clinical evidence yet.
BiomniPeer-reviewed (Science 2026)A general biomedical agent with wet-lab case studies [7].Case studies only.
AlphaEvolveWhite papers (arXiv 2025); machine-checkableProvably correct new algorithms, e.g. 4×4 complex matrix multiplication with 48 scalar multiplications [36]; with outside mathematicians, improved several best-known constructions [55].Not peer-reviewed, but checkable.
OpenScholarPeer-reviewed (Nature 2026)Citation accuracy on par with human experts; GPT-4o hallucinated (invented) citations 78–90% of the time [6].Measures citation accuracy only.
Review feedbackPeer-reviewed (Nat Mach Intell 2026; NEJM AI 2024)In a randomized trial on 20,000+ ICLR 2025 reviews, 27% of reviewers given LLM feedback updated their reviews [34]. GPT-4 comments overlap humans about as much as two humans do (30.85% vs 28.58%) [33].Measures change, not review quality.
AI ScientistPeer-reviewed (Nature 2026)An AI-written paper passed first-round review at a workshop [1].The workshop accepted 70%: a low bar.
X-ray beamline controlPeer-reviewed (Nat Mach Intell 2026)An LLM agent found reference reflections and the orientation matrix (crystal orientation on the diffractometer) on a real SSRL beamline [12].A human relayed each command, unmodified, for safety.
AFM controlPeer-reviewed (Nat Commun 2025)An agent on a real atomic force microscope scored 88.3% on documentation tasks but 33.3% on analysis [56].It "sleepwalked" off its instructions.

Claims awaiting validation

Numbers are the claimants' own.

ClaimSource typeWhat would validate it
Kosmos (Edison): independent scientists judged 79.4% of statements accurate, only 57.9% of synthesis statements; claims seven discoveries [2].Preprint (arXiv 2025)Peer review, outside replication, and false-positive rates across all runs.
Zochi (Intology): says its AI-written paper was accepted at ACL 2025 [57].Company report (GitHub)An audited record of human versus AI work.
Periodic Labs Neon: 55.3% on an internal test of 134 multiphase XRD patterns, graded by an LLM judge (a model grading answers); claimed to beat frontier models [58].Company blog (2026)Release the test set, or score against expert Rietveld refinements of public data.
Lila Sciences: an AI-guided loop screened 2,942 catalysts; InMnPdOx stayed below 0.5 V overpotential for 1,000 h in acid [59].Preprint (arXiv, 2026-09)Independent synthesis and durability tests.
OpenAI GPT-5 cases: four new math results, checked only by the human co-authors [60].Preprint (arXiv 2025)Refereed publication.
OpenAI Navier–Stokes (2026-09): Lean (proof-checker) certificates claim finite-time blowup (Clay alternatives C/D) [61].Code release (GitHub, 2026)Experts confirm the formal statements match the Clay problem, then refereed publication.

What the benchmarks show

"—" means no human baseline was reported.

BenchmarkMeasuresBest reported (system, date)Human/expert baseline
GPQA Diamond [8]Graduate-level science multiple choice95.8% (GPT-6 Astra, 2026-09, per the Epoch AI benchmarking hub) [9]Experts 65% on the full question pool (74% discounting clear mistakes); the paper puts expert accuracy on Diamond between 65% and 81%
HLE (Humanity's Last Exam) [62]Expert-written closed questions54.8% (GPT 6 Astra, 2026-09-09, per the Scale Labs leaderboard) [63]None
FrontierScience [64]Olympiad / research tasks77% / 25% (GPT-5.2, 2026-01)—
CritPt [65]Unpublished physics research problems32.3% (GPT-5.6 Sol, 2026-07, per the Epoch AI benchmarking hub) [9]—
ChemBench [66]Chemistry questionsBest models beat the best chemist surveyed (2025)Surveyed chemists
MaCBench [49]Chemistry/materials images (XRD, AFM)XRD intensity ranking 0.28 (2025)—
LAB-Bench [67]Practical biologyClaude 3.5 Sonnet (2024-07)Experts clearly ahead
LabSafety Bench [68]Lab hazards<70% hazard identification (2026)—
SciCode [69]Research code10.8% main problems (official leaderboard) [70]; 66.9% subproblems (Claude Opus 5.5, 2026-09, per Artificial Analysis) [71]—
ScienceAgentBench [72]Code for data-driven discovery42.2% (o1-preview, 3 tries, 2024-10)—
CORE-Bench Hard [73]Reproduce results from code95.5% with manual validation (77.8% automated), declared "solved" (Opus 4.5 + Claude Code, per the HAL leaderboard) [74]—
BixBench [75]Bioinformatics17% (2025-02)—
DiscoveryWorld [76]Simulated discovery≤18% completion (2024)Human scientists (MSc or PhD) 66%
MLE-bench [77]Kaggle machine-learning competitionsMedal in 64.4% (Famou-Agent 2.0, 2026-02, per the MLE-bench leaderboard) [78]Kaggle leaderboards
MLAgentBench [79]ML experiments37.5% (Claude 3 Opus, 2024)—
RE-Bench [80]ML research and developmentAgents 4× the expert score with a 2 h budget (2024-11)Experts 2× the agent score with 32 h
PaperBench [10]Replicate ICML papers21.0% (Claude 3.5 Sonnet, 2025-04)ML PhDs 41.4% (subset)
EXP-Bench [11]Full AI experiments0.5% (2025-05)—
SciReplicate-Bench [81]Implement algorithms from papers39% (2025)—
ReplicationBench [82]Astrophysics paper replication~20% (Claude Sonnet 4.5, 2025-10)—
AstaBench [28]Research assistance53.0% (Asta v0, 2025-08) [83]—
CURIE [84]Long-context science32% (2025-03)—
FIRE-Bench [85]Rediscover ML findings<50 F1 (2026)—
Collider-Bench [86]Reproduce LHC analysesNone reliably beats the baseline (2026-05)Physicist-in-the-loop
AFMBench [56]Real AFM hardware33.3% analysis (GPT-4o, 2025)—
APEXA-Bench [22]Synchrotron data reduction (58 tasks)Not yet scored (2026-09)—

Reading the benchmarks

  • Closed science exams no longer separate models. Experts score about 65–81% on GPQA Diamond [8], against a best reported 95.8% [9]. HLE's answer key also has problems: by FutureHouse's estimate about 29% of its text-only chemistry and biology answers conflict with the literature, and HLE's own expert re-review found about 18% of a subset problematic. New HLE scores also carry contamination flags (the questions may be in training data) [63, 87].
  • Well-specified computing tasks are nearly solved (CORE-Bench Hard, MLE-bench) [74, 78]; humans still lead RE-Bench at long time budgets [80].
  • Open-ended research is weak: 25% on FrontierScience-Research, about 20% on ReplicationBench, 0.5% on EXP-Bench [11, 64, 82].
  • Lab-facing skills lag: experts lead LAB-Bench, BixBench tops out at 17%, and no model exceeds 70% on hazard identification [67, 68, 75].
  • X-ray and neutron relevance is thin. On XRD, models find the highest peak (0.74 accuracy) but rank intensities at only 0.28; shown crystal structures, they assign space groups at only 0.45 [49]. AFMBench is a rare real-instrument benchmark [56]. In APEXA, a frontier model fabricated a calibration report for commands that never ran [22]. No public benchmark covers closed-loop x-ray or neutron beamtime (diffraction, spectroscopy, imaging); the SSRL work is a demonstration, not a benchmark [12].
  • Read scores with care. Many rest on LLM judges or private test sets, and saturated (near-ceiling) benchmarks still have construct-validity problems: they may not measure the skill they name [10, 58, 88].

Q3 · Infrastructure: models, frameworks, protocols and the path to HPC

Answer. Science agents are built from interchangeable parts: a vendor or center-hosted model, a general agent framework, the Model Context Protocol (MCP) for tools, and Globus Compute or Parsl for HPC. A facility's main job is therefore to expose the services it already runs (Bluesky/EPICS, Tiled, IRI APIs) as guarded, logged tools [17, 19, 22, 89].

Models

Frontier models such as GPT-5 [90], called through vendor APIs, are still the default engine; Co-Scientist, for example, is built on Gemini, and Biomni's setup expects Claude [4, 91]. But frameworks can usually swap them: ChemGraph accepts OpenAI, Anthropic, Google, Argonne's Argo gateway or a local Ollama server [89].

Open-weight models, whose weights can be downloaded and run locally, make self-hosting practical. gpt-oss-120b is Apache-2.0 licensed and tuned for tool use [92]. OLCF's inference service offers it [18]. Science-trained models also compete: ether0, a 24B chemistry reasoning model trained by reinforcement learning on 640,730 problems, beat frontier models and experts on molecular design [93].

Where the model runs is often a data-handling decision. ALCF built FIRST for private, secure inference; it lets researchers generate "billions of tokens daily on-premises" without commercial cloud [17]. DOE's National Nuclear Security Administration (NNSA) runs frontier AI models on the classified network of its Venado supercomputer at Los Alamos [94].

Agent frameworks and science toolkits

Frameworks typically build on ReAct, the loop in which the model reasons, calls a tool (acts), reads the result (observes) and repeats [95]. General orchestrators include LangGraph [96]; AutoGen is now in maintenance mode, with new users pointed to Microsoft Agent Framework [97]. Vendor SDKs package the same loop; the Claude Agent SDK, for example, adds hooks, subagents and permission controls [98].

A science toolkit is usually a thin layer over one of these orchestrators plus a registry of domain tools. ChemGraph combines LangGraph with ASE, RDKit and MCP, and runs jobs through Parsl or Globus Compute on Aurora and Polaris [89]. Academy runs stateful, asynchronous agents across HPC, experimental facilities and data repositories [99]. The value is in the tools, not the loop.

Tool and agent protocols

The Model Context Protocol (MCP) is an open standard that lets an agent call tools and read data through one common interface. A2A (Agent2Agent) links agents to each other; its maintainers call MCP the "vertical" agent-to-tool layer and A2A the "horizontal" agent-to-agent layer [100]. Both now sit in the Linux Foundation's Agentic AI Foundation: MCP since 9 Dec 2025, A2A (v1.0) since Aug 2026 [16, 100].

MCP is still changing. Its 2026-07-28 revision made the protocol stateless, moved long-running "tasks" into an official extension, and deprecated Roots, Sampling and Logging [101]. A lighter option is a "skill": a folder with a SKILL.md instruction file that the agent loads only when needed [102]. ALS staff write skills for the Lightfall beamline-control platform, and at the SNS one skill document lets a Slack bot run EQSANS reductions [103, 104].

The tool layer is also an attack surface. In tool poisoning, hidden instructions in a tool's description redirect the agent. MCPTox tested 45 live MCP servers: o1-mini's attack success rate was 72.8%, and no model refused more than 3% of attacks [105]. On HPC systems, a "hijacked authorized agent" acts with a user's valid credentials [106].

Connecting to workflows, HPC and instruments

One request passes through six layers, mostly existing DOE software:

  1. Model. A vendor API or a center-hosted open model. ALCF (FIRST, with Globus Auth) and OLCF (S3M tokens) serve models with vLLM, an open-source serving engine, behind OpenAI-compatible endpoints. These accept OpenAI-style requests, so code switches by changing an address and key [17, 18].
  2. Agent framework. An agent built with LangGraph, a vendor SDK or Academy reasons and picks the next tool [95, 99].
  3. MCP server or skill. The call reaches a thin wrapper around an existing service, such as Globus Labs' MCP servers for Globus Transfer, Compute and Search and for facility status APIs [19].
  4. Execution fabric. Globus Compute (formerly funcX) or Parsl runs the work on HPC [107, 108]. On Aurora, a gpt-oss-120b agent swarm had a planner split work among executors that shared one Parsl-backed MCP server [109]. The Integrated Research Infrastructure (IRI) Facility API standardizes status, account, compute/jobs, filesystem, storage and task endpoints, with instances at NERSC, ALCF and ESnet [110]; SLAC hosts a fork [111].
  5. Instrument and data. Bluesky and its Ophyd device layer, EPICS, and the Tiled data service are the plug points [112, 113, 114]. At SSRL, an agent with MCP tools wrote SPEC commands, relayed unmodified by a human, to align a single crystal [12]. At NSLS-II, VISION ran a voice-controlled scattering experiment by sending generated code to Bluesky [115]. ORNL agents steered an additive-manufacturing workflow linking OLCF with a simulated version of its Manufacturing Demonstration Facility [116].
  6. Provenance. Prompts, responses and decisions are logged with the workflow record, as PROV-AGENT does by extending W3C PROV [117].

Sandboxing and permissions. Biomni runs LLM-generated code with full system privileges by default [91]. NERSC recommends sandboxes and workspace-write mode (the agent may change files only inside its working directory) for coding agents, and holds users responsible for agent actions [118]. Facility projects put the limits in code, not prompts. APEXA's deterministic guard allowed 0/200 adversarial motor violations, against 15/200 for a safety prompt, and caught a fabricated calibration report [22]. NeuDiff, at the SNS TOPAZ instrument, uses allowlisted tools and fail-closed gates [15]. EnvTrace checks agent-written control code against a beamline digital twin before it runs [119].

Stack at a glance

LayerCommon choicesWhat facilities should watch
ModelsGPT-5 and other vendor APIs; self-hosted gpt-oss, DeepSeek-R1; science-tuned ether0 [92, 93, 120]Center-hosted OpenAI-compatible endpoints [17, 18]; data-sensitivity rules [94]
FrameworksLangGraph, vendor SDKs, Academy, ChemGraph [96, 98, 99]AutoGen churn [97]; code-execution privileges [91]
ProtocolsMCP, A2A, skills [16, 100, 102]Stateless MCP and tasks extension [101]; tool poisoning [105]
Workflow and HPCGlobus Compute, Parsl, IRI Facility API [107, 108, 110]MCP-wrapped facility APIs [19]; agents acting under user credentials [106]
Instruments and dataBluesky/Ophyd, EPICS, SPEC, Tiled [12, 112, 113, 114]Guards in code, not prompts [15, 22]; provenance [117]

Implications for facility deployments

Q4 · Facilities: agents at DOE x-ray and neutron user facilities and national labs

Answer. LLM agents at DOE facilities have moved from documentation assistants and first instrument-control prototypes (2023–24) to supervised control of real beamlines, accelerators and microscopes, peer-reviewed in 2025–26, but steering is still mostly single-campaign demonstrations; deployed systems are mainly knowledge assistants, data-reduction agents and one accelerator control-room framework (Osprey, at the ALS) [12, 14, 22, 115, 121].

Beamline, instrument and accelerator steering

This subsection covers agents that steer beamlines, instruments or accelerators, grouped by lab. Most agents that reduce or analyze data, or answer staff and user questions, are listed under Facility data analysis and knowledge assistants; that is where the SNS agents, and APEXA, APS-RAG and PEAR at the APS, appear.

SLAC

BNL (NSLS-II, CFN)

Argonne (APS, CNM)

LBNL (ALS, Molecular Foundry, A-Lab)

ORNL (CNMS, SNS)

Other accelerators and international facilities

Pre-LLM autonomous-experiment precursors

These closed loops use Bayesian or machine-learning methods, not language models. Most predate the agents above.

  1. Gaussian-process autonomous x-ray scattering, the gpCAM lineage [138].
  2. CAMEO: closed-loop Bayesian active learning at SSRL found a new Ge-Sb-Te phase-change material [139].
  3. ANDiE: autonomous neutron diffraction determined Néel temperatures 5-fold more efficiently [140].
  4. Argonne FAST: autonomous scanning microscopy needed under 25% of the sample [141].
  5. Bluesky-hosted agents coordinated multi-beamline measurements at NSLS-II [142, 143].
  6. LCLS physics-guided Bayesian optimization of a 12-parameter split-and-delay optic, tested in simulation [122]; ORNL edge-to-exascale steering at SNS, still a proof of concept [144].

Facility data analysis and knowledge assistants

DOE programs and policy

Status, integration points and facility constraints

Facility agent systems

SystemFacility/labTechniqueWhat the agent controls or analyzesStatusEvidence
AI X-ray scientistSLAC SSRLCrystal diffractionDiffractometer motors, detectorDemonstratedPeer-reviewed [12]
VISIONBNL NSLS-IIX-ray scatteringBluesky motors, detectorDemonstratedPeer-reviewed [115]
Learn-on-the-job agentsArgonne APS/CNMNanoprobe, roboticsMulti-task workflowsDemonstratedPeer-reviewed [126]
EAAArgonne APSX-ray imagingFocusing, feature searchDemonstratedPeer-reviewed [127]
ALS machine-physics agentLBNL ALSAcceleratorPlans multistage machine-physics experiments; archive retrieval, control-channel resolutionDemonstratedPeer-reviewed [13]
Osprey frameworkLBNL ALSAccelerator control roomPlan-first orchestration of control channels, with human review before hardware actionsDeployedPeer-reviewed [14]
LightfallLBNL ALSCoherent scatteringDevices, GP scansIn testingPreprint [103]
TEM AgentLBNL Molecular FoundryTEMMicroscope, HPCDemonstratedPeer-reviewed [130]
NeuDiff AgentORNL SNSNeutron crystallographyReduction to CIFDemonstratedPeer-reviewed [15]
SasAgent / EQSANS-CLIORNL; SNS EQ-SANSSANSFitting, reductionDemonstratedPeer-reviewed; preprint [104, 145]
APEXAArgonne APSDiffractionCalibration, integrationDeployedPreprint [22]
APS-RAGArgonne APSOperationsLogbooks, control dataDeployedPreprint [121]
ChatEEDSLACAccelerator operationsE-log retrievalDemonstratedPeer-reviewed workshop [124]
RongzaiCSNSNeutron powder diffractionRietveld refinementDeployedPreprint [136]
Genesis platformDOEAllAutonomous experimentationIn developmentExecutive order; White House release [23, 153]

Integration points

Facility-specific constraints

Q5 · Failure modes and risks

Answer. Documented failures include irreproducible runs, fabricated references and results, benchmarks that overstate real performance, unsafe or hijacked tool actions, and unclear credit; the mitigations in use are full logging, independent checks, sandboxed tools and disclosure rules [22, 174, 175, 176].

Reproducibility

For facilities: Pin model versions. Log prompts and tool calls with the data provenance; the NeuDiff Agent at SNS keeps full provenance from raw data to a validated CIF [15]. Have an expert check any structure refinement made by an agent.

Fabricated citations or results

For facilities: Check references automatically. Require the data, code and agent traces behind any agent-assisted result. At the APS, a frontier model fabricated a calibration report for commands that never ran, so an agent's report must be checked against what actually executed [22].

Evaluation gaps

For facilities: Test agents on held-out, in-house beamtime data graded by experts. Do not rely on a single LLM judge.

Safety and dual use

For facilities: Sandbox agents and allow only listed tools; NeuDiff at SNS uses allowlisted tools and fail-closed gates [15]. Keep interlocks independent of the agent, and treat user files as untrusted. At SSRL a human relayed each agent command unmodified [12]; at the ALS, Osprey requires human review of a complete plan before any hardware action [14]. At the APS, a deterministic guard allowed 0 of 200 motor-control violations against a simulated IOC, versus 15 of 200 with a safety prompt alone [22].

Credit and authorship

For facilities: Require AI-use disclosure in beamtime proposals, and do not allow external AI tools in proposal review panels.

Risk register

RiskExampleSignal from evidenceMitigation seen in the literature
Irreproducible resultsModel version driftUp to 15% swing between runs [177]; 84% to 51% across versions [178]Pin versions; shared logs of every run [179]
Overclaimed discoveryA-Lab materialsAbout two-thirds likely known phases [42]Expert crystallography review [42]
Fake referencesNeurIPS 2025 papersAbout 1 in 20 papers [20]Automated reference checks [20]; desk rejection [182]
Fabricated execution reportBeamline calibration agentReport for commands that never ran [22]Execution-integrity enforcement [22]
Reward hacking or selective reportingResearch-pipeline tasks30.5% unprompted [21]Review traces and code; multiverse checks [183, 184]
Benchmark–lab mismatchScience data tasksBest agents solve 32.4–42.2% [72]Benchmark checklist; in-house tasks [186]
Biased LLM judgesAI reviewersMean scores 2.30–4.23 by model [40]Final human review [40]
Prompt injectionHidden review prompts18 manuscripts [192]Treat inputs as untrusted [190]; misconduct rules [198]
Unsafe tool actionsHigh-stakes tools; motor moves23.9% failures [175]; 15 of 200 violations with a safety prompt alone [22]Sandboxing; deterministic guards; human plan review [14, 22]
Physical lab hazardsRobot–human collision16 hazards in one lab [194]; no model above 70% on hazard identification [68]ISO-based risk assessment [194]; human oversight [195]
Dual useToxin design40,000 molecules in under 6 h [193]Controlled-substance checks [35]
Undisclosed AI in reviewLLM-written reviews6.5–16.9% of text [200]Disclosure; canary prompts; bans [199, 203]

Q6 · Discussion: open questions for the session

Answer. The open questions for facilities are no longer whether agents can touch instruments (they already align crystals, run accelerator experiments and reduce neutron data) but where humans must stay in the loop, how agent results are verified and credited, and what shared infrastructure DOE facilities should build [12, 13, 15, 24].

Each prompt below states the evidence it rests on and links to the section where that evidence is discussed. The prompts are ordered from beamline-level to program-level questions.

Suggested running order for a 60–90 minute session. Three blocks of about 20 minutes each: at the beamline (prompts 1, 2 and 4), shared infrastructure and data (prompts 3, 5 and 6), and programs and people (prompts 7, 8 and 9). If time is short, use prompts 1, 2, 4, 6 and 8. Each block works best if it ends with one concrete output, for example a list of beamline actions that always need human approval, one candidate joint benchmark, or a first joint deliverable.

  1. Where should the human stay in the loop at a beamline? In its first real-beamline trial, the SSRL agent had a person relay every command [12], and the ALS accelerator framework requires human review of a complete plan before any hardware action [14] (Q4 steering). Outside facilities, targeted human interventions beat both full autonomy and step-by-step oversight [45], and accepted Agents4Science papers had more human guidance [40] (Q1 trends). For discussion: which actions (moving motors, changing sample environment, discarding data, ending a run) need explicit approval, and could facilities agree on a shared autonomy scale for beamtime, like the L1–L4 scale used here?
  2. Should safety be enforced in the tool layer rather than in the prompt? At APS, a deterministic execution guard allowed 0 of 200 adversarial motor-control violations, against 15 of 200 with a safety prompt alone [22]; NeuDiff at SNS uses allowlisted tools and fail-closed gates [15] (Q3 workflow and HPC). Tool poisoning succeeds against many models on real MCP servers [105] (Q3 protocols), and no model tested scored above 70% on lab hazard identification [68] (Q5 safety). For discussion: should DOE facilities define a common "agent-safe" instrument interface, with allowlisted tools, simulation-first execution and interlocks that do not depend on the agent?
  3. What evaluation would convince facility scientists? Closed-form science benchmarks are saturated [8, 9], while models rank XRD peak intensities poorly [49] and no public benchmark covers closed-loop x-ray or neutron beamtime (Q2 benchmarks). A beamline digital twin already scores LLM-written control code [119] (Q4 steering), and benchmark design flaws can badly mis-estimate agent performance [186] (Q5 evaluation gaps). For discussion: could SLAC, ORNL, BNL and university partners build a shared, held-out benchmark from archived beamtime (diffraction, SAXS/SANS, spectroscopy, imaging) with expert grading, plus digital twins for control tasks?
  4. How are agent-produced analyses verified before they reach a paper? An outside analysis disputed the A-Lab's claims of new materials and called automated Rietveld analysis of powder XRD not yet reliable [42] (Q2 established results). A frontier model fabricated a calibration report for commands that never ran [22], and research agents reward-hacked without being asked in open-ended tasks [21] (Q5 fabrication). A 2026 survey calls verification the field's central bottleneck [24]. For discussion: what provenance should be required for agent-assisted results from facility data (prompts, tool calls, code, model version, raw-data hashes), and who checks it: the user, the beamline scientist, or the journal? Provenance tooling for agents already exists [117].
  5. Which models may touch user data, and where do they run? Center-hosted, OpenAI-compatible inference on open-weight models already runs at ALCF and OLCF [17, 18] (Q3 models). VISION keeps its beamline models local [115], NERSC tells users not to place credentials in external AI services [118], and the Genesis Mission order requires security and export-control compliance [23] (Q3 workflow and HPC; Q4 constraints). For discussion: what is the policy for proprietary or sensitive user data: commercial APIs, center-hosted open-weight models, or both? Who pays for inference during beamtime?
  6. What should facilities standardize on: MCP servers, skills, or the APIs they already have? MCP moved to a Linux Foundation body in December 2025 [16], and its July 2026 revision changed core behavior [101] (Q3 protocols). MCP servers already wrap Globus and facility services [19], the IRI Facility API standardizes job and data endpoints [110], and Bluesky is shared across facilities [169] (Q3 workflow and HPC; Q4 status). Anthropic's DOE partnership describes MCP servers for instruments [165]. For discussion: should facilities jointly publish and maintain agent interfaces for Bluesky, EPICS, Tiled and IRI, rather than each beamline building its own? Who owns them as the protocols change?
  7. How should agent use be disclosed and credited in proposals, papers and review? Journals and ICMJE bar AI authorship and require disclosure [176, 196]. ICML 2026 found 795 reviews that broke its no-LLM policy and desk-rejected 497 papers from the reviewers involved [199], and NIH does not count applications substantially developed by AI as original [204] (Q5 credit). At SNS, LLM proposal rankings correlated with human rankings at ρ≈0.2–0.8, at over 100× lower cost [149] (Q4 analysis). For discussion: should beamtime proposals and facility-data papers declare agent use? Should LLMs help triage proposals, and if so, with what disclosure to proposers?
  8. Where will Genesis Mission funding meet facility operations? DOE's autonomous-laboratories challenge names user facilities as its nucleus [151]. The July 2026 announcement lists more than $5 billion in federal commitments, and DOE selected 278 projects [153]; the American Science Cloud aims to unify access to the 17 DOE labs and 28 user facilities that ESnet already links [158] (Q4 programs). Most facility agent systems are still single-campaign demonstrations (Q4 status). For discussion: what is a realistic first joint deliverable, for example one agent-steered experiment that spans a light or neutron source and an HPC center, and how do teams avoid one-off demonstrations that are never maintained?
  9. How do users and staff keep their expertise as agents take over routine steps? A facility chatbot already serves visiting EQ-SANS users [148] (Q4 analysis), and AI can create "illusions of understanding" and narrow the range of methods scientists use [189] (Q5 evaluation gaps). LLM-generated research ideas lost their apparent edge once experts carried them out [44] (Q1 trends). For discussion: which skills (alignment, reduction, refinement) must users still show before an agent acts for them, and how should university partners train students for agent-assisted beamtime?

Appendix

System catalog

All 83 systems placed in the Q1 taxonomy, sorted by stage and then year. Autonomy uses the L1–L4 scale in the glossary. Evidence type describes the cited source: peer-reviewed, peer-reviewed with independent validation, preprint, workshop, product or vendor report, or lab or project page. The one-line description under each name summarizes what the cited source reports. Facility and national-lab systems (Q4) are included alongside general-science systems.

SystemOrganizationYearStageAutonomyEvidence typeReference
CALMSRetrieval- and tool-augmented LLM answering facility-documentation questions and conversationally operating instruments.Argonne (APS, CNM, ALCF)2023Literature & ideationspans: Experiment design & self-driving labsL2 Tool-using agentPeer-reviewed[125]
ESACRetrieval-augmented chatbot serving as an interactive instrument reference for visiting EQ-SANS users.ORNL (SNS EQ-SANS)2024Literature & ideationspans: Experiment design & self-driving labsL1 AssistantPeer-reviewed[148]
GAIAReAct-style LLM assistant combining logbook and document retrieval with control-script generation for accelerator operations.DESY2024Literature & ideationspans: Experiment design & self-driving labsL2 Tool-using agentPreprint[206]
Gemini Deep ResearchConsumer deep-research agent that browses the web on the user's behalf and returns a cited, exportable report; launched December 2024.Google2024Literature & ideationL2 Tool-using agentProduct/vendor report[207]
LLM ideation agent vs. 100+ NLP researchersBlinded human study: agent ideas judged more novel (p<0.05) than experts' ideas but slightly weaker on feasibility.Stanford University2024Literature & ideationspans: Hypothesis generationL2 Tool-using agentPeer-reviewed[43]
OpenScholarRetrieval-augmented LM over 45M open-access papers; citation accuracy on par with experts; experts preferred OpenScholar-8B answers over expert-written ones 51% of the time.University of Washington / Ai22024Literature & ideationL2 Tool-using agentPeer-reviewed[6]
PaperQA2Literature-search agent that matched or exceeded subject experts on three literature tasks; found 2.34 contradictions per biology paper, 70% expert-validated.FutureHouse2024Literature & ideationL2 Tool-using agentPreprint[27]
ResearchAgentProposes problems, methods and experiment designs from a core paper plus academic graph, iteratively refined by LLM reviewing agents.KAIST / Microsoft Research2024Literature & ideationspans: Hypothesis generationL2 Tool-using agentPeer-reviewed[208]
STORMResearches a topic via multi-perspective question asking, then writes cited Wikipedia-like articles; 25% absolute gain in judged organization over a retrieval baseline.Stanford University2024Literature & ideationspans: Writing & reviewL2 Tool-using agentPeer-reviewed[209]
Ai2 Asta agents + AstaBenchScience-agent ecosystem with a 2,400+ problem benchmark; evaluation of 57 agents concludes AI remains far from solving science research assistance.Ai22025Literature & ideationspans: Data analysis & simulationL2 Tool-using agentPeer-reviewed[28]
Ai2 Scholar QAFree, open-source literature-synthesis app producing organized reports with attribution; outperformed competing systems on a recent scientific QA benchmark.Ai22025Literature & ideationL2 Tool-using agentPeer-reviewed[210]
ChatEEDAgentic retrieval assistant over electronic logbooks and wikis for accelerator operators.SLAC (accelerator operations)2025Literature & ideationL1 AssistantPeer-reviewed[124]
FutureHouse Platform (Crow, Falcon, Owl, Phoenix)Web platform of literature agents plus a ChemCrow-based chemistry planner; vendor claims better precision than PhD-level researchers on literature search.FutureHouse2025Literature & ideationspans: Experiment design & self-driving labsL2 Tool-using agentProduct/vendor report[211]
APS-RAGDeployed agentic GraphRAG over logbooks, documents, wikis and control data for APS staff; 70.3% vs 63.8% strict recall.Argonne (APS)2026Literature & ideationL2 Tool-using agentPreprint[121]
FunSearchLLM paired with an automated evaluator in an evolutionary program search; found cap-set constructions beyond the best known and better bin-packing heuristics.Google DeepMind2023Hypothesis generationspans: Data analysis & simulationL3 Closed-loopPeer-reviewed[212]
SciMONRetrieves literature inspirations and iteratively revises generated research ideas to increase novelty relative to prior work.UIUC / Ai22023Hypothesis generationspans: Literature & ideationL2 Tool-using agentPeer-reviewed[213]
MOOSE-ChemInspiration-retrieval, composition and ranking agents rediscovered many hypotheses from 51 high-impact 2024 chemistry papers using a pre-2024 LLM.Nanyang Technological University / Shanghai AI Laboratory2024Hypothesis generationspans: Literature & ideationL2 Tool-using agentPeer-reviewed[214]
SciAgentsMulti-agent LLMs reason over ontological knowledge graphs to generate, critique and refine research hypotheses for bio-inspired materials.MIT2024Hypothesis generationspans: Literature & ideationL2 Tool-using agentPeer-reviewed[215]
Co-ScientistGemini multi-agent generate-critique-refine tournament; AML drug-repurposing candidates validated in vitro; separately matched an external lab's unpublished gene-transfer mechanism.Google2025Hypothesis generationspans: Literature & ideationL2 Tool-using agentPeer-reviewed + independent validation[4]
DeepScientistMonth-long hypothesize-verify-analyze loop; ~5,000 ideas, ~1,100 tested, beat human state of the art on three AI tasks (183.7%, 1.9%, 7.9%).Westlake University2025Hypothesis generationspans: Data analysis & simulationL3 Closed-loopPeer-reviewed[39]
InternAgent (NovelSeek)Closed-loop hypothesis-to-verification framework across 12 tasks; e.g., reaction-yield prediction improved from 27.6% to 35.4% in 12 hours.Shanghai AI Laboratory2025Hypothesis generationspans: Data analysis & simulationL3 Closed-loopPreprint[216]
POPPERAgents design and run falsification tests with Type-I error control; matched human scientists validating biological hypotheses in one-tenth the time.Stanford University2025Hypothesis generationspans: Data analysis & simulationL2 Tool-using agentPeer-reviewed[217]
RobinLab-in-the-loop multi-agent system; proposed ripasudil for dry macular degeneration and analyzed follow-up RNA-seq; humans executed the wet-lab experiments.FutureHouse2025Hypothesis generationspans: Literature & ideation; Experiment design & self-driving labs; Data analysis & simulationL3 Closed-loopPeer-reviewed[5]
AletheiaGemini Deep Think math agent that generates, verifies and revises proofs; one paper produced without human intervention; four open Erdős problems solved autonomously.Google DeepMind2026Hypothesis generationspans: Writing & reviewL4 End-to-endPreprint[3]
A-LabRobotic solid-state synthesis with computation, literature-trained recipe models and active learning; reported 41 of 58 targets in 17 days; 2026 correction confirms 36.Lawrence Berkeley National Laboratory2023Experiment design & self-driving labsspans: Data analysis & simulationL3 Closed-loopPeer-reviewed[50]
ChemCrowGPT-4 agent with 18 expert-designed chemistry tools; autonomously planned and executed syntheses of an insect repellent and three organocatalysts.EPFL / University of Rochester / IBM Research2023Experiment design & self-driving labsspans: Data analysis & simulationL2 Tool-using agentPeer-reviewed[35]
CoscientistGPT-4 agent with search, code and lab-automation tools that designed, planned and ran experiments, including optimizing palladium-catalysed cross-couplings.Carnegie Mellon University2023Experiment design & self-driving labsspans: Literature & ideationL3 Closed-loopPeer-reviewed[29]
BioDiscoveryAgentIterative design of genetic perturbation screens; 21% better hit prediction than Bayesian-optimization baselines across six datasets.Stanford University / Arc Institute2024Experiment design & self-driving labsspans: Hypothesis generationL3 Closed-loopPeer-reviewed[218]
ChemAgentsOn-board Llama-3.1-70B hierarchical agents (literature, design, computation, robot) ran seven robotic chemistry tasks with minimal human intervention.University of Science and Technology of China2024Experiment design & self-driving labsspans: Literature & ideation; Data analysis & simulationL3 Closed-loopPeer-reviewed[219]
CRISPR-GPTGene-editing design and analysis agent; guided knockout of four genes (Cas12a) and activation of two genes (dCas9) in human cell lines.Stanford University / Princeton University2024Experiment design & self-driving labsspans: Data analysis & simulationL2 Tool-using agentPeer-reviewed[52]
k-agentsLLM agents encoding lab knowledge ran a superconducting quantum processor for hours, producing entangled states at the level of human scientists.University of Oxford / University of Toronto2024Experiment design & self-driving labsspans: Data analysis & simulationL3 Closed-loopPeer-reviewed[220]
LLM accelerator tuningLLM tunes an accelerator subsystem from an operator's natural-language prompt, benchmarked against Bayesian optimization and RL-trained optimizers.DESY2024Experiment design & self-driving labsL3 Closed-loopPeer-reviewed[135]
LLM-assisted scanning probe microscopyGPT-4 with instrument APIs converts experimental workflow ideas into microscope code; limited for in-depth experimental design.ORNL (CNMS)2024Experiment design & self-driving labsspans: Data analysis & simulationL1 AssistantPeer-reviewed[132]
Mobile-robot synthesis labMobile robots operate shared synthesis, UPLC-MS and benchtop NMR equipment; a heuristic decision-maker picks and re-checks hits in exploratory chemistry.University of Liverpool2024Experiment design & self-driving labsspans: Data analysis & simulationL3 Closed-loopPeer-reviewed[221]
ORGANALLM-interfaced robotic assistant for solubility, pH, recrystallization and electrochemistry; user study reported 80.3% average time saved.University of Toronto2024Experiment design & self-driving labsL3 Closed-loopPeer-reviewed[222]
Virtual LabLLM 'PI' agent leads specialist agents to build a nanobody design pipeline; 92 designed, two with improved binding to recent SARS-CoV-2 variants.Stanford University / Chan Zuckerberg Biohub2024Experiment design & self-driving labsspans: Hypothesis generation; Data analysis & simulationL2 Tool-using agentPeer-reviewed[31]
VISIONModular LLM assistant turning voice/text into Bluesky code at NSLS-II 11-BM; first voice-controlled X-ray scattering experiment.BNL (CFN, NSLS-II)2024Experiment design & self-driving labsspans: Data analysis & simulationL2 Tool-using agentPeer-reviewed[115]
AI X-ray scientistLLM agent with MCP tools and SPEC commands aligned a single crystal on SSRL BL17-2; developed in a virtual diffractometer; human relayed commands.SLAC (SSRL)2025Experiment design & self-driving labsspans: Data analysis & simulationL3 Closed-loopPeer-reviewed[12]
ALS accelerator agentic AIPlan-first agent ran multistage machine-physics experiments on the ALS accelerator; preparation time cut ~100x under operator safety constraints.LBNL (ALS)2025Experiment design & self-driving labsspans: Data analysis & simulationL2 Tool-using agentPeer-reviewed[13]
AutoLabsSelf-correcting multi-agent system turning natural-language goals into liquid-handler protocols; F1 > 0.89 versus expert procedures on multi-plate syntheses.Pacific Northwest National Laboratory2025Experiment design & self-driving labsL2 Tool-using agentPeer-reviewed[223]
CREStMultimodal-model-guided Bayesian optimization with robotics; 900+ chemistries and 3,500 tests in 3 months found a catalyst with 9.3-fold cost-specific gain.MIT2025Experiment design & self-driving labsspans: Data analysis & simulationL3 Closed-loopPeer-reviewed[224]
Instrument agents that learn on the jobHuman-in-the-loop LLM agents orchestrating an X-ray nanoprobe beamline and an autonomous robotic materials station, improving through iterative feedback.Argonne (APS, CNM)2025Experiment design & self-driving labsspans: Data analysis & simulationL2 Tool-using agentPeer-reviewed[126]
NSLS-II multi-beamline AI agentsNon-LLM agents inside Bluesky run on-the-fly reduction, GP modeling and Bayesian-optimized XRD/XAFS mapping across two beamlines.BNL (NSLS-II)2025Experiment design & self-driving labsspans: Data analysis & simulationL3 Closed-loopPreprint[143]
OspreyControl-room agent framework with human plan review before hardware actions; production deployment across hundreds of thousands of ALS control channels.LBNL (ALS)2025Experiment design & self-driving labsL2 Tool-using agentPeer-reviewed[14]
TEM AgentMCP tools let a commercial LLM control TEM subsystems, data management and HPC through text instructions, without extra training.LBNL (Molecular Foundry)2025Experiment design & self-driving labsspans: Data analysis & simulationL2 Tool-using agentPeer-reviewed[130]
A-Lab GPSS agentic reasoningAgentic AI proposed 352 air-free solid-state syntheses of halide spinel conductors; success fraction rose from 1.33% to 5.33%.LBNL (A-Lab)2026Experiment design & self-driving labsspans: Hypothesis generationL3 Closed-loopPreprint[131]
ALD reactor agentLLM agent converts user queries into JSON-encoded atomic layer deposition processes executed on a real reactor through an MCP-compatible interface.Argonne (Applied Materials Division)2026Experiment design & self-driving labsL2 Tool-using agentPeer-reviewed[128]
Experiment Automation Agents (EAA)Vision-language agents with MCP tools automate zone-plate focusing and natural-language feature search at an APS imaging beamline.Argonne (APS)2026Experiment design & self-driving labsspans: Data analysis & simulationL3 Closed-loopPeer-reviewed[127]
LightfallAPI-first beamline control platform with embedded LLM agent and staff-authored skills; in testing at ALS COSMIC-Scattering.LBNL (ALS)2026Experiment design & self-driving labsL2 Tool-using agentPreprint[103]
MARS19 LLM agents and 16 domain tools coordinate robotic synthesis, characterization and analysis for perovskite nanocrystal materials.Shenzhen Institute of Advanced Technology, CAS2026Experiment design & self-driving labsspans: Literature & ideation; Data analysis & simulationL3 Closed-loopPeer-reviewed[30]
AtomAgentsPhysics-aware multimodal multi-agent system that retrieves knowledge, runs atomistic simulations and analyzes results to design alloys.MIT2024Data analysis & simulationspans: Hypothesis generationL2 Tool-using agentPeer-reviewed[225]
Data InterpreterHierarchical task-graph planning with verified code generation for end-to-end data science; InfiAgent-DABench accuracy rose from 75.9% to 94.9%.DeepWisdom / MetaGPT2024Data analysis & simulationL2 Tool-using agentPeer-reviewed[226]
DS-AgentCase-based reasoning over Kaggle solutions to build and train ML models; 100% development-stage success with GPT-4 at about $1.60 per run.Jilin University2024Data analysis & simulationL2 Tool-using agentPeer-reviewed[227]
LLaMPHierarchical ReAct agents query Materials Project data and run atomistic simulations, reducing LLM errors on bulk moduli, band gaps and formation energies.UC Berkeley / Lawrence Berkeley National Laboratory2024Data analysis & simulationspans: Literature & ideationL2 Tool-using agentPeer-reviewed[25]
PEARMultiple LLM agents handle ptychography knowledge retrieval, code generation, parameter recommendation and image reasoning.Argonne (APS)2024Data analysis & simulationL2 Tool-using agentPreprint[147]
Agent LaboratoryTakes a human idea through literature review, experiments and report writing; 84% cheaper than prior autonomous methods; human feedback improved quality.AMD / Johns Hopkins University2025Data analysis & simulationspans: Literature & ideation; Writing & reviewL4 End-to-endPeer-reviewed[228]
AlphaEvolveEvolutionary Gemini coding agent with automated evaluators; found 4x4 complex matrix multiplication using 48 scalar multiplications, first improvement in 56 years.Google DeepMind2025Data analysis & simulationspans: Hypothesis generationL3 Closed-loopPreprint[36]
BiomniGeneral biomedical agent whose tool environment is mined from thousands of papers across 25 domains; composes code workflows and wet-lab protocols.Stanford University2025Data analysis & simulationspans: Experiment design & self-driving labsL2 Tool-using agentPeer-reviewed[7]
ChemGraphAgentic computational-chemistry workflows from ML potentials to DFT; on 13 tasks, multi-agent decomposition let small LLMs match GPT-4o in some cases.Argonne National Laboratory2025Data analysis & simulationL2 Tool-using agentPeer-reviewed[26]
DenarioModular multi-agent research assistant from idea to drafted and reviewed paper; showcased expert-scored AI-generated papers across many disciplines, including astrophysics.Flatiron Institute / University of Cambridge et al.2025Data analysis & simulationspans: Literature & ideation; Hypothesis generation; Writing & reviewL4 End-to-endPreprint[229]
El Agente QHierarchical-memory multi-agent system that writes, submits and debugs quantum-chemistry workflows from prompts; >87% average task success on course exercises.University of Toronto2025Data analysis & simulationL2 Tool-using agentPeer-reviewed[230]
Empirical Research Assistance (ERA)LLM plus tree search writes score-maximizing scientific software; 40 single-cell methods beat top leaderboard entries; 14 COVID-19 models beat the CDC ensemble.Google2025Data analysis & simulationL3 Closed-loopPreprint[231]
KosmosUp to 12-hour runs of data-analysis and literature agents sharing a world model; independent scientists judged 79.4% of report statements accurate.Edison Scientific (FutureHouse spinout)2025Data analysis & simulationspans: Literature & ideation; Hypothesis generation; Writing & reviewL4 End-to-endPreprint[2]
MDCrowAgent using 40+ expert-designed tools to set up, run and analyze molecular dynamics simulations; evaluated on 25 tasks of varying difficulty.FutureHouse / University of Rochester2025Data analysis & simulationL2 Tool-using agentPeer-reviewed[232]
SasAgentCoordinator plus three specialist LLM agents use SasView tools for SLD calculation, synthetic data and small-angle scattering fitting.ORNL (SNS)2025Data analysis & simulationL2 Tool-using agentPeer-reviewed[145]
APEXADeployed 61-tool multi-agent calibration and integration system; guard refuses results not backed by executed tool calls.Argonne (APS)2026Data analysis & simulationL2 Tool-using agentPreprint[22]
AutoResearchClawMulti-agent research pipeline with self-healing execution and seven human-intervention modes; targeted human input beat both full autonomy and step-by-step oversight.AIMING Lab, UNC Chapel Hill et al.2026Data analysis & simulationspans: Hypothesis generation; Writing & reviewL4 End-to-endPreprint[45]
EQSANS-CLIAgent-addressable SANS reduction tool; an external agent with one skill document drives complete reductions through a Slack bot.ORNL (SNS)2026Data analysis & simulationL2 Tool-using agentPreprint[104]
NeuDiff AgentGoverned agent takes single-crystal neutron data to validated CIF with allowlisted tools and fail-closed gates; 4.6-5.0x faster than manual.ORNL (SNS TOPAZ)2026Data analysis & simulationspans: Writing & reviewL2 Tool-using agentPeer-reviewed[15]
Rongzai agentLLM plus GSAS-II agent performs neutron Rietveld refinement from natural-language task to report; deployed at CSNS for external users.CSNS (China)2026Data analysis & simulationspans: Writing & reviewL3 Closed-loopPreprint[136]
GPT-4 paper feedback studyGPT-4 comments on full papers overlapped with human reviewers as much as reviewers overlap each other; 57.4% of 308 researchers found it helpful.Stanford University2023Writing & reviewL1 AssistantPeer-reviewed[33]
AgentReviewLLM-agent simulation of reviewers, authors and area chairs; reviewer biases alone produced 37.1% variation in paper decisions.Georgia Tech / William & Mary et al.2024Writing & reviewL1 AssistantPeer-reviewed[233]
CycleResearcher / CycleReviewerOpen LLMs trained with automated-reviewer feedback draft full papers whose experimental results are fabricated, not run; CycleReviewer cut score error 26.89% versus individual reviewers.Westlake University2024Writing & reviewspans: Hypothesis generationL1 AssistantPeer-reviewed[234]
data-to-paperBackward-traceable pipeline from annotated data to full paper; simple-goal autopilot runs recapitulated published findings in about 80-90% of cases.Technion2024Writing & reviewspans: Hypothesis generation; Data analysis & simulationL4 End-to-endPeer-reviewed[32]
The AI Scientist (v1, v2)Generates ideas, code, experiments, full manuscript and self-review; one paper passed first-round review at an ICLR 2025 workshop (70% acceptance rate).Sakana AI / University of Oxford / UBC2024Writing & reviewspans: Literature & ideation; Hypothesis generation; Data analysis & simulationL4 End-to-endPeer-reviewed[1]
AgentRxivShared preprint server where agent labs upload and build on each other's reports; collaborating labs gained 13.7% relative on MATH-500.Johns Hopkins University / ETH Zurich2025Writing & reviewspans: Literature & ideationL4 End-to-endPreprint[235]
AI-ResearcherPipeline from literature review and hypotheses to implementation and manuscript; introduces Scientist-Bench; authors report near-human paper quality.University of Hong Kong2025Writing & reviewspans: Literature & ideation; Hypothesis generation; Data analysis & simulationL4 End-to-endPeer-reviewed[236]
DeepReviewMulti-stage reviewer with literature retrieval and evidence-based argumentation; DeepReviewer-14B outperforms CycleReviewer-70B while using fewer tokens.Westlake University2025Writing & reviewL2 Tool-using agentPeer-reviewed[237]
LLM proposal rankingPairwise LLM ranking of proposals from three SNS beamlines correlates with human rankings (Spearman 0.2-0.8) at over 100x lower cost.ORNL (SNS)2025Writing & reviewL1 AssistantPeer-reviewed[149]
Review Feedback Agent (ICLR 2025)Randomized trial on 20,000+ ICLR 2025 reviews; 27% of reviewers given LLM feedback updated reviews, incorporating 12,000+ suggestions.Stanford University / ICLR2025Writing & reviewL1 AssistantPeer-reviewed[34]
Stanford Agentic ReviewerFree web reviewer grounding feedback in arXiv searches; Spearman correlation with a human reviewer 0.42 versus 0.41 human-human on ICLR 2025 papers.Stanford University2025Writing & reviewspans: Literature & ideationL2 Tool-using agentLab/project page[238]
ZochiCommercial autonomous research agent; vendor reports a Zochi-generated paper was accepted to the ACL 2025 main conference.Intology2025Writing & reviewspans: Hypothesis generation; Data analysis & simulationL4 End-to-endProduct/vendor report[38]
PrismFree AI-native LaTeX writing and collaboration workspace with an OpenAI model working inside the document context; launched January 2026.OpenAI2026Writing & reviewL1 AssistantProduct/vendor report[47]

Glossary

AI agent
A system in which a language model chooses and calls tools (search, code, databases, simulators, instruments) in a loop, observes the results and decides the next step, rather than answering once.
Autonomy level (L1–L4)
This report's four-step scale. L1 assistant: answers or drafts on request. L2 tool-using agent: plans and runs several tool calls inside one task. L3 closed-loop: repeats plan → act → observe against real instruments, robots, simulations or code over many cycles, with human checkpoints. L4 end-to-end: goes from a goal to experiments or analysis to a written paper or report with little human help.
Self-driving laboratory (SDL)
A laboratory where robots or instruments run experiments chosen by an algorithm that learns from earlier results. Older SDLs use Bayesian optimization; newer ones add a language-model agent for planning and interfaces.
Bayesian optimization / active learning
Methods that choose the next measurement from a statistical model of the results so far, trading off exploration and exploitation. They are used by most pre-LLM autonomous experiments at light and neutron sources.
Retrieval-augmented generation (RAG)
Giving a model documents it retrieves at run time (papers, manuals, logbooks) so that its answers can point to sources.
Deep research
Products and agents that run many web or literature searches, read the results and write a cited report.
Multi-agent system
Several model instances with different roles (for example planner, critic, coder) that pass messages to each other.
Tool calling / function calling
A model API feature: the model returns a structured request to run a named function with arguments, and the host program runs it.
Model Context Protocol (MCP)
An open protocol, introduced in November 2024, for connecting models to tools, data sources and prompts through a standard client–server interface.
Agent2Agent (A2A)
An open protocol for agents built by different vendors or teams to find and message each other.
Reasoning model
A model trained to produce long intermediate reasoning before answering. These models do better on multi-step math, code and science tasks.
Frontier model
One of the most capable current models from the leading AI companies (for example OpenAI, Anthropic and Google), usually used through a paid online API.
Token
The unit of text a model reads and writes, roughly a word or part of a word. Model use, speed and cost are counted in tokens.
Rollout
One complete attempt by an agent at a task, from its first step to its final answer.
Open-weight model
A model whose weights are published, so it can run on site (for example on a laboratory cluster) without sending data to a vendor.
Benchmark contamination
When test items, or close copies of them, appear in a model's training data, so that scores overstate real ability.
LLM-as-judge
Using a language model to grade outputs such as ideas, papers or reviews. It is cheap but can be biased or easy to game.
Hallucination / fabricated citation
Fluent output that is not true: invented references, numbers or results.
Prompt injection
Hidden instructions in data an agent reads (a web page, a PDF, a tool result) that take over the agent's behavior.
Human-in-the-loop
A design in which a person approves or edits key steps, for example before an agent moves a motor or submits a job.
Reward hacking
When an agent raises the score, or passes the check it is judged by, without doing the intended task: for example by editing tests, using leaked answers or reporting only favorable runs.
Agent trace
The step-by-step record of an agent's run: its prompts, reasoning, tool calls and tool results. It shows what the agent actually did, which its final report may not.
Allowlist
A fixed list of the tools or commands an agent may use; anything not on the list is refused.
Fail-closed gate
A check that stops the workflow when it fails or cannot run, instead of letting the action go ahead.
Deterministic guard
A rule written in ordinary code, outside the model, that checks each agent action or report (for example against motor limits, or against the record of commands actually run) and blocks it if the check fails. Unlike a safety instruction in the prompt, the model cannot argue its way past it.
Digital twin
A simulation of an instrument or beamline detailed enough to run and score control code before it touches real hardware.
Bluesky / Ophyd
Python libraries for experiment orchestration (Bluesky) and hardware abstraction (Ophyd), used at NSLS-II and many other light sources.
EPICS
The Experimental Physics and Industrial Control System: the control-system layer (process variables) under most DOE accelerators and beamlines.
Tiled
A data-access service that serves array and table data (for example beamline data) over HTTP, with search and slicing.
SPEC
A long-established command-line program for diffractometer and beamline control, used at SSRL and many other synchrotrons.
CIF / checkCIF
The crystallographic information file (CIF) is the standard text format for a crystal structure and its refinement. checkCIF is the International Union of Crystallography's validation service; its level A and B alerts flag the most serious problems.
Globus Compute
A federated function-as-a-service platform that runs Python functions on remote computers, including HPC systems, through a cloud-managed service.
Parsl
A Python library for parallel scripting: it runs many tasks as one workflow on laptops, clusters or HPC systems.
Integrated Research Infrastructure (IRI)
A DOE Office of Science program to connect experimental facilities, HPC centers and networks so that data can move between them and be analyzed as it is taken.
ALCF, OLCF, NERSC, ESnet
DOE Office of Science computing and network facilities: the Argonne and Oak Ridge Leadership Computing Facilities, the National Energy Research Scientific Computing Center at Berkeley Lab, and the Energy Sciences Network that links the labs and user facilities.
Facilities named in this report
SLAC: SSRL (Stanford Synchrotron Radiation Lightsource) and LCLS (Linac Coherent Light Source, an x-ray free-electron laser). BNL: NSLS-II (National Synchrotron Light Source II) and CFN (Center for Functional Nanomaterials). Argonne: APS (Advanced Photon Source) and CNM (Center for Nanoscale Materials). LBNL: ALS (Advanced Light Source) and the Molecular Foundry. ORNL: SNS (Spallation Neutron Source) and CNMS (Center for Nanophase Materials Sciences). Outside DOE: CSNS (China Spallation Neutron Source) and DESY (Germany), which runs the PETRA III synchrotron.
ICLR, ICML, NeurIPS, ACL
The leading machine-learning and natural-language-processing conferences. In these fields, peer-reviewed conference papers, not journal articles, are the main publications.
Desk rejection
Rejection of a submitted paper by the editors or program chairs without full peer review, for example for breaking submission rules.

Method notes

  • Run date and scope. The survey reflects sources available on 2026-09-25. It covers AI agents built on language models, plus the closely related autonomous-experiment systems they build on, across the research workflow. The focus is 2024–2026; a few landmark 2023 systems are included for context.
  • How the work was split. An orchestrating agent split the survey into parallel research lanes (Q1 catalog, Q2 evidence and benchmarks, Q3 infrastructure, Q4 facilities, Q5 risks) and a tooling lane (report builder and checker), then merged and cross-checked the results in later iterations. The iteration ledger and claims file are in .goal/.
  • Sources. Discovery used web search together with the arXiv API, Crossref, OpenAlex, Semantic Scholar and the OSTI.GOV API (for DOE-funded work). Primary sources (papers, preprints, specifications, lab and project pages) were preferred over press coverage. When a preprint has a peer-reviewed version, the published version is cited.
  • Evidence labels. Peer-reviewed: published in a refereed journal or archival proceedings. Peer-reviewed + independent validation: also checked by a group other than the developers. Preprint: not yet refereed. Product/vendor report and Lab/project page: self-reported. In Q2 the first two classes are kept apart from claims.
  • Autonomy scale. Systems are placed on the L1–L4 scale defined in the glossary, based on what the cited source shows the system doing, not on what it is said to be able to do in future.
  • Reference checks. check_report.py parses this file and confirms, in the same run: every DOI through Crossref (or doi.org for DataCite DOIs), with the fetched title matching the listed title; every arXiv ID through the arXiv API, with the title matching; every other URL by fetching it. It also checks that every reference is cited, that every citation points to a listed reference, and that the page loads no external resources.
  • Fact checks. Before publication, independent audit agents checked the text against primary sources in three passes: 79 headline numbers (12 corrected or re-attributed), 211 further statements across all six questions (36 corrected), and the labels of all 83 catalog rows (20 changes, mostly first-release years and four autonomy levels). The audit files are in notes/.
  • Limitations. The field moves quickly, and many results exist only as preprints or company announcements. Title checks confirm that a reference exists and is the one listed; they do not confirm every sentence attributed to it. Facility work that is internal or unpublished is not visible to this method. The audience is described only by institution and field.

References

Numbered in order of first citation. Each entry shows the evidence type in brackets. DOIs link through doi.org, preprints through arXiv, and other sources are linked directly.

  1. Lu et al. Towards end-to-end automation of AI research. Nature (2026). doi:10.1038/s41586-026-10265-5 [Peer-reviewed]
  2. Mitchener et al. Kosmos: An AI Scientist for Autonomous Discovery. arXiv preprint (2025). arXiv:2511.02824 [Preprint]
  3. Feng et al. Towards Autonomous Mathematics Research. arXiv preprint (2026). arXiv:2602.10177 [Preprint]
  4. Gottweis et al. Accelerating scientific discovery with Co-Scientist. Nature (2026). doi:10.1038/s41586-026-10644-y [Peer-reviewed + independent validation]
  5. Ghareeb et al. A multi-agent system for automating scientific discovery. Nature (2026). doi:10.1038/s41586-026-10652-y [Peer-reviewed]
  6. Asai et al. Synthesizing scientific literature with retrieval-augmented language models. Nature (2026). doi:10.1038/s41586-025-10072-4 [Peer-reviewed]
  7. Huang et al. Autonomous biomedical research with an artificial intelligence agent. Science (2026). doi:10.1126/science.adz4351 [Peer-reviewed]
  8. Rein et al. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. arXiv (2023). arXiv:2311.12022 [Benchmark (preprint)]
  9. Epoch AI. AI Benchmarks & Capabilities | Epoch AI. Epoch AI Benchmarking Hub (2026). link [Lab/project page]
  10. Starace et al. PaperBench: Evaluating AI's Ability to Replicate AI Research. arXiv (2025). arXiv:2504.01848 [Benchmark (preprint)]
  11. Kon et al. EXP-Bench: Can AI Conduct AI Research Experiments? arXiv (2025). arXiv:2505.24785 [Benchmark (preprint)]
  12. Chen et al. An agentic artificially intelligent X-ray scientist. Nature Machine Intelligence (2026). doi:10.1038/s42256-026-01261-5 [Peer-reviewed]
  13. Hellert et al. Agentic artificial intelligence for multistage physics experiments at a large-scale user facility particle accelerator. Physical Review Research (2026). doi:10.1103/jtqy-9jz1 [Peer-reviewed]
  14. Hellert et al. Osprey: Production-ready agentic AI for safety-critical control systems. APL Machine Learning (2026). doi:10.1063/5.0306302 [Peer-reviewed]
  15. Xiao et al. NeuDiff Agent: a governed AI workflow for single-crystal neutron crystallography. Journal of Applied Crystallography (2026). doi:10.1107/S1600576726004474 [Peer-reviewed]
  16. MCP maintainers. MCP joins the Agentic AI Foundation | Model Context Protocol Blog. MCP blog (2025). link [Official docs/repo]
  17. Tanikanti et al. (Argonne). FIRST: Federated Inference Resource Scheduling Toolkit for Scientific AI Model Access. SC '25 Workshops (2025). doi:10.1145/3731599.3767346 [Peer-reviewed]
  18. ORNL OLCF. OLCF Inference Service Documentation — OLCF User Documentation. docs.olcf.ornl.gov (2026). link [Official docs/repo]
  19. Pan, Chard et al. (Globus Labs/Argonne). Experiences with Model Context Protocol Servers for Science and High Performance Computing. arXiv (2025). arXiv:2508.18489 [Preprint]
  20. Russinovich, Kumar, Salem. Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences. arXiv (2026). arXiv:2607.00738 [Preprint]
  21. Huang et al. Reward Hacking Challenges Oversight of Autonomous Research Agents. arXiv (2026). arXiv:2609.28614 [Preprint]
  22. Tripathi et al. APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction. arXiv (2026). arXiv:2609.24165 [Preprint]
  23. The White House. Launching the Genesis Mission – The White House. Executive order (2025). link [Lab/project page]
  24. Ding et al. Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap. arXiv preprint (survey) (2026). arXiv:2608.05179 [Preprint]
  25. Chiang et al. LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval. EMNLP 2025 (2025). doi:10.18653/v1/2025.emnlp-main.1280 [Peer-reviewed]
  26. Pham et al. ChemGraph as an agentic framework for computational chemistry workflows. Communications Chemistry (2026). doi:10.1038/s42004-025-01776-9 [Peer-reviewed]
  27. Skarlinski et al. Language agents achieve superhuman synthesis of scientific knowledge. arXiv preprint (2024). arXiv:2409.13740 [Preprint]
  28. Bragg et al. AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite. ICLR 2026 (2025). arXiv:2510.21652 [Peer-reviewed]
  29. Boiko et al. Autonomous chemical research with large language models. Nature (2023). doi:10.1038/s41586-023-06792-0 [Peer-reviewed]
  30. Shi et al. Knowledge-driven autonomous materials research via collaborative multi-agent and robotic system. Matter (2026). doi:10.1016/j.matt.2025.102577 [Peer-reviewed]
  31. Swanson et al. The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies. Nature (2025). doi:10.1038/s41586-025-09442-9 [Peer-reviewed]
  32. Ifargan et al. Autonomous LLM-Driven Research — from Data to Human-Verifiable Research Papers. NEJM AI (2025). doi:10.1056/AIoa2400555 [Peer-reviewed]
  33. Liang et al. Can Large Language Models Provide Useful Feedback on Research Papers? A Large-Scale Empirical Analysis. NEJM AI (2024). doi:10.1056/AIoa2400196 [Peer-reviewed]
  34. Thakkar et al. A large-scale randomized study of large language model feedback in peer review. Nature Machine Intelligence (2026). doi:10.1038/s42256-026-01188-x [Peer-reviewed]
  35. Bran et al. Augmenting large language models with chemistry tools. Nature Machine Intelligence (2024). doi:10.1038/s42256-024-00832-8 [Peer-reviewed]
  36. Novikov et al. AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv (white paper) (2025). arXiv:2506.13131 [Preprint]
  37. FutureHouse. Announcing Edison Scientific | FutureHouse. FutureHouse announcement (2025). link [Product/vendor report]
  38. Intology. Zochi Publishes A* Paper | Intology. Intology blog (2025). link [Product/vendor report]
  39. Weng et al. DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively. ICLR 2026 (2025). arXiv:2509.26603 [Peer-reviewed]
  40. Bianchi et al. Exploring the use of AI authors and reviewers at Agents4Science. Nature Biotechnology (2025). doi:10.1038/s41587-025-02963-8 [Peer-reviewed]
  41. Penadés et al. AI mirrors experimental science to uncover a mechanism of gene transfer crucial to bacterial evolution. Cell (2025). doi:10.1016/j.cell.2025.08.018 [Peer-reviewed]
  42. Leeman et al. Challenges in High-Throughput Inorganic Materials Prediction and Autonomous Synthesis. PRX Energy (2024). doi:10.1103/PRXEnergy.3.011002 [Peer-reviewed]
  43. Si et al. Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers. ICLR 2025 (2025). arXiv:2409.04109 [Peer-reviewed]
  44. Si et al. The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas. ICLR 2026 (2025). arXiv:2506.20803 [Peer-reviewed]
  45. Liu et al. AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration. arXiv preprint (2026). arXiv:2605.20025 [Preprint]
  46. Feng et al. Aletheia tackles FirstProof autonomously. arXiv preprint (2026). arXiv:2602.21201 [Preprint]
  47. OpenAI. Prism - AI LaTeX Editor. Product page (2026). link [Product/vendor report]
  48. Bisht et al. Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery. arXiv preprint (position paper) (2026). arXiv:2605.08956 [Preprint]
  49. Alampara et al. Probing the limitations of multimodal language models for chemistry and materials research. Nature Computational Science (2025). doi:10.1038/s43588-025-00836-3 [Benchmark (peer-reviewed)]
  50. Szymanski et al. An autonomous laboratory for the accelerated synthesis of inorganic materials. Nature (2023). doi:10.1038/s41586-023-06734-w [Peer-reviewed]
  51. Szymanski et al. Author Correction: An autonomous laboratory for the accelerated synthesis of inorganic materials. Nature (2026). doi:10.1038/s41586-025-09992-y [Peer-reviewed]
  52. Qu et al. CRISPR-GPT for agentic automation of gene-editing experiments. Nature Biomedical Engineering (2025). doi:10.1038/s41551-025-01463-z [Peer-reviewed]
  53. He et al. Chimeric infective particles expand species boundaries in phage-inducible chromosomal island mobilization. Cell (2025). doi:10.1016/j.cell.2025.08.019 [Peer-reviewed]
  54. Guan et al. AI‐Assisted Drug Re‐Purposing for Human Liver Fibrosis. Advanced Science (2025). doi:10.1002/advs.202508751 [Peer-reviewed + independent validation]
  55. Georgiev et al. Mathematical exploration and discovery at scale. arXiv (2025). arXiv:2511.02864 [Preprint]
  56. Mandal et al. Evaluating large language model agents for automation of atomic force microscopy. Nature Communications (2025). doi:10.1038/s41467-025-64105-7 [Peer-reviewed]
  57. Intology. GitHub - IntologyAI/Zochi: Repository for Zochi's Research · GitHub. GitHub (2025). link [Product/vendor report]
  58. Periodic Labs. Nature Is Our Learning Environment – Periodic Labs. Periodic Labs blog (2026). link [Product/vendor report]
  59. Jenewein et al. AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution. arXiv (2026). arXiv:2609.30133 [Preprint]
  60. Bubeck et al. Early science acceleration experiments with GPT-5. arXiv preprint (2025). arXiv:2511.16072 [Preprint]
  61. OpenAI. GitHub - openai/NavierStokesAndEuler: Lean certificates accompanying Navier-Stokes and Euler results · GitHub. GitHub (2026). link [Product/vendor report]
  62. Phan et al. A benchmark of expert-level academic questions to assess AI capabilities. Nature (2026). doi:10.1038/s41586-025-09962-4 [Benchmark (peer-reviewed)]
  63. Scale AI. Scale Labs Leaderboard: Humanity's Last Exam. Scale Labs leaderboard (2026). link [Lab/project page]
  64. Wang et al. FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks. arXiv (2026). arXiv:2601.21165 [Benchmark (preprint)]
  65. Zhu et al. Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark. arXiv (2025). arXiv:2509.26574 [Benchmark (preprint)]
  66. Mirza et al. A framework for evaluating the chemical knowledge and reasoning abilities of large language models against the expertise of chemists. Nature Chemistry (2025). doi:10.1038/s41557-025-01815-x [Benchmark (peer-reviewed)]
  67. Laurent et al. LAB-Bench: Measuring Capabilities of Language Models for Biology Research. arXiv (2024). arXiv:2407.10362 [Benchmark (preprint)]
  68. Zhou et al. Benchmarking large language models on safety risks in scientific laboratories. Nature Machine Intelligence (2026). doi:10.1038/s42256-025-01152-1 [Benchmark (peer-reviewed)]
  69. Tian et al. SciCode: A Research Coding Benchmark Curated by Scientists. arXiv (2024). arXiv:2407.13168 [Benchmark (preprint)]
  70. SciCode team. Leaderboard - SciCode Benchmark. SciCode leaderboard (2025). link [Lab/project page]
  71. Artificial Analysis. SciCode Benchmark Leaderboard | Artificial Analysis. Artificial Analysis leaderboard (2026). link [Lab/project page]
  72. Chen et al. ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery. ICLR 2025 (2024). arXiv:2410.05080 [Benchmark (peer-reviewed)]
  73. Siegel et al. CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark. TMLR (2024). arXiv:2409.11363 [Benchmark (peer-reviewed)]
  74. Princeton HAL. HAL: CORE-Bench Hard Leaderboard. HAL leaderboard (2026). link [Lab/project page]
  75. Mitchener et al. BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology. arXiv (2025). arXiv:2503.00096 [Benchmark (preprint)]
  76. Jansen et al. DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents. arXiv (2024). arXiv:2406.06769 [Benchmark (preprint)]
  77. Chan et al. MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering. arXiv (2024). arXiv:2410.07095 [Benchmark (preprint)]
  78. OpenAI. GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHub. GitHub leaderboard (2026). link [Lab/project page]
  79. Huang et al. MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation. arXiv (2023). arXiv:2310.03302 [Benchmark (preprint)]
  80. Wijk et al. RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts. arXiv (2024). arXiv:2411.15114 [Benchmark (preprint)]
  81. Xiang et al. SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers. arXiv (2025). arXiv:2504.00255 [Benchmark (preprint)]
  82. Ye et al. ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers? arXiv (2025). arXiv:2510.24591 [Benchmark (preprint)]
  83. Ai2. AstaBench: Rigorous benchmarking of AI agents with a holistic scientific research suite | Ai2. Ai2 blog (2025). link [Lab/project page]
  84. Cui et al. CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning. arXiv (2025). arXiv:2503.13517 [Benchmark (preprint)]
  85. Wang et al. FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights. arXiv (2026). arXiv:2602.02905 [Benchmark (preprint)]
  86. Faroughy et al. Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction. arXiv (2026). arXiv:2605.13950 [Benchmark (preprint)]
  87. FutureHouse. About 30% of Humanity’s Last Exam Answers are Wrong | FutureHouse. FutureHouse blog (2025). link [Lab/project page]
  88. Nadgir et al. Life After Benchmark Saturation: A Case Study of CORE-Bench. arXiv (2026). arXiv:2606.26158 [Benchmark (preprint)]
  89. Argonne ALCF. GitHub - argonne-lcf/ChemGraph: Agentic framework for computational chemistry and materials science workflows · GitHub. GitHub (2026). link [Official docs/repo]
  90. OpenAI (Singh et al.). OpenAI GPT-5 System Card. arXiv (2025). arXiv:2601.03267 [Product/vendor report]
  91. Stanford SNAP. GitHub - snap-stanford/Biomni: Biomni: a general-purpose biomedical AI agent · GitHub. GitHub (2026). link [Official docs/repo]
  92. OpenAI. gpt-oss-120b & gpt-oss-20b Model Card. arXiv (2025). arXiv:2508.10925 [Product/vendor report]
  93. Narayanan et al. (FutureHouse). Training a Scientific Reasoning Model for Chemistry. Advances in Neural Information Processing Systems 38 (2025). doi:10.52202/085713-5269 [Peer-reviewed]
  94. NNSA. NNSA's Los Alamos National Laboratory launches frontier AI models on the Venado supercomputer | Department of Energy. DOE/NNSA news (2025). link [Government document]
  95. Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023 (2022). arXiv:2210.03629 [Peer-reviewed]
  96. LangChain. GitHub - langchain-ai/langgraph: Build resilient agents. · GitHub. GitHub (2026). link [Official docs/repo]
  97. Microsoft. GitHub - microsoft/autogen: A programming framework for agentic AI · GitHub. GitHub (2026). link [Official docs/repo]
  98. Anthropic. Agent SDK overview - Claude Code Docs. Official docs (2026). link [Official docs/repo]
  99. Kamatar et al. (Globus Labs/Argonne/UChicago). Empowering Scientific Workflows with Federated Agents. IPDPS 2026 (2026). doi:10.1109/ipdps65963.2026.00114 [Peer-reviewed]
  100. A2A project. A New Chapter for A2A: Joining the Agentic AI Foundation - A2A Protocol. a2a-protocol.org blog (2026). link [Official docs/repo]
  101. MCP maintainers. Key Changes - Model Context Protocol. modelcontextprotocol.io (2026). link [Specification]
  102. Agent Skills project. Specification - Agent Skills. agentskills.io (2026). link [Specification]
  103. Pandolfi et al. Lightfall: An API-first, LLM-addressable control platform for synchrotron beamlines. arXiv (2026). arXiv:2606.06711 [Preprint]
  104. Do. EQSANS-CLI: A natural-language, agent-ready command-line tool for small-angle neutron scattering data reduction at EQ-SANS. arXiv (2026). arXiv:2605.00651 [Preprint]
  105. Wang et al. MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers. arXiv (2025). arXiv:2508.14925 [Preprint]
  106. Li. Trusted Credentials, Untrusted Behavior: Benchmarking LLM-Agent Security in High-Performance Computing. arXiv (2026). arXiv:2607.18485 [Preprint]
  107. Chard et al. funcX: A Federated Function Serving Fabric for Science. HPDC '20 (2020). doi:10.1145/3369583.3392683 [Peer-reviewed]
  108. Babuji et al. Parsl. HPDC '19 (Parsl: Pervasive Parallel Programming in Python) (2019). doi:10.1145/3307681.3325400 [Peer-reviewed]
  109. Pham et al. (Argonne). Multi-Agent Orchestration for High-Throughput Materials Screening on a Leadership-Class System. arXiv (2026). arXiv:2604.07681 [Preprint]
  110. DOE IRI. GitHub - doe-iri/iri-facility-api-python: The IRI Facility API reference implementation (Python) · GitHub. GitHub (2026). link [Official docs/repo]
  111. SLAC. GitHub - slaclab/iri-facility-api-python: The IRI Facility API reference implementation (Python) · GitHub. GitHub (2026). link [Official docs/repo]
  112. Bluesky collaboration. Bluesky Project. blueskyproject.io (2026). link [Official docs/repo]
  113. EPICS collaboration. EPICS - Experimental Physics and Industrial Control System. epics-controls.org (2026). link [Official docs/repo]
  114. Bluesky Project. Tiled documentation — tiled main documentation. Project documentation (2026). link [Lab/project page]
  115. Mathur et al. VISION: a modular AI assistant for natural human-instrument interaction at scientific user facilities. Machine Learning: Science and Technology (2025). doi:10.1088/2632-2153/add9e4 [Peer-reviewed]
  116. Rosendo et al. (ORNL). AI Agents for Enabling Autonomous Experiments at ORNL's HPC and Manufacturing User Facilities. SC '25 Workshops (2025). doi:10.1145/3731599.3767592 [Peer-reviewed]
  117. Souza et al. (ORNL). PROV-AGENT: Unified Provenance for Tracking AI Agent Interactions in Agentic Workflows. IEEE eScience 2025 (2025). doi:10.1109/escience65000.2025.00093 [Peer-reviewed]
  118. NERSC. Overview - NERSC Documentation. docs.nersc.gov (2026). link [Official docs/repo]
  119. van der Vleuten et al. EnvTrace: Simulation-Based Semantic Evaluation of LLM Code via Execution Trace Alignment -- Demonstrated at Synchrotron Beamlines. arXiv (2025). arXiv:2511.09964 [Preprint]
  120. DeepSeek-AI (Guo et al.). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645:633-638 (2025). doi:10.1038/s41586-025-09422-z [Peer-reviewed]
  121. Sainju et al. A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility. arXiv (2026). arXiv:2607.24663 [Preprint]
  122. Mishra et al. Domain Knowledge Guided Bayesian Optimization For Autonomous Alignment Of Complex Scientific Instruments. arXiv (2026). arXiv:2602.10670 [Preprint]
  123. Segal et al. A Multi-Scale Cognitive Interaction Model of Instrument Operations at the Linac Coherent Light Source. arXiv (2024). arXiv:2408.04734 [Preprint]
  124. Reed et al. ChatEED: An agentic retrieval assistant for accelerator operators. Proceedings of the SC '25 Workshops (2025). doi:10.1145/3731599.3767408 [Peer-reviewed]
  125. Prince et al. Opportunities for retrieval and tool augmented large language models in scientific facilities. npj Computational Materials (2024). doi:10.1038/s41524-024-01423-2 [Peer-reviewed]
  126. Vriza et al. Operating advanced scientific instruments with AI agents that learn on the job. npj Computational Materials (2026). doi:10.1038/s41524-026-02005-0 [Peer-reviewed]
  127. Du et al. Experiment automation agents: automating materials characterization with vision language model agents. npj Computational Materials (2026). doi:10.1038/s41524-026-02213-8 [Peer-reviewed]
  128. Yanguas-Gil et al. Design and performance of AI agents interfacing with an atomic layer deposition tool. Review of Scientific Instruments (2026). doi:10.1063/5.0318770 [Peer-reviewed]
  129. Hellert et al. From natural language to control signals: a conceptual framework for semantic channel finding in complex experimental infrastructure. Machine Learning: Science and Technology (2026). doi:10.1088/2632-2153/ae8219 [Peer-reviewed]
  130. Wall et al. TEM Agent: enhancing transmission electron microscopy with modern AI tools. npj Computational Materials (2026). doi:10.1038/s41524-026-02103-z [Peer-reviewed]
  131. Fei et al. Agentic LLM Reasoning in a Self-Driving Laboratory for Air-Sensitive Lithium Halide Spinel Conductors. arXiv (2026). arXiv:2604.11957 [Preprint]
  132. Liu et al. Synergizing human expertise and AI efficiency with language model for microscopy operation and automated experiment design *. Machine Learning: Science and Technology (2024). doi:10.1088/2632-2153/ad52e9 [Peer-reviewed]
  133. Brim et al. A Microservices Architecture Toolkit for Interconnected Science Ecosystems. SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis (2024). doi:10.1109/SCW63240.2024.00259 [Peer-reviewed]
  134. Sulc et al. eLog analysis for accelerators: status and future outlook. arXiv (IPAC'25) (2025). arXiv:2506.12949 [Preprint]
  135. Kaiser et al. Large language models for human-machine collaborative particle accelerator tuning through natural language. Science Advances (2025). doi:10.1126/sciadv.adr4173 [Peer-reviewed]
  136. Li et al. Rongzai agent: A Large Language Model-Based Autonomous Assistant for Rietveld Refinement of Neutron Diffraction Data. arXiv (2026). arXiv:2605.13911 [Preprint]
  137. Fuchs et al. From Overload to Insights: How AI Agents Can Support Scientists in Analyzing Complex Data. arXiv preprint (2026). arXiv:2607.16845 [Preprint]
  138. Noack et al. A Kriging-Based Approach to Autonomous Experimentation with Applications to X-Ray Scattering. Scientific Reports (2019). doi:10.1038/s41598-019-48114-3 [Peer-reviewed]
  139. Kusne et al. On-the-fly closed-loop materials discovery via Bayesian active learning. Nature Communications (2020). doi:10.1038/s41467-020-19597-w [Peer-reviewed]
  140. McDannald et al. On-the-fly autonomous control of neutron diffraction via physics-informed Bayesian active learning. Applied Physics Reviews (2022). doi:10.1063/5.0082956 [Peer-reviewed]
  141. Kandel et al. Demonstration of an AI-driven workflow for autonomous high-resolution scanning microscopy. Nature Communications (2023). doi:10.1038/s41467-023-40339-1 [Peer-reviewed]
  142. Maffettone et al. Self-driving Multimodal Studies at User Facilities. arXiv (2023). arXiv:2301.09177 [Preprint]
  143. Corrao et al. A modular framework for collaborative human-AI, multi-modal and multi-beamline synchrotron experiments. arXiv (2025). arXiv:2509.22959 [Preprint]
  144. Yin et al. Integrated edge-to-exascale workflow for real-time steering in neutron scattering experiments. Structural Dynamics (2024). doi:10.1063/4.0000279 [Peer-reviewed]
  145. Ding et al. SasAgent : multi-agent artificial intelligence system for small-angle scattering data analysis. Journal of Applied Crystallography (2026). doi:10.1107/S1600576726001299 [Peer-reviewed]
  146. An. Automated Data Reduction on VULCAN. OSTI technical report (ORNL) (2026). doi:10.2172/3413708 [Government document]
  147. Yin et al. PEAR: A Robust and Flexible Automation Framework for Ptychography Enabled by Multiple Large Language Model Agents. arXiv (2024). arXiv:2410.09034 [Preprint]
  148. Do et al. ESAC (EQ-SANS Assisting Chatbot): Application of large language models and retrieval-augmented generation for enhanced user experience at EQ-SANS. SoftwareX (2025). doi:10.1016/j.softx.2025.102191 [Peer-reviewed]
  149. Ding et al. LLMs can assist with proposal selection at large user facilities. Scientific Reports (2026). doi:10.1038/s41598-026-60226-1 [Peer-reviewed]
  150. US DOE. Energy Department Announces 26 Genesis Mission Science and Technology Challenges to Accelerate AI-Enabled American Innovation and Leadership | Department of Energy. DOE press release (2026). link [Government document]
  151. US DOE. Achieving AI-Driven Autonomous Laboratories | Department of Energy. DOE web page (2026). link [Government document]
  152. US DOE. Energy Department Announces Collaboration Agreements with 24 Organizations to Advance the Genesis Mission | Department of Energy. DOE press release (2025). link [Government document]
  153. The White House. Trump Administration Announces More Than $5 Billion for the Genesis Mission, a National Mission on AI for Science – The White House. Press release (2026). link [Government document]
  154. US DOE. U.S. Department of Energy Announces More Than $800 Million in Partner Commitments to the Genesis Mission | Department of Energy. DOE press release (2026). link [Government document]
  155. Garimella et al. (SCAC subcommittee). Genesis Mission Frameworks for AI-Accelerated National Breakthroughs: Report from the SCAC Subcommittee on the Genesis Mission. OSTI technical report (2026). doi:10.2172/3387411 [Government document]
  156. US Congress. Public Law 119 - 21 - An act to provide for reconciliation pursuant to title II of H. Con. Res. 14. - PLAW-119publ21 | Content Details | GovInfo. Public Law 119-21 (2025). link [Government document]
  157. DOE Office of Science (ASCR). GRANTS The American Science Clou... | U.S. DOE Office of Science(SC). Lab announcement LAB 25-3555 (2025). link [Government document]
  158. LBNL. How the Genesis Mission’s American Science Cloud Advances Innovation – Berkeley Lab News Center. Berkeley Lab News Center (2026). link [Lab/project page]
  159. Argonne National Laboratory (via Newswise). Real‑time AI engine poised to revolutionize large‑scale imaging data at national labs. Lab news release (2026). link [Lab/project page]
  160. US DOE. Energy Department Advances Investments in AI for Science | Department of Energy. DOE press release (2025). link [Government document]
  161. DOE Office of Science (ASCR). GRANTS Robotics and Automation T... | U.S. DOE Office of Science(SC). Lab announcement LAB 26-3601 (2026). link [Government document]
  162. US DOE. DOE Announces Roadmap for New Initiative for Artificial Intelligence in Science, Security and Technology | Department of Energy. DOE press release (2024). link [Government document]
  163. LLNL. 1,000 Scientist AI Jam Session explores AI-driven scientific discovery | Lawrence Livermore National Laboratory. LLNL news (2025). link [Lab/project page]
  164. LANL. Los Alamos Lab partners with OpenAI to boost national security | LANL. LANL news (2025). link [Lab/project page]
  165. Anthropic. Working with the US Department of Energy \ Anthropic. Company news (2025). link [Product/vendor report]
  166. US DOE. Energy Department Announces New Partnership with NVIDIA and Oracle to Build Largest DOE AI Supercomputer | Department of Energy. DOE press release (2025). link [Government document]
  167. US DOE. Energy Department Announces New Public-Private Partnership Model, Two Supercomputers, to Accelerate American Dominance in Science and Technology | Department of Energy. DOE press release (2025). link [Government document]
  168. Miller et al. Integrated Research Infrastructure Architecture Blueprint Activity (Final Report 2023). OSTI technical report (2023). doi:10.2172/1984466 [Government document]
  169. Allan et al. Bluesky's Ahead: A Multi-Facility Collaboration for an a la Carte Software Project for Data Acquisition and Management. Synchrotron Radiation News (2019). doi:10.1080/08940886.2019.1608121 [Peer-reviewed]
  170. Koepp et al. Toward Unified Autonomous Scattering Experiments: A Cross-Facility Case Study at ALS and PETRA III. Photon Science (2026). doi:10.1021/photonsci.5c00044 [Peer-reviewed]
  171. Yin et al. Toward Agentic HPC: Serving and Evaluating LLM-Powered Agents for Scientific Applications on Leadership-Class Platforms. ISC High Performance 2026 Research Paper Proceedings (2026). doi:10.23919/ISC.2026.11520499 [Peer-reviewed]
  172. Le Houx. Benchmarking Autonomy in Scientific Experiments: A Hierarchical Taxonomy for Autonomous Large-Scale Facilities. arXiv (2026). arXiv:2601.06978 [Preprint]
  173. Ferreira da Silva et al. Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap. OSTI technical report (ORNL) (2026). doi:10.2172/3377973 [Government document]
  174. Kapoor et al. AI Agents That Matter. arXiv (2024). arXiv:2407.01502 [Preprint]
  175. Ruan et al. Identifying the Risks of LM Agents with an LM-Emulated Sandbox. ICLR 2024 (2024). arXiv:2309.15817 [Peer-reviewed]
  176. ICMJE. ICMJE | Recommendations | Preparing a Manuscript for Submission to a Medical Journal. ICMJE Recommendations (2025). link [Editorial/policy]
  177. Atil et al. Non-Determinism of "Deterministic" LLM Settings. arXiv (2024). arXiv:2408.04667 [Preprint]
  178. Chen, Zaharia, Zou. How Is ChatGPT’s Behavior Changing Over Time? Harvard Data Science Review (2024). doi:10.1162/99608f92.5317da47 [Peer-reviewed]
  179. Kapoor et al. Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation. arXiv (2025). arXiv:2510.11977 [Preprint]
  180. Beel, Kan, Baumgart. Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future? ACM SIGIR Forum (2025). doi:10.1145/3769733.3769747 [Peer-reviewed]
  181. Walters, Wilder. Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports (2023). doi:10.1038/s41598-023-41032-5 [Peer-reviewed]
  182. ICLR 2026 Program Chairs. ICLR 2026 Response to LLM-Generated Papers and Reviews – ICLR Blog. ICLR Blog (2025). link [Official page]
  183. Luo, Kasirzadeh, Shah. The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems. arXiv (2025). arXiv:2509.08713 [Preprint]
  184. Miao, Pritchard, Zou. The Agentic Garden of Forking Paths. arXiv (2026). arXiv:2607.01507 [Preprint]
  185. Zhang et al. A Careful Examination of Large Language Model Performance on Grade School Arithmetic. Advances in Neural Information Processing Systems 37 (NeurIPS 2024) (2024). doi:10.52202/079017-1485 [Peer-reviewed]
  186. Zhu et al. Establishing Best Practices for Building Rigorous Agentic Benchmarks. arXiv (2025). arXiv:2507.02825 [Preprint]
  187. Cemri et al. Why Do Multi-Agent LLM Systems Fail? Advances in Neural Information Processing Systems 38 (NeurIPS 2025) (2025). doi:10.52202/085713-4082 [Peer-reviewed]
  188. Zheng et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems 36 (NeurIPS 2023) (2023). doi:10.52202/075280-2020 [Peer-reviewed]
  189. Messeri, Crockett. Artificial intelligence and illusions of understanding in scientific research. Nature (2024). doi:10.1038/s41586-024-07146-0 [Peer-reviewed]
  190. Greshake et al. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. AISec '23 (ACM) (2023). doi:10.1145/3605764.3623985 [Peer-reviewed]
  191. Lu et al. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv (2024). arXiv:2408.06292 [Preprint]
  192. Lin. Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review. Communications of the ACM (2026). doi:10.1145/3779116 [Peer-reviewed]
  193. Urbina et al. Dual use of artificial-intelligence-powered drug discovery. Nature Machine Intelligence (2022). doi:10.1038/s42256-022-00465-9 [Peer-reviewed]
  194. Mariano et al. Robotic safety in self-driving laboratories. SLAS Technology (2026). doi:10.1016/j.slast.2026.100463 [Peer-reviewed]
  195. Ferreira da Silva et al. (ORNL). Shaping the Future of Self-Driving Autonomous Laboratories Workshop. OSTI technical report (DOE) (2024). doi:10.2172/2481197 [Official page]
  196. Nature editorial. Tools such as ChatGPT threaten transparent science; here are our ground rules for their use. Nature (2023). doi:10.1038/d41586-023-00191-1 [Editorial/policy]
  197. NeurIPS 2025. LLM Policy. NeurIPS website (2025). link [Editorial/policy]
  198. ICLR 2026 Program Chairs. Policies on Large Language Model Usage at ICLR 2026 – ICLR Blog. ICLR Blog (2025). link [Editorial/policy]
  199. ICML 2026 Program Chairs. On Violations of LLM Review Policies – ICML Blog. ICML Blog (2026). link [Official page]
  200. Liang et al. Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews. ICML 2024 (2024). arXiv:2403.07183 [Peer-reviewed]
  201. Pangram Labs. Pangram Predicts 21% of ICLR Reviews are AI-Generated | Pangram. Pangram blog (2025). link [Product/vendor report]
  202. Agents4Science organizers (Stanford). Open Conference of AI Agents for Science: 2025. Conference website (2025). link [Official page]
  203. NIH. NOT-OD-23-149: The Use of Generative Artificial Intelligence Technologies is Prohibited for the NIH Peer Review Process. NIH Guide notice (2023). link [Editorial/policy]
  204. NIH. NOT-OD-25-132: Supporting Fairness and Originality in NIH Research Applications. NIH Guide notice (2025). link [Editorial/policy]
  205. NSF. Notice to Research Community: Use of Generative Artificial Intelligence Technology in the NSF Merit Review Process - Policies | NSF - U.S. National Science Foundation. NSF policy notice (2023). link [Editorial/policy]
  206. Mayet. GAIA: A General AI Assistant for Intelligent Accelerator Operations. arXiv (2024). arXiv:2405.01359 [Preprint]
  207. Google. Gemini: Try Deep Research and Gemini 2.0 Flash Experimental. Google blog (The Keyword) (2024). link [Product/vendor report]
  208. Baek et al. ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models. NAACL 2025 (2025). doi:10.18653/v1/2025.naacl-long.342 [Peer-reviewed]
  209. Shao et al. Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models. NAACL 2024 (2024). doi:10.18653/v1/2024.naacl-long.347 [Peer-reviewed]
  210. Singh et al. Ai2 Scholar QA: Organized Literature Synthesis with Attribution. ACL 2025 System Demonstrations (2025). doi:10.18653/v1/2025.acl-demo.49 [Peer-reviewed]
  211. FutureHouse. FutureHouse Platform: Superintelligent AI Agents for Science | FutureHouse. FutureHouse announcement (2025). link [Product/vendor report]
  212. Romera-Paredes et al. Mathematical discoveries from program search with large language models. Nature (2023). doi:10.1038/s41586-023-06924-6 [Peer-reviewed]
  213. Wang et al. SciMON: Scientific Inspiration Machines Optimized for Novelty. ACL 2024 (2024). doi:10.18653/v1/2024.acl-long.18 [Peer-reviewed]
  214. Yang et al. MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses. ICLR 2025 (2025). arXiv:2410.07076 [Peer-reviewed]
  215. Ghafarollahi & Buehler. SciAgents: Automating Scientific Discovery Through Bioinspired Multi‐Agent Intelligent Graph Reasoning. Advanced Materials (2024). doi:10.1002/adma.202413523 [Peer-reviewed]
  216. InternAgent Team et al. InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification. arXiv preprint (2025). arXiv:2505.16938 [Preprint]
  217. Huang et al. Automated Hypothesis Validation with Agentic Sequential Falsifications. ICML 2025 (2025). arXiv:2502.09858 [Peer-reviewed]
  218. Roohani et al. BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments. ICLR 2025 (2025). arXiv:2405.17631 [Peer-reviewed]
  219. Song et al. A Multiagent-Driven Robotic AI Chemist Enabling Autonomous Chemical Research On Demand. Journal of the American Chemical Society (2025). doi:10.1021/jacs.4c17738 [Peer-reviewed]
  220. Cao et al. Automating quantum computing laboratory experiments with an agent-based AI framework. Patterns (2025). doi:10.1016/j.patter.2025.101372 [Peer-reviewed]
  221. Dai et al. Autonomous mobile robots for exploratory synthetic chemistry. Nature (2024). doi:10.1038/s41586-024-08173-7 [Peer-reviewed]
  222. Darvish et al. ORGANA: A robotic assistant for automated chemistry experimentation and characterization. Matter (2025). doi:10.1016/j.matt.2024.10.015 [Peer-reviewed]
  223. Panapitiya et al. AutoLabs: cognitive multi-agent systems with self-correction for autonomous chemical experimentation. Scientific Reports (2026). doi:10.1038/s41598-026-45593-z [Peer-reviewed]
  224. Zhang et al. A multimodal robotic platform for multi-element electrocatalyst discovery. Nature (2025). doi:10.1038/s41586-025-09640-5 [Peer-reviewed]
  225. Ghafarollahi & Buehler. Automating alloy design and discovery with physics-aware multimodal multiagent AI. PNAS (2025). doi:10.1073/pnas.2414074122 [Peer-reviewed]
  226. Hong et al. Data Interpreter: An LLM Agent for Data Science. Findings of ACL 2025 (2025). doi:10.18653/v1/2025.findings-acl.1016 [Peer-reviewed]
  227. Guo et al. DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning. ICML 2024 (2024). arXiv:2402.17453 [Peer-reviewed]
  228. Schmidgall et al. Agent Laboratory: Using LLM Agents as Research Assistants. Findings of EMNLP 2025 (2025). doi:10.18653/v1/2025.findings-emnlp.320 [Peer-reviewed]
  229. Villaescusa-Navarro et al. The Denario project: Deep knowledge AI agents for scientific discovery. arXiv preprint (2025). arXiv:2510.26887 [Preprint]
  230. Zou et al. El Agente: An autonomous agent for quantum chemistry. Matter (2025). doi:10.1016/j.matt.2025.102263 [Peer-reviewed]
  231. Aygün et al. An AI system to help scientists write expert-level empirical software. arXiv preprint (2025). arXiv:2509.06503 [Preprint]
  232. Campbell et al. MDCrow: automating molecular dynamics workflows with large language models. Machine Learning: Science and Technology (2026). doi:10.1088/2632-2153/ae4b07 [Peer-reviewed]
  233. Jin et al. AgentReview: Exploring Peer Review Dynamics with LLM Agents. EMNLP 2024 (2024). doi:10.18653/v1/2024.emnlp-main.70 [Peer-reviewed]
  234. Weng et al. CycleResearcher: Improving Automated Research via Automated Review. ICLR 2025 (2025). arXiv:2411.00816 [Peer-reviewed]
  235. Schmidgall & Moor. AgentRxiv: Towards Collaborative Autonomous Research. arXiv preprint (2025). arXiv:2503.18102 [Preprint]
  236. Tang et al. AI-Researcher: Autonomous Scientific Innovation. NeurIPS 2025 (2025). arXiv:2505.18705 [Peer-reviewed]
  237. Zhu et al. DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process. ACL 2025 (2025). doi:10.18653/v1/2025.acl-long.1420 [Peer-reviewed]
  238. Jiang & Ng (Stanford). Tech Overview - Stanford Agentic Reviewer. Project page (2025). link [Lab/project page]