Large language models do not replace chemists in a closed-loop catalysis experiment

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Artificial intelligence (AI) is reshaping scientific research and laboratory automation1. Large language models (LLMs) can perform aspects of scientific reasoning, which could in principle reduce human decision-making as a rate-limiting step in closed-loop automated experiments2-7. Yet the reasoning performance of LLMs compared with human experts in complex, noisy, long-running laboratory experiments is largely unexplored. Here we benchmarked LLM reasoning head-to-head against a team of human domain experts for a noisy 25-dimensional closed-loop colloidal catalysis problem8 explored using a mobile robot9. The LLM (GPT-5.1) navigated the available chemical space across a 528-experiment campaign, using background information and experimental data to propose an anionic surfactant that gave the largest single gain in catalyst activity, also adapting to a mid-campaign change in the measurement set-up. In a like-for-like final phase of 160 experiments, the human experts found a catalyst formulation that was, on average, more active than the best two formulations found in independent LLM runs. At the same time, the LLM reasoned 35 times faster and was estimated to be around 1,900 times less expensive than human reasoning. Inspection of the LLM reasoning traces found them mostly sound, but with some costly silent errors, logical inconsistencies, and apparent memory limitations, suggesting that LLMs are not out-of-the-box replacements for expert reasoning in problems of this complexity10. This points to a need to design closed-loop systems that combine the speed of machine reasoning with expert oversight.

Article activity feed