The AI Scientist Has A Metrology Problem – OpEd
Stronger lab AI can hide what the experiment did not measure. Casimir-force inversion is ill-conditioned: many material models give similar force curves, so a network can output parameters that come from the training prior, not from the data. Identifiability must come before a bigger model.
In a gold-sphere / thin-film example, five material parameters reduced to about two informative combinations. Nanometer uncertainty in the gap swamped force precision. Optimal sampling improved information only about twofold. The bottleneck was the ruler, not the algorithm.
Autonomous labs need a hierarchy: calibration and model validity before inference and optimization. Benchmarks should include geometry, nuisance parameters and priors—and allow the answer “not enough information.” Better metrology can matter more than another trillion parameters.
AI can optimize inference, but it cannot recover information that an instrument never measured. A Casimir-force inverse problem shows why the next frontier of scientific AI is better metrology, not merely larger models.
Artificial intelligence is rapidly becoming part of the laboratory.
It selects observations, fits spectra, reconstructs images, proposes molecular structures, identifies anomalies, and solves inverse problems that once required painstaking human analysis. The natural assumption is that as the algorithms become more powerful, scientific measurements will become correspondingly more informative.
Sometimes the opposite is true.
A sufficiently capable algorithm can make the weakest part of an experiment harder to see.
I recently encountered this problem in an unlikely place: the Casimir effect, the tiny interaction produced by electromagnetic fluctuations between closely separated surfaces. The physics has been studied for decades. The forward problem, predicting the force from material properties and geometry, is well developed. More recently, researchers have begun asking whether the process can be reversed: if the Casimir force is measured at different separations, can algorithms infer properties of the materials that produced it?
The idea is attractive. The Casimir interaction depends on electromagnetic response across a broad range of frequencies. Change the distance between two surfaces and the weighting of those frequencies changes. In principle, a sequence of force measurements contains information about the material.
On August 20, Hideo Iizuka and Shanhui Fan reported in Physical Review Applied that machine learning could infer thin-film thickness and broadband material-response parameters from modeled Casimir-force measurements. That result is important because it moves the question from speculation to computation.
But it also exposes the harder question.
Not: can a machine-learning system return an answer?
Rather: which........
