Physics-checked AI

AI calibration that refuses impossible data

It learns only what the physics leaves open, and refuses data that breaks physical limits instead of passing it on.

It is not a large language model and is not trained on text. The correction models are small and physical, and the learning is Bayesian.

a reading outside physical limitsrefused,not “corrected”a correction a law allowslearnedthe calibratorphysics priors · a physics-designed calibrationa store that checks its own history · fleet poolingPhylaws, units and physical limitsfeeds the calibrator directly
Why it is needed

Data-only AI is weakest exactly where it matters

Quantum sensors carry many small physical errors: tilt, temperature, the rotation of the Earth, the ground underneath, the vehicle they ride in. AI is now used to correct them.

Data-only correctionThe physics-checked approachEvidence
Fails silently where its data stopsCorrections are laws with their own variables (latitude enters as the law says), so they transfer+0.05 E against +0.44 E or more
Can output the impossibleEvery quantity has physical limits; a reading outside them is refused, not “corrected”Dead-sensor windows refused 100 %
Cannot say whyEvery correction traces to a named lawFour mechanisms found, constants within 1–8 %
Uncertainty is not honestIntervals from physics plus calibration; coverage measured and reported92–98 % coverage where claimed; misses reported
Needs lots of data per devicePhysics fixes most of the model; data fills the rest9.9 pT against 26.1 pT from one short session
Trusts its inputsEach calibration is checked against its own physical historyPoisoned day: 5.8 pT against 45.3 pT naive
Measured · four pre-registered studiesStudies 1–3 simulated · Study 4 real recordings

22 of 33 claims passed. Every failure is reported.

26.117.09.97.16.3no calibrationterms pickedby fitphysics priorsphysics-designedcalibrationoracle, knowsthe true termspT after one short session

Study 2: calibration from little data

An NV-diamond heart magnetometer, simulated. After one short session, physics priors and a physics-designed calibration came close to an oracle that knows the true terms.

Dead-sensor windows were refused 100 %, none served. Poisoned calibrations were refused 10 of 10; clean ones were refused 0 of 10.

Study 1: discovery and robustness

Gravity gradiometer, simulated with known truth. It found the right physical mechanisms (Coriolis, tilt × temperature) with constants within 1–8 % of the truth, and no false mechanism, even in a world that had none. Moved to a new latitude it added +0.05 E of error, against +0.44 to +0.47 E for every data-only model. At a new site: 16.6 E against 26.1 E for a black-box model. Injected faults: 100 % flagged.

Study 3: real reference data, one calibrator for a fleet

A real US Geological Survey magnetic observatory agreed with the IGRF-14 world field model (total field −0.21 %). 105,366 real weather temperatures were all inside the physical limits. 300 gradiometer stations pooled with the physics gave 6.67 E, close to 20,000 stations of physics alone (6.36 E); 300 stations alone gave 10.27 E.

Study 4: real sensor recordings

Real optically pumped magnetometer heart recordings (Kiel, PhysioNet) peaked at 16.3–55.3 pT, inside the predicted physical bounds. On real soil temperatures the heat equation was refuted, not forced: the damping depth came out at 11 cm from the amplitude and 29 cm from the phase. The model recorded that the law holds only for a homogeneous medium.

Limits
  • At the extremes it ties a good generic physics-ML library (6.96 against 7.00 E). Accuracy alone is not the advantage; trust is.
  • 11 of 33 claims failed, including interval coverage short of its bar in three places, and reuse across instruments: the heart sensor without the Study 2 method was worse than no correction.
  • Studies 1–3 used simulated instruments, and the real-data tests are small.
  • One team designed and ran the studies; an independent re-run would carry more weight.
  • It is not a medical, navigation or safety product. The heart recordings were checked against physical bounds, not used for diagnosis.
How it differs from PINNs, neural operators and sparse equation discovery
  • each correction traces to a named law;
  • units are checked;
  • physical limits are enforced on inputs and outputs;
  • it refuses a law the data contradicts;
  • it transfers to new conditions through the law’s own variables;
  • it compiles to small device firmware.

Those approaches are strong and tie it on accuracy. The lasting value is the curated, checked library of instrument mechanisms and the toolchain around it, not the physics itself, which is public.