Abstract
Scientific discovery seeks mechanisms that explain observations and predict beyond the measurements that produced them. Whereas large language models (LLMs) encode knowledge implicitly, mechanistic world models organise it as a parsimonious set of explicit, modular, reusable mechanisms whose predictions can be scored against data. In biology, where measurements are sparse and noisy, mechanistic world modelling must discover latent states and governing equations jointly, yet the prior knowledge that could constrain this search is largely unstructured. We introduce gemot, an agentic framework for mechanistic world modelling, and evaluate it on 18 published biological problems spanning molecular biology, epidemiology, and immune-cell differentiation, with up to 1,755 training measurements, 65 observables, 170 experimental conditions, and 3 data modalities per problem. gemot is auditable by construction: a semantic layer records each hypothesis, and a Model Context Protocol layer separates hypothesis formulation from numerical evaluation, so hypothesis scoring cannot be fabricated. On every problem, autonomously constructed models match or exceed the reference models in fit and parsimony, and generalise where held-outs exist. We demonstrate that gemot can formulate novel, biologically plausible mechanistic hypotheses. This work will enable decoding of interventional biological data into competing mechanistic hypotheses and design experiments that distinguish between them, and may serve as a blueprint for mechanistic world modelling in other domains.