Classical molecular mechanics (MM) force fields remain the cornerstone of molecular modeling across biomolecular structural biology, materials science, and computer-aided drug discovery. While polarizable and off-plane virtual-site augmented force fields significantly improve physical realism for aromatic stacking and anisotropic electrostatic interactions, parameterizing these potentials against condensed-phase target data presents a high-dimensional, non-convex inverse problem.
Recently, Košťál and co-workers introduced a Bayesian framework for force field learning, employing Gaussian Process (GP) emulators to map force field parameters $\boldsymbol{\theta} \in \mathbb{R}^D$ to condensed-phase Quantities of Interest (QoI), including radial distribution functions $g(r)$ and hydrogen bond coordination numbers $N_{\text{HB}}$. While statistically principled, the baseline methodology exhibits critical computational and mathematical bottlenecks:
The complete closed-loop architecture of our serverless multi-fidelity active learning platform is illustrated in Figure 1, outlining the five sequential stages from initial space-filling exploration and SVD basis extraction to uncertainty-penalized plausibility screening and thermodynamic reweighting.
We validate our platform on two benchmark condensed-phase systems:
Rather than fitting 200 independent scalar Gaussian processes across discrete radial bins $r_k$, we model continuous radial distribution functions $g(r; \boldsymbol{\theta})$ via Functional Principal Component Analysis (FPCA). Given an empirical ensemble of centered RDF curves $\{ g_i(r) - \bar{g}(r) \}_{i=1}^N$, we perform Singular Value Decomposition (SVD): $$g(r; \boldsymbol{\theta}) = \bar{g}(r) + \sum_{j=1}^{M} c_j(\boldsymbol{\theta}) \phi_j(r), \quad \text{s.t.} \quad \int \phi_j(r) \phi_l(r) \, dr = \delta_{jl}$$ where $\{ \phi_j(r) \}_{j=1}^M$ are orthonormal spatial basis modes and $c_j(\boldsymbol{\theta}) \sim \mathcal{GP}(0, K_j(\boldsymbol{\theta}, \boldsymbol{\theta}'))$ are independent GPyTorch Matérn-5/2 Gaussian processes. With only $M=5$ components, the FPCA surrogate captures 88.23% of empirical spatial variance while reducing surrogate parameter dimensionality by 97.5%. The leading spatial SVD eigenmodes $\phi_j(r)$ and the continuous RDF reconstructions are shown in Figure 2.
For scalar observable peaks $\mathcal{O}(\boldsymbol{\theta})$, fitting a GP on logarithmic observations $z_i = \ln \hat{\mathcal{O}}_i$ introduces finite-sampling bias because $\mathbb{E}[\ln \hat{\mathcal{O}}] < \ln \mathbb{E}[\hat{\mathcal{O}}]$ by Jensen's inequality. Applying the Delta method, we derive the heteroscedastic noise transformation: $$\sigma_{\ln}^2(\boldsymbol{\theta}) = \ln\left( 1 + \frac{\sigma_{\text{MD}}^2(\boldsymbol{\theta})}{N_{\text{eff}} \cdot \hat{\mathcal{O}}(\boldsymbol{\theta})^2} \right) \approx \frac{\sigma_{\text{MD}}^2(\boldsymbol{\theta})}{N_{\text{eff}} \cdot \hat{\mathcal{O}}(\boldsymbol{\theta})^2}$$ where $N_{\text{eff}} = N_{\text{frames}} / (2\tau_{\text{corr}} + 1)$ accounts for MD auto-correlation. The unbiased physical observable expectation is analytically recovered via: $$\mathbb{E}[\mathcal{O}(\boldsymbol{\theta})] = \exp\left( \mu_z(\boldsymbol{\theta}) + \frac{1}{2}\sigma_z^2(\boldsymbol{\theta}) \right)$$
Standard Mahalanobis plausibility gating functions $[\mu(\boldsymbol{\theta}) - \mu_0]^2 / [\sigma_{\text{GP}}^2(\boldsymbol{\theta}) + \sigma_0^2] \le \tau^2$ suffer from the "ignorance paradox": when $\boldsymbol{\theta}$ lies far outside the sampled domain, epistemic uncertainty explodes ($\sigma_{\text{GP}}(\boldsymbol{\theta}) \to \infty$), driving the ratio to zero and spuriously classifying unphysical regions as plausible.
We resolve this by formulating PlausibilityGateV2 with conservative confidence bounds and Sliced-Wasserstein distribution distances: $$\Omega_{\text{ref}} = \left\{ \boldsymbol{\theta} \in \Theta \; \middle| \; |\mu(\boldsymbol{\theta}) - \mu_0| + \kappa \sigma_{\text{GP}}(\boldsymbol{\theta}) \le \tau \sigma_0 \right\}$$ Setting $\kappa = 2.0$ guarantees that regions of high ignorance are strictly penalized and excluded from $\Omega_{\text{ref}}$ until confirmed by physical simulations, as visualized in Figure 3.
To eliminate the requirement of launching full MD simulations for every proposed parameter step, we deploy Multistate Bennett Acceptance Ratio (MBAR) reweighting. Within a local trust region around simulated anchor $\boldsymbol{\theta}_0$, the reweighted observable expectation and thermodynamic susceptibility metric tensor $\mathbf{C}_{ij}$ are computed via: $$\langle \mathcal{O} \rangle_{\boldsymbol{\theta}} = \sum_{t=1}^T w_t(\boldsymbol{\theta}) \mathcal{O}(\mathbf{x}_t), \quad w_t(\boldsymbol{\theta}) = \frac{\exp(-\beta \Delta U(\mathbf{x}_t; \boldsymbol{\theta}))}{\sum_{t'} \exp(-\beta \Delta U(\mathbf{x}_{t'}; \boldsymbol{\theta}))}$$ $$\mathbf{C}_{ij} = \frac{\partial^2 \ln Z}{\partial \theta_i \partial \theta_j} = \beta^2 \operatorname{Cov}_0\left( \frac{\partial U}{\partial \theta_i}, \frac{\partial U}{\partial \theta_j} \right)$$ Using Kish's Effective Sample Size ($\text{ESS} \ge 20\%$), local perturbations are evaluated in 0.0015 seconds ($400,000\times$ faster than MD).
Evaluating PlausibilityGateV2 across 3,007 candidate parameters admitted 495 physical candidates (16.46% manifold volume) while strictly rejecting spurious high-uncertainty boundary regions, as confirmed in Figure 3.
Executing the complete 2-stage active learning campaign on Modal Cloud generated 40 Stage 1 screening trajectories and 10 Stage 2 active refinement trajectories, completing in 181.8 seconds (3.0 minutes). The full empirical active discovery campaign and the resulting structural RDF agreement with physical MD ground truth are illustrated in Figure 4.
| Force Field Parameter | Posterior Mean ± 1σ | 95% Credible Interval | Prior Bounds |
|---|---|---|---|
| Carbonyl Carbon $q(\text{C1})$ | -0.2223 ± 0.1714 e | [-0.4865, +0.0810] e | [-0.50, +0.10] e |
| Carboxyl Oxygens $q(\text{O1, O2})$ | -0.6581 ± 0.1315 e | [-0.8851, -0.4163] e | [-0.90, -0.40] e |
| Methyl Carbon $q(\text{C2})$ | +0.7385 e | Implicit (Charge Conserved) | [+0.30, +0.90] e |
The inferred posterior parameter distributions and the decomposed Fisher information stiff/sloppy anisotropy modes ($1.71\times$ anisotropy ratio) are depicted in Figure 5, demonstrating robust convergence across all MCMC chains.
Compared to 5.0 hours sequentially on a workstation and 2.5 hours on a traditional Slurm cluster, Modal Cloud execution completed the full 50-simulation campaign in 3.0 minutes ($250\times$ faster) with 0.0% unphysical simulation waste, as benchmarked in Figure 6.
A systematic sampling strategy ablation and sensitivity sweep across plausibility thresholds $\tau \in [3.5, 4.5]$ are presented in Figure 7, confirming optimal viable yield and minimal computational waste. Furthermore, Figure 8 demonstrates the local 0.0015s thermodynamic reweighting trust region evaluated via Kish Effective Sample Size alongside the complete SHA-256 cryptographic provenance audit across all test suites.
We have presented an end-to-end, serverless multi-fidelity active learning platform for Bayesian force field optimization. By unifying Functional PCA spatial basis surrogates, Jensen-corrected Log-GPs, PlausibilityGateV2, Multistate MBAR thermodynamic reweighting, and NumPyro NUTS Hamiltonian MCMC, we resolve the major mathematical and computational bottlenecks of classical force field parameterization. Deployed live on Modal Cloud, our platform completes full 50-simulation campaigns in 3.0 minutes for under $1.00, achieving a $250\times$ wall-clock speedup with verified statistical convergence and cryptographic provenance.