Appendix E: Rigorous Proof Sketches for Ontological and Metabolic Laws#
TLDR#
This appendix contains rigorous proof sketches backing the ontology/metabolism results in the later cognition chapters.
Read it when you want the mathematical spine behind the narrative statements; otherwise treat it as a reference.
This appendix collects calculations for the cognition and multi-agent chapters. The multi-agent proofs track the operator, measure, and variational functional used in each result, so that changes of representation can be checked directly.
We operate on the latent Riemannian manifold \((\mathcal{Z}, G)\) with belief measures \(\rho \in \mathcal{P}(\mathcal{Z})\).
E.1 Proof of Theorem Theorem 21#
Statement: The ontology should expand from \(N_c\) to \(N_c + 1\) charts if and only if \(\Xi > \Xi_{\text{crit}}\) and \(\Delta V_{\text{proj}} > \mathcal{C}_{\text{complexity}}\).
Hypothesis: Let \(\mathcal{S}[N_c] = \inf_{\theta} \mathcal{S}_{\text{onto}}(\theta, N_c)\) be the value function of the ontological action for \(N_c\) charts.
Proof
Consider the discrete variation \(\Delta \mathcal{S} = \mathcal{S}[N_c + 1] - \mathcal{S}[N_c]\). By the definition of the Ontological Action (Section 30.3):
where \(\mathcal{S}_{\text{task}} = \mathbb{E}[\langle V \rangle]\) is the expected task value.
Expanding \(\mathcal{S}_{\text{task}}\) via a first-order Taylor approximation in the space of representations:
The marginal utility of a new chart is \(\frac{\partial \langle V \rangle}{\partial N_c} = \Delta V_{\text{proj}}\). The complexity cost is \(\mu_{\text{size}}\). Therefore:
The transition \(N_c \to N_c + 1\) is the global minimizer iff \(\Delta \mathcal{S} < 0\), which yields:
The condition \(\Xi > \Xi_{\text{crit}}\) ensures that the second variation of the texture-entropy functional \(\delta^2 H(z_{\text{tex}})\) is negative-definite at the vacuum. This precludes the absorption of the signal into the existing noise floor: if \(\Xi \le \Xi_{\text{crit}}\), the texture residual \(z_{\text{tex}}\) is truly unpredictable noise, and adding a chart provides no informational benefit. \(\square\)
E.2 Proof of Theorem Theorem 22#
Statement: The emergence of a new chart follows a supercritical pitchfork bifurcation with control parameter \(\mu = \Xi - \Xi_{\text{crit}}\).
Hypothesis: The potential \(\Phi_{\text{onto}}(r)\) is \(SO(n)\)-invariant near \(r=0\), where \(r = \|q_* - q_{\text{parent}}\|\) is the radial distance of the new query from the parent.
Proof
Let \(f(\Xi) = \Xi - \Xi_{\text{crit}}\) be the control parameter. By \(SO(n)\) symmetry, the Ontological Action can only depend on even powers of \(r\) near the origin. We expand in a power series:
where \(\beta > 0\) for stability (the quartic term must be positive for bounded energy).
The stationarity condition \(\frac{\partial \mathcal{S}}{\partial r} = 0\) yields:
This has solutions:
\(r = 0\) (trivial, no new chart)
\(r^2 = f(\Xi)/\beta\) (symmetry-broken state)
Analysis of stability:
For \(f(\Xi) < 0\) (i.e., \(\Xi < \Xi_{\text{crit}}\)): The Hessian at \(r=0\) is \(\frac{\partial^2 \mathcal{S}}{\partial r^2}|_{r=0} = -f(\Xi) > 0\). Thus \(r=0\) is a stable minimum.
For \(f(\Xi) > 0\) (i.e., \(\Xi > \Xi_{\text{crit}}\)): The Hessian at \(r=0\) becomes \(-f(\Xi) < 0\) (unstable). New minima appear at \(r^* = \sqrt{f(\Xi)/\beta}\).
Since \(r \ge 0\) is a radial coordinate, this constitutes a supercritical pitchfork bifurcation where the symmetry-broken state \(r^* > 0\) becomes the unique stable equilibrium for \(\Xi > \Xi_{\text{crit}}\).
The bifurcation diagram: for \(\Xi < \Xi_{\text{crit}}\), the system has a single stable fixed point at \(r=0\); for \(\Xi > \Xi_{\text{crit}}\), the origin becomes unstable and two symmetric branches (in the full space, a sphere of radius \(r^*\)) emerge. \(\square\)
E.3 Proof of Theorem Theorem 29#
Statement. Under the compact-domain, positive-density, no-flux, mass-preservation, and calibration hypotheses in the theorem, the conditional Landauer-form estimate is \(\dot{\mathcal M}(s)\ge T_c\lvert dH(\rho_s)/ds\rvert\).
Proof
Proof. The WFR equation is \(\partial_s\rho=-\nabla\!\cdot(\rho v)+\rho r\). For \(H(\rho)=-\int\rho\ln\rho\,d\mu_G\), differentiation gives
The no-flux condition removes the boundary contribution. Mass preservation removes the term \(\int\rho r\,d\mu_G\), so
Set
Cauchy–Schwarz in \(L^2(\rho d\mu_G)\) gives
Adding the estimates yields \(|dH/ds|\le\sqrt{I_\rho E_v}+\sqrt{J_\rho E_r}\). The theorem’s explicit calibration compares this right-hand side with \(\dot{\mathcal M}=\sigma_{\mathrm{met}}(E_v+\lambda^2E_r)\) and therefore gives the claimed Landauer-form inequality. No de Bruijn identity or universal physical identification is required. \(\square\)
E.4 Proof of Theorem Theorem 30#
Statement: The optimal computation budget \(S^*\) satisfies \(\frac{d}{ds} \langle V \rangle_{\rho_s}|_{s=S^*} = \dot{\mathcal{M}}(S^*)\).
Hypothesis: \(S^*\) is an interior point of \([0, S_{\max}]\).
Proof
Define the deliberation functional:
The necessary condition for an extremum is \(\mathcal{F}'(S) = 0\). By the Leibniz integral rule:
Using the result that \(\partial_s \rho\) is governed by the WFR operator \(\mathcal{L}_{\text{WFR}}\):
By the adjoint property of the WFR operator (the formal \(L^2(\rho)\) adjoint):
where \(\mathcal{L}_{\text{WFR}}^* V = -\langle \nabla V, v \rangle_G + Vr\) (transport-adjoint plus reaction).
For gradient flows in the covariant case, \(v = -G^{-1}\nabla_A V\) with \(\nabla_A V := \nabla V - A\):
Thus:
In the conservative case (\(A=0\)), \(G^{-1}(\nabla V, \nabla_A V) = \|\nabla V\|_G^2\), the power dissipated by the value-gradient flow. The stationarity condition \(\mathcal{F}'(S^*) = 0\) gives:
This states that the optimal stopping time \(S^*\) is reached when the power dissipated by the value-gradient flow exactly matches the metabolic cost rate. \(\square\)
E.5 Proof of Theorem Definition 167#
Statement: \(F_{\text{total}} = -G^{-1}\nabla_A V + \beta_{\text{exp}} G^{-1}\nabla\Psi_{\text{causal}}\).
Hypothesis: The agent’s path minimizes \(\mathcal{S} = \int L(z, \dot{z}) \, dt\) with Lagrangian \(L = \frac{1}{2}\|\dot{z}\|_G^2 - (V + \beta_{\text{exp}}\Psi_{\text{causal}})\).
Proof
The Euler-Lagrange equations for the functional are:
Computing the momentum:
Time derivative of momentum:
Potential gradient:
Euler-Lagrange equation:
Recognizing the Christoffel symbols of the first kind \([ij, k] = \frac{1}{2}(\partial_i G_{jk} + \partial_j G_{ik} - \partial_k G_{ij})\):
Contracting with \(G^{mk}\) and using \(\Gamma^m_{ij} = G^{mk}[ij, k]\):
This is the geodesic equation with forcing terms. In the overdamped limit (Section 22.3), inertia is negligible and the acceleration term vanishes, leaving:
The drift field \(F_{\text{total}}\) is the first-order velocity approximation, proving the additive force of curiosity. \(\square\)
E.6 Proof of Theorem Definition 168#
Statement: The macro-ontology \(K\) is interventionally closed iff \(I(K_{t+1}; Z_{\text{micro}, t} | K_t, do(K^{\text{act}}_t)) = 0\).
Hypothesis: Let \(\mathcal{M}\) be a Markov Blanket for \(K\).
Proof
We compare the mutual information under the observational measure \(P\) and the interventional measure \(P_{do(K^{\text{act}})}\).
Observational case: By the Causal Enclosure condition (Section 2.8):
This states that the macro-state \(K_{t+1}\) is conditionally independent of the micro-texture \(Z_{\text{micro}, t}\) given the current macro-state and action.
Interventional case: The \(do(K^{\text{act}}_t)\) operator performs a graph surgery that removes all incoming edges to \(K^{\text{act}}_t\) while preserving all other mechanisms. By Pearl’s Causal Markov Condition [Pearl, 2009]:
This is because the mechanism \(P(K_{t+1} | \text{parents}(K_{t+1}))\) is a structural equation that does not depend on how \(K^{\text{act}}_t\) was generated.
Combining the conditions: If the observational distribution satisfies \(I = 0\), then:
Since the mechanism is invariant under intervention:
Therefore, \(I(K_{t+1}; Z_{\text{micro}, t} | K_t, do(K^{\text{act}}_t)) = 0\).
Contrapositive (violation): If \(I > 0\) under \(do(K^{\text{act}}_t)\), there exists a back-door path through \(Z_{\text{micro}, t}\):
This path was confounded in observational data (the correlation between \(Z_{\text{micro}}\) and \(K_{t+1}\) was screened by the policy generating \(K^{\text{act}}_t\)). The intervention breaks this screening, exposing the hidden variable. The remedy is Ontological Expansion (Section 30): promote the relevant component of \(Z_{\text{micro}}\) to a new macro-variable in \(K\). \(\square\)
E.7 Scalar spectral and barrier calculations#
Definition 325 (The scalar metric realization)
Use the compact connected smooth scalar manifold and its positive strategic metric from the existing Appendix E.7 setup. The same spectral margin is \(\|G^{-1/2}hG^{-1/2}\|_{\mathrm{op}}<1\) for sign-indefinite perturbations; positive perturbations preserve positivity directly. The joint volume is \(w=\sqrt{\det\widetilde G}\), as in Theorem 63. Compactness makes a fixed smooth positive metric uniformly elliptic. This statement concerns the compact realization already used here, not a replacement of unbounded confinement by an artificial compact domain.
Definition 326 (Scalar form realization)
For the existing real \(C^2\) potential \(U\) bounded below, set \(q_\sigma[u]=\tfrac{\sigma^2}{2}\int|\nabla u|^2+\int U|u|^2\). The form domain is \(H^1\) on a closed manifold or for the Neumann realization, and \(H_0^1\) for the Dirichlet realization. Its self-adjoint operator is \(H_\sigma=-\sigma^2\Delta_{\widetilde G}/2+U\) with the selected boundary condition. The domain is the operator domain associated to this form, not unrestricted \(H^2\) on a domain with boundary. Compact embedding gives compact resolvent for this fixed model.
Definition 327 (Forbidden and allowed regions)
For an energy \(E\) define \(A_E=\{U\le E\}\) and \(K_E=\{U>E\}\). These are sets of a scalar potential. Their relation to payoff basins is tested separately by the unilateral inequalities of Theorem 48.
Theorem 125 (Fixed scalar ground state)
The compact connected nonmagnetic scalar realization has a simple ground eigenvalue \(E_0\) and an eigenfunction positive in the interior.
Proof. Compact resolvent and the lower form bound give a minimizer of the normalized Rayleigh quotient. Replacing it by its modulus does not increase its gradient energy, so a nonnegative minimizer exists. Interior elliptic regularity and the strong maximum principle make it strictly positive there. The scalar heat kernel is positivity improving on a connected domain with the selected Dirichlet or Neumann realization; a bounded real potential preserves this through its positive Feynman–Kac weight. For compact positive time the leading eigenspace of this positivity-improving self-adjoint semigroup is one-dimensional, hence so is the ground eigenspace. In the Dirichlet case the eigenfunction vanishes on the boundary. Every nonempty interior open set contains a relatively compact ball on which its continuous square has a positive minimum, giving positive mass in that open set. This proves stationary positivity, not a dynamical crossing rate. Compact resolvent and simplicity also give \(E_1>E_0\) for this fixed operator. \(\square\)
Definition 328 (Barrier metric at a fixed energy)
Define \(g_E=2(U-E)_+\widetilde G\) and \(d_E(x,A)=\inf_{\gamma:A\to x}\int\sqrt{2(U-E)_+}\,|\dot\gamma|_{\widetilde G}\). The factor two matches the kinetic normalization \(-\sigma^2\Delta/2\). Distances may vanish within an allowed connected region.
Theorem 126 (Exact weighted eigenfunction identity)
For an eigenfunction \((H_\sigma-E)u=0\) in the scalar realization and a bounded smooth real weight \(f\), put \(v=e^{f/\sigma}u\). Then
All integrals use \(d\mu_{\widetilde G}\); the weight is taken constant in the normal direction for the Neumann case. Compactly supported cutoffs give the local form, with their explicit derivative terms retained.
Proof. Test the weak eigenvalue equation against \(e^{2f/\sigma}u\) and take real parts. The gradient product is
Substitution proves the identity. For \(0<\epsilon<1\) and a bounded Lipschitz approximation to \(f=(1-\epsilon)d_E(\cdot,A_E)\), its gradient satisfies \(|\nabla f|^2\le2(1-\epsilon)^2(U-E)_+\) almost everywhere. Weak approximation preserves the resulting energy inequality. On the forbidden region the remaining potential coefficient is at least \((2\epsilon-\epsilon^2)(U-E)\); the allowed-region negative term controls the weighted integral. For a region where \(U-E\ge\delta>0\) and \(f\ge a\),
This is a weighted stationary estimate, including its constants. No dimension-independent \(H^1\to L^\infty\) embedding or uniform pointwise prefactor is used. \(\square\)
Corollary 48 (Metric comparison at fixed potential and energy)
For \(g_1\succeq g_0\) and the same \(U,E\), every path obeys \(\int\sqrt{2(U-E)_+}|\dot\gamma|_{g_1} \ge\int\sqrt{2(U-E)_+}|\dot\gamma|_{g_0}\). Taking infima gives \(d_E^{g_1}\ge d_E^{g_0}\). This compares barrier actions at the same energy. Changing a Hamiltonian generally changes its ground energy as well, so one cannot substitute two different ground energies into this fixed-energy inequality or deduce an ordering of transition probabilities from two upper bounds. \(\square\)
Theorem 127 (Feynman–Kac clock and normalization)
Let \(X_s\) have generator \(\Delta_{\widetilde G}/2\), killed at a Dirichlet boundary or reflected for the Neumann realization. Then
with the survival indicator in the killed case. To obtain the ground vector,
Proof. The diffusion generator and multiplication weight give \(\partial_tu=(\Delta/2-U/\sigma^2)u=-H_\sigma u/\sigma^2\) with the same initial and boundary data, which is the Feynman–Kac semigroup. The spectral expansion multiplies each eigencomponent by \(e^{-t(E_n-E_0)/\sigma^2}\); dominated convergence leaves exactly its ground projection. Divide by the nonzero overlap to recover \(u_0\). The original WFR generator is compared separately; Brownian motion here is the process associated to this displayed scalar semigroup. \(\square\)
Corollary 49 (Action-length inequality)
For any absolutely continuous path in the forbidden region, set \(a=|\dot\gamma|_{\widetilde G}\), \(b=\sqrt{2(U-E)}\). Then \(a^2/2+U-E-ab=(a-b)^2/2\ge0\). Integration gives \(\int(|\dot\gamma|^2/2+U-E)dt \ge\int\sqrt{2(U-E)}|\dot\gamma|dt\). Equality is attained on a parametrization with \(a=b\) wherever that parametrization is defined. This relates an action to barrier length. An asymptotic probability additionally belongs to a specified stochastic law and event; it is not furnished by this algebraic inequality.
E.8 Proof of Corollary Remark 18#
Statement: \(V_H(z) = T_c^2 \frac{\partial H(\pi)}{\partial T_c}\).
Hypothesis: Let \(\pi(a|z) = \frac{1}{Z} \exp\left(\frac{Q(z,a)}{T_c}\right)\) be the policy, where \(Z = \sum_a \exp(Q/T_c)\) is the partition function. Let \(\beta_{\text{ent}} = 1/T_c\) be the inverse Definition 73.
Proof
Step 1: Express Entropy in terms of \(\beta_{\text{ent}}\). The entropy of the policy is:
Substituting \(\ln \pi(a) = \beta_{\text{ent}} Q(a) - \ln Z\):
Step 2: Derivative of Entropy w.r.t. \(\beta_{\text{ent}}\). Differentiating with respect to \(\beta_{\text{ent}}\):
Using the identity \(\frac{\partial \ln Z}{\partial \beta_{\text{ent}}} = \mathbb{E}_\pi[Q]\):
where we used \(\frac{\partial \mathbb{E}[Q]}{\partial \beta_{\text{ent}}} = \mathrm{Var}(Q)\) (standard fluctuation-response relation).
Step 3: Relate \(\mathrm{Var}(Q)\) to Varentropy. Recall \(\mathcal{I}(a) = -\ln \pi(a) = -\beta_{\text{ent}} Q(a) + \ln Z\). The variance of the surprisal is:
Step 4: Change of variables to \(T_c\). We have \(V_H = \beta_{\text{ent}}^2 \mathrm{Var}(Q)\) and \(\frac{\partial H}{\partial \beta_{\text{ent}}} = -\beta_{\text{ent}} \mathrm{Var}(Q)\). Therefore \(V_H = -\beta_{\text{ent}} \frac{\partial H}{\partial \beta_{\text{ent}}}\).
Using the chain rule \(\frac{\partial}{\partial T_c} = -\frac{1}{T_c^2} \frac{\partial}{\partial \beta_{\text{ent}}}\):
Final Result: Rearranging yields:
This proves that Varentropy equals the heat capacity and measures the sensitivity of the entropy to temperature fluctuations. \(\square\)
E.9 Proof of Corollary Corollary 13#
Statement: For a bimodal policy on a value ridge, \(V_H\) is significant, distinguishing it from uniform noise.
Hypothesis: Let \(\pi\) be a mixture of two dominant modes with values \(Q_1, Q_2\) and a background of \(N-2\) negligible modes.
Proof
Step 1: Variance of Surprisal Form.
Since \(\mathcal{I} = -\beta Q + \ln Z\), we have \(V_H = \beta^2 \mathrm{Var}(Q)\).
Step 2: Two-Point Statistics. Consider two actions \(a_1, a_2\) with probabilities \(p, 1-p\). The variance of a Bernoulli variable taking values \(Q_1, Q_2\) is:
Thus:
For equally weighted modes (\(p = 1/2\)), this simplifies to:
Step 3: Interpretation of \(\Delta Q\). \(\Delta Q\) is the value gap between the two modes.
Perfect Symmetry (The Ridge): If \(Q_1 = Q_2\) exactly, then \(\Delta Q = 0 \implies V_H = 0\).
Structural Instability: When the agent is slightly off-center or when sampling includes the tails, the effective \(\Delta Q > 0\).
Step 4: Distinguishing Structure from Noise. For a distribution with structure (peaks and valleys), \(\mathrm{Var}(Q) > 0\). For a flat distribution (noise), \(\mathrm{Var}(Q) = 0\).
Specifically, on a ridge, the agent samples \(a_{\text{left}}\) and \(a_{\text{right}}\) (high \(Q\)) but also transitively samples the separating region (lower \(Q\)) during exploration. The variance of \(Q\) along the trajectory corresponds to \(V_H\):
This proves that \(V_H\) detects the topological feature (the valley) that distinguishes a fork from a flat plane. \(\square\)
E.10 Proof of Corollary Corollary 12#
Statement: To maintain stability, the cooling rate must satisfy \(|\dot{T}_c| \ll T_c / \sqrt{V_H}\).
Hypothesis: We require the probability distribution \(\pi_t\) to remain close to the equilibrium Boltzmann distribution \(\pi^*_{T_c(t)}\) during annealing. This is the Adiabatic Condition.
Proof
Step 1: Thermodynamic Speed. The rate of change of the policy distribution with respect to temperature is measured by the Fisher Information metric \(g_{TT}\) on the statistical manifold parameterized by \(T_c\):
Step 2: Relate Fisher Metric to Varentropy. Recall \(\ln \pi = \frac{Q}{T_c} - \ln Z\). Then:
Substituting into the Fisher definition:
Using \(V_H = \frac{\mathrm{Var}(Q)}{T_c^2}\) (from Proof E.8):
Step 3: Thermodynamic Length. The “distance” traversed in probability space for a small temperature change \(dT_c\) is \(ds^2 = g_{TT} dT_c^2\):
Step 4: Adiabatic Condition. For the system to relax to equilibrium (stay in the basin of attraction), the speed of change in distribution space must be bounded:
Substituting \(ds/dt\):
Solving for the cooling rate:
Conclusion: When Varentropy \(V_H\) is large (phase transition/critical point), the permissible cooling rate goes to zero. The Governor must apply the “Varentropy Brake” to prevent quenching the system into a suboptimal metastable state. \(\square\)
E.11 Proof of Corollary Remark 36#
Statement: \(\nabla \Psi_{\text{causal}} \propto \nabla \mathbb{E}_{z'} [ V_H[P(\theta_W | z, a, z')] ]\).
Hypothesis: We define \(\Psi_{\text{causal}}\) as the Expected Information Gain (EIG) about model parameters \(\theta\) given a transition \((z, a) \to z'\).
Proof
Step 1: Definition of EIG.
This is the Total Predictive Entropy minus the Expected Aleatoric Entropy.
Step 2: Decomposition of Uncertainty. For the “noisy TV” case (outcomes are stochastic noise independent of \(\theta\)):
Step 3: Varentropy as Structure Detector. The varentropy \(V_H(z' | z, a)\) measures the variance of log-probabilities.
Uniform noise: \(V_H^{\text{noise}} \to 0\) (all outcomes equally likely).
Structured uncertainty: \(V_H^{\text{structured}} > 0\) (some outcomes much more likely).
Step 4: Connection to Multimodality. If the model is uncertain about structure (\(\theta\)), the predictive distribution \(p(z')\) is a mixture of distinct hypotheses \(p(z'|\theta_1), p(z'|\theta_2)\). As established in Proof E.9, a mixture of distinct modes has high Varentropy compared to a broad unimodal distribution (noise).
Step 5: Operational Equivalence. Thus, maximizing EIG is functionally equivalent to maximizing the Varentropy of the expected outcome, provided the aleatoric noise floor is constant:
Conclusion: The agent should seek states where the World Model’s prediction has high Varentropy (conflicting hypotheses), as these offer the maximum potential for falsification (reduction of parameter variance). \(\square\)
E.12 Bellman generator and scalar wave variation#
Proof
For the diffusion and discount already used in Theorem 12, write \(\mathcal L=b\cdot\nabla+T_c\Delta_G\) and \(\gamma_h=e^{-\lambda h}\). The smooth Bellman equation has continuous-time form
In its stationary zero-drift sector, \((-\Delta_G+\lambda/T_c)V=r/T_c\); denote this screening coefficient by \(\kappa_B^2=\lambda/T_c\). The scalar field action used in this chapter instead defines the wave operator
The stationary operators coincide under the coefficient identification \(\kappa^2=\kappa_B^2\) and the same sources and boundary realization.
Proof. Generator consistency gives \(\mathbb E[V(Z_h,t+h)]=V+h(\partial_t+\mathcal L)V+o(h)\). Insert this and \(e^{-\lambda h}=1-\lambda h+o(h)\) into \(V=rh+e^{-\lambda h}\mathbb E[V(Z_h,t+h)]\), cancel \(V\), and divide by \(h\). The \(\partial_t^2V\) Taylor term has coefficient \(h/2\) after division and vanishes. Finite signal speed does not change that coefficient. For the wave model vary
Integration by parts against a compactly supported variation \(\eta\) gives \(\delta S=\int\eta[-\Box_gV-\kappa^2V+\rho_r]\sqrt{|g|}\,dx\). Stationarity proves the field equation. For a fixed product metric \(g=\operatorname{diag}(-c^2,G)\), \(\Box_g=c^{-2}\partial_t^2-\Delta_G\). These are explicit equations for two defined evolutions; equality of their stationary operators is the comparison established here. \(\square\)
E.13 Polar WFR calculation with exact signs#
Proof
On a smooth positive-density chart with the fixed spatial metric of the WFR equations, put \(R=\sqrt\rho\), \(p=dV-B\), \(v=G^{-1}p\), \(D=\nabla-iB/\sigma\), and \(Q_B=-\sigma^2\Delta_GR/(2R)\). For the equations already stated, \(\partial_s\rho+\operatorname{div}_G(\rho v)=r\rho\) and \(\partial_sV+|p|_G^2/2+\Phi_{\mathrm{eff}}=0\), the exact amplitude equation is
Proof. The product rule gives
Thus the kinetic real part is \(Q_B+|p|_G^2/2\). The \(-Q_B\) term cancels it to the given classical HJB expression. The imaginary part is \(-\sigma\operatorname{div}_G(\rho v)/(2\rho)+\sigma r/2\), equal to \(\sigma\partial_s\rho/(2\rho)\) by continuity. The time derivative is \(i\sigma\partial_s\psi/\psi=i\sigma\partial_s\rho/(2\rho)-\partial_sV\), so both parts agree. Reading these two parts backwards proves the local equivalence. At zeros use the density/current equations without division by \(R\); a global phase additionally retains its existing circulation data.
The compensating \(Q_B\) depends on \(|\psi|\), making this amplitude equation nonlinear. Omitting the compensation gives the distinct linear Schrödinger model with a \(+Q_B\) term in its Hamilton–Jacobi equation. A curl-modified mobility must be substituted into its own continuity equation; the above Laplacian yields precisely the canonical velocity \(G^{-1}p\). For time-dependent volume density \(w_s\), conservation reads \(\partial_s(w_s\rho)+\partial_i(w_s\rho v^i)=w_sr\rho\); the corresponding amplitude equation acquires \(-i\sigma\partial_s\log w_s/2\). \(\square\)
E.14 Complete-history conditional kernels#
Proof
On the standard measurable path spaces of the specified process, the complete history state \((t,\mathsf H_t)\) is Markov with its conditional extension kernel. Neither a positive delay alone nor the reward occupation screen determines whether a smaller state is Markov.
Proof. The sigma-algebra generated by \((t,\mathsf H_t)\) contains the entire modeled history through \(t\). Let \(K_{t,u}(h,\cdot)\) be the regular conditional law of the extended history through \(u\) given \(\mathsf H_t=h\). For a bounded history functional \(F\),
The tower property gives \(K_{t,u}=K_{t,v}K_{v,u}\) on realized histories. Including \(t\) in the state accounts for time-inhomogeneous coefficients. This proves the representation without identifying a finite compression. For the occupation screen take \(\alpha=0\): every history has screen zero, while histories with the same current position can have different \(z_{t-\tau}\) and therefore different delayed drifts. Conversely a delayed signal with zero coupling leaves a Markov local process Markov. These examples prove both limitations of the smaller-state claims. \(\square\)
E.15 Time averages and unilateral variations#
Proof
For a bounded differentiable density trajectory, \(T^{-1}\int_0^T\partial_t\rho\,dt=(\rho(T)-\rho(0))/T\to0\). This holds for many nonequilibrium trajectories and does not test unilateral payoff improvements. Moreover \(\langle\rho v\rangle\) need not vanish when \(\langle v\rangle=0\): on a periodic clock take \(v=\sin t\) and \(\rho=1+\epsilon\sin t\), \(0<\epsilon<1\), giving \(\langle\rho v\rangle=\epsilon/2\). Standing-wave expansions describe solutions of the defined wave operator; Nash conditions are the payoff inequalities in Theorem 48. \(\square\)
E.16 Strategic Hessian and response composition#
Proof
Use the smooth local best-response branch and Strategic Jacobian already specified in Definition 329. With the intrinsic connection on agent \(j\)’s manifold define the covariant tensor
No additional lowering of the Hessian indices is applied. The strategic metric prescription is \(\widetilde G^{(i)}=G^{(i)}+h^{(i)}\) with \(h^{(i)}=\sum_{j\ne i}\beta_{ij}\mathcal G^{(i)}_{ij}\). Its positive-definite domain is checked using the spectral margin already specified in Definition 325. The curvature equation Theorem 6 remains a separate differential identity; this algebraic prescription is not its solution.
For a \(C^2\) response \(y=b(x)\), direct differentiation gives
Thus the pulled-back \(H_{jj}\) is one contribution, not the full Hessian of the composed value. The final term vanishes at a stationary point in \(y\). For the positive metric \(\widetilde G=G+h\), subtraction of the two metric-compatible torsion-free connections gives the exact identity
Replacing \(\widetilde G^{-1}\) by \(G^{-1}\) gives its first-order expansion, with a remainder controlled by the inverse-metric identity \(\widetilde G^{-1}-G^{-1}=-G^{-1}h\widetilde G^{-1}\).
Definition 329 (Strategic Jacobian on the established response branch)
On the smooth, nondegenerate local best-response branch specified in the original strategic construction, differentiate \(\nabla_jV_j(b_j(z_{-j}),z_{-j})=0\). The chain rule gives \(H_{jj}^{(j)}\mathcal J_{ji}+H_{ji}^{(j)}=0\), hence \(\mathcal J_{ji}=-(H_{jj}^{(j)})^{-1}H_{ji}^{(j)}\). This identifies the branch derivative on the domain where the declared inverse exists. It is not a global selection of a multivalued best-response correspondence. Its pullback action on covariant Hessians is exactly that used in Definition 208.
E.17 Curvature and Bianchi identity#
Proof
For the defined connection, the adjoint covariant derivative is \(\mathcal D_\rho F_{\mu\nu}=\partial_\rho F_{\mu\nu}-ig[A_\rho,F_{\mu\nu}]\). On a test section \(u\), expansion gives \([D_\rho,F_{\mu\nu}]u=(\mathcal D_\rho F_{\mu\nu})u\). Insert \([D_\mu,D_\nu]=-igF_{\mu\nu}\) into \([D_\rho,[D_\mu,D_\nu]]+\mathrm{cyclic}=0\). Dividing by \(-ig\) gives
For \(g=0\) the same identity is \(d(dA)=0\). The geometric Christoffel terms cancel under cyclic antisymmetrization for the torsion-free connection. The identity holds in each smooth gauge chart and is preserved under transition functions by conjugation; nontrivial bundle topology does not violate it. \(\square\)
E.18 Vacuum expansion in the declared representation#
Proof
For the stable quartic potential already specified, \(\lambda>0\) and \(\mu^2<0\), write \(\Phi_0=(v/\sqrt2)n\), \(n^\dagger n=1\) and \(v^2=-\mu^2/\lambda\). The radial mass is \(m_h^2=2\lambda v^2\). The gauge quadratic term is \(-\tfrac12A_\mu^a(M^2)_{ab}A^{\mu b}\) with
Proof. Put \(q=\Phi^\dagger\Phi\). The minimum solves \(\mu^2+2\lambda q=0\). Substitute \(q=(v+h)^2/2\); the coefficient of \(h^2\) in \(U\) is \(\lambda v^2=m_h^2/2\). For a constant vacuum, \(D_\mu\Phi_0=-igA_\mu^aT_a\Phi_0\); the symmetric product of the commuting coefficients \(A^aA^b\) gives the displayed anticommutator. For any real \(u^a\), \(u^a(M^2)_{ab}u^b=2g^2\|(u^aT_a)\Phi_0\|^2\ge0\). Its kernel is the stabilizer Lie algebra of the vacuum. For a single \(SU(2)\) doublet with \(T_a=\sigma_a/2\), this evaluates to \((M^2)_{ab}=g^2v^2\delta_{ab}/4\). It is this representation that gives \(m_A=gv/2\). Other declared representations are evaluated by the same matrix formula. If \(\lambda\le0\), the stated stable quartic expansion does not apply; the potential itself reveals the failure. \(\square\)
E.19 Spectral minimization and the Nash comparison#
Proof
The Rayleigh quotient of the scalar Hamiltonian minimizes its single joint energy. The Nash test instead compares each agent’s own payoff under a unilateral change. Their equality must be checked by differentiating the actual objectives and by evaluating their global inequalities.
Proof by explicit comparison. Let \(V_1(x,y)=-(x-y)^2\) and \(V_2(x,y)=-(y-1)^2-Kx\) with \(K>0\). Both are strictly concave in their own coordinate; their best responses are \(x=y\) and \(y=1\), hence the unique Nash profile is \((1,1)\). The sum of costs is \(U=(x-y)^2+(y-1)^2+Kx\). Its \(x\) derivative at \((1,1)\) is \(K\), so Nash is not a stationary point of this joint energy. This example meets the smooth nondegenerate best-response structure of the Strategic Jacobian. Thus that machinery cannot justify the former general identification of Nash with a joint ground state. The ground-state and variational calculations retain their meaning for the scalar operator actually defined. \(\square\)
References#
Juan A. Ace\'bron, L. L. Bonilla, Conrad J. Pérez Vicente, Félix Ritort, and Renato Spigler. The Kuramoto model: a simple paradigm for synchronization phenomena. Reviews of Modern Physics, 77(1):137–185, 2005. doi:10.1103/RevModPhys.77.137.
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel. Constrained policy optimization. Proceedings of the 34th International Conference on Machine Learning, pages 22–31, 2017.
Shmuel Agmon. Lectures on Exponential Decay of Solutions of Second-Order Elliptic Equations: Bounds on Eigenfunctions of N-Body Schrödinger Operators. Volume 29 of Mathematical Notes. Princeton University Press, 1982. doi:10.1515/9781400853076.
Alexander A. Alemi, Ian Fischer, Joshua V. Dillon, and Kevin Murphy. Deep variational information bottleneck. arXiv preprint arXiv:1612.00410, 2016.
Eitan Altman. Constrained Markov Decision Processes. Chapman and Hall/CRC, 1999.
Shun-ichi Amari. Differential-Geometrical Methods in Statistics. Volume 28 of Lecture Notes in Statistics. Springer-Verlag, 1985. doi:10.1007/978-1-4612-5056-2.
Shun-ichi Amari. Natural gradient works efficiently in learning. Neural Computation, 10(2):251–276, 1998.
Shun-ichi Amari. Information Geometry and Its Applications. Springer, 2016.
W. Ambrose and I. M. Singer. A theorem on holonomy. Transactions of the American Mathematical Society, 75(3):428–443, 1953. doi:10.2307/1990687.
Vladimir I. Arnold. Mathematical Methods of Classical Mechanics. Springer-Verlag, 2nd edition, 1989.
Marshall Ball, Alon Rosen, Manuel Sabin, and Prashant Nalini Vasudevan. Proofs of useful work. In Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. 2017.
Adrien Bardes, Jean Ponce, and Yann LeCun. VICReg: variance-invariance-covariance regularization for self-supervised learning. In International Conference on Learning Representations. 2022.
Jacob D. Bekenstein. Black holes and entropy. Physical Review D, 7(8):2333–2346, 1973.
Marc G. Bellemare, Will Dabney, and Rémi Munos. A distributional perspective on reinforcement learning. In Proceedings of the 34th International Conference on Machine Learning, 449–458. 2017.
Richard Bellman. Dynamic Programming. Princeton University Press, Princeton, NJ, 1957.
Dionigi M. T. Benincasa and Fay Dowker. The scalar curvature of a causal set. Physical Review Letters, 104(18):181301, 2010. doi:10.1103/PhysRevLett.104.181301.
Charles H. Bennett. The thermodynamics of computation—a review. International Journal of Theoretical Physics, 21(12):905–940, 1982.
Felix Berkenkamp, Matteo Turchetta, Angela Schoellig, and Andreas Krause. Safe model-based reinforcement learning with stability guarantees. In Advances in Neural Information Processing Systems 30, 908–918. 2017.
Nicole Berline, Ezra Getzler, and Michèle Vergne. Heat Kernels and Dirac Operators. Volume 298 of Grundlehren der mathematischen Wissenschaften. Springer-Verlag, 1992. doi:10.1007/978-3-642-58088-8.
M. V. Berry. Quantal phase factors accompanying adiabatic changes. Proceedings of the Royal Society of London A, 392(1802):45–57, 1984. doi:10.1098/rspa.1984.0023.
Christopher M. Bishop. Pattern Recognition and Machine Learning. Information Science and Statistics. Springer, 2006. ISBN 978-0387310732.
David Bohm. A suggested interpretation of the quantum theory in terms of “hidden” variables. I. Physical Review, 85(2):166–179, 1952. doi:10.1103/PhysRev.85.166.
Luca Bombelli, Joohan Lee, David Meyer, and Rafael D. Sorkin. Space-time as a causal set. Physical Review Letters, 59(5):521–524, 1987. doi:10.1103/PhysRevLett.59.521.
Silvere Bonnabel. Stochastic gradient descent on riemannian manifolds. IEEE Transactions on Automatic Control, 58(9):2217–2229, 2013.
Michael M. Bronstein, Joan Bruna, Taco Cohen, and Petar Veličković. Geometric deep learning: grids, groups, graphs, geodesics, and gauges. arXiv preprint arXiv:2104.13478, 2021.
Romeo Brunetti, Klaus Fredenhagen, and Rainer Verch. The generally covariant locality principle: A new paradigm for local quantum physics. Communications in Mathematical Physics, 237:31–68, 2003. doi:10.1007/s00220-003-0815-7.
Krzysztof Burdzy, Robert Hołyst, and Peter March. Fleming-Viot processes and related distributions. Journal of Applied Probability, 37(1):131–144, 2000. doi:10.1239/jap/1014842272.
L. Lorne Campbell. An extended Čencov characterization of the information metric. Proceedings of the American Mathematical Society, 98(1):135–141, 1986.
Ran Canetti, Ben Riva, and Guy N. Rothblum. Practical delegation of computation using multiple servers. In Proceedings of the 18th ACM Conference on Computer and Communications Security, 445–454. 2011.
Eric A Carlen and Jan Maas. Gradient flow and entropy inequalities for quantum Markov semigroups with detailed balance. Journal of Functional Analysis, 273(5):1810–1869, 2017. doi:10.1016/j.jfa.2017.05.003.
Eric A. Carlen and Jan Maas. An analog of the 2-wasserstein metric in non-commutative probability under which the fermionic fokker-planck equation is gradient flow for the entropy. Communications in Mathematical Physics, 331:887–926, 2014.
Gunnar Carlsson. Topology and data. Bulletin of the American Mathematical Society, 46(2):255–308, 2009.
Kathryn Chaloner and Isabella Verdinelli. Bayesian experimental design: a review. Statistical Science, 10(3):273–304, 1995.
Bo Chang, Minmin Chen, Eldad Haber, and Ed H. Chi. AntisymmetricRNN: a dynamical system view on recurrent neural networks. International Conference on Learning Representations, 2019.
Sydney Chapman and T. G. Cowling. The Mathematical Theory of Non-uniform Gases. Cambridge University Press, 3rd edition, 1990. ISBN 978-0-521-40844-8.
Jeff Cheeger, Werner Muller, and Robert Schrader. On the curvature of piecewise flat spaces. Communications in Mathematical Physics, 92(3):405–454, 1984. doi:10.1007/BF01218616.
Lénaïc Chizat, Gabriel Peyré, Bernhard Schmitzer, and François-Xavier Vialard. An interpolating distance between optimal transport and Fisher-Rao metrics. Foundations of Computational Mathematics, 18(1):1–44, 2018. doi:10.1007/s10208-016-9331-y.
Lénaïc Chizat, Gabriel Peyré, Bernhard Schmitzer, and François-Xavier Vialard. Scaling algorithms for unbalanced optimal transport problems. Mathematics of Computation, 87(314):2563–2609, 2018.
Lénaïc Chizat, Gabriel Peyré, Bernhard Schmitzer, and François-Xavier Vialard. Unbalanced optimal transport: dynamic and kantorovich formulations. Journal of Functional Analysis, 274(11):3090–3123, 2018.
Shui-Nee Chow, Wen Huang, Yao Li, and Haomin Zhou. Fokker-Planck equations for a free energy functional or Markov process on a graph. Archive for Rational Mechanics and Analysis, 203(3):969–1008, 2012. doi:10.1007/s00205-011-0471-6.
Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh. A lyapunov-based approach to safe reinforcement learning. Advances in Neural Information Processing Systems 31, pages 8092–8101, 2018.
Taco S. Cohen, Maurice Weiler, Berkay Kicanaoglu, and Max Welling. Gauge equivariant convolutional networks and the icosahedral CNN. In Proceedings of the 36th International Conference on Machine Learning (ICML), volume 97 of Proceedings of Machine Learning Research, 1321–1330. PMLR, 2019. URL: http://proceedings.mlr.press/v97/cohen19d.html.
Taco S. Cohen and Max Welling. Group equivariant convolutional networks. In Proceedings of the 33rd International Conference on Machine Learning, 2990–2999. 2016.
Thomas M. Cover and Joy A. Thomas. Elements of Information Theory. John Wiley & Sons, Hoboken, NJ, 2nd edition, 2006.
Gavin E. Crooks. Entropy production fluctuation theorem and the nonequilibrium work relation for free energy differences. Physical Review E, 60(3):2721–2726, 1999.
Marco Cuturi. Sinkhorn distances: lightspeed computation of optimal transport. Advances in Neural Information Processing Systems 26, pages 2292–2300, 2013.
Will Dabney, Mark Rowland, Marc G. Bellemare, and Rémi Munos. Distributional reinforcement learning with quantile regression. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32. 2018.
Pim de Haan, Maurice Weiler, Taco Cohen, and Max Welling. Gauge equivariant mesh CNNs: anisotropic convolutions on geometric graphs. In International Conference on Learning Representations (ICLR). 2021. URL: https://openreview.net/forum?id=Jnspzp-oIZE.
Pierre Del Moral. Feynman-Kac Formulae: Genealogical and Interacting Particle Systems with Applications. Probability and Its Applications. Springer, 2004. doi:10.1007/978-1-4684-9393-1.
Pierre Del Moral. Feynman-Kac Formulae: Genealogical and Interacting Particle Systems with Applications. Probability and its Applications. Springer, 2004. doi:10.1007/978-1-4684-9393-1.
Amir Dembo and Ofer Zeitouni. Large Deviations Techniques and Applications. Volume 38 of Applications of Mathematics. Springer, 2nd edition, 1998. ISBN 978-0-387-98406-8.
Manfredo Perdigão do Carmo. Riemannian Geometry. Birkhäuser, 1992.
Albert Einstein. Über die von der molekularkinetischen theorie der wärme geforderte bewegung von in ruhenden flüssigkeiten suspendierten teilchen. Annalen der Physik, 322(8):549–560, 1905. English translation: On the Motion of Small Particles Suspended in a Stationary Liquid. doi:10.1002/andp.19053220806.
Albert Einstein. Die Feldgleichungen der Gravitation. Sitzungsberichte der Königlich Preussischen Akademie der Wissenschaften, pages 844–847, 1915.
Lawrence C. Evans. Partial Differential Equations. American Mathematical Society, 2nd edition, 2010.
Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. Proceedings of the 34th International Conference on Machine Learning, pages 1126–1135, 2017.
Luca Franceschi, Paolo Frasconi, Saverio Salzo, Riccardo Grazzi, and Massimiliano Pontil. Bilevel programming for hyperparameter optimization and meta-learning. Proceedings of the 35th International Conference on Machine Learning, pages 1568–1577, 2018.
Karl Friston. The free-energy principle: a unified brain theory? Nature Reviews Neuroscience, 11(2):127–138, 2010.
Karl Friston, Thomas FitzGerald, Francesco Rigoli, Philipp Schwartenbeck, and Giovanni Pezzulo. Active inference: a process theory. Neural Computation, 29(1):1–49, 2017.
Fabian Fuchs, Daniel E. Worrall, Volker Fischer, and Max Welling. SE(3)-transformers: 3D roto-translation equivariant attention networks. In Advances in Neural Information Processing Systems (NeurIPS), volume 33, 1970–1981. 2020. URL: https://proceedings.neurips.cc/paper/2020/hash/15231a7ce4ba789d13b722cc5c955834-Abstract.html.
Drew Fudenberg and Jean Tirole. Game Theory. MIT Press, 1991.
Octavian-Eugen Ganea, Gary Bécigneul, and Thomas Hofmann. Hyperbolic neural networks. In Advances in Neural Information Processing Systems 31, 5345–5355. 2018.
Sheldon L. Glashow. Partial-symmetries of weak interactions. Nuclear Physics, 22(4):579–588, 1961.
David E. Goldberg. Genetic Algorithms in Search, Optimization, and Machine Learning. Addison-Wesley, 1989. ISBN 978-0201157673.
Jeffrey Goldstone. Field theories with superconductor solutions. Il Nuovo Cimento, 19(1):154–164, 1961. doi:10.1007/BF02812722.
Vittorio Gorini, Andrzej Kossakowski, and E. C. G. Sudarshan. Completely positive dynamical semigroups of N-level systems. Journal of Mathematical Physics, 17(5):821–825, 1976.
Vittorio Gorini, Andrzej Kossakowski, and E. C. G. Sudarshan. Completely positive dynamical semigroups of N-level systems. Journal of Mathematical Physics, 17:821–825, 1976.
Sam Greydanus, Misko Dzamba, and Jason Yosinski. Hamiltonian neural networks. Advances in Neural Information Processing Systems 32, pages 15379–15389, 2019.
Alexander Grigor'yan. Heat Kernel and Analysis on Manifolds. American Mathematical Society, 2009.
Mikhail Gromov. Metric Structures for Riemannian and Non-Riemannian Spaces. Volume 152 of Progress in Mathematics. Birkhäuser, Boston, 1999. ISBN 978-0-8176-3898-5. Based on the 1981 French original.
David J. Gross and Frank Wilczek. Ultraviolet behavior of non-Abelian gauge theories. Physical Review Letters, 30(26):1343–1346, 1973. doi:10.1103/PhysRevLett.30.1343.
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron Courville. Improved training of wasserstein gans. Advances in Neural Information Processing Systems 30, pages 5767–5777, 2017.
David Ha and Jürgen Schmidhuber. World models. In Advances in Neural Information Processing Systems 31. 2018.
Rudolf Haag. Local Quantum Physics: Fields, Particles, Algebras. Springer, 2nd edition, 1992. doi:10.1007/978-3-642-61458-3.
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th International Conference on Machine Learning, 1861–1870. 2018.
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. Dream to control: learning behaviors by latent imagination. In International Conference on Learning Representations. 2020.
Brian C. Hall. Lie Groups, Lie Algebras, and Representations: An Elementary Introduction. Volume 222 of Graduate Texts in Mathematics. Springer, 2nd edition, 2015. ISBN 978-3-319-13466-6.
Richard S. Hamilton. Three-manifolds with positive Ricci curvature. Journal of Differential Geometry, 17(2):255–306, 1982.
W. K. Hastings. Monte carlo sampling methods using markov chains and their applications. Biometrika, 57(1):97–109, 1970. doi:10.1093/biomet/57.1.97.
Stephen W. Hawking. Particle creation by black holes. Communications in Mathematical Physics, 43(3):199–220, 1975.
Stephen W. Hawking and George F. R. Ellis. The Large Scale Structure of Space-Time. Cambridge University Press, 1973. ISBN 978-0-521-09906-6.
Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (GELUs). arXiv preprint arXiv:1606.08415, 2016. URL: https://arxiv.org/abs/1606.08415.
Sergio Hernández, Guillem Durán, and José M. Amigó. General algorithmic search. 2017. Also published in: Recent Trends in Chaotic, Nonlinear and Complex Dynamics, World Scientific, 2017. arXiv:1705.08691.
Sergio Hernández Cerezo and Guillem Durán Ballester. Fractal AI: a fragile theory of intelligence. 2018. arXiv:1803.05049.
Sergio Hernández Cerezo, Guillem Durán Ballester, and Spiros Baxevanakis. Solving Atari games using fractals and entropy. 2018. arXiv:1807.01081.
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems 33, 6840–6851. 2020.
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
John H. Holland. Adaptation in Natural and Artificial Systems: An Introductory Analysis with Applications to Biology, Control, and Artificial Intelligence. MIT Press, 2nd edition, 1992. ISBN 978-0262581110.
Gerard 't Hooft. Dimensional reduction in quantum gravity. arXiv:gr-qc/9310026, 1993.
Roger A. Horn and Charles R. Johnson. Matrix Analysis. Cambridge University Press, 2nd edition, 2012. ISBN 978-0-521-54823-6.
Hannes Hornischer, Paul J. Pritz, Johannes Pritz, Marco G. Mazza, and Margarete Boos. Modeling of human group coordination. Physical Review Research, 4(2):023037, 2022. doi:10.1103/PhysRevResearch.4.023037.
Timothy Hospedales, Antreas Antoniou, Paul Micaelli, and Amos Storkey. Meta-learning in neural networks: a survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(9):5149–5169, 2022.
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel. VIME: variational information maximizing exploration. Advances in Neural Information Processing Systems 29, pages 1109–1117, 2016.
Shengyi Huang and others. The 37 implementation details of proximal policy optimization. ICLR Blog Track, 2022. https://iclr-blog-track.github.io/2022/03/25/ppo-implementation-details/.
Aapo Hyvärinen and Hiroshi Morioka. Nonlinear ICA of temporally dependent stationary sources. Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, pages 460–469, 2017.
John David Jackson. Classical Electrodynamics. Wiley, 3rd edition, 1999. ISBN 978-0-471-30932-1.
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton. Adaptive mixtures of local experts. Neural Computation, 3(1):79–87, 1991.
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z. Leibo, David Silver, and Koray Kavukcuoglu. Reinforcement learning with unsupervised auxiliary tasks. arXiv preprint arXiv:1611.05397, 2017.
Max Jaderberg, Karen Simonyan, Andrew Zisserman, and Koray Kavukcuoglu. Spatial transformer networks. In Advances in Neural Information Processing Systems 28, 2017–2025. 2015.
John B. Johnson. Thermal agitation of electricity in conductors. Physical Review, 32(1):97–109, 1928.
Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra. Planning and acting in partially observable stochastic domains. Artificial Intelligence, 101(1-2):99–134, 1998.
Daniel Kahneman. Thinking, Fast and Slow. Farrar, Straus and Giroux, 2011.
Hilbert J. Kappen. Path integrals and symmetry breaking for optimal control theory. Journal of Statistical Mechanics: Theory and Experiment, 2005(11):P11011, 2005.
Alex Kendall, Yarin Gal, and Roberto Cipolla. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 7482–7491. 2018.
James Kennedy and Russell Eberhart. Particle swarm optimization. In Proceedings of ICNN'95 - International Conference on Neural Networks, volume 4, 1942–1948. IEEE, 1995. doi:10.1109/ICNN.1995.488968.
Hassan K. Khalil. Nonlinear Systems. Prentice Hall, Upper Saddle River, NJ, 3rd edition, 2002.
Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. In Advances in Neural Information Processing Systems 33, 18661–18673. 2020.
Michael Kirchhoff, Thomas Parr, Ensor Palacios, Karl Friston, and Julian Kiverstein. The Markov blankets of life: autonomy, active inference and the free energy principle. Journal of The Royal Society Interface, 2018.
Scott Kirkpatrick, C. Daniel Gelatt, and Mario P. Vecchi. Optimization by simulated annealing. Science, 220(4598):671–680, 1983. doi:10.1126/science.220.4598.671.
Shoshichi Kobayashi and Katsumi Nomizu. Foundations of Differential Geometry, Volume 1. Volume 1. Interscience Publishers, 1963.
J. Zico Kolter and Aleksander Mądry. Adversarial robustness - theory and practice. NeurIPS Tutorial, 2018.
Vijay R. Konda and John N. Tsitsiklis. Actor-critic algorithms. In Advances in Neural Information Processing Systems 12, 1008–1014. 2000.
Ilya Kostrikov, Denis Yarats, and Rob Fergus. Image augmentation is all you need: regularizing deep reinforcement learning from pixels. In International Conference on Learning Representations. 2021.
Ryogo Kubo. The fluctuation-dissipation theorem. Reports on Progress in Physics, 29(1):255–284, 1966. doi:10.1088/0034-4885/29/1/306.
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning. In Advances in Neural Information Processing Systems 33, 1179–1191. 2020.
L. D. Landau and E. M. Lifshitz. Statistical Physics, Part 1. Volume 5 of Course of Theoretical Physics. Butterworth-Heinemann, 3rd edition, 1980. ISBN 978-0-7506-3372-7.
Rolf Landauer. Irreversibility and heat generation in the computing process. IBM Journal of Research and Development, 5(3):183–191, 1961.
J. P. LaSalle. Stability theory for ordinary differential equations. Journal of Differential Equations, 4(1):57–65, 1968. doi:10.1016/0022-0396(68)90048-X.
Joseph P. LaSalle. The extent of asymptotic stability. Proceedings of the National Academy of Sciences, 46(3):363–365, 1960.
Michael Laskin, Aravind Srinivas, and Pieter Abbeel. CURL: contrastive unsupervised representations for reinforcement learning. In Proceedings of the 37th International Conference on Machine Learning, 5639–5650. 2020.
John M. Lee. Introduction to Smooth Manifolds. Springer, 2nd edition, 2012.
John M. Lee. Introduction to Riemannian Manifolds. Springer, 2nd edition, 2018.
John M. Lee. Introduction to Riemannian Manifolds. Volume 176 of Graduate Texts in Mathematics. Springer, 2nd edition, 2018. doi:10.1007/978-3-319-91755-9.
Ben Leimkuhler and Charles Matthews. Molecular Dynamics: With Deterministic and Stochastic Numerical Methods. Springer, 2015.
Ben Leimkuhler and Charles Matthews. Molecular Dynamics: With Deterministic and Stochastic Numerical Methods. Volume 39 of Interdisciplinary Applied Mathematics. Springer, 2015. doi:10.1007/978-3-319-16375-8.
Ben Leimkuhler and Charles Matthews. Efficient molecular dynamics using geodesic integration and solvent-solute splitting. Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences, 472(2189):20160138, 2016. doi:10.1098/rspa.2016.0138.
Leonid A. Levin. Universal sequential search problems. Problems of Information Transmission, 9(3):265–266, 1973.
David K. Lewis. Convention: A Philosophical Study. Harvard University Press, Cambridge, MA, 1969. ISBN 978-0-674-17003-4.
Matthias Liero, Alexander Mielke, and Giuseppe Savaré. Optimal entropy-transport problems and a new hellinger-kantorovich distance between positive measures. Inventiones Mathematicae, 211:969–1117, 2018.
Long-Ji Lin. Self-improving reactive agents based on reinforcement learning, planning and teaching. Machine Learning, 8(3-4):293–321, 1992.
Long-Ji Lin. Reinforcement learning for robots using neural networks. PhD thesis, Carnegie Mellon University, 1993.
Göran Lindblad. On the generators of quantum dynamical semigroups. Communications in Mathematical Physics, 48(2):119–130, 1976.
Göran Lindblad. On the generators of quantum dynamical semigroups. Communications in Mathematical Physics, 48:119–130, 1976.
Dennis V. Lindley. On a measure of the information provided by an experiment. The Annals of Mathematical Statistics, 27(4):986–1005, 1956.
Michael L. Littman. Markov games as a framework for multi-agent reinforcement learning. Proceedings of the 11th International Conference on Machine Learning, pages 157–163, 1994.
Aaron Lou, Isay Katsman, Qingxuan Jiang, Serge Belongie, Ser-Nam Lim, and Christopher De Sa. Differentiating through the Fréchet mean. Proceedings of the 37th International Conference on Machine Learning, pages 6393–6403, 2020.
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems 30, 6379–6390. 2017.
Christian Léonard. A survey of the schrödinger problem and some of its connections with optimal transport. Discrete and Continuous Dynamical Systems, 34(4):1533–1574, 2014.
Jan Maas. Gradient flows of the entropy for finite Markov chains. Journal of Functional Analysis, 261(8):2250–2292, 2011. doi:10.1016/j.jfa.2011.06.009.
Erwin Madelung. Quantentheorie in hydrodynamischer Form. Zeitschrift für Physik, 40(3–4):322–326, 1927. doi:10.1007/BF01400372.
James Martens. New insights and perspectives on the natural gradient method. Journal of Machine Learning Research, 21(146):1–76, 2020.
James Martens and Roger Grosse. Optimizing neural networks with kronecker-factored approximate curvature. Proceedings of the 32nd International Conference on Machine Learning, pages 2408–2417, 2015.
Humberto R. Maturana and Francisco J. Varela. Autopoiesis and Cognition: The Realization of the Living. Volume 42 of Boston Studies in the Philosophy of Science. D. Reidel Publishing Company, 1980.
James Clerk Maxwell. Theory of Heat. Longmans, Green, and Co., London, 1871.
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas. Communication-efficient learning of deep networks from decentralized data. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, 1273–1282. 2017.
Pankaj Mehta and David J. Schwab. An exact mapping between the variational renormalization group and deep learning. arXiv preprint arXiv:1410.3831, 2014.
Nicholas Metropolis, Arianna W. Rosenbluth, Marshall N. Rosenbluth, Augusta H. Teller, and Edward Teller. Equation of state calculations by fast computing machines. The Journal of Chemical Physics, 21(6):1087–1092, 1953. doi:10.1063/1.1699114.
David A. Meyer. The Dimension of Causal Sets. PhD thesis, Massachusetts Institute of Technology, 1988.
Sean P. Meyn and Richard L. Tweedie. Markov Chains and Stochastic Stability. Cambridge University Press, 2nd edition, 2012. doi:10.1017/CBO9780511626630.
Alexander Mielke. A gradient structure for reaction-diffusion systems and for energy-drift-diffusion systems. Nonlinearity, 24(4):1329, 2011. doi:10.1088/0951-7715/24/4/016.
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. Spectral normalization for generative adversarial networks. International Conference on Learning Representations, 2018.
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 2015.
Volodymyr Mnih and others. Human-level control through deep reinforcement learning. Nature, 518:529–533, 2015.
Jan Myrheim. Statistical geometry. Technical Report CERN-TH-2538, CERN, 1978.
Mikio Nakahara. Geometry, Topology and Physics. Taylor & Francis, 2nd edition, 2003. ISBN 978-0-7503-0606-5.
Satoshi Nakamoto. Bitcoin: a peer-to-peer electronic cash system. https://bitcoin.org/bitcoin.pdf, 2008.
Radford M. Neal. MCMC using hamiltonian dynamics. In Steve Brooks, Andrew Gelman, Galin Jones, and Xiao-Li Meng, editors, Handbook of Markov Chain Monte Carlo, chapter 5, pages 113–162. Chapman and Hall/CRC, 2011.
Andrew Y. Ng, Daishi Harada, and Stuart Russell. Policy invariance under reward transformations: theory and application to reward shaping. In Proceedings of the 16th International Conference on Machine Learning, 278–287. 1999.
Maximilian Nickel and Douwe Kiela. Poincaré embeddings for learning hierarchical representations. In Advances in Neural Information Processing Systems 30, 6338–6347. 2017.
Jorge Nocedal and Stephen J. Wright. Numerical Optimization. Springer, 2nd edition, 2006.
Harry Nyquist. Thermal agitation of electric charge in conductors. Physical Review, 32(1):110–113, 1928.
Reza Olfati-Saber and Richard M. Murray. Consensus problems in networks of agents with switching topology and time-delays. IEEE Transactions on Automatic Control, 49(9):1520–1533, 2004. doi:10.1109/TAC.2004.834113.
Lars Onsager and Stefan Machlup. Fluctuations and irreversible processes. Physical Review, 91(6):1505–1512, 1953.
Pedro A. Ortega and Daniel A. Braun. Thermodynamics as a theory of decision-making with information-processing costs. Proceedings of the Royal Society A, 2013.
Konrad Osterwalder and Robert Schrader. Axioms for Euclidean Green's functions. Communications in Mathematical Physics, 31(2):83–112, 1973. doi:10.1007/BF01645738.
Konrad Osterwalder and Robert Schrader. Axioms for Euclidean Green's functions II. Communications in Mathematical Physics, 42(3):281–305, 1975. doi:10.1007/BF01608978.
Pierre-Yves Oudeyer, Frédéric Kaplan, and Verena V. Hafner. Intrinsic motivation systems for autonomous mental development. IEEE Transactions on Evolutionary Computation, 11(2):265–286, 2007.
Juan M. R. Parrondo, Jordan M. Horowitz, and Takahiro Sagawa. Thermodynamics of information. Nature Physics, 11(2):131–139, 2015.
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell. Curiosity-driven exploration by self-supervised prediction. In Proceedings of the 34th International Conference on Machine Learning, 2778–2787. 2017.
Judea Pearl. Causality: Models, Reasoning, and Inference. Cambridge University Press, 2nd edition, 2009.
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli. Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice. Advances in Neural Information Processing Systems 30, pages 4785–4795, 2017.
Grigori Perelman. The entropy formula for the Ricci flow and its geometric applications. arXiv:math/0211159, 2002.
Michael E. Peskin and Daniel V. Schroeder. An Introduction to Quantum Field Theory. Addison-Wesley, 1995.
F. Peter and H. Weyl. Die Vollständigkeit der primitiven Darstellungen einer geschlossenen kontinuierlichen Gruppe. Mathematische Annalen, 97:737–755, 1927. doi:10.1007/BF01447892.
Dean A. Pomerleau. Efficient training of artificial neural networks for autonomous navigation. Neural Computation, 3(1):88–97, 1991.
David Premack and Guy Woodruff. Does the chimpanzee have a theory of mind? Behavioral and Brain Sciences, 1(4):515–526, 1978. doi:10.1017/S0140525X00076512.
Lawrence R. Rabiner. A tutorial on hidden Markov models and selected applications in speech recognition. Proceedings of the IEEE, 77(2):257–286, 1989.
Carl Edward Rasmussen and Christopher K. I. Williams. Gaussian Processes for Machine Learning. MIT Press, 2006.
Amal Kumar Raychaudhuri. Relativistic cosmology. I. Physical Review, 98(4):1123–1126, 1955. doi:10.1103/PhysRev.98.1123.
Tullio Regge. General relativity without coordinates. Il Nuovo Cimento, 19(3):558–571, 1961. doi:10.1007/BF02733251.
Hannes Risken. The Fokker-Planck Equation: Methods of Solution and Applications. Springer-Verlag, 2nd edition, 1996.
Christian P. Robert and George Casella. Monte Carlo Statistical Methods. Springer, 2nd edition, 2004. doi:10.1007/978-1-4757-4145-2.
Gareth O. Roberts and Richard L. Tweedie. Exponential convergence of langevin distributions and their discrete approximations. Bernoulli, 2(4):341–363, 1996. doi:10.2307/3318418.
Steven Rosenberg. The Laplacian on a Riemannian Manifold. Cambridge University Press, 1997.
Abdus Salam. Weak and electromagnetic interactions. In N. Svartholm, editor, Elementary Particle Theory: Proceedings of the Nobel Symposium, 367–377. Almqvist & Wiksell, 1968.
Andrew M. Saxe, James L. McClelland, and Surya Ganguli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. arXiv preprint arXiv:1312.6120, 2014.
Jürgen Schmidhuber. Formal theory of creativity, fun, and intrinsic motivation (1990–2010). IEEE Transactions on Autonomous Mental Development, 2(3):230–247, 2010.
Florian Schroff, Dmitry Kalenichenko, and James Philbin. FaceNet: a unified embedding for face recognition and clustering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 815–823. 2015.
John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, and Pieter Abbeel. Trust region policy optimization. Proceedings of the 32nd International Conference on Machine Learning, pages 1889–1897, 2015.
Max Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm, Aaron Courville, and Philip Bachman. Data-efficient reinforcement learning with self-predictive representations. In International Conference on Learning Representations. 2021.
Mark R. Sepanski. Compact Lie Groups. Volume 235 of Graduate Texts in Mathematics. Springer, 2007. ISBN 978-0-387-30263-8.
Jean-Pierre Serre. Linear Representations of Finite Groups. Volume 42 of Graduate Texts in Mathematics. Springer, 1977. ISBN 978-0-387-90190-9.
Lloyd S. Shapley. Stochastic games. Proceedings of the National Academy of Sciences, 39(10):1095–1100, 1953.
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: the sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017.
Yuhui Shi and Russell Eberhart. A modified particle swarm optimizer. In 1998 IEEE International Conference on Evolutionary Computation Proceedings, 69–73. IEEE, 1998. doi:10.1109/ICEC.1998.699146.
Herbert A. Simon. A behavioral model of rational choice. The Quarterly Journal of Economics, 69(1):99–118, 1955.
Herbert A. Simon. Models of Bounded Rationality. MIT Press, Cambridge, MA, 1982.
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. Proceedings of the 32nd International Conference on Machine Learning, pages 2256–2265, 2015.
Rafael D. Sorkin. Causal sets: discrete gravity. In Andrés Gomberoff and Donald Marolf, editors, Lectures on Quantum Gravity, Series of the Centro de Estudios Científicos, 305–327. Springer, 2005. doi:10.1007/0-387-24992-3_7.
Peter Spirtes, Clark Glymour, and Richard Scheines. Causation, Prediction, and Search. MIT Press, 2nd edition, 2000.
Aart J. Stam. Some inequalities satisfied by the quantities of information of Fisher and Shannon. Information and Control, 2(2):101–112, 1959.
Susanne Still, David A. Sivak, Anthony J. Bell, and Gavin E. Crooks. Thermodynamics of prediction. Physical Review Letters, 109(12):120604, 2012.
Adam Stooke, Joshua Achiam, and Pieter Abbeel. Responsive safety in reinforcement learning by PID lagrangian methods. Proceedings of the 37th International Conference on Machine Learning, pages 9133–9143, 2020.
Robert S. Strichartz. Analysis of the Laplacian on the complete Riemannian manifold. Journal of Functional Analysis, 52(1):48–79, 1983. doi:10.1016/0022-1236(83)90090-3.
Steven H. Strogatz. Nonlinear Dynamics and Chaos. Westview Press, 2nd edition, 2015.
Steven H. Strogatz. Nonlinear Dynamics and Chaos. CRC Press, 2nd edition, 2018.
Andrew Strominger and Cumrun Vafa. Microscopic origin of the Bekenstein-Hawking entropy. Physics Letters B, 379(1-4):99–104, 1996.
Leonard Susskind. The world as a hologram. Journal of Mathematical Physics, 36(11):6377–6396, 1995.
Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. MIT Press, 2nd edition, 2018.
Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. MIT Press, 2nd edition, 2018. ISBN 978-0262039246.
Leo Szilard. Über die Entropieverminderung in einem thermodynamischen System bei Eingriffen intelligenter Wesen. Zeitschrift für Physik, 53:840–856, 1929.
Alain-Sol Sznitman. Topics in propagation of chaos. In École d'Été de Probabilités de Saint-Flour XIX — 1989, volume 1464 of Lecture Notes in Mathematics, pages 165–251. Springer, 1991. doi:10.1007/BFb0085169.
Naftali Tishby, Fernando C. Pereira, and William Bialek. The information bottleneck method. arXiv preprint physics/0004057, 2000.
Emanuel Todorov. Efficient computation of optimal actions. Proceedings of the National Academy of Sciences, 106(28):11478–11483, 2009.
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning. In Advances in Neural Information Processing Systems 30, 6306–6315. 2017.
Soledad Villar, David W. Hogg, Kate Storey-Fisher, Weichi Yao, and Ben Blum-Smith. Scalars are universal: equivariant machine learning, structured like classical physics. In Advances in Neural Information Processing Systems (NeurIPS), volume 34, 28848–28863. 2021. URL: https://proceedings.neurips.cc/paper/2021/hash/f1b0775946bc0329b35b823b86eeb5f5-Abstract.html.
Robert M. Wald. General Relativity. University of Chicago Press, Chicago, 1984.
Maurice Weiler and Gabriele Cesa. General E(2)-equivariant steerable CNNs. In Advances in Neural Information Processing Systems (NeurIPS), volume 32. 2019. URL: https://proceedings.neurips.cc/paper/2019/hash/45d6637b718d0f24a237069fe41b0db4-Abstract.html.
Steven Weinberg. A model of leptons. Physical Review Letters, 19(21):1264–1266, 1967.
Steven Weinberg. The Quantum Theory of Fields, Volume I: Foundations. Cambridge University Press, 1995.
Steven Weinberg. The Quantum Theory of Fields, Volume II: Modern Applications. Cambridge University Press, 1996. doi:10.1017/CBO9781139644174.
Gregor Wentzel. Eine Verallgemeinerung der Quantenbedingungen für die Zwecke der Wellenmechanik. Zeitschrift für Physik, 38(6–7):518–529, 1926. doi:10.1007/BF01397171.
Hermann Weyl. The Classical Groups: Their Invariants and Representations. Princeton University Press, 2nd edition, 1946. ISBN 978-0-691-05756-9.
Hassler Whitney. Differentiable manifolds. Annals of Mathematics, 37(3):645–680, 1936.
Arthur S. Wightman. Quantum field theory in terms of vacuum expectation values. Physical Review, 101(2):860–866, 1956. doi:10.1103/PhysRev.101.860.
Kenneth G. Wilson. Confinement of quarks. Physical Review D, 10(8):2445–2459, 1974. doi:10.1103/PhysRevD.10.2445.
Eugene Wong. Stochastic processes in information and dynamical systems. Dover Publications, 2012.
Chen-Ning Yang and Robert L. Mills. Conservation of isotopic spin and isotopic gauge invariance. Physical Review, 96(1):191–195, 1954. doi:10.1103/PhysRev.96.191.
Hideki Yukawa. On the interaction of elementary particles. I. Proceedings of the Physico-Mathematical Society of Japan, 17:48–57, 1935. doi:10.11429/ppmsj1919.17.0_48.
Paul L. Zador. Asymptotic quantization error of continuous signals and the quantization dimension. IEEE Transactions on Information Theory, 28(2):139–149, 1982.
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. Barlow twins: self-supervised learning via redundancy reduction. In Proceedings of the 38th International Conference on Machine Learning, 12310–12320. 2021.
Jiaheng Zhang, Zhiyong Fang, Yupeng Zhang, and Dawn Song. Zero knowledge proofs for decision tree predictions and accuracy. In Proceedings of the ACM SIGSAC Conference on Computer and Communications Security, 2039–2053. 2020.
Bernt Øksendal. Stochastic Differential Equations: An Introduction with Applications. Springer-Verlag, 6th edition, 2003.
Nikolai N. Čencov. Statistical Decision Rules and Optimal Inference. American Mathematical Society, 1982.
Mars Climate Orbiter Mishap Investigation Board. Mars climate orbiter mishap investigation board phase i report. Technical Report, NASA, November 1999. ftp://ftp.hq.nasa.gov/pub/pao/reports/1999/MCO_report.pdf.
N. G. van Kampen. Stochastic Processes in Physics and Chemistry. North-Holland, revised edition, 1992. ISBN 978-0-444-89349-8.