Appendix F: Loss Terms Reference#
TLDR#
Centralizes 37 loss functions and objectives defined throughout Volume 1 for quick engineering reference.
Curved formulations: All distance-based losses use the Riemannian metric \(G_{ij}\) from the capacity-constrained geometry (Section 18). Each entry includes a “Flat limit” showing the standard ML formula recovered when \(G_{ij} \to \delta_{ij}\).
Organized by domain: TopoEncoder losses, supervised topology, control objectives, reward fields, multi-agent, barriers, belief dynamics, self-supervised, imitation, geometric consistency, and meta-learning.
Each loss includes: formula, parameters, units, purpose, and source section cross-reference.
Use this as a single reference when implementing the Fragile Agent training pipeline.
F.1 TopoEncoder Losses#
These losses train the Attentive Atlas representation stack (Section 3.2).
Definition 330 (F.1.1 (Reconstruction Loss))
Parameters:
\(x, \hat{x}\) – original and reconstructed observations
\(G_{ij}^{\text{obs}}(x)\) – metric tensor on observation space (learned or fixed)
Purpose: Ensures charted latents collectively preserve information for reconstruction, with metric-weighted distances.
Units: \([\mathrm{nat}]\) (when scaled appropriately) or metric-weighted MSE.
Flat limit: When \(G_{ij}^{\text{obs}} = \delta_{ij}\) (identity), recovers standard MSE: \(\|x - \hat{x}\|^2\).
Source: Section 3.2, Definition Definition 27
Definition 331 (F.1.2 (Vector Quantization Loss))
Parameters:
\(z_e\) – encoder output (pre-quantization)
\(z_q\) – quantized code embedding \(e_{K}\) (per-chart)
\(G_{ij}(z)\) – metric tensor on latent space
\(\beta\) – commitment weight (default 0.25)
Purpose: Stabilizes per-chart codebooks. The first term updates code vectors; the second term encourages the encoder to commit to nearby codes.
Flat limit: When \(G_{ij} = \delta_{ij}\), recovers standard VQ: \(\|z_q - z_e\|^2 + \beta\|z_e - z_q\|^2\).
Source: Section 3.2, Definition Definition 27
Definition 332 (F.1.3 (Routing Entropy Loss))
Parameters:
\(w_{bk}\) – router weights over charts
\(N_c\) – number of charts
Purpose: Raises per-sample routing entropy and discourages chart collapse. Batch-level usage and diversity losses are still needed to prevent dead charts.
Units: \([\mathrm{nat}]\)
Source: Section 3.2
Definition 333 (F.1.4 (Consistency Loss))
Parameters:
\(w^{\text{enc}}\) – encoder router weights
\(w^{\text{dec}}\) – decoder router weights
Purpose: Aligns encoder and decoder chart usage.
Units: \([\mathrm{nat}]\)
Source: Section 3.2
Definition 334 (F.1.5 (Window Loss / Grounding))
Parameters:
\(I(X;K) = H(K) - H(K|X)\) – mutual information between input and chart assignment
\(\epsilon_{\text{ground}}\) – grounding threshold
Purpose: Enforces the stable learning window by requiring chart assignments to carry information about inputs.
Units: \([\mathrm{nat}]\)
Source: Section 3.2
Definition 335 (F.1.6 (Per-Chart Code Entropy Loss))
Parameters:
\(C\) – code index within each chart
\(K\) – number of codes per chart
Purpose: Encourages each chart to use its codebook uniformly rather than collapsing to a subset.
Units: \([\mathrm{nat}]\)
Source: Section 3.2
Definition 336 (F.1.7 (Total TopoEncoder Loss))
Purpose: Compound loss enforcing sharp routing, charted quantization, and stable geometry.
Source: Section 3.2, Definition Definition 27
Definition 337 (F.1.8 (Jump Consistency Loss))
Parameters:
\(z_n^{(i)}\) – nuisance coordinate from chart \(i\)
\(\mathcal{J}_{i \to j}\) – learned jump operator from chart \(i\) to chart \(j\)
Purpose: Enforces consistency in chart overlaps by learning transitions between chart-local nuisance coordinates.
Units: Dimensionless (metric-weighted embedding distance).
Source: Section 7
F.2 Supervised Topology Losses#
These losses enforce geometric coherence of classification (Section 25).
Definition 338 (F.2.1 (Purity Loss / Conditional Entropy))
Parameters:
\(P(K=k) = \mathbb{E}_{x \sim \mathcal{D}}[w_k(x)]\) – marginal chart probability
\(H(Y \mid K=k) = -\sum_y P(Y=y \mid K=k) \log P(Y=y \mid K=k)\) – class entropy within chart \(k\)
Purpose: Measures how well charts separate classes. Low purity loss means each chart is associated with a single class. Equivalent to maximizing mutual information \(I(K; Y)\) since \(\mathcal{L}_{\text{purity}} = H(Y) - I(K; Y)\).
Units: \([\mathrm{nat}]\)
Source: Section 25.4, Definition Definition 113
Definition 339 (F.2.2 (Load Balance Loss))
Parameters:
\(\bar{w} = \mathbb{E}_{x \sim \mathcal{D}}[w(x)]\) – average router weight vector
\(N_c\) – number of charts
Purpose: Prevents “dead charts” (collapse to few charts). Encourages all charts to be used productively, addressing the expert-collapse problem in mixture-of-experts systems.
Units: \([\mathrm{nat}]\)
Source: Section 25.4, Definition Definition 114
Definition 340 (F.2.3 (Metric Contrastive Loss))
Parameters:
\(\mathcal{P}\) – set of sample pairs in batch
\(w_i, w_j\) – router weight vectors
\(m > 0\) – margin (minimum desired geodesic separation)
\(d_G^{\text{jump}}(z_i, z_j)\) – minimum geodesic jump cost between samples under metric \(G\)
Purpose: Enforces that different-class samples are geometrically far apart in geodesic jump distance. The weighting \(w_i^\top w_j\) focuses penalty on hard examples (high routing overlap despite different classes). The geodesic distance respects the curved manifold structure.
Units: \([\mathrm{nat}]\)
Flat limit: When \(G_{ij} = \delta_{ij}\), reduces to Euclidean jump distance.
Source: Section 25.4, Definition Definition 115
Definition 341 (F.2.4 (Route Alignment Loss))
Parameters:
\(w_k(x)\) – router weights for sample \(x\) and chart \(k\)
\(P(Y=\cdot \mid K=k)\) – per-chart class distributions
\(\text{CE}\) – cross-entropy loss
Purpose: Primary classification loss. The predicted class distribution (router-weighted average of per-chart distributions) must match the true label.
Units: \([\mathrm{nat}]\)
Source: Section 25.4, Definition Definition 116
Definition 342 (F.2.5 (Combined Supervised Topology Loss))
Typical hyperparameters:
Weight |
Typical Value |
Role |
|---|---|---|
\(\lambda_{\text{pur}}\) |
0.1 |
Chart purity |
\(\lambda_{\text{bal}}\) |
0.01 |
Load balancing |
\(\lambda_{\text{met}}\) |
0.01 |
Metric separation |
Purpose: Weighted combination enforcing chart purity, balanced usage, geometric separation, and prediction accuracy.
Source: Section 25.4, Definition Definition 117
Definition 343 (F.2.6 (Hierarchical Supervised Loss))
Parameters:
\(\ell \in \{0, \ldots, L\}\) – scale levels (bulk to boundary)
\(\alpha_\ell\) – per-scale weights (often \(\alpha_\ell = 1\) or decaying)
\(\mathcal{Y}_\ell\) – label space at scale \(\ell\) (coarse to fine)
Purpose: Enforces classification at multiple scales via stacked TopoEncoders. Coarse (bulk) layers distinguish broad categories; fine (boundary) layers distinguish leaf categories.
Units: \([\mathrm{nat}]\)
Source: Section 25.6, Definition Definition 120
F.3 Control and Value Objectives#
These objectives define control, value, and reward structure (Section 2, Section 18).
Definition 344 (F.3.1 (Cumulative Cost Functional))
Parameters:
\(\mathcal{L}_{\text{control}}\) – control/effort cost (KL penalty, action magnitude)
\(C(z_t, a_t)\) – task cost
Purpose: General optimal control objective under information/effort constraints. Specializes to KL-control and entropy-regularized RL.
Units: \([\mathrm{nat}]\)
Source: Section 2
Definition 345 (F.3.2 (Instantaneous Regularized Objective))
Parameters:
\(V(Z_t)\) – task-aligned cost-to-go (critic estimate)
\(\beta_K(-\log p_\psi(K_t))\) – macro codelength penalty (Occam’s razor for discrete state)
\(\beta_n, \beta_{\text{tex}}\) – residual regularization weights
\(T_c\) – cognitive temperature
\(\pi_0\) – prior policy
Purpose: Trades off task cost, representation complexity, and control effort, all in consistent units (nats).
Units: \([\mathrm{nat}]\)
Source: Section 3.2, Definition Definition 345
Definition 346 (F.3.3 (Monotonicity Surrogate Loss))
Purpose: Penalizes increases in the instantaneous objective \(F_t\) from one step to the next, encouraging trajectories that smoothly descend the objective landscape.
Units: \([\mathrm{nat}^2]\)
Source: Section 3.2
Definition 347 (F.3.4 (Closure Ratio Diagnostic))
Interpretation:
Ratio |
Meaning |
Action |
|---|---|---|
\(\ll 1\) |
Strong predictive law learned |
Success |
\(\approx 1\) |
No predictive law |
Increase model capacity |
\(> 1\) |
Worse than baseline |
Bug/degeneracy |
Purpose: Measures how much better the macro dynamics model predicts \(K_{t+1}\) compared to a marginal baseline. The gap estimates predictive information \(I(K_{t+1}; K_t, a_t)\).
Units: Dimensionless.
Source: Runtime routing diagnostics, Definition Definition 347
Definition 348 (F.3.5 (Causal Information Potential))
Parameters:
\(z, a\) – current state and action
\(z'\) – next state sampled from world model
\(\theta_W\) – world model parameters
\(\bar{P}\) – world model transition distribution
Purpose: Measures the Expected Information Gain about world model parameters from executing action \(a\) at state \(z\). High \(\Psi_{\text{causal}}\) indicates the outcome will resolve significant uncertainty about dynamics. Drives intrinsic motivation for exploration via Bayesian experimental design.
Units: \([\mathrm{nat}]\)
Source: Section 29, Definition Definition 166
F.4 Reward Field Objectives#
These define reward as a geometric object (Section 18).
Definition 349 (F.4.1 (Hodge Decomposition of Reward))
Components:
\(d\Phi\) – Gradient/conservative reward component (the control-loop cost critic is \(V=-\Phi\))
\(\delta\Psi\) – Solenoidal/rotational component (cyclic reward structure)
\(\eta\) – Harmonic component (topological cycles from manifold holes)
Purpose: Decomposes reward 1-form into orthogonal components. Separates optimizable value from inherently cyclic structure (e.g., Rock-Paper-Scissors).
Units: \([\Phi] = \mathrm{nat}\), \([\Psi] = \mathrm{nat} \cdot [\text{length}]^2\), \([\eta] = \mathrm{nat}/[\text{length}]\).
Source: Section 18.2, Theorem Theorem 11
Definition 350 (F.4.2 (Value Curl / Vorticity))
Properties:
Antisymmetric: \(\mathcal{F}_{ij} = -\mathcal{F}_{ji}\)
Satisfies Bianchi identity: \(d\mathcal{F} = 0\)
Gauge-invariant under \(\mathcal{R} \to \mathcal{R} + d\chi\)
Purpose: Detects non-conservative reward structure. Non-zero curl indicates orbiting strategies may be optimal. Diagnostic: \(\oint_\gamma \mathcal{R} = \int_\Sigma \mathcal{F} \, d\Sigma \neq 0\) implies non-conservative rewards.
Units: \([\mathcal{F}] = \mathrm{nat}/[\text{length}]^2\)
Source: Section 18.2, Definition Definition 97
Definition 351 (F.4.3 (Class-Conditioned Potential))
Parameters:
\(P(Y=y \mid K) = \text{softmax}(\Theta_{K,:})_y\) – learnable chart-to-class affinities
\(V_{\text{base}}(z, K)\) – unconditioned critic
\(\beta_{\text{class}} > 0\) – class temperature (inverse semantic diffusion)
Purpose: Shapes potential landscape so class-\(y\) regions become energy minima. Used for both classification (relaxation inference) and generation (Langevin sampling).
Units: \([V_y] = \mathrm{nat}\)
Source: Section 25.2, Definition Definition 109
F.5 Multi-Agent and Gauge Losses#
These govern multi-agent alignment (the inter-subjective metric chapter).
Definition 352 (F.5.1 (Synchronization Potential))
Parameters:
\(\beta\) – coupling strength
\(\mathcal{F}_{AB}\) – Locking curvature of the selected relative connection
Purpose: Penalizes curvature of the selected relative connection and can drive gauge locking under the hypotheses of the conditional strong-coupling proposition. It does not by itself synchronize the private metrics.
Units: \([\mathrm{nat}]\)
F.6 Metabolic and Information Losses#
These relate to computation cost and information bounds (Section 36, Section 33).
Definition 353 (F.6.1 (Ontological Stress))
Parameters:
\(z_{\text{tex}}^{(\ell)}\) – texture embedding at scale \(\ell\)
\(G_{ij}^{(\ell)}\) – metric tensor at scale \(\ell\)
Purpose: Measures predictability within texture across scales. High stress indicates ontological inadequacy—texture contains compressible structure that should have been captured by macro/nuisance. Dual to closure defect: closure measures micro-to-macro leakage; ontological stress measures within-texture predictability.
Units: Dimensionless (metric-weighted embedding norm).
Flat limit: When \(G_{ij}^{(\ell)} = \delta_{ij}\), recovers \(\sum_\ell \|z_{\text{tex}}^{(\ell)}\|^2\).
Source: Section 33
F.7 Barrier and Regularization Losses#
These losses enforce fundamental limits and trade-offs (Section 4).
Definition 354 (F.7.1 (Gradient Penalty Loss))
Parameters:
\(\hat{s}\) – interpolated samples between real and generated
\(V(\hat{s})\) – critic value at sample
\(G^{ij}(\hat{s})\) – inverse metric tensor (contravariant) at sample
\(K\) – target gradient norm (typically 1)
Purpose: Enforces Lipschitz constraint on the critic using the metric-induced norm. The covariant gradient norm \(\|\nabla_A V\|_G\) measures the gauge-invariant rate of change along geodesics. Prevents vanishing gradients in flat value regions (BarrierGap) and ensures smooth value landscape for stable learning.
Units: Dimensionless.
Flat limit: When \(G^{ij} = \delta^{ij}\) and \(A=0\), recovers \((\|\nabla_A V\|_2 - K)^2\).
Source: Section 4
Definition 355 (F.7.2 (Information-Control Loss))
Parameters:
\(\beta_K, \beta_n, \beta_{\text{tex}}\) – compression weights for macro/nuisance/texture
\(\mathfrak{D}(Z, A)\) – actuation cost (KL-control or action norm)
\(\gamma\) – control effort weight
Purpose: Balances the Information-Control Tradeoff (BarrierScat vs BarrierCap). High compression removes details needed for fine control; this loss finds the Pareto frontier.
Units: \([\mathrm{nat}]\)
Source: Section 4
Definition 356 (F.7.3 (Elastic Weight Consolidation))
Parameters:
\(F_i\) – diagonal Fisher Information for parameter \(i\)
\(\theta_i\) – current parameter value
\(\theta^*_{i,\text{old}}\) – parameter value from previous task
Purpose: Addresses the Stability-Plasticity Dilemma (BarrierVac vs BarrierPZ). High-sensitivity weights (large \(F_i\)) are constrained to preserve past learning; low-sensitivity weights can adapt freely.
Units: Dimensionless (parameter space distance, Fisher-weighted).
Source: Section 4
Definition 357 (F.7.4 (Bode Magnitude Loss))
Parameters:
\(\mathcal{F}(e_t)\) – Fourier transform of error signal
\(W(\omega)\) – frequency weighting function
Purpose: Addresses the Bode Sensitivity Integral (BarrierBode). Suppressing error in one frequency band amplifies it in another (waterbed effect). This loss explicitly chooses where to be sensitive vs. blind.
Units: \([\mathrm{nat}^2]\)
Source: Section 4
F.8 Belief Dynamics Losses#
These losses enforce belief update constraints (Section 11).
Definition 358 (F.8.1 (Quantum Speed Limit Loss))
Parameters:
\(d_G(z_{t+1}, z_t)\) – geodesic distance traveled in one step
\(v_{\max}\) – maximum allowed velocity in latent space
Purpose: Enforces the Quantum Speed Limit: belief cannot change faster than the Mandelstam-Tamm bound allows. Prevents unrealistic jumps in belief state.
Units: \([\mathrm{nat}^2]\)
Source: Section 11
Definition 359 (F.8.2 (Joint Prediction Loss))
Expanded in coordinates:
Parameters:
\(\hat{x}_{t+1}^A, \hat{x}_{t+1}^B\) – predictions from agents \(A\) and \(B\)
\(x_{t+1}\) – actual next observation
\(G_{ij}^{\text{obs}}\) – metric tensor on observation space
\(\Psi_{\text{sync}}\) – synchronization potential
\(\beta\) – coupling strength
Purpose: Multi-agent world model training. Both agents must predict accurately (measured under the observation-space metric), and their representations must synchronize (gauge lock).
Units: Metric-weighted prediction error + \([\mathrm{nat}]\).
Flat limit: When \(G_{ij}^{\text{obs}} = \delta_{ij}\), recovers \(\|\hat{x}^A - x\|^2 + \|\hat{x}^B - x\|^2 + \beta\Psi_{\text{sync}}\).
F.9 Self-Supervised and Contrastive Losses#
These losses prevent representation collapse without requiring labels (Section 3).
Definition 360 (F.9.1 (VICReg Loss))
Components:
Parameters:
\(z, z'\) – embeddings of two augmented views of the same input
\(G_{ij}(z)\) – metric tensor on embedding space
\(\gamma\) – variance threshold (typically 1)
\(\lambda, \mu, \nu\) – component weights (typically \(\lambda = 25\), \(\mu = \nu = 1\))
Purpose: Prevents representation collapse without negative samples. Invariance pulls augmented views together (geodesic distance); variance prevents dimension collapse (metric-aware); covariance decorrelates dimensions (whitened by metric).
Units: Dimensionless.
Flat limit: When \(G_{ij} = \delta_{ij}\), recovers standard VICReg with \(\|z - z'\|^2\).
Source: Section 3 (GeomCheck, Node 6)
Definition 361 (F.9.2 (InfoNCE Loss))
Parameters:
\(z_t, z_{t+k}\) – embeddings at current and future timesteps
\(z_j\) – negative samples (other timesteps or other sequences)
\(d_G(z, z')\) – geodesic distance under metric \(G\)
\(\tau\) – temperature parameter (squared distance scale)
Purpose: Contrastive predictive coding with geodesic similarity. Anchors macro latents to temporal structure by maximizing mutual information between present and future representations. The geodesic kernel \(k_G(z, z') = \exp(-d_G(z,z')^2/\tau)\) respects the curved geometry of the latent manifold.
Units: \([\mathrm{nat}]\)
Flat limit: When \(G_{ij} = \delta_{ij}\) and using \(\text{sim}(z,z') = -\|z-z'\|^2\), recovers standard InfoNCE with Gaussian kernel.
Source: Section 3 (GeomCheck, Node 6)
F.10 Imitation and Distillation Losses#
These losses train policies from demonstrations or teacher models (Section 25).
Definition 362 (F.10.1 (Behavior Cloning Loss))
Parameters:
\((s, a^*)\) – state-action pairs from expert demonstrations
\(\mathcal{D}_{\text{expert}}\) – expert demonstration dataset
\(\pi(a \mid s)\) – learned policy
Purpose: Supervised policy learning. Trains the policy to match expert actions via maximum likelihood.
Units: \([\mathrm{nat}]\)
Source: Section 25
F.11 Geometric Consistency Losses#
These losses enforce geometric laws derived from capacity constraints (Section 18, Section 20).
Definition 363 (F.11.1 (Metric-law residual))
Let
The coordinate-invariant metric-law loss is
or its minibatch approximation. It measures violation of the curvature–risk stationarity identity; it does not replace the separate capacity diagnostic.
Source: Capacity-Constrained Metric Law: Geometry from Interface Limits, Theorem Theorem 6
Definition 364 (F.11.2 (WFR Consistency Loss))
Parameters:
\(\rho_t\) – belief density at time \(t\)
\(r_t\) – reaction rate (birth/death)
\(v_t\) – transport velocity field
\(\Delta t\) – timestep
Purpose: Enforces Wasserstein-Fisher-Rao consistency. Penalizes deviations from the unbalanced continuity equation in cone-space formulation.
Units: \([\mathrm{nat}^2]\)
Source: Section 20
Definition 365 (F.11.2 (Critic TD Loss with PDE Regularization))
In the control-loop cost convention, write \(c_t:=-r_t\) and \(\rho_c:=-\rho_r\). The critic loss is
Parameters:
TD-Error \(= c + \gamma V(s') - V(s)\) – cost-convention temporal difference error
\(\Delta_G\) – Laplace-Beltrami operator on manifold
\(\lambda = -\ln\gamma/\Delta t\) and \(\kappa^2=\lambda/T_c\) – stationary-diffusion screening coefficient from the discount factor
\(\rho_c=-\rho_r\) – cost density (the reward-side equation uses \(\Phi=-V\) and \(\rho_r\))
Purpose: Combines TD learning with Helmholtz PDE regularization. The PDE term enforces that the critic satisfies the continuum Bellman equation.
Units: \([\mathrm{nat}^2]\)
Source: Section 18
F.12 Consensus Losses (Proof of Useful Work)#
These support the PoUW consensus mechanism (Section 38).
Definition 366 (F.12.1 (Waste Quotient))
Parameters:
\(\Delta I_{\text{world}}\) – mutual information gained about world
\(\dot{\mathcal{M}}(t)\) – metabolic flux (energy dissipation rate)
Interpretation:
Protocol |
Waste Quotient |
Meaning |
|---|---|---|
Bitcoin PoW |
\(W_{\text{BTC}} \approx 1\) |
Energy produces zero world knowledge |
Target PoUW |
\(W_{\text{PoUW}} \to 0\) |
Energy produces useful learning |
Purpose: Measures efficiency of consensus protocol. Low waste quotient means energy dissipation produces useful information gain.
Units: Dimensionless.
Source: Section 38.1, Definition Definition 303
F.13 Meta-Learning Losses#
These losses train the Governor (hyperparameter controller) via bilevel optimization (Section 26).
Definition 367 (F.13.1 (Governor Training Regret))
Parameters:
\(\phi\) – Governor parameters
\(\mathcal{T}\) – task from task distribution
\(\mathcal{L}_{\text{task}}(\theta_t)\) – task loss at training step \(t\)
\(C_k(\theta_t)\) – constraint \(k\) value (negative when satisfied)
\(\gamma_{\text{viol}}\) – constraint violation penalty weight
Purpose: Meta-learning objective for the Governor. Minimizes cumulative task loss (convergence speed) plus squared constraint violations (feasibility). The Governor learns to set hyperparameters \(\Lambda\) that lead to fast, stable training across diverse tasks.
Units: \([\mathrm{nat}]\)
Source: Section 26, Definition Definition 128
F.14 Summary Table#
All distance-based losses use the metric tensor \(G_{ij}\). Flat limits recover standard ML formulas when \(G_{ij} \to \delta_{ij}\).
Loss |
Sec |
Formula Key (Curved) |
Purpose |
|---|---|---|---|
VICReg |
3 |
\(d_G(z,z')^2 + \text{var}_G + \text{cov}_G\) |
Collapse prevention |
InfoNCE |
3 |
\(\exp(-d_G^2/\tau)\) kernel |
Contrastive prediction |
Reconstruction |
3.2 |
\((x-\hat{x})^i G_{ij}^{\text{obs}} (x-\hat{x})^j\) |
Information preservation |
VQ |
3.2 |
\((z_q-z_e)^i G_{ij} (z_q-z_e)^j\) |
Discrete symbol stability |
Closure |
3.2 |
\(-\log p(K_{t+1} \mid K_t, a)\) |
Causal enclosure |
Slowness |
3.2 |
\(d_G(e_{K_t}, e_{K_{t-1}})^2\) |
Anti-symbol-churn |
Nuisance KL |
3.2 |
\(D_{\text{KL}}(q(z_n) | \mathcal{N})\) |
Structured residual prior |
Texture KL |
3.2 |
\(D_{\text{KL}}(q(z_{\text{tex}}) | \mathcal{N})\) |
Reconstruction residual |
Monotonicity |
3.2 |
\(\text{ReLU}(F_{t+1} - F_t)^2\) |
Objective descent |
Gradient Penalty |
4 |
\((|\nabla_A V|_G - K)^2\) |
Lipschitz constraint |
InfoControl |
4 |
Compression + Control effort |
Information-control tradeoff |
EWC |
4 |
\(\sum_i F_i (\theta_i - \theta^*_i)^2\) |
Stability-plasticity balance |
Bode |
4 |
\(|\mathcal{F}(e_t) W(\omega)|^2\) |
Frequency sensitivity |
Overlap Consistency |
7 |
\((z_n^{(j)} - L_{i→j})^k G^{(j)}_{k\ell} (\cdot)^\ell\) |
Chart transition coherence |
QSL |
11 |
\(\text{ReLU}(d_G - v_{\max})^2\) |
Quantum speed limit |
Hodge Decomp |
18 |
\(\mathcal{R} = d\Phi + \delta\Psi + \eta\) |
Reward structure |
Value Curl |
18 |
\(\mathcal{F} = d\mathcal{R}\) |
Non-conservative detection |
Critic TD+PDE |
18 |
\(|\text{TD}|^2 + \lambda|\Delta_G V - \cdots|^2\) |
Value function learning |
WFR Consistency |
20 |
Cone-space continuity |
Transport-reaction balance |
Purity |
25 |
\(H(Y \mid K)\) |
Chart semantic purity |
Balance |
25 |
\(D_{\text{KL}}(\bar{w} | U)\) |
Prevent chart collapse |
Metric Contrastive |
25 |
\(\max(0, m - d_G^{\text{jump}})^2\) |
Geometric separation |
Route Alignment |
25 |
\(\text{CE}(\sum_k w_k P(Y \mid K), y)\) |
Classification accuracy |
Class Potential |
25 |
\(-\beta \log P(Y \mid K) + V_b\) |
Class basin formation |
Behavior Cloning |
25 |
\(-\log \pi(a^* \mid s)\) |
Imitation learning |
Hierarchical |
25 |
\(\sum_\ell \alpha_\ell \mathcal{L}^{(\ell)}\) |
Multi-scale classification |
Governor Regret |
26 |
\(\sum_t (\mathcal{L}_{\text{task}} + \gamma \text{ReLU}(C_k)^2)\) |
Meta-learning objective |
Causal Info |
29 |
\(\mathbb{E}[D_{\text{KL}}(p(\theta_W \mid z') | p(\theta_W))]\) |
Exploration via EIG |
Ontological Stress |
33 |
\((z_{\text{tex}}^{(\ell)})^i G^{(\ell)}_{ij} (z_{\text{tex}}^{(\ell)})^j\) |
Texture predictability |
Sync Potential |
Gauge chapter |
\(\beta\Psi_{\text{sync}}\) on \(\mathcal{D}_{AB}\) |
Relative-gauge alignment |
Joint Prediction |
37 |
\(d_G(\hat{x}^A, x)^2 + d_G(\hat{x}^B, x)^2\) |
Multi-agent world model |
Waste Quotient |
38 |
\(1 - \Delta I / \int \dot{\mathcal{M}} dt\) |
Consensus efficiency |