Intrinsic Motivation: Maximum-Entropy Exploration#

TLDR#

  • Maximum-entropy (MaxEnt) exploration is not random dithering: it is a control objective that preserves future reachability (“keep options open”).

  • In practice this becomes entropy/KL-regularized control on the macro state-action trajectory space: a temperature trades reward for diversity.

  • The key identity is a finite-horizon duality: maximizing causal policy entropy under reward constraints is equivalent to soft optimal control with a KL penalty toward a reference policy when the kernel and discount conventions are matched.

  • Use Sieve diagnostics to prevent MaxEnt failure modes (chattering/Zenoness, over-mixing, loss of grounding).

  • This chapter sets up the belief-dynamics and coupling-window chapters: exploration pressure must remain within the information-stability window.

Roadmap#

  1. Path entropy and exploration gradients (what is optimized).

  2. Duality with soft optimality (why MaxEnt control is “just” KL control).

  3. Practical guidance: temperatures, horizons, and diagnostic failure modes.

Let me tell you about a beautiful idea that connects two things you might not have thought were related: exploring the world and keeping your options open.

When you learn reinforcement learning, exploration often gets treated as a necessary evil. You need to try different things so you don’t get stuck in a local optimum, but exploration is just noise you add to your policy, right? Random actions to shake things loose.

Wrong. There’s a deeper way to think about exploration, and it leads to much better algorithms.

Here’s the key insight: exploration isn’t about randomness for its own sake. It’s about maintaining reachability. A good explorer is an agent that can still get to many different places in the future. An agent that’s painted itself into a corner—even if that corner looks pretty good right now—has lost something valuable. It’s lost the ability to change course.

Maximum-entropy exploration formalizes this. Instead of asking “what action gives me the highest expected reward right now?”, we ask “what action gives me the highest expected reward while preserving my ability to reach many future states?” The causal entropy of your future macro state-action trajectories—the randomness injected by the policy—is a measure of that ability; environmental transition noise is accounted for by the path law but is not credited to the exploration objective.

And here’s the beautiful thing, with an important qualification: under the finite-horizon, deterministic-kernel convention developed below, maximizing causal entropy and maximizing soft (KL-regularized) reward are the same problem viewed from different angles. Outside those matched hypotheses, the relation is a useful guide rather than an automatic identity.

Researcher Bridge: Max-Entropy Exploration in Macro Space

This is the MaxEnt RL idea applied to discrete macro state-action trajectories. Instead of adding a scalar bonus, we maximize the entropy of reachable macro state-action trajectories, which is the discrete version of “keep options open.”

The previous layers define representation (\(K,z_n,z_{\text{tex}}\)), predictive dynamics (\(\bar{P}\)), and stability/value constraints (\(V,G\), Sieve checks). This layer formalizes an intrinsic exploration pressure on the discrete macro register: prefer policies that keep the agent’s future macro state-action choices diverse, as measured by causal policy entropy, which supports reachability/controllability and reduces brittle overcommitment to narrow state-action paths.

Path Entropy and Exploration Gradients#

We’re going to work on the macro model—the discrete register of high-level states. Why discrete? Because entropy is well-defined for discrete distributions. There’s no ambiguity, no reference measures, no gauge choices. The entropy of a distribution over a finite set is just what Shannon said it is.

We start with the macro Markov kernel: the learned transition probabilities \(\bar{P}(k' | k, a)\) that tell us how macro-states evolve under actions. This is the causal enclosure we demanded earlier—the macro-symbol alone (plus action) is sufficient to predict the next macro-symbol.

We work on the macro model (the discrete register). Assume a macro Markov kernel

\[ \bar{P}(k'\mid k,a),\qquad k,k'\in\mathcal{K},\ a\in\mathcal{A},\]

which is the learned effective dynamics demanded by Causal Enclosure (Conditional Independence and Sufficiency (Causal Enclosure)).

In this chapter \(\mathcal A\) denotes the finite motor-macro alphabet, the range of \(K^{\mathrm{act}}_t\), and we write \(a\equiv K^{\mathrm{act}}_t\); motor nuisance and texture are marginalized out of \(\bar P\) and \(\mathcal R\).

Definition 34 (Macro Path Distribution)

Fix a horizon \(H\in\mathbb{N}\) and a (possibly stochastic) policy \(\pi(a\mid k)\). The induced distribution over length-\(H\) macro state-action trajectories

\[ \xi := (K_t, A_t, K_{t+1}, A_{t+1}, \dots, A_{t+H-1}, K_{t+H}) \in \mathcal{K}\times(\mathcal{A}\times\mathcal{K})^H\]

conditioned on \(K_t=k\) is

\[ P_\pi(\xi\mid k) := \prod_{h=0}^{H-1}\pi(A_{t+h}\mid K_{t+h})\ \bar{P}(K_{t+h+1}\mid K_{t+h},A_{t+h}).\]

(For continuous \(\mathcal{A}\), interpret \(P_\pi(\xi\mid k)\) as a density with respect to the action reference measure.)

What is this definition saying? Fix where you are now (macro-state \(k\)) and fix a policy (how you choose actions). Then there’s a distribution over full \(H\)-step state-action trajectories: which actions you might take and which macrostates follow. That’s \(P_\pi(\xi | k)\)—the probability of each possible state-action trajectory.

Notice the factorization. The policy term \(\pi(a_t\mid k_t)\) selects actions, and the dynamics term \(\bar{P}(k_{t+1}\mid k_t,a_t)\) advances the macrostate. Even if the dynamics were deterministic, a stochastic policy would still induce a distribution over state-action paths, and vice versa. When we define causal path entropy, we only credit randomness from the policy term. Stochasticity in \(\bar{P}\) is not under the agent’s control, so it does not contribute.

Definition 35 (Causal Path Entropy)

The causal path entropy at \((k,H)\) under \(\pi\) is the cumulative policy entropy along paths \(\xi\in\Gamma_H(k)\) induced by \(\pi\) and \(\bar{P}\):

\[ S_c(k,H;\pi) := \sum_{h=0}^{H-1} \mathbb{E}_{\xi\sim P_\pi(\cdot\mid k)} \left[ \mathcal H\!\left(\pi(\cdot\mid K_{t+h})\right) \right].\]

Only policy randomness contributes; stochasticity in \(\bar{P}\) does not add entropy credit. The expectation is taken under the path law induced by \(\pi\) and \(\bar{P}\). This quantity is well-typed because the macro register is discrete; for continuous \(\mathcal{A}\), interpret \(\mathcal H(\pi(\cdot\mid k))\) as a differential entropy with respect to the action reference measure.

Now we’re getting somewhere. \(S_c(k, H; \pi)\) measures how much randomness the agent injects into its future action choices along the trajectory. High causal entropy means the policy stays spread out at the states it expects to visit; low entropy means you commit to a narrow action plan.

The key distinction is control. Environmental noise does not count toward \(S_c\) because it is not chosen by the agent. Causal entropy only credits the randomness the agent injects through \(\pi\).

I want to emphasize why discreteness matters here. If \(\mathcal{A}\) were continuous, we’d have to talk about differential entropy, which depends on your choice of reference measure. Different reference measures give different entropies. With discrete actions (and a discrete macro register), there’s no such ambiguity. Entropy is entropy. This is one of the payoffs for the VQ-VAE architecture that quantizes latent states.

Definition 36 (Exploration Gradient, metric form)

Let \(z_{\text{macro}}=e_k\in\mathbb{R}^{d_m}\) denote the code embedding of \(k\) (The Shutter as a VQ-VAE (Discrete Macro, Continuous Micro)), and let \(G\) be the relevant metric on the macro chart (Second-Order Sensitivity: Value Defines a Local Metric). Assume smooth policy and kernel heads \(\pi_\theta(a\mid z)\) and \(\bar P_\phi(k'\mid z,a)\), and let \(\widetilde S_c(z,H;\pi)\) be the causal-entropy formula above with the initial code \(k\) replaced by \(z\) and these heads evaluated at \(z\).

\[ \mathbf{g}_{\text{expl}}(e_k) := T_c\,G(e_k)^{-1}\nabla_z\widetilde S_c(z,H;\pi)\big|_{z=e_k},\]

where \(T_c>0\) is the cognitive temperature (Definition 73). The straight-through VQ estimator transports this continuous gradient to the pre-quantization coordinates. In the strictly symbolic limit there is no tangent vector; use the separate preference ordering obtained by ranking \(S_c(k,H;\pi)\) over \(k\).

Interpretation (Exploration / Reachability). \(S_c(k,H;\pi)\) measures how much action-level randomness the agent injects along trajectories from \(k\) under \(\pi\). Increasing \(S_c\) preserves agent-controlled reachability: the policy avoids committing to a narrow action sequence, independent of environmental stochasticity.

The exploration gradient \(\mathbf{g}_{\text{expl}}\) tells you which direction in state space increases your future optionality. It’s like a compass pointing toward freedom. The cognitive temperature \(T_c\) controls how strongly you weight exploration versus exploitation—high temperature means exploration dominates, low temperature means you mostly follow reward gradients.

Now here’s a subtle point. The macro-state \(K\) is discrete, but we take gradients through its continuous embedding \(e_k\). How does that work? Through the straight-through estimator from the VQ-VAE. In the forward pass, you quantize to discrete codes. In the backward pass, you pretend the quantization was differentiable and flow gradients to the pre-quantization coordinates. It’s a hack, but it works beautifully in practice.

MaxEnt Duality: Utility + Entropy Regularization#

We’ve defined causal path entropy as a measure of future reachability. Now let’s connect this to standard reinforcement learning by showing, in a specified finite-horizon setting, how maximizing that policy-controlled entropy becomes a soft reward problem.

The setup is familiar: you have an instantaneous reward function \(\mathcal{R}(k, a)\) and a discount factor \(\gamma\). The standard utility-plus-entropy objective can use discounting, but the exact path-space equivalence proved later fixes a finite horizon and the undiscounted convention \(\gamma=1\). Keeping those cases separate prevents the entropy bookkeeping from changing halfway through the argument.

Definition 37 (MaxEnt RL objective on macrostates)

Let \(\mathcal{R}(k,a)\) be an instantaneous reward/cost-rate term (Re-typing Standard RL Primitives as Interface Signals, The HJB Correspondence (Costs as Value Updates)) and let \(\gamma\in(0,1)\) be the discount factor (dimensionless). The maximum-entropy objective is

\[ J_{T_c}(\pi) := \mathbb{E}_\pi\left[\sum_{t\ge 0}\gamma^t\left(\mathcal{R}(K_t,K^{\text{act}}_t) + T_c\,\mathcal{H}(\pi(\cdot\mid K_t))\right)\right],\]

where \(\mathcal{H}\) is Shannon entropy. This is the standard “utility + entropy regularization” objective.

Regimes.

  • \(T_c\to 0\): \(\pi\) collapses toward determinism; behavior can be brittle under distribution shift.

  • \(T_c\to\infty\): \(\pi\) approaches maximal entropy; behavior becomes overly random and may degrade grounding (BarrierScat).

  • The useful regime is intermediate: enough entropy to remain robust, enough utility to remain directed.

Let me make sure this is clear. The objective \(J_{T_c}(\pi)\) has two terms at each timestep:

  1. Reward: \(\mathcal{R}(K_t, K^{\text{act}}_t)\)—how good is the immediate outcome?

  2. Policy entropy: \(T_c \cdot \mathcal{H}(\pi(\cdot | K_t))\)—how spread out is your action distribution?

The cognitive temperature \(T_c\) trades off these two concerns. When \(T_c\) is large, you care a lot about keeping your options open (high entropy policy). When \(T_c\) is small, you care mostly about reward and your policy becomes more deterministic.

The extreme cases are instructive:

  • \(T_c \to 0\): You become a pure reward maximizer. Your policy converges to the greedy action at each state. This can be brittle—if your environment shifts, you have no backup plans.

  • \(T_c \to \infty\): You become uniformly random. Every action is equally likely. This is maximally robust but you get no actual work done.

The sweet spot is somewhere in between, and finding it is part of the art of RL algorithm design.

Proposition 10 (Soft Bellman form, discrete actions)

Assume finite \(\mathcal{A}\). Define the soft state value

\[ V^*(k) := \max_{\pi} \ \mathbb{E}\Big[\sum_{t\ge 0}\gamma^t(\mathcal{R}+T_c\mathcal{H})\ \Big|\ K_0=k\Big].\]

Then \(V^*\) satisfies the entropic Bellman fixed point

\[ V^*(k) = T_c \log \sum_{a\in\mathcal{A}} \exp\!\left(\frac{1}{T_c}\left(\mathcal{R}(k,a)+\gamma\,\mathbb{E}_{k'\sim\bar{P}(\cdot\mid k,a)}[V^*(k')]\right)\right),\]

and the corresponding optimal policy is the softmax policy

\[ \pi^*(a\mid k)\propto \exp\!\left(\frac{1}{T_c}\left(\mathcal{R}(k,a)+\gamma\,\mathbb{E}[V^*(k')]\right)\right).\]

Proof sketch. Standard convex duality / log-sum-exp variational identity: maximizing expected reward plus entropy yields a softmax (exponential-family) distribution; substituting back produces the log-partition recursion. (This is the “soft”/MaxEnt Bellman equation used in SAC-like methods.)

Consequence. The same mathematics can be read as:

  1. maximize reward while retaining policy entropy (MaxEnt RL), or

  2. maximize reachability/diversity of future macro state-action trajectories (intrinsic motivation).

The soft Bellman equation is gorgeous. Let me walk you through it.

In regular dynamic programming, the Bellman equation says: the value of a state is the maximum over actions of (immediate reward + discounted future value). You pick the best action and that determines your value.

In soft dynamic programming, you replace the maximum with a “soft maximum”—the log-sum-exp. Instead of picking one action, you average over all actions, weighted by their exponentiated values. This is the same as maximizing expected value plus entropy.

The log-sum-exp function has a beautiful property: it’s a smooth approximation to the maximum. When \(T_c\) is small, \(T_c \log \sum_a \exp(Q_a / T_c)\) approaches \(\max_a Q_a\). When \(T_c\) is large, it approaches the average plus some entropy term. The temperature interpolates between “be greedy” and “be uncertain.”

The optimal policy falls out automatically: it’s the softmax over Q-values (which are immediate reward plus discounted future value). High-value actions get more probability, but every action gets some probability, weighted by temperature.

The Two Faces of MaxEnt

The same mathematics has two interpretations that are useful in different contexts:

  1. MaxEnt RL view: “I want high reward, but I also want to hedge my bets by keeping my policy from being too deterministic.”

  2. Intrinsic motivation view: “I want to stay in regions of state space where many futures are reachable, because that gives me flexibility to adapt.”

Under the finite-horizon deterministic-kernel hypotheses below, these are the same objective, just explained differently. The first emphasizes reward-seeking behavior with entropy as a regularizer. The second emphasizes exploration and reachability with reward as a guide. For stochastic kernels or discounted objectives, the corresponding statement needs the restricted policy-induced laws or a discounted KL convention stated in the theorem.

For building intuition: if you’re optimizing for a known reward function, think MaxEnt RL. If you’re trying to build an agent that can adapt to changing goals, think intrinsic motivation.

Duality of Exploration and Soft Optimality#

Researcher Bridge: Soft RL Equals Exploration Duality

If you know SAC or KL control, this section formalizes why maximizing entropy and optimizing soft value are the same problem. The exploration gradient is just the covariant form of that duality.

Now we’re going to make the duality between exploration and soft optimality precise under explicit hypotheses. This is beautiful mathematics, and it has practical consequences for how you think about policy learning.

The claim is strong but deliberately scoped: for a finite horizon, a deterministic enclosure-consistent macro kernel, and a matched reference policy, maximizing causal path entropy subject to expected reward constraints is exactly the same problem as maximizing expected reward with a KL penalty toward that reference. Different objective functions, the same optimal solution. With stochastic dynamics or discounting, the path tilt and the policy-induced KL require a different statement.

Formal Definitions (Path Space, Causal Entropy, Exploration Gradient)#

Definition 38 (Causal Path Space)

For a macrostate \(k\in\mathcal{K}\) and horizon \(H\), define the future macro state-action path space

\[ \Gamma_H(k) := \left\{(k_0,a_0,k_1,a_1,\dots,a_{H-1},k_H)\in\mathcal{K}^{H+1}\times\mathcal{A}^H : k_0 = k\right\}.\]

This is just notation. \(\Gamma_H(k)\) is the set of all possible \(H\)-step state-action trajectories starting from \(k\). If \(\mathcal{A}\) is finite, \(\Gamma_H(k)\) has \(|\mathcal{A}|^H|\mathcal{K}|^H\) elements.

Definition 39 (Path Probability)

\(P_\pi(\xi\mid k)\) is the induced state-action path probability from Definition 34.

Definition 40 (Causal Entropy)

\(S_c(k,H;\pi)\) is the causal path entropy from Definition 35, i.e., the cumulative policy entropy along the induced path measure \(P_\pi(\cdot\mid k)\).

Definition 41 (Exploration gradient, covariant form)

On a macro chart with metric \(G\) (Second-Order Sensitivity: Value Defines a Local Metric),

\[ \mathbf{g}_{\text{expl}}(e_k) := T_c\,G(e_k)^{-1}\nabla_z\widetilde S_c(z,H;\pi)\big|_{z=e_k},\]

Why “covariant form”? Because we’re taking gradients with respect to the metric \(G\), not with respect to Euclidean coordinates. The metric gradient \(\nabla_G\) accounts for the local geometry of state space. This matters when your state space is curved or has non-uniform sensitivity—the natural direction of steepest ascent depends on the metric.

The Equivalence Theorem (Duality of Causal Regulation)#

Now for the main event. Under the finite-horizon deterministic-kernel assumptions stated next, the following theorem says that three apparently different ways of stating the optimal control problem are equivalent. They give the same optimal policy, and their objective values are related by simple transformations.

This is not a “they’re approximately the same” result within that setting: it is an exact equivalence. The same optimal policy arises whether you think about it as finite-horizon MaxEnt control, KL-regularized trajectory optimization with the kernel held fixed, or finite-horizon soft Bellman dynamic programming. Do not silently extend the claim to arbitrary stochastic or discounted path laws.

Theorem 5 (Finite-Horizon Equivalence for a Deterministic Macro Kernel)

Assume:

  1. finite macro alphabet \(\mathcal{K}\) and (for simplicity) finite action set \(\mathcal{A}\),

  2. a deterministic enclosure-consistent macro kernel \(\bar{P}(k'\mid k,a)\),

  3. bounded reward flux \(\mathcal{R}(k,a)\),

  4. a finite horizon \(H\) and undiscounted objective (\(\gamma=1\)).

Then the following are equivalent characterizations of the same finite-horizon optimal control law:

  1. Finite-horizon MaxEnt control: \(\pi^*\) maximizes \(\mathbb E_\pi[\sum_{h=0}^{H-1}(\mathcal R(K_{t+h},A_{t+h})+T_c\mathcal H(\pi(\cdot\mid K_{t+h})))]\) from the initial state \(K_t=k\).

  2. Exponentially tilted trajectory measure (KL-regularization). Fix a uniform reference (prior) policy \(\pi_0(a\mid k)\). For the deterministic kernel, the length-\(H\) optimal path law admits the exponential-family form relative to the reference measure induced by \(\pi_0\) and \(\bar P\):

    \[ P^*(\omega\mid K_t=k)\ \propto\ P_0(\omega \mid k)\, \exp\!\left(\frac{1}{T_c}\sum_{h=0}^{H-1}\mathcal{R}(K_{t+h},A_{t+h})\right),\]

    where \(P_0(\omega \mid k) := \prod_{h=0}^{H-1}\pi_0(A_{t+h}\mid K_{t+h})\,\bar{P}(K_{t+h+1}\mid K_{t+h},A_{t+h})\) is the finite-horizon reference measure.

  3. Finite-horizon soft Bellman optimality: with \(V_H^*\equiv0\),

    \[ V_h^*(k)=T_c\log\sum_{a\in\mathcal A}\exp\!\left(\frac{\mathcal R(k,a)+\mathbb E_{k'\sim\bar P(\cdot\mid k,a)}V_{h+1}^*(k')}{T_c}\right), \]

    and \(\pi_h^*(a\mid k)\) is the corresponding softmax policy.

Moreover, for the kernel-consistent family \(P=P_\pi\) and a uniform prior \(\pi_0\), the link is the KL-regularized variational identity

\[ \log Z_H(k) = \sup_{\pi} \left\{ \frac{1}{T_c}\,\mathbb{E}_{P_\pi}\!\left[\sum_{h=0}^{H-1}\mathcal{R}\right] -D_{\mathrm{KL}}(P_\pi\Vert P_0) \right\},\]

and the optimizer is the policy-induced law in item 2. For uniform \(\pi_0\), \(D_{\mathrm{KL}}(P_\pi\Vert P_0)=H\log|\mathcal A|-S_c(k,H;\pi)\), so this is maximization of expected reward plus \(T_c\) times the causal path entropy. The normalization satisfies \(T_c\log Z_H(k)=V_0^*(k)-T_cH\log|\mathcal A|\).

For stochastic kernels or discounted objectives, this equivalence does not hold in this form. One must either restrict the admissible laws to \(P=\pi\cdot\bar P\) and use the discounted per-step policy KL, or treat the full path tilt as a distinct risk-sensitive control problem.

Proof sketch. For deterministic dynamics, the path law is determined by its action probabilities. Applying the finite-space Gibbs variational identity to the action path gives the exponential tilt and the KL expression; backward conditioning yields the displayed soft Bellman recursion. The uniform-prior identity follows by cancellation of the fixed dynamics factors.

Let me unpack why this theorem matters, keeping its scope visible.

Form 1 is how you think about MaxEnt RL day-to-day: maximize reward plus entropy. SAC, Soft Q-Learning, and similar algorithms use related soft objectives, although their stochastic and often discounted settings are not automatically covered by this finite-horizon theorem.

Form 2 is the path-space view: instead of thinking about policies, think about distributions over entire state-action trajectories. The optimal state-action trajectory distribution is an exponential tilt of the reference distribution, where the tilt factor is the exponential of cumulative reward. High-reward state-action trajectories get exponentially more probability.

Form 3 is the dynamic programming view: the soft Bellman equation gives you a recursive way to compute optimal values, and the optimal policy is a softmax over Q-values.

Under the theorem’s hypotheses, these are all the same. If you solve one, you’ve solved them all for that finite-horizon problem: the optimal policy \(\pi^*\) is identical whether you derive it from Form 1, Form 2, or Form 3.

The key equation is the variational identity at the end: the finite-horizon log-normalizer equals the maximum over kernel-consistent state-action trajectory laws of expected reward minus KL to the reference. This is convex optimization, and the KL penalty is what gives a smooth log-sum-exp. The reference normalization contributes the explicit uniform-prior offset shown in the theorem.

Why the Log-Normalizer Matters

For the finite-horizon uniform-prior convention of the theorem, \(T_c\log Z_H(k)=V_0^*(k)-T_cH\log|\mathcal A|\). The additive constant does not affect policies or gradients, but it matters when \(Z_H\) is used as a numerical value.

In statistical mechanics, the log-normalizer of the Boltzmann distribution is the free energy. Here, the same Gibbs structure appears for a finite macro trajectory problem, with the explicit uniform-prior offset above.

For practical algorithms, an estimate of \(Z_H(k)\) can provide policy gradients, provided the finite-horizon deterministic-kernel hypotheses are satisfied.

Connection to RL #22: KL-Regularized Policies as Degenerate Exploration Duality

The Finite-Horizon Deterministic-Kernel Law: Under the hypotheses of Theorem Theorem 5, MaxEnt control is equivalent to an Exponentially Tilted Trajectory Measure:

\[ P^*(\omega|K_t=k) \propto P_0(\omega|k) \exp\!\left(\frac{1}{T_c}\sum_{h=0}^{H-1} \mathcal{R}(K_{t+h}, A_{t+h})\right)\]

The path-space log-normalizer obeys \(T_c\log Z_H(k)=V_0^*(k)-T_cH\log|\mathcal A|\) (Theorem Theorem 5). This is a KL-control (Gibbs/exponential-tilt) variational principle; a Schrödinger bridge additionally prescribes terminal marginals.

The Degenerate Limit: Use single-step KL penalty instead of path-space tilting. Ignore the trajectory structure.

The Special Case (Standard RL):

\[ J(\pi) = \mathbb{E}[R] - \lambda D_{\mathrm{KL}}(\pi \| \pi_0)\]

This recovers KL-Regularized Policy Gradient and exponential family policies.

What the generalization offers:

  • Path-space view: The optimal policy is a Gibbs/KL-control tilt of a prior trajectory measure

  • Trajectory entropy: Explores future macro state-action trajectories \(\omega = (K^{\text{act}}_t, K_{t+1}, K^{\text{act}}_{t+1}, \ldots)\), not just single actions

  • Variational principle: The finite-horizon log-partition function differs from the soft value by the explicit uniform-prior constant above

  • Causal entropy: \(S_c(k, H; \pi)\) measures future reachability under causal interventions

The connection to standard KL-regularized policies is instructive. When you add a KL penalty \(D_{\text{KL}}(\pi \| \pi_0)\) to your policy gradient objective, you’re using a one-step or per-step version of what we’re describing here. The full picture is path-space: you’re regularizing toward a reference state-action trajectory distribution, not just a reference action distribution.

This matters when your MDP has temporal structure. Single-step KL regularization does not encode the whole path law, while path-space KL regularization does. A Schrödinger bridge goes one step further by imposing endpoint marginals; it is useful here as an optional path-space analogy, not as a replacement for the finite-horizon control theorem.

For discrete macro-states, this is computationally tractable because the path space is finite. For continuous states, you’d need approximations—which is why practical algorithms like SAC use the single-step version and rely on temporal-difference learning to propagate future information backward.

Failure Modes and Diagnostics#

MaxEnt exploration is powerful, but it has characteristic failure modes. In the Fragile Agent, these are meant to be detectable and actionable:

  • Chattering / Zenoness: entropy pressure can cause rapid action switching. Monitor Zeno-style switching checks and add explicit switching penalties or reduce temperature/horizon.

  • Over-mixing (loss of macro identity): excessive entropy destroys stable macrostates. Monitor mixing/compactness and closure diagnostics; enforce the coupling window rather than increasing entropy indefinitely.

  • Premature collapse (no exploration): temperature too low collapses the policy to a brittle mode. Monitor entropy / reachability metrics and reintroduce exploration pressure when coverage shrinks.

  • Ungrounded exploration: exploring in the internal model without boundary support leads to hallucinated reachability. Monitor grounding/closure synchronization; intervene by tightening closure losses or shortening open-loop rollouts.

The intended workflow is: choose \(T_c\) (and horizon) as a knob, then let diagnostics decide when that setting is safe for the current regime.