Chapter 4: A Priori Estimates
Chapter 3 proved that the vector field of system (L1) is locally Lipschitz and that the Picard-Lindelöf theorem gives a unique local solution from every initial condition in the phase space
. That solution lives on a maximal interval
, and the blow-up alternative tells us that if the solution stays in a compact set, then
.
The question is: does the solution stay in a compact set?
In the Chapter 0 example, the answer was visible by inspection. The actor
stayed in
because the damping
vanished at the boundary, silencing the drift before
could escape. The distribution
stayed in
because the relaxation equation
is a convex combination that preserves the simplex. And the critic
stayed bounded because the equation
has a coercive linear term that pulls
back toward the origin. Together, these three mechanisms trapped every trajectory in the compact absorbing set
.
This chapter proves that all three confinement mechanisms extend to the general model. We will establish them in order — actor-box invariance, simplex invariance, critic coercivity — and then combine them to build the compact absorbing set. Once the absorbing set is in hand, the blow-up alternative from Chapter 3 gives global existence, completing the semiflow construction.
The proof plan has five steps:
Actor-box invariance (Section 4.1). The damping matrix
vanishes on the boundary of
, so a scalar barrier argument confines
to the box for all time.Simplex invariance (Section 4.2). The law equation has an explicit variation-of-constants formula that represents
as a convex combination of simplex elements.Critic coercivity (Section 4.3). The uniform positive definiteness of the critic matrix gives an energy estimate that bounds
and identifies the level toward which the critic norm decays; the absorbing radius
is chosen strictly above that level.The compact absorbing set (Section 4.4). Combining the three invariance results yields the explicit compact absorbing set
, which is forward invariant and absorbs every bounded subset of
.Global existence (Section 4.5). The a priori bounds prevent blow-up, and the blow-up alternative from Chapter 3 extends the local semiflow to all positive time.
4.1 Actor-Box Forward Invariance
The actor equation from Chapter 2 is

where
is the damping matrix and
is the raw actor drift. The damping factor
in the
-th coordinate vanishes when
, which is the boundary of the actor box
. The question is whether this vanishing is enough to prevent
from escaping.
For the dynamical-systems reader. The damping
acts as a soft projection: instead of discontinuously reflecting the trajectory at the boundary of
, it smoothly reduces the drift to zero as the boundary is approached. This is a design choice in the RL algorithm — the actor parameter space is bounded, and the damping enforces that bound without introducing discontinuities that would break the Lipschitz regularity proved in Chapter 3. The resulting invariance is stronger than what a hard projection would give: the boundary of
is an invariant face of the box, and a trajectory either stays away from it or stays on it forever.
Proposition 4.1 (Invariance of the actor box). Let
be a local solution of system (L1) (Definition 2.10) on an interval
. Then

More precisely: if
for some coordinate
, then
for all
. If
, then
for all
.
The proof uses a scalar barrier argument based on the arctanh function. The idea is that the change of variable
absorbs the damping factor and transforms the actor ODE into one with a bounded right-hand side. Since arctanh maps
to all of
, a trajectory that starts in the interior can never reach the boundary in finite time — it would need to push the arctanh coordinate to infinity, which the bounded right-hand side prevents.
Proof. Fix a coordinate
. Write

The function
is continuous on
because the solution is continuous and
is continuous (Lemma 3.2). The
-th actor coordinate satisfies the scalar ODE

Case 1: Interior start (
). Define the first exit time

with the convention
. We will show that
, meaning the trajectory never reaches the boundary.
For every
, we have
for all
, so the damping factor is strictly positive and we may divide both sides of the ODE by
. Using the identity

we obtain the transformed equation (the chain rule contributes a factor
from the argument
):

Integrating from
to
gives

Taking absolute values:

Now suppose for contradiction that
. Then
is continuous on the compact interval
, so the quantity

is finite. Therefore, for every
,

Applying the hyperbolic tangent (which is increasing and satisfies
for all finite
):

By continuity of
, the same bound holds at
:

This contradicts the definition of
as the first time
. Therefore
, and

Case 2: Boundary start (
). Define
. Then
, and

The coefficient
is continuous on
. The function

therefore has derivative zero. Evaluating at
gives

The exponential factor is strictly positive. Therefore
for every
, which means
for all
.
Case 3: Boundary start (
). Define
. Then
, and the analogous computation (with the integrating factor of opposite sign) gives

The function

has derivative zero. Since
and the exponential is strictly positive, we conclude
for all
, hence
on
.
Since the argument applies to every coordinate
, we conclude

The proof is complete. 
Verification in the Chapter 0 example. In Chapter 0, the actor is a scalar with
, so the actor box is
and the damping is
. The actor equation is

At
, the damping factor
kills the drift regardless of
,
, or
. The arctanh transformation becomes

which is a bounded function on
. In particular, consider a trajectory that starts in
with
; Chapter 0 verified that
is forward invariant, so
for all time. Then
, and the arctanh coordinate grows at most linearly in time: at rate bounded by
. Since
and
on
, the arctanh coordinate is bounded by
. The hyperbolic tangent of
is always strictly less than
, confirming that
never reaches
from an interior starting point.
The boundary cases are even simpler: if
, then
, the right-hand side of the ODE is zero, and
. The uniqueness argument in Case 2 shows that
for all time. Likewise
gives
for all time. These are the boundary equilibria observed in Chapter 0.
Takeaway. The actor stays in its box because the damping matrix
vanishes at the boundary, and the arctanh barrier argument shows that no interior trajectory can reach the boundary in finite time. The boundary itself is invariant: any trajectory that starts there remains there forever. This confines the first
coordinates of the solution to the compact set
.
4.2 Simplex Forward Invariance
The law equation from Chapter 2 is

where
is the relaxation rate and
is the prescribed closure map. The question is whether a distribution that starts in the simplex
stays there.
In the Chapter 0 example, the distribution variable was a single scalar
(with
), and the law equation was
. We observed that the relaxation pulls
toward a target in
, and the exponential decay
in the explicit solution ensures that
stays in
if it starts there. The general argument is essentially the same: the explicit solution of the law equation represents
as a convex combination of simplex elements.
Proposition 4.2 (Invariance of the state simplex). Let
be a local solution of system (L1) on
. Then

Proof. The law equation is a linear, non-autonomous ODE in
:

This is a first-order linear equation with constant coefficient
and time-dependent forcing
. The unique solution is given by the variation-of-constants formula: multiply both sides by the integrating factor
, observe that

and integrate from
to
:

Dividing by
gives the explicit representation

We now verify that this is a convex combination of elements of
.
Nonnegativity and normalization of the coefficients. The coefficient of
is
. The integrand in the second term has coefficient
for
. The total weight is

Membership in
. The initial condition
by hypothesis. Every value
belongs to
because
maps into
by definition (Definition 2.9). Since the coefficients are nonnegative and sum to
, the representation expresses
as a convex combination of points of
— an integral average rather than a finite convex combination — so we check membership coordinatewise. Each coordinate

is nonnegative because every term on the right is nonnegative, and summing over
— using that
and each
have coordinates summing to one — reproduces the total-weight computation above:
. We conclude

The proof is complete. 
Verification in the Chapter 0 example. In Chapter 0 the law equation is

with
. The variation-of-constants formula gives

Since
by Proposition 4.1, the target
lies in
. The initial condition
. The coefficients are nonnegative and sum to
, so
is a convex combination of values in
, which means
for all
.
We can also check the simplex property directly at the level of the vector field. The full two-component law equation is

Since
and
, adding the two equations gives

The total mass is conserved, confirming that the simplex constraint
is preserved by the flow.
Takeaway. The distribution stays in the simplex because the variation-of-constants solution is a convex combination of simplex elements, with nonnegative coefficients that sum to one. This confines the last
coordinates of the solution to the compact set
.
4.3 Critic Coercivity and the Energy Estimate
Propositions 4.1 and 4.2 confine the actor to
and the distribution to
. The actor box and the simplex are both compact, so these two coordinates cannot cause blow-up. The remaining question is whether the critic
stays bounded.
The critic equation from Chapter 2 is

where

The term
acts as a restoring force that pulls the critic back toward the origin, while
is a bounded forcing term. The strength of the restoring force is controlled by the coercivity constant
, which appears in the definition of
as the scalar multiple of the identity.
For the RL reader. The coercivity constant
above is the regularization strength in the critic update. In a typical linear temporal-difference learning rule, the data-driven sum
of rank-one feature-outer products can be degenerate — it may fail to be positive definite if the critic features do not span all of
. The added term
prevents this degeneracy. From the dynamical-systems perspective, this regularization is what turns the estimation loop into a dissipative mechanism: without it, the critic coordinate could grow without bound, and the system would not have a compact absorbing set.
Hypothesis (Critic coercivity). Throughout this chapter we use the following uniform coercivity hypothesis: there exists
such that for every
and every
,

The structural input from Chapter 2 is the bound
in Definition 2.4 (the critic coefficients, where the regularized matrix
is defined); that bound is what makes the symmetric part of
uniformly positive definite, with a constant that does not depend on
. Proposition 4.3 below verifies the displayed inequality from this structural input.
We will use the bounded reward constant
and the bounded feature constant
, both finite because
and
are finite. The product
will serve as the forcing bound.
We first establish the coercivity bound and the forcing bound, then derive the energy estimate.
Proposition 4.3 (Coercivity of the critic matrix). For every
and every
,

Moreover, the forcing vector satisfies

where
is the forcing bound introduced above.
Proof. By definition of
,

Each term
is nonnegative because the occupancy
and the square is nonnegative. Therefore

For the forcing bound, recall that
is a probability weight on
when
(the occupancy entries are nonnegative and sum to one). Therefore

The proof is complete. 
Two immediate consequences follow. First,
is positive definite (and therefore invertible) for every
. Second, the frozen critic equilibrium
exists and satisfies

because
.
We now derive the energy estimate, which is the main tool for controlling the critic norm.
Proposition 4.4 (Energy estimate for the critic). Let
be a local solution of system (L1) on
. Then for all
,

In particular,
for all
.
The proof has three steps: differentiate
, apply the coercivity and forcing bounds, and solve the resulting scalar differential inequality.
Proof. Step 1 (Differentiate the energy). Compute

Step 2 (Apply the bounds). By Proposition 4.3,

and

Combining gives the raw energy inequality:

The right-hand side of
is a quadratic in
that is negative whenever
. This already shows that the critic norm cannot grow past
from below, but we want an explicit decay estimate. To obtain one, we absorb the
term using the elementary inequality
with
and
:

Substituting:

Step 3 (Solve the differential inequality). Set
. Then

Therefore

The function
is nonincreasing. In particular,
for all
, which gives

Unfolding the definition of
:

Rearranging:

For the “in particular” bound, note that the right-hand side is a convex combination of
and
(with coefficients
and
), hence it is bounded by
. Taking square roots gives

The proof is complete. 
The energy estimate tells us two things. First, the critic norm can never exceed its initial value or the threshold
, whichever is larger. This immediately prevents finite-time blow-up of the critic. Second, regardless of how large
is, the critic norm eventually falls below
for any
— the exponential decay of the first term pulls
toward
from above. This decay is the absorption mechanism that we will use in Section 4.4.
Verification in the Chapter 0 example. In Chapter 0, the critic equation is

with
(a scalar) and
.
Coercivity. The coercivity constant is
, because
for all
. (The extra term
only increases the coercivity.) Therefore
.
Forcing bound. The forcing is
, with
for all
. Therefore
.
Asymptotic critic level. The frozen equilibrium is
, which satisfies
. The level toward which the critic norm decays is therefore
. The absorbing radius fixed in Chapter 2 is strictly larger:
, the value used in Chapter 0. The strict gap between the level and the radius is what makes absorption work in finite time — see Section 4.4.
Energy estimate. The energy estimate becomes

For
, we have
. At
, this gives
; as
, it gives
. The critic relaxes exponentially toward the ball of radius
.
We can also solve the Chapter 0 critic equation directly. For frozen
, the equation
has the explicit solution

Since
, the decay rate is at least
, consistent with the general energy estimate.
Takeaway. The critic is controlled by the coercivity constant
: the energy
decays exponentially toward the level
, with rate at least
. This prevents the critic from blowing up and identifies the level
toward which the critic norm decays. The absorbing radius
is then chosen strictly above that level — Chapter 2 fixes
— as Section 4.4 explains.
4.4 The Compact Absorbing Set
Propositions 4.1, 4.2, and 4.4 together confine the three groups of variables: the actor stays in
, the distribution stays in
, and the critic norm decays toward
. We now combine these results to build a single compact absorbing set for the entire system. In the Chapter 0 example, this set was
; we now construct its general counterpart.
Recall the definitions from Chapter 1. A set
is forward invariant if
for every
: trajectories starting in
stay in
. A set
absorbs a subset
if there exists a time
such that
for all
: every trajectory from
eventually enters
and never leaves. A compact absorbing set is one that is compact, forward invariant, and absorbs every bounded subset of
.
Choosing the absorbing radius. The energy estimate (Proposition 4.4) shows that
decays toward
. For the absorbing set to work, we need a radius
that is strictly larger than
, so that there is a gap between the ultimate critic bound and the boundary of the ball. This gap is what makes absorption work in finite time: trajectories that start with large
need time to decay into the ball, and the strict inequality ensures they eventually get inside.
We use the radius fixed in Chapter 2:

The definition gives both
and
, so the gap satisfies

The lower guard
in the maximum is not decoration. The model permits
— for instance, when all rewards vanish — and in that case the unguarded candidate
would close the gap entirely: the ball
is forward invariant, but the energy
only decays toward zero and never reaches it, so no bounded set would be absorbed in finite time. The guard keeps the gap at least
no matter how small
is.
With this radius, define

In the statement and proof below,
denotes the solution map, used for all
in anticipation of Proposition 4.6: the a priori bounds of Propositions 4.1, 4.2, and 4.4 hold on every interval of existence, and Proposition 4.6 uses only those bounds — not Proposition 4.5 — to show that every solution is global, so there is no circularity.
Proposition 4.5 (Compact forward invariant absorbing set). The set
is compact and forward invariant under the semiflow. Moreover,
absorbs every bounded subset of
: for every bounded set
, there exists
such that
for all
.
Proof. The proof has three parts: compactness, forward invariance, and absorption.
Compactness. The actor box
is compact. The closed ball
is compact (closed and bounded in finite dimensions). The simplex
is compact (closed and bounded). The product of finitely many compact sets is compact, so
is compact.
Forward invariance. Let
, and let
be the solution from this initial condition. By Proposition 4.1,
for all
. By Proposition 4.2,
for all
. By Proposition 4.4,

Since
and
, we have
and
, so

Therefore
for all
, and
. Since this holds for every initial condition in
, the set
is forward invariant.
Absorption. Let
be bounded, and define
. We need to find
such that
for all
and all initial conditions in
. (The actor and distribution coordinates are already confined to
and
by Propositions 4.1 and 4.2, so the only issue is the critic.)
By the energy estimate,

We want this to be at most
. Since
, we need

This holds for all
, where

where the maximum is taken before the logarithm. (If
, then
and
is already inside the absorbing ball.) For
, every trajectory from
satisfies
, and therefore
. 
Verification in the Chapter 0 example. In Chapter 0 the three invariance results combine into a single absorbing set; we now check the gap and the absorption time for an exterior trajectory. With
and
, the radius is
, and the absorbing set is

The gap is
.
For a trajectory starting at
, the absorption time is

After time
, the energy estimate guarantees
, so
. The trajectory has entered the absorbing set.
For a trajectory starting at
, we have
, so
: the trajectory is inside the absorbing ball from the start.
Takeaway. The compact absorbing set
combines the three confinement mechanisms: actor damping, simplex convexity, and critic coercivity. Every bounded orbit eventually enters
and stays there. The entry time depends only on the initial critic norm and the gap
.
4.5 Global Existence and the Semiflow
With the absorbing set in hand, we can close the argument that Chapter 3 left open: the maximal existence time is infinite for every initial condition in
, and the solution map defines a continuous semiflow.
Proposition 4.6 (Global well-posedness). For every initial datum
, system (L1) has a unique global solution

with values in
. The solution map

defines a continuous semiflow in the sense of Definition 1.1.
Proof. Fix an initial datum
. By the Picard-Lindelöf theorem (Proposition 3.6 and its blow-up alternative), the system has a unique maximal solution on
, and if
, the solution leaves every compact set.
By Propositions 4.1 and 4.2, the actor and distribution remain in
and
for all
. By the “in particular” bound in Proposition 4.4, the critic satisfies

Therefore the entire trajectory remains in the compact set

The blow-up alternative from Picard-Lindelöf (Proposition 3.6) states: if
, the solution must leave every compact subset of the ambient space as
. But the trajectory stays in the compact set
, which is a contradiction. Therefore
.
The semiflow properties —
,
, and continuous dependence on initial data — follow from the uniqueness part of the Picard-Lindelöf theorem and the global existence we have just established. The semigroup property
holds because the system is autonomous: the vector field does not depend on time, so restarting from
and running for time
gives the same result as running from
for time
. Continuous dependence on initial data is part of the Picard-Lindelöf theorem and extends to all
because the solution remains in a compact set. 
This completes the construction that Chapters 1 and 3 set up. System (L1) generates a continuous semiflow
, with the compact absorbing set
from Proposition 4.5. We now have every ingredient that Definition 1.3 required for an absorbing set:
is compact, forward invariant, and absorbs every bounded subset of
.
The semiflow and its absorbing set are the starting point for Chapter 5, where we will extract the global attractor — the smallest compact set that captures the long-time behavior of every trajectory.
4.6 Summary and Bridge Forward
This chapter has established the a priori estimates that make system (L1) globally well-posed and dissipative. The main results are:
Actor-box invariance (Proposition 4.1). The damping matrix
vanishes at the boundary of
, and the arctanh barrier argument shows that no interior trajectory can reach the boundary in finite time. Boundary trajectories stay on the boundary forever. The actor parameter is confined to
for all time.Simplex invariance (Proposition 4.2). The variation-of-constants solution of the law equation expresses
as a convex combination of simplex elements with nonnegative coefficients summing to one. The distribution stays in
for all time.Critic coercivity (Propositions 4.3 and 4.4). The coercivity constant
gives a uniform lower bound on the symmetric part of the critic matrix. The raw energy inequality
from Proposition 4.4 exposes the quadratic structure
; after Young’s inequality the resulting energy estimate shows that
decays exponentially toward
, with rate at least
.Compact absorbing set (Proposition 4.5). The set
is compact, forward invariant, and absorbs every bounded subset of
, with an explicit absorption time that depends on the initial critic norm.Global well-posedness (Proposition 4.6). The a priori bounds prevent blow-up, extending the local semiflow of Chapter 3 to all positive time. The solution map
is a continuous semiflow.
In the Chapter 0 example, these results recover the absorbing set
with
,
, and
.
The system now has everything needed for the next step: a continuous semiflow on a metric phase space with a compact absorbing set. What we do not yet have is the global attractor — the smallest compact invariant set that attracts every bounded subset. Extracting the attractor from the absorbing set requires one more ingredient (asymptotic compactness or, in finite dimensions, the simpler fact that a continuous map on a compact set has a compact omega-limit set). That construction is the subject of Chapter 5.
Exercises
Exercise 4.1 (Scalar barrier argument at the boundary). In the Chapter 0 model, carry the scalar barrier argument for the actor at
. Specifically:
(a) Write the actor ODE at the boundary: show that
and that the right-hand side is zero when
.
(b) Define
. Show that
.
(c) Conclude that if
, then
for all
.
Exercise 4.2 (Direct simplex check). In the Chapter 0 model, verify directly (without using the variation-of-constants formula) that the law equation preserves total mass.
(a) Write the full two-component law equation:
and
.
(b) Show that
, so that
is constant.
(c) Show that if
, then the variation-of-constants formula gives
for all
.
Exercise 4.3 (Coercivity constants). In the Chapter 0 model:
(a) Verify that the critic matrix is
and that
.
(b) Verify that the forcing is
and that
.
(c) Solve the energy inequality
explicitly. Verify that the solution is
.
(d) For
, compute the smallest time
such that
.
Exercise 4.4 (Absorbing-set entry time). In the Chapter 0 model with
:
(a) Compute the gap
.
(b) For a trajectory starting at
, compute the absorption time
.
(c) Verify numerically that the energy estimate gives
.
Exercise 4.5 (What happens without coercivity). Set
in the critic matrix, so that
.
(a) Give an example of critic features
in
for which the resulting matrix
is not invertible. (Hint: choose all features to point in the same direction.)
(b) Explain why the energy estimate breaks down when
. What specific step in the proof of Proposition 4.4 fails?
(c) Describe qualitatively what could happen to the critic trajectory
without the coercivity bound.
Exercise 4.6 (Removing the damping). Replace the damping function
with
(constant damping, equivalent to no damping).
(a) Is the resulting vector field still locally Lipschitz?
(b) Does local existence still hold?
(c) Explain why actor-box invariance fails. Construct a specific scenario in the Chapter 0 model where
escapes
in finite time. (Hint: if the raw drift
has a definite sign near the boundary, there is nothing to stop
from crossing it.)
Exercise 4.7 (Absorbing set under parameter changes). The absorbing set
depends on the data through
,
, and
.
(a) If we shrink the actor-box radius
(keeping all other data fixed), does the absorbing radius
change? Why or why not?
(b) If we increase
(stronger regularization), what happens to
? Note what happens once
drops below
.
(c) Explain why increasing
makes the system “more dissipative” and leads to faster absorption.
Exercise 4.8 (Nonlinear critic). Suppose the critic equation were nonlinear:
, where
is a smooth map satisfying
for some
and all
large enough.
(a) Can you still derive an energy estimate of the form
for large
?
(b) Does the absorbing-set construction still work? What is the main difference from the linear case?
(c) Explain why the linear structure of the critic in system (L1) gives a cleaner estimate than the nonlinear alternative.