Part II – Reliability Foundations
10 min read

Chapter 5 – The Bounded Refinement Loop: Convergence, Thrashing, and Finite Stopping

This chapter switches from Part I’s documentation examples to a tax calculation because schema and rounding failures provide distinct, comparable findings across attempts. Chapter 6 continues the same case to show how bounded context supports correction.

One bounded trace shows the pattern:

Stage Evidence
Attempt 1 Schema and tax-rounding checks fail; the candidate changes two files
Attempt 2 Only the tax-rounding check fails; the candidate changes one file
Attempt 3 All fast checks pass; one smaller change remains
Final checks All required checks and declared output checks pass
Result The run ends complete with a supported proposal; the Run Record is sealed

The same fast checks and versions ran on the first three attempts, so their findings are comparable. A prior failure disappears, the actual footprint contracts, and proposal closure occurs only after all required evidence is available.

This does not prove that every retry loop converges. It shows what evidence justifies calling this trace convergent. A circuit breaker can guarantee finite execution. Convergence requires more: reachable acceptance conditions, feedback that can guide a useful correction, and measurements stable enough to distinguish progress from newly exposed failure.

The Bounded Refinement Loop makes that policy explicit:

Finite adaptation without changing authority

The Bounded Refinement Loop

Findings may guide another proposal. They do not change the Mission, grant new effects, or remove the stopping rules.

1. Propose

Produce one candidate from the activated work contract and current working state.

2. Check

Run the required checks and retain normalized findings, omissions, and attempted effects.

3. Decide

Select one transition permitted by the Mission from the findings and remaining budgets.

4. Refine or seal

Use specific findings for another attempt within budget, or reach one terminal result and seal the Run Record.

complete

The run met its completion conditions. A change candidate is now a supported proposal; Admission or Adoption remains separate.

no_change

The required property already held and no candidate effect exists.

blocked

The Run Record names the external authority decision, trustworthy-oracle restoration, or environment repair needed for a new linked run.

failed

A breaker or failure condition ended execution, and no named external action can resolve it.

Working state

Latest candidate, comparable findings, completed checks, recent normalized states, and remaining budgets.

Reconstruction evidence

The logically append-only account of each attempt, effect, finding, and transition, retained outside candidate authority.

Each attempt may use the previous candidate and normalized Validator findings. The Mission’s authority, permitted effects, budgets, and acceptance criteria remain fixed. Validator pass and failure results are evidence, not lifecycle outcomes.

Every run ends in one of four ways: the candidate satisfies the completion conditions; the required property already holds and no change is needed; the work stops pending a named external decision or repair; or execution stops unsuccessfully and preserves its evidence.

The state machine names those outcomes complete, no_change, blocked, and failed.

From One Attempt to Refinement

Chapter 2 bounded one attempt as Prep, Model, and Validation. Refinement adds retained state and a transition policy:

  1. Propose: the bounded change step receives the fixed Mission, scope, and permitted prior findings, then produces one candidate.
  2. Check: Validators inspect that candidate and emit structured findings. Declared output checks run here as required evidence.
  3. Decide: when evidence is complete enough for the current stage, the Judge selects one allowed transition.
  4. Refine: if another attempt remains justified, normalized findings inform the next candidate without changing authority or scope.

An empty finding set is not enough for acceptance. Every required check must have run, and every declared output must have been checked. Required approvals must also be present, and the retained records must agree.

The Judge Is a Role

Chapter 1 needed no separate Judge because each result of its scope-matched check had one fixed destination. A Judge becomes useful when complete evidence can support more than one allowed route or when routing requires interpretation.

The role may be implemented by a routing table, deterministic heuristic, bounded model, human, or composition.

A model-based Judge inherits Chapter 3’s probabilistic-evaluation contract. Its output is schema-constrained, provenance-bearing, and limited to declared decisions.

Before applying any proposed transition, the runtime checks that required evidence is complete, declared outputs have passed their checks, the candidate stayed in scope, the Mission and budgets remain unchanged, and the transition is allowed. It rejects any decision that invents findings, widens scope, changes the Mission or budget, selects an undeclared transition, or attempts admission.

Human escalation should not become a routine retry stage. It is justified only when the run can name the external decision or repair required: missing authority, a novel policy conflict, an exception bound to this Mission and state, a high-consequence choice reserved for human authority, or the absence of a trustworthy oracle. The current run ends blocked and its Run Record is sealed; a resolution may authorize a new linked run but never reopen the old one.

State, Results, and Sealing

The state machine is small:

flowchart TD
  P[Propose] --> V[Check]
  V --> D{Decide}
  D -->|retry| R[Refine from findings]
  R --> P
  D -->|terminal| T[Seal Run Record]

The terminal node seals a Run Record with exactly one result: complete, no_change, blocked, or failed.

An early-stage failure does not need results from checks that could not yet run. It does need a complete account of why they did not run. A parse failure may therefore prevent later checks from running when each omission has a deterministic reason. For proposal closure, however, every required result must be complete and the applicable policy-bound decision rule, including any exact-exception condition, must be satisfied.

The workflow keeps two kinds of memory:

The next model request should contain bounded working state rather than the entire run history. Detailed evidence may remain reference-backed as long as retention preserves reconstruction. The Ledger guarantee is conditional on the protected sink and trust root developed in Chapter 12.

Validator results and terminal outcomes answer different questions:

Seal every terminal Run Record. Later intervention creates a linked run. A revert is a workspace transition back to the baseline; failed is the terminal result of the unsuccessful run, with evidence retained in the Run Record.

The Judge may choose the accept transition only after every required result is complete, the applicable policy-bound decision rule—including any exact-exception condition—is satisfied, and the evidence required for proposal closure is complete and internally consistent. That transition closes the proposal. It does not admit the candidate into Terrain or adopt it as a Map or policy.

Feedback Drives Adaptation; It Does Not Change Authority

Inside one run, the model’s weights do not change. The workflow adapts because normalized findings inform the next candidate while authority, scope, budgets, and acceptance criteria remain fixed. Feedback guides the search; it does not redefine the target or grant permission.

The feedback contract should identify the previous candidate, name specific findings, and require one output shape. Unbounded logs followed by “try again” make measurements harder to compare and invite scope leak.

Convergence and Thrashing

Convergence is a property of a trace: the refinement policy reaches the declared acceptance conditions before a breaker fires. Thrashing is also a property of a trace: attempts accumulate while comparable evidence shows no net progress.

Convergence is always relative to one fixed activated Mission. The run may use only its declared routes and correction strategies; it cannot change the target, authority, scope, budgets, or required checks to manufacture progress. If adopted intent or the relevant Terrain changes materially, later alignment requires a new linked Mission. If run evidence suggests that the workflow or policy itself should change, that evidence may support a separately governed proposal and Adoption. Neither change belongs inside the current refinement loop.

A finite run is not necessarily convergent. Stopping an impossible or unproductive loop as failed is correct termination. no_change is different again: the starting state already satisfied the request.

Useful progress signals include:

Failure counts are comparable only when the completed-check sets match. Fixing a parse failure may reveal ten semantic failures and still advance the evidence frontier. Progress need not occur on every attempt, but it must appear inside a declared window.

Reachability and Feedback Quality

Acceptance conditions state what a candidate must satisfy to become a supported proposal: parsing and scope, declared intent, all required checks, output evidence, approvals, and internally consistent records. They do not prove that a candidate can reach those conditions.

Three diagnoses matter:

If a human-written candidate cannot pass the fixed checks inside the declared scope, generation is not the problem.

Oscillation and Failure Routing

Thrashing often appears as A-to-B-to-A behavior: fix one property, regress another, then repeat an earlier candidate and finding set. Retaining the normalized candidate state, check plan, completed check set, and failing reason codes makes that recurrence detectable. If the same combination recurs inside the declared progress window, the run must stop; the workflow has returned to the same state.

Not every failure justifies another candidate:

Evidence class Allowed response Unresolved outcome
malformed output or ordinary correctable finding retry with the specific finding while progress and attempt budgets remain failed
ordinary out-of-scope effect reject the full candidate; retry only if policy permits correction without widening scope failed on repetition
effect on a protected surface deny and record the attempted effect; use a separately authorized workflow if the change is intended failed
missing approval or contradictory Mission stop candidate generation and identify the authority decision or Mission repair that could resolve it blocked only if that named decision or repair can resolve the condition; otherwise failed
transient check infrastructure failure retry the check without regenerating the candidate blocked only if a named environment repair can restore trustworthy evidence; otherwise failed
missing or contradictory evidence fail closed and preserve the integrity finding; name the external repair or decision if one can resolve it blocked only if that named action can resolve the condition; otherwise failed

Fail fast means stop the current validation path. It does not mean every failure has the same next transition.

Staged Checks

Expensive checks need not run on every attempt. A stable fast set supports comparable refinement; every remaining required check must still run before acceptance. A candidate that clears parsing, scope, lint, types, and narrow tests may advance to integration, environment, performance, or security checks. Failure during final checks follows the same routing rules; it does not erase the earlier evidence or justify a second Judge.

Finite Stopping

Circuit breakers make stopping deterministic even when candidate generation is stochastic:

Breaker Stop condition Result
attempts, time, or spend declared budget consumed failed
footprint actual effects exceed policy reject and route by evidence class
minimum progress no comparable progress inside the declared window failed
oscillation the same normalized candidate and check results recur failed
infrastructure retries the check cannot produce trustworthy evidence within budget blocked if a named environment repair can restore trustworthy evidence; otherwise failed
evidence integrity required records are missing or contradictory blocked if a named external repair or decision can resolve the condition; otherwise failed

There are no universal limits. Set them from consequence, task surface, review capacity, and prior evidence before final review and Activation. They remain fixed for the run; a different limit requires a new Mission and linked run.

That is enough to distinguish convergence from luck, newly exposed failure from regression, and finite failure from uncontrolled retry. The generator may remain probabilistic; the workflow keeps authority fixed, makes feedback comparable, and knows when another attempt is no longer justified.

Refinement is therefore part of the enforceable change process, not an agent being told to keep trying. An organization can delegate more correction work while retaining fixed limits on what may change, what counts as progress, and when continued effort must stop.

Share