Appendix A: Failure Modes and Limits
A governed loop offers conditional guarantees, not universal correctness. Under its stated trust assumptions, it can establish that effects stayed within named limits, required checks ran, and transitions followed policy. It can also establish that the Run Record was sealed with retained evidence and the declared handoff followed its rules. It cannot establish that the Map was wise, that the Validators covered every relevant property, that hidden dependencies did not correlate two evaluators, or that the governance cost was justified.
The following failure modes mark that boundary. Some identify conditions under which the architectural guarantees fail; others identify limits the architecture cannot remove or cases in which the controlled path is not worth its cost.
Mis-scoped Context
A Context Packet can be too small to contain a necessary contract or too large to preserve relevance. In the first case, the generator works without a fact needed to satisfy the Mission. In the second, unrelated material competes with the active constraints and enlarges the surface on which plausible but unauthorized structure can appear.
Deterministic extraction does not make a bad slice sufficient. It makes the slice reproducible. The evidence must therefore distinguish omitted authority, irrelevant inclusion, and unauthorized effects rather than collapsing all three into “the model failed.”
An effect boundary outside candidate authority can reject an
overbroad candidate. It cannot recover an omitted truth. If the active
authority cannot support a sufficient Context Packet, the run ends
blocked only when it names an external repair or decision;
otherwise it ends failed. Either outcome seals the Run
Record. The workflow cannot silently widen its own scope.
Non-convergence
Refinement does not guarantee convergence. Findings may repeat, alternate, or disappear without a meaningful improvement. Constraints may conflict. A requested state may be unreachable under the active budgets and effect envelope. A probabilistic evaluator may also vary enough to create the appearance of progress where none exists.
The relevant evidence spans attempts: candidate identities,
normalized findings, state signatures, evaluator protocol identities,
and progress measures. A finite loop responds to absent progress by
stopping. Ordinary exhaustion, oscillation, or repeated failure ends as
failed. blocked is reserved for a named
external repair or decision, such as granting missing authority,
restoring a trustworthy oracle, resolving contradictory Mission terms,
or repairing an environment the run cannot change. Without such a named
action, the run ends as failed.
Finite failure is part of reliability. A loop may fail to establish progress and still behave reliably by reaching a terminal result and sealing reconstructable evidence. Continuing until a favorable answer appears is the less reliable behavior.
Map Contamination
Map Contamination occurs when generated structure is mistaken for observed or adopted structure. A model invents a route, signature, schema field, dependency, or policy statement; a later step records that invention as if a Sensor had extracted it from Terrain or an authorized process had adopted it. Subsequent runs then optimize against a self-confirming fiction.
The Skeleton-First rule prevents one common form of this failure in two directions. In Map-to-Terrain work, an adopted design supplies a prescriptive minimum scaffold and permitted expansion rules. In Terrain-to-Map work, deterministic extraction records descriptive observed structure. Derived views remain traceable to their observations; judgment-heavy summaries remain candidates. A derived view is authoritative only for the selected descriptive property and source revisions named by the Mission. A judgment-heavy summary cannot overwrite prescriptive Map content or become Map authority without separate Adoption.
The deeper invariant is authority direction. A Terrain-to-Map workflow may report observed structure but may not make that structure intended policy. A Map-to-Terrain workflow may propose implementation from adopted intent but may not rewrite the intent that judges it.
Post-processing cannot erase an authority violation. If a candidate attempted undeclared effects, the complete attempted footprint remains evidence even when a deterministic cleanup could produce an apparently compliant diff. The candidate is rejected as attempted; a new candidate may be generated under corrected authority and context.
Validator Error and False Confidence
A Validator can be wrong. It can reject an acceptable candidate, accept a defective one, encode a stale rule, or test a property too weakly. Completing every required check therefore means only that the declared evidence requirements were satisfied. It does not mean that every important property was declared or that the resulting software is good in every relevant sense.
Probabilistic evaluation adds instability and correlation risk. Shared models, prompts, retrieval, fixtures, or context can make generator and evaluator agree on the same wrong answer. Chapter 3’s protected protocol makes that uncertainty more visible; it does not abolish it.
For a high-consequence property, a probabilistic evaluator cannot be the sole oracle unless adopted policy explicitly accepts its measured error characteristics for that Mission class and keeps the evaluation protocol outside candidate authority. Repeated false-positive or false-negative evidence may justify a Validator-change proposal. The active candidate cannot soften the rule that evaluates it.
Local Success Without Joint Coherence
The proposed model ordinarily assumes that concurrent workflows affect surfaces of shared Terrain that are independent for the property in question. Declared interfaces define cross-surface semantics; later reads, events, and Sensors carry versioned observations of admitted changes. Under declared delivery, freshness, and read-availability assumptions, dependent workflows gain eventual visibility of relevant changes. This is a limited analogy to eventual consistency, not a claim of one global snapshot or replicated-state convergence.
That assumption fails when concurrent work depends on observations another Admission can invalidate, non-commuting effects, changing interface semantics, spanning invariants, or limited capacity. Each workflow may satisfy its own Mission while their combination creates a collision or aggregate harm that neither Mission measures. Individual eligibility is not evidence of joint eligibility.
Chapter 12 owns the coordination rule for this failure. When
independence ends, current policy must govern the shared property
against current state at the protected Admission or effect boundary.
Without that path, the combined result remains unsupported even if each
local run is complete. Eventual visibility permits delay
across independent surfaces; it cannot replace atomic or joint judgment
where the invariant requires it.
Completion Without Admission
Chapter 10’s boundary remains decisive. For a proposal-producing run,
complete records proposal closure: the final artifact is a
supported proposal, and the Run Record is sealed. It does not make the
candidate Terrain. Admission revalidates the current base, policy,
identity, approvals, and the integrity and provenance of required
evidence. Any changed or missing predicate can deny the proposal without
reopening the sealed Run Record.
Automatically applying every supported proposal from a
complete run would collapse validation into authorization.
The same rule applies to Maps and policy: history may support a
proposal, but only separate Adoption changes authority.
Trust-Root Failure
The Ledger is contradiction-resistant only under a protected sink and trust root. Hash chains, storage administered outside runner authority, integrity checkpoints, and access control can expose missing, reordered, or conflicting evidence. They cannot prevent a sufficiently privileged actor from replacing both the evidence and the mechanism used to verify it.
Likewise, a Scope Guard is meaningful only when the runtime, credentials, execution isolation, and admission path are outside candidate control. If an ordinary Mission can rewrite its own Validator, conceal network effects, replace its evidence, or bypass the Mission Gate, the governed-loop claim no longer applies. The architecture makes the trust boundary explicit; it does not make that boundary invulnerable.
Economic Failure
A governed loop can cost more than it saves. Validator maintenance, false-positive handling, evidence retention, protected execution, and review all consume engineering capacity. For low-risk or infrequent work, ownership may cost more than the recurring work justifies.
The architecture does not guarantee productivity or strategic advantage. A workflow that is reliable but uneconomic should remain narrow, run with less autonomy, or not exist.
The Claim That Remains
Software Development as Code does not promise perfect specifications, complete oracles, inevitable convergence, tamper-proof history, or cheaper delivery everywhere. Its narrower claim is stronger because it is testable: under declared authority and trust assumptions, selected effects, evidence requirements, transitions, and Admission or Adoption decisions can be made explicit and enforceable.
That changes the shape of failure. Missing evidence stops the run. Scope violations remain visible. Non-convergence ends finitely. Supported proposals remain separate from admitted Terrain and adopted Maps or policies. History may inform future authority but cannot install it.