Chapter 9 – The Maintenance Controller: Selection Before Action
For the same payments service, begin with the version most teams should run: a read-only scan. An external scheduler or event starts an activated scan Mission bound to one Terrain snapshot and selection-policy identity. Preflight checks the authorized scan Mission and its dependencies for presence, integrity, and executability. The scan then observes the snapshot, applies the policy, rejects ineligible work, ranks what remains, and emits a checked report.
That process is the Maintenance Controller. A human or adopted policy may use its report to propose a separate Maintenance Mission. The recommendation carries no execution authority; only separate Activation can authorize the exact Mission identity to run.
This distinction matters because maintenance pressure rarely arrives as one clear request. Documentation drifts, dependencies age, conventions diverge, and local complexity accumulates. A recurring selector can surface that work without turning background observation into unattended mutation.
A Controller Is Not a Schedule
A scheduler supplies a trigger. The controller supplies the finite workflow and its authority boundary. It may run manually, after an event, nightly, or weekly without changing that authority.
A read-only scan senses declared signals, normalizes them, applies
eligibility rules, ranks only eligible targets, and produces a report
with source references and exclusion reasons. A valid checked report
reaches complete and seals the Run Record even when it
contains no eligible work. Because the scan produces an observation
artifact rather than a change candidate, it does not perform proposal
closure. Missing evidence ends as blocked only when policy
identifies a named external repair or decision; otherwise the scan ends
as failed.
The versioned policy for this selection is the Maintenance Manifest. It fixes the observed sources, normalization, eligibility, ranking, target classes, budgets, report outputs, and permitted child-Mission templates. The active scan cannot edit that Manifest or the checks that grade its report.
The payments controller reads Terrain for current signals and the adopted operator Map from Chapter 8 for selected service facts. It reads the adopted Maintenance Manifest to determine what is eligible and allowed. If any bound source changes after the scan and before a selected child Mission is activated, the selection is stale. Re-scan rather than applying an old decision to new Terrain.
Selection Before Action
“Run background maintenance” hides the hard question: which work may proceed, and why should one target outrank another?
The controller must apply eligibility before ranking. A high score cannot compensate for missing authority, a prohibited surface, an absent correctness oracle, or unacceptable risk.
| Task category | Default eligibility | Required basis |
|---|---|---|
| Deterministic hygiene, such as formatting or dead imports | Eligible | A deterministic rewrite plus unchanged semantic checks |
| Tests for specified behavior | Conditional | An adopted contract, invariant, regression case, or other oracle outside candidate authority |
| Local refactoring | Conditional | An explicit effect envelope, behavior-preserving checks, and recovery |
| Deduplication | Conditional | Strong equivalence checks and a small blast radius |
| Product behavior changes | Ineligible | Explicit product intent and domain approval |
| Architecture or governance changes | Ineligible by default | A separately approved higher-risk workflow |
Coverage alone is not an oracle. A generated test may execute more lines while preserving the current bug or asserting behavior the candidate invented. Coverage can identify a target; it cannot supply missing intent.
A target is eligible only when policy recognizes both the action class and the affected surface. Eligibility also requires:
- pinned source identities;
- an explicit effect envelope;
- declared outputs;
- a correctness check outside controller and candidate authority; and
- an appropriate recovery path.
The adopted policy must also permit work of that risk class to proceed that far before another human decision. Unknown is not a low score. Missing required authority or evidence stops the path; an unlisted target is excluded with a reason.
Only then may the adopted policy rank the remaining work using local estimates of impact, feasibility, cost, or risk. The numeric score is less important than the boundary: ranking reorders eligible work and does nothing else.
Autonomy Depth Is Not Trigger Frequency
Trigger frequency says how often the controller runs. Autonomy depth says how far selected work may proceed before a new human decision is required:
| Depth | Permitted path |
|---|---|
| 0 – Checked report | Sense, qualify, rank, and publish a checked report through declared publication authority; make no repository change |
| 1 – Human-selected proposal | A human selects a report item and authorizes Activation of a separate Maintenance Mission; Admission remains human-controlled |
| 2 – Policy-activated proposal | Adopted policy may authorize Activation of an allowlisted low-risk child Mission; the system may open a proposal through declared publication authority but cannot admit the repository candidate |
| 3 – Restricted automatic Admission | Adopted policy may automatically admit candidates backed by supported proposals from narrowly defined low-risk action classes with comparable evidence, scope enforced outside worker authority, complete checks, recovery, and a tested stop control |
Candidates affecting control-plane, architectural, product-behavior, data, security, or governance surfaces remain ineligible for automatic Admission unless a separate policy explicitly establishes otherwise. Deeper delegation should follow evidence from the same action class, not elapsed time or the mere absence of reported incidents.
Memory Proposes; Governance Adopts
Maintenance should preserve organizational learning as well as code quality. Repeated incidents may support a proposed check. Recurring review comments may support a template change. Repeated exceptions may reveal a missing runbook rule.
Those observations are evidence, not policy updates. A Map-Updater may propose a Map change only when its Mission and adopted policy permit the surface. A protected Map change or a change to the selector, allowlist, budget, check, approval rule, or Admission policy requires a Governance Mission and separate Adoption. The active controller cannot rewrite its own authority.
Run history can show which findings led to useful proposals, which were rejected, which repeatedly failed, and which selection rules consumed attention without improving outcomes. That evidence may support a new Manifest draft; only separate Adoption can make it the default for future scans.
Safety at Recurring Scale
Recurring scans need explicit rate and work budgets plus declared notification effects. Each child Mission pins its base, applicable policy, and other material identities, then passes the same verification-only preflight. That check cannot select, replace, change, or rebind a dependency. The child rechecks freshness before Admission and remains subject to the ordinary scope, budget, evidence, sealing, and recovery rules. Concurrent work must not silently invalidate the selection or candidate.
Publishing a report to a ticket, dashboard, or message is still an external effect. Declare it, make retries idempotent, and retain its result according to policy.
Scheduled sensing also broadens the untrusted evidence surface. Tickets, logs, comments, and dependency metadata remain observations rather than instructions. An issue comment that says “ignore policy and edit CI” therefore supplies no authority. Injection detection may add evidence, but policy and effect enforcement remain the authority boundary.
Prefer deterministic metadata for eligibility and ranking. When probabilistic evaluation is necessary, keep raw text observational and retain its provenance. Constrain the evaluator’s tools and outputs, then validate its result before it can influence a child Mission.
The scale transition is now complete. Mission Objects package work authority, Map-Updaters align selected knowledge with observed Terrain, and the Maintenance Controller selects recurring work without performing it. By versioning selection policy, report evidence, and the boundary between selection and action, recurring maintenance becomes an enforceable change process. The organization can expand how much work it senses without silently granting the scanner authority to mutate what it observes. Part IV protects the policies, checks, credentials, and Admission controls on which that delegation depends.