Chapter 4 – The Stochastic Generator: Variance, Validation, and Feedback
Two representative candidates come from the same fixed request and Context Packet:
Propose one short usage example from the supplied signatures and pinned behavioral tests. Change only the Map’s
Usageblock. Do not invent behavior.
Run 1:
country = normalize_country(" is ")
tax = calculate_tax(100, country, 24)
assert country == "IS"
assert tax == 24Run 2:
country = normalize_country(" is ")
tax = calculate_tax(100, country, 0.24)
assert country == "IS"
assert tax == 24Both are plausible enough to pass weak checks for valid Markdown,
valid function names, and permitted scope. Only Run 2 survives
behavioral execution against the pinned source. Run 1 turns
24% into the numeric rate 24, so its assertion
fails.
That gap is the chapter’s subject. The same task and apparent intent produced candidates with different relationships to the governing evidence. This is drift on a bounded surface. It is why the governed Map surface and observed Terrain must remain distinct, with authority declared for the specific property.
A Map is the structured representation being governed. It may express adopted intent or record selected observed structure. Terrain is the system state the workflow observes or proposes to change.
Part I established the mechanism before the theory. Part II explains why it needs protected checks and enforced limits, which the book calls Physics. Lower-level evidence may inform higher-level Maps, while adopted Maps constrain the workflows below. Variance is the starting condition those workflows must manage.
Large Language Models (LLMs) are Stochastic Generators: probability models that produce plausible outputs, not guaranteed ones. Reliable automation puts deterministic checks around generation. The full Software Development as Code (SDaC) runtime is the SDaC Engine; the model is one component inside it.
LLMs Are Probability Generators
At its core, an autoregressive LLM assigns probabilities to possible next tokens. The decoder then selects or samples a token according to its configuration. Small differences early in that sequence can produce materially different candidates.
Three implications matter in SDaC:
No guarantee of identical output: With sampling enabled, the same task request can produce different candidates across runs. Even when a provider supports greedy decoding or a temperature of
0, exact replay is warranted only when the full inference path provides that guarantee.Contextual sensitivity: Small changes in input or context can shift the distribution and change the output.
Small variations can cause distant effects: A minor output change can alter another interface, file, or behavior without producing a simple local failure path.
This is why SDaC treats generation as one component in a system, not as the plan itself: variance is measured, constrained, and gated.
Lehman’s Categories Depend on the Boundary
Lehman’s Specifiable / Problem-solving / Evolutionary (S/P/E) frame classifies a program by its relationship to its specification and problem environment, not by the tool or implementation technique used.1 The analogy is useful here when applied to a declared task or system boundary:
- S-Type (Specifiable): A formal specification completely defines the problem and its criterion of correctness. A schema check against a pinned schema can be treated as S-Type for that selected property; this does not make every use of the validator or the surrounding system S-Type.
- P-Type (Problem-solving): A practical solution requires approximation, and its acceptability is judged against the problem environment rather than a complete formal specification. Open-ended generation often performs P-Type work; a tightly specified generation subtask need not.
- E-Type (Evolutionary): The deployed system becomes part of the real-world activity it supports. The system and that environment influence each other, so continued usefulness requires continued adaptation.
Lehman’s First Law says an E-Type system must be continually adapted, or it becomes progressively less satisfactory. Raw one-shot generation does not solve that evolutionary problem: the system and its environment continue to change while the generator sees only a bounded, partial representation.
Many raw AI-assistant tasks therefore ask a Stochastic Generator to perform P-Type work directly on an E-Type system: an approximate change step touching a codebase embedded in changing reality.
An unconstrained Stochastic Generator produces variants, some of which violate the realities of the E-Type environment. The variance is stochastic drift; a violated constraint is a validation failure.
In that common case, SDaC places repeatable stages around approximate generation and enforces its authority and effects separately.
The SDaC Synthesis: P-Type Work Between Deterministic Stages
For a fully specified property, a Validator can be treated as an S-Type component: a parser, pinned-schema check, type boundary, or other exact check. Rule-based and behavioral grounding add authored constraints and observed cases where formal completeness is impossible.
This gives the Deterministic Sandwich from Chapter 2 its role. In the
minimal pattern, deterministic Prep and
Validation make input and output handling repeatable.
Declared authority, effect limits, and runtime enforcement contain the
work. None of those choices makes every task, check, or system
permanently one type.
Every Validator supplies findings at an enforceable boundary. Its grounding may be formal, rule-based, or behavioral, but its recorded result is still explicit evidence tied to the measured property and available for routing, including by a Judge when one is needed.
Generative AI Changes the Rate, Not the Evolutionary Problem
The S/P/E mismatch describes the risk around one generated candidate. Lehman’s broader laws describe pressures that emerge as an E-Type system changes repeatedly.2 They apply to the evolving system and its development process, not to the model in isolation. The applications below are the author’s synthesis.
The laws matter here as three pressures:
- Demand: The First Law, Continuing Change, and the Sixth, Continuing Growth, say an E-Type system must keep adapting and developing useful capability. Generative AI makes candidates cheaper, but it does not establish that they reflect current intent, fit the system, preserve existing obligations, or deserve Admission or Adoption.
- Integration cost: The Second Law, Increasing Complexity, and the Seventh, Declining Quality, describe how repeated change accumulates abstractions, dependencies, and exceptions while the operating environment moves away from old assumptions. Deliberate simplification and refreshed evidence must therefore accompany generation; otherwise, stale checks let a generator optimize an obsolete target.
- Absorption and feedback: The Fourth Law, Conservation of Organisational Stability, and the Fifth, Conservation of Familiarity, describe limits in activity and shared understanding. The Third Law, Self-Regulation, and the Eighth, Feedback System, describe the resulting multi-level feedback process. Faster candidate production can return as review queues, failed checks, reversions, or delay. Explicit budgets and stopping rules expose those pressures, while the Ledger keeps observation, candidate production, validation, Admission, and Adoption distinct.
These laws are empirical tendencies, not formal invariants, and the evidence behind them was uneven, particularly for the Fourth and Seventh Laws. Generative AI may change their rates. It does not abolish the underlying need to keep an E-Type system useful under changing conditions.
Chapter 14 returns to the strategic consequence: candidate-generation speed matters only as part of the full time to a governed response. It names and measures that interval as adaptation latency, then develops the author’s competitive hypothesis for environments where other organizations also accelerate search and implementation.
Operating a Stochastic Generator
Before final review and Activation, bind every selectable inference input that materially affects execution to the Mission: model and provider identity, decoding controls, request-template identity, context-source identities and selection rules, and tool versions. During execution, record the exact request and Context Packet, the actual model, provider, and tool identifiers, and any provider-supplied run identifier. A mismatch with the bindings fails; it cannot update the active contract. Measure that bound workflow rather than copying a universal temperature or seed. A tight output contract matters more than a fashionable parameter value.
Better models and prompts can improve candidate quality. Model and prompt selection should favor the least expensive combination that reliably reaches the required checks within the declared attempt, time, and cost budgets. Mechanical edits under strict gates may need little capability; ambiguous cross-surface work may need a stronger model or a human decision. Alternative comparisons are meaningful only inside the same bounded workflow and checks.
Internal Reasoning Does Not Replace the Loop
Models may deliberate before producing an answer, and some APIs expose summaries of that reasoning.3 Unrecorded private deliberation is not evidence. A recorded summary is evidence of the rationale expressed, not independent corroboration of the candidate’s substantive claim. In a tool-enabled runtime the model may request a compiler or test run; evidence of that run comes from the recorded tool result, not from the model’s prediction of it.
Reasoning capability may improve the change step. It does not by itself corroborate the candidate or make “thinking hard” equivalent to a test passing.
The Drift Experiment: Measure Structural Variance
The experiment repeats one bounded task under fixed inputs and controls. Its analysis distinguishes textual variation from structural variation, failures, and the time or attempts needed to reach a supported proposal. Two patches may differ in whitespace yet make the same structural change; two polished patches may touch different interfaces or files.
This is a local operating observation, not a general model benchmark. If unwanted structural variation or failure remains high, narrow the context and edit region, strengthen the checks, or use a more capable generator. The point is to make variance visible enough to govern.
Why Ad-hoc Instructions Stop Scaling
More specific instructions can reduce variation, but they cannot eliminate sampling or enforce their own interpretation. Extract the skeleton and edit region deterministically, generate only inside that boundary, and reject unwanted structural variance with Validators. Instructions shape a candidate; Physics decides whether it may advance.
When Validators turn variance into specific failures, the next question is whether correction converges or merely repeats. Chapter 5 makes that correction process bounded and measurable.
Meir M. Lehman, “Programs, Life Cycles, and Laws of Software Evolution,” Proceedings of the IEEE 68, no. 9 (1980). Paper.↩︎
Meir M. Lehman and Juan F. Ramil, “Rules and Tools for Software Evolution Planning and Management,” Annals of Software Engineering 11 (2001), 15–44. The authors reported that support remained under investigation for the Fourth Law and that the available data neither supported nor negated the Seventh. Paper.↩︎
Provider APIs distinguish private reasoning from exposed summaries and recorded tool outputs. OpenAI model documentation, Google Gemini documentation.↩︎