The merge verdict isn't the LLM's to give
In my CI pipeline, an LLM reviews every pull request: It reads the diff, cites the project’s authority docs, and writes up findings with severities and suggested fixes. What it does not do is decide whether the PR merges. The decider that gates the merge ( REQUEST_CHANGES or APPROVE ) is computed by deterministic TypeScript that the model cannot override, no matter what verdict it puts in its own output.
This post is about why that line exists, where it sits in the code, and the production failures that taught me the gate has to be plain code.
The problem: a model that finds three bugs can still approve
LLM reviewers are persuadable and inconsistent in a specific and annoying way. The model can enumerate three real defects in its findings list and then, in the same response, talk itself into verdict: approve. The findings and the verdict come out of the same forward pass, but nothing forces them to be consistent. Sometimes the diff description is soothing. Sometimes the PR body says “minor cleanup” and the model believes it. Sometimes it just drifts.
If the verdict field of the model’s JSON is what the merge gate reads, you have a gate that can be argued with. With today’s “AI” (read: LLMs/Agents) that simply does not work as the gate we typically expect.
The fix to this is small, boring, and literally dumb: the verdict is derived from the findings by code, not asserted by the model. In my merge-decider worker:
export function enforceFindingsVerdict(
llmVerdict: MergeDecision['verdict'],
findings: readonly { severity?: string }[],
): MergeDecision['verdict'] {
if (countBlocking(findings) > 0) {
return 'request_changes';
}
// No blocking findings: a suggestion-only set resolves to approve,
// because the merge gate clears ONLY on APPROVED — a lingering
// COMMENT would leave the PR un-mergeable over nits.
if (llmVerdict === 'request_changes') {
return 'approve';
}
if (llmVerdict === 'comment' && findings.length > 0) {
return 'approve';
}
return llmVerdict;
}
The model’s self-assigned verdict is an input, not the answer. One or more blocking findings forces request_changes regardless of what the model said. No blocking findings forces the verdict into a state the merge gate can actually act on. The prompt also instructs the model to follow this rule. But while the prompt is a request, this function is the guarantee.
The floor is enforced twice, deliberately. The merge-decider applies it when it produces the decision event, and the poster (the component that actually calls the GitHub review API) applies it again at the sink:
export function resolveReviewEvent(
verdict: Review['verdict'],
findings: readonly { severity?: string }[],
): GhReviewEvent {
if (findings.some((f) => isBlockingSeverity(f.severity))) return 'REQUEST_CHANGES';
return mapVerdict(verdict);
}
SuperNerd Level: even if a post command somehow reaches the poster without going through the decider, the review state that lands on GitHub is still consistent with the finding severities. The invariant lives at the decision point, and at the point of external effect. Nowhere does an LLM get a vote.
The shape around the gate
The pipeline is a NATS JetStream bus with deterministic adapters at the edges and stateless LLM workers in the middle. The design doc’s one-line rule: the LLM never holds a socket to an upstream. GitHub is spoken to by a webhook receiver on the way in and a poster on the way out; the model only ever sees a task envelope.
A PR review flows like this:
GitHub webhook → gh-webhook-receiver → NATS JetStream
→ review dispatcher (deterministic router, ConfigMap rules)
→ worker-da-reviewer-code (Sonnet) ┐ single-turn LLM calls,
→ worker-da-reviewer-docs (Haiku) ┘ authority docs injected
→ aggregator (joins both legs by correlation id, deterministic)
→ merge-decider (LLM drafts the summary; code computes the verdict)
→ github-poster-review (deterministic; posts as the HK-47 GitHub App)
The reviewer workers are single-turn claude --print invocations with the project’s citable authority documents injected into the prompt. Every finding has to carry a cited_principle pointing at a real doc entry, which is validated by schema, not by trust. Reviews land in about 3 minutes. The fleet as a whole is ~39 components, and the four bot personas (reviewer, CI doctor, merge gatekeeper, operator secretary) are four separate GitHub Apps whose private keys live in one mint sidecar; workers never hold GitHub credentials themselves, they request short-lived installation tokens over the bus.
Envelopes on the bus are slim — under a kilobyte, carrying a payload_ref pointer into Postgres rather than the payload itself. Every external mutation is recorded in a side-effects table keyed on task id, so a crash between “posted the review” and “acked the message” replays as a no-op instead of a duplicate review.
If you’re familiar with the project’s design principles you’ll recognize the sentence this all implements: LLM proposes, engine disposes. The model authors content; deterministic code decides what that content is allowed to do.
Two failures that earned their invariants
The truncation bug. An earlier iteration of this system put full upstream payloads. Including webhook bodies and diffs directly in the bus messages. A large payload got silently truncated in flight, downstream consumers parsed the mangled tail, and the failure surfaced far from its cause. The fix became a hard rule in the design doc: envelopes carry a Postgres reference, never the payload. It’s listed under “what to defend under future pushback,” because every few months an agent or someone proposes inlining the payload again “for convenience,” and the answer is the same bug returning with a new face.
The subject-overlap bug. The token-mint sidecar answers requests over core NATS request/reply. Plain request in, token out. Its request subject originally matched a wildcard that a JetStream stream was also configured to capture. So the stream intercepted every mint request, acknowledged it with a publish-ack, and the reply never reached the caller. Worse: this was latent for the sidecar’s entire life, because nothing exercised the path until a downstream worker hit a real CI failure and needed a token. The fix is a boot-time check — the sidecar queries the JetStream manager for any stream whose subject filter overlaps its request subject, and refuses to start if one matches. The lesson generalizes: when two messaging semantics share a namespace, an overlap is silent until the day it isn’t, so assert the invariant at boot instead of discovering it in an incident.
There’s a third one worth sharing, and it’s less a bug than a decision I can show you the evidence for.
The verdict floor started as findings.length > 0 ⇒ CHANGES_REQUESTED: any finding at all blocks the merge. The obvious worry is convergence. A one-line style nit blocks a PR, the coding agent pushes a fix, the re-review finds a different nit, and a trivial change grinds through round after round without ever landing. That fear is written into the code as a comment rather than recorded as a measurement, which is worth saying plainly (I’m quoting a worry here, not a PR I actually watched do it).
So the gate became tunable. BLOCKING_SEVERITY_THRESHOLD takes low, medium, high, or critical. Findings at or above it force request_changes, and a suggestion-only set approves with the suggestions carried in the review body. The severity-aware path is real code with unit tests behind it, not a hope.
But it’s also switched off. The default is low (the file itself calls that “everything blocks”), and not one deployed worker sets the variable. So the predicate actually running in production today is still the strict one I started with. The tunable is an escape hatch I built and haven’t needed to open yet.
The advantage to anyone building something similar: because the gate is a pure function with unit tests, tightening or loosening it is a reviewable diff and one config value. It’s not a prompt tweak I’d have to go re-verify by vibes.
And one from last week, since honesty is the genre: the merge-decider is the sole final-verdict step in the pipeline, which makes it the highest-blast-radius place for a swallowed error — if it dies silently, the PR gets no verdict at all and the merge gate fails closed. A real gateway 429 (“no deployments available, try again in 5332 seconds”) hit exactly that spot and was dead-lettered on the first attempt, because the catch block predated the shared rate-limit classification and treated a retryable provider error as terminal. The fix was shared error taxonomy across all the LLM workers: rate-limit and gateway-overload errors NAK and redeliver with backoff; only genuinely deterministic failures dead-letter.
What’s actually novel here (not much, and that’s fine)
The overall shape converges with positions other people have published. HumanLayer’s 12-Factor Agents argues for small, stateless, single-purpose LLM steps embedded in deterministic control flow that you own — that is this system, factor for factor: the workers are stateless per task, the control flow is routers and reducers in TypeScript, the LLM call is a pure function from envelope to structured output. CNCF’s July 2026 post Is a Pod the right deployment unit for an AI agent? lands near the same place from the infrastructure side; my answer happens to be “one Deployment per skill, generic worker image, prompt from a ConfigMap,” which is just the Kubernetes-native reading of the same instinct.
The part I haven’t found published anywhere is narrower: a code-review pipeline where the review verdict is structurally not the model’s output but a hard, non-LLM floor between the findings and the merge gate, enforced at both the decision point and the API sink. Plenty of writing says “keep a human in the loop” or “validate LLM output.” I haven’t seen anyone describe making the verdict itself a deterministic function of the findings, with the model’s own verdict demoted to a tiebreaker for the no-blocking-findings case. If someone has, I’d genuinely like to read it.
The whole thing runs live. The fleet observatory is at whatis.droidkluster.com — you can watch HK-47 post a CHANGES_REQUESTED verdict and the coding agent push rework in real time, gate and all.