Why single-LLM PR reviewers fail as CI/CD merge gates

Most AI code review bots ask a single chat model to do two incompatible jobs in one prompt: write conversational feedback for the author AND emit numeric scores or pass/fail verdicts for branch protection.

When a single generative prompt is responsible for both prose and numerical evaluation, temperature variance, prompt length, and diff formatting cause the same pull request to score 84/100 on one run and 68/100 on the next. Engineering teams quickly disable blocking gates when scores cannot be trusted.

// Traditional single-LLM failure mode: scores coupled to prose generation
// Run #1 -> Risk: 85, Security: 78 (PASS)
// Run #2 -> Risk: 72, Security: 61 (BLOCKED - identical diff!)

How CodeOtter’s Dual-Engine pipeline splits typed rubrics from prose

In CodeOtter, when a System One model is configured (either hosted via S1_PROVIDERS or locally via the ggmlc laya runtime), it owns every score in S1_SCORES and every boolean gate in S1_GATES. The language model is not asked for numeric scores at all.

Both engines execute concurrently against the pull request diff and repository guidelines (AGENTS.md and CLAUDE.md). Because the local laya runtime processes typed questions in deterministic chunks (chunked to maxQuestions = 8 per request), rubric evaluation finishes in milliseconds without hallucinated JSON schemas.

POST /v1/systemone
{
  "state": "<normalized PR diff + AGENTS.md + CLAUDE.md context>",
  "questions": [
    { "id": "correctness_risk", "kind": "score", "rubric": "S1_SCORES.correctness_risk" },
    { "id": "security",     "kind": "score",  "rubric": "S1_SCORES.security" },
    { "id": "blast_radius", "kind": "score",  "rubric": "S1_SCORES.blast_radius" },
    { "id": "title_gate",   "kind": "choice", "options": ["yes", "no"] },
    { "id": "secrets_gate", "kind": "choice", "options": ["yes", "no"] }
  ]
}

The six 0–100 gauges and binary merge gates

Every pull request evaluated by CodeOtter produces a standardized, organization-scoped scorecard stored in PocketBase:

• Quality (0–100): Evaluates control-flow regressions, off-by-one boundaries, unhandled promise rejections, and null-safety across modified hunks. • Blast Radius (0–100): Quantifies how many downstream modules, database migrations, or shared API contracts are impacted if the PR fails in production. • Risk (0–100): Audits authentication boundaries, SQL/command injection vectors, cross-org data leaks, and dependency surface changes. • Tests (0–100): Verifies whether new branches and failure modes have corresponding unit or integration assertions. • Readability (0–100): Measures cyclomatic complexity, adherence to repository conventions, and dead-code accumulation. • PR Hygiene (0–100): Rates title, description, scope focus, and review-comment hygiene.

Enforcing untrusted model normalization with judge()

CodeOtter treats all model output as untrusted input. Before any review record is written to PocketBase or rendered in the UI, the server-side judge() pipeline clamps scores to [0, 100], validates gate enums, sanitizes file walkthrough paths, and computes the final composite verdict.

This guarantees that your dashboard, GitHub status checks, and organization analytics always reflect clean, schema-verified engineering telemetry.