The seven questions below mirror the FAQ pinned on Show HN. Same answers, same wording, no varnish. If something is missing, open an issue on GitHub.
Frequently Asked Questions
Seven questions the CodeOtter team gets most often — on Show HN, on Reddit, in private email, and on community calls. Same answers, same wording, no varnish.
01What license does CodeOtter use?
CodeOtter is licensed under Elastic License 2.0 (ELv2). It is free to self-host, modify, and run forever on your own hardware. It is not OSI-OSS — the licence restricts hosting a competing SaaS that re-sells the unmodified binary. We say so up-front; we never call it OSI-OSS, and we never describe the licence that way. Full licence text: `codeotter.io/legal/terms`.
02What models can I run?
System 1 (calibrated scores + 9 merge gates, deterministic): Laya Typed-Decisions (455 MB Q4 GGUF) and Kev 0.8B (828 MB), shipped with CodeOtter, served by the bundled ggmlc laya runtime. The /v1/systemone HTTP endpoint also accepts any OpenAI-compatible System 1 model (TypeSafe Jev, OpenRouter-hosted, etc.). System 2 (narrative review): bring-your-own — Claude, OpenAI, Anthropic, OpenRouter, or any OpenAI-compatible endpoint. For 100% offline System 2, Admin → Local models downloads Qwen2.5-Coder 1.5B or 7B GGUF (served by llama.cpp's llama-server) or Microsoft CodeReviewer (Python sidecar). All sidecars spin up on demand and shut down after 15 idle minutes.
03What hardware do I need?
Minimum: a single x86_64 or Apple-silicon box with 8 GB RAM. Recommended for fully local: 16 GB RAM + Vulkan-capable discrete GPU OR Apple-silicon with unified memory. The 1.5B Qwen2.5-Coder GGUF fits in 8 GB; the 7B Qwen2.5-Coder needs 16 GB. Microsoft CodeReviewer is the smallest (~250 MB) and runs on CPU at acceptable speed. GPU detection is automatic (Vulkan / Metal); the runtime picks the best path without intervention.
04Why is the first review slow?
The first review on a fresh box pays a one-time model-load cost (~5 min on a mid-range laptop for the 7B GGUF, ~30 s for 1.5B). Subsequent reviews on the same box are fast (~10–30 s for typical PRs). Sidecars stay resident for 15 idle minutes after each review, so back-to-back reviews are warm. This is the trade-off for 100% offline — no warm model in someone else's data centre.
05I can ask ChatGPT to review a PR. Why do I need CodeOtter?
CodeOtter is the platform that hosts the review — the GitHub PR integration, the 9 binary merge gates, the AGENTS.md auto-detection + per-repo injection, the 1-click "Dismiss & remember rule" learnings, the inline committable suggestion blocks, and the incremental commit-by-commit delta reviews. ChatGPT (or any LLM) is one provider of the System 2 narrative engine. You can plug ChatGPT into CodeOtter via BYOK. But you cannot get the gates, AGENTS.md enforcement, inline suggestions, or delta reviews from ChatGPT alone — those live in the platform, not in the model.
06How is the System 1 score different from asking an LLM for a 1–100 score?
System 1 is a tiny classifier that returns strict JSON over a fixed rubric. The questions are closed-form: choice, score, or noul. Because the schema is fixed and the model is small, scores don't hallucinate and the 9 binary gates don't drift. Chat LLMs cluster every score between 75 and 85 because they are not trained to output calibrated numbers — they are trained to produce plausible prose. System 1 outputs a calibrated 0–100 because the model literally cannot output anything else.
07Does CodeOtter phone home?
No. Zero outbound calls from the review pipeline. PostHog JS on codeotter.io (the marketing site) is a separate concern and is opt-out via Do-Not-Track. The self-hosted product has zero telemetry by design — verifiable in your egress firewall logs during a full PR review run.
Something missing?
Open an issue on GitHub