Skip to content

Reference architecture

Automated code review and security scanner AI agent

Automated pull request review that catches injection risks, leaked secrets and architectural anti-patterns before a human reviewer starts.

Design constraints

What the design has to hold to

Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.

  • 01Inline GitHub review comments posted within 60 seconds of a pull request opening
  • 02Low false-positive rate: no comments on the team's established conventions
  • 03Deterministic SAST rules combined with LLM architectural critique
  • 04No retention of private source code by third-party model providers

Component topology

System components and technologies

Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.

Stack topology

Automated code review and security scanner AI agent

Illustrative reference architecture

  1. 01

    GitHub webhook receiver

    Receives pull_request events, changed files and diffs

    FastAPI + Octokit

  2. 02

    Static analysis (SAST) scanner

    Deterministic vulnerability and secret scanning

    Semgrep + TruffleHog

  3. 03

    Semantic code reviewer

    Reviews architectural patterns, race conditions and test coverage

    Claude API / self-hosted code model

  4. 04

    PR comment bot

    Posts actionable review comments with suggested changes

    GitHub App

Subsystem 01

GitHub webhook receiver

Receives pull_request events, changed files and diffs

Typical stack

FastAPI + Octokit

Subsystem 02

Static analysis (SAST) scanner

Deterministic vulnerability and secret scanning

Typical stack

Semgrep + TruffleHog

Subsystem 03

Semantic code reviewer

Reviews architectural patterns, race conditions and test coverage

Typical stack

Claude API / self-hosted code model

Subsystem 04

PR comment bot

Posts actionable review comments with suggested changes

Typical stack

GitHub App

Data lifecycle

End-to-end data flow

  1. A developer opens a pull request; the webhook sends the payload to the FastAPI review dispatcher.

  2. The scanner fetches the diff and runs Semgrep for known vulnerabilities and TruffleHog for leaked secrets.

  3. Changed files and related context are packaged into an AST-aware review prompt.

  4. The LLM reviews the change against the team's architecture decision records (ADRs) and style guides.

  5. The bot posts inline review comments with one-click 'Apply suggestion' blocks.

Reliability and resilience

Failure modes and how each is contained

Failure mode 01

Pedantic comments that developers learn to ignore

Mitigation

Strict confidence thresholds (above 90%), with comments limited to security, correctness and performance.

Failure mode 02

Oversized pull requests exceeding the context window

Mitigation

Skip generated files (lockfiles, minified bundles) and review large PRs in chunked passes.

Failure mode 03

Source code leaking to model providers

Mitigation

Zero-data-retention enterprise API endpoints, or self-hosted models inside a private VPC.

Questions

What teams ask about this design

Yes. It formats fixes with GitHub's suggestion syntax, so a developer can commit one with a single click.

We index your architecture decision records (ADRs) and coding standards and supply the relevant ones as context in each review prompt.

Planning a system like this?

Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.