Products

One control plane for every AI system.

Eight product areas — from PR scans and runtime protection through evaluations, governance, mesh, and cost — on the same tenant policies.

Scan Free – Enterprise

Static scans that comment on the pull request

Install the GitHub App and every pull request is scanned for authorization gaps, tenant isolation, injection, JWT mistakes, and AI-specific patterns. Findings land as inline review comments and a Check Run, and one click turns them into an executive summary.

GitHub-native

PR comments and Check Runs — no tokens to paste, no repo URL to copy.

Actionable findings

Every finding carries file:line and a remediation sentence your editor or AI agent can apply.

Executive reports

Generate executive, technical, and compliance reports from the same findings.

Read the scans docs →

Run Pro+

A fail-closed gateway, firewall, and DLP on live traffic

Point your app at the OpenAI-compatible gateway, or wrap calls with the SDK. Prompts, RAG chunks, tool results, and outputs are checked against your policies and allowed, warned, blocked, or redacted. If a check errors, the request is blocked — never silently passed through.

Prompt firewall

Injection and jailbreak screening on input and output, on the same policies as PR scans.

DLP

Redaction for PII and secrets before model output reaches a user or a tool.

Agent controls

Human-in-the-loop approval for high-agency tools like email, shell, and file share.

curl -s https://app.belfrylabs.ai/api/v1/gateway/v1/chat/completions \
  -H "Authorization: Bearer bak_…" -H "X-Tenant-ID: $TENANT" \
  -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"…"}]}'

Read the runtime docs →

Prove Pro+ · Enterprise for live Assessor

Red team campaigns and a human-confirmed live pen-test

Seed adversarial campaigns from the product and replay them as policies change. When you need proof for a buyer or an audit, the Enterprise Assessor runs live probes against your deployment — only after a human confirms the rules of engagement.

Attack corpus

Campaigns mapped to MITRE ATLAS techniques, versioned with your policies.

Purple-team loop

Runtime events seed the next campaign, so production lessons become tests.

Live Assessor

Enterprise pen-test with CVSS-normalized findings behind a signed RoE.

Read the red team docs →

Evaluate Pro+

Evaluations and model comparison on your own datasets

Run quality, safety, latency, and cost evaluations across OpenAI, Anthropic, Google, and more. Golden datasets stay inside your tenant, and every run is reproducible from a saved config.

Side-by-side

Compare models on the same prompts before you swap a provider.

Golden datasets

Regression checks that stay in your tenant, never a public leaderboard.

One engine

Scans, benchmarks, and red-team runs share a single evaluation path.

Read the evaluations docs →

Inventory Enterprise

Model inventory with expiry and degradation tracking

Every model, dataset, tool, and knowledge base in one registry. Belfry tracks lifecycle state, flags expiring artifacts, and watches golden-set scores for degradation so a silent provider-side model change shows up before your users notice.

Asset registry

Models, datasets, tools, and MCP servers with owners and lifecycle state.

Expiry alerts

Days-remaining reminders before a model or artifact ages out of policy.

Degradation tracking

Golden-set drift detection that catches silent model swaps.

Read the inventory docs →

Govern Pro+ · Enterprise

Compliance scorecards, SBOM/AIBOM, and approvals

Map current findings onto EU AI Act, NIST AI RMF, and ISO/IEC 42001 as a point-in-time scorecard of your system — not a certification of Belfry Labs. Export evidence packs, generate an SBOM/AIBOM with SLSA provenance, and route changes through approval workflows.

Scorecards

Point-in-time posture with evidence export for buyers and auditors.

SBOM / AIBOM

CycloneDX, SPDX, and SARIF with SLSA provenance for supply chain review.

Approvals

Human sign-off and an audit trail on every lifecycle state change.

Read the governance docs →

Connect Enterprise · Pro add-on

An MCP mesh you register, health-check, and govern

Register tenant MCP servers and A2A peers, run health checks, and make guarded test calls through the same runtime policies as chat traffic. Belfry itself is also an MCP server, so Cursor and Claude can list projects, run scans, and build executive reports.

Registry

Lifecycle state, trust score, and health for every MCP server and A2A peer.

Impersonation checks

Lookalike-server and descriptor-poisoning detection on every descriptor.

Belfry-as-MCP

Scans, findings, and executive reports exposed as MCP tools — no pentest tools.

Read the mesh docs →

Optimize Pro+

Cost tracking and cost-aware routing

See estimated spend per project, per model, and per route. When a cheaper eligible model meets your quality bar, the gateway can route to it — and optional monthly budget caps stop runaway spend at inference time.

Spend trends

Per-project and per-model cost with forecasts.

Cost-aware routing

Cheaper eligible models when quality allows — you set the floor.

Budget caps

Optional monthly caps enforced at inference when set.

Read the cost docs →

Start with the free taste

Ten GitHub PR scans and one architecture review, no card. Paid plans are per project — talk to us to upgrade.