GitHub-native
PR comments and Check Runs — no tokens to paste, no repo URL to copy.
Products
Eight product areas — from PR scans and runtime protection through evaluations, governance, mesh, and cost — on the same tenant policies.
Scan Free – Enterprise
Install the GitHub App and every pull request is scanned for authorization gaps, tenant isolation, injection, JWT mistakes, and AI-specific patterns. Findings land as inline review comments and a Check Run, and one click turns them into an executive summary.
PR comments and Check Runs — no tokens to paste, no repo URL to copy.
Every finding carries file:line and a remediation sentence your editor or AI agent can apply.
Generate executive, technical, and compliance reports from the same findings.
Run Pro+
Point your app at the OpenAI-compatible gateway, or wrap calls with the SDK. Prompts, RAG chunks, tool results, and outputs are checked against your policies and allowed, warned, blocked, or redacted. If a check errors, the request is blocked — never silently passed through.
Injection and jailbreak screening on input and output, on the same policies as PR scans.
Redaction for PII and secrets before model output reaches a user or a tool.
Human-in-the-loop approval for high-agency tools like email, shell, and file share.
curl -s https://app.belfrylabs.ai/api/v1/gateway/v1/chat/completions \
-H "Authorization: Bearer bak_…" -H "X-Tenant-ID: $TENANT" \
-d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"…"}]}'
Prove Pro+ · Enterprise for live Assessor
Seed adversarial campaigns from the product and replay them as policies change. When you need proof for a buyer or an audit, the Enterprise Assessor runs live probes against your deployment — only after a human confirms the rules of engagement.
Campaigns mapped to MITRE ATLAS techniques, versioned with your policies.
Runtime events seed the next campaign, so production lessons become tests.
Enterprise pen-test with CVSS-normalized findings behind a signed RoE.
Evaluate Pro+
Run quality, safety, latency, and cost evaluations across OpenAI, Anthropic, Google, and more. Golden datasets stay inside your tenant, and every run is reproducible from a saved config.
Compare models on the same prompts before you swap a provider.
Regression checks that stay in your tenant, never a public leaderboard.
Scans, benchmarks, and red-team runs share a single evaluation path.
Inventory Enterprise
Every model, dataset, tool, and knowledge base in one registry. Belfry tracks lifecycle state, flags expiring artifacts, and watches golden-set scores for degradation so a silent provider-side model change shows up before your users notice.
Models, datasets, tools, and MCP servers with owners and lifecycle state.
Days-remaining reminders before a model or artifact ages out of policy.
Golden-set drift detection that catches silent model swaps.
Govern Pro+ · Enterprise
Map current findings onto EU AI Act, NIST AI RMF, and ISO/IEC 42001 as a point-in-time scorecard of your system — not a certification of Belfry Labs. Export evidence packs, generate an SBOM/AIBOM with SLSA provenance, and route changes through approval workflows.
Point-in-time posture with evidence export for buyers and auditors.
CycloneDX, SPDX, and SARIF with SLSA provenance for supply chain review.
Human sign-off and an audit trail on every lifecycle state change.
Connect Enterprise · Pro add-on
Register tenant MCP servers and A2A peers, run health checks, and make guarded test calls through the same runtime policies as chat traffic. Belfry itself is also an MCP server, so Cursor and Claude can list projects, run scans, and build executive reports.
Lifecycle state, trust score, and health for every MCP server and A2A peer.
Lookalike-server and descriptor-poisoning detection on every descriptor.
Scans, findings, and executive reports exposed as MCP tools — no pentest tools.
Optimize Pro+
See estimated spend per project, per model, and per route. When a cheaper eligible model meets your quality bar, the gateway can route to it — and optional monthly budget caps stop runaway spend at inference time.
Per-project and per-model cost with forecasts.
Cheaper eligible models when quality allows — you set the floor.
Optional monthly caps enforced at inference when set.
Ten GitHub PR scans and one architecture review, no card. Paid plans are per project — talk to us to upgrade.