AiT AI GatewayIntelligent IT · MSP control plane

Multi-tenant AI gateway for IG MSP services.

One endpoint. Per-tenant virtual keys. Hard budget caps. Edge prompt caching. Audit log. PII redaction. Model failover. Built on LiteLLM + Cloudflare AI Gateway. Run by Manny, the Intelligent IT assistant.

Tenants ->Policies ->Audit log ->Virtual keys ->Public demo dashboard ->Token budgets ->Prompt templates ->
Tenants
8
active LiteLLM teams
Spend MTD
$8,668 / $13,150 cap
65.9% of combined cap
Requests · last 24h
28.1K
across all tenants
Policy violations · 30d
59
10 active policies · 22 active keys

What ships in the box

Per-tenant isolation

One LiteLLM team per tenant, one virtual key, one budget cap, one model allowlist. Tenants never see each other's traffic, keys, or spend. The master key only ever lives in GSM.

Server-side enforcement

Caps and content blocks are enforced at the gateway, before any provider call. PII / PHI / PCI never reaches a model. "Ignore previous instructions" never moves a budget.

Audit by default

Every request gets an immutable audit row: tenant, user, model, tokens, cost, latency, decision, policy hit. Useful for SOC2, HIPAA, and the "why did our bill spike" conversation.

Model failover & caching

Cloudflare AI Gateway sits in front for edge prompt caching (~30% hit rate on repeat workloads) and provider failover. When Anthropic stutters, traffic shifts to OpenAI inside the same policy.

Loading budgets…

Gateway Health Monitoring

Last 24h · 6 endpoints

SLA Summary

Overall Uptime
99.34%
Worst Endpoint
Groq LLaMA
highest error rate
Best Endpoint
Cerebras
highest uptime
Errors · Last 24h
0
across all endpoints
Avg P95 Latency
1.5s
all endpoints

Endpoint Status

DS Chathealthy
P50
420ms
P95
980ms
P99
1.6s
Error Rate
0.4%
Uptime
99.82%
DS Reasonerhealthy
P50
1.9s
P95
4.2s
P99
7.1s
Error Rate
0.7%
Uptime
99.61%
Cerebrashealthy
P50
190ms
P95
510ms
P99
890ms
Error Rate
0.2%
Uptime
99.91%
Groq LLaMAdegraded
P50
280ms
P95
820ms
P99
2.1s
Error Rate
3.8%
Uptime
97.43%
Gemini Prohealthy
P50
680ms
P95
1.6s
P99
3.0s
Error Rate
0.9%
Uptime
99.55%
Whisper v3healthy
P50
340ms
P95
710ms
P99
1.1s
Error Rate
0.6%
Uptime
99.72%

Latency Heatmap · 30-min slots

DS ChatDS ReasonerCerebrasGroq LLaMAGemini ProWhisper v3<300ms<600ms<1s<2s>4s

P95 Latency Trend · Last 24h

0ms125ms250ms375ms500ms

7-day Uptime by Endpoint

Endpoint
DS Chat
DS Reasoner
Cerebras
Groq LLaMA
Gemini Pro
Whisper v3