AiT AI GatewayIntelligent IT · MSP control plane

Semantic Cache — deduplicate prompts, cut token spend.

Similar prompts are matched against cached responses using cosine similarity. When a match exceeds the configured threshold, the cached response is served instantly — no model call, no token cost. Track hit rates, cost savings, and manage cache entries per tenant from this dashboard.

50 seed entries5 tenant groups0.92 default threshold$0.001 / 1K tokensTrigram similarity
Loading semantic cache data...