Similar prompts are matched against cached responses using cosine similarity. When a match exceeds the configured threshold, the cached response is served instantly — no model call, no token cost. Track hit rates, cost savings, and manage cache entries per tenant from this dashboard.