Agentic Cache
Overview
Section titled “Overview”Agentic Cache is the workspace reporting surface for cache reuse and savings. It attributes hit/miss outcomes and the tokens, cost, and latency they saved to a workspace, a Virtual Key, and a cache kind, so one console answers “how much is caching actually saving us, and for whom”.
Caching behaviour itself is configured elsewhere: Semantic caching decides what is reused on the LLM path, and Provider prompt caching handles the static prompt prefix at the upstream provider. Turning something on in Agentic Cache reports on reuse; it does not create it.
Key benefits:
- Reconciled savings - reuse is attributed to exact-response and semantic
kinds with saved-token and saved-cost counters, so you read one
$/tokensaved figure rather than two competing numbers. - Scoped reporting - every recorded outcome is tenant- and workspace-scoped and can be filtered by Virtual Key.
- Explicit reuse boundaries - the console is clear about where reuse is not permitted, notably governed MCP tool results.
How it differs from the other caches
Section titled “How it differs from the other caches”DeepIntShield exposes three related cache surfaces. They are not one shared cache.
| Capability | What it does | Scope |
|---|---|---|
| Agentic Cache (this page) | Reports hit/miss and savings, reconciled across kinds | Workspace and Virtual Key reporting |
| Semantic caching | Reuses LLM responses by exact and vector similarity | Per model/provider, with Virtual Key scoping |
| Provider prompt caching | Caches the static prefix of a prompt at the provider | Per provider |
When to use it
Section titled “When to use it”- You want one workspace view of cache hit rate and savings instead of reading them per plugin.
- You need savings broken down by Virtual Key to attribute spend back to a team or an application.
- You need to show an auditor which cache kinds are in use and which reuse is refused on the governed path.
Configuration
Section titled “Configuration”Per-workspace settings live on Agentic Cache → Settings at
/workspace/agentic-cache/settings. You can also reach them from
Agentic → Activity → Secure caches → Settings. Every cache page includes
the Settings navigation link. Master enablement, safety options, and TTLs are
on Settings; per-kind enable switches are on Agentic Caches at
/workspace/agentic-cache/agentic-caches.
-
Open Cost Optimization → Agentic Cache → Settings.
-
Under Master, turn on Agentic cache enabled to record cache outcomes for this workspace.
-
Under Semantic & safety and TTLs (seconds), review the per-kind switches. Use them to record how this workspace intends its caches to behave.
-
Click Save.
-
Go to Agentic Cache → Agentic Caches to see per-kind hit rates.
The Security Caches page lists the read-mostly caches that reduce authorization-path latency (decision/verdict, policy, key config). Cache lookup and enforcement still have measurable overhead, and misses require backing-store or policy work. These caches are invalidated structurally and on revocation push - see Virtual keys for the decision cache.
Settings reference
Section titled “Settings reference”| Field | Type | Default | Description |
|---|---|---|---|
enabled | boolean | true | Records cache outcomes for this workspace. |
response_enabled | boolean | true | Reports exact-response reuse. |
semantic_enabled | boolean | true | Reports semantic reuse. |
tool_result_enabled | boolean | true | Recorded intent. Governed MCP tool execution is not result-cached. |
embedding_enabled | boolean | true | Recorded intent. No embedding/RAG reuse in this release. |
mcp_discovery_enabled | boolean | true | Recorded intent. No tools/list reuse in this release. |
semantic_threshold | number (0-1) | 0.92 | Recorded intent. Set the effective threshold on the semantic cache. |
semantic_read_only | boolean | true | Recorded intent for semantic reuse. |
never_cache_high_risk | boolean | true | Recorded intent. Governed tool results are uncached regardless. |
encrypt_at_rest | boolean | true | Recorded intent for cached payload storage. |
honor_obligations | boolean | true | Recorded intent for serving a cached response under an obligation. |
response_ttl_seconds | integer | 3600 | Recorded intent. Set the effective TTL on the semantic cache. |
semantic_ttl_seconds | integer | 1800 | Recorded intent. Set the effective TTL on the semantic cache. |
tool_result_ttl_seconds | integer | 600 | Recorded intent for the legacy per-tool MCP cache. |
Cache kinds
Section titled “Cache kinds”| Kind | What it reports |
|---|---|
| Exact response | Reuse of an identical earlier request. |
| Semantic | Reuse of a similar earlier request by vector match. |
| Tool result | MCP tool-result reuse. Zero on the governed path, by design. |
| Embedding | Embedding/RAG reuse. Not available in this release. |
| MCP discovery | tools/list reuse. Not available in this release. |
Per-virtual-key cache scope
Section titled “Per-virtual-key cache scope”This section configures the semantic cache. Semantic reuse is bounded by the scope that distinguishes one caller from another inside a shared virtual key. By default the gateway scopes automatically; you can pin a scope mode on each virtual key for tighter control.
The available scope modes are:
| Mode | A cached entry is shared across… |
|---|---|
virtual_key | All callers using the same virtual key. |
user | The same end user (resolved from request user / governance identity). |
use_case | The same use_case metadata value. |
session | The same session. |
custom_metadata | The same value of the metadata keys you nominate. |
none | No reuse - caching is effectively off for the key. |
On a virtual key (Access & Credentials → Virtual API Keys → edit), the cache controls are:
- Automatic Cache - enable or disable automatic cache scoping for the key.
- Semantic Cache - enable or disable semantic matching for the key.
- Scope mode - pick the scope (
virtual_key,user,use_case,session,custom_metadata, ornone) and, for a custom scope, the metadata keys that define it (for exampleuse_case, orsession_idfor a session-scoped key). - Allow semantic reuse on unscoped requests - leave off if several end users share one key and you don’t want one user’s response served to another. When off, semantic lookups on a key with no per-caller scope are suppressed.
- Cache Key - an optional fixed cache key for the key. Leave it empty to use
automatic scoping; a request-level
x-deepintshield-cache-keyheader still overrides it.
Legacy per-tool MCP cache controls
Section titled “Legacy per-tool MCP cache controls”Outside the governed Agentic path, the legacy direct MCP cache can be explicitly
enabled at workspace/client level and its TTL overridden per tool, including
0s to disable a tool. Each entry maps a "<server>-<tool>" name to a duration
string:
{ "search-web_search": "5m", "db-run_query": "0s"}The direct cache defaults off and only explicitly cacheable tools are eligible.
It does not infer which tools are writes or high risk: exclude those tools from
cacheable_tools and avoid the * wildcard. These controls do not make a
governed call cacheable; see MCP tool execution for the
enforced boundary.
Monitoring savings
Section titled “Monitoring savings”The Agentic Cache → Overview page combines several counters:
- Decision-cache hit rate from the security decision cache.
- Response and Semantic hit rates.
- Saved (24h) - dollars, tokens, and latency saved.
Savings are credited after the response, with the saved tokens and cost attached when the provider reports them. Open Agent Insights → Caching for the response and semantic time series. Kinds listed above as unavailable in this release report zero and are not savings sources.
Examples
Section titled “Examples”Conservative workspace profile - record cache intent and keep MCP tool results explicitly off. Configure the semantic cache separately:
{ "enabled": true, "response_enabled": true, "semantic_enabled": true, "tool_result_enabled": false, "semantic_read_only": true, "never_cache_high_risk": true, "honor_obligations": true, "semantic_threshold": 0.92, "response_ttl_seconds": 3600, "semantic_ttl_seconds": 1800, "tool_result_ttl_seconds": 600}High-throughput chatbot with shared key, scoped per user - semantic reuse within a single end user only:
{ "cache_enabled": true, "semantic_cache_enabled": true, "cache_scope_mode": "user", "cache_allow_semantic_when_unscoped": false}Next steps
Section titled “Next steps”- Semantic caching - the response-similarity cache behind the response and semantic savings reported here.
- Provider prompt caching - cache the static prompt prefix at the upstream provider.
- Virtual keys - the boundary that scopes every cached entry, plus the per-key cache controls.
- MCP tool execution - tool authorization and its result-cache bypass.