Skip to content

Agentic Cache

Agentic Cache is the workspace reporting surface for cache reuse and savings. It attributes hit/miss outcomes and the tokens, cost, and latency they saved to a workspace, a Virtual Key, and a cache kind, so one console answers “how much is caching actually saving us, and for whom”.

Caching behaviour itself is configured elsewhere: Semantic caching decides what is reused on the LLM path, and Provider prompt caching handles the static prompt prefix at the upstream provider. Turning something on in Agentic Cache reports on reuse; it does not create it.

Key benefits:

  • Reconciled savings - reuse is attributed to exact-response and semantic kinds with saved-token and saved-cost counters, so you read one $/token saved figure rather than two competing numbers.
  • Scoped reporting - every recorded outcome is tenant- and workspace-scoped and can be filtered by Virtual Key.
  • Explicit reuse boundaries - the console is clear about where reuse is not permitted, notably governed MCP tool results.

DeepIntShield exposes three related cache surfaces. They are not one shared cache.

CapabilityWhat it doesScope
Agentic Cache (this page)Reports hit/miss and savings, reconciled across kindsWorkspace and Virtual Key reporting
Semantic cachingReuses LLM responses by exact and vector similarityPer model/provider, with Virtual Key scoping
Provider prompt cachingCaches the static prefix of a prompt at the providerPer provider
  • You want one workspace view of cache hit rate and savings instead of reading them per plugin.
  • You need savings broken down by Virtual Key to attribute spend back to a team or an application.
  • You need to show an auditor which cache kinds are in use and which reuse is refused on the governed path.

Per-workspace settings live on Agentic Cache → Settings at /workspace/agentic-cache/settings. You can also reach them from Agentic → Activity → Secure caches → Settings. Every cache page includes the Settings navigation link. Master enablement, safety options, and TTLs are on Settings; per-kind enable switches are on Agentic Caches at /workspace/agentic-cache/agentic-caches.

  1. Open Cost Optimization → Agentic Cache → Settings.

  2. Under Master, turn on Agentic cache enabled to record cache outcomes for this workspace.

  3. Under Semantic & safety and TTLs (seconds), review the per-kind switches. Use them to record how this workspace intends its caches to behave.

  4. Click Save.

  5. Go to Agentic Cache → Agentic Caches to see per-kind hit rates.

The Security Caches page lists the read-mostly caches that reduce authorization-path latency (decision/verdict, policy, key config). Cache lookup and enforcement still have measurable overhead, and misses require backing-store or policy work. These caches are invalidated structurally and on revocation push - see Virtual keys for the decision cache.

FieldTypeDefaultDescription
enabledbooleantrueRecords cache outcomes for this workspace.
response_enabledbooleantrueReports exact-response reuse.
semantic_enabledbooleantrueReports semantic reuse.
tool_result_enabledbooleantrueRecorded intent. Governed MCP tool execution is not result-cached.
embedding_enabledbooleantrueRecorded intent. No embedding/RAG reuse in this release.
mcp_discovery_enabledbooleantrueRecorded intent. No tools/list reuse in this release.
semantic_thresholdnumber (0-1)0.92Recorded intent. Set the effective threshold on the semantic cache.
semantic_read_onlybooleantrueRecorded intent for semantic reuse.
never_cache_high_riskbooleantrueRecorded intent. Governed tool results are uncached regardless.
encrypt_at_restbooleantrueRecorded intent for cached payload storage.
honor_obligationsbooleantrueRecorded intent for serving a cached response under an obligation.
response_ttl_secondsinteger3600Recorded intent. Set the effective TTL on the semantic cache.
semantic_ttl_secondsinteger1800Recorded intent. Set the effective TTL on the semantic cache.
tool_result_ttl_secondsinteger600Recorded intent for the legacy per-tool MCP cache.
KindWhat it reports
Exact responseReuse of an identical earlier request.
SemanticReuse of a similar earlier request by vector match.
Tool resultMCP tool-result reuse. Zero on the governed path, by design.
EmbeddingEmbedding/RAG reuse. Not available in this release.
MCP discoverytools/list reuse. Not available in this release.

This section configures the semantic cache. Semantic reuse is bounded by the scope that distinguishes one caller from another inside a shared virtual key. By default the gateway scopes automatically; you can pin a scope mode on each virtual key for tighter control.

The available scope modes are:

ModeA cached entry is shared across…
virtual_keyAll callers using the same virtual key.
userThe same end user (resolved from request user / governance identity).
use_caseThe same use_case metadata value.
sessionThe same session.
custom_metadataThe same value of the metadata keys you nominate.
noneNo reuse - caching is effectively off for the key.

On a virtual key (Access & Credentials → Virtual API Keys → edit), the cache controls are:

  • Automatic Cache - enable or disable automatic cache scoping for the key.
  • Semantic Cache - enable or disable semantic matching for the key.
  • Scope mode - pick the scope (virtual_key, user, use_case, session, custom_metadata, or none) and, for a custom scope, the metadata keys that define it (for example use_case, or session_id for a session-scoped key).
  • Allow semantic reuse on unscoped requests - leave off if several end users share one key and you don’t want one user’s response served to another. When off, semantic lookups on a key with no per-caller scope are suppressed.
  • Cache Key - an optional fixed cache key for the key. Leave it empty to use automatic scoping; a request-level x-deepintshield-cache-key header still overrides it.

Outside the governed Agentic path, the legacy direct MCP cache can be explicitly enabled at workspace/client level and its TTL overridden per tool, including 0s to disable a tool. Each entry maps a "<server>-<tool>" name to a duration string:

{
"search-web_search": "5m",
"db-run_query": "0s"
}

The direct cache defaults off and only explicitly cacheable tools are eligible. It does not infer which tools are writes or high risk: exclude those tools from cacheable_tools and avoid the * wildcard. These controls do not make a governed call cacheable; see MCP tool execution for the enforced boundary.

The Agentic Cache → Overview page combines several counters:

  • Decision-cache hit rate from the security decision cache.
  • Response and Semantic hit rates.
  • Saved (24h) - dollars, tokens, and latency saved.

Savings are credited after the response, with the saved tokens and cost attached when the provider reports them. Open Agent Insights → Caching for the response and semantic time series. Kinds listed above as unavailable in this release report zero and are not savings sources.

Conservative workspace profile - record cache intent and keep MCP tool results explicitly off. Configure the semantic cache separately:

{
"enabled": true,
"response_enabled": true,
"semantic_enabled": true,
"tool_result_enabled": false,
"semantic_read_only": true,
"never_cache_high_risk": true,
"honor_obligations": true,
"semantic_threshold": 0.92,
"response_ttl_seconds": 3600,
"semantic_ttl_seconds": 1800,
"tool_result_ttl_seconds": 600
}

High-throughput chatbot with shared key, scoped per user - semantic reuse within a single end user only:

{
"cache_enabled": true,
"semantic_cache_enabled": true,
"cache_scope_mode": "user",
"cache_allow_semantic_when_unscoped": false
}
  • Semantic caching - the response-similarity cache behind the response and semantic savings reported here.
  • Provider prompt caching - cache the static prompt prefix at the upstream provider.
  • Virtual keys - the boundary that scopes every cached entry, plus the per-key cache controls.
  • MCP tool execution - tool authorization and its result-cache bypass.