Delegated MCP authentication
Resource-bound token exchange and identity-isolated sessions.
DeepIntShield exposes resource-oriented provider operations without adding a model-catalog or database lookup to ordinary inference. Provider and key selection use the existing in-memory routing snapshot; a valid operation then makes the required upstream request.
POST /v1/responses is also available for providers whose adapters translate
their native inference output into Responses. Creating a Responses result
does not imply support for provider-side storage, retrieval, or compaction;
the lifecycle table below has narrower coverage.
With stream: true, consume typed Responses events. Translated streams retain
output-item IDs and indexes, text and function-argument deltas, reasoning or
refusal items when present, and the assembled output in the terminal response.
Native OpenAI clients can accumulate these events with their Responses stream
helpers. Each terminal outcome has a different meaning:
| Event | Application handling |
|---|---|
response.completed | Read the final response output and available usage. Output can contain tool calls requiring another turn. |
response.incomplete | Inspect incomplete_details.reason; token limits or content filtering can leave only a partial answer. |
response.failed | Inspect the response error and handle the failed request. |
| Connection closes without a terminal event | Treat the stream as interrupted; partial text is not a completed answer. |
Gemini and Chat-compatible translated streams map output-token limits to
max_output_tokens and supported safety/filter outcomes to content_filter.
Unknown error finish reasons fail rather than certify completion. Gemini EOF
without a provider finish reason also becomes a failed response.
For client-managed continuation, keep the input history and append the complete
returned output items before the next user message or tool results. A function
result is a function_call_output with the original call_id. Preserve opaque
reasoning content and signatures with their associated items, and continue
with the same provider and model. Reconstructing history from output_text
alone loses the tool and reasoning state.
Use store: false when managing history in your application. Where supported,
request reasoning.encrypted_content in include to retain opaque reasoning
state for later turns. Provider-managed previous_response_id continuation
requires that provider’s storage support and the original configured-key
context; it is not a portable identifier across providers.
See OpenAI integration for native client setup and MCP tool execution for the gateway’s Responses-format tool result.
| Operation | DeepIntShield route | OpenAI | Bedrock Mantle |
|---|---|---|---|
| Create | POST /v1/responses | ✅ | ✅ |
| Retrieve | GET /v1/responses/{response_id} | ✅ | ✅ |
| Retrieve as SSE | GET /v1/responses/{response_id}?stream=true | ✅ | ✅ |
| Delete | DELETE /v1/responses/{response_id} | ✅ | ✅ |
| Cancel | POST /v1/responses/{response_id}/cancel | ✅ | ✅ |
| List input items | GET /v1/responses/{response_id}/input_items | ✅ | ❌ |
| Compact | POST /v1/responses/compact | ✅ | ✅ |
Resource routes without a model require a provider selector. Use the
provider query parameter or x-model-provider header:
curl "$DEEPINTSHIELD_URL/v1/responses/resp_123?provider=openai" \ -H "Authorization: Bearer $DEEPINTSHIELD_VIRTUAL_KEY"Compaction is model-bearing and accepts a provider-qualified model:
curl "$DEEPINTSHIELD_URL/v1/responses/compact" \ -H "Authorization: Bearer $DEEPINTSHIELD_VIRTUAL_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "openai/gpt-5.6", "input": [{"role": "user", "content": "Condense our decisions."}] }'An upstream resource belongs to the provider account that created it. When
more than one configured key can serve the model, send
X-DeepIntShield-API-Key-ID or X-DeepIntShield-API-Key on creation and every
later resource call. DeepIntShield rejects an ambiguous model-less request
instead of trying another account.
OpenAI’s upstream lifecycle semantics are documented in the official Responses API reference.
Gemini and Vertex implement create, list, retrieve metadata, update expiry, and delete using either route spelling:
/v1/cached_contents/v1beta/cachedContentsCreate requests carry a provider-qualified model. List, get, patch, and delete
use provider=gemini, provider=vertex, or the equivalent x-model-provider
header.
curl "$DEEPINTSHIELD_URL/v1/cached_contents" \ -H "Authorization: Bearer $DEEPINTSHIELD_VIRTUAL_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gemini/models/gemini-2.5-flash", "displayName": "policy-context", "ttl": "3600s", "contents": [{"role": "user", "parts": [{"text": "Reference material"}]}] }'For updates, provide exactly one of ttl or expireTime. If supplied,
updateMask may name only the matching field. Google returns cache metadata on
get/list; it does not return the cached body. See the official Gemini context
caching guide.
Use the OpenAI provider for the complete Realtime transport and control surface. The WebSocket route is intended for server-to-server relaying:
wss://gateway.example/v1/realtime?model=openai/gpt-realtimeDeepIntShield authenticates, selects a provider key, and dials upstream once during the handshake. It then relays frames as opaque bytes in both directions, without per-frame JSON conversion, model discovery, or database I/O.
For browser or client media, use the WebRTC/SDP control routes:
| Route | Purpose |
|---|---|
POST /v1/realtime/client_secrets | Create an ephemeral client secret |
POST /v1/realtime/sessions | Negotiate a Realtime session |
POST /v1/realtime/transcription_sessions | Negotiate transcription |
POST /v1/realtime/translations/client_secrets | Create a translation client secret |
POST /v1/realtime/calls | Create a call |
POST /v1/realtime/calls/{call_id}/{action} | accept, hangup, refer, or reject |
JSON, multipart form data, and raw application/sdp are accepted where the
upstream operation supports them. The raw SDP body is preserved byte-for-byte;
multipart framing may be re-encoded. The negotiation body is limited to 2 MiB,
each WebSocket message to 16 MiB, and concurrent sessions to the configured
WebSocket admission limit.
OpenAI recommends WebRTC for browser/client connections and WebSocket for server-to-server use. Review the official Realtime guide before exposing a session to an untrusted client.
Mistral OCR is available at POST /v1/ocr. Supply a provider-qualified OCR
model and exactly one document variant (document_url, image_url, or file
content in a supported provider-native shape).
curl "$DEEPINTSHIELD_URL/v1/ocr" \ -H "Authorization: Bearer $DEEPINTSHIELD_VIRTUAL_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "mistral/mistral-ocr-latest", "document": { "type": "document_url", "document_url": "https://example.com/report.pdf" }, "include_image_base64": false }'Invalid page bounds, conflicting document variants, and inconsistent annotation options fail before the network call. Supported formats and model-specific features can change; check the Mistral OCR reference and the DeepIntShield Mistral guide.
Delegated MCP authentication
Resource-bound token exchange and identity-isolated sessions.
Durable webhooks
At-least-once event delivery and payload object storage.