Skip to content

OpenAI

Using GPT-6 Astra? Switching models can require changing the endpoint and request/response handling. Use Astra with Responses for Python, JavaScript, curl, and MCP examples. Astra function tools require Responses, and none is not a supported reasoning effort.

OpenAI is the baseline schema for DeepIntShield. When using OpenAI directly, parameters are passed through with minimal conversion - mostly validation and filtering of OpenAI-specific features.

OperationNon-StreamingStreamingEndpoint
Chat Completions✅✅/v1/chat/completions
Responses API✅✅/v1/responses
Text Completions✅✅/v1/completions
Embeddings✅-/v1/embeddings
Speech (TTS)✅✅/v1/audio/speech
Transcriptions (STT)✅✅/v1/audio/transcriptions
Image Generation✅✅/v1/images/generations
Image Edit✅✅/v1/images/edits
Image Variation✅-/v1/images/variations
Files✅-/v1/files
Batch✅-/v1/batches
Video Generation✅-/v1/videos
List Models✅-/v1/models

OpenAI supports retrieve, streaming retrieve, delete, cancel, input-item listing, and compaction through DeepIntShield. It also provides the full Realtime WebSocket frame relay and WebRTC/SDP control passthrough. See Protocol operations for routes, configured-key affinity, payload bounds, and current call-control limits.

Request Parameters

ParameterTypeRequiredNotes
modelstring✅Model identifier
messagesarray✅ChatMessage array with roles (docs)
temperaturefloat❌Sampling temperature (0-2)
top_pfloat❌Nucleus sampling parameter
stopstring/array❌Stop sequences
max_completion_tokensint❌Min 16, max output tokens
frequency_penaltyfloat❌Frequency penalty (-2 to 2)
presence_penaltyfloat❌Presence penalty (-2 to 2)
logit_biasobject❌Token logit adjustments
logprobsbool❌Include log probabilities
top_logprobsint❌Number of log probabilities per token
seedint❌Reproducibility seed
response_formatobject❌Output format (docs)
toolsarray❌Tool objects (docs)
tool_choicestring/object❌"auto", "none", "required", or specific tool
parallel_tool_callsbool❌Allow multiple simultaneous tool calls
stream_optionsobject❌Streaming options (docs)
reasoningobject❌Reasoning parameters (DeepIntShield docs, OpenAI docs)
userstring❌Truncated to 64 chars
metadataobject❌Custom metadata
storebool❌Filtered for non-OpenAI routing
service_tierstring❌Filtered for non-OpenAI routing
prompt_cache_keystring❌Filtered for non-OpenAI routing
predictionobject❌Predicted output for acceleration
audioobject❌Audio output config
modalitiesarray❌Response modalities (text, audio)

  • Reasoning: For the OpenAI SDK’s Chat Completions API, use reasoning_effort; DeepIntShield also accepts reasoning.effort and serializes it as reasoning_effort upstream. Supported efforts depend on the model. reasoning.max_tokens is not forwarded to OpenAI. See DeepIntShield reasoning docs.
  • Messages: All message roles are supported: system, user, assistant, tool, developer (treated as system). Content types: text, images via URL (image_url), audio input (input_audio). Tool messages include a tool_call_id.
  • Tools: Standard OpenAI tool format with strict mode support. Tool choice: "auto", "none", "required", or specific tool by name.
  • Responses: Passed through in standard OpenAI format. Finish reasons: stop, length, tool_calls, content_filter. Usage includes token counts and optionally cached/reasoning token details.
  • Streaming: Server-Sent Events format with delta.content, delta.tool_calls, finish_reason, and usage (final chunk only, automatically included by DeepIntShield). stream_options: { include_usage: true } is set by default for all streaming calls.
  • Cache Control: cache_control fields are stripped from messages, their content blocks, and tools before sending.
  • Token Enforcement: max_completion_tokens is enforced to have a minimum of 16. Values below 16 are automatically set to 16.
  • Special handling: user field is truncated to 64 characters; prompt_cache_key, store, service_tier are filtered when routing to non-OpenAI providers

Troubleshooting model access and parameters

Section titled “Troubleshooting model access and parameters”

Playground uses Responses for bundled GPT-5.6 Sol and GPT-6 Astra profiles. It translates the controls to Responses fields, supports JSON and SSE output, and saves opaque reasoning items for continuations on the same provider and model. An explicit preferred_api in synced model metadata remains authoritative.

An OpenAI model_not_found error can mean the configured upstream project lacks access to the requested model. Listing a model in DeepIntShield’s catalog does not grant that project access. Select a model available to that project or configure a provider key from a project with access.

A DeepIntShield model_blocked error instead means the virtual key’s allowed models exclude the request. Check Governance Hub → Virtual Keys: the virtual key policy, provider key model list, and OpenAI project permissions must all permit the model. Provider key changes do not expand a virtual key’s allowlist.

Virtual keys with MCP bindings can add function tools to a request even when the SDK call does not supply tools. If OpenAI rejects function tools with reasoning for gpt-5.6-sol on Chat Completions, set reasoning_effort="none":

import os
from openai import OpenAI
client = OpenAI(
base_url="https://app.deepintshield.com/openai",
api_key=os.environ["DEEPINTSHIELD_API_KEY"],
)
completion = client.chat.completions.create(
model="gpt-5.6-sol",
reasoning_effort="none",
messages=[{"role": "user", "content": "What is a transformer?"}],
)
print(completion.choices[0].message.content)

To retain reasoning with tools, use the Responses API instead:

response = client.responses.create(
model="gpt-5.6-sol",
reasoning={"effort": "medium"},
input="What is a transformer?",
)
print(response.output_text)

Omit temperature when the provider reports that only its default value is supported. If the error names a temperature that your SDK did not send, check the gateway’s hallucination-control temperature clamp and strict response-consistency settings. Gateway sampling overrides skip incompatible reasoning models; other configured hallucination-control techniques still apply. In Playground, load the model’s supported parameters before running a saved session; temperature from another model must not carry over. GPT-6 Astra does not support none reasoning, and its tool calling requires Responses. See the OpenAI model guide and GPT-5.6 Sol model documentation.

The Responses API is OpenAI’s structured output API.

Request Parameters

ParameterTypeRequiredNotes
modelstring✅Model identifier
inputstring/array✅Text or ContentBlock array (docs)
max_output_tokensint✅Maximum output length
backgroundbool❌Run request in background mode
conversationstring❌Conversation ID for continuing a conversation
includearray❌Array of fields to include in response (e.g., "web_search_call.action.sources")
instructionsstring❌System instructions
max_tool_callsint❌Maximum number of tool calls
metadataobject❌Custom metadata
parallel_tool_callsbool❌Allow multiple simultaneous tool calls
previous_response_idstring❌ID of previous response to continue from
prompt_cache_keystring❌Prompt caching key
reasoningobject❌ResponsesParametersReasoning configuration (DeepIntShield docs)
safety_identifierstring❌Safety identifier for content filtering
service_tierstring❌Service tier for the request
stream_optionsobject❌ResponsesStreamOptions configuration
storebool❌Store the response for later retrieval
temperaturefloat❌Sampling temperature
textobject❌ResponsesTextConfig for output formatting
top_logprobsint❌Number of log probabilities to return per token
top_pfloat❌Nucleus sampling parameter
tool_choicestring/object❌ResponsesToolChoice strategy
toolsarray❌ResponsesTool objects (docs)
truncationstring❌Truncation strategy (auto or off)
userstring❌Truncated to 64 chars

Special Message Handling (gpt-oss vs other models):

OpenAI models handle reasoning differently depending on the model family:

  • Non-gpt-oss models (GPT-4o, o1, etc.): Send reasoning as summaries. Reasoning-only messages (with no summary and only content blocks) are filtered out since these models don’t support reasoning content blocks in the request format.
  • gpt-oss models: Send reasoning as content blocks. Reasoning summaries in the request are converted to content blocks since gpt-oss expects reasoning as structured blocks, not summaries.

This conversion ensures compatibility across different model architectures for the structured Responses API. See DeepIntShield reasoning docs for detailed reasoning handling.

Token & Parameter Enforcement:

  • max_output_tokens is enforced to have a minimum of 16. Values below 16 are automatically set to 16.
  • reasoning.max_tokens field is automatically removed from JSON output (OpenAI Responses API doesn’t accept it).

Other conversions:

  • Action types zoom and region are converted to screenshot
  • cache_control fields are stripped from messages and tools
  • Unsupported tool types are silently filtered (only these are supported: function, file_search, computer_use_preview, web_search, mcp, code_interpreter, image_generation, local_shell, custom, web_search_preview)

Response: Includes id, status (completed, incomplete, pending, error), output array with message content, and token usage.

Streaming: Server-Sent Events with types: response.created, response.in_progress, response.output_item.added, response.content_part.added, response.output_text.delta, response.function_call_arguments.delta, response.completed, response.incomplete. stream_options: { include_usage: true } is set by default for all streaming calls.


Request Parameters

ParameterTypeRequiredNotes
modelstring✅Model identifier
promptstring/array✅Completion prompt(s)
max_tokensint❌Maximum output tokens
temperaturefloat❌Sampling temperature
top_pfloat❌Nucleus sampling
stopstring/array❌Stop sequences
userstring❌Truncated to 64 chars

  • Array prompts generate multiple completions. Finish reasons: stop or length. Streaming uses SSE format. stream_options: { include_usage: true } is set by default for streaming calls.
  • user field is truncated to 64 characters or set to nil if it exceeds the limit.

Request Parameters

ParameterTypeRequiredNotes
modelstring✅Model identifier
inputstring/array✅Text(s) to embed (docs)
encoding_formatstring❌float or base64
dimensionsint❌Output embedding dimensions
userstring❌NOT truncated (unlike chat/text)

  • No streaming support. Returns embedding array with usage counts.

Request Parameters

ParameterTypeRequiredNotes
modelstring✅tts-1 or tts-1-hd
inputstring✅Text to convert to speech
voicestring✅alloy, echo, fable, onyx, nova, shimmer
response_formatstring❌mp3, opus, aac, flac, wav, pcm
speedfloat❌0.25 to 4.0 (default 1.0)

  • Returns raw binary audio. Streaming supported in SSE format (base64 chunks), but not all models support streaming. stream_options: { include_usage: true } is set by default for streaming calls.

Request Parameters

ParameterTypeRequiredNotes
filebinary✅Audio file (multipart form-data)
modelstring✅whisper-1
languagestring❌ISO-639-1 language code
promptstring❌Optional prompt for context
temperaturefloat❌Sampling temperature
response_formatstring❌json, text, srt, vtt, verbose_json

  • Supported audio formats: mp3, mp4, mpeg, mpga, m4a, wav, webm
  • Response: Includes text, task, language, duration, and optionally word-level timing. Streaming supported in SSE format. stream_options: { include_usage: true } is set by default for streaming calls.

Request Parameters

ParameterTypeRequiredNotes
modelstring✅Model identifier (e.g., gpt-image-1; the legacy dall-e-3 / dall-e-2 are also accepted where your OpenAI org still has access)
promptstring✅Text description of the image to generate
nint❌Number of images to generate (1-10)
sizestring❌Image size: "256x256", "512x512", "1024x1024", "1792x1024", "1024x1792", "1536x1024", "1024x1536", "auto"
qualitystring❌Image quality: "auto", "high", "medium", "low", "hd", "standard"
stylestring❌Image style: "natural", "vivid" - DALL·E 3 only
response_formatstring❌Response format: "url" or "b64_json" - DALL·E models only; gpt-image-1 always returns b64_json
backgroundstring❌Background: "transparent", "opaque", "auto"
output_formatstring❌Output format: "png", "webp", "jpeg"
output_compressionint❌Compression level (0-100%)
partial_imagesint❌Number of partial images (0-3)
moderationstring❌Moderation level: "low", "auto"
userstring❌User identifier

Request Behavior

OpenAI is the baseline schema for image generation. Parameters pass through with minimal conversion:

  • Model, Prompt & Parameters: Sent to OpenAI as-is - no field mapping or transformation is performed.
  • Streaming: Set "stream": true in the request body to stream partial images.

Response Behavior

  • Non-streaming: OpenAI responses are returned as-is in the DeepIntShield image-generation response format.

  • Streaming: Streaming responses use Server-Sent Events (SSE) with these event types:

    • image_generation.partial_image: Intermediate image chunks with b64_json data
    • image_generation.completed: Final chunk for each image with usage information
    • error: Error events

    Each chunk includes:

    • type: Event type
    • sequence_number: Sequence number of the chunk
    • partial_image_index: Image index (0-N) for partial images
    • b64_json: Base64-encoded image data (may be null)
    • usage: Token usage (only in completed events)
    • created_at, size, quality, background, output_format: Additional metadata

    Chunks are delivered in order per image, with usage attached to the completed chunk for each image.

Endpoint: /v1/images/generations


Request Parameters

ParameterTypeRequiredNotes
modelstring✅Model identifier
promptstring✅Text description of the edit
image[]binary✅Image file(s) to edit (multipart form-data, supports multiple images)
maskbinary❌Mask image file (multipart form-data)
nint❌Number of images to generate (1-10)
sizestring❌Image size: "256x256", "512x512", "1024x1024", "1536x1024", "1024x1536", "auto"
qualitystring❌Image quality: "auto", "high", "medium", "low", "standard"
response_formatstring❌Response format: "url" or "b64_json" - DALL·E models only; gpt-image-1 always returns b64_json
backgroundstring❌Background: "transparent", "opaque", "auto"
input_fidelitystring❌Input fidelity: "low", "high"
partial_imagesint❌Number of partial images (0-3)
output_formatstring❌Output format: "png", "webp", "jpeg"
output_compressionint❌Compression level (0-100%)
userstring❌User identifier
streambool❌Enable streaming response

Request Behavior

  • Inputs: Send your model, prompt, and one or more images. Parameters pass through to OpenAI as-is.
  • Multipart Form Data: The request is sent as multipart/form-data:
    • Model & Prompt: Sent as form fields (model, prompt)
    • Images: Each image is sent as a separate image[] field with automatic MIME type detection (image/jpeg, image/webp, image/png)
    • Mask: If present, sent as a mask field with MIME type detection
    • Optional Parameters: All optional parameters (n, size, quality, response_format, background, input_fidelity, partial_images, output_format, output_compression, user) are sent as form fields
    • Streaming: Set stream: true to stream partial images

Response Behavior

  • Non-streaming: OpenAI responses are returned as-is in the DeepIntShield image-generation response format.

  • Streaming: Streaming responses use Server-Sent Events (SSE) with these event types:

    • image_edit.partial_image: Intermediate image chunks with b64_json data
    • image_edit.completed: Final chunk for each image with usage information
    • error: Error events

    Each chunk includes:

    • type: Event type (image_edit.partial_image or image_edit.completed)
    • sequence_number: Sequence number of the chunk
    • partial_image_index: Image index (0-N) for partial images
    • b64_json: Base64-encoded image data (may be null)
    • usage: Token usage (only in completed events)

    Chunks are delivered in order per image, with usage attached to the completed chunk for each image.

Endpoint: /v1/images/edits


Request Parameters

ParameterTypeRequiredNotes
modelstring✅Model identifier
imagebinary✅Image file to create variations from (multipart form-data)
nint❌Number of images to generate (1-10)
sizestring❌Image size: "256x256", "512x512", "1024x1024", "1792x1024", "1024x1792", "1536x1024", "1024x1536", "auto"
response_formatstring❌Response format: "url" or "b64_json" - DALL·E models only; gpt-image-1 always returns b64_json
userstring❌User identifier

Request Behavior

  • Inputs: Send your model and a single image. Parameters pass through to OpenAI as-is.
  • Multipart Form Data: The request is sent as multipart/form-data:
    • Model: Sent as a form field (model)
    • Image: Sent as an image field with automatic MIME type detection (image/jpeg, image/webp, image/png), defaulting to image/png if the type can’t be detected
    • Optional Parameters: All optional parameters (n, size, response_format, user) are sent as form fields
  • Single Image Only: OpenAI’s image variation API accepts only one input image. If you supply more, only the first is used.

Response Behavior

  • Non-streaming: OpenAI responses are returned as-is in the DeepIntShield image-variation response format.
  • Streaming: Not supported for image variation requests.

Endpoint: /v1/images/variations


Request Parameters

ParameterTypeRequiredNotes
filebinary✅File to upload (multipart form-data)
purposestring✅batch, fine-tune, or assistants
filenamestring❌Custom filename (defaults to file.jsonl)

Response: FileObject with id, bytes, created_at, filename, purpose, status (docs)

Query Parameters

ParameterTypeRequiredNotes
purposestring❌Filter by purpose
limitint❌Results per page
afterstring❌Pagination cursor
orderstring❌asc or desc

Cursor-based pagination with has_more flag.

Operations:

  • GET /v1/files/{file_id} - Retrieve file metadata
  • DELETE /v1/files/{file_id} - Delete file
  • GET /v1/files/{file_id}/content - Download file content

Request Parameters

ParameterTypeRequiredNotes
input_file_idstringConditionalFile ID OR requests array (not both)
requestsarrayConditionalBatchRequestItem objects (converted to JSONL)
endpointstring✅Target endpoint (e.g., /v1/chat/completions)
completion_windowstring❌24h (default)
metadataobject❌Custom metadata

Response: DeepIntShieldBatchCreateResponse with id, endpoint, input_file_id, status, created_at, request_counts (docs). Statuses: BatchStatus (validating, failed, in_progress, finalizing, completed, expired, cancelling, cancelled)

Query Parameters

ParameterTypeRequiredNotes
limitint❌Results per page
afterstring❌Pagination cursor

Operations:

  • GET /v1/batches/{batch_id} - Get batch DeepIntShieldBatchRetrieveResponse (docs)
  • POST /v1/batches/{batch_id}/cancel - Cancel batch (docs)
  1. Batch must be completed (has output_file_id)
  2. Download output file via Files API
  3. Parse JSONL - each BatchResultItem: {id, custom_id, response: {status_code, body}}

GET /v1/models - Lists available models with metadata. Model IDs in DeepIntShield responses are prefixed with openai/ (e.g., openai/gpt-4o). Results are aggregated from all configured API keys. No request body or parameters required.


Request Parameters

ParameterTypeRequiredNotes
modelstring✅e.g., sora-2
promptstring✅Text description of the video
input_referencestring❌Input image for image-to-video. Must be a base64 data URL (e.g., data:image/png;base64,...). Plain URLs are not accepted.
secondsstring❌Duration in seconds
sizestring❌Resolution: 720x1280 (default), 1280x720, 1024x1792, 1792x1024

Response: DeepIntShieldVideoGenerationResponse - id, status, model, prompt, created_at

Job Statuses: queued → in_progress → completed / failed

Retrieve / Download / Delete / List / Remix

Section titled “Retrieve / Download / Delete / List / Remix”
OperationEndpointNotes
Get statusGET /v1/videos/{id}Poll until status: completed
DownloadGET /v1/videos/{id}/contentReturns raw video bytes
DeleteDELETE /v1/videos/{id}Removes video job
List jobsGET /v1/videosQuery params: after, limit, order
RemixPOST /v1/videos/{id}/remixBody: {"prompt": "..."}

HTTP Status → Error Type mapping:

  • 400 - invalid_request_error
  • 401 - authentication_error
  • 403 - permission_error
  • 404 - not_found_error
  • 429 - rate_limit_error
  • 500 - api_error