Skip to content

Multimodal Support

Use DeepIntShield 2.8.3 with the native OpenAI client for text, images, PDFs, speech, transcription and supported video jobs. Select a model that supports the operation and is available through your provider account and virtual key.

WorkloadNative methodConsiderations
Image understandingchat.completions.create() / responses.create()Send image content parts to a vision-capable model.
PDF understandingresponses.create()Send an input_file part to a model/adapter that accepts PDFs.
Image generation/editingimages.generate() / images.edit()Reference inputs, masks and output formats are model-specific.
Speechaudio.speech.create()Requires a valid model/voice; returns audio bytes.
Transcriptionaudio.transcriptions.create()Upload an accepted audio format to a transcription model.
Video generationvideos.create() and status/content methodsPoll within a deadline and download only after completion.
Uploaded filesfiles.create(), files.retrieve(), files.delete()Lifecycle support does not imply every inference model accepts file IDs.

The Playground uses provider/model metadata to show operation-specific controls. Unknown models require an explicit operation choice. Provider capabilities describe adapters; model availability and task support further restrict them.

Terminal window
pip install "deepintshield==2.8.3"
export DEEPINTSHIELD_BASE_URL="https://app.deepintshield.com"
export DEEPINTSHIELD_VIRTUAL_KEY="sk-ds-your-virtual-key"

shield.openai() returns a native OpenAI client pointed at the gateway. Applications can also configure their own client with the virtual key and <gateway>/v1. Example model names illustrate request shapes; choose models configured and allowed in your deployment.

import base64
from pathlib import Path
from deepintshield import DeepintShield
image = base64.b64encode(Path("diagram.png").read_bytes()).decode("ascii")
with DeepintShield.from_env() as shield:
with shield.openai() as client:
result = client.chat.completions.create(
model="openai/gpt-4o-mini",
messages=[{"role": "user", "content": [
{"type": "text", "text": "Describe the shapes in this diagram."},
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{image}"}},
]}],
)
print(result.choices[0].message.content)

Multiple images use multiple content parts. Remote URLs may be accepted by a model, but the gateway attachment extractor does not fetch remote image bytes for policy inspection.

Use inline bytes with the actual filename and MIME type. Sending a filename or a provider file ID alone is a different operation from sending the bytes.

import base64
from pathlib import Path
from deepintshield import DeepintShield
document = Path("document.pdf")
encoded = base64.b64encode(document.read_bytes()).decode("ascii")
with DeepintShield.from_env() as shield:
with shield.openai() as client:
result = client.responses.create(
model="openai/gpt-4o-mini",
input=[{"role": "user", "content": [
{"type": "input_text", "text": "Summarize this document."},
{"type": "input_file", "filename": document.name,
"file_data": f"data:application/pdf;base64,{encoded}"},
]}],
)
if result.status != "completed" or result.error is not None:
raise RuntimeError(f"Document request did not complete: {result.status}")
print(result.output_text)

Supported Anthropic and Gemini PDF requests can also use this common route. For streaming, inspect final state and refusals as described in streaming responses.

Image Generation: Generating Images with AI

Section titled “Image Generation: Generating Images with AI”

For an image model that returns base64 output:

import base64
from pathlib import Path
from deepintshield import DeepintShield
with DeepintShield.from_env() as shield:
with shield.openai() as client:
result = client.images.generate(
model="openai/gpt-image-1",
prompt="A blue square on a white background",
size="1024x1024",
)
if not result.data or not result.data[0].b64_json:
raise RuntimeError("No base64 image was returned by the selected model.")
with Path("generated-image.png").open("xb") as output:
output.write(base64.b64decode(result.data[0].b64_json, validate=True))

Other models can return URLs or require different reference inputs and output parameters. Use the selected model’s contract. Runway image generation and editing have provider-specific requirements.

Speech returns binary audio. Transcription uploads audio to its own endpoint.

from pathlib import Path
from deepintshield import DeepintShield
with DeepintShield.from_env() as shield:
with shield.openai() as client:
with client.audio.speech.with_streaming_response.create(
model="openai/tts-1", input="Hello from DeepIntShield.",
voice="alloy", response_format="mp3",
) as response:
with Path("speech.mp3").open("xb") as output:
for chunk in response.iter_bytes():
output.write(chunk)
with Path("speech.mp3").open("rb") as audio:
transcript = client.audio.transcriptions.create(
model="openai/whisper-1", file=audio,
)
print(transcript.text)

Voice identifiers, languages, audio formats and streaming support depend on provider/model. ElevenLabs and Sarvam require their own valid voice and operation parameters; an OpenAI voice name does not select an account-specific voice.

Create a video job, poll with videos.retrieve() until terminal status or your deadline, and download content only after successful completion. Preserve opaque resource IDs and close binary download responses. Some models require an input reference image.

The SDK multimodal guide links runnable video and file lifecycle examples with bounded polling and explicit output checks. Delete only files your workflow created. Provider file IDs are references; the guardrail extractor does not resolve them to bytes.

InputCurrent inspection boundary
Text parts and eligible prompts/transcriptsEvaluated by applicable input/output policies.
Inline PDFSupported embedded text is extracted under resource limits; image-only pages do not gain OCR.
Inline imageSupported textual metadata can be extracted; pixels are not OCR-scanned by this extractor.
Remote URL or provider file IDThe extractor does not fetch or resolve the referenced bytes.
Raw audio or videoThis path does not automatically transcribe audio or inspect video frames.

Dedicated image/audio/video operation checks require the server’s GUARDRAILS_MULTIMODAL=true setting. This enables implemented policy paths; it does not add OCR, transcription, frame analysis or file resolution. An enforcing policy that requires redaction inside a binary attachment blocks when a safe rewrite is unavailable. Supply a sanitized attachment.

See guardrails, PII redaction and semantic caching for enforcement and reuse rules.