Skip to content

Streaming Responses

DeepIntShield 2.8.3 routes streaming requests through the selected model’s adapter. Inspect the final outcome: receiving text or reaching the end of a connection does not establish that the request completed successfully.

Terminal window
pip install "deepintshield==2.8.3"
export DEEPINTSHIELD_BASE_URL="https://app.deepintshield.com"
export DEEPINTSHIELD_VIRTUAL_KEY="sk-ds-your-virtual-key"

The base URL is the gateway origin. shield.openai() supplies the /openai endpoint and returns a native OpenAI client. Existing clients can instead use <gateway>/v1 with the virtual key as their API key.

Read delta.content as it arrives, handle refusals, and check finish_reason. Empty choices can occur on a usage-only chunk.

from deepintshield import DeepintShield
with DeepintShield.from_env() as shield:
with shield.openai() as client:
stream = client.chat.completions.create(
model="openai/gpt-4o-mini",
messages=[{"role": "user", "content": "Write a haiku about the ocean."}],
stream=True,
)
finish_reason = None
try:
for chunk in stream:
if not chunk.choices:
continue
choice = chunk.choices[0]
if choice.delta.refusal:
raise RuntimeError("The model refused the request.")
if choice.delta.content:
print(choice.delta.content, end="", flush=True)
if choice.finish_reason is not None:
finish_reason = choice.finish_reason
finally:
stream.close()
if finish_reason != "stop":
raise RuntimeError(f"Text answer did not finish: {finish_reason!r}")

This example expects a complete text answer. An application using tools must handle tool_calls as a request to execute its tool workflow. length and content_filter are separate outcomes. An interrupted stream with no terminal chunk has no confirmed completion.

For raw HTTP, use curl --no-buffer with stream: true in the request body. Chat streams use data: frames and a [DONE] terminator. Inspect the JSON chunks for finish reasons and errors; [DONE] alone does not describe the answer’s outcome. Usage and content fields depend on the provider and request.

Responses uses typed events for output items, text, reasoning, tool arguments, refusals and terminal state. In 2.8.3, supported gateway conversions preserve the terminal output snapshot so native final-response helpers can reconstruct the result, including translated Anthropic and Gemini streams.

from deepintshield import DeepintShield
with DeepintShield.from_env() as shield:
with shield.openai() as client:
with client.responses.stream(
model="openai/gpt-4o-mini",
input="Tell me one interesting fact about Mars.",
) as stream:
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
final = stream.get_final_response()
if final.status != "completed" or final.error is not None:
raise RuntimeError(f"Response did not complete: {final.status}")
refused = any(
getattr(part, "type", None) == "refusal"
for item in final.output
for part in (getattr(item, "content", None) or [])
)
if refused:
raise RuntimeError("The model refused the request.")
OutcomeApplication handling
response.completedRead the final output items; these can include tools or a refusal instead of ordinary text.
response.incompleteInspect incomplete_details; handle partial work explicitly.
response.failed or an error eventHandle the error; previously received deltas do not make the answer complete.
Connection ends before terminal stateTreat the result as interrupted. A native final-response helper may raise instead of returning a response.

Responses SSE uses typed event:/data: frames. Its terminal response event determines completion; transport EOF is insufficient. Preserve structured reasoning, tool-call and tool-result items when continuing a conversation. See protocol operations and tool calling.

shield.async_openai() returns a native AsyncOpenAI client. Use async with for its lifetime, await for requests, and async for for event streams. This is ordinary nonblocking client I/O. The separate gateway async job API submits work for later polling and does not support streaming.

Use the same Chat or Responses flow for image/PDF understanding on a model that supports the input. Speech and transcription use dedicated endpoints; available binary/SSE formats depend on the selected model. Chat streaming support does not establish audio or video streaming support.

Output inspection is incremental: already delivered text cannot be recalled, and redacting a terminal snapshot does not sanitize earlier deltas. Choose nonstreaming inference or an application buffering boundary when the complete output must receive a policy verdict before delivery.

Provider timeouts and connection failures remain observable during streaming. Close streams when a user cancels or stops reading, and handle native SDK exceptions alongside terminal events.