Streaming Responses
DeepIntShield 2.8.3 routes streaming requests through the selected model’s adapter. Inspect the final outcome: receiving text or reaching the end of a connection does not establish that the request completed successfully.
pip install "deepintshield==2.8.3"export DEEPINTSHIELD_BASE_URL="https://app.deepintshield.com"export DEEPINTSHIELD_VIRTUAL_KEY="sk-ds-your-virtual-key"The base URL is the gateway origin. shield.openai() supplies the /openai
endpoint and returns a native OpenAI client. Existing clients can instead use
<gateway>/v1 with the virtual key as their API key.
Streaming Chat Responses
Section titled “Streaming Chat Responses”Read delta.content as it arrives, handle refusals, and check finish_reason.
Empty choices can occur on a usage-only chunk.
from deepintshield import DeepintShield
with DeepintShield.from_env() as shield: with shield.openai() as client: stream = client.chat.completions.create( model="openai/gpt-4o-mini", messages=[{"role": "user", "content": "Write a haiku about the ocean."}], stream=True, ) finish_reason = None try: for chunk in stream: if not chunk.choices: continue choice = chunk.choices[0] if choice.delta.refusal: raise RuntimeError("The model refused the request.") if choice.delta.content: print(choice.delta.content, end="", flush=True) if choice.finish_reason is not None: finish_reason = choice.finish_reason finally: stream.close() if finish_reason != "stop": raise RuntimeError(f"Text answer did not finish: {finish_reason!r}")This example expects a complete text answer. An application using tools must
handle tool_calls as a request to execute its tool workflow. length and
content_filter are separate outcomes. An interrupted stream with no terminal
chunk has no confirmed completion.
For raw HTTP, use curl --no-buffer with stream: true in the request body.
Chat streams use data: frames and a [DONE] terminator. Inspect the JSON
chunks for finish reasons and errors; [DONE] alone does not describe the
answer’s outcome. Usage and content fields depend on the provider and request.
Responses API Streaming
Section titled “Responses API Streaming”Responses uses typed events for output items, text, reasoning, tool arguments, refusals and terminal state. In 2.8.3, supported gateway conversions preserve the terminal output snapshot so native final-response helpers can reconstruct the result, including translated Anthropic and Gemini streams.
from deepintshield import DeepintShield
with DeepintShield.from_env() as shield: with shield.openai() as client: with client.responses.stream( model="openai/gpt-4o-mini", input="Tell me one interesting fact about Mars.", ) as stream: for event in stream: if event.type == "response.output_text.delta": print(event.delta, end="", flush=True) final = stream.get_final_response() if final.status != "completed" or final.error is not None: raise RuntimeError(f"Response did not complete: {final.status}") refused = any( getattr(part, "type", None) == "refusal" for item in final.output for part in (getattr(item, "content", None) or []) ) if refused: raise RuntimeError("The model refused the request.")| Outcome | Application handling |
|---|---|
response.completed | Read the final output items; these can include tools or a refusal instead of ordinary text. |
response.incomplete | Inspect incomplete_details; handle partial work explicitly. |
response.failed or an error event | Handle the error; previously received deltas do not make the answer complete. |
| Connection ends before terminal state | Treat the result as interrupted. A native final-response helper may raise instead of returning a response. |
Responses SSE uses typed event:/data: frames. Its terminal response event
determines completion; transport EOF is insufficient. Preserve structured
reasoning, tool-call and tool-result items when continuing a conversation. See
protocol operations and
tool calling.
Native async clients
Section titled “Native async clients”shield.async_openai() returns a native AsyncOpenAI client. Use async with
for its lifetime, await for requests, and async for for event streams. This
is ordinary nonblocking client I/O. The separate
gateway async job API submits work for later
polling and does not support streaming.
Multimodal streams and policy checks
Section titled “Multimodal streams and policy checks”Use the same Chat or Responses flow for image/PDF understanding on a model that supports the input. Speech and transcription use dedicated endpoints; available binary/SSE formats depend on the selected model. Chat streaming support does not establish audio or video streaming support.
Output inspection is incremental: already delivered text cannot be recalled, and redacting a terminal snapshot does not sanitize earlier deltas. Choose nonstreaming inference or an application buffering boundary when the complete output must receive a policy verdict before delivery.
Provider timeouts and connection failures remain observable during streaming. Close streams when a user cancels or stops reading, and handle native SDK exceptions alongside terminal events.