Skip to content

Chat and guardrails

with shield.openai() as client:
response = client.chat.completions.create(
model="openai/gpt-4o-mini",
messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)

The native OpenAI client is the primary inference path in SDK 2.8.3 and is included in the core package. It retains native response types, streams, errors, retries, and request arguments. Choose the API and parameters supported by the selected gateway provider and model.

Use provider/model-id to select another configured provider. Preserve the provider’s model or deployment ID and omit optional sampling/reasoning settings unless that model supports them. File, batch, media, and Responses operations have their own gateway coverage; see the operation guide.

For a model with Responses support:

with shield.openai() as client:
response = client.responses.create(
model="openai/gpt-4o-mini",
input="Explain why the sky is blue in one sentence.",
store=False,
)
if response.status != "completed":
raise RuntimeError(f"Response did not complete: {response.status}")
print(response.output_text)

GPT-6 Astra tool workflows require Responses, including gateway-injected MCP tools. Follow Astra with Responses for its model-specific arguments.

The native streaming API returns typed events. Consume the terminal outcome as well as text deltas:

with shield.openai() as client:
completed = False
with client.responses.create(
model="openai/gpt-4o-mini",
input="Explain the result.",
store=False,
stream=True,
) as stream:
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
elif event.type in {"response.failed", "response.incomplete", "error"}:
raise RuntimeError(f"Response did not complete: {event.type}")
elif event.type == "response.completed":
completed = event.response.status == "completed"
if not completed:
raise RuntimeError("Stream ended without a completed response")

For native Chat Completions, use chat.completions.create(..., stream=True) and read chunk.choices[*].delta; inspect finish_reason for truncation, filtering, or a tool handoff. A text-only printer does not execute function calls. Responses likewise has separate function-argument, reasoning, refusal, and output-item events; handle the event types your application uses. Native helpers such as client.responses.stream(...) retain their installed SDK’s behavior, including final-response assembly.

Native async streams use await client.responses.create(..., stream=True), then async with stream and async for event in stream. Preserve the same terminal-outcome checks and close the stream when cancelling or stopping early.

Set stream=True to receive the exported synchronous ChatCompletionStream:

from deepintshield import ChatCompletionStream
stream: ChatCompletionStream = shield.chat(
model="openai/gpt-4o-mini",
messages=[{"role": "user", "content": "Explain the result"}],
stream=True,
)
with stream:
for chunk in stream:
print(chunk)

The call opens the HTTP response and validates its initial status and required SSE content type before it returns, but it does not buffer a successful response body. Iteration yields one decoded dict[str, Any] for each SSE data: event. SSE comments and events without data are ignored, multiple data: lines in one event are joined with a newline, and the required terminal data: [DONE] event is consumed without being yielded. An EOF before [DONE] is treated as a truncated stream. The parser caps each successful SSE event at 1 MiB before JSON decoding. An oversized event closes the response and raises chat_stream_invalid_event with details.reason == "event_too_large".

The stream is a single-consumer, closeable iterator and context manager. It closes on [DONE], a truncated EOF, an event or transport failure, explicit/context close, and early exit from a for loop. close() is idempotent, and closed reports whether the response has been released. The stream retains its owning client until it closes, so a temporary expression such as DeepintShield(...).chat(stream=True) remains valid. Prefer with, or call stream.close() when using next(stream) or when ownership crosses a function boundary, so partial consumption releases the connection deterministically.

Initial HTTP and transport failures normally raise from chat() before an iterator is returned. A non-success HTTP response retains at most 64 KiB of its body for diagnostics. During iteration:

ConditionBehavior
Missing/wrong SSE content type, malformed UTF-8/JSON, oversized event, non-object payload, or EOF before [DONE]DeepintShieldError with chat_stream_invalid_event
SSE error event or decoded object carrying an errorDeepintShieldError; a recognized gateway code is normalized, otherwise the fallback is chat_request_failed
Mid-stream HTTP/timeout failureNative httpx exception with transport_error or transport_timeout metadata
Owning DeepintShield client closed before the next readDeepintShieldError with client_closed

The stream itself is synchronous. Use shield.async_openai() when you need a native asynchronous inference iterator. See Error codes for structured handling.

Use evaluate_guardrail() when you want a verdict object without automatic exception behavior:

result = shield.evaluate_guardrail(
stage="input",
input="Please summarize this ticket",
model="gpt-4o-mini",
provider="openai",
metadata={"ticket_id": "T-100"},
)
print(result.decision, result.reason, result.mode)
if result.blocked:
# The application chooses what happens next.
...

The SDK posts to POST /api/guardrails/evaluate and converts the response into GuardrailResult:

AttributeMeaning
decisionLowercase gateway decision.
stageStage supplied by the caller.
reasonGateway reason, possibly empty.
modeReported enforcement mode such as sync, enforce, shadow, or async; possibly empty.
rawComplete decoded gateway response. Treat as sensitive.
allowedIn SDK 2.8.5 and later, True for allow, allow_with_redaction, redact, or monitor.
blockedLogical inverse of allowed.

Policy actions and returned decisions use different names. A blocking policy uses action=block, while an enforced rejection returns decision="deny" and blocked=True. A redaction policy uses action=redact, while an allowed, redacted request returns decision="allow_with_redaction" and blocked=False. Use the returned decision names in test expectations. deny, sandbox, approval decisions, and unknown decisions remain blocking.

Upgrade to SDK 2.8.5 or later before checking redaction results: earlier versions incorrectly classify allow_with_redaction as blocked. The SDK preserves the gateway’s decision; it does not rename deny to block.

The allowed property is decision-based; it does not reinterpret mode. When operating policies in shadow/advisory mode, inspect both fields and validate the behavior in a staging workspace.

StagePrimary content fieldsTypical helper
inputinputshield.agent.check_input(text)
outputoutputshield.agent.check_output(text)
actiontool_input, tool_name, action_class, domainsshield.agent.evaluate_tool(...) without a server
mcpaction fields plus server_labelshield.agent.evaluate_tool(...) with a server
ragRAG-specific endpoint is preferred for retrieved chunksshield.rag.evaluate(...)

Other evaluation fields include actor type/id/role/customer/team, model, provider, application and agent names, metadata, and persist. The gateway currently normalizes an unrecognized stage to input; applications should still treat the five values above as the supported contract and validate stage names before calling.

guard() evaluates and raises DeepintShieldBlockedError by default:

from deepintshield import DeepintShieldBlockedError
try:
shield.guard(stage="output", output=answer)
except DeepintShieldBlockedError as exc:
print(exc.code, exc.stage, exc.decision, exc.reason)

Set raise_on_block=False to always receive the result. The blocking exception retains the structured stage, decision, reason, status/payload fields, and the central guardrail_blocked error code. Do not display payload or raw policy reasons directly to end users.

shield.agent is a lightweight façade over the same guardrail endpoint:

from deepintshield import ToolInvocation
shield.agent.check_input(user_text)
shield.agent.evaluate_tool(
ToolInvocation(
tool_name="write_invoice",
tool_input={"invoice_id": "INV-7"},
action_class="write",
domains=["billing"],
)
)
shield.agent.check_output(model_text)

@shield.agent.tool(...) performs this explicit pre-call check around one synchronous function. For identity-aware PDP decisions, durable approvals, registration, and automatic framework boundaries, use the Agentic surface instead.

guard_turn() evaluates input, each listed tool call, and optional output in that order and returns a dictionary keyed by stage/tool. It stops at the first blocking exception unless raise_on_block=False.

Requests made through a provider client can also be guarded transparently by policies attached to the virtual key. Those failures normally surface through the provider SDK’s HTTP exception type, not DeepintShieldBlockedError, because the provider SDK owns response decoding. Use the response headers and structured gateway body when the provider exception exposes them, then map the code through the central error catalog.

Attachment and dedicated media inspection have different coverage from plain text. Inline PDF text and supported image metadata can participate in selected policies; remote URLs, file IDs, image pixels, and raw audio/video do not automatically become inspected text. See Multimodal inference for supported inputs, binary-redaction behavior, and the server multimodal guardrail setting.

Guardrail accuracy is policy- and detector-dependent. Test known-safe, known-bad, boundary, multilingual, and multimodal samples before enforcing a policy, and monitor false-positive/false-negative rates after deployment.