Skip to content

LiteLLM SDK

DeepIntShield 2.8.3 routes LiteLLM completions through the gateway’s OpenAI-compatible transport. Provider credentials and model translation stay on the gateway. LiteLLM owns its native completion responses and stream objects.

Terminal window
pip install "deepintshield[litellm]==2.8.3"
export DEEPINTSHIELD_VIRTUAL_KEY="sk-ds-your-virtual-key"
export DEEPINTSHIELD_BASE_URL="https://app.deepintshield.com"

Configure the chosen provider on the gateway and authorize its model for the Virtual Key. Completion calls use the gateway’s normal inference protections; they do not require an Agentic blueprint or registration.

from deepintshield import DeepintShield
with DeepintShield.from_env() as shield:
client = shield.litellm()
response = client.completion(
model="anthropic/claude-sonnet-4-5",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

The wrapper sends the original anthropic/claude-sonnet-4-5 model ID to the gateway at /litellm/chat/completions. It configures LiteLLM’s OpenAI transport and adds an outer openai/ model prefix internally. Pass the gateway’s provider/model directly to shield.litellm(); do not add that outer prefix yourself when using the wrapper.

The same pattern applies to configured Gemini, Azure, Bedrock, local, and custom provider models that support Chat Completions. Preserve the gateway provider prefix, deployment name, and any nested model path. LiteLLM’s local provider catalog does not determine which models your gateway makes available.

for chunk in client.completion(
model="anthropic/claude-sonnet-4-5",
messages=[{"role": "user", "content": "Write a short greeting."}],
stream=True,
):
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="", flush=True)

From an async function:

response = await client.acompletion(
model="anthropic/claude-sonnet-4-5",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

For async streaming, await client.acompletion(..., stream=True) and iterate over the returned stream using async for. Reuse the wrapper while its parent SDK client remains open.

Without the DeepIntShield wrapper, explicitly choose LiteLLM’s OpenAI transport and give it an outer openai/ prefix:

import os
from litellm import completion
response = completion(
model="openai/anthropic/claude-sonnet-4-5",
custom_llm_provider="openai",
api_base="https://app.deepintshield.com/litellm",
api_key=os.environ["DEEPINTSHIELD_VIRTUAL_KEY"],
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

LiteLLM removes the first openai/ before transmission, leaving the gateway’s anthropic/... identifier intact. A native bedrock/..., gemini/..., or anthropic/... selection in LiteLLM chooses that provider’s transport locally; changing api_base alone does not ensure the desired gateway protocol or credential behavior. The explicit OpenAI transport avoids local cloud credential discovery for this common inference path.

The existing /litellm prefix continues to expose selected native provider formats for applications that deliberately use them. Validate each native client’s route, authentication, and model semantics before using that path.

Pass supported native options such as timeout, stream, tools, extra_headers, and provider-supported parameters into .completion() or .acompletion(). The wrapper supplies the Virtual Key and merges gateway headers. It also preserves explicit OpenAI parameters for gateway model IDs that LiteLLM does not recognize locally. The selected model and gateway adapter still decide which parameters are valid.

LiteLLMShield exposes completion and asynchronous completion methods; it does not add Responses, embeddings, image, audio, file, or video methods. Use the native OpenAI integration or a provider-native integration for other operations. Streaming policy limits also apply; see streaming responses.

LiteLLM is an inference SDK. completion() and acompletion(), including their streaming forms, route through normal gateway inference without a synthetic llm.completion Agentic decision. This also applies inside shield.agentic.run(). Gateway authentication, Virtual Key permissions, and configured inference policies still apply.

Executable application tools and workflows keep their supported framework or service-side authorization boundary. When Agentic Security is enabled, those boundaries still require the applicable identity, blueprint, grants, and policy decision. Explicit shield.agentic.decide(...) calls also retain their decision contract. See Agents and Agentic governance for runtime settings and SDK error codes for handling failures.