LiteLLM SDK
DeepIntShield 2.8.3 routes LiteLLM completions through the gateway’s OpenAI-compatible transport. Provider credentials and model translation stay on the gateway. LiteLLM owns its native completion responses and stream objects.
Install and configure
Section titled “Install and configure”pip install "deepintshield[litellm]==2.8.3"export DEEPINTSHIELD_VIRTUAL_KEY="sk-ds-your-virtual-key"export DEEPINTSHIELD_BASE_URL="https://app.deepintshield.com"Configure the chosen provider on the gateway and authorize its model for the Virtual Key. Completion calls use the gateway’s normal inference protections; they do not require an Agentic blueprint or registration.
Use the SDK wrapper
Section titled “Use the SDK wrapper”from deepintshield import DeepintShield
with DeepintShield.from_env() as shield: client = shield.litellm() response = client.completion( model="anthropic/claude-sonnet-4-5", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content)The wrapper sends the original anthropic/claude-sonnet-4-5 model ID to the
gateway at /litellm/chat/completions. It configures LiteLLM’s OpenAI transport
and adds an outer openai/ model prefix internally. Pass the gateway’s
provider/model directly to shield.litellm(); do not add that outer prefix
yourself when using the wrapper.
The same pattern applies to configured Gemini, Azure, Bedrock, local, and custom provider models that support Chat Completions. Preserve the gateway provider prefix, deployment name, and any nested model path. LiteLLM’s local provider catalog does not determine which models your gateway makes available.
Streaming and asynchronous calls
Section titled “Streaming and asynchronous calls”for chunk in client.completion( model="anthropic/claude-sonnet-4-5", messages=[{"role": "user", "content": "Write a short greeting."}], stream=True,): if chunk.choices: print(chunk.choices[0].delta.content or "", end="", flush=True)From an async function:
response = await client.acompletion( model="anthropic/claude-sonnet-4-5", messages=[{"role": "user", "content": "Hello!"}],)print(response.choices[0].message.content)For async streaming, await client.acompletion(..., stream=True) and iterate
over the returned stream using async for. Reuse the wrapper while its parent
SDK client remains open.
Existing native LiteLLM code
Section titled “Existing native LiteLLM code”Without the DeepIntShield wrapper, explicitly choose LiteLLM’s OpenAI transport
and give it an outer openai/ prefix:
import osfrom litellm import completion
response = completion( model="openai/anthropic/claude-sonnet-4-5", custom_llm_provider="openai", api_base="https://app.deepintshield.com/litellm", api_key=os.environ["DEEPINTSHIELD_VIRTUAL_KEY"], messages=[{"role": "user", "content": "Hello!"}],)print(response.choices[0].message.content)LiteLLM removes the first openai/ before transmission, leaving the gateway’s
anthropic/... identifier intact. A native bedrock/..., gemini/..., or
anthropic/... selection in LiteLLM chooses that provider’s transport locally;
changing api_base alone does not ensure the desired gateway protocol or
credential behavior. The explicit OpenAI transport avoids local cloud
credential discovery for this common inference path.
The existing /litellm prefix continues to expose selected native provider
formats for applications that deliberately use them. Validate each native
client’s route, authentication, and model semantics before using that path.
Request options and operation limits
Section titled “Request options and operation limits”Pass supported native options such as timeout, stream, tools,
extra_headers, and provider-supported parameters into .completion() or
.acompletion(). The wrapper supplies the Virtual Key and merges gateway
headers. It also preserves explicit OpenAI parameters for gateway model IDs
that LiteLLM does not recognize locally. The selected model and gateway adapter
still decide which parameters are valid.
LiteLLMShield exposes completion and asynchronous completion methods; it does
not add Responses, embeddings, image, audio, file, or video methods. Use the
native OpenAI integration or a
provider-native integration for other
operations. Streaming policy limits also apply; see
streaming responses.
Governed execution
Section titled “Governed execution”LiteLLM is an inference SDK. completion() and acompletion(), including their
streaming forms, route through normal gateway inference without a synthetic
llm.completion Agentic decision. This also applies inside
shield.agentic.run(). Gateway authentication, Virtual Key permissions, and
configured inference policies still apply.
Executable application tools and workflows keep their supported framework or
service-side authorization boundary. When Agentic Security is enabled, those
boundaries still require the applicable identity, blueprint, grants, and policy
decision. Explicit shield.agentic.decide(...) calls also retain their decision
contract. See Agents and Agentic governance for runtime
settings and SDK error codes for handling failures.