Reasoning
Using GPT-6 Astra? Switching models can require changing the endpoint and request/response handling. Use Astra with Responses for Python, JavaScript, curl, and MCP examples. Astra function tools require Responses, and none is not a supported reasoning effort.
Overview
Section titled “Overview”Reasoning controls how a model allocates computation before answering. Depending on the provider and model, responses may include a reasoning summary, thinking content, or opaque continuation state. These are different response formats; they do not guarantee access to the model’s complete internal reasoning.
Provider Support Matrix
Section titled “Provider Support Matrix”| Provider | Request Field | Response Field | Min Budget | Effort Levels | Streaming |
|---|---|---|---|---|---|
| OpenAI | reasoning | reasoning_details | None | Model-specific | ✅ |
| Anthropic | thinking, output_config.effort | Content blocks | 1024 tokens for explicit budgets | Adaptive or budget-derived, by model | ✅ |
| Bedrock (Anthropic) | thinking, output_config.effort | Content blocks | 1024 tokens for explicit budgets | Adaptive or budget-derived, by model | ✅ |
| Gemini 2.5+ | thinking_config | thought parts | 1024 | Budget-only | ✅ |
| Gemini 3.0+ | thinking_config | thought parts | 1024 | minimal, low, medium, high + Budget | ✅ |
Request Configuration
Section titled “Request Configuration”Chat Completions API
Section titled “Chat Completions API”Add a reasoning object to your chat completions request body:
{ "model": "provider/model-name", "messages": [...], "reasoning": { "effort": "high", "max_tokens": 4096 }}from deepintshield import DeepintShield
shield = DeepintShield.from_env() # defaults to https://app.deepintshield.com
response = shield.chat( model="openai/o4-mini", messages=[{"role": "user", "content": "Explain quantum computing"}], reasoning={"effort": "high", "max_tokens": 4096},)curl --location 'https://app.deepintshield.com/v1/chat/completions' \--header 'Authorization: Bearer sk-ds-...' \--header 'Content-Type: application/json' \--data '{ "model": "openai/o4-mini", "messages": [{"role": "user", "content": "Explain quantum computing"}], "reasoning": {"effort": "high", "max_tokens": 4096}}'Responses API
Section titled “Responses API”The Responses API accepts the same reasoning object and adds an optional summary parameter:
{ "model": "provider/model-name", "input": [...], "reasoning": { "effort": "high", "max_tokens": 4096, "summary": "detailed" }}Parameter Reference
Section titled “Parameter Reference”Chat Completions API Parameters
Section titled “Chat Completions API Parameters”| Parameter | Type | Description |
|---|---|---|
effort | string | Reasoning intensity level |
max_tokens | int | Maximum tokens for reasoning (budget) |
Responses API Parameters
Section titled “Responses API Parameters”| Parameter | Type | Description |
|---|---|---|
effort | string | Reasoning intensity level |
max_tokens | int | Maximum tokens for reasoning (budget) |
summary | string | Summary level: brief, detailed, or json |
Provider-Specific Behavior
Section titled “Provider-Specific Behavior”OpenAI
Section titled “OpenAI”OpenAI uses effort-based reasoning. Supply reasoning.effort directly. If you supply only reasoning.max_tokens, the gateway derives an effort level for you.
Supported effort levels depend on the selected model. Use its catalog parameter profile rather than applying one fixed list to every OpenAI model. See the OpenAI provider guide for model-specific sampling, token-limit, and tool-calling behavior.
Anthropic
Section titled “Anthropic”The Anthropic adapter chooses thinking mode from the request and the recognized model ID:
| Request | Gateway conversion |
|---|---|
Explicit reasoning.max_tokens | Manual thinking.type: "enabled" with that budget; it takes priority over effort. |
| Effort only, recognized adaptive-thinking model | thinking.type: "adaptive" and output_config.effort. |
| Effort only, Opus 4.5 | Native effort plus a derived manual thinking budget. |
| Effort only, older model | A derived manual thinking budget. |
effort: "none", without a budget | thinking.type: "disabled". |
Adaptive recognition covers the supported Claude Opus, Sonnet, Fable, and Mythos model IDs and their supported dated, Vertex, and Bedrock forms. It does not infer capabilities from an arbitrary deployment name or a future version. Use the selected model’s parameter profile for its supported effort choices.
Thinking content is returned through reasoning_details. Preserve returned
signatures with their content when continuing a conversation; they are opaque
provider state.
The adapter applies thinking constraints after translating both normalized and native request parameters. When thinking is active, it omits incompatible temperature and top-k values and removes top-p values outside the supported range. Recognized fixed-sampling models omit all three. For recognized Claude models that prohibit temperature and top-p together, temperature takes precedence when thinking is off.
An explicit manual budget must be below the output-token limit unless the
provider’s supported interleaved-thinking mode permits otherwise. The gateway
never raises your output-token limit to accommodate the budget. Adaptive-only
models reject manual budgets; use reasoning.effort on those models. Manual
thinking also rejects forced tool use (any or a named tool); choose automatic
tool selection, no tool use, or disable thinking. These constraints also apply
to Claude routed through Bedrock and Vertex.
Dynamic Budget Handling:
| Input Value | Behavior |
|---|---|
-1 (dynamic) | Uses the minimum budget of 1024 |
< 1024 | Error |
>= 1024 | Used as-is |
Bedrock (Anthropic Models)
Section titled “Bedrock (Anthropic Models)”For Bedrock Claude models, an explicit reasoning.max_tokens selects a manual
budget. An effort-only request selects adaptive thinking for recognized
adaptive model IDs, or a derived budget for older models. effort: "none"
without a budget disables thinking.
Bedrock (Nova Models)
Section titled “Bedrock (Nova Models)”Bedrock Nova models use effort-based reasoning. Supply reasoning.effort.
| Effort | Notes |
|---|---|
minimal, low | Normal parameters allowed |
medium | Normal parameters allowed |
high | max_tokens, temperature, and top_p are not applied |
Notable differences from Anthropic on Bedrock:
- No minimum token budget constraint
- Uses effort levels instead of token budgets
- At
higheffort, conflicting sampling parameters are not sent
Gemini
Section titled “Gemini”Gemini supports both token budgets (reasoning.max_tokens) and effort levels (reasoning.effort), depending on the model version.
Model Version Support
Section titled “Model Version Support”| Gemini Version | Token Budget | Effort Level | Notes |
|---|---|---|---|
| 2.5+ | ✅ | ⚠️ (treated as a budget) | Budget-based models |
| 3.0+ | ✅ | ✅ | Support both budget and effort levels |
Effort levels on Pro models
Section titled “Effort levels on Pro models”Gemini Pro models support a narrower set of effort levels. When routed to a Pro model, the following adjustments are applied automatically:
| Effort | Non-Pro Models | Pro Models |
|---|---|---|
"none" | Disables thinking | Disables thinking |
"minimal" | minimal | low |
"low" | low | low |
"medium" | medium | high |
"high" | high | high |
Special Values
Section titled “Special Values”| Value | Field | Behavior |
|---|---|---|
0 | max_tokens | Disables reasoning |
-1 | max_tokens | Dynamic budget (Gemini decides) |
"none" | effort | Disables reasoning |
// Dynamic budget - let Gemini decide{ "reasoning": { "max_tokens": -1 } }
// Disable reasoning (either form works){ "reasoning": { "max_tokens": 0 } }{ "reasoning": { "effort": "none" } }Reasoning output is returned in the normalized reasoning_details array, the same as every other provider.
Two Reasoning Methods: Effort vs. Max Tokens
Section titled “Two Reasoning Methods: Effort vs. Max Tokens”Providers use one of two reasoning styles. You can use a single, consistent reasoning object regardless of which one the target provider expects.
| Style | Providers | Request Field |
|---|---|---|
| Effort-Based | OpenAI, AWS Bedrock Nova | reasoning.effort |
| Model-Dependent | Anthropic/Bedrock Claude, Gemini | reasoning.effort or reasoning.max_tokens |
| Budget-Based | Cohere | reasoning.max_tokens |
You can send effort and max_tokens together. The gateway uses whichever field is native to the target provider and translates the other for you, so you do not have to know each provider’s native format:
- Anthropic and Gemini: an explicit
max_tokensbudget takes priority. An effort-only request uses the model’s adaptive/level-based path when supported, otherwise a derived budget. - Cohere: an explicit
max_tokensbudget takes priority over a budget derived fromeffort. - Effort-based providers (OpenAI, Bedrock Nova): if
effortis present it is used; otherwise an effort level is derived frommax_tokens.
Omitting both fields leaves reasoning to the provider’s default behavior.
Provider-Specific Constraints
Section titled “Provider-Specific Constraints”Different providers enforce different minimum reasoning budgets:
| Provider | Minimum Budget |
|---|---|
| Anthropic | 1024 for explicit manual budgets |
| Bedrock Anthropic | 1024 for explicit manual budgets |
| Bedrock Nova | 1 |
| Cohere | 1 |
| Gemini | 1024 |
Requests below a provider’s minimum budget are clamped up to that minimum, except where a hard error applies (see the Anthropic constraint above).
Request Examples
Section titled “Request Examples”You can always send the same reasoning object; the gateway applies it to the target provider for you.
Effort on a budget-based provider (Anthropic) - works even though Anthropic is budget-based:
{ "model": "anthropic/claude-3-5-sonnet", "messages": [{"role": "user", "content": "..."}], "reasoning": {"effort": "high"}}Budget on an effort-based provider (Bedrock Nova) - works even though Nova is effort-based:
{ "model": "bedrock/us.amazon.nova-pro-v1:0", "messages": [{"role": "user", "content": "..."}], "reasoning": {"max_tokens": 2000}}Both fields provided - the field native to the target provider wins. For Anthropic (budget-based), max_tokens is used and effort is ignored:
{ "model": "anthropic/claude-3-5-sonnet", "messages": [{"role": "user", "content": "..."}], "reasoning": {"effort": "medium", "max_tokens": 2500}}Response Format
Section titled “Response Format”DeepIntShield Standard Response
Section titled “DeepIntShield Standard Response”For Chat Completions, supported reasoning output is represented in a normalized
reasoning_details array:
{ "choices": [{ "message": { "role": "assistant", "content": "Final response text", "reasoning_details": [ { "index": 0, "type": "text", "text": "Step-by-step reasoning content...", "signature": "optional_signature_for_verification" } ] } }]}Reasoning Details Fields
Section titled “Reasoning Details Fields”| Field | Type | Description | Present In |
|---|---|---|---|
index | int | Position in reasoning sequence | All |
type | string | Content type (text, encrypted, summary) | All |
text | string | Reasoning content | Chat Completions |
summary | string | Reasoning summary | Responses API |
signature | string | Opaque provider continuation state to preserve with its content | Anthropic, Bedrock |
Type Mappings
Section titled “Type Mappings”| Reasoning Type | When Used | Source |
|---|---|---|
reasoning.text | Direct thinking/reasoning content | Anthropic, Gemini, Bedrock |
reasoning.encrypted | Signature-verified reasoning | Anthropic, Bedrock Nova |
reasoning.summary | Summarized reasoning (Responses API) | All providers |
Streaming
Section titled “Streaming”Native Responses streams
Section titled “Native Responses streams”Responses requests use native Responses events and output items. They do not
use the Chat choices[].delta.reasoning_details shape shown below. Keep the
returned reasoning items, encrypted content, signatures, and function-call
items together when continuing with the same provider and model.
For translated Anthropic, Gemini, and Chat-compatible Responses streams, the
gateway preserves item identifiers, output indexes, argument deltas, and final
output snapshots. Treat response.completed, response.incomplete, and
response.failed as different outcomes; do not infer success from text deltas
or a closed connection. See Responses events and continuation.
Stream Event Types
Section titled “Stream Event Types”| Provider | Reasoning Event | Signature Event |
|---|---|---|
| OpenAI | reasoning (top-level) | N/A |
| Anthropic | thinking_delta | signature_delta |
| Bedrock | thinking_delta | signature_delta |
| Gemini | thought (in content) | thought_signature |
Anthropic Streaming Example
Section titled “Anthropic Streaming Example”// Stream eventsevent: content_block_startdata: {"type": "content_block_start", "content_block": {"type": "thinking"}}
event: content_block_deltadata: {"type": "content_block_delta", "delta": {"type": "thinking_delta", "thinking": "Let me"}}
event: content_block_deltadata: {"type": "content_block_delta", "delta": {"type": "thinking_delta", "thinking": " analyze..."}}
event: content_block_deltadata: {"type": "content_block_delta", "delta": {"type": "signature_delta", "signature": "EqoB..."}}
event: content_block_stopdata: {"type": "content_block_stop"}DeepIntShield Stream Response
Section titled “DeepIntShield Stream Response”// Thinking delta{ "choices": [{ "delta": { "reasoning_details": [{ "index": 0, "type": "text", "text": "Let me analyze..." }] } }]}
// Signature delta{ "choices": [{ "delta": { "reasoning_details": [{ "index": 0, "signature": "EqoB..." }] } }]}Caveats Summary
Section titled “Caveats Summary”Minimum Budget (Anthropic/Bedrock)
Severity: High
Behavior: An explicit manual reasoning.max_tokens must be >= 1024
Impact: Requests with lower values fail with error
Workaround: Use a supported manual budget or an effort-only adaptive request
Dynamic Budget Not Supported
Severity: Medium
Behavior: reasoning.max_tokens = -1 converted to 1024
Impact: The -1 sentinel does not select adaptive thinking
Workaround: Use effort only on a supported adaptive model, or a manual budget
Effort Level Normalization
Severity: Low
Behavior: Anthropic converts minimal to low; other adapters use model-specific mappings
Impact: Choose an effort supported by the destination model
Signature Field Provider-Specific
Severity: Low Behavior: Anthropic/Bedrock signatures and Gemini thought signatures are opaque provider continuation state Impact: Preserve them with the corresponding thinking or tool-call items when continuing; do not treat them as portable verification data
Anthropic Thinking Mode Selection
Severity: Low
Behavior: Explicit budgets select enabled; effort-only adaptive models select adaptive; none without a budget selects disabled
Impact: Remove an explicit budget before selecting adaptive or disabled thinking
Gemini: Only One Parameter Used
Severity: Medium
Behavior: When both effort and max_tokens are provided, max_tokens is used and effort is ignored
Impact: Effort value has no effect when max_tokens is present
Workaround: Provide only the parameter you want to use
Gemini: Model Version Differences
Severity: Medium Behavior: Gemini 2.5 is budget-based; 3.0+ supports both budgets and effort levels Impact: On Gemini 2.5, effort-only requests behave as a budget; on 3.0+ they use native effort levels
Gemini Pro: Limited Effort Levels
Severity: Low
Behavior: Pro models support only low and high effort levels
Impact: minimal behaves as low, and medium behaves as high on Pro models
Note: Non-Pro models support all four levels: minimal, low, medium, high
Complete Provider Comparison
Section titled “Complete Provider Comparison”Reasoning Model
Section titled “Reasoning Model”| Provider | Model Type | Budget Type | Min Budget | Signature Support |
|---|---|---|---|---|
| OpenAI | Effort-based | Effort-based | None | ❌ |
| Anthropic | Thinking blocks | Adaptive effort or manual budget | 1024 for manual budgets | ✅ |
| Bedrock (Anthropic) | Reasoning config | Adaptive effort or manual budget | 1024 for manual budgets | ✅ |
| Bedrock (Nova) | Reasoning config | Effort-based | None | ❌ |
| Gemini 2.5+ | Thinking config | Token budget | 1024 | ✅ |
| Gemini 3.0+ | Thinking config | Dual (budget + level) | 1024 | ✅ |
Parameter Support
Section titled “Parameter Support”| Provider | effort | max_tokens | summary | Streaming |
|---|---|---|---|---|
| OpenAI | Model-specific | Converted to effort when effort is absent | Model-specific | ✅ |
| Anthropic | Adaptive or budget-derived, by model | Manual mode | Model-specific | ✅ |
| Bedrock (Anthropic) | Adaptive or budget-derived, by model | Manual mode | Model-specific | ✅ |
| Bedrock (Nova) | ✅ (3 levels) | ⚠️ (ignored) | ❌ | ✅ |
| Gemini 2.5+ | ⚠️ (converts to budget) | ✅ | ❌ | ✅ |
| Gemini 3.0+ | ✅ (4 levels) | ✅ | ❌ | ✅ |
Troubleshooting
Section titled “Troubleshooting”Anthropic: “reasoning.max_tokens must be >= 1024”
Section titled “Anthropic: “reasoning.max_tokens must be >= 1024””Cause: Sending an explicit manual budget with max_tokens < 1024
Solution: Use reasoning.max_tokens >= 1024 for a model that supports manual
thinking, or omit the budget and use a supported effort for adaptive thinking.
// ❌ Invalid{"reasoning": {"effort": "high", "max_tokens": 500}}
// ✅ Valid{"reasoning": {"effort": "high", "max_tokens": 1024}}OpenAI: Model doesn’t support reasoning
Section titled “OpenAI: Model doesn’t support reasoning”Cause: Using an older model that doesn’t support reasoning (e.g., gpt-4-turbo)
Solution: Use OpenAI reasoning models: o4-mini, o3, o1, or the gpt-5 series. gpt-4o and gpt-4o-mini are not reasoning models and will reject the reasoning parameter.
Bedrock Nova: max_tokens parameter being ignored
Section titled “Bedrock Nova: max_tokens parameter being ignored”Expected Behavior: Bedrock Nova uses effort-based reasoning only
Solution: Provide effort parameter instead of max_tokens for Nova models
// ✅ Correct for Nova{"reasoning": {"effort": "high"}}