Skip to content

Reasoning

Using GPT-6 Astra? Switching models can require changing the endpoint and request/response handling. Use Astra with Responses for Python, JavaScript, curl, and MCP examples. Astra function tools require Responses, and none is not a supported reasoning effort.

Reasoning controls how a model allocates computation before answering. Depending on the provider and model, responses may include a reasoning summary, thinking content, or opaque continuation state. These are different response formats; they do not guarantee access to the model’s complete internal reasoning.


ProviderRequest FieldResponse FieldMin BudgetEffort LevelsStreaming
OpenAIreasoningreasoning_detailsNoneModel-specific✅
Anthropicthinking, output_config.effortContent blocks1024 tokens for explicit budgetsAdaptive or budget-derived, by model✅
Bedrock (Anthropic)thinking, output_config.effortContent blocks1024 tokens for explicit budgetsAdaptive or budget-derived, by model✅
Gemini 2.5+thinking_configthought parts1024Budget-only✅
Gemini 3.0+thinking_configthought parts1024minimal, low, medium, high + Budget✅

Add a reasoning object to your chat completions request body:

{
"model": "provider/model-name",
"messages": [...],
"reasoning": {
"effort": "high",
"max_tokens": 4096
}
}

The Responses API accepts the same reasoning object and adds an optional summary parameter:

{
"model": "provider/model-name",
"input": [...],
"reasoning": {
"effort": "high",
"max_tokens": 4096,
"summary": "detailed"
}
}
ParameterTypeDescription
effortstringReasoning intensity level
max_tokensintMaximum tokens for reasoning (budget)
ParameterTypeDescription
effortstringReasoning intensity level
max_tokensintMaximum tokens for reasoning (budget)
summarystringSummary level: brief, detailed, or json

OpenAI uses effort-based reasoning. Supply reasoning.effort directly. If you supply only reasoning.max_tokens, the gateway derives an effort level for you.

Supported effort levels depend on the selected model. Use its catalog parameter profile rather than applying one fixed list to every OpenAI model. See the OpenAI provider guide for model-specific sampling, token-limit, and tool-calling behavior.


The Anthropic adapter chooses thinking mode from the request and the recognized model ID:

RequestGateway conversion
Explicit reasoning.max_tokensManual thinking.type: "enabled" with that budget; it takes priority over effort.
Effort only, recognized adaptive-thinking modelthinking.type: "adaptive" and output_config.effort.
Effort only, Opus 4.5Native effort plus a derived manual thinking budget.
Effort only, older modelA derived manual thinking budget.
effort: "none", without a budgetthinking.type: "disabled".

Adaptive recognition covers the supported Claude Opus, Sonnet, Fable, and Mythos model IDs and their supported dated, Vertex, and Bedrock forms. It does not infer capabilities from an arbitrary deployment name or a future version. Use the selected model’s parameter profile for its supported effort choices.

Thinking content is returned through reasoning_details. Preserve returned signatures with their content when continuing a conversation; they are opaque provider state.

The adapter applies thinking constraints after translating both normalized and native request parameters. When thinking is active, it omits incompatible temperature and top-k values and removes top-p values outside the supported range. Recognized fixed-sampling models omit all three. For recognized Claude models that prohibit temperature and top-p together, temperature takes precedence when thinking is off.

An explicit manual budget must be below the output-token limit unless the provider’s supported interleaved-thinking mode permits otherwise. The gateway never raises your output-token limit to accommodate the budget. Adaptive-only models reject manual budgets; use reasoning.effort on those models. Manual thinking also rejects forced tool use (any or a named tool); choose automatic tool selection, no tool use, or disable thinking. These constraints also apply to Claude routed through Bedrock and Vertex.

Dynamic Budget Handling:

Input ValueBehavior
-1 (dynamic)Uses the minimum budget of 1024
< 1024Error
>= 1024Used as-is

For Bedrock Claude models, an explicit reasoning.max_tokens selects a manual budget. An effort-only request selects adaptive thinking for recognized adaptive model IDs, or a derived budget for older models. effort: "none" without a budget disables thinking.


Bedrock Nova models use effort-based reasoning. Supply reasoning.effort.

EffortNotes
minimal, lowNormal parameters allowed
mediumNormal parameters allowed
highmax_tokens, temperature, and top_p are not applied

Notable differences from Anthropic on Bedrock:

  • No minimum token budget constraint
  • Uses effort levels instead of token budgets
  • At high effort, conflicting sampling parameters are not sent

Gemini supports both token budgets (reasoning.max_tokens) and effort levels (reasoning.effort), depending on the model version.

Gemini VersionToken BudgetEffort LevelNotes
2.5+✅⚠️ (treated as a budget)Budget-based models
3.0+✅✅Support both budget and effort levels

Gemini Pro models support a narrower set of effort levels. When routed to a Pro model, the following adjustments are applied automatically:

EffortNon-Pro ModelsPro Models
"none"Disables thinkingDisables thinking
"minimal"minimallow
"low"lowlow
"medium"mediumhigh
"high"highhigh
ValueFieldBehavior
0max_tokensDisables reasoning
-1max_tokensDynamic budget (Gemini decides)
"none"effortDisables reasoning
// Dynamic budget - let Gemini decide
{ "reasoning": { "max_tokens": -1 } }
// Disable reasoning (either form works)
{ "reasoning": { "max_tokens": 0 } }
{ "reasoning": { "effort": "none" } }

Reasoning output is returned in the normalized reasoning_details array, the same as every other provider.


Two Reasoning Methods: Effort vs. Max Tokens

Section titled “Two Reasoning Methods: Effort vs. Max Tokens”

Providers use one of two reasoning styles. You can use a single, consistent reasoning object regardless of which one the target provider expects.

StyleProvidersRequest Field
Effort-BasedOpenAI, AWS Bedrock Novareasoning.effort
Model-DependentAnthropic/Bedrock Claude, Geminireasoning.effort or reasoning.max_tokens
Budget-BasedCoherereasoning.max_tokens

You can send effort and max_tokens together. The gateway uses whichever field is native to the target provider and translates the other for you, so you do not have to know each provider’s native format:

  • Anthropic and Gemini: an explicit max_tokens budget takes priority. An effort-only request uses the model’s adaptive/level-based path when supported, otherwise a derived budget.
  • Cohere: an explicit max_tokens budget takes priority over a budget derived from effort.
  • Effort-based providers (OpenAI, Bedrock Nova): if effort is present it is used; otherwise an effort level is derived from max_tokens.

Omitting both fields leaves reasoning to the provider’s default behavior.


Different providers enforce different minimum reasoning budgets:

ProviderMinimum Budget
Anthropic1024 for explicit manual budgets
Bedrock Anthropic1024 for explicit manual budgets
Bedrock Nova1
Cohere1
Gemini1024

Requests below a provider’s minimum budget are clamped up to that minimum, except where a hard error applies (see the Anthropic constraint above).


You can always send the same reasoning object; the gateway applies it to the target provider for you.

Effort on a budget-based provider (Anthropic) - works even though Anthropic is budget-based:

{
"model": "anthropic/claude-3-5-sonnet",
"messages": [{"role": "user", "content": "..."}],
"reasoning": {"effort": "high"}
}

Budget on an effort-based provider (Bedrock Nova) - works even though Nova is effort-based:

{
"model": "bedrock/us.amazon.nova-pro-v1:0",
"messages": [{"role": "user", "content": "..."}],
"reasoning": {"max_tokens": 2000}
}

Both fields provided - the field native to the target provider wins. For Anthropic (budget-based), max_tokens is used and effort is ignored:

{
"model": "anthropic/claude-3-5-sonnet",
"messages": [{"role": "user", "content": "..."}],
"reasoning": {"effort": "medium", "max_tokens": 2500}
}

For Chat Completions, supported reasoning output is represented in a normalized reasoning_details array:

{
"choices": [{
"message": {
"role": "assistant",
"content": "Final response text",
"reasoning_details": [
{
"index": 0,
"type": "text",
"text": "Step-by-step reasoning content...",
"signature": "optional_signature_for_verification"
}
]
}
}]
}
FieldTypeDescriptionPresent In
indexintPosition in reasoning sequenceAll
typestringContent type (text, encrypted, summary)All
textstringReasoning contentChat Completions
summarystringReasoning summaryResponses API
signaturestringOpaque provider continuation state to preserve with its contentAnthropic, Bedrock
Reasoning TypeWhen UsedSource
reasoning.textDirect thinking/reasoning contentAnthropic, Gemini, Bedrock
reasoning.encryptedSignature-verified reasoningAnthropic, Bedrock Nova
reasoning.summarySummarized reasoning (Responses API)All providers

Responses requests use native Responses events and output items. They do not use the Chat choices[].delta.reasoning_details shape shown below. Keep the returned reasoning items, encrypted content, signatures, and function-call items together when continuing with the same provider and model.

For translated Anthropic, Gemini, and Chat-compatible Responses streams, the gateway preserves item identifiers, output indexes, argument deltas, and final output snapshots. Treat response.completed, response.incomplete, and response.failed as different outcomes; do not infer success from text deltas or a closed connection. See Responses events and continuation.

ProviderReasoning EventSignature Event
OpenAIreasoning (top-level)N/A
Anthropicthinking_deltasignature_delta
Bedrockthinking_deltasignature_delta
Geminithought (in content)thought_signature
// Stream events
event: content_block_start
data: {"type": "content_block_start", "content_block": {"type": "thinking"}}
event: content_block_delta
data: {"type": "content_block_delta", "delta": {"type": "thinking_delta", "thinking": "Let me"}}
event: content_block_delta
data: {"type": "content_block_delta", "delta": {"type": "thinking_delta", "thinking": " analyze..."}}
event: content_block_delta
data: {"type": "content_block_delta", "delta": {"type": "signature_delta", "signature": "EqoB..."}}
event: content_block_stop
data: {"type": "content_block_stop"}
// Thinking delta
{
"choices": [{
"delta": {
"reasoning_details": [{
"index": 0,
"type": "text",
"text": "Let me analyze..."
}]
}
}]
}
// Signature delta
{
"choices": [{
"delta": {
"reasoning_details": [{
"index": 0,
"signature": "EqoB..."
}]
}
}]
}

Minimum Budget (Anthropic/Bedrock)

Severity: High Behavior: An explicit manual reasoning.max_tokens must be >= 1024 Impact: Requests with lower values fail with error Workaround: Use a supported manual budget or an effort-only adaptive request

Dynamic Budget Not Supported

Severity: Medium Behavior: reasoning.max_tokens = -1 converted to 1024 Impact: The -1 sentinel does not select adaptive thinking Workaround: Use effort only on a supported adaptive model, or a manual budget

Effort Level Normalization

Severity: Low Behavior: Anthropic converts minimal to low; other adapters use model-specific mappings Impact: Choose an effort supported by the destination model

Signature Field Provider-Specific

Severity: Low Behavior: Anthropic/Bedrock signatures and Gemini thought signatures are opaque provider continuation state Impact: Preserve them with the corresponding thinking or tool-call items when continuing; do not treat them as portable verification data

Anthropic Thinking Mode Selection

Severity: Low Behavior: Explicit budgets select enabled; effort-only adaptive models select adaptive; none without a budget selects disabled Impact: Remove an explicit budget before selecting adaptive or disabled thinking

Gemini: Only One Parameter Used

Severity: Medium Behavior: When both effort and max_tokens are provided, max_tokens is used and effort is ignored Impact: Effort value has no effect when max_tokens is present Workaround: Provide only the parameter you want to use

Gemini: Model Version Differences

Severity: Medium Behavior: Gemini 2.5 is budget-based; 3.0+ supports both budgets and effort levels Impact: On Gemini 2.5, effort-only requests behave as a budget; on 3.0+ they use native effort levels

Gemini Pro: Limited Effort Levels

Severity: Low Behavior: Pro models support only low and high effort levels Impact: minimal behaves as low, and medium behaves as high on Pro models Note: Non-Pro models support all four levels: minimal, low, medium, high


ProviderModel TypeBudget TypeMin BudgetSignature Support
OpenAIEffort-basedEffort-basedNone❌
AnthropicThinking blocksAdaptive effort or manual budget1024 for manual budgets✅
Bedrock (Anthropic)Reasoning configAdaptive effort or manual budget1024 for manual budgets✅
Bedrock (Nova)Reasoning configEffort-basedNone❌
Gemini 2.5+Thinking configToken budget1024✅
Gemini 3.0+Thinking configDual (budget + level)1024✅
Providereffortmax_tokenssummaryStreaming
OpenAIModel-specificConverted to effort when effort is absentModel-specific✅
AnthropicAdaptive or budget-derived, by modelManual modeModel-specific✅
Bedrock (Anthropic)Adaptive or budget-derived, by modelManual modeModel-specific✅
Bedrock (Nova)✅ (3 levels)⚠️ (ignored)❌✅
Gemini 2.5+⚠️ (converts to budget)✅❌✅
Gemini 3.0+✅ (4 levels)✅❌✅

Anthropic: “reasoning.max_tokens must be >= 1024”

Section titled “Anthropic: “reasoning.max_tokens must be >= 1024””

Cause: Sending an explicit manual budget with max_tokens < 1024

Solution: Use reasoning.max_tokens >= 1024 for a model that supports manual thinking, or omit the budget and use a supported effort for adaptive thinking.

// ❌ Invalid
{"reasoning": {"effort": "high", "max_tokens": 500}}
// ✅ Valid
{"reasoning": {"effort": "high", "max_tokens": 1024}}

Cause: Using an older model that doesn’t support reasoning (e.g., gpt-4-turbo)

Solution: Use OpenAI reasoning models: o4-mini, o3, o1, or the gpt-5 series. gpt-4o and gpt-4o-mini are not reasoning models and will reject the reasoning parameter.

Bedrock Nova: max_tokens parameter being ignored

Section titled “Bedrock Nova: max_tokens parameter being ignored”

Expected Behavior: Bedrock Nova uses effort-based reasoning only

Solution: Provide effort parameter instead of max_tokens for Nova models

// ✅ Correct for Nova
{"reasoning": {"effort": "high"}}