Skip to content

Prometheus

DeepIntShield exposes Prometheus metrics via two methods:

  1. Pull-based (Scraping): Traditional /metrics endpoint that Prometheus can scrape
  2. Push-based (Push Gateway): Push metrics to a Prometheus Push Gateway for cluster deployments

The metrics cover HTTP transport, upstream provider calls, token usage, cost in USD, cache hits, and streaming latency. Collection runs outside the response path: request-path instrumentation and enqueueing have bounded overhead, and the work that runs after the response does not delay it. Benchmark the enabled labels and exporters with your traffic shape before setting latency objectives.


DeepIntShield automatically exposes a /metrics endpoint when the telemetry plugin is enabled (enabled by default). No additional configuration is needed.

Add DeepIntShield to your Prometheus prometheus.yml:

scrape_configs:
- job_name: 'deepintshield'
static_configs:
- targets: ['deepintshield-host:8080']
scrape_interval: 15s

If DeepIntShield authentication is enabled, add basic_auth to your scrape config:

scrape_configs:
- job_name: 'deepintshield'
static_configs:
- targets: ['deepintshield-host:8080']
scrape_interval: 15s
basic_auth:
username: '<admin_username>'
password: '<admin_password>'
GET /metrics

Returns metrics in Prometheus exposition format.


For multi-node cluster deployments, the Prometheus plugin pushes metrics to a Prometheus Push Gateway. This ensures all nodes’ metrics are captured regardless of load balancer routing.

FieldTypeRequiredDefaultDescription
push_gateway_urlstring✅ Yes-Push Gateway URL (e.g., http://pushgateway:9091)
job_namestring❌ NodeepintshieldJob label for pushed metrics
instance_idstring❌ NohostnameInstance identifier for metric grouping
push_intervalinteger❌ No15Push interval in seconds (1-300)
basic_authobject❌ No-Basic auth credentials
FieldTypeRequiredDescription
usernamestring✅ YesBasic auth username
passwordstring✅ YesBasic auth password

  1. Navigate to Analytics → Integrations → Prometheus in the DeepIntShield dashboard
  2. The /metrics endpoint is shown at the top for scraping configuration
  3. To enable Push Gateway:
    • Enter the Push Gateway URL
    • Configure Job Name and Push Interval as needed
    • Optionally set a custom Instance ID
    • Enable Basic Authentication if required
    • Toggle Enable Push Gateway on
    • Click Save Prometheus Configuration

The following metrics are available from both the /metrics endpoint and Push Gateway.

These metrics track all incoming HTTP requests to DeepIntShield:

MetricTypeDescription
http_requests_totalCounterTotal number of HTTP requests
http_request_duration_secondsHistogramDuration of HTTP requests
http_request_size_bytesHistogramSize of incoming HTTP requests
http_response_size_bytesHistogramSize of outgoing HTTP responses

Labels: path, method, status, plus any custom labels you configure.

These metrics track requests forwarded to AI providers:

MetricTypeDescriptionLabels
deepintshield_upstream_requests_totalCounterTotal requests forwarded to upstream providersBase labels, custom labels
deepintshield_success_requests_totalCounterTotal successful requests to upstream providersBase labels, custom labels
deepintshield_error_requests_totalCounterTotal failed requests to upstream providersBase labels, reason, custom labels
deepintshield_upstream_latency_secondsHistogramLatency of upstream provider requestsBase labels, is_success, custom labels
deepintshield_input_tokens_totalCounterTotal input tokens sent to upstream providersBase labels, custom labels
deepintshield_output_tokens_totalCounterTotal output tokens received from upstream providersBase labels, custom labels
deepintshield_cache_hits_totalCounterTotal cache hits by type (direct/semantic)Base labels, cache_type, custom labels
deepintshield_cost_totalCounterTotal cost in USD for upstream provider requestsBase labels, custom labels
MetricTypeDescriptionLabels
deepintshield_stream_first_token_latency_secondsHistogramTime from request start to first streamed tokenBase labels
deepintshield_stream_inter_token_latency_secondsHistogramLatency between subsequent streamed tokensBase labels

All DeepIntShield metrics include these labels:

  • provider - AI provider name (e.g., openai, anthropic, azure)
  • model - Model name (e.g., gpt-4o-mini, claude-3-sonnet)
  • method - Request type (chat, text, embedding, speech, transcription)
  • virtual_key_id / virtual_key_name - Virtual key identifiers
  • team_id / team_name - Team identifiers (if governance enabled)
  • customer_id / customer_name - Customer identifiers (if governance enabled)
  • routing_engines_used - Comma-separated routing engines used (routing-rule, governance, loadbalancing)
  • routing_rule_id / routing_rule_name - Routing rule that matched the request
  • selected_key_id / selected_key_name - Provider key actually used
  • number_of_retries - Retry count
  • fallback_index - Fallback position (0 for the first attempt, 1 for the second, and so on)

Add your own dimensions for filtering and analysis.

  1. Open the dashboard and go to Settings → Observability.
  2. Under Prometheus Labels, enter a comma-separated list, for example team, environment, organization, project.
  3. Save. Every configured label becomes a dimension on the metrics above.

Supply the value for a configured label per request with x-deepintshield-prom-* headers:

Terminal window
curl -X POST https://app.deepintshield.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "x-deepintshield-vk: sk-ds-your-virtual-key" \
-H "x-deepintshield-prom-team: engineering" \
-H "x-deepintshield-prom-environment: production" \
-H "x-deepintshield-prom-organization: my-org" \
-H "x-deepintshield-prom-project: my-project" \
-d '{
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Hello!"}]
}'

Header format: prefix x-deepintshield-prom-, followed by the label name; the header value becomes the label value.


rate(deepintshield_success_requests_total[5m]) /
rate(deepintshield_upstream_requests_total[5m]) * 100
# Input tokens per minute by model
increase(deepintshield_input_tokens_total[1m])
# Output tokens per minute by model
increase(deepintshield_output_tokens_total[1m])
# Token efficiency (output/input ratio)
rate(deepintshield_output_tokens_total[5m]) /
rate(deepintshield_input_tokens_total[5m])
# Cost per second by provider
sum by (provider) (rate(deepintshield_cost_total[1m]))
# Daily cost estimate
sum by (provider) (increase(deepintshield_cost_total[1d]))
# Cost per request by provider and model
sum by (provider, model) (rate(deepintshield_cost_total[5m])) /
sum by (provider, model) (rate(deepintshield_upstream_requests_total[5m]))
# Cache hit rate by type
rate(deepintshield_cache_hits_total[5m]) /
rate(deepintshield_upstream_requests_total[5m]) * 100
# Direct vs semantic cache hits
sum by (cache_type) (rate(deepintshield_cache_hits_total[5m]))
# Error rate by provider
rate(deepintshield_error_requests_total[5m]) /
rate(deepintshield_upstream_requests_total[5m]) * 100
# Errors by model
sum by (model) (rate(deepintshield_error_requests_total[5m]))

High error rate:

- alert: DeepIntShieldHighErrorRate
expr: sum by (provider) (rate(deepintshield_error_requests_total[5m])) / sum by (provider) (rate(deepintshield_upstream_requests_total[5m])) > 0.05
for: 2m
labels:
severity: warning
annotations:
summary: "High error rate detected for provider {{ $labels.provider }} ({{ $value | humanizePercentage }})"

High cost:

- alert: DeepIntShieldHighCosts
expr: sum by (provider) (increase(deepintshield_cost_total[1d])) > 100 # $100/day threshold
for: 10m
labels:
severity: warning
annotations:
summary: "Daily cost for provider {{ $labels.provider }} exceeds $100 ({{ $value | printf \"%.2f\" }})"

Low cache hit rate:

- alert: DeepIntShieldLowCacheHitRate
expr: sum by (provider) (rate(deepintshield_cache_hits_total[15m])) / sum by (provider) (rate(deepintshield_upstream_requests_total[15m])) < 0.1
for: 5m
labels:
severity: info
annotations:
summary: "Cache hit rate for provider {{ $labels.provider }} below 10% ({{ $value | humanizePercentage }})"

If you don’t have a Push Gateway running, deploy one:

Terminal window
docker run -d -p 9091:9091 prom/pushgateway
Terminal window
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm install pushgateway prometheus-community/prometheus-pushgateway

Configure Prometheus to Scrape Push Gateway

Section titled “Configure Prometheus to Scrape Push Gateway”

Add to your prometheus.yml:

scrape_configs:
- job_name: 'pushgateway'
honor_labels: true
static_configs:
- targets: ['pushgateway:9091']

ScenarioRecommended Method
Single DeepIntShield instancePull (scraping)
Multiple instances, direct accessPull (scraping)
Multiple instances behind load balancerPush (Push Gateway)
Kubernetes with service meshPull or Push
Serverless / ephemeral instancesPush (Push Gateway)

When multiple DeepIntShield instances run behind a load balancer:

  1. Scraping randomness: Each scrape may hit different nodes, missing metrics from others
  2. Instance tracking: Push Gateway properly tracks per-instance metrics via instance label
  3. Aggregation: Downstream tools (Grafana, Datadog) can aggregate across all instances

failed to push metrics to push gateway: connection refused
  • Verify the Push Gateway URL is correct and reachable from DeepIntShield
  • Check firewall rules between DeepIntShield and Push Gateway
  • Ensure Push Gateway is running: curl http://pushgateway:9091/metrics
  • Verify the telemetry plugin is enabled (required for metrics collection)
  • Check DeepIntShield logs for push errors
  • Verify Prometheus is scraping the Push Gateway with honor_labels: true
  • Double-check username and password
  • Ensure basic auth is configured on the Push Gateway side
  • Check for special characters that may need escaping