Prometheus
Overview
Section titled “Overview”DeepIntShield exposes Prometheus metrics via two methods:
- Pull-based (Scraping): Traditional
/metricsendpoint that Prometheus can scrape - Push-based (Push Gateway): Push metrics to a Prometheus Push Gateway for cluster deployments
The metrics cover HTTP transport, upstream provider calls, token usage, cost in USD, cache hits, and streaming latency. Collection runs outside the response path: request-path instrumentation and enqueueing have bounded overhead, and the work that runs after the response does not delay it. Benchmark the enabled labels and exporters with your traffic shape before setting latency objectives.
Pull-based Scraping
Section titled “Pull-based Scraping”DeepIntShield automatically exposes a /metrics endpoint when the telemetry plugin is enabled (enabled by default). No additional configuration is needed.
Prometheus Configuration
Section titled “Prometheus Configuration”Add DeepIntShield to your Prometheus prometheus.yml:
scrape_configs: - job_name: 'deepintshield' static_configs: - targets: ['deepintshield-host:8080'] scrape_interval: 15sIf DeepIntShield authentication is enabled, add basic_auth to your scrape config:
scrape_configs: - job_name: 'deepintshield' static_configs: - targets: ['deepintshield-host:8080'] scrape_interval: 15s basic_auth: username: '<admin_username>' password: '<admin_password>'Endpoint
Section titled “Endpoint”GET /metricsReturns metrics in Prometheus exposition format.
Push-based (Push Gateway)
Section titled “Push-based (Push Gateway)”For multi-node cluster deployments, the Prometheus plugin pushes metrics to a Prometheus Push Gateway. This ensures all nodes’ metrics are captured regardless of load balancer routing.
Configuration
Section titled “Configuration”| Field | Type | Required | Default | Description |
|---|---|---|---|---|
push_gateway_url | string | ✅ Yes | - | Push Gateway URL (e.g., http://pushgateway:9091) |
job_name | string | ❌ No | deepintshield | Job label for pushed metrics |
instance_id | string | ❌ No | hostname | Instance identifier for metric grouping |
push_interval | integer | ❌ No | 15 | Push interval in seconds (1-300) |
basic_auth | object | ❌ No | - | Basic auth credentials |
Basic Auth Configuration
Section titled “Basic Auth Configuration”| Field | Type | Required | Description |
|---|---|---|---|
username | string | ✅ Yes | Basic auth username |
password | string | ✅ Yes | Basic auth password |
- Navigate to Analytics → Integrations → Prometheus in the DeepIntShield dashboard
- The
/metricsendpoint is shown at the top for scraping configuration - To enable Push Gateway:
- Enter the Push Gateway URL
- Configure Job Name and Push Interval as needed
- Optionally set a custom Instance ID
- Enable Basic Authentication if required
- Toggle Enable Push Gateway on
- Click Save Prometheus Configuration
{ "plugins": [ { "name": "prometheus", "enabled": true, "config": { "push_gateway_url": "http://pushgateway:9091", "job_name": "deepintshield", "push_interval": 15 } } ]}With Basic Auth
Section titled “With Basic Auth”{ "plugins": [ { "name": "prometheus", "enabled": true, "config": { "push_gateway_url": "http://pushgateway:9091", "job_name": "deepintshield", "push_interval": 15, "instance_id": "deepintshield-node-1", "basic_auth": { "username": "admin", "password": "secret" } } } ]}Available Metrics
Section titled “Available Metrics”The following metrics are available from both the /metrics endpoint and Push Gateway.
HTTP Transport Metrics
Section titled “HTTP Transport Metrics”These metrics track all incoming HTTP requests to DeepIntShield:
| Metric | Type | Description |
|---|---|---|
http_requests_total | Counter | Total number of HTTP requests |
http_request_duration_seconds | Histogram | Duration of HTTP requests |
http_request_size_bytes | Histogram | Size of incoming HTTP requests |
http_response_size_bytes | Histogram | Size of outgoing HTTP responses |
Labels: path, method, status, plus any custom labels you configure.
Upstream Provider Metrics
Section titled “Upstream Provider Metrics”These metrics track requests forwarded to AI providers:
| Metric | Type | Description | Labels |
|---|---|---|---|
deepintshield_upstream_requests_total | Counter | Total requests forwarded to upstream providers | Base labels, custom labels |
deepintshield_success_requests_total | Counter | Total successful requests to upstream providers | Base labels, custom labels |
deepintshield_error_requests_total | Counter | Total failed requests to upstream providers | Base labels, reason, custom labels |
deepintshield_upstream_latency_seconds | Histogram | Latency of upstream provider requests | Base labels, is_success, custom labels |
deepintshield_input_tokens_total | Counter | Total input tokens sent to upstream providers | Base labels, custom labels |
deepintshield_output_tokens_total | Counter | Total output tokens received from upstream providers | Base labels, custom labels |
deepintshield_cache_hits_total | Counter | Total cache hits by type (direct/semantic) | Base labels, cache_type, custom labels |
deepintshield_cost_total | Counter | Total cost in USD for upstream provider requests | Base labels, custom labels |
Streaming Metrics
Section titled “Streaming Metrics”| Metric | Type | Description | Labels |
|---|---|---|---|
deepintshield_stream_first_token_latency_seconds | Histogram | Time from request start to first streamed token | Base labels |
deepintshield_stream_inter_token_latency_seconds | Histogram | Latency between subsequent streamed tokens | Base labels |
Base Labels
Section titled “Base Labels”All DeepIntShield metrics include these labels:
provider- AI provider name (e.g.,openai,anthropic,azure)model- Model name (e.g.,gpt-4o-mini,claude-3-sonnet)method- Request type (chat,text,embedding,speech,transcription)virtual_key_id/virtual_key_name- Virtual key identifiersteam_id/team_name- Team identifiers (if governance enabled)customer_id/customer_name- Customer identifiers (if governance enabled)routing_engines_used- Comma-separated routing engines used (routing-rule,governance,loadbalancing)routing_rule_id/routing_rule_name- Routing rule that matched the requestselected_key_id/selected_key_name- Provider key actually usednumber_of_retries- Retry countfallback_index- Fallback position (0 for the first attempt, 1 for the second, and so on)
Custom Labels
Section titled “Custom Labels”Add your own dimensions for filtering and analysis.
Configured labels
Section titled “Configured labels”- Open the dashboard and go to Settings → Observability.
- Under Prometheus Labels, enter a comma-separated list, for example
team, environment, organization, project. - Save. Every configured label becomes a dimension on the metrics above.
Dynamic label values
Section titled “Dynamic label values”Supply the value for a configured label per request with x-deepintshield-prom-* headers:
curl -X POST https://app.deepintshield.com/v1/chat/completions \ -H "Content-Type: application/json" \ -H "x-deepintshield-vk: sk-ds-your-virtual-key" \ -H "x-deepintshield-prom-team: engineering" \ -H "x-deepintshield-prom-environment: production" \ -H "x-deepintshield-prom-organization: my-org" \ -H "x-deepintshield-prom-project: my-project" \ -d '{ "model": "gpt-4o-mini", "messages": [{"role": "user", "content": "Hello!"}] }'Header format: prefix x-deepintshield-prom-, followed by the label name; the header value becomes the label value.
Example Queries
Section titled “Example Queries”Success rate by provider
Section titled “Success rate by provider”rate(deepintshield_success_requests_total[5m]) /rate(deepintshield_upstream_requests_total[5m]) * 100Token usage
Section titled “Token usage”# Input tokens per minute by modelincrease(deepintshield_input_tokens_total[1m])
# Output tokens per minute by modelincrease(deepintshield_output_tokens_total[1m])
# Token efficiency (output/input ratio)rate(deepintshield_output_tokens_total[5m]) /rate(deepintshield_input_tokens_total[5m])Cost tracking
Section titled “Cost tracking”# Cost per second by providersum by (provider) (rate(deepintshield_cost_total[1m]))
# Daily cost estimatesum by (provider) (increase(deepintshield_cost_total[1d]))
# Cost per request by provider and modelsum by (provider, model) (rate(deepintshield_cost_total[5m])) /sum by (provider, model) (rate(deepintshield_upstream_requests_total[5m]))Cache performance
Section titled “Cache performance”# Cache hit rate by typerate(deepintshield_cache_hits_total[5m]) /rate(deepintshield_upstream_requests_total[5m]) * 100
# Direct vs semantic cache hitssum by (cache_type) (rate(deepintshield_cache_hits_total[5m]))Error rate
Section titled “Error rate”# Error rate by providerrate(deepintshield_error_requests_total[5m]) /rate(deepintshield_upstream_requests_total[5m]) * 100
# Errors by modelsum by (model) (rate(deepintshield_error_requests_total[5m]))Alerting Examples
Section titled “Alerting Examples”High error rate:
- alert: DeepIntShieldHighErrorRate expr: sum by (provider) (rate(deepintshield_error_requests_total[5m])) / sum by (provider) (rate(deepintshield_upstream_requests_total[5m])) > 0.05 for: 2m labels: severity: warning annotations: summary: "High error rate detected for provider {{ $labels.provider }} ({{ $value | humanizePercentage }})"High cost:
- alert: DeepIntShieldHighCosts expr: sum by (provider) (increase(deepintshield_cost_total[1d])) > 100 # $100/day threshold for: 10m labels: severity: warning annotations: summary: "Daily cost for provider {{ $labels.provider }} exceeds $100 ({{ $value | printf \"%.2f\" }})"Low cache hit rate:
- alert: DeepIntShieldLowCacheHitRate expr: sum by (provider) (rate(deepintshield_cache_hits_total[15m])) / sum by (provider) (rate(deepintshield_upstream_requests_total[15m])) < 0.1 for: 5m labels: severity: info annotations: summary: "Cache hit rate for provider {{ $labels.provider }} below 10% ({{ $value | humanizePercentage }})"Push Gateway Setup
Section titled “Push Gateway Setup”If you don’t have a Push Gateway running, deploy one:
Docker
Section titled “Docker”docker run -d -p 9091:9091 prom/pushgatewayKubernetes (Helm)
Section titled “Kubernetes (Helm)”helm repo add prometheus-community https://prometheus-community.github.io/helm-chartshelm install pushgateway prometheus-community/prometheus-pushgatewayConfigure Prometheus to Scrape Push Gateway
Section titled “Configure Prometheus to Scrape Push Gateway”Add to your prometheus.yml:
scrape_configs: - job_name: 'pushgateway' honor_labels: true static_configs: - targets: ['pushgateway:9091']Pull vs Push: When to Use Each
Section titled “Pull vs Push: When to Use Each”| Scenario | Recommended Method |
|---|---|
| Single DeepIntShield instance | Pull (scraping) |
| Multiple instances, direct access | Pull (scraping) |
| Multiple instances behind load balancer | Push (Push Gateway) |
| Kubernetes with service mesh | Pull or Push |
| Serverless / ephemeral instances | Push (Push Gateway) |
Why Push for Clusters?
Section titled “Why Push for Clusters?”When multiple DeepIntShield instances run behind a load balancer:
- Scraping randomness: Each scrape may hit different nodes, missing metrics from others
- Instance tracking: Push Gateway properly tracks per-instance metrics via
instancelabel - Aggregation: Downstream tools (Grafana, Datadog) can aggregate across all instances
Troubleshooting
Section titled “Troubleshooting”Push Gateway Connection Failed
Section titled “Push Gateway Connection Failed”failed to push metrics to push gateway: connection refused- Verify the Push Gateway URL is correct and reachable from DeepIntShield
- Check firewall rules between DeepIntShield and Push Gateway
- Ensure Push Gateway is running:
curl http://pushgateway:9091/metrics
Metrics Not Appearing
Section titled “Metrics Not Appearing”- Verify the telemetry plugin is enabled (required for metrics collection)
- Check DeepIntShield logs for push errors
- Verify Prometheus is scraping the Push Gateway with
honor_labels: true
Authentication Failed
Section titled “Authentication Failed”- Double-check username and password
- Ensure basic auth is configured on the Push Gateway side
- Check for special characters that may need escaping
Next Steps
Section titled “Next Steps”- Built-in Observability - Request and response logging
- OpenTelemetry - Distributed tracing and platform integrations