Request volume, reliability, latency percentiles, and estimated model usage for the current filters.
Cost method: every provider attempt is counted. Provider reported cost is preferred, otherwise the dashboard uses the provider token rate active when it calculates the view. Historical requests without attempt rows are counted once using their final served model; request timestamps are not repriced across rate changes. Latency: percentiles use nearest-rank over sampled latency_ms values (up to 50k).
Application response-cache outcomes and latency. These metrics are separate from provider prompt-cache token usage.
Token note: provider prompt-cache read/write tokens remain under token usage and do not determine response-cache HIT, MISS, or BYPASS.
Background inference latency and retry health. Duplicate-suppressed jobs are excluded from totals; provider prompt-cache tokens remain separate.
Estimated users are unique app installations or browser clients. They are not authenticated accounts.
How traffic is distributed across endpoints, apps, providers, platforms, and models.
Estimated unique users over time and their latest known platform and country.
| app | estimated active users | requests | request error rate |
|---|
Production request state from request_lifecycle. AI outcomes stay separate from routing/client errors, authorization errors, and observability storage errors. The raw HTTP 4xx / 5xx total remains visible for reconciliation.
Click a row to open the request detail and inspect its transition history.
| received | request | app | category | phase | result | failure | provider | bytes | correlation |
|---|
| date / time | id | app | endpoint | status | latency | cache | cost | media | model | model | country | gw | err | ★ |
|---|
Open a request to inspect every image and swipe between them.
| date / time | app | device | type | rating | message | contact | request | workflow |
|---|