AI Cost Center
AI 成本中心Token-level spend attribution across providers, models and products, reconciled against the portfolio AI budget.
Demo data· $50.3K / month run-rate
Month to date
$46.2K
+11%vs last month run-rate
14.3M requests billed this month
Projected month end
$51.3K
$1.8K of headroom left
Budget utilisation
86.9%
$53.2K budget · 3 days remaining
Total tokens
41.2B
$1.22 per million tokens
Cache hit rate
47.1%
14.2B input tokens served from cache
AI cost / revenue
11.6%
Gross-margin guardrail: 15%
Daily spend by provider
Stacked provider contribution over the last 30 days
5 providers
Spend by provider
Share of recomputed monthly AI budget
Spend by product
Which products consume the inference budget
Month-to-date trajectory
Actual daily spend in the current billing month
Model ledger
Top models by attributed cost, latency and cache efficiency
Tool calls
13.9M
Reasoning tokens
3.1B
Output tokens
7.9B
Avg tokens / request
2.9K
Input tokens
30.1B
Recomputed total
$50.3K
| Model | Provider | Cost | Share | Requests | Cache hit | P95 latency |
|---|---|---|---|---|---|---|
| gpt-5 | OpenAI | $21.2K | 42.1% | 1.5M | 51.9% | 5.4s |
| gpt-5-mini | OpenAI | $12.5K | 24.8% | 7.7M | 45.5% | 1.8s |
| claude-4.5-sonnet | Anthropic | $7.5K | 15% | 325.7K | 39.1% | 3.8s |
| claude-4.5-opus | Anthropic | $3.6K | 7.3% | 21.9K | 39.4% | 8.2s |
| deepseek-v4 | DeepSeek | $2K | 4% | 2M | 43.4% | 2.9s |
| gemini-3-pro | $1.8K | 3.6% | 278K | 60.6% | 3.5s | |
| gemini-3-flash | $1.1K | 2.3% | 2.1M | 51.7% | 864ms | |
| qwen-3-max | Qwen | $501.88 | 1% | 386.6K | 35.5% | 2.4s |