Executive view
The bill, at a glance
This is the most defensible estimate for paying for the observed workload through the APIs—not a subscription-price comparison.
Provider share
Caching changed the economics
Token profile
Why the volume looks enormous
Long-running agents repeatedly replay histories, repository context, and tool results. Almost all of that input was served from cache.
OpenAI
Codex usage profile
89,210 unique usage events after deduplication
Anthropic
Claude usage profile
53,080 unique API-like assistant requests after deduplication
Cost concentration
Which models drove the bill
Opus-class and Fable usage dominate Anthropic; GPT-5.6 Sol dominates OpenAI.
OpenAI model mix
Anthropic model mix
| Model | Input | Cached | Output | API estimate | No-cache control |
|---|---|---|---|---|---|
| GPT-5.6 Sol | 6,695.2M | 6,523.8M | 18.46M | $4,724.64 | $34,471.35 |
| GPT-5.5 | 2,248.9M | 2,129.4M | 8.02M | $1,902.70 | $11,485.00 |
| GPT-5.6 Terra | 2,794.8M | 2,737.7M | 3.89M | $892.52 | $7,102.79 |
| GPT-5.6 Luna | 79.3M | 74.1M | 0.07M | $16.48 | $101.68 |
| GPT-5.4 mini | 5.4M | 4.6M | 0.04M | $1.14 | $4.23 |
| Total | 11,823.7M | 11,469.5M | 30.47M | $7,537.47 | $53,165.06 |
| Model | Requests | Cache reads | Output | API estimate | No-cache control |
|---|---|---|---|---|---|
| Claude Fable 5 | 10,255 | 3,415.6M | 8.48M | $5,075.43 | $35,208.76 |
| Claude Opus 4.8 | 19,788 | 6,027.0M | 21.34M | $4,558.41 | $31,186.44 |
| Claude Opus 5 | 12,821 | 3,191.1M | 12.13M | $2,249.08 | $16,433.80 |
| Claude Sonnet 4.6 | 8,411 | 684.7M | 5.56M | $469.33 | $2,228.26 |
| Claude Haiku 4.5 | 1,175 | 65.7M | 0.45M | $16.14 | $72.49 |
| Claude Sonnet 5 · promo | 401 | 34.0M | 0.14M | $13.48 | $71.96 |
| Synthetic / non-billable | 229 | 0 | 0 | $0.00 | $0.00 |
| Total | 53,080 | 13,418.1M | 48.09M | $12,381.87 | $85,201.71 |
Attribution
Largest identifiable workloads
Names were inferred only from invocation-folder paths. Prompt and response text was not inspected.
OpenAI workloads
Anthropic workloads
Rate card
Pricing used in the estimate
USD per million tokens (MTok), using standard global, non-batch API rates.
OpenAI model rates
| Model | Input | Cached input | Output |
|---|---|---|---|
| GPT-5.6 Sol | $5.00 | $0.50 | $30.00 |
| GPT-5.6 Terra | $2.50 | $0.25 | $15.00 |
| GPT-5.6 Luna | $1.00 | $0.10 | $6.00 |
| GPT-5.5 | $5.00 | $0.50 | $30.00 |
| GPT-5.4 mini | $0.75 | $0.075 | $4.50 |
Anthropic model rates
| Model | Base input | 5m write | 1h write | Cache hit | Output |
|---|---|---|---|---|---|
| Claude Fable 5 | $10.00 | $12.50 | $20.00 | $1.00 | $50.00 |
| Claude Opus 5 / 4.8 | $5.00 | $6.25 | $10.00 | $0.50 | $25.00 |
| Claude Sonnet 4.6 | $3.00 | $3.75 | $6.00 | $0.30 | $15.00 |
| Claude Haiku 4.5 | $1.00 | $1.25 | $2.00 | $0.10 | $5.00 |
| Claude Sonnet 5 · through Aug 31 | $2.00 | $2.50 | $4.00 | $0.20 | $10.00 |
| Claude Sonnet 5 · from Sep 1 | $3.00 | $3.75 | $6.00 | $0.30 | $15.00 |
Action plan
Four practical cost controls
Use high as the normal ceiling
Reserve xhigh for bounded, high-value work. Historical Sol/xhigh usage
alone represents $4,054 at API-equivalent prices.
Protect cacheable prompt prefixes
Stable histories and workspace prefixes are the difference between roughly $20K and the $138K no-cache control.
Gate Fable and Opus-class work
Fable 5 plus Opus 4.8/5 account for approximately $11,883 of the $12,382 Anthropic estimate.
Install a monthly API-style meter
Record model, input, cache writes, cache reads, output, effort, and invocation so future reports become a repeatable budget control.
Audit trail
Method, quality, and exclusions
Only usage metadata was parsed. No prompt, response, email, document, or tool-result content was included.
Method and data quality
- OpenAI: 3,454 candidate rollout files inventoried; 520 recently modified files could contain in-window events.
- OpenAI parser retained 89,210 unique cumulative events, removed 230 repeats, and ignored 22 malformed JSONL lines.
- Anthropic: all 795 Claude project JSONL files were streamed; 53,080 unique requests remained after global message-ID deduplication.
- OpenAI long-context premiums were applied to 450 requests above 272K input tokens.
- Anthropic base input, five-minute writes, one-hour writes, cache reads, and output were priced separately.
- Three Claude messages had TTL subtotals exceeding the top-level cache-creation field by 274,418 tokens (0.12% of cache writes). TTL subtotals were used; the conservative impact is under $3.
Uncertainty and exclusions
- Local usage logs may not match provider invoice aggregation around retries or product-side orchestration.
- Deleted, off-device, or otherwise unavailable session logs are not included.
- US-only Anthropic inference would add 10% to the Anthropic portion.
- Batch discounts were not assumed because these were interactive agent workloads.
- Taxes, marketplace markups, separately billed tools, storage, containers, hosted search, images, and fast-mode premiums are excluded.