v0.0.16
🐛 Bug Fixes
- Fixed double-counted streaming usage for Anthropic (issue #13):
- Under the current Messages API, both
message_startandmessage_deltacarry cumulative usage for the whole request (message_deltarepeatsmessage_start's input/cache fields and adds the finaloutput_tokens). - The old code accumulated with
+=as if each event were incremental, soinput_tokens/cache_read_input_tokenswere counted twice — typically ~50% over-billing (e.g. 4522 vs the correct 3000 on claude-3.5-sonnet). - Events are now merged by max of the cumulative values, and the legacy shape (delta with only
output_tokens, zeroed input) still keepsmessage_start's input count.
- Under the current Messages API, both
- Fixed cache-read semantic mismatch causing under-billing: Claude's
input_tokensis disjoint fromcache_read/cache_creation_input_tokens, yet the shared billing formula assumes OpenAI's "cached ⊆ prompt" (input×(prompt−cached)). Cache reads were therefore subtracted twice — charged at read price while also dropping one input-price charge. Reads are now folded intoPromptTokens, so the formula yields the correctinput×input + readPrice×read. - Fixed AWS Bedrock Claude channels never billing cache: neither the streaming nor the non-streaming path wrote
cache_read_input_tokensintoPromptTokensDetails.CachedTokens; both now do. - Fixed cache-creation tokens being parsed but never billed:
cache_creation_input_tokenswas unmarshalled yet dropped from usage; it is now folded intoPromptTokensand billed at input price (actual write price is ≈1.25×input — a documented approximation, see #13).
🔧 Refactor
- New
relay/adaptor/anthropic/usage.go:ClaudeUsage2OpenAI(Claude→OpenAI usage normalization) andMergeClaudeUsage(max-merge of cumulative stream events). - The three consumer paths — native Anthropic, Vertex AI Claude, and AWS Bedrock Claude — now share these helpers, removing three divergent hand-written accumulators.
🧪 Tests
relay/adaptor/anthropic/main_test.go: themessage_deltafixture was updated from the legacy zeroed-input_tokensshape to the real cumulative shape so the double-count bug can no longer hide.- Added
TestClaudeUsage2OpenAI/TestMergeClaudeUsagebehavior tests covering cumulative sequences, legacy-shape compatibility, read/creation folding, and thecached ⊆ promptinvariant.
⚠️ Upgrade Notes
- Zero database migration and zero configuration changes.
- Billing behavior: Anthropic channels no longer double count and charge correctly;
prompt_tokensin usage logs now includes folded cache read/creation, consistent with OpenAI semantics (historical logs are not retroactively adjusted). - Verification:
go build ./...;go test ./relay/adaptor/anthropic/ ./relay/billing/ratio/ ./model/ ./controller/ ./middleware/.