Skip to content

v0.0.16 ​

🐛 Bug Fixes ​

  • Fixed double-counted streaming usage for Anthropic (issue #13):
    • Under the current Messages API, both message_start and message_delta carry cumulative usage for the whole request (message_delta repeats message_start's input/cache fields and adds the final output_tokens).
    • The old code accumulated with += as if each event were incremental, so input_tokens / cache_read_input_tokens were counted twice — typically ~50% over-billing (e.g. 4522 vs the correct 3000 on claude-3.5-sonnet).
    • Events are now merged by max of the cumulative values, and the legacy shape (delta with only output_tokens, zeroed input) still keeps message_start's input count.
  • Fixed cache-read semantic mismatch causing under-billing: Claude's input_tokens is disjoint from cache_read/cache_creation_input_tokens, yet the shared billing formula assumes OpenAI's "cached ⊆ prompt" (input×(prompt−cached)). Cache reads were therefore subtracted twice — charged at read price while also dropping one input-price charge. Reads are now folded into PromptTokens, so the formula yields the correct input×input + readPrice×read.
  • Fixed AWS Bedrock Claude channels never billing cache: neither the streaming nor the non-streaming path wrote cache_read_input_tokens into PromptTokensDetails.CachedTokens; both now do.
  • Fixed cache-creation tokens being parsed but never billed: cache_creation_input_tokens was unmarshalled yet dropped from usage; it is now folded into PromptTokens and billed at input price (actual write price is ≈1.25×input — a documented approximation, see #13).

🔧 Refactor ​

  • New relay/adaptor/anthropic/usage.go: ClaudeUsage2OpenAI (Claude→OpenAI usage normalization) and MergeClaudeUsage (max-merge of cumulative stream events).
  • The three consumer paths — native Anthropic, Vertex AI Claude, and AWS Bedrock Claude — now share these helpers, removing three divergent hand-written accumulators.

🧪 Tests ​

  • relay/adaptor/anthropic/main_test.go: the message_delta fixture was updated from the legacy zeroed-input_tokens shape to the real cumulative shape so the double-count bug can no longer hide.
  • Added TestClaudeUsage2OpenAI / TestMergeClaudeUsage behavior tests covering cumulative sequences, legacy-shape compatibility, read/creation folding, and the cached ⊆ prompt invariant.

⚠️ Upgrade Notes ​

  • Zero database migration and zero configuration changes.
  • Billing behavior: Anthropic channels no longer double count and charge correctly; prompt_tokens in usage logs now includes folded cache read/creation, consistent with OpenAI semantics (historical logs are not retroactively adjusted).
  • Verification: go build ./...; go test ./relay/adaptor/anthropic/ ./relay/billing/ratio/ ./model/ ./controller/ ./middleware/.