summaryrefslogtreecommitdiffhomepage
path: root/packages/transport-http/src/logic.ts
diff options
context:
space:
mode:
authorAdam Malczewski <[email protected]>2026-06-11 14:11:13 +0900
committerAdam Malczewski <[email protected]>2026-06-11 14:11:13 +0900
commit7ffb6b28f5b6bdbfc53ebed94fc68af557612189 (patch)
treee66d9ea9d326ef771cc473d81ca5716ff78b08a8 /packages/transport-http/src/logic.ts
parent763e5fb1c7fbfb4c7bbd43ffb935e42e5f5b5a42 (diff)
downloaddispatch-7ffb6b28f5b6bdbfc53ebed94fc68af557612189.tar.gz
dispatch-7ffb6b28f5b6bdbfc53ebed94fc68af557612189.zip
fix(cache-warming): accurate cache rate + expectedCacheRate (retention) metric
The Claude cache % read 100% whenever anything was cached, because the metric's denominator (inputTokens) excluded cached tokens on Anthropic. Fixed upstream in ../claude/provider-anthropic (inputTokens = total prompt); this commit adds the companion retention metric and exposes it: - transport-contract: WarmResponse += expectedCacheRate - transport-http: POST /chat/warm returns expectedCacheRate = cacheRead/(cacheRead+cacheWrite) - cache-warming: computeExpectedCacheRate + a per-conversation 'cache retention' surface stat - handoff: documents the fix + cache-rate vs expected-cache (cross-turn) for the FE Live-verified vs claude haiku: real turn cache rate 61% (was inflated 100%); warm within TTL expectedCacheRate=100%, after expiry=0%.
Diffstat (limited to 'packages/transport-http/src/logic.ts')
-rw-r--r--packages/transport-http/src/logic.ts9
1 files changed, 9 insertions, 0 deletions
diff --git a/packages/transport-http/src/logic.ts b/packages/transport-http/src/logic.ts
index bb827e2..bddedf0 100644
--- a/packages/transport-http/src/logic.ts
+++ b/packages/transport-http/src/logic.ts
@@ -113,3 +113,12 @@ export function computeCachePct(inputTokens: number, cacheReadTokens: number): n
if (inputTokens <= 0) return 0;
return Math.round(Math.max(0, Math.min(1, cacheReadTokens / inputTokens)) * 100);
}
+
+export function computeExpectedCacheRate(
+ cacheReadTokens: number,
+ cacheWriteTokens: number,
+): number {
+ const denom = cacheReadTokens + cacheWriteTokens;
+ if (denom <= 0) return 0;
+ return Math.round((cacheReadTokens / denom) * 100);
+}