Skip to main content
In a long session, most of the token bill is not generation — it is re-sending the same history every turn. Provider-side prefix caching bills the unchanged prefix at cache rates, provided two things hold: the prefix is byte-for-byte identical, and the request lands on the same cache shard. The sidecar has a deliberate design for both.

Routing: keep one session on one shard

OpenAI-style APIs shard by prompt_cache_key. The sidecar derives it from the session id (length-clamped) and injects it into the outbound payload through a hook — same session, same key, same shard.
Behind a gateway or relay, the infrastructure also has to route the session to the same upstream. Every request carries x-session-affinity and X-Session-Id.
Without these, requests land on random shards and an identical prefix is recomputed in full anyway — and the symptom is silent: the bill notices before you do.

The long-retention switch

Standard TTL by default; the switch opts into each provider’s long cache tier. The compat layer gates it and degrades automatically to the standard tier when a model does not support it, instead of erroring.

A compaction summary request must be shaped like a chat request

The most hidden and most valuable rule in this design. Compaction itself sends a model request to produce the summary. If that request’s system, tools, and parameters differ from the chat path, the server-side prefix cache hits nothing — compaction becomes a full-price recompute. This was measured in production: a summary request hit only 128 tokens (one cache block) before diverging. So the summary request is deliberately assembled identically to chat:
1

Same system (messages[0]), same tools, same affinity headers and prompt_cache_key

The only difference is one summary instruction appended at the tail — so the prefix cache can hit the entire conversation.
2

Same thinking parameters

Some endpoints fold parameters into the cache key (Anthropic keys on the thinking budget derived from max_tokens), so mismatched parameters break an otherwise identical prefix.
3

Same long-retention switch

Shares PI_CACHE_RETENTION with the chat path.

Payload fingerprints: locating a divergence

The two assembly paths (chat via the convertToLlm chain, summaries via their own assembly) build request bodies separately, so “looks the same, differs by bytes” is possible. Each of the three likeliest divergence points gets a short hash plus a size — tools (count + fingerprint), the first system message (length + fingerprint), and the message count — logged together with reasoning / max_tokens and friends at event level. Comparing two log lines pins down which segment diverged. No more guessing.

Counting hits: two standards

A turn counts as a miss when input ≥ 2,000 and ≥ 5% of the prompt. The flaw: it cannot tell “new content this turn” (tool results are new by definition) from “the prefix was recomputed” — about 90% of reports were false positives. Kept only as the legacy panel denominator.
Recomputed tokens for a turn = min(previous prompt, current prompt) − cacheRead, floored at 1,024. That is what the turn actually overpaid, with lastMiss attribution: idle past TTL, or a model switch.
The hit rate itself is cacheRead / (input + cacheRead + cacheWrite), and the panel shows a recent 10-request window — cumulative numbers get diluted by cold starts and rebuilds; only a short window shows steady state.
Two deliberate exemptions: a cold-start first turn has no prefix that “should have hit”, so it does not count; and a session whose provider never reported cache activity is excluded from recompute accounting entirely — otherwise sessions on providers that do not report caching would all register as misses.

Compaction and caching pull against each other

Compaction replaces history with a summary, so the prefix necessarily changes and the first turn after compaction is expected to recompute in full. Statistics track those as rebuilds, separate from real misses. Compaction’s own invariants (see the header note in apps/sidecar/pi-agent/src/agent/context.ts):
  • It happens only at turn boundaries; the summary replaces the entire history with no original tail kept;
  • The checkpoint row persists to JSONL while full message rows are retained, so UI history is unaffected;
  • If summary generation fails, a fresh_window fallback kicks in — no model request, a fixed rollover marker is loaded, and the session continues.
A known approximation: the transcript cannot distinguish “first turn after restoring a session post-restart” from “first turn after compaction”, so a restore right after a checkpoint is counted as a rebuild — off by at most one turn, in exchange for zero new persistence. A deliberate trade-off recorded in the code comments, not a bug.

Next

With the cache accounting settled, goal mode keeps a ledger of its own.