Routing: keep one session on one shard
OpenAI-style APIs shard by
prompt_cache_key. The sidecar derives it from the
session id (length-clamped) and injects it into the outbound payload through a hook —
same session, same key, same shard.Behind a gateway or relay, the infrastructure also has to route the session to the
same upstream. Every request carries
x-session-affinity and X-Session-Id.The long-retention switch
A compaction summary request must be shaped like a chat request
The most hidden and most valuable rule in this design. Compaction itself sends a model request to produce the summary. If that request’s system, tools, and parameters differ from the chat path, the server-side prefix cache hits nothing — compaction becomes a full-price recompute. This was measured in production: a summary request hit only 128 tokens (one cache block) before diverging. So the summary request is deliberately assembled identically to chat:1
Same system (messages[0]), same tools, same affinity headers and prompt_cache_key
The only difference is one summary instruction appended at the tail — so the prefix
cache can hit the entire conversation.
2
Same thinking parameters
Some endpoints fold parameters into the cache key (Anthropic keys on the thinking
budget derived from max_tokens), so mismatched parameters break an otherwise
identical prefix.
3
Same long-retention switch
Shares
PI_CACHE_RETENTION with the chat path.Payload fingerprints: locating a divergence
The two assembly paths (chat via the convertToLlm chain, summaries via their own assembly) build request bodies separately, so “looks the same, differs by bytes” is possible. Each of the three likeliest divergence points gets a short hash plus a size — tools (count + fingerprint), the first system message (length + fingerprint), and the message count — logged together with reasoning / max_tokens and friends at event level. Comparing two log lines pins down which segment diverged. No more guessing.Counting hits: two standards
A turn counts as a miss when input ≥ 2,000 and ≥ 5% of the prompt. The flaw: it
cannot tell “new content this turn” (tool results are new by definition) from “the
prefix was recomputed” — about 90% of reports were false positives. Kept only as the
legacy panel denominator.
Recomputed tokens for a turn = min(previous prompt, current prompt) − cacheRead,
floored at 1,024. That is what the turn actually overpaid, with lastMiss attribution:
idle past TTL, or a model switch.
cacheRead / (input + cacheRead + cacheWrite), and the panel
shows a recent 10-request window — cumulative numbers get diluted by cold starts and
rebuilds; only a short window shows steady state.
Two deliberate exemptions: a cold-start first turn has no prefix that “should have
hit”, so it does not count; and a session whose provider never reported cache activity
is excluded from recompute accounting entirely — otherwise sessions on providers that
do not report caching would all register as misses.
Compaction and caching pull against each other
Compaction replaces history with a summary, so the prefix necessarily changes and the first turn after compaction is expected to recompute in full. Statistics track those asrebuilds, separate from real misses.
Compaction’s own invariants (see the header note in
apps/sidecar/pi-agent/src/agent/context.ts):
- It happens only at turn boundaries; the summary replaces the entire history with no original tail kept;
- The checkpoint row persists to JSONL while full message rows are retained, so UI history is unaffected;
- If summary generation fails, a fresh_window fallback kicks in — no model request, a fixed rollover marker is loaded, and the session continues.
Next
With the cache accounting settled, goal mode keeps a ledger of its own.
