Span-level analysis of every conversation this agent runs – the token economics behind each answer and the tool-call orchestration underneath it. Levers span model routing, semantic model routing, prompt caching, context pruning, schema projection, request batching, retrieval (RAG) tuning, speculative decoding and token budgeting – cutting cost and latency without moving quality.
| Tool / span | Calls/wk | Avg tokens | p95 | Redundant | Lever |
|---|---|---|---|---|---|
VVector store (retrieval) | 8,420 | 2,600 | 0.9s | 4% | Cacheable |
CConfluence API | 3,100 | 1,800 | 1.4s | 18% | Over-fetch |
| 2,400 | 1,500 | 1.2s | 9% | Cacheable | |
| 1,900 | 1,100 | 0.8s | 3% | Healthy | |
| 1,500 | 3,200 | 1.7s | 22% | Over-fetch |
Summarise long incident threads before answering
Threads over 20 messages are passed verbatim into context. Summarise-then-answer to cut input tokens with no loss of grounding.
Tier single-fact lookups to a smaller model
~40% of queries are single-fact lookups that don't need Opus. Route them to Haiku and reserve Opus for multi-step reasoning.
Project AWS S3 doc responses to matched sections
S3 doc fetches pull whole files; request only the matched sections (schema projection) to drop 22% redundant tool-result tokens.