# Building Prompt Caching into Production: How I Cut Claude API Costs by 85%
When AdminOS went live with 5 AI agents handling WhatsApp conversations for multiple businesses, my first invoice from Anthropic was a wake-up call. I had built something that worked. I had not yet built something that was economical at scale.
Enter prompt caching. Two weeks and one architectural review later: 85% cost reduction. Same quality. Better latency on repeated requests.
Here's exactly how I did it.
What Prompt Caching Actually Is
When you call the Claude API, every token in your request is processed and charged. Your system prompt (which might be 2,000 words describing the AI's role, context, and rules) gets processed identically on every single call.
Prompt caching tells the API: this block of content is static, cache it at the API level. Subsequent calls that use the same static block hit the cache instead of re-processing. Cache read tokens cost 10% of standard input tokens.
The math: a 4,000-token system prompt called 1,000 times per day costs 4M tokens at full rate. With caching after the first call, that's 4,000 tokens at full rate + 3,996,000 at 10% rate. For Sonnet pricing, that's a daily saving of 90% on system prompt tokens.
The Implementation
The API change is minimal: it's the architectural thinking that's the work.
const response = await anthropic.messages.create({
model: 'claude-sonnet-4-6',
max_tokens: 1024,
system: [
{
type: 'text',
text: STATIC_SYSTEM_PROMPT, // ≥4096 tokens to trigger caching
cache_control: { type: 'ephemeral' } // Mark this block as cacheable
}
],
messages: [
{
role: 'user',
content: userMessage // Dynamic, never cached
}
]
});
The critical constraint: the cacheable block must be ≥4096 tokens. Smaller blocks don't qualify. This forced me to think about system prompt design differently.
Designing for Cacheability
The 4096-token minimum isn't a bug: it's a forcing function for better prompt design. Here's how I structured system prompts to hit the threshold while staying coherent:
Layer 1: Role and personality (~500 tokens), Who the AI is, what it does, how it speaks.
Layer 2: Domain knowledge (~2000 tokens). Everything the AI needs to know about the business context, product, users.
Layer 3: Rules and constraints (~1000 tokens). What it must never do, edge cases, fallback behaviors.
Layer 4: Examples (~600+ tokens), Few-shot examples of ideal responses in various scenarios.
Total: comfortably over 4096. Every subsequent call reads from cache.
What Changed Per Application
AdminOS (5 agents, WhatsApp-native): Each of the 5 specialist agents has a cached system prompt. The Debt Recovery agent's prompt includes the full escalation framework: 14 stages, tone guidelines, legal compliance notes. That's 5,200 tokens cached once per conversation thread.
JarvisOS (15 wings): Each wing's system prompt includes my full personal context: work history, current projects, health baseline, financial situation, goals. Static across all queries to that wing. Cached.
VarsityOS (Nova AI, 6 agents): Nova's system prompt includes complete SA student context, NSFAS structure, university calendar, crisis protocols, 11 SA languages. 6,800 tokens. Cached.
The Unexpected Benefit: Better Latency
Cache hits are faster than full processing. On repeated conversations with the same agent, response times improved by 15-30%. For a WhatsApp-native product where users expect quick replies, this matters more than I initially appreciated.
What I'd Tell Anyone Building AI Products
Cache your system prompts. Full stop. If you're calling the same AI with the same role context more than a few times a day, you are leaving significant money on the table.
The implementation takes one afternoon. The architectural thinking (designing system prompts to be simultaneously comprehensive enough to be useful and modular enough to cache) is the real investment. Make it.
85% is not an edge case. It's what good prompt architecture looks like.
Are you caching your system prompts in production?
Reader Insights
0 responses
No insights yet. Be the first!