Hello friends! The other day, while putting together the article on Xiaomi’s MiMo 2.6, I ran into a curious line in the price table: “Input (cache hit): $0.0036”. Right below it, the normal price: $0.435. More than a hundred times the difference for the same input token! If you use AI through an API, or you are just curious why it sometimes answers so fast, the cache hit is a concept worth understanding.
Let me explain with the diner on the corner. You walk in, the guy at the counter looks at you and says “the usual?”. Done, the order comes out fast and nobody writes anything down. That is a cache hit. Now imagine a new guy is working the counter: you explain everything again, the bread, how you like the burger, no onions. That is a cache miss. And there is a catch: if you disappear for too long, even the old guy forgets. I will come back to this diner at the end.
1. What a cache is, in one sentence
A cache is a place to keep things you will need again, close to where you will need them, so you do not have to fetch or compute everything all over.
Your browser does this with images from sites you visit often. Your phone does it with app data. And AI models started doing it with the text you send them.
2. Cache hit vs. cache miss
When you ask for something and it is already stored in the cache, that is a cache hit: the answer comes from the shortcut. When it is not, that is a cache miss: the system has to do the full work and, usually, stores the result for next time.
There is even a formula to check whether the cache is doing its job, the cache hit ratio. Cloudflare puts it like this: divide the number of hits by the total number of requests, hits plus misses. If 80 out of 100 requests came from the cache, your ratio is 80%.
3. What a cache hit does in AI
AI does not remember your conversation. With every new message, the app sends everything that came before all over again: the system instructions, the documents you attached and the entire history. In a long chat, or with an agent that works on its own for hours, that is a lot of text being reprocessed over and over.
Prompt caching solves this. If the beginning of what you send is identical to the previous call, the model reuses the work already done on that part instead of processing it all again. In Anthropic’s documentation and in OpenAI’s, the two benefits show up together: it gets cheaper and the response starts sooner.
The “cheaper” part is very concrete. Look at the price of Claude Opus 5.5 per million input tokens:
Claude Opus 5.5: price per million input tokens
The same token, four different prices. Scale in US dollars.
Notice that writing to the cache costs a bit more than the normal price. It is an investment: you pay a little extra the first time so you pay much less on the following ones.
4. In practice: how much it saves
I ran a simple simulation with the official prices. Imagine an assistant that loads a 50,000-token manual (about 75 pages) and answers 100 questions about it in a row, without letting the cache expire.
Simulation: a 50,000-token manual, 100 questions in a row
Cost of the manual’s input alone on Claude Opus 5.5. Scale in US dollars.
Without the cache, you pay for the manual’s 50,000 tokens 100 times. With it, you pay for the write once and then only for the reads, which cost 5% of the normal price on Opus 5.5. Same work, a fraction of the price. That is why agent tools, which resend the same context dozens of times, depend on this so much.
And every company prices the cache hit its own way:
What a cache hit costs
Cache read price as a share of the normal input price. Lower means a bigger discount.
5. How to get more cache hits
If you build with AI, these tips come straight from the documentation:
- Put what does not change at the start. Instructions, tools and documents first; the user’s question last. The cache works by prefix, so any change at the start invalidates everything after it.
- Watch out for things that change without you noticing. A date and time at the top of the instructions, for example, changes on every call and kills the cache.
- Respect the lifetime. At Anthropic, the default cache lasts 5 minutes (you can pay for 1 hour); at OpenAI, from GPT-5.6 onward, it is 30 minutes. After that, the guy at the counter forgot.
- Check that it is working. The API response tells you how many tokens came from the cache:
cache_read_input_tokensat Anthropic andcached_tokensat OpenAI. If that number is always zero, something is breaking the prefix. - Remember the minimum size. Short texts do not get cached: the minimum is 512 tokens on Anthropic’s newest models and 1,024 at OpenAI, from GPT-5.6 onward.
Fixed context up front, variable at the end. If you keep only one rule, keep that one.
6. Frequently asked questions
What is a cache hit, in a few words?
It is when what you asked for was already stored in the cache and the system serves it from there, without redoing the work. The opposite is a cache miss.
Does a cache hit make the AI smarter?
No. The model receives exactly the same text, with or without the cache. What changes is the cost and the time until the response starts.
Do I need to turn on prompt caching?
At OpenAI it is automatic on supported models. At Anthropic, you mark in the call how far you want to cache, or use the automatic mode. If you only use AI through the app, you do not need to do anything.
Back to the diner: always order the same way, and come back before the guy at the counter forgets.
Cheers, Fellipe Soares
Subscribe
Get the next articles by email.