Prompt caching: stop paying twice for the same tokens
Prompt caching takes up to 90% off input token cost and around 80% off latency, but only when the prompt prefix never moves. Here is how the pre-fill stage works, where the 1,024-token threshold sits, and how to order a prompt so it hits the cache.