Official docs (Prompt Caching): documents that Groq caches prompt prefixes automatically (volatile memory, expires after ~2 hours idle), with a 50% discount on cached input tokens and cached tokens not counting toward rate limits. Cache hits need EXACT prefix matches, a minimum cacheable length of 128-1024 tokens depending on model, and it's currently only on gpt-oss-20b, gpt-oss-120b, and gpt-oss-safeguard-20b. You track it via usage.prompt_tokens_details.cached_tokens.

## What to do
To actually get Groq prompt-cache hits, order your prompt with all static content first (system instructions, few-shot examples, tool definitions, schemas) and dynamic content last (user query, timestamps, session data). One stray timestamp or user ID at the top breaks the exact prefix match and nothing caches. Keep tool_choice and image usage identical across calls, since changing them invalidates the cache. Check usage.prompt_tokens_details.cached_tokens on each response and compute hit rate as cached_tokens / prompt_tokens; if it stays near zero, your prefixes are drifting. Note it can't be turned off and expires automatically, so don't build logic that depends on a cache existing. Batch requests still cache but the discount doesn't stack with the 50% batch discount.