TL;DR: Prompt caching (also called context caching) reuses the already-processed part of a prompt across LLM requests, so repeated context is billed at a fraction of the normal input price. With Amazon Bedrock prompt caching, that can reduce input costs by up to 90% and latency by up to 85%. Karini AI now supports it for prompts and AI agents, making it one of the simplest LLM cost optimization levers for enterprise agentic AI.
Most of what an AI agent reads, it has already read. Every iteration of an agent loop, and every question asked of the same long document, sends the model the same system prompt, the same tool definitions and the same reference material, and the enterprise pays full price for it each time.
Over the past year, Karini AI has built a set of levers for governing LLM token economics. At build time, teams compare models in the Model Hub to find the best price-to-performance for each task, optimize agents during design to cut unnecessary token consumption, and distill larger workflows onto smaller, more economical language models. In production, cost metering prices every call at the token level, and Budgets & Controls enforces spend limits and rate limits by organization, role, recipe, copilot and user.
Each of those levers changes which tokens you spend or how many you are allowed to spend. None of them changes what a repeated token costs. When a 12,000-token system prompt and toolset is re-sent on every step of a ten-step agent run, the model reprocesses it from scratch ten times, and you are billed for it ten times.
What if the model could remember the part of the prompt that never changes, and you paid a fraction of the price to reuse it?
Today we are launching Prompt Caching on the Karini AI platform.
What Is a KV Cache, and How Does Prompt Caching Work?
A large language model processes a request in two phases. In prefill, it reads the entire prompt and computes, for every token at every layer, a pair of attention tensors called keys and values. In decode, it generates the response one token at a time, and each new token attends to all of those keys and values.
Recomputing keys and values for the whole sequence at every generated token would be wasteful, so models keep them in memory for the life of the request. That store is the KV cache. It is why a model can write a 1,000-token answer without re-reading the prompt 1,000 times.
The KV cache has one limitation: it disappears when the request ends. The next request, even one that begins with exactly the same 10,000 tokens, runs prefill again from token one. Prefill is where most input cost and most time to first token (TTFT) come from on long prompts.
Prompt caching keeps that KV state alive across requests. Some providers call this context caching or prefix caching. When a new request begins with a prefix the provider has already processed, the model loads the stored keys and values instead of recomputing them, and runs prefill only on the new tokens. On Amazon Bedrock, AWS reports this can reduce costs by up to 90% and latency by up to 85% for supported models.
Three properties of prompt caching shape how you design for it:
- It is a prefix match. The cached portion must be identical, token for token, from the start of the prompt. Change one character early on and everything after it is a miss. Static content has to come first; variable content last.
- Reads are cheap; writes carry a premium. On Bedrock, a cache read typically costs about 10% of a standard input token, while writing a new cache entry costs about 125%. Caching pays off from the first reuse.
- It expires. Entries live for a time-to-live (TTL) window, commonly five minutes and refreshed on each hit, with longer windows on some models. Caching rewards workloads that reuse the same prefix in quick succession.
Why Are AI Agents and Long Documents So Expensive?
For a single short chat turn, prompt caching barely matters. It matters enormously for the two workloads that dominate enterprise agentic AI. If you are trying to reduce LLM costs or AI agent costs, these are the first places to look.
Agent Iterations: Pay Once, Reuse Every Step (~74% Lower Input Cost)
An agent does not make one model call; it makes a loop of them. Each iteration of a reason-act cycle re-sends the system prompt, the tool definitions, any reference context, and the full history of prior thoughts, tool calls and tool results, then appends one more step.
That means input tokens grow roughly quadratically with the number of steps. Take an RFP response agent with 12,000 tokens of instructions, tool schemas and company boilerplate, adding about 1,500 tokens of tool calls and results per step. Over ten iterations, it sends 187,500 input tokens, and more than 90% of them are an exact repeat of something the model saw seconds earlier.
With caching, each iteration reads everything up to the previous step from cache and pays full price only for the newest step:
| Agent run (10 iterations) | Input tokens sent | Billed input, standard-token equivalent |
|---|---|---|
| Without prompt caching | 187,500 | 187,500 |
| With prompt caching | 187,500 | ~48,100 (~74% lower) |
Illustrative: 12,000-token static prefix, 1,500 new tokens per step, cache reads at 10% and writes at 125% of the standard input rate. Output tokens are unchanged.
The latency effect compounds the same way. Later iterations, which carry the longest prompts, are exactly the ones that skip the most prefill.
Long Documents: Read Once, Query Many Times (~67% Lower Input Cost)
The second pattern is a large document interrogated more than once: a 40,000-token contract summarized for an executive, then scanned for risks, obligations, renewal terms and key dates. Without caching, every one of those prompts pays to re-read the whole contract.
With the document cached, the first request writes it once and each follow-up reads it at a fraction of the cost:
| Five prompts on one 40,000-token document | Billed input, standard-token equivalent |
|---|---|
| Without prompt caching | 211,000 |
| With prompt caching | ~70,300 (~67% lower) |
Illustrative: 2,000 tokens of instructions and examples plus the 40,000-token document cached, a 200-token question per prompt.
The same logic holds for long policy manuals behind a copilot, retrieval-augmented generation (RAG) prompts that share the same retrieved context, few-shot libraries, large JSON schemas and multimodal prompts carrying images.
What Does Prompt Caching on Karini AI Include?
Prompt Caching on Karini AI lets teams cache the stable prefix of any prompt or agent, so repeated context is billed at cache-read rates instead of full input rates. Three things make it more than a checkbox:
- You control the boundary. A caching breakpoint, inserted from the prompt playground with live color-coded highlighting, marks exactly where the static part of your prompt ends. Most platforms cache whatever the provider guesses.
- It works across every execution path. Single-shot prompts, ReAct agents, graph-based agents, deep agents and their subagents, streaming and non-streaming.
- Cached tokens are billed correctly. Base input, cache reads and cache writes are metered as three separate legs at their own rates, everywhere cost is reported, so the savings you see are the savings you get.
Caching is opt-in per request and per model endpoint. When it is off, execution is unchanged.
How Do You Place a Prompt Caching Breakpoint? A Summarization Example
Because Bedrock caches everything before a cache point, where you put the boundary is the whole game. Put it after your variable content and every call writes a new cache entry and reads nothing back.
Karini makes the boundary explicit. In the prompt playground, the Prompt Caching Breakpoint button sits alongside the other prompt-authoring tools and inserts a marker at your cursor. A summarization prompt might use two:
You are a contracts analyst. Follow the style guide and output schema below.
<style guide, output schema, few-shot examples>
<<PROMPT_CACHING_BREAKPOINT>>
Contract:
{contract_text}
<<PROMPT_CACHING_BREAKPOINT>>
Task: {summary_request}
The first segment (instructions and examples) is reused across every contract. The second (the contract itself) is reused across every question about that contract. Only the short task at the end is billed at full input rates.
A prompt with no breakpoint at all is cached in full, which is exactly right for a fully static prompt.
How Does Prompt Caching Work in AI Agents? An RFP Response Agent Example
Agents are where caching earns the most, and where it is hardest to get right by hand, because the prompt changes shape on every iteration.
Consider an RFP response agent. It carries a long system prompt (response guidelines, tone, compliance rules), a set of tools (knowledge-base search, past-proposal lookup, pricing tables) and company boilerplate. For each RFP section, it searches, drafts, checks requirements, and revises, iterating until the answer is complete.
When caching is enabled on the agent's model endpoint, Karini places cache points automatically on each iteration:
- on the system prompt, which never changes during the run;
- on the tool definitions, which are often thousands of tokens of JSON schema;
- on a rolling checkpoint at the newest message, so each step reads the entire prior conversation from cache and writes only what it just added.
No breakpoints to place, no changes to the agent's logic. In our testing, a warm agent run read over 9,300 tokens from cache and wrote none: the full prefix was served from cache.
This works across ReAct agents, graph-based agents, and deep agents including their subagents, in both streaming and non-streaming mode. Fallback models never cache by design: a fallback fires only when the primary has failed, and paying a cache-write premium on the recovery path buys nothing.
Cutting costs should never cost quality. Pair prompt caching with Agent Evaluation to confirm your agents' answers hold steady as token spend drops.
What Are the Benefits of Prompt Caching?
- Lower input cost: Repeated context is billed at roughly a tenth of the standard input rate.
- Faster responses: Cached prefixes skip reprocessing, cutting time-to-first-token on long prompts and late agent iterations.
- Cheaper agents at scale: Each agent step reuses everything before it from cache, so cost stops ballooning as runs get longer.
- Precise control: Breakpoints and live highlighting put the cache boundary exactly where the author wants it.
- Every cached token billed at its own rate: Standard input, cache reads and cache writes are metered separately, so the savings you see in Karini match your invoice, and FinOps teams can trust the token costs they report.
- Proven on live traffic: One prompt on Bedrock cost $0.010862 uncached, $0.013584 on the first run that writes the cache (+25%), and $0.003048 on every run that reads it (-72%). Caching pays for itself on the first reuse.
- Your rates, not list price: Cached tokens are priced at the rates on your model endpoint, including negotiated discounts, and flow into the same usage ledger as Budgets & Controls.
- No silent failures: Clear warnings when a prompt is too short to cache, a breakpoint splits nothing, or a price is missing.
- Broad model coverage: 34 Bedrock models across Anthropic Claude, Amazon Nova and OpenAI, with the right settings for each built in.
- Low-disruption rollout: Opt-in per endpoint and per request; with caching off, nothing changes.
Put Prompt Caching to Work
Any workload that sends the same context again and again is a candidate. Start with the ones where repeated tokens add up fastest:
- RFP and proposal response agents: Response guidelines, company boilerplate and tool definitions are reused on every step and every section.
- Contract, policy and compliance review: Cache the document once, then summarize it, extract obligations and check risks without paying to re-read it.
- Support and knowledge copilots: Long policy manuals, product catalogs and troubleshooting guides behind every customer question.
- Maintenance and operations assistants: Equipment manuals, SOPs and plant standards reused across technician queries.
- Document extraction pipelines: Fixed instructions, output schemas and few-shot examples applied to thousands of invoices, forms or reports.
- Multi-step research and analysis agents: Long-running loops where history grows with every tool call.
Getting started takes minutes:
- Pick one workload: your longest-running agent or your most-queried document.
- Turn on caching: enable Prompt Caching on its model endpoint in the Model Hub and choose a TTL.
- Place a breakpoint: for prompts, insert a breakpoint after the static content; agents cache automatically.
- Compare the cost: run it twice in the playground and watch cache reads replace full-price input.
Model comparison, agent optimization, cost metering and Budgets & Controls help you spend fewer tokens and govern how they are spent. Prompt caching makes the tokens you repeat cheaper and faster. Prompt Caching is available today on the Karini AI platform. Try it on your next agent run, or book a demo to see it on your own workloads.
FAQ: Prompt Caching on Karini AI
What is prompt caching?
Prompt caching stores the model's processed state (the KV cache) for the beginning of a prompt, so later requests that start with the same content reuse it instead of reprocessing it. Reused tokens cost a fraction of standard input tokens and return faster.
How is this different from caching a model's response?
Response caching returns a stored answer to an identical question. Prompt caching reuses only the processed input; the model still generates a fresh response to each request, so it works when questions differ but context repeats.
Which models support prompt caching?
34 models on Amazon Bedrock across Anthropic Claude, Amazon Nova and OpenAI. Each model's supported TTLs, minimum cacheable prefix and cache-point budget are built into the Model Hub.
Do I need to change my agents?
No. Enable caching on the agent's model endpoint and Karini caches the system prompt, tool definitions and conversation history automatically. Breakpoints are only needed for prompts where you want to control the boundary yourself.
Where should I place a breakpoint?
After the content that stays the same across calls and before the content that changes. Instructions and examples first, then reference documents, then the specific question.
How long does the cache last?
It depends on the model: 5 minutes (refreshed on each hit) is standard, some Anthropic models support 1 hour, and some OpenAI models on Bedrock cache for 30 minutes automatically. You choose the TTL on the endpoint from the options the model supports.
Does caching ever cost more?
Writing to cache costs about 25% more than a standard input token. If a prefix is reused even once within the TTL, caching is cheaper overall. For prompts that never repeat, leave it off.
How do I see my savings?
Cache-read and cache-write tokens and their costs appear in the playground, recipe and chatbot history, traces, and CSV exports, and they flow into the same usage ledger used by Budgets & Controls.
How much does prompt caching save?
On Amazon Bedrock, cache reads cost about 10% of standard input tokens, so the cached part of a prompt is up to 90% cheaper. In the examples above, a 10-step AI agent run used ~74% less input spend and five questions on one long document used ~67% less. Actual savings depend on how much of the prompt repeats and how often.
Does prompt caching change the model's output?
No. The model receives the same prompt and produces its response the same way; caching only skips recomputing work it has already done for the repeated prefix.
What is the minimum prompt size for caching?
It depends on the model. On Amazon Bedrock, the cached prefix must be at least 512 to 4,096 tokens, with 1,024 tokens the most common minimum. Karini warns you when a prefix falls below the model's minimum.
Does prompt caching work with RAG?
Yes, whenever retrieved context repeats across requests, such as follow-up questions on the same documents. Put stable instructions and shared context before the breakpoint, and per-query retrieved chunks after it.
What is the difference between a KV cache and prompt caching?
A KV cache stores attention keys and values inside a single request so the model can generate text efficiently. Prompt caching keeps that state for a prompt's prefix across requests, so later requests that start the same way skip reprocessing it.
Is prompt caching the same as context caching?
Yes. Providers use different names, including prompt caching, context caching and prefix caching, for reusing processed prompt content across LLM requests.





