← 返回事件
持续讨论科技

Optimizing cost and latency with Amazon Bedrock prompt caching

图:AWS Machine Learning

发生了什么

Prompt caching in Amazon Bedrock can cut input token costs by up to 90% when you repeatedly send the same context to foundation models. This post walks through six practical prompt caching scenarios using the Converse API: message content, system prompt, tool definition, mixed TTL, tenant isolation, and LangChain integration.

摘要按规则整理自下方来源原文

为什么在扩散

来源

一手来源