OpenAI GPT-5.6 Models Hit Amazon Bedrock with Explicit Prompt Caching

AWS ML Blog · rss · 2026-07-31

OpenAI's GPT-5.6 models (Sol, Terra, and Luna) are now generally available on Amazon Bedrock, covering capability tiers from complex reasoning to high-volume tasks.

Alongside the models, Bedrock introduced explicit prompt caching for GPT-5.6. This feature allows precise control over which prompt portions are cached and reused, offering a 90% discount on cached input tokens with a 30-minute retention period. It is particularly effective for agentic workflows with repetitive system instructions or tool definitions.

The post also details how to configure clients via the OpenAI-compatible Responses API, adjust reasoning effort, call tools, and optimize workloads using both implicit and explicit caching modes.

Original post →

More from Infra

Infra channel →