OpenAI GPT-5.6 Models Hit Amazon Bedrock with Explicit Prompt Caching
AWS ML Blog · rss · 2026-07-31
OpenAI's GPT-5.6 models (Sol, Terra, and Luna) are now generally available on Amazon Bedrock, covering capability tiers from complex reasoning to high-volume tasks.
Alongside the models, Bedrock introduced explicit prompt caching for GPT-5.6. This feature allows precise control over which prompt portions are cached and reused, offering a 90% discount on cached input tokens with a 30-minute retention period. It is particularly effective for agentic workflows with repetitive system instructions or tool definitions.
The post also details how to configure clients via the OpenAI-compatible Responses API, adjust reasoning effort, call tools, and optimize workloads using both implicit and explicit caching modes.
More from Infra
- Together AI Webinar: Deploying Open-Weight Models in Production — togethercompute · 2026-07-31
- AI Data Center Noise Hits 105 Decibels, Sparking Lawsuits Against Tech Giants — mkheck · 2026-07-31
- Yahoo Enhances Search Retargeting with Amazon Bedrock, Boosting Keyword Expansion by 600x — AWS ML Blog · 2026-07-31
- Analyst: AI Fundamentals Unchanged Amidst Severe Compute Deficit and Low Penetration — BenBajarin · 2026-07-31
- AI Buildout Validates Kleiner Perkins' Cleantech Fund 20 Years Later — matt_slotnick · 2026-07-31
- China's DUV Lithography: Not an ASML Killer, But an Iteration Loop — demian_ai · 2026-07-31