AWS walkthrough: agentic retrieval splits multi-intent queries into sub-searches on Bedrock

AWS ML Blog · rss · 2026-10-05

AWS ML Blog demonstrates agentic retrieval on Amazon Bedrock Managed Knowledge Base with LangChain. The core problem: standard RAG packs all user intents (e.g., comparing two products across three dimensions = six sub-questions) into a single query vector, so retrieved chunks look relevant but cover only part of the ask.

Agentic retrieval plans the retrieval instead: it decomposes the question into sub-queries, runs them, judges whether evidence suffices, and searches again if not. The post compares the Retrieve API (single hybrid search) against the AgenticRetrieveStream API (multi-step planning loop streaming trace events), covering cost differences between the two paths and when the cheaper one is right.

Practical details: requires boto3 ≥1.43.32 and langchain-aws ≥1.6.3; IAM must separate the knowledge base service role from the caller identity, noting bedrock:AgenticRetrieveStream and bedrock:InvokeModelWithResponseStream cannot be scoped to a knowledge base ARN; the often-missed bedrock:GetDocumentContent permission is required when the planner's FullDocumentExpansion step pulls a whole document, or queries fail midway.

Original post →

More from coding & agent

coding & agent channel →