AWS Tutorial: Persistent Agent Memory with S3 Vectors and NVIDIA NeMo

AWS ML Blog · rss · 2026-10-02

An AWS blog post gives a full implementation guide for using Amazon S3 Vectors as the persistent memory layer inside the NVIDIA NeMo Agent Toolkit (NAT), deployed on Amazon EKS, with a multi-agent investment research use case.

What NAT is: an open source, framework-agnostic framework for building, profiling and optimizing AI agents, working with Strands Agents, LangChain, LlamaIndex, CrewAI and custom implementations. It offers agent orchestration (nat run / nat serve), profiling (tokens, latency, throughput), evaluation (answer accuracy, context relevance, groundedness, agent trajectory) and automated hyperparameter tuning (temperature, topp, maxtokens).

NAT's memory module: key components are the MemoryEditor abstract interface (additems / search / removeitems), the MemoryItem data model (conversation history, tags, metadata, userid, memory text), MemoryBaseConfig (a Pydantic base discovered via the type field in YAML), and the automemoryagent wrapper. Built-in providers: Mem0, MemMachine, Redis, Zep.

Why S3 Vectors: semantic retrieval (cosine/euclidean), filterable metadata, strong write consistency (memories visible immediately), up to 2 billion vectors per index with no capacity planning, pay-per-storage/write/query with no idle compute, and IAM access control per bucket and index with per-tenant isolation.

Three implementation steps: (1) create the vector bucket and index (1024 dimensions matching Amazon Titan Text Embeddings V2, with content marked non-filterable); (2) implement and register a custom MemoryEditor plugin (boto3 calls to s3vectors and bedrock-runtime, uuid4 to avoid same-second key collisions, truncating content for metadata size limits); (3) configure the agent workflow. Prerequisites include an AWS account, an EKS cluster, NAT 1.6, Python 3.11/3.12, kubectl and Docker.

Original post →

More from coding & agent

coding & agent channel →