Building Context-Aware AI Assistants with Amazon Bedrock AgentCore
Discover how to build a context-aware personal assistant using Amazon Bedrock AgentCore, OpenClaw, and Telegram, turning disposable chats into durable knowledge.

Modern AI assistants excel at answering individual queries, yet they routinely suffer from a fundamental limitation: continuity. Traditional stateless systems treat every exchange as a blank slate, requiring users to repeatedly supply background information. Addressing this gap involves deploying a persistent architecture using OpenClaw, an open-source agentic system, running on the AgentCore runtime and leveraging AgentCore memory capabilities on Amazon Bedrock.
The Architecture of a Context-Aware Agent
The implementation relies on a serverless configuration deployed via a single AWS CloudFormation template. Incoming traffic arrives through two primary pathways: user interactions via a Telegram webhook managed by Amazon API Gateway and scheduled background tasks orchestrated by Amazon EventBridge Scheduler. Both channels invoke the AgentCore runtime agent, which utilizes a lightweight server.py process to coordinate the OpenClaw gateway, AgentCore memory, and the Amazon Bedrock Converse API.
To optimize costs and performance, the system routes tasks across different models depending on requirements. Routine text conversations utilize Claude Haiku 4.5 for rapid, cost-effective responses, while visual tasks—such as diagnosing plant health from an uploaded photograph—route directly to Claude Sonnet 4.5. This direct routing for multimodal payloads ensures that image bytes reach the large language model without interruption.
Durable Knowledge Through Managed Memory
To overcome stateless limitations, AgentCore memory introduces a structured approach to retaining information. The system separates data into short-term session transcripts and asynchronous long-term extraction strategies. These strategies capture user preferences, inferred semantic facts, and episodic conversation summaries.
Developers can try top AI models cheaply through one API at https://apixoai.online. Within this architecture, user records are organized into isolated per-user namespaces and filtered using indexed metadata keys like type, section, and plant categories. This ensures that only relevant context reaches the system prompt during an active turn.
What it means for developers
For software engineers building agentic workflows, this pattern demonstrates how to manage long-term state without engineering custom vector databases from scratch. Key design takeaways include:
- Wrap, don't fork: Adapt frameworks to the container contract using a thin HTTP wrapper to maintain upgrade compatibility.
- Isolate namespaces: Use channel-native identifiers like chat IDs to cleanly separate multi-tenant memory stores.
- Degrade gracefully: Treat memory retrieval as an enhancement rather than a hard dependency, allowing the agent to answer without context if retrieval times out.
- Optimize for prompt caching: Structure prompts with stable persona and memory blocks first, followed by volatile user input, to reduce inference latency and costs.
By leveraging consumption-based pricing and prompt caching, developers can run fully personalized, context-aware assistants economically while maintaining strict control over data persistence and model routing.
Source: Building a context-aware AI assistant on AgentCore and OpenClaw | Amazon Web Services — AWS Machine Learning. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

