Weaviate Engram is a fully managed AI memory and context service that gives AI agents persistent memory — so they can learn anything and keep it, instead of starting over every session.
Get a free cloud cluster Learn MoreEngram is a memory server by Weaviate which gives AI agents a place to store and recall AI memory — across conversations, users, and sessions, running on top of Weaviate’s vector database, so your agent can learn anything and keep it.
Most agents forget everything the moment a session ends. Making the context window longer isn’t a real fix either — LLMs get “lost in the middle” of long prompts, and every extra token adds cost and latency to every message you send. This is the real difference between context engineering and memory engineering: stuffing more into the context window versus actively maintaining what’s remembered.
So instead, you send Engram raw data — a conversation, a plain event, or facts you’ve already extracted — and it processes that asynchronously through a pipeline architecture, so your agentic workflow never loses memory between steps: first it extracts the individual facts, then it transforms them by deduplicating and merging with what it already knows, then it commits the result to storage. You get a run ID back immediately and don’t have to wait for any of it to finish.
Later, when your agent needs context memory, you search with a plain-language query using vector similarity, keyword (BM25), or hybrid retrieval, and Engram returns the matching memories ranked by relevance.
Every memory belongs to a topic — a category like “user preferences” that defines what gets extracted — and a scope, which controls who can see it: shared across a project, locked to one user, or isolated further by custom properties like a conversation ID. That’s what keeps one user’s data from ever leaking into another’s.
A 50-turn conversation can easily exceed 10,000 input tokens per request — and every LLM call re-sends the entire history, so turn 1’s messages get billed again at turn 2, 3, and every turn after. Engram replaces that growing list with a search: query for relevant memories and send only those, keeping context size flat no matter how long the conversation runs.
LLMs get “lost in the middle” of long prompts, and effective context length is still far below 100%. Cramming more history in doesn’t just fail to help — it degrades accuracy while increasing latency and cost on every request, since the whole conversation gets re-paid for each time.
Adding to Engram doesn’t block your app. You send the raw event and get a run ID back immediately, while Engram processes it asynchronously in the background — extracting facts, deduplicating them against what it already knows, and committing the result — so your user gets their response at full speed.
As context grows, agents run into the same failure modes: hallucinated details that compound on themselves, getting buried under irrelevant history, reaching for the wrong tool, or getting stuck between contradictory facts. Engram’s pipeline actively reconciles and deduplicates memories instead of letting a raw transcript pile up unmanaged.
A single agentic workflow can span multiple agents and context windows. Engram’s scopes let that memory be shared project-wide, locked to one user, or isolated further by custom properties — instead of getting stuck in one chat thread nobody else can see.
In Weaviate’s own internal testing, an agent without grounded memory fabricated the same plausible-sounding URL in two separate runs when context was incomplete. With Engram recalling the actual prior context, both fabrications were avoided.
Even with extended context windows, LLM performance degrades with long inputs. Sending longer conversation histories to models increases latency, inflates costs, and produces less grounded answers.
Memory is actively maintained, not accumulated. Engram doesn’t pile up an ever-growing context blob — pipelines extract relevant information and reconcile it against what’s already known, handling deduplication, preference changes, and time-evolving facts, then persist a clean, structured memory state.
Hybrid retrieval is optimized with Weaviate at scale. Memory retrieval inherits Weaviate’s vector, keyword, and topic-filtered search on the same production retrieval stack customers already trust, with no separate query layer to operate.
User interactions with agents, especially conversations, are noisy and can produce facts that evolve over time. Instead of accumulating raw facts and relying on LLMs to reconcile them, memory needs active management: extraction, deduplication, and consolidation.
Pipelines are durable and asynchronous — they run in the background and do not block the application. Buffers aggregate across runs and flush on time- or data-based triggers, enabling rollups and windowed processing.
Composable pipeline primitives give you four step types to fit different use cases: extract, transform, buffer, and commit — mixed and matched depending on what your pipeline needs to do.
When a single request is spread out across several agents, built-in memory patterns collapse. Shared, persistent memory is required for orchestrating multiple agents against the same task.
Templates get you started quickly and are customizable for deeper control. Common use cases — personalization, continual learning, multi-agent state — ship as templates, and teams that outgrow the template drop down to direct pipeline control without leaving the platform.
An intuitive organizational framework structures memories end to end: topics describe what’s worth remembering, scopes and properties slice and facet memories, and groups collect topics and pipelines into deployable units.
# Install the Engram plugin export ENGRAM_API_KEY="eng_..." /plugin marketplace add weaviate/engram-plugins /plugin install engram@weaviate-engram # That's it — memory starts working on your next prompt
# Install and configure pip install hermes-weaviate-engram hermes memory setup # choose weaviate_engram when prompted
from engram import EngramClient client = EngramClient(api_key=os.environ["ENGRAM_API_KEY"]) client.memories.add("...", user_id="alice") client.memories.search(query="...", user_id="alice")
curl -X POST https://api.engram.weaviate.io/v1/memories \ -H "Authorization: Bearer $ENGRAM_API_KEY" \ -H "Content-Type: application/json" \ -d '{"input": {"string": {"content": ["..."]}}, "user_id": "alice"}'
Six steps make up that architecture, from the moment you send raw input to the moment it comes back as a searchable memory.
You send raw text, a conversation, or pre-extracted facts to the API, along with scope parameters like user_id.
Input is routed to a group — a bundle of topics (what to remember) and a pipeline (how to process it) for one use case.
The pipeline’s extract step pulls individual facts from the input that match your configured topics.
Transform steps retrieve related existing memories and decide what to do with each — merge, rewrite, or discard duplicates.
The commit step persists the resulting create, update, and delete operations to Weaviate as vector-embedded memories.
Later, you query with vector, BM25, or hybrid retrieval, and Engram returns the matching memories ranked by relevance.