Agents forget.
Weaviate Engram doesn’t.

Weaviate Engram is a fully managed AI memory and context service that gives AI agents persistent memory — so they can learn anything and keep it, instead of starting over every session.

Get a free cloud cluster Learn More
Skip the blank-page problem Start from a pipeline built for your use case — personalization, continual learning, or shared agent state — instead of designing memory from scratch.
Point it at your data Conversations, events, or facts you’ve already extracted go straight to the API, exactly as they are.
Let the pipeline take it from there Extraction, deduplication, and commits run on their own in the background — infrastructure you’d otherwise have to build and maintain yourself.
Give your agent memory it can trust Query that memory back in real time, so answers stay grounded instead of drifting from one session to the next.

What is Engram?

An AI memory server for LLM agents and applications.

Engram is a memory server by Weaviate which gives AI agents a place to store and recall AI memory — across conversations, users, and sessions, running on top of Weaviate’s vector database, so your agent can learn anything and keep it.

Most agents forget everything the moment a session ends. Making the context window longer isn’t a real fix either — LLMs get “lost in the middle” of long prompts, and every extra token adds cost and latency to every message you send. This is the real difference between context engineering and memory engineering: stuffing more into the context window versus actively maintaining what’s remembered.

So instead, you send Engram raw data — a conversation, a plain event, or facts you’ve already extracted — and it processes that asynchronously through a pipeline architecture, so your agentic workflow never loses memory between steps: first it extracts the individual facts, then it transforms them by deduplicating and merging with what it already knows, then it commits the result to storage. You get a run ID back immediately and don’t have to wait for any of it to finish.

Later, when your agent needs context memory, you search with a plain-language query using vector similarity, keyword (BM25), or hybrid retrieval, and Engram returns the matching memories ranked by relevance.

Every memory belongs to a topic — a category like “user preferences” that defines what gets extracted — and a scope, which controls who can see it: shared across a project, locked to one user, or isolated further by custom properties like a conversation ID. That’s what keeps one user’s data from ever leaking into another’s.

Why your AI agent needs Weaviate Engram.

The context window was never designed to be AI memory.

Full conversation history gets expensive, fast.

A 50-turn conversation can easily exceed 10,000 input tokens per request — and every LLM call re-sends the entire history, so turn 1’s messages get billed again at turn 2, 3, and every turn after. Engram replaces that growing list with a search: query for relevant memories and send only those, keeping context size flat no matter how long the conversation runs.

A longer context window doesn’t fix it.

LLMs get “lost in the middle” of long prompts, and effective context length is still far below 100%. Cramming more history in doesn’t just fail to help — it degrades accuracy while increasing latency and cost on every request, since the whole conversation gets re-paid for each time.

Storing Engram memory is fire-and-forget.

Adding to Engram doesn’t block your app. You send the raw event and get a run ID back immediately, while Engram processes it asynchronously in the background — extracting facts, deduplicating them against what it already knows, and committing the result — so your user gets their response at full speed.

Raw transcripts break agents in predictable ways.

As context grows, agents run into the same failure modes: hallucinated details that compound on themselves, getting buried under irrelevant history, reaching for the wrong tool, or getting stuck between contradictory facts. Engram’s pipeline actively reconciles and deduplicates memories instead of letting a raw transcript pile up unmanaged.

Agentic workflows need memory that isn’t trapped in one conversation.

A single agentic workflow can span multiple agents and context windows. Engram’s scopes let that memory be shared project-wide, locked to one user, or isolated further by custom properties — instead of getting stuck in one chat thread nobody else can see.

Grounded recall stops agents from making things up.

In Weaviate’s own internal testing, an agent without grounded memory fabricated the same plausible-sounding URL in two separate runs when context was incomplete. With Engram recalling the actual prior context, both fabrications were avoided.

What problems does Engram solve and how?

Three structural failure modes prevent agentic applications from delivering consistent outcomes. Weaviate Engram solves each one at the infrastructure level.

Long-context degradation

Even with extended context windows, LLM performance degrades with long inputs. Sending longer conversation histories to models increases latency, inflates costs, and produces less grounded answers.

Fix

Memory is actively maintained, not accumulated. Engram doesn’t pile up an ever-growing context blob — pipelines extract relevant information and reconcile it against what’s already known, handling deduplication, preference changes, and time-evolving facts, then persist a clean, structured memory state.

Fix

Hybrid retrieval is optimized with Weaviate at scale. Memory retrieval inherits Weaviate’s vector, keyword, and topic-filtered search on the same production retrieval stack customers already trust, with no separate query layer to operate.

Messy raw data from interactions

User interactions with agents, especially conversations, are noisy and can produce facts that evolve over time. Instead of accumulating raw facts and relying on LLMs to reconcile them, memory needs active management: extraction, deduplication, and consolidation.

Fix

Pipelines are durable and asynchronous — they run in the background and do not block the application. Buffers aggregate across runs and flush on time- or data-based triggers, enabling rollups and windowed processing.

Fix

Composable pipeline primitives give you four step types to fit different use cases: extract, transform, buffer, and commit — mixed and matched depending on what your pipeline needs to do.

Multi-agent context fragmentation

When a single request is spread out across several agents, built-in memory patterns collapse. Shared, persistent memory is required for orchestrating multiple agents against the same task.

Fix

Templates get you started quickly and are customizable for deeper control. Common use cases — personalization, continual learning, multi-agent state — ship as templates, and teams that outgrow the template drop down to direct pipeline control without leaving the platform.

Fix

An intuitive organizational framework structures memories end to end: topics describe what’s worth remembering, scopes and properties slice and facet memories, and groups collect topics and pipelines into deployable units.

How do you integrate Engram?

Bring Weaviate Engram’s AI memory to Claude Code, Cursor, Copilot, and anywhere else your agents run.
Claude CodeOfficial plugin
# Install the Engram plugin
export ENGRAM_API_KEY="eng_..."

/plugin marketplace add weaviate/engram-plugins
/plugin install engram@weaviate-engram

# That's it — memory starts working on your next prompt
Hermes AgentOfficial plugin
# Install and configure
pip install hermes-weaviate-engram
hermes memory setup

# choose weaviate_engram when prompted
Python SDKAny Python app — Cursor, Copilot, custom agents
from engram import EngramClient

client = EngramClient(api_key=os.environ["ENGRAM_API_KEY"])
client.memories.add("...", user_id="alice")
client.memories.search(query="...", user_id="alice")
REST APIAny language, any tool
curl -X POST https://api.engram.weaviate.io/v1/memories \
  -H "Authorization: Bearer $ENGRAM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"input": {"string": {"content": ["..."]}}, "user_id": "alice"}'

Why move from old-age memory to Engram for AI agents?

What Engram replaces.
Today

Sending the whole conversation back to the model on every turn works fine at first. Past a point it doesn’t — the model gets “lost in the middle” of long inputs, and every extra message costs more and answers slower, because the entire history gets re-paid for each time.

With Engram

Engram replaces the growing history with a search — memory that stays flat in size no matter how long the conversation runs.

Today

A hand-curated file of stable facts works fine for small, fixed budgets. It has nowhere to put reasoning chains, weeks of context, or state that needs to stay separate per user — everything gets flattened into the same file whether it belongs together or not.

With Engram

Engram organizes memory into topics, scopes, and groups instead of one flat file — so multi-user state and reasoning history each have a real home.

Today

Storing memories in a vector database yourself is the easy part. The extraction, deduplication, scoping, async processing, and durability that come with Engram out of the box are what you end up rebuilding on your own.

With Engram

Engram ships that entire pipeline — extract, transform, buffer, commit — so you’re not rebuilding infrastructure you’d otherwise have to maintain yourself.

Today

A separate memory service means running a second system alongside your retrieval stack — its own scaling behavior, its own uptime to worry about, its own thing that can go down independently of everything else.

With Engram

Engram unifies memory and retrieval on Weaviate, so you inherit the same query infrastructure you already trust — a natural extension if you’re already on Weaviate, not a parallel deployment.

Today

Agents that should get better with use stay flat instead. Every session starts from zero, and the same intermediate problems get solved over and over — eventually the quality plateau becomes visible to users.

With Engram

Engram gives your agent a way to learn anything — and keep it, instead of solving the same problem from scratch every session.

What is Engram’s architecture?

How Weaviate Engram’s architecture actually works.

Six steps make up that architecture, from the moment you send raw input to the moment it comes back as a searchable memory.

1

Send input

You send raw text, a conversation, or pre-extracted facts to the API, along with scope parameters like user_id.

2

Route to a group

Input is routed to a group — a bundle of topics (what to remember) and a pipeline (how to process it) for one use case.

3

Extract

The pipeline’s extract step pulls individual facts from the input that match your configured topics.

4

Transform

Transform steps retrieve related existing memories and decide what to do with each — merge, rewrite, or discard duplicates.

5

Commit

The commit step persists the resulting create, update, and delete operations to Weaviate as vector-embedded memories.

6

Search & retrieve

Later, you query with vector, BM25, or hybrid retrieval, and Engram returns the matching memories ranked by relevance.

// ready when you are

Want Engram to power up your AI agent with super persistent memory?

Get a Free Cloud Cluster Today