What does it mean that pipeline steps form a composable graph?

Short answer: Pipelines are flexible graphs of steps with multiple entry points, not one fixed chain for every input.

Different content needs different paths: immediate process, buffer first, or skip extraction for pre-extracted facts. Separate entries can converge into shared transform and commit, or stay on distinct routes when that fits better. Multi-pass patterns like extract-commit-buffer-transform again only work with composable chaining. Engram lets teams compose steps to match domain needs.

This Part has now covered every individual stage a pipeline can contain, extraction, transformation, buffering, and commit. None of these stages exists in isolation, though, they connect together into an actual working sequence, and the way they connect turns out to be more flexible than a single, fixed chain. This chapter looks at pipelines as a composable graph, and what that flexibility actually makes possible.

Why Wouldn’t a Single, Fixed Sequence of Steps Be Enough to Handle Every Kind of Memory Processing a System Might Need?

A fixed, one-size-fits-all sequence would force every kind of content through identical handling regardless of what that content actually needs. Some content genuinely benefits from immediate processing straight through to storage, while other content benefits from pausing to accumulate alongside related material first. Some content needs its own dedicated extraction logic entirely distinct from other content types, while later stages of processing might reasonably be shared across all of them once each has already been converted into a common, comparable form. A single fixed chain can’t accommodate this kind of genuine variation, which is exactly why a pipeline is structured as a graph of composable steps rather than one rigid, unchangeable sequence.

What Does It Actually Mean for a Pipeline to Have Multiple Separate Entry Points Rather Than Just One?

Each distinct shape of incoming content, conversation, plain text, or already-extracted facts, enters the pipeline through its own dedicated starting step suited specifically to that shape, rather than all content being forced through one single, generic entry point regardless of its actual structure. This matters because the earlier chapters in this Part on conversation-aware extraction and on the differences between input shapes both depend on exactly this branching structure existing in the first place, each shape genuinely needs its own specific handling at the very start of processing, even if much of what happens afterward ends up being shared.

How Do These Separate Entry Points Actually Come Back Together Into Shared Processing Later On?

Once each entry point’s dedicated extraction step has done its shape-specific work, the results, now in a common form regardless of what shape the original input happened to take, can converge into the exact same downstream transformation and commit steps. This convergence is what avoids needless duplication, a system doesn’t need entirely separate transformation logic for conversation-derived facts versus plain-text-derived facts, since by the time either one reaches the transform stage, both are simply memories awaiting the same kind of integration decision the transform stage always makes.

Does Convergence Into Shared Steps Mean Every Pipeline Has to Route All Content Through the Exact Same Downstream Path?

Not necessarily, and this flexibility runs in both directions. A pipeline can just as easily route different content types toward entirely separate downstream handling when that separation genuinely makes sense for a specific use case, rather than being forced to converge everything into one shared path regardless of whether that convergence actually fits. The graph structure supports both patterns, shared convergence where it helps and continued separation where it doesn’t, leaving that choice to whoever is actually configuring the pipeline for a specific domain’s needs.

How Does This Graph Structure Actually Support the More Elaborate, Multi-Stage Patterns Covered Earlier in This Part?

The daily-rollup pattern covered in this Part’s discussion of buffering, extract, then transform, then commit, then buffer, then transform again, then commit again, is only possible because steps can be chained flexibly rather than being limited to one pass straight through a fixed sequence. A graph structure lets a pipeline revisit transformation and commit a second time after a buffer has accumulated a full day’s worth of already-committed memories, producing a higher-level daily summary from material that was already fully processed once before. A rigid, single-pass chain would have no way to express this kind of legitimate, multi-stage refinement.

How Does Weaviate Engram’s Pipeline Structure Let a System Actually Compose These Steps to Fit Its Own Specific Needs?

Weaviate Engram defines pipelines as a graph with dedicated entry points per input shape, converging into shared or separate downstream steps as a project’s configuration actually calls for. Consider a multi-channel customer feedback platform, gathering input from live chat conversations, emailed complaints, and a structured post-purchase survey all at once:

from engram import EngramClient

client = EngramClient(api_key=os.environ["ENGRAM_API_KEY"])

client.memories.add(
    [
        {"role": "user", "content": "The replacement part arrived damaged again, this is the second time."},
        {"role": "assistant", "content": "I'm escalating this to our quality team right away."},
    ],
    properties={"customer_id": "customer-feedback-7734"},
)

client.memories.add(
    "Emailed complaint: packaging insulation appears insufficient for fragile electronics shipments.",
    properties={"customer_id": "customer-feedback-9012"},
)

The chat conversation enters through its dedicated conversation extraction step while the emailed complaint enters through the plain-text extraction step, each handled according to its own actual structure, before both converge into the exact same shared transformation and commit logic that reconciles each customer’s feedback against whatever’s already on file:

results = client.memories.search(
    query="Are there recurring packaging or shipping damage complaints?",
    properties={"customer_id": "customer-feedback-7734"},
)

A quality team reviewing recurring packaging issues benefits from this shared convergence, since every channel’s feedback ends up integrated through identical deduplication and reconciliation logic regardless of whether it originally arrived as a live chat exchange or a plain emailed complaint. This is exactly the value a composable pipeline graph delivers: handling each input shape correctly at the start while still applying one single, consistent standard for how that content ultimately gets integrated into memory.

Composing extraction, transformation, buffering, and commit into a flexible graph gives a pipeline the structure to handle genuinely varied content correctly. One property this graph structure has to preserve carefully, especially once retries and durable execution enter the picture, is making sure the exact same piece of content never accidentally gets committed to memory more than once. Our next chapter, What is idempotency in memory writes?, takes up exactly that concern.