Home

Graph Engineering

Graph Engineering
RoomSystems
FieldAgent architecture, systems engineering, control theory
Known forLoop→graph topology, verifier nodes, fan-out/fan-in, routing & self-correction
Key figuresHarrison Chase, Shann Holmberg, Andrew Ng, Carlos E. Perez, Peter Steinberger

Graph Engineering — Systems Master Brief

Graph engineering is the discipline of turning a single autonomous loop into a directed graph of typed, verifiable nodes — so that agentic pipelines can verify their own inputs, fan out into parallel work, fan back in through checks, route on the outcome, and act without a human in the loop. The loop finds the path; the graph defines it. This page is the baseline reference: what it is, the canonical patterns, the audit checklist, the tool landscape, and how the pattern maps onto the enc1/MemPalace/Nordisk stack.

Definition


Graph engineering is the discipline of structuring an autonomous agent so that work moves through a fixed, typed graph of nodes rather than a single free-running loop. Each node does one job, declares its inputs and outputs, and is connected to the next by explicit edges. Branches allow parallel fan-out; verifier nodes gate whether a branch may proceed; router nodes choose the next edge based on the outcome. The result is a machine that can start without you, work until a stop condition, survive failure, and leave evidence.


The core contrast comes from Shann Holmberg’s framing (July 2026): "the loop finds the path; the graph defines the path." A loop is one linear execution that iterates until it finds the correct answer for the current run. A graph is the topology every run traverses — the set of possible paths, the branch conditions, and the gates between them. Engineering the loop optimizes the run. Engineering the graph optimizes the machine that produces runs.


Harrison Chase — creator of LangGraph — pushed the field to earn the name when he asked publicly whether "graph engineering" is a real category or simply "langgraph". The honest answer that emerged is that the graph is a system design layer, not a library: LangGraph, Temporal, Dagster, and Airflow implement it; the skill is designing the nodes, contracts, verifiers, and gates independent of the runtime.


Loop vs Graph


Loop simpleGraph topology
One linear execution pathFixed topology of nodes and edges
Iterates to find the answer per runEvery run traverses the defined structure
No branch, no verification stepBranches, verifiers, conditional edges
Cheap to build, hard to trustCostlier to build, auditable and survivable
Good for a single fluent taskGood for production multi-step machines

The field’s own shorthand: a one-shot agent is a demo; a production agent is a graph. When outputs matter and failures are costly, you reach for verifiers and gates. When the task is deterministic and single-step, a loop is enough.


Canonical Patterns (Anthropic)


Anthropic’s graph engineering guidance defines five patterns that recur across every well-built agent graph. Treat these as the vocabulary of the discipline.


1. Node Contract


Every node declares three things up front:


  • IN — the typed parameters the node accepts and requires
  • OUT — the typed results the node guarantees to emit
  • CONTRACT — the invariants that must hold on inputs and outputs

Contracts are what let a graph fail loudly and predictably. If a node cannot meet its contract, it returns a bounded error (never an untyped crash) and the graph either retries, routes, or halts. Example:


NODE: revenue-trend IN: { period: str, db_path: str } OUT: { total_revenue: float, order_count: int, aov: float, mom_pct: float, top_products: list, anomalies: list } CONTRACT: - total_revenue must match orders-table aggregation - anomalies only when deviation > 2σ from 90-day mean - if no data → return empty arrays, do not fail

2. Verifier Node


A verifier is a guard that checks actual output against expected output. It answers one question: did the upstream node actually do what it claimed? It can compare to a schema, an invariant, or a downstream confirmation (e.g., a rank check in search). If it fails, the graph retries, escalates, or stops — it never silently passes broken output downstream. This is the node that turns "trust the agent" into "trust but verify."


3. Fan-In Guard


When a graph fans out into parallel work, the fan-in point must guard what it recombines: cap the number of concurrent branches, merge redundant or partial results, and order or dedupe sub-fragments before synthesis. Without a guard, parallel fan-out produces unbounded token usage, duplicate work, and incoherent merges.


4. Conditional Edge


Edges are not always fixed. A conditional edge routes to the next node based on the content of the current output, not a hard-coded order. This is what makes a graph an adaptive decision machine rather than a fixed pipeline. The router node in the Nordisk revenue graph is a conditional edge: it reads the metrics and picks diagnose, alert, create-task, or log-brief.


5. Layered Fan-In


For large outputs, fuse results hierarchically rather than all at once: group branches by category, summarize each batch, then synthesize a final report from the batch summaries. This controls token usage and keeps large fan-out outputs coherent. This is the dedupe → synthesize pair in a many-agent graph.


Six-Step Playbook


  1. Build a loop. Get one linear execution correct first.
  2. Wire the loop into a graph. Make the steps explicit nodes with typed IN/OUT/CONTRACT.
  3. Add fan-out. Parallelize independent branches. Add a fan-in guard.
  4. Add verification. Insert verifier nodes at each risky boundary.
  5. Add human gates. Give humans an approve/reject/revise decision at high-consequence steps.
  6. Add dedupe + caps + report. Harden the edges: hard caps, partial-failure detection, and an evidence report on every run.

This is the order the content-ops graph and the Nordisk revenue graph both followed: build the loop, then level it up pattern by pattern.


Architecture Gallery


Revenue-Ops Graph (Nordisk)


[Scheduler] | -------------+------------- | | | [Harvest [Harvest [Harvest GSC] Shopify] Klaviyo] \ | / \-----------v-----------/ +---------------------------+ | Verifier Node | | harvest ok? data fresh? | | no → retry/log/alert | +------------+--------------+ | +------------v--------------+ | Fan-Out Node | | spawn parallel analysts: | | GSC | Revenue | Ads | Inv | +----+--------+-----+-------+ | | | v v v [Agent] [Agent] [Agent] (parallel) | | | +----+---+-----+ v +---------------------------+ | Dedupe + Cap + Report | | flatMap, hard cap, | | partial-failure detect | +------------+--------------+ v +---------------------------+ | Router Node | | revenue_drop → alert | | ads_zero → pause | | gsc_opp → content task | | normal → log + brief | +----+----------------+-----+ | | [Alert (TG)] [Action / Report]

Content-Ops Graph (verifier-gated autopublish)


Observe → Discover → Research (parallel fan-out) → Merge → Draft → Verify (3 skeptic verifiers) |-- pass → autopublish --> Record |-- revise → back to Draft →-- escalate → human gate Circuit-breaker: revert on rank-drop. Content KG in MemPalace.

Applied live to the 47-post Nordisk Shopify blog on a weekly cron. All nodes are implemented under /root/loops/nordisk-content-graph/: observe, discover, research, draft, verify, publish, revert, and record, plus research_node.py, test_nodes.py, spec.md, and a .circuit_breaker guard.


Typed DAG Scheduling (MoH Harness)


For the heavier scheduling end of the discipline, see the MoH page: a typed DAG with heterogeneous executors, runtime-schedulable as a graph rather than a fixed pipe. Graph engineering and typed DAG scheduling share their spine — nodes, contracts, edges, and gates — but a DAG is compile-time scheduled while a graph is runtime-routed.


Node Contract Template


NODE: <name> PURPOSE: <one sentence> PARENT: <graph this belongs to> DEPENDS_ON: <nodes that must complete first> IN: <param>: <type> — <description> OUT: <field>: <type> — <description> CONTRACT: - <invariant that must hold> - <validation rule> ERRORS: - <error condition> → <what happens> - <error condition> → <what happens>

Pitfall Catalog


Every audit of a graph should check for these recurring failure modes:


Stale-decision bugA conditional edge routes on data that has gone stale since it was read. Fix: timestamp every decision input and verify freshness.
Phantom nodesNodes that exist in the manifest but are never reached, or edges that point to deleted nodes. Fix: validate graph topology at load.
Hash-based change-detection bugsChange detection keyed on content hashes that ignores semantic change. Fix: verify the actual semantic invariant, not the hash.
Non-atomic state writesState written in pieces so a crash leaves the graph in a half-applied state. Fix: make each state transition atomic (single transaction).
Silent verifier passA verifier that always returns true because its invariant is weak or never exercised. Fix: unit-test verifiers with known-bad inputs.
Unbounded fan-outParallel branches with no cap or merging, blowing the token budget. Fix: hard cap + fan-in guard.

Audit Checklist


To review any agent graph, run it through the five-pattern lens and the playbook order. Two audit modes are standard:


  • Mode A (inline): one pass over every node, checking that IN/OUT/CONTRACT are explicit, verifiers are real, fan-in is guarded, edges are conditional where needed, and layered fan-in is used on large outputs.
  • Mode B (three parallel MoE reviewers): dispatch three independent reviewers over the graph, then synthesize their findings into a single verdict with evidence.

Good graph: every node declares a contract, every risky boundary has a verifier, fan-out is capped, edges route on content, and every run writes an evidence report.

Bad returns: contract-less nodes, verifiers that always pass, unbounded parallel work, fixed edges, and state writes that are not atomic.

When Graph, When Loop


Use a graphUse a loop
Many heterogeneous stepsSingle fluent task
You want verification / cross-checkingDeterministic single answer
Parallel work with fan-outSequential by nature
Complex shared stateStateless, ephemeral
Failure is costly, must surviveLow cost of a bad run

Most production machines need a graph. Most individual agent calls need only a loop. The skill is knowing when a run must become a machine.


Tool Landscape


CategoryToolsRole
Agent graph runtimesLangGraph, Meridian, PrefectNode/edge execution, conditional routing
Durable orchestrationTemporal, Dagster, Airflow, AWS Step FunctionsLong-running, retrying, scheduled DAGs
Declarative graph enginesFluxtion, FLAMECompile graphs statically for performance
Graph DBs / knowledge graphsNeo4j, Memgraph, MemPalace, enc1Store the knowledge graph agents traverse

"Graph engineering needs a compiler" (Fluxtion, Jul 2026) argues the next step is statically compiling a graph’s nodes and edges into a runtime that is not re-parsed on every run. That is the frontier: from interpreted graph to compiled graph.


Your Stack (enc1 · MemPalace · Nordisk)


Graph engineering is not abstract here — it is the load-bearing pattern behind the whole system. The mapping:


Knowledge graph: the enc1 ambient D3 graph (597 nodes / 918 links) and MemPalace store the shared memory that agents structure their work over. This page is itself a node in that graph.
Agent graph: You (Maestro) route the task → pi does tool work, Prime Agent does autonomous reasoning, Hermes does research → a verifier double-checks critical claims → outcomes land in the MemPalace knowledge graph.
Applied graphs: the content-ops graph (verifier-gated Shopify autopublish, live on weekly cron) and the revenue-ops target graph (schedule → harvest → verify → fan-out → analyze → dedupe → synthesize → router → act).
Layered kit: framework (graph-engineering skill) → patterns (graph-pipeline-audit skill) → applied example (content-ops graph) → working runtime (4-layer orchestrator) → Nordisk targets → outer loop (GraphOpt).

GraphOpt — The Self-Re-Orchestrating Outer Loop


The frontier of graph engineering — and what makes this stack self-learning — is GraphOpt: the outer loop that watches run history and re-tunes how the next run is orchestrated. A graph defines the topology of one run; GraphOpt learns the topology from the accumulated ledger of runs. This is the concrete implementation of Karpathy's "two loops" and Gloqo's "outer hill-climb from production traces."


Layer 2 GRAPHOPT ---- watches run history ---- | re-tunes topology / params Layer 1 GRAPH = one run's node + edge plan | Layer 0 LOOP = one linear execution GraphOpt meta-loop (one pass): OBSERVE → DRIFT-check → ANALYZE → GUARD → REPLAY (old vs new on real traces) → APPLY (compile to live-graph vN+1) → REPORT All gated behind the safety fence (below).

The meta-loop


  1. OBSERVE — ingest per-node run records from the ledger (append-only JSONL).
  2. DRIFT — the drift anchor re-checks the graph's success signal against ground truth. If reality diverges, auto-apply freezes.
  3. ANALYZE — a rules engine turns ledger evidence into structured topology deltas.
  4. GUARD — classifies each delta: tame (auto-eligible) vs structural (human review).
  5. REPLAY — auto deltas must WIN against the old topology on real, already-observed holdout traces before applying (the proof gate).
  6. APPLY — accepted deltas compile to live-graph.json version N+1.
  7. REPORT — metrics written for dashboards; a pending-review queue holds structural proposals.

Where it lands on this stack: the revenue-ops graph (harvest → verify → fan-out → analyze → dedupe → synthesize → router → act) currently has its retry counts, fan-out width, and alert thresholds hardcoded in the script. GraphOpt learns those from run history — e.g. "the verifier needed a retry on 4/5 runs, so raise retry_max;" "fan-out cost blew the token cap, so lower fan_out_width." The same applies to the content-ops and skillopt graphs.


The safety fence (immutable)


GraphOpt is designed for safe self-modification. Three fences rest on the self-modifying-systems literature (arXiv 2606.23075 guardrail erosion; eval-collapse / drift analysis):


1. Tame auto-applies, but only if proven. Only TUNE deltas on safe tunables (retries, fan-out width, timeouts, caps) — and only when the replay guard shows the new topology wins on real traces.
2. Structural changes need a human. Adding/removing/editing nodes or edges is parked in pending-review.json for explicit approval. A self-rewriting system must not rewrite its own guardrails.
3. The reward signal is off-limits. GraphOpt can never change what a node contract checks, what a threshold means, or what counts as "success." A machine that redefines its own reward is a machine that learns to game itself. Contract edits are rejected outright.
4. Drift anchor freezes everything. If claimed successes stop being confirmed by ground truth (ratio below 0.70), auto-apply freezes until a human confirms the signal.

Implementation: /root/skill-opt/graphopt/ — lib/ledger.py, lib/deltas.py, lib/analyzer.py, lib/guard.py, lib/replay.py, lib/drift_anchor.py, and orchestrator.py. Skill: graphopt (installed in both Hermes compound-engineering and Pi). Run python3 orchestrator.py --demo to see the full loop, including the three safety fences.


Key Sources

  • Shann Holmberg (Jul 2026): "the loop finds the path, the graph defines the path" — 1.9K-bookmarked thread
  • Harrison Chase (@hwchase17): is graph engineering a real category or just langgraph?
  • Anthropic: graph engineering write-up + video (Jul 2026) — 5 patterns + playbook + audit
  • Carlos E. Perez: From Loop Engineering to Graph Engineering
  • Peter Steinberger: "Are we still talking loops or did we shift to graphs yet?"
  • Gloqo AI (Jul 2026): Loop Engineering, Graph Engineering, and the Hill-Climb — the outer hill-climb from production traces
  • iii.dev (Jul 2026): Loops, Graphs, and the Layer That Matters
  • Fluxtion (Jul 2026): Graph Engineering Needs a Compiler
  • ChrisLema (Jul 2026): What ‘Loops to Graphs’ Looks Like in Production
  • Karpathy (via r/PromptEngineering, Aug 2026): "Two autonomous agent loops made my loop 1000x better with graph engineering"
  • arXiv 2606.23075: Safety in Self-Evolving LLM Agent Systems — guardrail-erosion collapse, the basis for GraphOpt's safety fence
  • Eigent.ai: Graph Engineering for AI Agents — sensor/eval drift, "scores keep climbing while reality diverges"

Connections

  • GraphOpt (outer loop) — this page, above, at GraphOpt — The Self-Re-Orchestrating Outer Loop; implementation in /root/skill-opt/graphopt/
  • MoH — Typed DAG Scheduling
  • Complex Adaptive Systems
  • General Systems Theory
  • Systems Dynamics
  • Retroduction
  • First Principles Thinking
  • Cliodynamics Deep Brief
  • Vannevar Bush
  • J.C.R. Licklider
  • Andrey Kolmogorov
  • Peter Turchin

  • See also

  • General Systems Theory
  • Complex Adaptive Systems
  • Systems Dynamics
  • Retroduction
  • First Principles Thinking
  • MoH — Typed DAG Scheduling