rotate for the full layout

Graph Engineering

0–2Where this sits

AI engineering

Accumulation, not replacement. Each layer fixes what the one below it could not.
01PromptOne request. Role, examples and the output contract — everything you can say in a single turn.
02ContextEverything it sees at call time — what you retrieve, what you rank first, and what you cut.
03HarnessEverything outside the model for one run: tools, state, timeouts, retries, approval gates.
04LoopPlan, act, verify, repeat — and it stops when a test passes, not when the output looks right.
05GraphMany nodes over one shared state — and the edges between them written in code, not improvised.
2–4The whole idea, in one picture

You already know this shape. It is triage.

Two shapes: one node with an edge back to itself, or many nodes with a test in the middle. NHS Accident & Emergency has solved this.

Loop engineering — the lone doctor

One agent, one context window. Plan, act, verify, repeat — then go round again.

one node · one context · one thing at a time one doctor · one patient · one thing at a time first · then · then not verified? go round again still not sure? go round again plan history act examine verify one test

Graph engineering — the NHS A&E team

Many nodes over one shared state, with the edges between them written in code.

three workers · one superstep three tests · ordered together, not in turn split triage worker bloods worker X-ray worker ECG validate the result deterministic objective human consultant acknowledges decides a code node settles it · the graph stops at a human the test settles it · a person still decides
4–6The definition

The loop lives inside a node. The graph lives between them.

One definition, and the three parts it is made of.
NODE NODE a loop inside a loop inside EDGE reads / writes reads / writes SHARED STATE one record · every node reads and writes it

Graph

The boxes and the arrows — and you draw both, before anything runs. The model does the work inside a box. The graph decides which box is next.

Node

One job with one worker on it — a model with its own context, tools and permissions, or plain code. Inside, it may loop as long as it likes. Outside, it is one input and one output.

Edge

A dependency — not “and then”. This node’s output is that node’s input. The test: can you name the variable that crosses it?

Five kinds, and only five: sequential · fan-out · join · conditional · human
Fan-out and join are where the parallelism lives — it comes from the edges, not from the nodes.

Shared state

One typed record every node reads and writes, so evidence accumulates in one place instead of four transcripts. Its fields merge rather than overwrite — which is what lets three nodes write in the same superstep without racing.

An actual use case.

8–9Why this was built

MuleSoft Community Edition. No vendor support.

We run 4.6.1, on-prem, two instances. Every escalation path has a rung above it. Ours does not.
THE ESCALATION PATH incident L1 L2 WHERE VENDOR SUPPORT WOULD BE Mule Community Edition — there is no contract, and no ticket to raise Mule 4 Triage Command Graph three specialists in parallel · a real build that settles it · a human gate it produces the evidence-backed RCA a vendor would have produced

No ticket to raise

No vendor RCA, no backport to ask for, no one to escalate to. The path ends with us — at 2am, on the night the interface is down.

No platform telemetry

No Runtime Manager, no Anypoint Monitoring. The evidence is files on a box — logs, a pom, per-env YAML — not a dashboard that correlates them for you.

No vendor knowledge base

Every known error is one we wrote down ourselves. Miss that, and the next occurrence is researched from scratch by whoever is on call.

This is the escalation we cannot raise — evidence gathered in parallel, a conclusion a real build has already tested, and a person who still decides.
9–10Why Mule incidents are hard by hand

One incident. Six places the answer could be.

Different files, different owners, different tools — and no one person reads all six.

Logs

Build and application output — and, just as often, which log does not exist.

build.log · app.log

Previous incidents

Has this signature been seen before, and was it ever written down?

known-errors/*.md

Codebase

Flow XML: which attributes were set, and — harder — which were quietly omitted.

src/main/mule/*.xml

Dependencies

Connector and runtime versions, enforcer rules, token substitution.

pom.xml

Environment config

The prod-versus-uat delta nobody diffs until something breaks.

*-prod.yaml · *-uat.yaml

MuleSoft runtime

What the platform assumes when you say nothing at all.

mule-artifact.json
github.com/mulesoft/mule @ 4.6.1
By hand you read those one after another, and the first plausible story usually wins.
The failure mode is not ignorance. It is premature conviction — and a single loop is a machine for producing it faster.
10–12The program

This is the program — not a picture of one.

The triage agent’s flow, compiled — what the spec asks for by discipline, the graph enforces by topology. Drawn from the declaration the code compiles; a test fails if they disagree.
Tool
IngestAlert
no log, no triage — reads the manifest, stops if nothing reads
Agent
Discover
sanitised signature · known errors · onset vs deploy · absence is evidence
fan-out
Agent
FlowTrace
flow XML — every attribute set, and every one omitted
Agent
BuildDependency
pom — runtime, connector and plugin versions · enforcer rules
Agent
ConfigConnector
per-env yaml delta · runtime defaults verified from source
ONE STEP
join
Deterministic
Diagnosis
hypothesis matrix, elimination — no model guess
Validator
Validation
real Maven, on a disposable copy
conditional edge — Validation reads the state and picks: on refutation, back up to Diagnosis (the only cycle)
in the agent: Opus 4.8 first pass · a second model and a judge only on SEV1/2 or contested
Judgment
Judgment
cause · routing · rollback · regression guard — recommend only
always (policy) — no path skips this
Gate
HumanGate
acknowledge the diagnosis package
on acknowledge
Tool
OutcomePack
shareable brief · Jira draft · known-error record for next time
12–19One incident, end to end

Watch it run. One superstep at a time.

EVIDENCE LEDGER · ONE SHARED STATE SUPERSTEP1 / 8EVIDENCE1 rowsHYPOTHESES—INGESTINGMEANWHILE, THE SAME WORK AS A LOOPreading source 1 of 6TOOLIngestAlert2msinbuild.log · deploy-history.txt · MANIFEST.md loaded SUPERSTEP2 / 8EVIDENCE3 rowsHYPOTHESES4 openDISCOVERINGMEANWHILE, THE SAME WORK AS A LOOPreading source 2 of 6AGENTDiscover3 evSIGmaven-enforcer RequireProperty failedCPdeploy inside burst window — shared build template rolled out SUPERSTEP3 / 8EVIDENCE10 rowsHYPOTHESES4 open3 AGENTS ‖ ONE SUPERSTEPMEANWHILE, THE SAME WORK AS A LOOPreading source 3 of 6 — one at a timeFlowTraceBuildDepConfigConnF1unsubstituted token @http.source.service@ → mainflow.xml:11B1enforcer requireProperty fail=true → pom.xml:89C1prod vs uat delta → adapter-httptojms-prod.yamlE2ABSENCE — no runtime evidence exists → app.log, gc.log SUPERSTEP4 / 8EVIDENCE12 rowsHYPOTHESES1 survivesELIMINATINGMEANWHILE, THE SAME WORK AS A LOOPreading source 5 of 6DETERMINISTICDiagnosis<1msH1survives — token never substituted · H2 H3 H4 eliminated SUPERSTEP5 / 8EVIDENCE12 rowsHYPOTHESES1 survivesREAL MAVEN BUILDMEANWHILE, THE SAME WORK AS A LOOPfirst plausible cause — unverifiedVALIDATORValidation1.5smvngenerate-resources → exit 1 · incident REPRODUCEDmvn-Dhttp.source.service → exit 0 · solution VALIDATED SUPERSTEP6 / 8EVIDENCE12 rowsHYPOTHESES1 survivesJUDGINGMEANWHILE, THE SAME WORK AS A LOOPrestating it, more confident each passJUDGMENTJudgment<1msRCAbuild template stopped passing the property · High · route: pipeline owner SUPERSTEP7 / 8EVIDENCE12 rowsHYPOTHESES1 survivesHALTED — AWAITING HUMANMEANWHILE, THE SAME WORK AS A LOOPno gate — it would simply proceedGATEHumanGatepress AGATErun halted · implementation_performed = false SUPERSTEP8 / 8EVIDENCE12 rowsHYPOTHESES1 survivesCOMPLETE — ADVICE ONLYMEANWHILE, THE SAME WORK AS A LOOPstopped: budget exhausted, not solvedTOOLOutcomePack2msOUTRCA brief · Jira draft NOT raised · residual unknowns published
19–20Close
Loops make one worker reliable.
Graphs make a team reliable.
Appendix · the pivot, if the live run could not happen

All we do differently is write it down.

You have been running graphs all year. Nobody drew one, nobody tested one, and nobody could point at one afterwards.
PHASE 1 your one ask PHASE 2 a plan PHASE 3 reads & searches PHASE 4 edits & tool calls PHASE 5 one result nobody selected a single one of these
The harness built that graph and routed the work.
Written down, it can be pointed at, tested, and gated. That is the rest of this talk.
Appendix · why this is safe to run

Open the gate on blast radius, not on confidence.

Confidence is the weakest input in that decision — it is the only one the model can influence.

Reversible and contained

A copy change, a test, an isolated function with coverage. One bad merge costs a revert.

→ This lane can open first.

Reversible but wide

A shared utility, a schema addition, anything a dozen callers touch.

→ Opens on deterministic checks plus a clean trajectory.

Hard to reverse

Migrations, deletions, anything writing to production data — or a change to a live integration between SAP and the storefront.

→ This lane does not open.
Mule integration changes sit in the third lane. So diagnose-only is not caution — it is the correct lane for the blast radius.
And it is enforced in code, not in a prompt. A threshold is a number someone eventually adjusts. A closed lane is not.
Appendix · the design, decision by decision

Six decisions — and not one of them is a prompt.

Each is a property of the graph, so it happens on every run — not when the model remembers to.

Evidence is parallel

Flow, build and config are independent sources, so they are one superstep. A loop has one cursor — and the order it picks quietly becomes a hypothesis.

FlowTrace ‖ BuildDependency ‖ ConfigConnector

Elimination is structural

A hypothesis survives only if supporting evidence was found and nothing refutes it. That is a property of the node, so it happens on every run.

Diagnosis — the 0.0s you will see live

Validation is unavoidable

It sits on the only path from diagnosis to judgment. A root cause you cannot reproduce is a hypothesis, and the output says so.

Validation — a real Maven build

A refutation comes back

The one cycle in the graph. Return the unit, not the batch — the specialists do not re-run, because nothing they saw changed.

Validation → Diagnosis

The gate is an edge, not a mood

The run stops because there is no edge forward until a person acknowledges. Confidence does not unlock it.

route_after_judgment — unconditional

Policy runs between steps

Supersteps are discrete, so the protected trees are re-hashed after every one. A node that edited the application stops the run.

policy.py · 127 tests
The most important edge is the one that is not there: nothing leads from Judgment to an implementation.
Appendix · two RCAs and a judge

Two RCAs and a judge — on escalation, not by default.

Both passes read the same Evidence Ledger, and neither sees the other’s conclusion before the judge does.

RCA-A — the default

One pass. Hypothesis matrix, eliminate with evidence including absence, one surviving cause, and the smallest layer-scoped fix. Most SEV3/4 faults close here.

Opus 4.8

RCA-B — the challenger

A separate, independent RCA from the same ledger. It never sees RCA-A — no cross-talk, no anchoring. Runs only when a trigger fires.

GPT-5.5 · SEV1/2 · contested reading

Judge — the consolidator

Evidence over eloquence. Adopts where they agree, resolves disagreement by the stronger verifiable evidence — never averages. May demote an unverified claim, or send a pass back to read a missing source.

Sonnet 4.5 · owns the final verdict
The cost gate: below-High confidence is almost always a missing input, not insufficient reasoning. Name it, enrich the ledger, re-run RCA-A once.
Two models over the same thin ledger produce two equally thin answers — and a judge forced to pick one.
Appendix

Run it yourself.

Start the demo

cd graph-engineering
./mule-triage-command-graph/run.sh
# http://localhost:7080

Prove the graph is real

cd mule-triage-command-graph
# compiled nodes and edges
../.venv/bin/python runner/inspect_graph.py

# LangGraph’s own checkpoints
../.venv/bin/python runner/inspect_graph.py \
  --db .triage-work/graph.sqlite

Sources

Loops and Graphs — @hanakoxbt, Aug 2026. The correction edge, and gating on blast radius rather than confidence.
Loop vs Graph engineering — @Sumanth_077, Sep 2026. The loop/graph distinction, and a loop as one node inside a graph.
From Loops to Graphs — @kirillk_web3 / Kimi K3, Sep 2026. “One node does the work, another watches it.”
从 Loop 到 Graph — 15-minute walkthrough, Sep 2026. The five-layer stack.
Internal Confluence Issues & Solutions, space SYSINT.

How diagnose-only is actually enforced

Hash. Protected trees are content-hashed after every superstep; any change raises PolicyViolation and stickily halts the run.
Scan. The outcome is rejected if it carries an applied, merged or deployed marker, or if implementation_performed is not false.
Confine. Writes are restricted to a scratch directory; validation runs on a copy, because the archetype’s own build rewrites source in place.
Prove. A test asserts the drawn topology equals the compiled graph — that is how we found an advertised gate bypass.

127 tests · runner/policy.py

Known caveat — say it if asked

The primary scenario MT-001 is a documented fallback. The originally intended primary — the last Mule 4 issue validation — could not be recovered from session memory, which returned no rows. MT-001 was chosen deliberately, not silently.

Speaker notes

All slides — click to jump

Enter→Space next ← back S speaker notes D fallback view O overview F full screen T light Home restart