Home

MoH — Typed DAG Scheduling with Heterogeneous Executors

MoH — Mixture of Harnesses
RoomSystems
TypeArchitecture Pattern
StatusDesign Phase — Iteration 4
Date2026-08-16
SupersedesHarness MoE (2026-08-13), MoH Iteration 2 (2026-08-15), MoH Iteration 3 (2026-08-16)
Prior ArtTemporal · Dagster · Bazel · Blackboard (HEARSAY-II) · Airflow · Nix · AWS Step Functions
Full Analysismoh-architecture.md (Iteration 2, superseded)
What changed in Iteration 4: Added adopt-vs-build decision section with explicit comparison to Temporal/Dagster. Added v0.5 thesis experiment (Hermes + pi heterogeneous composition). Replaced empty merge with typed assemble_inputs. Sticky degraded propagation as formal rule. Fixed dedup identity (stable occurrence key, not rule_version). Added Layer 2/3 boundary (structural checks only in runtime). Stated coverage proof limits explicitly. Distinguishes reviewer evidence by independence tier.

MoH is a typed DAG scheduling architecture with heterogeneous executors. Its unit of composition is a work contract — a typed description of what's needed, what authority is granted, and what acceptance criteria must pass — not a token stream or a gated routing decision.

The architecture routes between components of fundamentally different kinds — deterministic local execution, stateful analysis, and external research — whose outputs (file diffs, analytical decisions, and evidence bundles) cannot be averaged, blended, or gated like neural logits. The correct operation is decomposition into a typed DAG, not classification to a single expert.


Lineage and Prior Art

MoH stands on solved ground. The following prior art is explicitly acknowledged:

SystemWhat MoH borrowsWhat MoH does differentlyRisk if ignored
Temporal Deterministic orchestration, activity workers, Saga compensation, event-sourced state, idempotency keys, heartbeat, retry/lease semantics Typed plan representation with pre-execution validation; capability-based dispatch instead of task-queue routing Partial-failure rollback requires Saga; deterministic replay requires hermeticity barrier
Dagster Typed op I/O, asset checks (user-defined validation that blocks downstream), Resources for infrastructure abstraction, type-check before execution Heterogeneous executor routing (not just data ops); typed contracts between unlike harnesses Without typed I/O validation, type mismatches become runtime failures
Bazel Hermetic actions with declared reads/writes/effects, content-addressable output caching, effect tracking via dependency declarations Non-hermetic actions are first-class (LLM calls, web research) with explicit evidence labeling Non-hermetic nodes cannot be cached or deterministically replayed; must be labeled as such
HEARSAY-II (Blackboard) Typed knowledge sources with typed contributions to a shared blackboard; data-driven scheduling Contract-first static plan (vs. fully data-driven); typed receipts instead of shared mutable blackboard Blackboard scheduler bottleneck proved that central dispatch needs meta-reasoning or becomes intractable
Airflow Static DAG as plan representation; XCom for data passing; trigger rules for downstream dependencies Typed contracts instead of untyped XCom; capability-based executor routing instead of queue names Static DAG cannot replan; missing nodes are silent; CeleryKubernetesExecutor hybrid was abandoned
AWS Step Functions Explicit error handling with Catch/Retry on each task; Saga compensation via state machine modeling Typed join semantics; provenance tracking across service boundaries Compensation must be modeled in the plan, not assumed by the runtime

Three Layers, Separated

MoH is not a single system. It is three distinct layers that should be built and reasoned about independently:

LayerWhat it doesWhere LLMs livev0 scope
1. Plan representation Typed work contract → validated DAG. Schema for nodes, edges, types, contracts, effects, authority. Optional: LLM for contract-to-graph decomposition. Validator is deterministic. Static plan templates only. Validator checks coverage and type compatibility.
2. Workflow runtime Dispatch → receipt collection → retry → state machine → compensation → stop. Append-only event log. None. This is deterministic infrastructure. Envelope + receipt format. Single sequential pipeline. No retry, no compensation.
3. Agentic decision layer Decomposition, synthesis, re-planning, conflict resolution. The meta-reasoning around the graph. Primary. Plan(), join(), check() are judgment calls that use LLMs in practice. DEFERRED. Not needed until v1.

Claude (critic): "No fourth agent" isn't true. Plan(), join(), and check() are all judgment calls that will be LLM calls in practice. You've got a fourth agent named "control plane." Better to admit it and scope it explicitly — least privilege applies to the planner too.

Accepted. Layer 3 is the fourth agent. It is explicitly named and deferred. The control plane (Layer 2) is deterministic.

The Three Harnesses

The three current harnesses are roles, not kinds. The architecture routes on declared capabilities + authority, not on identity. If pi gains network access, or Hermes gains a file-write capability, the topology changes — and that's fine, as long as capability manifests are explicit.

HarnessNatural JobCapabilitiesOutput TypesFailure if Misused
pi Deterministic local execution read, write, shell, test (no network) file_diff, command_result, test_receipt Makes architectural judgment or produces weak research
Prime Agent Planning, reasoning, stateful analysis read, write, bash, analysis (persistent kernel) plan, report, dataset, decision Over-engineers simple edits; unbounded autonomy
Hermes External research and signal extraction web_search, fetch_content, bash (read-only), no write evidence_bundle, source_collection, findings Changes local state or treats weak signals as facts

Claude (critic): Route on declared capabilities + authority and the design survives harness #4, or pi gaining a web tool.

Accepted. Harnesses are capability manifests, not proper nouns. The three listed here are the current set. Adding a fourth means writing a capability manifest and an adapter, not changing architecture.

MoH Architecture — The Five Primitives


┌──────────────────────────────────────────────────────────────┐ │ META-HARNESS CONTROL PLANE │ │ (Deterministic — Layer 2) │ │ │ │ Five primitives: │ │ │ │ plan(contract) → graph │ │ Decompose a work contract into a typed task graph. │ │ v0: static template + validator. v1: LLM optional. │ │ │ │ run(node, envelope) → receipt │ │ Dispatch one graph node to the right harness. │ │ Envelope: scoped authority, inputs, budget, acceptance. │ │ Receipt: outputs, duration, provenance, errors. │ │ │ │ compose(receipts) → typed artifact │ │ Structural composition of unlike outputs with provenance. │ │ NOT a join/vote/average — different types compose │ │ through explicit type rules. │ │ │ │ commit(artifact, workspace) → effect │ │ Apply a typed effect to the workspace (file diff, │ │ notification, report write). Separated from compose │ │ because composing evidence ≠ applying a mutation. │ │ │ │ check(artifact, acceptance) → verdict │ │ Verify against acceptance criteria. │ │ Modeled after Dagster Asset Checks. │ │ Returns pass/fail/degraded. │ │ │ │ stop(reason, workspace_state) → terminal │ │ Clean or dirty termination. Records workspace state. │ │ Does NOT claim "no partial state." │ │ v0: records dirty workspace. v1: compensation per node. │ │ │ │ ──────────────────────────────────────────────────────────── │ │ State: append-only event/receipt log + typed artifact refs │ │ Policy: declarative config (routing, retry, gates, auth) │ │ Human: confirmation before irreversible mutations │ └──────────────────────────────────────────────────────────────┘

Type Algebra: Compose ≠ Commit

A core finding from the Iteration 3 review: join() was doing double duty. It was both "combine data" and "apply effects." These are separate operations with separate failure modes.

MoH defines a small type algebra with five composition operations:

OperationInput TypesOutput TypeSemanticsFailures
assemble_inputs(a, b, ...) Named slots {gsc, shopify, klaviyo, ads} each typed T snapshot Collect named source receipts into a record preserving each source's schema, status, freshness, and provenance. Missing slot is visible omission, not empty union. Missing required slot → incomplete snapshot. Degraded source → degraded snapshot. Never silently zero.
chain(a → b) A, B (different) C Feed A's output as input to B's transformation. Like Temporal's ExecuteActivity chains. Type mismatch at boundary → validation error before execution.
attach(evidence, claim) evidence_bundle, claim supported_claim Attach source citations to a claim. Produces a supported_claim with provenance chain. Evidence doesn't actually support claim → mark as "claim unsupported" rather than lying.
apply(diff, workspace) file_diff, workspace_id effect_receipt Apply a file diff to the workspace. This is a COMMIT, not a compose. Conflicting diffs → stop and merge. Workspace dirty on failure.
wrap(receipts, metadata) [receipt], metadata final_package Package multiple receipts into a deliverable with aggregated metadata, costs, and provenance. Missing receipt → incomplete package (flagged, not hidden).

Key rule: assemble_inputs, chain, attach, and wrap are compositions — they combine data without side effects. apply is a commit — it mutates workspace state. The runtime must distinguish these and require confirmation gates on commits. Homogeneous merge is intentionally absent: four harvest receipts share an envelope but are semantically incommensurable; each source must be handled as a named input.


Known Gaps and Remediation

The following issues were identified during Iteration 3 review (August 2026). They are ranked by remediation priority.

P0 — Ship-blocking

IssueDetailFixSource
Sticky degraded propagation No formal rule for how degraded status flows through the graph. A degraded harvest arrives at route_alert looking fine. Zero-vs-unavailable bug reappears one layer up. Degraded is sticky by default: any node consuming a degraded input returns degraded unless its contract explicitly declares tolerance. Tolerance must be named, typed, observable. Node receipt lists degraded input IDs and affected output fields. route_alert may emit data-quality alert but not business alert from partial data unless policy permits. Claude v2 + Prime v2
Credential in graph script API key embedded in source code in git history Rotate key. Load from environment/secret manager. Remove from git history. Prime audit
JSON/stdout mixing Machine JSON and human status text emitted on stdout; downstream JSON consumers break Strict JSON to stdout, human diagnostics to stderr Prime audit
Zero vs unavailable ambiguity Zero-row query returns same shape as missing/unavailable source; alert rules can't distinguish Required-source failure produces degraded state, not zero rows. Harvest receipt includes row count, max timestamp, error class. Prime audit + pi

P1 — Design-level

IssueDetailStatusSource
Plan omission undetectable If plan() omits a needed node, check() passes and output is wrong. Coverage proof catches this only after the deliverable is declared. If the contract itself under-declares what's needed, coverage is vacuously satisfied and everything passes. Coverage proves graph/contract completeness relative to declared intent. Does NOT prove intent is complete. Humans own intent completeness. Add intent-review step: review acceptance criteria, enumerate required decisions, map risks to required evidence. Claude 2 + Claude v2
No replan primitive Static DAG can't absorb surprises (stale data, revoked auth, schema drift). Temporal, Dagster, and Step Functions all handle this. v0: bounded restart (fresh snapshot, one re-execution). v1: dynamic graph expansion with depth budget. Claude 3 + Hermes research
stop() overclaims "No partial state" is impossible when pi has written files or run processes. Needs compensation per node. v0: stop(reason, dirty_workspace). Records what's incomplete. v1: Saga compensation registry per node type. Claude 5 + Temporal Saga research
Deterministic replay overclaimed LLM planners and live web research are not reproducible. Audit log ≠ replay. Rename to "audit trace." Simulation replay only for deterministic subgraphs. Label non-hermetic nodes explicitly. Claude 6 + Bazel hermeticity research
Missing fast path A one-line edit shouldn't pay full DAG overhead (contract → planner → validate → dispatch → compose → check) v0: if task fits ≤2 pi nodes, skip planner entirely. Direct dispatch. Claude 9 + pi
Source contradiction ranking "Rank by authority/date" for contradicting sources will hide quiet errors. Needs explicit source policy with uncertainty modeling. v0: flag contradiction, don't silently rank. v1: source policy DSL. Claude 10 + HEARSAY-II lesson
Layer 3 leaking into check() Citation quality, evidence-support checks, and anomaly-basis judgment are semantic. If they become runtime primitives, the deterministic control plane quietly acquires a model. Hard rule: Layer 2 checks are structural only (receipt status, schema, dedup keys, hashes, provenance links, artifact existence). Semantic judgment is always a graph node dispatched to a harness, producing a typed receipt. Claude v2
Adopt-vs-build undecided Gap list (compensation, heartbeat, versioning, retry, event sourcing, asset checks) describes Temporal + Dagster. No explicit rationale for building Layers 1+2 vs adopting. See adopt-vs-build decision section. Default: adopt for execution, own the semantic type system. Claude v2

P2 — Nice to have

IssueDetailStatus
Compensation for non-idempotent nodes Some harness effects (e.g., "send email") cannot be undone. Partial-failure compensation needs explicit modeling per effect type. Deferred to v1. v0: mark non-compensatable nodes and require human confirmation before dispatch.
Worker versioning Temporal's "many-versions problem" — when harness code changes mid-workflow. Not addressed in current MoH design. Deferred. Version receipts include harness version hash. Detection = feasible, resolution = deferred.
Heartbeat / liveness No detection of hung harness processes. Step Functions has heartbeat timeout; MoH doesn't. Deferred. v0: wall-clock budget only. v1: per-node heartbeat.

v0 Boundary: The Honest Pipeline

Build this first. Nothing else matters until this works.

The v0 is a deterministic, read-only revenue-alert graph. No autonomous mutation. No LLM planner. No multi-agent fan-out. No ad campaign pausing.

v0 contract:

  • Wrap each existing harvester in an envelope (start/end timestamps, row counts, max source timestamp, error class)
  • Emit a typed receipt for each harvest, not raw stdout
  • Required-source failure produces degraded — never silently zero
  • compute_metrics reads typed receipts, not raw stdout
  • Telegram/console output is idempotent (stable occurrence key: condition_id + entity_id + period; rule_version is evaluation metadata, not part of dedup identity)
  • All nodes emit strict JSON on stdout, human text on stderr
  • Secrets loaded from environment, not embedded in source

v0 nodes (fixed catalog — no LLM planner):

harvest_gsc → harvest_shopify → harvest_klaviyo → harvest_ads ↓ ↓ ↓ ↓ receipt receipt receipt receipt └──────────────┴───────────────┴──────────────┘ ↓ compute_metrics ↓ route_alert / write_report ↓ notify_telegram

v0 non-goals:

  • No autonomous campaign mutation or pausing
  • No dynamic LLM planner — static plan templates only
  • No multi-agent synthesis — report is a direct composition of typed receipts
  • No compensation or rollback — stop() records dirty workspace but doesn't clean it
  • No dynamic graph expansion — all nodes known at declaration time

v0 falsifiable success criteria:

  • 10+ successful fixture runs with schema-valid outputs
  • All source and node failures visible (no silent zero)
  • Reproducible audit reports from versioned artifacts
  • Zero duplicate notifications/effects in injected retry/crash tests
  • No claim of success without verified postconditions or complete provenance
  • p95 runtime and cost tracked per node type
Important: v0 tests the control-plane substrate — envelope, receipt, credential hygiene, stdout discipline, degraded propagation. It does NOT test the MoH thesis (heterogeneous composition of unlike outputs). That is v0.5 below.

v0.5 Thesis Experiment: Heterogeneous Composition

The smallest experiment that can falsify or support the MoH thesis: one graph where unlike outputs must genuinely compose.

1. CONTRACT Deliverable: proposed_change (typed artifact combining evidence + diff) Authority: network read (Hermes), local config read (pi). Forbidden: deploy. Acceptance: every claim has source provenance; every diff change maps to evidence. 2. NODES Hermes research external retention guidance → evidence_bundle pi inspect current retention config → file_diff (current state) Prime reconcile evidence + current config → implementation_plan pi apply approved plan to sandbox config → test_receipt Prime produce final report → decision_report 3. COMPOSE (the actual test) attach(evidence_bundle → claims in implementation_plan) → supported_claims chain(supported_claims → file_diff) → proposed_change Check: proposed_change status = degraded if any claim lacks provenance. Check: proposed_change status = unavailable if evidence_bundle unavailable. 4. COMMIT Blocked by default. Requires human approval or sandbox-only apply. 5. FAILURE CASES - Hermes returns degraded (one source unavailable): proposed_change must be degraded - Hermes returns evidence that contradicts pi's current config: compose must flag conflict - Hermes returns empty bundle: compose must reject, not silently pass - pi returns malformed diff: compose must reject with validation error

v0.5 falsifiable criteria:

  • Both harnesses invoked through same execution envelope without pretending outputs are identical
  • compose preserves provenance to each input artifact
  • At least one intentionally conflicting/mismatched input is represented explicitly, not dropped
  • A degraded or failed input propagates via sticky degradation rule
  • Same fixture and versions produce reproducible structural result
  • Single-harness baseline recorded for latency, cost, failure rate
  • Commit blocked when proposed_change status ≠ ok

If this cannot be implemented without ad hoc exceptions, the current contract design is falsified.


Adopt vs Build: Layers 1+2

Claude v2 (critic): Your gap list — Saga compensation, heartbeat, worker versioning, retry/lease, event sourcing, typed I/O validation, asset checks — is a description of Temporal plus Dagster. If Layers 1+2 converge on Dagster-with-worse-tooling, the argument for building them needs to be explicit and it isn't in the doc.

Accepted. This section makes the decision explicit.

Default recommendation: adopt for execution, own the semantic types. Use existing orchestrators for mature mechanics. Focus bespoke engineering on the MoH-specific layer: envelope/receipt type system, artifact provenance, capability declarations, harness adapters, compose semantics, and structural checks.

Temporal is the stronger candidate for durable, long-running, retryable workflows with Saga compensation. Dagster is stronger where assets, lineage, schedules, and typed op I/O dominate. A thin local adapter is useful for development and offline tests.

ConcernAdopt fromMoH-owned boundaryBuild bespoke only if
Retries/leasesTemporal/DagsterReceipt status + retry provenanceCross-harness semantics cannot be represented
HeartbeatsTemporalNode liveness in receiptLocal/air-gapped requirement
Worker versioningTemporal/DagsterSchema compat + version pinsHarness versions cannot be pinned
Typed I/OPydantic/JSON SchemaEnvelope + semantic adaptersMoH needs cross-harness contracts
Event historyTemporal/DagsterImmutable receipt referencesOne portable log is mandatory
Asset checksDagsterEvidence/claim checksSemantic checks must remain Layer 3 nodes
Dead letter / retry queuesTemporal/Step FunctionsDegraded propagation policyLocal-first with no daemon

Constraints that justify a bespoke Layer 2:

  • Must run local-first with no daemon or service dependency
  • Must operate offline or in sovereign/air-gapped environments
  • Must avoid hosted control-plane cost or data egress
  • Must have substantially smaller deployment footprint than Temporal/Dagster
  • Must support file-based deterministic replay without a service runtime

Each claimed constraint needs a test. Run v0.5 in a disconnected environment, measure deployment footprint, demonstrate an artifact type that cannot be represented as an activity/op without loss. Do not build full Layer 2 before this decision is validated.


The Envelope and Receipt Formats

The thin waist of the system. Every invocation uses one common envelope and returns one typed receipt.

Invocation envelope:

envelope: contract_id: unique work-order id node_id: node within the plan role: evidence_gatherer | mutator | validator | synthesizer objective: what this node must accomplish inputs: typed references to upstream receipts/artifacts authority: { network: read_only, filesystem: none } budget: { seconds: 90, calls: 8 } expected_type: evidence_bundle | file_diff | test_receipt | decision_report acceptance: [list of criteria this node's output must satisfy] trace: { correlation_id, causation_id }

Return receipt:

receipt: node_id: matching the invocation status: succeeded | degraded | failed | unknown output: type: evidence_bundle | file_diff | test_receipt | decision_report value: ... (type-specific structure) effects: [list of side effects performed] observations: ["one source was inaccessible"] provenance: [{ source: url, quote: "...", retrieved_at: timestamp, authority: primary }] validation: checks: [list] passed: true | false costs: seconds: 41 network_calls: 6 harness: name: hermes | pi | prime version: ...

Example Flow: Read-Only Revenue Alert (v0)

1. CONTRACT Deliverables: alert_report (one typed artifact) Authority: network read only. Forbidden: filesystem write, deploy. Acceptance: all sources confirmed fresh; no zero-masquerading-as-unavailable. 2. PLAN (fixed template, no LLM) harvest_gsc ──→ receipt_gsc harvest_shopify ─→ receipt_shopify harvest_klaviyo ─→ receipt_klaviyo harvest_ads ──→ receipt_ads compute_metrics (reads 4 typed receipts) ──→ metrics route_alert (reads metrics) ──→ decision notify_telegram (reads decision) ──→ notification_receipt 3. DISPATCH All four harvests run in parallel (all read-only). compute_metrics waits for all four receipts. route_alert and notify_telegram are sequential. 4. COMPOSE assemble_inputs(gsc=receipt_gsc, shopify=receipt_shopify, klaviyo=receipt_klaviyo, ads=receipt_ads) → harvest_snapshot wrap(harvest_summary + metrics + decision + notification_receipt) → final_package 5. CHECK Verify: all receipts status != failed Verify: no source returned degraded without alert Verify: notification was sent (idempotent dedup key checked) 6. STOP Record workspace state: dirty? no (read-only run). Emit final receipt to event log.

Iteration 4 Review Summary

This iteration incorporated structured reviews from all three harnesses plus an external critic. Key conclusions:

ReviewerRoleKey insight that changed the design
Claude External critic (independent) Lineage is workflow engines + build systems + blackboard, not MoE. Plan omission is the highest-risk failure. Join() is underspecified and doing double duty. Stop() can't guarantee clean state. Three harnesses are roles, not kinds.
Prime Agent Implementation audit (internal) Current graph has credential leak, mixed JSON/stdout, zero-vs-unavailable ambiguity. Separated three layers. Found concrete bugs. Specified v0 falsifiable criteria.
Hermes Prior-art research (internal + external evidence) HEARSAY-II scheduler bottleneck (external evidence): directly applies to plan() scaling risk. Temporal Saga + Dagster Asset Checks (external): closest prior art. Confirmed 10 Claude points against real systems (internal corroboration, not independent — same prompt lineage).
pi Implementation review (internal) Envelope-first v0 is the right starting point. Credential leak and JSON/stdout mixing are P0. Fast path must skip planner for simple edits. Route on capabilities, not proper nouns. Stop() must report dirty workspace.
Evidence independence: Claude is the only fully independent reviewer (no shared prompt lineage). Hermes' prior-art findings (HEARSAY-II, Temporal Saga) are external evidence; its point-by-point confirmation of Claude's critique is internal corroboration. Prime and pi are internal implementation reviews. Agreement among the three internal views is useful for consistency but is not independent validation. The central risk — plan omission via under-declared contract — requires adversarial review with no shared design context.

Design Principles

  • Thin waist: stable envelope format, capability manifests, typed receipts. The meta-harness doesn't grow its own agent — it stays a control plane.
  • Least privilege: every node gets scoped authority and a short-lived context. pi can't do research, Hermes can't touch files.
  • Lazy activation: start Hermes/Prime Agent only when a graph node requires them. Don't keep all three running.
  • Explicit policy: routing rules, retry policies, confirmation gates, and conflict resolution live in declarative config, not in agent prompts.
  • Streaming artifacts: pass snapshots, diffs, and evidence references — not giant context dumps.
  • Audit trace, not deterministic replay: event log reconstructs what was decided, not the decision itself. Non-hermetic nodes are labeled.
  • Human gates: require confirmation before irreversible mutation or low-confidence routing decisions.
  • Fail closed: missing provenance, type mismatch, budget exhaustion, failed checks — stop the graph.
  • Bounded autonomy: time, calls, depth, and repair iterations are hard limits, not suggestions.
  • Fast path always: if the task fits ≤2 pi nodes, skip the planner entirely. Zero overhead for simple work.
  • Layer 2 structural only: Runtime checks validate schema, status, hashes, dedup keys, provenance links, and artifact existence. Citation quality, evidence support, contradiction resolution, and alert-basis judgment are Layer 3 graph nodes — never hidden in the control plane.

Connections


  • MoH Architecture — Iteration 2 (superseded) — Original 353-line deep dive from Prime Agent. Superseded by this page but preserved for attribution.
  • Complex Adaptive Systems — MoH as a heterarchical control system
  • J.C.R. Licklider — Man-Computer Symbiosis

  • See also


    Categories: Home › Systems