| MoH — Mixture of Harnesses | |
|---|---|
| Room | Systems |
| Type | Architecture Pattern |
| Status | Design Phase — Iteration 4 |
| Date | 2026-08-16 |
| Supersedes | Harness MoE (2026-08-13), MoH Iteration 2 (2026-08-15), MoH Iteration 3 (2026-08-16) |
| Prior Art | Temporal · Dagster · Bazel · Blackboard (HEARSAY-II) · Airflow · Nix · AWS Step Functions |
| Full Analysis | moh-architecture.md (Iteration 2, superseded) |
merge with typed assemble_inputs. Sticky degraded propagation as formal rule. Fixed dedup identity (stable occurrence key, not rule_version). Added Layer 2/3 boundary (structural checks only in runtime). Stated coverage proof limits explicitly. Distinguishes reviewer evidence by independence tier.
MoH is a typed DAG scheduling architecture with heterogeneous executors. Its unit of composition is a work contract — a typed description of what's needed, what authority is granted, and what acceptance criteria must pass — not a token stream or a gated routing decision.
The architecture routes between components of fundamentally different kinds — deterministic local execution, stateful analysis, and external research — whose outputs (file diffs, analytical decisions, and evidence bundles) cannot be averaged, blended, or gated like neural logits. The correct operation is decomposition into a typed DAG, not classification to a single expert.
MoH stands on solved ground. The following prior art is explicitly acknowledged:
| System | What MoH borrows | What MoH does differently | Risk if ignored |
|---|---|---|---|
| Temporal | Deterministic orchestration, activity workers, Saga compensation, event-sourced state, idempotency keys, heartbeat, retry/lease semantics | Typed plan representation with pre-execution validation; capability-based dispatch instead of task-queue routing | Partial-failure rollback requires Saga; deterministic replay requires hermeticity barrier |
| Dagster | Typed op I/O, asset checks (user-defined validation that blocks downstream), Resources for infrastructure abstraction, type-check before execution | Heterogeneous executor routing (not just data ops); typed contracts between unlike harnesses | Without typed I/O validation, type mismatches become runtime failures |
| Bazel | Hermetic actions with declared reads/writes/effects, content-addressable output caching, effect tracking via dependency declarations | Non-hermetic actions are first-class (LLM calls, web research) with explicit evidence labeling | Non-hermetic nodes cannot be cached or deterministically replayed; must be labeled as such |
| HEARSAY-II (Blackboard) | Typed knowledge sources with typed contributions to a shared blackboard; data-driven scheduling | Contract-first static plan (vs. fully data-driven); typed receipts instead of shared mutable blackboard | Blackboard scheduler bottleneck proved that central dispatch needs meta-reasoning or becomes intractable |
| Airflow | Static DAG as plan representation; XCom for data passing; trigger rules for downstream dependencies | Typed contracts instead of untyped XCom; capability-based executor routing instead of queue names | Static DAG cannot replan; missing nodes are silent; CeleryKubernetesExecutor hybrid was abandoned |
| AWS Step Functions | Explicit error handling with Catch/Retry on each task; Saga compensation via state machine modeling | Typed join semantics; provenance tracking across service boundaries | Compensation must be modeled in the plan, not assumed by the runtime |
MoH is not a single system. It is three distinct layers that should be built and reasoned about independently:
| Layer | What it does | Where LLMs live | v0 scope |
|---|---|---|---|
| 1. Plan representation | Typed work contract → validated DAG. Schema for nodes, edges, types, contracts, effects, authority. | Optional: LLM for contract-to-graph decomposition. Validator is deterministic. | Static plan templates only. Validator checks coverage and type compatibility. |
| 2. Workflow runtime | Dispatch → receipt collection → retry → state machine → compensation → stop. Append-only event log. | None. This is deterministic infrastructure. | Envelope + receipt format. Single sequential pipeline. No retry, no compensation. |
| 3. Agentic decision layer | Decomposition, synthesis, re-planning, conflict resolution. The meta-reasoning around the graph. | Primary. Plan(), join(), check() are judgment calls that use LLMs in practice. | DEFERRED. Not needed until v1. |
The three current harnesses are roles, not kinds. The architecture routes on declared capabilities + authority, not on identity. If pi gains network access, or Hermes gains a file-write capability, the topology changes — and that's fine, as long as capability manifests are explicit.
| Harness | Natural Job | Capabilities | Output Types | Failure if Misused |
|---|---|---|---|---|
| pi | Deterministic local execution | read, write, shell, test (no network) | file_diff, command_result, test_receipt | Makes architectural judgment or produces weak research |
| Prime Agent | Planning, reasoning, stateful analysis | read, write, bash, analysis (persistent kernel) | plan, report, dataset, decision | Over-engineers simple edits; unbounded autonomy |
| Hermes | External research and signal extraction | web_search, fetch_content, bash (read-only), no write | evidence_bundle, source_collection, findings | Changes local state or treats weak signals as facts |
A core finding from the Iteration 3 review: join() was doing double duty. It was both "combine data" and "apply effects." These are separate operations with separate failure modes.
MoH defines a small type algebra with five composition operations:
| Operation | Input Types | Output Type | Semantics | Failures |
|---|---|---|---|---|
assemble_inputs(a, b, ...) |
Named slots {gsc, shopify, klaviyo, ads} each typed T | snapshot | Collect named source receipts into a record preserving each source's schema, status, freshness, and provenance. Missing slot is visible omission, not empty union. | Missing required slot → incomplete snapshot. Degraded source → degraded snapshot. Never silently zero. |
chain(a → b) |
A, B (different) | C | Feed A's output as input to B's transformation. Like Temporal's ExecuteActivity chains. | Type mismatch at boundary → validation error before execution. |
attach(evidence, claim) |
evidence_bundle, claim | supported_claim | Attach source citations to a claim. Produces a supported_claim with provenance chain. | Evidence doesn't actually support claim → mark as "claim unsupported" rather than lying. |
apply(diff, workspace) |
file_diff, workspace_id | effect_receipt | Apply a file diff to the workspace. This is a COMMIT, not a compose. | Conflicting diffs → stop and merge. Workspace dirty on failure. |
wrap(receipts, metadata) |
[receipt], metadata | final_package | Package multiple receipts into a deliverable with aggregated metadata, costs, and provenance. | Missing receipt → incomplete package (flagged, not hidden). |
Key rule: assemble_inputs, chain, attach, and wrap are compositions — they combine data without side effects. apply is a commit — it mutates workspace state. The runtime must distinguish these and require confirmation gates on commits. Homogeneous merge is intentionally absent: four harvest receipts share an envelope but are semantically incommensurable; each source must be handled as a named input.
The following issues were identified during Iteration 3 review (August 2026). They are ranked by remediation priority.
| Issue | Detail | Fix | Source |
|---|---|---|---|
| Sticky degraded propagation | No formal rule for how degraded status flows through the graph. A degraded harvest arrives at route_alert looking fine. Zero-vs-unavailable bug reappears one layer up. | Degraded is sticky by default: any node consuming a degraded input returns degraded unless its contract explicitly declares tolerance. Tolerance must be named, typed, observable. Node receipt lists degraded input IDs and affected output fields. route_alert may emit data-quality alert but not business alert from partial data unless policy permits. | Claude v2 + Prime v2 |
| Credential in graph script | API key embedded in source code in git history | Rotate key. Load from environment/secret manager. Remove from git history. | Prime audit |
| JSON/stdout mixing | Machine JSON and human status text emitted on stdout; downstream JSON consumers break | Strict JSON to stdout, human diagnostics to stderr | Prime audit |
| Zero vs unavailable ambiguity | Zero-row query returns same shape as missing/unavailable source; alert rules can't distinguish | Required-source failure produces degraded state, not zero rows. Harvest receipt includes row count, max timestamp, error class. |
Prime audit + pi |
| Issue | Detail | Status | Source |
|---|---|---|---|
| Plan omission undetectable | If plan() omits a needed node, check() passes and output is wrong. Coverage proof catches this only after the deliverable is declared. If the contract itself under-declares what's needed, coverage is vacuously satisfied and everything passes. | Coverage proves graph/contract completeness relative to declared intent. Does NOT prove intent is complete. Humans own intent completeness. Add intent-review step: review acceptance criteria, enumerate required decisions, map risks to required evidence. | Claude 2 + Claude v2 |
| No replan primitive | Static DAG can't absorb surprises (stale data, revoked auth, schema drift). Temporal, Dagster, and Step Functions all handle this. | v0: bounded restart (fresh snapshot, one re-execution). v1: dynamic graph expansion with depth budget. | Claude 3 + Hermes research |
| stop() overclaims | "No partial state" is impossible when pi has written files or run processes. Needs compensation per node. | v0: stop(reason, dirty_workspace). Records what's incomplete. v1: Saga compensation registry per node type. | Claude 5 + Temporal Saga research |
| Deterministic replay overclaimed | LLM planners and live web research are not reproducible. Audit log ≠ replay. | Rename to "audit trace." Simulation replay only for deterministic subgraphs. Label non-hermetic nodes explicitly. | Claude 6 + Bazel hermeticity research |
| Missing fast path | A one-line edit shouldn't pay full DAG overhead (contract → planner → validate → dispatch → compose → check) | v0: if task fits ≤2 pi nodes, skip planner entirely. Direct dispatch. | Claude 9 + pi |
| Source contradiction ranking | "Rank by authority/date" for contradicting sources will hide quiet errors. Needs explicit source policy with uncertainty modeling. | v0: flag contradiction, don't silently rank. v1: source policy DSL. | Claude 10 + HEARSAY-II lesson |
| Layer 3 leaking into check() | Citation quality, evidence-support checks, and anomaly-basis judgment are semantic. If they become runtime primitives, the deterministic control plane quietly acquires a model. | Hard rule: Layer 2 checks are structural only (receipt status, schema, dedup keys, hashes, provenance links, artifact existence). Semantic judgment is always a graph node dispatched to a harness, producing a typed receipt. | Claude v2 |
| Adopt-vs-build undecided | Gap list (compensation, heartbeat, versioning, retry, event sourcing, asset checks) describes Temporal + Dagster. No explicit rationale for building Layers 1+2 vs adopting. | See adopt-vs-build decision section. Default: adopt for execution, own the semantic type system. | Claude v2 |
| Issue | Detail | Status |
|---|---|---|
| Compensation for non-idempotent nodes | Some harness effects (e.g., "send email") cannot be undone. Partial-failure compensation needs explicit modeling per effect type. | Deferred to v1. v0: mark non-compensatable nodes and require human confirmation before dispatch. |
| Worker versioning | Temporal's "many-versions problem" — when harness code changes mid-workflow. Not addressed in current MoH design. | Deferred. Version receipts include harness version hash. Detection = feasible, resolution = deferred. |
| Heartbeat / liveness | No detection of hung harness processes. Step Functions has heartbeat timeout; MoH doesn't. | Deferred. v0: wall-clock budget only. v1: per-node heartbeat. |
The v0 is a deterministic, read-only revenue-alert graph. No autonomous mutation. No LLM planner. No multi-agent fan-out. No ad campaign pausing.
v0 contract:
degraded — never silently zerocompute_metrics reads typed receipts, not raw stdoutcondition_id + entity_id + period; rule_version is evaluation metadata, not part of dedup identity)v0 nodes (fixed catalog — no LLM planner):
v0 non-goals:
v0 falsifiable success criteria:
The smallest experiment that can falsify or support the MoH thesis: one graph where unlike outputs must genuinely compose.
v0.5 falsifiable criteria:
compose preserves provenance to each input artifactIf this cannot be implemented without ad hoc exceptions, the current contract design is falsified.
Default recommendation: adopt for execution, own the semantic types. Use existing orchestrators for mature mechanics. Focus bespoke engineering on the MoH-specific layer: envelope/receipt type system, artifact provenance, capability declarations, harness adapters, compose semantics, and structural checks.
Temporal is the stronger candidate for durable, long-running, retryable workflows with Saga compensation. Dagster is stronger where assets, lineage, schedules, and typed op I/O dominate. A thin local adapter is useful for development and offline tests.
| Concern | Adopt from | MoH-owned boundary | Build bespoke only if |
|---|---|---|---|
| Retries/leases | Temporal/Dagster | Receipt status + retry provenance | Cross-harness semantics cannot be represented |
| Heartbeats | Temporal | Node liveness in receipt | Local/air-gapped requirement |
| Worker versioning | Temporal/Dagster | Schema compat + version pins | Harness versions cannot be pinned |
| Typed I/O | Pydantic/JSON Schema | Envelope + semantic adapters | MoH needs cross-harness contracts |
| Event history | Temporal/Dagster | Immutable receipt references | One portable log is mandatory |
| Asset checks | Dagster | Evidence/claim checks | Semantic checks must remain Layer 3 nodes |
| Dead letter / retry queues | Temporal/Step Functions | Degraded propagation policy | Local-first with no daemon |
Constraints that justify a bespoke Layer 2:
Each claimed constraint needs a test. Run v0.5 in a disconnected environment, measure deployment footprint, demonstrate an artifact type that cannot be represented as an activity/op without loss. Do not build full Layer 2 before this decision is validated.
The thin waist of the system. Every invocation uses one common envelope and returns one typed receipt.
Invocation envelope:
Return receipt:
This iteration incorporated structured reviews from all three harnesses plus an external critic. Key conclusions:
| Reviewer | Role | Key insight that changed the design | |
|---|---|---|---|
| Claude | External critic (independent) | Lineage is workflow engines + build systems + blackboard, not MoE. Plan omission is the highest-risk failure. Join() is underspecified and doing double duty. Stop() can't guarantee clean state. Three harnesses are roles, not kinds. | |
| Prime Agent | Implementation audit (internal) | Current graph has credential leak, mixed JSON/stdout, zero-vs-unavailable ambiguity. Separated three layers. Found concrete bugs. Specified v0 falsifiable criteria. | |
| Hermes | Prior-art research (internal + external evidence) | HEARSAY-II scheduler bottleneck (external evidence): directly applies to plan() scaling risk. Temporal Saga + Dagster Asset Checks (external): closest prior art. Confirmed 10 Claude points against real systems (internal corroboration, not independent — same prompt lineage). | |
| pi | Implementation review (internal) | Envelope-first v0 is the right starting point. Credential leak and JSON/stdout mixing are P0. Fast path must skip planner for simple edits. Route on capabilities, not proper nouns. Stop() must report dirty workspace. | |
| Evidence independence: Claude is the only fully independent reviewer (no shared prompt lineage). Hermes' prior-art findings (HEARSAY-II, Temporal Saga) are external evidence; its point-by-point confirmation of Claude's critique is internal corroboration. Prime and pi are internal implementation reviews. Agreement among the three internal views is useful for consistency but is not independent validation. The central risk — plan omission via under-declared contract — requires adversarial review with no shared design context. | |||