Companion Document to Research Note 013
Technical Specification & Supplementary Materials
July 2026 — Companion to Research Note 013 — EM Foundation
Primary paper: Recursive Epistemic Elimination: A Failure Mode in Multi-Agent Deliberation Systems
What This Document Is
This is a working supplement to Research Note 013, not a standalone publication. It carries the executive abstract, the list of claims still requiring external citation, the specific statements traceable directly to the three MMDE source logs, the figure inventory, and the engineering specification for the MMDE team. Content here is more provisional and more implementation-facing than the main paper; treat the main paper as authoritative on the research claims and this document as authoritative on what remains to be verified or built.
Deliberative Provenance & Source Traceability
Session identifiers for the three source MMDE deliberations are retained here for internal traceability (mapped to the anonymized Phase A/B/C labels used in the main paper) but are not published elsewhere on the site, consistent with the main paper's confidentiality commitments. See Deliverable 4 below for the phase-to-session mapping and the full list of statements traceable directly to log content, each with its specific source citation.
Deliverable 2: One-Page Executive Abstract
Recursive Epistemic Elimination: A Failure Mode in Multi-Agent Deliberation Systems (alternative title: The Epistemic Collapse Threshold — Recursive Elimination of Action Under Certainty-Only Optimization; working title retained for continuity: When Deliberation Prevents Decision)
A multi-model deliberation system (“MMDE”), used internally by EM Foundation to reason about complex decision problems through adversarial, role-structured review by five language models, was applied across three linked deliberation phases to the same abstracted commercial dispute — first optimizing for strategic value (Phase A), then for settlement value (Phase B), then for certainty (Phase C). Phases A and B produced rich, actionable recommendation sets. Phase C, instructed to audit every factual assumption behind those recommendations and “remove every recommendation that depends upon an unverified assumption,” converged over three deliberation rounds on the conclusion that zero recommendations survived. A subsequent adversarial round then showed that the small set of “zero-assumption” fallback actions the system had proposed — narrow requests for missing documents — were not in fact free of assumptions, and would themselves be eliminated under the same rule applied consistently.
This document terms this pattern Recursive Epistemic Elimination: the iterative removal of candidate actions because each depends on an unverified assumption, extended until it reaches the assumptions embedded in the verification actions proposed to resolve the original uncertainty. This document defines the Epistemic Collapse Threshold as the point at which a system’s uncertainty filter has eliminated every action-bearing recommendation, leaving only recursive verification, deferral, or inaction. This document shows formally that a certainty-only retention rule — retain an action only if every one of its assumptions is verified — produces action-space contraction that decays roughly exponentially in the number of assumptions per action (Proposition 1), while a graded decision rule is not structurally forced to eliminate an action every time a new assumption is discovered, since a discovered assumption changes its score by an amount proportional to materiality rather than by a discontinuous jump to rejection (Proposition 2, deliberately stated as a weaker claim than convergence); and that because verification actions are themselves actions with their own assumption sets, the certainty-only rule has no internal stopping point short of the empty set unless something external supplies one.
A further, subtler finding: adversarial red-teaming — ordinarily assumed to improve decision quality by finding flaws — instead accelerates collapse under a certainty-only rule, because discovering an additional assumption mechanically shrinks the surviving action space rather than informing a graded judgment about it. The Foundation treats this as, if it survives literature review, the paper’s most distinctive individual claim.
The Foundation is explicit about the paper’s central limitation: the eliminating instruction was stated directly and literally in the prompt, so the observed collapse is plausibly correct instruction-following applied to a genuinely assumption-heavy recommendation set, not a spontaneous pathology. The Foundation argues this reading strengthens rather than weakens the paper’s relevance, because it identifies a reachable, structural consequence of a specific, reasonable-sounding class of safety instruction (“do not act on unverified premises”) that a system designer might otherwise adopt without realizing it has no natural stopping rule.
The Foundation proposes a corrected decision architecture that separates epistemic confidence from seven other scored factors — conditional utility, downside severity, reversibility, information value, option value, direct cost, and delay cost — with regret computed afterward as a derived diagnostic rather than scored independently, and routes every candidate action, including verification actions, through a seven-state machine (ADOPT, ADOPT WITH SAFEGUARDS, RUN LOW-COST VERIFICATION, DEFER PENDING CRITICAL FACT, REJECT, EXPIRED DUE TO DELAY, ESCALATE FOR HUMAN REVIEW) rather than a binary verified/rejected gate. Downside severity and reversibility are kept as two separate scores rather than one combined “risk” axis, so that an action dangerous because the harm is severe reads differently from one dangerous because it cannot be undone. For the specific class of low-cost, reversible information-gathering actions, this document argues the routing logic is better understood through option preservation (the value of a cheap, reversible right to act further, familiar from real-options theory) than through expected value in the ordinary sense, since the magnitude of a genuine information action’s upside is frequently unknowable in advance by construction. The central design correction: unverified must not automatically equal reject. This document closes with testable predictions and a proposed five-configuration experimental comparison for validating the corrected architecture against the observed failure mode.
Deliverable 3: Claims Requiring External Research or Citation
The following claims in the paper are stated from general domain knowledge and should be verified against primary literature before publication. None are load-bearing for the paper’s formal argument (Sections 4–8), which is self-contained; all are load-bearing for Section 12’s novelty assessment.
- Value of information / Bayesian decision theory — the formal treatment of when to gather information before deciding versus act immediately, and its standard mathematical form, needs a specific citation (e.g., to the foundational decision-analysis literature) rather than the generic reference given in Section 12.
- Herbert Simon, bounded rationality — needs citation to the specific originating work(s) and a check that this document’s characterization (“real decision-makers cannot achieve exhaustive verification”) is a fair summary rather than a simplification that elides later refinements of the concept.
- Frank Knight, risk vs. uncertainty (Knightian uncertainty) — needs citation to the originating work and a check on whether this document’s three/four-tier verified/inferred/unverified/contradicted taxonomy maps cleanly onto Knight’s risk/uncertainty distinction or is better treated as a distinct, more granular scheme.
- The precautionary principle and its paralysis critique — needs citation both to a standard statement of the principle and to specific published critiques arguing it can produce paralysis, ideally from policy/regulatory-science literature rather than informal commentary.
- “Analysis paralysis” — needs a citation tracing the term’s use in management/decision-science literature, and a check for whether any existing formalization already captures the recursive-regress mechanism in Section 4.2 (if so, the novelty claim in Section 12 needs revision).
- “Epistemic learned helplessness” — needs citation to its originating usage and a careful check on whether this document’s characterization (an individual reasoner’s phenomenon) versus the multi-agent, red-team-amplified pattern documented here is an accurate distinction or an overstated one.
- Corrigibility and safe-exploration literature in AI safety — needs a literature search (this was explicitly flagged in the assignment as an area to check against) to confirm no existing published work already documents a recursive-elimination collapse of this kind in multi-agent deliberation or debate settings; this is the single most important citation gap for the novelty claim in Section 12.
- Multi-agent debate and deliberative-cascade literature — needs citation and a check against this document’s characterization of adversarial review as generally assumed, in that literature, to improve rather than degrade decision quality — this framing needs a supporting citation or should be softened to “commonly assumed” without attribution.
- Conservative optimization (as a named AI-safety concept, distinct from this document’s informal use of “conservative” in Section 9) — needs a citation check to ensure this document is not conflating a term of art with its own informal usage.
- Robust decision-making literature — mentioned in the assignment’s prior-art list but not substantively addressed in the current draft; needs to be either incorporated into Section 13 with a citation or explicitly noted as reviewed-and-distinguished.
- Real-options theory — Section 4.5 and Section 9 now lean on option-preservation as a first-class quantity for information-gathering actions, borrowing the “right without the obligation to act” framing from real-options valuation in finance and operations research. This needs a specific citation to the originating and standard treatments of real-options theory, and a check on whether this document’s informal use of OV(a) is consistent with, or a simplification of, the formal option-pricing machinery in that literature (e.g., whether a Black-Scholes-style or binomial-lattice treatment would be more defensible than the additive term currently used).
- Proposition 2 (Section 4.4) — already renamed and narrowed to “Absence of Structurally Forced Monotonic Contraction Under a Graded Decision Rule,” a deliberately weaker claim than convergence. Before publication, confirm this weaker framing is in fact defensible as stated (i.e., that boundedness/proportionality of Δ DV(a) updates really does follow from the graded-rule construction without additional unstated assumptions), and keep the stronger convergence question explicitly reserved for the companion paper on decision under residual uncertainty rather than reintroduced here.
Deliverable 4: Statements Supported Directly by the Attached MMDE Logs
The following statements in the paper are source-derived — traceable to specific content in the three audit logs — as distinct from the theoretical/formal material, which is original analysis built on top of the logs. This document is internal-use only and retains the original session identifiers for traceability; the public paper refers to these same sessions only as Phase A, Phase B, and Phase C, per the confidentiality and framing revisions below.
| Phase A |
EMF-1785539837336 |
Litigation-strategy optimization |
| Phase B |
EMF-1785542312323 |
Settlement-economics optimization |
| Phase C |
EMF-1785543113151 |
Certainty optimization / assumption audit |
- All three sessions used the same five-model, four-role (Analyst, Critic, Synthesizer, Reviewer), red-team-enabled deliberation structure. (Source: log headers, all three sessions.)
- Phase A’s objective was explicitly to design a sequenced, multi-track campaign, not merely to recommend filing suit, optimizing simultaneously across multiple named probability and value dimensions with per-recommendation benefit/risk/dependency/alternative/confidence fields. (Source: Phase A research prompt and Round 0/1 outputs.)
- Phase A’s adversarial round identified real gaps (unmodeled defendant coordination, one-sided regulatory-outcome modeling, insufficiently tested claim viability) but these resulted in caveats and refinements to the existing recommendation set, not removal of recommendations. (Source: Phase A adversarial-review section and consensus synthesis.)
- Phase B’s objective was explicitly to ignore litigation and pleadings and optimize purely for settlement/negotiation economics, with a stated formula for when a counterparty’s cost of continuing exceeds its cost of resolving. (Source: Phase B research prompt and consensus synthesis.)
- Phase B’s adversarial round specifically challenged the assumption that a regulatory filing is a controllable, reversible leverage tool, on the grounds that a regulator, once engaged, controls the outcome of its own proceeding. (Source: Phase B adversarial-review section, Grok 4.3 and GLM 5.2 responses.)
- Phase C’s objective explicitly instructed the models to audit every assumption behind prior recommendations, identify verification status and cost, determine what happens if false, determine survival, produce a dependency graph, and — verbatim — “Remove every recommendation that depends upon an unverified assumption.” (Source: Phase C research prompt.)
- Phase C, Round 0 produced an explicit four-tier assumption classification scheme (verified / inferred / unverified / contradicted) applied to prior recommendations. (Source: Phase C, Claude Sonnet 4.6 Round 0 response, “PRELIMINARY NOTE ON METHOD.”)
- Phase C, Round 1 reached an explicit terminal conclusion, subsequently adopted across models, that zero recommendations survive the audit. (Source: Phase C Round 1, multiple model responses referencing this conclusion as reached by “Grok” and then adopted; also stated directly in the consensus synthesis: “zero recommendations survive.”)
- Phase C, Round 2 proposed reclassifying a small number of narrow document requests as “zero-assumption” actions and then, on reflection within the same round, relabeled them as “verification preconditions” rather than “surviving recommendations,” with reinstatement of eliminated recommendations gated on individual re-verification of their assumption chains. (Source: Phase C Round 2 responses and “Surviving Zero-Assumption Actions” section.)
- Phase C’s adversarial round directly challenged the “zero-assumption” framing, with at least two models (identified in the log as an Adversarial Reviewer instance and a further Adversarial Reviewer instance) independently arguing that a document request presupposes unverified facts including counterparty possession of the record, willingness to treat the request as legitimate, absence of a legal or contractual bar to informal disclosure, and standing to request the document outside formal discovery. (Source: Phase C adversarial-review section, “OBJECTION 1: The ‘Zero-Assumption’ Framing of Document Requests Is False,” and the parallel objection beginning “The emerging consensus has several dangerous weaknesses… ‘Verification step’ is being treated as harmless when it still encodes a theory.”)
- Phase C’s adversarial round also raised, independently, an explicit objection that the dependency graph had no time axis and that a certainty-optimized plan producing zero actionable recommendations while deadline-type periods run is “strategically self-defeating,” not merely epistemically cautious. (Source: Phase C adversarial-review section, “OBJECTION 5: The Dependency Graph Has No Time Axis.”)
- Phase C’s final consensus synthesis explicitly records the convergence on “zero recommendations survive” as a genuine, multi-round, cross-model epistemic refinement rather than an isolated single-model claim, describing the Round 1→Round 2→adversarial-round progression as “earned, not manufactured.” (Source: Phase C Variance Report section.)
- Phase C’s final consensus synthesis produced a structured dependency-graph output distinguishing a small verified fact base, a larger unverified assumption layer, a list of removed recommendations (spanning regulatory filings, a pre-suit discovery mechanism, a tort theory, a statutory claim, demand letters, litigation sequencing, and expert engagement), and a four-item verification precondition layer explicitly labeled as distinct from recommendations. (Source: Phase C consensus synthesis dependency-graph block.)
- Phase C’s consensus synthesis included a “reinstatement condition” rule stating that a removed recommendation becomes eligible for reinstatement only when its specific supporting assumption moves from unverified/inferred/contradicted to verified status, with reinstatement requiring re-auditing the full assumption chain — and a separate adversarial-round objection (Objection 3) that an earlier version of this reinstatement condition was circular, specifying only that reinstatement occurs “when the document is produced” without defining sufficiency. (Source: Phase C consensus synthesis and adversarial-review “OBJECTION 3: Reinstatement Conditions Are Circular or Empty.”)
- In all three phases, two of the five models (identified in the logs as DeepSeek V4 Flash and, in two of the three sessions, GLM 5.2) returned empty, blank, or truncated (“null”) responses in at least one round, meaning the effective deliberating group in practice was smaller than five models for portions of each session. (Source: direct observation across all three logs, e.g., Phase A “DeepSeek V4 Flash (Adversarial Reviewer)” section containing only a token count with no substantive text, and “GLM 5.2 (Adversarial Reviewer)” containing only the word “null.”)
Five figures have been generated and are embedded directly in the paper, numbered consecutively (figures/ subfolder: fig1_progression.png, fig2_contraction_curve.png, fig3_regress.png, fig4_state_machine.png, fig5_boolean_vs_bayesian.png). The three additional figures originally sketched (dependency-graph schematic, six/eight-axis radar chart, experimental design matrix) are noted at the end of this deliverable as not-yet-produced rather than left as numbered gaps in the paper.
- Figure 1 — Three-phase progression schematic (realized). Left-to-right flow: Phase A (Strategic optimization → phased action set) → Phase B (Settlement optimization → counterparty-triggered action set) → Phase C (Certainty optimization → Round 0/1 audit → “zero recommendations survive” → Round 2 repair → red-team discovery that the repair layer also contains assumptions → Epistemic Collapse Threshold). Embedded in Section 6.
- Figure 2 — Action-space contraction curve (realized). Plot of P(Retain(a)) = (1-p)n against n (number of material assumptions per action) for several values of p (per-assumption unverified probability: 0.1, 0.2, 0.3), illustrating the exponential decay described in Section 4.1 and formalized in Proposition 1 (Section 4.4). Purely illustrative/formal, not derived from the case study’s specific numeric assumption counts. Embedded in Section 4.1.
- Figure 3 — The recursive verification regress (realized). Layered diagram: V0 (original action set) → assumptions found → V1 (verification layer) → assumptions found in V1 → an explicit branch point showing where an external stopping rule (Section 9’s scoring architecture) intercepts the regress versus where it would continue unchecked to V2. Embedded in Section 4.2.
- Figure 4 — Recommendation state machine (realized, regenerated in landscape orientation for legibility). Full seven-state routing logic over the eight scored factors (Ê, U, H, Rev, IV, OV, Cd, Cτ), with H(a) and Rev(a) kept as separate decision points rather than one combined risk axis. Embedded in Section 9.
- Figure 5 — Boolean verification logic versus graded decision logic (realized). Side-by-side comparison of the certainty-only filter’s binary keep/discard pipeline, which conflates severity and reversibility into one unscored notion of “risk,” against the eight-factor architecture’s graded, seven-state pipeline with H and Rev kept explicitly separate. Embedded in Section 9.
Not yet produced (would need dedicated production time; noted here rather than left as gaps in the paper’s figure numbering): - A dependency-graph schematic reproducing, in fully generic anonymized form, the shape of Phase C’s own verified-fact-base → unverified-assumption-layer → removed-recommendations output (Section 6’s table already covers this content in tabular form). - An eight-factor radar/spider chart contrasting two example action profiles (a high-adopt-confidence profile and a defer/reject profile), illustrating Appendix C’s rubric visually. - An experimental design matrix figure showing the five proposed MMDE configurations (Section 15) against the ten measured outcomes, for use as a pre-registration or protocol figure.
Deliverable 6: Concise Technical Specification for the MMDE Engineering Team
Purpose. Translate the corrected decision architecture (paper Section 9) into an implementable change to the MMDE deliberation behavior, addressing the specific failure mode documented in Sections 2–7.
A note on what is and is not established about the current implementation. This specification describes what has been observed in the deliberation transcripts — the output behavior the prompts produced — not confirmed facts about the MMDE codebase, orchestration repository, or internal pipeline stages. At the time of writing, the actual MMDE orchestration repository has not been located or inspected, so statements below are phrased in terms of observed behavior that should be preserved or changed, not in terms of specific functions, schemas, or pipeline stages known to exist. Repository inspection should confirm or correct every architectural assumption below before implementation begins.
Problem being fixed. The observed deliberation behavior permits a session objective to instruct (explicitly or implicitly) a binary retention rule: an action is kept only if all its supporting assumptions are verified. Applied with enough adversarial rigor, this rule eliminates all actions, including fallback “verification” actions, because those actions are treated in the transcripts as exempt from scoring rather than scored. Nothing in the observed behavior weighs the cost of producing zero output.
Required changes:
- Separate the assumption-audit behavior from the adoption decision. The assumption-audit behavior observed in Phase C — the four-tier verified/inferred/unverified/contradicted classification — should be preserved as an input-producing step that yields Ê(a) per action. Whatever currently determines final output should no longer treat that classification, by itself, as determining whether an action appears in the final recommendation output; a distinct adoption step should consume Ê(a) as one of eight scored inputs (see item 2), not as a standalone veto. Confirm against the actual repository whether this separation already exists structurally or needs to be introduced.
- Implement the eight-factor scoring model. For every candidate action surfaced by any deliberation role (Analyst, Critic, Synthesizer, Reviewer, red team), compute or elicit: Ê(a), U(a) (conditional utility), H(a) (downside severity), Rev(a) (reversibility), IV(a) (information value), OV(a) (option value), Cd(a) (direct cost), and Cτ(a) (delay cost). Keep H(a) and Rev(a) as two separate fields, not one combined “risk” score — a single combined score cannot distinguish an action that is dangerous because the harm is severe from one that is dangerous because it cannot be undone, and the two call for different mitigations. Each should be a model-elicited score (e.g., a 0–1 scale with a required one-sentence justification per factor) rather than a free-text judgment, to keep scores comparable across actions and across deliberation rounds. Compute Regret(a) afterward from H(a), Rev(a), and Δ DV(a) (paper Section 4.3) rather than eliciting it as a ninth independent score.
- Implement the seven-state state machine (Appendix B). Replace the observed implicit binary keep/discard behavior with explicit state assignment: ADOPT, ADOPT WITH SAFEGUARDS, RUN LOW-COST VERIFICATION, DEFER PENDING CRITICAL FACT, REJECT, EXPIRED DUE TO DELAY, ESCALATE FOR HUMAN REVIEW. Every action in the final synthesis output must carry one of these seven state labels, not a binary survived/removed flag.
- Do not exempt verification actions from scoring. Any action proposed specifically to resolve an assumption (a “precondition,” “verification step,” or similarly labeled action) must be run through the same eight-factor scoring as any other candidate action. It is expected that well-designed verification actions will typically score high IV/OV, low Cd, and high Rev, and therefore route to RUN LOW-COST VERIFICATION on their merits — but this must be computed, not assumed, per Section 7’s core finding.
- Add a mandatory, non-defaultable delay-cost field. Cτ(a) must be an explicit, required field for every action, including a structured check for any stated deadline, decay, or window-closure risk in the fact pattern. If no such risk is identified, the field should record “none identified” explicitly rather than being left blank, so that its absence is visible and auditable rather than silently defaulted to zero.
- Add an explicit “cost of the null action” computation. At the synthesis stage, alongside the list of adopted/deferred/rejected actions, compute and report DV(null) and the resulting Δ DV(a) for each action (paper Section 4.3) — the estimated cost of the system recommending no action at all — so that a zero-recommendation outcome is never presented without an accompanying statement of what that outcome itself costs.
- Add a regress-depth guard. If a verification action proposed to resolve an assumption is itself found, in adversarial review, to depend on a further unverified assumption, cap the number of recursive verification layers considered (configurable; default suggestion: 2) before the process is required to route the deepest-layer action through RUN LOW-COST VERIFICATION or ESCALATE FOR HUMAN REVIEW rather than continuing to spawn further verification layers indefinitely.
- Add a Deliberative Confidence Collapse monitor. Track, across rounds within a session, whether the size of the adopted-action set is monotonically shrinking round over round with no round in which it grows; if so, flag the session output with an explicit warning banner noting the pattern, so a human reviewer can distinguish “the deliberation converged because it should have” from “the deliberation ratcheted toward the empty set.”
- Synthesis-output format change. The final consensus synthesis should report, per action: its eight factor scores, its derived Regret(a), its assigned state, and (for anything not in the ADOPT or ADOPT WITH SAFEGUARDS state) a one-line statement of what would need to change for its state to improve — replacing the observed free-text “removed recommendations” list with a structured, machine-parseable and human-auditable table.
- Backward compatibility. Existing session objectives that ask the system to “optimize certainty” or similar should be reinterpreted as setting Ê(a)’s weight higher relative to U(a) in the adoption step, rather than as instructing a literal binary retention rule — i.e., “optimize for certainty” becomes a weighting instruction to the eight-factor model, not a bypass of it.
Out of scope for this specification: changes to the underlying five-model roster, role structure (Analyst/Critic/Synthesizer/Reviewer), or the decision to include an adversarial red-team pass — all of which the case study and this paper treat as sound design choices that surfaced a real problem, not as sources of the problem themselves. Also out of scope until repository inspection is complete: any claim about specific function names, schemas, or existing code structure — items 1–10 above should be read as target behavior, to be mapped onto whatever the actual implementation turns out to look like.