Why Fusion spends more tokens on purpose
Fusion uses more AI sessions than a single-agent build. Our own runs show what that extra verification caught, where it helped, and when one builder is the better buy.
read the entryFusion uses more AI sessions than a single-agent build. Our own runs show what that extra verification caught, where it helped, and when one builder is the better buy.
read the entry
The first AI coding agents usually produced credible implementations. The biggest improvement came when fresh reviewers challenged the work, combined the best ideas, and tested the finished result through paths the builders missed.
A Verified Receipt means an objective gate executed and every hard criterion of a frozen contract passed, with evidence pinned to the exact commit. Not a summary, not a vibe: a checkable record.
A jury drawn from rival model vendors breaks family bias: a model grading its sibling's homework is a conflict of interest. Cross-vendor votes, blind packets, and preserved dissent make the verdict neutral.
The RunSpec is the one contract the build, jury, integrators, QA, and the receipt share: a versioned spec that freezes the objective, scope, criteria, risks, and proof meaning before agents build. Freeze it before implementation and a changed goal becomes visible instead of convenient.
A candidate can change implementation files, but it cannot redefine the protected checks that decide whether the result passed. Fusion restores gate-feeding inputs from the pinned base, verifies their fingerprints, then runs the objective command against the candidate SHA.
The first jury finds the strongest candidate ideas. The second jury verifies that synthesis actually produced a better final. Integration changes the artifact, so it must not inherit the candidate verdict.
Picking the best candidate leaves value on the table, but combining candidates creates new failure modes: dropped grafts, flattened dissent, new regressions. Fusion makes configured integrators produce competing complete finals, re-gates each one, and returns them to jury pass 2.
Planning work cannot pass an execution gate, so Fusion issues an Accepted Receipt instead of a Verified one. The ladder of commitment says exactly how much proof stands behind each artifact.
The agents closest to a build share its context and its blind spots. Fusion brings in fresh QA after jury pass 2: first a quick ship-readiness filter, then a required independent deep audit that looks for defects the builders, judges, and integrators may all have learned to overlook.
"Failed" is too small a word for a verification run. Fusion records typed stop states so the operator can distinguish bad code, stale evidence, jury contention, infrastructure failure, and a retryable publication problem. Each stop state needs a different recovery route.
A Fusion Receipt gives every frozen criterion its own evidence row spanning six stages: gate, jury pass 1, Integrator disposition, jury pass 2, QA, and final status. The row follows one promise across the whole run without hiding absent proof.
Fusion separates verification into stages because executable tests, comparative judgment, and adversarial release review establish different kinds of evidence. The gate decides what executes, the jury decides which gate-passer best serves the intent, and fresh QA decides whether the selected final is safe to hand off.