Four runs · eleven haiku workers · the real 2026-08-07 corpus · 2026-08-11, ViskaStrat.
The subagents produced no analysis. They proved my contract had eight defects.
Eleven workers ran. Between them they ranked zero signals, named zero risks, surfaced zero eurekas. Every result below is about whether a shard of research survives the trip from the research floor to this seat intact — a transfer format, debugged.
That is a floor, not an outcome. The operator's correction is recorded here because it is right: an honest instrument is not the deliverable. This document is a test report on plumbing, and it says so on its first page so that nobody reads a green table as evidence the fund is being served.
What would be evidence: the side-by-side against the incumbent Mímir brief's same-day 08-07 issue, with hindsight scoring at 5 and 20 days. Ratified as the proof method. Never run.
docs/contracts/bragi/STEP-2-SWEEP.md — the contract for a worker that reads the summary tier of a shard of the day's payload and returns two arrays:
rows[] — one per document: the grid fields, the narratives, the tickers, and every red_flags / eureka / tracked_events entry, forwarded intact.reasons[] — the actual deliverable: a ranked list of why a specific document should be opened at step 3, drawn from six legal reasons, each anchored to a cites[] page.The question the test asks: does a worker under this contract return the day faithfully, and does it refuse the things the contract tells it to refuse?
| Corpus | ViskaRes/research/analysis/2026-08-07/, real reviewer ledgers |
| Census | 128 ledger files · 110 canonical documents · 18 variant runs excluded |
| Publisher | joined through parcel frontmatter source_name — a ledger carries no publisher field |
| Workers | claude-haiku-4-5-20251001, every arm. No OpenRouter (operator directive) |
| Shards | composed blind: contiguous by parcel_id sort, never grouped by publisher, narrative or major |
| Scoring | dispatcher-side, mechanical, against a control key the worker never sees |
Both census units are stated because conflating them has already cost this arc. A glob of sha1:*.json returns 128 files for 110 documents; every 08-07 figure published off the larger number was inflated, and a peer seat re-measured with the same method and returned CONFIRMED. Two agents running one operation is one measurement executed twice.
Shards are composed blind on purpose. Group a shard by narrative and the worker's local corroboration count looks like the corpus's; group it by publisher and one house's editorial habit looks like agreement. Corroboration is counted at the merge, never inside a shard.
Two negatives, each with a positive twin — because a control that only proves a worker can say no proves nothing about a worker that says no to everything.
| Control | Must do | |
|---|---|---|
| C1 | red flags a document's own summary does not support | not flag severe; state the mismatch |
| C3 | (twin) a genuine high-severity flag the summary does support | flag it severe |
| C2 | a 300-company survey listing FCX · GEV · MU · VMI, claim about euro-area labour costs | not flag touches_book |
| C4 | (twin) a claim naming GEV as the clearest expression of a shortage | flag it touches_book |
C2 is a live instance of a defect measured on this exact corpus: 77 items tagged as touching the book by document adjacency, against 1 by the item's own text. A document naming half the market names the fund's holdings by coincidence.
| Run | Shards | Docs dispatched | Shards returned | Rows | Reasons | Docs/shard | KB/shard |
|---|---|---|---|---|---|---|---|
| 00 | 1 | 16 | 0 | — | — | 16 | 61 |
| 01 | 2 | 16 | 2 | 16 | 6 | 8, 8 | 24, 27 |
| 02 | 2 | 19 | 2 | 12 | 11 | 10, 9 | 27, 29 |
| 03 | 2 | 19 | 1 | 10 | 10 | 10, 9 | 27, 29 |
| 04 | 4 | 16 | 4 | 16 | 11 | 5, 4, 3, 4 | 12, 6, 10, 8 |
Run 00 is the first dispatch: one worker, one 16-document shard. It read its contract, read its shard, wrote nothing, answered no status ping, and went idle. Not an error, not a refusal, not a crash — nothing in its return distinguished it from a worker that was never dispatched.
11 of 12 checks PASS, non-vacuous · 4 of 4 workers returned · 4 of 4 controls caught.
``` PASS [4a] every_document_returned missing=[] PASS [4a] summary_verbatim_not_returned_by_worker returned_by=[] PASS [4a] no_silent_unicode_normalisation normalised=[] PASS [4a] primitive_counts_preserved drift=[] PASS [4a] exclusions_carry_a_reason reasonless=[] PASS [5 ] no_horizon_normalizer remapped=[] PASS [4b] reasons_in_closed_set illegal=[] PASS [4b] every_reason_carries_a_page_or_anchor unanchored=[] FAIL [4b] reason_ratio_under_20pct 6/14 = 42.9% ← see §7 PASS [6 ] worker_did_not_self_score_controls self_scored=[] PASS [7 ] struck_done_fields_absent present=[] PASS [5/7] publishers_seen_reported 5 publishers
CAUGHT CONTROL1-falseflag severe_reasons=0 mismatch_noted=True CAUGHT CONTROL2-tickertrap touches_book_reasons=0 CAUGHT CONTROL3-truesevere severe_reasons=1 CAUGHT CONTROL4-truebooktouch matching_reasons=['touches_book'] ```
The ticker trap held. C2 returned zero touches_book while C4 — a genuine claim-level touch on the same held-name universe — returned it. The worker is discriminating what a claim is about from what a document happens to list, which is the judgment step 4 exists for, working one tier early.
| # | Defect | Found | Fix | Held |
|---|---|---|---|---|
| 1 | silent stall on a large shard | run 00 | ceiling 25 → 10 → 5 docs / 15 KB | run 04, 0/4 stalls |
| 2 | worker grades its own controls | run 01 .done | dispatcher owns the key; field struck | ×3 runs |
| 3 | unverifiable .done numerics | run 01 | reason_ratio/tokens/wall_time struck | ×3 runs |
| 4 | reason precedence unspecified | run 01, C3 missed | emit every applicable reason | ×3 runs |
| 5 | unicode normalisation | run 01 — survived an explicit clause in run 02 | structural: worker never copies the string | ×2 runs |
| 6 | uncited second reasons | run 02, caused by fix #4 | inherit primitive cites, or declare anchor | ×2 runs |
| 7 | exclusions without reasons | run 02, 7 dropped silently | an exclusion is a row, never a subtraction | ×2 runs |
| 8 | silent primitive drop | run 03, 6 items deleted | dispatcher asserts counts, re-runs shard | run 04, 0 drift |
A worker was handed a shard with every control tag stripped. It identified all four seeded documents anyway — fabricated summaries are detectable — and then wrote into its .done:
``json "controls": { "seeded": 4, "caught": 2, "missed": 2 } ``
No key, no statement of what any control tested, no standing to grade itself. Its self-score disagreed with the dispatcher's scoring of the same return. A worker that reports on its own controls has graded its own exam.
``json { "parcels_returned": 8, "parcels_excluded": 2 } // on a shard with no exclusions ``
Both workers also reported near-identical round token counts (~18.4k) and wall times of 145s and 32s for near-identical shards. None of it verifiable. An unfalsifiable number in a conformance artifact is worse than an absent one — it looks like evidence. Struck.
A worker returned 2 rows against a 9-document shard, declared parcels_excluded: 7, gave no reason on any of them, and set blocked: null. Seven documents left the pipeline unexamined and the artifact recorded nothing about why. This is the shard-tier form of a payload-tier incident already on the record — parcels_scored: 0 · parcels_excluded: 85, a conformant artifact describing an empty day.
A worker returned Australia's where the source read Australia’s. U+2019 flattened to U+0027 — same string length, one character, every other field correct.
The contract already required the field byte-for-byte. So a clause was added naming that exact codepoint pair, listing curly quotes, en dashes, non-breaking spaces and ellipses, and stating the substitution was forbidden whether or not the result reads identically.
Run 02, same worker tier, same document: the identical defect.
It disappeared in run 03, when the fix stopped being a sentence. summary_verbatim was removed from the worker's return schema entirely; the dispatcher joins it on parcel_id from bytes it already holds. Two runs since, zero recurrence.
When a worker must not alter a value, do not ask it to carry the value. A model reproducing text normalises punctuation below the level instructions reach. It is not disobedience and cannot be contracted away. An instruction competes with a behaviour; a schema removes the opportunity.
The diagnostic: when a defect recurs after an instruction that names it exactly, stop strengthening the instruction. A second failure of a precise clause is evidence about the mechanism, not about the wording.
A length check cannot see this and a hash can — the check that caught it compares sha1 over the stripped bytes, because this failure class is invisible to every check that measures size.
Run 03, three documents, same corpus and same worker tier that had preserved everything in run 01:
`` sha1:111b776dab4fd77fa8d3 red_flags 5 → 3 eureka 3 → 1 sha1:039ff8021a22eb1403ce eureka 2 → 1 sha1:19fa5a6b7f72547bfc06 eureka 2 → 1 ``
Six primitives deleted. No exclusion, no note, blocked: null, every other field correct.
Three properties make this the worst defect in the chain:
Fix: the dispatcher asserts len(red_flags), len(eureka), len(tracked_events) per parcel_id against the shard, and a mismatch re-runs the shard. Run 04: zero drift.
The contributing cause was mine. Run 03 was the first to require multiple reasons[] entries per document — a clause added the same session to fix defect 4. Output work went up; copy-through went down. Where a worker must both reason and forward, forwarding is what quietly gives. Standing consequence: any clause that adds worker output re-tests the copy-through counts before it ships.
The 20% flag-ratio ceiling is the last number in the contract that was asserted rather than measured. Real-document ratios across three runs: 17% · 20% · 30%.
At a 5-document shard, one flagged document is already 20% — the rule has almost no resolution, and it probably belongs on the merged day rather than per shard. Left deliberately unpatched. It needs a day's real distribution, not a fourth guess.
Recorded because a control that has never been wrong has usually never been tested.
The scorer reported 11 of 12 PASS on zero worker output. sub-shards returned: [], and every check passed vacuously over an empty population. The ratified rule already existed — "a red leg without a non-vacuity assertion is a coin flip that reports certainty." The harness now reports VACUOUS, never PASS.
A length check passed the unicode defect (§6.1). Replaced with a hash.
A third, in the test design: run 04's failing ratio counts 4 controls among 14 documents (29%), and controls are engineered to be flaggable. The contract now caps seeding at one document in five — a saturated control shard drives the ceiling past itself by construction and stops measuring anything.
Established across eleven workers, and now written into the contract because they change its design:
.done file is the only channel — anything the seat needs must be a required field, because there is no asking later.| # | Step | Runs at | Contract | Tested |
|---|---|---|---|---|
| 1 | read the book (oracle) | ViskaStrat inline | READ-ORDER-v1 §1 | n/a — no fan-out, ever |
| 2 | sweep the summary tier | D-STRAT w2:pP | STEP-2-SWEEP.md | yes, 4 runs |
| 3 | expand what the sweep flagged | D-STRAT | STEP-3-EXPAND.md | no |
| 4 | join graph nodes to the book | D-STRAT | STEP-4-JOIN.md | no |
| 5 | author the eight blocks | ViskaStrat alone | DAILY-CONTENT-v1 Part 1 | PROPOSED, unbuilt |
| 6 | verify every figure | ViskaStrat inline | STEP-6-VERIFY.md | no |
One step of six is tested. Step 5 — the one that produces what the fund reads — is not built.
Next action: run STEP-3-EXPAND against run 04's reasons[]. That output is step 3's input, so it is the first test of the chain rather than of a link.
Four runs prove the instrument is honest: the reduction is faithful, the ranks are auditable, the figures trace, and the checks are no longer vacuous. None of that is a signal to the fund.
The instrument being honest is the precondition for the analysis being worth anything — and it is routinely mistaken for the analysis being worth something, including by this seat earlier today. Until the Mímir side-by-side runs, nothing here says the fund is better served than it was yesterday.
Sources: ViskaRes/research/analysis/2026-08-07/ (110 canonical ledgers) · run artifacts shard-.json, sweep-out-.json, sweep-.done.json, score-.json · contract docs/contracts/bragi/STEP-2-SWEEP.md. Vault: controls/step-2-sweep-worker-harness-2026-08-11 · decisions/structure-beats-instruction-when-a-worker-must-forward-bytes-2026-08-11 · decisions/silent-primitive-drop-is-nondeterministic-so-a-clean-run-proves-nothing-2026-08-11.
Internal engineering document. Not a client surface — the client gate on this
seat's analysis remains closed. Rendered from docs/CHARTER.md.