---
title: "Bills are grounded by a mixed-instrument regime · Agentic Atlas"
description: "Bills are grounded by a mixed-instrument regime: Two splits fall straight out. A token denomination spanning both hard instruments is split into two columns."
canonical: "https://agentic-atlas.dev/decisions/0017-mixed-instrument-bill-grounding"
last-updated: "2026-08-19"
---

Manifest-selected Decision

Canonical id

0017-mixed-instrument-bill-grounding

Status

accepted

Catalog order

9

Catalog revision `35263c4c415da742953d0462804fb14424e2244dae4c63efd27e468988de70ab`

# 0017: Bills are grounded by a mixed-instrument regime

**Status:** accepted **Date:** 2026-08-13

## Context

Three nodes in the corpus run on a **bill** — denominations billed per arm against a declared baseline — and the two that are `stable` were never measured. `session-bill` and `one-guide-three-bills` both stand on declared scenario inputs plus recomputable arithmetic, and each disclaims measurement in its own epigraph. Neither cites a research file. Neither ran anything.

One Review, Five Bills declared a different gate for itself when it landed at `drafting`: measured runs populate the bill table. That gate was written without checking what grounding the corpus had ever performed, and the corpus has never performed this kind. Its one measurement memo (`research/2026-08-02-deferred-context-measurement.md`) is deterministic token counting — `count.js`, one pass, no model call — and ADR 0013's *Open measurements* files as out of reach precisely what the example promised: quality deltas need "live A/B against scored results," latency is unpublished, trigger fidelity is unsampled.

So a node had promised a gate no node in the tree has cleared, using an instrument the corpus's own measurement ADR had declared it did not hold. The pressure is to replace that gate with one that can actually be met without leaving ADR 0013's discipline: **count what is countable, sample what must be sampled, declare the rest open.** The full working — fixtures, seed strata, parity controls, budget — is the workshop spec `docs/superpowers/specs/2026-08-13-one-review-five-bills-grounding-protocol.md`; what follows is the part that binds nodes other than the one that provoked it.

## Decision

### A bill's denominations are assigned to instruments, and the instruments never blend

Three instruments, assigned per denomination and named wherever a figure appears:

- **counted** — deterministic token counts over authored artifacts (`count.js`, no model call): input tokens, orchestrator window.
- **sampled** — figures read off live executions at a declared n: generated tokens, wall-clock, critical-path tokens, recall.
- **open** — declared unmeasured, in ADR 0013's voice.

Two splits fall straight out. **A token denomination spanning both hard instruments is split into two columns.** An arm's *authored* half — instruction, pinned diff, data package, declared return shape — is countable today with no model; its *generated* half — reasoning, prose, synthesis — is unknowable without running. They become `input tokens (counted)` and `generated tokens (sampled)`, because the whole value of a split gate is that a reader sees at a glance which numbers are hard.

**Wall-clock is reported twice.** Elapsed time on an API-backed harness is contaminated by queueing, retries and rate limits, which bite hardest on precisely the parallel arm and can invert a headline for reasons that have nothing to do with architecture. So elapsed is reported as environment-attributable, and **critical-path generated tokens** — the longest single chain of generation, computable from transcripts — carry the architectural reading, because that chain is what parallelism actually shortens.

Residency stays available as a cross-check column and is not the primary denomination of a measured bill. Residency is derived over a *declared session shape*; promoting it would re-import declared-input softness into the one bill whose purpose is to be harder-grounded than its siblings.

### Marking is conditional on the bill, not universal

**A bill whose cells come from more than one instrument marks every cell; a uniform-instrument bill declares its instrument once, in its epigraph.**

The conditional form is the point. It forces no churn on the two `stable` siblings — they are uniform-instrument, so the epigraph disclaimer each already carries is exactly what per-cell marks would say — while stating *why* they are conformant rather than merely leaving them alone. The rule lands in AUTHORING beside the examples rule it serves.

### The cut rule is per cell, not per rung

A wrong predicted direction is the most valuable output a measurement run produces: a corrected cell teaches more than a confirmed one, and cutting the rung erases the correction. A cell the run cannot settle is marked **indeterminate** and shown. Cutting is reserved for a rung with no surviving benefit in any denomination.

**Reduce a published table by fixture, never by result.** Dropping an indeterminate cell launders a null into an absence, which is the exact move ADR 0013 exists to prevent; the research memo holds every cell raw regardless.

Where a rung predicts against itself in most denominations, its handling under a null is **declared in the node before any arm executes**. A handling written up afterward is a result with a backdated timestamp — the same failure ADR 0013 was written to close.

### Seeds are frozen before prompts are authored

A soft denomination becomes countable by planting a known catalogue into the fixture, so the denominator is known. Three constraints on that conversion:

- **Ordering.** The catalogue is frozen **before** any arm's prompts are written. The same person writes both, so ordering is the only defence against prompts drifting toward the seeds, and it costs nothing.
- **No self-scoring.** An arm grading its own depth would violate verification asymmetry in the corpus's own showcase.
- **A judge blind to both.** The cross-checking judge is blind to arm identity **and** to the catalogue. A judge holding the seed list is not a cross-check — it is a second run of the same instrument wearing independence as a costume. Its output is compared to the list afterward: agreement validates the seeds' representativeness, and **disagreement is the finding**, either seeds that missed what matters or an arm that found what no list held.

### Discharge is tied to reportable cells, and *reportable* means determinate in either direction

A backlog item a worked example stands in for closes on **reportable cells**, never on the example's status rung. A cell is reportable when the run settled it in either direction: a corrected cell — the law measurably hurt — discharges exactly as a confirming one does, because the obligation is to measure, not to win.

Status alone is not enough because the per-cell cut rule lets a rung promote while returning nothing, which would discharge a measurement obligation with a measurement that showed nothing. Two guards ride on the discharge itself:

- **Against chance.** Non-overlap at small n is a weak instrument, and a few spuriously reportable cells are expected under the null across a wide matrix. The qualifying cell must therefore be reportable on the **large fixture** and **uncontradicted** by the small one. An indeterminate small-fixture cell does not block — an effect may genuinely vanish at small scale — but a reportable small-fixture cell pointing the other way does.
- **Non-discharge is an expected outcome, on purpose.** Parity controls chosen to flatter the declared baseline make every rung's win *harder* to earn. An item still open after the run is the bias doing its job, not the run failing, and is read that way.

## Consequences

- The example's gate becomes meetable and its annotation becomes `drafting — fleshed gate: measurement run per ADR 0017`. Its *Verification* section pre-declares every rubric — seed protocol, parity controls, sample discipline, null handlings, discharge guards — and its *Result* section carries the marked bill form. Both land before any arm executes; if the run is never bought, the node still states honestly what would settle it.
- AUTHORING carries the conditional marking rule and the billed-ladder form vocabulary; GLOSSARY carries the denominations nodes reason with (`orchestrator window`, `seeded recall`, `critical-path tokens`).
- The two `stable` siblings take no edit. The conditional form is what buys that, and it is why the rule is written conditionally rather than as a blanket per-cell requirement.
- The backlog's context-envelope pair no longer discharges on the example reaching `fleshed`; it discharges on reportable cells attributable to the envelope's law, under both guards above.
- **Executable measurement code enters the prose repo.** The arm prompts and the runner that executes them land under `research/`, beside `count.js` — the first precedent since that counter for code in this tree that calls a model rather than counting bytes. This is a consequence of the sampled instrument, not a pressure of its own: a sampled figure whose harness is unpreserved is not preserved method, and the citation chain has to stay inside the corpus for the node's Verification to mean anything. The counting half of the regime already recomputes this way, which is the shape being extended rather than invented.
- No figure is authorized by this record. The regime says what a bill's numbers must be produced by and marked as; producing them is the run.

## Related

- ADR 0013 — Deferred context is grounded in measured skills — the discipline this record extends from counting to sampling, and the source of the *declare it open* voice.
- One Review, Five Bills — the node whose gate provoked the regime and the first bill marked under it.
- `docs/superpowers/specs/2026-08-13-one-review-five-bills-grounding-protocol.md` — the workshop spec: fixtures, seed strata, parity controls, budget, and the full decision index.
- `research/count.js` — the counted instrument, and the neighbour the run's harness lands beside.

[Browse all selected Decisions](https://agentic-atlas.dev/decisions)