Planning bug: an acceptance criterion asserted a PIPELINE's outcome ("the PR is green") when the card owned one STAGE of it — the other stages were never verified
Filed by the motir run of MOTIR-2927 (2026-08-17). The correction is already applied — the card's disposition is recorded as a comment on it, and the missing stage is MOTIR-2937. This card is the telemetry.
Lesson: notes.html #298 (motir-meta#221, merged 2026-08-17T20:30:47Z), corrected by motir-meta#223 — see the close-out below.
What the plan asserted
MOTIR-2927 is a one-line fix: add --pass-with-no-tests to test:e2e:acceptance. Its mechanism section is exemplary — the flag, the paths: filter, the sibling gate that cannot help, all measured on a named commit with --list transcripts on both sides.
Its second acceptance criterion then reached one level past that mechanism:
A PR that deletes the lane's only acceptance spec is green — verified by removing a probe spec on the branch, not by reasoning about the flag.
Why that is a different claim than it looks
"The lane's test step stops failing" is a claim about the STAGE this card owns. "The PR is green" is a claim about the whole job — a five-step pipeline of which the flag is step three. The criterion even anticipates the failure mode it is guarding against ("not by reasoning about the flag") and still guards only the flag: it asks the run to MEASURE step three and then asserts the outcome of steps four and five, which nothing had looked at.
Measured on probe PR #17: the paths: filter fired on the deletion, Run pnpm test:e2e:acceptance passed — and the job went red two steps later at Publish the acceptance video, on a TemplateValidationException in a vendored composite action that has, on the evidence, never loaded in either repo. The card's mechanism was right; its terminal sentence was not attainable by it.
A second, smaller instance in the same card
The same card's rationale says the failure becomes reachable "the first time a project completes the receipt lifecycle the starter's own docs/acceptance-video.md teaches." It did not teach it. grep -in 'retire\|promote\|lifecycle' docs/ over origin/main returned exactly one hit, about a receipt superseding its predecessor — never about a spec leaving the lane. The card's own AC 5 is what creates that guidance, so the two sentences contradict each other: one treats the doc as an existing premise, the other as a deliverable.
The rule this is telemetry for
A criterion states the evidence a reviewer needs. When that evidence is the STATE OF A PIPELINE the card contributes one stage to, the plan owes a verification of the OTHER stages — or the criterion has to be re-stated at the altitude the card actually controls ("the lane's test step no longer fails on an empty lane"), with the pipeline-level claim carried by whichever card owns the rest. Neither was done here, and the gap is invisible at plan time precisely because the mechanism analysis is so thorough: a card that measured its own stage that carefully reads as one that measured everything.
Close-out (2026-08-17) — SETTLED: this stays a LESSON. No rule changes.
This section replaces the paragraph the card originally ended on, which read: "Two candidate homes if this recurs: plan-rules/core.md gate 14 (the AC axes) already asks whether a criterion is checkable — this is the narrower question of whether it is checkable BY THIS CARD; and run.md's suite-wide-criterion rule is the same shape one altitude over. One instance is a LESSON, not yet a RULE — logged in notes.html accordingly." That is the notes.html #221 trap — a conditional prescription cannot record that it was later discharged, so the next reader finds a real recurrence and demands a change that was already refused. Do not restore the conditional form.
The card's own premise was false, and that is the finding
"Gate 14 already asks whether a criterion is checkable — this is the narrower question of whether it is checkable BY THIS CARD." Gate 14 asks precisely the narrower question. Its governing sentence (plan-rules/core.md:296) is not about checkability at all:
Every AC satisfiable INSIDE this card's own boundary — on THREE axes: WHO performs the crossing step, WHICH SURFACE it is discharged on, and WHEN it becomes true.
and its (a) ACTOR axis carries this criterion's tell verbatim:
The tell: a criterion phrased as a WORLD-STATE ("the live project now shows Y", "the run reaches step Z") rather than an OUTPUT ("this code does Z") — a world-state is legitimate only when this card's deliverable is the LAST step before the state holds.
"A PR that deletes the lane's only acceptance spec is green" is a world-state; the flag is step three of five.
Verdict — LESSON, on three grounds
- The governing sentence covers the case verbatim — the question the card and #298 both say no gate asks is gate 14's title.
- The ACTOR axis's tell matches the criterion word for word, including the qualifier that decides it (LAST step before the state holds).
- ORDER test — 13 days. Gate 14 landed
8f9668a, 2026-08-04 11:12:50 -0700; MOTIR-2927 was authored 2026-08-17T19:09:50Z, same planner. It sits incore.md, which EVERY pass loads (MANIFEST.mdroutes it intomotir log-bug, the pass that authored the card), and it is mirrored inmotir-ai'sAC_SATISFIABLE_INSIDE_THE_BOUNDARY. A rule that was loaded, applicable and skipped is a diligence problem, not a trigger problem (the MOTIR-2280 discriminator).
The widening that was considered and refused: axis (a) fires "whenever a card declares a scope boundary" while axis (c) fires unconditionally, so "make (a) unconditional too" looks like a trigger that stops one item short. It is not this incident's gap — MOTIR-2927 declares a boundary twice (its opening paragraph assigns the trigger question to MOTIR-2909; the criterion's own "not by reasoning about the flag" names the flag as the card's limit), so the axis was applicable here. Promoting on a gap the incident does not exhibit is the MOTIR-1917 over-reach.
What replaced the false claim in notes.html #298 — the camouflage
The ACTOR axis fires on a declared scope boundary, and this card's boundary declaration is not in its scope paragraph: it is inside the offending criterion. "not by reasoning about the flag" names the flag as the card's own limit in the very sentence that steps past it. A reader auditing for a does NOT do X line finds none, and the one clause that should have triggered the audit reads as rigour instead.
Close-out verification (rung 2, re-executed)
| claim | verdict |
|---|---|
| Correction applied — MOTIR-2937 filed, disposition on MOTIR-2927 | held (2937 todo; the disposition comment is exact and complete) |
| Lesson landed | held — notes.html #298 on origin/main, count line Across the 298 mistakes consistent, no duplicate or non-numeric mistake-num, no renumber collision |
"grep -in 'retire|promote|lifecycle' docs/ returned exactly one hit" | EXACT — 1 at the card's own base 256e70d^ (docs/acceptance-video.md:61), and 5 on origin/main today because criterion 5 shipped the guidance |
| Probe PR #17's failure step | held — Publish the acceptance video, TemplateValidationException, two steps after the test step |
| "the card's disposition is recorded as a comment on it" | literally true and insufficient — MOTIR-2927's DESCRIPTION still carried criterion 2 and the false doc premise verbatim. A description is the spec and a comment is not, so both were amended on the record, marked AMENDED 2026-08-17 (MOTIR-2938) with the original wording quoted inline. |
Advisory disposition
validate_work_item MOTIR-2938 → valid: true, one advisory naming MOTIR-2937 (todo). Not consumed, no edge owed: 2937 is the remedy card this record FILED, not a prerequisite it reads. A record card's close condition is correction verified in the tenant + lesson on origin/main — blocking the report on its own subject would hold a 15-minute card open for another card's whole run.