@chrono-meta/fh-gate 1.4.89 β 1.4.91
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/judgment_circuits.txt +14 -0
- package/.claude/rules/fh_4axis_gate.md +7 -0
- package/.claude-plugin/marketplace.json +2 -2
- package/AGENTS.md +25 -0
- package/CLAUDE.md +202 -12
- package/knowledge/shared/harness-core/dispatch_conditional_prohibition.md +105 -0
- package/knowledge/shared/harness-core/fh_three_layer_canon.md +307 -0
- package/knowledge/shared/harness-core/harness_incubator_doctrine.md +100 -0
- package/knowledge/shared/harness-core/onboarding_acceleration_autopilot.md +3 -1
- package/knowledge/shared/harness-core/ship_readiness_gate.md +181 -13
- package/knowledge/shared/learnings/subagent_invocations_log.yaml +640 -0
- package/knowledge/shared/rules/multi_session_close_protocol.md +118 -0
- package/package.json +23 -2
- package/plugins/fh-commons/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/.claude-plugin/plugin.json +1 -1
- package/plugins/fh-meta/skills/auto-decorrelation/SKILL.md +56 -8
- package/plugins/fh-meta/skills/install-wizard/SKILL.md +33 -0
- package/plugins/fh-meta/skills/install-wizard/SKILL_detail.md +104 -15
- package/scripts/branch_claim.sh +266 -0
- package/scripts/chamber_run.sh +64 -2
- package/scripts/chamber_witness.sh +439 -0
- package/scripts/compaction_probe.sh +456 -0
- package/scripts/digest_landing_check.sh +385 -0
- package/scripts/directional_diff_gate.sh +459 -0
- package/scripts/fh_env_delta_scan.sh +15 -0
- package/scripts/fh_session_load.sh +34 -1
- package/scripts/field_canon_preload.sh +129 -0
- package/scripts/hook_source_lib.sh +41 -0
- package/scripts/judgment_circuit_lint.sh +239 -0
- package/scripts/novelty_claim_check.sh +193 -0
- package/scripts/relay_channel.sh +645 -0
- package/scripts/reviewer_capability_corpus.tsv +124 -0
- package/scripts/selfcheck.sh +106 -0
- package/scripts/session_close_check.sh +134 -1
- package/scripts/test_branch_claim_lanes.sh +231 -0
- package/scripts/test_dispatch_log_lanes.sh +35 -1
- package/scripts/test_field_canon_lanes.sh +142 -0
- package/scripts/test_hook_source_gate_lanes.sh +81 -0
- package/scripts/test_marker_crossfamily_lanes.sh +132 -0
- package/scripts/test_marker_floor_lanes.sh +9 -8
- package/scripts/test_relay_channel_lanes.sh +583 -0
- package/scripts/test_reviewer_capability_conformance.sh +173 -0
- package/scripts/test_wizard_snippet_merge_lanes.sh +104 -11
- package/scripts/utterance_landing_check.sh +209 -0
- package/templates/.git-hooks/pre-commit +297 -13
- package/templates/settings.Compaction.snippet.json +56 -0
- package/templates/settings.FieldCanon.snippet.json +51 -0
|
@@ -15,13 +15,32 @@ but zero real emits. So the gate scores **identities by evidence**, and the hone
|
|
|
15
15
|
| Status | Meaning | Bar |
|
|
16
16
|
|---|---|---|
|
|
17
17
|
| π’ GREEN | **REALIZED** | a concrete track-record artifact proves the identity fired for real, nβ₯1 (a real gate block, a measured probe, a real orchestration record) β *not* a doc that describes it |
|
|
18
|
+
| π΅ RC | **RELEASE-CANDIDATE** | implemented **and** calibrated on a known pair **and** its own self-test green β but it has not yet fired in a real situation. All three legs, each with an evidence line; two out of three is π‘ |
|
|
18
19
|
| π‘ YELLOW | **PARTIAL** | pieces work but no single closed track record (e.g. two half-pipelines that never connected end-to-end) |
|
|
19
20
|
| π΄ RED | **μ΄μλ‘ (ideal-only)** | documented aspiration, never actually run; or the source itself says "not built yet / named target" |
|
|
20
21
|
|
|
21
|
-
**All-green rule**: ship + tag only when **every** identity is π’. A π‘ or π΄ blocks the tag β and names
|
|
22
|
+
**All-green rule**: ship + tag only when **every** identity is π’. A π΅, π‘ or π΄ blocks the tag β and names
|
|
22
23
|
exactly what real run is missing. The remedy for RED is never to relabel it green; it is to **run it and
|
|
23
24
|
leave the artifact** (the operator's standing rule: "μ΄μλ‘ μ΄λ©΄ μ€μ λ‘ λλ €λ΄μ μ€μ μ λ¨κ²¨μΌ νλ€").
|
|
24
25
|
|
|
26
|
+
**π΅ RC is deliberately not green.** It is the rung for "we built it, we proved the instrument works, and
|
|
27
|
+
our own tests pass" β a real and reportable milestone, and *still* short of the bar, because a self-test
|
|
28
|
+
is authored by the same party it tests. The boundary is the one named in
|
|
29
|
+
`[[feedback_adversarial_review_not_substitute_for_first_use]]`: **passing your own tests is not firing in
|
|
30
|
+
the real situation**, and the first real use is repeatedly what invalidates the design. So RC never opens
|
|
31
|
+
a `v1.0.0`; only REALIZED does.
|
|
32
|
+
|
|
33
|
+
> The origin note (2026-08-08 session card) drew this ladder with RC and REALIZED **both** marked π’. That
|
|
34
|
+
> shorthand is fine in a card and breaks here: this gate's rule reads literally as "every identity π’ β
|
|
35
|
+
> ship", so a green RC would open the tag that the same note says RC must not open. Distinct symbol, same
|
|
36
|
+
> intent.
|
|
37
|
+
|
|
38
|
+
**RC is self-reported by construction β so it carries evidence, not a claim.** Each RC status names, on
|
|
39
|
+
one line, *what ran and what came out* (the non-vacuity requirement borrowed from the 4-axis marker's
|
|
40
|
+
`axis2-evidence`: a recorded verdict, count or fixture result β never "it works"). An RC without that line
|
|
41
|
+
is π‘. Where an instrument could not be calibrated against a real case, the leg ships labelled
|
|
42
|
+
**`UNCALIBRATED`** rather than silently counted (`not found β 0`).
|
|
43
|
+
|
|
25
44
|
## Dominance, not concession β the AlexNet bar
|
|
26
45
|
A harness earns the right to say "we compose with other harnesses" **only from proven dominance**, never
|
|
27
46
|
as a humble concession. The reference is AlexNet: on data it had never seen, it did not *participate* β it
|
|
@@ -102,26 +121,175 @@ notes claim more green than the audit shows is the defect the gate exists to pre
|
|
|
102
121
|
the non-all-green status (β’β€ π’, β£ π‘, β β‘ π΄), elected to tag `v0.1.0` as this honest baseline; the
|
|
103
122
|
decision is logged here and the tag's notes state the real status.
|
|
104
123
|
|
|
124
|
+
## The four engines β what has to run for an identity to be reachable at all
|
|
125
|
+
|
|
126
|
+
An identity is what a harness *claims*; an **engine** is a capability the harness must actually possess
|
|
127
|
+
for that claim to be reachable. They are different axes, and scoring only identities hides *why* one is
|
|
128
|
+
stuck: the failure shows up in the identity and the cause sits in the engine.
|
|
129
|
+
|
|
130
|
+
| Engine | What it is | Why an identity needs it |
|
|
131
|
+
|---|---|---|
|
|
132
|
+
| **external-grounding** | Asking the world on its own initiative β reaching outside the repo before asserting novelty or settling a design, without being told to | Anything **new** has no known answer inside; asserting `net-new` from an internal grep alone is how a phantom is born |
|
|
133
|
+
| **judgment-circuit** | A forged decision circuit: what counts as success, which way to lean under uncertainty, what is out of scope, what never happens | Anything **autonomous** has no direction without one; the harness fills the vacuum with volume instead |
|
|
134
|
+
| **ship-gate** | Mechanical blocking before an irreversible surface β commit, publish, delete, rewrite | Anything that **ships** needs a last line that does not depend on remembering |
|
|
135
|
+
| **context-continuity** | Not losing the thread mid-run β across compaction, sub-agents, machines and sessions | Anything **long** loses its own premises first, and the loss is silent |
|
|
136
|
+
|
|
137
|
+
**Naming rule β do not translate `judgment-circuit` as "soul".** In Korean the operator's word is μνΌ, but
|
|
138
|
+
the English word reads as *persona*, and the single largest finding of the 105-run measurement behind this
|
|
139
|
+
engine was precisely that **an identity declaration is not a judgment circuit** ("λλ ~μ΄λ€" measured as a
|
|
140
|
+
net loss; removing it recovered +0.67 on the weak tier). A one-word translation re-fuses exactly what the
|
|
141
|
+
measurement separated.
|
|
142
|
+
|
|
143
|
+
**Why engines gate the advertised capabilities**: the harness's most-advertised surfaces β incubating a new
|
|
144
|
+
project, orchestrating a multi-harness cluster β are simultaneously *long, autonomous, novel and shipping*.
|
|
145
|
+
They therefore load all four engines at once, which is why a harness with a mature ship-gate and little
|
|
146
|
+
else appears to fail *at* those surfaces while the cause is underneath them.
|
|
147
|
+
|
|
148
|
+
> β οΈ **The identityβengine mapping below was composed by the AI, not measured.** It is a structural
|
|
149
|
+
> hypothesis, not established causality. The way to test it is to bring **one** engine up a rung and watch
|
|
150
|
+
> whether the mapped identity moves; until then, read the column as a claim about *what to try*, not about
|
|
151
|
+
> what is known. (The counterweight matters here: a mapping that looks tidy is the easiest thing to start
|
|
152
|
+
> citing as a finding.)
|
|
153
|
+
|
|
105
154
|
## FH's own status (2026-07-14) β NOT yet all-green
|
|
106
155
|
|
|
107
|
-
|
|
156
|
+
Engine column added 2026-08-08 (mapping is the unverified hypothesis flagged above; Status column is
|
|
157
|
+
unchanged and keeps its own 2026-07-14 evidence).
|
|
158
|
+
|
|
159
|
+
| # | Identity | Engines it loads | Status | Evidence / what's missing |
|
|
160
|
+
|---|---|---|---|---|
|
|
161
|
+
| β’ | κ±°λ²λμ€ κ²μ΄νΈ (governance) | ship-gate | π’ GREEN | pre-commit/pre-push physically block; moat measured 3β4 family blind (HITL 8/8 ABSENT); cross-family caught a real companion-store-name leak 2026-07-14 (fail-closed) |
|
|
162
|
+
| β€ | μ¦νμ (amplifier) | judgment-circuit | π’ GREEN | short-intentβliterature-groundingβultimate-doc real instances; rules-diet β18.2k measured; intent-routing probe 94% (below) |
|
|
163
|
+
| β£ | νλ°ν°μ΄βμ‘°μ§ μ ν (**π΅ RC, 2026-08-09**) | external-grounding | π΅ RC | frontier-digest launchd auto + AX submission docs both real, but digestβorg never closed as ONE pipeline. **2026-08-09**: the missing link was built β `scripts/digest_landing_check.sh` extracts the digest's candidate table into probes and reuses the existing landing checker (no second verifier). Self-test 8 lanes green. **π΅ RC (2026-08-09)**: the mtime defect that initially held it back is closed β the since-filter now splits two axes (git-tracked β commit time via `git log --since`; gitignored `tracks/**` β mtime, the only evidence that axis has; dirty-tracked β `UNMEASURED`), and **two lanes pin that split**: a file with only a fresh mtime is *not* counted, and a file with only a fresh commit *is* counted even when its mtime is stale. The second lane matters β without it the fix degenerates into "discard all tracked files so only negatives pass" (named by the cross-family reviewer). Self-test **10 lanes** green. **What remains is a named residual, not a calibration gap**: `file-change β token-introduction` β a file committed after the digest may carry the token from before (closing it needs token-level diff, which does not fit the checker's interface). The instrument therefore prints, and this row states, that it is a **screener, not an adjudicator**: hits must be opened. Four real runs, four hand-verifications, four defects found |
|
|
164
|
+
| β | λ©ν°νλ€μ€ ν΄λ¬μ€ν° (**π΅ RC, 2026-08-09**) | context-continuity | π΅ RC | routing already ran for real (17 nodes, sidecar-orchestrator, Skill Bus). **The relay half is now built rather than specified**: `capability_composition_contract.md` (2026-08-02) was a complete spec with **zero implementing code** β the β blocker was missing wiring, not missing design ([[feedback_built_but_not_wired]]). `scripts/relay_channel.sh` executes it (strictest-wins merge Β· typed invocation Β· checks 1/2/3 Β· short-circuit Β· causal binding), `scripts/test_relay_channel_lanes.sh` carries **64 lanes, BLOCK/PASS symmetric**, and three arms ran across **two real field harnesses** (pmh-dev Β· qasp-dev) on FH's own assets. β **The measured result is the divergence arm, and its mechanism is not what the first draft of this row said.** On `templates/.git-hooks`, `qasp` alone returns exit 0 β a single-node pass would have shipped it β and the composition returns `BLOCKED` because `pmh` returns `FINDINGS`. But `qasp`'s exit 0 is `degrade-scan: no scannable (py/sh) target files`: **zero files were scanned.** The qasp copy predates pmh's 2026-07-28 shebang pass, so extension-less hook files are invisible to it, and its exit 0 means *no target*, not *clean*. So the composition did not catch a substantive disagreement between two harnesses β it caught **a single node rendering an unmeasured surface as a pass**, which is `[[feedback_not_found_is_not_zero_family]]`, and structurally the spec's own Β§β.4 B1 ("the exit 0 that means I never started"). That is a *stronger* result than the first framing and a narrower one: it demonstrates the union catching a blind spot, not decorrelated judgment. **Correction also to the order claim**: both orders return `rc=2`, but in the pmh-first order the chain short-circuits at node 1 and qasp never runs β only the qasp-first order actually exercises the union. Non-decorative: reverting each wiring line reddens lanes and no reversion passes silently. **Why this is RC and not π’**: (a) the row's *other* half, external-harness recommend, is still parked; (b) the spec's own named gap `scripts/capability_registry_check.sh` (M1βM5 registration + M4 pair) still does not exist β only the call moment is closed; (c) the two nodes are **copies of one scanner at different staleness** (all three copies β pmh 237 ln, qasp 121 ln, FH 269 ln β share a byte-identical 12-line header; the clean arm's two `out_sha` were identical), so the run proves the channel turns and that composing unequal copies has value, not that two independent judgments were decorrelated. Artifact: `tracks/_meta/identity_audit_2026-08-09_relay_channel.md` |
|
|
165
|
+
| β‘ | νλ‘μ νΈ μΈνλ² μ΄ν° (**π΅ RC, 2026-08-09**) | context-continuity + judgment-circuit | π΅ RC | **RC μΈ λ€λ¦¬κ° μ°λ€** β (a) ꡬν: `chamber_run.sh` 6λ¨κ³ κ²μ΄νΈ (b) known-pair: λ¬λ κ²μ΄νΈ **18 λ μΈ**(`test_chamber_run_lanes.sh`, BLOCK/PASS λμΉ β PASS arm μ΄ μμ΄μΌ "μ λΆ λ§λ κ²μ΄νΈ"λ κ±Έλ¦°λ€) + μμ μ¦μΈ **16 λ μΈ**(`chamber_witness.sh`) (c) self-test μ΄λ‘. **μ€μν© λ°ν λκΈ° = formal chamber EMIT μμ§ 0** β κ·Έκ²μ΄ RC κ° π’ μ΄ μλ μ΄μ μ΄μ RC μ μ κ·Έ μ체λ€. β οΈ **κ·Έ 0 μ ν΄μμ΄ 2026-08-09 μ λ°λμλ€**: μ§κΈκΉμ§ *"μ±λ²κ° μ격ν΄μ"* λ‘ μ½μμΌλ, KILL λ ν보 λ€μκ° **λ©ν-ν** μ΄κ³ μ μΌν EMIT(`forge-wiki`)λ§ **νλ-ν** μ΄λ€ β μ¦ *λ³μ μ μμλ* κ² μλλΌ **μ μ΄μ λμμ΄ μλ νλ³΄κ° λ€μ΄μμ** κ°λ₯μ±μ΄ μλ€. νλ β₯ λ©ν νλ‘νμΌκ³Ό μ¨μ(precocial) κΈ°μ€ μ μ: `harness_incubator_doctrine.md Β§3-a`. β οΈ κ·Έ λΆλ₯λ **μ¬νμ μ΄λ€μ‘κ³ n=9** λΌ κ°μ€μ΄λ€ β μ¬μ λ±λ‘ ν λ€μ λ°μ μμΈ‘ν΄μΌ κ²°κ³Όκ° λλ€. μλ μ νμ μ€μ μ΄λ ₯μΌλ‘ λ¨κΈ΄λ€ |
|
|
166
|
+
| β‘-old | (μ΄λ ₯) νλ‘μ νΈ μΈνλ² μ΄ν° | context-continuity + judgment-circuit | π‘ PARTIAL | incubation is running β **stockbattle is being incubated now** (S1 built, mid-flight) + qasp/pmh spin-out precedent + scaffold-emit shipped (doctrine: "emit shipped today as scaffold+approval; the chamber flow is the named target"). **Corrected 2026-08-08** (the old text read "6 runs, 6 KILL β¦ 0/6", which was stale on both counts, and the ledger itself was missing a run): hand-counted from `tracks/_chamber/INDEX.md` β **9 full runs (#2β#10), 8 KILL, 1 EMIT** (#1 is a trigger probe, not a full run). Runs #5β#6 *measured* the emit-worthiness criterion (net-new β§ artifact-shaped β§ real-data-precision-adequate β§ hub-state-independent); run #6 confirmed the graduation-order principle β hub-internal proof before standalone extraction, never the reverse. **The π‘ is now held for a different reason than before.** The old reason ("no closed emit-via-incubation yet") is false: run #9 `forge-wiki` emitted and shipped publicly under operator approval with the Pre-Publish gate passed. What is *not* proven is that the **formal chamber flow** produced it β that run's workspace holds only an `EMISSION_VERDICT.md`, with no `INTENT.md`, `BUDGET.md` or `SIM_NOTES.md`, so the intent/budget/blind-persona gates have no artifact and the verdict was written after the fact. The first run to complete the formal flow end-to-end is #10 (2026-08-08, 3 blind isolated personas) and it KILLed. So: **the identity has fired once, the mechanism has not yet been shown to be what fired it**, and the dominance result every π’ owes is still outstanding β π‘ |
|
|
167
|
+
|
|
168
|
+
### β‘ promotion criteria β and what the criteria themselves turned out not to be able to check
|
|
169
|
+
|
|
170
|
+
β‘'s π‘ has been re-argued on different grounds each round, every round re-deriving the bar from scratch.
|
|
171
|
+
This section exists so the next round starts from a stated condition. **A first draft of it was refuted by
|
|
172
|
+
cross-family review before it was committed**, and the refutation is more useful than the draft was, so
|
|
173
|
+
both are recorded.
|
|
174
|
+
|
|
175
|
+
**What the draft got wrong.** It scored run #9 `forge-wiki` as **P1 FAIL** on the grounds that its
|
|
176
|
+
workspace holds only `EMISSION_VERDICT.md` β no `INTENT.md`, `BUDGET.md`, `SIM_NOTES.md`. But that
|
|
177
|
+
verdict file *contains* the substance those files would hold: the net-new determination (two survey
|
|
178
|
+
generations, 15+ systems / 6 standards cross-checked), the artifact-shaped determination, and the
|
|
179
|
+
real-code precision leg with a raw-data anchor (`forge-wiki/tests/sim_data_2026-07-18/`, N=50 concurrent
|
|
180
|
+
writers, A/B/C design contrast, reps=3, contaminated reps voided and re-run). Absent **files** were read
|
|
181
|
+
as an absent **gate** β `[[feedback_not_found_is_not_zero_family]]`, committed by the very section citing
|
|
182
|
+
the rule it broke. The honest score for #9 is **UNKNOWN**, not FAIL.
|
|
183
|
+
|
|
184
|
+
**What actually holds β‘ at π‘, once the formalism is stripped out.** Not the missing filenames β the
|
|
185
|
+
missing **ordering witness**. The claim that would promote β‘ is *the mechanism screened this, and then it
|
|
186
|
+
emitted*; what #9 can show is *it emitted, and a verdict describes screening*. Nothing distinguishes a
|
|
187
|
+
gate that ran before the outcome from a record written after it.
|
|
188
|
+
|
|
189
|
+
**And that witness cannot currently be produced.** `tracks/**` is gitignored (`.gitignore:40` β verified
|
|
190
|
+
per file with `git check-ignore -v`), so no chamber artifact is under version control, and mtimes are the
|
|
191
|
+
only ordering evidence there is. Mtimes are trivially forgeable. So the draft's own check β "written
|
|
192
|
+
*before* the verdict, compare mtimes" β **cannot be satisfied by any run, honest or not**. It was an
|
|
193
|
+
unreachable condition, which is the shape that trains people to delete the thing being counted
|
|
194
|
+
(`[[feedback_unreachable_done_when_trains_evasion]]`).
|
|
195
|
+
|
|
196
|
+
**So the promotion condition is one thing, and it is a build, not a check:**
|
|
197
|
+
|
|
198
|
+
| | Condition | Check class | Status |
|
|
108
199
|
|---|---|---|---|
|
|
109
|
-
|
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
200
|
+
| **P1** | An EMIT run leaves an ordering record that does not depend on trusting the author β the intent/budget/sim record committed, hashed, or otherwise witnessed **outside** the gitignored workspace, before the verdict | mandatory-pass | **channel now exists (2026-08-08)** β `scripts/chamber_witness.sh`, wired into `chamber_run.sh` steps 2β5. **Still unsatisfied**: no run holds a witness yet |
|
|
201
|
+
|
|
202
|
+
**P1's channel was built, and that is not the same as P1 passing.** The row above said *not buildable
|
|
203
|
+
today*; that is no longer true, and the reason it was true is worth keeping because it names the shape of
|
|
204
|
+
the fix. The blocker was never "we lack a checker" β it was that `tracks/**` is gitignored, so the only
|
|
205
|
+
ordering evidence was mtime, which is trivially forgeable. The channel takes the **second form P1's own
|
|
206
|
+
sentence already permitted β `hashed`**: the artifacts stay in the private workspace and only their
|
|
207
|
+
SHA-256 goes into a tracked ledger (`knowledge/shared/learnings/chamber_ordering_witness.yaml`). Content
|
|
208
|
+
disclosure is zero, and the commit graph carries the ordering.
|
|
209
|
+
|
|
210
|
+
**What makes the witness bind here specifically**: `main` runs a `non_fast_forward` ruleset, so pushed
|
|
211
|
+
history cannot be rewritten β the commit order is not something the author can retroactively change. That
|
|
212
|
+
is a *precondition*, not a property of the script: if the ruleset is ever relaxed, this witness weakens
|
|
213
|
+
with it, and that dependency is recorded in the script header rather than assumed.
|
|
214
|
+
|
|
215
|
+
**What it does NOT prove β stated because the failure to state it is the defect this section exists for.**
|
|
216
|
+
A commitment proves *this content was fixed at this time*. It does not prove *the author was honest*: one
|
|
217
|
+
can still decide the outcome first and write an INTENT to match before committing. The attack it closes is
|
|
218
|
+
**retroactive rewriting** β producing a verdict and then editing INTENT to claim it was screened β and
|
|
219
|
+
`verify` returns `TAMPERED` for exactly that (calibrated lane, not a claim).
|
|
220
|
+
|
|
221
|
+
**Known-pair calibrated, 16 lanes**: correct order β `WITNESSED(0)` Β· verdict-committed-first β
|
|
222
|
+
`UNORDERED(1)` Β· post-hoc rewrite β `TAMPERED(1)` Β· unrecorded run β `UNWITNESSED(2)` Β· recorded but
|
|
223
|
+
uncommitted β `PENDING(2)` Β· missing artifact / bad slug charset / ledger-write failure β `rc=10` Β·
|
|
224
|
+
verdict hash absent β `INCOMPLETE(2)` Β· gates and verdict in the **same commit** β `UNORDERED(1)` Β·
|
|
225
|
+
only *some* gate artifacts before the verdict β not a pass Β· a duplicate hash from another run β
|
|
226
|
+
does not hide `PENDING` Β· single-arg call β survives Β· control β still 0.
|
|
227
|
+
**`2` is not a pass** β `not found β 0` is enforced in the exit code itself, so a caller cannot read
|
|
228
|
+
"no record" as "no problem".
|
|
229
|
+
|
|
230
|
+
**Seven of those lanes exist because a cross-family audit found the first version fail-open**, and the
|
|
231
|
+
finding rate is the point: the author's own review produced **zero** of them. codex (gpt-5.5) returned 11
|
|
232
|
+
defects with source lines and **reproduced four of them by execution** β a ledger write to an invalid path
|
|
233
|
+
still printed `witnessed` and returned 0; gates and verdict in one commit passed; `INTENT` alone before the
|
|
234
|
+
verdict passed while `BUDGET`/`SIM_NOTES` landed after it; and a hash reused from another run masked an
|
|
235
|
+
uncommitted entry. Worst of all, **a run with no verdict hash at all returned `0`, which the runner rendered
|
|
236
|
+
as "usable as identity β‘ promotion evidence"** β the witness channel issuing a green with no witness, which
|
|
237
|
+
is the exact failure it was built to prevent. Each fix carries a lane, and two were proven non-decorative by
|
|
238
|
+
revert (reverting either reddens exactly one lane, 1/16).
|
|
239
|
+
|
|
240
|
+
**The nine historical runs stay `UNWITNESSED`, and are not back-filled.** Hashing them now would record
|
|
241
|
+
the artifacts as they are *today*, after their verdicts β a record written after the outcome, which is
|
|
242
|
+
precisely the thing the witness exists to distinguish. Back-filling would produce a ledger that looks
|
|
243
|
+
witnessed and proves nothing. So run #9 `forge-wiki` remains **UNKNOWN** on ordering, as the section above
|
|
244
|
+
already concluded, and P1 is first satisfiable by the **next** chamber run.
|
|
245
|
+
|
|
246
|
+
**P2 (dominance) is deliberately NOT listed**, and the reason is a finding about the gate rather than
|
|
247
|
+
about β‘. The draft required it, citing Β§"Gate consequence". Checked against the table: β’ does carry a
|
|
248
|
+
dominance result (moat measured 3β4 family blind, HITL 8/8 ABSENT), but **β€ is π’ on `intent-routing
|
|
249
|
+
probe 94%` β a self-measurement, not a head-to-head**. Requiring dominance of β‘ while β€ holds π’ without
|
|
250
|
+
it is a bar invented for one row. The inconsistency is real and it is **the gate's, not β‘'s**: either
|
|
251
|
+
Β§Gate-consequence binds every π’ and β€ is over-scored, or it is advisory and β‘ must not be held to it.
|
|
252
|
+
Resolving that is a separate change to the status definitions β flagged here, not silently settled by
|
|
253
|
+
scoring β‘ against a rule the table does not apply uniformly.
|
|
254
|
+
|
|
255
|
+
**Recurrence count, stated precisely because the draft muddled it.** The π‘ has been re-argued **three**
|
|
256
|
+
times; "promotion attempted without stated criteria" has been *recognized as a problem* **once** (today).
|
|
257
|
+
Those count different things, and the draft cited N=1 while asserting three re-arguments in the same
|
|
258
|
+
paragraph. Neither number licenses a checker right now β P1 is not implementable at all until the
|
|
259
|
+
ordering channel exists, so there is nothing to mechanize yet.
|
|
114
260
|
|
|
115
261
|
**Cross-cutting measured (intent-based autonomous completion)**: blind floor-tier Sonnet trigger-accuracy
|
|
116
262
|
probe (n=10, 2026-07-14): **should-fire 7.5/8 (94%), false-fire 0/2**. One weak trigger (simulate-first /
|
|
117
263
|
incubator entry absorbed into deep-clarify) β the identity-β‘ weakness surfaces in routing too.
|
|
118
264
|
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
265
|
+
> **μ΄ νκ° λ€λ£¨λ κ²μ 3μΈ΅ μ€ ν μΈ΅μ΄λ€.** 5λ μ 체μ±μ΄ 무μμ λ°μΉκ³ (4λ μμ§) 무μμΌλ‘
|
|
266
|
+
> λ²Όλ €μ§λμ§(3λ¨ κ³΅μ )λ `fh_three_layer_canon.md` κ° μ λ³Έμ΄λ€. **μ΄ νμ `engine` μ΄μ΄ κ³§
|
|
267
|
+
> 4λ μμ§**μ΄λ©°(judgment-circuit=μνΌ Β· ship-gate=νμ§κ²μ΄νΈ Β· context-continuity=λ§₯λ½μ μ§ Β·
|
|
268
|
+
> external-grounding=μ§λ¬ΈνκΈ°), κ·Έ λμμ μλ‘ λ§λ κ²μ΄ μλλΌ μ΄ νμ μ΄λ―Έ μλ κ²μ΄λ€.
|
|
269
|
+
|
|
270
|
+
**Verdict (2026-08-09 β supersedes the 2026-07-14 line)**: FH is tagged **`v0.1.0` = honest baseline**,
|
|
271
|
+
not all-green. β’β€ are π’, **β β‘β£ are π΅ RC**, **none π΄** β the `v0.1.0` notes state this and make no
|
|
272
|
+
all-green claim (per the refined 0.xβ1.0 mapping above). **`v1.0.0` remains the all-green target.**
|
|
273
|
+
|
|
274
|
+
> *Why this paragraph is being rewritten rather than edited in place*: it read **"β β‘β£ π‘"** for three
|
|
275
|
+
> sessions **after** the rows above had moved β β‘ to RC on 2026-08-09 (PR #281), β£ on 2026-08-09
|
|
276
|
+
> (PR #283), β in this run. Each session corrected its own row and left the summary alone, which is
|
|
277
|
+
> `[[feedback_half_fix_propagation_boundary]]` inside a single file: the propagation boundary is not
|
|
278
|
+
> only "other files", it is **every place in this file that restates the same fact**. A summary that
|
|
279
|
+
> contradicts its own table is worse than no summary, because it is the line a reader quotes.
|
|
280
|
+
|
|
281
|
+
What now blocks `v1.0` is **closing the π΅s** β RC means the mechanism stands in the lab, π’ means it
|
|
282
|
+
walked outside:
|
|
283
|
+
|
|
284
|
+
```
|
|
285
|
+
β external-harness recommend (cluster-wizard, still parked) Β· capability_registry_check.sh
|
|
286
|
+
(registration-moment M1βM5, named by the spec, still absent) Β· a run across two genuinely
|
|
287
|
+
different capabilities rather than two forks of one scanner
|
|
288
|
+
β‘ a formal chamber EMIT β the mechanism firing in a real situation, not a retrofitted verdict
|
|
289
|
+
β£ file-change β token-introduction β the instrument is a screener, not an adjudicator
|
|
290
|
+
```
|
|
291
|
+
|
|
292
|
+
Each remedy is a run that leaves an artifact, tracked in `tracks/_meta/identity_audit_*.md`.
|
|
125
293
|
|
|
126
294
|
> **β β‘ correction (2026-07-14)**: an earlier pass marked β β‘ π΄ by collapsing each identity onto its most
|
|
127
295
|
> advanced *single mechanism* β β‘ onto the formal chamber EMIT (0/5), β onto the continuous-relay channel.
|