@chrono-meta/fh-gate 1.4.41 → 1.4.43

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (33) hide show
  1. package/AGENTS.md +5 -3
  2. package/CATALOG.md +6 -0
  3. package/CLAUDE.md +65 -130
  4. package/docs/CONTRIBUTING.md +2 -2
  5. package/knowledge/shared/dialogue/ai_dialogue_playbook.md +137 -0
  6. package/knowledge/shared/dialogue/claude_code_runtime_flow.md +170 -0
  7. package/knowledge/shared/dialogue/memory_intent_recall.md +209 -0
  8. package/knowledge/shared/harness-core/claude_md_gate_details.md +170 -0
  9. package/knowledge/shared/harness-core/companion_store_pluggable_cross_audit_2026-06-11.md +118 -0
  10. package/knowledge/shared/harness-core/crucible_mode.md +112 -0
  11. package/knowledge/shared/harness-core/deep_research_capability_ladder.md +122 -0
  12. package/knowledge/shared/harness-core/fh_detail_protocols.md +163 -0
  13. package/knowledge/shared/harness-core/fh_ecosystem_positioning.md +147 -0
  14. package/knowledge/shared/harness-core/fh_opencode_governance_wrapper.md +163 -0
  15. package/knowledge/shared/harness-core/fh_synergy_playbook.md +217 -0
  16. package/knowledge/shared/harness-core/gate_locality_principle.md +57 -0
  17. package/knowledge/shared/harness-core/goal_quench_anthropic_issue.md +104 -0
  18. package/knowledge/shared/harness-core/harness_6axis_framework.md +136 -0
  19. package/knowledge/shared/harness-core/harness_design_decision_lens.md +108 -0
  20. package/knowledge/shared/harness-core/harness_frontier_diagnosis_2026-06-02.md +102 -0
  21. package/knowledge/shared/harness-core/hub_compounding_loop.md +109 -0
  22. package/knowledge/shared/harness-core/hub_maturity_roadmap.md +201 -0
  23. package/knowledge/shared/harness-core/hybrid_orchestration_architecture_roadmap.md +196 -0
  24. package/knowledge/shared/harness-core/live_surface_automation_pattern.md +110 -0
  25. package/knowledge/shared/harness-core/measurement-integrity-checklist.md +59 -0
  26. package/knowledge/shared/harness-core/meta_harness_engineering_definition.md +116 -0
  27. package/knowledge/shared/harness-core/multi_model_sidecar_strategy.md +651 -0
  28. package/knowledge/shared/harness-core/persona_container_schema.md +172 -0
  29. package/knowledge/shared/harness-core/return_path_gate.md +120 -0
  30. package/knowledge/shared/harness-core/self_evolution_routine.md +268 -0
  31. package/knowledge/shared/harness-core/skill_quality_rubric.md +71 -0
  32. package/knowledge/shared/harness-core/tpa_schema.md +136 -0
  33. package/package.json +3 -2
@@ -0,0 +1,201 @@
1
+ ---
2
+ name: Hub Maturity 3-Phase Roadmap Frame (template)
3
+ description: Long-term evolution path frame for the hub. Phase I (entering maturity) → Phase II (frontier following) → Phase III (frontier leading) 3-stage model. Fixes the gap, output, completion criteria, and transition conditions for each phase as a quarterly re-diagnosis reference document. The maturity axis is the parent frame referenced by monthly level snapshots and quarterly re-diagnosis. **This file is a template — in actual hub operation, add project-specific information (timing, asset names, identifiers) to write an operating copy**.
4
+ type: reference
5
+ date: 2026-04-28
6
+ tags: [harness, maturity, roadmap, 3-phase, frontier-tracking, evolution, compounding, long-term, hub, strategic, template-frame]
7
+ scope: hub-template
8
+ ---
9
+
10
+ # Hub Maturity 3-Phase Roadmap (frame)
11
+
12
+ ## Why this document exists
13
+
14
+ If the monthly level snapshot captures **"where are we now"**, this document captures **"where are we going and when do we transition"**. If the quarterly re-diagnosis captures **current position vs industry frontier**, this document captures **the evolution of the relationship with the frontier itself** (follower → leader).
15
+
16
+ **Core vision**:
17
+
18
+ > The hub model matures enough to **enter the maturity phase**, **follow** the frontier's advancement direction periodically, and after that, the path where **we become the frontier**.
19
+
20
+ This vision is made explicit as 3 phases so that at each quarterly re-diagnosis, "which phase are we in · what are the conditions for the next phase transition" can be judged.
21
+
22
+ ---
23
+
24
+ ## 1. 3-Phase Overview
25
+
26
+ | Phase | Scope | Core gap (example) | Representative output (example) | Standard duration (example) |
27
+ |---|---|---|---|---|
28
+ | **I. Entering maturity** | Infrastructure/asset building + routine establishment | Operational gap axes remain · automation not established · 0 external propagation | Weekly audit automation · operations guide · 1-2 external assets · sub-agent judgment · self-diagnosis warning reduction | 3-6 months |
29
+ | **II. Frontier following** | Quarterly re-diagnosis routine + gap auto-detection | Single/dual input sources · no auto gap detection · cadence not confirmed | Quarterly frontier diagnosis 2 times + monthly brief established + external scan cadence confirmed | 5-8 months |
30
+ | **III. Frontier leading** | Self-invented outbound propagation + self-evolving | 0 open-source/presentation record · no self-evolving loop · 0 industry citations | Public refactor · blog/presentation 10+ per year · industry citation case accumulation | Ongoing (no completion) |
31
+
32
+ ---
33
+
34
+ ## 2. Current Position (fill in operating copy)
35
+
36
+ Fill in the following format in the operating copy:
37
+
38
+ ```
39
+ **Phase X · ~N% progress** — infrastructure/asset accumulation status · N remaining completion gaps
40
+ ```
41
+
42
+ ### Achievement table (example format)
43
+
44
+ | Area | Status | Basis |
45
+ |---|---|---|
46
+ | 6-axis framework + feedback loop | ✅ | (corresponding asset path) |
47
+ | Package structure alignment | ✅ | (realignment session) |
48
+ | ... | ... | ... |
49
+
50
+ ### Gaps remaining
51
+
52
+ See §3 Phase I completion criteria.
53
+
54
+ ---
55
+
56
+ ## 3. Phase I Completion Criteria (5 measurable criteria)
57
+
58
+ Phase II entry gate passed when all 5 are met.
59
+
60
+ | # | Condition | Measurement method (frame) |
61
+ |---|---|---|
62
+ | 1 | **Weekly audit automation established** | audit skill run 3+ times + manual N min → auto N min measured |
63
+ | 2 | **Axis N operations guide established** | orchestration guide draft + mode switch 3+ cases accumulated (if gap axis exists) |
64
+ | 3 | **External propagation N cases** | Distributed to external channel + external response received (generally 2 cases recommended) |
65
+ | 4 | **Sub-agent pilot promotion/deprecation judgment** | 2+ week observation + invocation log-based judgment (`accepted ≥ 60%` / `rejected ≥ 40%` / `invocation count ≥ N`) |
66
+ | 5 | **Self-diagnosis warning reduction** | Quarterly self-diagnosis warnings N items → 1 or fewer |
67
+
68
+ ### Completion checklist derived Decision
69
+
70
+ - Phase I completion date = earliest date all 5 criteria are met
71
+ - Completion confirmation event = **Phase II entry meeting** (1 separate session, §4 entry condition check)
72
+ - No Phase II output (e.g., external GitHub scan) before Phase I completion — simplification principle violation
73
+
74
+ ---
75
+
76
+ ## 4. Phase II (Frontier Following) — Entry conditions, cadence, outputs
77
+
78
+ ### 4.1 Entry conditions
79
+
80
+ - Phase I completion 5 criteria **all** met
81
+ - Monthly level snapshot updated 2+ consecutive times
82
+ - §2 current position re-judged as "Phase II · 0%"
83
+
84
+ ### 4.2 Cadence 3-option comparison
85
+
86
+ | Option | Cycle | Pros | Cons |
87
+ |---|---|---|---|
88
+ | **(a) Quarterly only** | 3 months | Simple | Gap detection delayed 3 months (slow if frontier moves fast) |
89
+ | **(b) Quarterly + monthly light scan** | 3 months + 4 weeks | Gap detection within 1 month + 10 min addition to existing monthly routine (minimum invasive) | (None — recommended) |
90
+ | (c) Trigger-based | When stagnation detected | Resource efficient | Stagnation detection criteria + auto-alerts + trigger tags all require new infra. **Simplification principle violation risk** |
91
+
92
+ → **Recommended: (b) quarterly + monthly**. Joining existing monthly routine = minimum invasive. Consistent with simplification principle.
93
+
94
+ ### 4.3 Representative outputs (frame)
95
+
96
+ | Output | Cycle | Content |
97
+ |---|---|---|
98
+ | **Quarterly frontier diagnosis #N** | 3 months | 6-axis level change vs previous edition + 3-5 new frontier techniques integrated |
99
+ | **Monthly brief** | 4 weeks | 10-min scan of GitHub trending, Anthropic/OpenAI blog, and similar — 3-5 line summary |
100
+ | **External GitHub scan actual operation** | Monthly | Monthly 10-min scan of public repos. Sprint Contract in 5 lines |
101
+ | **Auto gap detection** | Scanner feature addition | Flag "frontier diagnosis not updated > 90 days" |
102
+
103
+ ### 4.4 Completion conditions
104
+
105
+ - Quarterly re-diagnosis 2 times + monthly brief 6+ consecutive productions
106
+ - 1+ time **own methodology back-referenced from frontier** found
107
+ - 2+ axes in monthly level snapshot sustained as "leading" judgment
108
+
109
+ ---
110
+
111
+ ## 5. Phase III (Frontier Leading) — Entry conditions, 6 indicators, outputs
112
+
113
+ ### 5.1 Entry conditions
114
+
115
+ - Phase II completion 3 conditions all met
116
+ - 3+ of N self-invented assets **observed as original concepts** for similar industry concepts
117
+ - Public seed repository external contribution record started (at least one of fork/issue/star)
118
+
119
+ ### 5.2 6 Leading indicators
120
+
121
+ | # | Indicator | Measurement |
122
+ |---|---|---|
123
+ | 1 | Public seed repository record | star/fork/issue count |
124
+ | 2 | Blog/presentation | 2-3 per quarter · 10+ per year |
125
+ | 3 | Industry citation cases | External articles/seminars citing this hub/methodology |
126
+ | 4 | External organization adoption | Other companies/departments |
127
+ | 5 | Self-evolving loop demonstration | Skill generates skills |
128
+ | 6 | Self-invented industry original recognition | Adopted as industry term/frame |
129
+
130
+ ### 5.3 Representative outputs
131
+
132
+ - **Public seed repository refactor** — from personal seed to collaboratable template. Team customization + common protocol separation
133
+ - **10+ blog/presentations per year** — externalize 1 self-invented asset per quarter
134
+ - **Self-evolving loop MVP** — audit skill proposes own skill generation + user approval → auto-generate → usage observation → deprecate/improve
135
+
136
+ ### 5.4 Phase III has no "completion"
137
+
138
+ Phase III is an ongoing state. Instead of completion criteria, **3 maintenance conditions**:
139
+ - 3+ of 6 indicators continuously rising
140
+ - "Frontier level maintained" judgment in quarterly re-diagnosis (no regression)
141
+ - Self-diagnosis 8-item checklist failure signal 0 maintained
142
+
143
+ ---
144
+
145
+ ## 6. Common principles for transitions
146
+
147
+ ### 6.1 Consistent simplification principle
148
+
149
+ Phase transitions are **methodology/abstraction level rises**, not **increases in file/skill/rule count**. At each phase transition, do the following first:
150
+
151
+ - [ ] Self-diagnosis 8-item checklist check — confirm 0 failure signals
152
+ - [ ] Files unreferenced 6+ months → move to `archive/` or consolidate
153
+ - [ ] Clean up previous phase temporary outputs made unnecessary by new phase entry
154
+ - [ ] Confirm CLAUDE.md and CATALOG within 200 lines
155
+ - [ ] Re-confirm optimization principle: field harness → simpler over time; meta-harness → complexity earns its scope (purge orphaned/redundant/decorative units)
156
+
157
+ ### 6.2 Transition deferred on simplification failure
158
+
159
+ Above checklist not passed → **phase transition deferred**. Perform previous phase remaining work + simplification work for 1 additional month then re-check.
160
+
161
+ ### 6.3 Phase regression possible
162
+
163
+ **Regression** from Phase II to Phase I also allowed. E.g., frontier following routine missed 2 consecutive times → "manual re-establishment" re-perform part of Phase I. Do not force linear progression.
164
+
165
+ ---
166
+
167
+ ## 7. This roadmap update cycle
168
+
169
+ | Trigger | Update content |
170
+ |---|---|
171
+ | Quarterly re-diagnosis | §2 current position re-judgment + §3·§4·§5 criteria change check |
172
+ | Phase transition event | Record corresponding phase completion + initialize next phase progress |
173
+ | Self-invented asset recognized as industry original | §5.1 entry condition counter update |
174
+ | Simplification principle violation detected | §6 transition deferral activation record |
175
+
176
+ ---
177
+
178
+ ## 8. Operating copy writing guide (for template users)
179
+
180
+ When moving this frame to an operating copy:
181
+
182
+ 1. **Fill §2 current position** — add asset list and basis material paths
183
+ 2. **Concretize §3 5 criteria** — add `current` column + `target date` column (e.g., `2026-05-17`)
184
+ 3. **Recommend §4.2 cadence (b)** — if adopting other option, need to prove §6.1 simplification gate passed
185
+ 4. **Add current counter to §5 6 indicators** — starting from all 0 is natural
186
+ 5. **Update §2 at each quarterly re-diagnosis** + record cumulative changes in §7 trigger table
187
+
188
+ ---
189
+
190
+ ## 9. Related assets
191
+
192
+ - 6-axis framework — `knowledge/shared/harness-core/harness_6axis_framework.md` (frame premise)
193
+ - Monthly level snapshot — `knowledge/shared/harness-core/harness_level_snapshot_*.md` (current position)
194
+ - Quarterly re-diagnosis — `knowledge/shared/harness-core/harness_frontier_diagnosis_*.md` (following basis)
195
+ - Feedback automation — `knowledge/shared/harness-core/hub_compounding_loop.md` (Phase transition gate input)
196
+
197
+ ---
198
+
199
+ ## 10. One-line conclusion
200
+
201
+ **Phase I completion → Phase II frontier following (b)cadence → Phase III leading. Simplification principle is the common gate for each transition**.
@@ -0,0 +1,196 @@
1
+ ---
2
+ name: hybrid-orchestration-architecture-roadmap
3
+ description: Proposed (not-yet-implemented) architecture roadmap for fh as an intelligent hybrid orchestration engine — Claude Code as main driver, other CLIs/APIs as runtime-discovered sidecars, with Zero-Config standalone fallback for plugin-only users. Design intent + reconciliation against existing FH assets.
4
+ type: roadmap
5
+ date: 2026-06-09
6
+ status: proposed (design intent — NOT implemented; see §Status & Reconciliation)
7
+ tags: [hybrid-orchestration, sidecar, zero-config, jit-probing, install-wizard, roadmap, proposed]
8
+ ---
9
+
10
+ # Hybrid Orchestration Architecture — Roadmap (Proposed)
11
+
12
+ > **Status: PROPOSED design intent, not current behavior.** This document captures a
13
+ > forward-looking architecture for fh. Several components it describes — a `config.json`
14
+ > engine-topology file, runtime JIT engine probing, an install-wizard that auto-builds an
15
+ > engine map — **do not exist in FH today**. Read §Status & Reconciliation first to see
16
+ > what is already shipped vs what is aspirational. Nothing here should be cited as a
17
+ > current feature.
18
+ >
19
+ > **Source**: operator design doc (2026-06-09), reconciled to FH conventions on ingest —
20
+ > pinned model versions → `{model-name}` placeholders (per `multi_model_sidecar_strategy.md`
21
+ > §Generalization #4); illustrative pseudo-code marked as such (FH's real layer is markdown
22
+ > methodology + Claude-native automation, not a Python runtime engine).
23
+
24
+ ## 1. Overview
25
+
26
+ fh aims to be a **hybrid orchestration** harness: **Claude Code as the main driver (core
27
+ controller)**, with third-party CLIs / APIs mapped in as **sidecars** according to whatever
28
+ the user's local environment makes available. The goal is to escape single-model dependency
29
+ and **progressively enhance** output by discovering the user's own resources (subscription
30
+ CLIs, API keys) at runtime — while still working in a fully **Zero-Config** state for users
31
+ who copy only a single plugin/skill rather than installing the whole framework.
32
+
33
+ This is the design north star. The mechanism for *delegating to sidecars* is already
34
+ validated and shipped (see `multi_model_sidecar_strategy.md`); the *automatic
35
+ discovery/topology* layer below is the proposed addition.
36
+
37
+ ## 2. Model roles (provider-agnostic)
38
+
39
+ FH convention forbids pinning version numbers or benchmark figures into durable assets
40
+ (they age and become phantom claims), so roles are stated by function, not by pinned model:
41
+
42
+ | Role | Engine slot | Function | Maps to FH skill family |
43
+ |---|---|---|---|
44
+ | **Core Driver** | `{primary-model}` (strongest available; CC host) | Full code edit, architecture, final synthesis | steel-quench, refactor, design consensus |
45
+ | **Context/Analysis Sidecar** | `{large-context-sidecar}` | Bulk log/API-spec parsing → distilled clues | harvest-loop bulk scan |
46
+ | **Test/Util Sidecar** | `{fast-util-sidecar}` | Boilerplate, unit tests, repetitive drops | speed-run / scaffolding |
47
+
48
+ > Capability/benchmark comparisons between providers change every release cycle — consult a
49
+ > live source (`/frontier-digest`, or the `claude-api` skill for Claude specifics) at decision
50
+ > time rather than trusting any number frozen into this doc. No SWE-bench figures are recorded
51
+ > here by design.
52
+
53
+ ## 3. Main driver + sidecar dual channel
54
+
55
+ ```
56
+ ┌─────────────────────────────┐
57
+ │ user terminal (fh) │
58
+ └──────────────┬──────────────┘
59
+ │ [intelligent routing]
60
+ ┌───────────────────┼───────────────────┐
61
+ ▼ (bulk / low-cost) ▼ (final reasoning) ▼ (fast unit-test / repeat)
62
+ ┌─────────────────┐ ┌──────────────────┐ ┌─────────────────┐
63
+ │ context sidecar │ │ Claude Code (main)│ │ util sidecar │
64
+ └────────┬────────┘ └────────┬─────────┘ └────────┬────────┘
65
+ └────────────────────┼─────────────────────┘
66
+
67
+ [final quality verification + apply]
68
+ ```
69
+
70
+ - **Claude Code (main)** keeps full project context/architecture and is the only writer of
71
+ final code. This matches FH's existing rule: *"Host is always single"*
72
+ (`multi_model_sidecar_strategy.md` §Mechanism).
73
+ - **Sidecars** are stateless one-shot `Bash`-invoked processes whose stdout is folded back in
74
+ by the calling skill — **already the shipped mechanism**, not new.
75
+
76
+ ## 4. 3-Tier routing protocol (proposed ordering)
77
+
78
+ When a skill needs a sidecar engine, probe in priority order:
79
+
80
+ ```
81
+ [sidecar request]
82
+ ├── Tier 1: subscription CLI (zero marginal cost)
83
+ │ discover logged-in binaries: `claude -p`, `aider`, `gemini`, `codex`, `gh copilot`
84
+ ├── Tier 2: native API call (pay-per-use)
85
+ │ env keys: GEMINI_API_KEY, OPENAI_API_KEY, ANTHROPIC_API_KEY
86
+ └── Tier 3: self-contained sub-agent split (fallback)
87
+ no external sidecar → Claude Code spawns isolated sub-agent / prompt-chunking
88
+ ```
89
+
90
+ > **Relation to the shipped fallback chain**: `multi_model_sidecar_strategy.md`
91
+ > §Implementation-Patterns already defines a 3-tier *fallback* chain (Copilot CLI → corporate
92
+ > endpoint → direct Gemini/Codex). The ordering here is framed by **cost/access tier** rather
93
+ > than network reachability. These are two lenses on the same fan-out; if implemented, they
94
+ > should be unified into one routing table, not maintained as two competing lists. Tier 3
95
+ > ("Claude internal sub-agent fallback") is the genuinely new contribution — it guarantees the
96
+ > chain never hard-fails even with zero external resources.
97
+
98
+ ## 5. Static inheritance + dynamic JIT fallback (proposed)
99
+
100
+ ```
101
+ execution request (skill / agent)
102
+
103
+
104
+ [config.json present?] ← PROPOSED file; does NOT exist in FH today
105
+ ├── Yes → static-inheritance mode → load priority map → run immediately
106
+ └── No → Zero-Config standalone mode → runtime JIT probing:
107
+ 1. local subscription CLI scan (Tier 1)
108
+ 2. system API-key env scan (Tier 2)
109
+ 3. main-driver sub-agent fallback (Tier 3)
110
+ ```
111
+
112
+ **5.1 Static config inheritance** — if an install-wizard run had persisted an engine map,
113
+ skills would adopt it as default and skip per-call scan overhead. *(Proposed: install-wizard
114
+ does not generate such a file today.)*
115
+
116
+ **5.2 Runtime JIT probing** — if no config exists (or a tool was added after install), the
117
+ harness probes the live environment on-the-fly via the Tier 1→2→3 protocol and binds the best
118
+ path. All probes failing → safe descent to Tier 3 (main model handles it).
119
+
120
+ **5.3 Zero-Config standalone self-reliance** — the key requirement for **Mode C** (plugin/skill
121
+ copied without the full framework, see `.claude/rules/modes_and_value.md`). With no
122
+ `config.json`, the module must **not error** — it switches to Zero-Config standalone mode and
123
+ JIT-probes for whatever local tools exist, then proceeds quietly. Falls back to Tier 3 if none.
124
+
125
+ ```python
126
+ # ILLUSTRATIVE resolution logic — NOT shipped code.
127
+ # FH's actual layer is markdown methodology + Claude-native automation (skills/rules/hooks),
128
+ # not a Python engine. This sketch only shows the intended decision order.
129
+ def resolve_sidecar_engine(preferred="{util-sidecar}"):
130
+ cfg = load_static_config_safe() # Step 1: static config (if any)
131
+ if cfg and cfg.get("mapped_engine"):
132
+ return cfg["mapped_engine"]
133
+ if shutil.which("aider") or shutil.which("{cli}"): # Tier 1: subscription CLI
134
+ return "CLI_WRAPPER_MODE"
135
+ if os.environ.get("{PROVIDER}_API_KEY"): # Tier 2: API key
136
+ return "NATIVE_API_MODE"
137
+ return "CLAUDE_SUBAGENT_FALLBACK" # Tier 3: peaceful fallback
138
+ ```
139
+
140
+ ## 6. Intelligent install-wizard (proposed extension)
141
+
142
+ Today `install-wizard` sets up the periodic-audit notification structure (zshrc hook +
143
+ sentinels + session-start mtime detection). It does **not** build an engine topology. The
144
+ proposed extension would add, at install time:
145
+
146
+ - **Binary probing** — scan `$PATH` for available CLIs/dependencies.
147
+ - **Profile mapping** — write the discovered hybrid config (the proposed `config.json`).
148
+ - **Sanity check** — a light sidecar call to confirm the pipeline actually works.
149
+
150
+ Onboarding UX phrasing (proposed):
151
+ - Minimal env → *"Configured a Claude-Code-only harness. All skills run safely via context
152
+ splitting, no API key needed."*
153
+ - Expansion nudge → *"To enable cheaper large-analysis skills, register a sidecar API key later;
154
+ fh will auto-detect it on the next run and apply sidecar orchestration."*
155
+
156
+ ## 7. sim-conductor & core-skill improvement directions (proposed)
157
+
158
+ - **7.1 `claude -p` non-interactive pipe runtime** — exploit CC's non-interactive flag; pipe
159
+ stdout/stdin between agents; reuse session cache via a harness wrapper.
160
+ - **7.2 Hybrid context bridge** — normalize heterogeneous sidecar output (API JSON vs CLI
161
+ streamed text); embed an extractor (code-block / JSON-structure parser) in a bridge layer so
162
+ callers always get structured responses regardless of invocation path.
163
+ - **7.3 Built-in sub-agent injection** — for minimal envs with no third-party CLI/API: place
164
+ worker prompt specs (e.g. `steel-quench-worker.md`, `harvest-loop-worker.md`) into
165
+ `.claude/agents/`, so the harness spawns Claude's own isolated-context sub-agent pool for
166
+ parallel work. *(Note: per `operations.md`, personal agents in shared repos should be kept
167
+ local via `.git/info/exclude` — an installer-placed agent must respect that boundary.)*
168
+
169
+ ---
170
+
171
+ ## Status & Reconciliation (read this before citing anything)
172
+
173
+ | Component | In FH today? | Where / note |
174
+ |---|---|---|
175
+ | Sidecar delegation via `Bash` (stateless, host-single) | ✅ Shipped + validated | `multi_model_sidecar_strategy.md` (Experiment 1·2) |
176
+ | Cross-provider perspective-diversity rationale | ✅ Shipped | same doc + `steel-quench` Wave 5 |
177
+ | 3-tier *fallback* chain (network-reachability lens) | ✅ Shipped | same doc §Implementation-Patterns |
178
+ | 3-tier *routing* by cost/access (Tier1 CLI→Tier2 API→Tier3 subagent) | 🟡 Partial | reorders the shipped chain; Tier-3 subagent fallback is new |
179
+ | `config.json` engine-topology file | ❌ Proposed | does not exist |
180
+ | Runtime JIT engine probing | ❌ Proposed | does not exist |
181
+ | Zero-Config standalone auto-heal (Mode C) | ❌ Proposed | concept aligns with Mode C in `modes_and_value.md` |
182
+ | install-wizard builds engine topology / sanity-check | ❌ Proposed | install-wizard today = zshrc hook + sentinels only |
183
+ | Hybrid context bridge / built-in sub-agent injection | ❌ Proposed | sim-conductor roadmap item |
184
+
185
+ **Conflicts resolved on ingest** (per FH conventions): pinned model versions + SWE-bench
186
+ figures dropped → `{model-name}` placeholders (avoids phantom/stale claims); Python
187
+ `engine_resolver.py` marked ILLUSTRATIVE (FH has no Python runtime layer); every non-shipped
188
+ component tagged ❌ Proposed above so this roadmap can never be mistaken for current behavior.
189
+
190
+ ## References
191
+
192
+ - `knowledge/shared/harness-core/multi_model_sidecar_strategy.md` — shipped sidecar mechanism + fallback chain (this roadmap extends, does not replace, it)
193
+ - `plugins/fh-meta/skills/install-wizard/SKILL.md` — current install behavior (the topology extension would build on this)
194
+ - `.claude/rules/modes_and_value.md` — Mode C (plugin/skill-only) that Zero-Config self-reliance targets
195
+ - `.claude/rules/operations.md` — sub-agent boundary rules the built-in-injection item must respect
196
+ - `plugins/fh-meta/skills/frontier-digest/SKILL.md` — live source for current model capability/benchmark comparison (do not freeze numbers here)
@@ -0,0 +1,110 @@
1
+ ---
2
+ name: live-surface-automation-pattern
3
+ description: The capability pattern FH routes to when a mapping project needs an agent to drive a live UI surface (web/mobile) — observe-act-verify over the running app, not code/data fetch. FH routes drivers (no-reinvention); the value is the cross-platform observe-act-verify contract, the Appium-less principle, and the hybrid-WebView vision-synthesis rule. Generalizable across mapping projects; not tied to any one project.
4
+ date: 2026-06-14
5
+ tags: [live-surface, ui-automation, observe-act-verify, appium-less, hybrid-webview, no-reinvention, mapping-acceleration]
6
+ ---
7
+
8
+ # Live-Surface Automation Pattern
9
+
10
+ When a mapping project's work lives on a **live UI surface** — a running mobile app or web page whose
11
+ state cannot be reached by reading code or fetching data — FH's posture is the same as for any
12
+ capability it does not own: **detect the need and route to the best driver present, do not build a UI
13
+ engine** (no-reinvention). The value FH adds is the *contract* below, not a new automation tool.
14
+
15
+ This is the acceleration axis behind the "③ map-project acceleration" door's live-surface capability
16
+ (`[[fh-live-surface-acceleration]]`). It was first validated on web (Playwright MCP); the 2026-06-14
17
+ validation extended it to mobile (Android + iOS) and surfaced the two principles that make it work.
18
+ A 2026-06-19 two-surface session then measured the **logged-in-web** (claude-in-chrome) and
19
+ **native-desktop** (computer-use) lanes against the same six-primitive contract — splitting the web
20
+ row into isolated vs logged-in and adding a native-desktop row (see Driver routing).
21
+
22
+ ## The observe-act-verify contract (driver-agnostic)
23
+
24
+ Every driver fills the same six primitives; the runner above them is identical regardless of backend:
25
+
26
+ | primitive | role | verdict impact |
27
+ |---|---|---|
28
+ | launch | bring the app/page up | failure → BLOCKED |
29
+ | screenshot | vision-channel capture | evidence |
30
+ | observe | structured element tree (XML/AX/DOM) → normalized elements | candidate ranking input |
31
+ | tap | element/coordinate action | act |
32
+ | input_text | text entry | act |
33
+ | screen_signature | pre/post comparison | re-observe |
34
+
35
+ A TC/step's intent text is ranked against observed elements (text · description · role · clickable ·
36
+ bounds), the top candidate is acted on, and the post-action observation is checked **mechanically**
37
+ (target text present/absent + candidate count + adapter exit) — no judge LLM decides the verdict.
38
+
39
+ ## Driver routing (no-reinvention)
40
+
41
+ | surface | driver | note |
42
+ |---|---|---|
43
+ | web — isolated (no login) | Playwright MCP (Claude-usable) / Stagehand | isolated session, no operator cookies; generic-coding-agent usable — covers users without a proprietary browser-agent app |
44
+ | web — logged-in (operator's authenticated session) | authenticated-session browser driver — currently claude-in-chrome (`/chrome`) | reaches SSO-behind SPAs the isolated lane can't; `find` → DOM element ref. Current-instance caveats in footnote † |
45
+ | native desktop screen | screen-control driver — currently computer-use (screenshot + coordinate, vision channel) | reaches non-browser native app windows the browser lanes can't; `observe` is vision-coordinate, not an AX/DOM tree. ⚠️ local-session-bound (remote-control), no mobile; `input_text` unverified (5/6 — see note) |
46
+ | Android (emulator **or** real device) | `adb` + `uiautomator dump` + `input tap` | same code path for both — see Portability below |
47
+ | iOS simulator | `idb ui describe-all` + `idb ui tap` (idb-companion + fb-idb) | native AX tree |
48
+ | hybrid WebView region | **vision channel** (screenshot + element detection) | native AX is blind here — see Hybrid rule |
49
+
50
+ FH routes; it does not reimplement these. Drivers are pluggable; the contract is fixed.
51
+
52
+ **Observe channel + measurement (2026-06-19).** computer-use fills `observe` via the **vision channel**
53
+ (screenshot + coordinate) — the same channel the hybrid-WebView rule already mandates (below), now for
54
+ native desktop; the browser and mobile drivers fill it with a structured tree (DOM ref / AX). Measured
55
+ that session: claude-in-chrome closed **all six** primitives on both a no-login and an operator-logged-in
56
+ surface; computer-use closed **five** (launch / screenshot / observe(vision) / tap / screen_signature) —
57
+ `input_text` is keyboard-capable but was not exercised. The lanes cover **different surfaces, not
58
+ redundant ones** — calculator success on computer-use does not imply chrome, and only chrome reaches the
59
+ logged-in session.
60
+
61
+ † **Current-instance caveats (point-in-time, not routing inputs).** claude-in-chrome and computer-use are
62
+ built-in MCPs — verify presence via the `/mcp` UI, not `claude mcp list`. claude-in-chrome needs a direct
63
+ first-party plan (unavailable on Bedrock/Vertex) and a per-site permission grant. Tool surface as measured
64
+ 2026-06-19: claude-in-chrome ~22, computer-use ~24 (counts drift across releases — not a routing input).
65
+
66
+ ## ★ Principle 1 — Appium-less is the enabler
67
+
68
+ Going **through Appium is the failure mode**, not the solution. Appium's WebView-context switching is
69
+ flaky on hybrid apps; in one field case it burned a large automation budget without completing a single
70
+ hybrid flow (repeated run→fail→rerun). The **direct path** (`uiautomator dump` / `idb describe-all` +
71
+ coordinate tap + vision) bypasses that failure mode entirely. Appium is therefore *one regression-stage
72
+ adapter*, not the exploration/execution substrate.
73
+
74
+ ## ★ Principle 2 — Hybrid WebView is opaque to native accessibility → synthesize vision
75
+
76
+ Live-measured (2026-06-14, iOS hybrid sandbox): a WKWebView's **web content does not appear in the
77
+ native accessibility tree** — only native chrome and bridge-triggered native overlays do. A tap on a
78
+ web button fired the web→native bridge and the resulting native picker *did* appear in the AX tree
79
+ (4→10 elements), confirming the boundary precisely.
80
+
81
+ → For hybrid apps, the observe channel must be **synthesized**: native AX for native targets + **vision
82
+ (screenshot + element detection) for the WebView interior** + optionally a web-layer inspector. An
83
+ XML/AX-only agent silently misses the entire web form. This converges with the external frontier
84
+ finding that SOTA UI automation is hybrid (AX/accessibility for action targets, vision for grounding
85
+ and verification, deterministic probes for assertions).
86
+
87
+ ## Portability — design on simulator, run on real device (Appium-less)
88
+
89
+ Because the Android driver uses the same `adb -s <serial>` interface for an emulator and a USB-connected
90
+ real device, a structure authored and validated on the **simulator runs unchanged on the real device**
91
+ — no per-device manual element extraction. Live-measured: a flow validated on the emulator auto-extracted
92
+ the full native element set on a real-device banking sandbox with zero code change. This portability —
93
+ sim-authored → real-device execution without Appium — is itself a meaningful capability step (a project
94
+ may ship it before any AI-prospective layer).
95
+
96
+ ## Caveats (honest scope)
97
+
98
+ - Element-ranking *accuracy* (intent → correct element) is unmeasured until a project supplies real
99
+ fixtures from its own app; the contract and loop are validated, the ranking quality is per-project.
100
+ - Native-only screens work out of the box; hybrid screens need the vision channel wired (a pluggable
101
+ detector — FH routes to a vision model, does not build one).
102
+ - Korean / IME text entry on Android real devices via `adb input text` is unreliable (measured) — needs
103
+ an IME-broadcast keyboard, not raw input.
104
+
105
+ ## Cross-refs
106
+
107
+ `[[fh-live-surface-acceleration]]` (the capability-axis memory) · `multi_model_sidecar_strategy.md`
108
+ (surface routing) · `deep_research_capability_ladder.md` (the sibling "route, don't build" pattern for
109
+ research). The 2026-06-14 frontier survey backing the hybrid-SOTA convergence lives in the private
110
+ companion store (`paper-signals/frontier_computer_use_ui_automation_2026-06-14.md`).
@@ -0,0 +1,59 @@
1
+ # Measurement-Integrity Checklist — cross-model measurement pre-flight
2
+
3
+ > A cross-model measurement is only trustworthy if its **instrument** is verified first.
4
+ > Measurement integrity is a *precondition*, not a result. Three observed failure modes, each with a
5
+ > concrete countermeasure. Consult this before any FH measurement that compares models (sims, sidecar
6
+ > comparisons, capability-equalizer runs, the-bible model panels, A6-class experiments).
7
+
8
+ This is a **checklist a measurement consults**, not a gate with triggers and not a dispatch surface —
9
+ hence a knowledge doc, the lightest asset that holds it (a harness gets simpler over time). If it ever
10
+ becomes a gate other skills invoke, revisit the weight.
11
+
12
+ ## The three failure modes + countermeasures
13
+
14
+ | # | Failure mode (observed) | Countermeasure |
15
+ |---|---|---|
16
+ | 1 | **Silent model fallback** — passing a model *slug* silently resolved to a weaker model (e.g. an `agy` slug fell back to Flash) instead of the intended one. The run *looks* like the named model but isn't. | **Pin the display name, not the slug** (e.g. `"Gemini 3.1 Pro (High)"`, not a bare slug). Confirm the resolved identity, don't assume the slug binds. |
17
+ | 2 | **Non-deterministic borderline verdicts** — contested/borderline cases flip across runs (observed: haiku 4/4 flip; flagship models flip too — flipping is **not** a tier signal). A single draw is noise, not a measurement. | **reps ≥ 3 on any borderline/contested verdict.** A single run on a contested case is inadmissible. Report the flip pattern (STABLE vs FLIP), not just the modal verdict. |
18
+ | 3 | **Generic self-identity probe** — a probe any model passes ("are you working? → OK") proves nothing about *which* model answered. | **Use a discriminating probe** — one that two different models answer *differently*. A generic-pass probe is invalid. The probe is a **pattern, not a fixed string**: a probe that discriminates Opus 4.8 from Sonnet 4.6 today may both-pass a future model generation, so **re-validate the probe each model generation** (same staleness class `memory-hygiene` exists to catch). |
19
+
20
+ ## Why these are entangled (and why they matter beyond their own scope)
21
+
22
+ Item #2 (reps≥3) is the discipline that **retracted half the evidence** for the
23
+ `[[feedback_correlated_blindspot_union_over_majority]]` finding — one of its two supporting cases
24
+ turned out to be non-deterministic borderline flipping, not a stable correlated blind spot. So this
25
+ checklist is the **prerequisite** for any "correlated error" claim: you cannot call an error correlated
26
+ (and prescribe union-over-majority) until reps≥3 has distinguished a stable correlated error from a
27
+ single-draw artifact. Item #3 (discriminating probe) **embodies** the judge-robustness /
28
+ mechanical-anchor principle — don't trust self-reported identity, prove it discriminatingly
29
+ ([[feedback_judge_robustness_mechanical_anchor]]).
30
+
31
+ ## Done When
32
+
33
+ - The checklist enumerates all three failure modes, each with its countermeasure.
34
+ *Check class: mandatory-pass (binary — three items present, each with a countermeasure).*
35
+ - The probe item specifies a **discriminating** test and rejects generic probes.
36
+ *Check class: judged, pair: a probe that two different models both pass must FAIL this check; a
37
+ discriminating one must distinguish them.*
38
+ - Any FH cross-model measurement records which checklist items it ran.
39
+ *Check class: measured (count of items applied) — closes the predict-verify loop for future audit.*
40
+
41
+ ## Optional hardening (when a measurement feeds a published / paper claim)
42
+
43
+ Escalate item #1 (display-name pin) and item #3 (identity verification) from prose to a **logged
44
+ mechanical assertion**: the measurement harness records the *verified* model identity it observed, not
45
+ the requested slug. Prose discipline is sufficient for internal dogfooding; a published claim earns the
46
+ mechanical log.
47
+
48
+ > **External dogfood (a second, field-layer instance — n=1 external, a signal not a settled frontier):**
49
+ > the sister skill `ponytail` ships a runnable instance of the precondition behind all three modes —
50
+ > a `--selftest` that proves each instrument (`good===true && bad===false`) before any API spend, and
51
+ > two caught instrument contaminations. Detail + pinned citations: `tracks/_audit/session_2026_06_24_ponytail-lazy-senior-dev.md` §2-C (single source).
52
+
53
+ ---
54
+
55
+ **Origin** (2026-06-22 harvest-loop): three failure modes observed across the-bible L2 model panel
56
+ (agy slug→Flash silent fallback; reps=3 non-determinism) and prior multi-model sims (generic-probe
57
+ ambiguity). Sister findings: [[feedback_correlated_blindspot_union_over_majority]] (reps≥3 prerequisite),
58
+ [[feedback_judge_robustness_mechanical_anchor]] (discriminating-probe = mechanical anchor),
59
+ [[reference_agy_model_catalog]] (display-name pin — agy slug fallback documented there).