@chrono-meta/fh-gate 1.4.41 → 1.4.43
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +5 -3
- package/CATALOG.md +6 -0
- package/CLAUDE.md +65 -130
- package/docs/CONTRIBUTING.md +2 -2
- package/knowledge/shared/dialogue/ai_dialogue_playbook.md +137 -0
- package/knowledge/shared/dialogue/claude_code_runtime_flow.md +170 -0
- package/knowledge/shared/dialogue/memory_intent_recall.md +209 -0
- package/knowledge/shared/harness-core/claude_md_gate_details.md +170 -0
- package/knowledge/shared/harness-core/companion_store_pluggable_cross_audit_2026-06-11.md +118 -0
- package/knowledge/shared/harness-core/crucible_mode.md +112 -0
- package/knowledge/shared/harness-core/deep_research_capability_ladder.md +122 -0
- package/knowledge/shared/harness-core/fh_detail_protocols.md +163 -0
- package/knowledge/shared/harness-core/fh_ecosystem_positioning.md +147 -0
- package/knowledge/shared/harness-core/fh_opencode_governance_wrapper.md +163 -0
- package/knowledge/shared/harness-core/fh_synergy_playbook.md +217 -0
- package/knowledge/shared/harness-core/gate_locality_principle.md +57 -0
- package/knowledge/shared/harness-core/goal_quench_anthropic_issue.md +104 -0
- package/knowledge/shared/harness-core/harness_6axis_framework.md +136 -0
- package/knowledge/shared/harness-core/harness_design_decision_lens.md +108 -0
- package/knowledge/shared/harness-core/harness_frontier_diagnosis_2026-06-02.md +102 -0
- package/knowledge/shared/harness-core/hub_compounding_loop.md +109 -0
- package/knowledge/shared/harness-core/hub_maturity_roadmap.md +201 -0
- package/knowledge/shared/harness-core/hybrid_orchestration_architecture_roadmap.md +196 -0
- package/knowledge/shared/harness-core/live_surface_automation_pattern.md +110 -0
- package/knowledge/shared/harness-core/measurement-integrity-checklist.md +59 -0
- package/knowledge/shared/harness-core/meta_harness_engineering_definition.md +116 -0
- package/knowledge/shared/harness-core/multi_model_sidecar_strategy.md +651 -0
- package/knowledge/shared/harness-core/persona_container_schema.md +172 -0
- package/knowledge/shared/harness-core/return_path_gate.md +120 -0
- package/knowledge/shared/harness-core/self_evolution_routine.md +268 -0
- package/knowledge/shared/harness-core/skill_quality_rubric.md +71 -0
- package/knowledge/shared/harness-core/tpa_schema.md +136 -0
- package/package.json +3 -2
|
@@ -0,0 +1,201 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: Hub Maturity 3-Phase Roadmap Frame (template)
|
|
3
|
+
description: Long-term evolution path frame for the hub. Phase I (entering maturity) → Phase II (frontier following) → Phase III (frontier leading) 3-stage model. Fixes the gap, output, completion criteria, and transition conditions for each phase as a quarterly re-diagnosis reference document. The maturity axis is the parent frame referenced by monthly level snapshots and quarterly re-diagnosis. **This file is a template — in actual hub operation, add project-specific information (timing, asset names, identifiers) to write an operating copy**.
|
|
4
|
+
type: reference
|
|
5
|
+
date: 2026-04-28
|
|
6
|
+
tags: [harness, maturity, roadmap, 3-phase, frontier-tracking, evolution, compounding, long-term, hub, strategic, template-frame]
|
|
7
|
+
scope: hub-template
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# Hub Maturity 3-Phase Roadmap (frame)
|
|
11
|
+
|
|
12
|
+
## Why this document exists
|
|
13
|
+
|
|
14
|
+
If the monthly level snapshot captures **"where are we now"**, this document captures **"where are we going and when do we transition"**. If the quarterly re-diagnosis captures **current position vs industry frontier**, this document captures **the evolution of the relationship with the frontier itself** (follower → leader).
|
|
15
|
+
|
|
16
|
+
**Core vision**:
|
|
17
|
+
|
|
18
|
+
> The hub model matures enough to **enter the maturity phase**, **follow** the frontier's advancement direction periodically, and after that, the path where **we become the frontier**.
|
|
19
|
+
|
|
20
|
+
This vision is made explicit as 3 phases so that at each quarterly re-diagnosis, "which phase are we in · what are the conditions for the next phase transition" can be judged.
|
|
21
|
+
|
|
22
|
+
---
|
|
23
|
+
|
|
24
|
+
## 1. 3-Phase Overview
|
|
25
|
+
|
|
26
|
+
| Phase | Scope | Core gap (example) | Representative output (example) | Standard duration (example) |
|
|
27
|
+
|---|---|---|---|---|
|
|
28
|
+
| **I. Entering maturity** | Infrastructure/asset building + routine establishment | Operational gap axes remain · automation not established · 0 external propagation | Weekly audit automation · operations guide · 1-2 external assets · sub-agent judgment · self-diagnosis warning reduction | 3-6 months |
|
|
29
|
+
| **II. Frontier following** | Quarterly re-diagnosis routine + gap auto-detection | Single/dual input sources · no auto gap detection · cadence not confirmed | Quarterly frontier diagnosis 2 times + monthly brief established + external scan cadence confirmed | 5-8 months |
|
|
30
|
+
| **III. Frontier leading** | Self-invented outbound propagation + self-evolving | 0 open-source/presentation record · no self-evolving loop · 0 industry citations | Public refactor · blog/presentation 10+ per year · industry citation case accumulation | Ongoing (no completion) |
|
|
31
|
+
|
|
32
|
+
---
|
|
33
|
+
|
|
34
|
+
## 2. Current Position (fill in operating copy)
|
|
35
|
+
|
|
36
|
+
Fill in the following format in the operating copy:
|
|
37
|
+
|
|
38
|
+
```
|
|
39
|
+
**Phase X · ~N% progress** — infrastructure/asset accumulation status · N remaining completion gaps
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
### Achievement table (example format)
|
|
43
|
+
|
|
44
|
+
| Area | Status | Basis |
|
|
45
|
+
|---|---|---|
|
|
46
|
+
| 6-axis framework + feedback loop | ✅ | (corresponding asset path) |
|
|
47
|
+
| Package structure alignment | ✅ | (realignment session) |
|
|
48
|
+
| ... | ... | ... |
|
|
49
|
+
|
|
50
|
+
### Gaps remaining
|
|
51
|
+
|
|
52
|
+
See §3 Phase I completion criteria.
|
|
53
|
+
|
|
54
|
+
---
|
|
55
|
+
|
|
56
|
+
## 3. Phase I Completion Criteria (5 measurable criteria)
|
|
57
|
+
|
|
58
|
+
Phase II entry gate passed when all 5 are met.
|
|
59
|
+
|
|
60
|
+
| # | Condition | Measurement method (frame) |
|
|
61
|
+
|---|---|---|
|
|
62
|
+
| 1 | **Weekly audit automation established** | audit skill run 3+ times + manual N min → auto N min measured |
|
|
63
|
+
| 2 | **Axis N operations guide established** | orchestration guide draft + mode switch 3+ cases accumulated (if gap axis exists) |
|
|
64
|
+
| 3 | **External propagation N cases** | Distributed to external channel + external response received (generally 2 cases recommended) |
|
|
65
|
+
| 4 | **Sub-agent pilot promotion/deprecation judgment** | 2+ week observation + invocation log-based judgment (`accepted ≥ 60%` / `rejected ≥ 40%` / `invocation count ≥ N`) |
|
|
66
|
+
| 5 | **Self-diagnosis warning reduction** | Quarterly self-diagnosis warnings N items → 1 or fewer |
|
|
67
|
+
|
|
68
|
+
### Completion checklist derived Decision
|
|
69
|
+
|
|
70
|
+
- Phase I completion date = earliest date all 5 criteria are met
|
|
71
|
+
- Completion confirmation event = **Phase II entry meeting** (1 separate session, §4 entry condition check)
|
|
72
|
+
- No Phase II output (e.g., external GitHub scan) before Phase I completion — simplification principle violation
|
|
73
|
+
|
|
74
|
+
---
|
|
75
|
+
|
|
76
|
+
## 4. Phase II (Frontier Following) — Entry conditions, cadence, outputs
|
|
77
|
+
|
|
78
|
+
### 4.1 Entry conditions
|
|
79
|
+
|
|
80
|
+
- Phase I completion 5 criteria **all** met
|
|
81
|
+
- Monthly level snapshot updated 2+ consecutive times
|
|
82
|
+
- §2 current position re-judged as "Phase II · 0%"
|
|
83
|
+
|
|
84
|
+
### 4.2 Cadence 3-option comparison
|
|
85
|
+
|
|
86
|
+
| Option | Cycle | Pros | Cons |
|
|
87
|
+
|---|---|---|---|
|
|
88
|
+
| **(a) Quarterly only** | 3 months | Simple | Gap detection delayed 3 months (slow if frontier moves fast) |
|
|
89
|
+
| **(b) Quarterly + monthly light scan** | 3 months + 4 weeks | Gap detection within 1 month + 10 min addition to existing monthly routine (minimum invasive) | (None — recommended) |
|
|
90
|
+
| (c) Trigger-based | When stagnation detected | Resource efficient | Stagnation detection criteria + auto-alerts + trigger tags all require new infra. **Simplification principle violation risk** |
|
|
91
|
+
|
|
92
|
+
→ **Recommended: (b) quarterly + monthly**. Joining existing monthly routine = minimum invasive. Consistent with simplification principle.
|
|
93
|
+
|
|
94
|
+
### 4.3 Representative outputs (frame)
|
|
95
|
+
|
|
96
|
+
| Output | Cycle | Content |
|
|
97
|
+
|---|---|---|
|
|
98
|
+
| **Quarterly frontier diagnosis #N** | 3 months | 6-axis level change vs previous edition + 3-5 new frontier techniques integrated |
|
|
99
|
+
| **Monthly brief** | 4 weeks | 10-min scan of GitHub trending, Anthropic/OpenAI blog, and similar — 3-5 line summary |
|
|
100
|
+
| **External GitHub scan actual operation** | Monthly | Monthly 10-min scan of public repos. Sprint Contract in 5 lines |
|
|
101
|
+
| **Auto gap detection** | Scanner feature addition | Flag "frontier diagnosis not updated > 90 days" |
|
|
102
|
+
|
|
103
|
+
### 4.4 Completion conditions
|
|
104
|
+
|
|
105
|
+
- Quarterly re-diagnosis 2 times + monthly brief 6+ consecutive productions
|
|
106
|
+
- 1+ time **own methodology back-referenced from frontier** found
|
|
107
|
+
- 2+ axes in monthly level snapshot sustained as "leading" judgment
|
|
108
|
+
|
|
109
|
+
---
|
|
110
|
+
|
|
111
|
+
## 5. Phase III (Frontier Leading) — Entry conditions, 6 indicators, outputs
|
|
112
|
+
|
|
113
|
+
### 5.1 Entry conditions
|
|
114
|
+
|
|
115
|
+
- Phase II completion 3 conditions all met
|
|
116
|
+
- 3+ of N self-invented assets **observed as original concepts** for similar industry concepts
|
|
117
|
+
- Public seed repository external contribution record started (at least one of fork/issue/star)
|
|
118
|
+
|
|
119
|
+
### 5.2 6 Leading indicators
|
|
120
|
+
|
|
121
|
+
| # | Indicator | Measurement |
|
|
122
|
+
|---|---|---|
|
|
123
|
+
| 1 | Public seed repository record | star/fork/issue count |
|
|
124
|
+
| 2 | Blog/presentation | 2-3 per quarter · 10+ per year |
|
|
125
|
+
| 3 | Industry citation cases | External articles/seminars citing this hub/methodology |
|
|
126
|
+
| 4 | External organization adoption | Other companies/departments |
|
|
127
|
+
| 5 | Self-evolving loop demonstration | Skill generates skills |
|
|
128
|
+
| 6 | Self-invented industry original recognition | Adopted as industry term/frame |
|
|
129
|
+
|
|
130
|
+
### 5.3 Representative outputs
|
|
131
|
+
|
|
132
|
+
- **Public seed repository refactor** — from personal seed to collaboratable template. Team customization + common protocol separation
|
|
133
|
+
- **10+ blog/presentations per year** — externalize 1 self-invented asset per quarter
|
|
134
|
+
- **Self-evolving loop MVP** — audit skill proposes own skill generation + user approval → auto-generate → usage observation → deprecate/improve
|
|
135
|
+
|
|
136
|
+
### 5.4 Phase III has no "completion"
|
|
137
|
+
|
|
138
|
+
Phase III is an ongoing state. Instead of completion criteria, **3 maintenance conditions**:
|
|
139
|
+
- 3+ of 6 indicators continuously rising
|
|
140
|
+
- "Frontier level maintained" judgment in quarterly re-diagnosis (no regression)
|
|
141
|
+
- Self-diagnosis 8-item checklist failure signal 0 maintained
|
|
142
|
+
|
|
143
|
+
---
|
|
144
|
+
|
|
145
|
+
## 6. Common principles for transitions
|
|
146
|
+
|
|
147
|
+
### 6.1 Consistent simplification principle
|
|
148
|
+
|
|
149
|
+
Phase transitions are **methodology/abstraction level rises**, not **increases in file/skill/rule count**. At each phase transition, do the following first:
|
|
150
|
+
|
|
151
|
+
- [ ] Self-diagnosis 8-item checklist check — confirm 0 failure signals
|
|
152
|
+
- [ ] Files unreferenced 6+ months → move to `archive/` or consolidate
|
|
153
|
+
- [ ] Clean up previous phase temporary outputs made unnecessary by new phase entry
|
|
154
|
+
- [ ] Confirm CLAUDE.md and CATALOG within 200 lines
|
|
155
|
+
- [ ] Re-confirm optimization principle: field harness → simpler over time; meta-harness → complexity earns its scope (purge orphaned/redundant/decorative units)
|
|
156
|
+
|
|
157
|
+
### 6.2 Transition deferred on simplification failure
|
|
158
|
+
|
|
159
|
+
Above checklist not passed → **phase transition deferred**. Perform previous phase remaining work + simplification work for 1 additional month then re-check.
|
|
160
|
+
|
|
161
|
+
### 6.3 Phase regression possible
|
|
162
|
+
|
|
163
|
+
**Regression** from Phase II to Phase I also allowed. E.g., frontier following routine missed 2 consecutive times → "manual re-establishment" re-perform part of Phase I. Do not force linear progression.
|
|
164
|
+
|
|
165
|
+
---
|
|
166
|
+
|
|
167
|
+
## 7. This roadmap update cycle
|
|
168
|
+
|
|
169
|
+
| Trigger | Update content |
|
|
170
|
+
|---|---|
|
|
171
|
+
| Quarterly re-diagnosis | §2 current position re-judgment + §3·§4·§5 criteria change check |
|
|
172
|
+
| Phase transition event | Record corresponding phase completion + initialize next phase progress |
|
|
173
|
+
| Self-invented asset recognized as industry original | §5.1 entry condition counter update |
|
|
174
|
+
| Simplification principle violation detected | §6 transition deferral activation record |
|
|
175
|
+
|
|
176
|
+
---
|
|
177
|
+
|
|
178
|
+
## 8. Operating copy writing guide (for template users)
|
|
179
|
+
|
|
180
|
+
When moving this frame to an operating copy:
|
|
181
|
+
|
|
182
|
+
1. **Fill §2 current position** — add asset list and basis material paths
|
|
183
|
+
2. **Concretize §3 5 criteria** — add `current` column + `target date` column (e.g., `2026-05-17`)
|
|
184
|
+
3. **Recommend §4.2 cadence (b)** — if adopting other option, need to prove §6.1 simplification gate passed
|
|
185
|
+
4. **Add current counter to §5 6 indicators** — starting from all 0 is natural
|
|
186
|
+
5. **Update §2 at each quarterly re-diagnosis** + record cumulative changes in §7 trigger table
|
|
187
|
+
|
|
188
|
+
---
|
|
189
|
+
|
|
190
|
+
## 9. Related assets
|
|
191
|
+
|
|
192
|
+
- 6-axis framework — `knowledge/shared/harness-core/harness_6axis_framework.md` (frame premise)
|
|
193
|
+
- Monthly level snapshot — `knowledge/shared/harness-core/harness_level_snapshot_*.md` (current position)
|
|
194
|
+
- Quarterly re-diagnosis — `knowledge/shared/harness-core/harness_frontier_diagnosis_*.md` (following basis)
|
|
195
|
+
- Feedback automation — `knowledge/shared/harness-core/hub_compounding_loop.md` (Phase transition gate input)
|
|
196
|
+
|
|
197
|
+
---
|
|
198
|
+
|
|
199
|
+
## 10. One-line conclusion
|
|
200
|
+
|
|
201
|
+
**Phase I completion → Phase II frontier following (b)cadence → Phase III leading. Simplification principle is the common gate for each transition**.
|
|
@@ -0,0 +1,196 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: hybrid-orchestration-architecture-roadmap
|
|
3
|
+
description: Proposed (not-yet-implemented) architecture roadmap for fh as an intelligent hybrid orchestration engine — Claude Code as main driver, other CLIs/APIs as runtime-discovered sidecars, with Zero-Config standalone fallback for plugin-only users. Design intent + reconciliation against existing FH assets.
|
|
4
|
+
type: roadmap
|
|
5
|
+
date: 2026-06-09
|
|
6
|
+
status: proposed (design intent — NOT implemented; see §Status & Reconciliation)
|
|
7
|
+
tags: [hybrid-orchestration, sidecar, zero-config, jit-probing, install-wizard, roadmap, proposed]
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# Hybrid Orchestration Architecture — Roadmap (Proposed)
|
|
11
|
+
|
|
12
|
+
> **Status: PROPOSED design intent, not current behavior.** This document captures a
|
|
13
|
+
> forward-looking architecture for fh. Several components it describes — a `config.json`
|
|
14
|
+
> engine-topology file, runtime JIT engine probing, an install-wizard that auto-builds an
|
|
15
|
+
> engine map — **do not exist in FH today**. Read §Status & Reconciliation first to see
|
|
16
|
+
> what is already shipped vs what is aspirational. Nothing here should be cited as a
|
|
17
|
+
> current feature.
|
|
18
|
+
>
|
|
19
|
+
> **Source**: operator design doc (2026-06-09), reconciled to FH conventions on ingest —
|
|
20
|
+
> pinned model versions → `{model-name}` placeholders (per `multi_model_sidecar_strategy.md`
|
|
21
|
+
> §Generalization #4); illustrative pseudo-code marked as such (FH's real layer is markdown
|
|
22
|
+
> methodology + Claude-native automation, not a Python runtime engine).
|
|
23
|
+
|
|
24
|
+
## 1. Overview
|
|
25
|
+
|
|
26
|
+
fh aims to be a **hybrid orchestration** harness: **Claude Code as the main driver (core
|
|
27
|
+
controller)**, with third-party CLIs / APIs mapped in as **sidecars** according to whatever
|
|
28
|
+
the user's local environment makes available. The goal is to escape single-model dependency
|
|
29
|
+
and **progressively enhance** output by discovering the user's own resources (subscription
|
|
30
|
+
CLIs, API keys) at runtime — while still working in a fully **Zero-Config** state for users
|
|
31
|
+
who copy only a single plugin/skill rather than installing the whole framework.
|
|
32
|
+
|
|
33
|
+
This is the design north star. The mechanism for *delegating to sidecars* is already
|
|
34
|
+
validated and shipped (see `multi_model_sidecar_strategy.md`); the *automatic
|
|
35
|
+
discovery/topology* layer below is the proposed addition.
|
|
36
|
+
|
|
37
|
+
## 2. Model roles (provider-agnostic)
|
|
38
|
+
|
|
39
|
+
FH convention forbids pinning version numbers or benchmark figures into durable assets
|
|
40
|
+
(they age and become phantom claims), so roles are stated by function, not by pinned model:
|
|
41
|
+
|
|
42
|
+
| Role | Engine slot | Function | Maps to FH skill family |
|
|
43
|
+
|---|---|---|---|
|
|
44
|
+
| **Core Driver** | `{primary-model}` (strongest available; CC host) | Full code edit, architecture, final synthesis | steel-quench, refactor, design consensus |
|
|
45
|
+
| **Context/Analysis Sidecar** | `{large-context-sidecar}` | Bulk log/API-spec parsing → distilled clues | harvest-loop bulk scan |
|
|
46
|
+
| **Test/Util Sidecar** | `{fast-util-sidecar}` | Boilerplate, unit tests, repetitive drops | speed-run / scaffolding |
|
|
47
|
+
|
|
48
|
+
> Capability/benchmark comparisons between providers change every release cycle — consult a
|
|
49
|
+
> live source (`/frontier-digest`, or the `claude-api` skill for Claude specifics) at decision
|
|
50
|
+
> time rather than trusting any number frozen into this doc. No SWE-bench figures are recorded
|
|
51
|
+
> here by design.
|
|
52
|
+
|
|
53
|
+
## 3. Main driver + sidecar dual channel
|
|
54
|
+
|
|
55
|
+
```
|
|
56
|
+
┌─────────────────────────────┐
|
|
57
|
+
│ user terminal (fh) │
|
|
58
|
+
└──────────────┬──────────────┘
|
|
59
|
+
│ [intelligent routing]
|
|
60
|
+
┌───────────────────┼───────────────────┐
|
|
61
|
+
▼ (bulk / low-cost) ▼ (final reasoning) ▼ (fast unit-test / repeat)
|
|
62
|
+
┌─────────────────┐ ┌──────────────────┐ ┌─────────────────┐
|
|
63
|
+
│ context sidecar │ │ Claude Code (main)│ │ util sidecar │
|
|
64
|
+
└────────┬────────┘ └────────┬─────────┘ └────────┬────────┘
|
|
65
|
+
└────────────────────┼─────────────────────┘
|
|
66
|
+
▼
|
|
67
|
+
[final quality verification + apply]
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
- **Claude Code (main)** keeps full project context/architecture and is the only writer of
|
|
71
|
+
final code. This matches FH's existing rule: *"Host is always single"*
|
|
72
|
+
(`multi_model_sidecar_strategy.md` §Mechanism).
|
|
73
|
+
- **Sidecars** are stateless one-shot `Bash`-invoked processes whose stdout is folded back in
|
|
74
|
+
by the calling skill — **already the shipped mechanism**, not new.
|
|
75
|
+
|
|
76
|
+
## 4. 3-Tier routing protocol (proposed ordering)
|
|
77
|
+
|
|
78
|
+
When a skill needs a sidecar engine, probe in priority order:
|
|
79
|
+
|
|
80
|
+
```
|
|
81
|
+
[sidecar request]
|
|
82
|
+
├── Tier 1: subscription CLI (zero marginal cost)
|
|
83
|
+
│ discover logged-in binaries: `claude -p`, `aider`, `gemini`, `codex`, `gh copilot`
|
|
84
|
+
├── Tier 2: native API call (pay-per-use)
|
|
85
|
+
│ env keys: GEMINI_API_KEY, OPENAI_API_KEY, ANTHROPIC_API_KEY
|
|
86
|
+
└── Tier 3: self-contained sub-agent split (fallback)
|
|
87
|
+
no external sidecar → Claude Code spawns isolated sub-agent / prompt-chunking
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
> **Relation to the shipped fallback chain**: `multi_model_sidecar_strategy.md`
|
|
91
|
+
> §Implementation-Patterns already defines a 3-tier *fallback* chain (Copilot CLI → corporate
|
|
92
|
+
> endpoint → direct Gemini/Codex). The ordering here is framed by **cost/access tier** rather
|
|
93
|
+
> than network reachability. These are two lenses on the same fan-out; if implemented, they
|
|
94
|
+
> should be unified into one routing table, not maintained as two competing lists. Tier 3
|
|
95
|
+
> ("Claude internal sub-agent fallback") is the genuinely new contribution — it guarantees the
|
|
96
|
+
> chain never hard-fails even with zero external resources.
|
|
97
|
+
|
|
98
|
+
## 5. Static inheritance + dynamic JIT fallback (proposed)
|
|
99
|
+
|
|
100
|
+
```
|
|
101
|
+
execution request (skill / agent)
|
|
102
|
+
│
|
|
103
|
+
▼
|
|
104
|
+
[config.json present?] ← PROPOSED file; does NOT exist in FH today
|
|
105
|
+
├── Yes → static-inheritance mode → load priority map → run immediately
|
|
106
|
+
└── No → Zero-Config standalone mode → runtime JIT probing:
|
|
107
|
+
1. local subscription CLI scan (Tier 1)
|
|
108
|
+
2. system API-key env scan (Tier 2)
|
|
109
|
+
3. main-driver sub-agent fallback (Tier 3)
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
**5.1 Static config inheritance** — if an install-wizard run had persisted an engine map,
|
|
113
|
+
skills would adopt it as default and skip per-call scan overhead. *(Proposed: install-wizard
|
|
114
|
+
does not generate such a file today.)*
|
|
115
|
+
|
|
116
|
+
**5.2 Runtime JIT probing** — if no config exists (or a tool was added after install), the
|
|
117
|
+
harness probes the live environment on-the-fly via the Tier 1→2→3 protocol and binds the best
|
|
118
|
+
path. All probes failing → safe descent to Tier 3 (main model handles it).
|
|
119
|
+
|
|
120
|
+
**5.3 Zero-Config standalone self-reliance** — the key requirement for **Mode C** (plugin/skill
|
|
121
|
+
copied without the full framework, see `.claude/rules/modes_and_value.md`). With no
|
|
122
|
+
`config.json`, the module must **not error** — it switches to Zero-Config standalone mode and
|
|
123
|
+
JIT-probes for whatever local tools exist, then proceeds quietly. Falls back to Tier 3 if none.
|
|
124
|
+
|
|
125
|
+
```python
|
|
126
|
+
# ILLUSTRATIVE resolution logic — NOT shipped code.
|
|
127
|
+
# FH's actual layer is markdown methodology + Claude-native automation (skills/rules/hooks),
|
|
128
|
+
# not a Python engine. This sketch only shows the intended decision order.
|
|
129
|
+
def resolve_sidecar_engine(preferred="{util-sidecar}"):
|
|
130
|
+
cfg = load_static_config_safe() # Step 1: static config (if any)
|
|
131
|
+
if cfg and cfg.get("mapped_engine"):
|
|
132
|
+
return cfg["mapped_engine"]
|
|
133
|
+
if shutil.which("aider") or shutil.which("{cli}"): # Tier 1: subscription CLI
|
|
134
|
+
return "CLI_WRAPPER_MODE"
|
|
135
|
+
if os.environ.get("{PROVIDER}_API_KEY"): # Tier 2: API key
|
|
136
|
+
return "NATIVE_API_MODE"
|
|
137
|
+
return "CLAUDE_SUBAGENT_FALLBACK" # Tier 3: peaceful fallback
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
## 6. Intelligent install-wizard (proposed extension)
|
|
141
|
+
|
|
142
|
+
Today `install-wizard` sets up the periodic-audit notification structure (zshrc hook +
|
|
143
|
+
sentinels + session-start mtime detection). It does **not** build an engine topology. The
|
|
144
|
+
proposed extension would add, at install time:
|
|
145
|
+
|
|
146
|
+
- **Binary probing** — scan `$PATH` for available CLIs/dependencies.
|
|
147
|
+
- **Profile mapping** — write the discovered hybrid config (the proposed `config.json`).
|
|
148
|
+
- **Sanity check** — a light sidecar call to confirm the pipeline actually works.
|
|
149
|
+
|
|
150
|
+
Onboarding UX phrasing (proposed):
|
|
151
|
+
- Minimal env → *"Configured a Claude-Code-only harness. All skills run safely via context
|
|
152
|
+
splitting, no API key needed."*
|
|
153
|
+
- Expansion nudge → *"To enable cheaper large-analysis skills, register a sidecar API key later;
|
|
154
|
+
fh will auto-detect it on the next run and apply sidecar orchestration."*
|
|
155
|
+
|
|
156
|
+
## 7. sim-conductor & core-skill improvement directions (proposed)
|
|
157
|
+
|
|
158
|
+
- **7.1 `claude -p` non-interactive pipe runtime** — exploit CC's non-interactive flag; pipe
|
|
159
|
+
stdout/stdin between agents; reuse session cache via a harness wrapper.
|
|
160
|
+
- **7.2 Hybrid context bridge** — normalize heterogeneous sidecar output (API JSON vs CLI
|
|
161
|
+
streamed text); embed an extractor (code-block / JSON-structure parser) in a bridge layer so
|
|
162
|
+
callers always get structured responses regardless of invocation path.
|
|
163
|
+
- **7.3 Built-in sub-agent injection** — for minimal envs with no third-party CLI/API: place
|
|
164
|
+
worker prompt specs (e.g. `steel-quench-worker.md`, `harvest-loop-worker.md`) into
|
|
165
|
+
`.claude/agents/`, so the harness spawns Claude's own isolated-context sub-agent pool for
|
|
166
|
+
parallel work. *(Note: per `operations.md`, personal agents in shared repos should be kept
|
|
167
|
+
local via `.git/info/exclude` — an installer-placed agent must respect that boundary.)*
|
|
168
|
+
|
|
169
|
+
---
|
|
170
|
+
|
|
171
|
+
## Status & Reconciliation (read this before citing anything)
|
|
172
|
+
|
|
173
|
+
| Component | In FH today? | Where / note |
|
|
174
|
+
|---|---|---|
|
|
175
|
+
| Sidecar delegation via `Bash` (stateless, host-single) | ✅ Shipped + validated | `multi_model_sidecar_strategy.md` (Experiment 1·2) |
|
|
176
|
+
| Cross-provider perspective-diversity rationale | ✅ Shipped | same doc + `steel-quench` Wave 5 |
|
|
177
|
+
| 3-tier *fallback* chain (network-reachability lens) | ✅ Shipped | same doc §Implementation-Patterns |
|
|
178
|
+
| 3-tier *routing* by cost/access (Tier1 CLI→Tier2 API→Tier3 subagent) | 🟡 Partial | reorders the shipped chain; Tier-3 subagent fallback is new |
|
|
179
|
+
| `config.json` engine-topology file | ❌ Proposed | does not exist |
|
|
180
|
+
| Runtime JIT engine probing | ❌ Proposed | does not exist |
|
|
181
|
+
| Zero-Config standalone auto-heal (Mode C) | ❌ Proposed | concept aligns with Mode C in `modes_and_value.md` |
|
|
182
|
+
| install-wizard builds engine topology / sanity-check | ❌ Proposed | install-wizard today = zshrc hook + sentinels only |
|
|
183
|
+
| Hybrid context bridge / built-in sub-agent injection | ❌ Proposed | sim-conductor roadmap item |
|
|
184
|
+
|
|
185
|
+
**Conflicts resolved on ingest** (per FH conventions): pinned model versions + SWE-bench
|
|
186
|
+
figures dropped → `{model-name}` placeholders (avoids phantom/stale claims); Python
|
|
187
|
+
`engine_resolver.py` marked ILLUSTRATIVE (FH has no Python runtime layer); every non-shipped
|
|
188
|
+
component tagged ❌ Proposed above so this roadmap can never be mistaken for current behavior.
|
|
189
|
+
|
|
190
|
+
## References
|
|
191
|
+
|
|
192
|
+
- `knowledge/shared/harness-core/multi_model_sidecar_strategy.md` — shipped sidecar mechanism + fallback chain (this roadmap extends, does not replace, it)
|
|
193
|
+
- `plugins/fh-meta/skills/install-wizard/SKILL.md` — current install behavior (the topology extension would build on this)
|
|
194
|
+
- `.claude/rules/modes_and_value.md` — Mode C (plugin/skill-only) that Zero-Config self-reliance targets
|
|
195
|
+
- `.claude/rules/operations.md` — sub-agent boundary rules the built-in-injection item must respect
|
|
196
|
+
- `plugins/fh-meta/skills/frontier-digest/SKILL.md` — live source for current model capability/benchmark comparison (do not freeze numbers here)
|
|
@@ -0,0 +1,110 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: live-surface-automation-pattern
|
|
3
|
+
description: The capability pattern FH routes to when a mapping project needs an agent to drive a live UI surface (web/mobile) — observe-act-verify over the running app, not code/data fetch. FH routes drivers (no-reinvention); the value is the cross-platform observe-act-verify contract, the Appium-less principle, and the hybrid-WebView vision-synthesis rule. Generalizable across mapping projects; not tied to any one project.
|
|
4
|
+
date: 2026-06-14
|
|
5
|
+
tags: [live-surface, ui-automation, observe-act-verify, appium-less, hybrid-webview, no-reinvention, mapping-acceleration]
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# Live-Surface Automation Pattern
|
|
9
|
+
|
|
10
|
+
When a mapping project's work lives on a **live UI surface** — a running mobile app or web page whose
|
|
11
|
+
state cannot be reached by reading code or fetching data — FH's posture is the same as for any
|
|
12
|
+
capability it does not own: **detect the need and route to the best driver present, do not build a UI
|
|
13
|
+
engine** (no-reinvention). The value FH adds is the *contract* below, not a new automation tool.
|
|
14
|
+
|
|
15
|
+
This is the acceleration axis behind the "③ map-project acceleration" door's live-surface capability
|
|
16
|
+
(`[[fh-live-surface-acceleration]]`). It was first validated on web (Playwright MCP); the 2026-06-14
|
|
17
|
+
validation extended it to mobile (Android + iOS) and surfaced the two principles that make it work.
|
|
18
|
+
A 2026-06-19 two-surface session then measured the **logged-in-web** (claude-in-chrome) and
|
|
19
|
+
**native-desktop** (computer-use) lanes against the same six-primitive contract — splitting the web
|
|
20
|
+
row into isolated vs logged-in and adding a native-desktop row (see Driver routing).
|
|
21
|
+
|
|
22
|
+
## The observe-act-verify contract (driver-agnostic)
|
|
23
|
+
|
|
24
|
+
Every driver fills the same six primitives; the runner above them is identical regardless of backend:
|
|
25
|
+
|
|
26
|
+
| primitive | role | verdict impact |
|
|
27
|
+
|---|---|---|
|
|
28
|
+
| launch | bring the app/page up | failure → BLOCKED |
|
|
29
|
+
| screenshot | vision-channel capture | evidence |
|
|
30
|
+
| observe | structured element tree (XML/AX/DOM) → normalized elements | candidate ranking input |
|
|
31
|
+
| tap | element/coordinate action | act |
|
|
32
|
+
| input_text | text entry | act |
|
|
33
|
+
| screen_signature | pre/post comparison | re-observe |
|
|
34
|
+
|
|
35
|
+
A TC/step's intent text is ranked against observed elements (text · description · role · clickable ·
|
|
36
|
+
bounds), the top candidate is acted on, and the post-action observation is checked **mechanically**
|
|
37
|
+
(target text present/absent + candidate count + adapter exit) — no judge LLM decides the verdict.
|
|
38
|
+
|
|
39
|
+
## Driver routing (no-reinvention)
|
|
40
|
+
|
|
41
|
+
| surface | driver | note |
|
|
42
|
+
|---|---|---|
|
|
43
|
+
| web — isolated (no login) | Playwright MCP (Claude-usable) / Stagehand | isolated session, no operator cookies; generic-coding-agent usable — covers users without a proprietary browser-agent app |
|
|
44
|
+
| web — logged-in (operator's authenticated session) | authenticated-session browser driver — currently claude-in-chrome (`/chrome`) | reaches SSO-behind SPAs the isolated lane can't; `find` → DOM element ref. Current-instance caveats in footnote † |
|
|
45
|
+
| native desktop screen | screen-control driver — currently computer-use (screenshot + coordinate, vision channel) | reaches non-browser native app windows the browser lanes can't; `observe` is vision-coordinate, not an AX/DOM tree. ⚠️ local-session-bound (remote-control), no mobile; `input_text` unverified (5/6 — see note) |
|
|
46
|
+
| Android (emulator **or** real device) | `adb` + `uiautomator dump` + `input tap` | same code path for both — see Portability below |
|
|
47
|
+
| iOS simulator | `idb ui describe-all` + `idb ui tap` (idb-companion + fb-idb) | native AX tree |
|
|
48
|
+
| hybrid WebView region | **vision channel** (screenshot + element detection) | native AX is blind here — see Hybrid rule |
|
|
49
|
+
|
|
50
|
+
FH routes; it does not reimplement these. Drivers are pluggable; the contract is fixed.
|
|
51
|
+
|
|
52
|
+
**Observe channel + measurement (2026-06-19).** computer-use fills `observe` via the **vision channel**
|
|
53
|
+
(screenshot + coordinate) — the same channel the hybrid-WebView rule already mandates (below), now for
|
|
54
|
+
native desktop; the browser and mobile drivers fill it with a structured tree (DOM ref / AX). Measured
|
|
55
|
+
that session: claude-in-chrome closed **all six** primitives on both a no-login and an operator-logged-in
|
|
56
|
+
surface; computer-use closed **five** (launch / screenshot / observe(vision) / tap / screen_signature) —
|
|
57
|
+
`input_text` is keyboard-capable but was not exercised. The lanes cover **different surfaces, not
|
|
58
|
+
redundant ones** — calculator success on computer-use does not imply chrome, and only chrome reaches the
|
|
59
|
+
logged-in session.
|
|
60
|
+
|
|
61
|
+
† **Current-instance caveats (point-in-time, not routing inputs).** claude-in-chrome and computer-use are
|
|
62
|
+
built-in MCPs — verify presence via the `/mcp` UI, not `claude mcp list`. claude-in-chrome needs a direct
|
|
63
|
+
first-party plan (unavailable on Bedrock/Vertex) and a per-site permission grant. Tool surface as measured
|
|
64
|
+
2026-06-19: claude-in-chrome ~22, computer-use ~24 (counts drift across releases — not a routing input).
|
|
65
|
+
|
|
66
|
+
## ★ Principle 1 — Appium-less is the enabler
|
|
67
|
+
|
|
68
|
+
Going **through Appium is the failure mode**, not the solution. Appium's WebView-context switching is
|
|
69
|
+
flaky on hybrid apps; in one field case it burned a large automation budget without completing a single
|
|
70
|
+
hybrid flow (repeated run→fail→rerun). The **direct path** (`uiautomator dump` / `idb describe-all` +
|
|
71
|
+
coordinate tap + vision) bypasses that failure mode entirely. Appium is therefore *one regression-stage
|
|
72
|
+
adapter*, not the exploration/execution substrate.
|
|
73
|
+
|
|
74
|
+
## ★ Principle 2 — Hybrid WebView is opaque to native accessibility → synthesize vision
|
|
75
|
+
|
|
76
|
+
Live-measured (2026-06-14, iOS hybrid sandbox): a WKWebView's **web content does not appear in the
|
|
77
|
+
native accessibility tree** — only native chrome and bridge-triggered native overlays do. A tap on a
|
|
78
|
+
web button fired the web→native bridge and the resulting native picker *did* appear in the AX tree
|
|
79
|
+
(4→10 elements), confirming the boundary precisely.
|
|
80
|
+
|
|
81
|
+
→ For hybrid apps, the observe channel must be **synthesized**: native AX for native targets + **vision
|
|
82
|
+
(screenshot + element detection) for the WebView interior** + optionally a web-layer inspector. An
|
|
83
|
+
XML/AX-only agent silently misses the entire web form. This converges with the external frontier
|
|
84
|
+
finding that SOTA UI automation is hybrid (AX/accessibility for action targets, vision for grounding
|
|
85
|
+
and verification, deterministic probes for assertions).
|
|
86
|
+
|
|
87
|
+
## Portability — design on simulator, run on real device (Appium-less)
|
|
88
|
+
|
|
89
|
+
Because the Android driver uses the same `adb -s <serial>` interface for an emulator and a USB-connected
|
|
90
|
+
real device, a structure authored and validated on the **simulator runs unchanged on the real device**
|
|
91
|
+
— no per-device manual element extraction. Live-measured: a flow validated on the emulator auto-extracted
|
|
92
|
+
the full native element set on a real-device banking sandbox with zero code change. This portability —
|
|
93
|
+
sim-authored → real-device execution without Appium — is itself a meaningful capability step (a project
|
|
94
|
+
may ship it before any AI-prospective layer).
|
|
95
|
+
|
|
96
|
+
## Caveats (honest scope)
|
|
97
|
+
|
|
98
|
+
- Element-ranking *accuracy* (intent → correct element) is unmeasured until a project supplies real
|
|
99
|
+
fixtures from its own app; the contract and loop are validated, the ranking quality is per-project.
|
|
100
|
+
- Native-only screens work out of the box; hybrid screens need the vision channel wired (a pluggable
|
|
101
|
+
detector — FH routes to a vision model, does not build one).
|
|
102
|
+
- Korean / IME text entry on Android real devices via `adb input text` is unreliable (measured) — needs
|
|
103
|
+
an IME-broadcast keyboard, not raw input.
|
|
104
|
+
|
|
105
|
+
## Cross-refs
|
|
106
|
+
|
|
107
|
+
`[[fh-live-surface-acceleration]]` (the capability-axis memory) · `multi_model_sidecar_strategy.md`
|
|
108
|
+
(surface routing) · `deep_research_capability_ladder.md` (the sibling "route, don't build" pattern for
|
|
109
|
+
research). The 2026-06-14 frontier survey backing the hybrid-SOTA convergence lives in the private
|
|
110
|
+
companion store (`paper-signals/frontier_computer_use_ui_automation_2026-06-14.md`).
|
|
@@ -0,0 +1,59 @@
|
|
|
1
|
+
# Measurement-Integrity Checklist — cross-model measurement pre-flight
|
|
2
|
+
|
|
3
|
+
> A cross-model measurement is only trustworthy if its **instrument** is verified first.
|
|
4
|
+
> Measurement integrity is a *precondition*, not a result. Three observed failure modes, each with a
|
|
5
|
+
> concrete countermeasure. Consult this before any FH measurement that compares models (sims, sidecar
|
|
6
|
+
> comparisons, capability-equalizer runs, the-bible model panels, A6-class experiments).
|
|
7
|
+
|
|
8
|
+
This is a **checklist a measurement consults**, not a gate with triggers and not a dispatch surface —
|
|
9
|
+
hence a knowledge doc, the lightest asset that holds it (a harness gets simpler over time). If it ever
|
|
10
|
+
becomes a gate other skills invoke, revisit the weight.
|
|
11
|
+
|
|
12
|
+
## The three failure modes + countermeasures
|
|
13
|
+
|
|
14
|
+
| # | Failure mode (observed) | Countermeasure |
|
|
15
|
+
|---|---|---|
|
|
16
|
+
| 1 | **Silent model fallback** — passing a model *slug* silently resolved to a weaker model (e.g. an `agy` slug fell back to Flash) instead of the intended one. The run *looks* like the named model but isn't. | **Pin the display name, not the slug** (e.g. `"Gemini 3.1 Pro (High)"`, not a bare slug). Confirm the resolved identity, don't assume the slug binds. |
|
|
17
|
+
| 2 | **Non-deterministic borderline verdicts** — contested/borderline cases flip across runs (observed: haiku 4/4 flip; flagship models flip too — flipping is **not** a tier signal). A single draw is noise, not a measurement. | **reps ≥ 3 on any borderline/contested verdict.** A single run on a contested case is inadmissible. Report the flip pattern (STABLE vs FLIP), not just the modal verdict. |
|
|
18
|
+
| 3 | **Generic self-identity probe** — a probe any model passes ("are you working? → OK") proves nothing about *which* model answered. | **Use a discriminating probe** — one that two different models answer *differently*. A generic-pass probe is invalid. The probe is a **pattern, not a fixed string**: a probe that discriminates Opus 4.8 from Sonnet 4.6 today may both-pass a future model generation, so **re-validate the probe each model generation** (same staleness class `memory-hygiene` exists to catch). |
|
|
19
|
+
|
|
20
|
+
## Why these are entangled (and why they matter beyond their own scope)
|
|
21
|
+
|
|
22
|
+
Item #2 (reps≥3) is the discipline that **retracted half the evidence** for the
|
|
23
|
+
`[[feedback_correlated_blindspot_union_over_majority]]` finding — one of its two supporting cases
|
|
24
|
+
turned out to be non-deterministic borderline flipping, not a stable correlated blind spot. So this
|
|
25
|
+
checklist is the **prerequisite** for any "correlated error" claim: you cannot call an error correlated
|
|
26
|
+
(and prescribe union-over-majority) until reps≥3 has distinguished a stable correlated error from a
|
|
27
|
+
single-draw artifact. Item #3 (discriminating probe) **embodies** the judge-robustness /
|
|
28
|
+
mechanical-anchor principle — don't trust self-reported identity, prove it discriminatingly
|
|
29
|
+
([[feedback_judge_robustness_mechanical_anchor]]).
|
|
30
|
+
|
|
31
|
+
## Done When
|
|
32
|
+
|
|
33
|
+
- The checklist enumerates all three failure modes, each with its countermeasure.
|
|
34
|
+
*Check class: mandatory-pass (binary — three items present, each with a countermeasure).*
|
|
35
|
+
- The probe item specifies a **discriminating** test and rejects generic probes.
|
|
36
|
+
*Check class: judged, pair: a probe that two different models both pass must FAIL this check; a
|
|
37
|
+
discriminating one must distinguish them.*
|
|
38
|
+
- Any FH cross-model measurement records which checklist items it ran.
|
|
39
|
+
*Check class: measured (count of items applied) — closes the predict-verify loop for future audit.*
|
|
40
|
+
|
|
41
|
+
## Optional hardening (when a measurement feeds a published / paper claim)
|
|
42
|
+
|
|
43
|
+
Escalate item #1 (display-name pin) and item #3 (identity verification) from prose to a **logged
|
|
44
|
+
mechanical assertion**: the measurement harness records the *verified* model identity it observed, not
|
|
45
|
+
the requested slug. Prose discipline is sufficient for internal dogfooding; a published claim earns the
|
|
46
|
+
mechanical log.
|
|
47
|
+
|
|
48
|
+
> **External dogfood (a second, field-layer instance — n=1 external, a signal not a settled frontier):**
|
|
49
|
+
> the sister skill `ponytail` ships a runnable instance of the precondition behind all three modes —
|
|
50
|
+
> a `--selftest` that proves each instrument (`good===true && bad===false`) before any API spend, and
|
|
51
|
+
> two caught instrument contaminations. Detail + pinned citations: `tracks/_audit/session_2026_06_24_ponytail-lazy-senior-dev.md` §2-C (single source).
|
|
52
|
+
|
|
53
|
+
---
|
|
54
|
+
|
|
55
|
+
**Origin** (2026-06-22 harvest-loop): three failure modes observed across the-bible L2 model panel
|
|
56
|
+
(agy slug→Flash silent fallback; reps=3 non-determinism) and prior multi-model sims (generic-probe
|
|
57
|
+
ambiguity). Sister findings: [[feedback_correlated_blindspot_union_over_majority]] (reps≥3 prerequisite),
|
|
58
|
+
[[feedback_judge_robustness_mechanical_anchor]] (discriminating-probe = mechanical anchor),
|
|
59
|
+
[[reference_agy_model_catalog]] (display-name pin — agy slug fallback documented there).
|