@staix/agent-hub 0.12.6 → 0.12.8
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +15 -0
- package/LICENSES/Apache-2.0.txt +204 -0
- package/THIRD_PARTY_NOTICES.md +19 -0
- package/docs/cooperbench.md +1 -1
- package/docs/events.md +32 -0
- package/docs/operations.md +27 -5
- package/docs/specs/2026-09-19-agent-hub-design.md +10 -0
- package/docs/specs/2026-10-04-switchyard-source-port-design.md +309 -0
- package/docs/verification/2026-10-03-0.12.6.md +9 -1
- package/docs/verification/2026-10-03-0.12.7.md +58 -0
- package/package.json +4 -2
- package/plugins/agent-hub/.claude-plugin/plugin.json +1 -1
- package/plugins/agent-hub/server.js +635 -957
- package/src/adapters/codex-appserver.ts +2 -2
- package/src/adapters/local-worker.ts +138 -26
- package/src/adapters/pi.ts +8 -0
- package/src/hub/daemon.ts +55 -14
- package/src/hub/events.ts +5 -0
- package/src/hub/inference.ts +20 -0
- package/src/hub/progress.ts +251 -0
- package/src/hub/project.ts +5 -4
- package/src/hub/routing.ts +4 -0
- package/src/local/tools.ts +9 -0
- package/src/models/relay.ts +10 -1
- package/src/models/route/advisor.ts +102 -0
- package/src/models/route/config.ts +41 -0
- package/src/models/route/escalation.ts +82 -0
- package/src/models/route/judge.ts +23 -0
- package/src/models/route/labels.ts +167 -0
- package/src/models/route/normalize.ts +103 -0
- package/src/models/route/plan-execute.ts +52 -0
- package/src/models/route/prompts.ts +8 -0
- package/src/models/route/relay-selector.ts +102 -0
- package/src/models/route/runtime.ts +156 -0
- package/src/models/route/signals.ts +234 -0
- package/src/models/route/stage.ts +93 -0
- package/src/models/route/state.ts +54 -0
- package/src/models/route/text.ts +61 -0
- package/src/omniroute/client.ts +8 -1
- package/templates/routing.toml +28 -0
|
@@ -0,0 +1,309 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Switchyard 라우팅 로직의 소스 수준 도입 계획 - agent-hub TypeScript 이식
|
|
3
|
+
date: 2026-10-04
|
|
4
|
+
project: agent-hub
|
|
5
|
+
status: approved (implementation requested 2026-10-04)
|
|
6
|
+
related: 261004-review-agent-hub-switchyard-comparison.md, 261004-plan-agent-hub-research-v2.md
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
# Switchyard 라우팅 로직의 소스 수준 도입 계획
|
|
10
|
+
|
|
11
|
+
## 요약
|
|
12
|
+
|
|
13
|
+
- **방식**: Switchyard 바이너리(sidecar)를 띄우는 대신, 필요한 라우팅 로직을 agent-hub의 TypeScript 소스로 옮겨 hub 프로세스 안에서 실행함.
|
|
14
|
+
- 근거
|
|
15
|
+
- agent-hub는 runtime 의존성 0이 원칙인데 Switchyard 서버는 "Demo" 등급, pre-1.0, 릴리스 바이너리가 없음(cargo 빌드).
|
|
16
|
+
- hub가 이미 모델 호출 루프를 직접 가진 곳이 두 군데 있음: local worker, Pi용 model relay.
|
|
17
|
+
- 그 루프 안에서는 HTTP 프록시도 프로토콜 변환도 필요 없음.
|
|
18
|
+
- **이식 대상** (Switchyard main `c8848511`, 2026-10-02)
|
|
19
|
+
- 1순위 순수 로직
|
|
20
|
+
- 도구 신호 추출 + Stage 점수기.
|
|
21
|
+
- Plan/Execute.
|
|
22
|
+
- Advisor gate(트리거, 검토 예산, transcript, 판정 해석, REDO).
|
|
23
|
+
- Escalation(요약기, 확인 연속·latch).
|
|
24
|
+
- 공용 도구: middle-drop, 판정 해석, 세션 상태.
|
|
25
|
+
- 2순위: capability 분류기, 호출 재시도·cooldown·deadline 정책의 일부.
|
|
26
|
+
- 이식하지 않음: decision judge(확률 응답 API 필요), 스트림 재생, Terminus 전용 중복 제거, Codex Responses 원시 블록, OTel 계측, 사용자 정의 ToolSemantics.
|
|
27
|
+
- **착륙 지점**
|
|
28
|
+
- local worker의 모델 호출(`src/adapters/local-worker.ts`의 `call()`).
|
|
29
|
+
- Pi model relay의 backend 선택(`src/models/relay.ts`의 `selectBackend`).
|
|
30
|
+
- hub inference(`src/hub/inference.ts`).
|
|
31
|
+
- peer 관측(`src/hub/facts.ts`, `codex-appserver.ts`의 항목).
|
|
32
|
+
- **바로 쓸 수 있는 tier**: OmniRoute 조합 `coding`(capable)과 `fast`(efficient), Mac 로컬 MLX(Qwen3.5 4B, 8k 컨텍스트). 새 모델 없이 시작 가능.
|
|
33
|
+
- **규모 추정**: 이식 코드 약 1,700~2,100줄, 연결 약 500~800줄, 테스트 약 1,500줄 이상. Switchyard 테스트 중 순수 함수 표 테스트 약 100건을 golden으로 옮김.
|
|
34
|
+
- **단계**
|
|
35
|
+
- 0단계: spec·라이선스·뼈대.
|
|
36
|
+
- 1단계: 순수 핵심 이식과 golden 테스트.
|
|
37
|
+
- 2단계: local worker 연결.
|
|
38
|
+
- 3단계: Pi relay 연결.
|
|
39
|
+
- 4단계: peer 진행 신호와 정체 판정(연구용).
|
|
40
|
+
- 5단계: sidecar 존치 결정.
|
|
41
|
+
- **연구와의 관계**
|
|
42
|
+
- 4단계의 peer 진행 신호는 연구 계획 v2의 NC3(실행 신호로 조정 방식 전환) 자료.
|
|
43
|
+
- Advisor gate 이식은 #42 AC2 감독자 regime과 NC2 commit gate의 시제품 기반.
|
|
44
|
+
|
|
45
|
+
## 이식 대상 분석 (소스 기준)
|
|
46
|
+
|
|
47
|
+
- 출처: Switchyard `crates/libsy/src`(아래 `A/`는 `algorithms/`). 줄 수는 원 저장소 기준.
|
|
48
|
+
- 모든 `.rs` 파일 머리에 `SPDX-License-Identifier: Apache-2.0`가 있음. 프롬프트 `.md`·`.json`에는 머리말이 없음.
|
|
49
|
+
|
|
50
|
+
| 구성 요소 | 원 로직(raw/code) | 원 테스트(건) | 순수 여부 | 턴당 외부 호출 | TS 추정 | 난이도 |
|
|
51
|
+
|---|---|---|---|---|---|---|
|
|
52
|
+
| Stage(`A/stage.rs`, `A/util/stage.rs`, `A/util/tool_signals.rs`) | 2,288 / 1,711 | 111 | 순수(판정기 선택) | 0, 모호할 때 판정기 1 | 950~1,100 | 중 |
|
|
53
|
+
| Advisor gate(`A/advisor_gate*`) | 1,195 / 882 | 47 | 혼합 | 검토 시 advisor 1, REDO 시 executor 1 | 300~400 | 하 |
|
|
54
|
+
| Escalation(`A/escalation.rs`, `A/util/escalation.rs`) | 1,008 / 825 | 29 | 혼합 | latch 전 efficient 1 + 판정기 1 | 450~550 | 중 |
|
|
55
|
+
| Capability 분류기(`A/llm_class.rs` 일부) | 일부 | 33 | 혼합 | 판정기 1 | 250~350 | 하 |
|
|
56
|
+
| Plan/Execute(`A/plan_execute.rs`) | 202 / 171 | 4 | 순수 | 0 | 80~120 | 하 |
|
|
57
|
+
| 공용(`llm_judge`, `robustness`, `affinity`, `fall_through`, `prompts`) | 1,282 / 911 | 57 | 판정 호출 외 순수 | - | 200~300 | 하 |
|
|
58
|
+
| 호출 정책(재시도, cooldown, deadline) | llm-client 쪽 | - | - | 모든 호출 | 120~180 | 하~중 |
|
|
59
|
+
|
|
60
|
+
### Stage 점수기 (원문 확인)
|
|
61
|
+
|
|
62
|
+
- **신호 추출**: 대화의 도구 호출·결과를 훑어 계산함.
|
|
63
|
+
- 쓰기·편집·읽기·계획·새 호출 수(최근 3회 창).
|
|
64
|
+
- 결과 창의 오류 심각도: SOFT 0.3, HARD 0.7, CRITICAL 1.0. 오류 문자열 표 약 11종(traceback, import error, assertion, timeout, out of memory 등)과 컴파일 오류·런타임 예외·panic 탐지.
|
|
65
|
+
- 같은 오류 지문의 반복, 연속 무오류 수, 테스트 통과 판정(통과·실패 문구와 숫자 키워드).
|
|
66
|
+
- **도구 어휘**: 하네스별 도구 이름 표.
|
|
67
|
+
- 편집: edit, multiedit, apply_patch 등. 쓰기: write, create_file 등.
|
|
68
|
+
- 읽기: read, grep, glob 등. 계획: todowrite, update_plan 등. 셸: bash, shell_command, exec_command 등.
|
|
69
|
+
- 셸 명령을 분해해 쓰기·편집·읽기를 추정함(`sed -i`, `cat >`, `git apply`, 포매터 등).
|
|
70
|
+
- **점수**
|
|
71
|
+
- 차원: severity, spinning(깊이 8 이상에서 최근 쓰기·읽기·계획 없음), exploring(최근 읽기·계획만), production_intensity(최근 쓰기+편집 비율).
|
|
72
|
+
- `score = tanh(5 · 0.1 · (severity/0.7 + spinning + exploring − production_intensity))`, 신뢰도 = |score|.
|
|
73
|
+
- 무조건 capable: 컨텍스트 압축 흔적, 심각도 1.0, 같은 오류 반복.
|
|
74
|
+
- 임계값 t의 모호 구간(|score| ≤ t)은 분류기 또는 picker 기본 tier로.
|
|
75
|
+
- capable 결정 뒤 2회 유지(hold). 깨끗한 테스트 통과가 유지를 해제.
|
|
76
|
+
- **상태**: 세션별(1시간 TTL). hold 키는 agent id별로 나뉠 수 있음.
|
|
77
|
+
- **주의**: t = 0.5에서는 Dimensions 경로만으로 efficient를 고를 수 없음(최소가 −1 단위, 점수 −0.46). efficient는 picker 기본값으로만 선택됨.
|
|
78
|
+
|
|
79
|
+
### Advisor gate (원문 확인)
|
|
80
|
+
|
|
81
|
+
- **트리거**
|
|
82
|
+
- `no_tool_call`: 도구 호출 없는 turn이면서 도구 결과가 `gate_min_tool_results` 이상.
|
|
83
|
+
- `pattern`: 정규식.
|
|
84
|
+
- 정체 검사: assistant turn이 `gate_stall_turns` 이상일 때 대화당 1회.
|
|
85
|
+
- **예산**: 범위(세션)별 `max_reviews`. 판정기 실패는 예산을 돌려주고, 실패 3회면 그 범위의 검토를 끔. 추적 범위 상한 1,024.
|
|
86
|
+
- **transcript**: 대화 JSON 뒤에 executor의 마지막 turn을 붙임. 상한을 넘으면 앞 1/4과 뒤 3/4만 남기고 가운데를 표시와 함께 버림(middle-drop).
|
|
87
|
+
- **판정 해석**: 첫 단어 APPROVE 또는 REDO(정규식, 대소문자 무시, "verdict:" 접두 허용). REDO 뒤 나머지가 계획.
|
|
88
|
+
- **REDO**: executor 응답을 assistant로, `redo_feedback_prefix + 계획`을 user로 붙여 executor를 다시 부름.
|
|
89
|
+
- **실패 처리**: fail-open(기본)이면 보류한 turn을 그대로 통과.
|
|
90
|
+
- **원문의 문서와 코드 불일치**: 문서는 "Rust runner가 fail_open과 무관하게 HTTP 실패에서 멈춘다"고 하나, 코드는 `fail_open`을 따름. 이식은 코드를 기준으로 함.
|
|
91
|
+
|
|
92
|
+
### Escalation (원문 확인)
|
|
93
|
+
|
|
94
|
+
- **기본값**: confirmations 2, recent_turn_window 28, window_message_chars 500.
|
|
95
|
+
- **판정 결과**: `escalate`, `category`(none, repetition, false_progress, drift, desperation, capability_gap), `new_evidence`, `reason`.
|
|
96
|
+
- **확인 연속**: 같은 범주의 새 증거면 +1, 범주가 바뀌면 1부터, 거절이나 낡은 증거면 0. 해석 실패는 연속을 유지.
|
|
97
|
+
- **latch**: 확인 수에 도달하면 capable로 고정. de-escalation은 선택.
|
|
98
|
+
- **요약기**: 판정기 입력은 지시문(1,000자), 첫 사용자 요청(4,000자), 최근 창. 전체 상한 18,000자.
|
|
99
|
+
- **실패 처리**: Rust 호스트에서는 판정기 HTTP 실패가 요청을 멈춤(fail-closed). 이식에서는 agent-hub inference 규칙(fail-open, backoff)을 따르도록 바꿈. 이는 의도한 차이로 spec에 적음.
|
|
100
|
+
|
|
101
|
+
### Plan/Execute (원문 확인)
|
|
102
|
+
|
|
103
|
+
- 첫 변경(편집·쓰기, 셸로 추정한 것 포함)이 보이기 전까지 capable로 계획. 변경을 보면 그 세션을 실행 단계로 고정하고 efficient로.
|
|
104
|
+
- 계획 단계에는 계획용 시스템 프롬프트를 앞에 붙임.
|
|
105
|
+
|
|
106
|
+
## agent-hub 착륙 지점
|
|
107
|
+
|
|
108
|
+
- **local worker** (`src/adapters/local-worker.ts`)
|
|
109
|
+
- `call()`이 매 호출의 메시지(`system` + `history` + 이번 turn)와 도구 목록을 만들고, sidecar 경유 route 또는 OmniRoute `fixed_model`로 호출함.
|
|
110
|
+
- 도구: `read`, `write`, `edit`, `bash`(+ 작업 도구). Stage 어휘의 read/write/edit/bash와 이름이 같음.
|
|
111
|
+
- 호출은 스트리밍이 아님(`ChatResult`). advisor gate의 응답 보류가 자연스러움.
|
|
112
|
+
- 작업 class별 route 선택이 이미 있음(`turnPolicy`). PII turn은 on-prem 경로 확인(`onCampus()`) 규칙이 있음.
|
|
113
|
+
- **Pi model relay** (`src/models/relay.ts`)
|
|
114
|
+
- Pi의 OpenAI 호환 요청을 받아 MLX 또는 DGX(OmniRoute)로 보냄. 별칭은 `dgx/coding`, `dgx/fast`, `mlx/fast`이고 MLX 실패 시 `dgx/fast`로 폴백.
|
|
115
|
+
- `selectBackend(request)` 훅이 정의돼 있으나 지금은 쓰이지 않음(`daemon.ts`가 넘기지 않음).
|
|
116
|
+
- 응답을 SSE로 그대로 흘려보냄. 따라서 advisor gate에는 스트림 보류가 필요함.
|
|
117
|
+
- **hub inference** (`src/hub/inference.ts`)
|
|
118
|
+
- digest 요약과 작업 분류(닫힌 목록)를 함.
|
|
119
|
+
- 규칙: fail-open, 8초 제한, 실패 시 5분 backoff. 출력은 길이 제한 텍스트이거나 닫힌 목록과 대조한 값만.
|
|
120
|
+
- **peer 관측**
|
|
121
|
+
- Codex: `item/completed`의 `fileChange`, `commandExecution`, `mcpToolCall`(`codex-appserver.ts`).
|
|
122
|
+
- Claude: turn-free 프로젝트에서만 hook으로 도구 호출을 받음(Edit, Write, MultiEdit, Read, Bash 등).
|
|
123
|
+
- local worker와 Pi: hub가 모든 호출을 앎.
|
|
124
|
+
|
|
125
|
+
## 설계 결정 (권고안)
|
|
126
|
+
|
|
127
|
+
- **E1 모듈 위치**
|
|
128
|
+
- `src/models/route/`(모델 경로의 정책, 기존 `src/models/relay.ts`와 같은 층).
|
|
129
|
+
- 작업 배정 정책인 `src/hub/routing.ts`와 이름이 겹치지 않게 함.
|
|
130
|
+
- 파일: `signals.ts`, `stage.ts`, `plan-execute.ts`, `advisor.ts`, `escalation.ts`, `classifier.ts`(2순위), `judge.ts`(판정 해석, fence 제거, 닫힌 목록 대조), `text.ts`(middle-drop, truncate-middle, append-note), `state.ts`(세션별 상태, TTL, 상한), `prompts.ts`.
|
|
131
|
+
- **E2 호스트 주도 계약**
|
|
132
|
+
- Switchyard의 `Step::CallModel`/`Done` 구조를 따름. 알고리즘은 "이 tier로" 또는 "이 판정기 요청을 대신 호출해 달라"를 돌려줌.
|
|
133
|
+
- 호출은 호스트(local worker, relay, inference)가 기존 OmniRoute client로 수행. 순수 함수 테스트가 쉬워지고 runtime 의존성 0을 유지.
|
|
134
|
+
- **E3 충실도**
|
|
135
|
+
- 상수, 공식, 어휘, 우선순위를 원문대로 옮김. 파일 머리에 원 경로와 커밋을 적음.
|
|
136
|
+
- 의도한 차이는 spec에 목록으로 둠:
|
|
137
|
+
- Escalation 판정기 실패를 fail-open으로.
|
|
138
|
+
- REDO 피드백을 local worker 이력에 남김.
|
|
139
|
+
- Responses·Codex 원시 블록 제외.
|
|
140
|
+
- Escalation의 여러 작업 앵커가 전체 입력 상한을 넘을 때 최신 실행 창을 보존하도록 앵커를 함께 축약. 상류의 최신 활동 누락 버그를 의도적으로 수정.
|
|
141
|
+
- 문자 수는 Rust의 Unicode scalar 기준이므로 JS에서는 code point 단위(`[...s]`)로 셈. 정규식은 `u` 플래그.
|
|
142
|
+
- **E4 입력 정규화**
|
|
143
|
+
- OpenAI chat 메시지를 Switchyard 내부 형태로 맞추는 어댑터.
|
|
144
|
+
- `role: "tool"`은 사용자 역할의 도구 결과.
|
|
145
|
+
- 인자 문자열은 JSON으로 해석하고, 실패하면 `{raw}`.
|
|
146
|
+
- system 메시지는 지시문으로 분리.
|
|
147
|
+
- peer 관측용 어댑터: Codex `commandExecution`(명령, exit code, 출력)과 `fileChange`, Claude hook 도구 호출을 같은 신호 구조로 변환.
|
|
148
|
+
- **E5 설정**
|
|
149
|
+
- `routing.toml`에 `[hub_routes."<id>"]` 표를 새로 둠. 유형은 `stage`, `plan_execute`, `advisor`, `escalation`, tier는 OmniRoute 모델 id나 relay 별칭.
|
|
150
|
+
- Switchyard sidecar용 `[routes.*]`와 분리해 sidecar 설정 생성기가 옮기지 않게 함.
|
|
151
|
+
- local worker의 route가 `hub/` 접두면 hub 안에서, `sy/` 접두면 sidecar(있을 때)로 감.
|
|
152
|
+
- **E6 실패 처리**
|
|
153
|
+
- 라우팅 판단이나 판정기 호출이 실패하면 기존 경로(`fixed_model`)로 폴백. turn과 전달을 늦추지 않음.
|
|
154
|
+
- 판정기 호출에 deadline을 둠(inference와 같은 원칙).
|
|
155
|
+
- **E7 PII**
|
|
156
|
+
- PII turn의 판정기·advisor 호출은 on-prem 경로가 확인될 때만(`onCampus()`). 아니면 검토 없이 진행.
|
|
157
|
+
- peer 진행 판정은 PII 작업을 제외. 이벤트에 작업 텍스트를 넣지 않음(`publicView` 규칙).
|
|
158
|
+
- **E8 이벤트와 결정 라벨** (이슈 #126, 2026-10-04 추가 요구사항 반영)
|
|
159
|
+
- 기본 `route`: peer, route, tier, source(override, dimensions, hold, classifier, default), score, ms.
|
|
160
|
+
- local worker의 모든 `route`에 추가: 고유 `decision` id, `turn_start` / `turn_end`와 같은 `turn` id, 작업이 있는 경우 숫자 `task`, boolean `pii`, 결정 시점의 `severity`, `spinning`, `exploring`, `production`.
|
|
161
|
+
- `route_outcome`: `decision`으로 연결하고 각 결정당 정확히 1회 기록. 해당 turn이 완료되거나 실패할 때까지 보류해 이후 escalation latch도 반영.
|
|
162
|
+
- `next`: 그 호출 응답에 대한 도구 결과의 severity(0, 0.3, 0.7, 1.0), tests(pass, fail, none), repeat(이전 오류 지문과 반복 여부). 최종 응답처럼 다음 도구 결과가 없으면 생략.
|
|
163
|
+
- `advisor`: 해당 호출의 최종 응답을 검토했을 때 approve, redo, failed 중 하나.
|
|
164
|
+
- `turn`: completed 또는 failed. 원래 turn id는 `route.turn`을 통해 연결하며 outcome의 `turnId`도 같은 id를 기록.
|
|
165
|
+
- `latched`: 같은 turn에서 이후 capable escalation이 latch됐는지 여부.
|
|
166
|
+
- 성공, 모델 실패, 예산 중단, watchdog 취소, stop을 포함한 모든 local turn 종료에서 미정 outcome을 정리. 종료한 turn의 늦은 완료가 새 turn의 라벨을 오염시키거나 중복 outcome을 만들지 않도록 세대 확인.
|
|
167
|
+
- 숫자, id, boolean, 닫힌 목록만 기록. 메시지·도구 출력·작업 텍스트·오류 지문 원문은 라벨에 넣지 않음. PII도 같은 형태에 `pii: true`로 기록해 export에서 제외 가능.
|
|
168
|
+
- 선택적 telemetry. sink 예외가 turn, 모델 선택, 결과 전달에 영향을 주지 않음.
|
|
169
|
+
- `advisor`: trigger, verdict, 버린 문자 수. `progress`: peer, task, severity, spinning, exploring, production. `stuck`: peer, task, category, streak, latched.
|
|
170
|
+
- task review 결과는 offline join: `route.task`와 기존 `task` 이벤트의 approved / changes_requested를 연결. `docs/events.md`에 절차 기록. 새 명령은 추가하지 않음.
|
|
171
|
+
- Pi relay 라벨은 #126 범위 밖. 기존 Pi `route` 형태와 고정 backend 동작을 보존하며, 다음 요청의 도구 결과를 쓰는 확장은 후속 작업.
|
|
172
|
+
- `events.jsonl`의 추가 필드·새 이벤트만 변경. control WS 모양은 유지하므로 `PROTOCOL`은 올리지 않음.
|
|
173
|
+
- **E9 라이선스**
|
|
174
|
+
- 이식한 파일마다 머리말을 둠: 원 SPDX 두 줄을 남기고, "Ported to TypeScript from NVIDIA NeMo Switchyard `<path>` at `c8848511`, modified"를 적음.
|
|
175
|
+
- 저장소 루트에 `THIRD_PARTY_NOTICES.md`(Switchyard NOTICE 귀속 문구)와 `LICENSES/Apache-2.0.txt`.
|
|
176
|
+
- `package.json` `files`와 `scripts/check.sh`의 npm tarball 내용 검사에 추가.
|
|
177
|
+
- 프롬프트는 `prompts.ts`의 문자열 상수로 넣고 같은 머리말을 붙임.
|
|
178
|
+
- **E10 Switchyard 추적**
|
|
179
|
+
- 고정 커밋 기준으로 이식. 상류 변경은 상수·어휘·프롬프트 파일의 diff를 보는 점검 스크립트로 분기마다 확인.
|
|
180
|
+
- 자동 동기화는 하지 않음.
|
|
181
|
+
|
|
182
|
+
## 단계별 계획
|
|
183
|
+
|
|
184
|
+
### 0단계: spec과 뼈대
|
|
185
|
+
|
|
186
|
+
- 저장소 이슈 본문 = spec. 크기상 `docs/specs/`에 설계 파일을 둠(설계 spec의 L2 절 개정 포함).
|
|
187
|
+
- 라이선스 파일(E9), 모듈 뼈대, 프롬프트 상수, golden 테스트 틀.
|
|
188
|
+
- 완료 기준: `scripts/check.sh` 통과, npm tarball에 고지 파일 포함.
|
|
189
|
+
|
|
190
|
+
### 1단계: 순수 핵심 이식과 golden 테스트 (동작 변화 없음)
|
|
191
|
+
|
|
192
|
+
- **이식**
|
|
193
|
+
- `signals.ts`: 추출, 어휘, 심각도, 지문, 테스트 통과 판정.
|
|
194
|
+
- `stage.ts`: 차원, 점수, `pick_tier`, hold.
|
|
195
|
+
- `plan-execute.ts`.
|
|
196
|
+
- `text.ts`, `judge.ts`, `state.ts`.
|
|
197
|
+
- `advisor.ts`: 트리거, 예산, transcript, 판정 해석, REDO 메시지 구성.
|
|
198
|
+
- `escalation.ts`: 요약기, 연속·latch 상태기계.
|
|
199
|
+
- **golden 테스트** (Switchyard 테스트를 옮김)
|
|
200
|
+
- tool_signals 표 테스트 약 60건(셸 분해, 포매터, 테스트 통과, 반복 실패 등).
|
|
201
|
+
- `pick_tier`·override·note 동기 테스트 13건.
|
|
202
|
+
- advisor 판정 해석 표 12건, middle-drop 정확 문자열, 예산 5건, 트리거 4건.
|
|
203
|
+
- escalation 요약·창·상한·truncate 테스트 약 10건(Terminus 전용 5건 제외).
|
|
204
|
+
- Plan/Execute 4건.
|
|
205
|
+
- **선택: 차등 검증**
|
|
206
|
+
- 실제 agent-hub 대화(local worker·Pi 기록)를 모아 Rust 원본과 TS 이식의 결정을 비교.
|
|
207
|
+
- Rust 빌드가 필요하므로 무거운 작업 규칙(한 번에 하나)을 따르고, 결과 fixture만 저장소에 넣음.
|
|
208
|
+
- **완료 기준**: golden 전부 통과, 판단 경로 p95가 메시지 200개 이력에서 수 ms 이내, `scripts/check.sh` 통과.
|
|
209
|
+
|
|
210
|
+
### 2단계: local worker 연결
|
|
211
|
+
|
|
212
|
+
- `call()`에 hub route를 추가.
|
|
213
|
+
- stage: 매 호출 tier 선택.
|
|
214
|
+
- plan_execute: 첫 변경 전 capable, 뒤로 efficient.
|
|
215
|
+
- escalation: efficient 시작, 확인된 정체에서 capable로 latch.
|
|
216
|
+
- advisor: 도구 호출 없는 마지막 응답을 상위 모델이 검토. REDO면 피드백을 이번 turn 메시지에 붙여 계속(`maxSteps` 안에서).
|
|
217
|
+
- 기본 tier: efficient `fast`, capable `coding`(OmniRoute 조합). 설정으로 바꿈.
|
|
218
|
+
- 이벤트(E8), 사용량 기록에 선택한 tier 표시.
|
|
219
|
+
- 기존 규칙 유지: 한 turn의 메시지는 turn 끝에 history로 합류(중간 push 금지), 부작용 있는 도구 실행 뒤 실패한 turn은 재전달하지 않음.
|
|
220
|
+
- 테스트: `test/fakes/model-server.ts`로 도구 결과를 각본화. 오류 반복 → capable, 테스트 통과 → hold 해제, APPROVE/REDO, 판정기 장애 → 폴백, PII → 판정기 생략.
|
|
221
|
+
- 완료 기준: 위 시나리오 통과, 라우팅 실패가 turn 실패로 번지지 않음, 문서(`docs/operations.md`, `templates/routing.toml`) 갱신.
|
|
222
|
+
|
|
223
|
+
### 3단계: Pi relay 연결
|
|
224
|
+
|
|
225
|
+
- **3a**: relay에 가상 별칭 `hub/auto`를 추가하고 `selectBackend`에서 stage 점수로 `dgx/coding`, `dgx/fast`, `mlx/fast` 중 선택.
|
|
226
|
+
- MLX는 8k 컨텍스트이므로 입력이 크면 MLX를 후보에서 뺌. 기존 MLX→DGX 폴백 유지.
|
|
227
|
+
- **3b** (선택): Pi의 advisor gate. SSE 응답을 보류했다가 재생해야 하므로 스트림 버퍼(`buffered_response` 약 80줄)를 함께 이식.
|
|
228
|
+
- 완료 기준: 별칭 선택 테스트, 폴백 테스트, Pi 실사용 smoke 1회.
|
|
229
|
+
|
|
230
|
+
### 4단계: peer 진행 신호와 정체 판정 (연구용)
|
|
231
|
+
|
|
232
|
+
- **진행 신호**
|
|
233
|
+
- E4의 관측 어댑터로 peer·작업별 신호를 계산해 `progress` 이벤트로 남김.
|
|
234
|
+
- 대상: Codex, local, Pi, Claude(turn-free hook이 있을 때만).
|
|
235
|
+
- **정체 판정**
|
|
236
|
+
- inference에 escalation 판정을 추가함. 프롬프트는 "모델 tier"를 "다른 peer로 이관"으로 바꿔 고쳐 씀. 결과는 닫힌 목록과 대조.
|
|
237
|
+
- 판정기 호출은 진행 신호가 문제를 보일 때만(오류 반복, spinning) 해서 비용을 줄임.
|
|
238
|
+
- 연속·latch는 peer·작업별.
|
|
239
|
+
- **출력**: 콘솔 알림과 `stuck` 이벤트, 재배정 제안까지. 자동 이관은 하지 않음(측정 뒤 결정).
|
|
240
|
+
- **연구 쪽**: 벤치마크 ledger에 진행 신호 시계열을 넣어 NC3 분석 자료로 씀(관측 범위가 peer마다 다르다는 점을 함께 기록).
|
|
241
|
+
- 완료 기준: Codex 각본 테스트, PII 제외 테스트, 판정기 장애 시 fail-open.
|
|
242
|
+
|
|
243
|
+
### 5단계: sidecar 존치 결정
|
|
244
|
+
|
|
245
|
+
- 2·3단계 뒤 hub 자체 트래픽이 in-process 라우팅으로 충분하면, Switchyard sidecar를 선택 기능으로 낮춤. 이식하지 않은 알고리즘(composite 분류기, sub-agent)이 필요할 때만 사용.
|
|
246
|
+
- 구독 에이전트 측정 프록시(비교 보고서의 D4)는 이 계획 밖. 별도 결정.
|
|
247
|
+
|
|
248
|
+
## 규모와 일정 (추정)
|
|
249
|
+
|
|
250
|
+
- 0단계: 0.5일.
|
|
251
|
+
- 1단계: 2~3일. 이식 약 1,700~2,100줄, 테스트 약 1,500줄.
|
|
252
|
+
- 2단계: 1.5~2일. 연결 약 300~500줄, 테스트.
|
|
253
|
+
- 3단계: 3a 0.5일, 3b 1일.
|
|
254
|
+
- 4단계: 2일.
|
|
255
|
+
- 각 단계는 별도 PR. Full PR 게이트(OCR 리뷰, CI 두 플랫폼) 적용.
|
|
256
|
+
- ICSE Tool Demo(10-23) 일정과 겹치므로 착수 시점은 결정 대기.
|
|
257
|
+
|
|
258
|
+
## 위험
|
|
259
|
+
|
|
260
|
+
- **보정 차이**: Switchyard의 임계값과 어휘는 그쪽 벤치마크 모델(GPT-5.6 Luna/Sol 등)로 맞춘 것. 우리 tier(OmniRoute `coding`·`fast`, MLX 4B)에서는 다시 맞춰야 함. 기본은 efficient-first, 임계값은 측정 뒤 조정.
|
|
261
|
+
- **휴리스틱의 오판**: 원문에서도 `ls x 2> /dev/null`이 쓰기로 분류되는 등 부분 문자열 규칙의 한계가 있음. 그대로 옮기고 golden으로 고정하되, 고칠 때는 의도한 차이로 기록.
|
|
262
|
+
- **하네스 어휘 변화**: 벤더 도구 이름이 바뀌면 신호가 틀어짐. 상류 점검(E10)과 peer 관측 어댑터의 테스트로 대응.
|
|
263
|
+
- **컨텍스트 한도**: MLX 8k에서 도구 이력이 길면 넘침. relay의 입력 추정과 폴백을 유지하고, 컨텍스트 초과 문구 탐지(Switchyard 호출 정책)를 2순위로 이식.
|
|
264
|
+
- **REDO와 이력**: Switchyard는 REDO를 클라이언트 이력에 남기지 않지만, 우리는 hub가 이력을 가지므로 남김. 같은 피드백이 다음 turn 판단에 영향을 줄 수 있음. 테스트로 확인.
|
|
265
|
+
- **라이선스**: 고지 누락 시 Apache-2.0 위반. tarball 검사에 넣어 막음.
|
|
266
|
+
- **관측 범위의 비대칭**: Claude는 turn-free hook이 있을 때만 도구 호출이 보임. 4단계 자료를 비교할 때 이 차이를 명시.
|
|
267
|
+
|
|
268
|
+
## 이슈 구성안 (등록은 승인 후)
|
|
269
|
+
|
|
270
|
+
- 상위: "Port Switchyard routing logic in-process (stage, plan/execute, advisor gate, escalation)". spec 파일 동반.
|
|
271
|
+
- 하위 A: "Pure routing core ported from Switchyard with golden tests and license notices" (0·1단계).
|
|
272
|
+
- 하위 B: "In-process model routes for the local worker" (2단계).
|
|
273
|
+
- 하위 C: "hub/auto alias in the Pi model relay" (3a, 3b는 선택).
|
|
274
|
+
- 하위 D: "Peer progress signals and stuck verdicts" (4단계).
|
|
275
|
+
- 각 이슈 본문 = spec. 착수는 이슈별 지시 후.
|
|
276
|
+
|
|
277
|
+
## 결정이 필요한 사항
|
|
278
|
+
|
|
279
|
+
- 이식 범위 A~D 승인과 이슈 등록.
|
|
280
|
+
- tier 대응: efficient `fast`, capable `coding`, MLX의 역할.
|
|
281
|
+
- REDO 피드백을 local worker 이력에 남기는 정책.
|
|
282
|
+
- 4단계의 자동 재배정 여부(권고: 하지 않음, 측정 뒤 결정).
|
|
283
|
+
- 1단계의 차등 검증(Rust 빌드 필요) 실행 여부.
|
|
284
|
+
- Switchyard sidecar의 장래(5단계).
|
|
285
|
+
- 착수 시점(ICSE Tool Demo 마감 10-23 이후 권고).
|
|
286
|
+
|
|
287
|
+
## 근거
|
|
288
|
+
|
|
289
|
+
- Switchyard(main `c8848511a7e2e1d605070c7a68905bdc24c6481a`, 2026-10-02, 로컬 shallow clone)
|
|
290
|
+
- `crates/libsy/src/algorithms/stage.rs`, `util/stage.rs`(상수 35~47행, 점수 335~347행), `util/tool_signals.rs`(심각도 31~33행, 어휘 100~224행).
|
|
291
|
+
- `advisor_gate.rs`, `advisor_gate/{budget,trigger,transcript,turn,signals}.rs`(middle-drop 51~73행, 실패 상한 18행).
|
|
292
|
+
- `escalation.rs`, `util/escalation.rs`(기본값 134~143행), `plan_execute.rs`, `llm_class.rs`, `util/{llm_judge,robustness,affinity,prompts}.rs`, `fall_through.rs`.
|
|
293
|
+
- `crates/libsy/src/prompts/*`, `crates/libsy-llm-client/src/{client,backend,run}.rs`(호출 정책).
|
|
294
|
+
- 소스 분석 보고(조사 에이전트, 2026-10-04)와 주요 상수·공식의 직접 확인.
|
|
295
|
+
- agent-hub(main `ab16c7b`)
|
|
296
|
+
- `src/adapters/local-worker.ts`(`call()`, `commit()`), `src/local/tools.ts`(도구 이름), `src/models/relay.ts`(`selectBackend`, 폴백), `src/hub/daemon.ts`(relay 시작, Pi 설정), `src/hub/inference.ts`, `src/hub/facts.ts`, `src/adapters/codex-appserver.ts`, `test/fakes/model-server.ts`, `docs/events.md`, `AGENTS.md`(inference·history·PII 규칙).
|
|
297
|
+
- 기본 설정: Pi `dgx_coding = "coding"`, `dgx_fast = "fast"`, MLX `qwen3.5:4b-mlx`(8k).
|
|
298
|
+
|
|
299
|
+
## Implementation contract (approved 2026-10-04)
|
|
300
|
+
|
|
301
|
+
- Required scope: phases A (0/1), B (2), C (3a), D (4).
|
|
302
|
+
- Defaults: efficient `fast`, capable `coding`; MLX only within its configured input and output context limit.
|
|
303
|
+
- REDO feedback stays in the completed local turn history. Failed turns never leave unmatched tool calls.
|
|
304
|
+
- PII bypasses optional judges unless the exact on-campus gateway is positively confirmed; progress tracking is disabled while PII work is open.
|
|
305
|
+
- Optional Pi advisor SSE replay and Rust differential build are deferred. No automatic peer reassignment.
|
|
306
|
+
- Existing sidecar routes remain optional and backward compatible; `[hub_routes]` never enters generated sidecar TOML. The sidecar retirement decision remains evidence-driven.
|
|
307
|
+
- Each required phase is a separate reviewable commit and PR. No release or production configuration changes are implied.
|
|
308
|
+
|
|
309
|
+
Implementation tracking: [#124](https://github.com/STAIxBWLB/agent-hub/issues/124), core [#125](https://github.com/STAIxBWLB/agent-hub/issues/125), local [#126](https://github.com/STAIxBWLB/agent-hub/issues/126), Pi [#127](https://github.com/STAIxBWLB/agent-hub/issues/127), observations [#128](https://github.com/STAIxBWLB/agent-hub/issues/128).
|
|
@@ -47,7 +47,15 @@ Cost of the group stops: each ACP or Pi stop now reads the process table twice (
|
|
|
47
47
|
| V | solo Codex with a visitor in the fixture | 31.0 s into the work | SIGTERM to the runner | `incomplete_or_unknown`, the visitor left alone and recorded as `program: sleep` | the visitor | withheld; `restore.ts` refused while it ran, restored once it left |
|
|
48
48
|
| S2 | advisory | native readiness | SIGINT to the group | `clean` (452 ms) | nothing | inputs, locks, trust |
|
|
49
49
|
|
|
50
|
-
Runs A, V and S2 were repeated on `fdc884e` with the same outcomes. Every one of these runs registered its fixtures with Orca, since the runner still added them then. The user has since directed that no fixture is added to Orca without an explicit registration (#117), and the registrations these runs left are removed under #117. Since `70b1daf` the runner only looks registrations up, and since `b03b035` it refuses an unregistered selection up front. A live check of the latter registered nothing. A run prepared afresh (`/private/tmp/ahub-0126-refusal`, 40 fixtures, none registered) was started with the runner of `b03b035`. It refused the whole run before anything was locked or recorded, naming every unregistered fixture. That lookup was later replaced by #118's read-only helper (`scripts/benchmarks/orca-workspace.ts`, merged into this branch from main), which refuses at the first fixture missing and checks each arm's identity against the preflight; #118 records its own checks, and this refusal check was not repeated on it. Orca's repository list was the same before and after (15 repositories), the upstream checkout kept its mode, and no attempt record was written.
|
|
50
|
+
Runs A, V and S2 were repeated on `fdc884e` with the same outcomes. Every one of these runs registered its fixtures with Orca, since the runner still added them then. The user has since directed that no fixture is added to Orca without an explicit registration (#117), and the registrations these runs left are removed under #117. Since `70b1daf` the runner only looks registrations up, and since `b03b035` it refuses an unregistered selection up front. A live check of the latter registered nothing. A run prepared afresh (`/private/tmp/ahub-0126-refusal`, 40 fixtures, none registered) was started with the runner of `b03b035`. It refused the whole run before anything was locked or recorded, naming every unregistered fixture. That lookup was later replaced by #118's read-only helper (`scripts/benchmarks/orca-workspace.ts`, merged into this branch from main), which refuses at the first fixture missing and checks each arm's identity against the preflight; #118 records its own checks, and this refusal check was not repeated on it. Orca's repository list was the same before and after (15 repositories), the upstream checkout kept its mode, and no attempt record was written. Before the release, no arm had run on the lookup-only runner, which needs fixtures registered explicitly first: the live arms above ran on the runner that still added them. Every record carries `platform: darwin`. The Codex skills condition was 177 user and 5 system skills in every arm above with Codex. Arms on the released runner follow.
|
|
51
|
+
|
|
52
|
+
After the release, with the user's authorization for the Orca registrations, three of these arms ran on the released runner (`44e04e5`, Bun 1.3.14, which 0.12.6's CI used). Each run's four case-0 fixtures were registered by hand before the run and removed after it, and Orca's repository list ended as it began (16 repositories, the same ids; the one more than at the refusal check is an unrelated project added meanwhile). Every record carries `platform: darwin`; the Codex skills condition was 177 user and 4 system skills in each.
|
|
53
|
+
|
|
54
|
+
| Run | Arm | Stopped | Cleanup | Left running | Restoration |
|
|
55
|
+
|---|---|---|---|---|---|
|
|
56
|
+
| A | solo Codex | SIGTERM to the runner, 30.8 s into the work | `clean` (313 ms) | nothing | inputs, sibling locks |
|
|
57
|
+
| B | advisory | SIGINT to the group, 40.7 s into the work | `clean` (447 ms) | nothing | inputs, locks, trust |
|
|
58
|
+
| V | solo Codex with a visitor in the fixture | SIGTERM to the runner, 30.6 s into the work | `incomplete_or_unknown`, the visitor left alone and recorded as `program: sleep` | the visitor | withheld; `restore.ts` refused while it ran, restored once it left |
|
|
51
59
|
|
|
52
60
|
## Review
|
|
53
61
|
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# Agent-hub 0.12.7 verification
|
|
2
|
+
|
|
3
|
+
Implementation and candidate verification for issue #120 (the CooperBench run's restoration state), performed on 2026-10-03, with Bun 1.4.2 (#121). Native CLIs: Codex 0.160.0, Claude Code 2.1.288. The hub runtime is unchanged from 0.12.6 apart from the plugin bundle, which is built with Bun 1.4.2. Production application is verified separately after publication.
|
|
4
|
+
|
|
5
|
+
## What changed
|
|
6
|
+
|
|
7
|
+
- The runner writes `restoration.json` as `{restored: false, runner}` right before it locks anything, and its outcome replaces it at the end. A runner killed in between leaves a file that sends the recovery in.
|
|
8
|
+
- The runner refuses a run directory whose earlier run is not restored (`reuseProblem`), before it reads any mode.
|
|
9
|
+
- `restore.ts` trusts `restored: true` only when the restoration ledger agrees (`unrestored`). It waits while a runner named by either file still runs. With no ledger, nothing was locked; with neither file, there is nothing to recover.
|
|
10
|
+
|
|
11
|
+
## Automated gate
|
|
12
|
+
|
|
13
|
+
`scripts/check.sh` on Bun 1.4.2, at the head of the pull request, with the fixes of three review rounds: 759 tests passed, 0 failed, 3,848 expectations across 68 files in 132 s; `check: OK` (757 and 3,833 before them). Mutation checks (each test fails with its part removed): the old early return on `restored: true`, the runner named by `restoration.json` ignored, and a ledger required.
|
|
14
|
+
|
|
15
|
+
## Live checks
|
|
16
|
+
|
|
17
|
+
Orca lookups in these runs were answered by a read-only stand-in for the `orca` CLI. It answered `repo list` and `worktree list` for the run's own fixtures, passed `worktree current` to the real Orca, and refused everything else. Nothing was registered in Orca. The arms were real: Codex 0.160.0 through the hub, with the inputs and sibling artifacts really locked.
|
|
18
|
+
|
|
19
|
+
- **SIGKILL mid-arm** (`/private/tmp/ahub-0127-K2`, solo Codex, killed 30 s into the work):
|
|
20
|
+
- Before the kill, `restoration.json` was `restored: false` and named the runner. It stayed so after the kill. The protected root and a sibling fixture were mode 000, and the arm's daemon and Codex app-server were still running.
|
|
21
|
+
- A re-run in the same directory was refused (exit 1), but by its first read of the locked private inputs (a bare EACCES), before the reuse check. The review moved the check first; see K4.
|
|
22
|
+
- `restore.ts` refused while the daemon, the app-server and what ran below them were running (exit 1, each named). After `ahub kill` it restored the run (exit 0): the modes were back, and `restoration.json` said `recovered`.
|
|
23
|
+
- **Reuse refused** (`/private/tmp/ahub-0127-K3`, a fresh preparation whose `restoration.json` said `restored: false`): the real runner refused it with "an earlier run here is not restored (restoration.json)". It refused before it read any mode or wrote a ledger or `cohort.json`; the protected root's mode was unchanged. `restore.ts` then found no ledger and a runner that was gone, and marked the directory restored.
|
|
24
|
+
- **SIGKILL mid-arm again, after the review** (`/private/tmp/ahub-0127-K4`, the same run as K2; its log names `58c8db0`, since the fixes were not committed yet, but its `prepared.json` and `cohort.json` pin the sources of `bdc5581`): the re-run in the same directory was now refused by the reuse check, naming `restore.ts` ("an earlier run here is not restored (restoration.json): run bun scripts/benchmarks/restore.ts --run ... first"). The recovery was held while the arm's processes ran, and restored the run after `ahub kill`.
|
|
25
|
+
- **SIGKILL mid-arm on the final runner** (`/private/tmp/ahub-0127-K5`, after the third review round; its log names `ff2480e`, since the fixes were not committed yet, but its `prepared.json` pins the final `native.ts` and `teardown.ts`; and so does its `cohort.json`; `restore.ts` is not pinned, and this run's ledger names its runner, so `owed` returns it unchanged and the run does not exercise the pre-0.12.5 rule): the same outcome as K4. The marker stayed, the re-run was refused by the reuse check naming `restore.ts`, the recovery was held while the arm's processes ran, and it restored the run after `ahub kill`.
|
|
26
|
+
- **A first attempt** (`/private/tmp/ahub-0127-K`) stopped before the marker, because the stand-in did not yet answer `worktree current`. Nothing was locked or written. The recovery of that directory, which had neither file, blocked on "the ledger does not name the runner"; it now reports nothing to recover (a test covers it).
|
|
27
|
+
- **Upgrade:** `scripts/smoke-recovery-09-10.ts` recovered the published 0.12.6 package into the candidate. Both run protocol 13, and the source's integrity came from the registry. Operation `b7643429-7f0c-4f70-8128-857bb965a5d2`: the queue was preserved, and the task digest is the same as in 0.12.6's check.
|
|
28
|
+
|
|
29
|
+
## Known limits
|
|
30
|
+
|
|
31
|
+
- One runner per run directory at a time: the reuse check runs again right before the marker, but two runners started on one directory at the same moment can both pass it (a `ponytail:` in `native.ts` names the exclusive claim that would close it).
|
|
32
|
+
- Grading still trusts `restoration.json` alone. With the marker, a run whose runner died says `restored: false`, so grading refuses it; a stale `restored: true` can only come from a run before 0.12.7.
|
|
33
|
+
- Deferred, as listed in #116: a recovery of a 0.12.3 or 0.12.4 ledger killed mid restore leaves a temp file no later recovery looks for, and a 0.12.5 runner whose rename threw left its trust temp file behind a `restored: true` run.
|
|
34
|
+
|
|
35
|
+
## Review
|
|
36
|
+
|
|
37
|
+
Independent read-only reviews (Claude Code OCR delegation), each on the head of the time. The first, on `58c8db0`, went through every state a run directory can be in, and through a runner and a recovery started together. It found no Critical or Important defect. Its Minor findings were fixed:
|
|
38
|
+
|
|
39
|
+
- The reuse check runs first, before the runner resolves its inputs, so a directory whose inputs a dead runner locked is refused with a message that names `restore.ts`, not a bare EACCES (K4). It runs again right before the marker. The limit for two runners started together is recorded above.
|
|
40
|
+
- `restore.ts` reads the process table before the files, so the runners the files name are checked again right before anything is restored. A test with a runner missing from the first table covers it.
|
|
41
|
+
- A ledger from before 0.12.5 (no runner identity) under `restored: true` keeps its trust entry as it was then. Those versions left a concurrently changed trust entry at `written`, and taking it back would remove the user's entry. A test covers it.
|
|
42
|
+
- `unrestored` says a `not_written` entry whose temp file is left is still unrestored, which is what the code did; the spec and the comment now say so, with table rows.
|
|
43
|
+
- The runner's outcome uses `unrestored`, so it and the recovery's agreement check cannot disagree. Smaller wording fixes.
|
|
44
|
+
|
|
45
|
+
The second round, on `bdc5581`, found no Critical or Important defect. Its Minor finding was fixed. The exception for ledgers from before 0.12.5 covered the whole ledger, so a re-run that died over such a directory kept its inputs locked. Meanwhile the reuse check sent the operator to a recovery that then did nothing. Now only the trust entry is kept as it was then (`owed` in `teardown.ts`); the runner's reuse check and the recovery use the same rule, and the locks are restored. A test covers it and fails with the earlier rule. Its nits were fixed too:
|
|
46
|
+
|
|
47
|
+
- Nothing is written to the run directory before the marker, so a run that died between the two checks keeps its `cohort.json`.
|
|
48
|
+
- The unused table parameter is gone.
|
|
49
|
+
- The reasons read "still unrestored: ...".
|
|
50
|
+
|
|
51
|
+
The third round, on `ff2480e`, found no Critical or Important defect. Its Minor finding was fixed: a 0.12.7 runner that accepted a directory with a pre-0.12.5 ledger, wrote its marker and died before a ledger of its own left that old ledger under a marker. The recovery then took the old trust entry back, because the rule keyed on `restored: true`. `owed` now also applies under a marker that names a runner; a test covers it and fails without it. Its nits were fixed too:
|
|
52
|
+
|
|
53
|
+
- The rule lives in `owed` alone, and the recovery settles the trust entry `owed` returns.
|
|
54
|
+
- The run directory's writes are inside the `try`, so a failure there still ends with an outcome.
|
|
55
|
+
- `AGENTS.md` names `owed`, and the spec wording is fixed.
|
|
56
|
+
- K5 ran the final runner.
|
|
57
|
+
|
|
58
|
+
The fourth round, on `68a0f9b`, found no Critical or Important defect. It ran the reuse check and the recovery over 35 directory shapes, and in none did the reuse check accept a directory where the recovery then restored anything. Its nits, all wording, were fixed: the `AGENTS.md` rule, the CHANGELOG and the docs now state the condition under which the old trust entry is kept, and K5's provenance is noted.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@staix/agent-hub",
|
|
3
|
-
"version": "0.12.
|
|
3
|
+
"version": "0.12.8",
|
|
4
4
|
"description": "Native multi-agent hub: Claude Code, Codex, Kimi Code, Pi and local inference as peers in one project",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
|
@@ -29,7 +29,9 @@
|
|
|
29
29
|
"docs",
|
|
30
30
|
"README.md",
|
|
31
31
|
"LICENSE",
|
|
32
|
-
"CHANGELOG.md"
|
|
32
|
+
"CHANGELOG.md",
|
|
33
|
+
"THIRD_PARTY_NOTICES.md",
|
|
34
|
+
"LICENSES"
|
|
33
35
|
],
|
|
34
36
|
"repository": {
|
|
35
37
|
"type": "git",
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "agent-hub",
|
|
3
|
-
"version": "0.12.
|
|
3
|
+
"version": "0.12.8",
|
|
4
4
|
"description": "Channel between Claude Code and the agent-hub daemon: peer messages from Codex, Kimi and the local worker arrive as channel events; hub_send replies.",
|
|
5
5
|
"author": {
|
|
6
6
|
"name": "Young Joon Lee",
|