@uzysjung/agent-harness 26.160.1 → 26.162.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.ko.md +77 -70
- package/README.md +64 -57
- package/dist/{chunk-ZXWOY5VO.js → chunk-3G34MEGZ.js} +132 -15
- package/dist/chunk-3G34MEGZ.js.map +1 -0
- package/dist/index.js +8121 -4513
- package/dist/index.js.map +1 -1
- package/dist/trust-tier-drift.js +1 -1
- package/package.json +5 -4
- package/templates/codex/config.toml.template +8 -20
- package/templates/hooks/session-start.sh +13 -3
- package/templates/scripts/check-absence.sh +5 -4
- package/templates/skills/audit-harness-fit/README.md +1 -0
- package/templates/skills/audit-harness-fit/SKILL.md +3 -1
- package/templates/skills/audit-harness-fit/evals/scenarios.yaml +18 -0
- package/templates/skills/audit-harness-fit/references/audit.md +37 -0
- package/templates/skills/audit-service-gaps/SKILL.md +7 -12
- package/templates/skills/compaction-handoff/SKILL.md +5 -6
- package/templates/skills/external-model-consult/SKILL.md +9 -13
- package/templates/skills/gh-issue-workflow/SKILL.md +7 -12
- package/templates/skills/model-orchestration/SKILL.md +57 -17
- package/templates/skills/multi-persona-review/SKILL.md +6 -10
- package/templates/skills/north-star/SKILL.md +7 -13
- package/templates/skills/objective-brief/SKILL.md +7 -12
- package/templates/skills/recurrence-prevention/SKILL.md +7 -11
- package/templates/skills/self-hosted-github-runner/SKILL.md +6 -7
- package/templates/skills/ui-visual-review/SKILL.md +7 -1
- package/dist/chunk-ZXWOY5VO.js.map +0 -1
package/README.ko.md
CHANGED
|
@@ -1,6 +1,8 @@
|
|
|
1
1
|
# uzys-agent-harness
|
|
2
2
|
|
|
3
|
-
AI 코딩 도구에 꼭 필요한 룰·훅·스킬만 남기고 나머지는
|
|
3
|
+
AI 코딩 도구에 꼭 필요한 룰·훅·스킬만 남기고 나머지는 모두 덜어낸다. 위저드를 한 번 실행하면 내 기술 스택에 맞는 검증된 도구 모음을 프로젝트 환경에 설치할 수 있다.
|
|
4
|
+
|
|
5
|
+
**Claude Code** · **Codex** · **OpenCode** · **Antigravity** 에서 사용할 수 있다.
|
|
4
6
|
|
|
5
7
|
[](LICENSE)
|
|
6
8
|
[](https://github.com/uzysjung/uzys-agent-harness/tags)
|
|
@@ -12,121 +14,126 @@ AI 코딩 도구에 꼭 필요한 룰·훅·스킬만 남기고 나머지는 걷
|
|
|
12
14
|
|
|
13
15
|
---
|
|
14
16
|
|
|
15
|
-
##
|
|
17
|
+
## 빠른 시작
|
|
16
18
|
|
|
17
|
-
Node 20 이상이 필요하다.
|
|
19
|
+
Node 20 이상이 필요하다. 프로젝트 폴더에서 다음 명령을 실행한다:
|
|
18
20
|
|
|
19
21
|
```bash
|
|
20
22
|
npx -y @uzysjung/agent-harness
|
|
21
23
|
```
|
|
22
24
|
|
|
23
|
-
위저드가
|
|
25
|
+
위저드가 다음 다섯 단계를 안내한다:
|
|
24
26
|
|
|
25
27
|
```
|
|
26
|
-
1/
|
|
27
|
-
2/
|
|
28
|
-
3/
|
|
29
|
-
4/
|
|
30
|
-
5/
|
|
31
|
-
6/6 Installing
|
|
28
|
+
1/5 Tracks 무엇을 만드는지(스택) — 항목을 미리 체크해 줄 뿐이다
|
|
29
|
+
2/5 CLI claude / codex / opencode / antigravity — 여러 개 가능
|
|
30
|
+
3/5 Install items 트랙에 맞는 항목이 미리 체크되어 있다. 원치 않는 것은 해제
|
|
31
|
+
4/5 Confirm 요약과 함께, 이 선택이 매 세션에 얹는 컨텍스트 크기를 보여 준다
|
|
32
|
+
5/5 Installing
|
|
32
33
|
```
|
|
33
34
|
|
|
34
|
-
|
|
35
|
+
그다음 같은 폴더에서 AI 코딩 도구를 실행하면, 첫 세션부터 룰과 스킬이 바로 적용된다:
|
|
35
36
|
|
|
36
37
|
```bash
|
|
37
38
|
claude # 또는 codex / opencode / agy
|
|
38
39
|
```
|
|
39
40
|
|
|
40
|
-
|
|
41
|
+
**처음 할 일.** 설치가 끝나면 *내 프로젝트*에 대한 정보를 채울 수 있는 빈칸이 포함된 `CLAUDE.md`(또는 `AGENTS.md`) 파일이 생긴다. 에이전트에게 `audit-harness-fit` 스킬을 한 번 실행해 달라고 요청하면, 에이전트가 저장소를 분석해 코드에 기반한 내용으로 빈칸을 채워 준다. 나중에 이 스킬을 다시 실행하면 현재 프로젝트 상황에 하네스가 여전히 잘 맞는지 점검해 준다.
|
|
41
42
|
|
|
42
|
-
|
|
43
|
+
위저드를 띄울 수 없는 환경(CI · 컨테이너 · 스크립트 등)이라면 플래그를 사용해 설치할 수 있다. 필수 플래그는 `install --track <name>` 하나뿐이다 — [비대화형 설치](docs/USAGE.md#non-interactive-install). Claude Code 플러그인을 설치하려면 시스템 PATH 에 `claude` 명령이 있어야 한다. 만약 없다면 경고 메시지만 남기고 플러그인 설치를 건너뛴다.
|
|
43
44
|
|
|
44
|
-
|
|
45
|
+
## 무엇이 들어 있나
|
|
45
46
|
|
|
46
|
-
|
|
|
47
|
-
|
|
48
|
-
|
|
|
49
|
-
|
|
|
47
|
+
| 구성 | 무엇인가 | 에이전트가 언제 읽나 |
|
|
48
|
+
|---|---|---|
|
|
49
|
+
| **룰** | 짧은 파일 6개: git 정책 · 변경 관리 · 문서 · 테스트 · 배포 · CLI 개발. 개발 트랙은 5개를, `tooling` · `full` 은 6개 모두를, 비즈니스 트랙은 어느 프로젝트에나 맞는 3개를 받는다 | 매 세션 |
|
|
50
|
+
| **훅** | CLI 가 스스로 실행하는 스크립트. Claude Code 에는 2개가 들어간다: 하나는 세션을 시작할 때 스펙과 변경 기록을 불러오고, 다른 하나는 `.env` · lock 파일 · 인증서 편집을 막는다 — 하네스에서 "안 된다"고 제한하는 유일한 장치이며, 편집을 막을 때마다 로그를 한 줄 남긴다 | 자동으로 — 세션을 시작할 때, 편집하기 직전에 |
|
|
51
|
+
| **스킬** | 작업에 필요할 때 에이전트가 열어보는 단계별 절차서. 이 저장소에서 관리하는 방법론 스킬과 트랙에 필요한 기술 스택 스킬이 들어 있다(예: `csr-supabase` 의 React · shadcn · Supabase · Postgres) | 필요할 때만 — 한 줄짜리 설명만 항상 띄워 두고, 본문은 사용할 때만 읽는다 |
|
|
52
|
+
| **에이전트** | 메인 에이전트가 작업을 맡기는 조수. 모든 트랙에 독립 검증자 `reviewer` 가 있고, 개발 트랙에는 `implementer` 가 있으며, 그것을 쓰는 트랙에만 `data-analyst` · `strategist` 가 들어간다 | 메인 에이전트가 작업을 위임할 때 |
|
|
53
|
+
| **앵커** | CLI 가 매 세션마다 읽는 작업 원칙 파일. 원래 있던 내 `CLAUDE.md` 파일은 그대로 유지된다 — 하네스는 import 구문 한 줄만 추가할 뿐 나머지는 건드리지 않는다([어느 파일이 누구 것인가](docs/CONTEXT-FILES.md)) | 매 세션 |
|
|
50
54
|
|
|
51
|
-
|
|
55
|
+
모든 트랙에 공통으로 들어가는 방법론 스킬은 4가지다: `north-star` · `objective-brief` · `gh-issue-workflow` · `audit-harness-fit`. 번들 스킬은 `--with` 나 `--without` 뒤에 이름을 적어 추가하거나 뺄 수 있다.
|
|
52
56
|
|
|
53
|
-
|
|
57
|
+
어떤 CLI 에 어떤 구성 요소가 들어갈까:
|
|
54
58
|
|
|
55
|
-
|
|
59
|
+
| CLI | 룰 | 스킬 | 훅 | 플러그인 |
|
|
60
|
+
|---|---|---|---|---|
|
|
61
|
+
| Claude Code | ✓ | ✓ | ✓ | ✓ |
|
|
62
|
+
| Codex | ✓ (`AGENTS.md` 안에 포함) | ✓ | 세션 시작할 때만 | — |
|
|
63
|
+
| OpenCode | ✓ (`AGENTS.md` 안에 포함) | ✓ | — | — |
|
|
64
|
+
| Antigravity | ✓ | ✓ | — | — |
|
|
56
65
|
|
|
57
|
-
|
|
66
|
+
플러그인은 Claude Code 고유의 기능이므로 Claude 전용으로만 제공된다. 스킬과 룰은 같은 원본 파일을 바탕으로 네 가지 CLI 용으로 각각 만들어지므로, 사용하는 도구를 바꿔도 내용은 똑같이 유지된다.
|
|
58
67
|
|
|
59
|
-
|
|
68
|
+
## 트랙 고르기
|
|
60
69
|
|
|
61
|
-
|
|
70
|
+
**트랙**은 프로젝트의 성격에 맞게 구성된 초기 도구 묶음이다. 위저드 3단계에서 필요한 항목을 미리 체크해 주는 역할만 하므로, 원하지 않는 항목은 체크를 해제할 수 있으며 트랙을 여러 개 골라도 괜찮다.
|
|
62
71
|
|
|
63
|
-
-
|
|
72
|
+
- **스택 미정** — `base`: 기본 원칙 · 방법론 스킬 · 테스트 룰을 제공한다. 특정 기술 스택에 종속된 내용은 없다(다른 모든 개발 트랙에도 기본으로 포함된다).
|
|
73
|
+
- **프론트엔드 + 백엔드** — `csr-supabase` · `csr-fastify` · `csr-fastapi` · `ssr-nextjs` · `ssr-htmx`
|
|
74
|
+
- **데이터** — `data`
|
|
75
|
+
- **비즈니스** — `executive` · `project-management` · `growth-marketing`
|
|
76
|
+
- **메타** — `tooling`: 앱 기술 스택이 따로 없는 Bash 나 Markdown 프로젝트용
|
|
77
|
+
- **전부** — `full`
|
|
64
78
|
|
|
65
|
-
|
|
79
|
+
[트랙별 설치 항목 확인하기 →](docs/TRACKS.md)
|
|
66
80
|
|
|
67
|
-
|
|
81
|
+
## 매일 쓰는 명령
|
|
68
82
|
|
|
69
|
-
|
|
83
|
+
| 하고 싶은 일 | 실행할 명령 |
|
|
84
|
+
|---|---|
|
|
85
|
+
| 이 프로젝트에 설치된 항목 확인하기 | `npx -y @uzysjung/agent-harness list` |
|
|
86
|
+
| 최신 버전으로 업데이트하기 | `npx -y @uzysjung/agent-harness update` |
|
|
87
|
+
| 나중에 다른 CLI 추가하기 | `npx -y @uzysjung/agent-harness install --track <내 트랙> --cli <새 CLI>` |
|
|
88
|
+
| 전체 · 특정 CLI · 일부 자산 지우기 | `npx -y @uzysjung/agent-harness uninstall` (`--cli <name>`, `--only <id>`, `--dry-run` 옵션으로 미리 볼 수 있다) |
|
|
70
89
|
|
|
71
|
-
|
|
72
|
-
- **훅** — Claude Code 에 2개. 하나는 세션 시작 때 스펙과 변경 이력을 읽어 오고, 하나는 `.env` · lock 파일 · 인증서 편집을 막는다. 이 둘째 훅이 하네스에서 유일하게 "안 된다"고 말하는 것이고, 막을 때마다 로그에 한 줄을 남긴다. Codex 는 세션 시작 훅만 받는다 — Codex 의 훅 API 는 파일 편집을 가로채지 못한다.
|
|
73
|
-
- **스킬** — 이 리포에서 직접 쓰고 관리하는 방법론 스킬과, 트랙이 부르는 스택 스킬. 4종은 모든 트랙에 포함된다(`north-star` · `objective-brief` · `gh-issue-workflow` · `audit-harness-fit`). 개발 트랙은 방법론 스킬 5종과 사고 대응 런북 하나를 더 받고, 스택이 있는 트랙에는 `frontend-design` 과 스택별 스킬(예: `csr-supabase` 면 React · shadcn · Supabase · Postgres)이 설치된다. 번들 스킬 13종은 `--with` / `--without` 으로 이름을 지정할 수 있다.
|
|
74
|
-
- **에이전트** — 모든 트랙에 독립 `reviewer`, 개발 트랙에 `implementer`, `data-analyst` 와 `strategist` 는 그것을 쓰는 트랙에만.
|
|
75
|
-
- **작업 원칙 앵커** — CLI 가 매 세션 읽는 파일 하나. 내 `CLAUDE.md` 는 내 것으로 남고, 하네스는 import 한 줄만 얹고 나머지는 건드리지 않는다. 어느 파일이 누구 것인지: [docs/CONTEXT-FILES.md](docs/CONTEXT-FILES.md).
|
|
90
|
+
`update` 명령은 하네스가 설치한 파일을 최신으로 갱신하고, 새 버전에 추가된 스킬을 설치하며, 새로운 훅처럼 다시 설치해야 할 항목이 있으면 안내해 준다. 선택하지 않은 CLI 는 새로 설치하지 않는다 — CLI 환경은 설치할 때만 추가되고 `uninstall` 로만 지울 수 있다. `update --only skills` 처럼 특정 묶음만 업데이트할 수도 있다.
|
|
76
91
|
|
|
77
|
-
CLI
|
|
92
|
+
`uninstall` 은 터미널에서 CLI 하나 · 고른 자산 · 전부 셋 중에서 고르게 하고, `--dry-run` 으로 계획만 먼저 볼 수 있다. `.claude/` · `.codex/` · `.opencode/` 는 지우지 않고 `<dir>.backup-<ts>` 로 옮기므로, 거기 직접 둔 파일은 백업에 남는다.
|
|
78
93
|
|
|
79
|
-
|
|
80
|
-
|---|---|---|---|---|
|
|
81
|
-
| Claude Code | ✓ | ✓ | ✓ | ✓ |
|
|
82
|
-
| Codex | ✓ (`AGENTS.md` 안에) | ✓ | 세션 시작만 | — |
|
|
83
|
-
| OpenCode | ✓ (`AGENTS.md` 안에) | ✓ | — | — |
|
|
84
|
-
| Antigravity | ✓ | ✓ | — | — |
|
|
85
|
-
|
|
86
|
-
플러그인은 Claude Code 자체 메커니즘이라 Claude 전용이다. 스킬과 룰은 같은 원본에서 네 CLI 에 맞게 변환되므로 서로 어긋날 일이 없다.
|
|
94
|
+
**기존 프로젝트에 적용해도 안전하다.** 내가 직접 수정한 파일을 교체해야 할 때는 타임스탬프를 붙여 백업본을 남기고 그 경로를 알려 준다. 내가 작성하거나 수정한 파일은 백업 없이 지우지 않으며, 기존에 있던 `.mcp.json` 서버 설정은 덮어쓰지 않고 내용을 병합한다([기존 프로젝트에 설치하기](docs/USAGE.md#installing-into-an-existing-project)).
|
|
87
95
|
|
|
88
|
-
|
|
96
|
+
**현재 프로젝트에만 설치된다.** `~/.codex/` · `~/.opencode/` · `~/.gemini/` 디렉터리나 전역 npm 환경에는 아무것도 설치하지 않는다. 유일한 예외는 Claude Code 플러그인이다 — `claude` CLI 는 플러그인 캐시를 `~/.claude/plugins/` 디렉터리에 저장하며 프로젝트는 메타데이터로 구분하기 때문이다. `.claude/` 디렉터리 밖에는 `.mcp.json`, `.gitignore` 파일에 추가하는 몇 줄(파일이 이미 있을 때), Supabase 트랙의 경우 `.env.example`, 그리고 설치 기록을 남기는 `.uzys-agent-harness/` 디렉터리만 생성한다 — [전체 설치 목록 보기](docs/USAGE.md#what-the-harness-writes).
|
|
89
97
|
|
|
90
|
-
|
|
98
|
+
## 다른 도구를 이미 쓰고 있다면
|
|
91
99
|
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
- **메타** — `tooling`: 앱 스택 없는 Bash · Markdown 프로젝트
|
|
97
|
-
- **전체** — `full`
|
|
100
|
+
| 원하는 작업 | 방법 |
|
|
101
|
+
|---|---|
|
|
102
|
+
| 하네스 전체 설치 — 룰 · 훅 · 에이전트 · 스택에 맞는 스킬 | 위에서 설명한 위저드 사용 |
|
|
103
|
+
| 이 저장소에서 제공하는 특정 스킬 하나만 설치 | `npx skills add uzysjung/uzys-agent-harness --skill <id> -a claude-code` |
|
|
98
104
|
|
|
99
|
-
|
|
105
|
+
이 저장소에서 제공하는 스킬은 [skills CLI](https://github.com/vercel-labs/skills) 를 사용해 하나씩 설치할 수도 있다 — 위저드가 복사하는 파일과 똑같은 파일이며, `references/` 폴더의 참고 자료도 함께 다운로드된다. `npx skills add uzysjung/uzys-agent-harness --list` 명령을 실행하면 스킬 id 목록을 볼 수 있다([자세한 설명](docs/USAGE.md#one-skill-without-the-harness)). [skills.sh/uzysjung/uzys-agent-harness](https://skills.sh/uzysjung/uzys-agent-harness) 에도 배포되어 있다.
|
|
100
106
|
|
|
101
|
-
##
|
|
107
|
+
## 왜 이렇게 만들었나
|
|
102
108
|
|
|
103
|
-
|
|
104
|
-
npx -y @uzysjung/agent-harness list # 이 프로젝트에 무엇이 깔렸나
|
|
105
|
-
npx -y @uzysjung/agent-harness update # 현재 릴리즈로 갱신
|
|
106
|
-
npx -y @uzysjung/agent-harness uninstall # 골라서 제거하거나 전부 제거
|
|
107
|
-
```
|
|
109
|
+
AI 코딩 도구에 룰과 스킬을 무작정 늘린다고 해서 결과가 좋아지지는 않는다. 항상 읽어 들여야 하는 지시문은 매 세션마다 불필요하게 컨텍스트 용량을 차지하며, 이미 잘 알고 있는 내용을 똑똑한 모델에게 다시 가르치면 오히려 작업 속도만 느려진다. 그래서 하네스는 모델이 실수하기 쉬운 지점에만 원칙을 세우고, 모델의 성능이 개선되면 그 원칙마저 다시 덜어내는 방식을 취한다.
|
|
108
110
|
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
111
|
+
| 흔한 방식 | 이 하네스 |
|
|
112
|
+
|---|---|
|
|
113
|
+
| "X 하지 마라" · "항상 Y 하라" 같은 룰을 많이 둔다 | 지침은 *최종 결과가 어때야 하는지(무엇이 참이어야 하는가)*만 알려주고 *어떻게* 달성할지는 모델에게 맡긴다. 우리 룰 44문장 중에서 에이전트의 행동을 바꾼 것이 관측된 문장은 0개였다 — 사고를 잡은 것은 안전장치(게이트), 테스트 코드, 그리고 독립적인 검증자였다 |
|
|
114
|
+
| 금지 사항을 문장으로 길게 적는다 | 되돌릴 수 없는 치명적인 실수는 시스템 장치로 막는다: 훅이 `.env` 파일이나 인증 키 파일 편집을 차단하고, 하네스가 적용을 돕는 GitHub 룰셋이 기본 브랜치를 안전하게 보호한다 |
|
|
115
|
+
| 설정과 자산이 계속 쌓이기만 한다 | "이것이 없으면 에이전트가 느려지거나 실수한다"는 관측이 있는 자산만 남겨둔다. 그렇지 않은 자산은 필요할 때만 불러오는 스킬로 강등되거나 완전히 삭제된다 — 그렇기 때문에 `update` 는 새로 더한 것과 *뺀 것*을 함께 가져온다 |
|
|
116
|
+
| 프로젝트에 맞는지 설치할 때 한 번만 확인한다 | `audit-harness-fit` 스킬이 프로젝트 상태를 계속 점검한다: 불필요한 질문, 반복되는 검사, 서로 모순되는 결정, 더 똑똑해진 모델에게는 필요 없어진 절차를 찾아내서 개선안을 제시한다 |
|
|
117
|
+
| 코드를 작성할 때마다 매번 꼼꼼히 검증한다 | 사용자에게 보여줄 화면이나 기능이 완성되었을 때, 코드를 직접 작성하지 않은 별도의 에이전트가 코드를 실행하고 테스트해 본다 — 무작정 보호 장치를 늘리는 대신, 핵심을 확실하게 지킬 수 있는 가장 가벼운 방법을 선택한다 |
|
|
118
|
+
| 파일 이름이나 함수 단위로 설명을 적는다 | 에이전트가 *내 서비스를 사용할 사용자*의 관점에서 먼저 설명한다(`user-centered-explanation`). 사용자의 승인이 필요한 작업이라면 상황 설명(맥락) → 직면한 문제 → 가능한 선택지 → 에이전트의 추천 순서로 보고한다 |
|
|
112
119
|
|
|
113
|
-
|
|
120
|
+
어떤 자산을 추가할지 혹은 제거할지 결정할 때 던지는 첫 질문은 하나다 — *이 자산이 AI 코딩 도구를 활용해 개발을 더 잘할 수 있도록 돕는가?* 자세한 철학은 [docs/NORTH_STAR.md](docs/NORTH_STAR.md) 에서 확인할 수 있다.
|
|
114
121
|
|
|
115
122
|
## 검증
|
|
116
123
|
|
|
117
|
-
외부
|
|
124
|
+
외부 자산을 **검증된 항목**으로 등록하려면 세 가지 조건을 만족해야 한다: GitHub 스타 1,000개 이상일 것, 보관 처리된(archived) 저장소가 아닐 것, 격리된 환경에서 설치 명령을 실제로 실행해 정상 동작을 확인했을 것. 매달 실행되는 두 개의 CI 작업이 저장소의 스타 수와 설치 경로 유효성을 다시 점검한다. 이 검증은 코드를 한 줄 한 줄 분석하는 보안 감사가 **아니며**, 자산 내용에 포함된 프롬프트 인젝션 공격 여부까지 검사하지는 않는다. npm 과 npx 자산은 버전을 고정해서 사용하고, 플러그인과 스킬 자산은 원본 저장소의 최신 커밋(upstream HEAD)을 가져온다.
|
|
118
125
|
|
|
119
|
-
3단계에서 `★ official`
|
|
126
|
+
위저드 3단계에서 보이는 `★ official` 태그는 Anthropic 공식 마켓플레이스 자산과 하네스에서 자체 제공하는 자산을 의미하며, `⚠ experimental` 태그는 스타가 1,000개 미만인 자산을 뜻한다 — 이 항목들은 기본적으로 체크되어 있지 않으므로 직접 선택해야 설치할 수 있다. 검증을 통과한 자산에는 따로 태그가 붙지 않는다. 이러한 등급은 상태를 알려 주기 위한 용도일 뿐 설치를 막지는 않는다. 설치된 자산은 다른 서드파티 의존성 패키지와 동일한 기준으로 취급한다: [SECURITY.md](SECURITY.md).
|
|
120
127
|
|
|
121
128
|
## 문서
|
|
122
129
|
|
|
123
|
-
- [사용
|
|
124
|
-
- [트랙](docs/TRACKS.md) —
|
|
125
|
-
- [호환성
|
|
126
|
-
- [어느 파일이 누구 것인가](docs/CONTEXT-FILES.md) — `CLAUDE.md` · 앵커 · `AGENTS.md`
|
|
127
|
-
- [워크플로
|
|
128
|
-
- [보안](SECURITY.md) —
|
|
129
|
-
- [
|
|
130
|
+
- [사용 안내](docs/USAGE.md) — 설치 플래그 · 설치 범위 · update · uninstall · CLI 환경별 세부 사항 · 생성되는 파일의 용도와 위치
|
|
131
|
+
- [트랙](docs/TRACKS.md) — 트랙을 선택할 때 미리 체크되는 항목들
|
|
132
|
+
- [호환성 표](docs/COMPATIBILITY.md) — 각 자산의 설치 방식 · 지원하는 CLI · 검증 방법
|
|
133
|
+
- [어느 파일이 누구 것인가](docs/CONTEXT-FILES.md) — `CLAUDE.md` · 앵커 · `AGENTS.md` 등 컨텍스트 관련 파일 설명
|
|
134
|
+
- [워크플로 안내](docs/WORKFLOWS.md) — 선택해서 사용할 수 있는 워크플로 묶음 비교와, 굳이 사용하지 않아도 되는 상황 안내
|
|
135
|
+
- [보안](SECURITY.md) — 검증 과정에서 확인하는 것과 확인하지 않는 것, 취약점 신고 방법
|
|
136
|
+
- [North Star](docs/NORTH_STAR.md) · [결정 기록](docs/decisions/) — 하네스가 왜 이런 철학과 구조를 가지게 되었는지에 대한 배경
|
|
130
137
|
|
|
131
138
|
## License
|
|
132
139
|
|
package/README.md
CHANGED
|
@@ -1,6 +1,8 @@
|
|
|
1
1
|
# uzys-agent-harness
|
|
2
2
|
|
|
3
|
-
Give your AI coding tool the few rules, hooks, and skills it actually needs — and nothing it doesn't. One wizard installs a vetted set for your stack
|
|
3
|
+
Give your AI coding tool the few rules, hooks, and skills it actually needs — and nothing it doesn't. One wizard installs a vetted set for your stack, scoped to your project.
|
|
4
|
+
|
|
5
|
+
Works with **Claude Code** · **Codex** · **OpenCode** · **Antigravity**.
|
|
4
6
|
|
|
5
7
|
[](LICENSE)
|
|
6
8
|
[](https://github.com/uzysjung/uzys-agent-harness/tags)
|
|
@@ -12,67 +14,45 @@ Give your AI coding tool the few rules, hooks, and skills it actually needs —
|
|
|
12
14
|
|
|
13
15
|
---
|
|
14
16
|
|
|
15
|
-
##
|
|
17
|
+
## Quick start
|
|
16
18
|
|
|
17
|
-
You need Node 20 or newer.
|
|
19
|
+
You need Node 20 or newer. Run this in your project folder:
|
|
18
20
|
|
|
19
21
|
```bash
|
|
20
22
|
npx -y @uzysjung/agent-harness
|
|
21
23
|
```
|
|
22
24
|
|
|
23
|
-
The wizard
|
|
25
|
+
The wizard asks five things:
|
|
24
26
|
|
|
25
27
|
```
|
|
26
|
-
1/
|
|
27
|
-
2/
|
|
28
|
-
3/
|
|
29
|
-
4/
|
|
30
|
-
5/
|
|
31
|
-
6/6 Installing
|
|
28
|
+
1/5 Tracks what you are building (a stack) — this only pre-checks items
|
|
29
|
+
2/5 CLI claude / codex / opencode / antigravity — one or more
|
|
30
|
+
3/5 Install items everything is pre-checked for your track; uncheck what you don't want
|
|
31
|
+
4/5 Confirm summary, plus how much context your selection adds to each session
|
|
32
|
+
5/5 Installing
|
|
32
33
|
```
|
|
33
34
|
|
|
34
|
-
Then open your
|
|
35
|
+
Then open your AI coding tool in the same folder. The rules and skills are live from the first session:
|
|
35
36
|
|
|
36
37
|
```bash
|
|
37
38
|
claude # or codex / opencode / agy
|
|
38
39
|
```
|
|
39
40
|
|
|
40
|
-
**
|
|
41
|
-
|
|
42
|
-
The wizard needs a terminal. For CI, containers, or onboarding scripts, use the flag form: `install --track <name>` is the only required flag — see [non-interactive install](docs/USAGE.md#non-interactive-install).
|
|
43
|
-
|
|
44
|
-
### Two ways to install
|
|
45
|
-
|
|
46
|
-
| You want… | Do this |
|
|
47
|
-
|---|---|
|
|
48
|
-
| The harness — rules, hooks, agents, and the skills your stack calls for, curated by track | The wizard above |
|
|
49
|
-
| One skill, nothing else — no harness, no track | `npx skills add uzysjung/uzys-agent-harness --skill <id> -a claude-code` |
|
|
50
|
-
|
|
51
|
-
Every skill this repo ships is installable on its own with the [skills CLI](https://github.com/vercel-labs/skills); `npx skills add uzysjung/uzys-agent-harness --list` shows the ids. You get the same files the installer copies, `references/` included. Re-run the command to refresh ([details](docs/USAGE.md#one-skill-without-the-harness)). They are listed on [skills.sh/uzysjung/uzys-agent-harness](https://skills.sh/uzysjung/uzys-agent-harness) as well.
|
|
52
|
-
|
|
53
|
-
## Philosophy — keep only the frame the model needs
|
|
41
|
+
**First thing to do.** The install leaves a `CLAUDE.md` (or `AGENTS.md`) with blank sections about *your* project. Ask your agent to run the `audit-harness-fit` skill once — it reads the repository and fills those sections from the code. Run it again later to check whether the harness still fits.
|
|
54
42
|
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
- **"Don't do X" and "always do Y" degrade today's models.** A pile of prohibitions and mandates removes the model's room to judge. The moment a situation differs slightly from the rule, the model gets stuck or works around it. Once guards start guarding other guards, development slows down while quality stays flat. So the harness writes guidance as *what must be true* — the outcome, the invariant, the boundary — and leaves *how* to the model. Prohibitions are kept for irreversible damage — secrets and shared history — and even those are enforced by a mechanism rather than a sentence: a hook blocks edits to `.env` and key files, and a GitHub ruleset the harness helps you apply protects the default branch. This direction came from measurement, not taste: of the 44 sentences in our own rules, the number observed to change an agent's behaviour was zero. What actually caught incidents was gates, tests, and an independent reviewer.
|
|
58
|
-
|
|
59
|
-
- **Skills improve as models improve.** Each asset is kept only when there is an observation behind it: "without this, the agent is measurably slower or wrong." Guidance without that evidence is removed from the always-loaded set — it becomes a skill that loads only when needed, or it is retired. So `update` brings you the current judgement — what was added, and what was cut — not just more.
|
|
60
|
-
|
|
61
|
-
- **`audit-harness-fit` keeps checking.** Run once after install and it fills your project context from the repository. Run later and it finds needless questions, repeated checks, decisions that contradict each other, and procedures that a better model no longer needs — then proposes the edit. Whether the harness fits your project is a question you keep asking, not one you settle at install time.
|
|
62
|
-
|
|
63
|
-
- **Speed and quality are not a trade-off.** Irreversible damage is blocked by a hook, repeated mistakes are caught by a short rule, and everything else is left to the model. Verification is not "everything, every time": it runs when a user-facing scene is complete, by a separate agent that did not write the code and actually executes it. The goal is the lightest protection that reliably keeps what matters, not more guards.
|
|
64
|
-
|
|
65
|
-
- **Think and explain from the customer's side of the service you are building.** The agent describes a problem, a change, or a choice first as what the user does and sees, not as file names and functions (`user-centered-explanation`). When a decision needs your approval, it arrives as context → problem → options → recommendation, so you can decide without re-reading the conversation.
|
|
66
|
-
|
|
67
|
-
Those five are the direction, and the first question for any asset — in or out — is *does this help a person build better with an AI coding tool?* The long form is [docs/NORTH_STAR.md](docs/NORTH_STAR.md).
|
|
43
|
+
No terminal for the wizard (CI, containers, scripts)? Use flags: `install --track <name>` is the only required one — see [non-interactive install](docs/USAGE.md#non-interactive-install). Claude Code plugins need the `claude` command on your PATH; without it they are skipped with a warning.
|
|
68
44
|
|
|
69
45
|
## What you get
|
|
70
46
|
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
47
|
+
| Piece | What it is | When your agent reads it |
|
|
48
|
+
|---|---|---|
|
|
49
|
+
| **Rules** | Six short files: git policy, change management, documentation, testing, shipping, CLI development. Dev tracks get five; `tooling` and `full` add the sixth; business tracks get the three that apply to any project | Every session |
|
|
50
|
+
| **Hooks** | Scripts your CLI runs on its own. Two, on Claude Code: one loads your spec and change log at session start; one blocks edits to `.env`, lock files, and certificates — the only thing in the harness that says "no", and it logs one line each time | Automatically, at session start or before an edit |
|
|
51
|
+
| **Skills** | Step-by-step playbooks the agent opens when a task calls for them — the method skills written in this repo, plus the stack skills your track needs (for example React, shadcn, Supabase, Postgres on `csr-supabase`) | Only when relevant — a one-line description stays loaded, the body loads on use |
|
|
52
|
+
| **Agents** | Helpers the main agent can hand work to: an independent `reviewer` on every track, `implementer` on dev tracks, `data-analyst` and `strategist` on the tracks that use them | When the main agent delegates |
|
|
53
|
+
| **Anchor** | One working-principles file your CLI reads every session. Your own `CLAUDE.md` stays yours — the harness adds one import line and never touches the rest ([which file is whose](docs/CONTEXT-FILES.md)) | Every session |
|
|
54
|
+
|
|
55
|
+
Four method skills go to every track: `north-star`, `objective-brief`, `gh-issue-workflow`, `audit-harness-fit`. Bundled skills can be added or dropped by name with `--with` / `--without`.
|
|
76
56
|
|
|
77
57
|
What reaches which CLI:
|
|
78
58
|
|
|
@@ -83,40 +63,67 @@ What reaches which CLI:
|
|
|
83
63
|
| OpenCode | ✓ (in `AGENTS.md`) | ✓ | — | — |
|
|
84
64
|
| Antigravity | ✓ | ✓ | — | — |
|
|
85
65
|
|
|
86
|
-
Plugins are Claude Code's own mechanism, so they are Claude-only
|
|
66
|
+
Plugins are Claude Code's own mechanism, so they are Claude-only. Skills and rules render for all four from the same source, so they stay consistent across tools.
|
|
87
67
|
|
|
88
|
-
##
|
|
68
|
+
## Pick a track
|
|
89
69
|
|
|
90
|
-
|
|
70
|
+
A **track** is a starting set for what you are building. It only pre-checks items at step 3 — you can uncheck anything, and pick more than one track.
|
|
91
71
|
|
|
92
|
-
- **No stack yet** — `base`: principles, method skills, and testing rules; nothing stack-specific
|
|
72
|
+
- **No stack yet** — `base`: principles, method skills, and testing rules; nothing stack-specific (every dev track already includes it)
|
|
93
73
|
- **Frontend + backend** — `csr-supabase` · `csr-fastify` · `csr-fastapi` · `ssr-nextjs` · `ssr-htmx`
|
|
94
74
|
- **Data** — `data`
|
|
95
75
|
- **Business** — `executive` · `project-management` · `growth-marketing`
|
|
96
76
|
- **Meta** — `tooling`: Bash and Markdown projects with no app stack
|
|
97
77
|
- **Everything** — `full`
|
|
98
78
|
|
|
99
|
-
|
|
79
|
+
[What each track installs →](docs/TRACKS.md)
|
|
100
80
|
|
|
101
81
|
## Day to day
|
|
102
82
|
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
npx -y @uzysjung/agent-harness
|
|
106
|
-
npx -y @uzysjung/agent-harness
|
|
107
|
-
|
|
83
|
+
| You want to… | Run |
|
|
84
|
+
|---|---|
|
|
85
|
+
| See what this project got | `npx -y @uzysjung/agent-harness list` |
|
|
86
|
+
| Bring it to the current release | `npx -y @uzysjung/agent-harness update` |
|
|
87
|
+
| Add another CLI later | `npx -y @uzysjung/agent-harness install --track <your track> --cli <new cli>` |
|
|
88
|
+
| Remove everything, one CLI, or single assets | `npx -y @uzysjung/agent-harness uninstall` (`--cli <name>` · `--only <id>` · `--dry-run` to preview) |
|
|
89
|
+
|
|
90
|
+
`update` refreshes what the harness installed, adds skills a newer release introduced, and tells you when something (a new hook) needs a reinstall instead. It never installs a CLI you did not choose — installing adds a CLI, and only `uninstall` takes one away. `update --only skills` limits it to one group.
|
|
91
|
+
|
|
92
|
+
In a terminal, `uninstall` offers three choices — one CLI, selected assets, or everything — and `--dry-run` shows the plan first. It never deletes `.claude/`, `.codex/`, or `.opencode/`: each is moved aside as `<dir>.backup-<ts>`, so files you put there yourself stay in the backup.
|
|
93
|
+
|
|
94
|
+
**Safe on an existing project.** Before replacing a file you edited, the harness writes a timestamped backup next to it and prints the path. Nothing you wrote or edited is deleted without a backup beside it, and your existing `.mcp.json` servers are merged, not replaced ([installing into an existing project](docs/USAGE.md#installing-into-an-existing-project)).
|
|
95
|
+
|
|
96
|
+
**Your project only.** Nothing goes to `~/.codex/`, `~/.opencode/`, `~/.gemini/`, or global npm. Claude Code plugins are the one exception: the `claude` CLI keeps its plugin cache under `~/.claude/plugins/` and isolates projects by metadata. Besides `.claude/`, install writes `.mcp.json`, a few `.gitignore` lines (when that file exists), an `.env.example` on Supabase tracks, and its own record at `.uzys-agent-harness/` — [the full list](docs/USAGE.md#what-the-harness-writes).
|
|
108
97
|
|
|
109
|
-
|
|
98
|
+
## Already using another tool?
|
|
110
99
|
|
|
111
|
-
|
|
100
|
+
| You want… | Use |
|
|
101
|
+
|---|---|
|
|
102
|
+
| The whole harness — rules, hooks, agents, and the skills your stack calls for | The wizard above |
|
|
103
|
+
| One skill from this repo, nothing else | `npx skills add uzysjung/uzys-agent-harness --skill <id> -a claude-code` |
|
|
104
|
+
|
|
105
|
+
Every skill this repo ships is installable on its own with the [skills CLI](https://github.com/vercel-labs/skills) — the same files the installer copies, `references/` included. `npx skills add uzysjung/uzys-agent-harness --list` shows the ids ([details](docs/USAGE.md#one-skill-without-the-harness)); they are also listed on [skills.sh/uzysjung/uzys-agent-harness](https://skills.sh/uzysjung/uzys-agent-harness).
|
|
106
|
+
|
|
107
|
+
## Why it is built this way
|
|
108
|
+
|
|
109
|
+
Piling rules and skills onto an AI coding tool does not make it better: every always-loaded instruction costs context in every session, and telling a capable model how to do what it already does well only slows it down. So the harness keeps a principle only where the model is likely to slip, and takes it back out as models improve.
|
|
110
|
+
|
|
111
|
+
| Common approach | This harness |
|
|
112
|
+
|---|---|
|
|
113
|
+
| Many "don't do X" and "always do Y" rules | Guidance states *what must be true* and leaves *how* to the model. Of the 44 sentences in our own rules, the number observed to change an agent's behaviour was zero — gates, tests, and an independent reviewer caught the incidents |
|
|
114
|
+
| Prohibitions written as sentences | Irreversible damage is blocked by a mechanism: a hook stops edits to `.env` and key files, and a GitHub ruleset the harness helps you apply protects the default branch |
|
|
115
|
+
| Assets pile up and stay | An asset stays only with an observation that the agent is slower or wrong without it; otherwise it moves to an on-demand skill or is retired — so `update` brings what was added *and* what was cut |
|
|
116
|
+
| Fit is decided once, at install | `audit-harness-fit` keeps checking: needless questions, repeated checks, conflicting decisions, procedures a better model no longer needs — then proposes the edit |
|
|
117
|
+
| Verify everything, every time | Verification runs when a user-facing scene is complete, by a separate agent that did not write the code and actually runs it — the lightest protection that reliably keeps what matters, not more guards |
|
|
118
|
+
| Explanations in file names and functions | The agent explains from your user's side first (`user-centered-explanation`); approval requests arrive as context → problem → options → recommendation |
|
|
112
119
|
|
|
113
|
-
|
|
120
|
+
The first question for any asset, in or out, is *does this help a person build better with an AI coding tool?* The long form is [docs/NORTH_STAR.md](docs/NORTH_STAR.md).
|
|
114
121
|
|
|
115
122
|
## Vetting
|
|
116
123
|
|
|
117
124
|
An external asset is **vetted** when it has at least 1,000 GitHub stars, is not archived, and its install command has been run and checked in an isolated environment. Two monthly CI jobs re-check the stars and the install path. Vetting is **not** a line-by-line security audit and does not scan asset contents for prompt injection. npm and npx assets are pinned to a version; plugin and skill assets resolve to upstream HEAD.
|
|
118
125
|
|
|
119
|
-
At step 3, `★ official` marks Anthropic-official marketplaces and this harness's own assets, `⚠ experimental` marks assets under 1,000 stars — never pre-checked, added only by you. Vetted assets carry no badge. Tiers inform; they never block. Treat installed assets like any other third-party dependency: [SECURITY.md](SECURITY.md).
|
|
126
|
+
At step 3, `★ official` marks Anthropic-official marketplaces and this harness's own assets, and `⚠ experimental` marks assets under 1,000 stars — never pre-checked, added only by you. Vetted assets carry no badge. Tiers inform; they never block. Treat installed assets like any other third-party dependency: [SECURITY.md](SECURITY.md).
|
|
120
127
|
|
|
121
128
|
## Docs
|
|
122
129
|
|