agent-bios 0.19.2 → 0.19.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/DEPENDENCIES.md +32 -4
- package/README.md +54 -4
- package/claude/guides/cli-multi-model-workflow.md +1 -1
- package/claude/guides/coding-staged-workflow.md +17 -0
- package/claude/guides/llm-capability-boundary.md +7 -1
- package/claude/guides/tooling-gotchas.md +20 -1
- package/claude/guides/ui-design/visual-direction.md +88 -0
- package/claude/guides/ui-design.md +90 -0
- package/claude/guides/verification-discipline.md +10 -1
- package/claude/hooks/tooling-gotchas-hook.py +41 -0
- package/codex/guides/cli-multi-model-workflow.md +1 -1
- package/codex/guides/coding-staged-workflow.md +17 -0
- package/codex/guides/llm-capability-boundary.md +7 -1
- package/codex/guides/tooling-gotchas.md +20 -1
- package/codex/guides/ui-design/visual-direction.md +88 -0
- package/codex/guides/ui-design.md +90 -0
- package/codex/guides/verification-discipline.md +10 -1
- package/compose/app_bridge/scripts/bridge.py +23 -6
- package/compose/app_desktop/server.py +250 -0
- package/compose/domains.json +1 -0
- package/compose/host_platform.py +121 -0
- package/compose/instructions-state.py +5 -2
- package/compose/instructions_app.py +281 -48
- package/compose/instructions_catalog.py +2 -2
- package/compose/instructions_import.py +19 -11
- package/compose/instructions_install.py +77 -11
- package/compose/instructions_session.py +6 -6
- package/compose/instructions_setup.py +41 -4
- package/compose/instructions_setup_cli.py +31 -10
- package/compose/instructions_setup_i18n.py +3 -0
- package/compose/instructions_store.py +6 -6
- package/compose/instructions_transaction.py +8 -6
- package/compose/instructions_ui_runtime.py +2 -1
- package/compose/native_cli.py +52 -0
- package/compose/runtime_entry.py +58 -0
- package/compose/windows_deploy.py +719 -0
- package/docs/instructions.md +1 -0
- package/docs/releases/0.19.3.md +107 -0
- package/docs/session-model.md +8 -0
- package/docs/setup.md +36 -0
- package/docs/windows.md +99 -0
- package/install.sh +1 -0
- package/launch/agent-launch.py +12 -4
- package/launch/agent-launch.zsh +11 -1
- package/package.json +8 -3
- package/provenance.json +1 -1
package/DEPENDENCIES.md
CHANGED
|
@@ -113,7 +113,7 @@ execution evidence have narrower scope:
|
|
|
113
113
|
by the current `--version` reports. Authenticated instructions-agent execution and post-fix
|
|
114
114
|
authenticated resume remain unverified.
|
|
115
115
|
|
|
116
|
-
## Codex app
|
|
116
|
+
## Codex app
|
|
117
117
|
|
|
118
118
|
The Codex desktop app is a separate host from the Codex CLI. Its optional bridge needs
|
|
119
119
|
native skill discovery for `~/.agents/skills/agent-bios`, the explicit-invocation policy in
|
|
@@ -146,6 +146,30 @@ requires explicit selection, enables no native hooks or agents, and Off cannot r
|
|
|
146
146
|
previously returned text. Registration is off by default and owns only its discovery
|
|
147
147
|
link, preserving global instruction files and foreign entries.
|
|
148
148
|
|
|
149
|
+
## Claude Desktop
|
|
150
|
+
|
|
151
|
+
Claude Desktop has no launch to project into and no skill root an installer can reach,
|
|
152
|
+
so it pulls through a local MCP server. `agent-bios app desktop` writes a `.mcpb` bundle
|
|
153
|
+
under the private state root; the user installs it through Desktop's own dialog, and
|
|
154
|
+
agent-bios writes nothing into Desktop's directories or `claude_desktop_config.json`.
|
|
155
|
+
The bundle's server (`compose/app_desktop/server.py`) is standard-library Python speaking
|
|
156
|
+
MCP 2025-11-25 over stdio, with no network service. Desktop resolves a bare `python3`
|
|
157
|
+
through the user's login-shell `PATH` and ships no Python of its own, so the manifest names
|
|
158
|
+
the absolute interpreter that generated it; that interpreter must stay installed, and the
|
|
159
|
+
bundle is regenerated after it changes. The server resolves the confirmed private release
|
|
160
|
+
on every call.
|
|
161
|
+
|
|
162
|
+
Desktop sends no conversation identity with a tool call, so `app session --host
|
|
163
|
+
claude-desktop` mints one on preview or use. Measured on Desktop 2.16120.0 (macOS,
|
|
164
|
+
2026-09-30): the bundle route, protocol 2025-11-25, a 140,666-byte result delivered
|
|
165
|
+
inline, and delivery with end-marker confirmation in both chat and the Code tab. The Code
|
|
166
|
+
tab runs Claude Code on the user's own `~/.claude`, but its extension server still starts
|
|
167
|
+
outside the session's working directory, so project scope stays excluded there too.
|
|
168
|
+
Windows and `roots/list` are unverified. Local tests do not establish that a given Desktop
|
|
169
|
+
version loads the bundle.
|
|
170
|
+
|
|
171
|
+
## Instruction import
|
|
172
|
+
|
|
149
173
|
Local instruction import also needs no model SDK or parser framework. Standard-library
|
|
150
174
|
code discovers fixed instruction filenames at known global and explicit project roots,
|
|
151
175
|
captures redacted evidence through `learn/redact.py`, and validates source digests,
|
|
@@ -189,9 +213,12 @@ Agent-bios does not install credentials or infer a model account from dependency
|
|
|
189
213
|
`codex-helm.sh` and `claude-run.sh` from the selected package. Their command paths,
|
|
190
214
|
binding, reach and fallback are declared in the launch contract. An unavailable or
|
|
191
215
|
unauthenticated opposite-family route is reported, not credited as a completed review.
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
216
|
+
<!-- mcp-inventory:start -->
|
|
217
|
+
<!-- facts: {"bundlers": {"compose/app_desktop/server.py": ["compose/instructions_app.py"]}, "capabilities": [], "servers": ["compose/app_desktop/server.py"]} -->
|
|
218
|
+
- **MCP servers** — derived from source by `gates/check-mcp-inventory.py`.
|
|
219
|
+
Shipped servers: `compose/app_desktop/server.py` (bundled by `compose/instructions_app.py`).
|
|
220
|
+
Launch capabilities offering `mcp-stdio-v1`: none, so no shipped review method requires MCP; the launcher registers a user-specific server only through a selected capability that declares it.
|
|
221
|
+
<!-- mcp-inventory:end -->
|
|
195
222
|
- **spreadsheet-processing** — an optional skill referenced by the spreadsheet rule
|
|
196
223
|
when those instructions are selected. If unavailable, its inline plain-tools/code and real
|
|
197
224
|
spreadsheet-engine validation fallback applies.
|
|
@@ -220,6 +247,7 @@ Inspect the installed private state using this checkout's runtime:
|
|
|
220
247
|
bash install.sh verify
|
|
221
248
|
bash install.sh instructions status --json
|
|
222
249
|
bash install.sh app status --json
|
|
250
|
+
bash install.sh app desktop --dry-run --json
|
|
223
251
|
```
|
|
224
252
|
|
|
225
253
|
Provision packages only when explicitly requested for learning validation, compatibility
|
package/README.md
CHANGED
|
@@ -5,7 +5,7 @@
|
|
|
5
5
|
Build an instruction library you can inspect, edit, and reuse. Choose what each
|
|
6
6
|
CLI session or Codex app task uses, while preserving your existing global instruction files.
|
|
7
7
|
|
|
8
|
-
[0.19.
|
|
8
|
+
[0.19.3 release notes](docs/releases/0.19.3.md)
|
|
9
9
|
|
|
10
10
|
## Purpose
|
|
11
11
|
|
|
@@ -16,6 +16,10 @@ their roles within the standards and context of their team, industry, and organi
|
|
|
16
16
|
|
|
17
17
|
A team is the basic unit for selecting, adopting, and sharing a work environment.
|
|
18
18
|
|
|
19
|
+
The goal is efficient work and continuity. Shared environments provide reusable
|
|
20
|
+
defaults that workers can adapt to their repository and personal working needs;
|
|
21
|
+
they are not a mechanism for enforcing uniform working methods.
|
|
22
|
+
|
|
19
23
|
Workers should be able to select an environment appropriate to their role and
|
|
20
24
|
team, collaborate from shared standards and decision context, and continue
|
|
21
25
|
the work when a worker, model, session, or device changes. Shared context does
|
|
@@ -54,6 +58,16 @@ accepts that older command name. See [compatibility](docs/instructions-compatibi
|
|
|
54
58
|
This is an instruction and launch layer, not a replacement for either host CLI.
|
|
55
59
|
It does not train the model or guarantee that the model follows every instruction.
|
|
56
60
|
|
|
61
|
+
## Windows native preview
|
|
62
|
+
|
|
63
|
+
A Windows x64 installer with a bundled Python runtime is built by the Windows
|
|
64
|
+
workflow. It provides `agent-bios.exe` and `agent-launch.exe` without npm or WSL.
|
|
65
|
+
The same workflow also qualifies a script distribution that uses an approved or
|
|
66
|
+
provisioned CPython 3.13 with signed PowerShell commands and no custom EXE.
|
|
67
|
+
Each Windows script release page carries the one-line PowerShell command for
|
|
68
|
+
that release, and the fixed address below serves an explicitly promoted one.
|
|
69
|
+
See [Windows installation and validation limits](docs/windows.md).
|
|
70
|
+
|
|
57
71
|
## Quick start
|
|
58
72
|
|
|
59
73
|
Use the terminal installer or set up through a Codex app conversation. Both offer
|
|
@@ -62,11 +76,26 @@ See [setup and prerequisites](docs/setup.md).
|
|
|
62
76
|
|
|
63
77
|
### In a terminal
|
|
64
78
|
|
|
65
|
-
|
|
66
|
-
|
|
79
|
+
One line per platform. Each downloads from the project's installation page, which
|
|
80
|
+
serves an explicitly promoted release rather than a moving latest.
|
|
81
|
+
|
|
82
|
+
On **Windows** — from PowerShell, the Command Prompt or the Run box. The current
|
|
83
|
+
release is an unsigned preview, which is what the trailing flag accepts:
|
|
84
|
+
|
|
85
|
+
```text
|
|
86
|
+
powershell -NoProfile -ExecutionPolicy Bypass -Command "iwr 'https://kangminlee-maker.github.io/agent-bios/install.ps1' -UseBasicParsing -OutFile (Join-Path ([IO.Path]::GetTempPath()) 'agent-bios-install.ps1') -ErrorAction Stop; & (Join-Path ([IO.Path]::GetTempPath()) 'agent-bios-install.ps1') -AcceptUnsignedPreview"
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
On **macOS or Linux**, where **Python 3.11+** and **Node.js 18+ with npm** must
|
|
90
|
+
already be present:
|
|
91
|
+
|
|
92
|
+
```text
|
|
93
|
+
curl -fsSL https://kangminlee-maker.github.io/agent-bios/install.sh | bash
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
Then deploy the work environment:
|
|
67
97
|
|
|
68
98
|
```bash
|
|
69
|
-
npm install -g agent-bios@0.19.0
|
|
70
99
|
agent-bios install
|
|
71
100
|
```
|
|
72
101
|
|
|
@@ -83,6 +112,12 @@ your project:
|
|
|
83
112
|
|
|
84
113
|
Use `codex` instead of `claude` for a Codex CLI session.
|
|
85
114
|
|
|
115
|
+
The full path is deliberate: `agent-bios install` deploys the launcher to
|
|
116
|
+
`~/.local/bin`, which your shell may not search. Add that directory to `PATH` to
|
|
117
|
+
type `agent-launch` directly, and run `agent-bios shell restore` to make a bare
|
|
118
|
+
`claude` or `codex` open the launcher (`agent-bios shell remove` undoes it).
|
|
119
|
+
Installation reports both when they are not in place.
|
|
120
|
+
|
|
86
121
|
1. Choose a Builder preset or **Custom**.
|
|
87
122
|
2. Review the model, review setup, and permissions. **Some presets request
|
|
88
123
|
permission bypass**; select settings appropriate for your project.
|
|
@@ -132,6 +167,21 @@ To add instructions to a task, explicitly ask it to use your chosen instructions
|
|
|
132
167
|
Installation and opening Instructions Studio do not activate task context.
|
|
133
168
|
[App use, off, and personal instruction import →](docs/setup.md#use-instructions-in-a-codex-app-task)
|
|
134
169
|
|
|
170
|
+
### In Claude Desktop
|
|
171
|
+
|
|
172
|
+
Claude Desktop opens conversations without a launcher, so agent-bios reaches it as a
|
|
173
|
+
local extension you install once. After installing agent-bios on the same Mac, run:
|
|
174
|
+
|
|
175
|
+
```bash
|
|
176
|
+
agent-bios app desktop
|
|
177
|
+
```
|
|
178
|
+
|
|
179
|
+
Open the `.mcpb` file it reports with Claude Desktop and confirm the installation.
|
|
180
|
+
In a conversation, ask for your agent-bios instructions: the extension returns your
|
|
181
|
+
saved selection as context for that conversation only. Nothing is added until you ask,
|
|
182
|
+
and a conversation already in progress receives nothing on its own.
|
|
183
|
+
[Desktop use, confirmation and limits →](docs/setup.md#use-instructions-in-claude-desktop)
|
|
184
|
+
|
|
135
185
|
## Your instruction library
|
|
136
186
|
|
|
137
187
|
Open **Instructions Studio** in the launcher, or run `agent-bios instructions`.
|
|
@@ -67,7 +67,7 @@ Delegate execution, not decisions. A unit is delegable only when it is decision-
|
|
|
67
67
|
- Use a resident teammate only for dependent slices in one burst. Verify that the CLI preserves its model and context; resume-after-completion may silently change both. Retire after the burst or cache TTL, and persist durable knowledge in files.
|
|
68
68
|
- After a discard or direction change, respawn once a routine round costs about as much as a fresh slice. Recover unique in-flight state to files first.
|
|
69
69
|
- Redirects to busy workers may queue rather than preempt. Check artifacts before destructive redirects, phrase them conditionally, and stop an actively harmful worker by scoped PID/worktree authority.
|
|
70
|
-
- Idle/progress notifications are hypotheses; verify repo artifacts before re-dispatch. An idle signal is liveness decoupled from the report: a subagent can go idle without ever delivering its result, so idle-without-report is not done — request the report explicitly rather than waiting. Cross-reset state belongs in files, not task boards or transcripts. When polling concurrent async jobs, pin the exact id/handle received at dispatch — a "latest" convenience selector can silently point at a sibling job and return plausible-but-wrong results.
|
|
70
|
+
- Idle/progress notifications are hypotheses; verify repo artifacts before re-dispatch. An idle signal is liveness decoupled from the report: a subagent can go idle without ever delivering its result, so idle-without-report is not done — request the report explicitly rather than waiting. A delivered report can also arrive cut off with no marker — mid-table or mid-finding — and asking for it again inline truncates the same way. When a report is longer than a short summary or must outlive the turn, name an output file in the brief: the worker writes the full report there and returns the path and a one-line summary. Treat a report that ends mid-item as incomplete and switch to the file rather than re-requesting inline. Cross-reset state belongs in files, not task boards or transcripts. When polling concurrent async jobs, pin the exact id/handle received at dispatch — a "latest" convenience selector can silently point at a sibling job and return plausible-but-wrong results.
|
|
71
71
|
- Give reviewers/subagents a read-only diff, snapshot, or isolated worktree — not the live tree the main is editing — and forbid destructive git ops (checkout --, reset --hard, stash, clean) on any tree with uncommitted work; re-verify tree integrity before trusting results produced mid-edit.
|
|
72
72
|
- Codex `spawn_agent` decides how much of the parent crosses: `fork_turns` defaults to `all`, and takes `none` or a turn count. A `SubagentStart` hook there receives `agent_type` and may return `continue: false`, so a tier rule can be enforced rather than stated.
|
|
73
73
|
- No per-spawn instructions suppression exists on either host: the subagent definition carries model and effort, not scope. Excluding the standing instructions is a process-level act — `claude --setting-sources ''`, or `CODEX_HOME` pointed at a directory holding only `auth.json` — and it removes the tier definitions with them, so a reader without those instructions and a pinned tier cannot come from one process. An emptied `CODEX_HOME` without `auth.json` fails 401; skills still load.
|
|
@@ -25,6 +25,11 @@ change introduced. The stages below are for work that outgrows that sentence.
|
|
|
25
25
|
|
|
26
26
|
When the user asks to "설계" or design, stay in design mode. Focus on high-level design and implementation-process design, then present the plan, tradeoffs, review gates, and implementation trigger. Move to implementation after the user asks to implement or approves the plan.
|
|
27
27
|
|
|
28
|
+
For operational user-interface flows, information layout, visual hierarchy or
|
|
29
|
+
interaction design, use `${CLAUDE_CONFIG_DIR:-$HOME/.claude}/guides/ui-design.md`.
|
|
30
|
+
Keep bounded corrections on the lightweight path; this pointer does not require
|
|
31
|
+
a full UI redesign.
|
|
32
|
+
|
|
28
33
|
## When To Use
|
|
29
34
|
|
|
30
35
|
- Use this workflow for architecture changes, new features, cross-module behavior changes, ontology changes, review-driven fixes, or work that affects user-visible behavior, authority, lifecycle, validation, failure handling, or roadmap commitments.
|
|
@@ -88,6 +93,18 @@ blast radius. Re-mapping a level obliges enumerating every reader of that field,
|
|
|
88
93
|
commonly gates shipping, repair, retry, and display at once. A deferred defect is pinned as a
|
|
89
94
|
strict expected failure, never a silent pass.
|
|
90
95
|
|
|
96
|
+
**Choose the failure posture by what the next step reads.** Across these instructions, an
|
|
97
|
+
unqualified instruction to fail loud or fail clearly means surface the problem and reject the bad
|
|
98
|
+
local result; it does not by itself decide whether a whole production run stops. In development,
|
|
99
|
+
tests, and gates, stop and name the problem, because a silent pass hides a defect. On a production
|
|
100
|
+
runtime path, warn loudly and halt only when continuing would contaminate what the next step
|
|
101
|
+
reads: an input is stale or marked not reusable (partial, failed, blocked); a contract-failing
|
|
102
|
+
value is about to be recorded as valid; or an external write has an unknown outcome. Otherwise —
|
|
103
|
+
an isolated item failure, a tripped breaker, an exhausted budget — warn and continue: work already
|
|
104
|
+
done stays valid, work not done is recorded as not done, and the next run picks it up. This
|
|
105
|
+
default assumes recoverable state kept in artifacts rather than only in the running process. The
|
|
106
|
+
system's owner can redefine it. In no stage does a problem pass silently.
|
|
107
|
+
|
|
91
108
|
## Review Loop
|
|
92
109
|
|
|
93
110
|
- At each stage, run review loops as appropriate: self review, subagent review when available, and structured multi-lens review when the repository or domain supports one (concrete tool: Environment Binding below).
|
|
@@ -110,7 +110,13 @@ Important levers:
|
|
|
110
110
|
grounding checks, provenance checks, citation checks, static checks, E2E
|
|
111
111
|
checks, and semantic quality gates.
|
|
112
112
|
- Retry/fail policy: retry transient generation failures; fail clearly when the
|
|
113
|
-
available route cannot enforce the required contract.
|
|
113
|
+
available route cannot enforce the required contract. A contract-failing item
|
|
114
|
+
is rejected loudly and never recorded as a valid result; record its failed or
|
|
115
|
+
not-done status in artifact state. On a production path the whole run halts
|
|
116
|
+
only when continuing would contaminate what the next step reads — a stale or
|
|
117
|
+
not-reusable input, a contract-failing value about to be recorded as valid,
|
|
118
|
+
or an external write with an unknown outcome; otherwise it warns loudly and
|
|
119
|
+
continues, and the next run picks up the work not done.
|
|
114
120
|
- Observability: prompt packet snapshot, model/provider version, schema hash,
|
|
115
121
|
source snapshot, validator decision, retry reason, and artifact lineage.
|
|
116
122
|
|
|
@@ -256,7 +256,26 @@ depends on it, pin it explicitly instead of trusting the environment.
|
|
|
256
256
|
and check for a front-side cache/CDN separately. Aim the probe at an
|
|
257
257
|
in-unit sentinel the app answers without credentials: a denial the app
|
|
258
258
|
produces anyway passes with the control off, and a redirect into it is a
|
|
259
|
-
bypass.
|
|
259
|
+
bypass. The address an IP-based rule compares is the one selected at its
|
|
260
|
+
enforcement point, and configuration alone may not establish which: behind
|
|
261
|
+
a proxy or CDN it can be the client address the configured forwarded-header
|
|
262
|
+
trust chain hands the app, and a managed platform can route particular
|
|
263
|
+
destinations outside an otherwise configured NAT path — on GCP, Google API
|
|
264
|
+
traffic did not use the Cloud NAT address despite
|
|
265
|
+
`privateIpGoogleAccess: false`. Before writing or editing an allowlist or
|
|
266
|
+
perimeter rule, observe the address for the real workload and destination
|
|
267
|
+
at the enforcement point, with a control path that should show a different
|
|
268
|
+
one.
|
|
269
|
+
- **Locating a credential must not print it**: a command meant only to find
|
|
270
|
+
where a secret lives, or to check that it is set, exposes the value if it
|
|
271
|
+
emits it into the transcript — `cat` of an env file, `printenv`,
|
|
272
|
+
`echo $TOKEN`, a keychain read with `-w`, or a grep whose match is the
|
|
273
|
+
secret line. Check existence, length, or shape without emitting any secret
|
|
274
|
+
characters (`[ -n "${X:-}" ]`, `wc -c < file`, a key-name-only listing, a
|
|
275
|
+
nonprinting format check), and let the consumer read the value directly
|
|
276
|
+
from the environment or credential store only when it needs it. Do not
|
|
277
|
+
print even a partial prefix. A value that reached the transcript is
|
|
278
|
+
exposed: rotate it.
|
|
260
279
|
- **Tightening exposure is a behavior change for external clients**: switching
|
|
261
280
|
ingress mode, adding an allowlist, or requiring auth is not safe when
|
|
262
281
|
callers live outside your redeploy. Enumerate which clients reach the
|
|
@@ -0,0 +1,88 @@
|
|
|
1
|
+
---
|
|
2
|
+
language: en
|
|
3
|
+
status: active
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Choosing a visual direction for an operational interface
|
|
7
|
+
|
|
8
|
+
Use when a new visual direction or a material arrangement decision remains
|
|
9
|
+
unresolved in the requested work. For an established direction or a bounded copy
|
|
10
|
+
correction, use the existing system and verify the affected result proportionately.
|
|
11
|
+
|
|
12
|
+
## Start from the task and supported system
|
|
13
|
+
|
|
14
|
+
Identify the content people must read, compare, select, edit or act on. Preserve
|
|
15
|
+
the working scope, authority and state distinctions that matter to that task.
|
|
16
|
+
Inspect the project's existing components and tokens before introducing a new
|
|
17
|
+
visual vocabulary. An external reference does not itself justify replacing the
|
|
18
|
+
project's framework or component library.
|
|
19
|
+
|
|
20
|
+
Consult a reference when it helps resolve the remaining visual decision; reuse
|
|
21
|
+
sufficient existing evidence. Research more when the decision needs it or the
|
|
22
|
+
user requests research. Prefer task-relevant product views or official pattern
|
|
23
|
+
documentation. Distinguish live observations, release screenshots, historical
|
|
24
|
+
images and explanatory artwork. Record the source and date, what it informs,
|
|
25
|
+
and its applicability limits in the ordinary design note. Do not infer token
|
|
26
|
+
values or usability from an image alone.
|
|
27
|
+
|
|
28
|
+
## Choose only the patterns the task needs
|
|
29
|
+
|
|
30
|
+
| Task | Useful arrangement | Meaning to preserve |
|
|
31
|
+
| --- | --- | --- |
|
|
32
|
+
| Reading and review | Give the artifact sufficient space; keep necessary evidence nearby when switching would obstruct comparison | The artifact, source, applicable conditions and review target |
|
|
33
|
+
| Repeated record comparison | Align comparable values; distinguish whole-list, selected-set and individual-record actions | Units, comparison basis, exceptions and each action's scope |
|
|
34
|
+
| Selection-driven editing | Keep relevant properties bound to the selected object | Selection, draft and resulting changes remain attached to the same target |
|
|
35
|
+
| Input and result confirmation | Keep input, consequential effects and results close enough to follow | Entered, saved, approved and actually executed states stay distinct |
|
|
36
|
+
|
|
37
|
+
These are conditional patterns, not a required navigation structure or panel
|
|
38
|
+
count. Choose persistent supporting panels or transient overlays according to
|
|
39
|
+
their role. Give prose a readable measure and comparisons sufficient width.
|
|
40
|
+
Adapt the arrangement to the supported viewport while preserving the current
|
|
41
|
+
selection, return context, reading order and keyboard focus order. Compact
|
|
42
|
+
navigation when needed to keep the main task reachable on a narrow screen.
|
|
43
|
+
|
|
44
|
+
## Define roles before values
|
|
45
|
+
|
|
46
|
+
Explain the proposed reading order, emphasis, density and grouping briefly.
|
|
47
|
+
Map foregrounds, surfaces, action states, spacing relationships and any required
|
|
48
|
+
depth to the existing semantic tokens. Keep focus, selection, feedback and
|
|
49
|
+
disabled states distinct even when values initially coincide. Add a pattern
|
|
50
|
+
token only when it represents an independent decision. The role of a token
|
|
51
|
+
must survive theme changes; check its actual foreground/background pairing.
|
|
52
|
+
|
|
53
|
+
For information-dense operational work, start with typography, alignment and
|
|
54
|
+
spacing to organize content; add containers, color or elevation when they
|
|
55
|
+
communicate a needed relationship or interaction. Large introductory areas,
|
|
56
|
+
repeated summary cards, badges and icon backgrounds need a task or brand purpose.
|
|
57
|
+
Ordinary content need not look raised. Use depth when it clarifies a floating,
|
|
58
|
+
movable or otherwise distinct object. Preserve useful brand and user preferences.
|
|
59
|
+
|
|
60
|
+
Do not impose a reference palette, fixed pixel scale, permanently dim navigation,
|
|
61
|
+
universal card ban or identical layout across tasks. When no system exists, mark
|
|
62
|
+
initial values as project proposals. Define consistent text roles and spacing
|
|
63
|
+
within and between groups, then check the resulting density. Resolve crowded
|
|
64
|
+
content through grouping, widths and wrapping before shrinking text. Keep words
|
|
65
|
+
and meaningful phrases together where the language permits, with a fallback for
|
|
66
|
+
unbroken strings. Check representative Korean text when the product uses Korean.
|
|
67
|
+
|
|
68
|
+
## Compare only enough to resolve the decision
|
|
69
|
+
|
|
70
|
+
Use representative content, long labels, important exceptions and relevant
|
|
71
|
+
states at the target sizes. Keep task, content and capability constant while
|
|
72
|
+
comparing visual alternatives; use a credible baseline. Distinguish token
|
|
73
|
+
changes from changes to regions, information order and grouping. Typography
|
|
74
|
+
and spacing can still alter wrapping and geometry. Where narrow layouts
|
|
75
|
+
converge, do not claim a difference the user cannot see.
|
|
76
|
+
|
|
77
|
+
Inspect the actual rendered result when appearance is being judged. Check
|
|
78
|
+
contrast, clipping, wrapping, state distinctions, action targets and focus
|
|
79
|
+
visibility under the applicable accessibility requirements. Verify that
|
|
80
|
+
rearrangement preserves semantic reading and keyboard order. For interactive
|
|
81
|
+
changes, exercise selection, draft retention and the relevant failure/recovery
|
|
82
|
+
paths; styling does not establish permission or execution.
|
|
83
|
+
|
|
84
|
+
Record visual preference separately from observed task performance. Rendered
|
|
85
|
+
samples support visual inspection; interaction and user-performance claims need
|
|
86
|
+
their own evidence. Preserve the result and remaining uncertainty in ordinary
|
|
87
|
+
team artifacts. No prescribed number of variants or new research report is
|
|
88
|
+
needed for every task.
|
|
@@ -0,0 +1,90 @@
|
|
|
1
|
+
---
|
|
2
|
+
guide_id: ui-design
|
|
3
|
+
language: en
|
|
4
|
+
status: active
|
|
5
|
+
description: For designing, reviewing, or changing task flows, information layout, visual hierarchy, or interaction in operational user interfaces; keep bounded corrections proportionate.
|
|
6
|
+
use_when:
|
|
7
|
+
- designing or reviewing a work application, operations tool, or interactive work surface
|
|
8
|
+
- changing information arrangement, visual hierarchy, design tokens, or interaction states
|
|
9
|
+
- preserving scope, evidence, authority, and continuity through a user interface
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
# Operational interface design and delivery
|
|
13
|
+
|
|
14
|
+
Use for interfaces where people inspect evidence, compose work, make decisions
|
|
15
|
+
or act on records. Match the deliverable to the user's requested stage and scope.
|
|
16
|
+
A clear text correction needs that correction and a proportionate check.
|
|
17
|
+
|
|
18
|
+
1. **Establish what the sources mean now.** Separate implemented behavior,
|
|
19
|
+
accepted requirements, proposed design choices and illustrative states.
|
|
20
|
+
Reconcile amendments and decisions before reusing an older screen. A newer
|
|
21
|
+
date alone does not establish authority. Keep material contradictions and
|
|
22
|
+
unknowns visible instead of silently choosing a convenient interpretation.
|
|
23
|
+
|
|
24
|
+
2. **Preserve requested coverage; bound the work at the right level.** Identify
|
|
25
|
+
the operator, outcome and evidence that would complete this deliverable.
|
|
26
|
+
Distinguish people consuming work context from those authoring or managing
|
|
27
|
+
it; do not invent their visit frequency or force a common starting screen.
|
|
28
|
+
For a complete design request, cover the required operations even when their
|
|
29
|
+
backend is not built; specify their subjects, inputs, effects, dependencies
|
|
30
|
+
and recovery rather than claiming implementation. For a bounded change,
|
|
31
|
+
complete the requested path before expanding neighboring features. Include
|
|
32
|
+
adjacent repairs when that path depends on them or this change causes a
|
|
33
|
+
regression, and state why. An unavailable implementation does not erase a
|
|
34
|
+
requested design contract.
|
|
35
|
+
|
|
36
|
+
3. **Organize around a meaningful next decision.** Show the working scope,
|
|
37
|
+
evidence to compare, current state, available action and resulting next step
|
|
38
|
+
together. Internal modules or protocol phases are not automatically menus
|
|
39
|
+
or buttons. Combine preparation steps behind one authorized intent when no
|
|
40
|
+
user choice is needed, while keeping distinct outcomes understandable. Use
|
|
41
|
+
the team's visual language to assign emphasis according to the task, and
|
|
42
|
+
make typography, spacing, color and boundaries express meaningful
|
|
43
|
+
relationships. Test whether actual task content is prominent at the target
|
|
44
|
+
size; decorative summaries must earn their space. When layout uncertainty
|
|
45
|
+
matters, compare arrangements at the lowest useful fidelity, explaining the
|
|
46
|
+
task tradeoff and what evidence would change the choice.
|
|
47
|
+
|
|
48
|
+
When a new visual direction or a material arrangement decision remains
|
|
49
|
+
unresolved, use
|
|
50
|
+
`${CLAUDE_CONFIG_DIR:-$HOME/.claude}/guides/ui-design/visual-direction.md`.
|
|
51
|
+
An established direction or a bounded correction does not require new
|
|
52
|
+
reference research or multiple designs.
|
|
53
|
+
|
|
54
|
+
4. **Keep qualifications with the thing they qualify.** Preserve applicable
|
|
55
|
+
conditions, exceptions, units, source/version, time meaning and affected
|
|
56
|
+
scope through summaries, edits, comparisons and exports. Keep independent
|
|
57
|
+
dimensions separate. Missing evidence is not an empty result or a zero, and
|
|
58
|
+
an index or relationship view is not proof of its underlying body or meaning.
|
|
59
|
+
Expand technical detail when useful without hiding a consequential limit.
|
|
60
|
+
|
|
61
|
+
5. **Make each action's authority and effect precise.** Name the exact target,
|
|
62
|
+
base and changed content when reviewing a change. Recheck a changed base or
|
|
63
|
+
permission instead of carrying approval onto different input. Distinguish
|
|
64
|
+
selection, saved draft, approval, execution and recipient use as applicable;
|
|
65
|
+
one positive state does not prove the next. Preserve request identity when
|
|
66
|
+
an outcome is uncertain and reconcile its actual result before retrying.
|
|
67
|
+
UI controls reflect the existing authoritative policy; they do not create it.
|
|
68
|
+
If policy is unsettled, label proposed choices and the responsible decision
|
|
69
|
+
owner rather than implying permission or omitting the required design.
|
|
70
|
+
|
|
71
|
+
6. **Carry the same work across views and people.** Preserve the draft, target,
|
|
72
|
+
review/request identity and return context when moving between surfaces.
|
|
73
|
+
Equivalent outcomes need suitable representations, not identical layouts or
|
|
74
|
+
duplicated business rules. Provide keyboard and structured alternatives for
|
|
75
|
+
necessary spatial interactions. Verify evidence for the actual recipient;
|
|
76
|
+
another worker's receipt does not establish their access or delivery.
|
|
77
|
+
|
|
78
|
+
7. **Verify this deliverable, then finish it.** Walk a design's concrete scenario
|
|
79
|
+
and relevant exceptions against its sources; inspect its proposed composition.
|
|
80
|
+
For working changes, exercise the changed runtime path, relevant failures,
|
|
81
|
+
keyboard and recovery. When a visual change is material, inspect the affected
|
|
82
|
+
composition with representative content and relevant states at the target
|
|
83
|
+
size, distinguishing token changes from changes to information arrangement.
|
|
84
|
+
Check that visual, reading and focus order remain coherent across layouts.
|
|
85
|
+
Report visual preference separately from observed task performance. Keep
|
|
86
|
+
simulation, observed runtime and user evidence distinct. Once the required
|
|
87
|
+
checks are complete, leave the result, material decisions, source bindings
|
|
88
|
+
and remaining work in ordinary team artifacts. Keep private facts within
|
|
89
|
+
their permitted boundary. Choose dependencies only when a requirement
|
|
90
|
+
demands a decision, using the supported environment first.
|
|
@@ -77,6 +77,7 @@ cheapest one to write.
|
|
|
77
77
|
- A/B or on/off measurements: before accepting a null result, verify the arms actually received different treatment in the mechanism under test — a shared default or unconditional upstream step can silently apply the treatment to both arms.
|
|
78
78
|
- Multi-stage pipelines with nondeterministic stages: a final-output diff cannot attribute an effect or a regression to a stage — it conflates the change with run-to-run variance. Persist every stage's output, tabulate what each creates, may edit, and only guards, restrict the suspects to the stages that edit the content in question, and find the first stage where the intended effect disappears or the defect appears. Fix there, preferring a structural recheck over another prompt-level instruction that already failed.
|
|
79
79
|
- Before/after comparisons: pin the input to an immutable copy — a snapshot or versioned artifact — and run both arms against it, because a live artifact (a growing log, a regenerated upstream stage) drifts between runs and any diff over it, a matching one included, is evidence of nothing; when the arms are metered, restore the baseline's exact upstream inputs and re-run only the changed stage. This is input identity, not the separate unit-and-denominator basis rule.
|
|
80
|
+
- Cost figures from provider usage records: before pricing, map every provider's token fields onto one schema — uncached input, cache read, cache write, output. Some providers report input as a total that already includes cached tokens: treating that total as uncached input and then adding the cache-read field again counts the cached portion twice, while charging the inclusive total once at the full rate misprices that portion instead. Other providers exclude cached tokens from the input field. When a provider reports or prices cache writes separately, keep them in their own field. Confirm each provider's field meaning against its own usage documentation or a known sample, never by analogy with another provider, and derive the cache hit rate from the normalized fields.
|
|
80
81
|
- Model-behavior guardrails: verify by changed behavior, not recitation — a staged battery from named-trigger cases through disguised, deconfounded, category-wide, and single-variable framings; a clean pass means "no known defect", so re-run the battery when the model changes.
|
|
81
82
|
- Branch/version test builds against real data: explicitly separate every state sink the app touches (files, DB, OS-level stores that ignore env overrides), confirm the launch path propagates the isolation to child processes, and back up live data before the first run — a mismatched schema that drops unknown fields on write is data loss, not a no-op.
|
|
82
83
|
- Sandbox, replay, or re-adjudication runs on production-derived config: enumerate every outbound channel the stage can reach — publish, upload, notify, external write — and disable or redirect each one before the run, proving each disarm fires as you would prove a path guard; a guard on the input or target path alone leaves egress armed. Fingerprint every external destination before the run and diff it after, so an escaped write is caught by the run rather than by a recipient.
|
|
@@ -151,6 +152,11 @@ produce a green with no evidence behind it:
|
|
|
151
152
|
- **The suspiciously fast or empty run.** When a check goes green unexpectedly quickly, or reports
|
|
152
153
|
nothing at all, dump what it actually ran over before believing it. A harness that crashed early
|
|
153
154
|
and one that found nothing produce the same exit code.
|
|
155
|
+
- **The verdict that ran before its checks.** If a runner computes or prints its overall verdict
|
|
156
|
+
before all checks finish, later failures cannot change it. Accumulate one failure count across
|
|
157
|
+
every check, then compute the verdict and exit status once, after the last check. Read a
|
|
158
|
+
multi-check run from that final count and an exit status verified to derive from it — never from
|
|
159
|
+
`tail` or collapsed last lines, which hide failures printed earlier.
|
|
154
160
|
- **The control that failed by crashing.** A negative control is evidence only when it fails
|
|
155
161
|
through the assertion it names: a traceback and a caught violation share an exit code, and an
|
|
156
162
|
early crash can pre-empt every control after it. Treat each traceback in a control run as a
|
|
@@ -173,7 +179,10 @@ produce a green with no evidence behind it:
|
|
|
173
179
|
- **The mutant that never ran.** A mutation verdict counts only if the mutant is valid: it
|
|
174
180
|
compiled, sits on a path the exercised test traverses, and changes the guarded behavior, not
|
|
175
181
|
healed downstream or coinciding with a default. The runner must report build failure,
|
|
176
|
-
unreachable, and
|
|
182
|
+
unreachable, equivalent, and timed out distinctly from KILLED and SURVIVED, and halt on a moved
|
|
183
|
+
anchor. A timeout is not a kill: it moves with machine load, so set the limit well above the
|
|
184
|
+
unmutated suite's runtime and accept a verdict only when repeated runs agree on which mutants
|
|
185
|
+
timed out.
|
|
177
186
|
More tests red than the mutation should touch indicts it; classify a survivor (rebuild,
|
|
178
187
|
discard, genuine gap) before writing a test. **Equivalent is a verdict about the probe as
|
|
179
188
|
much as the mutant**: a probe that is dead or returns a constant reports every mutant as
|
|
@@ -18,7 +18,29 @@ GUIDE = "guides/tooling-gotchas.md"
|
|
|
18
18
|
# Both hosts receive the same command payload and additionalContext response.
|
|
19
19
|
# The guide remains available when native hooks are disabled, untrusted, or do
|
|
20
20
|
# not cover a tool path. --self-test checks every anchor against both guide trees.
|
|
21
|
+
_SECRET_FILE = (r"(?:\.env(?:\.\w+)?|\.netrc|\.pgpass|\.npmrc|\.pypirc|credentials(?:\.json)?"
|
|
22
|
+
r"|auth\.json|secrets?\.(?:ya?ml|json|toml))\b")
|
|
21
23
|
RULES = [
|
|
24
|
+
# First because it is the one reminder whose miss is not re-work but exposure. Each
|
|
25
|
+
# branch is a command whose ordinary output IS the value: a keychain read with -w/-g,
|
|
26
|
+
# printenv or a bare env, a token-printing CLI, an echo of a credential-named variable
|
|
27
|
+
# (`\w*TOKEN\b` leaves $MAX_TOKENS alone), and a display or match over a secret-bearing
|
|
28
|
+
# file. A grep with -l/-L/-c/-q prints names, counts or nothing, so it stays quiet.
|
|
29
|
+
("credential-print",
|
|
30
|
+
re.compile(r"(?i)\bsecurity\s+find-(?:generic|internet)-password\b[^|;&]*\s-[a-z]*[wg]\b"
|
|
31
|
+
r"|\bprintenv\b(?!\s+(?:PATH|HOME|SHELL|USER|PWD|LANG|TERM|TMPDIR)\b)"
|
|
32
|
+
r"|(?:^|[|;&]\s*)env\s*(?:$|[|;&])"
|
|
33
|
+
r"|\bgcloud\s+auth\s+(?:application-default\s+)?print-(?:access|identity)-token\b"
|
|
34
|
+
r"|\bgh\s+auth\s+token\b"
|
|
35
|
+
r"|\bkubectl\s+get\s+secrets?\b[^|;&]*-o\s*=?\s*(?:ya?ml|json|jsonpath)"
|
|
36
|
+
r"|\b(?:echo|printf)\b[^|;&]*\$\{?\w*(?:TOKEN|SECRET|PASSWORD|PASSWD|API_?KEY"
|
|
37
|
+
r"|PRIVATE_?KEY|ACCESS_?KEY|CREDENTIALS?)\b"
|
|
38
|
+
r"|\b(?:cat|head|tail|less|more|bat)\b[^|;&]*" + _SECRET_FILE +
|
|
39
|
+
r"|\b(?:grep|rg)\b(?![^|;&]*\s-[a-zA-Z]*[lLcq])[^|;&]*" + _SECRET_FILE),
|
|
40
|
+
"Checking where a credential lives or that it is set must not print it — test "
|
|
41
|
+
"existence, length, or shape without emitting secret characters; a value that "
|
|
42
|
+
f"reached the transcript is exposed, rotate it ({GUIDE}).",
|
|
43
|
+
"Locating a credential must not print it"),
|
|
22
44
|
("reserved-shell-names",
|
|
23
45
|
re.compile(r"\b(UID|EUID|GID|PPID)="),
|
|
24
46
|
"Assigning reserved shell names (UID/EUID/GID/PPID) can invoke the bound "
|
|
@@ -280,6 +302,7 @@ def self_test() -> int:
|
|
|
280
302
|
# matches() never reaches the caller either — both leave the rule inert while this file
|
|
281
303
|
# goes on reporting that all of them fire.
|
|
282
304
|
FIXTURES = {
|
|
305
|
+
"credential-print": "cat .env",
|
|
283
306
|
"reserved-shell-names": "UID=0 echo hi",
|
|
284
307
|
"git-diff-two-dot": "git diff main..HEAD",
|
|
285
308
|
"git-pull-dirty": "git pull origin main",
|
|
@@ -417,6 +440,24 @@ def self_test() -> int:
|
|
|
417
440
|
if "grep-binary-heuristic" in [h for h, _ in matches("(grep -a needle payload)", limit=None)]:
|
|
418
441
|
problems.append("grep-binary-heuristic: fired although the subshell-grouped "
|
|
419
442
|
"grep carries its own text-mode flag")
|
|
443
|
+
# Every branch of credential-print is a separate way the value reaches the transcript,
|
|
444
|
+
# so each gets its own firing case — a regex that loses one alternative still fires on
|
|
445
|
+
# `cat .env` and would pass the fixture above. The quiet cases are the look-alikes the
|
|
446
|
+
# rule must leave alone: existence and length checks, name-only and count-only greps,
|
|
447
|
+
# a variable named for LLM token counts, a keychain lookup without -w, and a printenv
|
|
448
|
+
# of a non-secret variable.
|
|
449
|
+
for cmd in ("printenv", "env | grep KEY", "echo $GITHUB_TOKEN", 'echo "${OPENAI_API_KEY}"',
|
|
450
|
+
"printf '%s' $DB_PASSWORD", "security find-generic-password -s svc -w",
|
|
451
|
+
"gcloud auth print-access-token", "gh auth token", "grep API_KEY .env",
|
|
452
|
+
"kubectl get secret db -o yaml", "head -3 credentials.json", "tail ~/.netrc"):
|
|
453
|
+
if "credential-print" not in [h for h, _ in matches(cmd, limit=None)]:
|
|
454
|
+
problems.append(f"credential-print: does not fire on a value-printing command ({cmd!r})")
|
|
455
|
+
for cmd in ('[ -n "${GH_TOKEN:-}" ]', "wc -c < .env", "grep -c API_KEY .env", "grep -l token .env",
|
|
456
|
+
"echo $MAX_TOKENS", "security find-generic-password -s svc", "gh auth status",
|
|
457
|
+
"gcloud auth list", "cat README.md", "env FOO=1 python3 x.py", "printenv PATH",
|
|
458
|
+
"grep -rn token src/"):
|
|
459
|
+
if "credential-print" in [h for h, _ in matches(cmd, limit=None)]:
|
|
460
|
+
problems.append(f"credential-print: fired on a command that prints no secret ({cmd!r})")
|
|
420
461
|
for envcmd in ("LC_ALL=C grep needle payload", "cat payload | LC_ALL=C grep needle"):
|
|
421
462
|
if "grep-binary-heuristic" not in [h for h, _ in matches(envcmd, limit=None)]:
|
|
422
463
|
problems.append(f"grep-binary-heuristic: an env-assignment prefix hid the "
|
|
@@ -67,7 +67,7 @@ Delegate execution, not decisions. A unit is delegable only when it is decision-
|
|
|
67
67
|
- Use a resident teammate only for dependent slices in one burst. Verify that the CLI preserves its model and context; resume-after-completion may silently change both. Retire after the burst or cache TTL, and persist durable knowledge in files.
|
|
68
68
|
- After a discard or direction change, respawn once a routine round costs about as much as a fresh slice. Recover unique in-flight state to files first.
|
|
69
69
|
- Redirects to busy workers may queue rather than preempt. Check artifacts before destructive redirects, phrase them conditionally, and stop an actively harmful worker by scoped PID/worktree authority.
|
|
70
|
-
- Idle/progress notifications are hypotheses; verify repo artifacts before re-dispatch. An idle signal is liveness decoupled from the report: a subagent can go idle without ever delivering its result, so idle-without-report is not done — request the report explicitly rather than waiting. Cross-reset state belongs in files, not task boards or transcripts. When polling concurrent async jobs, pin the exact id/handle received at dispatch — a "latest" convenience selector can silently point at a sibling job and return plausible-but-wrong results.
|
|
70
|
+
- Idle/progress notifications are hypotheses; verify repo artifacts before re-dispatch. An idle signal is liveness decoupled from the report: a subagent can go idle without ever delivering its result, so idle-without-report is not done — request the report explicitly rather than waiting. A delivered report can also arrive cut off with no marker — mid-table or mid-finding — and asking for it again inline truncates the same way. When a report is longer than a short summary or must outlive the turn, name an output file in the brief: the worker writes the full report there and returns the path and a one-line summary. Treat a report that ends mid-item as incomplete and switch to the file rather than re-requesting inline. Cross-reset state belongs in files, not task boards or transcripts. When polling concurrent async jobs, pin the exact id/handle received at dispatch — a "latest" convenience selector can silently point at a sibling job and return plausible-but-wrong results.
|
|
71
71
|
- Give reviewers/subagents a read-only diff, snapshot, or isolated worktree — not the live tree the main is editing — and forbid destructive git ops (checkout --, reset --hard, stash, clean) on any tree with uncommitted work; re-verify tree integrity before trusting results produced mid-edit.
|
|
72
72
|
- Codex `spawn_agent` decides how much of the parent crosses: `fork_turns` defaults to `all`, and takes `none` or a turn count. A `SubagentStart` hook there receives `agent_type` and may return `continue: false`, so a tier rule can be enforced rather than stated.
|
|
73
73
|
- No per-spawn instructions suppression exists on either host: the subagent definition carries model and effort, not scope. Excluding the standing instructions is a process-level act — `claude --setting-sources ''`, or `CODEX_HOME` pointed at a directory holding only `auth.json` — and it removes the tier definitions with them, so a reader without those instructions and a pinned tier cannot come from one process. An emptied `CODEX_HOME` without `auth.json` fails 401; skills still load.
|
|
@@ -25,6 +25,11 @@ change introduced. The stages below are for work that outgrows that sentence.
|
|
|
25
25
|
|
|
26
26
|
When the user asks to "설계" or design, stay in design mode. Focus on high-level design and implementation-process design, then present the plan, tradeoffs, review gates, and implementation trigger. Move to implementation after the user asks to implement or approves the plan.
|
|
27
27
|
|
|
28
|
+
For operational user-interface flows, information layout, visual hierarchy or
|
|
29
|
+
interaction design, use `${CODEX_HOME:-$HOME/.codex}/guides/ui-design.md`.
|
|
30
|
+
Keep bounded corrections on the lightweight path; this pointer does not require
|
|
31
|
+
a full UI redesign.
|
|
32
|
+
|
|
28
33
|
## When To Use
|
|
29
34
|
|
|
30
35
|
- Use this workflow for architecture changes, new features, cross-module behavior changes, ontology changes, review-driven fixes, or work that affects user-visible behavior, authority, lifecycle, validation, failure handling, or roadmap commitments.
|
|
@@ -88,6 +93,18 @@ blast radius. Re-mapping a level obliges enumerating every reader of that field,
|
|
|
88
93
|
commonly gates shipping, repair, retry, and display at once. A deferred defect is pinned as a
|
|
89
94
|
strict expected failure, never a silent pass.
|
|
90
95
|
|
|
96
|
+
**Choose the failure posture by what the next step reads.** Across these instructions, an
|
|
97
|
+
unqualified instruction to fail loud or fail clearly means surface the problem and reject the bad
|
|
98
|
+
local result; it does not by itself decide whether a whole production run stops. In development,
|
|
99
|
+
tests, and gates, stop and name the problem, because a silent pass hides a defect. On a production
|
|
100
|
+
runtime path, warn loudly and halt only when continuing would contaminate what the next step
|
|
101
|
+
reads: an input is stale or marked not reusable (partial, failed, blocked); a contract-failing
|
|
102
|
+
value is about to be recorded as valid; or an external write has an unknown outcome. Otherwise —
|
|
103
|
+
an isolated item failure, a tripped breaker, an exhausted budget — warn and continue: work already
|
|
104
|
+
done stays valid, work not done is recorded as not done, and the next run picks it up. This
|
|
105
|
+
default assumes recoverable state kept in artifacts rather than only in the running process. The
|
|
106
|
+
system's owner can redefine it. In no stage does a problem pass silently.
|
|
107
|
+
|
|
91
108
|
## Review Loop
|
|
92
109
|
|
|
93
110
|
- At each stage, run review loops as appropriate: self review, subagent review when available, and structured multi-lens review when the repository or domain supports one (concrete tool: Environment Binding below).
|
|
@@ -110,7 +110,13 @@ Important levers:
|
|
|
110
110
|
grounding checks, provenance checks, citation checks, static checks, E2E
|
|
111
111
|
checks, and semantic quality gates.
|
|
112
112
|
- Retry/fail policy: retry transient generation failures; fail clearly when the
|
|
113
|
-
available route cannot enforce the required contract.
|
|
113
|
+
available route cannot enforce the required contract. A contract-failing item
|
|
114
|
+
is rejected loudly and never recorded as a valid result; record its failed or
|
|
115
|
+
not-done status in artifact state. On a production path the whole run halts
|
|
116
|
+
only when continuing would contaminate what the next step reads — a stale or
|
|
117
|
+
not-reusable input, a contract-failing value about to be recorded as valid,
|
|
118
|
+
or an external write with an unknown outcome; otherwise it warns loudly and
|
|
119
|
+
continues, and the next run picks up the work not done.
|
|
114
120
|
- Observability: prompt packet snapshot, model/provider version, schema hash,
|
|
115
121
|
source snapshot, validator decision, retry reason, and artifact lineage.
|
|
116
122
|
|
|
@@ -256,7 +256,26 @@ depends on it, pin it explicitly instead of trusting the environment.
|
|
|
256
256
|
and check for a front-side cache/CDN separately. Aim the probe at an
|
|
257
257
|
in-unit sentinel the app answers without credentials: a denial the app
|
|
258
258
|
produces anyway passes with the control off, and a redirect into it is a
|
|
259
|
-
bypass.
|
|
259
|
+
bypass. The address an IP-based rule compares is the one selected at its
|
|
260
|
+
enforcement point, and configuration alone may not establish which: behind
|
|
261
|
+
a proxy or CDN it can be the client address the configured forwarded-header
|
|
262
|
+
trust chain hands the app, and a managed platform can route particular
|
|
263
|
+
destinations outside an otherwise configured NAT path — on GCP, Google API
|
|
264
|
+
traffic did not use the Cloud NAT address despite
|
|
265
|
+
`privateIpGoogleAccess: false`. Before writing or editing an allowlist or
|
|
266
|
+
perimeter rule, observe the address for the real workload and destination
|
|
267
|
+
at the enforcement point, with a control path that should show a different
|
|
268
|
+
one.
|
|
269
|
+
- **Locating a credential must not print it**: a command meant only to find
|
|
270
|
+
where a secret lives, or to check that it is set, exposes the value if it
|
|
271
|
+
emits it into the transcript — `cat` of an env file, `printenv`,
|
|
272
|
+
`echo $TOKEN`, a keychain read with `-w`, or a grep whose match is the
|
|
273
|
+
secret line. Check existence, length, or shape without emitting any secret
|
|
274
|
+
characters (`[ -n "${X:-}" ]`, `wc -c < file`, a key-name-only listing, a
|
|
275
|
+
nonprinting format check), and let the consumer read the value directly
|
|
276
|
+
from the environment or credential store only when it needs it. Do not
|
|
277
|
+
print even a partial prefix. A value that reached the transcript is
|
|
278
|
+
exposed: rotate it.
|
|
260
279
|
- **Tightening exposure is a behavior change for external clients**: switching
|
|
261
280
|
ingress mode, adding an allowlist, or requiring auth is not safe when
|
|
262
281
|
callers live outside your redeploy. Enumerate which clients reach the
|