codex-orchestrator 2.0.3 → 2.0.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (131) hide show
  1. package/CHANGELOG.md +28 -427
  2. package/README.md +135 -37
  3. package/dist/src/index.d.ts +1 -1
  4. package/dist/src/index.d.ts.map +1 -1
  5. package/dist/src/v2/adapters/gh-issue-adapter.d.ts.map +1 -1
  6. package/dist/src/v2/adapters/gh-issue-adapter.js +6 -7
  7. package/dist/src/v2/adapters/gh-issue-adapter.js.map +1 -1
  8. package/dist/src/v2/cli-contract.d.ts +3 -3
  9. package/dist/src/v2/cli-contract.d.ts.map +1 -1
  10. package/dist/src/v2/cli-contract.js +1 -3
  11. package/dist/src/v2/cli-contract.js.map +1 -1
  12. package/dist/src/v2/cli.d.ts +24 -0
  13. package/dist/src/v2/cli.d.ts.map +1 -0
  14. package/dist/src/v2/{candidate-cli.js → cli.js} +18 -29
  15. package/dist/src/v2/cli.js.map +1 -0
  16. package/dist/src/v2/code-review-report.d.ts +1 -1
  17. package/dist/src/v2/code-review-report.d.ts.map +1 -1
  18. package/dist/src/v2/code-review-report.js +2 -2
  19. package/dist/src/v2/code-review-report.js.map +1 -1
  20. package/dist/src/v2/codex-process.d.ts.map +1 -1
  21. package/dist/src/v2/codex-process.js +12 -1
  22. package/dist/src/v2/codex-process.js.map +1 -1
  23. package/dist/src/v2/config.d.ts +2 -2
  24. package/dist/src/v2/config.d.ts.map +1 -1
  25. package/dist/src/v2/config.js.map +1 -1
  26. package/dist/src/v2/contained-report-operation.d.ts +2 -2
  27. package/dist/src/v2/contained-report-operation.d.ts.map +1 -1
  28. package/dist/src/v2/contained-report-operation.js +1 -1
  29. package/dist/src/v2/contained-report-operation.js.map +1 -1
  30. package/dist/src/v2/containment.d.ts +4 -0
  31. package/dist/src/v2/containment.d.ts.map +1 -1
  32. package/dist/src/v2/containment.js +9 -0
  33. package/dist/src/v2/containment.js.map +1 -1
  34. package/dist/src/v2/direct-delivery.d.ts +5 -10
  35. package/dist/src/v2/direct-delivery.d.ts.map +1 -1
  36. package/dist/src/v2/direct-delivery.js +25 -90
  37. package/dist/src/v2/direct-delivery.js.map +1 -1
  38. package/dist/src/v2/proof-report.d.ts.map +1 -1
  39. package/dist/src/v2/proof-report.js +55 -29
  40. package/dist/src/v2/proof-report.js.map +1 -1
  41. package/dist/src/v2/run-issue.d.ts +6 -9
  42. package/dist/src/v2/run-issue.d.ts.map +1 -1
  43. package/dist/src/v2/run-issue.js +98 -43
  44. package/dist/src/v2/run-issue.js.map +1 -1
  45. package/dist/src/v2/run-store.d.ts +4 -4
  46. package/dist/src/v2/run-store.d.ts.map +1 -1
  47. package/dist/src/v2/run-store.js +25 -40
  48. package/dist/src/v2/run-store.js.map +1 -1
  49. package/dist/src/v2/runtime.d.ts +3 -3
  50. package/dist/src/v2/runtime.d.ts.map +1 -1
  51. package/dist/src/v2/runtime.js +125 -47
  52. package/dist/src/v2/runtime.js.map +1 -1
  53. package/dist/src/v2/setup-cli.d.ts.map +1 -1
  54. package/dist/src/v2/setup-cli.js +4 -11
  55. package/dist/src/v2/setup-cli.js.map +1 -1
  56. package/dist/src/v2/setup-runtime.d.ts.map +1 -1
  57. package/dist/src/v2/setup-runtime.js +1 -61
  58. package/dist/src/v2/setup-runtime.js.map +1 -1
  59. package/dist/src/v2/setup-store.d.ts +0 -5
  60. package/dist/src/v2/setup-store.d.ts.map +1 -1
  61. package/dist/src/v2/setup-store.js +3 -106
  62. package/dist/src/v2/setup-store.js.map +1 -1
  63. package/dist/src/v2/setup.d.ts +6 -46
  64. package/dist/src/v2/setup.d.ts.map +1 -1
  65. package/dist/src/v2/setup.js +11 -293
  66. package/dist/src/v2/setup.js.map +1 -1
  67. package/dist/src/v2/workflow-assets.d.ts +19 -11
  68. package/dist/src/v2/workflow-assets.d.ts.map +1 -1
  69. package/dist/src/v2/workflow-assets.js +132 -40
  70. package/dist/src/v2/workflow-assets.js.map +1 -1
  71. package/docs/deep-dive.md +272 -56
  72. package/internal-workflow/docs/agents/bugfix-quality-gate.md +11 -0
  73. package/internal-workflow/docs/agents/coding-skill-routing.md +116 -196
  74. package/internal-workflow/docs/agents/review-gates.md +32 -39
  75. package/internal-workflow/docs/agents/review-protocol.md +75 -147
  76. package/internal-workflow/evals/coding-skill-evals.json +66 -0
  77. package/internal-workflow/manifest.json +1 -1
  78. package/internal-workflow/operations/acceptance-proof/SKILL.md +7 -1
  79. package/internal-workflow/operations/ambiguity-review/SKILL.md +2 -0
  80. package/internal-workflow/operations/code-review/SKILL.md +21 -1
  81. package/internal-workflow/operations/implementation/SKILL.md +22 -1
  82. package/internal-workflow/operations/spec-author/SKILL.md +10 -1
  83. package/internal-workflow/operations/spec-review/SKILL.md +10 -1
  84. package/internal-workflow/operations/triage/SKILL.md +10 -1
  85. package/internal-workflow/schemas/code-review-v1.json +1 -1
  86. package/internal-workflow/schemas/proof-report-v1.json +1 -1
  87. package/internal-workflow/skills/agent-auto/SKILL.md +6 -1
  88. package/internal-workflow/skills/code-debugger/SKILL.md +122 -0
  89. package/internal-workflow/skills/code-debugger/agents/openai.yaml +7 -0
  90. package/internal-workflow/skills/code-review/SKILL.md +33 -11
  91. package/internal-workflow/skills/code-review/references/cleanup-lens.md +52 -0
  92. package/internal-workflow/skills/implementation-spec-maker/SKILL.md +15 -6
  93. package/internal-workflow/skills/implementation-spec-maker/references/spec-template.md +2 -2
  94. package/internal-workflow/skills/implementation-spec-review/SKILL.md +108 -204
  95. package/internal-workflow/skills/implementation-spec-review/evals/evals.json +24 -0
  96. package/internal-workflow/skills/implementation-spec-review/references/review-loop.md +93 -0
  97. package/internal-workflow/skills/small-task-implementer/SKILL.md +15 -8
  98. package/internal-workflow/skills/spec-implementer/SKILL.md +101 -172
  99. package/internal-workflow/skills/spec-implementer/evals/evals.json +30 -0
  100. package/internal-workflow/skills/spec-implementer/references/review-loop.md +94 -0
  101. package/internal-workflow/skills/tdd/SKILL.md +15 -2
  102. package/internal-workflow/skills/tdd/agents/openai.yaml +2 -2
  103. package/package.json +9 -6
  104. package/dist/src/v2/adapters/target-activity-fence.d.ts +0 -23
  105. package/dist/src/v2/adapters/target-activity-fence.d.ts.map +0 -1
  106. package/dist/src/v2/adapters/target-activity-fence.js +0 -249
  107. package/dist/src/v2/adapters/target-activity-fence.js.map +0 -1
  108. package/dist/src/v2/candidate-cli.d.ts +0 -26
  109. package/dist/src/v2/candidate-cli.d.ts.map +0 -1
  110. package/dist/src/v2/candidate-cli.js.map +0 -1
  111. package/dist/src/v2/legacy-cutover.d.ts +0 -52
  112. package/dist/src/v2/legacy-cutover.d.ts.map +0 -1
  113. package/dist/src/v2/legacy-cutover.js +0 -87
  114. package/dist/src/v2/legacy-cutover.js.map +0 -1
  115. package/internal-workflow/docs/agents/artifact-review-loop.md +0 -267
  116. package/internal-workflow/docs/agents/implementation-review-loop.md +0 -302
  117. package/internal-workflow/operations/cleanup-review/SKILL.md +0 -3
  118. package/internal-workflow/operations/spec-implementation/SKILL.md +0 -3
  119. package/internal-workflow/profiles/implementer_deep.toml +0 -9
  120. package/internal-workflow/profiles/researcher_standard.toml +0 -9
  121. package/internal-workflow/profiles/reviewer_fast.toml +0 -9
  122. package/internal-workflow/skills/cleanup-review/SKILL.md +0 -84
  123. package/internal-workflow/skills/cleanup-review/agents/openai.yaml +0 -6
  124. package/internal-workflow/skills/codebase-design/DEEPENING.md +0 -35
  125. package/internal-workflow/skills/codebase-design/DESIGN-IT-TWICE.md +0 -50
  126. package/internal-workflow/skills/codebase-design/SKILL.md +0 -82
  127. package/internal-workflow/skills/codebase-design/agents/openai.yaml +0 -6
  128. package/internal-workflow/skills/research/SKILL.md +0 -107
  129. package/internal-workflow/skills/research/agents/openai.yaml +0 -6
  130. package/internal-workflow/skills/ui-evidence-proof/SKILL.md +0 -123
  131. package/internal-workflow/skills/ui-evidence-proof/agents/openai.yaml +0 -6
@@ -1,84 +0,0 @@
1
- ---
2
- name: "cleanup-review"
3
- description: "Run an independent post-implementation simplification review only when a concrete cleanup risk cannot fit final code review. Find removable duplication, obsolete paths, workarounds, and unjustified abstractions without re-reviewing correctness."
4
- ---
5
-
6
- # Cleanup Review
7
-
8
- Use this skill after implementation has settled and before final `$code-review`.
9
- It is a cleanup-only Adapter: its job is to reduce maintenance surface without
10
- changing required observable behavior.
11
-
12
- Do not invoke it separately from implementation size or risk classification.
13
- The final `$code-review` spec/standards lens owns bounded hygiene for every
14
- profile, with a dedicated parallel reviewer track for `high`. Use this Adapter
15
- only when the user, approved source, or repo policy names a concrete evidenced
16
- cleanup risk that cannot be covered proportionately by that lens.
17
-
18
- When an Implementation Review Module schedules this Adapter, follow
19
- `../../docs/agents/implementation-review-loop.md` for topology, state, defect
20
- lifecycle, and Closure. Otherwise follow
21
- `../../docs/agents/review-gates.md`. This skill does not create its own loop.
22
-
23
- ## Invocation Contract
24
-
25
- - Use one profile-selected independent reviewer. Root never reviews or certifies cleanup inline.
26
- - One logical activation spans the Full pass and every lineage-preserving Closure; protocol session rotation creates neither a new activation nor Full pass.
27
- - Run one Full cleanup review only after the complete implementation diff and required validation have settled. Never run it per slice or re-enter cleanup after final code review starts.
28
- - The implementation owner integrates accepted fixes. The reviewer remains read-only and returns evidence and bounded changes.
29
- - If independent review is unavailable, report the gate as unavailable; do not substitute root self-review.
30
-
31
- ## In Scope
32
-
33
- - duplicated logic, sources of truth, registrations, or old/new paths kept in parallel
34
- - dead helpers, flags, adapters, branches, comments, tests, or documentation left by the change
35
- - workaround-shaped conditionals, magic ordering, symptom patches, and unnecessary state
36
- - compatibility or fallback behavior without current repository, source-authority, or production evidence
37
- - new services, events/listeners, adapters, or indirection with one current consumer and only speculative reuse
38
- - ownership placement only when moving or deleting code restores an existing owner without redesigning the system
39
- - stale or duplicated tests/docs only when their cleanup debt is itself the finding
40
-
41
- ## Out Of Scope
42
-
43
- - functional correctness, acceptance completeness, security, privacy, performance, or product decisions
44
- - missing regression coverage or weak behavior proof by itself; route it to the scheduled code-review/test-quality lens
45
- - broad architecture redesign, future extensibility, or a request for a new abstraction unrelated to removing current duplication
46
- - discovery of new sibling edge cases in unchanged behavior
47
- - style, naming, or formatting preferences without concrete maintenance cost
48
-
49
- ## Review Method
50
-
51
- 1. Pin the target revision/diff and source authority supplied by the Review Plan or root.
52
- 2. Inventory material additions, replacements, compatibility paths, and new runtime owners.
53
- 3. Classify each material complexity decision:
54
- - `KEEP`: required by a current invariant and supported by evidence or a boundary proof.
55
- - `SIMPLIFY`: required behavior can use a smaller existing seam or fewer states/branches.
56
- - `REMOVE`: no current behavior, authority, consumer, or compatibility evidence requires it.
57
- 4. Report a finding only when it names exact evidence, concrete maintenance cost, and a behavior-preserving simplification. Uncertain removal becomes follow-up, not a guessed blocker.
58
-
59
- A one-producer/one-consumer abstraction defaults to `SIMPLIFY` unless a concrete
60
- lifecycle, transaction, dependency-direction, or fanout invariant requires the
61
- boundary. Do not request a new abstraction unless it reduces current duplication
62
- or restores an existing owner now.
63
-
64
- For Module-scheduled Closure, inspect only accepted cleanup repairs and their
65
- causal fan-out. Do not restart broad discovery or convert proof-only gaps into
66
- cleanup defects.
67
-
68
- In Module mode, reuse supplied stable cleanup IDs and return each as `verified`
69
- or `reopened`. Hand settled decisions/IDs to final code review; that reviewer
70
- rechecks hygiene only for a concrete regression caused by its own repair.
71
-
72
- ## Output
73
-
74
- Return exactly:
75
-
76
- 1. `Verdict: Clean | Cleanup Needed | Cleanup Blocked`
77
- 2. `Final Coverage: cleanup-only`
78
- 3. `Decisions`: concise `KEEP | SIMPLIFY | REMOVE` records for material complexity, each with evidence and protected invariant or maintenance cost
79
- 4. `Findings`: actionable cleanup defects with supplied ID/lifecycle update, exact path/line evidence, and a behavior-preserving fix; otherwise `None`
80
- 5. `Follow-up Needed`: only risky removals or missing authority/evidence; otherwise `None`
81
-
82
- `Cleanup Blocked` means the cleanup decision requires unavailable evidence or an
83
- explicit product/ownership choice. It never means that correctness or test-quality
84
- review should be performed inside this skill. Keep the result short and auditable.
@@ -1,6 +0,0 @@
1
- interface:
2
- display_name: "Cleanup Review"
3
- short_description: "Exceptional independent simplification review"
4
- default_prompt: "Use $cleanup-review only when an explicit concrete cleanup risk cannot fit the final code-review standards lens."
5
- policy:
6
- allow_implicit_invocation: true
@@ -1,35 +0,0 @@
1
- # Deepening
2
-
3
- How to deepen a cluster of shallow modules safely, given its dependencies. Assumes the vocabulary in [SKILL.md](./SKILL.md) — `module`, `interface`, `seam`, `adapter`.
4
-
5
- ## Dependency categories
6
-
7
- When assessing a candidate for deepening, classify its dependencies. The category determines how the deepened module is tested across its seam.
8
-
9
- ### 1. In-process
10
-
11
- Pure computation, in-memory state, no I/O. Always deepenable — merge the modules and test through the new interface directly. No adapter needed.
12
-
13
- ### 2. Local-substitutable
14
-
15
- Dependencies that have local test stand-ins. Deepenable if the stand-in exists. The deepened module is tested with the stand-in running in the test suite. The seam is internal; no port at the module's external interface.
16
-
17
- ### 3. Remote but owned (Ports & Adapters)
18
-
19
- Your own services across a network boundary. Define a port at the seam. The deep module owns the logic; the transport is injected as an adapter. Tests use an in-memory adapter. Production uses a network adapter.
20
-
21
- ### 4. True external (Mock)
22
-
23
- Third-party services you don't control. The deepened module takes the external dependency as an injected port; tests provide a mock adapter.
24
-
25
- ## Seam discipline
26
-
27
- - **One adapter means a hypothetical seam. Two adapters means a real one.**
28
- - **Internal seams vs external seams.** A deep module can have internal seams used by its own tests. Don't expose internal seams through the interface just because tests use them.
29
-
30
- ## Testing strategy: replace, don't layer
31
-
32
- - Old unit tests on shallow modules become waste once tests at the deepened module's interface exist — delete them.
33
- - Write new tests at the deepened module's interface. The interface is the test surface.
34
- - Tests assert on observable outcomes through the interface, not internal state.
35
- - Tests should survive internal refactors — they describe behaviour, not implementation.
@@ -1,50 +0,0 @@
1
- # Design It Twice
2
-
3
- When the user wants to explore alternative interfaces for a chosen deepening candidate, use this parallel sub-agent pattern.
4
-
5
- Uses the vocabulary in [SKILL.md](./SKILL.md) — `module`, `interface`, `seam`, `adapter`, `leverage`.
6
-
7
- Run this workflow only when the candidate is selected, its constraints are known, and the interface decision is consequential enough to justify independent alternatives. Do not run it for a vocabulary question, one bounded recommendation, or routine helper extraction.
8
-
9
- This file defines the comparison protocol; the active root driver executes it and retains session ownership, user dialogue, delegation, and the final decision.
10
-
11
- ## Process
12
-
13
- ### 1. Frame the problem space
14
-
15
- The active root driver writes a user-facing explanation of the problem space for the chosen candidate:
16
-
17
- - The constraints any new interface would need to satisfy
18
- - The dependencies it would rely on, and which category they fall into
19
- - A rough illustrative code sketch to ground the constraints — not a proposal, just a way to make the constraints concrete
20
-
21
- Show this to the user, then immediately proceed to Step 2. The user reads and thinks while the sub-agents work in parallel.
22
-
23
- ### 2. Spawn sub-agents
24
-
25
- Delegate 3+ radically different interfaces to isolated subagent contexts using the exact named role selected by local coding-skill routing. Run them in parallel when supported. This shared reference must not invent a provider-specific tool or override the configured role, model, or effort.
26
-
27
- If independent subagents are unavailable, state that the normal comparison cannot run. As a fallback, produce three sequential inline alternatives labelled **non-independent**; do not claim that they provide independent design evidence.
28
-
29
- Prompt each sub-agent with a separate technical brief and a different design constraint:
30
-
31
- - Agent 1: minimize the interface
32
- - Agent 2: maximize flexibility
33
- - Agent 3: optimize for the most common caller
34
- - Agent 4 (if applicable): design around ports & adapters for cross-seam dependencies
35
-
36
- Include both [SKILL.md](./SKILL.md) vocabulary and `CONTEXT.md` vocabulary in the brief so each sub-agent names things consistently.
37
-
38
- Each sub-agent outputs:
39
-
40
- 1. Interface
41
- 2. Usage example
42
- 3. What the implementation hides behind the seam
43
- 4. Dependency strategy and adapters
44
- 5. Trade-offs
45
-
46
- ### 3. Present and compare
47
-
48
- The active root driver presents designs sequentially so the user can absorb each one, then compares them in prose. Contrast by depth, locality, and seam placement.
49
-
50
- After comparing, give your own recommendation. If elements from different designs would combine well, propose a hybrid.
@@ -1,82 +0,0 @@
1
- ---
2
- name: codebase-design
3
- description: Design or evaluate a named module interface, deep-module seam, or durable public test seam. Provides bounded design vocabulary; not a session driver, repository-wide scan, or implementation workflow.
4
- ---
5
-
6
- # Codebase Design
7
-
8
- Design deep modules: a lot of behaviour behind a small interface, placed at a clean seam, testable through that interface. Use this language and these principles wherever code is being designed or restructured. The aim is leverage for callers, locality for maintainers, and testability for everyone.
9
-
10
- ## Reference Contract
11
-
12
- Use this skill as the vocabulary owner and bounded decision lens. Answer the named design question through supplied or locally verified evidence. Do not start repository-wide exploration, implementation, writes, subagents, or an open-ended design session merely because this skill loaded.
13
-
14
- Let the active driver own workflow, checkpoints, user dialogue, and mutations. Use `$improve-codebase-architecture` for an explicitly requested architecture scan, `$grilling` for an interactive decision tree, and the normal plan/spec/TDD flows for design authority and implementation.
15
-
16
- Use Design It Twice only after a consequential deepening candidate is selected and its constraints are known.
17
-
18
- ## Glossary
19
-
20
- Use these terms exactly — don't substitute "component", "service", "API", or "boundary". Consistent language is the whole point.
21
-
22
- **Module** — anything with an interface and an implementation. Deliberately scale-agnostic: a function, class, package, or tier-spanning slice. Avoid: unit, component, service.
23
-
24
- **Interface** — everything a caller must know to use the module correctly: the type signature, but also invariants, ordering constraints, error modes, required configuration, and performance characteristics. Avoid: API, signature.
25
-
26
- **Implementation** — what's inside a module, its body of code. Distinct from **Adapter**: a thing can be a small adapter with a large implementation or a large adapter with a small implementation. Reach for "adapter" when the seam is the topic; "implementation" otherwise.
27
-
28
- **Depth** — leverage at the interface: the amount of behaviour a caller or test can exercise per unit of interface they have to learn. A module is **deep** when a large amount of behaviour sits behind a small interface, **shallow** when the interface is nearly as complex as the implementation.
29
-
30
- **Seam** — a place where you can alter behaviour without editing in that place; the location at which a module's interface lives. Avoid: boundary.
31
-
32
- **Adapter** — a concrete thing that satisfies an interface at a seam. Describes role, not substance.
33
-
34
- **Leverage** — what callers get from depth: more capability per unit of interface they learn.
35
-
36
- **Locality** — what maintainers get from depth: change, bugs, knowledge, and verification concentrate in one place rather than spreading across callers.
37
-
38
- ## Deep vs shallow
39
-
40
- **Deep module** = small interface + lots of implementation.
41
-
42
- **Shallow module** = large interface + little implementation.
43
-
44
- When designing an interface, ask:
45
-
46
- - Can I reduce the number of methods?
47
- - Can I simplify the parameters?
48
- - Can I hide more complexity inside?
49
-
50
- ## Principles
51
-
52
- - **Depth is a property of the interface, not the implementation.**
53
- - **The deletion test.** Imagine deleting the module. If complexity vanishes, it was a pass-through. If complexity reappears across N callers, it was earning its keep.
54
- - **The interface is the test surface.**
55
- - **One adapter means a hypothetical seam. Two adapters means a real one.**
56
-
57
- ## Designing for testability
58
-
59
- Good interfaces make testing natural:
60
-
61
- 1. Accept dependencies, don't create them.
62
- 2. Return results, don't produce side effects.
63
- 3. Keep the surface area small.
64
-
65
- ## Relationships
66
-
67
- - A **Module** has exactly one **Interface**.
68
- - **Depth** is a property of a **Module**, measured against its **Interface**.
69
- - A **Seam** is where a **Module**'s **Interface** lives.
70
- - An **Adapter** sits at a **Seam** and satisfies the **Interface**.
71
- - **Depth** produces **Leverage** for callers and **Locality** for maintainers.
72
-
73
- ## Rejected framings
74
-
75
- - **Depth as ratio of implementation-lines to interface-lines** — rewards padding the implementation.
76
- - **"Interface" as the TypeScript `interface` keyword or a class's public methods** — too narrow.
77
- - **"Boundary"** — overloaded with DDD's bounded context. Say **seam** or **interface**.
78
-
79
- ## Going deeper
80
-
81
- - **Deepening a cluster given its dependencies** — see [DEEPENING.md](./DEEPENING.md)
82
- - **Exploring alternative interfaces** — see [DESIGN-IT-TWICE.md](./DESIGN-IT-TWICE.md)
@@ -1,6 +0,0 @@
1
- interface:
2
- display_name: "Codebase Design"
3
- short_description: "Reference vocabulary for deep-module decisions"
4
- default_prompt: "Use $codebase-design inside the current workflow as a bounded reference to evaluate this named Module Interface or Seam; keep the active driver in control and do not start a broader workflow."
5
- policy:
6
- allow_implicit_invocation: true
@@ -1,107 +0,0 @@
1
- ---
2
- name: research
3
- description: Research material external API, SDK, specification, service, or source-code questions using primary sources and save one cited repository artifact. Use for requested durable/delegated research or multi-source contract uncertainty; not for narrow lookups, repo-only work, bug reproduction, or specialized docs tasks.
4
- ---
5
-
6
- # Research
7
-
8
- Resolve one external question into reusable evidence for downstream coding work.
9
- The invoked skill authorizes one `researcher_standard` child; root owns source
10
- verification, artifact integration, user communication, and later decisions.
11
-
12
- ## Route Proportionately
13
-
14
- - Read local evidence first: manifests, lockfiles, installed source, tests,
15
- configs, ADRs, and repository docs.
16
- - Keep one narrow documentation lookup inline unless the user explicitly requests delegation or a durable artifact. When the lookup stays inline, use the owning specialized docs skill or tool and answer in chat without creating an artifact.
17
- - Invoke this workflow when the user requests delegated reading or a saved research result, or when a material decision needs multi-source comparison, freshness checking, or external contract synthesis.
18
- - Use repo exploration or bug-diagnosis skills when the owning evidence is local
19
- code or runtime behavior. Research may supply one external contract input but
20
- never owns bug reproduction or implementation.
21
-
22
- ## Build The Research Capsule
23
-
24
- Before delegation, record:
25
-
26
- - the exact question and decision it must unblock;
27
- - relevant verified local context;
28
- - in-scope and excluded products, versions, environments, and claims;
29
- - allowed primary-source types and required freshness;
30
- - the repository output path.
31
-
32
- Use the repository's existing research-note convention. If none exists, choose
33
- `docs/research/YYYY-MM-DD/HHMM-<slug>.md`.
34
-
35
- ## Delegate One Bounded Question
36
-
37
- Launch one fresh `researcher_standard` child with the Research Capsule and no
38
- inherited conclusions. The child is read-only and must return:
39
-
40
- 1. a short answer;
41
- 2. a claim-to-source ledger for every material fact;
42
- 3. source version or publication/update date when available;
43
- 4. conflicts, uncertainty, and missing evidence;
44
- 5. clearly labelled inferences for the repository decision.
45
-
46
- While it reads, continue only independent local work. Do not make or implement
47
- the blocked decision before the research returns. If the named role is
48
- unavailable, perform the same bounded workflow inline and report the fallback;
49
- do not substitute an unrelated code explorer or reviewer.
50
-
51
- ## Source Standard
52
-
53
- Prefer the source that owns the claim:
54
-
55
- 1. official documentation or specifications;
56
- 2. first-party source code, changelogs, release notes, or issue trackers;
57
- 3. first-party APIs or published schemas.
58
-
59
- Use secondary material only to discover primary sources or to expose a disputed
60
- interpretation. Never promote it to authority when an owning source exists.
61
- Cite the exact page or repository location that supports each material claim.
62
- Separate sourced fact from inference, and state when current behavior cannot be
63
- confirmed.
64
-
65
- Use specialized source adapters when applicable: for example, `$openai-docs`
66
- for OpenAI products, Context7 for precise package documentation, and site
67
- parsers for extraction. Their output still must satisfy this source standard.
68
-
69
- ## Verify And Save
70
-
71
- Root must open and verify every source behind a claim that changes architecture,
72
- scope, implementation, security, cost, or compatibility. Repair unsupported or
73
- overstated claims, then save exactly one Markdown artifact:
74
-
75
- ```markdown
76
- # <Research question>
77
-
78
- ## Decision To Unblock
79
- <decision and relevant local context>
80
-
81
- ## Short Answer
82
- <concise answer>
83
-
84
- ## Findings
85
- | Claim | Primary Source | Version / Date | Confidence |
86
- | --- | --- | --- | --- |
87
- | ... | ... | ... | ... |
88
-
89
- ## Repository Implications
90
- <clearly labelled inferences and affected plans/specs/tickets>
91
-
92
- ## Conflicts And Unknowns
93
- <conflicting sources, stale evidence, and unresolved questions>
94
- ```
95
-
96
- Do not include credentials, private tokens, or copied secrets. Link or cite
97
- sources instead of reproducing long copyrighted passages.
98
-
99
- ## Downstream Contract
100
-
101
- - Return the saved path and the decision it now supports.
102
- - Let plans, PRDs, tickets, and implementation specs cite the artifact as their
103
- external Evidence Map instead of repeating the research.
104
- - Re-read only claims invalidated by changed versions, dates, contracts, or
105
- source conflicts.
106
- - Treat the artifact as evidence, not implementation authority. Behavior-changing
107
- work still follows the normal TDD, implementation, and review routes.
@@ -1,6 +0,0 @@
1
- interface:
2
- display_name: "Research"
3
- short_description: "Research from high-trust primary sources"
4
- default_prompt: "Use $research to investigate this external question and save a cited repository artifact."
5
- policy:
6
- allow_implicit_invocation: true
@@ -1,123 +0,0 @@
1
- ---
2
- name: ui-evidence-proof
3
- description: Prove UI, layout, copy, responsive, browser, mobile, or frontend changes with fresh visual artifacts. Use when app-facing completion requires workflow/viewport evidence and criterion-to-artifact mapping.
4
- ---
5
-
6
- # UI Evidence Proof
7
-
8
- Use this skill when the task needs visual proof in normal Codex app chat. It is a prompt-level proof workflow, not the package runner's machine-enforced Acceptance Proof. The goal is to stop "some screenshot exists" from becoming "the UI is proven."
9
-
10
- ## Activation
11
-
12
- Use this skill for:
13
-
14
- - explicit requests for visual proof, UI proof, screenshots, videos, browser proof, app proof, or layout checks;
15
- - frontend, web, mobile, app-facing, responsive, visual, copy, or styling changes before final handoff;
16
- - checking a screenshot/video/UI dump against an expected user workflow or visual complaint.
17
-
18
- If a stronger project-specific proof runner exists, use it when appropriate, but still use this skill to inspect and explain the evidence before claiming success.
19
-
20
- ## Non-Negotiable Standard
21
-
22
- Visual proof passes only when the current artifact visibly proves the requested user-facing state. A screenshot, video, or UI dump is not proof by itself.
23
-
24
- Before claiming success, produce or verify a UI Evidence Report with:
25
-
26
- 1. **Exact workflow** - the entrypoint and user path used to reach the state.
27
- 2. **Source inputs** - issue/request criteria, screenshot complaint, manual QA notes, implementation evidence, or reproduction signal used to derive checks.
28
- 3. **Viewport coverage** - the relevant desktop/mobile/tablet sizes and why they are sufficient.
29
- 4. **Artifact freshness** - the final post-change artifacts inspected after the last run/edit.
30
- 5. **Layout review** - spacing, padding, clipping, overlap, alignment, responsive placement, and the specific visual complaint.
31
- 6. **Copy review** - user-facing labels and absence of rejected technical terms when copy is part of the task.
32
- 7. **Criterion mapping** - which artifact proves which criterion.
33
-
34
- If any required dimension is missing, do not say proof passed. Continue the proof loop when in scope; otherwise report a blocker or residual risk.
35
-
36
- ## Workflow
37
-
38
- ### 1. Extract The Proof Target
39
-
40
- Identify the task-specific checks before opening tools:
41
-
42
- - user workflow: login/entrypoint, create/edit/detail/settings/modal state, relevant route or screen;
43
- - expected state: visible controls, content, copy, data, interaction result;
44
- - layout expectations: spacing, grouping, alignment, overflow, clipping, responsive behavior;
45
- - copy expectations: accepted wording and rejected implementation terms;
46
- - source inputs: issue, user request, PRD/spec, prior screenshot, manual QA plan, reproduction signal.
47
-
48
- If criteria are underspecified, derive the smallest explicit checklist from the request and state assumptions. Ask only when ambiguity would materially change the proof.
49
-
50
- ### 2. Drive The Real UI Path
51
-
52
- - Prefer the same entrypoint a user would use.
53
- - When credentials are available through the repo's documented environment or smoke setup, use the real login UI.
54
- - If you seed cookies/session state, record why the normal UI path was unavailable or irrelevant.
55
- - For web layout proof, include a wide desktop viewport unless the task clearly excludes desktop.
56
- - Include mobile only when the task mentions mobile/responsive behavior or the change is likely to affect small screens.
57
-
58
- ### 3. Capture Current Artifacts
59
-
60
- Capture artifacts after reaching the exact state:
61
-
62
- - screenshots for visual state;
63
- - video or step screenshots for interaction flows;
64
- - UI tree/DOM text dump when it helps verify labels or hidden overflow;
65
- - console/log output only as supporting evidence, not a visual substitute.
66
-
67
- Do not reuse old screenshots. Check timestamp/path or recapture after the final change. If a path is overwritten, inspect the current file before using it in the final report.
68
-
69
- ### 4. Inspect The Artifact
70
-
71
- Look at the artifact, not only DOM selectors or test assertions.
72
-
73
- Check:
74
-
75
- - correct workflow/state, not a nearby equivalent screen;
76
- - enough surrounding context, not a crop that hides the layout relationship;
77
- - spacing/padding/breathing room;
78
- - overlap, clipping, scroll traps, sticky headers/footers, and text truncation;
79
- - alignment/grouping against neighboring controls;
80
- - expected copy and absence of rejected technical wording;
81
- - obvious loading, error, empty, stale-data, or permission states.
82
-
83
- Use measurements when helpful, but do not let convenient geometry replace the actual visual complaint.
84
-
85
- ### 5. Decide And Act
86
-
87
- - If the proof passes, write the UI Evidence Report.
88
- - If proof reveals an in-scope defect and the user asked for implementation, fix it and rerun proof.
89
- - If proof reveals an out-of-scope defect, record it as residual risk or follow-up.
90
- - If required tools, credentials, or services are missing, report a concrete blocker with the proof attempted.
91
-
92
- ## UI Evidence Report
93
-
94
- Use this compact format in final answers or proof notes:
95
-
96
- ```markdown
97
- ## UI Evidence
98
-
99
- - **Workflow:** <entrypoint -> path -> exact screen state>
100
- - **Source inputs:** <request/issue/spec/reproduction/manual QA sources used>
101
- - **Viewports:** <viewport names and dimensions, with reason>
102
- - **Fresh artifacts:** <current screenshot/video/UI dump paths captured after final run>
103
- - **Layout review:** <spacing/padding/overlap/clipping/alignment findings>
104
- - **Copy review:** <accepted labels and rejected terms checked>
105
-
106
- | Criterion | Artifact | Evidence | Result |
107
- | --- | --- | --- | --- |
108
- | <criterion> | <path/url> | <what is visible and why it proves it> | Pass/Fail/Risk |
109
- ```
110
-
111
- Keep it short. Do not include the table when the task is tiny and prose is clearer, but still cover every proof dimension.
112
-
113
- ## Failure Conditions
114
-
115
- Treat these as proof failures:
116
-
117
- - artifact does not show the exact requested workflow or screen state;
118
- - only mobile/narrow proof exists for a desktop layout claim;
119
- - artifact was captured before the final change or may be stale;
120
- - screenshot exists but layout/copy was not inspected;
121
- - only selector/DOM existence is verified for a visual/layout claim;
122
- - proof uses technical labels the user rejected;
123
- - direct session seeding bypasses a relevant login/user path without explanation.
@@ -1,6 +0,0 @@
1
- interface:
2
- display_name: "UI Evidence Proof"
3
- short_description: "Prove UI changes with fresh visual evidence"
4
- default_prompt: "Use $ui-evidence-proof to verify this UI change with the exact workflow, relevant viewports, current artifacts, layout and copy review, and criterion-to-artifact mapping before claiming it is done."
5
- policy:
6
- allow_implicit_invocation: true