@lmliheng/agent-tuning 0.2.13

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE ADDED
@@ -0,0 +1,201 @@
1
+ Apache License
2
+ Version 2.0, January 2004
3
+ http://www.apache.org/licenses/
4
+
5
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
6
+
7
+ 1. Definitions.
8
+
9
+ "License" shall mean the terms and conditions for use, reproduction,
10
+ and distribution as defined by Sections 1 through 9 of this document.
11
+
12
+ "Licensor" shall mean the copyright owner or entity authorized by
13
+ the copyright owner that is granting the License.
14
+
15
+ "Legal Entity" shall mean the union of the acting entity and all
16
+ other entities that control, are controlled by, or are under common
17
+ control with that entity. For the purposes of this definition,
18
+ "control" means (i) the power, direct or indirect, to cause the
19
+ direction or management of such entity, whether by contract or
20
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
21
+ outstanding shares, or (iii) beneficial ownership of such entity.
22
+
23
+ "You" (or "Your") shall mean an individual or Legal Entity
24
+ exercising permissions granted by this License.
25
+
26
+ "Source" form shall mean the preferred form for making modifications,
27
+ including but not limited to software source code, documentation
28
+ source, and configuration files.
29
+
30
+ "Object" form shall mean any form resulting from mechanical
31
+ transformation or translation of a Source form, including but
32
+ not limited to compiled object code, generated documentation,
33
+ and conversions to other media types.
34
+
35
+ "Work" shall mean the work of authorship, whether in Source or
36
+ Object form, made available under the License, as indicated by a
37
+ copyright notice that is included in or attached to the work
38
+ (an example is provided in the Appendix below).
39
+
40
+ "Derivative Works" shall mean any work, whether in Source or Object
41
+ form, that is based on (or derived from) the Work and for which the
42
+ editorial revisions, annotations, elaborations, or other modifications
43
+ represent, as a whole, an original work of authorship. For the purposes
44
+ of this License, Derivative Works shall not include works that remain
45
+ separable from, or merely link (or bind by name) to the interfaces of,
46
+ the Work and Derivative Works thereof.
47
+
48
+ "Contribution" shall mean any work of authorship, including
49
+ the original version of the Work and any modifications or additions
50
+ to that Work or Derivative Works thereof, that is intentionally
51
+ submitted to Licensor for inclusion in the Work by the copyright owner
52
+ or by an individual or Legal Entity authorized to submit on behalf of
53
+ the copyright owner. For the purposes of this definition, "submitted"
54
+ means any form of electronic, verbal, or written communication sent
55
+ to the Licensor or its representatives, including but not limited to
56
+ communication on electronic mailing lists, source code control systems,
57
+ and issue tracking systems that are managed by, or on behalf of, the
58
+ Licensor for the purpose of discussing and improving the Work, but
59
+ excluding communication that is conspicuously marked or otherwise
60
+ designated in writing by the copyright owner as "Not a Contribution."
61
+
62
+ "Contributor" shall mean Licensor and any individual or Legal Entity
63
+ on behalf of whom a Contribution has been received by Licensor and
64
+ subsequently incorporated within the Work.
65
+
66
+ 2. Grant of Copyright License. Subject to the terms and conditions of
67
+ this License, each Contributor hereby grants to You a perpetual,
68
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
69
+ copyright license to reproduce, prepare Derivative Works of,
70
+ publicly display, publicly perform, sublicense, and distribute the
71
+ Work and such Derivative Works in Source or Object form.
72
+
73
+ 3. Grant of Patent License. Subject to the terms and conditions of
74
+ this License, each Contributor hereby grants to You a perpetual,
75
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
76
+ (except as stated in this section) patent license to make, have made,
77
+ use, offer to sell, sell, import, and otherwise transfer the Work,
78
+ where such license applies only to those patent claims licensable
79
+ by such Contributor that are necessarily infringed by their
80
+ Contribution(s) alone or by combination of their Contribution(s)
81
+ with the Work to which such Contribution(s) was submitted. If You
82
+ institute patent litigation against any entity (including a
83
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
84
+ or a Contribution incorporated within the Work constitutes direct
85
+ or contributory patent infringement, then any patent licenses
86
+ granted to You under this License for that Work shall terminate
87
+ as of the date such litigation is filed.
88
+
89
+ 4. Redistribution. You may reproduce and distribute copies of the
90
+ Work or Derivative Works thereof in any medium, with or without
91
+ modifications, and in Source or Object form, provided that You
92
+ meet the following conditions:
93
+
94
+ (a) You must give any other recipients of the Work or
95
+ Derivative Works a copy of this License; and
96
+
97
+ (b) You must cause any modified files to carry prominent notices
98
+ stating that You changed the files; and
99
+
100
+ (c) You must retain, in the Source form of any Derivative Works
101
+ that You distribute, all copyright, patent, trademark, and
102
+ attribution notices from the Source form of the Work,
103
+ excluding those notices that do not pertain to any part of
104
+ the Derivative Works; and
105
+
106
+ (d) If the Work includes a "NOTICE" text file as part of its
107
+ distribution, then any Derivative Works that You distribute must
108
+ include a readable copy of the attribution notices contained
109
+ within such NOTICE file, excluding those notices that do not
110
+ pertain to any part of the Derivative Works, in at least one
111
+ of the following places: within a NOTICE text file distributed
112
+ as part of the Derivative Works; within the Source form or
113
+ documentation, if provided along with the Derivative Works; or,
114
+ within a display generated by the Derivative Works, if and
115
+ wherever such third-party notices normally appear. The contents
116
+ of the NOTICE file are for informational purposes only and
117
+ do not modify the License. You may add Your own attribution
118
+ notices within Derivative Works that You distribute, alongside
119
+ or as an addendum to the NOTICE text from the Work, provided
120
+ that such additional attribution notices cannot be construed
121
+ as modifying the License.
122
+
123
+ You may add Your own copyright statement to Your modifications and
124
+ may provide additional or different license terms and conditions
125
+ for use, reproduction, or distribution of Your modifications, or
126
+ for any such Derivative Works as a whole, provided Your use,
127
+ reproduction, and distribution of the Work otherwise complies with
128
+ the conditions stated in this License.
129
+
130
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
131
+ any Contribution intentionally submitted for inclusion in the Work
132
+ by You to the Licensor shall be under the terms and conditions of
133
+ this License, without any additional terms or conditions.
134
+ Notwithstanding the above, nothing herein shall supersede or modify
135
+ the terms of any separate license agreement you may have executed
136
+ with Licensor regarding such Contributions.
137
+
138
+ 6. Trademarks. This License does not grant permission to use the trade
139
+ names, trademarks, service marks, or product names of the Licensor,
140
+ except as required for reasonable and customary use in describing the
141
+ origin of the Work and reproducing the content of the NOTICE file.
142
+
143
+ 7. Disclaimer of Warranty. Unless required by applicable law or
144
+ agreed to in writing, Licensor provides the Work (and each
145
+ Contributor provides its Contributions) on an "AS IS" BASIS,
146
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
147
+ implied, including, without limitation, any warranties or conditions
148
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
149
+ PARTICULAR PURPOSE. You are solely responsible for determining the
150
+ appropriateness of using or redistributing the Work and assume any
151
+ risks associated with Your exercise of permissions under this License.
152
+
153
+ 8. Limitation of Liability. In no event and under no legal theory,
154
+ whether in tort (including negligence), contract, or otherwise,
155
+ unless required by applicable law (such as deliberate and grossly
156
+ negligent acts) or agreed to in writing, shall any Contributor be
157
+ liable to You for damages, including any direct, indirect, special,
158
+ incidental, or consequential damages of any character arising as a
159
+ result of this License or out of the use or inability to use the
160
+ Work (including but not limited to damages for loss of goodwill,
161
+ work stoppage, computer failure or malfunction, or any and all
162
+ other commercial damages or losses), even if such Contributor
163
+ has been advised of the possibility of such damages.
164
+
165
+ 9. Accepting Warranty or Additional Liability. While redistributing
166
+ the Work or Derivative Works thereof, You may choose to offer,
167
+ and charge a fee for, acceptance of support, warranty, indemnity,
168
+ or other liability obligations and/or rights consistent with this
169
+ License. However, in accepting such obligations, You may act only
170
+ on Your own behalf and on Your sole responsibility, not on behalf
171
+ of any other Contributor, and only if You agree to indemnify,
172
+ defend, and hold each Contributor harmless for any liability
173
+ incurred by, or claims asserted against, such Contributor by reason
174
+ of your accepting any such warranty or additional liability.
175
+
176
+ END OF TERMS AND CONDITIONS
177
+
178
+ APPENDIX: How to apply the Apache License to your work.
179
+
180
+ To apply the Apache License to your work, attach the following
181
+ boilerplate notice, with the fields enclosed by brackets "[]"
182
+ replaced with your own identifying information. (Don't include
183
+ the brackets!) The text should be enclosed in the appropriate
184
+ comment syntax for the file format. We also recommend that a
185
+ file or class name and description of purpose be included on the
186
+ same "printed page" as the copyright notice for easier
187
+ identification within third-party archives.
188
+
189
+ Copyright [yyyy] [name of copyright owner]
190
+
191
+ Licensed under the Apache License, Version 2.0 (the "License");
192
+ you may not use this file except in compliance with the License.
193
+ You may obtain a copy of the License at
194
+
195
+ http://www.apache.org/licenses/LICENSE-2.0
196
+
197
+ Unless required by applicable law or agreed to in writing, software
198
+ distributed under the License is distributed on an "AS IS" BASIS,
199
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
200
+ See the License for the specific language governing permissions and
201
+ limitations under the License.
package/icon.svg ADDED
@@ -0,0 +1,7 @@
1
+ <svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.7" stroke-linecap="round" stroke-linejoin="round">
2
+ <path d="M5 15.5a7 7 0 0 1 14 0" />
3
+ <path d="M12 8.5v1.8" />
4
+ <path d="M7.05 10.55l1.27 1.27" />
5
+ <path d="M12 15.5l3.8-3.8" />
6
+ <circle cx="12" cy="15.5" r="1.1" />
7
+ </svg>
package/package.json ADDED
@@ -0,0 +1,21 @@
1
+ {
2
+ "name": "@lmliheng/agent-tuning",
3
+ "version": "0.2.13",
4
+ "description": "Agent tuning as a loop: initialize an agent from a requirement, design and calibrate a capability benchmark, run and score isolated cases, and improve the agent from measured results.",
5
+ "license": "Apache-2.0",
6
+ "repository": {
7
+ "type": "git",
8
+ "url": "git+https://github.com/lmliheng/Adelie.git",
9
+ "directory": "plugins/agent-tuning"
10
+ },
11
+ "files": [
12
+ "plugin.json",
13
+ "icon.svg",
14
+ "skills",
15
+ "hooks",
16
+ "LICENSE"
17
+ ],
18
+ "publishConfig": {
19
+ "access": "public"
20
+ }
21
+ }
package/plugin.json ADDED
@@ -0,0 +1,8 @@
1
+ {
2
+ "description": "Agent tuning as a loop: initialize an agent from a requirement, design and calibrate a capability benchmark, run and score isolated cases, and improve the agent from measured results.",
3
+ "description_zh": "闭环的 Agent 调优:依据需求初始化 Agent、设计并校准能力 Benchmark、隔离执行与评分、再根据测得结果改进 Agent。",
4
+ "short_description": "Initialize, benchmark, evaluate and improve agents.",
5
+ "short_description_zh": "初始化、评测、评分与改进 Agent。",
6
+ "version": "2026.09.30.1",
7
+ "category": "ai-app-development"
8
+ }
@@ -0,0 +1,140 @@
1
+ ---
2
+ name: agent-evaluation
3
+ description: Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.
4
+ ---
5
+
6
+ # Agent Evaluation
7
+
8
+ Handle one evaluation request from a `run_subagent` caller: run the specified Test Agent on one Benchmark Case once, score that execution privately, and return one protocol result.
9
+
10
+ The top-level Benchmark Designer or Optimizer owns all Case and Run loops, concurrency, and follow-up handling. This worker handles no other Case or Run, launches no evaluator or subagent, modifies no Agent or Benchmark, and never writes `scoreboard.yaml`. Use the Penguin CLI only to launch the specified Test Agent; do not use it to create another phase, designer, optimizer, or evaluator.
11
+
12
+ Operate silently. Call tools without progress messages. Across all streamed and final responses, the only worker-authored text must be the final plain protocol YAML. Emit no narration, headings, Markdown fences, summaries, private scoring details, or other text.
13
+
14
+ ## Before you start
15
+
16
+ Use this Skill only for a complete request from a `run_subagent` caller. If the request is incomplete or inconsistent, return `invalid_request` through the protocol instead of asking the user a question.
17
+
18
+ ## Contract
19
+
20
+ Require exactly one value for every field below:
21
+
22
+ ```text
23
+ protocol_version: 1
24
+ case_id: <case_id>
25
+ run: <1_based_run_index>
26
+ expected_version: <tested_agent_state_version>
27
+ test_agent_id: <test_agent_id>
28
+ benchmark_id: <benchmark_id>
29
+ provider: <provider>
30
+ model_id: <model_id>
31
+ ```
32
+
33
+ One request represents one Test Agent execution. The `run` value identifies that execution; it is not a repeat count. `provider` and `model_id` must both be non-empty and select that exact configured model. If a required field is missing, duplicated, or conflicting, return `invalid_request` without creating a Workspace or launching the Test Agent.
34
+
35
+ Return a **scored result** when the Test Agent ran and the Rubric could be applied. Wrong, malformed, or missing Test Agent output is still a scored result. Return an **evaluation failure** when the request, Benchmark, launch, version check, Trace binding, or scoring process prevents a valid score.
36
+
37
+ Resolve the Project, Test Agent, Benchmark, and Case only from the explicit request and Environment App Data Dir. Reject traversal, symlink escape, or any path outside the requested Test Agent and Benchmark. Never read a Project configuration file, credential, or vault.
38
+
39
+ ## Prepare
40
+
41
+ Use the `App Data Dir` from the Environment:
42
+
43
+ ```text
44
+ TEST_AGENT_DIR = <app_data_dir>/agents/<test_agent_id>
45
+ BENCHMARK_DIR = <app_data_dir>/benchmarks/<benchmark_id>
46
+ ```
47
+
48
+ The Benchmark is Project-level and is not owned by the Test Agent: it sits beside `agents/` and may evaluate several Agents. `test_agent_id` names the Agent this request evaluates; return it as `agent_id`.
49
+
50
+ Reject path traversal, symlink escape, or any resolved path outside the requested Test Agent and Benchmark. Inspect only the requested Agent State, Benchmark config and Case, isolated Test Workspace, and Traces needed to verify this execution. Do not inspect another Agent, Project secrets, hidden configuration, or unrelated Workspaces or Traces.
51
+
52
+ Require `agent_state/system_config.yaml`, `benchmark_config.toml`, `<case_id>/statement/README.md`, and `<case_id>/rubric/README.md`. Return `benchmark_invalid` when `benchmark_config.toml` says `status = "failed"`: a Benchmark whose calibration failed is not evaluated. Treat `run` only as the caller-owned label for this evaluation and return it unchanged; do not read or validate the total Run count. The top-level Agent State `version`, defaulting to 1, must equal `expected_version`; otherwise return `version_changed`. Read and snapshot `model.thinking_level` from this Target Agent config, using the normal Agent-config default `medium` only when the field is absent. This configured value is the evaluation `thinking_level`; do not require or read thinking metadata from a Trace.
53
+
54
+ Before launch, snapshot every file under the Case's `statement/` and `rubric/` directories. Require a usable Rubric whose scoring items total exactly 100 points. Create a unique Workspace under `<test_agent_dir>/workspaces/`, resolve it to an absolute canonical path, and verify that the resolved path remains under that directory. Copy only `statement/` into it. The Test Agent may see the Statement and its own State, but never the Rubric, Gold answers, scoring rules, or Evaluator reasoning.
55
+
56
+ ## Run and verify
57
+
58
+ Use an existing verified Penguin CLI or repository-local launcher. Do not install or probe a launcher. Snapshot the isolated Workspace and record the existing Trace files.
59
+
60
+ Resolve `PROJECT_DIR`, then derive and verify `PROJECT_ID`, then derive and verify `PENGUIN_HOME`. Perform these as separate shell statements in this order. Never compress the assignments onto one command line, derive a value before its input exists, or substitute another Penguin home. Before launch, confirm that `PROJECT_ID` equals the basename of `PROJECT_DIR` and `PENGUIN_HOME` equals its dirname.
61
+
62
+ Start one foreground execution with a fresh top-level Session. With an explicit pair, use:
63
+
64
+ ```bash
65
+ PROJECT_DIR="<app_data_dir>" # the App Data Dir value from your Environment section is the project root
66
+ PROJECT_ID="$(basename "$PROJECT_DIR")"
67
+ PENGUIN_HOME="$(dirname "$PROJECT_DIR")"
68
+ export PENGUIN_HOME
69
+ penguin run \
70
+ --message "Read README.md in the current Workspace and complete the task exactly as specified there." \
71
+ --provider "<provider>" --model-id "<model_id>" --project-id "$PROJECT_ID" \
72
+ --agent-id "<test_agent_id>" --workspace "<absolute_unique_workspace_path>" \
73
+ --approve allow-all --source benchmark
74
+ ```
75
+
76
+ `--source benchmark` files the Test Session under the Evaluations folder of the Web App's session list rather than the Test Agent's active conversations; never omit it.
77
+
78
+ Use the exact requested Agent, Project, absolute Workspace path, and model pair. Never omit either model flag and never fall back to a Project default. If a launch fails, retry only when unchanged Workspace and Trace evidence proves that the Test Agent did not start. Every retry must follow a new diagnosis and apply a specific correction; never repeat an unchanged launch. Do not impose a numeric retry limit while distinct safe repairs remain. Return `evaluation_failed` when no new repair remains, external configuration is required, or it is unclear whether the Test Agent started.
79
+
80
+ Verify after the run that the State version, configured `model.thinking_level`, and both directory snapshots are unchanged. Return `version_changed` when the State version or configured thinking level differs and `benchmark_invalid` when the Statement or Rubric differs.
81
+
82
+ Inspect only new or changed Traces. Bind exactly one root Test Trace whose Workspace, Agent State path, provider, and model match this request. Ignore unrelated parallel Traces and exclude the root Trace's directly referenced child Sessions. Return `evaluation_failed` if there is no unique match. Read the actual non-empty `provider` and `model_id` from the bound root Trace's `session_meta`; return `evaluation_failed` if either is unavailable. Use the unchanged Target Agent configuration snapshot—not Trace metadata—for `thinking_level`.
83
+
84
+ ## Score
85
+
86
+ Inspect only the isolated Workspace, the bound root Trace, its directly referenced child Traces, and the private Rubric. Apply every scoring item and allowed equivalent. Keep Rubric contents, Gold answers, per-item scoring, and scoring rationale private.
87
+
88
+ A wrong answer, missing artifact, malformed output, or task failure attributable to the Test Agent is scored behavior and returns `status: ok`. A launcher, Trace-binding, or Evaluator failure is not scored. Return `benchmark_invalid` when the Rubric cannot be applied and `evaluation_failed` when the score is non-finite or outside `0..100`.
89
+
90
+ Set `duration_ms` from the root Test Session. Compute cost only from reliable final cumulative usage or cost already recorded in that Session and directly referenced child Traces found in the same bounded pass. Never browse, query a pricing service, or infer cost from external model prices. If the required data is unavailable, return `cost: null`. Missing cost data must not invalidate a score.
91
+
92
+ Round `score` to two decimal places. Preserve a non-null `cost` at the precision recorded in the Trace; do not round it. Write `duration_ms` as a non-negative integer rounded to the nearest millisecond.
93
+
94
+ ## Return
95
+
96
+ Return the required YAML as the only worker-authored text. Do not wrap it in backticks or a Markdown fence.
97
+
98
+ If the caller reports that your response formatting was invalid, use the scored or failed result already present in this Session and resend only the clean protocol YAML. Do not call tools, relaunch the Test Agent, rescore, or add an explanation.
99
+
100
+ For a scored result:
101
+
102
+ ```text
103
+ protocol_version: 1
104
+ status: ok
105
+ case_id: <case_id>
106
+ run: <run>
107
+ agent_id: <test_agent_id>
108
+ expected_version: <version>
109
+ provider: <actual_provider>
110
+ model_id: <actual_model_id>
111
+ thinking_level: <configured_thinking_level>
112
+ score: <0_to_100>
113
+ cost: <number_or_null>
114
+ duration_ms: <non_negative_integer>
115
+ session_id: <test_session_id>
116
+ ```
117
+
118
+ For an evaluation failure, use `null` for an identity field that was missing or conflicting:
119
+
120
+ ```text
121
+ protocol_version: 1
122
+ status: failed
123
+ case_id: <case_id_or_null>
124
+ run: <run_or_null>
125
+ agent_id: <test_agent_id_or_null>
126
+ expected_version: <version_or_null>
127
+ provider: <provider_or_null>
128
+ model_id: <model_id_or_null>
129
+ thinking_level: <thinking_level_or_null>
130
+ failure_code: <stable_failure_code>
131
+ ```
132
+
133
+ Use four failure codes:
134
+
135
+ - `invalid_request`: the request is incomplete or inconsistent.
136
+ - `benchmark_invalid`: the Statement, Rubric, or scoring contract is invalid.
137
+ - `version_changed`: the Test Agent version does not match the request or changed during evaluation.
138
+ - `evaluation_failed`: launch could not be safely repaired, or Trace binding or scoring failed.
139
+
140
+ Never include score, cost, duration, Session id, private data, or optimization advice on failure.
@@ -0,0 +1,172 @@
1
+ ---
2
+ name: agent-initialization
3
+ description: Initialize or extend an Agent from a user requirement - write AGENTS.md, set identity metadata, and create, install or import Skills and hook packages (scripts the harness runs on every prompt, before tool calls, or after a task).
4
+ ---
5
+
6
+ # Agent Initialization
7
+
8
+ This skill initializes an agent's settings from a user requirement — plain files in the target agent's directory. It is also the reference for giving an agent, yourself included, a new Skill or a new hook: both are directories you write, and the harness picks them up from disk.
9
+
10
+ ## Before you start
11
+
12
+ If the user's message only invokes this skill (e.g. "use agent-initialization skill") without a concrete requirement, ask the user what agent they want and what it should do. But when the requirement is already concrete — even a single sentence like "an expert that answers questions about X" — do **not** ask follow-up questions: derive the role and rules from that sentence, apply the defaults below, and list your assumptions in the final reply.
13
+
14
+ ## Resolve the inherited runtime
15
+
16
+ Treat the current Agent as the **Builder**. Resolve the runtime before creating a new Agent:
17
+
18
+ - `provider` and `model_id` are one complete pair. If the user explicitly supplies both, use that pair. If the user supplies neither, inherit the current Builder Session's `Provider` and `Model ID` from the Environment. Reject a half pair.
19
+ - `thinking_level` is independent. If the user explicitly supplies it, use that value. Otherwise read `model.thinking_level` from the Builder's own `agent_state/system_config.yaml`; when the field is absent, use the normal Agent-config default `medium`.
20
+
21
+ Write the resolved `thinking_level` into a brand-new target Agent's `model.thinking_level`, preserving all other copied `model` fields. Penguin does not persist `provider` or `model_id` in Agent State, so never add either field to `system_config.yaml`. When the same request continues into Benchmark design, carry the resolved model pair forward explicitly so evaluation uses the Builder runtime instead of a Project default. When configuring an existing Agent, change `model.thinking_level` only when the user explicitly requests that runtime change.
22
+
23
+ ## Locate the target agent
24
+
25
+ All agents of this project live side by side under `agents/` in the App Data Dir:
26
+
27
+ ```bash
28
+ APP_DATA_DIR="<app_data_dir>" # the App Data Dir value from your Environment section
29
+ ls "$APP_DATA_DIR/agents" # existing agents (each is a folder here)
30
+ TARGET="$APP_DATA_DIR/agents/<agent_id>" # the agent to configure
31
+ ```
32
+
33
+ An agent directory contains `agent_state/` (`system_config.yaml`, `AGENTS.md`, `skills/`, `hooks/`, `memory/`, `tools/`) plus `scratchpad/` — and `traces/`, which appears once the agent has run at least once. `hooks/` exists once a hook package is installed; create it when you need it.
34
+
35
+ To extend the agent you are running as, the target is your own directory: `<app_data_dir>/agents/<agent_id>` with the Agent ID from your Environment section.
36
+
37
+ ## Write AGENTS.md
38
+
39
+ `agent_state/AGENTS.md` is injected into the agent's system prompt — it is where the user requirement becomes behavior. Keep `system_config.yaml`'s `system_prompt` untouched (that is the stable system layer); put everything requirement-specific in AGENTS.md:
40
+
41
+ - Role — what the agent is for, in one or two sentences.
42
+ - Domain guidance — the concrete rules, steps and constraints derived from the user requirement.
43
+
44
+ Be concise: AGENTS.md is prompt context, not documentation. For a domain expert that answers from a knowledge base, a good AGENTS.md is a few lines: the role sentence, "answer strictly from the provided context blocks", citation rules ("cite blocks inline as [1][2]"), a refusal rule for questions the context cannot answer, and "answer in the language of the question".
45
+
46
+ ## Skills
47
+
48
+ A Skill is a directory `agent_state/skills/<skill_name>/` containing a `SKILL.md`. The directory name is the Skill's name (letters, digits, `_`, `-`); nothing else registers it.
49
+
50
+ ```md
51
+ ---
52
+ name: <skill_name>
53
+ description: <one line: what it does and when to use it>
54
+ version: <YYYY.MM.DD.N, e.g. 2026.09.29.1 - today's date, N counts that day's changes>
55
+ ---
56
+
57
+ <skill_instructions>
58
+ ```
59
+
60
+ The `description` line is what the target agent sees in its system prompt, so it has to say when the Skill applies; the body is read only once the agent decides to use it. Optional `short_description` and `short_description_zh` lines give the UI a short blurb. Files the body refers to go beside it (`reference/<topic>.md`, scripts), linked by relative path.
61
+
62
+ There are three ways to get one in place:
63
+
64
+ - **Create.** Write the directory yourself. Keep the body to what a capable agent would not already know: the steps, the commands, the traps.
65
+ - **Install from the library.** Copy the whole `skills/<skill_name>/` directory from an agent that has it — `default_agent` ships the whole library. The user can also install from the Web App's Plugins page.
66
+ - **Import.** Fetch a Skill from a URL, a repository or a local path (`curl`, `git clone`, `unzip`), then place it under `skills/`. A Skill written for another tool usually needs only the frontmatter above. Read everything you fetched in full before installing, and tell the user what it does: a Skill becomes instructions the agent follows in every future session.
67
+
68
+ Do not register Skills in AGENTS.md; the frontmatter is injected automatically.
69
+
70
+ Common library bundles, so you don't under-equip the target:
71
+
72
+ - **App builder** (builds apps or web frontends): `penguin-sdk`, `web-design`, `agenthub-models`.
73
+ - **Knowledge expert** (answers questions over a document set): usually **no** harness agent is needed — build a RAG app with the penguin-sdk skill instead, and configure the app's embedded agent (below).
74
+ - **Evaluation loop**: `benchmark-design`, `agent-evaluation`, `agent-optimization`.
75
+
76
+ When creating a Test Agent, install only the capabilities it needs to solve ordinary tasks.
77
+
78
+ ## Hook packages
79
+
80
+ A hook is a script the harness itself runs at a fixed point of the agent loop. Use one when something must happen every time, whether or not the model remembers: adding context to every prompt, vetting a tool call, deciding that a finished task should continue. A rule the model can simply follow belongs in AGENTS.md or a Skill instead.
81
+
82
+ A hook package is a directory `agent_state/hooks/<package_name>/` holding a `hooks.json` and the scripts it names:
83
+
84
+ ```json
85
+ {
86
+ "name": "append-time",
87
+ "description": "Adds the current local time to every prompt.",
88
+ "description_zh": "为每条 Prompt 附上当前本地时间。",
89
+ "version": "2026.09.29.1",
90
+ "user_prompt": [{ "command": "time.mjs", "timeout": 10 }]
91
+ }
92
+ ```
93
+
94
+ ```js
95
+ // time.mjs - answers every prompt with the current local time.
96
+ const now = new Date();
97
+ const pad = (n) => String(n).padStart(2, "0");
98
+ const offset = -now.getTimezoneOffset();
99
+ const stamp =
100
+ `${now.getFullYear()}-${pad(now.getMonth() + 1)}-${pad(now.getDate())} ` +
101
+ `${pad(now.getHours())}:${pad(now.getMinutes())}:${pad(now.getSeconds())}`;
102
+ const utc = `UTC${offset < 0 ? "-" : "+"}${pad(Math.floor(Math.abs(offset) / 60))}:${pad(Math.abs(offset) % 60)}`;
103
+ const zone = Intl.DateTimeFormat().resolvedOptions().timeZone;
104
+ process.stdout.write(`${JSON.stringify({ context: `Current time: ${stamp} ${zone} (${utc})` })}\n`);
105
+ ```
106
+
107
+ The three hook points, one command list each in `hooks.json` (leave out the ones you do not use):
108
+
109
+ | Key | Runs | The script answers |
110
+ | --- | --- | --- |
111
+ | `user_prompt` | Every time the user submits a prompt | `{ "context": "<text>" }` — sent to the model right behind the user's message |
112
+ | `pre_tool_use` | Before each tool call is approved | `{ "decision": "allow" \| "deny", "reason": "<one line>" }` |
113
+ | `stop` | After every task ends | `{ "decision": "continue", "input": "<next user message>" }` to keep going, or `{ "decision": "stop" }` |
114
+
115
+ Every script is plain Node (`.mjs`, builtin modules only), run as `node <script>` with the package directory as its working directory. It receives one JSON object on stdin and prints one JSON object on stdout; printing nothing means no opinion. A non-zero exit, output that is not JSON, or running past `timeout` seconds (default 60) counts as a failure: it is recorded and ignored, and never stops the run. The full contract — every stdin field, the remaining answer fields, `trigger`, and how to convert another tool's hooks — is in [`reference/hooks.md`](reference/hooks.md); read it before writing a `pre_tool_use` or `stop` hook.
116
+
117
+ The same three ways apply:
118
+
119
+ - **Create.** Write the directory. Give `user_prompt` and `pre_tool_use` commands a small `timeout`: they run on the hot path.
120
+ - **Install from the library.** Copy the whole `hooks/<package_name>/` directory from an agent that has it, or have the user install the plugin from the Web App's Plugins page.
121
+ - **Import.** Fetch the source, read every script in full and review it for anything that exfiltrates data, touches files outside its purpose or runs unknown commands, then convert it to the layout above. Another tool's hook configuration (the `hooks` block of a Claude Code `settings.json`, for instance) maps point by point; `reference/hooks.md` has the table.
122
+
123
+ Test a script by hand before you rely on it — feed it the input it will get and check the exit code and the output:
124
+
125
+ ```bash
126
+ cd "$TARGET/agent_state/hooks/append-time"
127
+ echo '{"hook":"user_prompt","session_id":"test","scratchpad_dir":"/tmp","prompt":"hello"}' | node time.mjs; echo "exit $?"
128
+ ```
129
+
130
+ A hook runs on the user's machine, on every prompt, tool call or task, under the Session's sandbox: the same policy as the agent's commands, which with the sandbox off means the harness's own permissions. Tell the user what each hook you install does and at which point it fires. One switch turns all of an agent's hooks off without uninstalling them: `hooks.enabled: false` in `system_config.yaml`.
131
+
132
+ ## When a change takes effect
133
+
134
+ The harness reads an agent's configuration from disk each time a model context opens — AGENTS.md, `system_config.yaml`, Skills and hook packages alike:
135
+
136
+ - a **new conversation** starts with everything you wrote;
137
+ - a **conversation already running** — the one you are in, if you are changing yourself — picks the change up when its context is next compacted. The user can force that with `/compact`.
138
+
139
+ Nothing changes in the middle of a context, so a hook you just wrote will not fire on the next message of this conversation, and a Skill you just wrote is not in your own system prompt yet. You can still read a new SKILL.md directly and follow it now. Say which of the two cases applies when you report.
140
+
141
+ ## Set name and description
142
+
143
+ In the target's `agent_state/system_config.yaml`, set the top-level `name:` and `description:` fields so the agent is recognizable in lists. For an existing Agent, edit only these two fields unless the user explicitly requested a `thinking_level` change.
144
+
145
+ ## Creating a brand-new agent
146
+
147
+ Prefer configuring an agent the user already created. If the user requires a new Agent, confirm that `TARGET` does not exist. If it already exists, stop and tell the user; never silently overwrite, reinitialize, or reuse an existing Agent under the same id.
148
+
149
+ After confirming that the target is absent, pick a short id using letters, digits, `_`, or `-`, copy the default Agent's `system_config.yaml` as the base, and create the layout described above:
150
+
151
+ ```bash
152
+ mkdir -p "$TARGET/agent_state/skills" "$TARGET/agent_state/hooks" "$TARGET/agent_state/memory" "$TARGET/agent_state/tools" "$TARGET/scratchpad"
153
+ cp "$APP_DATA_DIR/agents/default_agent/agent_state/system_config.yaml" "$TARGET/agent_state/"
154
+ ```
155
+
156
+ Then set the top-level `name`, `description`, and `version: 1`, set `model.thinking_level` to the resolved value, write `agent_state/AGENTS.md` (it lives under `agent_state/`, not at the agent directory root), and install only the Skills and hook packages required by the user's requirement. Do not persist the resolved provider/model pair in the Agent State.
157
+
158
+ ## Validate and report
159
+
160
+ Before finishing:
161
+
162
+ - parse `agent_state/system_config.yaml` and confirm `name`, `description`, a positive integer `version`, and the expected `model.thinking_level`;
163
+ - confirm `agent_state/AGENTS.md` exists and is non-empty;
164
+ - confirm every installed Skill has a parseable `SKILL.md`, and its `name` matches its directory;
165
+ - confirm every installed hook package has a `hooks.json` that parses, that each command it lists is a file inside the package, and that each script ran by hand exits 0 and prints JSON or nothing;
166
+ - confirm no Agent outside `TARGET` was changed.
167
+
168
+ Report the target path, whether an existing Agent was configured or a new Agent was created, assumptions, installed Skills and hook packages (for each hook: what it does and at which point it fires), when the change takes effect, the resolved runtime and whether each value was user-specified or inherited, and validation results.
169
+
170
+ ## The embedded agent of an SDK app
171
+
172
+ An app built with the penguin-sdk skill carries its own agent inside the project (`createAgent({ root })` initializes `<app>/penguin_data/default_project/agents/default_agent/` on first run). That directory has exactly the layout described here, and everything in this skill applies to it: write the app's persona into its `agent_state/AGENTS.md` (the penguin-sdk recipe keeps the source of truth in the project's `persona.md` and copies it in during ingest), and set `name`/`description` in its `system_config.yaml` so the app is recognizable. This is how "the app becomes an expert on X": the persona lives in the embedded agent's AGENTS.md, not in application code.
@@ -0,0 +1,112 @@
1
+ # Hook packages: the full contract
2
+
3
+ A hook package is `agent_state/hooks/<package_name>/`: a `hooks.json` manifest and the scripts it names. The directory name is the package's name (letters, digits, `_`, `-`) and wins over the manifest's `name`.
4
+
5
+ ## `hooks.json`
6
+
7
+ ```json
8
+ {
9
+ "name": "<package_name>",
10
+ "description": "<one line, English>",
11
+ "description_zh": "<one line, Chinese - optional>",
12
+ "version": "<YYYY.MM.DD.N>",
13
+ "user_prompt": [{ "command": "prompt.mjs", "timeout": 10 }],
14
+ "pre_tool_use": [{ "command": "guard.mjs", "timeout": 5 }],
15
+ "stop": [{ "command": "stop.mjs", "timeout": 60 }]
16
+ }
17
+ ```
18
+
19
+ - A hook point you do not use may be left out or given `[]`.
20
+ - `command` is a path relative to the package directory and must stay inside it. An entry without one is ignored.
21
+ - `timeout` is in seconds; 60 when absent.
22
+ - `version` is today's date plus a counter, `2026.09.29.1`. Raise it whenever you change the package.
23
+ - Commands of one point run in the order listed; packages run in name order.
24
+
25
+ ### `trigger` (on `user_prompt` commands only)
26
+
27
+ | Value | The command runs |
28
+ | --- | --- |
29
+ | `"prompt"` (default) | On every prompt the user submits |
30
+ | `"host"` | Only when the host starts the package's own flow by name. Goal mode's start script is the shipped example; an ordinary hook never needs it |
31
+
32
+ Write `trigger` explicitly only for `"host"`. One thing to know when editing an older package: a manifest whose `version` is older than `2026.09.29.1` has its unmarked `user_prompt` commands read as `"host"`, so raise the version when you want them to run on every prompt.
33
+
34
+ ## What a script receives and answers
35
+
36
+ Every script gets one JSON object on stdin with these fields, plus the ones of its point:
37
+
38
+ | Field | Meaning |
39
+ | --- | --- |
40
+ | `hook` | The point: `user_prompt`, `pre_tool_use` or `stop` |
41
+ | `session_id` | The Session's id |
42
+ | `trace_path` | Absolute path of the Trace file being written — a JSONL record of the conversation so far. Absent when the Session keeps no Trace |
43
+
44
+ stdout is one JSON object, or nothing for "no opinion". Unknown fields are dropped.
45
+
46
+ ### `user_prompt`
47
+
48
+ Runs when a task starts from a message the user wrote. It does not run for messages the harness or the server injects: scheduled tasks, a `stop` hook's continuation, background-task reports, mid-run steering.
49
+
50
+ | stdin | |
51
+ | --- | --- |
52
+ | `prompt` | The user's text, with leading marker blocks (a skill invocation, an origin note) stripped |
53
+ | `scratchpad_dir` | The Session's scratchpad directory — the place for state the hook keeps between prompts |
54
+
55
+ | stdout | |
56
+ | --- | --- |
57
+ | `context` | Text for the model. It is sent as its own message right behind the user's, marked as harness-injected, and shown to the user as a collapsed card. Empty or absent adds nothing |
58
+
59
+ The context is appended after the user's message, so it never invalidates the prompt cache of the conversation before it. Keep it short: it is paid for on every prompt.
60
+
61
+ ### `pre_tool_use`
62
+
63
+ Runs once per tool call, before the approval step.
64
+
65
+ | stdin | |
66
+ | --- | --- |
67
+ | `tool_name` | The tool being called, e.g. `exec_command` |
68
+ | `tool_call_id` | The call's id |
69
+ | `arguments` | The call's arguments as the raw JSON string the model wrote — parse it yourself |
70
+
71
+ | stdout | |
72
+ | --- | --- |
73
+ | `decision` | `"deny"` refuses the call and the model reads `reason`; `"allow"` approves it without asking the user; absent leaves the call to the normal approval |
74
+ | `reason` | One line for people |
75
+ | `output` | Your own record: an object of string, number or boolean values |
76
+
77
+ The first decision among all hooks wins. An `allow` never overrides the Project's command policy: a command the policy forbids stays forbidden.
78
+
79
+ ### `stop`
80
+
81
+ Runs after every task: the model's final reply, or a cutoff (the user's abort, a failed request, the turn cap).
82
+
83
+ | stdout | |
84
+ | --- | --- |
85
+ | `decision` | `"continue"` starts another task with `input` as its user message; `"stop"` lets the run end |
86
+ | `input` | Required with `continue`: the next task's input text |
87
+ | `reason` | One line for people |
88
+ | `output` | Your own record: an object of string, number or boolean values |
89
+ | `subagent` | `{ "prompt": "<text>", "agent_id": "<optional>" }` — hands work to a detached background session |
90
+
91
+ The first `continue` wins. After a cutoff a `continue` is recorded but not run: the user's interruption outranks every hook. A `stop` hook that always continues never lets the agent finish — decide from state (the Trace, a file in the scratchpad), and make sure the state eventually says stop.
92
+
93
+ ## Failures
94
+
95
+ A non-zero exit, stdout that is not JSON, or a timeout is a failure. So is a script the Session's sandbox could not confine: it never runs, and its error starts with `sandbox:`. The harness records a failure as a `hook` event carrying the error (the tail of stderr for a non-zero exit) and carries on as if the hook had no opinion. To debug a hook that seems not to fire, look for `"type":"hook"` records in the Trace file, and run the script by hand:
96
+
97
+ ```bash
98
+ echo '{"hook":"stop","session_id":"test","trace_path":"/path/to/trace.jsonl"}' | node stop.mjs; echo "exit $?"
99
+ ```
100
+
101
+ Check three more things when nothing happens: the agent's `hooks.enabled` is not `false`, the conversation has opened a new context since you wrote the package (see "When a change takes effect" in SKILL.md), and the session is not a subagent's — child sessions run no hooks.
102
+
103
+ ## Converting hooks from Claude Code
104
+
105
+ | Claude Code | Here |
106
+ | --- | --- |
107
+ | `UserPromptSubmit` | `user_prompt` — print `{ "context": … }` instead of plain text |
108
+ | `PreToolUse` | `pre_tool_use` — `permissionDecision` becomes `decision`; filter on `tool_name` inside the script, there is no `matcher` |
109
+ | `Stop` | `stop` — `{"decision":"block","reason":…}` becomes `{"decision":"continue","input":…}` |
110
+ | `PostToolUse`, `SessionStart`, `Notification`, others | No counterpart. Say so rather than approximating |
111
+
112
+ A Claude Code hook is a shell command; here a command is a Node script. Wrap a shell command with `node:child_process` (`execFileSync`) when you need one, and keep the JSON on stdout yours.
@@ -0,0 +1,131 @@
1
+ ---
2
+ name: agent-optimization
3
+ description: Improve an Agent State through versioned scores and score-linked Traces from a frozen Benchmark.
4
+ ---
5
+
6
+ # Agent Optimization
7
+
8
+ Improve one Test Agent through an evidence → hypothesis → Candidate → evaluation → accept or rollback loop. Use public Statements, scores, and Test Traces as black-box feedback. Delegate every evaluation to an `agent-evaluation` subagent; never run or score the Test Agent directly.
9
+
10
+ ## Before you start
11
+
12
+ If the request does not identify the Test Agent, frozen Benchmark, desired target score, positive Run count, and round limit, ask for the missing inputs. When they are already supplied, proceed without asking the user to restate them.
13
+
14
+ ## Goal and contract
15
+
16
+ Require an explicit Test Agent, a frozen Benchmark with a complete valid Formal Baseline, a desired target score, a positive `runs` value, and a positive round limit. The Benchmark's `benchmark_config.toml` must say `status = "published"`; a `draft` Benchmark is still being built and is not frozen, and a `failed` one never finished calibrating, so stop and explain in either case. `runs` is the number of Runs per Case for every Candidate in this optimization Session. Freeze it for the Session; do not infer it from `benchmark_config.toml` or the Formal Baseline. Read the evaluation `(provider, model_id, thinking_level)` from the complete Evaluation that matches the current Agent State; do not require the user to repeat it. An Evaluation without any part of this runtime is incomplete and cannot be used as a Reference. The top-level Session must provide `run_subagent`, and the current Agent must have the `agent-evaluation` Skill. If a prerequisite is missing, stop and explain what is needed. Do not create the missing Agent, Benchmark, or Baseline, and do not evaluate the Test Agent directly.
17
+
18
+ A **Reference** is the Agent State currently kept as best, together with its complete Evaluation on the frozen Benchmark.
19
+
20
+ Each round starts from the Reference and tests a bounded, general **Candidate**. Evaluate every Candidate on the frozen Case set with the requested `runs` count and the Reference evaluation runtime. The initial Formal Baseline has one Run per Case; do not rerun or backfill it to the requested count. Compare each Candidate's stored top-level average directly with the current Reference score even when their Run counts differ. Accept the Candidate only when the change is admissible, its Evaluation is complete and valid, and its top-level `score` is strictly higher than the Reference Evaluation's `score`. An accepted Candidate and its Evaluation become the next Reference; otherwise restore the previous Reference. Stop early when the Reference reaches the desired target; otherwise run no more than the requested number of complete valid Candidate rounds.
21
+
22
+ ## Access and changes
23
+
24
+ Resolve paths from the Environment's App Data Dir without recursively discovering the Project:
25
+
26
+ ```text
27
+ PROJECT_DIR = <app_data_dir>
28
+ PROJECT_ID = <basename_of_project_dir>
29
+ PENGUIN_HOME = <parent_of_project_dir>
30
+ TARGET = <app_data_dir>/agents/<test_agent_id>
31
+ STATE = <target>/agent_state
32
+ TRACES = <target>/traces
33
+ BENCHMARK = <app_data_dir>/benchmarks/<benchmark_id>
34
+ SCOREBOARD = <benchmark>/scoreboard.yaml
35
+ SNAPSHOTS = <target>/snapshots
36
+ ```
37
+
38
+ The Benchmark is Project-level rather than owned by the Test Agent: it sits beside `agents/` and may evaluate several Agents. `test_agent_id` names the one this Session optimizes, and every Evaluation records it. Use only the Evaluations whose `agent_id` is that Agent as a Reference or for diagnosis.
39
+
40
+ Inspect only the requested Test Agent and Benchmark: the Agent State, public Statements, Scoreboard, and score-linked Test Traces or artifacts from the Baseline and this optimization, including rejected Candidates.
41
+
42
+ Do not inspect Rubrics, Gold answers, private scoring conditions, Evaluator State, Workspace, or Trace, other Agents, or Project secrets. If private evaluation information enters the Optimizer context, restore the active Candidate and stop as contaminated.
43
+
44
+ Modify only the Test Agent State and the versioned snapshot required to protect it. Do not change the frozen Benchmark, Test Traces, or Project configuration. The only Benchmark write is appending a complete accepted Candidate Evaluation to `scoreboard.yaml`.
45
+
46
+ ## Optimization loop
47
+
48
+ For each round:
49
+
50
+ 1. **Establish the Reference.** Confirm that its complete Evaluation covers the frozen Case set, uses the frozen evaluation runtime, and matches the current Agent State version. Do not require its Run count to equal the requested Candidate `runs` count.
51
+ 2. **Diagnose capability gaps.** Compare each Case's `runs[].score` on the fixed `0..100` scale; use the Evaluation's top-level average `score` only for whole-version comparison. Use public Statements, score-linked Test Traces, and prior accepted or rejected attempts to identify observable behaviors that general Agent State changes could improve. Use repeated Runs to distinguish stable behavior from variation.
52
+ 3. **State a falsifiable hypothesis.** Choose the related gaps to address, connect them to a bounded Candidate, and state which observable decisions or artifacts should change and why. A change that only adds analysis steps without predicting a behavioral change is not a useful hypothesis. If the current diagnosis is exhausted, use the remaining public evidence and prior attempts to construct a different admissible Candidate.
53
+ 4. **Create one Candidate from the Reference.** Apply the change and its Candidate version under the construction and rollback rules below. Do not carry rejected Candidate files into the next attempt.
54
+ 5. **Check admissibility.** Confirm that the change is general, uses no private evaluation information, and modifies only permitted Test Agent State.
55
+ 6. **Evaluate the Candidate.** Delegate the complete frozen Case set × requested `runs` matrix in parallel under the evaluation rules below and assemble all returned cells. Do not modify the Candidate while any cell is in flight.
56
+ 7. **Decide.** Accept the Candidate only when every cell is valid and its Evaluation's top-level average `score` is strictly higher than the Reference Evaluation's `score`. Otherwise restore the Reference. Record separately whether the predicted Case behavior changed; a higher Evaluation score accepts the Candidate even when the stated hypothesis was not supported.
57
+ 8. **Persist and continue.** Immediately append and verify every accepted Candidate Evaluation before starting another round. An accepted Candidate becomes the next Reference. Use valid results from rejected Candidates only as evidence for a later hypothesis. Stop when the Reference reaches the desired target. Otherwise complete the requested number of valid Candidate rounds unless infrastructure, contamination, concurrent State changes, or the inability to construct any admissible Candidate creates a concrete blocker. At the round limit, retain the highest-scoring accepted Reference.
58
+
59
+ A round counts only after one Candidate has a complete valid Evaluation. Corrected requests, validity repairs, and evaluation retries do not consume the round limit. A complete valid Evaluation of a rejected Candidate does count.
60
+
61
+ ## Build and roll back a Candidate
62
+
63
+ Create one Candidate per round from the current Reference. Put behavioral guidance in `AGENTS.md`, reusable target-owned capabilities in a focused Skill, and runtime limits in safe `system_config.yaml` fields. Do not edit `system_prompt` unless requested, modify library-provided Skills for target-specific behavior, or change `model.thinking_level`; the Reference Scoreboard fixes the evaluation thinking level.
64
+
65
+ Candidate version numbers only increase. Start with `Reference version + 1` and never reuse a rejected version. Before changing the Agent State, save the original contents and record any files the Candidate creates.
66
+
67
+ Before changing each Reference State, ensure `<target>/snapshots/v<Reference version>.tar.gz` exists. Reuse it when present. Otherwise create it yourself before editing by atomically archiving `agent_state/` while excluding `.vault.toml`; validate the archived version and never overwrite an existing same-version snapshot. If snapshot creation fails, stop before changing Agent State and report the failure.
68
+
69
+ Keep the exact original-file record for fast in-round rollback.
70
+
71
+ If the Candidate is rejected or cannot be evaluated, restore the Reference files and version, remove files created by the Candidate, and verify the restoration. If another process changes the Agent State, stop without overwriting it.
72
+
73
+ ## Delegate evaluation
74
+
75
+ For each frozen Case, dispatch exactly the requested number of Run cells, using one-based Run indices `1..runs`. Call `run_subagent` for each cell with:
76
+
77
+ ```text
78
+ Use the `agent-evaluation` Skill. Run the specified Test Agent on the specified Case exactly once, then score that single execution.
79
+ protocol_version: 1
80
+ case_id: <case_id>
81
+ run: <1_based_run_index>
82
+ expected_version: <test_agent_state_version>
83
+ test_agent_id: <test_agent_id>
84
+ benchmark_id: <benchmark_id>
85
+ provider: <provider>
86
+ model_id: <model_id>
87
+ ```
88
+
89
+ Inspect the complete streamed and final worker response. Before reading `status`, `score`, or any other protocol field, verify that the worker-authored text is exactly one plain protocol YAML document. Narration, headings, code fences, summaries, or scoring details are not valid protocol. Ask the same Evaluator to resend only the clean YAML from its existing result; do not rerun the Test Agent for a formatting repair and do not extract YAML from the invalid response yourself. Transport metadata added by `run_subagent` is not worker-authored text. If private evaluation information appears, follow the contamination rule above.
90
+
91
+ For every scored result, require its `agent_id` to equal the requested Test Agent and its actual `provider`, `model_id`, and `thinking_level` to equal the Reference runtime. A mismatch invalidates the Candidate matrix and stops optimization; never compare or record scores produced under a different runtime.
92
+
93
+ Correct and resend an `invalid_request`. Stop on `version_changed` or `benchmark_invalid`.
94
+
95
+ For `evaluation_failed`, keep the same Candidate and incomplete matrix. Ask the same Evaluator to diagnose and repair the failed cell, then rerun only that cell when evidence proves the Test Agent did not start. Every retry must apply a new, specific repair; never repeat an unchanged request or launch, and do not impose a numeric retry limit while distinct safe repairs remain. Do not inspect private Evaluator State or abandon the Candidate to design the next version. Stop when no new safe repair remains, external configuration is required, or it is unclear whether the Test Agent started.
96
+
97
+ ## Record and report
98
+
99
+ Append each complete accepted Candidate Evaluation to `scoreboard.yaml` immediately after acceptance and verify the stored Agent id, version, score, matrix, and Session ids before continuing. Obtain the current UTC timestamp from the environment, for example with `date -u +"%Y-%m-%dT%H:%M:%SZ"`, rather than inferring UTC from a displayed local time. Use the same field names as the Baseline:
100
+
101
+ ```yaml
102
+ - time: <ISO-8601 timestamp>
103
+ agent_id: <test_agent_id>
104
+ version: <Candidate version>
105
+ provider: <provider>
106
+ model_id: <model_id>
107
+ thinking_level: <thinking_level>
108
+ summary_title: >-
109
+ <public title>
110
+ summary: >-
111
+ <public summary>
112
+ score: <average of the Case scores>
113
+ cost: <average of known Case costs, or null when every Case cost is null>
114
+ duration_ms: <average of the Case durations>
115
+ cases:
116
+ - case: <case_id>
117
+ score: <average of the Run scores>
118
+ cost: <average of known Run costs, or null when every Run cost is null>
119
+ duration_ms: <average of the Run durations>
120
+ runs:
121
+ - score: <Run score>
122
+ cost: <Run cost or null>
123
+ duration_ms: <Run duration>
124
+ session_id: <Test Session id>
125
+ ```
126
+
127
+ After writing, parse the complete `scoreboard.yaml` and verify the appended Evaluation, including its `agent_id`, before reporting success or continuing.
128
+
129
+ Every Run and Case score is on the fixed `0..100` scale. Do not write `max_score`. Calculate and write every Case and Evaluation average directly in the Scoreboard: ignore `null` values when averaging cost and write `null` only when all contributing costs are unknown; round `score` averages to two decimal places, `cost` averages to six decimal places, and `duration_ms` averages to the nearest integer. These stored values are authoritative—do not add a server, frontend, script, or consistency check that recomputes or validates them. Do not add an `aggregate` object or use `case_id`, `mean_score`, `mean_cost`, or `mean_duration_ms`. Do not record rejected Candidates in the Scoreboard.
130
+
131
+ Report the Baseline and every fully evaluated Candidate with its score, Run count, version, change, decision, and Test Session ids. Make the one-Run Formal Baseline and requested Candidate `runs` count explicit. For each Candidate, distinguish the acceptance decision from whether its stated hypothesis was supported by the predicted Case behavior. Include the final retained version, stop reason, and known limitations. Never report a score for an Agent State that was not evaluated.
@@ -0,0 +1,206 @@
1
+ ---
2
+ name: benchmark-design
3
+ description: Design and calibrate a multi-Case capability Benchmark and establish a traceable Formal Baseline.
4
+ ---
5
+
6
+ # Benchmark Design
7
+
8
+ Build a multi-Case Benchmark for one Test Agent, calibrate its difficulty with one Run per Case, and record the selected frozen Pilot as the Formal Baseline.
9
+
10
+ This Skill changes the Benchmark, never the Test Agent. It does not run or score the Test Agent. Delegate every evaluation with `run_subagent`, and tell each worker to use `agent-evaluation`. Stop after the Baseline; do not begin optimization.
11
+
12
+ ## Before you start
13
+
14
+ If the request does not identify a Test Agent, target capability, desired baseline score, and Pilot iteration limit, ask for the missing inputs. When they are already supplied, proceed without asking the user to restate them. Treat the current Agent as the **Builder**. A user-specified evaluation `(provider, model_id)` takes priority; otherwise inherit the current Builder Session's complete `Provider` and `Model ID` from the Environment. Never use a Project default as an implicit evaluation runtime.
15
+
16
+ ## Workflow
17
+
18
+ - A **Pilot** is a one-Run-per-Case evaluation used to improve the Benchmark. Unselected Pilot results never enter the Scoreboard; the selected result becomes the Formal Baseline after Freeze.
19
+ - **Freeze** means the Benchmark revision and evaluation settings stop changing.
20
+ - A **Formal Baseline** is the accepted result of the selected complete valid Pilot revision, recorded after that exact revision is frozen on one unchanged Agent State version.
21
+
22
+ Follow this order:
23
+
24
+ 1. Validate the Test Agent, target capability, resolved evaluation Runtime, and evaluation access.
25
+ 2. Write a Capability Contract that defines the observable process to measure, common weaker behavior, and the general Agent State improvement the Benchmark should train.
26
+ 3. Plan the complete initial Case set and point allocation. For each Case, privately state the intended behavior, a plausible shortcut for a strong Test Agent, and how the Case distinguishes them. Write and leak-check the complete initial Benchmark.
27
+ 4. Complete one valid evaluation for every planned Case. Together these results form Pilot iteration 1; finish this complete set before refining any Case.
28
+ 5. For later Pilot iterations, use scores and Traces to reconstruct how the Test Agent solved each Case. A single iteration may refine multiple Cases or difficulty dimensions; rerun every affected Case.
29
+ 6. Freeze the first valid Pilot revision that meets the desired baseline score. If none does within the requested valid-iteration limit, restore and freeze the lowest-scoring valid Pilot revision. The desired score steers refinement; the publish gate is fixed at 85 on the `0..100` scale, so a frozen revision that scores below 85 is published even when it misses the desired score.
30
+ 7. Freeze the selected revision and record its complete one-Run-per-Case Pilot result as the Formal Baseline when every cell is valid, the Agent State version remains unchanged, and no known design defect remains. Do not rerun or backfill it. The Formal score does not determine validity.
31
+
32
+ ## Setup and access
33
+
34
+ Require a Test Agent id, target capability, desired baseline score on the fixed `0..100` scale, and a positive Pilot iteration limit. Derive a short semantic Benchmark id if needed. Resolve `(provider, model_id)` once before the first Pilot: use a user-supplied complete pair when present, otherwise inherit the current Builder Session's `Provider` and `Model ID` from the Environment. Reject a half pair or an unavailable inherited value. Read `thinking_level` from the Test Agent's `model.thinking_level` in `agent_state/system_config.yaml`, using the normal Agent-config default `medium` only when that field is absent. Do not read `thinking_level` from a Trace and do not inspect Project configuration.
35
+
36
+ The current Session must provide `run_subagent`, and the current Agent must have `agent-evaluation` installed. If either is missing, stop and explain what is needed.
37
+
38
+ Use the Environment's `App Data Dir` and the explicit Test Agent id:
39
+
40
+ ```text
41
+ TEST_AGENT_DIR = <app_data_dir>/agents/<test_agent_id>
42
+ BENCHMARK_DIR = <app_data_dir>/benchmarks/<benchmark_id>
43
+ SCOREBOARD = <benchmark_dir>/scoreboard.yaml
44
+ ```
45
+
46
+ A Benchmark lives beside `agents/`, not inside one: it belongs to the Project and may evaluate several Agents. `test_agent_id` names the one this request evaluates, and every Evaluation records it.
47
+
48
+ Access only the specified Test Agent and Benchmark: the Agent State, complete Benchmark, and Test Traces or artifacts from valid evaluations. Do not access other Agents, Project secrets, or Evaluator State, Workspace, or Trace.
49
+
50
+ Read the Agent State version from the top-level `version` in `agent_state/system_config.yaml`; use 1 only when it is absent.
51
+
52
+ ## Build the Benchmark
53
+
54
+ ```text
55
+ <benchmark_id>/
56
+ ├── benchmark_config.toml
57
+ ├── scoreboard.yaml
58
+ └── CASE-<nnn>-<semantic-name>/
59
+ ├── statement/
60
+ │ ├── README.md
61
+ │ └── <optional-public-materials>
62
+ └── rubric/
63
+ └── README.md
64
+ ```
65
+
66
+ Each Case contains:
67
+
68
+ - `statement/`, which is public to the Test Agent and defines the objective, available materials, and required artifact.
69
+ - `rubric/`, which is private and defines observable scoring items, points, and Gold answers.
70
+
71
+ Both directories require a `README.md` and may contain supporting files. Do not put Gold answers for evaluated instances, hidden mappings, or private scoring conditions in `statement/`.
72
+
73
+ Create `benchmark_config.toml` with `title`, `description`, `runs = 1`, and `status = "draft"`. Benchmark design always uses one Run per Case; do not ask for or accept another Run count. `status = "draft"` tells the Web App that the Benchmark is still being built — it shows the Benchmark masked, and nobody can use or open it until the status is `published`. A draft ends in one of two states: `published` once a Formal Baseline scoring below 85 is recorded, or `failed` when calibration produces no valid Pilot result to freeze or when the lowest-scoring valid revision still scores 85 or above at the iteration limit. Initialize `scoreboard.yaml` with `evaluations: []`.
74
+
75
+ Pass the resolved `(provider, model_id)` explicitly in every Pilot Evaluator request, starting with the first cell. Freeze that pair and the Test Agent's configured `thinking_level` for the complete Benchmark workflow. Every scored Evaluator result must report the requested pair and the same configured thinking level. A mismatch invalidates the matrix.
76
+
77
+ Before planning Cases, state the Capability Contract:
78
+
79
+ - the public evidence available to the Test Agent;
80
+ - the observable decisions, intermediate artifacts, and checks the capability requires;
81
+ - the weaker behaviors or shortcuts the Benchmark should distinguish; and
82
+ - the reusable Agent State behavior that could improve the measured capability.
83
+
84
+ Before writing each Case, privately record the required behavior, a plausible shortcut for a strong Test Agent, the chosen difficulty, the different scored decision or artifact each behavior should produce, and why the distinction measures the target capability. Design the Case so the measured capability affects the score. Do not optimize the Statement to help the Test Agent succeed or copy this design rationale into it.
85
+
86
+ The Statement presents the task, not the Benchmark's teaching or design intent. It describes the objective, available materials, option meanings, output format, and necessary constraints. It must not prescribe the reasoning sequence, identify decisive evidence, name the shortcut, or reveal private scoring preferences. When an auditable artifact is needed, request concise supporting evidence without prescribing how to obtain it.
87
+
88
+ Keep the evaluation contract well-defined, but do not require the public Statement to uniquely determine the Gold. Public information may be incomplete or conflicting, and the Rubric may encode a private decision standard or preference. Fix that private standard before evaluating the revision and never change its Gold after seeing the evaluated answer. The standard must remain tied to the target capability: it should express a stable reusable policy, priority, inference boundary, or other behavior that a better Agent State could apply across instances. Do not use a capability-irrelevant random hidden mapping merely to lower the score, and do not disclose every decisive premise or priority merely to make the public task complete.
89
+
90
+ The first complete revision is an exploratory probe. Use its Pilot to learn how the Test Agent interprets the tasks, forms candidate rules, and uses shortcuts; refine the Benchmark before treating it as calibrated. A later revision may intentionally add information gaps, conflicts, private preferences, or other capability-relevant distinctions in response to an earlier Trace, provided the next revision's Rubric is fixed before dispatch.
91
+
92
+ Every Case Rubric has a fixed maximum of 100 points, with observable scoring items and meaningful partial credit. Allocate most points within each Case to decisions or concise artifacts on which the intended behavior and plausible shortcut differ. Keep generic format compliance, evidence enumeration, and analysis completeness from creating a high score floor unless those are themselves the target capability. Allocate points from capability coverage before the first Pilot. Do not change scoring items solely to satisfy the desired score; when a redesign changes coverage, re-plan that Case's 100-point allocation before evaluating the revised Case set. When final choices do not distinguish the intended behavior from a shortcut, score a concise auditable artifact, but define only its required content or format—not the method used to produce it.
93
+
94
+ Before the first dispatch of every new or changed Case revision, run a consistency review:
95
+
96
+ - Confirm that the current Statement is internally coherent. Intentional conflicts must be presented as conflicts between sources, rules, or positions rather than as contradictory claims by the Benchmark itself.
97
+ - Confirm that the current Rubric is consistent with the current Statement and fixed private standard. It must be self-contained and must not refer to an earlier revision or missing context.
98
+ - Confirm that every scoring item applies to the Case's actual requested output and relies only on premises that are defined, provided, or explicitly private under the fixed standard.
99
+
100
+ This review does not require the public Statement to contain enough information to reproduce the private standard or uniquely derive every Gold answer. Unchanged Cases do not need another review during that iteration. Keep this review in Builder analysis and Trace; fix defects in the Case rather than creating a separate audit artifact.
101
+
102
+ Also compare all public files with the private Rubric. Confirm that no public file reveals Gold answers, private scoring conditions, or hints that identify the intended solution. This is the leak check.
103
+
104
+ ## Delegate evaluation
105
+
106
+ For each Case × Run cell, call `run_subagent` with the request below. Dispatch independent cells in parallel up to available concurrency.
107
+
108
+ ```text
109
+ Use the `agent-evaluation` Skill. Run the specified Test Agent on the specified Case exactly once, then score that single execution.
110
+ protocol_version: 1
111
+ case_id: <case_id>
112
+ run: <1_based_run_index>
113
+ expected_version: <test_agent_state_version>
114
+ test_agent_id: <test_agent_id>
115
+ benchmark_id: <benchmark_id>
116
+ provider: <provider>
117
+ model_id: <model_id>
118
+ ```
119
+
120
+ Inspect the complete streamed and final worker response. Before reading `status`, `score`, or any other protocol field, verify that the worker-authored text is exactly one plain protocol YAML document. Narration, headings, code fences, summaries, or scoring details are not valid protocol. Ask the same Evaluator to resend only the clean YAML from its existing result; do not rerun the Test Agent for a formatting repair and do not extract YAML from the invalid response yourself. Transport metadata added by `run_subagent` is not worker-authored text. A wrong or missing Test Agent artifact is a valid scored result and must not be retried.
121
+
122
+ For every scored result, require non-empty `agent_id`, `provider`, `model_id`, and `thinking_level`. Require `agent_id` to equal the requested Test Agent. Require the model pair to equal the explicitly resolved pair and the thinking level to equal the Test Agent configuration read before dispatch. Reject a Pilot result whose cells report mixed or mismatched runtimes. The Evaluator verifies provider/model from the root Trace and reports thinking from the unchanged Target Agent configuration; it does not require Trace metadata for thinking.
123
+
124
+ Correct and resend an `invalid_request`. For `benchmark_invalid`, repair and rerun the affected Case during Pilot. For `version_changed`, discard the current Pilot result and restart after the Agent version is stable.
125
+
126
+ For `evaluation_failed`, keep the same Benchmark revision and cell. Diagnose the failure and retry only when evidence proves the Test Agent did not start and the retry applies a new, specific repair. Do not set a numeric retry limit or repeat an unchanged launch. Stop when no new safe repair remains, external configuration is required, or it is unclear whether the Test Agent started. Never treat an evaluation failure as score zero.
127
+
128
+ ## Refine the Benchmark
129
+
130
+ Treat the first draft as a hypothesis. The first valid result from every planned Case together forms Pilot iteration 1. A later iteration starts after a difficulty refinement and completes when every affected Case has a valid new result. Request corrections, validity repairs, and evaluation reruns stay in the current iteration and do not consume the requested iteration budget. Use the recorded Agent State version and fixed evaluation runtime.
131
+
132
+ Keep unselected Pilot results out of the Scoreboard. During calibration, retain only one temporary restorable copy: the lowest-scoring complete valid revision seen so far, including its one-Run-per-Case result. Store it outside the Project's `benchmarks/`, replace it only when a lower valid revision completes, and never retain invalid revisions.
133
+
134
+ Use the Pilot to find the current Test Agent's capability boundary.
135
+
136
+ Before editing, distinguish a validity repair from a difficulty refinement. A validity repair fixes an unusable task or scoring contract and stays in the current Pilot iteration. A difficulty refinement changes what the valid Benchmark measures and completes the next iteration after every affected Case has a valid result.
137
+
138
+ Before editing, estimate how much of the score the planned refinements can affect. If the range is too small to materially approach the desired score, revise more affected Cases, use more than one difficulty dimension, or replace low-signal Cases.
139
+
140
+ Prefer refinements that create one or more scored separating decisions. A refinement may change the public task or evidence, introduce or preserve a reasonable information gap or conflict, or apply a fixed private standard. Adding another explicit rule, exception, source, or checklist is not a difficulty increase when the observed strategy can still follow it to the Gold. A Rubric-only refinement is allowed but not preferred when the public task already contains the relevant information, the current Rubric fails to distinguish merely mentioning it from handling it correctly, and the Builder can explain which reusable capability the new scoring distinction measures. Do not add points merely because the previous Test Agent omitted a phrase. Fix the revised Rubric before dispatch and treat it as a changed Case revision.
141
+
142
+ For each refinement iteration:
143
+
144
+ 1. **Observed strategy.** Reconstruct the Test Agent's actual solution method from its score, artifact, and Trace.
145
+ 2. **Missing behavior.** Identify the general behavior that the observed strategy skipped or simplified. Repair missing evidence, arbitrary mappings, ambiguity, or scoring defects before increasing difficulty.
146
+ 3. **Separating prediction.** Before dispatch, predict the decision or artifact the observed strategy will produce, the different result the desired behavior will produce, and the score range affected. If both behaviors are expected to reach the same scored result, choose another refinement.
147
+ 4. Update any number of diagnosed Cases or difficulty dimensions, run the consistency review and leak check for each changed revision, and rerun every affected Case.
148
+
149
+ Reuse a Pilot result only when the Case revision, scoring, Agent State version, and evaluation runtime are unchanged.
150
+
151
+ An information gap or supported alternative is not automatically a design defect. Treat it as a defect only when the task or fixed private standard is incoherent, changes after evaluation, leaks the answer, or no reusable Agent behavior could plausibly improve the score.
152
+
153
+ More rows, fields, distractors, files, near-duplicate examples, or explicit rule layers do not increase difficulty when the observed strategy still solves the Case. Base refinements on observed behavior and fix the Gold before each evaluation.
154
+
155
+ Freeze immediately when a complete valid Pilot iteration meets the desired baseline score and no known design defect remains. Do not run another difficulty refinement merely to create more score margin. Otherwise continue through the requested valid-iteration limit. If the desired score is still unmet, restore the temporary lowest-scoring valid revision and proceed to Freeze. The desired baseline score is a calibration target, not the publish gate: the gate is fixed at 85 on the `0..100` scale, and any frozen valid revision scoring below 85 is published, however far it stays from the desired score. Report `calibration_failed` only when no valid Pilot revision can be produced, evaluation failures prevent a valid selection, or the lowest-scoring valid revision still scores 85 or above at the iteration limit — a Test Agent that already scores that high leaves the Benchmark nothing to measure. Missing the desired score alone is never a failure.
156
+
157
+ ## Freeze and record the Formal Baseline
158
+
159
+ After selecting the Pilot revision, restore that exact revision and its complete result if needed. Run a complete consistency review and final leak check across every Case. If the review finds a defect, repair it and produce a complete valid one-Run-per-Case Pilot result for the repaired revision before selecting and freezing it. Freeze the Benchmark and record the current Agent State version. Do not launch a fresh Formal matrix, rerun the selected Pilot, or backfill it to another Run count.
160
+
161
+ Accept the selected Pilot result as the Formal Baseline when every Case has exactly one valid Run, every cell reports the frozen evaluation runtime, the Agent State version remains unchanged, the private scoring standard remained fixed, and every score loss reflects the Capability Contract. Record the Formal Baseline even when its score does not meet the desired baseline score; only a score of 85 or above blocks it.
162
+
163
+ Report `calibration_failed` only when no valid revision remains, evaluation failures prevent a complete selected Pilot result, or the selected revision scores 85 or above. Never record a partial, abandoned, invalid, or non-selected Pilot result as the Formal Baseline.
164
+
165
+ ## Record and finish
166
+
167
+ After validation, obtain the current UTC timestamp from the environment, for example with `date -u +"%Y-%m-%dT%H:%M:%SZ"`, rather than inferring UTC from a displayed local time. Append only the accepted Formal Baseline to `scoreboard.yaml` using exactly this structure:
168
+
169
+ ```yaml
170
+ evaluations:
171
+ - time: <ISO-8601 timestamp>
172
+ agent_id: <test_agent_id>
173
+ version: <Agent State version>
174
+ provider: <provider>
175
+ model_id: <model_id>
176
+ thinking_level: <thinking_level>
177
+ summary_title: >-
178
+ <public title>
179
+ summary: >-
180
+ <public summary>
181
+ score: <average of the Case scores>
182
+ cost: <average of known Case costs, or null when every Case cost is null>
183
+ duration_ms: <average of the Case durations>
184
+ cases:
185
+ - case: <case_id>
186
+ score: <average of the Run scores>
187
+ cost: <average of known Run costs, or null when every Run cost is null>
188
+ duration_ms: <average of the Run durations>
189
+ runs:
190
+ - score: <Run score>
191
+ cost: <Run cost or null>
192
+ duration_ms: <Run duration>
193
+ session_id: <Test Session id>
194
+ ```
195
+
196
+ After writing, parse the complete `scoreboard.yaml` and verify the appended Evaluation, including its `agent_id`, before reporting success or continuing.
197
+
198
+ Once the Formal Baseline is verified, set `status = "published"` in `benchmark_config.toml`. Change only that line, keep `title`, `description` and `runs` as they are, and parse the file again to confirm it is valid TOML. When the run ends in `calibration_failed`, set `status = "failed"` the same way — change only that line, keep the other fields, and parse the file again — so the Web App tells the user that this Benchmark failed to calibrate and has to be deleted and created again. Never leave a failed Benchmark on `draft`, and never write `failed` because the desired baseline score was missed: a Formal Baseline below 85 is published.
199
+
200
+ Every Run and Case score is on the fixed `0..100` scale. Do not write `max_score`. Calculate and write every Case and Evaluation average directly in the Scoreboard: ignore `null` values when averaging cost and write `null` only when all contributing costs are unknown; round `score` averages to two decimal places, `cost` averages to six decimal places, and `duration_ms` averages to the nearest integer. These stored values are authoritative—do not add a server, frontend, script, or consistency check that recomputes or validates them. Do not add an `aggregate` object or use `case_id`, `mean_score`, `mean_cost`, or `mean_duration_ms`.
201
+
202
+ Report the Benchmark path, configuration, Agent State version, Evaluation average and Case Run scores, Test Session ids, and known limitations. Include one compact row per Pilot iteration with its score, diagnosed capability gap, difficulty adjustment, and freeze or stop decision. Identify which one-Run-per-Case Pilot result was recorded as the Formal Baseline.
203
+
204
+ After the accepted Formal Baseline is recorded, delete the temporary lowest-revision copy and other Builder calibration scaffolding. Keep the frozen Benchmark, Scoreboard, evaluation Workspaces, and score-linked Traces.
205
+
206
+ Do not reveal Rubrics, Gold answers, latent rules, per-item scores, or private scoring information. Stop after reporting the Baseline; do not modify the Test Agent or begin optimization.