@jenga-ai/agent 1.0.1 → 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (38) hide show
  1. package/README.md +3 -4
  2. package/agents/scrum-master.md +75 -0
  3. package/mcp/router/embedder.js +1 -1
  4. package/mcp/training_runner/index.js +239 -0
  5. package/mcp/training_runner/package-lock.json +1065 -0
  6. package/mcp/training_runner/package.json +15 -0
  7. package/package.json +14 -16
  8. package/skills/close-story/SKILL.md +203 -0
  9. package/skills/close-story/scripts/check-story-closeable.sh +195 -0
  10. package/skills/close-story/scripts/compute-scope-divergence.sh +128 -0
  11. package/skills/close-story/scripts/extract-diff-stats.sh +48 -0
  12. package/skills/close-story/scripts/extract-task-diff-stats.sh +97 -0
  13. package/skills/close-story/scripts/update-task-frontmatter.sh +103 -0
  14. package/skills/commit/SKILL.md +18 -0
  15. package/skills/distribute/CONFIG_SCHEMA.md +90 -0
  16. package/skills/distribute/SKILL.md +173 -0
  17. package/skills/distribute/scripts/check-version.sh +74 -0
  18. package/skills/distribute/scripts/commit-version-bump.sh +108 -0
  19. package/skills/distribute/scripts/distribute-changes.sh +381 -0
  20. package/skills/do/SKILL.md +314 -0
  21. package/skills/do/assets/intent-vs-diff-prompt.md +69 -0
  22. package/skills/doc/assets/path-objectives.yaml +13 -0
  23. package/skills/init/SKILL.md +4 -3
  24. package/skills/init/assets/strategy_stub_template.md +38 -0
  25. package/skills/init/scripts/init.sh +6 -1
  26. package/skills/jenga/SKILL.md +51 -2
  27. package/skills/strategy/SKILL.md +312 -0
  28. package/templates/SCRUM_BOARD_SCHEMA.md +49 -0
  29. package/skills/train/SKILL.md +0 -116
  30. package/skills/train/assets/dashboard-templates/classifiers.html +0 -106
  31. package/skills/train/assets/dashboard-templates/nlp.html +0 -102
  32. package/skills/train/assets/dashboard-templates/transformers.html +0 -98
  33. package/skills/train/assets/results-parsers/__init__.py +0 -9
  34. package/skills/train/assets/results-parsers/classifiers.py +0 -84
  35. package/skills/train/assets/results-parsers/nlp.py +0 -88
  36. package/skills/train/assets/results-parsers/reporter.py +0 -154
  37. package/skills/train/assets/results-parsers/transformers.py +0 -120
  38. package/skills/train/train_cli.py +0 -786
@@ -0,0 +1,312 @@
1
+ ---
2
+ name: strategy
3
+ description: Walk through a guided conversation to capture or update docs/STRATEGY.md — covering Vision, Value Proposition, Scope, and Target Audience — one section at a time.
4
+ metadata:
5
+ prefered_agent: developer
6
+ keywords:
7
+ - strategy
8
+ - strategic brief
9
+ - vision
10
+ - value proposition
11
+ - target audience
12
+ examples:
13
+ - "/strategy"
14
+ - "let's capture the strategy"
15
+ - "update the strategy brief"
16
+ - "write the strategy doc"
17
+ ---
18
+
19
+ # Strategy — Guided Capture of docs/STRATEGY.md
20
+
21
+ ## Entry Point
22
+
23
+ Check whether `docs/STRATEGY.md` exists before proceeding.
24
+
25
+ - If the file does **not** exist → run **Phase A — New Capture Flow** below.
26
+ - If the file **does** exist → run **Phase B — Update Existing File** below.
27
+
28
+ ---
29
+
30
+ ## Phase A — New Capture Flow
31
+
32
+ Walk the user through four sections in order. Ask **one focused question per section**. Do not present a list of questions up front. Wait for the user's answer before moving to the next section.
33
+
34
+ ### Step 1 — Vision
35
+
36
+ Ask the user exactly:
37
+
38
+ > "What is the long-term direction or ambition for this project? Describe where you want it to be in 3–5 years, or what success ultimately looks like."
39
+
40
+ Wait for the answer. Record it as the **Vision** content.
41
+
42
+ ### Step 2 — Value Proposition
43
+
44
+ Ask the user exactly:
45
+
46
+ > "What makes this project uniquely valuable? What would users or teams miss most if it didn't exist?"
47
+
48
+ Wait for the answer. Record it as the **Value Proposition** content.
49
+
50
+ ### Step 3 — Scope: In Scope
51
+
52
+ Ask the user exactly:
53
+
54
+ > "What does this project explicitly cover? List the main capabilities or areas of responsibility."
55
+
56
+ Wait for the answer. Record it as the **In Scope** content.
57
+
58
+ ### Step 4 — Scope: Out of Scope
59
+
60
+ Ask the user exactly:
61
+
62
+ > "What does this project explicitly NOT cover? (Do not include revenue models, pricing, or competitive analysis — those are always out of scope.)"
63
+
64
+ Wait for the answer. Record it as the **Out of Scope** content. The model must also always append to the Out of Scope section (regardless of whether the user mentions them):
65
+
66
+ - Revenue model, pricing strategy, or monetisation plans
67
+ - Competitive analysis or feature comparison with other tools
68
+
69
+ Do not ask the user for these — include them silently.
70
+
71
+ ### Step 5 — Target Audience
72
+
73
+ Ask the user exactly:
74
+
75
+ > "Who is this built for? Describe the primary user personas or teams."
76
+
77
+ Wait for the answer. Record it as the **Target Audience** content.
78
+
79
+ ---
80
+
81
+ ## Pre-Write Summary and Confirmation
82
+
83
+ After collecting all five pieces of content (Vision, Value Proposition, In Scope, Out of Scope, Target Audience), surface a formatted summary. Use exactly this shape:
84
+
85
+ ```
86
+ Here is what will be written to docs/STRATEGY.md:
87
+
88
+ ## Vision
89
+ <vision content>
90
+
91
+ ## Value Proposition
92
+ <value proposition content>
93
+
94
+ ## Scope
95
+
96
+ ### In Scope
97
+ <in-scope content>
98
+
99
+ ### Out of Scope
100
+ <out-of-scope content (user answer + always-excluded items)>
101
+
102
+ ## Target Audience
103
+ <target audience content>
104
+
105
+ ---
106
+
107
+ Does this look right? Type **yes** to write `docs/STRATEGY.md`, or tell me what to change.
108
+ ```
109
+
110
+ - If the user says **yes** (or any clear affirmative) → proceed to the Write step.
111
+ - If the user requests a change → re-ask the relevant question(s) using the same prompts from Phase A, collect the updated answer(s), then loop back to this summary step. Do not write the file until the user explicitly confirms.
112
+
113
+ ---
114
+
115
+ ## Write the File
116
+
117
+ Generate `docs/STRATEGY.md` using **only the content collected in the conversation**. Do not invent facts, add filler text, or leave any placeholder text in the output.
118
+
119
+ The file must follow this exact structure:
120
+
121
+ ```markdown
122
+ # Strategy Brief
123
+
124
+ > **Audience:** Investors and strategic partners. This document does not include revenue models, pricing, or competitive analysis.
125
+
126
+ ---
127
+
128
+ ## Vision
129
+
130
+ <vision content>
131
+
132
+ ---
133
+
134
+ ## Value Proposition
135
+
136
+ <value proposition content>
137
+
138
+ ---
139
+
140
+ ## Scope
141
+
142
+ ### In Scope
143
+
144
+ <in-scope content>
145
+
146
+ ### Out of Scope
147
+
148
+ <out-of-scope content>
149
+
150
+ ---
151
+
152
+ ## Target Audience
153
+
154
+ <target audience content>
155
+ ```
156
+
157
+ Write the file to `docs/STRATEGY.md`. After writing, print exactly:
158
+
159
+ ```
160
+ Written to `docs/STRATEGY.md` — captured Vision, Value Proposition, Scope, and Target Audience.
161
+ ```
162
+
163
+ ---
164
+
165
+ ## Phase B — Update Existing File
166
+
167
+ `docs/STRATEGY.md` already exists. Do **not** run the new-capture flow. Instead, walk the user through each section so they can decide what to keep and what to replace.
168
+
169
+ ### Step 1 — Read the Existing File
170
+
171
+ Read `docs/STRATEGY.md` in full. Parse out the content under each of the four sections:
172
+
173
+ - **Vision** — content under `## Vision`
174
+ - **Value Proposition** — content under `## Value Proposition`
175
+ - **In Scope** — content under `### In Scope`
176
+ - **Out of Scope** — content under `### Out of Scope`
177
+ - **Target Audience** — content under `## Target Audience`
178
+
179
+ Store these as the current values for each section. They will be used as defaults unless the user explicitly chooses to update them.
180
+
181
+ ### Step 2 — Surface Each Section One at a Time
182
+
183
+ Work through the sections in this order: Vision, Value Proposition, Scope (In + Out together), Target Audience.
184
+
185
+ For **Vision**, present exactly:
186
+
187
+ > "Here is the current Vision:
188
+ >
189
+ > ---
190
+ > <existing vision content>
191
+ > ---
192
+ >
193
+ > Do you want to keep this or update it? (keep / update)"
194
+
195
+ - If "keep" → store the existing Vision content unchanged and move on.
196
+ - If "update" → ask exactly:
197
+ > "What is the long-term direction or ambition for this project? Describe where you want it to be in 3–5 years, or what success ultimately looks like."
198
+ Wait for the answer. Store it as the new Vision content.
199
+
200
+ For **Value Proposition**, present exactly:
201
+
202
+ > "Here is the current Value Proposition:
203
+ >
204
+ > ---
205
+ > <existing value proposition content>
206
+ > ---
207
+ >
208
+ > Do you want to keep this or update it? (keep / update)"
209
+
210
+ - If "keep" → store the existing Value Proposition content unchanged and move on.
211
+ - If "update" → ask exactly:
212
+ > "What makes this project uniquely valuable? What would users or teams miss most if it didn't exist?"
213
+ Wait for the answer. Store it as the new Value Proposition content.
214
+
215
+ For **Scope**, present the In Scope and Out of Scope content together:
216
+
217
+ > "Here is the current Scope:
218
+ >
219
+ > **In Scope**
220
+ > ---
221
+ > <existing in-scope content>
222
+ > ---
223
+ >
224
+ > **Out of Scope**
225
+ > ---
226
+ > <existing out-of-scope content>
227
+ > ---
228
+ >
229
+ > Do you want to keep this or update it? (keep / update)"
230
+
231
+ - If "keep" → store the existing In Scope and Out of Scope content unchanged and move on.
232
+ - If "update" → ask the two focused scope questions from Phase A in sequence:
233
+ 1. > "What does this project explicitly cover? List the main capabilities or areas of responsibility."
234
+ Wait for the answer. Store as the new In Scope content.
235
+ 2. > "What does this project explicitly NOT cover? (Do not include revenue models, pricing, or competitive analysis — those are always out of scope.)"
236
+ Wait for the answer. Store as the new Out of Scope content. Silently append the always-excluded items (revenue model, pricing strategy, competitive analysis) as in Phase A — do not ask the user for these.
237
+
238
+ For **Target Audience**, present exactly:
239
+
240
+ > "Here is the current Target Audience:
241
+ >
242
+ > ---
243
+ > <existing target audience content>
244
+ > ---
245
+ >
246
+ > Do you want to keep this or update it? (keep / update)"
247
+
248
+ - If "keep" → store the existing Target Audience content unchanged and move on.
249
+ - If "update" → ask exactly:
250
+ > "Who is this built for? Describe the primary user personas or teams."
251
+ Wait for the answer. Store it as the new Target Audience content.
252
+
253
+ ### Step 3 — Pre-Write Summary and Confirmation
254
+
255
+ After all four sections have been decided (either kept or updated), surface a full formatted preview of the final document. Use exactly this shape:
256
+
257
+ ```
258
+ Here is what will be written to docs/STRATEGY.md:
259
+
260
+ ## Vision
261
+ <vision content>
262
+
263
+ ## Value Proposition
264
+ <value proposition content>
265
+
266
+ ## Scope
267
+
268
+ ### In Scope
269
+ <in-scope content>
270
+
271
+ ### Out of Scope
272
+ <out-of-scope content>
273
+
274
+ ## Target Audience
275
+ <target audience content>
276
+
277
+ ---
278
+
279
+ Does this look right? Type **yes** to overwrite `docs/STRATEGY.md`, or tell me what to change.
280
+ ```
281
+
282
+ - If the user says **yes** (or any clear affirmative) → proceed to the Write step.
283
+ - If the user requests a change → re-surface only the relevant section(s) using the same keep/update flow above, collect the updated answer(s), then loop back to this preview step. Do not write the file until the user explicitly confirms.
284
+
285
+ ### Step 4 — Write the File
286
+
287
+ Overwrite `docs/STRATEGY.md` using exactly the same format as Phase A output — only real content from the conversation, no placeholder text. Apply all hard constraints from the Hard Constraints block below.
288
+
289
+ After writing, print exactly:
290
+
291
+ ```
292
+ Written to `docs/STRATEGY.md` — updated Vision, Value Proposition, Scope, and Target Audience.
293
+ ```
294
+
295
+ If some sections were kept unchanged and others were updated, list only the updated sections:
296
+
297
+ ```
298
+ Written to `docs/STRATEGY.md` — updated <list of updated sections>.
299
+ ```
300
+
301
+ ---
302
+
303
+ ## Hard Constraints
304
+
305
+ Enforce these rules throughout the entire skill session — at every step, in every question, and in the written output:
306
+
307
+ 1. **Never ask about** revenue model, pricing, monetisation, competitive analysis, market positioning, financial projections, or any topic outside Vision, Value Proposition, Scope, and Target Audience.
308
+ 2. **Never write the file** without receiving explicit user confirmation ("yes" or a clear affirmative) at the pre-write summary step.
309
+ 3. **Never include placeholder text** in the written output. Every section must contain real content from the conversation.
310
+ 4. **Never add sections** beyond the four defined ones (Vision, Value Proposition, Scope, Target Audience) and their required sub-sections.
311
+ 5. **Always include** the audience callout block at the top of the written file, verbatim.
312
+ 6. **Always include** the always-excluded Out of Scope items (revenue model, pricing, competitive analysis) even if the user omits them.
@@ -86,6 +86,7 @@ dates_previously_completed: # comma-separated list, e.g. 2026-01-15, 2026-03-22
86
86
  reopened_on: # comma-separated list, e.g. 2026-02-01, 2026-04-10
87
87
  reopened_reason: # comma-separated list, e.g. "Scope expanded", "Bug found post-release"
88
88
  docs: [] # optional list of repo-relative documentation paths, e.g. ["README.md", "docs/API.md"]
89
+ epic_scope_approval: false # set to true by the human operator only when any task in this epic has execution_scope: epic
89
90
  stories:
90
91
  - E##_S##
91
92
  - E##_S##
@@ -101,6 +102,8 @@ stories:
101
102
  - <Concrete, testable criterion>
102
103
  ```
103
104
 
105
+ > **`epic_scope_approval`** — This field is **set by the human operator only**. Neither the scrum-master nor the developer agent may set it to `true`. It must be `true` before any task with `execution_scope: epic` can be executed. Its absence is equivalent to `false`.
106
+
104
107
  ### Story — `E##_S##_<slug>.md`
105
108
 
106
109
  ```markdown
@@ -150,6 +153,12 @@ reopened_on: # comma-separated list, e.g. 2026-02-01, 2026-04-10
150
153
  reopened_reason: # comma-separated list, e.g. "Scope expanded", "Bug found post-release"
151
154
  assigned_to: developer | tester | scrum-master
152
155
  docs: [] # optional list of repo-relative documentation paths, e.g. ["README.md", "docs/API.md"]
156
+ execution_scope: task # task | story | epic | inline; omit for legacy tasks (defaults to task)
157
+ needs_docs: true # boolean; omit for legacy tasks (defaults to true)
158
+ scope_rationale: "" # required when execution_scope is set; must contain a numeric/file-count claim
159
+ jenga_assigned: true # boolean; true = machine-assigned, false = human override
160
+ override_justification: "" # required when jenga_assigned: false
161
+ epic_scope_approval: false # required (as true) when execution_scope: epic; set by human operator only
153
162
  ---
154
163
 
155
164
  # Task: <Title>
@@ -164,6 +173,8 @@ docs: [] # optional list of repo-relative documentation path
164
173
  - [ ] <Verifiable criterion>
165
174
  ```
166
175
 
176
+ > **Backward compatibility:** Tasks that omit all six new fields (`execution_scope`, `needs_docs`, `scope_rationale`, `jenga_assigned`, `override_justification`, `epic_scope_approval`) are treated as `execution_scope: task` / `needs_docs: true` and are processed without error. Agents must not reject board files that lack these fields.
177
+
167
178
  ---
168
179
 
169
180
  ## Story Format Standards
@@ -181,6 +192,44 @@ docs: [] # optional list of repo-relative documentation path
181
192
  - Optionality: existing board files remain valid when `docs` is omitted.
182
193
  - Format: use a YAML list. Inline (`docs: ["README.md"]`) and expanded list styles are both valid.
183
194
 
195
+ ## Execution Scope Fields (Task)
196
+
197
+ These six fields control the execution footprint of a task within the `/jenga` and `/do` workflows. They are **optional** — omitting all six is valid and equivalent to `execution_scope: task` / `needs_docs: true`.
198
+
199
+ **`execution_scope`**
200
+ - Valid values: `task` | `story` | `epic` | `inline`
201
+ - When required: optional; omit for legacy tasks (runtime default: `task`)
202
+ - Description: defines how broadly this task's implementation touches the codebase.
203
+ - `task` — standard single-task scope (default)
204
+ - `story` — task may touch files across multiple tasks in the same story
205
+ - `epic` — task may touch files across stories; requires `epic_scope_approval: true` on the parent epic
206
+ - `inline` — trivial change (e.g. config tweak, comment, schema doc); no execution plan or summary document is needed
207
+
208
+ **`needs_docs`**
209
+ - Valid values: `true` | `false`
210
+ - When required: optional; omit for legacy tasks (runtime default: `true`)
211
+ - Description: when `false`, the developer agent skips writing an execution plan and execution summary for this task. Automatically implied as `false` when `execution_scope: inline`.
212
+
213
+ **`scope_rationale`**
214
+ - Valid values: any non-empty string; must include a measurable claim (e.g. a file count or line count)
215
+ - When required: required when `execution_scope` is explicitly set to any value
216
+ - Description: a brief justification explaining why the chosen scope is appropriate. Validators reject a blank value when `execution_scope` is present.
217
+
218
+ **`jenga_assigned`**
219
+ - Valid values: `true` | `false`
220
+ - When required: optional; omit for legacy tasks (runtime default: `true`)
221
+ - Description: indicates whether the execution scope was assigned by the `/jenga` orchestrator (`true`) or overridden by a human (`false`). When `false`, `override_justification` is required.
222
+
223
+ **`override_justification`**
224
+ - Valid values: any non-empty string
225
+ - When required: required when `jenga_assigned: false`
226
+ - Description: explains why the human operator overrode the machine-assigned scope. Must be non-empty when present.
227
+
228
+ **`epic_scope_approval`**
229
+ - Valid values: `true` | `false`
230
+ - When required: required on the **epic** frontmatter (as `true`) before any task with `execution_scope: epic` may be executed; the task frontmatter carries this field for reference only
231
+ - Description: the authoritative value is always read from the **epic** frontmatter, not the task. This field is **set by the human operator only** — neither the scrum-master nor the developer agent may set it to `true`. Its absence on the epic is equivalent to `false`.
232
+
184
233
  These rules apply to every story file written or amended by the **Scrum Master**. They are enforced at write time by `scripts/validate-story-format.sh`.
185
234
 
186
235
  ### Acceptance Criteria (`## Acceptance Criteria`)
@@ -1,116 +0,0 @@
1
- ---
2
- name: train
3
- description: Scaffold and run ML training jobs. Use `new <type> <job-name>` to scaffold a job from a template, or `run <job-dir>` to execute the two-phase validate → train pipeline on an existing job.
4
- keywords:
5
- - train
6
- - ml training
7
- - model training
8
- - machine learning
9
- - fine-tune
10
- examples:
11
- - "scaffold a new training job"
12
- - "train a classifier model"
13
- ---
14
-
15
- # Train — ML Training Job Orchestrator
16
-
17
- ## Usage
18
-
19
- ```
20
- python skills/train/train_cli.py <subcommand> [args]
21
-
22
- Subcommands:
23
- new <type> <job-name> Scaffold a new training job from a template
24
- run <job-dir> Execute validate → train pipeline on an existing job
25
-
26
- Types: classifiers, transformers, nlp
27
- ```
28
-
29
- If an unrecognised subcommand is given, argparse prints the usage message above and
30
- exits with a non-zero code.
31
-
32
- ## Subcommands
33
-
34
- ### `/train new <type> <job-name>`
35
- Scaffold a new training job from the appropriate template (positional / scriptable mode).
36
-
37
- **Supported types:** `classifiers`, `transformers`, `nlp`
38
-
39
- **Optional flags:**
40
- - `--model <value>` — Override the model name/type in `config.yaml`
41
- - `--epochs <n>` — Override the number of training epochs/iterations
42
- - `--batch-size <n>` — Override the batch size
43
-
44
- **Steps:**
45
- 1. Validate `<type>` is one of: `classifiers`, `transformers`, `nlp`
46
- 2. Copy the template directory from `.training/template/<type>/` to `jobs/<job-name>/`
47
- - If `jobs/<job-name>/` already exists, auto-suffix: `jobs/<job-name>-1/`, `jobs/<job-name>-2/`, etc.
48
- 3. Apply any CLI flag overrides to the scaffolded `config.yaml`
49
- 4. Generate `start.sh` in the job directory (executable)
50
- 5. Read the `workflow:` block from `config.yaml` and print next-step instructions
51
- 6. Print: `✅ Scaffolded job '<job-name>' from template '<type>' at jobs/<job-name>/`
52
-
53
- ### `/train new --interactive [--full]`
54
- Launch a guided wizard that scaffolds the job **and** configures `config.yaml` interactively.
55
-
56
- **Flags:**
57
- - `--interactive` / `-i` — Enable the wizard. Job name and type are collected via prompts.
58
- - `--full` — Expand the wizard to prompt for **every** configurable field (not just the critical ones).
59
-
60
- **Wizard flow:**
61
- 1. Prompt for job name (freeform, used as directory name under `jobs/`)
62
- 2. Prompt for job type (`classifiers`, `transformers`, or `nlp`)
63
- 3. Scaffold the job directory from the template
64
- 4. Prompt for critical config fields (2–3 key model params + data path per type):
65
- - **classifiers**: train file, target column, model algorithm, n_estimators, test split
66
- - **transformers**: model name, task, train file, epochs, learning rate
67
- - **nlp**: model name, task, train file, iterations, learning rate
68
- 5. With `--full`: additionally prompt for every remaining field in `config.yaml`
69
- 6. Write user responses back into the scaffolded `config.yaml`
70
-
71
- Each prompt shows the field name, its current default, and an inline description hint.
72
-
73
- **Ctrl+C handling:**
74
- Pressing `Ctrl+C` at any point during the wizard interrupts cleanly. If a scaffold directory has already been created, the user is asked:
75
- ```
76
- Delete scaffolded directory '<path>'? [y/N]
77
- ```
78
- Answering `y` removes the directory; `N` (or Enter) keeps it.
79
-
80
- ### `/train run <job-dir>`
81
- Execute the two-phase pre-flight → smoke test pipeline on an existing job directory.
82
-
83
- **Steps:**
84
-
85
- #### Phase A — Pre-flight validation
86
- 1. Print: `[pre-flight] Running validate.py in <job-dir>...`
87
- 2. Run `validate.py` as a subprocess inside `<job-dir>`:
88
- ```bash
89
- cd <job-dir> && python validate.py
90
- ```
91
- 3. Stream all stdout/stderr output to the terminal in real time.
92
- 4. If `validate.py` exits non-zero:
93
- - Print: `❌ [pre-flight] FAILED — halting before smoke test.`
94
- - Surface the full error output.
95
- - **Stop here. Do not proceed to Phase B.**
96
- 5. If `validate.py` exits zero:
97
- - Print: `✅ [pre-flight] Passed.`
98
-
99
- #### Phase B — Smoke test
100
- 1. Print: `[smoke test] Running train.py --smoke in <job-dir>...`
101
- 2. Run `train.py --smoke` as a subprocess inside `<job-dir>`:
102
- ```bash
103
- cd <job-dir> && python train.py --smoke
104
- ```
105
- 3. Stream all stdout/stderr output to the terminal in real time.
106
- 4. If `train.py` exits non-zero:
107
- - Print: `❌ [smoke test] FAILED.`
108
- - Surface the full error output.
109
- 5. If `train.py` exits zero:
110
- - Print: `✅ [smoke test] Passed. Job '<job-dir>' completed successfully.`
111
-
112
- ## Notes
113
- - `train.py` and `validate.py` must have no awareness of each other — the `/train` skill owns the gate.
114
- - Both subprocess calls must stream output (not buffer) so the user sees progress in real time.
115
- - Phase B is **never** invoked if Phase A fails.
116
- - The `--interactive` wizard requires `pyyaml` (`pip install pyyaml`) to patch `config.yaml`.
@@ -1,106 +0,0 @@
1
- <!DOCTYPE html>
2
- <html lang="en">
3
- <head>
4
- <meta charset="UTF-8" />
5
- <meta name="viewport" content="width=device-width, initial-scale=1.0" />
6
- <title>Classifier Training Dashboard — {{job_name}}</title>
7
- <style>
8
- *, *::before, *::after { box-sizing: border-box; margin: 0; padding: 0; }
9
- body { font-family: system-ui, -apple-system, sans-serif; background: #f5f7fa; color: #1a1a2e; padding: 2rem; }
10
- header { margin-bottom: 2rem; }
11
- header h1 { font-size: 1.6rem; font-weight: 700; }
12
- header p { color: #555; margin-top: .25rem; font-size: .9rem; }
13
- .card { background: #fff; border-radius: 10px; box-shadow: 0 2px 8px rgba(0,0,0,.08); padding: 1.5rem; margin-bottom: 1.5rem; }
14
- .card h2 { font-size: 1rem; font-weight: 600; color: #333; margin-bottom: 1rem; text-transform: uppercase; letter-spacing: .05em; }
15
- .metrics-grid { display: grid; grid-template-columns: repeat(auto-fit, minmax(150px, 1fr)); gap: 1rem; }
16
- .metric-box { background: #f0f4ff; border-radius: 8px; padding: 1rem; text-align: center; }
17
- .metric-box .value { font-size: 2rem; font-weight: 700; color: #3b5bdb; }
18
- .metric-box .label { font-size: .8rem; color: #666; margin-top: .25rem; }
19
- .bar-chart { display: flex; flex-direction: column; gap: .6rem; }
20
- .bar-row { display: flex; align-items: center; gap .5rem; }
21
- .bar-label { width: 100px; font-size: .85rem; color: #444; flex-shrink: 0; }
22
- .bar-track { flex: 1; background: #e9ecef; border-radius: 4px; height: 20px; position: relative; }
23
- .bar-fill { height: 100%; border-radius: 4px; background: #3b5bdb; transition: width .4s; }
24
- .bar-val { width: 50px; text-align: right; font-size: .8rem; color: #444; }
25
- table { width: 100%; border-collapse: collapse; }
26
- th, td { padding: .6rem 1rem; text-align: left; font-size: .9rem; border-bottom: 1px solid #eee; }
27
- th { background: #f8f9fa; font-weight: 600; color: #555; }
28
- footer { text-align: center; font-size: .75rem; color: #999; margin-top: 2rem; }
29
- </style>
30
- </head>
31
- <body>
32
-
33
- <header>
34
- <h1>🤖 Classifier Training Dashboard</h1>
35
- <p>Job: <strong>{{job_name}}</strong> &nbsp;|&nbsp; Model: <strong>{{model_type}}</strong></p>
36
- </header>
37
-
38
- <div class="card">
39
- <h2>Key Metrics</h2>
40
- <div class="metrics-grid">
41
- <div class="metric-box">
42
- <div class="value">{{accuracy}}</div>
43
- <div class="label">Accuracy</div>
44
- </div>
45
- <div class="metric-box">
46
- <div class="value">{{f1_score}}</div>
47
- <div class="label">F1 Score</div>
48
- </div>
49
- <div class="metric-box">
50
- <div class="value">{{precision}}</div>
51
- <div class="label">Precision</div>
52
- </div>
53
- <div class="metric-box">
54
- <div class="value">{{recall}}</div>
55
- <div class="label">Recall</div>
56
- </div>
57
- </div>
58
- </div>
59
-
60
- <div class="card">
61
- <h2>Metric Comparison</h2>
62
- <div class="bar-chart">
63
- <div class="bar-row">
64
- <span class="bar-label">Accuracy</span>
65
- <div class="bar-track"><div class="bar-fill" style="width: calc({{accuracy}} * 100%)"></div></div>
66
- <span class="bar-val">{{accuracy}}</span>
67
- </div>
68
- <div class="bar-row">
69
- <span class="bar-label">F1 Score</span>
70
- <div class="bar-track"><div class="bar-fill" style="width: calc({{f1_score}} * 100%)"></div></div>
71
- <span class="bar-val">{{f1_score}}</span>
72
- </div>
73
- <div class="bar-row">
74
- <span class="bar-label">Precision</span>
75
- <div class="bar-track"><div class="bar-fill" style="width: calc({{precision}} * 100%)"></div></div>
76
- <span class="bar-val">{{precision}}</span>
77
- </div>
78
- <div class="bar-row">
79
- <span class="bar-label">Recall</span>
80
- <div class="bar-track"><div class="bar-fill" style="width: calc({{recall}} * 100%)"></div></div>
81
- <span class="bar-val">{{recall}}</span>
82
- </div>
83
- </div>
84
- </div>
85
-
86
- <div class="card">
87
- <h2>Summary Table</h2>
88
- <table>
89
- <thead>
90
- <tr><th>Metric</th><th>Value</th></tr>
91
- </thead>
92
- <tbody>
93
- <tr><td>Accuracy</td><td>{{accuracy}}</td></tr>
94
- <tr><td>F1 Score</td><td>{{f1_score}}</td></tr>
95
- <tr><td>Precision</td><td>{{precision}}</td></tr>
96
- <tr><td>Recall</td><td>{{recall}}</td></tr>
97
- <tr><td>Model Type</td><td>{{model_type}}</td></tr>
98
- <tr><td>Job Name</td><td>{{job_name}}</td></tr>
99
- </tbody>
100
- </table>
101
- </div>
102
-
103
- <footer>Generated by JengaAgent /train skill &nbsp;|&nbsp; {{job_name}}</footer>
104
-
105
- </body>
106
- </html>