@jenga-ai/agent 1.0.0 → 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (38) hide show
  1. package/README.md +28 -10
  2. package/agents/scrum-master.md +75 -0
  3. package/mcp/router/embedder.js +1 -1
  4. package/mcp/training_runner/index.js +239 -0
  5. package/mcp/training_runner/package-lock.json +1065 -0
  6. package/mcp/training_runner/package.json +15 -0
  7. package/package.json +13 -11
  8. package/skills/close-story/SKILL.md +203 -0
  9. package/skills/close-story/scripts/check-story-closeable.sh +195 -0
  10. package/skills/close-story/scripts/compute-scope-divergence.sh +128 -0
  11. package/skills/close-story/scripts/extract-diff-stats.sh +48 -0
  12. package/skills/close-story/scripts/extract-task-diff-stats.sh +97 -0
  13. package/skills/close-story/scripts/update-task-frontmatter.sh +103 -0
  14. package/skills/commit/SKILL.md +18 -0
  15. package/skills/distribute/CONFIG_SCHEMA.md +90 -0
  16. package/skills/distribute/SKILL.md +173 -0
  17. package/skills/distribute/scripts/check-version.sh +74 -0
  18. package/skills/distribute/scripts/commit-version-bump.sh +108 -0
  19. package/skills/distribute/scripts/distribute-changes.sh +381 -0
  20. package/skills/do/SKILL.md +314 -0
  21. package/skills/do/assets/intent-vs-diff-prompt.md +69 -0
  22. package/skills/doc/assets/path-objectives.yaml +13 -0
  23. package/skills/init/SKILL.md +4 -3
  24. package/skills/init/assets/strategy_stub_template.md +38 -0
  25. package/skills/init/scripts/init.sh +6 -1
  26. package/skills/jenga/SKILL.md +51 -2
  27. package/skills/strategy/SKILL.md +312 -0
  28. package/templates/SCRUM_BOARD_SCHEMA.md +49 -0
  29. package/skills/train/SKILL.md +0 -116
  30. package/skills/train/assets/dashboard-templates/classifiers.html +0 -106
  31. package/skills/train/assets/dashboard-templates/nlp.html +0 -102
  32. package/skills/train/assets/dashboard-templates/transformers.html +0 -98
  33. package/skills/train/assets/results-parsers/__init__.py +0 -9
  34. package/skills/train/assets/results-parsers/classifiers.py +0 -84
  35. package/skills/train/assets/results-parsers/nlp.py +0 -88
  36. package/skills/train/assets/results-parsers/reporter.py +0 -154
  37. package/skills/train/assets/results-parsers/transformers.py +0 -120
  38. package/skills/train/train_cli.py +0 -786
package/README.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # JengaAgent
2
2
 
3
- **A structured multi-agent development workflow built on [Claude Code](https://docs.anthropic.com/en/docs/claude-code).** Three specialised AI agents — Scrum Master, Developer, and Tester — collaborate through a shared scrum board, an event-driven trigger queue, and 28 slash-command skills to take a project from idea to verified, committed code — across as many sessions as it takes.
3
+ **A structured multi-agent development workflow that works with any AI agent or AI-native IDE.** Three specialised AI agents — Scrum Master, Developer, and Tester — collaborate through a shared scrum board, an event-driven trigger queue, and 28 slash-command skills to take a project from idea to verified, committed code — across as many sessions as it takes.
4
4
 
5
5
  > 📖 **Full reference:** [project/.wiki/documentation.md](project/.wiki/documentation.md) | [Intro Guide](project/.wiki/intro-guide.md)
6
6
 
@@ -12,7 +12,7 @@ Without a framework like JengaAgent, AI-assisted development has serious structu
12
12
 
13
13
  | Problem | Reality |
14
14
  |---|---|
15
- | **AI has no session memory** | Each Claude Code session starts from scratch — no awareness of open tasks, past decisions, or what was already tested |
15
+ | **AI has no session memory** | Every AI agent session starts from scratch — no awareness of open tasks, past decisions, or what was already tested |
16
16
  | **No role separation** | The AI writes *and* "tests" code in the same context, leading to hallucinated test results and self-affirming bugs |
17
17
  | **No structured planning** | Work happens ad-hoc — no Epic → Story → Task hierarchy to organise or track progress |
18
18
  | **No handoff protocol** | Switching from implementing to testing means manually re-explaining context every time |
@@ -30,11 +30,11 @@ JengaAgent solves each of these with structure: persistent board state, strict a
30
30
  **Without JengaAgent:**
31
31
  ```
32
32
  You: "Add user authentication"
33
- Claude: [writes auth code, declares it works, session ends]
33
+ AI agent: [writes auth code, declares it works, session ends]
34
34
 
35
35
  Next session:
36
36
  You: "What's the status of auth?"
37
- Claude: "I don't have context from the previous session."
37
+ AI agent: "I don't have context from the previous session."
38
38
  ```
39
39
 
40
40
  **With JengaAgent:**
@@ -136,7 +136,11 @@ Each agent is defined in `.agents/agents/`. They communicate exclusively through
136
136
 
137
137
  ## Getting Started
138
138
 
139
- **Prerequisites:** [Claude Code](https://docs.anthropic.com/en/docs/claude-code) installed and configured.
139
+ **Prerequisites:** An AI agent or AI-native IDE. Supported platforms include:
140
+ - [Claude Code](https://docs.anthropic.com/en/docs/claude-code)
141
+ - GitHub Copilot (VS Code extension or CLI)
142
+ - [Warp](https://www.warp.dev/)
143
+ - Codex CLI
140
144
 
141
145
  **Install from npm:**
142
146
  ```sh
@@ -155,9 +159,24 @@ Run `/status` at any time to see where the project stands.
155
159
 
156
160
  ---
157
161
 
162
+ ## Supported Platforms
163
+
164
+ JengaAgent has been tested with:
165
+
166
+ | Platform | Type |
167
+ |---|---|
168
+ | [Claude Code](https://docs.anthropic.com/en/docs/claude-code) | AI coding agent |
169
+ | [GitHub Copilot](https://github.com/features/copilot) | AI coding assistant |
170
+ | [Warp](https://www.warp.dev/) | AI-native terminal |
171
+ | Codex CLI | AI coding agent |
172
+
173
+ The framework is platform-agnostic by design — any AI agent that can read Markdown files and execute slash commands can use it.
174
+
175
+ ---
176
+
158
177
  ## Skills (Slash Commands)
159
178
 
160
- Skills live in `.agents/skills/<name>/SKILL.md`. Invoke with `/<name>` in any Claude Code session.
179
+ Skills live in `.agents/skills/<name>/SKILL.md`. Invoke with `/<name>` in your AI agent or IDE's command interface.
161
180
 
162
181
  ### Setup & Planning
163
182
 
@@ -216,11 +235,10 @@ Skills live in `.agents/skills/<name>/SKILL.md`. Invoke with `/<name>` in any Cl
216
235
 
217
236
  JengaAgent can propagate its workflow files to other projects on your machine via `/distribute`.
218
237
 
219
- 1. **Register consumer projects** — place `jenga.config.json` at each consumer root (copy from `templates/JENGA_CONFIG_TEMPLATE.json`).
220
- 2. **Add search paths** — add absolute directory paths to `.jenga_paths` (one per line, git-ignored).
221
- 3. **Run `/distribute`** — choose `major`, `minor`, `patch`, or `amend` release type; the skill handles versioning and syncs all registered projects.
238
+ 1. **Register consumer projects** — add each consuming project to `distribute.config.json` at the repo root (or pass a path directly: `/distribute /path/to/project`).
239
+ 2. **Run `/distribute`** — choose `major`, `minor`, `patch`, or `amend` release type; the skill handles versioning, dry-run preview, file copy, and a version bump commit.
222
240
 
223
- To exclude specific files per consumer project, add a `.jenga_ignore` at the consumer root (never overwritten by distribute).
241
+ Each consuming project holds a `jenga.config.json` tracking the distributed version, last distribution date, and source. To exclude specific files per consumer project, add a `.jenga_ignore` at the consumer root (never overwritten by distribute).
224
242
 
225
243
  ---
226
244
 
@@ -118,6 +118,81 @@ Used for smaller, more technical units of work — typically a sub-item within a
118
118
 
119
119
  ---
120
120
 
121
+ ## Execution Scope Assignment
122
+
123
+ When breaking down a story into tasks, assign `execution_scope` to each task using the heuristics below. Read all numeric thresholds from `project/configs/scope-thresholds.json` at breakdown time — do not embed literal values in these instructions. The relevant fields are `inline_max_files`, `inline_max_lines`, and `story_max_files`.
124
+
125
+ ### `inline` scope
126
+
127
+ Assign `inline` when **all** of the following are true:
128
+ - The task touches exactly `inline_max_files` file (per `project/configs/scope-thresholds.json`)
129
+ - No new tests are required
130
+ - The change is purely additive or config-level (no logic branches introduced)
131
+ - The estimated diff is `inline_max_lines` lines or fewer (per `project/configs/scope-thresholds.json`)
132
+
133
+ `needs_docs` for every `inline`-scoped task is always `false`.
134
+
135
+ ### `story` scope
136
+
137
+ Assign `story` when **all** of the following are true:
138
+ - All tasks in the story operate in the same module or directory
139
+ - No cross-story dependencies exist
140
+ - The total file count across all tasks in the story is fewer than `story_max_files` (per `project/configs/scope-thresholds.json`)
141
+ - The mandatory contention check passes (see below)
142
+
143
+ **Mandatory contention check (required before assigning `story` scope):** Before assigning `story` scope to any task, confirm that no two tasks in the story write to the same shared infrastructure file (e.g. `package.json`, `settings.json`, `pyproject.toml`, `distribute.config.json`). This check is required — it is not optional.
144
+
145
+ If contention exists between any two tasks, downgrade **both** conflicting tasks to `task` scope and document the conflict in each task's `scope_rationale` (e.g. `"downgraded from story: contention on package.json with T02"`). Do not assign `story` scope to either conflicting task.
146
+
147
+ ### `task` scope (default)
148
+
149
+ Use `task` scope when **any** of the following are true:
150
+ - Branching logic or non-trivial architecture is involved
151
+ - Tester validation is sensitive to the implementation approach
152
+ - Shared infrastructure is touched (e.g. `package.json`, `settings.json`)
153
+ - Cross-story dependencies exist
154
+ - The mandatory contention check fails for `story` scope
155
+ - Uncertainty makes a more optimistic scope assignment unjustifiable
156
+
157
+ When in doubt, default to `task`. `task` is the safe choice and imposes no penalty.
158
+
159
+ ### `epic` scope
160
+
161
+ **Never assign `epic` scope autonomously.** If a task appears to require epic-level scope, do **not** set `execution_scope: epic` in the frontmatter. Instead:
162
+ 1. Assign `execution_scope: task` in the frontmatter.
163
+ 2. Add a note in the story description or `scope_rationale` explaining that this task may require epic-level scope and why, directed at the human operator.
164
+ 3. The human operator sets `epic_scope_approval: true` when they are ready to authorise it. Never set `epic_scope_approval: true` autonomously.
165
+
166
+ ---
167
+
168
+ ## `needs_docs` Assessment
169
+
170
+ `needs_docs` is assessed **independently** from `execution_scope`. Do not derive one from the other (except for `inline`, which always sets `needs_docs: false`).
171
+
172
+ **Assign `needs_docs: false` when:**
173
+ - The task is `inline` scope (always false)
174
+ - The acceptance criteria are binary and self-evident from reading the diff (e.g. "add field X to schema")
175
+ - No non-obvious architectural decision is required
176
+
177
+ **Assign `needs_docs: true` when:**
178
+ - A non-obvious architectural decision is required
179
+ - Multiple valid implementation approaches exist and the chosen one needs justification
180
+ - The tester cannot verify correctness without understanding the implementation intent
181
+
182
+ ---
183
+
184
+ ## `scope_rationale` Requirement
185
+
186
+ Every task with `execution_scope` set in frontmatter **must** include a `scope_rationale` string. The rationale must reference at least one measurable or file-count criterion. Generic statements (e.g. "this task is small") are not acceptable.
187
+
188
+ **Acceptable example:** `"touches 1 file (SCRUM_BOARD_SCHEMA.md); purely additive schema documentation change, estimated under 30 lines"`
189
+
190
+ **Not acceptable:** `"this is a small change"`
191
+
192
+ If a measurable rationale cannot be constructed, default to `execution_scope: task` rather than guessing at a more optimistic scope.
193
+
194
+ ---
195
+
121
196
  ## Workflow
122
197
 
123
198
  ### 1. Intake & Mapping
@@ -1,4 +1,4 @@
1
- import { pipeline } from "@xenova/transformers";
1
+ import { pipeline } from "@huggingface/transformers";
2
2
 
3
3
  let _pipe = null;
4
4
 
@@ -0,0 +1,239 @@
1
+ #!/usr/bin/env node
2
+ /**
3
+ * training_runner — MCP server
4
+ *
5
+ * Exposes a single tool: run_training_job
6
+ * - Validates the job directory and config.yaml
7
+ * - Checks confirm_before_run flag
8
+ * - Executes bash start.sh, collecting stdout/stderr
9
+ * - Returns a structured result including exit_code, duration, and results.json contents
10
+ */
11
+ import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
12
+ import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
13
+ import { z } from "zod";
14
+ import { existsSync, readFileSync } from "fs";
15
+ import { join, resolve } from "path";
16
+ import { spawn } from "child_process";
17
+ import { load as yamlLoad } from "js-yaml";
18
+
19
+ const server = new McpServer({
20
+ name: "training_runner",
21
+ version: "1.0.0",
22
+ });
23
+
24
+ // ---------------------------------------------------------------------------
25
+ // Helpers
26
+ // ---------------------------------------------------------------------------
27
+
28
+ /**
29
+ * Load and parse config.yaml from the job directory.
30
+ * Returns { config, error } — error is a string if loading failed.
31
+ */
32
+ function loadConfig(jobDir) {
33
+ const configPath = join(jobDir, "input", "config.yaml");
34
+ if (!existsSync(configPath)) {
35
+ return { config: null, error: `config.yaml not found at: ${configPath}` };
36
+ }
37
+ try {
38
+ const raw = readFileSync(configPath, "utf8");
39
+ const config = yamlLoad(raw) || {};
40
+ return { config, error: null };
41
+ } catch (err) {
42
+ return { config: null, error: `Failed to parse config.yaml: ${err.message}` };
43
+ }
44
+ }
45
+
46
+ /**
47
+ * Validate required config fields.
48
+ * Returns an array of missing field names (empty = all present).
49
+ */
50
+ function validateRequiredFields(config) {
51
+ const missing = [];
52
+ const model = config.model || {};
53
+
54
+ // Every job type requires at least one of model.type or model.name
55
+ if (!model.type && !model.name) {
56
+ missing.push("model.type (classifiers) or model.name (transformers/nlp)");
57
+ }
58
+
59
+ return missing;
60
+ }
61
+
62
+ /**
63
+ * Run `bash start.sh` in jobDir; collect output lines and exit code.
64
+ * Returns a Promise<{ lines: string[], exitCode: number }>.
65
+ */
66
+ function runStartSh(jobDir) {
67
+ return new Promise((resolve_) => {
68
+ const lines = [];
69
+ const proc = spawn("bash", ["start.sh"], {
70
+ cwd: jobDir,
71
+ env: process.env,
72
+ });
73
+
74
+ const collect = (data) => {
75
+ const text = data.toString();
76
+ text.split("\n").forEach((line) => {
77
+ if (line !== "") lines.push(line);
78
+ });
79
+ };
80
+
81
+ proc.stdout.on("data", collect);
82
+ proc.stderr.on("data", collect);
83
+
84
+ proc.on("close", (exitCode) => {
85
+ resolve_({ lines, exitCode: exitCode ?? -1 });
86
+ });
87
+
88
+ proc.on("error", (err) => {
89
+ lines.push(`[runner error] ${err.message}`);
90
+ resolve_({ lines, exitCode: -1 });
91
+ });
92
+ });
93
+ }
94
+
95
+ /**
96
+ * Read results.json from the job directory, if it exists.
97
+ * Returns the parsed object or null.
98
+ */
99
+ function readResultsJson(jobDir) {
100
+ const resultsPath = join(jobDir, "results.json");
101
+ if (!existsSync(resultsPath)) return null;
102
+ try {
103
+ const raw = readFileSync(resultsPath, "utf8");
104
+ return JSON.parse(raw);
105
+ } catch {
106
+ return null;
107
+ }
108
+ }
109
+
110
+ // ---------------------------------------------------------------------------
111
+ // Tool: run_training_job
112
+ // ---------------------------------------------------------------------------
113
+ server.tool(
114
+ "run_training_job",
115
+ "Validate and run an ML training job by executing bash start.sh in the job directory.",
116
+ {
117
+ job_dir: z
118
+ .string()
119
+ .describe("Absolute or relative path to the job directory (must contain input/config.yaml and start.sh)."),
120
+ confirm: z
121
+ .boolean()
122
+ .optional()
123
+ .default(false)
124
+ .describe(
125
+ "Set to true to confirm execution when confirm_before_run is enabled in config.yaml."
126
+ ),
127
+ },
128
+ async ({ job_dir, confirm }) => {
129
+ const jobDir = resolve(job_dir);
130
+
131
+ // --- T02: Directory validation ---
132
+ if (!existsSync(jobDir)) {
133
+ return {
134
+ content: [
135
+ {
136
+ type: "text",
137
+ text: `❌ Job directory not found: ${jobDir}`,
138
+ },
139
+ ],
140
+ };
141
+ }
142
+
143
+ // --- T02: config.yaml validation ---
144
+ const { config, error: configError } = loadConfig(jobDir);
145
+ if (configError) {
146
+ return {
147
+ content: [{ type: "text", text: `❌ ${configError}` }],
148
+ };
149
+ }
150
+
151
+ // --- T02: Required field validation ---
152
+ const missingFields = validateRequiredFields(config);
153
+ if (missingFields.length > 0) {
154
+ return {
155
+ content: [
156
+ {
157
+ type: "text",
158
+ text: `❌ config.yaml is missing required fields:\n${missingFields.map((f) => ` - ${f}`).join("\n")}`,
159
+ },
160
+ ],
161
+ };
162
+ }
163
+
164
+ // --- T02: start.sh must exist ---
165
+ const startSh = join(jobDir, "start.sh");
166
+ if (!existsSync(startSh)) {
167
+ return {
168
+ content: [{ type: "text", text: `❌ start.sh not found in ${jobDir}. Re-scaffold with /train new to generate it.` }],
169
+ };
170
+ }
171
+
172
+ // --- T04: confirm_before_run gate ---
173
+ const workflow = config.workflow || {};
174
+ if (workflow.confirm_before_run === true && !confirm) {
175
+ return {
176
+ content: [
177
+ {
178
+ type: "text",
179
+ text: [
180
+ `⚠️ This job has confirm_before_run: true in its config.yaml.`,
181
+ ``,
182
+ `Job directory: ${jobDir}`,
183
+ ``,
184
+ `To proceed, re-invoke this tool with confirm: true.`,
185
+ ].join("\n"),
186
+ },
187
+ ],
188
+ };
189
+ }
190
+
191
+ // --- T03: Execute bash start.sh ---
192
+ const startTime = Date.now();
193
+ const { lines, exitCode } = await runStartSh(jobDir);
194
+ const durationSeconds = ((Date.now() - startTime) / 1000).toFixed(2);
195
+
196
+ // --- T04 / E01_S04_T02: Read results.json if present ---
197
+ const results = readResultsJson(jobDir);
198
+
199
+ // Build response
200
+ const outputText = lines.join("\n");
201
+ const status = exitCode === 0 ? "✅ Completed" : `❌ Failed (exit code ${exitCode})`;
202
+
203
+ const summary = [
204
+ `${status}`,
205
+ `Job directory : ${jobDir}`,
206
+ `Duration : ${durationSeconds}s`,
207
+ `Exit code : ${exitCode}`,
208
+ ``,
209
+ `--- Output ---`,
210
+ outputText || "(no output)",
211
+ ];
212
+
213
+ if (results !== null) {
214
+ summary.push(``, `--- results.json ---`, JSON.stringify(results, null, 2));
215
+ }
216
+
217
+ // Structured result as JSON (appended as second content block)
218
+ const structuredResult = {
219
+ exit_code: exitCode,
220
+ job_dir: jobDir,
221
+ duration_seconds: parseFloat(durationSeconds),
222
+ output: lines,
223
+ results: results,
224
+ };
225
+
226
+ return {
227
+ content: [
228
+ { type: "text", text: summary.join("\n") },
229
+ { type: "text", text: JSON.stringify(structuredResult, null, 2) },
230
+ ],
231
+ };
232
+ }
233
+ );
234
+
235
+ // ---------------------------------------------------------------------------
236
+ // Start server
237
+ // ---------------------------------------------------------------------------
238
+ const transport = new StdioServerTransport();
239
+ await server.connect(transport);