karajan-code 4.13.0 → 4.15.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (46) hide show
  1. package/README.md +2 -1
  2. package/docs/README.es.md +2 -1
  3. package/package.json +3 -3
  4. package/packages/ai-trash/src/cli.js +3 -1
  5. package/packages/hu-board/public/utils/formatters.js +1 -1
  6. package/packages/hu-board/src/cleanup-zombies.js +0 -1
  7. package/packages/hu-board/src/command-runner.js +0 -1
  8. package/packages/hu-board/src/config-yaml.js +1 -2
  9. package/packages/hu-board/src/ephemeral-cleaner.js +0 -2
  10. package/packages/hu-board/src/plan-mutations.js +0 -1
  11. package/packages/hu-board/src/preflight.js +1 -1
  12. package/packages/hu-board/src/routes/api.js +0 -15
  13. package/packages/hu-board/src/server.js +0 -1
  14. package/packages/hu-board/src/sync.js +1 -1
  15. package/packages/hu-board/src/token-store.js +0 -1
  16. package/packages/hu-board/src/zombie-reaper.js +0 -1
  17. package/src/agents/agy-agent.js +62 -0
  18. package/src/agents/claude-agent.js +3 -2
  19. package/src/agents/codex-agent.js +3 -1
  20. package/src/agents/index.js +4 -0
  21. package/src/agents/kimi-agent.js +45 -0
  22. package/src/checks/ai-surface.js +5 -2
  23. package/src/checks/binaries.js +29 -1
  24. package/src/cli/advanced-commands.js +1 -1
  25. package/src/cli/register-pipeline.js +26 -4
  26. package/src/commands/check.js +9 -1
  27. package/src/commands/harden.js +1 -1
  28. package/src/commands/tournament.js +85 -0
  29. package/src/config/overrides.js +8 -1
  30. package/src/harden/config-engine.js +27 -0
  31. package/src/harden/config-templates.js +8 -3
  32. package/src/harden/workflow-engine.js +11 -1
  33. package/src/harden/workflow-templates.js +41 -22
  34. package/src/prompts/diff-clip.js +27 -0
  35. package/src/prompts/reviewer.js +8 -2
  36. package/src/prompts/split-signal.js +82 -0
  37. package/src/review/one-shot-review.js +44 -10
  38. package/src/review/reviewer-fallback.js +125 -0
  39. package/src/roles/reviewer-role.js +9 -8
  40. package/src/tournament/crown.js +88 -0
  41. package/src/tournament/judge.js +151 -0
  42. package/src/tournament/run.js +142 -0
  43. package/src/tournament/scoreboard.js +124 -0
  44. package/src/utils/agent-detect.js +31 -2
  45. package/src/utils/os-detect.js +5 -0
  46. package/src/utils/role-env.js +2 -1
package/README.md CHANGED
@@ -22,7 +22,7 @@
22
22
 
23
23
  ---
24
24
 
25
- Your AI agent (Claude Code, Codex, Gemini CLI, Cursor…) writes the code. **Karajan governs how it happens**: it installs a method your agent follows on every task, and enforces it with git gates that make a false green structurally impossible.
25
+ Your AI agent (Claude Code, Codex, Copilot, Antigravity, Cursor…) writes the code. **Karajan governs how it happens**: it installs a method your agent follows on every task, and enforces it with git gates that make a false green structurally impossible.
26
26
 
27
27
  - **RAG before assuming** — `kj rag query` answers what your codebase does; no agent guesses. The install wires it as a native MCP tool (`kj_rag_query`) so querying the index is the agent's cheapest path. Works out of the box: local Ollama, or the built-in ONNX embedder when nothing can be installed; cloud embedders require an explicit sensitivity declaration and PII-redact every chunk. A distilled engineering canon rides along: `kj rag query --library` serves pattern cards (when it applies, when it does NOT, the canonical citation) so plans name a greenfield alternative instead of following the legacy line by inertia.
28
28
  - **Card first, on YOUR board** — every piece of work is tracked before it starts: kj's HU Board (`kj hu add|move|list`), the Planning Game, or the board the project already uses (Linear, Trello, Jira, GitHub Issues) via your agent's own MCP/tools. Declared, verified at install, never optional — Karajan does not run without a board. ADRs live in git (`kj adr add|list`).
@@ -34,6 +34,7 @@ Your AI agent (Claude Code, Codex, Gemini CLI, Cursor…) writes the code. **Kar
34
34
  - **Nothing personal ships** — every outbound boundary audits before it leaves the machine: the pre-commit rejects a staged diff carrying your denylisted personal data, hardcoded platform tokens (`ghp_`, `sk-`, `AKIA`…) block outright, `verify-pack`-style tarball scans guard the publish, and `kj privacy scan <dir>` audits any build output. Your denylist lives in `~/.karajan/privacy.yml` — the install asks and writes it for you.
35
35
  - **Installing IS activating** — `kj env install` performs the enforcement itself (git hooks, verdict gate, tool gate) instead of trusting the agent to run setup steps, and ends by printing the method into the very conversation that installed it. A commit outside the method is rejected, not narrated.
36
36
  - **The turn cannot end red — the Sentinel** — a deterministic supervisor (zero LLM) wired into the harness's synchronous hooks records the method state of the session as tools run, and a Stop hook blocks the agent from ending its turn while method violations are open. The program rules, the agent thinks. See [guarantee levels](#guarantee-levels-governed-vs-supervised).
37
+ - **The review panel never runs dry** — nine built-in agents (Claude Code, Codex, GitHub Copilot, Antigravity `agy` — the gemini successor —, Kimi Code, Qwen, OpenCode, Aider, Gemini legacy), and when the configured reviewer exhausts its quota, `kj review` switches to an authenticated candidate with a LOUD notice — or hands you the menu of candidates with their tier (free / subscription / local) and the exact login command. Never a silent failure, never the brain reviewing itself.
37
38
 
38
39
  This repo runs under its own environment: every commit to karajan-code carries a cross-AI verdict.
39
40
 
package/docs/README.es.md CHANGED
@@ -14,7 +14,7 @@
14
14
 
15
15
  ---
16
16
 
17
- Tu agente de IA (Claude Code, Codex, Gemini CLI, Cursor…) escribe el código. **Karajan gobierna cómo ocurre**: instala un método que tu agente sigue en cada tarea y lo hace cumplir con gates de git que hacen el falso verde estructuralmente imposible.
17
+ Tu agente de IA (Claude Code, Codex, Copilot, Antigravity, Cursor…) escribe el código. **Karajan gobierna cómo ocurre**: instala un método que tu agente sigue en cada tarea y lo hace cumplir con gates de git que hacen el falso verde estructuralmente imposible.
18
18
 
19
19
  - **RAG antes de suponer** — `kj rag query` responde qué hace tu código; ningún agente adivina. La instalación lo cablea como herramienta MCP nativa (`kj_rag_query`), de modo que consultar el índice sea el camino más barato del agente. Funciona de serie: Ollama local, o el embedder ONNX integrado cuando no se puede instalar nada; los embedders cloud exigen declarar la sensibilidad y redactan PII de cada chunk. Y viaja con un canon de ingeniería destilado: `kj rag query --library` sirve fichas de patrón (cuándo aplica, cuándo NO, la cita canónica) para que los planes nombren una alternativa greenfield en vez de seguir la línea del legacy por inercia.
20
20
  - **Card primero, en TU board** — todo trabajo se registra antes de empezar: el HU Board de kj (`kj hu add|move|list`), el Planning Game, o el board que el proyecto ya use (Linear, Trello, Jira, GitHub Issues) vía los MCP/tools de tu agente. Declarado, verificado en la instalación, jamás opcional — Karajan no funciona sin board. Los ADRs viven en git (`kj adr add|list`).
@@ -26,6 +26,7 @@ Tu agente de IA (Claude Code, Codex, Gemini CLI, Cursor…) escribe el código.
26
26
  - **Nada personal se publica** — cada boundary de salida se audita antes de dejar la máquina: el pre-commit rechaza un diff con tus datos vetados, los tokens de plataforma hardcodeados (`ghp_`, `sk-`, `AKIA`…) bloquean directamente, el scan del tarball guarda el publish, y `kj privacy scan <dir>` audita cualquier build. Tu denylist vive en `~/.karajan/privacy.yml` — la instalación pregunta y la escribe por ti.
27
27
  - **Instalar ES activar** — `kj env install` ejecuta él mismo el enforcement (hooks de git, gate de veredicto, tool gate) en vez de confiar en que el agente corra pasos de setup, y termina imprimiendo el método en la propia conversación que instaló. Un commit fuera del método se rechaza, no se narra.
28
28
  - **El turno no puede terminar en rojo — el Sentinel** — un supervisor determinista (cero LLM) cableado a los hooks síncronos del harness registra el estado del método de la sesión según corren las herramientas, y un hook Stop bloquea que el agente termine su turno mientras haya violaciones abiertas. El programa manda, el agente piensa. Ver [niveles de garantía](#niveles-de-garantía-gobernado-vs-supervisado).
29
+ - **El panel de revisión nunca se seca** — nueve agentes integrados (Claude Code, Codex, GitHub Copilot, Antigravity `agy` — el sucesor de gemini —, Kimi Code, Qwen, OpenCode, Aider, Gemini legacy), y cuando al reviewer configurado se le agota la cuota, `kj review` cambia a un candidato autenticado AVISANDO en alto — o te entrega el menú de candidatos con su tier (gratis / suscripción / local) y el comando de login exacto. Jamás un fallo mudo, jamás el brain revisándose a sí mismo.
29
30
 
30
31
  Este repo corre bajo su propio entorno: cada commit de karajan-code lleva un veredicto de IA cruzada.
31
32
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "karajan-code",
3
- "version": "4.13.0",
3
+ "version": "4.15.0",
4
4
  "description": "Local multi-agent coding orchestrator with TDD, SonarQube, and code review pipeline",
5
5
  "type": "module",
6
6
  "license": "AGPL-3.0",
@@ -76,8 +76,8 @@
76
76
  "test": "vitest run",
77
77
  "test:watch": "vitest",
78
78
  "test:coverage": "vitest run --coverage",
79
- "lint": "eslint src/",
80
- "lint:fix": "eslint src/ --fix",
79
+ "lint": "eslint src/ packages/",
80
+ "lint:fix": "eslint src/ packages/ --fix",
81
81
  "lint:syntax": "((command -v rg >/dev/null 2>&1 && rg --files src tests | rg '\\.js$') || find src tests -type f -name '*.js') | while IFS= read -r f; do node --check \"$f\"; done",
82
82
  "format:check": "prettier --check .",
83
83
  "format:fix": "prettier --write .",
@@ -113,7 +113,9 @@ async function cmdPurge(root, id, out, err) {
113
113
  async function cmdEmpty(root, flags, out) {
114
114
  await ensureRoot(root);
115
115
  const m = await loadManifest(root);
116
- let dropped = [];
116
+ // Assigned in every branch below (the trailing else covers the default),
117
+ // so an initializer would never be read (no-useless-assignment).
118
+ let dropped;
117
119
  if (flags["older-than-days"]) {
118
120
  const days = Number(flags["older-than-days"]);
119
121
  if (!Number.isFinite(days) || days < 0) throw new Error("--older-than-days must be >= 0");
@@ -117,7 +117,7 @@ function humaniseProjectName(id) {
117
117
  if (!id || typeof id !== 'string') return id || '';
118
118
  const tail = id.split(/[/_]/).filter(Boolean).pop() || id;
119
119
  const words = tail
120
- .split(/[\s\-]+/)
120
+ .split(/[\s-]+/)
121
121
  .map((w) => w.replace(/[^\p{L}\p{N}]/gu, ''))
122
122
  .filter((w) => w.length > 0);
123
123
  if (words.length === 0) return id;
@@ -23,7 +23,6 @@
23
23
 
24
24
  import fs from 'node:fs';
25
25
  import path from 'node:path';
26
- import { homedir } from 'node:os';
27
26
  import {
28
27
  getDb,
29
28
  getKjHome,
@@ -18,7 +18,6 @@
18
18
 
19
19
  import { mkdirSync, openSync } from "node:fs";
20
20
  import { join, dirname } from "node:path";
21
- import { homedir } from "node:os";
22
21
  import { spawn } from "node:child_process";
23
22
  import { fileURLToPath } from "node:url";
24
23
  import { randomUUID } from "node:crypto";
@@ -30,7 +30,6 @@
30
30
 
31
31
  import { existsSync, mkdirSync, readFileSync, writeFileSync, copyFileSync, renameSync } from 'node:fs';
32
32
  import { join, dirname } from 'node:path';
33
- import { tmpdir } from 'node:os';
34
33
  import yaml from 'js-yaml';
35
34
  import { getKjHome } from './db.js';
36
35
 
@@ -448,7 +447,7 @@ export function readConfig({ scope = 'global' } = {}) {
448
447
  parsed = yaml.load(readFileSync(p, 'utf8'), { json: true }) || {};
449
448
  exists = true;
450
449
  } catch (err) {
451
- throw new Error(`No se pudo parsear ${p}: ${err.message}`);
450
+ throw new Error(`No se pudo parsear ${p}: ${err.message}`, { cause: err });
452
451
  }
453
452
  }
454
453
  const fields = EDITABLE_FIELDS.map((f) => {
@@ -150,8 +150,6 @@ export function cleanupEphemeralProjects({
150
150
  const candidates = findEphemeralProjects(allProjects, { ...opts, now: now() });
151
151
  if (candidates.length === 0) return [];
152
152
 
153
- const countStories = db.prepare("SELECT COUNT(*) AS n FROM stories WHERE project_id = ?");
154
- const countSessions = db.prepare("SELECT COUNT(*) AS n FROM sessions WHERE project_id = ?");
155
153
  const delStories = db.prepare("DELETE FROM stories WHERE project_id = ?");
156
154
  const delSessions = db.prepare("DELETE FROM sessions WHERE project_id = ?");
157
155
  const delProject = db.prepare("DELETE FROM projects WHERE id = ?");
@@ -16,7 +16,6 @@
16
16
  */
17
17
  import { readFileSync, existsSync, readdirSync, mkdirSync, openSync } from 'node:fs';
18
18
  import { join, dirname } from 'node:path';
19
- import { homedir } from 'node:os';
20
19
  import { writeJsonAtomicSync } from 'karajan-core/atomic-write';
21
20
  import { spawn } from 'node:child_process';
22
21
  import { trackRun, untrack } from './run-tracker.js';
@@ -28,7 +28,7 @@
28
28
  */
29
29
 
30
30
  import { execSync } from 'node:child_process';
31
- import { existsSync, readFileSync, readdirSync, statSync } from 'node:fs';
31
+ import { existsSync, readFileSync, readdirSync } from 'node:fs';
32
32
  import { join } from 'node:path';
33
33
  import { getHuBoardPlansDirs } from './db.js';
34
34
 
@@ -21,7 +21,6 @@ export async function readPlanCached(p) {
21
21
  planFileCache.set(p, { mtimeMs, plan });
22
22
  return plan;
23
23
  }
24
- import os from 'node:os';
25
24
  import path from 'node:path';
26
25
  import { spawn as spawnChild } from 'node:child_process';
27
26
  import {
@@ -36,7 +35,6 @@ import {
36
35
  getKjHome,
37
36
  getHuBoardRunsDir,
38
37
  getHuBoardPlansDir,
39
- getHuBoardLegacyPlansDir,
40
38
  getHuBoardPlansDirs,
41
39
  getStoryRow,
42
40
  listPlanIdsForProject,
@@ -93,19 +91,6 @@ function huStoriesDir() {
93
91
  return path.join(getKjHome(), 'hu-stories');
94
92
  }
95
93
 
96
- /**
97
- * Best-effort removal of the hu-stories/<id>/ directory.
98
- */
99
- async function removeBatchDir(batchId) {
100
- try {
101
- const dir = path.join(huStoriesDir(), batchId);
102
- await fsp.rm(dir, { recursive: true, force: true });
103
- return true;
104
- } catch {
105
- return false;
106
- }
107
- }
108
-
109
94
  /**
110
95
  * GET /api/dashboard - Global dashboard statistics.
111
96
  */
@@ -4,7 +4,6 @@ import rateLimit from 'express-rate-limit';
4
4
  import { join, dirname } from 'node:path';
5
5
  import { fileURLToPath, pathToFileURL } from 'node:url';
6
6
  import { writeFileSync, rmSync, mkdirSync, realpathSync } from 'node:fs';
7
- import { homedir } from 'node:os';
8
7
  import { initDb, closeDb } from './db.js';
9
8
  import { fullScan, startWatcher } from './sync.js';
10
9
  import apiRoutes from './routes/api.js';
@@ -1,5 +1,5 @@
1
1
  import { watch } from 'chokidar';
2
- import { readFileSync, readdirSync, existsSync, statSync, rmSync } from 'node:fs';
2
+ import { readFileSync, readdirSync, existsSync, rmSync } from 'node:fs';
3
3
  import { dirname, join, basename } from 'node:path';
4
4
  import { homedir } from 'node:os';
5
5
  import {
@@ -12,7 +12,6 @@
12
12
 
13
13
  import { randomBytes } from "node:crypto";
14
14
  import { mkdirSync, readFileSync, writeFileSync, chmodSync, existsSync } from "node:fs";
15
- import { homedir } from "node:os";
16
15
  import { dirname, join } from "node:path";
17
16
 
18
17
  const TOKEN_BYTES = 32;
@@ -45,7 +45,6 @@
45
45
 
46
46
  import fs from "node:fs";
47
47
  import path from "node:path";
48
- import os from "node:os";
49
48
 
50
49
  const DEFAULT_CODING_HOURS = 6;
51
50
  const DEFAULT_PAUSED_HOURS = 24;
@@ -0,0 +1,62 @@
1
+ import { BaseAgent } from "./base-agent.js";
2
+ import { resolveBin } from "./resolve-bin.js";
3
+
4
+ /**
5
+ * Antigravity CLI (`agy`) — KJC-TSK-0729 fase 2b: the ninth built-in agent.
6
+ * Google's official successor to the retired gemini CLI (dead 2026-06-18),
7
+ * covered by the user's Google AI Pro/Ultra subscription — the natural
8
+ * third AI for solomon arbitration, back at last.
9
+ *
10
+ * Verified live against agy 1.1.10:
11
+ * - `-p <prompt>` runs print mode; `--output-format json` emits ONE JSON
12
+ * object: {status: "SUCCESS", response, usage}. Non-SUCCESS status can
13
+ * arrive with exit 0, so ok checks BOTH.
14
+ * - `--disable-slash-commands` keeps prompt content (diffs can contain
15
+ * "/lines") from expanding as slash commands in print mode.
16
+ * - Permission prompts are denied in print mode unless
17
+ * `--dangerously-skip-permissions` (coder mode only).
18
+ * - Auth lives under ~/.gemini/antigravity-cli (device/browser login).
19
+ */
20
+ export class AgyAgent extends BaseAgent {
21
+ async runTask(task) {
22
+ return this._exec(task, this.getRoleModel(task.role || "coder"), ["--dangerously-skip-permissions"]);
23
+ }
24
+
25
+ async reviewTask(task) {
26
+ return this._exec(
27
+ { ...task, prompt: `Do NOT use any tools. Answer directly from this prompt only.\n\n${task.prompt}` },
28
+ this.getRoleModel(task.role || "reviewer"),
29
+ [],
30
+ );
31
+ }
32
+
33
+ async _exec(task, model, extraArgs) {
34
+ const args = ["-p", task.prompt, "--output-format", "json", "--disable-slash-commands", ...extraArgs];
35
+ if (model) args.push("--model", model);
36
+ const res = await this.runCommand(resolveBin("agy"), args, {
37
+ onOutput: task.onOutput,
38
+ silenceTimeoutMs: task.silenceTimeoutMs,
39
+ timeout: task.timeoutMs,
40
+ cwd: task.cwd,
41
+ });
42
+ const parsed = extractAgyOutput(res.stdout);
43
+ const ok = res.exitCode === 0 && parsed?.status === "SUCCESS";
44
+ return {
45
+ ok,
46
+ output: parsed?.response ?? res.stdout,
47
+ error: ok ? res.stderr : parsed?.error || res.stderr || `agy status: ${parsed?.status || "unparseable"}`,
48
+ exitCode: res.exitCode,
49
+ };
50
+ }
51
+ }
52
+
53
+ /** Parse agy's single-object JSON print output ({status, response}), null if not JSON. */
54
+ export function extractAgyOutput(stdout) {
55
+ if (!stdout) return null;
56
+ try {
57
+ const obj = JSON.parse(stdout.trim());
58
+ return typeof obj === "object" && obj !== null ? obj : null;
59
+ } catch {
60
+ return null;
61
+ }
62
+ }
@@ -374,7 +374,8 @@ export class ClaudeAgent extends BaseAgent {
374
374
  onOutput: streamFilter,
375
375
  silenceTimeoutMs: task.silenceTimeoutMs,
376
376
  timeout: task.timeoutMs,
377
- env: task.env
377
+ env: task.env,
378
+ cwd: task.cwd
378
379
  }));
379
380
  const raw = pickOutput(res);
380
381
  const output = extractTextFromStreamJson(raw);
@@ -384,7 +385,7 @@ export class ClaudeAgent extends BaseAgent {
384
385
 
385
386
  // Without streaming, use json output to get structured response via stderr
386
387
  args.push("--output-format", "json");
387
- const res = await this.runCommand(resolveBin("claude"), args, cleanExecaOpts({ env: task.env }));
388
+ const res = await this.runCommand(resolveBin("claude"), args, cleanExecaOpts({ env: task.env, cwd: task.cwd }));
388
389
  const raw = pickOutput(res);
389
390
  const output = extractTextFromStreamJson(raw);
390
391
  const usage = extractUsageFromStreamJson(raw);
@@ -85,7 +85,9 @@ export class CodexAgent extends BaseAgent {
85
85
  onOutput: task.onOutput,
86
86
  silenceTimeoutMs: task.silenceTimeoutMs,
87
87
  timeout: task.timeoutMs,
88
- input: task.prompt
88
+ input: task.prompt,
89
+ // TOR-A (KJC-TSK-0723): tournament lanes run the coder INSIDE the lane.
90
+ cwd: task.cwd
89
91
  });
90
92
  const usage = extractCodexTokens(res.stdout);
91
93
  return { ok: res.exitCode === 0, output: res.stdout, error: res.stderr, exitCode: res.exitCode, ...usage };
@@ -5,6 +5,8 @@ import { AiderAgent } from "./aider-agent.js";
5
5
  import { OpenCodeAgent } from "./opencode-agent.js";
6
6
  import { QwenAgent } from "./qwen-agent.js";
7
7
  import { CopilotAgent } from "./copilot-agent.js";
8
+ import { KimiAgent } from "./kimi-agent.js";
9
+ import { AgyAgent } from "./agy-agent.js";
8
10
 
9
11
  const agentRegistry = new Map();
10
12
 
@@ -53,3 +55,5 @@ registerAgent("aider", AiderAgent, { bin: "aider", installUrl: "https://aider.ch
53
55
  registerAgent("opencode", OpenCodeAgent, { bin: "opencode", installUrl: "https://opencode.ai" });
54
56
  registerAgent("qwen", QwenAgent, { bin: "qwen", installUrl: "https://github.com/QwenLM/qwen-code" });
55
57
  registerAgent("copilot", CopilotAgent, { bin: "copilot", installUrl: "https://docs.github.com/copilot/how-tos/use-copilot-agents/use-copilot-cli" });
58
+ registerAgent("kimi", KimiAgent, { bin: "kimi", installUrl: "https://github.com/MoonshotAI/kimi-code" });
59
+ registerAgent("agy", AgyAgent, { bin: "agy", installUrl: "https://antigravity.google" });
@@ -0,0 +1,45 @@
1
+ import { BaseAgent } from "./base-agent.js";
2
+ import { resolveBin } from "./resolve-bin.js";
3
+
4
+ /**
5
+ * Kimi Code (Moonshot, binary `kimi`) — KJC-TSK-0729 fase 2: the eighth
6
+ * built-in agent and the free-tier reserve of the reviewer/solomon panel.
7
+ *
8
+ * Verified live against kimi-code 0.27.0:
9
+ * - `-p <prompt>` runs one prompt non-interactively and prints the
10
+ * response (text output by default) — the prompt travels as an argument,
11
+ * same E2BIG caveat as claude/copilot (KJC-BUG-0121) for huge diffs.
12
+ * - `-m <model>` picks a model alias; `-y` auto-approves actions.
13
+ * - Auth is the device-code login (`kimi login`), no env keys involved;
14
+ * "No model configured" on a logged machine means the login did not
15
+ * populate the managed provider (re-run `kimi login`).
16
+ */
17
+ export class KimiAgent extends BaseAgent {
18
+ async runTask(task) {
19
+ // Coder mode IS an agentic run: file edits and shell need approval.
20
+ return this._exec(task, this.getRoleModel(task.role || "coder"), ["-y"]);
21
+ }
22
+
23
+ async reviewTask(task) {
24
+ // Review/arbitration answers from the prompt alone — no auto-approval,
25
+ // and an explicit no-tools instruction so a headless run never stalls
26
+ // waiting for a permission prompt nobody can answer.
27
+ return this._exec(
28
+ { ...task, prompt: `Do NOT use any tools. Answer directly from this prompt only.\n\n${task.prompt}` },
29
+ this.getRoleModel(task.role || "reviewer"),
30
+ [],
31
+ );
32
+ }
33
+
34
+ async _exec(task, model, extraArgs) {
35
+ const args = ["-p", task.prompt, ...extraArgs];
36
+ if (model) args.push("-m", model);
37
+ const res = await this.runCommand(resolveBin("kimi"), args, {
38
+ onOutput: task.onOutput,
39
+ silenceTimeoutMs: task.silenceTimeoutMs,
40
+ timeout: task.timeoutMs,
41
+ cwd: task.cwd,
42
+ });
43
+ return { ok: res.exitCode === 0, output: res.stdout, error: res.stderr, exitCode: res.exitCode };
44
+ }
45
+ }
@@ -50,8 +50,11 @@ export function collectAiSurface({ projectDir = process.cwd(), home = os.homedir
50
50
  * Diff the current surface against the last-seen snapshot and persist the
51
51
  * new state. First run records the baseline silently.
52
52
  */
53
- export function checkAiSurface({ projectDir = process.cwd(), home = os.homedir(), statePath = path.join(os.homedir(), ".karajan", "ai-surface.json") } = {}) {
54
- const surface = collectAiSurface({ projectDir, home });
53
+ export function checkAiSurface({ projectDir = process.cwd(), home = os.homedir(), statePath = path.join(os.homedir(), ".karajan", "ai-surface.json"), extraSurface = [] } = {}) {
54
+ // KJC-TSK-0728: extraSurface carries entries the (async) caller collected
55
+ // outside config files — e.g. observed agent CLIs as "grok (cli)" — so
56
+ // they ride the same snapshot and the same "NEW since last check" drift.
57
+ const surface = [...new Set([...collectAiSurface({ projectDir, home }), ...extraSurface])].sort();
55
58
  let state = {};
56
59
  try { state = JSON.parse(readFileSync(statePath, "utf8")); } catch { /* first run or corrupt → baseline */ }
57
60
  const prev = Array.isArray(state[projectDir]) ? state[projectDir] : null;
@@ -12,7 +12,7 @@
12
12
  * the install command and can offer to run it).
13
13
  */
14
14
 
15
- import { checkBinary, KNOWN_AGENTS } from "../utils/agent-detect.js";
15
+ import { checkBinary, KNOWN_AGENTS, detectObservedAgents } from "../utils/agent-detect.js";
16
16
  import { runCommand } from "../utils/process.js";
17
17
  import { withDocLink } from "../utils/doc-links.js";
18
18
  import { getInstallHint, appliesToStack } from "../utils/install-hints.js";
@@ -38,6 +38,33 @@ function createAgentCheck(agent) {
38
38
  };
39
39
  }
40
40
 
41
+ /**
42
+ * KJC-TSK-0728 — ONE aggregate line for the observation census (agent CLIs
43
+ * kj sees but does not drive). Only FOUND CLIs are listed; the missing ones
44
+ * make no noise, and the check never fails. `detector` is injectable for
45
+ * tests.
46
+ */
47
+ export function createObservedAgentsCheck({ detector = detectObservedAgents } = {}) {
48
+ return {
49
+ name: "agents:observed",
50
+ label: "Agent CLIs observed (not driven)",
51
+ strategy: STRATEGY.MANUAL,
52
+ async detect() {
53
+ let found = [];
54
+ try {
55
+ found = (await detector()).filter((a) => a.available);
56
+ } catch { /* census is best-effort */ }
57
+ return {
58
+ ok: true,
59
+ severity: "info",
60
+ detail: found.length
61
+ ? found.map((a) => `${a.name} (${(a.version || "").split(" ").at(-1) || "?"})`).join(", ")
62
+ : "none detected",
63
+ };
64
+ },
65
+ };
66
+ }
67
+
41
68
  /**
42
69
  * Check the presence of a core binary (node, npm, git).
43
70
  */
@@ -165,6 +192,7 @@ export function getBinaryChecks() {
165
192
  for (const agent of KNOWN_AGENTS) {
166
193
  checks.push(createAgentCheck(agent));
167
194
  }
195
+ checks.push(createObservedAgentsCheck());
168
196
  for (const bin of ["node", "npm", "git"]) {
169
197
  checks.push(createCoreBinaryCheck(bin));
170
198
  }
@@ -27,7 +27,7 @@ export const META_COMMANDS = ["advanced", "help"];
27
27
  * so a newly-registered command can never silently vanish from `kj advanced`.
28
28
  */
29
29
  export const ADVANCED_GROUPS = [
30
- { title: "Pipeline (piezas sueltas)", commands: ["autorun", "code", "review", "solomon", "agent", "scan"] },
30
+ { title: "Pipeline (piezas sueltas)", commands: ["autorun", "code", "review", "solomon", "agent", "scan", "tournament"] },
31
31
  { title: "Análisis pre-run", commands: ["discover", "triage", "researcher", "architect", "onboard", "brief"] },
32
32
  { title: "Búsqueda / RAG", commands: ["rag", "qmd", "watch"] },
33
33
  { title: "Calidad / auditoría", commands: ["audit", "check", "mutate", "webperf", "sonar", "privacy", "release"] },
@@ -3,6 +3,8 @@ import { ollamaStartCommand, ollamaStopCommand, ollamaStatusCommand, ollamaPullC
3
3
  import { configCommand } from "../commands/config.js";
4
4
  import { codeCommand } from "../commands/code.js";
5
5
  import { reviewCommand } from "../commands/review.js";
6
+ import { tournamentCommand } from "../commands/tournament.js";
7
+ import { resolveTaskInput } from "../utils/task-file.js";
6
8
  import { reviewGateCommand, solomonCommand } from "../commands/review-gate.js";
7
9
  import { scanCommand } from "../commands/scan.js";
8
10
  import { installToolsCommand } from "../commands/install-tools.js";
@@ -144,7 +146,6 @@ export function registerPipeline(program, { pkgVersion }) {
144
146
  effectiveTask = plan.task;
145
147
  }
146
148
  }
147
- const { resolveTaskInput } = await import("../utils/task-file.js");
148
149
  const resolvedTask = await resolveTaskInput({ task: effectiveTask, taskFile: flags.taskFile, projectDir: config.projectDir, logger });
149
150
  const res = await runCommandHandler({ task: resolvedTask, config, logger, flags });
150
151
  // KJC-TSK-0674 (issue #1289): an aborted run (spec-review cancel,
@@ -163,7 +164,6 @@ export function registerPipeline(program, { pkgVersion }) {
163
164
  .option("--skip-spec-review", "Bypass the spec-reviewer pre-pipeline audit")
164
165
  .action(async (task, flags) => {
165
166
  const code = await withConfig(pkgVersion, "autorun", flags, async ({ config, logger }) => {
166
- const { resolveTaskInput } = await import("../utils/task-file.js");
167
167
  const { autorunCommand } = await import("../commands/autorun.js");
168
168
  const resolvedTask = await resolveTaskInput({ task, taskFile: flags.taskFile, projectDir: config.projectDir, logger });
169
169
  const res = await autorunCommand({ task: resolvedTask, config, logger, flags });
@@ -181,7 +181,6 @@ export function registerPipeline(program, { pkgVersion }) {
181
181
  .option("--coder-model <name>")
182
182
  .action(async (task, flags) => {
183
183
  await withConfig(pkgVersion, "code", flags, async ({ config, logger }) => {
184
- const { resolveTaskInput } = await import("../utils/task-file.js");
185
184
  const resolvedTask = await resolveTaskInput({ task, taskFile: flags.taskFile, projectDir: config.projectDir, logger });
186
185
  await codeCommand({ task: resolvedTask, config, logger });
187
186
  });
@@ -199,6 +198,30 @@ export function registerPipeline(program, { pkgVersion }) {
199
198
  });
200
199
  });
201
200
 
201
+ // KJC-TSK-0723 (TOR-A) — the governed tournament: same task, N coders,
202
+ // isolated lanes; the method picks the winner (TOR-B/C), not vibes.
203
+ program
204
+ .command("tournament")
205
+ .description("Fan the SAME task out to N coders in isolated worktree lanes and collect per-lane evidence (diff, suite, log)")
206
+ .argument("[task]", "Task description (REQUIRED — provide as argument or via --task-file)")
207
+ .option("--task-file <path>", "Read the task from a file (e.g. .md)")
208
+ .option("--coders <csv>", "Comma-separated coder agents (minimum 2), e.g. claude,codex,agy")
209
+ .option("--score <id>", "Score an EXISTING tournament from its artifacts (deterministic, zero LLM)")
210
+ .option("--judge <id>", "Cross-AI judge ranks the finalists (scoreboard in sight; solomon on tie/dispute)")
211
+ .option("--crown <id>", "Commit the judged winner in its lane THROUGH the real review gate")
212
+ .option("--card <ref>", "With --crown: card reference for the commit message (card-first gates)")
213
+ .option("--json", "With --score/--judge/--crown: emit stable JSON")
214
+ .action(async (task, flags) => {
215
+ await withConfig(pkgVersion, "tournament", flags, async ({ config, logger }) => {
216
+ if (flags.score || flags.judge || flags.crown) {
217
+ await tournamentCommand({ task: "", config, logger, flags });
218
+ return;
219
+ }
220
+ const resolvedTask = await resolveTaskInput({ task, taskFile: flags.taskFile, projectDir: config.projectDir, logger });
221
+ await tournamentCommand({ task: resolvedTask, config, logger, flags });
222
+ });
223
+ });
224
+
202
225
  program
203
226
  .command("review")
204
227
  .description("Run only reviewer (--staged/--check: v4 cross-AI gate with recorded verdict)")
@@ -220,7 +243,6 @@ export function registerPipeline(program, { pkgVersion }) {
220
243
  await reviewGateCommand({ config, logger, flags: { ...flags, task } });
221
244
  return;
222
245
  }
223
- const { resolveTaskInput } = await import("../utils/task-file.js");
224
246
  const resolvedTask = await resolveTaskInput({ task, taskFile: flags.taskFile, projectDir: config.projectDir, logger });
225
247
  await reviewCommand({ task: resolvedTask, config, logger, baseRef: flags.baseRef });
226
248
  });
@@ -7,6 +7,7 @@
7
7
  import { checkHarden } from "../harden/check.js";
8
8
  import { collectMethodStats, formatMethodStats } from "../checks/method.js";
9
9
  import { checkAiSurface, formatAiSurface } from "../checks/ai-surface.js";
10
+ import { detectObservedAgents } from "../utils/agent-detect.js";
10
11
 
11
12
  export async function checkCommand({ projectDir = process.cwd(), profile = "standard", json = false, logger = console } = {}) {
12
13
  const result = await checkHarden({ projectDir, profile });
@@ -14,8 +15,15 @@ export async function checkCommand({ projectDir = process.cwd(), profile = "stan
14
15
  // rides along in check output but never affects the exit code.
15
16
  const method = await collectMethodStats({ projectDir }).catch(() => null);
16
17
  // KJC-TSK-0694: same deal for the MCP inventory — a nudge, never a gate.
18
+ // KJC-TSK-0728: observed agent CLIs ride the same snapshot as "(cli)"
19
+ // entries, so a newly-appeared agent binary trips the same drift question.
17
20
  let aiSurface = null;
18
- try { aiSurface = checkAiSurface({ projectDir }); } catch { /* inventory is best-effort */ }
21
+ try {
22
+ const clis = (await detectObservedAgents().catch(() => []))
23
+ .filter((a) => a.available)
24
+ .map((a) => `${a.name} (cli)`);
25
+ aiSurface = checkAiSurface({ projectDir, extraSurface: clis });
26
+ } catch { /* inventory is best-effort */ }
19
27
 
20
28
  if (json) {
21
29
  logger.info?.(JSON.stringify({ ...result, method, aiSurface }));
@@ -184,7 +184,7 @@ export async function hardenCommand({
184
184
  const verb = dryRun ? "would install" : "installed";
185
185
  logger.info?.(`kj harden (${profile}) — ${verb} ${result.hooks.length} hook(s) → ${result.hooksPath}`);
186
186
  for (const h of result.hooks) logger.info?.(` • ${h.hook}: ${h.action}`);
187
- for (const c of out.configs) logger.info?.(` • ${c.file}: ${c.action}`);
187
+ for (const c of out.configs) logger.info?.(` • ${c.file}: ${c.action}${c.note ? ` (${c.note})` : ""}`);
188
188
  for (const w of out.workflows) logger.info?.(` • ${w.file}: ${w.action}`);
189
189
  for (const g of out.guidelines) logger.info?.(` • ${g.file}: ${g.action}`);
190
190
  if (!dryRun) logger.info?.("core.hooksPath set. Verify later with `kj check`.");