karajan-code 4.12.0 → 4.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -22,7 +22,7 @@
22
22
 
23
23
  ---
24
24
 
25
- Your AI agent (Claude Code, Codex, Gemini CLI, Cursor…) writes the code. **Karajan governs how it happens**: it installs a method your agent follows on every task, and enforces it with git gates that make a false green structurally impossible.
25
+ Your AI agent (Claude Code, Codex, Copilot, Antigravity, Cursor…) writes the code. **Karajan governs how it happens**: it installs a method your agent follows on every task, and enforces it with git gates that make a false green structurally impossible.
26
26
 
27
27
  - **RAG before assuming** — `kj rag query` answers what your codebase does; no agent guesses. The install wires it as a native MCP tool (`kj_rag_query`) so querying the index is the agent's cheapest path. Works out of the box: local Ollama, or the built-in ONNX embedder when nothing can be installed; cloud embedders require an explicit sensitivity declaration and PII-redact every chunk. A distilled engineering canon rides along: `kj rag query --library` serves pattern cards (when it applies, when it does NOT, the canonical citation) so plans name a greenfield alternative instead of following the legacy line by inertia.
28
28
  - **Card first, on YOUR board** — every piece of work is tracked before it starts: kj's HU Board (`kj hu add|move|list`), the Planning Game, or the board the project already uses (Linear, Trello, Jira, GitHub Issues) via your agent's own MCP/tools. Declared, verified at install, never optional — Karajan does not run without a board. ADRs live in git (`kj adr add|list`).
@@ -33,6 +33,8 @@ Your AI agent (Claude Code, Codex, Gemini CLI, Cursor…) writes the code. **Kar
33
33
  - **Least privilege for agents** — spawned agent subprocesses receive an env allowlist (their own CLI's auth, never your cloud keys or registry tokens), and `kj check` inventories every MCP the project can reach, flagging what appeared since the last check. Sensitive-surface tasks self-invoke `kj audit --security` — a zero-token pass (prompt-injection over the agent-context files + OSV + Semgrep + Sonar) — and remediate before review.
34
34
  - **Nothing personal ships** — every outbound boundary audits before it leaves the machine: the pre-commit rejects a staged diff carrying your denylisted personal data, hardcoded platform tokens (`ghp_`, `sk-`, `AKIA`…) block outright, `verify-pack`-style tarball scans guard the publish, and `kj privacy scan <dir>` audits any build output. Your denylist lives in `~/.karajan/privacy.yml` — the install asks and writes it for you.
35
35
  - **Installing IS activating** — `kj env install` performs the enforcement itself (git hooks, verdict gate, tool gate) instead of trusting the agent to run setup steps, and ends by printing the method into the very conversation that installed it. A commit outside the method is rejected, not narrated.
36
+ - **The turn cannot end red — the Sentinel** — a deterministic supervisor (zero LLM) wired into the harness's synchronous hooks records the method state of the session as tools run, and a Stop hook blocks the agent from ending its turn while method violations are open. The program rules, the agent thinks. See [guarantee levels](#guarantee-levels-governed-vs-supervised).
37
+ - **The review panel never runs dry** — nine built-in agents (Claude Code, Codex, GitHub Copilot, Antigravity `agy` — the gemini successor —, Kimi Code, Qwen, OpenCode, Aider, Gemini legacy), and when the configured reviewer exhausts its quota, `kj review` switches to an authenticated candidate with a LOUD notice — or hands you the menu of candidates with their tier (free / subscription / local) and the exact login command. Never a silent failure, never the brain reviewing itself.
36
38
 
37
39
  This repo runs under its own environment: every commit to karajan-code carries a cross-AI verdict.
38
40
 
@@ -67,6 +69,12 @@ Requires git and at least one AI agent CLI — two enables cross-AI review; thre
67
69
 
68
70
  Full method: [Work with your agent](https://karajancode.com/docs/v4/working-with-your-agent/) · [The gates](https://karajancode.com/docs/v4/gates/) · [Command reference](https://karajancode.com/docs/v4/commands/).
69
71
 
72
+ ## Guarantee levels: governed vs supervised
73
+
74
+ Karajan **governs** any agent with git gates — the false green is structurally impossible no matter who writes, because the gates live in the repository, not in the agent's goodwill. On top of that, the **Karajan Sentinel** adds **supervision inside the turn**: `kj harden` wires deterministic hooks into the harness — one records the method state of the session as tools run (sources edited vs tests touched, escapes used), and a Stop hook blocks the agent from ending its turn while violations are open: sources edited on the base branch, a branch without a card, code without a single test touched. Every block states the exact violation and its remediation; `kj sentinel status` shows what the supervisor sees. It fails open — and says so — rather than ever hanging a session, and every `KJ_ALLOW_*` escape is recorded.
75
+
76
+ Synchronous blocking hooks exist today only in Claude Code. That makes the supported setup explicit: **to guarantee a harness that controls the LLM, use Claude Code as the host** — Claude writes, Codex reviews (a review subprocess needs no hooks), and a third CLI arbitrates when available. On any other host Karajan still governs at the full git-gate level and tells you which level is active — it never pretends a supervision it cannot enforce.
77
+
70
78
  ## Headless mode
71
79
 
72
80
  The classic multiagent pipeline lives on for CI and automation: `kj run "<task>"` orchestrates coder/reviewer/tester subprocess roles unattended, with the same gates. Agents and CI pass `--non-interactive` (or `KJ_NON_INTERACTIVE=1`): safe gates auto-answer, FAIL findings stop the run with a real exit code. `kj advanced` lists the full surface. [Headless mode docs](https://karajancode.com/docs/v4/headless/).
package/docs/README.es.md CHANGED
@@ -14,7 +14,7 @@
14
14
 
15
15
  ---
16
16
 
17
- Tu agente de IA (Claude Code, Codex, Gemini CLI, Cursor…) escribe el código. **Karajan gobierna cómo ocurre**: instala un método que tu agente sigue en cada tarea y lo hace cumplir con gates de git que hacen el falso verde estructuralmente imposible.
17
+ Tu agente de IA (Claude Code, Codex, Copilot, Antigravity, Cursor…) escribe el código. **Karajan gobierna cómo ocurre**: instala un método que tu agente sigue en cada tarea y lo hace cumplir con gates de git que hacen el falso verde estructuralmente imposible.
18
18
 
19
19
  - **RAG antes de suponer** — `kj rag query` responde qué hace tu código; ningún agente adivina. La instalación lo cablea como herramienta MCP nativa (`kj_rag_query`), de modo que consultar el índice sea el camino más barato del agente. Funciona de serie: Ollama local, o el embedder ONNX integrado cuando no se puede instalar nada; los embedders cloud exigen declarar la sensibilidad y redactan PII de cada chunk. Y viaja con un canon de ingeniería destilado: `kj rag query --library` sirve fichas de patrón (cuándo aplica, cuándo NO, la cita canónica) para que los planes nombren una alternativa greenfield en vez de seguir la línea del legacy por inercia.
20
20
  - **Card primero, en TU board** — todo trabajo se registra antes de empezar: el HU Board de kj (`kj hu add|move|list`), el Planning Game, o el board que el proyecto ya use (Linear, Trello, Jira, GitHub Issues) vía los MCP/tools de tu agente. Declarado, verificado en la instalación, jamás opcional — Karajan no funciona sin board. Los ADRs viven en git (`kj adr add|list`).
@@ -25,6 +25,8 @@ Tu agente de IA (Claude Code, Codex, Gemini CLI, Cursor…) escribe el código.
25
25
  - **Mínimo privilegio para agentes** — los subprocesos de agente reciben un allowlist de entorno (la auth de su propio CLI, jamás tus claves cloud ni tokens de registro), y `kj check` inventaría cada MCP alcanzable del proyecto marcando lo aparecido desde el último check. Las tareas con superficie sensible se auto-invocan `kj audit --security` — pasada de cero tokens (prompt-injection sobre los ficheros de contexto del agente + OSV + Semgrep + Sonar) — y remedian antes del review.
26
26
  - **Nada personal se publica** — cada boundary de salida se audita antes de dejar la máquina: el pre-commit rechaza un diff con tus datos vetados, los tokens de plataforma hardcodeados (`ghp_`, `sk-`, `AKIA`…) bloquean directamente, el scan del tarball guarda el publish, y `kj privacy scan <dir>` audita cualquier build. Tu denylist vive en `~/.karajan/privacy.yml` — la instalación pregunta y la escribe por ti.
27
27
  - **Instalar ES activar** — `kj env install` ejecuta él mismo el enforcement (hooks de git, gate de veredicto, tool gate) en vez de confiar en que el agente corra pasos de setup, y termina imprimiendo el método en la propia conversación que instaló. Un commit fuera del método se rechaza, no se narra.
28
+ - **El turno no puede terminar en rojo — el Sentinel** — un supervisor determinista (cero LLM) cableado a los hooks síncronos del harness registra el estado del método de la sesión según corren las herramientas, y un hook Stop bloquea que el agente termine su turno mientras haya violaciones abiertas. El programa manda, el agente piensa. Ver [niveles de garantía](#niveles-de-garantía-gobernado-vs-supervisado).
29
+ - **El panel de revisión nunca se seca** — nueve agentes integrados (Claude Code, Codex, GitHub Copilot, Antigravity `agy` — el sucesor de gemini —, Kimi Code, Qwen, OpenCode, Aider, Gemini legacy), y cuando al reviewer configurado se le agota la cuota, `kj review` cambia a un candidato autenticado AVISANDO en alto — o te entrega el menú de candidatos con su tier (gratis / suscripción / local) y el comando de login exacto. Jamás un fallo mudo, jamás el brain revisándose a sí mismo.
28
30
 
29
31
  Este repo corre bajo su propio entorno: cada commit de karajan-code lleva un veredicto de IA cruzada.
30
32
 
@@ -59,6 +61,12 @@ Requiere git y al menos un CLI de agente de IA — con dos hay revisión cruzada
59
61
 
60
62
  Método completo: [Trabaja con tu agente](https://karajancode.com/docs/es/v4/working-with-your-agent/) · [Los gates](https://karajancode.com/docs/es/v4/gates/) · [Referencia de comandos](https://karajancode.com/docs/es/v4/commands/).
61
63
 
64
+ ## Niveles de garantía: gobernado vs supervisado
65
+
66
+ Karajan **gobierna** a cualquier agente con gates de git — el falso verde es estructuralmente imposible escriba quien escriba, porque los gates viven en el repositorio, no en la buena voluntad del agente. Encima de eso, el **Karajan Sentinel** añade **supervisión dentro del turno**: `kj harden` cablea hooks deterministas al harness — uno registra el estado del método de la sesión según corren las herramientas (fuentes editadas vs tests tocados, escapes usados), y un hook Stop bloquea que el agente termine su turno mientras haya violaciones abiertas: fuentes editadas en la rama base, rama sin card, código sin tocar un solo test. Cada bloqueo dice la violación exacta y su remediación; `kj sentinel status` muestra lo que ve el supervisor. Ante cualquier duda falla en abierto — y lo dice — antes que colgar una sesión, y cada escape `KJ_ALLOW_*` queda registrado.
67
+
68
+ Los hooks síncronos con capacidad de bloqueo hoy solo existen en Claude Code. Eso hace explícito el montaje soportado: **para garantizar un harness que controle al LLM, usa Claude Code como anfitrión** — Claude escribe, Codex revisa (un subproceso de revisión no necesita hooks), y un tercer CLI arbitra cuando está disponible. Con cualquier otro anfitrión Karajan sigue gobernando al nivel completo de gates de git y te dice qué nivel está activo — jamás finge una supervisión que no puede imponer.
69
+
62
70
  ## Modo headless
63
71
 
64
72
  El pipeline multiagente clásico sigue vivo para CI y automatización: `kj run "<tarea>"` orquesta roles coder/reviewer/tester en subprocesos sin humano delante, con los mismos gates. Agentes y CI pasan `--non-interactive` (o `KJ_NON_INTERACTIVE=1`): los gates seguros se auto-responden y los findings FAIL paran el run con exit code de verdad. `kj advanced` lista la superficie completa. [Doc del modo headless](https://karajancode.com/docs/es/v4/headless/).
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "karajan-code",
3
- "version": "4.12.0",
3
+ "version": "4.14.0",
4
4
  "description": "Local multi-agent coding orchestrator with TDD, SonarQube, and code review pipeline",
5
5
  "type": "module",
6
6
  "license": "AGPL-3.0",
@@ -76,8 +76,8 @@
76
76
  "test": "vitest run",
77
77
  "test:watch": "vitest",
78
78
  "test:coverage": "vitest run --coverage",
79
- "lint": "eslint src/",
80
- "lint:fix": "eslint src/ --fix",
79
+ "lint": "eslint src/ packages/",
80
+ "lint:fix": "eslint src/ packages/ --fix",
81
81
  "lint:syntax": "((command -v rg >/dev/null 2>&1 && rg --files src tests | rg '\\.js$') || find src tests -type f -name '*.js') | while IFS= read -r f; do node --check \"$f\"; done",
82
82
  "format:check": "prettier --check .",
83
83
  "format:fix": "prettier --write .",
@@ -113,7 +113,9 @@ async function cmdPurge(root, id, out, err) {
113
113
  async function cmdEmpty(root, flags, out) {
114
114
  await ensureRoot(root);
115
115
  const m = await loadManifest(root);
116
- let dropped = [];
116
+ // Assigned in every branch below (the trailing else covers the default),
117
+ // so an initializer would never be read (no-useless-assignment).
118
+ let dropped;
117
119
  if (flags["older-than-days"]) {
118
120
  const days = Number(flags["older-than-days"]);
119
121
  if (!Number.isFinite(days) || days < 0) throw new Error("--older-than-days must be >= 0");
@@ -117,7 +117,7 @@ function humaniseProjectName(id) {
117
117
  if (!id || typeof id !== 'string') return id || '';
118
118
  const tail = id.split(/[/_]/).filter(Boolean).pop() || id;
119
119
  const words = tail
120
- .split(/[\s\-]+/)
120
+ .split(/[\s-]+/)
121
121
  .map((w) => w.replace(/[^\p{L}\p{N}]/gu, ''))
122
122
  .filter((w) => w.length > 0);
123
123
  if (words.length === 0) return id;
@@ -23,7 +23,6 @@
23
23
 
24
24
  import fs from 'node:fs';
25
25
  import path from 'node:path';
26
- import { homedir } from 'node:os';
27
26
  import {
28
27
  getDb,
29
28
  getKjHome,
@@ -18,7 +18,6 @@
18
18
 
19
19
  import { mkdirSync, openSync } from "node:fs";
20
20
  import { join, dirname } from "node:path";
21
- import { homedir } from "node:os";
22
21
  import { spawn } from "node:child_process";
23
22
  import { fileURLToPath } from "node:url";
24
23
  import { randomUUID } from "node:crypto";
@@ -30,7 +30,6 @@
30
30
 
31
31
  import { existsSync, mkdirSync, readFileSync, writeFileSync, copyFileSync, renameSync } from 'node:fs';
32
32
  import { join, dirname } from 'node:path';
33
- import { tmpdir } from 'node:os';
34
33
  import yaml from 'js-yaml';
35
34
  import { getKjHome } from './db.js';
36
35
 
@@ -448,7 +447,7 @@ export function readConfig({ scope = 'global' } = {}) {
448
447
  parsed = yaml.load(readFileSync(p, 'utf8'), { json: true }) || {};
449
448
  exists = true;
450
449
  } catch (err) {
451
- throw new Error(`No se pudo parsear ${p}: ${err.message}`);
450
+ throw new Error(`No se pudo parsear ${p}: ${err.message}`, { cause: err });
452
451
  }
453
452
  }
454
453
  const fields = EDITABLE_FIELDS.map((f) => {
@@ -150,8 +150,6 @@ export function cleanupEphemeralProjects({
150
150
  const candidates = findEphemeralProjects(allProjects, { ...opts, now: now() });
151
151
  if (candidates.length === 0) return [];
152
152
 
153
- const countStories = db.prepare("SELECT COUNT(*) AS n FROM stories WHERE project_id = ?");
154
- const countSessions = db.prepare("SELECT COUNT(*) AS n FROM sessions WHERE project_id = ?");
155
153
  const delStories = db.prepare("DELETE FROM stories WHERE project_id = ?");
156
154
  const delSessions = db.prepare("DELETE FROM sessions WHERE project_id = ?");
157
155
  const delProject = db.prepare("DELETE FROM projects WHERE id = ?");
@@ -16,7 +16,6 @@
16
16
  */
17
17
  import { readFileSync, existsSync, readdirSync, mkdirSync, openSync } from 'node:fs';
18
18
  import { join, dirname } from 'node:path';
19
- import { homedir } from 'node:os';
20
19
  import { writeJsonAtomicSync } from 'karajan-core/atomic-write';
21
20
  import { spawn } from 'node:child_process';
22
21
  import { trackRun, untrack } from './run-tracker.js';
@@ -28,7 +28,7 @@
28
28
  */
29
29
 
30
30
  import { execSync } from 'node:child_process';
31
- import { existsSync, readFileSync, readdirSync, statSync } from 'node:fs';
31
+ import { existsSync, readFileSync, readdirSync } from 'node:fs';
32
32
  import { join } from 'node:path';
33
33
  import { getHuBoardPlansDirs } from './db.js';
34
34
 
@@ -21,7 +21,6 @@ export async function readPlanCached(p) {
21
21
  planFileCache.set(p, { mtimeMs, plan });
22
22
  return plan;
23
23
  }
24
- import os from 'node:os';
25
24
  import path from 'node:path';
26
25
  import { spawn as spawnChild } from 'node:child_process';
27
26
  import {
@@ -36,7 +35,6 @@ import {
36
35
  getKjHome,
37
36
  getHuBoardRunsDir,
38
37
  getHuBoardPlansDir,
39
- getHuBoardLegacyPlansDir,
40
38
  getHuBoardPlansDirs,
41
39
  getStoryRow,
42
40
  listPlanIdsForProject,
@@ -93,19 +91,6 @@ function huStoriesDir() {
93
91
  return path.join(getKjHome(), 'hu-stories');
94
92
  }
95
93
 
96
- /**
97
- * Best-effort removal of the hu-stories/<id>/ directory.
98
- */
99
- async function removeBatchDir(batchId) {
100
- try {
101
- const dir = path.join(huStoriesDir(), batchId);
102
- await fsp.rm(dir, { recursive: true, force: true });
103
- return true;
104
- } catch {
105
- return false;
106
- }
107
- }
108
-
109
94
  /**
110
95
  * GET /api/dashboard - Global dashboard statistics.
111
96
  */
@@ -4,7 +4,6 @@ import rateLimit from 'express-rate-limit';
4
4
  import { join, dirname } from 'node:path';
5
5
  import { fileURLToPath, pathToFileURL } from 'node:url';
6
6
  import { writeFileSync, rmSync, mkdirSync, realpathSync } from 'node:fs';
7
- import { homedir } from 'node:os';
8
7
  import { initDb, closeDb } from './db.js';
9
8
  import { fullScan, startWatcher } from './sync.js';
10
9
  import apiRoutes from './routes/api.js';
@@ -1,5 +1,5 @@
1
1
  import { watch } from 'chokidar';
2
- import { readFileSync, readdirSync, existsSync, statSync, rmSync } from 'node:fs';
2
+ import { readFileSync, readdirSync, existsSync, rmSync } from 'node:fs';
3
3
  import { dirname, join, basename } from 'node:path';
4
4
  import { homedir } from 'node:os';
5
5
  import {
@@ -12,7 +12,6 @@
12
12
 
13
13
  import { randomBytes } from "node:crypto";
14
14
  import { mkdirSync, readFileSync, writeFileSync, chmodSync, existsSync } from "node:fs";
15
- import { homedir } from "node:os";
16
15
  import { dirname, join } from "node:path";
17
16
 
18
17
  const TOKEN_BYTES = 32;
@@ -45,7 +45,6 @@
45
45
 
46
46
  import fs from "node:fs";
47
47
  import path from "node:path";
48
- import os from "node:os";
49
48
 
50
49
  const DEFAULT_CODING_HOURS = 6;
51
50
  const DEFAULT_PAUSED_HOURS = 24;
@@ -0,0 +1,61 @@
1
+ import { BaseAgent } from "./base-agent.js";
2
+ import { resolveBin } from "./resolve-bin.js";
3
+
4
+ /**
5
+ * Antigravity CLI (`agy`) — KJC-TSK-0729 fase 2b: the ninth built-in agent.
6
+ * Google's official successor to the retired gemini CLI (dead 2026-06-18),
7
+ * covered by the user's Google AI Pro/Ultra subscription — the natural
8
+ * third AI for solomon arbitration, back at last.
9
+ *
10
+ * Verified live against agy 1.1.10:
11
+ * - `-p <prompt>` runs print mode; `--output-format json` emits ONE JSON
12
+ * object: {status: "SUCCESS", response, usage}. Non-SUCCESS status can
13
+ * arrive with exit 0, so ok checks BOTH.
14
+ * - `--disable-slash-commands` keeps prompt content (diffs can contain
15
+ * "/lines") from expanding as slash commands in print mode.
16
+ * - Permission prompts are denied in print mode unless
17
+ * `--dangerously-skip-permissions` (coder mode only).
18
+ * - Auth lives under ~/.gemini/antigravity-cli (device/browser login).
19
+ */
20
+ export class AgyAgent extends BaseAgent {
21
+ async runTask(task) {
22
+ return this._exec(task, this.getRoleModel(task.role || "coder"), ["--dangerously-skip-permissions"]);
23
+ }
24
+
25
+ async reviewTask(task) {
26
+ return this._exec(
27
+ { ...task, prompt: `Do NOT use any tools. Answer directly from this prompt only.\n\n${task.prompt}` },
28
+ this.getRoleModel(task.role || "reviewer"),
29
+ [],
30
+ );
31
+ }
32
+
33
+ async _exec(task, model, extraArgs) {
34
+ const args = ["-p", task.prompt, "--output-format", "json", "--disable-slash-commands", ...extraArgs];
35
+ if (model) args.push("--model", model);
36
+ const res = await this.runCommand(resolveBin("agy"), args, {
37
+ onOutput: task.onOutput,
38
+ silenceTimeoutMs: task.silenceTimeoutMs,
39
+ timeout: task.timeoutMs,
40
+ });
41
+ const parsed = extractAgyOutput(res.stdout);
42
+ const ok = res.exitCode === 0 && parsed?.status === "SUCCESS";
43
+ return {
44
+ ok,
45
+ output: parsed?.response ?? res.stdout,
46
+ error: ok ? res.stderr : parsed?.error || res.stderr || `agy status: ${parsed?.status || "unparseable"}`,
47
+ exitCode: res.exitCode,
48
+ };
49
+ }
50
+ }
51
+
52
+ /** Parse agy's single-object JSON print output ({status, response}), null if not JSON. */
53
+ export function extractAgyOutput(stdout) {
54
+ if (!stdout) return null;
55
+ try {
56
+ const obj = JSON.parse(stdout.trim());
57
+ return typeof obj === "object" && obj !== null ? obj : null;
58
+ } catch {
59
+ return null;
60
+ }
61
+ }
@@ -5,6 +5,8 @@ import { AiderAgent } from "./aider-agent.js";
5
5
  import { OpenCodeAgent } from "./opencode-agent.js";
6
6
  import { QwenAgent } from "./qwen-agent.js";
7
7
  import { CopilotAgent } from "./copilot-agent.js";
8
+ import { KimiAgent } from "./kimi-agent.js";
9
+ import { AgyAgent } from "./agy-agent.js";
8
10
 
9
11
  const agentRegistry = new Map();
10
12
 
@@ -53,3 +55,5 @@ registerAgent("aider", AiderAgent, { bin: "aider", installUrl: "https://aider.ch
53
55
  registerAgent("opencode", OpenCodeAgent, { bin: "opencode", installUrl: "https://opencode.ai" });
54
56
  registerAgent("qwen", QwenAgent, { bin: "qwen", installUrl: "https://github.com/QwenLM/qwen-code" });
55
57
  registerAgent("copilot", CopilotAgent, { bin: "copilot", installUrl: "https://docs.github.com/copilot/how-tos/use-copilot-agents/use-copilot-cli" });
58
+ registerAgent("kimi", KimiAgent, { bin: "kimi", installUrl: "https://github.com/MoonshotAI/kimi-code" });
59
+ registerAgent("agy", AgyAgent, { bin: "agy", installUrl: "https://antigravity.google" });
@@ -0,0 +1,44 @@
1
+ import { BaseAgent } from "./base-agent.js";
2
+ import { resolveBin } from "./resolve-bin.js";
3
+
4
+ /**
5
+ * Kimi Code (Moonshot, binary `kimi`) — KJC-TSK-0729 fase 2: the eighth
6
+ * built-in agent and the free-tier reserve of the reviewer/solomon panel.
7
+ *
8
+ * Verified live against kimi-code 0.27.0:
9
+ * - `-p <prompt>` runs one prompt non-interactively and prints the
10
+ * response (text output by default) — the prompt travels as an argument,
11
+ * same E2BIG caveat as claude/copilot (KJC-BUG-0121) for huge diffs.
12
+ * - `-m <model>` picks a model alias; `-y` auto-approves actions.
13
+ * - Auth is the device-code login (`kimi login`), no env keys involved;
14
+ * "No model configured" on a logged machine means the login did not
15
+ * populate the managed provider (re-run `kimi login`).
16
+ */
17
+ export class KimiAgent extends BaseAgent {
18
+ async runTask(task) {
19
+ // Coder mode IS an agentic run: file edits and shell need approval.
20
+ return this._exec(task, this.getRoleModel(task.role || "coder"), ["-y"]);
21
+ }
22
+
23
+ async reviewTask(task) {
24
+ // Review/arbitration answers from the prompt alone — no auto-approval,
25
+ // and an explicit no-tools instruction so a headless run never stalls
26
+ // waiting for a permission prompt nobody can answer.
27
+ return this._exec(
28
+ { ...task, prompt: `Do NOT use any tools. Answer directly from this prompt only.\n\n${task.prompt}` },
29
+ this.getRoleModel(task.role || "reviewer"),
30
+ [],
31
+ );
32
+ }
33
+
34
+ async _exec(task, model, extraArgs) {
35
+ const args = ["-p", task.prompt, ...extraArgs];
36
+ if (model) args.push("-m", model);
37
+ const res = await this.runCommand(resolveBin("kimi"), args, {
38
+ onOutput: task.onOutput,
39
+ silenceTimeoutMs: task.silenceTimeoutMs,
40
+ timeout: task.timeoutMs,
41
+ });
42
+ return { ok: res.exitCode === 0, output: res.stdout, error: res.stderr, exitCode: res.exitCode };
43
+ }
44
+ }
@@ -50,8 +50,11 @@ export function collectAiSurface({ projectDir = process.cwd(), home = os.homedir
50
50
  * Diff the current surface against the last-seen snapshot and persist the
51
51
  * new state. First run records the baseline silently.
52
52
  */
53
- export function checkAiSurface({ projectDir = process.cwd(), home = os.homedir(), statePath = path.join(os.homedir(), ".karajan", "ai-surface.json") } = {}) {
54
- const surface = collectAiSurface({ projectDir, home });
53
+ export function checkAiSurface({ projectDir = process.cwd(), home = os.homedir(), statePath = path.join(os.homedir(), ".karajan", "ai-surface.json"), extraSurface = [] } = {}) {
54
+ // KJC-TSK-0728: extraSurface carries entries the (async) caller collected
55
+ // outside config files — e.g. observed agent CLIs as "grok (cli)" — so
56
+ // they ride the same snapshot and the same "NEW since last check" drift.
57
+ const surface = [...new Set([...collectAiSurface({ projectDir, home }), ...extraSurface])].sort();
55
58
  let state = {};
56
59
  try { state = JSON.parse(readFileSync(statePath, "utf8")); } catch { /* first run or corrupt → baseline */ }
57
60
  const prev = Array.isArray(state[projectDir]) ? state[projectDir] : null;
@@ -12,7 +12,7 @@
12
12
  * the install command and can offer to run it).
13
13
  */
14
14
 
15
- import { checkBinary, KNOWN_AGENTS } from "../utils/agent-detect.js";
15
+ import { checkBinary, KNOWN_AGENTS, detectObservedAgents } from "../utils/agent-detect.js";
16
16
  import { runCommand } from "../utils/process.js";
17
17
  import { withDocLink } from "../utils/doc-links.js";
18
18
  import { getInstallHint, appliesToStack } from "../utils/install-hints.js";
@@ -38,6 +38,33 @@ function createAgentCheck(agent) {
38
38
  };
39
39
  }
40
40
 
41
+ /**
42
+ * KJC-TSK-0728 — ONE aggregate line for the observation census (agent CLIs
43
+ * kj sees but does not drive). Only FOUND CLIs are listed; the missing ones
44
+ * make no noise, and the check never fails. `detector` is injectable for
45
+ * tests.
46
+ */
47
+ export function createObservedAgentsCheck({ detector = detectObservedAgents } = {}) {
48
+ return {
49
+ name: "agents:observed",
50
+ label: "Agent CLIs observed (not driven)",
51
+ strategy: STRATEGY.MANUAL,
52
+ async detect() {
53
+ let found = [];
54
+ try {
55
+ found = (await detector()).filter((a) => a.available);
56
+ } catch { /* census is best-effort */ }
57
+ return {
58
+ ok: true,
59
+ severity: "info",
60
+ detail: found.length
61
+ ? found.map((a) => `${a.name} (${(a.version || "").split(" ").at(-1) || "?"})`).join(", ")
62
+ : "none detected",
63
+ };
64
+ },
65
+ };
66
+ }
67
+
41
68
  /**
42
69
  * Check the presence of a core binary (node, npm, git).
43
70
  */
@@ -165,6 +192,7 @@ export function getBinaryChecks() {
165
192
  for (const agent of KNOWN_AGENTS) {
166
193
  checks.push(createAgentCheck(agent));
167
194
  }
195
+ checks.push(createObservedAgentsCheck());
168
196
  for (const bin of ["node", "npm", "git"]) {
169
197
  checks.push(createCoreBinaryCheck(bin));
170
198
  }
@@ -31,7 +31,7 @@ export const ADVANCED_GROUPS = [
31
31
  { title: "Análisis pre-run", commands: ["discover", "triage", "researcher", "architect", "onboard", "brief"] },
32
32
  { title: "Búsqueda / RAG", commands: ["rag", "qmd", "watch"] },
33
33
  { title: "Calidad / auditoría", commands: ["audit", "check", "mutate", "webperf", "sonar", "privacy", "release"] },
34
- { title: "Sesión / board", commands: ["resume", "report", "board", "hu", "adr", "worktree", "undo", "standby"] },
34
+ { title: "Sesión / board", commands: ["resume", "report", "board", "hu", "adr", "worktree", "undo", "standby", "sentinel"] },
35
35
  { title: "Infra / setup", commands: ["install-tools", "ollama", "skills", "roles", "agents", "env"] },
36
36
  { title: "Mantenimiento", commands: ["clean", "sync", "telemetry", "report-issue"] },
37
37
  ];
@@ -28,6 +28,10 @@ import { worktreeCommand } from "../commands/worktree.js";
28
28
  import { addAdr, listAdrs } from "../environment/adr.js";
29
29
  import { formatAdvancedIndex } from "../commands/advanced.js";
30
30
  import { withConfig } from "./_shared.js";
31
+ import { existsSync } from "node:fs";
32
+ import { join } from "node:path";
33
+ import { spawnSync } from "node:child_process";
34
+ import { verifySentinelScripts, resolveSentinelRoot } from "../harden/sentinel-hooks.js";
31
35
 
32
36
  /**
33
37
  * Register the "meta" / single-role / housekeeping commands: pre-pipeline
@@ -280,6 +284,29 @@ export function registerMeta(program, { pkgVersion }) {
280
284
  });
281
285
  });
282
286
 
287
+ // KJC-TSK-0713 — the Sentinel: the method state the harness hooks record.
288
+ const sentinel = program.command("sentinel").description("Deterministic method supervisor wired into the harness hooks (Claude Code)");
289
+ sentinel.command("status")
290
+ .description("Print the per-session method facts (sources/tests edited, escapes used) and any open violation the Stop gate would block on")
291
+ .action(() => {
292
+ const script = join(resolveSentinelRoot(), ".karajan", "harness", "stop.mjs");
293
+ if (!existsSync(script)) {
294
+ console.log("sentinel: not installed in this project — run `kj harden` (standard profile or higher)");
295
+ process.exitCode = 1;
296
+ return;
297
+ }
298
+ process.exitCode = spawnSync("node", [script, "--status"], { stdio: "inherit" }).status ?? 0;
299
+ });
300
+ sentinel.command("verify")
301
+ .description("Verify the harness scripts match this kj install — the root of trust is the installed package, not the project tree; exit 1 lists what was modified")
302
+ .option("--json", "Machine-readable result")
303
+ .action((flags) => {
304
+ const res = verifySentinelScripts();
305
+ if (flags.json) process.stdout.write(`${JSON.stringify(res)}\n`);
306
+ else console.log(res.ok ? "sentinel verify: scripts intactos" : `sentinel verify: modificados fuera de kj harden: ${res.mismatched.join(", ")} — restaura con kj harden`);
307
+ process.exitCode = res.ok ? 0 : 1;
308
+ });
309
+
283
310
  // KJC-TSK-0704 — the outbound privacy boundary: audit before anything ships.
284
311
  const privacy = program.command("privacy").description("Personal-data (PII) auditing of outbound boundaries: staged diffs, build outputs, docs trees");
285
312
  privacy.command("scan [paths...]")
@@ -7,6 +7,7 @@
7
7
  import { checkHarden } from "../harden/check.js";
8
8
  import { collectMethodStats, formatMethodStats } from "../checks/method.js";
9
9
  import { checkAiSurface, formatAiSurface } from "../checks/ai-surface.js";
10
+ import { detectObservedAgents } from "../utils/agent-detect.js";
10
11
 
11
12
  export async function checkCommand({ projectDir = process.cwd(), profile = "standard", json = false, logger = console } = {}) {
12
13
  const result = await checkHarden({ projectDir, profile });
@@ -14,8 +15,15 @@ export async function checkCommand({ projectDir = process.cwd(), profile = "stan
14
15
  // rides along in check output but never affects the exit code.
15
16
  const method = await collectMethodStats({ projectDir }).catch(() => null);
16
17
  // KJC-TSK-0694: same deal for the MCP inventory — a nudge, never a gate.
18
+ // KJC-TSK-0728: observed agent CLIs ride the same snapshot as "(cli)"
19
+ // entries, so a newly-appeared agent binary trips the same drift question.
17
20
  let aiSurface = null;
18
- try { aiSurface = checkAiSurface({ projectDir }); } catch { /* inventory is best-effort */ }
21
+ try {
22
+ const clis = (await detectObservedAgents().catch(() => []))
23
+ .filter((a) => a.available)
24
+ .map((a) => `${a.name} (cli)`);
25
+ aiSurface = checkAiSurface({ projectDir, extraSurface: clis });
26
+ } catch { /* inventory is best-effort */ }
19
27
 
20
28
  if (json) {
21
29
  logger.info?.(JSON.stringify({ ...result, method, aiSurface }));
@@ -19,6 +19,7 @@ import { installGuidelines } from "../harden/guidelines-engine.js";
19
19
  import { commandsForLanguage } from "../harden/hook-commands.js";
20
20
  import { installHooks } from "../harden/harden-engine.js";
21
21
  import { installHarnessHooks } from "../harden/harness-hooks.js";
22
+ import { installSentinelHooks } from "../harden/sentinel-hooks.js";
22
23
  import { detectStackRoots } from "../harden/stack-roots.js";
23
24
  import { installWorkflows } from "../harden/workflow-engine.js";
24
25
  import { detectTestFramework } from "../utils/project-detect.js";
@@ -146,6 +147,10 @@ export async function hardenCommand({
146
147
  if (profile !== "minimal" && !dryRun) {
147
148
  const hh = installHarnessHooks({ projectDir, logger });
148
149
  result.harnessHooks = hh.wired ? "wired" : "script-only";
150
+ // KJC-TSK-0713 — the Sentinel: method state + Stop gate (turn cannot
151
+ // end red). Same Claude-only harness surface as the tool gate.
152
+ const sh = installSentinelHooks({ projectDir, logger });
153
+ result.sentinelHooks = sh.wired ? "wired" : "script-only";
149
154
  }
150
155
  } catch (err) {
151
156
  if (json) logger.info?.(JSON.stringify({ ok: false, error: err.message }));
@@ -162,6 +162,9 @@ async function buildReport(dir, sessionId) {
162
162
  if (session.pg_task_id) report.pg_task_id = session.pg_task_id;
163
163
  if (session.pg_project_id) report.pg_project_id = session.pg_project_id;
164
164
  if (session.rtk_savings) report.rtk_savings = session.rtk_savings;
165
+ // The session may belong to another workspace: its snapshot knows the
166
+ // project dir the sentinel state lives in (same source solomon-rules uses).
167
+ if (session.config_snapshot?.projectDir) report.project_dir = session.config_snapshot.projectDir;
165
168
  return report;
166
169
  }
167
170
 
@@ -210,6 +213,36 @@ function printTextReport(report) {
210
213
  console.log(formatCommitsText(report.commits_generated));
211
214
  }
212
215
 
216
+ /**
217
+ * KJC-TSK-0715 — escapes are the user's decisions, never the agent's silent
218
+ * shortcuts: every KJ_ALLOW_* the sentinel honored surfaces in the report.
219
+ */
220
+ export function formatSentinelEscapes(state) {
221
+ const events = Array.isArray(state?.escape_events) ? state.escape_events : [];
222
+ if (events.length === 0) return null;
223
+ const lines = events.map(
224
+ (e) => ` ${e.ts ? new Date(e.ts).toISOString() : "?"} ${e.escape} (tool: ${e.tool || "?"}, session: ${e.sid || "?"})`
225
+ );
226
+ return `Sentinel escapes used (${events.length}):\n${lines.join("\n")}`;
227
+ }
228
+
229
+ // The sentinel state is a PROJECT artifact (<project>/.karajan/harness), not
230
+ // a session artifact: the `dir` the report handlers receive is the GLOBAL
231
+ // sessions root (~/.karajan/sessions) and never contains it. `kj report` is
232
+ // run from inside the project, so the project dir is the cwd — overridable
233
+ // for callers reporting on another workspace.
234
+ async function printSentinelEscapes({ projectDir = process.cwd(), format } = {}) {
235
+ if (format === "json") return;
236
+ try {
237
+ const raw = await fs.readFile(path.join(projectDir, ".karajan", "harness", "sentinel-state.json"), "utf8");
238
+ const text = formatSentinelEscapes(JSON.parse(raw));
239
+ if (text) {
240
+ console.log("");
241
+ console.log(text);
242
+ }
243
+ } catch { /* no sentinel state in this project — nothing to report */ }
244
+ }
245
+
213
246
  function formatDuration(ms) {
214
247
  if (ms === null || ms === undefined) return "-";
215
248
  if (ms < 1000) return `${ms}ms`;
@@ -384,6 +417,7 @@ async function handlePgTaskReport({ dir, pgTask, list, sessionId, format, trace,
384
417
  } else {
385
418
  printTextReport(report);
386
419
  }
420
+ await printSentinelEscapes({ projectDir: report.project_dir, format });
387
421
  }
388
422
 
389
423
  async function handleSingleSessionReport({ dir, entries, sessionId, format, trace, currency }) {
@@ -405,9 +439,10 @@ async function handleSingleSessionReport({ dir, entries, sessionId, format, trac
405
439
  if (trace) {
406
440
  const { cur, rate } = await resolveTraceOptions(currency);
407
441
  printTraceReport(report, cur, rate);
408
- return;
442
+ } else {
443
+ printTextReport(report);
409
444
  }
410
- printTextReport(report);
445
+ await printSentinelEscapes({ projectDir: report.project_dir, format });
411
446
  }
412
447
 
413
448
  export async function reportCommand({ list = false, sessionId = null, format = "text", trace = false, currency = "usd", pgTask = null }) {
@@ -81,7 +81,14 @@ function applyRoleOverrides(out, flags) {
81
81
  // boolean means the flag belongs to the command itself (`kj audit
82
82
  // --security`, KJC-TSK-0695) and must not become a provider.
83
83
  for (const [flag, role] of ROLE_PROVIDER_FLAGS) {
84
- if (typeof flags[flag] === "string" && flags[flag]) out.roles[role].provider = flags[flag];
84
+ if (typeof flags[flag] === "string" && flags[flag]) {
85
+ // A model pin belongs to the provider it was written for — switching
86
+ // provider by flag drops it (KJC-TSK-0729: codex's pinned mini reached
87
+ // agy, which answered with its model list instead of a verdict). An
88
+ // explicit --<role>-model flag re-pins below, after this loop.
89
+ if (out.roles[role].provider && out.roles[role].provider !== flags[flag]) out.roles[role].model = null;
90
+ out.roles[role].provider = flags[flag];
91
+ }
85
92
  }
86
93
  // coder/reviewer also update top-level aliases
87
94
  if (flags.coder) out.coder = flags.coder;