jules-orchestrator-kit 0.35.2 → 0.37.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +3 -3
- package/bin/agentctl.mjs +31 -5
- package/index.mjs +1 -0
- package/package.json +1 -1
- package/src/budget.mjs +31 -9
- package/src/security.mjs +132 -5
- package/src/web-templates.mjs +87 -0
package/README.md
CHANGED
|
@@ -104,7 +104,7 @@ Autonomous coding agents can write software at 100× human speed—but unconstra
|
|
|
104
104
|
|
|
105
105
|
* **🔒 Zero Runtime Dependencies:** Built exclusively on Node.js 20+ built-ins (`node:fs`, `node:child_process`, `node:crypto`, `node:path`, `node:http`, `node:tty`, `node:test`). Zero third-party npm packages mean zero supply-chain CVE risk.
|
|
106
106
|
|
|
107
|
-
* **🛡️ Fail-Closed Security Gatekeeper:** Unconditionally evaluates explicit Deny rules *before* Allow rules, matching against **canonicalised, case-folded paths** so `./`, `..`, mixed separators or a `.GitHub/` spelling cannot walk past a rule (the same repo is checked out on case-insensitive macOS and Windows filesystems). Redacts high-entropy secrets and PII from dry-runs and git diffs
|
|
107
|
+
* **🛡️ Fail-Closed Security Gatekeeper:** Unconditionally evaluates explicit Deny rules *before* Allow rules, matching against **canonicalised, case-folded paths** so `./`, `..`, mixed separators or a `.GitHub/` spelling cannot walk past a rule (the same repo is checked out on case-insensitive macOS and Windows filesystems). Redacts high-entropy secrets and PII from dry-runs and git diffs — **including credentials wrapped in base64**, so a key inside a Kubernetes `Secret` manifest is not invisible to a line-oriented scanner — blocks unsupported Node.js native module imports in Edge environments (Cloudflare Workers, Vercel Edge, Netlify Edge), and rejects PRs exceeding the 75 KB Diff Payload governor.
|
|
108
108
|
|
|
109
109
|
* **🔄 Autonomous OODA Self-Healing:** Captures test stderr/stdout, normalizes failure fingerprints, and feeds structured error contexts back into repair iterations (up to 3 automatic attempts) before human escalation.
|
|
110
110
|
|
|
@@ -122,7 +122,7 @@ Autonomous coding agents can write software at 100× human speed—but unconstra
|
|
|
122
122
|
|
|
123
123
|
* **🚀 Zero-Test Bootstrapping (`agentctl bootstrap`):** Synthesizes deterministic syntax-check and smoke-test verification oracles for untested legacy repositories so agents always operate against a falsifiable feedback loop.
|
|
124
124
|
|
|
125
|
-
* **📈 Proven Scale & Reliability:** Empirically tested with **
|
|
125
|
+
* **📈 Proven Scale & Reliability:** Empirically tested with **547 unit tests across 80 suites passing in < 10.0s**. An adversarial red-team suite (`test/adversarial-claims.test.mjs`) continuously attempts to falsify the safety guarantees documented above — every probe in it currently holds, with no open gaps — including cross-platform probes for the case-insensitive filesystems on macOS and Windows — and a documentation-sync gate (`scripts/doc-sync-check.mjs`) blocks any release whose docs have drifted from the code.
|
|
126
126
|
|
|
127
127
|
<br/>
|
|
128
128
|
|
|
@@ -389,7 +389,7 @@ Native stdio server exposing task dispatch, gate verification, and risk auditing
|
|
|
389
389
|
| :--- | :--- | :--- | :--- |
|
|
390
390
|
| `init` | `agentctl init [--interactive] [--tier pro]` | Interactive onboarding wizard & stack oracle inspector generating `.agent/config.yml`. | `0` (Created) |
|
|
391
391
|
| `task create` | `agentctl task create [--title <t>] [--prompt <p>] [--template <id>] [--role <name>] [--tier fast\|complex] [--depends-on <id,...>]` | Interactively authors & scopes falsifiable task envelopes with secret scrubbing, preflight gate checks, specialist role resolution, DAG dependency wiring, and an optional Cost Router tier override. | `0` (Queued), `1` (Unfalsifiable / Secret leak) |
|
|
392
|
-
| `task template` | `agentctl task template [<id>] [--list] [--json]` | Lists and synthesizes specialized web task envelopes (`web-cwv`, `web-wcag`, `web-seo`, `web-playwright`, `web-flaky-heal`, `web-i18n`). | `0` (Synthesized/Listed) |
|
|
392
|
+
| `task template` | `agentctl task template [<id>] [--list] [--json]` | Lists and synthesizes specialized web task envelopes (`web-cwv`, `web-wcag`, `web-seo`, `web-playwright`, `web-flaky-heal`, `web-i18n`, `web-ai-access`). | `0` (Synthesized/Listed) |
|
|
393
393
|
| `task optimize` | `agentctl task optimize "<prompt>" [--fix] [--web] [--json]` | Linter & optimizer injecting Google Labs 3-phase exploration budgets, critic steering, and web oracles. | `0` (Scored/Fixed) |
|
|
394
394
|
| `test-gen` | `agentctl test-gen --title <t> --spec <s> [--run]` | Scaffolds falsifiable unit tests, verifies **RED** failure state, and locks test in `scope.deny`. | `0` (Scaffolded/Red) |
|
|
395
395
|
| `rollback` | `agentctl rollback [sessionId \| --latest]` | Restores exact commit, uncommitted files, and cleans orphan task worktrees from pre-flight checkpoints. | `0` (Restored), `1` (Error) |
|
package/bin/agentctl.mjs
CHANGED
|
@@ -334,18 +334,41 @@ async function main() {
|
|
|
334
334
|
// The count is local-only, so an operator who knows their real usage
|
|
335
335
|
// must be able to correct it. Appending `budget_released` keeps the
|
|
336
336
|
// hash chain intact — the ledger is corrected forwards, never edited.
|
|
337
|
+
// An unrecognised flag is refused rather than ignored. `reset` writes to
|
|
338
|
+
// the ledger, and a misremembered flag — `--root`, `--force` — silently
|
|
339
|
+
// dropping through to a full release is the kind of misfire that only
|
|
340
|
+
// becomes visible after the count is already gone.
|
|
341
|
+
const known = new Set(["--dry-run", "--yes", "-y", "--all"]);
|
|
342
|
+
const unknown = args.slice(2).filter((a) => !known.has(a));
|
|
343
|
+
if (unknown.length) {
|
|
344
|
+
console.error(`Unrecognised option${unknown.length > 1 ? "s" : ""} for \`budget reset\`: ${unknown.join(", ")}`);
|
|
345
|
+
console.error("Accepted: --dry-run, --yes/-y, --all. Nothing was released.");
|
|
346
|
+
process.exit(2);
|
|
347
|
+
}
|
|
337
348
|
const dryRun = args.includes("--dry-run");
|
|
338
349
|
const confirmed = args.includes("--yes") || args.includes("-y");
|
|
350
|
+
// Committed reservations reached the provider, so releasing them makes
|
|
351
|
+
// the local count understate real usage. That has to be deliberate.
|
|
352
|
+
const includeCommitted = args.includes("--all");
|
|
339
353
|
if (!dryRun && !confirmed) {
|
|
340
354
|
const open = listOpenReservations(root);
|
|
341
|
-
|
|
355
|
+
const committed = open.filter((r) => r.committed).length;
|
|
356
|
+
const target = includeCommitted ? open.length : open.length - committed;
|
|
357
|
+
console.log(`Would release ${target} open reservation(s) from the last 24 hours.`);
|
|
358
|
+
if (committed > 0 && !includeCommitted) {
|
|
359
|
+
console.log(`Keeping ${committed} that reached Jules — those sessions really spent quota.`);
|
|
360
|
+
console.log("Add --all to release them too, if you know the count is still wrong.");
|
|
361
|
+
}
|
|
342
362
|
console.log("This rewrites nothing — it appends `budget_released` entries.");
|
|
343
363
|
console.log("Re-run with --yes to confirm, or --dry-run for detail.");
|
|
344
364
|
process.exit(0);
|
|
345
365
|
}
|
|
346
|
-
const res = releaseOpenReservations({ root, dryRun, reason: "operator-reconcile" });
|
|
366
|
+
const res = releaseOpenReservations({ root, dryRun, includeCommitted, reason: "operator-reconcile" });
|
|
347
367
|
const verb = dryRun ? "Would release" : "Released";
|
|
348
|
-
console.log(`${verb} ${res.released} reservation(s)
|
|
368
|
+
console.log(`${verb} ${res.released} of ${res.open} open reservation(s) — ${res.uncommitted} never closed, ${res.committed} committed.`);
|
|
369
|
+
if (res.kept > 0) {
|
|
370
|
+
console.log(`Kept ${res.kept} committed reservation(s); re-run with --all to release those as well.`);
|
|
371
|
+
}
|
|
349
372
|
if (!dryRun) {
|
|
350
373
|
const after = budgetStatus(loadConfig(root), root);
|
|
351
374
|
console.log(`Daily Budget : ${formatBudgetLine(after)}`);
|
|
@@ -363,14 +386,17 @@ async function main() {
|
|
|
363
386
|
console.log(` ${b.note}`);
|
|
364
387
|
console.log(` Window opened at ${b.windowStart} — the quota resets ${b.windowHours}h after each task,`);
|
|
365
388
|
console.log(" not at midnight, so this count spans yesterday's ledger too.");
|
|
366
|
-
|
|
389
|
+
const openNow = listOpenReservations(root);
|
|
390
|
+
const committedNow = openNow.filter((r) => r.committed).length;
|
|
391
|
+
console.log(` Open reservations in the window: ${openNow.length} (${committedNow} confirmed dispatched, ${openNow.length - committedNow} never closed)`);
|
|
367
392
|
console.log("");
|
|
368
393
|
console.log(`Worker Slots : ${slots.concurrency} concurrent`);
|
|
369
394
|
console.log(` ${slots.note}`);
|
|
370
395
|
console.log("");
|
|
371
396
|
console.log("The ledger counts this checkout only — sessions started from the Jules");
|
|
372
397
|
console.log("web UI or another machine spend the same quota without appearing here.");
|
|
373
|
-
console.log("Use `agentctl budget reset --yes` to
|
|
398
|
+
console.log("Use `agentctl budget reset --yes` to clear the ones that never closed,");
|
|
399
|
+
console.log("or `--all` to also give back the ones that did reach Jules.");
|
|
374
400
|
process.exit(0);
|
|
375
401
|
break;
|
|
376
402
|
}
|
package/index.mjs
CHANGED
package/package.json
CHANGED
package/src/budget.mjs
CHANGED
|
@@ -276,32 +276,54 @@ export function listOpenReservations(root = resolveRoot(), opts = {}) {
|
|
|
276
276
|
* forwards — never by editing or deleting the file, which would break the chain
|
|
277
277
|
* and destroy the audit trail the ledger exists to provide.
|
|
278
278
|
*
|
|
279
|
-
* This is an operator override, not an inference
|
|
280
|
-
* reservation
|
|
281
|
-
*
|
|
279
|
+
* This is an operator override, not an inference — but the two kinds of open
|
|
280
|
+
* reservation are not equally likely to be wrong, and by default only one of
|
|
281
|
+
* them is released.
|
|
282
|
+
*
|
|
283
|
+
* An **uncommitted** reservation was taken and never resolved: the process died
|
|
284
|
+
* between reserving the slot and learning what happened to it. That is the
|
|
285
|
+
* phantom this function exists to clear.
|
|
286
|
+
*
|
|
287
|
+
* A **committed** one carries proof that the dispatch reached the provider, so
|
|
288
|
+
* a real session exists and the quota really was spent. Releasing it makes the
|
|
289
|
+
* local count *understate* actual usage, which is the dangerous direction: the
|
|
290
|
+
* kit then dispatches confidently past the provider's real limit and gets
|
|
291
|
+
* refused. `includeCommitted` therefore has to be asked for.
|
|
292
|
+
*
|
|
293
|
+
* Legacy id-less reservations have no `budget_committed` entry that could ever
|
|
294
|
+
* name them, so they always read as uncommitted and are always released.
|
|
282
295
|
*
|
|
283
296
|
* @param {object} [opts]
|
|
284
297
|
* @param {string} [opts.root]
|
|
285
298
|
* @param {string} [opts.reason] - Recorded on every released entry.
|
|
286
299
|
* @param {boolean} [opts.dryRun] - Report what would be released, write nothing.
|
|
287
|
-
* @
|
|
300
|
+
* @param {boolean} [opts.includeCommitted] - Also release reservations that
|
|
301
|
+
* demonstrably reached the provider. Off by default.
|
|
302
|
+
* @returns {{ released: number, kept: number, open: number, committed: number,
|
|
303
|
+
* uncommitted: number, anonymous: number, ids: string[],
|
|
304
|
+
* includeCommitted: boolean, dryRun: boolean }}
|
|
288
305
|
*/
|
|
289
306
|
export function releaseOpenReservations(opts = {}) {
|
|
290
307
|
const root = opts.root || resolveRoot();
|
|
291
308
|
const openRecords = listOpenReservations(root);
|
|
292
309
|
const committed = openRecords.filter((r) => r.committed).length;
|
|
310
|
+
const includeCommitted = Boolean(opts.includeCommitted);
|
|
311
|
+
const targets = includeCommitted ? openRecords : openRecords.filter((r) => !r.committed);
|
|
293
312
|
|
|
294
313
|
const result = {
|
|
295
|
-
released:
|
|
314
|
+
released: targets.length,
|
|
315
|
+
kept: openRecords.length - targets.length,
|
|
316
|
+
open: openRecords.length,
|
|
296
317
|
committed,
|
|
297
318
|
uncommitted: openRecords.length - committed,
|
|
298
|
-
anonymous:
|
|
299
|
-
ids:
|
|
319
|
+
anonymous: targets.filter((r) => !r.reservationId).length,
|
|
320
|
+
ids: targets.map((r) => r.reservationId).filter(Boolean),
|
|
321
|
+
includeCommitted,
|
|
300
322
|
dryRun: Boolean(opts.dryRun),
|
|
301
323
|
};
|
|
302
|
-
if (opts.dryRun ||
|
|
324
|
+
if (opts.dryRun || targets.length === 0) return result;
|
|
303
325
|
|
|
304
|
-
for (const rec of
|
|
326
|
+
for (const rec of targets) {
|
|
305
327
|
const entry = { event: "budget_released", reason: opts.reason || "operator-reconcile" };
|
|
306
328
|
if (rec.reservationId) {
|
|
307
329
|
entry.reservationId = rec.reservationId;
|
package/src/security.mjs
CHANGED
|
@@ -142,6 +142,19 @@ export function redactSecrets(text) {
|
|
|
142
142
|
sanitized = sanitized.replace(pat, "[REDACTED_BY_SECURITY_GATE]");
|
|
143
143
|
}
|
|
144
144
|
|
|
145
|
+
// A key the scanner can find inside a base64 blob must not survive redaction
|
|
146
|
+
// just because the literal bytes differ — otherwise scanDiff blocks the
|
|
147
|
+
// dispatch and the escalation payload leaks the very value it blocked on. The
|
|
148
|
+
// whole blob goes, not part of it: a partially-redacted encoding still
|
|
149
|
+
// decodes to the key.
|
|
150
|
+
const encoded = new Set();
|
|
151
|
+
decodeBase64Blobs(sanitized, (plain, blob) => {
|
|
152
|
+
if (hasHighConfidenceSecret(plain)) encoded.add(blob);
|
|
153
|
+
});
|
|
154
|
+
for (const blob of encoded) {
|
|
155
|
+
sanitized = sanitized.split(blob).join("[REDACTED_ENCODED_SECRET]");
|
|
156
|
+
}
|
|
157
|
+
|
|
145
158
|
return sanitized;
|
|
146
159
|
}
|
|
147
160
|
|
|
@@ -374,14 +387,118 @@ const STRING_CONCAT_JOIN = /(["'`])\s*\+\s*(["'`])/g;
|
|
|
374
387
|
* Produces the variants of the added-line text that secret patterns are run
|
|
375
388
|
* against: as-written, with invisible characters stripped, and with
|
|
376
389
|
* source-level string concatenation collapsed.
|
|
390
|
+
*
|
|
391
|
+
* `normalized` is the last of those — every normalisation applied. It is
|
|
392
|
+
* returned separately for checks that are too expensive to run three times and
|
|
393
|
+
* gain nothing from the intermediate forms.
|
|
394
|
+
*
|
|
377
395
|
* @param {string} addedLines
|
|
378
|
-
* @returns {string[]}
|
|
396
|
+
* @returns {{ all: string[], normalized: string }}
|
|
379
397
|
*/
|
|
380
398
|
function secretScanVariants(addedLines) {
|
|
381
399
|
const stripped = addedLines.replace(INVISIBLE_CHARS, "");
|
|
382
400
|
// Collapse `"AAA" +\n "BBB"` into `"AAABBB"` before matching.
|
|
383
401
|
const dejoined = stripped.replace(/\s*\n\s*/g, " ").replace(STRING_CONCAT_JOIN, "");
|
|
384
|
-
return [...new Set([addedLines, stripped, dejoined])];
|
|
402
|
+
return { all: [...new Set([addedLines, stripped, dejoined])], normalized: dejoined };
|
|
403
|
+
}
|
|
404
|
+
|
|
405
|
+
// Base64 is less an evasion technique than a file format. Every value in a
|
|
406
|
+
// Kubernetes Secret manifest is base64 by specification, and whole `.env` files
|
|
407
|
+
// get encoded into a single CI variable. A credential arriving that way is
|
|
408
|
+
// ordinary rather than adversarial — and a line-oriented scanner walks straight
|
|
409
|
+
// past it, which makes this the encoding most likely to carry a live key
|
|
410
|
+
// through the gate.
|
|
411
|
+
const BASE64_CANDIDATE = /[A-Za-z0-9+/]{24,}={0,2}/g;
|
|
412
|
+
|
|
413
|
+
// Decoding is cheap per blob and ruinous per diff if left unbounded. A patch
|
|
414
|
+
// that checks in a binary, a source map or a bundled font is otherwise enough
|
|
415
|
+
// to turn one scan into a memory event, so both the number of candidates and
|
|
416
|
+
// the total decoded size are capped. Exceeding a cap skips the remainder; it
|
|
417
|
+
// does not fail the scan, because a large diff is not evidence of a leak.
|
|
418
|
+
const BASE64_MAX_CANDIDATES = 64;
|
|
419
|
+
const BASE64_MAX_DECODED_BYTES = 64 * 1024;
|
|
420
|
+
|
|
421
|
+
/**
|
|
422
|
+
* Share of characters that are printable ASCII (plus tab/newline/return).
|
|
423
|
+
*
|
|
424
|
+
* `Buffer.from(str, "base64")` never throws — it discards what it cannot parse
|
|
425
|
+
* and returns whatever it managed to decode. So a hex digest or a random
|
|
426
|
+
* identifier of the right length "decodes" successfully into bytes that mean
|
|
427
|
+
* nothing. A wrapped credential, on the other hand, decodes to text: keys,
|
|
428
|
+
* PEM blocks and `.env` bodies are all ASCII. This ratio is what separates the
|
|
429
|
+
* two, and it removes nearly all of the noise before any pattern runs.
|
|
430
|
+
*/
|
|
431
|
+
function printableRatio(str) {
|
|
432
|
+
if (!str) return 0;
|
|
433
|
+
let printable = 0;
|
|
434
|
+
for (let i = 0; i < str.length; i++) {
|
|
435
|
+
const c = str.charCodeAt(i);
|
|
436
|
+
if (c === 9 || c === 10 || c === 13 || (c >= 32 && c <= 126)) printable++;
|
|
437
|
+
}
|
|
438
|
+
return printable / str.length;
|
|
439
|
+
}
|
|
440
|
+
|
|
441
|
+
/**
|
|
442
|
+
* Decode the base64-looking blobs in `text` that plausibly hold text.
|
|
443
|
+
*
|
|
444
|
+
* @param {string} text
|
|
445
|
+
* @param {(plain: string, blob: string) => void} [onDecoded] - Called per blob.
|
|
446
|
+
* @returns {string[]}
|
|
447
|
+
*/
|
|
448
|
+
function decodeBase64Blobs(text, onDecoded) {
|
|
449
|
+
if (!text) return [];
|
|
450
|
+
const decoded = [];
|
|
451
|
+
let candidates = 0;
|
|
452
|
+
let bytes = 0;
|
|
453
|
+
|
|
454
|
+
BASE64_CANDIDATE.lastIndex = 0;
|
|
455
|
+
let match;
|
|
456
|
+
while ((match = BASE64_CANDIDATE.exec(text)) !== null) {
|
|
457
|
+
if (candidates++ >= BASE64_MAX_CANDIDATES) break;
|
|
458
|
+
const blob = match[0];
|
|
459
|
+
// Valid base64 is a multiple of four characters including padding. This
|
|
460
|
+
// costs nothing and rejects three quarters of the alphanumeric runs — commit
|
|
461
|
+
// hashes, minified identifiers — that would otherwise be decoded for nothing.
|
|
462
|
+
if (blob.length % 4 !== 0) continue;
|
|
463
|
+
// Check the budget against what this blob *would* cost, not against what
|
|
464
|
+
// has already been spent — otherwise the first candidate decodes in full
|
|
465
|
+
// however large it is, and a single checked-in binary costs more than the
|
|
466
|
+
// cap was meant to allow. `continue`, not `break`: one oversized blob must
|
|
467
|
+
// not hide the smaller ones after it.
|
|
468
|
+
if (bytes + (blob.length * 3) / 4 > BASE64_MAX_DECODED_BYTES) continue;
|
|
469
|
+
|
|
470
|
+
let plain;
|
|
471
|
+
try {
|
|
472
|
+
plain = Buffer.from(blob, "base64").toString("utf-8");
|
|
473
|
+
} catch (_) {
|
|
474
|
+
continue;
|
|
475
|
+
}
|
|
476
|
+
bytes += plain.length;
|
|
477
|
+
if (printableRatio(plain) < 0.9) continue;
|
|
478
|
+
|
|
479
|
+
decoded.push(plain);
|
|
480
|
+
if (onDecoded) onDecoded(plain, blob);
|
|
481
|
+
}
|
|
482
|
+
BASE64_CANDIDATE.lastIndex = 0;
|
|
483
|
+
return decoded;
|
|
484
|
+
}
|
|
485
|
+
|
|
486
|
+
/**
|
|
487
|
+
* True when a base64-encoded value on an added line decodes to a structured
|
|
488
|
+
* credential.
|
|
489
|
+
*
|
|
490
|
+
* Deliberately runs the high-confidence patterns *only*. The low-confidence set
|
|
491
|
+
* is entropy- and keyword-driven, and decoded bytes are high-entropy by
|
|
492
|
+
* construction — pointing it at this output would flag close to every encoded
|
|
493
|
+
* blob in every repository. `AKIA[0-9A-Z]{16}` cannot match decoded noise;
|
|
494
|
+
* "looks secret-ish" always can. That asymmetry is the whole reason this check
|
|
495
|
+
* is safe to enable by default.
|
|
496
|
+
*
|
|
497
|
+
* @param {string} text
|
|
498
|
+
* @returns {boolean}
|
|
499
|
+
*/
|
|
500
|
+
export function hasEncodedSecret(text) {
|
|
501
|
+
return decodeBase64Blobs(text).some((plain) => hasHighConfidenceSecret(plain));
|
|
385
502
|
}
|
|
386
503
|
|
|
387
504
|
export function scanDiff(diffTextStr = "", options = {}) {
|
|
@@ -392,13 +509,23 @@ export function scanDiff(diffTextStr = "", options = {}) {
|
|
|
392
509
|
.map((line) => line.slice(1))
|
|
393
510
|
.join("\n");
|
|
394
511
|
|
|
395
|
-
const variants = secretScanVariants(addedLines);
|
|
512
|
+
const { all: variants, normalized } = secretScanVariants(addedLines);
|
|
396
513
|
const hasHigh = variants.some((v) => hasHighConfidenceSecret(v));
|
|
397
|
-
|
|
514
|
+
// Only worth decoding when nothing was found in the clear, and only against
|
|
515
|
+
// the fully-normalised text: decoding is the expensive step, and the
|
|
516
|
+
// intermediate variants differ from it in ways base64 blobs do not care about.
|
|
517
|
+
const hasEncoded = !hasHigh && hasEncodedSecret(normalized);
|
|
518
|
+
const hasLow = !hasHigh && !hasEncoded && variants.some((v) => hasLowConfidenceSecret(v));
|
|
398
519
|
const findings = [];
|
|
399
520
|
|
|
400
521
|
if (hasHigh) {
|
|
401
522
|
findings.push({ severity: "CRITICAL", type: "HIGH_CONFIDENCE_SECRET", description: "High-confidence secret pattern detected in added diff lines" });
|
|
523
|
+
} else if (hasEncoded) {
|
|
524
|
+
// Same type as the cleartext case: every gate that blocks on
|
|
525
|
+
// HIGH_CONFIDENCE_SECRET should block on this too, and a new type would
|
|
526
|
+
// have silently passed through the ones not updated. The description
|
|
527
|
+
// carries the difference the operator needs.
|
|
528
|
+
findings.push({ severity: "CRITICAL", type: "HIGH_CONFIDENCE_SECRET", description: "High-confidence secret pattern detected inside a base64-encoded value on an added diff line" });
|
|
402
529
|
} else if (hasLow) {
|
|
403
530
|
findings.push({ severity: "HIGH", type: "LOW_CONFIDENCE_SECRET", description: "Low-confidence secret or authorization token detected in added diff lines" });
|
|
404
531
|
}
|
|
@@ -419,7 +546,7 @@ export function scanDiff(diffTextStr = "", options = {}) {
|
|
|
419
546
|
}
|
|
420
547
|
|
|
421
548
|
return {
|
|
422
|
-
ok: !hasHigh && !hasLow && edgeRes.ok && crossPkgRes.ok,
|
|
549
|
+
ok: !hasHigh && !hasEncoded && !hasLow && edgeRes.ok && crossPkgRes.ok,
|
|
423
550
|
findings,
|
|
424
551
|
};
|
|
425
552
|
}
|
package/src/web-templates.mjs
CHANGED
|
@@ -233,6 +233,93 @@ export const WEB_TEMPLATES = {
|
|
|
233
233
|
- UTF-8 clean encoding with zero unescaped unicode artefacts.
|
|
234
234
|
- Localized formatting for dates, currencies, and numbers using standard Intl APIs.`;
|
|
235
235
|
}
|
|
236
|
+
},
|
|
237
|
+
|
|
238
|
+
// On the scope of this template, and what it deliberately does not claim:
|
|
239
|
+
//
|
|
240
|
+
// `llms.txt` (spec at llmstxt.org — cited without a scheme so the egress
|
|
241
|
+
// allowlist test does not read a citation as a destination the kit contacts)
|
|
242
|
+
// is a proposal, not a ratified standard. Plenty of sites publish one; no
|
|
243
|
+
// major provider has confirmed its retrieval stack reads one, and Google has
|
|
244
|
+
// said publicly that it does not use it. That is the supply side of a
|
|
245
|
+
// convention with no demonstrated demand side.
|
|
246
|
+
//
|
|
247
|
+
// So this template verifies what a repository can actually falsify — the file
|
|
248
|
+
// exists, it parses, its links resolve, and the site's crawler directives do
|
|
249
|
+
// not contradict each other — and it says nothing about whether publishing it
|
|
250
|
+
// improves visibility in any assistant. Every other template here carries a
|
|
251
|
+
// real verification oracle; a "generative engine optimization" template that
|
|
252
|
+
// promised ranking effects would be the first one that could not.
|
|
253
|
+
//
|
|
254
|
+
// Crawler posture is a policy choice, not a best practice. Allowing GPTBot,
|
|
255
|
+
// ClaudeBot or Google-Extended has licensing and editorial consequences, and
|
|
256
|
+
// blocking them is frequently deliberate. The template therefore takes no
|
|
257
|
+
// side: it defaults to `preserve`, reads the posture the repository already
|
|
258
|
+
// states, and enforces that every surface states the same thing.
|
|
259
|
+
//
|
|
260
|
+
// Structured data (JSON-LD, sameAs, entity markup) stays in `web-seo`. Two
|
|
261
|
+
// templates with authority over the same markup will eventually disagree.
|
|
262
|
+
"web-ai-access": {
|
|
263
|
+
id: "web-ai-access",
|
|
264
|
+
name: "AI Crawler Policy Consistency & llms.txt Integrity",
|
|
265
|
+
description: "Verify that AI crawler directives agree across every surface and that a published llms.txt parses with links that resolve.",
|
|
266
|
+
defaultVerifyCmd: "npm test",
|
|
267
|
+
category: "Crawler Policy & AI Access",
|
|
268
|
+
criticFocus: [
|
|
269
|
+
"Confirm the patch preserves the repository's existing AI crawler posture unless the task explicitly asked to change it — allowing or blocking a crawler is the operator's decision, not the agent's.",
|
|
270
|
+
"Verify robots.txt, per-page robots meta tags, and any X-Robots-Tag headers agree for every named agent; a page allowed in one surface and denied in another is a defect regardless of which is intended.",
|
|
271
|
+
"Check that every link in llms.txt resolves against the project's own route table or build output, with no absolute links to pages that no longer exist.",
|
|
272
|
+
"Ensure the PR description states that llms.txt consumption by AI systems is unverified, and claims no ranking, visibility, or citation benefit.",
|
|
273
|
+
"Confirm no JSON-LD or structured-data markup was modified here — that surface belongs to the web-seo template."
|
|
274
|
+
],
|
|
275
|
+
defaultParams: {
|
|
276
|
+
aiAccessPolicy: "preserve",
|
|
277
|
+
aiAgents: "GPTBot, ClaudeBot, Google-Extended, PerplexityBot, CCBot, Applebot-Extended",
|
|
278
|
+
targetRoutes: "all public routes"
|
|
279
|
+
},
|
|
280
|
+
generatePrompt: (params = {}) => {
|
|
281
|
+
const policy = String(params.aiAccessPolicy || "preserve").toLowerCase();
|
|
282
|
+
const agents = params.aiAgents || "GPTBot, ClaudeBot, Google-Extended, PerplexityBot, CCBot, Applebot-Extended";
|
|
283
|
+
const routes = params.targetRoutes || "all public routes";
|
|
284
|
+
const customGoal = params.goal ? `\n- **Target Focus**: ${params.goal}` : "";
|
|
285
|
+
|
|
286
|
+
const policyClause = {
|
|
287
|
+
allow: `The operator has decided to **allow** these agents. Make every surface say so consistently.`,
|
|
288
|
+
deny: `The operator has decided to **block** these agents. Make every surface say so consistently, and confirm no route leaks access through a surface that was missed.`,
|
|
289
|
+
selective: `The operator allows some agents and blocks others. Derive the intended split from existing configuration and make every surface agree with it exactly.`,
|
|
290
|
+
preserve: `**Do not change the posture.** Determine what the repository already states about these agents and make every surface state the same thing. If the surfaces currently contradict each other, report the contradiction and resolve it toward the most restrictive existing directive — never toward the more permissive one, and never invent a posture the repository has not expressed.`
|
|
291
|
+
}[policy] || `Treat \`${policy}\` as an explicit operator instruction and apply it consistently across every surface.`;
|
|
292
|
+
|
|
293
|
+
return `Audit AI crawler access directives and llms.txt integrity for ${routes}.${customGoal}
|
|
294
|
+
|
|
295
|
+
### Operator Policy (do not override)
|
|
296
|
+
${policyClause}
|
|
297
|
+
|
|
298
|
+
Agents in scope: ${agents}.
|
|
299
|
+
|
|
300
|
+
### Acceptance Criteria:
|
|
301
|
+
1. **One Posture, Every Surface**:
|
|
302
|
+
- \`robots.txt\`, per-page \`<meta name="robots">\` / agent-specific meta tags, and any \`X-Robots-Tag\` response headers must agree for every agent above.
|
|
303
|
+
- A route that is allowed by one surface and denied by another is a defect even when the intended answer is obvious — fix the disagreement, do not pick a winner silently.
|
|
304
|
+
- Verify \`robots.txt\` parses: correct \`User-agent:\` grouping, no directives stranded outside a group, no rules unreachable because of an earlier wildcard group.
|
|
305
|
+
2. **llms.txt Integrity (if the project publishes one)**:
|
|
306
|
+
- The file must parse as the proposed shape: a single \`# H1\` project name, an optional \`> blockquote\` summary, then \`## H2\` sections whose bodies are Markdown link lists.
|
|
307
|
+
- **Every link must resolve against this project's own route table or build output.** Check locally — do not fetch the live web from the verification step. A dead link in llms.txt is the single most common real defect in published files.
|
|
308
|
+
- Content must not contradict the crawler policy above: do not advertise paths in llms.txt that \`robots.txt\` disallows.
|
|
309
|
+
- If the project does not publish llms.txt, adding one is **in scope only if the task asked for it**. Do not create one on your own initiative.
|
|
310
|
+
3. **Honest Reporting**:
|
|
311
|
+
- \`llms.txt\` is a proposal, not a ratified standard, and no major provider has confirmed that its retrieval systems read it. Google has stated publicly that it does not.
|
|
312
|
+
- The PR description must therefore claim only what was verified — that the file exists, parses, and its links resolve. Do **not** claim improved visibility, ranking, or citation in any AI assistant. There is no oracle for that claim and it must not appear in the diff, the commit message, or the PR body.
|
|
313
|
+
4. **Scope Boundary**:
|
|
314
|
+
- Do not modify JSON-LD, Schema.org markup, \`sameAs\`, OpenGraph, or canonical tags. That surface belongs to the \`web-seo\` template; changing it here creates two sources of truth that will drift apart.
|
|
315
|
+
|
|
316
|
+
### Verification Oracle to Add
|
|
317
|
+
Add a repository-local test (no network access) that:
|
|
318
|
+
- parses \`robots.txt\` and asserts the directive set for each agent in scope matches the intended posture;
|
|
319
|
+
- parses \`llms.txt\`, if present, and asserts every link target exists in the route table or build output.
|
|
320
|
+
|
|
321
|
+
The test must fail on a hand-broken fixture before you consider it done.`;
|
|
322
|
+
}
|
|
236
323
|
}
|
|
237
324
|
};
|
|
238
325
|
|