@hifullmoon/aicommit 2.6.2 → 2.6.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,16 @@ This file lists notable user-facing changes. Internal refactors, test-only chang
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
+ ## [2.6.3] - 2026-09-15
8
+
9
+ ### Changed
10
+
11
+ - Very large `auto` split inventories now bundle adjacent local candidates by module, file kind, and Git status before hierarchical model planning, substantially reducing repeated input tokens while retaining every file in the final plan.
12
+
13
+ ### Fixed
14
+
15
+ - Split execution now batches large path inventories instead of overflowing Node's Git output/argument buffers, and large plan/checkpoint files remain readable for recovery.
16
+
7
17
  ## [2.6.2] - 2026-09-14
8
18
 
9
19
  ### Fixed
@@ -245,7 +255,8 @@ This file lists notable user-facing changes. Internal refactors, test-only chang
245
255
  - Added file-level split planning and execution with Git-state concurrency checks.
246
256
  - Added provider presets and user/project configuration boundaries.
247
257
 
248
- [Unreleased]: https://github.com/hi-fullmoon/AICommit/compare/v2.6.2...HEAD
258
+ [Unreleased]: https://github.com/hi-fullmoon/AICommit/compare/v2.6.3...HEAD
259
+ [2.6.3]: https://github.com/hi-fullmoon/AICommit/releases/tag/v2.6.3
249
260
  [2.6.2]: https://github.com/hi-fullmoon/AICommit/releases/tag/v2.6.2
250
261
  [2.6.1]: https://github.com/hi-fullmoon/AICommit/releases/tag/v2.6.1
251
262
  [2.6.0]: https://github.com/hi-fullmoon/AICommit/releases/tag/v2.6.0
package/README.md CHANGED
@@ -492,7 +492,7 @@ The default `largeChange.strategy: "auto"` inventories every file locally, group
492
492
 
493
493
  A normal commit typically needs one model request, with no per-file AI calls or recursive model reduction. The inventory contains at most 16 representative groups under a UTF-8 byte budget, prioritizing coverage across code, configuration, tests, and other categories. It explicitly describes sampling limits. Both terminal and JSON output distinguish fully analyzed files, representative excerpts, and metadata-only files. Provider retries, response recovery, policy correction, and user-requested regeneration can still add requests.
494
494
 
495
- Split mode builds local candidates and sends them in bounded batches of at most `splitMaxPlanFiles`, then merges the batch plans hierarchically. Every file remains represented even when the complete candidate inventory cannot fit one request. If `deep` analysis exhausts its aggregate budget or the hierarchy cannot converge, interactive and dry-run flows produce one conservative all-files plan with an explicit warning instead of using incomplete model output. Non-interactive committing stops unless `--allow-single-fallback` explicitly authorizes that degradation. Small changes keep the existing request path.
495
+ Split mode builds local candidates and sends them in bounded batches of at most `splitMaxPlanFiles`, then merges the batch plans hierarchically. When an `auto` inventory exceeds four times that candidate limit, adjacent candidates are first bundled locally by top-level module, file kind, and Git status; the model receives compact counts, examples, and selected excerpts while the complete file mapping stays local. Every file remains represented even when the complete candidate inventory cannot fit one request. If `deep` analysis exhausts its aggregate budget or the hierarchy cannot converge, interactive and dry-run flows produce one conservative all-files plan with an explicit warning instead of using incomplete model output. Non-interactive committing stops unless `--allow-single-fallback` explicitly authorizes that degradation. Small changes keep the existing request path.
496
496
 
497
497
  For exhaustive chunk-by-chunk model analysis, opt in through personal configuration:
498
498
 
package/README.zh-CN.md CHANGED
@@ -494,7 +494,7 @@ exec zsh
494
494
 
495
495
  普通提交通常只需要一次模型请求,不会为每个文件调用 AI,也不会递归调用模型汇总。摘要最多包含 16 个代表组,并受 UTF-8 字节预算约束;优先覆盖代码、配置和测试等不同类别。摘要明确说明抽样范围,界面和 JSON 分别报告全文分析、代表片段和仅元数据的文件数,不把抽样视为完整理解。提供方重试、响应恢复、格式修正和用户重新生成仍可能增加请求。
496
496
 
497
- 批次提交先在本地建立候选组,再按每批最多 `splitMaxPlanFiles` 个候选发送,并分层合并各批计划;完整候选清单放不进一次请求时,仍会保留每个文件。如果 `deep` 分析耗尽总预算,或分层规划无法收敛,交互和 dry-run 流程会明确警告并生成一个覆盖全部文件的保守计划,而不会采用不完整的模型结果;非交互提交默认停止,只有显式传入 `--allow-single-fallback` 才允许该降级。小变更保持原有请求路径。
497
+ 批次提交先在本地建立候选组,再按每批最多 `splitMaxPlanFiles` 个候选发送,并分层合并各批计划。当 `auto` 清单规模超过该候选上限的四倍时,会先按顶层模块、文件类型和 Git 状态在本地归并相邻候选;模型只接收紧凑计数、少量路径示例和代表片段,完整文件映射始终留在本地。完整候选清单放不进一次请求时,仍会保留每个文件。如果 `deep` 分析耗尽总预算,或分层规划无法收敛,交互和 dry-run 流程会明确警告并生成一个覆盖全部文件的保守计划,而不会采用不完整的模型结果;非交互提交默认停止,只有显式传入 `--allow-single-fallback` 才允许该降级。小变更保持原有请求路径。
498
498
 
499
499
  确实需要逐块 AI 分析时,在个人配置中设置:
500
500
 
@@ -33,7 +33,7 @@ The release workflow uses npm Trusted Publishing without a long-lived `NPM_TOKEN
33
33
  ```bash
34
34
  workdir=$(mktemp -d)
35
35
  cd "$workdir"
36
- npm install --package-lock-only @hifullmoon/aicommit@2.6.2
36
+ npm install --package-lock-only @hifullmoon/aicommit@2.6.3
37
37
  npm audit signatures
38
38
  ```
39
39
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@hifullmoon/aicommit",
3
- "version": "2.6.2",
3
+ "version": "2.6.3",
4
4
  "description": "Safe, local-first AI commit message generator for Git workflows",
5
5
  "type": "module",
6
6
  "bin": {
@@ -195,6 +195,16 @@
195
195
  "type": "number",
196
196
  "minimum": 0
197
197
  },
198
+ "planningCandidatesOriginal": {
199
+ "type": "number",
200
+ "minimum": 0,
201
+ "description": "Local candidates before token-saving planning compaction."
202
+ },
203
+ "planningCandidates": {
204
+ "type": "number",
205
+ "minimum": 0,
206
+ "description": "Candidate bundles sent through hierarchical split planning."
207
+ },
198
208
  "requests": {
199
209
  "type": "number",
200
210
  "minimum": 0
@@ -12,6 +12,7 @@ import {
12
12
  isDeepAnalysis,
13
13
  isGeneratedFile,
14
14
  localOverview,
15
+ compactLocalPlanFacts,
15
16
  localPlanBatches,
16
17
  } from './local-analysis.js';
17
18
 
@@ -714,8 +715,15 @@ export async function summarizeChanges(config, facts) {
714
715
  }
715
716
 
716
717
  export async function planAnalyzedChanges(config, facts, coverage = null, stream = null) {
717
- let candidates = facts;
718
718
  const deep = isDeepAnalysis(config) && coverage?.strategy !== 'auto';
719
+ const compacted = deep
720
+ ? { facts, originalCandidates: facts.length, planningCandidates: facts.length }
721
+ : compactLocalPlanFacts(config, facts);
722
+ let candidates = compacted.facts;
723
+ if (coverage) {
724
+ coverage.planningCandidatesOriginal = compacted.originalCandidates;
725
+ coverage.planningCandidates = compacted.planningCandidates;
726
+ }
719
727
  const policy = normalizeCommitPolicy(config.commitPolicy, config.language);
720
728
  const cap = Math.min(
721
729
  config.splitMaxDiffChars || 16000,
@@ -761,7 +769,7 @@ export async function planAnalyzedChanges(config, facts, coverage = null, stream
761
769
  };
762
770
  const groups = await jsonCall(
763
771
  config,
764
- `Each commit message must follow this policy: ${JSON.stringify(policy)}.\nGroup related changes into logical commits, including related implementation and tests across directories. Return [{"ids":[input IDs],"summary":"factual combined change summary","subject":"commit subject","body":"optional commit body"}]. Assign every input ID exactly once. Do not merge unrelated changes just to reduce group count.`,
772
+ `Each commit message must follow this policy: ${JSON.stringify(policy)}.\nGroup related changes into logical commits, including related implementation and tests across directories. Return [{"ids":[input IDs],"summary":"factual combined change summary","subject":"commit subject","body":"optional commit body"}]. Assign every input ID exactly once. Do not merge unrelated changes just to reduce group count.${!deep && level === 0 ? ' Inputs are a compact local inventory, not full semantic summaries. One ID may represent multiple files and must remain atomic; use kind, module, status, fileCount, examples, and any representativeExcerpt conservatively.' : ''}`,
765
773
  batch,
766
774
  (candidate) => {
767
775
  validatePartition(candidate, ids);
package/src/git-spool.js CHANGED
@@ -11,7 +11,7 @@ process.once('exit', cleanupGitSpools);
11
11
 
12
12
  // Git writes directly to a private temporary file, never to a Node stdout
13
13
  // buffer. Consumers scan the captured bytes without reopening the worktree.
14
- export function spoolGit(commands, cwd) {
14
+ export function spoolGit(commands, cwd, { env = undefined } = {}) {
15
15
  const dir = mkdtempSync(join(tmpdir(), 'aicommit-diff-'));
16
16
  const path = join(dir, 'patch');
17
17
  let fd;
@@ -20,6 +20,7 @@ export function spoolGit(commands, cwd) {
20
20
  for (const args of commands) {
21
21
  execFileSync('git', ['--no-pager', ...args], {
22
22
  cwd,
23
+ ...(env ? { env } : {}),
23
24
  stdio: ['ignore', fd, 'pipe'],
24
25
  timeout: 120000,
25
26
  maxBuffer: 1024 * 1024,
@@ -79,6 +80,23 @@ export function spoolGit(commands, cwd) {
79
80
  }
80
81
  if (pending.length) yield pending.toString('utf8');
81
82
  },
83
+ *nulFields(maxFieldBytes = 16 * 1024) {
84
+ let pending = Buffer.alloc(0);
85
+ for (const buffer of this.buffers()) {
86
+ const data = pending.length ? Buffer.concat([pending, buffer]) : buffer;
87
+ let start = 0;
88
+ for (let end = data.indexOf(0); end !== -1; end = data.indexOf(0, start)) {
89
+ if (end - start > maxFieldBytes)
90
+ throw new Error(`NUL-delimited Git field exceeds ${maxFieldBytes} bytes.`);
91
+ yield data.subarray(start, end).toString('utf8');
92
+ start = end + 1;
93
+ }
94
+ pending = Buffer.from(data.subarray(start));
95
+ if (pending.length > maxFieldBytes)
96
+ throw new Error(`NUL-delimited Git field exceeds ${maxFieldBytes} bytes.`);
97
+ }
98
+ if (pending.length) throw new Error('NUL-delimited Git output was truncated.');
99
+ },
82
100
  text(limit) {
83
101
  if (size > limit) return null;
84
102
  return [...this.lines()].join('').trim();
@@ -251,6 +251,86 @@ export function analyzeLocally(config, capture, protect, previews) {
251
251
  };
252
252
  }
253
253
 
254
+ function planningBucketKey(fact) {
255
+ const topLevel = fact.module === '.' ? '.' : fact.module.split('/')[0];
256
+ return JSON.stringify([
257
+ topLevel,
258
+ fact.kind,
259
+ fact.status,
260
+ fact.modeChange || '',
261
+ Boolean(fact.binary),
262
+ ]);
263
+ }
264
+
265
+ function moduleSpan(facts) {
266
+ const modules = [...new Set(facts.map((fact) => fact.module))];
267
+ if (modules.length === 1) return modules[0];
268
+ const parts = modules.map((module) => module.split('/'));
269
+ const common = [];
270
+ for (let index = 0; index < parts[0].length; index++) {
271
+ const value = parts[0][index];
272
+ if (!parts.every((item) => item[index] === value)) break;
273
+ common.push(value);
274
+ }
275
+ return common.length ? `${common.join('/')}/*` : 'multiple modules';
276
+ }
277
+
278
+ // Extremely large auto-mode inventories cannot represent every independent
279
+ // file candidate to the model without hundreds of repeated requests. Bundle
280
+ // adjacent candidates locally while retaining their complete file mapping;
281
+ // the model groups bundle IDs, and execution still covers every original path.
282
+ export function compactLocalPlanFacts(config, facts) {
283
+ const maxItems = config.splitMaxPlanFiles || 100;
284
+ const threshold = maxItems * 4;
285
+ if (facts.length <= threshold) {
286
+ return { facts, originalCandidates: facts.length, planningCandidates: facts.length };
287
+ }
288
+
289
+ const target = Math.max(1, maxItems * 2);
290
+ const bundleSize = Math.ceil(facts.length / target);
291
+ const ordered = [...facts].sort(
292
+ (left, right) =>
293
+ planningBucketKey(left).localeCompare(planningBucketKey(right)) ||
294
+ left.module.localeCompare(right.module) ||
295
+ left.id.localeCompare(right.id),
296
+ );
297
+ const bundles = [];
298
+ let bucket = [];
299
+ let key = null;
300
+ const flush = () => {
301
+ if (!bucket.length) return;
302
+ const first = bucket[0];
303
+ const evidence = bucket.find((fact) => fact.evidence)?.evidence || '';
304
+ const modeChanges = new Set(bucket.map((fact) => fact.modeChange).filter(Boolean));
305
+ bundles.push({
306
+ id: `C${bundles.length + 1}`,
307
+ kind: first.kind,
308
+ module: moduleSpan(bucket),
309
+ status: first.status,
310
+ files: bucket.flatMap((fact) => fact.files),
311
+ additions: bucket.reduce((total, fact) => total + fact.additions, 0),
312
+ deletions: bucket.reduce((total, fact) => total + fact.deletions, 0),
313
+ modeChange: modeChanges.size === 1 ? [...modeChanges][0] : null,
314
+ binary: bucket.some((fact) => fact.binary),
315
+ evidence,
316
+ });
317
+ bucket = [];
318
+ };
319
+
320
+ for (const fact of ordered) {
321
+ const nextKey = planningBucketKey(fact);
322
+ if (bucket.length && (nextKey !== key || bucket.length >= bundleSize)) flush();
323
+ key = nextKey;
324
+ bucket.push(fact);
325
+ }
326
+ flush();
327
+ return {
328
+ facts: bundles,
329
+ originalCandidates: facts.length,
330
+ planningCandidates: bundles.length,
331
+ };
332
+ }
333
+
254
334
  export function localPlanBatches(config, facts, cap) {
255
335
  const error = () =>
256
336
  fail(ERROR_CATEGORIES.CONFIG, 'A split candidate is too large for one planning request.', {
@@ -260,10 +340,7 @@ export function localPlanBatches(config, facts, cap) {
260
340
  const batches = [];
261
341
  let batch = [];
262
342
  for (const fact of facts) {
263
- const item = {
264
- ...summaryOf(fact, false),
265
- summary: 'Local candidate only; verify grouping. Content not fully analyzed.',
266
- };
343
+ const item = summaryOf(fact, false);
267
344
  if (Buffer.byteLength(JSON.stringify([item])) > cap) throw error();
268
345
  const candidate = [...batch, item];
269
346
  if (
@@ -278,12 +355,15 @@ export function localPlanBatches(config, facts, cap) {
278
355
  if (batch.length) batches.push(batch);
279
356
 
280
357
  const factsById = new Map(facts.map((fact) => [fact.id, fact]));
358
+ let evidenceCount = 0;
281
359
  for (const items of batches) {
282
360
  for (const item of items) {
361
+ if (evidenceCount >= 16) break;
283
362
  const evidence = factsById.get(item.id)?.evidence;
284
363
  if (!evidence) continue;
285
364
  item.representativeExcerpt = evidence;
286
365
  if (Buffer.byteLength(JSON.stringify(items)) > cap) delete item.representativeExcerpt;
366
+ else evidenceCount++;
287
367
  }
288
368
  }
289
369
  return batches;
@@ -15,7 +15,10 @@ import { validateSplitPlanArtifact } from './split-plan.js';
15
15
 
16
16
  export const SPLIT_CHECKPOINT_KIND = 'aicommit-split-checkpoint';
17
17
  export const SPLIT_CHECKPOINT_VERSION = 1;
18
- const MAX_CHECKPOINT_BYTES = 2 * 1024 * 1024;
18
+ // Checkpoints contain both the validated plan and one object snapshot per real
19
+ // path, so their bound must be larger than the plan artifact bound. This still
20
+ // caps parsing of local metadata while allowing the supported 10,000 changes.
21
+ const MAX_CHECKPOINT_BYTES = 128 * 1024 * 1024;
19
22
  const OID_RE = /^(?:[0-9a-f]{40}|[0-9a-f]{64})$/;
20
23
 
21
24
  function object(value) {
@@ -197,7 +200,7 @@ export function readSplitCheckpoint(projectRoot) {
197
200
  if (!stat.isFile() || stat.isSymbolicLink()) {
198
201
  throw new Error('Split checkpoint must be a regular, non-symbolic-link file.');
199
202
  }
200
- if (stat.size > MAX_CHECKPOINT_BYTES) throw new Error('Split checkpoint exceeds 2 MiB.');
203
+ if (stat.size > MAX_CHECKPOINT_BYTES) throw new Error('Split checkpoint exceeds 128 MiB.');
201
204
  let parsed;
202
205
  try {
203
206
  parsed = JSON.parse(readFileSync(path, 'utf8'));
@@ -213,10 +216,14 @@ export function writeSplitCheckpoint(projectRoot, input) {
213
216
  ...input,
214
217
  updatedAt: new Date().toISOString(),
215
218
  });
219
+ const serialized = JSON.stringify(checkpoint, null, 2) + '\n';
220
+ if (Buffer.byteLength(serialized) > MAX_CHECKPOINT_BYTES) {
221
+ throw new Error('Split checkpoint exceeds the 128 MiB limit.');
222
+ }
216
223
  mkdirSync(dirname(path), { recursive: true, mode: 0o700 });
217
224
  const temporary = `${path}.${process.pid}.${Date.now()}.tmp`;
218
225
  try {
219
- writeFileSync(temporary, JSON.stringify(checkpoint, null, 2) + '\n', {
226
+ writeFileSync(temporary, serialized, {
220
227
  encoding: 'utf8',
221
228
  mode: 0o600,
222
229
  });
@@ -4,6 +4,8 @@ import { mkdtempSync, rmSync } from 'node:fs';
4
4
  import { tmpdir } from 'node:os';
5
5
  import { join } from 'node:path';
6
6
 
7
+ import { pathBatches } from './git.js';
8
+
7
9
  const HUNK_HEADER_RE = /^@@ -(\d+)(?:,(\d+))? \+(\d+)(?:,(\d+))? @@/;
8
10
 
9
11
  function git(projectRoot, args, { indexPath = null, input = undefined } = {}) {
@@ -215,9 +217,13 @@ export function validateHunkTransaction(projectRoot, plan, snapshots) {
215
217
  }
216
218
 
217
219
  const finalEntries = parseIndexEntries(
218
- git(projectRoot, ['ls-files', '--stage', '-z', '--', ...snapshots.map((item) => item.path)], {
219
- indexPath,
220
- }),
220
+ pathBatches(snapshots.map((item) => item.path))
221
+ .map((batch) =>
222
+ git(projectRoot, ['--literal-pathspecs', 'ls-files', '--stage', '-z', '--', ...batch], {
223
+ indexPath,
224
+ }),
225
+ )
226
+ .join(''),
221
227
  );
222
228
  for (const entry of snapshots) {
223
229
  if (!entriesEqual(finalEntries.get(entry.path) || null, entry.target)) {
package/src/split-plan.js CHANGED
@@ -5,7 +5,10 @@ import { normalizeCommitPolicy, validateCommitCandidate } from './policy.js';
5
5
 
6
6
  export const SPLIT_PLAN_KIND = 'aicommit-split-plan';
7
7
  export const SPLIT_PLAN_VERSION = 1;
8
- const MAX_PLAN_BYTES = 1024 * 1024;
8
+ // A file-level split can legitimately contain thousands of long paths. Keep
9
+ // reads bounded, but size the artifact limit for the validated 10,000-change
10
+ // ceiling instead of rejecting plans that this module can create itself.
11
+ const MAX_PLAN_BYTES = 64 * 1024 * 1024;
9
12
  const HASH_RE = /^(?:[0-9a-f]{40}|[0-9a-f]{64})$/;
10
13
  const FINGERPRINT_RE = /^[0-9a-f]{64}$/;
11
14
  const HUNK_HASH_RE = /^[0-9a-f]{64}$/;
@@ -241,13 +244,14 @@ export function validateSplitPlanArtifact(input) {
241
244
  normalizeGroup(group, index, allowedPaths, hunkCatalog, commitPolicy, input.language),
242
245
  );
243
246
  const assignedFiles = groups.flatMap((group) => group.files);
244
- if (new Set(assignedFiles).size !== assignedFiles.length) {
247
+ const assignedFileSet = new Set(assignedFiles);
248
+ if (assignedFileSet.size !== assignedFiles.length) {
245
249
  throw new Error('Split plan assigns a path to more than one group.');
246
250
  }
247
251
  const assignedHunks = new Set();
248
252
  for (const group of groups) {
249
253
  for (const assignment of group.hunks || []) {
250
- if (assignedFiles.includes(assignment.path)) {
254
+ if (assignedFileSet.has(assignment.path)) {
251
255
  throw new Error(`Split plan assigns both a whole file and hunks: ${assignment.path}`);
252
256
  }
253
257
  for (const id of assignment.ids) {
@@ -260,7 +264,7 @@ export function validateSplitPlanArtifact(input) {
260
264
  }
261
265
  }
262
266
  const missing = displayPaths.filter((path) => {
263
- if (assignedFiles.includes(path)) return false;
267
+ if (assignedFileSet.has(path)) return false;
264
268
  const ids = hunkCatalog.get(path);
265
269
  return !ids || [...ids].some((id) => !assignedHunks.has(`${path}\0${id}`));
266
270
  });
@@ -310,10 +314,14 @@ export function createSplitPlanArtifact({
310
314
  export async function writeSplitPlanArtifact(path, artifact) {
311
315
  const absolute = resolve(path);
312
316
  const validated = validateSplitPlanArtifact(artifact);
317
+ const serialized = JSON.stringify(validated, null, 2) + '\n';
318
+ if (Buffer.byteLength(serialized) > MAX_PLAN_BYTES) {
319
+ throw new Error('Split plan exceeds the 64 MiB limit.');
320
+ }
313
321
  await mkdir(dirname(absolute), { recursive: true, mode: 0o700 });
314
322
  const temporary = `${absolute}.${process.pid}.${Date.now()}.tmp`;
315
323
  try {
316
- await writeFile(temporary, JSON.stringify(validated, null, 2) + '\n', {
324
+ await writeFile(temporary, serialized, {
317
325
  encoding: 'utf8',
318
326
  mode: 0o600,
319
327
  flag: 'wx',
@@ -341,7 +349,7 @@ export async function readSplitPlanArtifact(path) {
341
349
  if (!stat.isFile() || stat.isSymbolicLink()) {
342
350
  throw new Error('Split plan must be a regular, non-symbolic-link file.');
343
351
  }
344
- if (stat.size > MAX_PLAN_BYTES) throw new Error('Split plan exceeds the 1 MiB limit.');
352
+ if (stat.size > MAX_PLAN_BYTES) throw new Error('Split plan exceeds the 64 MiB limit.');
345
353
  let parsed;
346
354
  try {
347
355
  parsed = JSON.parse(await readFile(absolute, 'utf8'));
package/src/split.js CHANGED
@@ -236,20 +236,33 @@ export function getSplitChangedFiles(cwd, scope = 'all') {
236
236
  // the planner must see the same latest content executeSplit will git add -A.
237
237
  function getWorkingTreeDiff(projectRoot, head, contextLines, paths = []) {
238
238
  const u = unifiedArg(contextLines);
239
- const pathArgs = paths.length ? ['--', ...paths] : [];
240
- if (head) {
241
- return readGit(['diff', u, 'HEAD', ...pathArgs], projectRoot).trim();
239
+ const batches = paths.length ? pathBatches(paths) : [[]];
240
+ const parts = [];
241
+ for (const batch of batches) {
242
+ const pathArgs = batch.length ? ['--', ...batch] : [];
243
+ if (head) {
244
+ parts.push(readGit(['diff', u, 'HEAD', ...pathArgs], projectRoot).trim());
245
+ continue;
246
+ }
247
+ parts.push(readGit(['diff', u, '--cached', ...pathArgs], projectRoot).trim());
248
+ parts.push(readGit(['diff', u, ...pathArgs], projectRoot).trim());
242
249
  }
243
-
244
- const cached = readGit(['diff', u, '--cached', ...pathArgs], projectRoot).trim();
245
- const unstaged = readGit(['diff', u, ...pathArgs], projectRoot).trim();
246
- return [cached, unstaged].filter(Boolean).join('\n');
250
+ return parts.filter(Boolean).join('\n');
247
251
  }
248
252
 
249
253
  export function getSplitDiff(projectRoot, head, contextLines, scope = 'all', paths = []) {
250
254
  if (scope === 'staged') {
251
- const pathArgs = paths.length ? ['--', ...paths] : [];
252
- return readGit(['diff', unifiedArg(contextLines), '--cached', ...pathArgs], projectRoot).trim();
255
+ const batches = paths.length ? pathBatches(paths) : [[]];
256
+ return batches
257
+ .map((batch) => {
258
+ const pathArgs = batch.length ? ['--', ...batch] : [];
259
+ return readGit(
260
+ ['diff', unifiedArg(contextLines), '--cached', ...pathArgs],
261
+ projectRoot,
262
+ ).trim();
263
+ })
264
+ .filter(Boolean)
265
+ .join('\n');
253
266
  }
254
267
  return getWorkingTreeDiff(projectRoot, head, contextLines, paths);
255
268
  }
@@ -261,7 +274,7 @@ export function getSplitStateFingerprint(projectRoot, head, files, scope = 'all'
261
274
  const hash = createHash('sha256');
262
275
  if (scope === 'staged') {
263
276
  hash.update(head ? readGit(['rev-parse', 'HEAD'], projectRoot).trim() : '<unborn>');
264
- hash.update(readGit(['ls-files', '--stage', '-z'], projectRoot));
277
+ updateGitHash(hash, ['ls-files', '--stage', '-z'], projectRoot);
265
278
  return hash.digest('hex');
266
279
  }
267
280
  files ||= getAllChangedFiles(projectRoot);
@@ -909,9 +922,9 @@ function getGroupDiff(
909
922
  return 'Changed files (new files, no diff available):\n' + parts.join('\n');
910
923
  }
911
924
 
912
- function parseStageZeroEntries(text) {
925
+ function parseStageZeroEntries(fields) {
913
926
  const entries = new Map();
914
- for (const field of text.split('\0')) {
927
+ for (const field of fields) {
915
928
  if (!field) continue;
916
929
  const match = field.match(/^(\d+) ([0-9a-f]+) (\d)\t([\s\S]+)$/);
917
930
  if (!match || match[3] !== '0') continue;
@@ -920,8 +933,37 @@ function parseStageZeroEntries(text) {
920
933
  return entries;
921
934
  }
922
935
 
923
- function readStageZeroEntries(projectRoot) {
924
- return parseStageZeroEntries(readGit(['ls-files', '--stage', '-z'], projectRoot));
936
+ function* filterGitFields(fields, wanted) {
937
+ for (const field of fields) {
938
+ const tab = field.indexOf('\t');
939
+ if (tab !== -1 && wanted.has(field.slice(tab + 1))) yield field;
940
+ }
941
+ }
942
+
943
+ function readStageZeroEntries(projectRoot, paths = null, indexPath = null) {
944
+ if (paths !== null && !paths.length) return new Map();
945
+ const wanted = paths === null ? null : new Set(paths);
946
+ const batches = paths === null ? null : pathBatches(paths);
947
+ if (batches && batches.length <= 8) {
948
+ const output = batches
949
+ .map((batch) => {
950
+ const args = ['--literal-pathspecs', 'ls-files', '--stage', '-z', '--', ...batch];
951
+ return indexPath
952
+ ? runGitWithIndex(args, projectRoot, indexPath)
953
+ : readGit(args, projectRoot);
954
+ })
955
+ .join('');
956
+ return parseStageZeroEntries(output.split('\0'));
957
+ }
958
+ const source = spoolGit([['--literal-pathspecs', 'ls-files', '--stage', '-z']], projectRoot, {
959
+ ...(indexPath ? { env: { ...process.env, GIT_INDEX_FILE: indexPath } } : {}),
960
+ });
961
+ try {
962
+ const fields = wanted ? filterGitFields(source.nulFields(), wanted) : source.nulFields();
963
+ return parseStageZeroEntries(fields);
964
+ } finally {
965
+ source.dispose();
966
+ }
925
967
  }
926
968
 
927
969
  function runGitWithIndex(args, projectRoot, indexPath, inherit = false, input = undefined) {
@@ -977,25 +1019,11 @@ function activeCommitHooks(projectRoot) {
977
1019
 
978
1020
  function submodulePaths(projectRoot, paths) {
979
1021
  const found = new Set();
980
- const index = pathBatches(paths)
981
- .map((batch) =>
982
- readGit(['--literal-pathspecs', 'ls-files', '--stage', '-z', '--', ...batch], projectRoot),
983
- )
984
- .join('');
985
- for (const entry of index.split('\0')) {
986
- const match = entry.match(/^160000 [0-9a-f]+ \d\t([\s\S]+)$/);
987
- if (match) found.add(match[1]);
988
- }
1022
+ for (const [path, entry] of readStageZeroEntries(projectRoot, paths))
1023
+ if (entry.mode === '160000') found.add(path);
989
1024
  if (hasHead(projectRoot)) {
990
- const tree = pathBatches(paths)
991
- .map((batch) =>
992
- readGit(['--literal-pathspecs', 'ls-tree', '-rz', 'HEAD', '--', ...batch], projectRoot),
993
- )
994
- .join('');
995
- for (const entry of tree.split('\0')) {
996
- const match = entry.match(/^160000 commit [0-9a-f]+\t([\s\S]+)$/);
997
- if (match) found.add(match[1]);
998
- }
1025
+ for (const [path, entry] of readHeadEntries(projectRoot, paths))
1026
+ if (entry.mode === '160000') found.add(path);
999
1027
  }
1000
1028
  return [...found];
1001
1029
  }
@@ -1112,7 +1140,7 @@ function resetCommittedPaths(projectRoot, groups, allFiles) {
1112
1140
  }
1113
1141
 
1114
1142
  function captureTargetEntries(projectRoot, scope, paths) {
1115
- if (scope === 'staged') return readStageZeroEntries(projectRoot);
1143
+ if (scope === 'staged') return readStageZeroEntries(projectRoot, paths);
1116
1144
  const tempDir = mkdtempSync(join(tmpdir(), 'aicommit-split-snapshot-'));
1117
1145
  const indexPath = join(tempDir, 'index');
1118
1146
  try {
@@ -1128,9 +1156,7 @@ function captureTargetEntries(projectRoot, scope, paths) {
1128
1156
  false,
1129
1157
  paths.join('\0') + '\0',
1130
1158
  );
1131
- return parseStageZeroEntries(
1132
- runGitWithIndex(['ls-files', '--stage', '-z'], projectRoot, indexPath),
1133
- );
1159
+ return readStageZeroEntries(projectRoot, paths, indexPath);
1134
1160
  } finally {
1135
1161
  rmSync(tempDir, { recursive: true, force: true });
1136
1162
  }
@@ -1139,7 +1165,7 @@ function captureTargetEntries(projectRoot, scope, paths) {
1139
1165
  function captureCheckpointSnapshots(projectRoot, scope, allFiles) {
1140
1166
  const paths = [...new Set(allFiles.flatMap((change) => change.addPaths))].sort();
1141
1167
  const target = captureTargetEntries(projectRoot, scope, paths);
1142
- const index = readStageZeroEntries(projectRoot);
1168
+ const index = readStageZeroEntries(projectRoot, paths);
1143
1169
  return paths.map((path) => ({
1144
1170
  path,
1145
1171
  target: target.get(path) || null,
@@ -1243,15 +1269,35 @@ function splitCommitFailure(projectRoot, message) {
1243
1269
  function readHeadEntries(projectRoot, paths) {
1244
1270
  const entries = new Map();
1245
1271
  if (!hasHead(projectRoot) || !paths.length) return entries;
1246
- const text = pathBatches(paths)
1247
- .map((batch) =>
1248
- readGit(['--literal-pathspecs', 'ls-tree', '-rz', 'HEAD', '--', ...batch], projectRoot),
1249
- )
1250
- .join('');
1251
- for (const field of text.split('\0')) {
1252
- if (!field) continue;
1253
- const match = field.match(/^(\d+) \S+ ([0-9a-f]+)\t([\s\S]+)$/);
1254
- if (match) entries.set(match[3], { mode: match[1], oid: match[2] });
1272
+ const wanted = new Set(paths);
1273
+ const batches = pathBatches(paths);
1274
+ let source = null;
1275
+ try {
1276
+ const fields =
1277
+ batches.length <= 8
1278
+ ? batches
1279
+ .map((batch) =>
1280
+ readGit(
1281
+ ['--literal-pathspecs', 'ls-tree', '-rz', 'HEAD', '--', ...batch],
1282
+ projectRoot,
1283
+ ),
1284
+ )
1285
+ .join('')
1286
+ .split('\0')
1287
+ : filterGitFields(
1288
+ (source = spoolGit(
1289
+ [['--literal-pathspecs', 'ls-tree', '-rz', 'HEAD']],
1290
+ projectRoot,
1291
+ )).nulFields(),
1292
+ wanted,
1293
+ );
1294
+ for (const field of fields) {
1295
+ if (!field) continue;
1296
+ const match = field.match(/^(\d+) \S+ ([0-9a-f]+)\t([\s\S]+)$/);
1297
+ if (match) entries.set(match[3], { mode: match[1], oid: match[2] });
1298
+ }
1299
+ } finally {
1300
+ source?.dispose();
1255
1301
  }
1256
1302
  return entries;
1257
1303
  }
@@ -1289,7 +1335,7 @@ function reconcileCompletedIndex(projectRoot, checkpoint) {
1289
1335
  const paths = [
1290
1336
  ...new Set(groups.flatMap((group) => expandGroupPaths(group, checkpoint.plan.changes))),
1291
1337
  ];
1292
- const current = readStageZeroEntries(projectRoot);
1338
+ const current = readStageZeroEntries(projectRoot, paths);
1293
1339
  const head = readHeadEntries(projectRoot, paths);
1294
1340
  const snapshots = new Map(checkpoint.snapshots.map((snapshot) => [snapshot.path, snapshot]));
1295
1341
  for (const path of paths) {
@@ -1826,6 +1872,13 @@ export async function splitFlow(
1826
1872
  analysis.coverage,
1827
1873
  planningStream,
1828
1874
  );
1875
+ if (
1876
+ analysis.coverage.planningCandidates < analysis.coverage.planningCandidatesOriginal
1877
+ ) {
1878
+ console.error(
1879
+ ` Planning input: bundled ${analysis.coverage.planningCandidatesOriginal} local candidates into ${analysis.coverage.planningCandidates} token-efficient groups.`,
1880
+ );
1881
+ }
1829
1882
  } catch (err) {
1830
1883
  const exhausted =
1831
1884
  err.data?.analysis?.exhausted ||