@hifullmoon/aicommit 2.6.2 → 2.6.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +19 -1
- package/README.md +1 -1
- package/README.zh-CN.md +1 -1
- package/docs/distribution.md +1 -1
- package/package.json +1 -1
- package/schemas/aicommit-output.schema.json +10 -0
- package/src/change-analysis.js +40 -9
- package/src/git-spool.js +19 -1
- package/src/local-analysis.js +84 -4
- package/src/split-checkpoint.js +10 -3
- package/src/split-hunks.js +9 -3
- package/src/split-plan.js +14 -6
- package/src/split.js +100 -47
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,22 @@ This file lists notable user-facing changes. Internal refactors, test-only chang
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
+
## [2.6.4] - 2026-09-19
|
|
8
|
+
|
|
9
|
+
### Fixed
|
|
10
|
+
|
|
11
|
+
- Large split planning now sizes batches for the model response limit and preserves valid file groupings when the provider omits a group summary.
|
|
12
|
+
|
|
13
|
+
## [2.6.3] - 2026-09-15
|
|
14
|
+
|
|
15
|
+
### Changed
|
|
16
|
+
|
|
17
|
+
- Very large `auto` split inventories now bundle adjacent local candidates by module, file kind, and Git status before hierarchical model planning, substantially reducing repeated input tokens while retaining every file in the final plan.
|
|
18
|
+
|
|
19
|
+
### Fixed
|
|
20
|
+
|
|
21
|
+
- Split execution now batches large path inventories instead of overflowing Node's Git output/argument buffers, and large plan/checkpoint files remain readable for recovery.
|
|
22
|
+
|
|
7
23
|
## [2.6.2] - 2026-09-14
|
|
8
24
|
|
|
9
25
|
### Fixed
|
|
@@ -245,7 +261,9 @@ This file lists notable user-facing changes. Internal refactors, test-only chang
|
|
|
245
261
|
- Added file-level split planning and execution with Git-state concurrency checks.
|
|
246
262
|
- Added provider presets and user/project configuration boundaries.
|
|
247
263
|
|
|
248
|
-
[Unreleased]: https://github.com/hi-fullmoon/AICommit/compare/v2.6.
|
|
264
|
+
[Unreleased]: https://github.com/hi-fullmoon/AICommit/compare/v2.6.4...HEAD
|
|
265
|
+
[2.6.4]: https://github.com/hi-fullmoon/AICommit/releases/tag/v2.6.4
|
|
266
|
+
[2.6.3]: https://github.com/hi-fullmoon/AICommit/releases/tag/v2.6.3
|
|
249
267
|
[2.6.2]: https://github.com/hi-fullmoon/AICommit/releases/tag/v2.6.2
|
|
250
268
|
[2.6.1]: https://github.com/hi-fullmoon/AICommit/releases/tag/v2.6.1
|
|
251
269
|
[2.6.0]: https://github.com/hi-fullmoon/AICommit/releases/tag/v2.6.0
|
package/README.md
CHANGED
|
@@ -492,7 +492,7 @@ The default `largeChange.strategy: "auto"` inventories every file locally, group
|
|
|
492
492
|
|
|
493
493
|
A normal commit typically needs one model request, with no per-file AI calls or recursive model reduction. The inventory contains at most 16 representative groups under a UTF-8 byte budget, prioritizing coverage across code, configuration, tests, and other categories. It explicitly describes sampling limits. Both terminal and JSON output distinguish fully analyzed files, representative excerpts, and metadata-only files. Provider retries, response recovery, policy correction, and user-requested regeneration can still add requests.
|
|
494
494
|
|
|
495
|
-
Split mode builds local candidates and sends them in bounded batches of at most `splitMaxPlanFiles`, then merges the batch plans hierarchically. Every file remains represented even when the complete candidate inventory cannot fit one request. If `deep` analysis exhausts its aggregate budget or the hierarchy cannot converge, interactive and dry-run flows produce one conservative all-files plan with an explicit warning instead of using incomplete model output. Non-interactive committing stops unless `--allow-single-fallback` explicitly authorizes that degradation. Small changes keep the existing request path.
|
|
495
|
+
Split mode builds local candidates and sends them in bounded batches of at most `splitMaxPlanFiles`, then merges the batch plans hierarchically. When an `auto` inventory exceeds four times that candidate limit, adjacent candidates are first bundled locally by top-level module, file kind, and Git status; the model receives compact counts, examples, and selected excerpts while the complete file mapping stays local. Every file remains represented even when the complete candidate inventory cannot fit one request. If `deep` analysis exhausts its aggregate budget or the hierarchy cannot converge, interactive and dry-run flows produce one conservative all-files plan with an explicit warning instead of using incomplete model output. Non-interactive committing stops unless `--allow-single-fallback` explicitly authorizes that degradation. Small changes keep the existing request path.
|
|
496
496
|
|
|
497
497
|
For exhaustive chunk-by-chunk model analysis, opt in through personal configuration:
|
|
498
498
|
|
package/README.zh-CN.md
CHANGED
|
@@ -494,7 +494,7 @@ exec zsh
|
|
|
494
494
|
|
|
495
495
|
普通提交通常只需要一次模型请求,不会为每个文件调用 AI,也不会递归调用模型汇总。摘要最多包含 16 个代表组,并受 UTF-8 字节预算约束;优先覆盖代码、配置和测试等不同类别。摘要明确说明抽样范围,界面和 JSON 分别报告全文分析、代表片段和仅元数据的文件数,不把抽样视为完整理解。提供方重试、响应恢复、格式修正和用户重新生成仍可能增加请求。
|
|
496
496
|
|
|
497
|
-
批次提交先在本地建立候选组,再按每批最多 `splitMaxPlanFiles`
|
|
497
|
+
批次提交先在本地建立候选组,再按每批最多 `splitMaxPlanFiles` 个候选发送,并分层合并各批计划。当 `auto` 清单规模超过该候选上限的四倍时,会先按顶层模块、文件类型和 Git 状态在本地归并相邻候选;模型只接收紧凑计数、少量路径示例和代表片段,完整文件映射始终留在本地。完整候选清单放不进一次请求时,仍会保留每个文件。如果 `deep` 分析耗尽总预算,或分层规划无法收敛,交互和 dry-run 流程会明确警告并生成一个覆盖全部文件的保守计划,而不会采用不完整的模型结果;非交互提交默认停止,只有显式传入 `--allow-single-fallback` 才允许该降级。小变更保持原有请求路径。
|
|
498
498
|
|
|
499
499
|
确实需要逐块 AI 分析时,在个人配置中设置:
|
|
500
500
|
|
package/docs/distribution.md
CHANGED
|
@@ -33,7 +33,7 @@ The release workflow uses npm Trusted Publishing without a long-lived `NPM_TOKEN
|
|
|
33
33
|
```bash
|
|
34
34
|
workdir=$(mktemp -d)
|
|
35
35
|
cd "$workdir"
|
|
36
|
-
npm install --package-lock-only @hifullmoon/aicommit@2.6.
|
|
36
|
+
npm install --package-lock-only @hifullmoon/aicommit@2.6.4
|
|
37
37
|
npm audit signatures
|
|
38
38
|
```
|
|
39
39
|
|
package/package.json
CHANGED
|
@@ -195,6 +195,16 @@
|
|
|
195
195
|
"type": "number",
|
|
196
196
|
"minimum": 0
|
|
197
197
|
},
|
|
198
|
+
"planningCandidatesOriginal": {
|
|
199
|
+
"type": "number",
|
|
200
|
+
"minimum": 0,
|
|
201
|
+
"description": "Local candidates before token-saving planning compaction."
|
|
202
|
+
},
|
|
203
|
+
"planningCandidates": {
|
|
204
|
+
"type": "number",
|
|
205
|
+
"minimum": 0,
|
|
206
|
+
"description": "Candidate bundles sent through hierarchical split planning."
|
|
207
|
+
},
|
|
198
208
|
"requests": {
|
|
199
209
|
"type": "number",
|
|
200
210
|
"minimum": 0
|
package/src/change-analysis.js
CHANGED
|
@@ -12,6 +12,7 @@ import {
|
|
|
12
12
|
isDeepAnalysis,
|
|
13
13
|
isGeneratedFile,
|
|
14
14
|
localOverview,
|
|
15
|
+
compactLocalPlanFacts,
|
|
15
16
|
localPlanBatches,
|
|
16
17
|
} from './local-analysis.js';
|
|
17
18
|
|
|
@@ -350,17 +351,17 @@ async function jsonCall(config, instruction, items, validate = null, stream = nu
|
|
|
350
351
|
}
|
|
351
352
|
}
|
|
352
353
|
|
|
353
|
-
export function validatePartition(groups, ids) {
|
|
354
|
+
export function validatePartition(groups, ids, requireSummary = true) {
|
|
354
355
|
const remaining = new Set(ids);
|
|
355
356
|
if (!Array.isArray(groups) || !groups.length)
|
|
356
357
|
throw fail(ERROR_CATEGORIES.RESPONSE_FORMAT, 'Analysis returned an empty partition.');
|
|
357
358
|
for (const group of groups) {
|
|
358
359
|
if (
|
|
360
|
+
!group ||
|
|
359
361
|
!Array.isArray(group.ids) ||
|
|
360
362
|
!group.ids.length ||
|
|
361
|
-
|
|
362
|
-
|
|
363
|
-
group.summary.length > 2000
|
|
363
|
+
(requireSummary &&
|
|
364
|
+
(typeof group.summary !== 'string' || !group.summary.trim() || group.summary.length > 2000))
|
|
364
365
|
)
|
|
365
366
|
throw fail(ERROR_CATEGORIES.RESPONSE_FORMAT, 'Analysis returned an invalid group.');
|
|
366
367
|
for (const id of group.ids)
|
|
@@ -714,22 +715,37 @@ export async function summarizeChanges(config, facts) {
|
|
|
714
715
|
}
|
|
715
716
|
|
|
716
717
|
export async function planAnalyzedChanges(config, facts, coverage = null, stream = null) {
|
|
717
|
-
let candidates = facts;
|
|
718
718
|
const deep = isDeepAnalysis(config) && coverage?.strategy !== 'auto';
|
|
719
|
+
const compacted = deep
|
|
720
|
+
? { facts, originalCandidates: facts.length, planningCandidates: facts.length }
|
|
721
|
+
: compactLocalPlanFacts(config, facts);
|
|
722
|
+
let candidates = compacted.facts;
|
|
723
|
+
if (coverage) {
|
|
724
|
+
coverage.planningCandidatesOriginal = compacted.originalCandidates;
|
|
725
|
+
coverage.planningCandidates = compacted.planningCandidates;
|
|
726
|
+
}
|
|
719
727
|
const policy = normalizeCommitPolicy(config.commitPolicy, config.language);
|
|
720
728
|
const cap = Math.min(
|
|
721
729
|
config.splitMaxDiffChars || 16000,
|
|
722
730
|
Math.floor(config.analysisBudget.limits.chunkInputTokens * 0.6),
|
|
723
731
|
);
|
|
732
|
+
// A model may keep every input as a separate commit. Limit candidate count
|
|
733
|
+
// by the response budget too, leaving room for a complete JSON plan.
|
|
734
|
+
const outputTokens = Math.min(2048, config.maxTokens || 1024);
|
|
735
|
+
const maxPlanItems = Math.max(
|
|
736
|
+
1,
|
|
737
|
+
Math.min(config.splitMaxPlanFiles || 100, Math.floor((outputTokens - 224) / 80)),
|
|
738
|
+
);
|
|
739
|
+
const batchConfig = { ...config, splitMaxPlanFiles: maxPlanItems };
|
|
724
740
|
for (let level = 0; level < 8; level++) {
|
|
725
741
|
const byId = new Map(candidates.map((x) => [x.id, x]));
|
|
726
742
|
const batches =
|
|
727
743
|
!deep && level === 0
|
|
728
|
-
? localPlanBatches(
|
|
744
|
+
? localPlanBatches(batchConfig, candidates, cap)
|
|
729
745
|
: packItems(
|
|
730
746
|
candidates.map(({ id, path, summary }) => ({ id, path, summary })),
|
|
731
747
|
cap,
|
|
732
|
-
|
|
748
|
+
maxPlanItems,
|
|
733
749
|
);
|
|
734
750
|
if (!deep && level === 0 && coverage) {
|
|
735
751
|
coverage.sampledFiles = batches.flat().filter((item) => item.representativeExcerpt).length;
|
|
@@ -761,10 +777,10 @@ export async function planAnalyzedChanges(config, facts, coverage = null, stream
|
|
|
761
777
|
};
|
|
762
778
|
const groups = await jsonCall(
|
|
763
779
|
config,
|
|
764
|
-
`Each commit message must follow this policy: ${JSON.stringify(policy)}.\nGroup related changes into logical commits, including related implementation and tests across directories. Return [{"ids":[input IDs],"summary":"factual combined change summary","subject":"commit subject","body":"optional commit body"}]. Assign every input ID exactly once. Do not merge unrelated changes just to reduce group count
|
|
780
|
+
`Each commit message must follow this policy: ${JSON.stringify(policy)}.\nGroup related changes into logical commits, including related implementation and tests across directories. Return [{"ids":[input IDs],"summary":"factual combined change summary","subject":"commit subject","body":"optional commit body"}]. Assign every input ID exactly once. Do not merge unrelated changes just to reduce group count.${finalPlan ? ' The summary field is optional in this final plan.' : ''}${!deep && level === 0 ? ' Inputs are a compact local inventory, not full semantic summaries. One ID may represent multiple files and must remain atomic; use kind, module, status, fileCount, examples, and any representativeExcerpt conservatively.' : ''}`,
|
|
765
781
|
batch,
|
|
766
782
|
(candidate) => {
|
|
767
|
-
validatePartition(candidate, ids);
|
|
783
|
+
validatePartition(candidate, ids, false);
|
|
768
784
|
if (candidate.some((group) => typeof group.subject !== 'string' || !group.subject.trim()))
|
|
769
785
|
throw fail(
|
|
770
786
|
ERROR_CATEGORIES.RESPONSE_FORMAT,
|
|
@@ -774,8 +790,23 @@ export async function planAnalyzedChanges(config, facts, coverage = null, stream
|
|
|
774
790
|
batchStream,
|
|
775
791
|
);
|
|
776
792
|
for (const group of groups) {
|
|
793
|
+
const sourceSummary = group.ids
|
|
794
|
+
.map((id) => {
|
|
795
|
+
const fact = byId.get(id);
|
|
796
|
+
return (
|
|
797
|
+
fact.summary ||
|
|
798
|
+
`${fact.status || 'Changed'} ${fact.kind || 'files'} in ${fact.module || fact.path || 'repository'} (${fact.files.length} files)`
|
|
799
|
+
).slice(0, 160);
|
|
800
|
+
})
|
|
801
|
+
.join('; ');
|
|
777
802
|
next.push({
|
|
778
803
|
...group,
|
|
804
|
+
summary:
|
|
805
|
+
typeof group.summary === 'string' &&
|
|
806
|
+
group.summary.trim() &&
|
|
807
|
+
group.summary.length <= 2000
|
|
808
|
+
? group.summary
|
|
809
|
+
: sourceSummary.slice(0, 2000),
|
|
779
810
|
id: `L${level}G${next.length}`,
|
|
780
811
|
files: group.ids.flatMap((id) => byId.get(id).files),
|
|
781
812
|
});
|
package/src/git-spool.js
CHANGED
|
@@ -11,7 +11,7 @@ process.once('exit', cleanupGitSpools);
|
|
|
11
11
|
|
|
12
12
|
// Git writes directly to a private temporary file, never to a Node stdout
|
|
13
13
|
// buffer. Consumers scan the captured bytes without reopening the worktree.
|
|
14
|
-
export function spoolGit(commands, cwd) {
|
|
14
|
+
export function spoolGit(commands, cwd, { env = undefined } = {}) {
|
|
15
15
|
const dir = mkdtempSync(join(tmpdir(), 'aicommit-diff-'));
|
|
16
16
|
const path = join(dir, 'patch');
|
|
17
17
|
let fd;
|
|
@@ -20,6 +20,7 @@ export function spoolGit(commands, cwd) {
|
|
|
20
20
|
for (const args of commands) {
|
|
21
21
|
execFileSync('git', ['--no-pager', ...args], {
|
|
22
22
|
cwd,
|
|
23
|
+
...(env ? { env } : {}),
|
|
23
24
|
stdio: ['ignore', fd, 'pipe'],
|
|
24
25
|
timeout: 120000,
|
|
25
26
|
maxBuffer: 1024 * 1024,
|
|
@@ -79,6 +80,23 @@ export function spoolGit(commands, cwd) {
|
|
|
79
80
|
}
|
|
80
81
|
if (pending.length) yield pending.toString('utf8');
|
|
81
82
|
},
|
|
83
|
+
*nulFields(maxFieldBytes = 16 * 1024) {
|
|
84
|
+
let pending = Buffer.alloc(0);
|
|
85
|
+
for (const buffer of this.buffers()) {
|
|
86
|
+
const data = pending.length ? Buffer.concat([pending, buffer]) : buffer;
|
|
87
|
+
let start = 0;
|
|
88
|
+
for (let end = data.indexOf(0); end !== -1; end = data.indexOf(0, start)) {
|
|
89
|
+
if (end - start > maxFieldBytes)
|
|
90
|
+
throw new Error(`NUL-delimited Git field exceeds ${maxFieldBytes} bytes.`);
|
|
91
|
+
yield data.subarray(start, end).toString('utf8');
|
|
92
|
+
start = end + 1;
|
|
93
|
+
}
|
|
94
|
+
pending = Buffer.from(data.subarray(start));
|
|
95
|
+
if (pending.length > maxFieldBytes)
|
|
96
|
+
throw new Error(`NUL-delimited Git field exceeds ${maxFieldBytes} bytes.`);
|
|
97
|
+
}
|
|
98
|
+
if (pending.length) throw new Error('NUL-delimited Git output was truncated.');
|
|
99
|
+
},
|
|
82
100
|
text(limit) {
|
|
83
101
|
if (size > limit) return null;
|
|
84
102
|
return [...this.lines()].join('').trim();
|
package/src/local-analysis.js
CHANGED
|
@@ -251,6 +251,86 @@ export function analyzeLocally(config, capture, protect, previews) {
|
|
|
251
251
|
};
|
|
252
252
|
}
|
|
253
253
|
|
|
254
|
+
function planningBucketKey(fact) {
|
|
255
|
+
const topLevel = fact.module === '.' ? '.' : fact.module.split('/')[0];
|
|
256
|
+
return JSON.stringify([
|
|
257
|
+
topLevel,
|
|
258
|
+
fact.kind,
|
|
259
|
+
fact.status,
|
|
260
|
+
fact.modeChange || '',
|
|
261
|
+
Boolean(fact.binary),
|
|
262
|
+
]);
|
|
263
|
+
}
|
|
264
|
+
|
|
265
|
+
function moduleSpan(facts) {
|
|
266
|
+
const modules = [...new Set(facts.map((fact) => fact.module))];
|
|
267
|
+
if (modules.length === 1) return modules[0];
|
|
268
|
+
const parts = modules.map((module) => module.split('/'));
|
|
269
|
+
const common = [];
|
|
270
|
+
for (let index = 0; index < parts[0].length; index++) {
|
|
271
|
+
const value = parts[0][index];
|
|
272
|
+
if (!parts.every((item) => item[index] === value)) break;
|
|
273
|
+
common.push(value);
|
|
274
|
+
}
|
|
275
|
+
return common.length ? `${common.join('/')}/*` : 'multiple modules';
|
|
276
|
+
}
|
|
277
|
+
|
|
278
|
+
// Extremely large auto-mode inventories cannot represent every independent
|
|
279
|
+
// file candidate to the model without hundreds of repeated requests. Bundle
|
|
280
|
+
// adjacent candidates locally while retaining their complete file mapping;
|
|
281
|
+
// the model groups bundle IDs, and execution still covers every original path.
|
|
282
|
+
export function compactLocalPlanFacts(config, facts) {
|
|
283
|
+
const maxItems = config.splitMaxPlanFiles || 100;
|
|
284
|
+
const threshold = maxItems * 4;
|
|
285
|
+
if (facts.length <= threshold) {
|
|
286
|
+
return { facts, originalCandidates: facts.length, planningCandidates: facts.length };
|
|
287
|
+
}
|
|
288
|
+
|
|
289
|
+
const target = Math.max(1, maxItems * 2);
|
|
290
|
+
const bundleSize = Math.ceil(facts.length / target);
|
|
291
|
+
const ordered = [...facts].sort(
|
|
292
|
+
(left, right) =>
|
|
293
|
+
planningBucketKey(left).localeCompare(planningBucketKey(right)) ||
|
|
294
|
+
left.module.localeCompare(right.module) ||
|
|
295
|
+
left.id.localeCompare(right.id),
|
|
296
|
+
);
|
|
297
|
+
const bundles = [];
|
|
298
|
+
let bucket = [];
|
|
299
|
+
let key = null;
|
|
300
|
+
const flush = () => {
|
|
301
|
+
if (!bucket.length) return;
|
|
302
|
+
const first = bucket[0];
|
|
303
|
+
const evidence = bucket.find((fact) => fact.evidence)?.evidence || '';
|
|
304
|
+
const modeChanges = new Set(bucket.map((fact) => fact.modeChange).filter(Boolean));
|
|
305
|
+
bundles.push({
|
|
306
|
+
id: `C${bundles.length + 1}`,
|
|
307
|
+
kind: first.kind,
|
|
308
|
+
module: moduleSpan(bucket),
|
|
309
|
+
status: first.status,
|
|
310
|
+
files: bucket.flatMap((fact) => fact.files),
|
|
311
|
+
additions: bucket.reduce((total, fact) => total + fact.additions, 0),
|
|
312
|
+
deletions: bucket.reduce((total, fact) => total + fact.deletions, 0),
|
|
313
|
+
modeChange: modeChanges.size === 1 ? [...modeChanges][0] : null,
|
|
314
|
+
binary: bucket.some((fact) => fact.binary),
|
|
315
|
+
evidence,
|
|
316
|
+
});
|
|
317
|
+
bucket = [];
|
|
318
|
+
};
|
|
319
|
+
|
|
320
|
+
for (const fact of ordered) {
|
|
321
|
+
const nextKey = planningBucketKey(fact);
|
|
322
|
+
if (bucket.length && (nextKey !== key || bucket.length >= bundleSize)) flush();
|
|
323
|
+
key = nextKey;
|
|
324
|
+
bucket.push(fact);
|
|
325
|
+
}
|
|
326
|
+
flush();
|
|
327
|
+
return {
|
|
328
|
+
facts: bundles,
|
|
329
|
+
originalCandidates: facts.length,
|
|
330
|
+
planningCandidates: bundles.length,
|
|
331
|
+
};
|
|
332
|
+
}
|
|
333
|
+
|
|
254
334
|
export function localPlanBatches(config, facts, cap) {
|
|
255
335
|
const error = () =>
|
|
256
336
|
fail(ERROR_CATEGORIES.CONFIG, 'A split candidate is too large for one planning request.', {
|
|
@@ -260,10 +340,7 @@ export function localPlanBatches(config, facts, cap) {
|
|
|
260
340
|
const batches = [];
|
|
261
341
|
let batch = [];
|
|
262
342
|
for (const fact of facts) {
|
|
263
|
-
const item =
|
|
264
|
-
...summaryOf(fact, false),
|
|
265
|
-
summary: 'Local candidate only; verify grouping. Content not fully analyzed.',
|
|
266
|
-
};
|
|
343
|
+
const item = summaryOf(fact, false);
|
|
267
344
|
if (Buffer.byteLength(JSON.stringify([item])) > cap) throw error();
|
|
268
345
|
const candidate = [...batch, item];
|
|
269
346
|
if (
|
|
@@ -278,12 +355,15 @@ export function localPlanBatches(config, facts, cap) {
|
|
|
278
355
|
if (batch.length) batches.push(batch);
|
|
279
356
|
|
|
280
357
|
const factsById = new Map(facts.map((fact) => [fact.id, fact]));
|
|
358
|
+
let evidenceCount = 0;
|
|
281
359
|
for (const items of batches) {
|
|
282
360
|
for (const item of items) {
|
|
361
|
+
if (evidenceCount >= 16) break;
|
|
283
362
|
const evidence = factsById.get(item.id)?.evidence;
|
|
284
363
|
if (!evidence) continue;
|
|
285
364
|
item.representativeExcerpt = evidence;
|
|
286
365
|
if (Buffer.byteLength(JSON.stringify(items)) > cap) delete item.representativeExcerpt;
|
|
366
|
+
else evidenceCount++;
|
|
287
367
|
}
|
|
288
368
|
}
|
|
289
369
|
return batches;
|
package/src/split-checkpoint.js
CHANGED
|
@@ -15,7 +15,10 @@ import { validateSplitPlanArtifact } from './split-plan.js';
|
|
|
15
15
|
|
|
16
16
|
export const SPLIT_CHECKPOINT_KIND = 'aicommit-split-checkpoint';
|
|
17
17
|
export const SPLIT_CHECKPOINT_VERSION = 1;
|
|
18
|
-
|
|
18
|
+
// Checkpoints contain both the validated plan and one object snapshot per real
|
|
19
|
+
// path, so their bound must be larger than the plan artifact bound. This still
|
|
20
|
+
// caps parsing of local metadata while allowing the supported 10,000 changes.
|
|
21
|
+
const MAX_CHECKPOINT_BYTES = 128 * 1024 * 1024;
|
|
19
22
|
const OID_RE = /^(?:[0-9a-f]{40}|[0-9a-f]{64})$/;
|
|
20
23
|
|
|
21
24
|
function object(value) {
|
|
@@ -197,7 +200,7 @@ export function readSplitCheckpoint(projectRoot) {
|
|
|
197
200
|
if (!stat.isFile() || stat.isSymbolicLink()) {
|
|
198
201
|
throw new Error('Split checkpoint must be a regular, non-symbolic-link file.');
|
|
199
202
|
}
|
|
200
|
-
if (stat.size > MAX_CHECKPOINT_BYTES) throw new Error('Split checkpoint exceeds
|
|
203
|
+
if (stat.size > MAX_CHECKPOINT_BYTES) throw new Error('Split checkpoint exceeds 128 MiB.');
|
|
201
204
|
let parsed;
|
|
202
205
|
try {
|
|
203
206
|
parsed = JSON.parse(readFileSync(path, 'utf8'));
|
|
@@ -213,10 +216,14 @@ export function writeSplitCheckpoint(projectRoot, input) {
|
|
|
213
216
|
...input,
|
|
214
217
|
updatedAt: new Date().toISOString(),
|
|
215
218
|
});
|
|
219
|
+
const serialized = JSON.stringify(checkpoint, null, 2) + '\n';
|
|
220
|
+
if (Buffer.byteLength(serialized) > MAX_CHECKPOINT_BYTES) {
|
|
221
|
+
throw new Error('Split checkpoint exceeds the 128 MiB limit.');
|
|
222
|
+
}
|
|
216
223
|
mkdirSync(dirname(path), { recursive: true, mode: 0o700 });
|
|
217
224
|
const temporary = `${path}.${process.pid}.${Date.now()}.tmp`;
|
|
218
225
|
try {
|
|
219
|
-
writeFileSync(temporary,
|
|
226
|
+
writeFileSync(temporary, serialized, {
|
|
220
227
|
encoding: 'utf8',
|
|
221
228
|
mode: 0o600,
|
|
222
229
|
});
|
package/src/split-hunks.js
CHANGED
|
@@ -4,6 +4,8 @@ import { mkdtempSync, rmSync } from 'node:fs';
|
|
|
4
4
|
import { tmpdir } from 'node:os';
|
|
5
5
|
import { join } from 'node:path';
|
|
6
6
|
|
|
7
|
+
import { pathBatches } from './git.js';
|
|
8
|
+
|
|
7
9
|
const HUNK_HEADER_RE = /^@@ -(\d+)(?:,(\d+))? \+(\d+)(?:,(\d+))? @@/;
|
|
8
10
|
|
|
9
11
|
function git(projectRoot, args, { indexPath = null, input = undefined } = {}) {
|
|
@@ -215,9 +217,13 @@ export function validateHunkTransaction(projectRoot, plan, snapshots) {
|
|
|
215
217
|
}
|
|
216
218
|
|
|
217
219
|
const finalEntries = parseIndexEntries(
|
|
218
|
-
|
|
219
|
-
|
|
220
|
-
|
|
220
|
+
pathBatches(snapshots.map((item) => item.path))
|
|
221
|
+
.map((batch) =>
|
|
222
|
+
git(projectRoot, ['--literal-pathspecs', 'ls-files', '--stage', '-z', '--', ...batch], {
|
|
223
|
+
indexPath,
|
|
224
|
+
}),
|
|
225
|
+
)
|
|
226
|
+
.join(''),
|
|
221
227
|
);
|
|
222
228
|
for (const entry of snapshots) {
|
|
223
229
|
if (!entriesEqual(finalEntries.get(entry.path) || null, entry.target)) {
|
package/src/split-plan.js
CHANGED
|
@@ -5,7 +5,10 @@ import { normalizeCommitPolicy, validateCommitCandidate } from './policy.js';
|
|
|
5
5
|
|
|
6
6
|
export const SPLIT_PLAN_KIND = 'aicommit-split-plan';
|
|
7
7
|
export const SPLIT_PLAN_VERSION = 1;
|
|
8
|
-
|
|
8
|
+
// A file-level split can legitimately contain thousands of long paths. Keep
|
|
9
|
+
// reads bounded, but size the artifact limit for the validated 10,000-change
|
|
10
|
+
// ceiling instead of rejecting plans that this module can create itself.
|
|
11
|
+
const MAX_PLAN_BYTES = 64 * 1024 * 1024;
|
|
9
12
|
const HASH_RE = /^(?:[0-9a-f]{40}|[0-9a-f]{64})$/;
|
|
10
13
|
const FINGERPRINT_RE = /^[0-9a-f]{64}$/;
|
|
11
14
|
const HUNK_HASH_RE = /^[0-9a-f]{64}$/;
|
|
@@ -241,13 +244,14 @@ export function validateSplitPlanArtifact(input) {
|
|
|
241
244
|
normalizeGroup(group, index, allowedPaths, hunkCatalog, commitPolicy, input.language),
|
|
242
245
|
);
|
|
243
246
|
const assignedFiles = groups.flatMap((group) => group.files);
|
|
244
|
-
|
|
247
|
+
const assignedFileSet = new Set(assignedFiles);
|
|
248
|
+
if (assignedFileSet.size !== assignedFiles.length) {
|
|
245
249
|
throw new Error('Split plan assigns a path to more than one group.');
|
|
246
250
|
}
|
|
247
251
|
const assignedHunks = new Set();
|
|
248
252
|
for (const group of groups) {
|
|
249
253
|
for (const assignment of group.hunks || []) {
|
|
250
|
-
if (
|
|
254
|
+
if (assignedFileSet.has(assignment.path)) {
|
|
251
255
|
throw new Error(`Split plan assigns both a whole file and hunks: ${assignment.path}`);
|
|
252
256
|
}
|
|
253
257
|
for (const id of assignment.ids) {
|
|
@@ -260,7 +264,7 @@ export function validateSplitPlanArtifact(input) {
|
|
|
260
264
|
}
|
|
261
265
|
}
|
|
262
266
|
const missing = displayPaths.filter((path) => {
|
|
263
|
-
if (
|
|
267
|
+
if (assignedFileSet.has(path)) return false;
|
|
264
268
|
const ids = hunkCatalog.get(path);
|
|
265
269
|
return !ids || [...ids].some((id) => !assignedHunks.has(`${path}\0${id}`));
|
|
266
270
|
});
|
|
@@ -310,10 +314,14 @@ export function createSplitPlanArtifact({
|
|
|
310
314
|
export async function writeSplitPlanArtifact(path, artifact) {
|
|
311
315
|
const absolute = resolve(path);
|
|
312
316
|
const validated = validateSplitPlanArtifact(artifact);
|
|
317
|
+
const serialized = JSON.stringify(validated, null, 2) + '\n';
|
|
318
|
+
if (Buffer.byteLength(serialized) > MAX_PLAN_BYTES) {
|
|
319
|
+
throw new Error('Split plan exceeds the 64 MiB limit.');
|
|
320
|
+
}
|
|
313
321
|
await mkdir(dirname(absolute), { recursive: true, mode: 0o700 });
|
|
314
322
|
const temporary = `${absolute}.${process.pid}.${Date.now()}.tmp`;
|
|
315
323
|
try {
|
|
316
|
-
await writeFile(temporary,
|
|
324
|
+
await writeFile(temporary, serialized, {
|
|
317
325
|
encoding: 'utf8',
|
|
318
326
|
mode: 0o600,
|
|
319
327
|
flag: 'wx',
|
|
@@ -341,7 +349,7 @@ export async function readSplitPlanArtifact(path) {
|
|
|
341
349
|
if (!stat.isFile() || stat.isSymbolicLink()) {
|
|
342
350
|
throw new Error('Split plan must be a regular, non-symbolic-link file.');
|
|
343
351
|
}
|
|
344
|
-
if (stat.size > MAX_PLAN_BYTES) throw new Error('Split plan exceeds the
|
|
352
|
+
if (stat.size > MAX_PLAN_BYTES) throw new Error('Split plan exceeds the 64 MiB limit.');
|
|
345
353
|
let parsed;
|
|
346
354
|
try {
|
|
347
355
|
parsed = JSON.parse(await readFile(absolute, 'utf8'));
|
package/src/split.js
CHANGED
|
@@ -236,20 +236,33 @@ export function getSplitChangedFiles(cwd, scope = 'all') {
|
|
|
236
236
|
// the planner must see the same latest content executeSplit will git add -A.
|
|
237
237
|
function getWorkingTreeDiff(projectRoot, head, contextLines, paths = []) {
|
|
238
238
|
const u = unifiedArg(contextLines);
|
|
239
|
-
const
|
|
240
|
-
|
|
241
|
-
|
|
239
|
+
const batches = paths.length ? pathBatches(paths) : [[]];
|
|
240
|
+
const parts = [];
|
|
241
|
+
for (const batch of batches) {
|
|
242
|
+
const pathArgs = batch.length ? ['--', ...batch] : [];
|
|
243
|
+
if (head) {
|
|
244
|
+
parts.push(readGit(['diff', u, 'HEAD', ...pathArgs], projectRoot).trim());
|
|
245
|
+
continue;
|
|
246
|
+
}
|
|
247
|
+
parts.push(readGit(['diff', u, '--cached', ...pathArgs], projectRoot).trim());
|
|
248
|
+
parts.push(readGit(['diff', u, ...pathArgs], projectRoot).trim());
|
|
242
249
|
}
|
|
243
|
-
|
|
244
|
-
const cached = readGit(['diff', u, '--cached', ...pathArgs], projectRoot).trim();
|
|
245
|
-
const unstaged = readGit(['diff', u, ...pathArgs], projectRoot).trim();
|
|
246
|
-
return [cached, unstaged].filter(Boolean).join('\n');
|
|
250
|
+
return parts.filter(Boolean).join('\n');
|
|
247
251
|
}
|
|
248
252
|
|
|
249
253
|
export function getSplitDiff(projectRoot, head, contextLines, scope = 'all', paths = []) {
|
|
250
254
|
if (scope === 'staged') {
|
|
251
|
-
const
|
|
252
|
-
return
|
|
255
|
+
const batches = paths.length ? pathBatches(paths) : [[]];
|
|
256
|
+
return batches
|
|
257
|
+
.map((batch) => {
|
|
258
|
+
const pathArgs = batch.length ? ['--', ...batch] : [];
|
|
259
|
+
return readGit(
|
|
260
|
+
['diff', unifiedArg(contextLines), '--cached', ...pathArgs],
|
|
261
|
+
projectRoot,
|
|
262
|
+
).trim();
|
|
263
|
+
})
|
|
264
|
+
.filter(Boolean)
|
|
265
|
+
.join('\n');
|
|
253
266
|
}
|
|
254
267
|
return getWorkingTreeDiff(projectRoot, head, contextLines, paths);
|
|
255
268
|
}
|
|
@@ -261,7 +274,7 @@ export function getSplitStateFingerprint(projectRoot, head, files, scope = 'all'
|
|
|
261
274
|
const hash = createHash('sha256');
|
|
262
275
|
if (scope === 'staged') {
|
|
263
276
|
hash.update(head ? readGit(['rev-parse', 'HEAD'], projectRoot).trim() : '<unborn>');
|
|
264
|
-
hash
|
|
277
|
+
updateGitHash(hash, ['ls-files', '--stage', '-z'], projectRoot);
|
|
265
278
|
return hash.digest('hex');
|
|
266
279
|
}
|
|
267
280
|
files ||= getAllChangedFiles(projectRoot);
|
|
@@ -909,9 +922,9 @@ function getGroupDiff(
|
|
|
909
922
|
return 'Changed files (new files, no diff available):\n' + parts.join('\n');
|
|
910
923
|
}
|
|
911
924
|
|
|
912
|
-
function parseStageZeroEntries(
|
|
925
|
+
function parseStageZeroEntries(fields) {
|
|
913
926
|
const entries = new Map();
|
|
914
|
-
for (const field of
|
|
927
|
+
for (const field of fields) {
|
|
915
928
|
if (!field) continue;
|
|
916
929
|
const match = field.match(/^(\d+) ([0-9a-f]+) (\d)\t([\s\S]+)$/);
|
|
917
930
|
if (!match || match[3] !== '0') continue;
|
|
@@ -920,8 +933,37 @@ function parseStageZeroEntries(text) {
|
|
|
920
933
|
return entries;
|
|
921
934
|
}
|
|
922
935
|
|
|
923
|
-
function
|
|
924
|
-
|
|
936
|
+
function* filterGitFields(fields, wanted) {
|
|
937
|
+
for (const field of fields) {
|
|
938
|
+
const tab = field.indexOf('\t');
|
|
939
|
+
if (tab !== -1 && wanted.has(field.slice(tab + 1))) yield field;
|
|
940
|
+
}
|
|
941
|
+
}
|
|
942
|
+
|
|
943
|
+
function readStageZeroEntries(projectRoot, paths = null, indexPath = null) {
|
|
944
|
+
if (paths !== null && !paths.length) return new Map();
|
|
945
|
+
const wanted = paths === null ? null : new Set(paths);
|
|
946
|
+
const batches = paths === null ? null : pathBatches(paths);
|
|
947
|
+
if (batches && batches.length <= 8) {
|
|
948
|
+
const output = batches
|
|
949
|
+
.map((batch) => {
|
|
950
|
+
const args = ['--literal-pathspecs', 'ls-files', '--stage', '-z', '--', ...batch];
|
|
951
|
+
return indexPath
|
|
952
|
+
? runGitWithIndex(args, projectRoot, indexPath)
|
|
953
|
+
: readGit(args, projectRoot);
|
|
954
|
+
})
|
|
955
|
+
.join('');
|
|
956
|
+
return parseStageZeroEntries(output.split('\0'));
|
|
957
|
+
}
|
|
958
|
+
const source = spoolGit([['--literal-pathspecs', 'ls-files', '--stage', '-z']], projectRoot, {
|
|
959
|
+
...(indexPath ? { env: { ...process.env, GIT_INDEX_FILE: indexPath } } : {}),
|
|
960
|
+
});
|
|
961
|
+
try {
|
|
962
|
+
const fields = wanted ? filterGitFields(source.nulFields(), wanted) : source.nulFields();
|
|
963
|
+
return parseStageZeroEntries(fields);
|
|
964
|
+
} finally {
|
|
965
|
+
source.dispose();
|
|
966
|
+
}
|
|
925
967
|
}
|
|
926
968
|
|
|
927
969
|
function runGitWithIndex(args, projectRoot, indexPath, inherit = false, input = undefined) {
|
|
@@ -977,25 +1019,11 @@ function activeCommitHooks(projectRoot) {
|
|
|
977
1019
|
|
|
978
1020
|
function submodulePaths(projectRoot, paths) {
|
|
979
1021
|
const found = new Set();
|
|
980
|
-
const
|
|
981
|
-
.
|
|
982
|
-
readGit(['--literal-pathspecs', 'ls-files', '--stage', '-z', '--', ...batch], projectRoot),
|
|
983
|
-
)
|
|
984
|
-
.join('');
|
|
985
|
-
for (const entry of index.split('\0')) {
|
|
986
|
-
const match = entry.match(/^160000 [0-9a-f]+ \d\t([\s\S]+)$/);
|
|
987
|
-
if (match) found.add(match[1]);
|
|
988
|
-
}
|
|
1022
|
+
for (const [path, entry] of readStageZeroEntries(projectRoot, paths))
|
|
1023
|
+
if (entry.mode === '160000') found.add(path);
|
|
989
1024
|
if (hasHead(projectRoot)) {
|
|
990
|
-
const
|
|
991
|
-
.
|
|
992
|
-
readGit(['--literal-pathspecs', 'ls-tree', '-rz', 'HEAD', '--', ...batch], projectRoot),
|
|
993
|
-
)
|
|
994
|
-
.join('');
|
|
995
|
-
for (const entry of tree.split('\0')) {
|
|
996
|
-
const match = entry.match(/^160000 commit [0-9a-f]+\t([\s\S]+)$/);
|
|
997
|
-
if (match) found.add(match[1]);
|
|
998
|
-
}
|
|
1025
|
+
for (const [path, entry] of readHeadEntries(projectRoot, paths))
|
|
1026
|
+
if (entry.mode === '160000') found.add(path);
|
|
999
1027
|
}
|
|
1000
1028
|
return [...found];
|
|
1001
1029
|
}
|
|
@@ -1112,7 +1140,7 @@ function resetCommittedPaths(projectRoot, groups, allFiles) {
|
|
|
1112
1140
|
}
|
|
1113
1141
|
|
|
1114
1142
|
function captureTargetEntries(projectRoot, scope, paths) {
|
|
1115
|
-
if (scope === 'staged') return readStageZeroEntries(projectRoot);
|
|
1143
|
+
if (scope === 'staged') return readStageZeroEntries(projectRoot, paths);
|
|
1116
1144
|
const tempDir = mkdtempSync(join(tmpdir(), 'aicommit-split-snapshot-'));
|
|
1117
1145
|
const indexPath = join(tempDir, 'index');
|
|
1118
1146
|
try {
|
|
@@ -1128,9 +1156,7 @@ function captureTargetEntries(projectRoot, scope, paths) {
|
|
|
1128
1156
|
false,
|
|
1129
1157
|
paths.join('\0') + '\0',
|
|
1130
1158
|
);
|
|
1131
|
-
return
|
|
1132
|
-
runGitWithIndex(['ls-files', '--stage', '-z'], projectRoot, indexPath),
|
|
1133
|
-
);
|
|
1159
|
+
return readStageZeroEntries(projectRoot, paths, indexPath);
|
|
1134
1160
|
} finally {
|
|
1135
1161
|
rmSync(tempDir, { recursive: true, force: true });
|
|
1136
1162
|
}
|
|
@@ -1139,7 +1165,7 @@ function captureTargetEntries(projectRoot, scope, paths) {
|
|
|
1139
1165
|
function captureCheckpointSnapshots(projectRoot, scope, allFiles) {
|
|
1140
1166
|
const paths = [...new Set(allFiles.flatMap((change) => change.addPaths))].sort();
|
|
1141
1167
|
const target = captureTargetEntries(projectRoot, scope, paths);
|
|
1142
|
-
const index = readStageZeroEntries(projectRoot);
|
|
1168
|
+
const index = readStageZeroEntries(projectRoot, paths);
|
|
1143
1169
|
return paths.map((path) => ({
|
|
1144
1170
|
path,
|
|
1145
1171
|
target: target.get(path) || null,
|
|
@@ -1243,15 +1269,35 @@ function splitCommitFailure(projectRoot, message) {
|
|
|
1243
1269
|
function readHeadEntries(projectRoot, paths) {
|
|
1244
1270
|
const entries = new Map();
|
|
1245
1271
|
if (!hasHead(projectRoot) || !paths.length) return entries;
|
|
1246
|
-
const
|
|
1247
|
-
|
|
1248
|
-
|
|
1249
|
-
|
|
1250
|
-
|
|
1251
|
-
|
|
1252
|
-
|
|
1253
|
-
|
|
1254
|
-
|
|
1272
|
+
const wanted = new Set(paths);
|
|
1273
|
+
const batches = pathBatches(paths);
|
|
1274
|
+
let source = null;
|
|
1275
|
+
try {
|
|
1276
|
+
const fields =
|
|
1277
|
+
batches.length <= 8
|
|
1278
|
+
? batches
|
|
1279
|
+
.map((batch) =>
|
|
1280
|
+
readGit(
|
|
1281
|
+
['--literal-pathspecs', 'ls-tree', '-rz', 'HEAD', '--', ...batch],
|
|
1282
|
+
projectRoot,
|
|
1283
|
+
),
|
|
1284
|
+
)
|
|
1285
|
+
.join('')
|
|
1286
|
+
.split('\0')
|
|
1287
|
+
: filterGitFields(
|
|
1288
|
+
(source = spoolGit(
|
|
1289
|
+
[['--literal-pathspecs', 'ls-tree', '-rz', 'HEAD']],
|
|
1290
|
+
projectRoot,
|
|
1291
|
+
)).nulFields(),
|
|
1292
|
+
wanted,
|
|
1293
|
+
);
|
|
1294
|
+
for (const field of fields) {
|
|
1295
|
+
if (!field) continue;
|
|
1296
|
+
const match = field.match(/^(\d+) \S+ ([0-9a-f]+)\t([\s\S]+)$/);
|
|
1297
|
+
if (match) entries.set(match[3], { mode: match[1], oid: match[2] });
|
|
1298
|
+
}
|
|
1299
|
+
} finally {
|
|
1300
|
+
source?.dispose();
|
|
1255
1301
|
}
|
|
1256
1302
|
return entries;
|
|
1257
1303
|
}
|
|
@@ -1289,7 +1335,7 @@ function reconcileCompletedIndex(projectRoot, checkpoint) {
|
|
|
1289
1335
|
const paths = [
|
|
1290
1336
|
...new Set(groups.flatMap((group) => expandGroupPaths(group, checkpoint.plan.changes))),
|
|
1291
1337
|
];
|
|
1292
|
-
const current = readStageZeroEntries(projectRoot);
|
|
1338
|
+
const current = readStageZeroEntries(projectRoot, paths);
|
|
1293
1339
|
const head = readHeadEntries(projectRoot, paths);
|
|
1294
1340
|
const snapshots = new Map(checkpoint.snapshots.map((snapshot) => [snapshot.path, snapshot]));
|
|
1295
1341
|
for (const path of paths) {
|
|
@@ -1826,6 +1872,13 @@ export async function splitFlow(
|
|
|
1826
1872
|
analysis.coverage,
|
|
1827
1873
|
planningStream,
|
|
1828
1874
|
);
|
|
1875
|
+
if (
|
|
1876
|
+
analysis.coverage.planningCandidates < analysis.coverage.planningCandidatesOriginal
|
|
1877
|
+
) {
|
|
1878
|
+
console.error(
|
|
1879
|
+
` Planning input: bundled ${analysis.coverage.planningCandidatesOriginal} local candidates into ${analysis.coverage.planningCandidates} token-efficient groups.`,
|
|
1880
|
+
);
|
|
1881
|
+
}
|
|
1829
1882
|
} catch (err) {
|
|
1830
1883
|
const exhausted =
|
|
1831
1884
|
err.data?.analysis?.exhausted ||
|