@lazyingart/agintiflow 0.20.336-integration.3 → 0.20.339-integration.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/aginti-work-examples/README.md +22 -0
- package/aginti-work-examples/deepseek-cli-fallback-20261001.md +82 -0
- package/bin/aginti-execution-worker.js +5 -0
- package/docs/DEEPSEEK_ARTIFACT_COMPLETION.md +62 -0
- package/docs/deepseek-cli-fallback.md +134 -0
- package/docs/general-agent-backend-research-2026-09-27.md +213 -0
- package/docs/integration-deep-research.md +126 -1
- package/docs/integration-hosted-model.md +162 -0
- package/docs/public-pdf-acquisition.md +458 -0
- package/docs/supervision-campaign-ledger.md +81 -0
- package/package.json +37 -8
- package/scripts/eval-deepseek-cli.mjs +158 -0
- package/scripts/smoke-cli-chat.js +2 -0
- package/scripts/smoke-coding-tools.js +8 -0
- package/scripts/smoke-execution-worker-systemd-boundary.js +26 -0
- package/scripts/smoke-integration-analysis-api-server.js +31 -0
- package/scripts/smoke-integration-analysis-planner.js +765 -22
- package/scripts/smoke-integration-analysis-session-service.js +195 -9
- package/scripts/smoke-integration-api.js +10 -0
- package/scripts/smoke-integration-grounded-search.js +88 -1
- package/scripts/smoke-integration-model-binding.js +109 -0
- package/scripts/smoke-public-pdf-download.js +30 -0
- package/scripts/smoke-scs-evidence-visibility.js +26 -0
- package/scripts/smoke-truthful-completion.js +151 -0
- package/scripts/test-document-worker-fixture.js +4 -2
- package/src/agent-runner.js +75 -39
- package/src/cli.js +7 -2
- package/src/command-policy.js +18 -2
- package/src/config.js +5 -1
- package/src/engineering-guidance.js +1 -1
- package/src/execution-worker-systemd-boundary.js +28 -11
- package/src/integration-acquired-paper-contract.js +199 -0
- package/src/integration-acquired-paper-transfer.js +110 -0
- package/src/integration-analysis-cli.js +24 -3
- package/src/integration-analysis-config.js +114 -26
- package/src/integration-analysis-planner.js +445 -145
- package/src/integration-analysis-server.js +44 -12
- package/src/integration-analysis-session-service.js +536 -82
- package/src/integration-api.js +9 -5
- package/src/integration-document-worker-config.js +8 -1
- package/src/integration-document-worker-server.js +47 -2
- package/src/integration-document-worker-service.js +48 -0
- package/src/integration-file-worker-client.js +132 -4
- package/src/integration-file-worker-store.js +230 -155
- package/src/integration-grounded-search.js +57 -22
- package/src/integration-model-binding.js +123 -0
- package/src/integration-paper-acquisition-contract.js +131 -0
- package/src/integration-paper-acquisition.js +74 -0
- package/src/integration-paper-checkpoint.js +52 -0
- package/src/integration-paper-selection.js +65 -0
- package/src/integration-policy.js +27 -2
- package/src/integration-research-synthesis.js +80 -0
- package/src/model-client.js +50 -66
- package/src/provider-runtime.js +3 -2
- package/src/public-pdf-download.js +233 -0
- package/src/scs-evidence.js +26 -9
- package/test/cli-fallback-policy.test.js +109 -0
- package/test/fixtures/acquired-paper.js +48 -0
- package/test/fixtures/paper-download.js +37 -0
- package/test/fixtures/paper-recovery-child.js +38 -0
- package/test/fixtures/paper-recovery-runner.js +55 -0
- package/test/fixtures/paper-source.js +19 -0
- package/test/fixtures/paper-worker.js +32 -0
- package/test/integration-acquired-paper-http.test.js +261 -0
- package/test/integration-acquired-paper-session.test.js +232 -0
- package/test/integration-acquired-paper-store.test.js +370 -0
- package/test/integration-acquired-paper-transfer.test.js +83 -0
- package/test/integration-file-worker-cancellation.test.js +93 -0
- package/test/integration-independent-vision.test.js +141 -0
- package/test/integration-paper-acquisition.test.js +351 -0
- package/test/integration-paper-recovery.test.js +246 -0
- package/test/integration-paper-selection.test.js +60 -0
- package/test/integration-research-synthesis.test.js +111 -0
- package/test/integration-vision-inference.test.js +43 -0
- package/test/model-request-lifecycle.test.js +171 -0
- package/test/provider-runtime-private-health.test.js +64 -0
- package/test/public-pdf-download.test.js +266 -0
|
@@ -0,0 +1,22 @@
|
|
|
1
|
+
# AgInTiFlow Work Examples
|
|
2
|
+
|
|
3
|
+
This directory records supervised AgInTiFlow homework runs. It is not a generic demo gallery; each entry should include evidence that AgInTiFlow actually completed the task.
|
|
4
|
+
|
|
5
|
+
## Rule
|
|
6
|
+
|
|
7
|
+
Do not record a task as complete only because the agent said it was complete. Each example needs durable proof such as:
|
|
8
|
+
|
|
9
|
+
- source files or commits,
|
|
10
|
+
- build/test logs,
|
|
11
|
+
- screenshots or PDFs,
|
|
12
|
+
- session id and tmux session name,
|
|
13
|
+
- artifact paths that still exist,
|
|
14
|
+
- notes about flaws found and AgInTiFlow upgrades made.
|
|
15
|
+
|
|
16
|
+
## Current Examples
|
|
17
|
+
|
|
18
|
+
| Example | Profile | Status | Evidence |
|
|
19
|
+
| --- | --- | --- | --- |
|
|
20
|
+
| `android-tipsplit` | `android` + `auto` capability hardening | Verified | Android app built/tested/installed/launched; durable screenshot copied here |
|
|
21
|
+
|
|
22
|
+
Future runs should add a subfolder per task, plus a row in `homework-ledger.md`.
|
|
@@ -0,0 +1,82 @@
|
|
|
1
|
+
# DeepSeek CLI fallback acceptance — 2026-10-01
|
|
2
|
+
|
|
3
|
+
## Scope
|
|
4
|
+
|
|
5
|
+
Improve everyday CLI work without a Codex wrapper: inspect files, create a
|
|
6
|
+
summary, repair a small project, run real tests, and continue the same session.
|
|
7
|
+
AgInTi performed all target-workspace edits through DeepSeek. The supervising
|
|
8
|
+
agent changed AgInTi itself, supplied synthetic fixtures and ordinary requests,
|
|
9
|
+
and independently checked outputs. No user's project was used as a test fixture.
|
|
10
|
+
|
|
11
|
+
The installed baseline was `0.20.331`. Development started from `486c56c`, which
|
|
12
|
+
already contained the undeployed request-lifecycle and read-only-input fixes on
|
|
13
|
+
top of integration `bf9c3e0`. The new release is `0.20.337-integration.0`, an opt-in
|
|
14
|
+
integration release rather than a promotion of the npm `latest` channel.
|
|
15
|
+
|
|
16
|
+
## Failures were reproduced, not inferred
|
|
17
|
+
|
|
18
|
+
| Observation | Evidence | Repair or conclusion |
|
|
19
|
+
| --- | --- | --- |
|
|
20
|
+
| Baseline summary request temporarily changed `risks.txt` | Live Flash tool events showed an unwanted patch, a compensating restoration, and then the output; 8 model turns / 14 seconds | Deploy the existing source/output contract repair. The same request took 4 turns / 6 seconds and never edited either source. Timings are individual observations, not a benchmark. |
|
|
21
|
+
| Fast/manual selection silently ran Main | Offline config reproduced Flash becoming Pro, and an explicit DeepSeek manual model becoming a separately configured OpenAI main model | Keep the selected executor while retaining SCS planning and completion validation. |
|
|
22
|
+
| Could enable wrappers but could not explicitly disable them on resume | `--no-wrappers` was rejected; a stored true value had no matching negative CLI patch | Add symmetric parsing and durable resume support, including an ambient-true regression. |
|
|
23
|
+
| Retry could outlive its deadline | Earlier deterministic reproduction and eight lifecycle regressions | Include the previous cancellation/deadline repair; do not claim provider token generation became faster. |
|
|
24
|
+
| Final read-only discovery stopped an otherwise fixed task | A real `find` pipeline omitted `-maxdepth`; the CLI paused asking for broad host permission | Keep the command blocked but classify it as recoverable discovery. Preserve denial through pipelines/sequences. The same session continued and completed with unchanged permissions. |
|
|
25
|
+
| Optional inline JavaScript stopped a second restricted-host run | A compound `npm test; node -e ...` command was not admitted | Keep that host restriction. Describe host tool limits to the model and recommend Normal Docker workspace for general coding. Do not expand arbitrary host execution just to make a test pass. |
|
|
26
|
+
| The prompt itself encouraged unrelated cache searches | Engineering guidance suggested unbounded Python-cache discovery irrespective of stack | Use relevant-stack, bounded inspection and scoped claims; no unrelated cleanup requirement. |
|
|
27
|
+
|
|
28
|
+
The first Pro code fix passed its supplied tests, but an independent `10.075`
|
|
29
|
+
rounding example failed. A normal follow-up in the same session produced a
|
|
30
|
+
decimal-string repair and a new regression, passing the independent check. This
|
|
31
|
+
is useful evidence of resumability, and also a reminder that passing supplied
|
|
32
|
+
tests alone does not establish complete correctness. No money-specific patch was
|
|
33
|
+
added to AgInTi core.
|
|
34
|
+
|
|
35
|
+
## Verification
|
|
36
|
+
|
|
37
|
+
- Full existing `npm test`: passed, with live provider-attribution probing off.
|
|
38
|
+
Existing occupied-port skips in document-worker checks remain skips, not passes
|
|
39
|
+
for those external services.
|
|
40
|
+
- Eight request-lifecycle regressions passed.
|
|
41
|
+
- Five new CLI routing/wrapper/discovery tests passed; the routing and discovery
|
|
42
|
+
reproductions failed before their repairs.
|
|
43
|
+
- Coding-tool policy, model-role, syntax, and CLI checks cover the changed paths.
|
|
44
|
+
- A denied discovery command is still denied. Mutating find actions, shell
|
|
45
|
+
execution, redirection, and destructive commands are not admitted by the repair.
|
|
46
|
+
- Live evaluator: `scripts/eval-deepseek-cli.mjs`. It retains each failed attempt
|
|
47
|
+
rather than overwriting it with a later pass. Its default follows the shipped
|
|
48
|
+
Normal Docker workspace; `AGINTIFLOW_EVAL_SANDBOX=host` exercises restricted host
|
|
49
|
+
behavior. A failed case exits nonzero.
|
|
50
|
+
|
|
51
|
+
Synthetic prompts, model results, events, independent test logs, and original
|
|
52
|
+
input hashes remain in a private acceptance directory. Credentials and raw
|
|
53
|
+
session history are not part of this public record or npm package.
|
|
54
|
+
|
|
55
|
+
The final fresh-workspace run passed all three cases with the Flash executor:
|
|
56
|
+
|
|
57
|
+
| Case | Model turns | Wall time | Independent result |
|
|
58
|
+
| --- | ---: | ---: | --- |
|
|
59
|
+
| Read two sources and write a summary | 5 | 9.1 s | Both sources unchanged; summary names both sources |
|
|
60
|
+
| Repair invoice utility and add regression coverage | 15 | 107.7 s | Project tests and 22 external oracle assertions passed |
|
|
61
|
+
| Resume for Chinese README instructions | 7 more | 17.3 s | Same session, next goal revision, implementation/tests byte-identical; tests and oracle still passed |
|
|
62
|
+
|
|
63
|
+
This is one successful run after the documented failures, not a success-rate or
|
|
64
|
+
speed claim. Model calls, planning, validation, and tool use all contribute to
|
|
65
|
+
wall time. No external Codex/Claude/Gemini agent wrapper was called. The existing
|
|
66
|
+
Docker image was reused; no GPU model, GUI desktop, or dependency rebuild was
|
|
67
|
+
needed. The test-created containers exited after each command.
|
|
68
|
+
|
|
69
|
+
The packaged CLI was then installed globally and tested again in a separate fresh
|
|
70
|
+
workspace, using the installed entry point rather than the source checkout. All
|
|
71
|
+
three cases passed again: summary in 6.8 seconds / 4 turns, code repair in 37.1
|
|
72
|
+
seconds / 5 turns, and Chinese follow-up in 16.7 seconds / 4 additional turns.
|
|
73
|
+
The same independent checks passed, including all 22 oracle assertions and
|
|
74
|
+
unchanged input/code/test bytes where required. The installed package also passed
|
|
75
|
+
the 13 focused lifecycle and CLI-policy regressions. These timings describe that
|
|
76
|
+
individual run and are not a throughput guarantee.
|
|
77
|
+
|
|
78
|
+
## Usage
|
|
79
|
+
|
|
80
|
+
See [DeepSeek CLI fallback](../docs/deepseek-cli-fallback.md) for start/resume,
|
|
81
|
+
permissions, model selection, and reproducible checks. No task is automatically
|
|
82
|
+
moved out of Codex when quota expires, and no Codex history is rewritten.
|
|
@@ -6,6 +6,11 @@ import {
|
|
|
6
6
|
} from "../src/execution-worker-server.js";
|
|
7
7
|
|
|
8
8
|
try {
|
|
9
|
+
if (Number.parseInt(process.versions.node, 10) < 22) {
|
|
10
|
+
throw Object.assign(new Error("The execution worker requires Node 22 or later."), {
|
|
11
|
+
code: "EXECUTION_WORKER_NODE_UNSUPPORTED",
|
|
12
|
+
});
|
|
13
|
+
}
|
|
9
14
|
const config = await loadExecutionWorkerServerConfig();
|
|
10
15
|
const runtime = await createProductionExecutionWorkerServer({ config });
|
|
11
16
|
installExecutionWorkerShutdownHandlers(runtime.server);
|
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
# DeepSeek Artifact Completion and Host Integration
|
|
2
|
+
|
|
3
|
+
Machine hosts should invoke both new and resumed AgInTi turns with
|
|
4
|
+
`--no-wrappers` when AgInTi is the fallback for an unavailable external agent.
|
|
5
|
+
Keep the provider explicit; credentials do not authorize a provider switch.
|
|
6
|
+
Use the host's existing routines rather than rewriting them in the agent loop.
|
|
7
|
+
|
|
8
|
+
## Source Changes Are Not Artifact Changes
|
|
9
|
+
|
|
10
|
+
A task that reads inputs and writes a separate summary must not be pushed into
|
|
11
|
+
editing those inputs. A task that creates or updates artifacts inside an exact
|
|
12
|
+
host-provided root must not be pushed into changing the generator's source.
|
|
13
|
+
Private task files deliberately do not increment the project-source revision.
|
|
14
|
+
|
|
15
|
+
Execution and completion now share bounded artifact-root resolution. A real
|
|
16
|
+
current-turn scoped write can satisfy artifact freshness without an unrelated
|
|
17
|
+
project edit. Its saved timestamp must follow the current execution contract;
|
|
18
|
+
old mutations, no-op patches, and workspace-wide roots do not qualify for this
|
|
19
|
+
exemption. Explicit input repairs still require their requested changes.
|
|
20
|
+
After a verified scoped write satisfies a prior missing-file/freshness repair,
|
|
21
|
+
the runtime also stops enforcing that retained repair's obsolete source-edit
|
|
22
|
+
phase. Actual source or artifact quality defects remain blocking. Keeping the
|
|
23
|
+
current contract correct without retiring the obsolete phase can still trap a
|
|
24
|
+
valid routine in repeated no-op patches.
|
|
25
|
+
File, command, requested-format, source-grounding, and quality evidence gates
|
|
26
|
+
remain active. In particular, a pre-existing PDF cannot satisfy a request to
|
|
27
|
+
materially revise and rebuild it.
|
|
28
|
+
|
|
29
|
+
Postfix input declarations such as `requirements.txt is a read-only input`
|
|
30
|
+
are recognized without accidentally excluding a neighboring output path.
|
|
31
|
+
Coordinated preservation instructions such as `leave a.txt and b.txt untouched`
|
|
32
|
+
exclude both inputs from deliverables, not just paths adjacent to the verb.
|
|
33
|
+
An output requested after that preservation clause remains required.
|
|
34
|
+
|
|
35
|
+
## Source-Free Answers
|
|
36
|
+
|
|
37
|
+
Refusing to invent a forecast is not a forecast. Chinese and English negated
|
|
38
|
+
forecast clauses are stripped before assessing whether an answer needs source
|
|
39
|
+
evidence. A real prediction later in the same sentence remains evidence-gated.
|
|
40
|
+
Likewise, saying an item is unverified, pending confirmation, or needs no repeat
|
|
41
|
+
confirmation is not a claim that it has been independently validated. Only the
|
|
42
|
+
negated phrase is removed; real validation claims in the remaining text still
|
|
43
|
+
require evidence.
|
|
44
|
+
|
|
45
|
+
## Validation
|
|
46
|
+
|
|
47
|
+
Run the evidence-visibility, truthful-completion, and scoped-artifact-research
|
|
48
|
+
smokes, then the full `npm test`. The LabCanvas host adds separate mocked
|
|
49
|
+
timeout/session/concurrency tests and opt-in live DeepSeek acceptance:
|
|
50
|
+
|
|
51
|
+
```bash
|
|
52
|
+
python scripts/evaluate_aginti_fallback.py --live --command /path/to/bin/aginti-cli.js
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
The host test uses synthetic inputs, isolated session registries, actual routine
|
|
56
|
+
commands, a compiled Chinese memo PDF, and resumed turns. It prohibits external
|
|
57
|
+
chat sends, public publication, and external agent wrappers. Saved inputs and
|
|
58
|
+
attempted tool writes are checked independently of the model's final answer.
|
|
59
|
+
|
|
60
|
+
This is acceptance for simple host tasks, not a claim that DeepSeek matches
|
|
61
|
+
Codex across all tasks. Keep model/provider limitations separate from agent-loop
|
|
62
|
+
or host-integration failures.
|
|
@@ -0,0 +1,134 @@
|
|
|
1
|
+
# DeepSeek as a practical CLI fallback
|
|
2
|
+
|
|
3
|
+
Use AgInTiFlow for small coding, documentation, and workspace tasks when Codex is
|
|
4
|
+
unavailable. It runs its own tool loop with DeepSeek API credentials; a Codex
|
|
5
|
+
subscription or a working Codex CLI is not required. DeepSeek API usage has its
|
|
6
|
+
own billing and availability.
|
|
7
|
+
|
|
8
|
+
## Start in the project you want to work on
|
|
9
|
+
|
|
10
|
+
Configure credentials once if needed:
|
|
11
|
+
|
|
12
|
+
```bash
|
|
13
|
+
aginti auth deepseek
|
|
14
|
+
```
|
|
15
|
+
|
|
16
|
+
From the project's terminal:
|
|
17
|
+
|
|
18
|
+
```bash
|
|
19
|
+
aginti --provider deepseek --routing fast --no-wrappers
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
Type an ordinary request, such as “Fix the failing tests and explain the change.”
|
|
23
|
+
The default Normal permission mode allows project edits and uses the Docker
|
|
24
|
+
workspace. It does not grant unrestricted host access. To use installed host
|
|
25
|
+
tools for a small, trusted project while keeping the existing command policy:
|
|
26
|
+
|
|
27
|
+
```bash
|
|
28
|
+
aginti --provider deepseek --routing fast --no-wrappers --sandbox-mode host
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
Host mode is not an OS sandbox. Commands still run as your user. Keep the project
|
|
32
|
+
scope narrow; use the Docker workspace when isolation is important. There is no
|
|
33
|
+
need to switch to Danger mode merely to read files, make patches, or run ordinary
|
|
34
|
+
tests. Arbitrary inline interpreter snippets can still stop in restricted host
|
|
35
|
+
mode; the normal Docker workspace is recommended for general coding. New
|
|
36
|
+
processes pick up an installed upgrade; no desktop reboot is needed.
|
|
37
|
+
|
|
38
|
+
One task, followed by a resumable exit:
|
|
39
|
+
|
|
40
|
+
```bash
|
|
41
|
+
aginti run --provider deepseek --routing fast --no-wrappers \
|
|
42
|
+
"Read notes.txt and risks.txt. Write summary.md. Leave the inputs unchanged."
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
Resume from the same project:
|
|
46
|
+
|
|
47
|
+
```bash
|
|
48
|
+
aginti resume
|
|
49
|
+
aginti resume latest --no-wrappers
|
|
50
|
+
aginti resume SESSION_ID --no-wrappers "Continue and run the tests."
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
AgInTi sessions and Codex sessions are separate. These commands resume AgInTi
|
|
54
|
+
history; they do not load a Codex JSONL file. For a quota handoff, leave a short
|
|
55
|
+
project note with the current task, files changed, checks already run, and next
|
|
56
|
+
step, then ask AgInTi to read it. Do not include account tokens or private keys.
|
|
57
|
+
|
|
58
|
+
To change a saved AgInTi session explicitly to DeepSeek Flash:
|
|
59
|
+
|
|
60
|
+
```bash
|
|
61
|
+
aginti resume SESSION_ID --provider deepseek --model deepseek-v4-flash \
|
|
62
|
+
--routing manual --route-provider deepseek --route-model deepseek-v4-flash \
|
|
63
|
+
--main-provider deepseek --main-model deepseek-v4-pro \
|
|
64
|
+
--spare-provider deepseek --spare-model deepseek-v4-pro --no-wrappers
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
Ordinary resume preserves saved runtime choices. An ambient environment change
|
|
68
|
+
does not silently switch the saved account/provider/model. `--no-wrappers`
|
|
69
|
+
explicitly disables Codex/Claude/Gemini-style external agent wrappers, including
|
|
70
|
+
when a saved session or environment previously enabled them. It does not disable
|
|
71
|
+
all shell commands or all separately configured external tools.
|
|
72
|
+
|
|
73
|
+
## Model control and limits
|
|
74
|
+
|
|
75
|
+
- `--routing fast` keeps the fast executor, currently DeepSeek v4 Flash.
|
|
76
|
+
- `--routing manual --model MODEL` keeps that executor model.
|
|
77
|
+
- `--routing smart` retains automatic task routing and can choose Pro.
|
|
78
|
+
- Planning and completion validation remain active when the task requires them.
|
|
79
|
+
Their separately configured model roles may use Pro even with a Flash executor;
|
|
80
|
+
Fast does not mean every auxiliary request uses Flash.
|
|
81
|
+
- Default local-first routing is unchanged. The commands above explicitly choose
|
|
82
|
+
DeepSeek, rather than silently switching an existing local session to a cloud API.
|
|
83
|
+
- Review the diff and test results. A successful sample is not a promise of Codex
|
|
84
|
+
parity, error-free code, or completion of every task without follow-up.
|
|
85
|
+
|
|
86
|
+
## Repairs in 0.20.337-integration.0
|
|
87
|
+
|
|
88
|
+
This integration build includes two earlier, previously undeployed fixes:
|
|
89
|
+
|
|
90
|
+
1. The model compatibility retry shares the original deadline and cancellation
|
|
91
|
+
boundary. A stuck retry cannot keep AgInTi waiting beyond its deadline or
|
|
92
|
+
turn cancellation into a late success.
|
|
93
|
+
2. Reading input files to create a separate output no longer forces mutations
|
|
94
|
+
to those inputs. Positive source-edit requests still require real edits.
|
|
95
|
+
|
|
96
|
+
It also adds:
|
|
97
|
+
|
|
98
|
+
3. Fast/manual executor selection is preserved when SCS planning and evidence
|
|
99
|
+
validation activate. Previously these modes could silently run the main model,
|
|
100
|
+
including a separately configured provider.
|
|
101
|
+
4. `--no-wrappers` works for a new run and as a persistent resume patch. Explicit
|
|
102
|
+
`--allow-wrappers` can still re-enable them.
|
|
103
|
+
5. Syntactically read-only `find` without a depth bound remains **blocked**, but
|
|
104
|
+
returns a recoverable discovery error. The agent can choose structured file
|
|
105
|
+
tools or add a bounded `-maxdepth`, without asking for destructive permission.
|
|
106
|
+
That denial survives pipelines and command sequences. `-delete`, `-exec`,
|
|
107
|
+
file output, unknown shell segments, and other writes remain under the existing
|
|
108
|
+
stronger policies. No permission is silently granted and no denied command runs.
|
|
109
|
+
|
|
110
|
+
## Reproduce the live acceptance check
|
|
111
|
+
|
|
112
|
+
This is opt-in, spends DeepSeek credits, and creates a new synthetic workspace.
|
|
113
|
+
It does not point the agent at your current project's source files.
|
|
114
|
+
|
|
115
|
+
```bash
|
|
116
|
+
AGINTIFLOW_REAL_DEEPSEEK=1 npm run eval:deepseek-cli
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
Optional: set `AGINTIFLOW_EVAL_ROOT` to a directory for retained private evidence,
|
|
120
|
+
or `AGINTIFLOW_EVAL_CLI` to an installed `bin/aginti-cli.js` to verify that package.
|
|
121
|
+
The default evaluation location is a fresh temporary directory. The evaluator
|
|
122
|
+
uses the default Normal Docker workspace; set `AGINTIFLOW_EVAL_SANDBOX=host` to
|
|
123
|
+
test the more restricted host policy instead. That is a distinct acceptance
|
|
124
|
+
configuration, and a blocked host run is not counted as a pass.
|
|
125
|
+
|
|
126
|
+
The evaluator checks a two-source summary, a small code repair with real tests,
|
|
127
|
+
and a Chinese README follow-up in the same session. Independent checks compare
|
|
128
|
+
input files, run tests again, exercise an external decimal-rounding oracle, verify
|
|
129
|
+
unchanged implementation/tests during the documentation follow-up, and inspect
|
|
130
|
+
actual model/tool events. Model summaries alone do not establish a pass. A failed
|
|
131
|
+
case produces a nonzero exit and a retained report; it is not relabeled successful.
|
|
132
|
+
|
|
133
|
+
See [the dated acceptance record](../aginti-work-examples/deepseek-cli-fallback-20261001.md)
|
|
134
|
+
for measured outcomes, failures encountered, and release/install details.
|
|
@@ -0,0 +1,213 @@
|
|
|
1
|
+
# General Agent Backend Research
|
|
2
|
+
|
|
3
|
+
Date: 2026-09-27. This is a source-based assessment plus an offline runtime
|
|
4
|
+
repair, not a production deployment or a live model-quality benchmark.
|
|
5
|
+
|
|
6
|
+
## Baseline and Scope
|
|
7
|
+
|
|
8
|
+
- AgInTiFlow: `bf9c3e0386f86502246e776d25dcf29d43d89216`,
|
|
9
|
+
`integration/deepseek-analysis-20260910`, `0.20.336-integration.4`.
|
|
10
|
+
- Implementation branch: `fix/provider-request-lifecycle-20260927`, in
|
|
11
|
+
`AgInTiFlow-worktrees/AgInTiFlow-runtime-responsiveness`.
|
|
12
|
+
- AgenticApp reference HEAD: `7465ec7eece9ef92111ab28284fd7d54c2b2c8be`.
|
|
13
|
+
Its current source and handoffs were read; its unrelated dirty work was preserved.
|
|
14
|
+
- EchoMind reference HEAD: `190aeb3325dec6f5d31c56be80ed4f0b35d86f13`.
|
|
15
|
+
Its provider factory, conversation, voice, and single-flight code were read.
|
|
16
|
+
- No private conversation bodies, credentials, or model weights were required.
|
|
17
|
+
No application services, GPU jobs, provider defaults, or public routes were changed.
|
|
18
|
+
|
|
19
|
+
The existing SQLite campaign was opened read-only. It contains 69 capability
|
|
20
|
+
rows (6 passed, 63 passed-after-fix) and 114 test rows (8 passed, 104
|
|
21
|
+
passed-after-fix, 2 historical failures). Its latest recorded scenarios use
|
|
22
|
+
0.20.331-era versions. This is substantial regression evidence, but does not
|
|
23
|
+
establish current integration performance or universal task coverage. The
|
|
24
|
+
new request-lifecycle regression is separate evidence, not a rewritten old pass.
|
|
25
|
+
|
|
26
|
+
## What Already Works Architecturally
|
|
27
|
+
|
|
28
|
+
AgenticApp's `src/agenticapp/workspace_agent.py::run_aginti_turn` locks its
|
|
29
|
+
conversation registry, selects a persistent AgInTi session, and invokes the
|
|
30
|
+
machine CLI. `_run_aginti_provider_chain` preserves that session when changing
|
|
31
|
+
provider. `_parse_aginti_machine_result` treats stopped, failed, or unresolved
|
|
32
|
+
tool-protocol output as failure. The host registers actual artifact files;
|
|
33
|
+
an assistant mentioning a file is insufficient.
|
|
34
|
+
|
|
35
|
+
Its `references/aginti-primary-labcanvas-agent-handoff-2026-08-18.md` explicitly
|
|
36
|
+
assigns mature domain routines to their owning applications. This is the right
|
|
37
|
+
general boundary: AgInTi interprets, chooses tools, retains execution context,
|
|
38
|
+
and verifies completion. Applications own account policy, domain APIs, delivery,
|
|
39
|
+
and authorization of irreversible actions. Their chat names, schedules, and
|
|
40
|
+
private transport details should not enter AgInTi's core.
|
|
41
|
+
|
|
42
|
+
AgInTi already has mechanisms worth retaining:
|
|
43
|
+
|
|
44
|
+
| Need | Existing implementation | Validation used in this slice |
|
|
45
|
+
| --- | --- | --- |
|
|
46
|
+
| Provider handoff without restarting the task | `src/provider-handoff.js`, `src/agent-runner.js` | `smoke:provider-handoff` |
|
|
47
|
+
| Provider-aware context budgets | `src/context-budget-controller.js` | `smoke:context-budget-recovery` |
|
|
48
|
+
| Bounded tool disclosure and convergence | `src/progressive-tool-selection.js` | `eval:local-first-agent` |
|
|
49
|
+
| Persistent session, inbox, and crash replay | `src/session-store.js`, `src/session-runtime.js` | `smoke:session-runtime`, `smoke:runtime-core`, `smoke:inbox` |
|
|
50
|
+
| Evidence-backed completion | `src/scs-evidence.js`, `src/agent-runner.js` | `smoke:truthful-completion` |
|
|
51
|
+
| Scoped research, computation, documents, and artifacts | `src/integration-analysis-planner.js`, `src/integration-analysis-session-service.js` | `smoke:integration-analysis-planner` |
|
|
52
|
+
| Server-owned public authority | `src/integration-api.js`, `src/integration-auth.js`, `src/integration-policy.js` | Retain existing readiness and ownership gates |
|
|
53
|
+
|
|
54
|
+
The public analysis planner and the broad workspace runtime have different tool
|
|
55
|
+
surfaces. EchoMind must consume the declared public capability surface; it must
|
|
56
|
+
not assume that a CLI tool available on a developer workstation is a public
|
|
57
|
+
application capability.
|
|
58
|
+
|
|
59
|
+
## Upstream Comparison
|
|
60
|
+
|
|
61
|
+
Existing clean reference checkouts were updated with fast-forward-only pulls,
|
|
62
|
+
with Git hooks and recursive submodule updates disabled. No upstream code was
|
|
63
|
+
copied into AgInTi and no upstream dependency installation was performed.
|
|
64
|
+
|
|
65
|
+
| Reference | Inspected revision | Relevant source and lesson |
|
|
66
|
+
| --- | --- | --- |
|
|
67
|
+
| OpenAI Codex | `41f9084b30812db321a0b592def4f500d1e79cf4` | `codex-rs/core/src/client.rs`, cancellation tokens and consumer-drop handling; `codex-rs/history/src/compaction_checkpoint.rs`, durable context boundaries |
|
|
68
|
+
| Claude Code | `7779afb12e3635f46f56ec823979d68350ae000b` | Public plugins and hook examples only; this checkout is not evidence of its proprietary runtime implementation |
|
|
69
|
+
| Gemini CLI | `2fe7c2d3f065dc40ad573d50b2091116f8a4aa18` | `packages/core/src/utils/retry.ts`, abort-aware retries; `packages/core/src/telemetry/`, stage-specific observations |
|
|
70
|
+
| Qwen Code | `36710ff6908c96a4b95db59c2f46c2b49278d02e` | `packages/core/src/utils/retry.ts`, cancellable delay and explicit retry policy |
|
|
71
|
+
| GitHub Copilot SDK | `d106d29dc6c5112da2abdae59008571b6692f12b` | `nodejs/src/session.ts::awaitWorkflowOperation`, guard before dispatch and abort race; `docs/features/session-persistence.md`, explicit session lifecycle |
|
|
72
|
+
| DeepSeek Harness | `99f6f02fecdb7dff40c3fbc9470f5907c29f74ca` | Pull timed out twice; existing `docs/agent-lifecycle.zh.md` describes durable session events separately from live coordination. This is a dated reference, not a verified current upstream snapshot. |
|
|
73
|
+
|
|
74
|
+
The reusable lesson is explicit lifecycle control, not more agents per task.
|
|
75
|
+
Anthropic's [context-engineering guidance](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents)
|
|
76
|
+
supports selective retrieval, small clear tool surfaces, and preserving important
|
|
77
|
+
state during compaction. AgInTi already implements parts of this; improve their
|
|
78
|
+
measured behavior before adding more prompts or parallel workers.
|
|
79
|
+
|
|
80
|
+
[Gemini's telemetry documentation](https://geminicli.com/docs/cli/telemetry/)
|
|
81
|
+
provides a useful reference for separating model, tool, and agent observations.
|
|
82
|
+
Apply that idea with private-content logging disabled, rather than copying a
|
|
83
|
+
telemetry service or transmitting user data.
|
|
84
|
+
|
|
85
|
+
## Reproduced Defect and Implemented Repair
|
|
86
|
+
|
|
87
|
+
`src/model-client.js::createChatCompletion` bounded the first SDK call using
|
|
88
|
+
`Promise.race`, but awaited its unsupported-`reasoning_effort` compatibility
|
|
89
|
+
retry directly. The deadline still aborted a signal, yet the caller could remain
|
|
90
|
+
waiting and accept a late success. Caller cancellation similarly depended on
|
|
91
|
+
the client settling itself, rather than ending AgInTi's wait.
|
|
92
|
+
|
|
93
|
+
Before editing, a controlled client reproduced a retry still pending 80 ms after
|
|
94
|
+
a 25 ms deadline and then returning success. Eight deterministic tests were
|
|
95
|
+
added; six failed on the original source. A real OpenAI SDK instance with an
|
|
96
|
+
offline fetch implementation also reproduced the problem when the compatibility
|
|
97
|
+
retry entered a 429 Retry-After backoff. No provider request was transmitted.
|
|
98
|
+
|
|
99
|
+
The repair uses one shared deadline and cancellation outcome for both attempts:
|
|
100
|
+
|
|
101
|
+
- Already-cancelled work makes zero client calls.
|
|
102
|
+
- Caller cancellation ends the wait and sends abort to the transport.
|
|
103
|
+
- Both attempts share the original deadline; retry does not restart the clock.
|
|
104
|
+
- Late resolution cannot replace cancellation or timeout with success.
|
|
105
|
+
- Cancellation cannot trigger another compatibility attempt.
|
|
106
|
+
- Timers and parent listeners are removed after completion or failure.
|
|
107
|
+
- The optional compatibility retry, normal payloads, and provider-error metadata
|
|
108
|
+
remain intact. Existing timeout classification can still drive permitted
|
|
109
|
+
same-session handoff.
|
|
110
|
+
|
|
111
|
+
This improves a demonstrated waiting failure. It does not make the underlying
|
|
112
|
+
model generate tokens faster, guarantee a remote provider stopped computing,
|
|
113
|
+
or eliminate an SDK-owned backoff timer. The latter is drained in the regression
|
|
114
|
+
and cannot issue another fetch after abort. It also does not change the separate
|
|
115
|
+
public analysis planner's own model-request lifecycle.
|
|
116
|
+
|
|
117
|
+
## Smallest EchoMind Integration
|
|
118
|
+
|
|
119
|
+
The inspected EchoMind `EchoMind/echomind/ai_client_factory.py` and
|
|
120
|
+
`mixed_ai_request.py` create provider clients. `voice_processor.py::get_ai_response`
|
|
121
|
+
uses schema-constrained responses, conversation context, and single-flight
|
|
122
|
+
deduplication; `process_audio` calls it through an executor. This is not yet a
|
|
123
|
+
durable AgInTi task interface, and replacing a provider URL cannot provide one.
|
|
124
|
+
|
|
125
|
+
Recommended boundary for a subsequent, separately tested integration:
|
|
126
|
+
|
|
127
|
+
```mermaid
|
|
128
|
+
flowchart LR
|
|
129
|
+
UI[EchoMind chat or voice] --> BFF[Authenticated EchoMind backend]
|
|
130
|
+
BFF --> API[AgInTi versioned session and run API]
|
|
131
|
+
API --> Runtime[AgInTi policy, context, tools, evidence]
|
|
132
|
+
Runtime --> Workers[Admitted execution and artifact workers]
|
|
133
|
+
Runtime --> Models[Explicit permitted provider routes]
|
|
134
|
+
API --> Events[Durable public events and verified artifacts]
|
|
135
|
+
Events --> BFF
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
Bind each user/conversation to an owned thread, preserve the same run/session
|
|
139
|
+
through reconnect and permitted provider handoff, and use mutation idempotency.
|
|
140
|
+
Use persisted run events for progress and replay. A stop request should cancel
|
|
141
|
+
the scoped run; a disconnected phone should not create a second task. Keep
|
|
142
|
+
speech recognition, TTS, language settings, notifications, and social features
|
|
143
|
+
in EchoMind. Add capabilities incrementally: text-only inference, source-backed
|
|
144
|
+
research, then admitted file/computation work.
|
|
145
|
+
|
|
146
|
+
`safe-chat` is a bounded stateless response service, not a substitute for this
|
|
147
|
+
session API. Its prose documentation predates the provider-general code, so
|
|
148
|
+
effective source configuration and capability probes must take precedence over
|
|
149
|
+
old DeepSeek-only examples. No EchoMind wiring or deployment was attempted here.
|
|
150
|
+
|
|
151
|
+
## Prioritized Next Work
|
|
152
|
+
|
|
153
|
+
1. **Measure latency without prompts in telemetry.** Correlate queue admission,
|
|
154
|
+
readiness, model attempts, retries, tools, validation, and terminal commit.
|
|
155
|
+
Existing `model.requested` events do not by themselves describe every SDK
|
|
156
|
+
attempt or first-token delay. Track p50/p95, request/tool counts, and repair
|
|
157
|
+
count per task class before claiming speed improvements.
|
|
158
|
+
2. **Turn routine reuse into explicit capability contracts.** Reuse the existing
|
|
159
|
+
skill/profile/tool mechanisms. Domain owners provide bounded inputs, scope,
|
|
160
|
+
side-effect class, cancellation, output manifest, and verifier. The runtime
|
|
161
|
+
should not rediscover an entire repository after receiving an authoritative
|
|
162
|
+
routine that already answers the request.
|
|
163
|
+
3. **Test the application adapter under real lifecycle failures.** Phone
|
|
164
|
+
reconnect, duplicate submit, cancellation, provider outage, worker restart,
|
|
165
|
+
ownership conflicts, and response loss must preserve one durable task and
|
|
166
|
+
verified outputs. Never replay a side-effecting tool turn as a fresh prompt.
|
|
167
|
+
4. **Make perceived progress honest.** Expose stage and verified tool progress
|
|
168
|
+
early. If draft text is streamed later, mark it provisional and retain the
|
|
169
|
+
existing evidence gates before accepting a final answer or artifact.
|
|
170
|
+
5. **Expand measured coverage, not permissions.** Reuse narrow application
|
|
171
|
+
routines for CAD, media, research, and documents. Unsupported host actions
|
|
172
|
+
should remain explicit capability limits. Worker readiness and optional-role
|
|
173
|
+
degradation must remain truthful.
|
|
174
|
+
|
|
175
|
+
Suggested acceptance prompts are ordinary, imperfect requests: summarize notes
|
|
176
|
+
without writing files; correct a prior answer in the same thread; inspect an
|
|
177
|
+
already-specified read-only routine; research a topic using both official web
|
|
178
|
+
and paper sources; calculate a result and create exactly the requested files;
|
|
179
|
+
cancel while a provider is retrying; reconnect after artifact commit without
|
|
180
|
+
repeating the work. Verify files and durable state from outside the agent.
|
|
181
|
+
|
|
182
|
+
No current live DeepSeek/LocalLLM latency or quality claim is made. Broad
|
|
183
|
+
production readiness still requires these app-level canaries and configured
|
|
184
|
+
capability proofs. A passing offline regression is evidence for its exact
|
|
185
|
+
contract, not a claim that all repositories or tasks now work perfectly.
|
|
186
|
+
|
|
187
|
+
## Validation Record
|
|
188
|
+
|
|
189
|
+
- `node --test test/model-request-lifecycle.test.js`: 8 passed. Before the
|
|
190
|
+
repair, the same file produced 6 failures and 2 passes.
|
|
191
|
+
- `npm run smoke:model-roles`: passed, including the new regression file.
|
|
192
|
+
- `npm run smoke:provider-handoff`: passed with persisted mock-provider runs.
|
|
193
|
+
- Separate focused passes: `smoke:local-failure-recovery`,
|
|
194
|
+
`smoke:context-budget-recovery`, `smoke:session-runtime`, `smoke:runtime-core`,
|
|
195
|
+
`smoke:inbox`, `smoke:truthful-completion`, and
|
|
196
|
+
`smoke:integration-analysis-planner`.
|
|
197
|
+
- `npm run eval:local-first-agent`: 18 passed, 0 failed, 0 skipped;
|
|
198
|
+
its offline guard observed zero network attempts.
|
|
199
|
+
- `npm run check`: 300 JavaScript files passed. The new test file also passed
|
|
200
|
+
its explicit `node --check`.
|
|
201
|
+
- `AGINTIFLOW_PROVIDER_ATTRIBUTION_LIVE=0 npm test`: exit 0, including pretest.
|
|
202
|
+
The document-worker-server and document-worker-cross-boundary scripts
|
|
203
|
+
reported occupied-port skips for `127.0.0.1:18102`; its existing listener was
|
|
204
|
+
left untouched. These two HTTP checks are not claimed as executed passes.
|
|
205
|
+
- Dry-run npm packaging includes the new regression and excludes private
|
|
206
|
+
session, credential, database, and dependency paths in the inspected file list.
|
|
207
|
+
- `git diff --check`: passed.
|
|
208
|
+
|
|
209
|
+
Tests reused the baseline dependency installation through a temporary symlink;
|
|
210
|
+
no dependencies were installed or upgraded. Only task-owned smoke autostart
|
|
211
|
+
servers were stopped during cleanup. Package version remains
|
|
212
|
+
`0.20.336-integration.4`; no release, deployment, installed-package change, or
|
|
213
|
+
EchoMind/AgenticApp source change was performed.
|