codex-adaptive-effort 0.1.0-alpha.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,34 @@
1
+ # Known limitations and acceptance boundaries
2
+
3
+ The dated [local acceptance record](LOCAL_ACCEPTANCE.md) describes the tested macOS CLI subset; [desktop acceptance](DESKTOP_ACCEPTANCE.md) covers an independent desktop instance on that installation. Neither establishes packaged desktop integration or general production readiness.
4
+
5
+ | Area | What is true now | What is not established |
6
+ |---|---|---|
7
+ | Code delivery | Runnable original Node ESM source, tests, offline demo | No claim of maintained production service |
8
+ | Proxy | Synthetic HTTP/SSE tests plus real CLI and isolated desktop ChatGPT-route acceptance | No WebSocket acceptance |
9
+ | Codex | Native model/list, isolated CLI, desktop plain-text off/locks/cancellation/recovery; CLI tools in off tested on one installation | Other client versions, packaged desktop installation, ordinary ChatGPT chats |
10
+ | ChatGPT auth | Native authentication worked for the recorded CLI run; CAE does not read login files | Other accounts, environments, attestation and client versions |
11
+ | API auth | Explicit separate api route; normal API key remains caller-owned | A subscription is not API credit; no live API request was made |
12
+ | Jev | Eight evaluator-only fixture calls and a second three-case desktop shadow run succeeded; two earlier desktop timeouts remain unexplained. Desktop Jev defaults to enforced off/shadow; a separate process-only auto opt-in retains the 8-call cap and default 1500 ms ceiling; an explicit process-only timeout parameter allows a 2000 ms experiment. Real desktop medium-to-high, medium-to-low and timeout fallback requests completed successfully; the timeout last observed request creation before any send-start event. See [shadow](JEV_SHADOW_ACCEPTANCE.md) and [auto](LOCAL_ACCEPTANCE.md) records | Benefit of the 2000 ms experiment (the successful downshift took only 665 ms), cause of pre-send timeout, reliable completion within the 1500 ms budget, representative Chinese classification quality, reliable uncertainty handling and task outcome quality |
13
+ | Model efforts | Supplied by actual capability probe or operator | No hardcoded assurance that any model supports all known effort names |
14
+ | Cache | Cache controls/history preserved; request-level effort only | No configuration_update insertion, cached-prefix preservation or measured savings |
15
+ | Context | Limited Unicode character excerpts and omission counts | Not token-exact, not all long-history requirements, no image/audio understanding |
16
+ | Errors | Explicit structured tool failure fields recognized | A plain text log can hide failure; content classification remains Jev's job |
17
+ | Bypass | Unsupported requests retain original bytes; structured tool-result history bypass was observed in a real session and is an accepted current boundary | Not every native history shape is tested; manual locks do not override the bypass |
18
+ | Sessions | Explicit header plus auth partition; otherwise no cross-call lease | Cannot infer a reliable session from cache keys or message similarity |
19
+ | Completion | Recognized SSE/JSON terminal metadata | A setting in a response does not prove actual reasoning allocation |
20
+ | Usage | Known counters and unknowns separated; evaluator attempts recorded | Dollar savings, account quota debits, quality equivalence and full-task costs |
21
+ | Other costs | Generation endpoint observed | Compaction costs, invisible retries, incomplete usage and external provider debits may be missing |
22
+ | Security | Local auth/Host/origin checks, bounded state, metadata logs, no redirects | Not resistant to a malicious same-user process; redaction is not total confidentiality |
23
+ | Recovery | No global config writes; normal Codex relaunch restores ordinary path | Proxy crash is not automatic failover |
24
+ | Publication | Existing public repository; baseline CI success verified | Every new commit needs its own CI result; never reuse an older commit's result |
25
+
26
+ Manual lock is instance-wide and only changes eligible `auto` requests. `off` and `shadow` never change effort, even with a lock set. To retain original execution without Jev cost use `off`; to remove the entire proxy from the path restart ordinary Codex.
27
+
28
+ History eligibility is conservative: recognized text messages, opaque reasoning/call records and string tool outputs are supported. Unknown item/content types, malformed text and non-string tool outputs bypass unchanged with no evaluator call; a manual lock does not override this boundary. See the [code review](CODE_REVIEW_2026-09-28.md).
29
+
30
+ The [desktop launcher](DESKTOP_LAUNCHER.md) is limited to the validated macOS arm64 app/CLI combination. It shares the native Codex home, does not migrate old threads, and requires a new local thread to use the configured provider. Normal stop and terminal interruption are tested; forced supervisor death and stale socket recovery require manual inspection. No native UI title patch or packaged plugin is provided.
31
+
32
+ A server process lifetime cap is not a spending guarantee. Services may bill requests that time out or are cancelled. No trial uses a real key unless the operator explicitly enables it.
33
+
34
+ Release blockers for any stable claim: independent security review; real native CLI acceptance for both intended auth paths; packaged desktop integration and upgrade compatibility; long-session/compaction acceptance; low/normal/hard task comparisons including Chinese short follow-ups; actual quality, cache, latency and cost observations. Until then, treat this as a protocol-level alpha.
@@ -0,0 +1,83 @@
1
+ # CLI and native validation guide
2
+
3
+ Validate in a separate process and a directory containing only synthetic, non-sensitive tasks. Keep an ordinary Codex session available. Never read/copy native login files or change global `~/.codex/config.toml`, approval policies or sandbox settings. For the validated desktop version, use the [desktop guide](DESKTOP_LAUNCHER.md).
4
+
5
+ ## 1. Offline checks and native metadata
6
+
7
+ ```bash
8
+ npm ci --ignore-scripts
9
+ npm run verify
10
+ node bin/cae.mjs doctor
11
+ node bin/cae.mjs probe > capabilities.local.json
12
+ ```
13
+
14
+ The tests and demo do not use real providers. `doctor` queries the native version; `probe` initializes app-server and queries `model/list` without generating output. Native Codex may consult online model metadata. If the CLI is absent from `PATH`, pass `--codex /path/to/trusted/codex` to doctor/probe and subsequent CLI launches. Do not replace or upgrade the desktop binary to make a check pass.
15
+
16
+ Record the source commit, time, OS/architecture, Node version, trusted CLI path/version, exit codes and actual model/effort set. Keep the raw capture local. A failed probe is not permission to substitute synthetic models or scrape credentials.
17
+
18
+ ## 2. Initialize an isolated configuration
19
+
20
+ Choose a real model from the capture. `$MODEL_ID` below means the exact value you selected:
21
+
22
+ ```bash
23
+ node bin/cae.mjs init --model "$MODEL_ID" --auth chatgpt --capabilities capabilities.local.json
24
+ node bin/cae.mjs launch-args --auth chatgpt
25
+ ```
26
+
27
+ Initialization refuses to overwrite `.cae`. For another experiment use `--dir PATH` and pass `--config PATH/config.json` to later commands. CAE's directory is not a replacement Codex home. Native Codex continues to handle login.
28
+
29
+ Preserve the existing auth route. ChatGPT subscription use requires `--auth chatgpt`. API use requires an explicitly selected `--auth api` configuration and the caller's normal `OPENAI_API_KEY`; it may be billed separately. Never silently fall back from subscription to API.
30
+
31
+ Review the printed non-secret arguments: fixed model, local Responses URL, auth route and environment variable names, with no approval/sandbox overrides. Model IDs and effort values are environment-specific.
32
+
33
+ ## 3. Off transport trial
34
+
35
+ Real requests need explicit authorization. Set `mode` to `off` in the isolated CAE configuration before starting. Leave normal Codex settings unchanged.
36
+
37
+ ```bash
38
+ # Terminal A: foreground proxy
39
+ node bin/cae.mjs serve --enable-upstream
40
+ # Terminal B: experimental CLI, retaining native authentication
41
+ node bin/cae.mjs codex --auth chatgpt --
42
+ ```
43
+
44
+ Submit a minimal synthetic text task, then test a read-only tool in the synthetic directory, a follow-up, cancellation and recovery. Off disables evaluation/changes, not model usage or network forwarding. Record errors such as 401/403 or protocol failures; do not export cookies, change billing routes or hide retries.
45
+
46
+ ## 4. Manual effort trial without Jev
47
+
48
+ Use efforts confirmed by the probe. The following values are examples, not universal capabilities:
49
+
50
+ ```bash
51
+ node bin/cae.mjs lock low
52
+ node bin/cae.mjs control auto
53
+ # Send an eligible synthetic text task and wait for completion.
54
+ node bin/cae.mjs lock high
55
+ # Send another eligible task and wait for completion.
56
+ node bin/cae.mjs control off
57
+ node bin/cae.mjs unlock
58
+ node bin/cae.mjs report
59
+ ```
60
+
61
+ Correlate prepared/sent/completed events. Unsupported shapes, including structured tool-result history, can bypass even a manual lock. Do not remove history to force eligibility. A request sent with a value does not prove the model's actual reasoning allocation.
62
+
63
+ ## 5. Jev shadow, then auto
64
+
65
+ Jev sends bounded task text to TypeSafe and may incur usage. Enable it only with explicit authorization for that external processing and the normal independent key.
66
+
67
+ For a standalone serve process: stop the proxy, set `judge.kind` to `typesafe` and `mode` to `shadow` in the isolated configuration, supply `TYPESAFE_API_KEY` through your normal environment setup, then run:
68
+
69
+ ```bash
70
+ node bin/cae.mjs serve --enable-upstream --enable-jev
71
+ ```
72
+
73
+ Use a small configured `judge.maxCalls` budget. First verify shadow recommendations and fallback. With separate authorization to change effort, use `control auto`. The [desktop launcher](DESKTOP_LAUNCHER.md) has its own process-only flags, eight-call cap and timeout experiment; do not assume standalone serve uses those desktop caps.
74
+
75
+ The evaluator-only [fixture runner](JEV_SHADOW_ACCEPTANCE.md) provides a separate bounded protocol check without executor generation. Neither fixtures nor a few successful requests establish representative accuracy or savings.
76
+
77
+ ## 6. Record evidence and recover
78
+
79
+ Use the [acceptance template](acceptance-template.md). Separate local tests, CI, native metadata, actual sends, completion, UI observations and task quality. Record unknown usage as unknown. For comparative evaluation, control task content, starting state, context, model, mode, versions and timing; a single later success is not an A/B result.
80
+
81
+ End the experimental CLI, stop the foreground proxy, then relaunch normal Codex. Desktop trials use `desktop stop`. `control off` alone does not remove the proxy. Do not rewrite login files as a recovery step.
82
+
83
+ Never commit `.cae`, captures, raw logs, credentials or private task histories. Public evidence must be a reviewed, sanitized summary.
package/docs/NPM.md ADDED
@@ -0,0 +1,73 @@
1
+ # Install and run the npm package
2
+
3
+ Requires Node.js **22.16+** and npm. Native integration also requires your existing Codex installation and normal login. The package contains original CAE runtime code and selected guides, with no runtime dependencies, install hooks or bundled native binaries.
4
+
5
+ ## Package availability
6
+
7
+ This revision prepares `codex-adaptive-effort@0.1.0-alpha.1` for npm distribution. **It has not been published to the npm registry by this change.** Do not infer registry availability from a source version or a successful package test.
8
+
9
+ Until the first registry release, install a reviewed Git commit with npm (requires Git):
10
+
11
+ ```bash
12
+ # Replace REVIEWED_COMMIT with the full tested Git commit SHA.
13
+ npm install --global --ignore-scripts github:ppxu/codex-adaptive-effort#REVIEWED_COMMIT
14
+ cae --version
15
+ cae --help
16
+ ```
17
+
18
+ After a maintainer publishes the alpha release to npm, the shorter installation command will be:
19
+
20
+ ```bash
21
+ npm install --global --ignore-scripts codex-adaptive-effort@alpha
22
+ ```
23
+
24
+ No administrator privileges should be necessary with a user-owned Node installation. If `cae` is not found, add npm's global executable directory to your `PATH` (`npm prefix --global`: the `bin` subdirectory on macOS/Linux, or the prefix itself on Windows). Do not put credentials in command arguments.
25
+
26
+ ## Use a dedicated working directory
27
+
28
+ Global installation makes the command available from any directory; configuration remains local to the directory you choose. It does not create or modify `~/.codex/config.toml`. Keep the same working directory in each terminal, or pass the same explicit `--config PATH`.
29
+
30
+ ```bash
31
+ mkdir cae-trial
32
+ cd cae-trial
33
+ cae doctor
34
+ cae probe > capabilities.local.json
35
+ ```
36
+
37
+ If the native CLI is not on `PATH`, use `--codex /path/to/trusted/codex` for doctor/probe. Select `MODEL_ID` from the actual capability capture. Do not use a model ID copied from a fixture.
38
+
39
+ For the [validated desktop version](DESKTOP_LAUNCHER.md):
40
+
41
+ ```bash
42
+ cae desktop start --model "$MODEL_ID" --auth chatgpt --enable-upstream
43
+ # In a second terminal, from the same cae-trial directory:
44
+ cae desktop status
45
+ cae desktop stop
46
+ ```
47
+
48
+ The first launch creates `.cae/desktop/config.json`. Later starts can omit `--model` and reuse that configuration. Startup is foreground and still enforces native signature/version/capability/provider checks. Use a new local chat in the experimental window. Default mode is shadow with the baseline evaluator; Jev remains separately opt-in. Starting the launcher does not submit a task. Sending a task uses the normal model allowance.
49
+
50
+ For CLI integration, follow [local validation](LOCAL_VALIDATION.md), replacing `node bin/cae.mjs` with `cae`. The [desktop guide](DESKTOP_LAUNCHER.md) uses the same substitution for manual controls, Jev opt-in and recovery. No repository checkout is needed after installation.
51
+
52
+ ## Upgrade, remove or install a local archive
53
+
54
+ Stop the experimental instance before replacing the package: generated desktop launchers reference the installed package path. After an authorized registry release, upgrade with the same `npm install --global ...@alpha` command. A CAE upgrade does not widen native desktop compatibility or change stored configuration.
55
+
56
+ ```bash
57
+ cae desktop stop
58
+ npm uninstall --global codex-adaptive-effort
59
+ ```
60
+
61
+ Removal leaves your project-local `.cae` directory intact. Stop the experiment and return to ordinary Codex to restore the normal route. Never remove native login files.
62
+
63
+ From a source checkout, a local archive can be built and installed without any registry publication:
64
+
65
+ ```bash
66
+ npm ci --ignore-scripts
67
+ npm run verify
68
+ npm pack --ignore-scripts
69
+ npm install --global --ignore-scripts ./codex-adaptive-effort-0.1.0-alpha.1.tgz
70
+ cae --version
71
+ ```
72
+
73
+ The explicit `package.json` file list excludes `.cae`, captures, dotenv, credentials, logs, source manifests, tests and maintenance scripts. Archives are ignored by Git. The automated package test builds and installs an archive into a temporary prefix offline, checks the actual command shim, version/help, synthetic initialization and launch arguments, and verifies that desktop startup still refuses missing authorization before native inspection. It never installs globally into the user's prefix or starts a real model.
package/package.json ADDED
@@ -0,0 +1,56 @@
1
+ {
2
+ "name": "codex-adaptive-effort",
3
+ "version": "0.1.0-alpha.1",
4
+ "description": "A fixed-model, opt-in reasoning-effort controller for local Codex Responses traffic.",
5
+ "type": "module",
6
+ "publishConfig": {"access": "public", "registry": "https://registry.npmjs.org/", "tag": "alpha"},
7
+ "files": [
8
+ "bin/cae.mjs",
9
+ "bin/cae-desktop-bridge.mjs",
10
+ "src/audit.mjs",
11
+ "src/codex.mjs",
12
+ "src/config.mjs",
13
+ "src/context.mjs",
14
+ "src/controller.mjs",
15
+ "src/desktop.mjs",
16
+ "src/judge-timing.mjs",
17
+ "src/judge.mjs",
18
+ "src/proxy.mjs",
19
+ "src/stream.mjs",
20
+ "src/util.mjs",
21
+ "src/version.mjs",
22
+ "README.en.md",
23
+ "README.zh-CN.md",
24
+ "LICENSE",
25
+ "THIRD_PARTY_NOTICES.md",
26
+ "docs/NPM.md",
27
+ "docs/DESKTOP_LAUNCHER.md",
28
+ "docs/LOCAL_VALIDATION.md",
29
+ "docs/LIMITATIONS.md"
30
+ ],
31
+ "engines": {"node": ">=22.16.0"},
32
+ "bin": {"cae": "./bin/cae.mjs"},
33
+ "scripts": {
34
+ "check": "node scripts/check.mjs",
35
+ "test": "node --test test/*.test.mjs",
36
+ "coverage": "node --experimental-test-coverage --test test/*.test.mjs",
37
+ "demo": "node scripts/demo.mjs",
38
+ "verify": "npm run check && npm test && npm run demo"
39
+ },
40
+ "license": "MIT",
41
+ "repository": {
42
+ "type": "git",
43
+ "url": "git+https://github.com/ppxu/codex-adaptive-effort.git"
44
+ },
45
+ "homepage": "https://github.com/ppxu/codex-adaptive-effort#readme",
46
+ "bugs": {
47
+ "url": "https://github.com/ppxu/codex-adaptive-effort/issues"
48
+ },
49
+ "keywords": [
50
+ "codex",
51
+ "reasoning-effort",
52
+ "local-proxy",
53
+ "jev",
54
+ "nodejs"
55
+ ]
56
+ }
package/src/audit.mjs ADDED
@@ -0,0 +1,71 @@
1
+ import { openSync, writeSync, closeSync, constants, fchmodSync } from 'node:fs';
2
+ import { CaeError, knownNumber, isObject } from './util.mjs';
3
+
4
+ const FIELDS = new Set(['event', 'requestId', 'model', 'mode', 'effort', 'source', 'reason', 'revision',
5
+ 'proposedEffort', 'incomingEffort', 'changed', 'lease', 'confidence', 'judgeMs', 'judgeInputTokens',
6
+ 'judgeKind', 'stateChars', 'httpStatus', 'terminal', 'inputTokens', 'cachedInputTokens', 'outputTokens',
7
+ 'reasoningTokens', 'durationMs', 'completed']);
8
+ const TIMING_FIELDS = ['judgeRequestCreatedMs', 'judgeSendStartMs', 'judgeRequestSentMs',
9
+ 'judgeResponseHeadersMs', 'judgeResponseBodyMs', 'judgeValidatedMs', 'judgeObservedMs'];
10
+ const TIMING_STAGES = new Set(['started', 'created', 'sending', 'waiting_headers', 'reading_body', 'validating', 'completed']);
11
+ export class Audit {
12
+ constructor(path) {
13
+ this.failed = false;
14
+ try {
15
+ this.fd = openSync(path, constants.O_WRONLY | constants.O_CREAT | constants.O_APPEND | (constants.O_NOFOLLOW ?? 0), 0o600);
16
+ if (process.platform !== 'win32') fchmodSync(this.fd, 0o600);
17
+ } catch { throw new CaeError('cannot_open_audit_log'); }
18
+ }
19
+ emit(record) {
20
+ const clean = { schema: 1, time: new Date().toISOString() };
21
+ for (const [key, value] of Object.entries(record)) {
22
+ if (TIMING_FIELDS.includes(key)) { clean[key] = knownNumber(value); continue; }
23
+ if (key === 'judgeStage') { if (TIMING_STAGES.has(value)) clean[key] = value; continue; }
24
+ if (!FIELDS.has(key)) continue;
25
+ if (value === null || typeof value === 'boolean' || (typeof value === 'number' && Number.isFinite(value))) clean[key] = value;
26
+ else if (typeof value === 'string' && value.length < 200) clean[key] = value;
27
+ }
28
+ try { writeSync(this.fd, JSON.stringify(clean) + '\n'); }
29
+ catch { this.failed = true; } // Do not interrupt a model stream for a logging failure.
30
+ }
31
+ close() { if (this.fd !== undefined) { closeSync(this.fd); this.fd = undefined; } }
32
+ }
33
+ export function usageOf(response) {
34
+ const u = response?.usage;
35
+ return { inputTokens: knownNumber(u?.input_tokens), cachedInputTokens: knownNumber(u?.input_tokens_details?.cached_tokens),
36
+ outputTokens: knownNumber(u?.output_tokens), reasoningTokens: knownNumber(u?.output_tokens_details?.reasoning_tokens) };
37
+ }
38
+ export function report(records) {
39
+ records = records.filter(isObject);
40
+ const outcomes = records.filter(r => r.event === 'upstream_outcome');
41
+ const decisions = records.filter(r => r.event === 'decision');
42
+ const summary = {
43
+ decisions: decisions.length, requests: outcomes.length,
44
+ completed: outcomes.filter(r => r.completed).length,
45
+ changedRequests: records.filter(r => r.event === 'request_sent' && r.changed).length,
46
+ sources: {}, tokenObservations: {}, measuredSavings: null,
47
+ note: 'Observed usage only. Reasoning tokens are a subset of output tokens; do not add them again. Missing is not zero. Shadow traffic still uses the original effort. No counterfactual or quality equivalence is established.',
48
+ };
49
+ for (const r of decisions) summary.sources[r.source] = (summary.sources[r.source] ?? 0) + 1;
50
+ for (const key of ['inputTokens', 'cachedInputTokens', 'outputTokens', 'reasoningTokens']) {
51
+ const values = outcomes.map(r => knownNumber(r[key])).filter(v => v !== null);
52
+ summary.tokenObservations[key] = { observedSum: values.length ? values.reduce((a, b) => a + b, 0) : null,
53
+ knownRequests: values.length, unknownRequests: outcomes.length - values.length };
54
+ }
55
+ const judgeOutcomes = records.filter(r => r.event === 'judge_finished');
56
+ const knownJudgeUsage = judgeOutcomes.map(r => knownNumber(r.judgeInputTokens)).filter(v => v !== null);
57
+ summary.evaluator = {
58
+ attempts: records.filter(r => r.event === 'judge_started').length,
59
+ externalAttempts: records.filter(r => r.event === 'judge_started' && r.judgeKind === 'typesafe').length,
60
+ observedInputTokens: knownJudgeUsage.length ? knownJudgeUsage.reduce((a, b) => a + b, 0) : null,
61
+ unknownInputUsage: judgeOutcomes.length - knownJudgeUsage.length,
62
+ observedDurationMs: judgeOutcomes.reduce((n, r) => n + (knownNumber(r.judgeMs) ?? 0), 0),
63
+ note: 'Cancelled/timed-out evaluator calls may be billable even when usage is unknown. No dollar estimate.'
64
+ };
65
+ summary.evaluator.timeoutStages = {};
66
+ for (const row of judgeOutcomes.filter(r => r.reason === 'judge_timeout')) {
67
+ const stage = TIMING_STAGES.has(row.judgeStage) ? row.judgeStage : 'unknown';
68
+ summary.evaluator.timeoutStages[stage] = (summary.evaluator.timeoutStages[stage] ?? 0) + 1;
69
+ }
70
+ return summary;
71
+ }
package/src/codex.mjs ADDED
@@ -0,0 +1,101 @@
1
+ import { spawn, spawnSync } from 'node:child_process';
2
+ import { createInterface } from 'node:readline';
3
+ import { CaeError, isObject } from './util.mjs';
4
+ import { EFFORT_NAMES } from './config.mjs';
5
+ import { VERSION } from './version.mjs';
6
+
7
+ export function nativeEnvironment(source = process.env) {
8
+ const env = { ...source }; delete env.TYPESAFE_API_KEY; return env;
9
+ }
10
+ export function doctor(binary = 'codex') {
11
+ const check = spawnSync(binary, ['--version'], { encoding: 'utf8', timeout: 5000, windowsHide: true, env: nativeEnvironment() });
12
+ return { node: process.version, platform: process.platform, architecture: process.arch,
13
+ codexFound: !check.error && check.status === 0,
14
+ codexVersion: /codex[^\r\n]{0,50}\d[^\r\n]{0,50}/i.exec(check.stdout ?? '')?.[0] ?? null,
15
+ credentialsReadByCAE: false, paidCallsMade: false,
16
+ desktopCompatibility: 'pending local acceptance', websocketAdapter: 'not implemented' };
17
+ }
18
+ export function normalizeModel(m) {
19
+ if (typeof m?.model !== 'string' || !/^[a-zA-Z0-9][a-zA-Z0-9._:/-]{0,159}$/.test(m.model)) return null;
20
+ if (!Array.isArray(m.supportedReasoningEfforts)) return null;
21
+ const efforts = m.supportedReasoningEfforts.map(e => e?.reasoningEffort);
22
+ if (!Array.isArray(efforts) || !efforts.length || efforts.some(e => !EFFORT_NAMES.includes(e)) ||
23
+ !efforts.includes(m.defaultReasoningEffort)) return null;
24
+ return { model: m.model, supportedEfforts: [...new Set(efforts)], baseline: m.defaultReasoningEffort };
25
+ }
26
+ /** Only initialize + model/list. No turns, tool calls, credentials API or file API. */
27
+ export async function probeModels(binary = 'codex', { args = ['app-server'], timeoutMs = 15000 } = {}) {
28
+ return new Promise((resolve, reject) => {
29
+ const child = spawn(binary, args, { stdio: ['pipe', 'pipe', 'ignore'], windowsHide: true, env: nativeEnvironment() });
30
+ const rl = createInterface({ input: child.stdout });
31
+ let settled = false, nextId = 2, expected = 1, pages = 0, outputBytes = 0;
32
+ const models = []; const cursors = new Set();
33
+ const timer = setTimeout(() => finish(new CaeError('codex_probe_timeout')), timeoutMs);
34
+ function finish(error) {
35
+ if (settled) return; settled = true; clearTimeout(timer); rl.close(); child.stdin.end(); child.kill();
36
+ const killTimer = setTimeout(() => { if (child.exitCode === null) child.kill('SIGKILL'); }, 1000);
37
+ killTimer.unref(); child.once('exit', () => clearTimeout(killTimer));
38
+ if (error) reject(error);
39
+ else resolve({ schema: 1, observedAt: new Date().toISOString(), source: 'codex app-server model/list', models,
40
+ skippedUnsupportedEntries: skipped, paidGenerations: 0 });
41
+ }
42
+ let skipped = 0;
43
+ function send(msg) { if (!settled) child.stdin.write(JSON.stringify(msg) + '\n'); }
44
+ function list(cursor) {
45
+ if (++pages > 20) { finish(new CaeError('codex_probe_page_limit')); return; }
46
+ expected = nextId++;
47
+ send({ id: expected, method: 'model/list', params: { limit: 100, includeHidden: false, ...(cursor ? { cursor } : {}) } });
48
+ }
49
+ child.stdout.on('data', b => { outputBytes += b.length; if (outputBytes > 4 * 1024 * 1024) finish(new CaeError('codex_probe_output_limit')); });
50
+ child.once('error', () => finish(new CaeError('codex_not_available')));
51
+ child.stdin.on('error', () => finish(new CaeError('codex_probe_pipe_error')));
52
+ child.once('exit', () => { if (!settled) finish(new CaeError('codex_probe_exited')); });
53
+ rl.on('line', line => {
54
+ if (settled) return;
55
+ let msg; try { msg = JSON.parse(line); } catch { finish(new CaeError('codex_probe_invalid_json')); return; }
56
+ if (!isObject(msg)) { finish(new CaeError('codex_probe_invalid_message')); return; }
57
+ if (msg.id !== expected) return;
58
+ if (msg.error) { finish(new CaeError('codex_probe_rpc_error')); return; }
59
+ if (expected === 1) { send({ method: 'initialized', params: {} }); list(); return; }
60
+ if (!Array.isArray(msg.result?.data)) { finish(new CaeError('codex_probe_invalid_models')); return; }
61
+ for (const entry of msg.result.data) { const m = normalizeModel(entry); if (m) models.push(m); else ++skipped; }
62
+ const cursor = msg.result.nextCursor;
63
+ if (cursor === null || cursor === undefined) { finish(); return; }
64
+ if (typeof cursor !== 'string' || cursors.has(cursor)) { finish(new CaeError('codex_probe_cursor_loop')); return; }
65
+ cursors.add(cursor); list(cursor);
66
+ });
67
+ send({ id: 1, method: 'initialize', params: { clientInfo: { name: 'codex_adaptive_effort', version: VERSION } } });
68
+ });
69
+ }
70
+ export function codexArgs(config, auth, passthrough = []) {
71
+ if (!['api', 'chatgpt'].includes(auth) || config.upstream.kind !== auth) throw new CaeError('auth_route_mismatch');
72
+ if (!config.port) throw new CaeError('fixed_port_required_for_codex');
73
+ const pairs = {
74
+ model: config.model,
75
+ model_reasoning_effort: config.baseline,
76
+ model_provider: 'cae',
77
+ 'model_providers.cae.name': 'CAE experimental local effort controller',
78
+ 'model_providers.cae.base_url': `http://127.0.0.1:${config.port}/v1`,
79
+ 'model_providers.cae.wire_api': 'responses',
80
+ 'model_providers.cae.supports_websockets': false,
81
+ 'model_providers.cae.request_max_retries': 0,
82
+ 'model_providers.cae.stream_max_retries': 0,
83
+ 'model_providers.cae.env_http_headers': null,
84
+ };
85
+ if (auth === 'chatgpt') pairs['model_providers.cae.requires_openai_auth'] = true;
86
+ else { pairs['model_providers.cae.requires_openai_auth'] = false; pairs['model_providers.cae.env_key'] = 'OPENAI_API_KEY'; }
87
+ const overrides = Object.entries(pairs).flatMap(([key, value]) => ['-c', `${key}=${value === null ? '{"x-cae-token"="CAE_LOCAL_TOKEN"}' : JSON.stringify(value)}`]);
88
+ // Desktop supplies root -c flags and app-server-local -c flags. Native CLI
89
+ // discards the root config collection when the subcommand has its own.
90
+ // Recognize the desktop prefix without mistaking an exec prompt or a config
91
+ // value for a subcommand; other launch shapes keep their existing ordering.
92
+ let command = 0;
93
+ while (command < passthrough.length) {
94
+ if (['-c', '--config'].includes(passthrough[command]) && command + 1 < passthrough.length) command += 2;
95
+ else if (passthrough[command].startsWith('--config=')) ++command;
96
+ else break;
97
+ }
98
+ if (passthrough[command] === 'app-server')
99
+ return [...passthrough.slice(0, command + 1), ...overrides, ...passthrough.slice(command + 1)];
100
+ return [...overrides, ...passthrough];
101
+ }
package/src/config.mjs ADDED
@@ -0,0 +1,74 @@
1
+ import { readFileSync, lstatSync } from 'node:fs';
2
+ import { resolve, dirname } from 'node:path';
3
+ import { CaeError, isObject } from './util.mjs';
4
+
5
+ export const EFFORT_NAMES = ['none', 'minimal', 'low', 'medium', 'high', 'xhigh', 'max', 'ultra'];
6
+ export const UPSTREAMS = Object.freeze({
7
+ api: 'https://api.openai.com/v1',
8
+ chatgpt: 'https://chatgpt.com/backend-api/codex',
9
+ });
10
+ export function validateConfig(c) {
11
+ if (!isObject(c) || c.version !== 1) throw new CaeError('config_version');
12
+ const allowed = ['version', 'model', 'supportedEfforts', 'baseline', 'mode', 'port', 'upstream',
13
+ 'tokenFile', 'logFile', 'judge', 'lease', 'maxBodyBytes', 'upstreamTimeoutMs', 'capabilitySource'];
14
+ if (Object.keys(c).some(k => !allowed.includes(k))) throw new CaeError('unknown_config_field');
15
+ if (typeof c.model !== 'string' || !/^[a-zA-Z0-9][a-zA-Z0-9._:/-]{0,159}$/.test(c.model)) throw new CaeError('config_model');
16
+ if (!Array.isArray(c.supportedEfforts) || c.supportedEfforts.length < 1 || c.supportedEfforts.length > EFFORT_NAMES.length ||
17
+ new Set(c.supportedEfforts).size !== c.supportedEfforts.length ||
18
+ c.supportedEfforts.some(v => !EFFORT_NAMES.includes(v))) throw new CaeError('config_efforts');
19
+ if (!c.supportedEfforts.includes(c.baseline)) throw new CaeError('config_baseline');
20
+ if (!['off', 'shadow', 'auto'].includes(c.mode)) throw new CaeError('config_mode');
21
+ if (!Number.isInteger(c.port) || c.port < 0 || c.port > 65535) throw new CaeError('config_port');
22
+ if (!isObject(c.upstream)) throw new CaeError('config_upstream');
23
+ if (c.upstream.kind === 'mock') {
24
+ let url;
25
+ try { url = new URL(c.upstream.baseUrl); } catch { throw new CaeError('config_mock_url'); }
26
+ if (url.protocol !== 'http:' || url.hostname !== '127.0.0.1' || !url.port || url.username || url.password || url.search || url.hash || !['', '/', '/v1'].includes(url.pathname))
27
+ throw new CaeError('config_mock_url');
28
+ } else if (typeof c.upstream.kind !== 'string' || !Object.hasOwn(UPSTREAMS, c.upstream.kind) || c.upstream.baseUrl !== undefined) throw new CaeError('config_upstream');
29
+ if (typeof c.tokenFile !== 'string' || !c.tokenFile || typeof c.logFile !== 'string' || !c.logFile) throw new CaeError('config_paths');
30
+ if (!isObject(c.judge) || !['baseline', 'typesafe'].includes(c.judge.kind)) throw new CaeError('config_judge');
31
+ if (typeof c.judge.model !== 'string' || !/^[a-zA-Z0-9][a-zA-Z0-9._:/-]{0,159}$/.test(c.judge.model)) throw new CaeError('config_judge_model');
32
+ for (const [v, lo, hi, code] of [
33
+ [c.judge.timeoutMs, 50, 30000, 'config_judge_timeout'],
34
+ [c.judge.maxCalls, 0, 100000, 'config_judge_budget'],
35
+ [c.lease?.maxGenerations, 1, 4, 'config_lease'],
36
+ [c.lease?.ttlMs, 100, 300000, 'config_lease_ttl'],
37
+ [c.maxBodyBytes, 1024, 32 * 1024 * 1024, 'config_body_limit'],
38
+ [c.upstreamTimeoutMs, 100, 3600000, 'config_upstream_timeout'],
39
+ ]) if (!Number.isInteger(v) || v < lo || v > hi) throw new CaeError(code);
40
+ return structuredClone(c);
41
+ }
42
+ export function defaultConfig(model, supportedEfforts, baseline) {
43
+ return validateConfig({
44
+ version: 1, model, supportedEfforts, baseline, mode: 'shadow', port: 4318,
45
+ upstream: { kind: 'api' }, tokenFile: 'local.key', logFile: 'events.jsonl',
46
+ judge: { kind: 'baseline', model: 'jev-latest', timeoutMs: 1500, maxCalls: 100 },
47
+ lease: { maxGenerations: 4, ttlMs: 30000 }, maxBodyBytes: 8 * 1024 * 1024,
48
+ upstreamTimeoutMs: 300000, capabilitySource: 'operator-supplied; not live-verified',
49
+ });
50
+ }
51
+ export function loadConfig(path) {
52
+ const absolute = resolve(path);
53
+ let value;
54
+ try { value = JSON.parse(readFileSync(absolute, 'utf8')); }
55
+ catch { throw new CaeError('cannot_read_config'); }
56
+ const config = validateConfig(value);
57
+ config.tokenFile = resolve(dirname(absolute), config.tokenFile);
58
+ config.logFile = resolve(dirname(absolute), config.logFile);
59
+ if (config.tokenFile === config.logFile) throw new CaeError('config_path_collision');
60
+ return config;
61
+ }
62
+ export function readLocalToken(path) {
63
+ try {
64
+ const st = lstatSync(path);
65
+ if (!st.isFile() || st.isSymbolicLink()) throw new Error();
66
+ if (process.platform !== 'win32' && (st.mode & 0o077) !== 0) throw new Error();
67
+ const token = readFileSync(path, 'utf8').trim();
68
+ if (!/^[a-f0-9]{64}$/.test(token)) throw new Error();
69
+ return token;
70
+ } catch { throw new CaeError('invalid_or_insecure_local_token'); }
71
+ }
72
+ export function upstreamBase(c) {
73
+ return c.upstream.kind === 'mock' ? c.upstream.baseUrl.replace(/\/$/, '') : UPSTREAMS[c.upstream.kind];
74
+ }
@@ -0,0 +1,87 @@
1
+ import { digest, isObject, shortText, stable } from './util.mjs';
2
+
3
+ /** Best-effort redaction; never a promise that an excerpt is safe to disclose. */
4
+ export function redact(text) {
5
+ return String(text)
6
+ .replace(/-----BEGIN [^-]*PRIVATE KEY-----[\s\S]*?-----END [^-]*PRIVATE KEY-----/g, '[REDACTED PRIVATE KEY]')
7
+ .replace(/\b(?:sk-|ghp_|github_pat_)[A-Za-z0-9_-]{12,}\b/g, '[REDACTED TOKEN]')
8
+ .replace(/\bBearer\s+[A-Za-z0-9._~+\/-]+=*/gi, 'Bearer [REDACTED]')
9
+ .replace(/((?:api[_-]?key|access[_-]?token|password|secret)\s*["']?\s*[:=]\s*["']?)[^\s,"'}]+/gi, '$1[REDACTED]')
10
+ .replace(/\/(?:Users|home)\/[^/\s]+/g, '/[HOME]');
11
+ }
12
+ function textOf(content) {
13
+ if (typeof content === 'string') return content;
14
+ if (!Array.isArray(content)) return '';
15
+ return content.filter(p => isObject(p) && ['input_text', 'output_text', 'text'].includes(p.type))
16
+ .map(p => typeof p.text === 'string' ? p.text : '').join('\n');
17
+ }
18
+ function hasMedia(content) {
19
+ return Array.isArray(content) && content.some(p => isObject(p) && /image|audio|video|file/.test(p.type ?? ''));
20
+ }
21
+ function supportedItem(item) {
22
+ if (!isObject(item)) return false;
23
+ if (item.type === undefined || item.type === 'message') {
24
+ return ['user', 'assistant', 'system', 'developer'].includes(item.role) &&
25
+ (typeof item.content === 'string' || (Array.isArray(item.content) && item.content.every(p =>
26
+ isObject(p) && ['input_text', 'output_text', 'text'].includes(p.type) && typeof p.text === 'string')));
27
+ }
28
+ if (['function_call_output', 'custom_tool_call_output'].includes(item.type)) return typeof item.output === 'string';
29
+ // Retain opaque reasoning and call records for history integrity, never evaluator text.
30
+ return ['reasoning', 'function_call', 'custom_tool_call'].includes(item.type);
31
+ }
32
+ function failed(item) {
33
+ if (item.is_error === true || item.success === false) return true;
34
+ let out = item.output;
35
+ if (typeof out === 'string') { try { out = JSON.parse(out); } catch { return false; } }
36
+ return isObject(out) && (out.is_error === true || out.success === false ||
37
+ (typeof out.exit_code === 'number' && out.exit_code !== 0) ||
38
+ (typeof out.exitCode === 'number' && out.exitCode !== 0));
39
+ }
40
+ export function inspectRequest(body) {
41
+ const items = typeof body.input === 'string' ? [{ role: 'user', content: body.input }] : body.input;
42
+ if (!Array.isArray(items)) return { eligible: false, reason: 'unsupported_input' };
43
+ if (items.length > 4096) return { eligible: false, reason: 'history_item_limit' };
44
+ if (body.previous_response_id || body.conversation) return { eligible: false, reason: 'delta_history' };
45
+ if (body.background === true) return { eligible: false, reason: 'background_response' };
46
+ if (items.some(i => isObject(i) && i.type === 'configuration_update')) return { eligible: false, reason: 'existing_configuration_update' };
47
+ if (items.some(i => isObject(i) && /compaction/.test(i.type ?? '')) || body.context_management)
48
+ return { eligible: false, reason: 'compaction_boundary' };
49
+ if (body.truncation && body.truncation !== 'disabled') return { eligible: false, reason: 'automatic_truncation' };
50
+ if (items.some(i => isObject(i) && (hasMedia(i.content) || (typeof i.output === 'object' && i.output !== null))))
51
+ return { eligible: false, reason: 'media_or_structured_tool_evidence' };
52
+ if (!items.every(supportedItem)) return { eligible: false, reason: 'unsupported_history_item' };
53
+ const users = items.filter(i => isObject(i) && i.role === 'user');
54
+ const latest = users.at(-1);
55
+ if (!latest || !textOf(latest.content).trim()) return { eligible: false, reason: 'missing_user_goal' };
56
+ const publicNotes = items.filter(i => isObject(i) && i.role === 'assistant').slice(-2);
57
+ const tools = items.filter(i => isObject(i) && ['function_call_output', 'custom_tool_call_output'].includes(i.type));
58
+ const errors = tools.filter(failed).map(i => digest(i));
59
+ const clean = (text, max) => shortText(redact(text), max);
60
+ const state = {
61
+ task: clean(textOf(users.length > 1 ? users.at(-2).content : latest.content), 1800),
62
+ latestUser: clean(textOf(latest.content), 2200),
63
+ publicNotes: publicNotes.map(i => clean(textOf(i.content), 600)),
64
+ recentTools: tools.slice(-3).map(i => ({
65
+ result: clean(typeof i.output === 'string' ? i.output : '', 600), failed: failed(i),
66
+ })),
67
+ omittedToolCalls: Math.max(0, tools.length - 3),
68
+ evidenceLimits: 'Bounded character excerpts; missing text remains unknown. No image understanding.',
69
+ };
70
+ // All integrity fingerprints are local only. No auth, session ids, encrypted reasoning,
71
+ // tool definitions or system/developer instructions enter evaluator state.
72
+ return {
73
+ eligible: true, state, itemCount: items.length,
74
+ itemHashes: items.map(digest),
75
+ taskKey: digest(users.map(i => textOf(i.content))),
76
+ constraintsKey: digest({ model: body.model, instructions: body.instructions ?? null, tools: body.tools ?? null }),
77
+ errorsKey: digest(errors),
78
+ reasoningKey: digest(body.reasoning ?? null),
79
+ stateChars: Array.from(stable(state)).length,
80
+ };
81
+ }
82
+ export function canContinue(previous, next) {
83
+ return previous && next.eligible && previous.taskKey === next.taskKey &&
84
+ previous.constraintsKey === next.constraintsKey && previous.errorsKey === next.errorsKey &&
85
+ previous.reasoningKey === next.reasoningKey && next.itemCount > previous.itemCount &&
86
+ previous.itemHashes.every((h, i) => next.itemHashes[i] === h);
87
+ }