agents-relay 1.0.5 → 1.0.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +14 -13
- package/dist/adapters.js +107 -257
- package/dist/cli.js +19 -8
- package/dist/events.js +1 -1
- package/dist/planner.js +9 -4
- package/dist/reconciler.js +132 -100
- package/dist/relayd.js +3 -3
- package/dist/store.js +1 -1
- package/package.json +1 -1
- package/skills/agents-relay/SKILL.md +58 -44
- package/skills/agents-relay/agents/planner.agent.md +12 -9
- package/skills/chatgpt-browser-worker/SKILL.md +56 -96
- package/skills/chatgpt-browser-worker/agents/browser-worker.agent.md +30 -101
- package/skills/chatgpt-browser-worker/scripts/_temporary_bh.py +171 -0
- package/skills/chatgpt-browser-worker/scripts/{create_bh.py → temporary_bh.py} +5 -8
- package/skills/chatgpt-browser-worker/references/contract.md +0 -132
- package/skills/chatgpt-browser-worker/references/orchestration.md +0 -101
- package/skills/chatgpt-browser-worker/scripts/_create_bh.py +0 -158
- package/skills/chatgpt-browser-worker/scripts/_operate_bh.py +0 -198
- package/skills/chatgpt-browser-worker/scripts/contract.py +0 -151
- package/skills/chatgpt-browser-worker/scripts/create.py +0 -84
- package/skills/chatgpt-browser-worker/scripts/operate_bh.py +0 -36
- package/skills/chatgpt-browser-worker/scripts/operations.py +0 -179
- package/skills/chatgpt-browser-worker/tests/fixtures/relay_lifecycle.json +0 -21
- package/skills/chatgpt-browser-worker/tests/test_agent_definition.py +0 -27
- package/skills/chatgpt-browser-worker/tests/test_contract.py +0 -315
|
@@ -1,102 +1,62 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: chatgpt-browser-worker
|
|
3
|
-
description:
|
|
3
|
+
description: Submit an isolated one-shot ChatGPT browser worker task through browser-harness with a required durable output destination.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# ChatGPT browser worker
|
|
7
7
|
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
result
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
browser-harness and verification contract.
|
|
64
|
-
|
|
65
|
-
The browser-worker agent's `create` intent starts a new ChatGPT chat when the
|
|
66
|
-
attached tab is already a conversation and emits JSON with the observed
|
|
67
|
-
`thread_id`, conversation URL, selected Project, and thinking observation.
|
|
68
|
-
`scripts/create_bh.py` is only a helper for that agent, not a caller-facing
|
|
69
|
-
runtime contract.
|
|
70
|
-
|
|
71
|
-
For implementation/debugging, the agent may use the helper equivalent for an
|
|
72
|
-
existing thread with the durable identity from persisted state:
|
|
73
|
-
|
|
74
|
-
```text
|
|
75
|
-
python3 scripts/operate_bh.py resume --thread-id ID --project NAME
|
|
76
|
-
python3 scripts/operate_bh.py continue --thread-id ID --project NAME --prompt TEXT
|
|
77
|
-
python3 scripts/operate_bh.py status --thread-id ID --project NAME
|
|
78
|
-
python3 scripts/operate_bh.py result --thread-id ID --project NAME
|
|
79
|
-
python3 scripts/operate_bh.py delete --thread-id ID --project NAME
|
|
80
|
-
```
|
|
81
|
-
|
|
82
|
-
Each command performs one serial browser-harness operation. The injected
|
|
83
|
-
semantic adapter in `scripts/operations.py` is the testable boundary; it does
|
|
84
|
-
not import MacBridge or a ChatGPT runtime API.
|
|
85
|
-
|
|
86
|
-
## Verification boundary
|
|
87
|
-
|
|
88
|
-
The worker may report `completed` only from an observed assistant message and
|
|
89
|
-
normalized result. A process exit, click, URL change, or tab title alone is not
|
|
90
|
-
proof that a thread was created, resumed, completed, or deleted. Authentication
|
|
91
|
-
walls, MFA, consent, ambiguous account/project selection, and unverified
|
|
92
|
-
thinking levels are `blocked` or `failed` conditions and must not be
|
|
93
|
-
self-healed.
|
|
94
|
-
|
|
95
|
-
Delete targets the exact durable conversation URL, requires an observed action
|
|
96
|
-
and confirmation, and verifies that the requested thread is no longer visible.
|
|
97
|
-
Repeated cleanup is idempotent (`not_found` is accepted); archive/undo UI
|
|
98
|
-
controls are evidence only and never revive a deleted tombstone.
|
|
99
|
-
|
|
100
|
-
Keep credentials, prompts containing private data, and browser runtime state
|
|
101
|
-
out of durable state and events. Runtime artifacts belong under the caller's
|
|
102
|
-
ignored `runs/` directory.
|
|
8
|
+
Use this skill to hand one self-contained task to ChatGPT through the authenticated
|
|
9
|
+
browser session. The caller does not manage ChatGPT threads, model selection,
|
|
10
|
+
thinking level, polling, or result retrieval.
|
|
11
|
+
|
|
12
|
+
## Output contract
|
|
13
|
+
|
|
14
|
+
Every task handed to this worker MUST declare exactly one durable output mode
|
|
15
|
+
before the worker is launched.
|
|
16
|
+
|
|
17
|
+
### 1. Task / PR output
|
|
18
|
+
|
|
19
|
+
Use this for repository or managed-task work.
|
|
20
|
+
|
|
21
|
+
The task prompt must identify the exact managed task/PR that owns the result.
|
|
22
|
+
The worker performs the requested work directly against that task/PR and records
|
|
23
|
+
its final outcome there.
|
|
24
|
+
|
|
25
|
+
Output:
|
|
26
|
+
task/PR
|
|
27
|
+
<exact managed task/PR identity and required final action>
|
|
28
|
+
|
|
29
|
+
### 2. File output
|
|
30
|
+
|
|
31
|
+
Use this for research, analysis, reports, or other artifact-producing work.
|
|
32
|
+
|
|
33
|
+
The task prompt must name the exact file path. That file is the authoritative
|
|
34
|
+
result.
|
|
35
|
+
|
|
36
|
+
Output:
|
|
37
|
+
file
|
|
38
|
+
<exact file path>
|
|
39
|
+
|
|
40
|
+
A Browser ChatGPT worker task without one of these output declarations is
|
|
41
|
+
invalid and must not be launched. The orchestrator chooses and writes the output
|
|
42
|
+
contract when it creates the worker task; the worker must not invent or change
|
|
43
|
+
the destination.
|
|
44
|
+
|
|
45
|
+
Progress and terminal success/failure are always reported through the normal
|
|
46
|
+
Neo event contract. Events are execution observability, not a third output mode.
|
|
47
|
+
|
|
48
|
+
## Runtime behavior
|
|
49
|
+
|
|
50
|
+
The worker is one-shot. It opens an isolated worker-owned ChatGPT tab, submits
|
|
51
|
+
the complete task, verifies that ChatGPT accepted the submission, closes its
|
|
52
|
+
owned tab, and returns. ChatGPT continues the task independently and publishes
|
|
53
|
+
the result through the declared output contract.
|
|
54
|
+
|
|
55
|
+
The worker must never interact through a pre-existing user ChatGPT tab.
|
|
56
|
+
Browser details, transient conversation identity, submission verification,
|
|
57
|
+
model/default-thinking behavior, and tab cleanup are implementation concerns of
|
|
58
|
+
the worker and are not part of the caller-facing contract.
|
|
59
|
+
|
|
60
|
+
The runtime entry point is the browser-worker agent. Callers launch that agent
|
|
61
|
+
with the complete task prompt and declared output; they do not call helper
|
|
62
|
+
scripts directly.
|
|
@@ -1,113 +1,42 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: browser-worker
|
|
3
|
-
description:
|
|
3
|
+
description: Submit one isolated ChatGPT browser task and return after the task is accepted.
|
|
4
4
|
type: worker
|
|
5
5
|
---
|
|
6
6
|
|
|
7
7
|
# Browser Worker Agent
|
|
8
8
|
|
|
9
|
-
You are the runtime owner for the `chatgpt-browser-worker` skill.
|
|
10
|
-
|
|
11
|
-
outer adapter launches you with one typed lifecycle request; Agents Relay
|
|
12
|
-
owns task correlation, retries, events, and run-local persistence, but does
|
|
13
|
-
not own browser execution.
|
|
9
|
+
You are the runtime owner for the `chatgpt-browser-worker` skill. Load and follow
|
|
10
|
+
that skill and the `browser-harness` skill before execution.
|
|
14
11
|
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
- Use `browser-harness` for every browser interaction. Do not use another
|
|
18
|
-
browser stack, ChatGPT API, MacBridge ChatGPT runtime, copied cookies, or
|
|
19
|
-
direct HTTP calls to ChatGPT.
|
|
20
|
-
- Preserve the exact observed ChatGPT `thread_id` and Project identity. A URL,
|
|
21
|
-
title, tab index, or inferred conversation is not an identity. A resumed
|
|
22
|
-
thread must remain in its recorded Project; reject a mismatch.
|
|
23
|
-
- Treat the scripts in `scripts/` as helpers only. You may use, repair, or
|
|
24
|
-
bypass `create_bh.py` and `operate_bh.py` when they are brittle, but preserve
|
|
25
|
-
the skill contract and its evidence requirements.
|
|
26
|
-
- When a selector, layout, label, or ordinary UI flow changes, re-observe the
|
|
27
|
-
current page through `browser-harness`, identify the semantic control, and
|
|
28
|
-
adapt the operation. Do not guess from stale selectors or coordinates.
|
|
29
|
-
- Stop with `blocked` for authentication, MFA, consent, ambiguous account or
|
|
30
|
-
Project selection, or any human decision. Do not bypass these gates.
|
|
31
|
-
- Stop with `blocked` or `failed` when the result cannot be verified. A process
|
|
32
|
-
exit, URL change, tab title, prior assistant bubble, or synthetic message ID
|
|
33
|
-
is not completion evidence.
|
|
34
|
-
- Return only observations from the current browser state and persist the last
|
|
35
|
-
known `thread_id` when an operation fails after creation.
|
|
36
|
-
|
|
37
|
-
## Typed adapter contract
|
|
38
|
-
|
|
39
|
-
The input is one JSON object. `operation` must be one of `create`, `resume`,
|
|
40
|
-
`continue`, `status`, `result`, or `delete`.
|
|
41
|
-
|
|
42
|
-
```json
|
|
43
|
-
{
|
|
44
|
-
"operation": "create",
|
|
45
|
-
"thread_id": null,
|
|
46
|
-
"project": {"name": "neo", "id": "project-optional"},
|
|
47
|
-
"prompt": "Run the assigned task.",
|
|
48
|
-
"thinking_level": "high",
|
|
49
|
-
"state": null
|
|
50
|
-
}
|
|
51
|
-
```
|
|
12
|
+
The outer orchestrator gives you one complete task. That task MUST already
|
|
13
|
+
contain exactly one output declaration defined by the skill:
|
|
52
14
|
|
|
53
|
-
|
|
15
|
+
- `task/PR` with the exact managed task/PR identity and required final action; or
|
|
16
|
+
- `file` with the exact authoritative output path.
|
|
54
17
|
|
|
55
|
-
|
|
56
|
-
- `resume`: requires `thread_id` and `project.name`; verify the recorded
|
|
57
|
-
Project before any other thread action.
|
|
58
|
-
- `continue`: requires `thread_id`, `project.name`, and `prompt`; open and
|
|
59
|
-
verify the exact thread before sending the follow-up.
|
|
60
|
-
- `status`, `result`, and `delete`: require the durable `thread_id`; use the
|
|
61
|
-
persisted `state` when supplied and never recreate a missing identity.
|
|
62
|
-
- `thinking_level` is `default`, `low`, `medium`, or `high`; preserve an
|
|
63
|
-
observed effective value of `unknown` instead of guessing.
|
|
18
|
+
If the output declaration is missing or ambiguous, do not invent one.
|
|
64
19
|
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
```json
|
|
68
|
-
{
|
|
69
|
-
"operation": "result",
|
|
70
|
-
"status": "completed",
|
|
71
|
-
"thread_id": "chatgpt-conversation-id",
|
|
72
|
-
"project": {"name": "neo", "id": "project-id"},
|
|
73
|
-
"result": {
|
|
74
|
-
"message_id": "observed-assistant-message-id",
|
|
75
|
-
"text": "The normalized assistant answer.",
|
|
76
|
-
"verified": true,
|
|
77
|
-
"observed_at": "2026-09-19T12:02:00Z"
|
|
78
|
-
},
|
|
79
|
-
"state": {
|
|
80
|
-
"schema_version": 1,
|
|
81
|
-
"thread_id": "chatgpt-conversation-id",
|
|
82
|
-
"status": "completed",
|
|
83
|
-
"project": {"name": "neo", "id": "project-id"}
|
|
84
|
-
},
|
|
85
|
-
"error": null
|
|
86
|
-
}
|
|
87
|
-
```
|
|
88
|
-
|
|
89
|
-
Every response must include `operation`, `status`, and the exact
|
|
90
|
-
`thread_id` when one is known. A successful `result` response must include an
|
|
91
|
-
observed assistant `message_id`, non-empty normalized `text`, and
|
|
92
|
-
`result.verified: true`. `status` is not a result and must not claim
|
|
93
|
-
completion. For `delete`, return a verified `deleted` or idempotent verified
|
|
94
|
-
`not_found` observation and retain a terminal `deleted` tombstone.
|
|
95
|
-
|
|
96
|
-
For blocked or failed operations, return the last known identity and a typed
|
|
97
|
-
error without credentials or private prompts:
|
|
98
|
-
|
|
99
|
-
```json
|
|
100
|
-
{
|
|
101
|
-
"operation": "resume",
|
|
102
|
-
"status": "blocked",
|
|
103
|
-
"thread_id": "chatgpt-conversation-id",
|
|
104
|
-
"error": {
|
|
105
|
-
"code": "authentication_required",
|
|
106
|
-
"message": "Sign-in or MFA requires user action.",
|
|
107
|
-
"retryable": false
|
|
108
|
-
}
|
|
109
|
-
}
|
|
110
|
-
```
|
|
20
|
+
## Runtime rules
|
|
111
21
|
|
|
112
|
-
|
|
113
|
-
|
|
22
|
+
- Use `browser-harness` for every ChatGPT browser interaction.
|
|
23
|
+
- Execute only a one-shot submission. Do not expose or operate a
|
|
24
|
+
create/resume/status/result/continue lifecycle for callers.
|
|
25
|
+
- Always create a fresh worker-owned tab. Never type, upload, click New chat,
|
|
26
|
+
select a Project, or submit through a pre-existing user ChatGPT tab.
|
|
27
|
+
- Use Temporary Chat and the account's existing default model/thinking
|
|
28
|
+
settings. Do not change model or thinking settings.
|
|
29
|
+
- Upload requested files, submit the complete task prompt, and verify that the
|
|
30
|
+
submission became a new user turn.
|
|
31
|
+
- Once submission is verified, close the worker-owned tab and return. Do not
|
|
32
|
+
wait for the assistant response and do not reopen or poll the conversation.
|
|
33
|
+
- Treat any observed conversation/thread identity only as diagnostic evidence,
|
|
34
|
+
never as a resumable handle.
|
|
35
|
+
- Progress and terminal success/failure for the actual delegated task are
|
|
36
|
+
reported through the normal Neo event contract by the executing task. The
|
|
37
|
+
final durable result goes to the declared task/PR or file output.
|
|
38
|
+
- Stop on authentication, MFA, consent, or ambiguous browser state rather than
|
|
39
|
+
bypassing it.
|
|
40
|
+
|
|
41
|
+
Helper scripts under `scripts/` are implementation details. Callers must not
|
|
42
|
+
invoke them directly.
|
|
@@ -0,0 +1,171 @@
|
|
|
1
|
+
import atexit
|
|
2
|
+
import json
|
|
3
|
+
import os
|
|
4
|
+
import re
|
|
5
|
+
import time
|
|
6
|
+
|
|
7
|
+
from browser_harness import *
|
|
8
|
+
|
|
9
|
+
|
|
10
|
+
CFG = json.load(open("__CFG_PATH__", encoding="utf-8"))
|
|
11
|
+
_OWNED_TABS = []
|
|
12
|
+
|
|
13
|
+
|
|
14
|
+
def _new_owned_tab(url):
|
|
15
|
+
target_id = cdp("Target.createTarget", url="about:blank", background=True)["targetId"]
|
|
16
|
+
switch_tab(target_id)
|
|
17
|
+
_OWNED_TABS.append(target_id)
|
|
18
|
+
if url != "about:blank":
|
|
19
|
+
goto_url(url)
|
|
20
|
+
return target_id
|
|
21
|
+
|
|
22
|
+
|
|
23
|
+
def _close_owned_tabs():
|
|
24
|
+
while _OWNED_TABS:
|
|
25
|
+
try:
|
|
26
|
+
close_tab(_OWNED_TABS.pop())
|
|
27
|
+
except Exception:
|
|
28
|
+
pass
|
|
29
|
+
|
|
30
|
+
|
|
31
|
+
atexit.register(_close_owned_tabs)
|
|
32
|
+
|
|
33
|
+
|
|
34
|
+
def _composer():
|
|
35
|
+
selector = js("""(() => {
|
|
36
|
+
const preferred = document.querySelector('#prompt-textarea');
|
|
37
|
+
if (preferred) {
|
|
38
|
+
const r=preferred.getBoundingClientRect();
|
|
39
|
+
if (r.width>0 && r.height>0) return '#prompt-textarea';
|
|
40
|
+
}
|
|
41
|
+
const fallback = [...document.querySelectorAll('textarea,[contenteditable="true"]')]
|
|
42
|
+
.find(e => { const r=e.getBoundingClientRect(); return r.width>0 && r.height>0 && !e.disabled; });
|
|
43
|
+
if (!fallback) return null;
|
|
44
|
+
return fallback.tagName === 'TEXTAREA' ? 'textarea' : '[contenteditable="true"]';
|
|
45
|
+
})()""")
|
|
46
|
+
if not selector:
|
|
47
|
+
raise RuntimeError("Temporary Chat composer was not observed")
|
|
48
|
+
return selector
|
|
49
|
+
|
|
50
|
+
|
|
51
|
+
def _composer_text(selector):
|
|
52
|
+
return js(f"""(() => {{
|
|
53
|
+
const e=document.querySelector({json.dumps(selector)});
|
|
54
|
+
return e ? ((e.innerText ?? e.value) || '') : '';
|
|
55
|
+
}})()""") or ""
|
|
56
|
+
|
|
57
|
+
|
|
58
|
+
def _attachment_names():
|
|
59
|
+
return js(r"""(() => [...document.querySelectorAll('button[aria-label^="Remove file"]')]
|
|
60
|
+
.map(b => (b.getAttribute('aria-label') || '').replace(/^Remove file\s+\d+:\s*/, ''))
|
|
61
|
+
.filter(Boolean))()""") or []
|
|
62
|
+
|
|
63
|
+
|
|
64
|
+
def _upload_files(paths):
|
|
65
|
+
if not paths:
|
|
66
|
+
return []
|
|
67
|
+
selector = js("""(() => {
|
|
68
|
+
for (const s of ['#upload-files','#upload-media','input[name="upload-media"]','input[type="file"]'])
|
|
69
|
+
if (document.querySelector(s)) return s;
|
|
70
|
+
return null;
|
|
71
|
+
})()""")
|
|
72
|
+
if not selector:
|
|
73
|
+
raise RuntimeError("ChatGPT file input was not observed")
|
|
74
|
+
for path in paths:
|
|
75
|
+
if not os.path.isfile(path):
|
|
76
|
+
raise RuntimeError(f"attachment does not exist: {path}")
|
|
77
|
+
upload_file(selector, path)
|
|
78
|
+
expected = os.path.basename(path)
|
|
79
|
+
deadline = time.time() + 20
|
|
80
|
+
while time.time() < deadline and expected not in _attachment_names():
|
|
81
|
+
time.sleep(.25)
|
|
82
|
+
if expected not in _attachment_names():
|
|
83
|
+
raise RuntimeError(f"attachment was not observed ready: {expected}")
|
|
84
|
+
return [os.path.basename(path) for path in paths]
|
|
85
|
+
|
|
86
|
+
|
|
87
|
+
def _user_turns():
|
|
88
|
+
return js("""(() => [...document.querySelectorAll('[data-message-author-role="user"]')]
|
|
89
|
+
.map(e => ({text:e.innerText.trim(), id:e.getAttribute('data-message-id') || ''}))
|
|
90
|
+
.filter(e => e.text))()""") or []
|
|
91
|
+
|
|
92
|
+
|
|
93
|
+
def _click_send():
|
|
94
|
+
ok = js("""(() => {
|
|
95
|
+
const b=document.querySelector('button[data-testid="send-button"],button[aria-label="Send prompt"]');
|
|
96
|
+
if (!b || b.disabled || b.getAttribute('aria-disabled') === 'true') return false;
|
|
97
|
+
b.click(); return true;
|
|
98
|
+
})()""")
|
|
99
|
+
if not ok:
|
|
100
|
+
raise RuntimeError("Temporary Chat Send prompt button was not observed ready")
|
|
101
|
+
|
|
102
|
+
|
|
103
|
+
def _wait_user_turn(before_count, prompt, timeout=20):
|
|
104
|
+
deadline = time.time() + timeout
|
|
105
|
+
while time.time() < deadline:
|
|
106
|
+
turns = _user_turns()
|
|
107
|
+
if len(turns) > before_count and prompt.strip() in turns[-1]["text"]:
|
|
108
|
+
return turns[-1]
|
|
109
|
+
time.sleep(.25)
|
|
110
|
+
raise RuntimeError("Temporary Chat prompt did not become an observed user turn")
|
|
111
|
+
|
|
112
|
+
|
|
113
|
+
def _diagnostic_thread_id():
|
|
114
|
+
match = re.search(r"/c/([^/?#]+)", page_info().get("url", ""))
|
|
115
|
+
return match.group(1) if match else None
|
|
116
|
+
|
|
117
|
+
|
|
118
|
+
_new_owned_tab("https://chatgpt.com/")
|
|
119
|
+
wait_for_load()
|
|
120
|
+
|
|
121
|
+
enabled = js("""(() => {
|
|
122
|
+
const matches=[...document.querySelectorAll('button')].filter(
|
|
123
|
+
b => (b.getAttribute('aria-label') || '').trim() === 'Temporary chat'
|
|
124
|
+
);
|
|
125
|
+
if (matches.length !== 1) return false;
|
|
126
|
+
matches[0].click();
|
|
127
|
+
return true;
|
|
128
|
+
})()""")
|
|
129
|
+
if not enabled:
|
|
130
|
+
raise RuntimeError("Temporary Chat toggle was not uniquely observed")
|
|
131
|
+
|
|
132
|
+
deadline = time.time() + 20
|
|
133
|
+
while time.time() < deadline:
|
|
134
|
+
if "temporary-chat=true" in page_info().get("url", ""):
|
|
135
|
+
try:
|
|
136
|
+
_composer()
|
|
137
|
+
break
|
|
138
|
+
except RuntimeError:
|
|
139
|
+
pass
|
|
140
|
+
time.sleep(.25)
|
|
141
|
+
else:
|
|
142
|
+
raise RuntimeError("Temporary Chat mode did not become ready")
|
|
143
|
+
|
|
144
|
+
attachments = _upload_files(CFG.get("file", []))
|
|
145
|
+
selector = _composer()
|
|
146
|
+
before_count = len(_user_turns())
|
|
147
|
+
fill_input(selector, CFG["prompt"], clear_first=True)
|
|
148
|
+
if CFG["prompt"].strip() not in _composer_text(selector):
|
|
149
|
+
info = js(f"""(() => {{
|
|
150
|
+
const e=document.querySelector({json.dumps(selector)});
|
|
151
|
+
const r=e.getBoundingClientRect();
|
|
152
|
+
return {{x:r.x+r.width/2,y:r.y+r.height/2}};
|
|
153
|
+
}})()""")
|
|
154
|
+
click_at_xy(info["x"], info["y"])
|
|
155
|
+
press_key("CTRL+A")
|
|
156
|
+
type_text(CFG["prompt"])
|
|
157
|
+
if CFG["prompt"].strip() not in _composer_text(selector):
|
|
158
|
+
raise RuntimeError("Temporary Chat prompt was not observed in the composer")
|
|
159
|
+
|
|
160
|
+
_click_send()
|
|
161
|
+
user_turn = _wait_user_turn(before_count, CFG["prompt"])
|
|
162
|
+
|
|
163
|
+
print(json.dumps({
|
|
164
|
+
"operation": "submit",
|
|
165
|
+
"status": "submitted",
|
|
166
|
+
"temporary": True,
|
|
167
|
+
"attachments": attachments,
|
|
168
|
+
"diagnostic_thread_id": _diagnostic_thread_id(),
|
|
169
|
+
"user_message_id": user_turn.get("id") or None,
|
|
170
|
+
"verified": True,
|
|
171
|
+
}, ensure_ascii=False))
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
#!/usr/bin/env python3
|
|
2
|
-
"""
|
|
2
|
+
"""Submit one isolated Temporary Chat task through browser-harness."""
|
|
3
3
|
from __future__ import annotations
|
|
4
4
|
|
|
5
5
|
import argparse
|
|
@@ -9,23 +9,20 @@ import subprocess
|
|
|
9
9
|
import tempfile
|
|
10
10
|
|
|
11
11
|
|
|
12
|
-
BH_SCRIPT = os.path.join(os.path.dirname(os.path.abspath(__file__)), "
|
|
12
|
+
BH_SCRIPT = os.path.join(os.path.dirname(os.path.abspath(__file__)), "_temporary_bh.py")
|
|
13
13
|
|
|
14
14
|
|
|
15
15
|
def main() -> int:
|
|
16
16
|
parser = argparse.ArgumentParser()
|
|
17
|
-
parser.add_argument("--project", required=True)
|
|
18
17
|
parser.add_argument("--prompt", required=True)
|
|
19
|
-
parser.add_argument("--
|
|
18
|
+
parser.add_argument("--file", action="append", default=[])
|
|
20
19
|
args = parser.parse_args()
|
|
21
|
-
config = {"project": {"name": args.project}, "prompt": args.prompt, "thinking_level": args.thinking_level}
|
|
22
20
|
with tempfile.NamedTemporaryFile(mode="w", suffix=".json", delete=False) as handle:
|
|
23
|
-
json.dump(
|
|
21
|
+
json.dump(vars(args), handle, ensure_ascii=False)
|
|
24
22
|
config_path = handle.name
|
|
25
23
|
try:
|
|
26
24
|
code = open(BH_SCRIPT, encoding="utf-8").read().replace("__CFG_PATH__", config_path)
|
|
27
|
-
|
|
28
|
-
return result.returncode
|
|
25
|
+
return subprocess.run(["browser-harness"], input=code, text=True, timeout=180).returncode
|
|
29
26
|
finally:
|
|
30
27
|
os.unlink(config_path)
|
|
31
28
|
|
|
@@ -1,132 +0,0 @@
|
|
|
1
|
-
# ChatGPT browser-worker contract
|
|
2
|
-
|
|
3
|
-
This document is the stable boundary between Neo orchestration and a
|
|
4
|
-
`browser-harness` adapter. It is intentionally independent of ChatGPT's DOM,
|
|
5
|
-
URL layout, or internal network calls.
|
|
6
|
-
|
|
7
|
-
## Requests
|
|
8
|
-
|
|
9
|
-
All requests contain an operation and no credentials:
|
|
10
|
-
|
|
11
|
-
```json
|
|
12
|
-
{
|
|
13
|
-
"operation": "create",
|
|
14
|
-
"project": {"name": "neo", "id": "project-optional"},
|
|
15
|
-
"prompt": "Run the assigned task.",
|
|
16
|
-
"thinking_level": "high"
|
|
17
|
-
}
|
|
18
|
-
```
|
|
19
|
-
|
|
20
|
-
`create` requires `project.name`, `prompt`, and `thinking_level`.
|
|
21
|
-
`resume` requires `thread_id` and `project`; its thinking level is optional and
|
|
22
|
-
defaults to the persisted request. `status`, `result`, and `delete` require
|
|
23
|
-
only `thread_id`.
|
|
24
|
-
|
|
25
|
-
Allowed requested thinking levels are `default`, `low`, `medium`, and `high`.
|
|
26
|
-
The adapter may expose a current UI label in an observation, but it must map it
|
|
27
|
-
to one of these values or `unknown` rather than guessing.
|
|
28
|
-
|
|
29
|
-
## Durable thread state
|
|
30
|
-
|
|
31
|
-
The state file is the only persisted worker identity. It may look like:
|
|
32
|
-
|
|
33
|
-
```json
|
|
34
|
-
{
|
|
35
|
-
"schema_version": 1,
|
|
36
|
-
"thread_id": "chatgpt-conversation-id",
|
|
37
|
-
"conversation_url": "https://chatgpt.com/c/chatgpt-conversation-id",
|
|
38
|
-
"project": {"id": "project-id", "name": "neo"},
|
|
39
|
-
"status": "awaiting_result",
|
|
40
|
-
"requested_thinking_level": "high",
|
|
41
|
-
"effective_thinking_level": "high",
|
|
42
|
-
"created_at": "2026-09-19T12:00:00Z",
|
|
43
|
-
"updated_at": "2026-09-19T12:01:00Z",
|
|
44
|
-
"last_error": null
|
|
45
|
-
}
|
|
46
|
-
```
|
|
47
|
-
|
|
48
|
-
Required fields are `schema_version`, `thread_id`, `project.name`, `status`,
|
|
49
|
-
`requested_thinking_level`, and `effective_thinking_level`. `project.id` and
|
|
50
|
-
`conversation_url` are optional because the UI may not expose them at every
|
|
51
|
-
boundary, but the adapter must preserve them when observed.
|
|
52
|
-
|
|
53
|
-
The only valid statuses are:
|
|
54
|
-
|
|
55
|
-
| Status | Meaning |
|
|
56
|
-
| --- | --- |
|
|
57
|
-
| `created` | identity exists; no prompt has been sent yet |
|
|
58
|
-
| `running` | a prompt was sent and work is in progress |
|
|
59
|
-
| `awaiting_result` | the UI indicates a response may be read |
|
|
60
|
-
| `completed` | an assistant result was observed and normalized |
|
|
61
|
-
| `failed` | the operation failed; recovery may resume this identity |
|
|
62
|
-
| `blocked` | human/authentication/ambiguity decision is required |
|
|
63
|
-
| `deleted` | cleanup was verified; identity must not be reused |
|
|
64
|
-
|
|
65
|
-
Valid transitions are:
|
|
66
|
-
|
|
67
|
-
```text
|
|
68
|
-
create -> created -> running -> awaiting_result -> completed
|
|
69
|
-
| |
|
|
70
|
-
+-> failed +-> running (follow-up)
|
|
71
|
-
any live state -> blocked
|
|
72
|
-
any live state -> deleted (only after verified cleanup)
|
|
73
|
-
failed/blocked -> running (only after an explicit recovery operation)
|
|
74
|
-
```
|
|
75
|
-
|
|
76
|
-
`deleted` is terminal. A new conversation gets a new `thread_id`.
|
|
77
|
-
|
|
78
|
-
## Results
|
|
79
|
-
|
|
80
|
-
```json
|
|
81
|
-
{
|
|
82
|
-
"thread_id": "chatgpt-conversation-id",
|
|
83
|
-
"status": "completed",
|
|
84
|
-
"text": "The assistant's normalized final answer.",
|
|
85
|
-
"message_id": "observed-message-id",
|
|
86
|
-
"observed_at": "2026-09-19T12:02:00Z"
|
|
87
|
-
}
|
|
88
|
-
```
|
|
89
|
-
|
|
90
|
-
`text` is required only for `completed`. Partial or streaming text is not a
|
|
91
|
-
completed result. A result with no observed assistant message is an error, not
|
|
92
|
-
success.
|
|
93
|
-
|
|
94
|
-
## Browser adapter port
|
|
95
|
-
|
|
96
|
-
The adapter is tested through a fake port with these semantic calls:
|
|
97
|
-
|
|
98
|
-
```text
|
|
99
|
-
select_project(project) -> observed_project
|
|
100
|
-
open_thread(thread_id) -> observed_thread
|
|
101
|
-
set_thinking_level(level) -> observation
|
|
102
|
-
send_prompt(prompt) -> observation
|
|
103
|
-
read_status() -> status_observation
|
|
104
|
-
read_result() -> result_observation
|
|
105
|
-
delete_thread() -> cleanup_observation
|
|
106
|
-
```
|
|
107
|
-
|
|
108
|
-
`cleanup_observation` must contain the requested `thread_id` (or omit it only
|
|
109
|
-
when the browser has verified the thread is absent), `outcome` equal to
|
|
110
|
-
`deleted` or `not_found`, and `verified: true`. `not_found` is a successful
|
|
111
|
-
idempotent retry. Archive/undo controls may be reported as observational
|
|
112
|
-
metadata, but they do not turn a delete into a recoverable state: a deleted
|
|
113
|
-
tombstone remains terminal and must never be reused.
|
|
114
|
-
|
|
115
|
-
These names describe the boundary, not a required Python class or browser
|
|
116
|
-
selector implementation. Each call must return observed values or a typed
|
|
117
|
-
failure. The contract layer must remain usable with a fake port and must not
|
|
118
|
-
import or launch `browser-harness` itself.
|
|
119
|
-
|
|
120
|
-
## Testable safety boundaries
|
|
121
|
-
|
|
122
|
-
- no request can omit the Project on `create` or `resume`;
|
|
123
|
-
- Project IDs are compared when both browser boundaries expose them; a
|
|
124
|
-
name-only observation remains valid because the UI may hide opaque IDs;
|
|
125
|
-
- no state can omit a non-empty `thread_id` or valid status;
|
|
126
|
-
- `resume` rejects an observed Project mismatch;
|
|
127
|
-
- requested and effective thinking levels are distinct fields;
|
|
128
|
-
- unknown effective level stays `unknown`;
|
|
129
|
-
- only an observed assistant message yields `completed`;
|
|
130
|
-
- delete is terminal and cannot be followed by resume;
|
|
131
|
-
- browser, login, and ambiguity failures preserve the thread identity and are
|
|
132
|
-
classified as `failed` or `blocked`.
|