agents-relay 1.0.5 → 1.0.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,102 +1,62 @@
1
1
  ---
2
2
  name: chatgpt-browser-worker
3
- description: Manage ChatGPT work threads through the authenticated browser-harness session, including Project selection, thinking-level requests, durable thread identity, resume/status/result retrieval, and cleanup.
3
+ description: Submit an isolated one-shot ChatGPT browser worker task through browser-harness with a required durable output destination.
4
4
  ---
5
5
 
6
6
  # ChatGPT browser worker
7
7
 
8
- This skill is the browser-backed ChatGPT worker contract for Neo. It uses the
9
- existing `browser-harness` command as its only browser interaction layer. It
10
- does not use the MacBridge ChatGPT runtime, ChatGPT APIs, copied cookies, or a
11
- second browser automation stack.
12
-
13
- ## Capability contract
14
-
15
- The worker boundary exposes these lifecycle operations:
16
-
17
- ```text
18
- create(project, prompt, thinking_level) -> ThreadState
19
- resume(thread_id, project) -> ThreadState
20
- status(thread_id) -> StatusObservation
21
- result(thread_id) -> ThreadResult
22
- delete(thread_id) -> DeletedThreadState
23
- ```
24
-
25
- `project.name` is required when creating or resuming a thread. `project.id` is
26
- optional because the UI may not expose it at every boundary. The adapter must
27
- select and verify that Project before sending the first prompt. A resumed
28
- thread keeps its recorded Project identity; callers cannot silently move it to
29
- another Project.
30
-
31
- Thinking levels are split into `requested_thinking_level` and
32
- `effective_thinking_level`. Requested values are `default`, `low`, `medium`,
33
- and `high`; effective values may additionally be `unknown`. `default` leaves
34
- the account's current/default setting unchanged. The adapter must preserve
35
- `unknown` rather than guessing from a label or silently downgrading a request.
36
-
37
- ## Durable identity and state
38
-
39
- Persist the returned state after every successful lifecycle transition. The
40
- durable key is the ChatGPT `thread_id`; a URL, title, or browser tab index is
41
- not an identity. Records retain a deleted thread as a tombstone so a stale
42
- caller cannot recreate or accidentally reuse it. The complete JSON schema and
43
- valid transition table are in [references/contract.md](references/contract.md).
44
-
45
- Pure validation and transition helpers live in `scripts/contract.py`, with
46
- focused tests in `tests/test_contract.py`. The browser-harness driver is
47
- intentionally injected through the semantic port described in the reference;
48
- the contract layer does not import or launch `browser-harness`.
49
- Run `python3 -m pytest skills/chatgpt-browser-worker/tests` from the repository
50
- root.
51
-
52
- For the Neo/Leo and Agents Relay handoff—task IDs, run-local state, event
53
- ownership, and the create/resume/result/delete invocation sequence—see
54
- [references/orchestration.md](references/orchestration.md). The outer
55
- orchestrator may use MacBridge as a transport, but this worker never treats
56
- MacBridge as the ChatGPT runtime.
57
-
58
- The runtime entry point is the
59
- [browser-worker agent](agents/browser-worker.agent.md). Callers launch that
60
- agent with a typed lifecycle intent; they must not directly call
61
- `create_bh.py` or `operate_bh.py`. Those scripts are helper implementations
62
- that the agent may use, repair, or bypass while preserving this skill's
63
- browser-harness and verification contract.
64
-
65
- The browser-worker agent's `create` intent starts a new ChatGPT chat when the
66
- attached tab is already a conversation and emits JSON with the observed
67
- `thread_id`, conversation URL, selected Project, and thinking observation.
68
- `scripts/create_bh.py` is only a helper for that agent, not a caller-facing
69
- runtime contract.
70
-
71
- For implementation/debugging, the agent may use the helper equivalent for an
72
- existing thread with the durable identity from persisted state:
73
-
74
- ```text
75
- python3 scripts/operate_bh.py resume --thread-id ID --project NAME
76
- python3 scripts/operate_bh.py continue --thread-id ID --project NAME --prompt TEXT
77
- python3 scripts/operate_bh.py status --thread-id ID --project NAME
78
- python3 scripts/operate_bh.py result --thread-id ID --project NAME
79
- python3 scripts/operate_bh.py delete --thread-id ID --project NAME
80
- ```
81
-
82
- Each command performs one serial browser-harness operation. The injected
83
- semantic adapter in `scripts/operations.py` is the testable boundary; it does
84
- not import MacBridge or a ChatGPT runtime API.
85
-
86
- ## Verification boundary
87
-
88
- The worker may report `completed` only from an observed assistant message and
89
- normalized result. A process exit, click, URL change, or tab title alone is not
90
- proof that a thread was created, resumed, completed, or deleted. Authentication
91
- walls, MFA, consent, ambiguous account/project selection, and unverified
92
- thinking levels are `blocked` or `failed` conditions and must not be
93
- self-healed.
94
-
95
- Delete targets the exact durable conversation URL, requires an observed action
96
- and confirmation, and verifies that the requested thread is no longer visible.
97
- Repeated cleanup is idempotent (`not_found` is accepted); archive/undo UI
98
- controls are evidence only and never revive a deleted tombstone.
99
-
100
- Keep credentials, prompts containing private data, and browser runtime state
101
- out of durable state and events. Runtime artifacts belong under the caller's
102
- ignored `runs/` directory.
8
+ Use this skill to hand one self-contained task to ChatGPT through the authenticated
9
+ browser session. The caller does not manage ChatGPT threads, model selection,
10
+ thinking level, polling, or result retrieval.
11
+
12
+ ## Output contract
13
+
14
+ Every task handed to this worker MUST declare exactly one durable output mode
15
+ before the worker is launched.
16
+
17
+ ### 1. Task / PR output
18
+
19
+ Use this for repository or managed-task work.
20
+
21
+ The task prompt must identify the exact managed task/PR that owns the result.
22
+ The worker performs the requested work directly against that task/PR and records
23
+ its final outcome there.
24
+
25
+ Output:
26
+ task/PR
27
+ <exact managed task/PR identity and required final action>
28
+
29
+ ### 2. File output
30
+
31
+ Use this for research, analysis, reports, or other artifact-producing work.
32
+
33
+ The task prompt must name the exact file path. That file is the authoritative
34
+ result.
35
+
36
+ Output:
37
+ file
38
+ <exact file path>
39
+
40
+ A Browser ChatGPT worker task without one of these output declarations is
41
+ invalid and must not be launched. The orchestrator chooses and writes the output
42
+ contract when it creates the worker task; the worker must not invent or change
43
+ the destination.
44
+
45
+ Progress and terminal success/failure are always reported through the normal
46
+ Neo event contract. Events are execution observability, not a third output mode.
47
+
48
+ ## Runtime behavior
49
+
50
+ The worker is one-shot. It opens an isolated worker-owned ChatGPT tab, submits
51
+ the complete task, verifies that ChatGPT accepted the submission, closes its
52
+ owned tab, and returns. ChatGPT continues the task independently and publishes
53
+ the result through the declared output contract.
54
+
55
+ The worker must never interact through a pre-existing user ChatGPT tab.
56
+ Browser details, transient conversation identity, submission verification,
57
+ model/default-thinking behavior, and tab cleanup are implementation concerns of
58
+ the worker and are not part of the caller-facing contract.
59
+
60
+ The runtime entry point is the browser-worker agent. Callers launch that agent
61
+ with the complete task prompt and declared output; they do not call helper
62
+ scripts directly.
@@ -1,113 +1,42 @@
1
1
  ---
2
2
  name: browser-worker
3
- description: Execute the ChatGPT browser-worker lifecycle through browser-harness while preserving verified thread and Project identity.
3
+ description: Submit one isolated ChatGPT browser task and return after the task is accepted.
4
4
  type: worker
5
5
  ---
6
6
 
7
7
  # Browser Worker Agent
8
8
 
9
- You are the runtime owner for the `chatgpt-browser-worker` skill. Before doing
10
- any work, load and follow both this skill and the `browser-harness` skill. The
11
- outer adapter launches you with one typed lifecycle request; Agents Relay
12
- owns task correlation, retries, events, and run-local persistence, but does
13
- not own browser execution.
9
+ You are the runtime owner for the `chatgpt-browser-worker` skill. Load and follow
10
+ that skill and the `browser-harness` skill before execution.
14
11
 
15
- ## Runtime rules
16
-
17
- - Use `browser-harness` for every browser interaction. Do not use another
18
- browser stack, ChatGPT API, MacBridge ChatGPT runtime, copied cookies, or
19
- direct HTTP calls to ChatGPT.
20
- - Preserve the exact observed ChatGPT `thread_id` and Project identity. A URL,
21
- title, tab index, or inferred conversation is not an identity. A resumed
22
- thread must remain in its recorded Project; reject a mismatch.
23
- - Treat the scripts in `scripts/` as helpers only. You may use, repair, or
24
- bypass `create_bh.py` and `operate_bh.py` when they are brittle, but preserve
25
- the skill contract and its evidence requirements.
26
- - When a selector, layout, label, or ordinary UI flow changes, re-observe the
27
- current page through `browser-harness`, identify the semantic control, and
28
- adapt the operation. Do not guess from stale selectors or coordinates.
29
- - Stop with `blocked` for authentication, MFA, consent, ambiguous account or
30
- Project selection, or any human decision. Do not bypass these gates.
31
- - Stop with `blocked` or `failed` when the result cannot be verified. A process
32
- exit, URL change, tab title, prior assistant bubble, or synthetic message ID
33
- is not completion evidence.
34
- - Return only observations from the current browser state and persist the last
35
- known `thread_id` when an operation fails after creation.
36
-
37
- ## Typed adapter contract
38
-
39
- The input is one JSON object. `operation` must be one of `create`, `resume`,
40
- `continue`, `status`, `result`, or `delete`.
41
-
42
- ```json
43
- {
44
- "operation": "create",
45
- "thread_id": null,
46
- "project": {"name": "neo", "id": "project-optional"},
47
- "prompt": "Run the assigned task.",
48
- "thinking_level": "high",
49
- "state": null
50
- }
51
- ```
12
+ The outer orchestrator gives you one complete task. That task MUST already
13
+ contain exactly one output declaration defined by the skill:
52
14
 
53
- Input requirements:
15
+ - `task/PR` with the exact managed task/PR identity and required final action; or
16
+ - `file` with the exact authoritative output path.
54
17
 
55
- - `create`: requires `project.name`, `prompt`, and `thinking_level`.
56
- - `resume`: requires `thread_id` and `project.name`; verify the recorded
57
- Project before any other thread action.
58
- - `continue`: requires `thread_id`, `project.name`, and `prompt`; open and
59
- verify the exact thread before sending the follow-up.
60
- - `status`, `result`, and `delete`: require the durable `thread_id`; use the
61
- persisted `state` when supplied and never recreate a missing identity.
62
- - `thinking_level` is `default`, `low`, `medium`, or `high`; preserve an
63
- observed effective value of `unknown` instead of guessing.
18
+ If the output declaration is missing or ambiguous, do not invent one.
64
19
 
65
- The output is one JSON object suitable for an outer adapter:
66
-
67
- ```json
68
- {
69
- "operation": "result",
70
- "status": "completed",
71
- "thread_id": "chatgpt-conversation-id",
72
- "project": {"name": "neo", "id": "project-id"},
73
- "result": {
74
- "message_id": "observed-assistant-message-id",
75
- "text": "The normalized assistant answer.",
76
- "verified": true,
77
- "observed_at": "2026-09-19T12:02:00Z"
78
- },
79
- "state": {
80
- "schema_version": 1,
81
- "thread_id": "chatgpt-conversation-id",
82
- "status": "completed",
83
- "project": {"name": "neo", "id": "project-id"}
84
- },
85
- "error": null
86
- }
87
- ```
88
-
89
- Every response must include `operation`, `status`, and the exact
90
- `thread_id` when one is known. A successful `result` response must include an
91
- observed assistant `message_id`, non-empty normalized `text`, and
92
- `result.verified: true`. `status` is not a result and must not claim
93
- completion. For `delete`, return a verified `deleted` or idempotent verified
94
- `not_found` observation and retain a terminal `deleted` tombstone.
95
-
96
- For blocked or failed operations, return the last known identity and a typed
97
- error without credentials or private prompts:
98
-
99
- ```json
100
- {
101
- "operation": "resume",
102
- "status": "blocked",
103
- "thread_id": "chatgpt-conversation-id",
104
- "error": {
105
- "code": "authentication_required",
106
- "message": "Sign-in or MFA requires user action.",
107
- "retryable": false
108
- }
109
- }
110
- ```
20
+ ## Runtime rules
111
21
 
112
- Do not report success until the relevant browser observation has been
113
- validated against [../references/contract.md](../references/contract.md).
22
+ - Use `browser-harness` for every ChatGPT browser interaction.
23
+ - Execute only a one-shot submission. Do not expose or operate a
24
+ create/resume/status/result/continue lifecycle for callers.
25
+ - Always create a fresh worker-owned tab. Never type, upload, click New chat,
26
+ select a Project, or submit through a pre-existing user ChatGPT tab.
27
+ - Use Temporary Chat and the account's existing default model/thinking
28
+ settings. Do not change model or thinking settings.
29
+ - Upload requested files, submit the complete task prompt, and verify that the
30
+ submission became a new user turn.
31
+ - Once submission is verified, close the worker-owned tab and return. Do not
32
+ wait for the assistant response and do not reopen or poll the conversation.
33
+ - Treat any observed conversation/thread identity only as diagnostic evidence,
34
+ never as a resumable handle.
35
+ - Progress and terminal success/failure for the actual delegated task are
36
+ reported through the normal Neo event contract by the executing task. The
37
+ final durable result goes to the declared task/PR or file output.
38
+ - Stop on authentication, MFA, consent, or ambiguous browser state rather than
39
+ bypassing it.
40
+
41
+ Helper scripts under `scripts/` are implementation details. Callers must not
42
+ invoke them directly.
@@ -0,0 +1,171 @@
1
+ import atexit
2
+ import json
3
+ import os
4
+ import re
5
+ import time
6
+
7
+ from browser_harness import *
8
+
9
+
10
+ CFG = json.load(open("__CFG_PATH__", encoding="utf-8"))
11
+ _OWNED_TABS = []
12
+
13
+
14
+ def _new_owned_tab(url):
15
+ target_id = cdp("Target.createTarget", url="about:blank", background=True)["targetId"]
16
+ switch_tab(target_id)
17
+ _OWNED_TABS.append(target_id)
18
+ if url != "about:blank":
19
+ goto_url(url)
20
+ return target_id
21
+
22
+
23
+ def _close_owned_tabs():
24
+ while _OWNED_TABS:
25
+ try:
26
+ close_tab(_OWNED_TABS.pop())
27
+ except Exception:
28
+ pass
29
+
30
+
31
+ atexit.register(_close_owned_tabs)
32
+
33
+
34
+ def _composer():
35
+ selector = js("""(() => {
36
+ const preferred = document.querySelector('#prompt-textarea');
37
+ if (preferred) {
38
+ const r=preferred.getBoundingClientRect();
39
+ if (r.width>0 && r.height>0) return '#prompt-textarea';
40
+ }
41
+ const fallback = [...document.querySelectorAll('textarea,[contenteditable="true"]')]
42
+ .find(e => { const r=e.getBoundingClientRect(); return r.width>0 && r.height>0 && !e.disabled; });
43
+ if (!fallback) return null;
44
+ return fallback.tagName === 'TEXTAREA' ? 'textarea' : '[contenteditable="true"]';
45
+ })()""")
46
+ if not selector:
47
+ raise RuntimeError("Temporary Chat composer was not observed")
48
+ return selector
49
+
50
+
51
+ def _composer_text(selector):
52
+ return js(f"""(() => {{
53
+ const e=document.querySelector({json.dumps(selector)});
54
+ return e ? ((e.innerText ?? e.value) || '') : '';
55
+ }})()""") or ""
56
+
57
+
58
+ def _attachment_names():
59
+ return js(r"""(() => [...document.querySelectorAll('button[aria-label^="Remove file"]')]
60
+ .map(b => (b.getAttribute('aria-label') || '').replace(/^Remove file\s+\d+:\s*/, ''))
61
+ .filter(Boolean))()""") or []
62
+
63
+
64
+ def _upload_files(paths):
65
+ if not paths:
66
+ return []
67
+ selector = js("""(() => {
68
+ for (const s of ['#upload-files','#upload-media','input[name="upload-media"]','input[type="file"]'])
69
+ if (document.querySelector(s)) return s;
70
+ return null;
71
+ })()""")
72
+ if not selector:
73
+ raise RuntimeError("ChatGPT file input was not observed")
74
+ for path in paths:
75
+ if not os.path.isfile(path):
76
+ raise RuntimeError(f"attachment does not exist: {path}")
77
+ upload_file(selector, path)
78
+ expected = os.path.basename(path)
79
+ deadline = time.time() + 20
80
+ while time.time() < deadline and expected not in _attachment_names():
81
+ time.sleep(.25)
82
+ if expected not in _attachment_names():
83
+ raise RuntimeError(f"attachment was not observed ready: {expected}")
84
+ return [os.path.basename(path) for path in paths]
85
+
86
+
87
+ def _user_turns():
88
+ return js("""(() => [...document.querySelectorAll('[data-message-author-role="user"]')]
89
+ .map(e => ({text:e.innerText.trim(), id:e.getAttribute('data-message-id') || ''}))
90
+ .filter(e => e.text))()""") or []
91
+
92
+
93
+ def _click_send():
94
+ ok = js("""(() => {
95
+ const b=document.querySelector('button[data-testid="send-button"],button[aria-label="Send prompt"]');
96
+ if (!b || b.disabled || b.getAttribute('aria-disabled') === 'true') return false;
97
+ b.click(); return true;
98
+ })()""")
99
+ if not ok:
100
+ raise RuntimeError("Temporary Chat Send prompt button was not observed ready")
101
+
102
+
103
+ def _wait_user_turn(before_count, prompt, timeout=20):
104
+ deadline = time.time() + timeout
105
+ while time.time() < deadline:
106
+ turns = _user_turns()
107
+ if len(turns) > before_count and prompt.strip() in turns[-1]["text"]:
108
+ return turns[-1]
109
+ time.sleep(.25)
110
+ raise RuntimeError("Temporary Chat prompt did not become an observed user turn")
111
+
112
+
113
+ def _diagnostic_thread_id():
114
+ match = re.search(r"/c/([^/?#]+)", page_info().get("url", ""))
115
+ return match.group(1) if match else None
116
+
117
+
118
+ _new_owned_tab("https://chatgpt.com/")
119
+ wait_for_load()
120
+
121
+ enabled = js("""(() => {
122
+ const matches=[...document.querySelectorAll('button')].filter(
123
+ b => (b.getAttribute('aria-label') || '').trim() === 'Temporary chat'
124
+ );
125
+ if (matches.length !== 1) return false;
126
+ matches[0].click();
127
+ return true;
128
+ })()""")
129
+ if not enabled:
130
+ raise RuntimeError("Temporary Chat toggle was not uniquely observed")
131
+
132
+ deadline = time.time() + 20
133
+ while time.time() < deadline:
134
+ if "temporary-chat=true" in page_info().get("url", ""):
135
+ try:
136
+ _composer()
137
+ break
138
+ except RuntimeError:
139
+ pass
140
+ time.sleep(.25)
141
+ else:
142
+ raise RuntimeError("Temporary Chat mode did not become ready")
143
+
144
+ attachments = _upload_files(CFG.get("file", []))
145
+ selector = _composer()
146
+ before_count = len(_user_turns())
147
+ fill_input(selector, CFG["prompt"], clear_first=True)
148
+ if CFG["prompt"].strip() not in _composer_text(selector):
149
+ info = js(f"""(() => {{
150
+ const e=document.querySelector({json.dumps(selector)});
151
+ const r=e.getBoundingClientRect();
152
+ return {{x:r.x+r.width/2,y:r.y+r.height/2}};
153
+ }})()""")
154
+ click_at_xy(info["x"], info["y"])
155
+ press_key("CTRL+A")
156
+ type_text(CFG["prompt"])
157
+ if CFG["prompt"].strip() not in _composer_text(selector):
158
+ raise RuntimeError("Temporary Chat prompt was not observed in the composer")
159
+
160
+ _click_send()
161
+ user_turn = _wait_user_turn(before_count, CFG["prompt"])
162
+
163
+ print(json.dumps({
164
+ "operation": "submit",
165
+ "status": "submitted",
166
+ "temporary": True,
167
+ "attachments": attachments,
168
+ "diagnostic_thread_id": _diagnostic_thread_id(),
169
+ "user_message_id": user_turn.get("id") or None,
170
+ "verified": True,
171
+ }, ensure_ascii=False))
@@ -1,5 +1,5 @@
1
1
  #!/usr/bin/env python3
2
- """Create one ChatGPT thread using the authenticated browser-harness session."""
2
+ """Submit one isolated Temporary Chat task through browser-harness."""
3
3
  from __future__ import annotations
4
4
 
5
5
  import argparse
@@ -9,23 +9,20 @@ import subprocess
9
9
  import tempfile
10
10
 
11
11
 
12
- BH_SCRIPT = os.path.join(os.path.dirname(os.path.abspath(__file__)), "_create_bh.py")
12
+ BH_SCRIPT = os.path.join(os.path.dirname(os.path.abspath(__file__)), "_temporary_bh.py")
13
13
 
14
14
 
15
15
  def main() -> int:
16
16
  parser = argparse.ArgumentParser()
17
- parser.add_argument("--project", required=True)
18
17
  parser.add_argument("--prompt", required=True)
19
- parser.add_argument("--thinking-level", choices=("default", "low", "medium", "high"), default="default")
18
+ parser.add_argument("--file", action="append", default=[])
20
19
  args = parser.parse_args()
21
- config = {"project": {"name": args.project}, "prompt": args.prompt, "thinking_level": args.thinking_level}
22
20
  with tempfile.NamedTemporaryFile(mode="w", suffix=".json", delete=False) as handle:
23
- json.dump(config, handle, ensure_ascii=False)
21
+ json.dump(vars(args), handle, ensure_ascii=False)
24
22
  config_path = handle.name
25
23
  try:
26
24
  code = open(BH_SCRIPT, encoding="utf-8").read().replace("__CFG_PATH__", config_path)
27
- result = subprocess.run(["browser-harness"], input=code, text=True, timeout=180)
28
- return result.returncode
25
+ return subprocess.run(["browser-harness"], input=code, text=True, timeout=180).returncode
29
26
  finally:
30
27
  os.unlink(config_path)
31
28
 
@@ -1,132 +0,0 @@
1
- # ChatGPT browser-worker contract
2
-
3
- This document is the stable boundary between Neo orchestration and a
4
- `browser-harness` adapter. It is intentionally independent of ChatGPT's DOM,
5
- URL layout, or internal network calls.
6
-
7
- ## Requests
8
-
9
- All requests contain an operation and no credentials:
10
-
11
- ```json
12
- {
13
- "operation": "create",
14
- "project": {"name": "neo", "id": "project-optional"},
15
- "prompt": "Run the assigned task.",
16
- "thinking_level": "high"
17
- }
18
- ```
19
-
20
- `create` requires `project.name`, `prompt`, and `thinking_level`.
21
- `resume` requires `thread_id` and `project`; its thinking level is optional and
22
- defaults to the persisted request. `status`, `result`, and `delete` require
23
- only `thread_id`.
24
-
25
- Allowed requested thinking levels are `default`, `low`, `medium`, and `high`.
26
- The adapter may expose a current UI label in an observation, but it must map it
27
- to one of these values or `unknown` rather than guessing.
28
-
29
- ## Durable thread state
30
-
31
- The state file is the only persisted worker identity. It may look like:
32
-
33
- ```json
34
- {
35
- "schema_version": 1,
36
- "thread_id": "chatgpt-conversation-id",
37
- "conversation_url": "https://chatgpt.com/c/chatgpt-conversation-id",
38
- "project": {"id": "project-id", "name": "neo"},
39
- "status": "awaiting_result",
40
- "requested_thinking_level": "high",
41
- "effective_thinking_level": "high",
42
- "created_at": "2026-09-19T12:00:00Z",
43
- "updated_at": "2026-09-19T12:01:00Z",
44
- "last_error": null
45
- }
46
- ```
47
-
48
- Required fields are `schema_version`, `thread_id`, `project.name`, `status`,
49
- `requested_thinking_level`, and `effective_thinking_level`. `project.id` and
50
- `conversation_url` are optional because the UI may not expose them at every
51
- boundary, but the adapter must preserve them when observed.
52
-
53
- The only valid statuses are:
54
-
55
- | Status | Meaning |
56
- | --- | --- |
57
- | `created` | identity exists; no prompt has been sent yet |
58
- | `running` | a prompt was sent and work is in progress |
59
- | `awaiting_result` | the UI indicates a response may be read |
60
- | `completed` | an assistant result was observed and normalized |
61
- | `failed` | the operation failed; recovery may resume this identity |
62
- | `blocked` | human/authentication/ambiguity decision is required |
63
- | `deleted` | cleanup was verified; identity must not be reused |
64
-
65
- Valid transitions are:
66
-
67
- ```text
68
- create -> created -> running -> awaiting_result -> completed
69
- | |
70
- +-> failed +-> running (follow-up)
71
- any live state -> blocked
72
- any live state -> deleted (only after verified cleanup)
73
- failed/blocked -> running (only after an explicit recovery operation)
74
- ```
75
-
76
- `deleted` is terminal. A new conversation gets a new `thread_id`.
77
-
78
- ## Results
79
-
80
- ```json
81
- {
82
- "thread_id": "chatgpt-conversation-id",
83
- "status": "completed",
84
- "text": "The assistant's normalized final answer.",
85
- "message_id": "observed-message-id",
86
- "observed_at": "2026-09-19T12:02:00Z"
87
- }
88
- ```
89
-
90
- `text` is required only for `completed`. Partial or streaming text is not a
91
- completed result. A result with no observed assistant message is an error, not
92
- success.
93
-
94
- ## Browser adapter port
95
-
96
- The adapter is tested through a fake port with these semantic calls:
97
-
98
- ```text
99
- select_project(project) -> observed_project
100
- open_thread(thread_id) -> observed_thread
101
- set_thinking_level(level) -> observation
102
- send_prompt(prompt) -> observation
103
- read_status() -> status_observation
104
- read_result() -> result_observation
105
- delete_thread() -> cleanup_observation
106
- ```
107
-
108
- `cleanup_observation` must contain the requested `thread_id` (or omit it only
109
- when the browser has verified the thread is absent), `outcome` equal to
110
- `deleted` or `not_found`, and `verified: true`. `not_found` is a successful
111
- idempotent retry. Archive/undo controls may be reported as observational
112
- metadata, but they do not turn a delete into a recoverable state: a deleted
113
- tombstone remains terminal and must never be reused.
114
-
115
- These names describe the boundary, not a required Python class or browser
116
- selector implementation. Each call must return observed values or a typed
117
- failure. The contract layer must remain usable with a fake port and must not
118
- import or launch `browser-harness` itself.
119
-
120
- ## Testable safety boundaries
121
-
122
- - no request can omit the Project on `create` or `resume`;
123
- - Project IDs are compared when both browser boundaries expose them; a
124
- name-only observation remains valid because the UI may hide opaque IDs;
125
- - no state can omit a non-empty `thread_id` or valid status;
126
- - `resume` rejects an observed Project mismatch;
127
- - requested and effective thinking levels are distinct fields;
128
- - unknown effective level stays `unknown`;
129
- - only an observed assistant message yields `completed`;
130
- - delete is terminal and cannot be followed by resume;
131
- - browser, login, and ambiguity failures preserve the thread identity and are
132
- classified as `failed` or `blocked`.