@awebai/oats 0.39.4 → 0.40.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -47,8 +47,8 @@ override; report the risk you accepted). It never touches the checkout it
47
47
  runs from: it exports the SHA into a detached worktree under the system
48
48
  temporary directory (recorded in `MANIFEST.json`, `--export <dir>` overrides)
49
49
  and runs every build step there. The bumped manifests exist only in that
50
- export; the version-bump commit to `main` remains the workflow's job, or a
51
- manual PR.
50
+ export; the version-bump commit reaches `main` through the pull request the
51
+ workflow opens (or a manual one), merged by a maintainer.
52
52
 
53
53
  `publish-npm` and `release-github` print their plan and refuse without
54
54
  `--yes`. `tag` creates the local tag without `--yes` but pushes only with
@@ -117,8 +117,12 @@ Everything after `build` reads `MANIFEST.json` and the files already staged:
117
117
  runs `release.yml`, whose steps are idempotent: it skips the live npm
118
118
  versions, re-uploads the same assets, and attaches the attestations. That
119
119
  later pass is the way to add provenance; nothing is republished.
120
- - **The version-bump PR.** The workflow's final step; open it by hand if the
121
- workflow does not run.
120
+ - **The version-bump PR.** The workflow's final step opens it and stops: the
121
+ run never merges into `main`. A maintainer reviews that the diff is the
122
+ version lines only and merges it. The PR is opened with the workflow token,
123
+ so no checks run on it by themselves; close and reopen it to run them, and
124
+ do not read missing checks as green. Open the PR by hand if the workflow
125
+ does not run.
122
126
  - **Legs for hosts you do not have.** The Linux AppImage/DEB need a Linux
123
127
  host; the lane says so and `stage` lists what is missing.
124
128
 
@@ -0,0 +1,181 @@
1
+ # OATS 0.40.0
2
+
3
+ ## Added
4
+
5
+ - **"Needs input": an instance can say it is blocked on a human** (feature
6
+ `waiting-on-you`). The `waitingOnYou` fact that `oats instance events` has
7
+ carried since 0.24.12 finally has producers. It shows on every running
8
+ instance row of `oats status --json` and on `oats session inspect --json`
9
+ as `waitingOnYou: {since, producer, reason, message} | null`, and plain
10
+ `oats status` prints `! needs input (<reason>): <message>` under the row.
11
+ Claims are display-only: nothing in the kernel acts on them, and
12
+ `oats session input` is unchanged. See docs/desktop-cli-api.md, "Waiting
13
+ on you".
14
+ - `oats instance waiting <set|clear> --producer <id> [--reason
15
+ permission|question|attention] [--message <text>]` records a producer's
16
+ claim as a `waiting` event, appended only when the claim changes.
17
+ - `oats instance attention [--message "<one line>"] [--clear]` is the
18
+ agent's own claim, run from its home (`$OATS_INSTANCE_HOME` only). The
19
+ oats.core instructions now teach it: when an agent has asked a human
20
+ something and cannot continue without the answer, it runs
21
+ `oats instance attention --message "…"` and ends its turn, then
22
+ `--clear` once it has the answer. Only the agent or a session boundary
23
+ clears an agent claim.
24
+ - `--message` is one line of 1 to 200 characters. Refused, at write and
25
+ again at read (a refused stored message reads as `null`): control
26
+ characters, U+2028/2029, the bidi controls U+202A–202E and
27
+ U+2066–2069, U+200B, U+2060, U+FEFF and the tag characters
28
+ U+E0000–E007F. Everything else is allowed, ZWJ/ZWNJ (U+200C/D) and
29
+ LRM/RLM/ALM (U+200E/F, U+061C) included. The Desktop uses the same set.
30
+ - A claim is live only after the incarnation's latest kernel `launched`,
31
+ `restarted` or `stopped` row, so a crashed or restarted session never
32
+ shows a stale claim. `waitingOnYou` and `waitingClaims[]` entries gain
33
+ `message` (a string or `null`); the shape is otherwise unchanged.
34
+
35
+ - **oats.core 2.4.0 (oats.framework 1.6.0) reports Claude Code permission
36
+ prompts and questions.** For a Claude instance, oats.core manages its own
37
+ hook entries in `<home>/.claude/settings.json`. A permission prompt sets
38
+ `waitingOnYou` with reason `permission`; AskUserQuestion or an MCP
39
+ elicitation sets `question`; the claim clears when the session moves on.
40
+ The hooks always exit 0, print nothing and time out in seconds, so a failing
41
+ emitter means "unknown", never a blocked tool. Every other key and entry in
42
+ that file is left alone. oats.core now requires OATS 0.40.0. See
43
+ docs/capabilities.md, "oats.core: needs input".
44
+ - **Desktop: the sidebar shows which agents are waiting on you.** When an
45
+ agent is waiting for a tool approval, a question or an attention request,
46
+ its row shows a "Needs input" mark next to its name. Hover over the row, or
47
+ focus it from the keyboard, to see what it is waiting for, its message and
48
+ how long it has waited. A collapsed parent shows "N below" for waiting agents
49
+ hidden under it, and an open terminal tab shows the mark as well. The mark
50
+ appears only while the agent is running and the installed OATS reports
51
+ `waiting-on-you`. It goes on the next roster refresh after the wait ends.
52
+ - **Schedules take an optional `description`**
53
+ ([#546](https://github.com/awebai/oats/issues/546)): what a job is for, in
54
+ words. `oats schedule add|update --file` accepts it as one line of 1 to 200
55
+ characters with no control characters; anything else is
56
+ `E_SCHEDULE_INVALID` naming `field: "description"`. It is stored as given
57
+ and returned by `schedule list --json` and `show --json` in each row's
58
+ `description` (`null` when there is none), which is how the Desktop reads
59
+ it. It is informational only and never reaches a run. Jobs without one are
60
+ unchanged. Earlier kernels drop the key without refusing it. See
61
+ [schedules](../schedules.md#kinds).
62
+ - **`oats doctor` warns about unresolved schedule attempts**
63
+ (`schedule-unresolved`, [#545](https://github.com/awebai/oats/issues/545)).
64
+ This covers each job, in this deployment and in the other deployments this
65
+ host ticks, whose launch has no recorded result. The warning gives the job,
66
+ the attempt's scheduledFor and startedAt and its age, whether it holds a
67
+ host slot, the first error, and the remedy: `oats schedule reconcile <id>`,
68
+ with `--clear` when a command's effects can't be proven. In `--json` it is
69
+ a `problems[]` item with `severity: "warning"`; the exit status is
70
+ unchanged. See [schedules](../schedules.md#what-a-run-reports).
71
+
72
+ ## Changed
73
+
74
+ - **An unknown command or operation run no longer stalls the host's other
75
+ jobs** ([#545](https://github.com/awebai/oats/issues/545)). Once the kernel
76
+ has observed the command's process exit (it returned, or was stopped at the
77
+ timeout), its unresolved attempt stops counting against `maxConcurrent`.
78
+ The job itself still waits for `oats schedule reconcile`, so it never runs
79
+ twice on unproven effects. Its attempt now records `exited`, `exitStatus`
80
+ and `exitSignal`, and `show` reports `running: false`. Spawn jobs keep
81
+ their slot, and so does an attempt recorded by an earlier kernel, until
82
+ reconcile.
83
+ - **An unresolved attempt keeps its first error.** Later due ticks used to
84
+ replace the cause ("command timed out…", "answered no valid envelope…",
85
+ "command outcome unconfirmed: …") with "launch attempt without a recorded
86
+ result" in `lastRun` and its history row. The cause is now kept on the
87
+ attempt (`attempt.error`) and in `lastRun`.
88
+ - **Schedule IDs may contain up to 100 characters**, locally and in workspace
89
+ definitions. Trigger IDs remain limited to 40. Spawn schedules still need
90
+ a short enough `purpose` to fit the 64-character instance-name limit.
91
+ - **Desktop timing edits preserve a schedule's description**, including its
92
+ exact whitespace and Unicode text. Existing jobs without descriptions
93
+ remain editable.
94
+ - **oats.engineering 1.8.1** (catalog and workspace pin, and the bundled
95
+ mirrors): a PR is reviewed only after its owner says the developer's review
96
+ loop converged. `/pr-review` starts on the owner's hand-over (the loop's
97
+ final verdict and rounds, the exact head, CI green on it), and a head that
98
+ moves afterwards waits for the next hand-over, after which only the delta is
99
+ reviewed and the verdict re-bound; a PR with no owning expert is ready when
100
+ its author marks it ready or asks for review at a named head.
101
+ Experts verify only converged branches and hand PRs over with all three
102
+ facts; developers report convergence, with the head, instead of
103
+ intermediate heads.
104
+ - **The Desktop accepts OATS CLIs `>=0.25.8 <0.41.0`**, so it runs against
105
+ this release's kernel; the Desktop 0.39.x refuses a 0.40 CLI.
106
+ - **`oats session start` and `restart` now record a `launched` event**
107
+ (`data.phase: "start"` or `"restart"`, `data.startId`), as spawn already
108
+ did, once per start and in each log: a start recovered from an interrupted
109
+ one's receipt adds none, copies the same row into a log that missed it, or
110
+ writes one with `phase: "recovered"` dated at its launch when the
111
+ interrupted start had written it to neither log. A consumer that counts a home's
112
+ events sees one more row per start.
113
+ - **oats.framework 1.6.0** (oats.core 2.4.0; catalog and workspace pin
114
+ `oats-framework/v1.6.0`): the package that carries the "needs input"
115
+ emitter and the `oats instance attention` protocol above. A workspace's
116
+ `oats.framework: v1.6.0` resolves to that tag through the official
117
+ catalog; workspaces pinned to v1.5.0 keep oats.core 2.3.0 until they move
118
+ the pin and sync.
119
+ - **Desktop: collapsing a row hides only rows in its own deployment
120
+ section.** A row whose parent name matched an agent in another deployment
121
+ was hidden when that agent was collapsed. It now stays visible.
122
+
123
+ ## Fixed
124
+
125
+ - **A scheduled command that ignores SIGTERM can no longer wedge the host
126
+ scheduler** ([#547](https://github.com/awebai/oats/issues/547)). Command and
127
+ operation runs, workspace-spawn launches and spawn previews now send SIGTERM
128
+ at the five-minute timeout, then SIGKILL to the remaining process group after
129
+ two seconds. The runner observes the direct child's exit and bounds output
130
+ pipe draining before returning. Timed-out commands remain unknown and wait
131
+ for reconcile, while their exited processes no longer hold host slots;
132
+ spawn jobs retain their existing slot semantics.
133
+
134
+ - **Missing harness-package diagnostics give usable direct install commands.**
135
+ Spawn and restart preflight no longer suggest the removed `oats install`
136
+ verb. Guidance names the operator's responsibility, quotes the selected
137
+ executable and context paths, preserves Pi's resource directory, and keeps
138
+ Claude marketplace registration before plugin installation. OATS still
139
+ verifies requirements without installing packages.
140
+
141
+ ## Notes
142
+
143
+ - **Claude Code honours only the last `--settings` flag.** On Claude Code
144
+ 2.1.288 a second `--settings` replaces the first wholesale, even when it
145
+ has no hooks. That is why oats.core writes the project
146
+ `<home>/.claude/settings.json`, which composes with the user's settings and
147
+ with a `--settings`. Any capability that needs Claude settings must manage
148
+ its own entries in that file, under its own marker, and never pass
149
+ `--settings`.
150
+ - **Claude instances spawned before the upgrade get the emitter only on
151
+ respawn.** A home's launch hook comes from its recorded module copy, so a
152
+ restart keeps the old oats.core. The `attention` verb works on any harness
153
+ as soon as the kernel is upgraded.
154
+ - If a user's Claude configuration sets `disableAllHooks` or
155
+ `allowManagedHooksOnly`, the emitter never runs and the field stays `null`.
156
+ - Refusing a Claude Code permission prompt fires no hook (Claude Code
157
+ 2.1.288), so after a "No" the claim stays, labelled `permission`, until the
158
+ human's next prompt; Claude is waiting for that prompt anyway.
159
+ - The emitter debounces only the per-tool-call clears; `UserPromptSubmit`,
160
+ `Stop` and `SessionEnd` are not debounced. A turn-boundary clear gets a
161
+ CLI call: from that hook, or, if another hook holds the lock, from that
162
+ holder if it still has time; otherwise from the next event. That bounds
163
+ what an unfenced kernel write can leave wrong
164
+ ([#568](https://github.com/awebai/oats/issues/568)): after a hook
165
+ suspended past 5 s mid-call (the machine slept) or a reaper killed at a
166
+ precise instant, a wrongly shown claim lasts until the end of the turn,
167
+ or, if the turn's last hook found a reconciliation out of time, until the
168
+ next event (the human's next prompt); a hidden question lasts until the
169
+ human answers it. A permission prompt can also stay hidden until it is
170
+ answered without any of that: when it opens while another hook's
171
+ reconciliation runs out of time (a slow clear on a loaded machine).
172
+ Parallel main-thread tool calls may clear a claim early. See
173
+ docs/capabilities.md, oats.core: needs input, Limits.
174
+ - Tool calls by background or parallel subagents in the same Claude session
175
+ do not clear the claim: the emitter skips tool events that carry a
176
+ subagent's `agent_id`. A permission prompt never says who asked, so after
177
+ a human approves a subagent's prompt the claim stays until the next
178
+ main-thread event (its next tool call, the subagent's completion, a stop or
179
+ a prompt). For a foreground subagent the main thread makes no call until
180
+ the subagent finishes, so the claim can last the whole subagent run: too
181
+ long, never hiding a real block.
@@ -0,0 +1,20 @@
1
+ # OATS 0.40.1
2
+
3
+ 0.40.0 was tagged but never published because its Desktop build failed in the release workflow. 0.40.1 is the first published 0.40 release and includes everything in the [0.40.0 notes](v0.40.0.md).
4
+
5
+ ## Changed
6
+
7
+ - **Scheduled jobs default to a host concurrency of five.** The registry stores
8
+ explicit choices only. Legacy stored one migrates to the default once; an
9
+ explicit `oats schedule host install --max-concurrent 1` after upgrade stays
10
+ one. Use `--max-concurrent N|default` and the independent
11
+ `--triggers-max-concurrent N|none` to configure caps through the CLI
12
+ ([#550](https://github.com/awebai/oats/issues/550)). Host status reports both
13
+ caps; triggers remain uncapped by the host unless explicitly limited.
14
+
15
+ ## Fixed
16
+
17
+ - **Desktop tests run with only Desktop dependencies installed.** The schedule
18
+ description round-trip integration test now lives in the root suite, where
19
+ kernel dependencies are available. Desktop's standalone release test suite
20
+ retains its unit coverage without importing the kernel.
@@ -0,0 +1,120 @@
1
+ # OATS 0.40.2
2
+
3
+ ## Added
4
+
5
+ - **Unconfirmed spawn and operation failures carry structural evidence**
6
+ ([#548](https://github.com/awebai/oats/issues/548)). Incomplete keyed spawns
7
+ and failed spawn compensation set `error.details.unconfirmed: true`; the
8
+ operation wrapper promotes a provider's literal true marker while retaining
9
+ its complete envelope. Completed compensation stays unmarked. Existing
10
+ text fallbacks and scheduler slot/reconcile behavior remain during this
11
+ additive migration; older copied providers are not upgraded by this change.
12
+
13
+ ## Fixed
14
+
15
+ - **Schedule registry reads are lock-free and do not mutate state**
16
+ ([#579](https://github.com/awebai/oats/issues/579)). Status and other readers
17
+ use coherent atomic-file snapshots, including during registry writes. Legacy
18
+ cap migration is projected in memory and persisted only by locked mutations,
19
+ with a one-slot restoration notice on stderr. Invalid stored schedule caps
20
+ now refuse instead of silently increasing capacity; explicit CLI cap options
21
+ can repair them. Old hand-set and implicit ones cannot be distinguished, nor
22
+ can current kernels identify a one reintroduced by an older writer sharing
23
+ the registry; stop mixed-version writes and explicitly set the desired cap.
24
+ Shared directory locks wait through transient owner publication/removal within
25
+ their retry deadline without stealing locks.
26
+ - **OATS Desktop: "Needs input" shows on a remote instance, and a withheld
27
+ note reads as the reason.** Desktop hid every remote claim: a remote roster
28
+ row carries no `runtimeState`, and Desktop read "not reported" as "not
29
+ running" (#582). An unreported state no longer hides the claim; a row
30
+ reported stopped, unreachable or unsupported, or one whose server did not
31
+ answer, still shows none. When a claim's note holds a link or text such as
32
+ `token: …`, Desktop does not show it. The row's description, the terminal
33
+ tab's title and the hover card now show the reason in words, such as "Asked
34
+ you a question", where they showed "[Detail withheld]" (#584). The activity
35
+ view still marks such a note as withheld, and `oats status` prints it in
36
+ full.
37
+ - **A remote instance's "needs input" reaches the roster**
38
+ ([#582](https://github.com/awebai/oats/issues/582)). `oats server roster`
39
+ built a remote row from a fixed list of facts and dropped the host's
40
+ `waitingOnYou`. A row now carries it when, and only when, the host's kernel
41
+ reports it: a row from a host before 0.40.0, and a saved route the host did
42
+ not list, have no such key, so "not reported" stays distinct from `null`
43
+ ("no claim"). The value passes the kernel's read rule again on this side;
44
+ anything that is not a claim is `null`. See
45
+ [the remote roster](../desktop-cli-api.md#the-remote-roster-oats-server-roster---json).
46
+ - **"Needs input" in a deployment addressed through a symlink**
47
+ ([#583](https://github.com/awebai/oats/issues/583)). A home there has two
48
+ spellings: `oats status --dir <symlink>` addressed it lexically, while a
49
+ started or restarted session carries the real path. Event rows are keyed by
50
+ the home, so the session's claims never showed on status, a restart did not
51
+ void a claim written under the other spelling, and a clear through one
52
+ spelling found nothing. Rows are now stored and matched under the home's
53
+ real path, whichever spelling a writer or reader uses, and both spellings
54
+ use one workspace log, the deployment's own `.agents/events`. Answers keep
55
+ the spelling they were asked in: the `home` in `oats instance events --json`
56
+ (the top-level field and each row's) and in the `oats instance waiting` and
57
+ `oats instance attention` answers is the home as the caller addressed it.
58
+ Deployments not reached through a symlink see no change.
59
+
60
+ Rows an earlier kernel wrote under the lexical spelling are not rewritten
61
+ and not matched: after the upgrade they are foreign, counted in
62
+ `integrity.foreignRows`. In a symlinked deployment that means:
63
+ - a claim that was live under the lexical spelling (a session that was
64
+ spawned and never restarted) stops showing. It shows again only when it
65
+ is made anew: at the agent's next permission prompt, question or
66
+ `oats instance attention`. A restart starts a new session with no claim;
67
+ - a claim that stayed set because its restart was recorded under the real
68
+ path is gone, not stuck;
69
+ - the rows a started or restarted session wrote under the real path are
70
+ read now. They cannot show a stale claim: such a session wrote its clears
71
+ and its session boundaries under the real path too.
72
+ - **oats.core 2.4.1: the Claude Code emitter no longer speaks for another
73
+ instance** ([#584](https://github.com/awebai/oats/issues/584)). Its hook
74
+ script took the home from `$OATS_INSTANCE_HOME`. A Claude process that
75
+ loaded one home's `.claude/settings.json` while carrying another instance's
76
+ environment (a nested `claude -p`, a `claude -p` started with its working
77
+ directory in another home, a pane that inherited the variables) set and
78
+ cleared that other instance's claim, and its `Stop` and `SessionEnd` clears
79
+ could erase a real one. The launch hook now writes the home's real path
80
+ into each hook command, and the script acts only when `$OATS_INSTANCE_HOME`
81
+ names that home, through any spelling; otherwise it does nothing. Ships in
82
+ oats.framework 1.6.1 (compatibility unchanged; catalog and workspace pin
83
+ `oats-framework/v1.6.1`): a workspace's `oats.framework: v1.6.1` resolves
84
+ to that tag through the official catalog, and a workspace pinned to v1.6.0
85
+ keeps oats.core 2.4.0 until it moves the pin and syncs. A home picks it up
86
+ at its next spawn. See [capabilities](../capabilities.md).
87
+ - **A waiting row whose time is not a date is no claim**
88
+ ([#584](https://github.com/awebai/oats/issues/584)). The reader validated a
89
+ stored claim's producer, reason and message but passed its time through.
90
+ And `oats help` lists `[--dir <d>]` for `oats instance waiting`.
91
+ - **Catchable scheduler-supervisor shutdown cleans up its owned child group**
92
+ ([#580](https://github.com/awebai/oats/issues/580)). SIGINT, SIGTERM and SIGHUP
93
+ enter the existing bounded TERM/KILL cleanup once, preserving observed child
94
+ exit evidence and treating interrupted envelopes as unconfirmed. The group
95
+ is checked when its leader exits and never signalled again after it is seen
96
+ empty, reducing but not eliminating process-group ID reuse races. Escaped
97
+ sessions and unrecoverable supervisor deaths such as SIGKILL/OOM remain
98
+ outside that cleanup guarantee; no receipt means no proven child exit, so
99
+ the unknown attempt and host slot remain pending reconciliation.
100
+ - **Workspace schedule opt-outs and named trust accept 100-character IDs**
101
+ ([#581](https://github.com/awebai/oats/issues/581)). Host configuration now
102
+ accepts the full schedule name limit in `schedules.disabled` and
103
+ `automations.trust`, so enabling/disabling and trusting schedules with
104
+ 41–100-character IDs works through the CLI and workspace placement checks.
105
+ Trigger definitions and `triggers.disabled` retain their 40-character limit;
106
+ member syntax, allowed characters and uniqueness rules are unchanged.
107
+
108
+ ## Changed
109
+
110
+ - **Session input restores single-paste, single-Enter terminal semantics**
111
+ ([#562](https://github.com/awebai/oats/issues/562)). Corrects 0.39.4's
112
+ screen-derived extra Enters and `enter-not-taken` refusal after successful
113
+ terminal commands. Pre-Enter settling remains, with a shared monotonic
114
+ 2-second observation budget; at most two read-only post-Enter looks share
115
+ 1 second. Each probe and sleep is limited by its remaining budget.
116
+ `submitted: true` reports terminal-operation success; `verified` reports
117
+ display change only, and false never authorizes retry. Neither proves model
118
+ acceptance, pending draft state or absence of effects. Command errors and
119
+ authority checks are unchanged; busy-pane submission, exactly-once delivery
120
+ and generic-error retry uncertainty are not solved by this correction.
package/docs/schedules.md CHANGED
@@ -38,7 +38,7 @@ The design is in the
38
38
  | `<deployment>/oats-schedules.json` | This machine's local definitions, `{version: 1, jobs: {<id>: …}}`, local triggers included (`kind: "trigger"`). |
39
39
  | `<deployment>/.agents/automations/snapshot.json` | The workspace definitions discovered from the members. |
40
40
  | `<deployment>/.agents/schedules/` | Run state: `state.json` (last minute and recent runs per job), `triggers.json` (polls, pending events, fired keys) and one lock directory per running job. |
41
- | `~/.oats/schedules/registry.json` | The deployments this host ticks, `maxConcurrent` (default 1: running scheduled jobs) and `triggersMaxConcurrent` (absent: no host cap on trigger-spawned live instances). The two caps are separate. |
41
+ | `~/.oats/schedules/registry.json` | The deployments this host ticks, `maxConcurrent` (absent: default 5 running scheduled jobs) and `triggersMaxConcurrent` (absent: no host cap on trigger-spawned live instances). The two caps are separate. |
42
42
 
43
43
  One host lock serializes ticks, run-now, reconcile and remove. It is never
44
44
  reclaimed by another process: a lock whose owner is gone is reported with the
@@ -48,7 +48,18 @@ directory to remove.
48
48
 
49
49
  Every definition carries `id`, `enabled`, `cron`, `tz` and `kind`. `cron` has
50
50
  five fields (minute hour day month weekday) and `tz` is a required IANA zone;
51
- both are evaluated by the croner library.
51
+ both are evaluated by the croner library. Schedule IDs use lowercase letters,
52
+ digits and dashes, from 1 to 100 characters. Spawn schedules with a long ID
53
+ need an explicit shorter `purpose` to fit the instance-name limit below.
54
+
55
+ Any kind may carry `description`: what the job is for, in words, for the
56
+ people reading `oats schedule list`, `show` and the Desktop. It is one line of
57
+ 1 to 200 characters with no control characters (no CR, LF, TAB or any other
58
+ C0 or C1 character, nor a Unicode line or paragraph separator); anything else
59
+ is `E_SCHEDULE_INVALID` with `field: "description"`. It is stored as given and
60
+ is informational only: it never reaches a run's argv, environment, task or
61
+ reconcile. A capability that registers jobs (knowledge harvest's `run-source`
62
+ jobs) sets it so that its command jobs can be told apart.
52
63
 
53
64
  - **spawn** `{…, agent, agentsRoot?, repo?, backend?, purpose?, task,
54
65
  launchConfig?, harness?, model?, yolo?, wake?}` — every due minute launches
@@ -70,7 +81,24 @@ both are evaluated by the croner library.
70
81
  no shell; `oats schedule` itself is refused) in `cwd`, an existing directory
71
82
  inside the deployment. The runner tracks any instance the command's
72
83
  envelope names, including an independent worker it reports, until its home
73
- is gone. A command's return is not task completion.
84
+ is gone. A command's return is not task completion. Command/operation runs and
85
+ workspace-spawn launches have a five-minute child timeout. The supervisor
86
+ sends SIGTERM to the child's process group, allows two seconds for cleanup,
87
+ then sends SIGKILL if the group remains. It observes the direct child's exit
88
+ before returning; inherited output pipes cannot hold the tick indefinitely.
89
+ SIGINT, SIGTERM or SIGHUP received by the supervisor enters that same cleanup
90
+ once; repeated signals do not bypass it. A timeout or interrupted supervisor
91
+ leaves effects unconfirmed even if the child printed an envelope. Spawn
92
+ previews use the same bounded runner.
93
+
94
+ Cleanup covers the owned process group. A descendant that creates its own
95
+ session can escape it; inherited pipes are bounded but that escaped process
96
+ is not terminated by this group cleanup. The supervisor checks the group when
97
+ the leader exits and never signals it after observing it empty. This reduces
98
+ the group-ID reuse window; it does not eliminate PID reuse races. SIGKILL,
99
+ OOM and other unrecoverable supervisor deaths cannot run JavaScript handlers:
100
+ cleanup is not guaranteed then. Without a private supervisor receipt, the
101
+ scheduler keeps the unknown attempt and its slot until reconciliation.
74
102
  - **wake** `{…, home, message}` — every due minute inspects the instance at
75
103
  `home`. Running: `message` is delivered once as terminal input (bracketed
76
104
  paste plus Enter), never an interrupt. Not running: the home is started with
@@ -220,7 +248,8 @@ owner: github.com/ana
220
248
  is an `E_AUTOMATION_SCHEMA` problem, never silently skipped.
221
249
  - **The id** is `id:`, else the filename stem. The same id twice in one member
222
250
  for one kind is `E_AUTOMATION_DUPLICATE`, naming both paths; the second file
223
- is not listed. A trigger and a schedule may share an id. A member named
251
+ is not listed. Schedule IDs allow 1 to 100 lowercase letters, digits and dashes; trigger
252
+ IDs allow 1 to 40. A trigger and a schedule may share an id. A member named
224
253
  `local` is refused, because `local/<id>` names this host's own definitions.
225
254
  - **A workspace schedule is `run: spawn` or `run: command`.** A command's
226
255
  `cwd` is relative to the deployment and must stay inside it. `wake` and
@@ -255,9 +284,12 @@ add` / `oats schedule add` definitions need no trust.
255
284
 
256
285
  **Opting out on one host.** `oats trigger disable <member>/<id>` writes
257
286
  `triggers.disabled`, and `oats schedule disable <member>/<id>` writes
258
- `schedules.disabled`, in `oats-local.yaml`; `enable` removes the entry. A
259
- workspace definition is never edited or removed from the CLI (`update` and
260
- `remove` answer `E_AUTOMATION_WORKSPACE`): change the file in Git.
287
+ `schedules.disabled`, in `oats-local.yaml`; `enable` removes the entry. The
288
+ schedule part of a qualified ID accepts up to 100 characters in both
289
+ `schedules.disabled` and named `automations.trust` entries. Trigger definitions
290
+ and `triggers.disabled` keep their 40-character limit; the shared trust list
291
+ does not widen trigger IDs. A workspace definition is never edited or removed
292
+ from the CLI (`update` and `remove` answer `E_AUTOMATION_WORKSPACE`): change the file in Git.
261
293
 
262
294
  **Refresh.**
263
295
 
@@ -288,6 +320,8 @@ oats schedule test <id> # dry run: where it runs, whether its soul res
288
320
  oats schedule tick [--dry-run] # evaluate this deployment now; --dry-run launches nothing
289
321
  oats schedule reconcile <id> [--clear] # resolve an attempt whose result was never recorded
290
322
  oats schedule host install # register this deployment and install the one host timer (idempotent)
323
+ oats schedule host install --max-concurrent 3 --triggers-max-concurrent 2
324
+ oats schedule host install --max-concurrent default --triggers-max-concurrent none
291
325
  oats schedule host status | uninstall
292
326
  oats spawn <agent> ... --wake-every 15 --wake-message "Anything new?" # or --wake-file spec.json
293
327
  ```
@@ -295,7 +329,41 @@ oats spawn <agent> ... --wake-every 15 --wake-message "Anything new?" # or --w
295
329
  `<id>` is `local/<id>` (or the bare id) or `<member>/<id>`. `host uninstall`
296
330
  unregisters the deployment and removes the timer once none is registered.
297
331
  Every `oats schedule` subcommand takes `--server <id>` instead of `--dir` to
298
- run on that registered server.
332
+ run on that registered server. Setting or resetting host caps remotely requires
333
+ the destination to advertise both `schedule` and `schedule-host-caps`; an older
334
+ or unknown peer is refused with `E_REMOTE_INCOMPATIBLE` before host install is
335
+ forwarded. Upgrade OATS on the destination to use these options. A remote
336
+ install without cap options retains its existing behavior.
337
+
338
+ The host allows five running scheduled jobs by default. Use `host install
339
+ --max-concurrent N` to choose a positive integer, or `--max-concurrent default`
340
+ to restore the default. `--triggers-max-concurrent N` independently limits live
341
+ trigger-spawned instances; `none` removes that cap. Omitted flags preserve the
342
+ current choices. Invalid values fail with `E_BAD_ARGS` before registration or
343
+ timer changes. `host status` reports the effective `maxConcurrent` and
344
+ `triggersMaxConcurrent` (`null` when uncapped).
345
+
346
+ The registry stores explicit choices only. Reading status or the registry takes
347
+ no registry lock and creates or rewrites no files or directories. A reader
348
+ interprets a pre-migration stored `maxConcurrent: 1` as the default of five in
349
+ memory. Registration, unregistration and cap updates persist that migration
350
+ once, under the registry lock, even if the workspace membership is unchanged.
351
+ Other explicit values and the independent trigger cap survive.
352
+
353
+ A pre-migration hand-set one is indistinguishable from the old implicit one;
354
+ both follow that compatibility choice. When a write migrates one to the default,
355
+ it prints a notice on stderr with `oats schedule host install --max-concurrent 1`
356
+ to restore one if needed. Pure reads stay silent and JSON stdout is unchanged.
357
+ A later explicit one survives writes by current kernels. Older binaries sharing
358
+ the registry can write one back while preserving `capsVersion: 2`; there is no
359
+ provenance to distinguish that from a new explicit one. Stop mixed-version
360
+ writes and explicitly set the desired cap with the supported CLI.
361
+
362
+ A present invalid `maxConcurrent` is `E_SCHEDULE_INVALID`, not a fallback to
363
+ five. Correct it with `oats schedule host install --max-concurrent N` (or
364
+ `--max-concurrent default`) from the deployment, or with `--dir <deployment>`.
365
+ An absent value still means five. Set caps through the CLI; do not edit the
366
+ registry by hand.
299
367
 
300
368
  `oats schedule list --json` answers:
301
369
 
@@ -306,14 +374,14 @@ run on that registered server.
306
374
  schedules: [ <row> ],
307
375
  triggers: { count, command: "oats trigger list" },
308
376
  snapshot: { takenAt, problems } | null,
309
- scheduler: { installed, active, unit?, lastTick, maxConcurrent, tickIntervalSec,
377
+ scheduler: { installed, active, unit?, lastTick, maxConcurrent, triggersMaxConcurrent, tickIntervalSec,
310
378
  workspace, registered, workspaces, live } }
311
379
  ```
312
380
 
313
381
  Each row is the stored definition plus `id` (bare for a local schedule,
314
- `<member>/<id>` for a workspace one), `qualifiedId`, `origin`, `owner`,
315
- `runsOn`, `runsHere`, `reason`, `enabledHere`, `soul`, `nextDue`, `lastRun`,
316
- `recentRuns` and `running`; an unreadable row carries `unreadable: { code,
382
+ `<member>/<id>` for a workspace one), `qualifiedId`, `description` (`null`
383
+ when there is none), `origin`, `owner`, `runsOn`, `runsHere`, `reason`,
384
+ `enabledHere`, `soul`, `nextDue`, `lastRun`, `recentRuns` and `running`; an unreadable row carries `unreadable: { code,
317
385
  message }` instead of failing the list. `scheduler.active` is what the OS
318
386
  reports about the timer. `oats trigger list --json` carries the same
319
387
  `scheduler`. The field-level contract is in
@@ -328,12 +396,42 @@ you), `launch-failed`, `unknown`, and for wake jobs `delivered`, `started` or
328
396
  `skipped`. The kernel never claims a task succeeded.
329
397
 
330
398
  `unknown` means the launch's side effects are unconfirmed: a command timed out
331
- or answered no envelope, or an attempt was never recorded. The job keeps its
332
- slot and is skipped until `oats schedule reconcile <id>`, which adopts only an
333
- attributable receipt (a spawn job's instance, named for its minute, or the
334
- instance a command's answer named). When nothing is attributable, check the
335
- roster and the host by hand, then `reconcile <id> --clear` records
336
- `launch-failed` and frees the slot.
399
+ or answered no envelope, or an attempt was never recorded. The job is skipped
400
+ until `oats schedule reconcile <id>`, which adopts only an attributable
401
+ receipt (a spawn job's instance, named for its minute, or the instance a
402
+ command's answer named). When nothing is attributable, check the roster and
403
+ the host by hand, then `reconcile <id> --clear` records `launch-failed` and
404
+ frees the slot. The unresolved attempt shows in `show` as `attempt:
405
+ {scheduledFor, startedAt, error?, exited?, exitStatus?, exitSignal?}`;
406
+ `error` is the first run's cause, which later skipped ticks keep in `lastRun`.
407
+
408
+ `oats doctor` warns about each unresolved attempt, in this deployment and in
409
+ the other deployments this host ticks (they share its slots). The text form
410
+ is `! schedule-unresolved: …`; in `--json` it is a `problems[]` item:
411
+
412
+ ```text
413
+ { code: "schedule-unresolved", severity: "warning", scope, id, kind,
414
+ scheduledFor, startedAt, ageSeconds, holdsSlot, exited, error, remedy, message }
415
+ ```
416
+
417
+ `holdsSlot` says whether the job counts against `maxConcurrent`. `remedy` is
418
+ `oats schedule reconcile <id>`, with `--clear` for a command or operation
419
+ whose effects no named home proves (check the roster and the host by hand
420
+ first), and with `--dir <scope>` for another deployment. A warning never
421
+ changes doctor's exit status. Workspace command kinds come from the last
422
+ saved automations snapshot, without a refresh or an account lookup. Unresolved
423
+ workspace state is reported even if its definition is no longer available.
424
+
425
+ Whether an `unknown` job keeps its host slot depends on what is still running:
426
+
427
+ - A `command` or `operation` job whose process exit the kernel observed (it
428
+ returned, or was stopped at the five-minute timeout) holds no slot: the
429
+ process runs nothing any more, and an instance it spawned has its own
430
+ lifecycle. Its attempt shows `exited: true` with the exit status or
431
+ signal, and other jobs keep running.
432
+ - A `spawn` job keeps its slot, which stands for the instance it may have
433
+ launched. So does a command whose exit was not observed (the runner threw,
434
+ the process never started) and a legacy attempt without exit evidence (no `exited`), until reconcile.
337
435
 
338
436
  **Slots.** A wake job that starts a stopped home holds a launch slot until the
339
437
  harness is proven stopped or the home is gone; delivering to a running home
@@ -344,7 +442,7 @@ continues.
344
442
 
345
443
  **Changing a job.** `disable` never stops anything. `update` never touches a
346
444
  running instance, and while a job holds a slot or has an unresolved attempt
347
- only `cron`, `tz` and `enabled` can change. `remove` refuses while the job's
445
+ only `cron`, `tz`, `enabled` and `description` can change. `remove` refuses while the job's
348
446
  instance is tracked or its effects are unresolved (`--force` forgets the job
349
447
  without stopping anything). Retiring an instance removes the wake jobs bound
350
448
  to its home.
package/docs/servers.md CHANGED
@@ -298,7 +298,9 @@ the host's own facts from its `status --json`: `identity`,
298
298
  `identityAddress`, `teams`, `startedAt`, `createdAt`, `model`,
299
299
  `runtimeState`, `parentInstance`, `siblingInstance`, `relation`,
300
300
  `relativeTo` and `spawnOrigin`. A fact the host does not supply is `null`
301
- (an older host, or a saved route the host no longer lists). A removed or edited registration keeps
301
+ (an older host, or a saved route the host no longer lists). A row also
302
+ carries `waitingOnYou` (needs input) when the host's kernel reports it, and
303
+ only then: an older host's row has no such key. A removed or edited registration keeps
302
304
  its group from the saved routes. State is pulled on every call within
303
305
  `--per-target` (default 20 s) of a total `--budget` (default 45 s); a group
304
306
  not reached is reported with `E_ROSTER_BUDGET`. `--server <id>` narrows it.
@@ -54,7 +54,7 @@ members: # repo refs, NO @revision (E_WORKSPAC
54
54
  - git:github.com/acme/tools # a member that ALSO publishes a package (see below)
55
55
 
56
56
  packages: # the ONLY versioned things
57
- oats.framework: v1.5.0 # bare version → resolves through the official catalog
57
+ oats.framework: v1.6.1 # bare version → resolves through the official catalog
58
58
  oats.okf: v4.1.1
59
59
  acme.tools: git:github.com/acme/tools@v0.4.0 # outside the catalog → git:<repo>@<tag|OID>; still a package
60
60
 
@@ -45,6 +45,8 @@ export const PLACEMENT_REASONS = Object.freeze(["host-unnamed", "assigned-elsewh
45
45
  export const KIND_NAMES = Object.freeze(["trigger", "schedule"]);
46
46
  const NEVER_SCANNED = new Set(["oats-package", ".git", "node_modules"]);
47
47
  export const AUTOMATION_ID_RE = /^[a-z0-9-]{1,40}$/;
48
+ /** Schedule definition names have a wider bound than trigger names. */
49
+ export const SCHEDULE_NAME_RE = /^[a-z0-9-]{1,100}$/;
48
50
  /** A model id an automation may pass to `oats spawn --model`: a provider/model id, never an option.
49
51
  * The first character is alphanumeric (or the spawn CLI's own `@native-default`), so no value can
50
52
  * be read as a flag (re-review B #1). `[` `]` admit a harness's context-size alias (`opus[1m]`). */
@@ -106,7 +108,8 @@ export function parseAutomationFile(desc, { stem, path, bytes, member, repoKey,
106
108
  const unknown = Object.keys(doc).filter((k) => !allowed.includes(k));
107
109
  if (unknown.length) return problem(`unknown field${unknown.length > 1 ? "s" : ""} ${unknown.join(", ")} (a ${desc.fileKind} carries ${allowed.join(", ")})`, unknown[0]);
108
110
  const name = doc.id === undefined ? stem : doc.id;
109
- if (typeof name !== "string" || !AUTOMATION_ID_RE.test(name)) return problem(`the id (${doc.id === undefined ? "the filename stem" : "id:"} ${JSON.stringify(name)}) must be lowercase letters, digits and dashes, 1 to 40 characters`, "id");
111
+ const namePattern = desc.kind === "schedule" ? SCHEDULE_NAME_RE : AUTOMATION_ID_RE;
112
+ if (typeof name !== "string" || !namePattern.test(name)) return problem(`the id (${doc.id === undefined ? "the filename stem" : "id:"} ${JSON.stringify(name)}) must be lowercase letters, digits and dashes, 1 to ${desc.kind === "schedule" ? 100 : 40} characters`, "id");
110
113
  if (typeof doc.runsOn !== "string" || !HOST_NAME_RE.test(doc.runsOn)) return problem("runsOn: the host name that runs it (oats-local.yaml host.name: lowercase letters, digits and dashes)", "runsOn");
111
114
  const owner = parseOwner(doc.owner);
112
115
  if (!owner) return problem("owner: the GitHub account it acts as, <host>/<login> (e.g. github.com/acme-kb-bot)", "owner");