@noetaris/harness 0.4.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -156,6 +156,157 @@ const resumed = run.resume(response, interruptId)
156
156
  const resumed = agent.resume(response, sessionId, interruptId)
157
157
  ```
158
158
 
159
+ **Resume replays the interrupted step from the top.** On resume, the step that called
160
+ `ctx.interrupt()` runs again from its first line; each `ctx.interrupt()` call it reaches
161
+ returns the stored response instead of pausing. State updates from the paused attempt are
162
+ not kept — a step's update is applied only when `run` returns. So any side effect a step
163
+ performs *before* `ctx.interrupt()` (an HTTP call, a DB write, a sub agent run) happens
164
+ again on resume. Making that safe is up to the step: guard the side effect so a replay
165
+ skips or reuses it.
166
+
167
+ **Inside a fork, only paused branches replay.** When a fork pauses, each branch is saved
168
+ separately. On resume, a branch that already finished is not run again — its saved state
169
+ goes straight to the join. Every branch that is still paused runs again from its paused
170
+ step, even if the response you sent was for a different branch; a branch with no response
171
+ yet reaches its `ctx.interrupt()` again and re-pauses. The session stays paused until
172
+ every pending interrupt (see `agent.status(sessionId).pendingInterrupts`) is answered.
173
+ If any branch fails, the fork fails fast: the other branches are aborted and paused
174
+ branches are dropped. The same happens when a branch throws a graph-definition error
175
+ (for example a signal with no matching `.on()`): the other branches are aborted, and the
176
+ run rejects with that error only after they have stopped.
177
+
178
+ **Stopping a fork.** `stop()` while a fork runs stops each branch at its next step
179
+ boundary. What happens next depends on whether any branch is waiting on an interrupt:
180
+
181
+ | At the stop | Run settles with | To continue |
182
+ |---|---|---|
183
+ | No branch is waiting on an interrupt | `signal: null` — a plain stop; `agent.status()` shows `paused` with no pending interrupts | `agent.run()` on the same session continues the fork: finished branches are not run again, stopped branches continue from where they stopped. `resume()` throws `NoInterruptError`. |
184
+ | At least one branch (at any depth) is waiting on an interrupt | `signal: '$interrupt'` | `resume()` answers it and re-runs every paused branch, including the ones that were only stopped |
185
+
186
+ Input passed to the `agent.run()` that continues a stopped fork is merged into the main
187
+ state only. The branches that continue keep their own saved state and do not see it; it
188
+ is visible from the join onward.
189
+
190
+ #### Calling sub agents from a step
191
+
192
+ A sub agent is just another agent with its own session. When it pauses, its `run()`
193
+ resolves with `signal: '$interrupt'` — it does not throw — so the calling step must pass
194
+ the interrupt up with `ctx.interrupt()` and, after resume, call the sub agent's
195
+ `resume()`. Because the step replays, the call must be idempotent. Checking the sub
196
+ agent's session phase is enough:
197
+
198
+ ```ts
199
+ async function callSubAgent(agent: Agent, sessionId: string, input: object, ctx: any) {
200
+ const status = await agent.status(sessionId)
201
+ if (status.phase === 'completed') return loadResult(sessionId) // your own lookup — don't run again
202
+ let outcome
203
+ if (status.phase === 'paused') {
204
+ const pending = status.pendingInterrupts[0]!
205
+ const answer = await ctx.interrupt(pending.prompt, `${sessionId}:${pending.interruptId}`)
206
+ outcome = await agent.resume(answer, sessionId, pending.interruptId)
207
+ } else {
208
+ outcome = await agent.run(input, { sessionId })
209
+ }
210
+ if (outcome.signal === '$interrupt') {
211
+ const { interruptId, prompt } = outcome.state.$interrupt
212
+ await ctx.interrupt(prompt, `${sessionId}:${interruptId}`) // pauses the main agent
213
+ }
214
+ return outcome.state
215
+ }
216
+ ```
217
+
218
+ Use a stable sub agent `sessionId` (for example derived from the main `sessionId`) so a
219
+ replay finds the same session. To run several sub agents in parallel, put each in its own
220
+ fork branch rather than `Promise.all` in one step: then, when one pauses, the finished ones
221
+ are not run again on resume.
222
+
223
+ ```ts
224
+ h.loop(l =>
225
+ l.start()
226
+ .fork('research')
227
+ .branch('web', b => b.start()
228
+ .step('callWeb', {
229
+ run: async (s, ctx) => ({ web: await callSubAgent(webAgent, `${ctx.sessionId}:web`, {}, ctx) }),
230
+ route: () => 'done',
231
+ })
232
+ .on('done').end()) // inside a branch, .end() ends the branch
233
+ .branch('docs', b => b.start()
234
+ .step('callDocs', {
235
+ run: async (s, ctx) => ({ docs: await callSubAgent(docsAgent, `${ctx.sessionId}:docs`, {}, ctx) }),
236
+ route: () => 'done',
237
+ })
238
+ .on('done').end())
239
+ .join('merge', { run: mergeResults, route: () => 'done' })
240
+ .on('done').end()
241
+ )
242
+ ```
243
+
244
+ If a fork fails fast while a sub agent is paused, that sub agent's own session stays
245
+ paused in its store — clean it up or reuse it the next time the step runs.
246
+
247
+ ### Observers
248
+
249
+ An `Observer` receives telemetry hooks for a run. Pass it as `observer` in the resources
250
+ of `agent.run()` or `agent.resume()`. Combine several with `composeObservers([a, b])`;
251
+ a hook that throws is reported to `onObserverError` (or `console.error`) and never stops
252
+ the run. Every hook is optional.
253
+
254
+ | Hook | Fires |
255
+ |---|---|
256
+ | `onRunStart(ctx)` | once when a run starts |
257
+ | `onRunEnd(ctx, { signal, durationMs })` | once when a run settles (`$stopped` for a stop) |
258
+ | `onStepStart(ctx)` | before each step |
259
+ | `onStepEnd(ctx, { durationMs })` | after a step's `run` succeeds, before its `route` |
260
+ | `onStepError(ctx, { error, durationMs })` | when a step's `run` throws |
261
+ | `onInterrupt(ctx, { prompt, interruptId })` | when a step calls `ctx.interrupt()` |
262
+ | `onEvent(ctx, type, payload)` | on `ctx.emit()` and adapter events such as `llm.response` |
263
+ | `onStepSettled(ctx, event)` | once per step, when its outcome and destination are final |
264
+
265
+ **`onStepSettled`** is the one hook that tells you, for every step, how it ended and
266
+ where the run goes next. It fires exactly once for each `onStepStart`, on every path:
267
+
268
+ ```
269
+ onStepStart → onStepEnd | onStepError | onInterrupt → onStepSettled → onRunEnd
270
+ ```
271
+
272
+ It also fires for a fork node and for each step inside a branch (`ctx.branchPath` names the
273
+ branch). A graph-definition error (`UnknownSignalError`, `NoNextStepError`,
274
+ `MissingReducerError`) settles with `next.kind === 'throw'` before the run rejects. It does
275
+ not fire when a run stops before starting a step, or for a finished branch that a resume
276
+ skips.
277
+
278
+ | `event` field | Meaning |
279
+ |---|---|
280
+ | `outcome` | `'ok'`, `'error'`, `'interrupt'`, or `'stopped'` (a fork whose branches were stopped) |
281
+ | `durationMs` | from `onStepStart` until the destination is known — includes `route` |
282
+ | `update` | what `run` returned, present only when it was applied to state. The raw object — copy it if you keep it |
283
+ | `signal` | what `route` returned, when it was called |
284
+ | `next` | `{ kind: 'step', name }`, `{ kind: 'end' }`, `{ kind: 'pause' }`, or `{ kind: 'throw' }` |
285
+ | `error` | present when `outcome` is `'error'` |
286
+ | `interrupt` | `{ interruptId, prompt }`, when a plain step paused on an interrupt |
287
+ | `fork` | on a fork node that ended done or paused: `{ branches: [{ name, status, touchedKeys }] }` |
288
+
289
+ ```ts
290
+ const logger: Observer = {
291
+ onStepSettled(ctx, e) {
292
+ const next = e.next.kind === 'step' ? `→ ${e.next.name}` : e.next.kind
293
+ console.log(`${ctx.stepName}: ${e.outcome} ${next}`, e.signal ? `(signal ${e.signal})` : '')
294
+ },
295
+ }
296
+
297
+ const h = createHarness()({ total: field({ default: () => 0 }) }).loop(l =>
298
+ l.start()
299
+ .step('fetch', { run: async () => ({ total: 3 }) })
300
+ .step('decide', { route: s => (s.total > 0 ? 'done' : 'empty') })
301
+ .on('done').end()
302
+ .on('empty').end()
303
+ )
304
+
305
+ await createAgent('demo', h, {}).run({}, { observer: logger })
306
+ // fetch: ok → decide
307
+ // decide: ok end (signal done)
308
+ ```
309
+
159
310
  ## API
160
311
 
161
312
  | Export | Description |
@@ -165,10 +316,12 @@ const resumed = agent.resume(response, sessionId, interruptId)
165
316
  | `field<T>(opts)` | Declares a state field with a default and optional reduce function. |
166
317
  | `required()` | Marks a provider slot as required at `createAgent()`. |
167
318
  | `runtime()` | Marks a provider slot as required at `agent.run()`. |
168
- | `composeObservers(...observers)` | Merges multiple `Observer` instances into one fan-out observer. |
319
+ | `composeObservers([a, b], onObserverError?)` | Merges multiple `Observer` instances into one fan-out observer; a throwing observer is isolated and reported to `onObserverError`. |
169
320
  | `SessionStore` | Interface for session persistence backends. |
170
321
  | `StoredRun` | Type for a persisted run snapshot. Includes `agentId`, `runId`, `sessionId`, `phase`, and state. |
171
- | `Observer` | Interface for telemetry hooks on run and step lifecycle events. |
322
+ | `Observer` | Interface for telemetry hooks on run and step lifecycle events (see [Observers](#observers)). |
323
+ | `StepSettledEvent` | Payload of `Observer.onStepSettled` — outcome, duration, update, signal, next, error, interrupt, fork. |
324
+ | `StepSettledNext` | Where the run goes after a step settles: `step`, `end`, `pause`, or `throw`. |
172
325
  | `ObserverAware` | Interface for provider objects that accept an `Observer` binding via `bindObserver()`. |
173
326
  | `RunContext` | Context passed to run-level observer hooks — `agentId`, `sessionId`. |
174
327
  | `StepContext` | Context passed to step-level observer hooks — `agentId`, `sessionId`, `stepName`. |
@@ -190,6 +343,8 @@ const resumed = agent.resume(response, sessionId, interruptId)
190
343
  - ESM only (`"type": "module"`)
191
344
  - Zero runtime dependencies
192
345
 
346
+ See [LIMITATIONS.md](LIMITATIONS.md) for known limitations.
347
+
193
348
  ## Related Packages
194
349
 
195
350
  - [`@noetaris/harness-store`](https://github.com/noetaris-lab/harness-store) — session store implementations (`InMemorySessionStore`, `LocalFileSessionStore`, etc.)
package/dist/index.d.ts CHANGED
@@ -145,6 +145,62 @@ interface StepContext {
145
145
  */
146
146
  readonly branchPath?: readonly string[];
147
147
  }
148
+ /**
149
+ * Where execution goes after a step settles.
150
+ *
151
+ * - `step` — the next step to run (a route target, `.next()`, the implicit next step, the
152
+ * `l.onError()` fallback, or the join after a fork).
153
+ * - `end` — the loop ended via `.end()` (inside a fork branch: the branch ended).
154
+ * - `pause` — the run paused (interrupt, stop, or an unhandled error) and can be continued later.
155
+ * - `throw` — a graph-definition error (`UnknownSignalError`, `NoNextStepError`,
156
+ * `MissingReducerError`) is about to reject the run.
157
+ */
158
+ type StepSettledNext = {
159
+ readonly kind: 'step';
160
+ readonly name: string;
161
+ } | {
162
+ readonly kind: 'end';
163
+ } | {
164
+ readonly kind: 'pause';
165
+ } | {
166
+ readonly kind: 'throw';
167
+ };
168
+ /**
169
+ * Payload of {@link Observer.onStepSettled}: everything known about one step once its
170
+ * outcome and destination are final. Optional keys are omitted, never set to `undefined`.
171
+ */
172
+ interface StepSettledEvent {
173
+ /** `'stopped'` only occurs on fork nodes whose branches were stopped without an interrupt. */
174
+ readonly outcome: 'ok' | 'error' | 'interrupt' | 'stopped';
175
+ /** From `onStepStart` until the destination is known — includes `route()`. */
176
+ readonly durationMs: number;
177
+ /**
178
+ * The value `run()` returned, present iff it was applied to state. The raw reference,
179
+ * unfiltered — it may be aliased into state, so copy it before the next step if you keep it.
180
+ */
181
+ readonly update?: Record<string, unknown>;
182
+ /** Present iff `route()` was called and returned a signal. */
183
+ readonly signal?: string;
184
+ readonly next: StepSettledNext;
185
+ /** Present iff `outcome` is `'error'`. */
186
+ readonly error?: Error;
187
+ /** Present iff `outcome` is `'interrupt'` on a plain (non-fork) step. */
188
+ readonly interrupt?: {
189
+ readonly interruptId: string;
190
+ readonly prompt: unknown;
191
+ };
192
+ /**
193
+ * Fork nodes only, when every branch ended done or paused (not when the fork failed):
194
+ * each branch in declaration order.
195
+ */
196
+ readonly fork?: {
197
+ readonly branches: readonly {
198
+ readonly name: string;
199
+ readonly status: 'done' | 'paused';
200
+ readonly touchedKeys: readonly string[];
201
+ }[];
202
+ };
203
+ }
148
204
  /**
149
205
  * Observability hook interface. All methods are optional — implement only
150
206
  * the hooks you need.
@@ -193,40 +249,47 @@ interface Observer {
193
249
  * by LLM adapters (e.g. `'llm.response'`).
194
250
  */
195
251
  onEvent?: (ctx: StepContext, type: string, payload: unknown) => void;
252
+ /**
253
+ * Called exactly once per step, on every path, once the step's outcome and destination
254
+ * are final: after `onStepEnd` / `onStepError` / `onInterrupt` (when that hook fires on the
255
+ * path) and before `onRunEnd`. Also fires for fork nodes and for steps inside fork branches
256
+ * (with `ctx.branchPath`). Not fired when a run stops before starting a step.
257
+ */
258
+ onStepSettled?: (ctx: StepContext, event: StepSettledEvent) => void;
196
259
  }
197
260
  /**
198
- * Implemented by resources (e.g. LLM adapters) that accept an {@link Observer}
199
- * at run time. The harness calls `bindObserver` on every slot value that
200
- * implements this interface before the first step runs.
261
+ * The telemetry a resource needs to attribute an `invoke()` call to the correct
262
+ * run and step: the run's {@link Observer} and the current {@link StepContext}.
201
263
  *
202
- * Optionally, the harness calls `setStepContext` at the start of each step for
203
- * every slot that exposes it. Adapters use this to attribute per-step telemetry
204
- * (e.g. `observer.onEvent`) to the correct step without manual calls from step code.
264
+ * The harness passes this per call rather than storing it on the resource instance,
265
+ * so concurrently-running fork branches that share one instance never race on it —
266
+ * each call carries its own context by value.
267
+ */
268
+ interface TelemetryContext {
269
+ readonly observer: Observer;
270
+ readonly stepContext: StepContext;
271
+ }
272
+ /**
273
+ * Implemented by resources (e.g. LLM adapters) that emit per-step telemetry.
274
+ *
275
+ * The harness calls {@link ObserverAware.withTelemetry} at the start of each step for every
276
+ * slot that exposes it, and installs the returned value as that slot on the step's `ctx`.
277
+ * Step code then calls the returned view (e.g. `ctx.llm.invoke(...)`) with no telemetry
278
+ * arguments, and the view forwards the {@link TelemetryContext} into the call it wraps — so
279
+ * telemetry travels *with the call* rather than living in mutable instance state.
280
+ *
281
+ * The return is typed `unknown` because core exports no LLM-domain type; a concrete adapter
282
+ * narrows it (e.g. to its `LLM` view). The harness treats the returned value as an opaque
283
+ * slot, exactly as it treats every provided slot.
205
284
  */
206
285
  interface ObserverAware {
207
286
  /**
208
- * Receive the run's observer. The harness calls this once per `agent.run()`
209
- * invocation before execution begins.
210
- */
211
- bindObserver(observer: Observer): void;
212
- /**
213
- * Receive the current step context. The harness calls this at the start of
214
- * each step for every slot that exposes this method, before calling `step.run`.
215
- *
216
- * Adapters that emit `observer.onEvent` in their `invoke()` method store the
217
- * provided `StepContext` and use it as the first argument to `onEvent`, so
218
- * events are attributed to the correct step automatically.
219
- *
220
- * **Known limitation (fork/join branches):** the store-in-a-field pattern above is safe only
221
- * because exactly one step runs at a time in a non-forked run. When a fork's branches run
222
- * concurrently, they typically share one `ObserverAware` slot instance (`h.provide('llm', ...)`
223
- * provides it once), so concurrent branches race on that stored field and `onEvent` calls may
224
- * be attributed to the wrong branch/step. This is a known v1 limitation of the built-in LLM
225
- * adapters (`@noetaris/harness-anthropic`/`-openai`/`-google`/`-ollama`), not fixed for this
226
- * release — a correct fix means passing `StepContext` through the call instead of storing it.
227
- * Treat per-step LLM telemetry as unreliable for concurrently-running branches until then.
287
+ * Return a view of this resource scoped to the given telemetry. The harness installs
288
+ * the returned value as the resource's slot on the step's `ctx` before calling `step.run`.
289
+ * Called once per step (and once per step per branch), so each concurrent branch gets its
290
+ * own scoped view and there is no shared field to race on.
228
291
  */
229
- setStepContext?(ctx: StepContext): void;
292
+ withTelemetry(ctx: TelemetryContext): unknown;
230
293
  }
231
294
  /**
232
295
  * Combine multiple {@link Observer} instances into one. Each hook on the
@@ -234,10 +297,27 @@ interface ObserverAware {
234
297
  *
235
298
  * @example
236
299
  * ```ts
237
- * const observer = composeObservers(otelObserver, metricsObserver)
300
+ * const observer = composeObservers([otelObserver, metricsObserver])
238
301
  * ```
239
302
  */
240
- declare function composeObservers(...observers: Observer[]): Observer;
303
+ /**
304
+ * Identifies which observer hook threw. Passed to an {@link ObserverErrorSink}.
305
+ */
306
+ interface ObserverErrorContext {
307
+ /** The observer hook that threw, e.g. `'onStepStart'`, `'onRunEnd'`, `'onEvent'`. */
308
+ readonly hookName: string;
309
+ }
310
+ /**
311
+ * Opt-in sink for errors thrown by observer callbacks. Supplied under the
312
+ * `'onObserverError'` key in `agent.run()` / `agent.resume()` resources.
313
+ *
314
+ * Telemetry is a side-channel: an observer that throws must never crash a run.
315
+ * The harness routes every such throw here and continues. The sink is called at
316
+ * most once per failed hook invocation and MUST NOT rethrow into the run — if it
317
+ * does, the harness contains that too (see {@link safeInvoke}).
318
+ */
319
+ type ObserverErrorSink = (error: unknown, ctx: ObserverErrorContext) => void;
320
+ declare function composeObservers(observers: Observer[], onObserverError?: ObserverErrorSink): Observer;
241
321
 
242
322
  type Cursor = string | ForkCursor;
243
323
  /**
@@ -664,11 +744,10 @@ interface LoopBuilder<S, Ctx> {
664
744
  * with .branch() immediately after. The step immediately following the fork automatically
665
745
  * waits for every branch to settle (the "join") — no separate join primitive exists.
666
746
  *
667
- * Known limitation: the built-in LLM adapters (`@noetaris/harness-anthropic`/`-openai`/
668
- * `-google`/`-ollama`) attribute per-step telemetry via a stored-field pattern that is not
669
- * safe when one adapter instance is shared across concurrently-running branches (the normal
670
- * setup). See `ObserverAware.setStepContext`'s doc comment (`agent/observer.ts`) for details
671
- * — not fixed for this release.
747
+ * Per-step telemetry is attributed correctly across concurrent branches: the harness scopes
748
+ * each branch's slots per step via {@link ObserverAware.withTelemetry}, so telemetry travels
749
+ * with the call rather than through shared instance state. See `ObserverAware`'s doc comment
750
+ * (`agent/observer.ts`).
672
751
  */
673
752
  fork(name: string): LoopBuilder<S, Ctx>;
674
753
  /**
@@ -967,4 +1046,4 @@ declare class LeaseExpiredError extends Error {
967
1046
  constructor(sessionId: string);
968
1047
  }
969
1048
 
970
- export { type Agent, type BranchCursor, type BranchDef, type ClaimOptions, type Cursor, type DeepWithMarkers, type FieldDefinition, type ForkCursor, type ForkDef, type FrameworkState, type Harness, type Lease, LeaseExpiredError, type LoopDefinition, type LoopNode, LoopNotDefinedError, NoInterruptError, type Observer, type ObserverAware, REQUIRED_TAG, RUNTIME_TAG, type RequiredMarker, type RouteFn, type RunContext, type RunFn, type RuntimeMarker, SessionBusyError, SessionInFlightError, SessionPendingInterruptError, type SessionStore, type SignalTransition, type StateFromSchema, type StepContext, type StepDef, type StepState, StoreLoadError, type StoredRun, type StoredRunMetadata, type TransitionTarget, composeObservers, createAgent, createHarness, field, isForkCursor, isForkDef, isRequiredMarker, isRuntimeMarker, required, runtime };
1049
+ export { type Agent, type BranchCursor, type BranchDef, type ClaimOptions, type Cursor, type DeepWithMarkers, type FieldDefinition, type ForkCursor, type ForkDef, type FrameworkState, type Harness, type Lease, LeaseExpiredError, type LoopDefinition, type LoopNode, LoopNotDefinedError, NoInterruptError, type Observer, type ObserverAware, type ObserverErrorContext, type ObserverErrorSink, REQUIRED_TAG, RUNTIME_TAG, type RequiredMarker, type RouteFn, type RunContext, type RunFn, type RuntimeMarker, SessionBusyError, SessionInFlightError, SessionPendingInterruptError, type SessionStore, type SignalTransition, type StateFromSchema, type StepContext, type StepDef, type StepSettledEvent, type StepSettledNext, type StepState, StoreLoadError, type StoredRun, type StoredRunMetadata, type TelemetryContext, type TransitionTarget, composeObservers, createAgent, createHarness, field, isForkCursor, isForkDef, isRequiredMarker, isRuntimeMarker, required, runtime };