humanish 0.88.1 → 0.88.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,57 @@
1
+ # Humanish 0.88.2: sequential studies honor model-spend thresholds
2
+
3
+ Sequential shared-world studies now enforce the `execution.caps.maxUsd` and
4
+ `maxTotalUsd` thresholds they previously accepted without applying. This covers
5
+ computer-use participants sharing a clone or local-tree subject with
6
+ `subject.topology: shared-world` and `execution.concurrency: 1`.
7
+
8
+ Each participant's reported usage feeds the per-participant threshold and the
9
+ shared study estimate. Final usage, including a closing report, is reconciled
10
+ before the next participant starts. A participant interrupted by a threshold
11
+ has `budget_reached` / `incomplete`; its reason distinguishes prior activity
12
+ from no material progress. Later participants blocked by the shared threshold
13
+ make no model requests and add no executed turn to the checkpoint timeline.
14
+ A closing report that crosses the threshold preserves the already recorded
15
+ completion condition while blocking subsequent participants.
16
+ The bundle retains all declared participants and an explicit blocked suffix.
17
+ Verification checks the suffix against its preceding interruption, participant
18
+ records and unchanged executed timeline; absent participants cannot masquerade
19
+ as budget-blocked seats.
20
+
21
+ These checks happen after a model response and before its actions or another
22
+ participant turn. The current request can overshoot a threshold. They estimate
23
+ model spend; they do not reserve future requests or cap provider invoices,
24
+ desktop compute, or target-app charges. A zero threshold can still allow the
25
+ first model request. Use an explicit dry-run for a path without provider calls.
26
+
27
+ An unpriced model with a declared threshold is refused before allocation. For
28
+ an otherwise completed response, missing or partial usage ends capped sequential
29
+ execution with `harness_error` and `usage_unreported` (“provider usage unavailable”).
30
+ The CUA loop sends no further participant request, retry or closing request.
31
+ An explicitly incomplete provider response retains its original interruption
32
+ cause first; its missing usage remains unknown. For these strict capped sessions,
33
+ the default OpenAI adapter makes one dispatch per requested turn, including
34
+ HTTP errors and policy negotiation. The loop cancels the request signal when
35
+ its timeout wins, so an outstanding transport cannot retry after the loop ends.
36
+ Injected providers must honor that signal and control their own dispatches;
37
+ cancellation cannot undo an already billed request. Known partial costs stay
38
+ in the trace, and an unknown shared budget blocks later participants.
39
+
40
+ Sequential traces now retain dated model estimates. The run total explicitly
41
+ marks desktop compute as unmeasured, so the model subtotal is a lower bound.
42
+ Observer labels a partial estimate as known cost with the total unknown.
43
+ A custom session returning a different model identity fails orchestration and
44
+ blocks later participants while preserving its original outcome and the
45
+ estimate for its returned model. This does not retrospectively enforce an
46
+ arbitrary custom runner that ignored its cap options.
47
+
48
+ Existing bundles are not rewritten. Uncapped sessions and other execution
49
+ routes retain their existing behavior. The stricter unknown-usage rule applies
50
+ to sequential capped studies; the sequential route still has no running
51
+ Observer usage stream.
52
+
53
+ Verification covers the actual participant loop with captured provider usage,
54
+ individual and shared thresholds, unstarted seats, zero and missing usage,
55
+ closing requests, model mismatches, checkpoint evidence and cleanup. These
56
+ contract checks establish the tested behavior; they do not establish persona
57
+ effectiveness, independent adoption, or exact provider billing.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "humanish",
3
- "version": "0.88.1",
3
+ "version": "0.88.2",
4
4
  "description": "Open-source-safe CLI for persona simulation, observer review, and public-safe feedback drafts.",
5
5
  "author": "Daniel G Wilson <daniel@danielgwilson.com>",
6
6
  "keywords": [