@stage5/lumine 0.2.68 → 0.2.70
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +23 -6
- package/lib/admin-featured-history.js +185 -0
- package/lib/admin-featured.js +33 -0
- package/lib/admin-runtime-logs.js +45 -6
- package/lib/admin-workflows.js +52 -2
- package/lib/admin.js +269 -87
- package/lib/assets.js +11 -8
- package/lib/build-review.js +609 -8
- package/lib/commands.js +18 -9
- package/lib/constants.js +26 -1
- package/lib/sdk.js +114 -4
- package/lib/sponsor-duty.js +15 -2
- package/lib/thumbnail.js +2 -2
- package/package.json +1 -1
- package/sdk/BUILD_SDK_INDEX.md +72 -6
- package/sdk/LUMINE_ADMIN.md +344 -21
package/sdk/LUMINE_ADMIN.md
CHANGED
|
@@ -484,6 +484,27 @@ it never authorizes unrelated daily work or generic recommendation commands.
|
|
|
484
484
|
completed all-Featured-comment review. Include refresh progress and any
|
|
485
485
|
carryovers as required in the `Featured rotation` report section below.
|
|
486
486
|
|
|
487
|
+
## Math Lab content duty in full daily management
|
|
488
|
+
|
|
489
|
+
Mikey added Math Lab question design and publishing to the full daily workflow
|
|
490
|
+
on 2026-09-08. Follow [Math Lab daily question publishing](../../agent-guides/math-lab-daily.md)
|
|
491
|
+
for the canonical Build 2460, owner account, 12-grade/36-question editorial
|
|
492
|
+
process, verification, repeat-run recovery, release gates, and final reporting.
|
|
493
|
+
This is not part of Featured-only or newspaper-only work and is not a new
|
|
494
|
+
scheduler, delegated API scope, or automatic extension of admin permissions.
|
|
495
|
+
Use the expressly authorized owner Build workflow for Math Lab; retain the
|
|
496
|
+
normal Zero/Ciel actor separation for other administration.
|
|
497
|
+
|
|
498
|
+
The initial private draft must remain unpublished until Mikey authorizes its
|
|
499
|
+
launch. After launch, routine content releases follow the standing duty but
|
|
500
|
+
cannot bypass a reward-enabled app's exact-version approval gate. Local edits
|
|
501
|
+
and draft saves do not require release approval. Real XP/Coins may be changed
|
|
502
|
+
only by the currently published, approved artifact through server-verified
|
|
503
|
+
reward claims; private builds, previews, local tests, unpublished branches, and
|
|
504
|
+
superseded versions cannot award real balances. These reward controls are a
|
|
505
|
+
required design contract, not a claim that a reward SDK or server enforcement
|
|
506
|
+
has already been implemented. See the guide before adding reward capabilities.
|
|
507
|
+
|
|
487
508
|
## Escalation to Mikey
|
|
488
509
|
|
|
489
510
|
A full daily management run is not finished when the mutations are done. Curation surfaces things only
|
|
@@ -891,6 +912,27 @@ type DailyRunComplete = Success<{
|
|
|
891
912
|
rotationAdvanced: false; // legacy field; calendar schedules never advance by run
|
|
892
913
|
}>;
|
|
893
914
|
type DailyRunFail = DailyRunComplete;
|
|
915
|
+
// `escalation add` echoes the same open-item shape `escalation list` returns,
|
|
916
|
+
// so the run.escalation audit ID that `escalation set <auditId>` needs is in
|
|
917
|
+
// the add response; no list round-trip is required. `recordedAt` (the audit
|
|
918
|
+
// row's timestamp) appears only in `list`.
|
|
919
|
+
type DailyRunEscalationAdd = Success<{
|
|
920
|
+
escalation: {
|
|
921
|
+
auditId: number; // run.escalation audit event ID = escalation identity
|
|
922
|
+
runId: number;
|
|
923
|
+
targetType: string | null;
|
|
924
|
+
targetId: number | null;
|
|
925
|
+
url: string | null;
|
|
926
|
+
summary: string;
|
|
927
|
+
severity: "attention" | "urgent";
|
|
928
|
+
status: "open";
|
|
929
|
+
decisionNote: null;
|
|
930
|
+
decisionAuditId: null;
|
|
931
|
+
decisionRevision: 0;
|
|
932
|
+
decisionUpdatedAt: null;
|
|
933
|
+
decisionByUserId: null;
|
|
934
|
+
};
|
|
935
|
+
}>;
|
|
894
936
|
```
|
|
895
937
|
|
|
896
938
|
### Build Workshop sponsor applications and integrity
|
|
@@ -1038,6 +1080,53 @@ migration cannot quietly treat unsupported telemetry as an empty list.
|
|
|
1038
1080
|
A Featured-only start instead returns an explicitly suppressed, empty handoff
|
|
1039
1081
|
and performs no todo reads, writes, capacity checks, or surfacing increments.
|
|
1040
1082
|
|
|
1083
|
+
Evidence-dependent investigations must have a durable collection plan. Do not
|
|
1084
|
+
carry “still no recycle-under-load evidence” forward without checking the
|
|
1085
|
+
collector and recording new observations. Full starts and reports return
|
|
1086
|
+
`runtimeEvidence` with a scoped read command; this handoff is suppressed for
|
|
1087
|
+
Featured/newspaper slices. During a full daily run with pending runtime
|
|
1088
|
+
investigations, use:
|
|
1089
|
+
|
|
1090
|
+
```sh
|
|
1091
|
+
lumine admin runtime evidence primary --days 7 --json
|
|
1092
|
+
# Only when the investigation also covers the configured target host:
|
|
1093
|
+
lumine admin runtime evidence target --days 7 --json
|
|
1094
|
+
```
|
|
1095
|
+
|
|
1096
|
+
This is a read-only, run-independent command. It does not acquire/clear a log
|
|
1097
|
+
lease, trigger a recycle, or authorize an unrelated management review. Host
|
|
1098
|
+
routing is explicit and never silently substitutes primary for target. An
|
|
1099
|
+
older API without this endpoint is unsupported, not “no incidents.”
|
|
1100
|
+
|
|
1101
|
+
The cluster primary samples existing worker health snapshots once per minute
|
|
1102
|
+
and records memory-guard/operator recycle lifecycle events. Records survive
|
|
1103
|
+
restarts in the health directory's `evidence/` subdirectory (seven UTC calendar
|
|
1104
|
+
days, at most 2 MiB/day; 256 KiB/day reserved for events). Work is recorded as
|
|
1105
|
+
counts by kind, not user IDs, request labels, messages or tokens. Evidence is
|
|
1106
|
+
passive: never force production load, a worker recycle or a restart merely to
|
|
1107
|
+
complete an experiment. New collector code requires activation of the new
|
|
1108
|
+
**primary generation**, not just a rolling worker reload; verify a fresh live
|
|
1109
|
+
sample before saying collection is running.
|
|
1110
|
+
|
|
1111
|
+
For each affected todo, persist the checked host/runtime identity, previous and
|
|
1112
|
+
new evidence cutoff, sample count/gaps/freshness, qualifying event IDs and the
|
|
1113
|
+
next safe action. `unavailable`, `stale` or `incomplete` means investigate the
|
|
1114
|
+
collection gap. A healthy collector with no qualifying recycle means keep
|
|
1115
|
+
collecting, not stall or claim a failure. Save the relevant summary in the todo
|
|
1116
|
+
before the seven-day retention window passes. Missing/stale active-work or OOM
|
|
1117
|
+
observations remain unknown; they are never zero or idle by default.
|
|
1118
|
+
|
|
1119
|
+
Set investigation-specific acceptance criteria before interpreting results.
|
|
1120
|
+
For memory/recycle investigations, distinguish a long-lived steady-memory
|
|
1121
|
+
baseline from a recycle observed under load. Require the requested/signalled/
|
|
1122
|
+
recovered sequence, fresh pre-recycle work evidence, replacement identity,
|
|
1123
|
+
bounded recovery time, and no OOM-counter increase across the same primary
|
|
1124
|
+
generation. “Topology recovered” does **not** prove each interrupted user task
|
|
1125
|
+
completed: correlate those tasks' existing canonical outcomes before closing
|
|
1126
|
+
a user-work continuity investigation. Never close based only on deployed code,
|
|
1127
|
+
a current healthy snapshot, or an unobserved event. Daily runs collect evidence
|
|
1128
|
+
and update todos; fixes and releases follow the project's existing authority rules.
|
|
1129
|
+
|
|
1041
1130
|
`kind` is `task` or `experiment`. New items may start `open`, `in_progress`, or
|
|
1042
1131
|
`blocked`; updates may also use `completed` or `cancelled`. A progress note is
|
|
1043
1132
|
required for every update. For experiments, put the acceptance criteria in the
|
|
@@ -1101,6 +1190,8 @@ lumine admin subjects candidates --effort unassigned --json
|
|
|
1101
1190
|
lumine admin subjects candidates --unviewed --json
|
|
1102
1191
|
lumine admin builds candidates --all --limit 50 --json
|
|
1103
1192
|
lumine admin builds review build:884 --output-dir ./build-review --json
|
|
1193
|
+
lumine admin builds review build:884 --output-dir ./build-review \
|
|
1194
|
+
--interact ./build-review/steps.json --json
|
|
1104
1195
|
```
|
|
1105
1196
|
|
|
1106
1197
|
Schemas:
|
|
@@ -1215,8 +1306,10 @@ verifies the request fingerprint and confirmed spool digest before requesting
|
|
|
1215
1306
|
another page. An interrupted request is never counted as queue coverage.
|
|
1216
1307
|
|
|
1217
1308
|
Recommendations default to `--since-run`: the server uses the previous
|
|
1218
|
-
completed run's start time
|
|
1219
|
-
|
|
1309
|
+
completed full run's start time, even if that gap exceeds 30 days. On a first
|
|
1310
|
+
run, the fallback begins seven days before the current run's stored start, so
|
|
1311
|
+
it cannot drift between pages. Insight reports retain their separate 30-day
|
|
1312
|
+
limit. That deliberate start-to-start overlap gives the queue
|
|
1220
1313
|
at-least-once coverage when content arrives after the prior snapshot but before
|
|
1221
1314
|
that run completes. `--after` supplies an explicit inclusive timestamp.
|
|
1222
1315
|
All-history traversal is deliberately available only through
|
|
@@ -1224,6 +1317,23 @@ All-history traversal is deliberately available only through
|
|
|
1224
1317
|
boundary for bounded modes, so deploying a new CLI against an older API cannot
|
|
1225
1318
|
silently fall back to a million-row historical scan.
|
|
1226
1319
|
|
|
1320
|
+
`recommendations list` is the Earn Recommend picker, not a "new comments"
|
|
1321
|
+
feed: it returns only comments on subjects with an assigned effort level,
|
|
1322
|
+
whose length exceeds that level's threshold (>100 / >250 / >450 / >700
|
|
1323
|
+
characters for effort ≤2 / 3 / 4 / 5), with no skip row and no existing
|
|
1324
|
+
recommendation from any effective Level 5+ user (1000+ AP or Teacher
|
|
1325
|
+
authority), Zero, or Ciel. Community-recommended and short comments are
|
|
1326
|
+
deliberately absent, so an empty page or an empty window is normal and is
|
|
1327
|
+
not evidence of a broken walk; Featured-subject comments are reviewed through
|
|
1328
|
+
the Featured comment scan instead. A successful `post recommend` does not
|
|
1329
|
+
imply the target was queue-eligible.
|
|
1330
|
+
|
|
1331
|
+
After upgrading the API and CLI to the stable run-start window, start a fresh
|
|
1332
|
+
`--since-run` scan with a new checkpoint path, without `--resume`. Older
|
|
1333
|
+
since-run checkpoints are rejected even when already exhausted: they may have
|
|
1334
|
+
captured the former 30-day reporting cap. They are left intact as evidence.
|
|
1335
|
+
Explicit `--after` and `--include-legacy` checkpoints retain their contracts.
|
|
1336
|
+
|
|
1227
1337
|
Subject candidates follow the same window contract. They default to the
|
|
1228
1338
|
previous completed full run's start (with the seven-day first-run fallback),
|
|
1229
1339
|
accept an explicit inclusive `--after`, and require `--include-legacy` for a
|
|
@@ -1231,9 +1341,16 @@ lifetime traversal. `--since-run`, `--after`, and `--include-legacy` are
|
|
|
1231
1341
|
mutually exclusive. The CLI also requires the API to echo the bounded Subject
|
|
1232
1342
|
window before accepting a page.
|
|
1233
1343
|
|
|
1234
|
-
`builds candidates`
|
|
1235
|
-
|
|
1236
|
-
|
|
1344
|
+
`builds candidates` uses the admin publication-window endpoint, ordered by
|
|
1345
|
+
`publishedAt` and Build ID, not workspace `updatedAt`. Like Subject discovery,
|
|
1346
|
+
it defaults to the previous completed full run's start (seven-day first-run
|
|
1347
|
+
fallback), accepts inclusive `--after`, and requires `--include-legacy` for
|
|
1348
|
+
all history. These flags are mutually exclusive. The first page freezes the
|
|
1349
|
+
time boundary and an artifact-version high-water mark; the cursor and local
|
|
1350
|
+
checkpoint retain both. An old public-browser checkpoint cannot be reused.
|
|
1351
|
+
The CLI fails closed if the API does not confirm this publication window.
|
|
1352
|
+
|
|
1353
|
+
It is available through the `admin` namespace only while a delegated run is active;
|
|
1237
1354
|
page until `pagination.exhausted`. Each item includes its canonical app URL,
|
|
1238
1355
|
published artifact version, and whether its code is pullable. This list does
|
|
1239
1356
|
not decide that an app deserves a comment. The management agent must open and
|
|
@@ -1241,6 +1358,12 @@ genuinely try the published runtime, or pull and read an open-source project,
|
|
|
1241
1358
|
before making that judgment. Direct API/persona automation is never a review
|
|
1242
1359
|
substitute.
|
|
1243
1360
|
|
|
1361
|
+
Only current public Main releases are candidates. Workspace-only saves and
|
|
1362
|
+
unchanged reactivations do not become new releases. A Build republished during
|
|
1363
|
+
paging can leave the snapshot; its new release is reconsidered by the next
|
|
1364
|
+
overlapping start-to-start window. Always recheck the current artifact before
|
|
1365
|
+
reviewing; this is not a frozen copy of an app or an immutable release archive.
|
|
1366
|
+
|
|
1244
1367
|
`builds review` is the managed runtime path: it fetches the current published
|
|
1245
1368
|
artifact identity, launches the app in an isolated temporary Chromium profile,
|
|
1246
1369
|
captures a screenshot and bounded console evidence, then fetches the identity
|
|
@@ -1253,6 +1376,53 @@ learned during that review. The receipt binds the draft to the exact reviewed
|
|
|
1253
1376
|
artifact without copying a version number by hand; the server owns the Build,
|
|
1254
1377
|
version, method, and review-time fields around that understanding.
|
|
1255
1378
|
|
|
1379
|
+
Without `--interact` the review captures only the start screen after
|
|
1380
|
+
`--wait-ms`. `--interact <steps.json>` adds a bounded, ordered interaction
|
|
1381
|
+
script that runs inside the app's runtime iframe after that start screenshot,
|
|
1382
|
+
so the receipt can show what happens when the app is actually used. The file is
|
|
1383
|
+
a JSON array (or `{ "steps": [...] }`) of at most **12** steps, each exactly one
|
|
1384
|
+
of:
|
|
1385
|
+
|
|
1386
|
+
```json
|
|
1387
|
+
[
|
|
1388
|
+
{ "click": "text=Start" },
|
|
1389
|
+
{ "wait": 1500 },
|
|
1390
|
+
{ "screenshot": "after-start" },
|
|
1391
|
+
{ "type": { "selector": "input[name=name]", "text": "Zero" } },
|
|
1392
|
+
{ "press": "Enter" },
|
|
1393
|
+
{ "press": "ArrowLeft" },
|
|
1394
|
+
{ "screenshot": "moved" }
|
|
1395
|
+
]
|
|
1396
|
+
```
|
|
1397
|
+
|
|
1398
|
+
- `click`: a CSS selector, or `text=<visible text>` (case-insensitive; the
|
|
1399
|
+
smallest visible element whose text matches, then a bounded contains match).
|
|
1400
|
+
Dispatched as a trusted mouse click at the element's centre, so canvas games
|
|
1401
|
+
and buttons both receive it.
|
|
1402
|
+
- `type`: `{ selector, text }` — clicks the element, then inserts single-line
|
|
1403
|
+
text (at most 200 characters) as trusted input.
|
|
1404
|
+
- `press`: `Enter`, `Space`, `Escape`, `Tab`, `Backspace`, `ArrowUp/Down/Left/Right`,
|
|
1405
|
+
a letter, or a digit. The frame is focused first if nothing was clicked yet.
|
|
1406
|
+
- `wait`: 1–5000 ms.
|
|
1407
|
+
- `screenshot`: a unique label (1–40 letters/digits/`-`/`_`, not `runtime` or
|
|
1408
|
+
`review`); saved as `<label>.png` beside `runtime.png`.
|
|
1409
|
+
|
|
1410
|
+
The whole script is capped at **60 s**; it stops at the first failed step
|
|
1411
|
+
(element not found/not visible, budget exhausted, frame unreachable). Console
|
|
1412
|
+
evidence stays bounded exactly as before. `review.json` gains
|
|
1413
|
+
`screenshots: [{ label, path, bytes }]` (script screenshots only; the start
|
|
1414
|
+
screen stays in `screenshot`) and
|
|
1415
|
+
`interaction: { path, stepsPlanned, stepsCompleted, status, failedStep, frame,
|
|
1416
|
+
elapsedMs, steps }` with a per-step record (coordinates for clicks, the saved
|
|
1417
|
+
path for screenshots, the error for a failed step). Both are `[]`/`null`
|
|
1418
|
+
without `--interact`. The review remains one receipt bound to one artifact: the
|
|
1419
|
+
published version is re-read after the script finishes, and a script that did
|
|
1420
|
+
not complete makes the receipt `failed` with
|
|
1421
|
+
`CLI_ADMIN_BUILD_REVIEW_INTERACTION_FAILED` (the completed steps and their
|
|
1422
|
+
screenshots are still listed) so a draft can never cite interactions that did
|
|
1423
|
+
not happen. `comment draft --review-receipt` accepts a receipt only when every
|
|
1424
|
+
listed screenshot still exists unchanged and the script completed.
|
|
1425
|
+
|
|
1256
1426
|
During every full daily management review, scan recent Build candidates back through the
|
|
1257
1427
|
run's review window alongside Subjects and the recommendation queue. An app
|
|
1258
1428
|
that is thin, broken, private, unchanged since a prior substantive bot
|
|
@@ -1551,6 +1721,16 @@ after the finalized coverage boundary. `null` means the subject predates provabl
|
|
|
1551
1721
|
coverage—never convert that unknown into “never Featured.” Both the website
|
|
1552
1722
|
editor and Lumine mutations write this append-only history in the same
|
|
1553
1723
|
transaction as the canonical board replacement.
|
|
1724
|
+
Each API read accepts up to **100 subject IDs**, independently of the
|
|
1725
|
+
20-subject delegated addition policy. `--all` automatically batches larger
|
|
1726
|
+
lists (up to 20,000 IDs), exhausts every batch's event pages, and retains all
|
|
1727
|
+
per-subject summaries. Use the exact command with `--resume` after interruption;
|
|
1728
|
+
confirmed pages are not replayed. Single-batch `--cursor` remains available,
|
|
1729
|
+
but cannot be combined with `--all`. Multi-batch results explicitly use
|
|
1730
|
+
`pagination.snapshotScope: "per-batch"`; events are ordered by input batch,
|
|
1731
|
+
then descending event ID within that batch—not by one global snapshot.
|
|
1732
|
+
`data.scan.batches` records each batch's coverage, snapshot and private spool.
|
|
1733
|
+
Deploy the matching API before using the expanded read bound.
|
|
1554
1734
|
For a retry whose board transaction committed but whose canonical detail reload
|
|
1555
1735
|
failed, the audit-linked history event is the durable receipt: the API re-reads
|
|
1556
1736
|
the current board and preserves the original changed-mutation accounting.
|
|
@@ -1637,7 +1817,8 @@ lumine admin featured comments scan --checkpoint featured-read.json --json
|
|
|
1637
1817
|
# On interruption: repeat with --resume, in the same active run.
|
|
1638
1818
|
# Read ALL pageFiles, including full root context and nested replies.
|
|
1639
1819
|
lumine admin featured comments acknowledge \
|
|
1640
|
-
--checkpoint featured-read.json --reviewed
|
|
1820
|
+
--checkpoint featured-read.json --reviewed \
|
|
1821
|
+
--decisions-template featured-decisions.json --json
|
|
1641
1822
|
```
|
|
1642
1823
|
|
|
1643
1824
|
The scan snapshots every current Featured Subject (up to 100) and its maximum
|
|
@@ -1658,6 +1839,26 @@ missing Subject IDs and comment counts, and distinguishes complete from partial
|
|
|
1658
1839
|
coverage. A download alone never counts as a read or grants recommendations.
|
|
1659
1840
|
After a resumed scan fills a gap, read those pages and acknowledge again.
|
|
1660
1841
|
|
|
1842
|
+
Acknowledge returns `data.reviewedCoverage` (the coverage receipt; its `id` is
|
|
1843
|
+
the `coverageId` the recommend step needs — it is a different audit row from
|
|
1844
|
+
the review ID) and a ready-to-fill `data.decisionsTemplate`
|
|
1845
|
+
`{ reviewId, coverageId, selections: [] }`. With `--decisions-template <file>`
|
|
1846
|
+
the CLI also writes that template as a private mode-0600 file; only
|
|
1847
|
+
`acknowledge --reviewed` writes it (a scan rejects the flag). Start the
|
|
1848
|
+
decisions file from the template rather than assembling the identifiers by
|
|
1849
|
+
hand. `readFeaturedSelections` refuses a file whose `coverageId` equals its
|
|
1850
|
+
`reviewId` before any request is sent, and the API answers a wrong receipt
|
|
1851
|
+
with `CLI_ADMIN_FEATURED_COVERAGE_MISMATCH` naming the expected receipt, e.g.
|
|
1852
|
+
`coverageId 5719 is not the acknowledged coverage receipt for review 5719
|
|
1853
|
+
(expected 5749; 5719 is the review ID itself).` with
|
|
1854
|
+
`details: { reviewId, suppliedCoverageId, expectedCoverageId,
|
|
1855
|
+
suppliedReceiptAction, suppliedCoverageReviewId }`, or
|
|
1856
|
+
`CLI_ADMIN_FEATURED_COVERAGE_MISSING` when the review has no acknowledged
|
|
1857
|
+
coverage in this run. Other receipt failures (a page ID from another run,
|
|
1858
|
+
operator, or actor) keep the generic
|
|
1859
|
+
`Receipt #<id> is not a completed <action> receipt of this active run,
|
|
1860
|
+
operator and actor.` rejection.
|
|
1861
|
+
|
|
1661
1862
|
Compose a decisions JSON file from the genuinely reviewed comments, using the
|
|
1662
1863
|
returned review ID, coverage receipt ID, and each selected comment's page ID:
|
|
1663
1864
|
|
|
@@ -2108,10 +2309,23 @@ type NewsSubmit = NewsStatus; // "success"; newspaper includes revisionNumber
|
|
|
2108
2309
|
lumine admin bot-output --json
|
|
2109
2310
|
lumine admin bot-output --days 3 --json
|
|
2110
2311
|
lumine admin bot-output --cursor '<pagination.nextCursor>' --json
|
|
2312
|
+
lumine admin bot-output context 3797910 --reason "Review reported bot conduct in its conversation context" --json
|
|
2313
|
+
lumine admin bot-output context 3797910 --reason "Continue the same bot-conduct review" --cursor '<pagination.nextCursor>' --json
|
|
2111
2314
|
```
|
|
2112
2315
|
|
|
2113
2316
|
**Every full daily review reads what Zero and Ciel themselves said since the
|
|
2114
2317
|
last completed full review.**
|
|
2318
|
+
|
|
2319
|
+
**Ordinary wrong answers and hallucinations are expected model limitations,
|
|
2320
|
+
not website incidents.** A factual error, mistaken puzzle answer, or imperfect
|
|
2321
|
+
reasoning alone does not warrant an escalation, engineering todo, or a code
|
|
2322
|
+
patch. Model quality improves through LLM upgrades; do not add hard-coded
|
|
2323
|
+
answer validators, secondary graders, forced research, correctness retry loops,
|
|
2324
|
+
or subject-specific rules to compensate. A normal conversational correction is
|
|
2325
|
+
enough when appropriate. This does not excuse actual harmful conduct or
|
|
2326
|
+
application failures, nor weaken security, permissions, billing, or canonical
|
|
2327
|
+
server-state checks: investigate those distinct problems on concrete evidence.
|
|
2328
|
+
|
|
2115
2329
|
The bots talk to children constantly — chat replies, Daily Reflection
|
|
2116
2330
|
responses, autonomous comment-assistant comments — and a harmful message must
|
|
2117
2331
|
never depend on a kid being brave enough to report it (real incident,
|
|
@@ -2120,7 +2334,9 @@ streak — "I'm telling you: Stop", guilt framing, ordering him to quit Daily
|
|
|
2120
2334
|
Reflections — and it surfaced only because the kid showed Mikey).
|
|
2121
2335
|
|
|
2122
2336
|
`bot-output` returns, windowed since the operator's last completed full run
|
|
2123
|
-
(`--days 1..30` overrides
|
|
2337
|
+
(`--days 1..30` overrides; the bare form sends no `days` parameter at all, so
|
|
2338
|
+
the API applies that default window — an older CLI wrongly validated the empty
|
|
2339
|
+
default and failed with "--days must be an integer"): `chatMessages` (every stored Zero/Ciel chat and
|
|
2124
2340
|
reflection reply, with full text and recipient metadata when its best-effort
|
|
2125
2341
|
prompt audit exists) and `comments`
|
|
2126
2342
|
(every public bot comment/reply). Individual utterances are returned in full;
|
|
@@ -2137,6 +2353,27 @@ right after the brief, and **read every row** — the tool deliberately does no
|
|
|
2137
2353
|
filtering, scoring, or keyword matching, because the judgment is the reviewing
|
|
2138
2354
|
agent's.
|
|
2139
2355
|
|
|
2356
|
+
Chat output now also includes `messageKind`, `attachment`, and the stored
|
|
2357
|
+
`generation` outcome (success, failure, cancelled, generating, or unresolved).
|
|
2358
|
+
This is stored-message evidence, not a live request-guard check: do not call
|
|
2359
|
+
an empty row an orphan solely from its text. Hidden attachment locations and
|
|
2360
|
+
arbitrary settings/request keys are never returned. `source: voice` identifies
|
|
2361
|
+
newly recorded voice transcripts, while typed input during a call says `typed`;
|
|
2362
|
+
older replies correctly say `not-recorded`
|
|
2363
|
+
because the historical schema did not distinguish typed text from voice.
|
|
2364
|
+
|
|
2365
|
+
`bot-output context <messageId>` is a private, **run-independent** investigation.
|
|
2366
|
+
It requires a 1–500 character reason and records a minimized access receipt,
|
|
2367
|
+
not the private message text, in the audit. It returns the specified existing
|
|
2368
|
+
Zero/Ciel output plus preceding messages in that bot's own two-person
|
|
2369
|
+
conversation and exact topic/subchannel, oldest first within each page.
|
|
2370
|
+
Default 20, maximum 40 messages per page; continue manually with the cursor.
|
|
2371
|
+
The complete scope is bounded to 100 prior rows, 24 hours, and 10,000 message
|
|
2372
|
+
IDs before the anchor. `boundedLimitReached` means stop and report that bound,
|
|
2373
|
+
not that all channel history was reviewed. Deleted messages and hidden
|
|
2374
|
+
attachments remain hidden. Group-channel browsing, `--all`, and `--days` are
|
|
2375
|
+
not supported. Never start an entire daily run just to investigate one reply.
|
|
2376
|
+
|
|
2140
2377
|
### API runtime-log review (same phase, every full daily review)
|
|
2141
2378
|
|
|
2142
2379
|
The bot-conduct review also owns a bounded production API log review. Bot
|
|
@@ -2145,6 +2382,24 @@ the HTTP layer while stdout records a degraded fallback/retry loop or stderr
|
|
|
2145
2382
|
records a side-effect failure. Reviewing only `bot-output` can therefore miss
|
|
2146
2383
|
the other half of what happened.
|
|
2147
2384
|
|
|
2385
|
+
For passive RSS/recycle investigations, use the read-only command independently
|
|
2386
|
+
of a daily run or production-log review:
|
|
2387
|
+
|
|
2388
|
+
```bash
|
|
2389
|
+
lumine admin runtime evidence primary --days 7 --output ./runtime-evidence.json --json
|
|
2390
|
+
# Use target explicitly only when investigating a configured second host.
|
|
2391
|
+
```
|
|
2392
|
+
|
|
2393
|
+
It performs no restart, log clear, review lease acquisition, or fallback to a
|
|
2394
|
+
different host. `collecting`, `incomplete`, `stale`, and `unavailable` describe
|
|
2395
|
+
evidence coverage, not a verdict that the system is healthy. A 404 means the
|
|
2396
|
+
API route is not deployed; an old primary generation can also lack collector
|
|
2397
|
+
samples after workers update. Record that activation gap and arrange an
|
|
2398
|
+
authorized release—do not silently close the investigation or force a recycle.
|
|
2399
|
+
An observed topology recovery alone does not prove interrupted user work
|
|
2400
|
+
survived. Keep the evidence cutoff, gaps and actual outcomes in the relevant
|
|
2401
|
+
todo so the next run can continue.
|
|
2402
|
+
|
|
2148
2403
|
The current API-side files are:
|
|
2149
2404
|
|
|
2150
2405
|
- `/home/ec2-user/server/logs/twinkle-api.err.log`
|
|
@@ -2157,6 +2412,27 @@ in scope so a later API-side worker is not silently omitted. Use the delegated,
|
|
|
2157
2412
|
run-independent workflow; it holds one server lease across the review and
|
|
2158
2413
|
writes private, digest-verified local artifacts:
|
|
2159
2414
|
|
|
2415
|
+
For the deploy-time two-API topology, `runtime-logs start primary` and
|
|
2416
|
+
`runtime-logs start target` explicitly select the host. Omitted host means
|
|
2417
|
+
primary; pre-migration NULL owners also mean primary. Use a separate private
|
|
2418
|
+
output/session directory for each completed review. Review-ID/session operations
|
|
2419
|
+
route back to the recorded owner; never treat a peer's files as that review's
|
|
2420
|
+
bytes. An unresolved start key cannot be replayed against a different host.
|
|
2421
|
+
The additive host-owner migration and compatible API must be live before this
|
|
2422
|
+
CLI capability is published.
|
|
2423
|
+
|
|
2424
|
+
Review every participating host, including primary private-helper logs. A
|
|
2425
|
+
primary review does not cover the target. Finish the exclusive review before
|
|
2426
|
+
that host is held; a held/unavailable owner returns an explicit retryable failure,
|
|
2427
|
+
not another host's snapshots or an independent log service. Do not abandon its
|
|
2428
|
+
lease merely to bypass a deployment guard. After a planned hold, final shutdown
|
|
2429
|
+
deltas are reviewed via management SSH outside any active lease, recorded, and
|
|
2430
|
+
API stderr is cleared only with the existing guarded `npm run logs:clear-errors`
|
|
2431
|
+
plus post-clear re-read. This is the deployment runbook's final boundary, not
|
|
2432
|
+
permission to bypass an active Lumine lease. A stopped target whose final logs
|
|
2433
|
+
were reviewed does not need to be started for daily management; starting EC2
|
|
2434
|
+
requires separate authority. See `twinkle-api/DEPLOY_TIME_HANDOFF.md`.
|
|
2435
|
+
|
|
2160
2436
|
```bash
|
|
2161
2437
|
lumine admin runtime-logs start --output-dir ./runtime-log-review --json
|
|
2162
2438
|
# Read every file under data.artifacts.latestSnapshot.snapshotPath.
|
|
@@ -2227,7 +2503,16 @@ checking, and in-place truncation occur on the same open descriptor. It then
|
|
|
2227
2503
|
returns `post_clear_review_required` with another immutable snapshot. That
|
|
2228
2504
|
snapshot also captures normal-output bytes that arrived after the prior
|
|
2229
2505
|
acknowledged cutoff, so routine stdout traffic cannot make the review infinite.
|
|
2230
|
-
Read it and run the same `finish --reviewed` command again.
|
|
2506
|
+
Read it and run the same `finish --reviewed` command again. **Only
|
|
2507
|
+
`data.completionStatus: "completed"` means the review is done.** The top-level
|
|
2508
|
+
`status` mirrors it: `"needs_review"` for both non-terminal outcomes
|
|
2509
|
+
(`needs_review` and `post_clear_review_required`, including a recovered
|
|
2510
|
+
pending snapshot), `"success"` only when `completed`, and `"already_done"` for
|
|
2511
|
+
a replay of an already-finished review. `ok` stays `true` in every case; a
|
|
2512
|
+
`needs_review` result is a valid response that requires another read plus
|
|
2513
|
+
finish, not an error. The CLI derives the top-level status from
|
|
2514
|
+
`completionStatus`, so it is correct against an API that still answers the
|
|
2515
|
+
older `success` envelope. A review clears
|
|
2231
2516
|
`twinkle-api.err.log` at most once. The lease closes when the reviewed error
|
|
2232
2517
|
boundary is still stable, i.e. every byte now in the API error log arrived
|
|
2233
2518
|
after that clear and was captured and acknowledged; errors that arrive before
|
|
@@ -2831,12 +3116,19 @@ farm-signal sections added that day; AI Card summon watch added 2026-08-24):
|
|
|
2831
3116
|
is never called a dodge. Offers without redemptions are a reason to inspect
|
|
2832
3117
|
sample size, event type, and offer age — not proof of a broken funnel by
|
|
2833
3118
|
themselves.
|
|
2834
|
-
- `goneQuiet` — the inverse of `notableCandidates
|
|
2835
|
-
|
|
2836
|
-
|
|
2837
|
-
|
|
2838
|
-
|
|
2839
|
-
|
|
3119
|
+
- `goneQuiet` — the inverse of `notableCandidates`, and window-relative: a
|
|
3120
|
+
user "went quiet" the moment 7 days of silence passed since their
|
|
3121
|
+
`lastActive`, and the section lists only the users whose quiet moment fell
|
|
3122
|
+
inside the brief's window (`lastActive` between `sinceTs - 7d` and
|
|
3123
|
+
`now - 7d`). A one-day brief therefore reports one day of crossings and
|
|
3124
|
+
contiguous daily runs list each user once; it never re-lists everyone seen
|
|
3125
|
+
in the past fortnight. `totals.wentQuiet` is that cohort; `previouslyRegular`
|
|
3126
|
+
is the subset with at least 7 distinct active days (completed daily tasks or
|
|
3127
|
+
Wordle plays) in the 30 days ending on their own last active day, and
|
|
3128
|
+
`users` is that subset ranked by `regularityScore` (`dailyTasks * 2 +
|
|
3129
|
+
wordlePlays`, both measured over the same per-user span), capped at 15 with
|
|
3130
|
+
`daysQuiet`. Use it for product signal (what did they stop doing?) and
|
|
3131
|
+
gentle outreach candidates; never guilt a child in public about absence.
|
|
2840
3132
|
- `newUserFunnel` — signups in the window with `activeOnDayOne` (any
|
|
2841
3133
|
XP-ledger event within 24h of joining) and `returnedAfterDayOne`
|
|
2842
3134
|
(`lastActive` beyond their first day), plus the newest few accounts.
|
|
@@ -3137,11 +3429,38 @@ cannot choose or override that identity. The authorization lasts ten minutes,
|
|
|
3137
3429
|
is bound to that exact comment, permits only reading that target and editing
|
|
3138
3430
|
it, and never runs daily duties, changes the Bangkok calendar assignment, or
|
|
3139
3431
|
contributes to a daily-run mutation count. A correction cannot target a human
|
|
3140
|
-
comment, notification record, deleted comment
|
|
3141
|
-
|
|
3142
|
-
the published
|
|
3432
|
+
comment, notification record, or deleted comment. A Build comment must belong
|
|
3433
|
+
to a public, published canonical owner Build; editing it requires a fresh review
|
|
3434
|
+
of the exact published version and private review context. Starting a newer
|
|
3435
|
+
correction supersedes an older active one.
|
|
3143
3436
|
Finish it explicitly after the canonical edit is confirmed.
|
|
3144
3437
|
|
|
3438
|
+
For a Build, inspect the current app and its full discussion first. Managed
|
|
3439
|
+
runtime review does not require a daily run and does not start one:
|
|
3440
|
+
|
|
3441
|
+
```bash
|
|
3442
|
+
lumine admin builds review build:884 --output-dir ./build-review --json
|
|
3443
|
+
lumine admin correction start 456 --json
|
|
3444
|
+
lumine admin comment edit 456 --file corrected.md \
|
|
3445
|
+
--review-receipt ./build-review/<returned-review-directory>/review.json \
|
|
3446
|
+
--review-context context.json --json
|
|
3447
|
+
lumine admin correction complete <sessionId> --json
|
|
3448
|
+
```
|
|
3449
|
+
|
|
3450
|
+
Use the exact receipt path returned by `builds review`, or pass manual
|
|
3451
|
+
`--reviewed-version <artifactId> --reviewed-via runtime|code` evidence instead.
|
|
3452
|
+
The private context file contains only `{"understanding":"What you actually reviewed"}`.
|
|
3453
|
+
The API locks the Build and comment, verifies the current version and ownership,
|
|
3454
|
+
and commits the text, mention updates, and a new immutable review record together.
|
|
3455
|
+
Only the edited comment's context link moves; older bot replies retain their
|
|
3456
|
+
historical review context. Changed/deleted comments, changed versions, and
|
|
3457
|
+
private/noncanonical Builds fail without a partial edit. A fresh review can be
|
|
3458
|
+
stored even if the public text is unchanged. The CLI requires the exact edited
|
|
3459
|
+
text plus `edit.buildReviewContextStored: true`, the reviewed version, and a
|
|
3460
|
+
canonical review record ID before claiming success. Older APIs that still block
|
|
3461
|
+
Build edits must be deployed first; do not silently substitute a duplicate reply
|
|
3462
|
+
when Mikey requested an edit.
|
|
3463
|
+
|
|
3145
3464
|
**Editing the bot's own comments.** `comment edit <commentId> --file
|
|
3146
3465
|
<comment.md>` replaces the text of a comment the acting bot itself authored —
|
|
3147
3466
|
for correcting a factual error, an unfulfillable claim, or outdated guidance
|
|
@@ -3153,7 +3472,8 @@ composed-comment rules (plain UTF-8, 10,000-character limit, truth about what
|
|
|
3153
3472
|
the session actually did) and publishes through the website's canonical
|
|
3154
3473
|
comment-edit pipeline — mentions are reprocessed (a newly added `@mikey`
|
|
3155
3474
|
notifies him), and Earn-candidate projections resync. Submitting identical
|
|
3156
|
-
text returns `already_done
|
|
3475
|
+
text returns `already_done` for non-Build comments; a Build edit can still save
|
|
3476
|
+
a fresh review without changing its text. It requires either the exact active correction
|
|
3157
3477
|
session above or the `comment:post` scope of a comment-mode `post` run, and is
|
|
3158
3478
|
audited as `comment.edit` with the previous content in `beforeState` and
|
|
3159
3479
|
`data.edit.previousContent`. Edit sparingly:
|
|
@@ -3399,9 +3719,12 @@ sponsor pays from their own battery. Mentions elsewhere in a Build, replies to
|
|
|
3399
3719
|
unlinked or legacy bot comments, replies to humans, and the other bot remain
|
|
3400
3720
|
ineligible. If the published version has changed, the responder is told the
|
|
3401
3721
|
stored understanding belongs to the reviewed older version and must say it has
|
|
3402
|
-
not checked behavior that could have changed.
|
|
3403
|
-
|
|
3404
|
-
|
|
3722
|
+
not checked behavior that could have changed. Editing the acting bot's own
|
|
3723
|
+
Build comment is supported with the same fresh reviewed-version/method and
|
|
3724
|
+
private-context flags, including managed review receipts. It updates the
|
|
3725
|
+
existing comment, not a duplicate reply. See the narrow correction workflow
|
|
3726
|
+
above; a version-bound follow-up remains appropriate when the conversation
|
|
3727
|
+
calls for an additional reply instead of an edit.
|
|
3405
3728
|
|
|
3406
3729
|
**Offer a Lumine prompt when the moment invites it (Mikey's direction,
|
|
3407
3730
|
2026-08-10).** Zero and Ciel may include one concrete, copy-pasteable Lumine
|