@flame0510/project-aether 1.6.2 → 1.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (52) hide show
  1. package/LICENSE +21 -0
  2. package/README.md +6 -2
  3. package/agent-templates/atlas/HEARTBEAT.md +1 -1
  4. package/app/agents/ImageDownloadBanner.tsx +171 -37
  5. package/app/api/agents/download-image/route.ts +29 -5
  6. package/app/api/agents/image-status/route.ts +17 -2
  7. package/app/api/assistant/route.ts +21 -5
  8. package/app/api/auth/login/route.ts +2 -2
  9. package/app/api/metrics/route.ts +126 -23
  10. package/app/api/setup/agent-image/route.ts +6 -4
  11. package/app/api/system-health/route.ts +25 -21
  12. package/app/components/DashboardToolbar.tsx +2 -2
  13. package/app/components/LineageGraphPage.tsx +4 -4
  14. package/app/components/SessionDrawer.tsx +8 -8
  15. package/app/components/Sidebar.tsx +10 -0
  16. package/app/components/Skeleton.tsx +4 -1
  17. package/app/components/SystemCockpit.tsx +83 -3
  18. package/app/components/ui/Meter.tsx +34 -0
  19. package/app/components/ui/TimeSeriesChart.tsx +226 -0
  20. package/app/components/ui/index.ts +2 -0
  21. package/app/globals.css +42 -1
  22. package/app/setup/PageClient.tsx +1 -1
  23. package/app/system/PageClient.tsx +263 -0
  24. package/app/system/SystemSkeleton.tsx +115 -0
  25. package/app/system/loading.tsx +13 -0
  26. package/app/system/page.tsx +5 -0
  27. package/bin/postinstall.js +5 -1
  28. package/daemon.js +274 -214
  29. package/docs/ARCHITECTURE.md +65 -34
  30. package/docs/CONTAINER-TERMINAL.md +17 -8
  31. package/docs/DESIGN-SYSTEM.md +22 -12
  32. package/docs/FRONTEND-ARCHITECTURE.md +8 -4
  33. package/docs/REV4A.md +37 -86
  34. package/docs/dev/API-REFERENCE.md +102 -19
  35. package/docs/dev/DATABASE.md +79 -28
  36. package/docs/dev/SESSION-MAINTENANCE-PLAN.md +6 -6
  37. package/docs/rag/DATA-FRESHNESS.md +31 -15
  38. package/docs/rag/GLOSSARY.md +8 -5
  39. package/docs/rag/REV4A-OVERVIEW.md +14 -7
  40. package/docs/rag/WHAT-I-CAN-ANSWER.md +4 -3
  41. package/lib/agent-images.ts +43 -14
  42. package/lib/buildAgentImage.ts +142 -6
  43. package/lib/db-bootstrap.mjs +0 -11
  44. package/lib/metrics-db.ts +48 -0
  45. package/lib/patterns/sessionPresentation.ts +4 -2
  46. package/lib/rev4a-auth.d.ts +1 -0
  47. package/lib/rev4a-auth.js +18 -2
  48. package/next.config.mjs +9 -1
  49. package/package.json +2 -2
  50. package/scripts/backup.sh +48 -54
  51. package/scripts/check-language.mjs +21 -3
  52. package/scripts/restore.sh +77 -59
@@ -1,6 +1,6 @@
1
1
  # Rev4a API Reference
2
2
 
3
- > **Last updated:** 2026-09-15
3
+ > **Last updated:** 2026-09-26
4
4
 
5
5
  All routes are under `/api/`. Authentication is required on every endpoint
6
6
  unless otherwise noted.
@@ -249,7 +249,10 @@ Returns current provider configuration state.
249
249
  |---|---|
250
250
  | `summary=1` | Returns only `{ provider, label, configured }` per provider — no `models`, no `pricing`, no `apiKey`. The full response awaits live pricing from OpenRouter, a network round trip costing ~2.4s on a cold cache; callers that only need to know whether *any* provider is configured should use this. |
251
251
 
252
- **Response:**
252
+ **Response:** each provider carries its models, and a model carries `params` (the total
253
+ parameter count) when its Hugging Face card publishes one — the Gateway row shows it next
254
+ to the price, and a model without a card simply omits it.
255
+
253
256
  ```json
254
257
  {
255
258
  "providers": [
@@ -1576,7 +1579,7 @@ nothing opens until an operator approves it.
1576
1579
 
1577
1580
  **Response:**
1578
1581
  ```json
1579
- { "url": "http://212.0.113.7:3710/#token=…", "loopbackHost": false }
1582
+ { "url": "http://203.0.113.7:3710/#token=…", "loopbackHost": false }
1580
1583
  ```
1581
1584
 
1582
1585
  - `url` is the plain token link. The host is the one the request reached Rev4a on
@@ -1737,6 +1740,17 @@ in flight, so polling this endpoint does not re-query the registry every time.
1737
1740
  {
1738
1741
  "downloading": false,
1739
1742
  "downloadingVersion": null,
1743
+ "downloadPercent": null,
1744
+ "downloadMessage": null,
1745
+ "canceled": false,
1746
+ "lastResult": {
1747
+ "version": "2026.9.3",
1748
+ "result": "succeeded",
1749
+ "source": "registry",
1750
+ "message": "OpenClaw 2026.9.3 image downloaded",
1751
+ "at": 1758700000000
1752
+ },
1753
+ "serverTime": 1758700004000,
1740
1754
  "exists": true,
1741
1755
  "needsUpdate": false,
1742
1756
  "localVersions": ["2026.9.3"],
@@ -1754,6 +1768,11 @@ in flight, so polling this endpoint does not re-query the registry every time.
1754
1768
  | `newestLocal` | The version new agents are created on |
1755
1769
  | `available` | Newest supported version on the registry that is not downloaded and is newer than `newestLocal`, or null |
1756
1770
  | `downloading` / `downloadingVersion` | A download running, and its version |
1771
+ | `downloadPercent` | Fraction of layers finished, from `docker pull`'s own per-layer lines (`getDownloadProgress()`), 0-99. Null until it can be computed: while the version is resolved, while docker is still listing the layers, on a pull where every layer is already present, and on the local-build fallback |
1772
+ | `downloadMessage` | The latest line of docker's output, verbatim (the banner shows it under the percent) |
1773
+ | `canceled` | Whether the last finished download ended as cancelled — derived from `lastResult`, so a cancel that arrived after the image was already tagged does not make a successful download read as cancelled |
1774
+ | `serverTime` | The server's clock (epoch ms) when it answered; clients age `lastResult.at` with it instead of their own clock |
1775
+ | `lastResult` | The last download's outcome (`getLastDownloadResult()`, `lib/buildAgentImage.ts`), or null before any download has run this process. `result` is `succeeded` / `failed` / `canceled`; `source` (`present` / `registry` / `build`) only on success. Kept regardless of whether a download is running now, so a client that missed the live transition — a page reload, a return from another screen — still sees it via `at` (epoch ms); overwritten by the next download, never cleared on its own. Same idea as `finished` in `lib/cold-backup.ts` |
1757
1776
  | `registryReachable` | The tag list could be read |
1758
1777
 
1759
1778
  ### `POST /api/agents/download-image`
@@ -1770,19 +1789,42 @@ newest supported version the registry publishes.
1770
1789
  { "started": true, "version": "2026.9.3" }
1771
1790
  ```
1772
1791
 
1773
- **Errors:** `400` unsupported version; `409` a download is already in progress.
1792
+ **Errors:** `400` unsupported version; `409` a download is already in progress; `503`
1793
+ Docker is not available (refused here, before `202`, so the client gets the reason).
1774
1794
 
1775
1795
  The download is `lib/buildAgentImage.ts` → `downloadAgentImage({ version, onEvent, signal })`:
1776
1796
  - nothing to do when `openclaw-agent-base:<version>` is already here
1777
1797
  - `docker pull <registry>:<version>`, tag `openclaw-agent-base:<version>`, untag the
1778
- registry reference (`pullAgentImage()` in `lib/agent-images.ts`)
1798
+ registry reference — also when the tag step fails or is cancelled, so no
1799
+ `<registry>:<version>` reference is left behind (`pullAgentImage()` in `lib/agent-images.ts`)
1779
1800
  - when the pull fails and the repository Dockerfile's `ARG OPENCLAW_VERSION` is that
1780
1801
  version, `docker build --build-arg OPENCLAW_VERSION=<version>` instead; any other
1781
1802
  version fails
1782
1803
  - one download at a time (in-process lock, `getIsDownloading()` /
1783
- `getDownloadingVersion()`); output goes to `/tmp/rev4a-download-<timestamp>.log`
1804
+ `getDownloadingVersion()`); docker's output also goes to
1805
+ `/tmp/rev4a-download-<timestamp>.log` (mode 0600; logs older than 24 h are removed when a download ends) for diagnosis
1806
+ - progress: a layer counts as finished at `Pull complete` or `Already exists` (not at
1807
+ `Download complete`, which still has the extraction ahead); the percent is withheld
1808
+ until the first layer starts transferring — before that the layer list may be
1809
+ incomplete — never goes down, and stays below 100: completion is reported by
1810
+ `lastResult`, not by the percent
1811
+
1812
+ ### `DELETE /api/agents/download-image`
1813
+ Cancels the running download, if any (`cancelDownload()`): the `docker pull`/`docker build`
1814
+ process is killed, or never started when the cancel arrives before it (`runDocker` refuses
1815
+ an already-aborted signal, and the download checks after every wait). The rejected
1816
+ `downloadAgentImage()` promise carries a distinct `'canceled'` message, so the route's
1817
+ background `.catch` does not log it as a failure; `lastResult.result` becomes `canceled`.
1784
1818
 
1785
- ### `GET /api/agents-active`
1819
+ **Auth:** browser cookie or bearer token
1820
+
1821
+ **Response (202):**
1822
+ ```json
1823
+ { "canceling": true }
1824
+ ```
1825
+
1826
+ **Errors:** `404` nothing is running — the download has not started yet, or has
1827
+ already ended.
1786
1828
 
1787
1829
  ### `GET /api/agents-active`
1788
1830
  Configured agents (from `openclaw.json`) enriched with recent session activity
@@ -1864,42 +1906,83 @@ Register a parent→child relationship between sessions.
1864
1906
  ## Metrics
1865
1907
 
1866
1908
  ### `GET /api/metrics`
1867
- System metrics: latest snapshot + 24 h history + 24 h aggregates.
1909
+ Machine metrics of the host Rev4a runs on — CPU, RAM, swap, storage — read from
1910
+ `metrics.db` (written by the daemon every 30 s, see `docs/dev/DATABASE.md`): the latest
1911
+ sample, live host facts, and a bucketed history for one range. The System page
1912
+ (`/system`) and the dashboard's Machine card use it.
1868
1913
 
1869
1914
  **Auth:** Bearer or browser cookie
1870
1915
 
1916
+ **Query:** `range` — `1h` (default), `24h`, `7d` or `30d`; anything else answers `400`.
1917
+ The range is cut into buckets of `bucket_s` seconds (about 120 points, never shorter than
1918
+ 60 s — two samples, so a timer that drifts a second leaves no false gaps), each with its
1919
+ average and peak: 60 s for `1h`, 720 s for `24h`, 5040 s for `7d`, 21 600 s for `30d`.
1920
+
1871
1921
  **Response:**
1872
1922
  ```json
1873
1923
  {
1874
- "latest": { "ts": 1749201000000, "cpu_percent": 12.5, "ram_used_mb": 1820, "ram_total_mb": 4096, "disk_used_gb": 38.2, "disk_total_gb": 100.0, "load_avg_1m": 0.42 },
1875
- "history": [ ...up to 288 rows (24h at 5min intervals)... ],
1876
- "stats_24h": { "avg_cpu": 14.2, "max_cpu": 68.0, "avg_ram_mb": 1750 }
1924
+ "latest": {
1925
+ "ts": 1790419256, "age_s": 12, "cpu": 21.5,
1926
+ "ram_used_mb": 4779, "ram_total_mb": 7851, "swap_used_mb": 2048, "swap_total_mb": 4096,
1927
+ "load_avg_1m": 1.35, "load_avg_5m": 1.56, "load_avg_15m": 1.21,
1928
+ "disks": [
1929
+ { "mount": "/", "roles": ["root", "docker", "data"], "used_mb": 61440, "avail_mb": 38912, "total_mb": 102400, "percent": 61 }
1930
+ ]
1931
+ },
1932
+ "host": { "hostname": "vps", "platform": "linux", "cores": 4, "uptime_s": 290000 },
1933
+ "interval_s": 30,
1934
+ "range": "1h",
1935
+ "bucket_s": 60,
1936
+ "history": [ { "ts": 1790419230, "cpu_avg": 21.5, "cpu_max": 29.1, "ram_avg": 60.9, "ram_max": 61.2 } ],
1937
+ "disk_history": [ { "mount": "/", "points": [ { "ts": 1790419230, "percent": 61 } ] } ]
1877
1938
  }
1878
1939
  ```
1879
1940
 
1941
+ - `ts` values are in **seconds**; a `history` point's `ts` is the start of its bucket.
1942
+ Buckets with no samples are absent (the daemon was not running), not zero.
1943
+ - `ram_used_mb` is total − available, the figure `free` shows (Linux `/proc/meminfo`); on
1944
+ macOS active + wired + compressed pages. `swap_*` are `null` where there is no swap
1945
+ figure. `ram_avg`/`ram_max` are percentages.
1946
+ - `disks`: each filesystem once — `/`, Docker's data root (`docker info`), and Rev4a's
1947
+ data directory, merged when they are the same device; `roles` says which of the three
1948
+ it is. `percent` is used / (used + available), as `df` computes it.
1949
+ - `latest` is `null` before the daemon's first sample (the database is created when the
1950
+ daemon starts, the first sample about 5 s later), and when `metrics.db` has no tables. `age_s` older than 3 × `interval_s`
1951
+ means the daemon is not sampling; the page says so.
1952
+ - `host` is read live by the API process, which runs on the same host.
1953
+
1880
1954
  ---
1881
1955
 
1882
1956
  ## System Health
1883
1957
 
1884
1958
  ### `GET /api/system-health`
1885
- Aggregated health checks with recommendations.
1959
+ Aggregated checks of the runtime, cron jobs and usage cost, with recommendations. It
1960
+ does not report CPU, RAM or disk (those are in `GET /api/metrics`). Its `runtime.heartbeat`
1961
+ check ("Daemon heartbeat") reads the age of the newest machine metrics sample in
1962
+ `metrics.db`: it shows the daemon is alive, not that the OpenClaw session poll ingests —
1963
+ the two run on separate timers. A `metrics.db` that cannot be read makes this one check
1964
+ say so; it does not fail the others.
1886
1965
 
1887
- **Auth:** browser cookie
1966
+ **Auth:** browser cookie or bearer token
1888
1967
 
1889
1968
  **Response:**
1890
1969
  ```json
1891
1970
  {
1892
1971
  "health": "ok",
1893
1972
  "checks": [
1894
- { "name": "daemon", "status": "ok", "detail": "last poll 18s ago" },
1895
- { "name": "disk", "status": "warn", "detail": "82% used" }
1973
+ { "id": "runtime.sessions", "label": "Runtime sessions", "health": "ok", "value": 1, "details": "last session activity 3m ago", "source": "runtime" }
1974
+ ],
1975
+ "recommendations": [
1976
+ { "id": "cost.review-usage-based", "severity": "warning", "source": "cost", "title": "High usage-based cost", "details": "…", "actionHref": "/gateway", "dismissible": true, "createdAt": "2026-09-26T10:00:00.000Z" }
1896
1977
  ],
1897
- "recommendations": [ "Consider archiving old sessions to reduce disk usage." ],
1898
- "generatedAt": 1749201500000
1978
+ "generatedAt": "2026-09-26T10:00:00.000Z"
1899
1979
  }
1900
1980
  ```
1901
1981
 
1902
- `health` values: `"ok"` / `"warn"` / `"error"`
1982
+ Checks: `runtime.sessions`, `runtime.heartbeat`, `runtime.errors24h`,
1983
+ `cost.usageBasedToday`, `cron.jobs`. `health` (overall and per check) is `"ok"` /
1984
+ `"warning"` / `"error"`; a recommendation's `severity` is `"info"` / `"warning"` /
1985
+ `"critical"`. `generatedAt` and `createdAt` are ISO strings.
1903
1986
 
1904
1987
  ---
1905
1988
 
@@ -2619,7 +2702,7 @@ the process that spawned it, so that file is the only record of what happened.
2619
2702
 
2620
2703
  **Response:** `{ "success": true, "message": "...", "log": "update.log" }`
2621
2704
 
2622
- The banner records what it asked for, polls `/api/version` until the installed version
2705
+ The banner records what it asked for, polls `/api/update-check` until the installed version
2623
2706
  moves (or two minutes pass), reloads, and on the next mount either confirms
2624
2707
  "Updated to vX" or reports that the update did not complete and points at the log.
2625
2708
 
@@ -1,14 +1,17 @@
1
1
  # Rev4a — Database Reference
2
2
 
3
- > **Last updated:** 2026-08-28
3
+ > **Last updated:** 2026-09-26
4
4
 
5
- Rev4a uses a single SQLite file (`events.db`) in WAL mode.
5
+ Rev4a keeps three SQLite files in WAL mode, by default all in the data directory's `data/`:
6
+ `events.db` (sessions and events, below), `metrics.db` (machine metrics, see
7
+ [Metrics Database](#metrics-database-metricsdb)) and `credentials.db` (see
8
+ [Credentials Database](#credentials-database-credentialsdb)).
6
9
 
7
10
  ## Location
8
11
 
9
12
  | Environment | Path |
10
13
  |---|---|
11
- | VPS (default) | `/path/to/rev4a/data/events.db` |
14
+ | Default | `~/.config/rev4a/data/events.db` (`$REV4A_DATA_DIR/data/events.db`) |
12
15
  | Custom | Set `REV4A_DB` env var |
13
16
 
14
17
  ## WAL Mode
@@ -19,8 +22,9 @@ PRAGMA synchronous = NORMAL;
19
22
  ```
20
23
 
21
24
  - The daemon writes; the Next.js API routes open read-only connections
22
- - `PRAGMA wal_checkpoint(PASSIVE)` runs after each daemon poll cycle
23
- - WAL files: `events.db-shm`, `events.db-wal` — do not delete while daemon is running
25
+ - `events.db`: `PRAGMA wal_checkpoint(PASSIVE)` after each daemon poll cycle, `FULL` every 10
26
+ - `metrics.db`: SQLite's automatic checkpoint (every 1000 pages) — no explicit one
27
+ - WAL files (`*.db-shm`, `*.db-wal`) — do not delete while the daemon is running
24
28
 
25
29
  ## Tables
26
30
 
@@ -88,24 +92,9 @@ CREATE TABLE cost_override (
88
92
 
89
93
  When a `cost_override` row exists for the current month, the UI displays the override value instead of the computed sum.
90
94
 
91
- ### `system_metrics` — infrastructure metrics
92
-
93
- ```sql
94
- CREATE TABLE system_metrics (
95
- id INTEGER PRIMARY KEY AUTOINCREMENT,
96
- ts INTEGER NOT NULL,
97
- cpu_percent REAL,
98
- ram_used_mb INTEGER,
99
- ram_total_mb INTEGER,
100
- disk_used_gb REAL,
101
- disk_total_gb REAL,
102
- load_avg_1m REAL
103
- );
104
-
105
- CREATE INDEX idx_metrics_ts ON system_metrics(ts);
106
- ```
107
-
108
- Rows older than 24 h are pruned automatically by the daemon.
95
+ The daemon also writes a `system_anomaly` event (session `system`) when the machine
96
+ crosses a threshold — see [Metrics Database](#metrics-database-metricsdb). No feature
97
+ consumes those events yet; they only pass through the generic event feeds.
109
98
 
110
99
  ### `agent_upgrades` — agent updates and rollbacks
111
100
 
@@ -235,14 +224,76 @@ SELECT * FROM tree;
235
224
  ## Backup
236
225
 
237
226
  ```bash
238
- # Manual snapshot
239
- sqlite3 events.db ".backup events.db.bak-$(date +%s)"
227
+ # Everything Rev4a owns: the whole data directory (.env, data/, shared/)
228
+ bash scripts/backup.sh # -> ~/rev4a-backups/rev4a-backup-<UTC time>.tar.gz
229
+ bash scripts/restore.sh ~/rev4a-backups/rev4a-backup-<UTC time>.tar.gz
230
+
231
+ # One database only, consistent while Rev4a runs (needs the sqlite3 CLI)
232
+ sqlite3 ~/.config/rev4a/data/events.db ".backup events-$(date +%s).db"
233
+ ```
234
+
235
+ `backup.sh` archives the Rev4a data directory (`$REV4A_DATA_DIR`, default
236
+ `~/.config/rev4a`) whole, so a new file there is included without changing the script.
237
+ The archive (mode 0600, directory 0700) holds every secret in plain text; keep a copy
238
+ off the machine. `restore.sh` refuses an archive with links in it and a running
239
+ `rev4a serve`, archives the current state into `~/rev4a-backups` first, stops
240
+ `rev4a.service` when it runs, and replaces the data directory as a whole (a leftover
241
+ SQLite `-wal` cannot mix into the restored state), then resets the permissions. Agent volumes are not included —
242
+ the dashboard's cold backups cover them.
243
+
244
+ ## Metrics Database (`metrics.db`)
245
+
246
+ Machine metrics of the host — CPU, RAM, swap, storage — kept apart from `events.db`, so
247
+ the event log holds sessions and events only and the history can be dropped on its own.
248
+
249
+ **Location:** `data/metrics.db`, next to `events.db` (it follows `REV4A_DB`'s directory; both
250
+ the daemon and the API resolve the data directory the same way, `REV4A_DATA_DIR` included).
251
+ **Created by** the daemon when it starts (`CREATE TABLE IF NOT EXISTS`), which is also its
252
+ only writer; `GET /api/metrics` and `GET /api/system-health` read it
253
+ (`lib/metrics-db.ts`). Until the daemon has run once the file does not exist, and the
254
+ API answers with `latest: null`. If the daemon cannot open it (corrupt, unwritable), it
255
+ logs `[METRICS] disabled` and keeps polling sessions without sampling.
256
+ **Sampling:** the first sample about 5 s after the daemon starts, then every 30 s, on a timer
257
+ of its own — independent of the OpenClaw session poll.
258
+ **Retention:** 30 days, pruned at every sample: ~86 400 `system_metrics` rows, and one
259
+ `system_disks` row per filesystem per sample (up to ~259 200 for three).
260
+
261
+ ```sql
262
+ CREATE TABLE system_metrics (
263
+ id INTEGER PRIMARY KEY AUTOINCREMENT,
264
+ ts INTEGER NOT NULL, -- Unix seconds
265
+ cpu_percent REAL, -- all cores, busy share since the previous sample
266
+ ram_used_mb INTEGER, -- total − available (Linux), as `free` shows it
267
+ ram_total_mb INTEGER,
268
+ swap_used_mb INTEGER, -- NULL where there is no swap figure
269
+ swap_total_mb INTEGER,
270
+ load_avg_1m REAL,
271
+ load_avg_5m REAL,
272
+ load_avg_15m REAL
273
+ );
274
+ CREATE INDEX idx_metrics_ts ON system_metrics(ts);
240
275
 
241
- # Automated (via backup.sh in repo root)
242
- bash /data/.openclaw/workspace-ops/rev4a-next-ts/backup.sh
276
+ -- One row per filesystem per sample: `/`, Docker's data root and Rev4a's data directory,
277
+ -- each device once; roles lists which of them it is (e.g. 'root,docker,data').
278
+ CREATE TABLE system_disks (
279
+ ts INTEGER NOT NULL, -- the same ts as the system_metrics row
280
+ mount TEXT NOT NULL,
281
+ roles TEXT,
282
+ used_mb INTEGER, -- blocks − free
283
+ avail_mb INTEGER, -- what a normal user can still write
284
+ total_mb INTEGER
285
+ );
286
+ CREATE INDEX idx_disks_ts ON system_disks(ts);
287
+ CREATE INDEX idx_disks_mount_ts ON system_disks(mount, ts);
243
288
  ```
244
289
 
245
- Backups are stored in `backups/` inside the repo directory.
290
+ **Anomalies** go to `events.db` as `system_anomaly` events: CPU above 85 % or RAM above
291
+ 90 % on two consecutive samples (at most once per 5 minutes per metric), and a filesystem
292
+ at 90 % or more once when it reaches the threshold, with no cooldown (again only after it
293
+ has dropped back under, or once after each daemon start: the state is kept in memory).
294
+
295
+ Installs from before this database may still have a `system_metrics` table in
296
+ `events.db`: nothing reads or writes it any more, and it can be dropped.
246
297
 
247
298
  ## Credentials Database (`credentials.db`)
248
299
 
@@ -597,11 +597,11 @@ POST /api/sessions/maintenance/cleanup
597
597
 
598
598
  ## 9. References
599
599
 
600
- - [Session maintenance](/concepts/session) — OpenClaw docs
601
- - [Session management deep dive](/reference/session-management-compaction) — store schema + maintenance rules
602
- - [Compaction](/concepts/compaction) — auto + manual compaction
603
- - [Session pruning](/concepts/session-pruning) — tool-result trimming
604
- - [Transcript hygiene](/reference/transcript-hygiene) — provider-specific fixups
605
- - [Cron jobs](/automation/cron-jobs) — cron.sessionRetention + run log config
600
+ - [Session maintenance](https://docs.openclaw.ai/concepts/session) — OpenClaw docs
601
+ - [Session management deep dive](https://docs.openclaw.ai/reference/session-management-compaction) — store schema + maintenance rules
602
+ - [Compaction](https://docs.openclaw.ai/concepts/compaction) — auto + manual compaction
603
+ - [Session pruning](https://docs.openclaw.ai/concepts/session-pruning) — tool-result trimming
604
+ - [Transcript hygiene](https://docs.openclaw.ai/reference/transcript-hygiene) — provider-specific fixups
605
+ - [Cron jobs](https://docs.openclaw.ai/automation/cron-jobs) — cron.sessionRetention + run log config
606
606
  - [Rev4a API Reference](API-REFERENCE.md)
607
607
  - [Rev4a Frontend Architecture](../FRONTEND-ARCHITECTURE.md)
@@ -1,6 +1,6 @@
1
1
  # Data Freshness in Rev4a
2
2
 
3
- > **Last updated:** 2026-09-15
3
+ > **Last updated:** 2026-09-26
4
4
 
5
5
  Understanding how current the data in Rev4a is — and what is truly real-time vs. periodically updated.
6
6
 
@@ -14,9 +14,8 @@ The daemon polls `openclaw sessions --json --all-agents` on a fixed 30-second
14
14
  timer. Session statuses, token counts, and costs are therefore up to 30 seconds
15
15
  behind reality.
16
16
 
17
- The code declares a shorter 15-second interval for when a session is `working`,
18
- but nothing uses it — the timer is always 30 seconds. Do not tell a user the
19
- dashboard speeds up during active work.
17
+ The interval is fixed: it does not speed up while a session is `working`. Do not tell
18
+ a user the dashboard refreshes faster during active work.
20
19
 
21
20
  **What this means for PULSE:** if you ask "is Argus working right now?", the
22
21
  answer reflects data that is up to 30 seconds old.
@@ -29,26 +28,31 @@ answer reflects data that is up to 30 seconds old.
29
28
 
30
29
  The `/api/stream` endpoint pushes updates to the dashboard every 5 seconds. It
31
30
  sends exactly three things:
32
- - New events (spawn, complete, error, tool_call)
31
+ - New events (spawn, complete, fail, spawn_timeout, system_anomaly)
33
32
  - The session list
34
33
  - Today's cost total
35
34
 
36
35
  Lineage is **not** pushed over the stream; the Lineage page fetches it itself.
37
36
 
38
37
  The event feed is as close to real-time as Rev4a gets. However, events are
39
- generated by the daemon's poll cycle, so an event that just happened may take up
38
+ generated by the daemon every 30 seconds — session events by its poll of OpenClaw,
39
+ system anomalies by its machine sampling — so an event that just happened may take up
40
40
  to 30 seconds to appear, plus up to 5 more to reach the browser.
41
41
 
42
42
  ---
43
43
 
44
44
  ## System metrics (CPU, RAM, disk)
45
45
 
46
- **Update frequency:** every daemon poll cycle (30 seconds)
46
+ **Update frequency:** every 30 seconds, sampled by the Rev4a daemon on its own timer
47
+ (it does not depend on OpenClaw being installed on the host). The System page refreshes
48
+ every 30 seconds too, so a figure is at most about a minute old.
47
49
 
48
- **Retention:** 30 days (older rows are pruned automatically each cycle)
50
+ **Retention:** 30 days, in their own database (`metrics.db`).
49
51
 
50
- These metrics are collected and stored, and served by `GET /api/metrics`. They
51
- are **not** shown on the dashboard — see the System Health note below.
52
+ Shown on the **System** page (`/system`) — current CPU, memory, swap and storage, and
53
+ their history over 1 hour, 24 hours, 7 days or 30 days — and, as three figures, on the
54
+ Dashboard's Machine card. If the newest sample is older than 90 seconds the page says
55
+ the collector is not running instead of showing stale numbers as current.
52
56
 
53
57
  ---
54
58
 
@@ -96,11 +100,22 @@ at most every 10 minutes, so a newly published version can take that long to app
96
100
  The banner polls the endpoint every 2 seconds for as long as the page is open, not only
97
101
  while something is running.
98
102
 
99
- The operation the banner triggers is a **`docker pull`**, not a build. While it
100
- runs the banner shows an indeterminate animated bar: there is no percentage, no
101
- byte count, and no log output in the UI. A log is written server-side to
103
+ The operation the banner triggers is a **`docker pull`**, not a build (unless the
104
+ pull fails and this repository's own Dockerfile builds that exact version, a
105
+ development fallback). While it runs the banner shows the percent of layers
106
+ finished and the latest line of docker's output; the bar is indeterminate until a
107
+ percent can be computed (while the version is resolved and the layers are listed,
108
+ and during the local-build fallback). The percent never goes down and stops at 99:
109
+ completion is announced by the outcome, not by 100%. A **Cancel** button, offered
110
+ once the download is running, stops it — the `docker pull`/`docker build` process is
111
+ killed, or never started if the cancel came first. A log is written server-side to
102
112
  `/tmp/rev4a-download-<timestamp>.log`, but nothing in the dashboard reads it.
103
113
 
114
+ The outcome comes from the server's own record of the last download: "ready" and
115
+ "canceled" show for 3 seconds, a failure stays until the next action. An outcome
116
+ that happened while the Agents page was closed or reloading is still reported if it
117
+ is less than 30 seconds old.
118
+
104
119
  ---
105
120
 
106
121
  ## Memory / context files
@@ -127,8 +142,9 @@ The tracked set is `USER.md`, `MEMORY.md`, `AGENTS.md`, `SOUL.md`, `HEARTBEAT.md
127
142
  ## How to check if data is fresh
128
143
 
129
144
  System health is a card on the **Dashboard**, not a page of its own. Its checks
130
- cover session runtime, ingestion freshness, errors in the last 24 hours, today's
131
- usage-based cost, and cron jobs. It reports no CPU, RAM, or disk figures.
145
+ cover session runtime, the daemon heartbeat, errors in the last 24 hours, today's
146
+ usage-based cost, and cron jobs. CPU, RAM and disk are on the System page (`/system`)
147
+ and the Dashboard's Machine card instead.
132
148
 
133
149
  To check directly on the server:
134
150
  ```bash
@@ -1,6 +1,6 @@
1
1
  # Rev4a Glossary
2
2
 
3
- > **Last updated:** 2026-09-15
3
+ > **Last updated:** 2026-09-26
4
4
 
5
5
  Terms you'll encounter while using the Rev4a dashboard.
6
6
 
@@ -43,10 +43,10 @@ A third-party service token (GitHub PAT, Trello API key, Vercel token, Supabase
43
43
  A scheduled task that runs an agent automatically at a fixed time (e.g. every night at 3:15 AM). Configured with standard cron syntax. Jobs can be enabled/disabled per entry.
44
44
 
45
45
  ## Dashboard
46
- The main page of Rev4a (`/`). Shows live sessions, cost summary, health cards for Rev4a's runtime, cron and lineage, and a real-time event feed. It shows no CPU, RAM, disk or load average, and there is no setup banner: an incomplete wizard redirects to `/wizard`.
46
+ The main page of Rev4a (`/`). Shows live sessions, cost summary, health cards for Rev4a's runtime, cron and lineage, and a real-time event feed. Its Machine card shows the server's CPU, RAM and fullest disk right now, with a link to the System page. There is no setup banner: an incomplete wizard redirects to `/wizard`.
47
47
 
48
48
  ## Event
49
- A lifecycle occurrence for a session: spawned, completed, errored, or a tool call made. Events appear in the live feed on the Dashboard and are pushed via SSE every 5 seconds.
49
+ A lifecycle occurrence for a session: spawned, completed, failed, or missing too long (spawn timeout); the daemon also records system anomalies (high CPU or RAM) as events. Events appear in the live feed on the Dashboard and are pushed via SSE every 5 seconds.
50
50
 
51
51
  ## Gateway
52
52
  The routing layer that connects agents to AI providers. The Gateway page manages provider API keys and the model catalogue, and pushes that catalogue to every agent. Which model a given agent runs is set elsewhere, in the Model section of that agent's detail panel.
@@ -130,7 +130,10 @@ Rev4a's store for third-party service tokens (GitHub, Trello, Vercel, Supabase,
130
130
  ## Workspace
131
131
  The directory where an agent's operational files live (config, memory files, skills, plugins, scripts). The Workspace page lets you browse, view, and edit files across the host workspace, which appears as `Local`, and all agent containers.
132
132
 
133
+ ## System (page)
134
+ The page with the machine Rev4a runs on: CPU (with cores and load), memory and swap, and each disk's used and free space, now and over 1 hour, 24 hours, 7 days or 30 days. A bar turns yellow and then red when it reaches a threshold (CPU and memory 85 % / 95 %, disk 80 % / 90 %), and the level is written next to it.
135
+
133
136
  ## System Health
134
- A card on the Dashboard reporting whether Rev4a itself is working, not the hardware it runs on. `/api/system-health` returns checks for session runtime, feed ingestion freshness, errors in the last 24 hours, today's usage-based cost, and cron jobs, each with a health of `ok`, `warning` or `error`, plus a list of recommendations.
137
+ A card on the Dashboard reporting whether Rev4a itself is working, not the hardware it runs on. `/api/system-health` returns checks for session runtime, the daemon heartbeat (how old its latest machine sample is), errors in the last 24 hours, today's usage-based cost, and cron jobs, each with a health of `ok`, `warning` or `error`, plus a list of recommendations.
135
138
 
136
- It reports no CPU, RAM, disk or load average. Those metrics are collected by the daemon and stored for 30 days, and are served by `/api/metrics`, but no page displays them.
139
+ It reports no CPU, RAM, disk or load average: those are on the System page (`/system`), sampled every 30 seconds by the daemon and kept for 30 days.
@@ -1,6 +1,6 @@
1
1
  # What is Rev4a?
2
2
 
3
- > **Last updated:** 2026-09-22
3
+ > **Last updated:** 2026-09-26
4
4
 
5
5
  Rev4a is the control panel for your AI agent infrastructure. It shows you everything your agents are doing, how much they cost, and whether the system is healthy — all in one dashboard.
6
6
 
@@ -9,7 +9,7 @@ Rev4a is the control panel for your AI agent infrastructure. It shows you everyt
9
9
  ### Dashboard (`/`)
10
10
  The main page. See live sessions (who's working right now), cost summary, an overall health indicator, and a real-time event feed. Click any session row to open the Session Drawer and see every tool call the agent made.
11
11
 
12
- The health cards cover Rev4a's own runtime, cron and watchdog, and lineage. They do **not** show CPU, RAM, disk, or load average: those metrics are collected and stored, but no page displays them. There is no wizard banner on the dashboard; an incomplete setup redirects to `/wizard` instead.
12
+ The health cards cover Rev4a's own runtime, cron and watchdog, and lineage. The Machine card shows the server's CPU, RAM and fullest disk right now, coloured when it reaches a threshold, and links to the System page for details. There is no wizard banner on the dashboard; an incomplete setup redirects to `/wizard` instead.
13
13
 
14
14
  ### First-run wizard (`/wizard`)
15
15
  First-run setup wizard that guides new users through configuration. Four steps:
@@ -58,10 +58,14 @@ went missing. Afterwards **Roll back to <version>** puts back the backup taken j
58
58
  the update, on the old version; anything the agent did after the update is lost. The agent
59
59
  must be running to start an update.
60
60
 
61
- That button runs a `docker pull` from the registry. It is a download, not a build.
62
- While it runs the banner shows an indeterminate animated bar with no percentage and
63
- no log output; there is no modal, no "View Progress", no "Run in Background", and
64
- no way to abort from the UI. Status is polled every 2 seconds.
61
+ The agent base image banner above (not this update — an agent's own **Update to
62
+ <version>** recreates it on an image already downloaded) is the one running the
63
+ `docker pull`. While it runs the banner shows the percent of layers finished and the
64
+ latest line of docker's output; a **Cancel** button, offered once the download is
65
+ running, stops it. There is no modal, no "View Progress", no "Run in Background". The
66
+ outcome ("ready" and "canceled" for 3 seconds, a failure until the next action) is
67
+ reported even when it happened while the Agents page was closed or reloading, if it is
68
+ less than 30 seconds old. Status is polled every 2 seconds.
65
69
 
66
70
  ### Create Agent (`/agents/create`)
67
71
  Wizard to spin up a new agent. Steps: choose a template (Prometheus, Argus, Atlas, etc.), name your agent, pick a model, set an optional port range (default: auto-assigned 10-port block, e.g. 3700-3709). The agent is created as a Docker container with a persistent volume — all config, workspace files, and credentials survive container restarts.
@@ -70,13 +74,16 @@ Wizard to spin up a new agent. Steps: choose a template (Prometheus, Argus, Atla
70
74
  Interactive graph showing session family trees — which agent spawned which child agent, across configurable time periods: 1d, 3d, 7d, 15d, 30d, and all. Click any node to inspect. Includes a live feed side panel.
71
75
 
72
76
  ### Gateway (`/gateway`)
73
- Connects agents to AI providers. Add and remove provider API keys, enable or disable individual models in the catalogue read from `models.config.json`, and push the resulting model list to every agent container. It does not assign models to agents: choosing which model an agent runs is done per agent, in the Model section of the agent's detail panel on the Agents page.
77
+ Connects agents to AI providers. Add and remove provider API keys, enable or disable individual models in the catalogue read from `models.config.json`, and push the resulting model list to every agent container. It does not assign models to agents: choosing which model an agent runs is done per agent, in the Model section of the agent's detail panel on the Agents page. Every model row has a **Details** button: the modal shows the model's description, price and provider, its architecture and parameter count when the weights are public, reasoning modes, capabilities, benchmarks with sources, licence and how much disk the weights take — the numbers come from the model card on Hugging Face and from curated entries that cite a source, never from estimates.
74
78
 
75
79
  ### Containers (`/containers`)
76
80
  Full list of all Docker containers on the server, including stopped ones. Each row shows name, status, image, IP, ports, and an agent pill where applicable. Click "Terminal" to open an interactive shell into that container.
77
81
 
78
82
  The page shows no CPU or memory figures and has no start, stop, or restart buttons: it is read-only apart from the terminal link. Container lifecycle is managed from the Agents page.
79
83
 
84
+ ### System (`/system`)
85
+ The machine Rev4a runs on. Three cards — CPU (percent, cores, load average), memory (used, available, swap) and storage (each disk with its used and free space, and whether it is the system disk, where Docker keeps agents, or where Rev4a keeps its data) — then charts of CPU, memory and each disk over 1 hour, 24 hours, 7 days or 30 days (for CPU and memory the average line with the peak shaded). Hover a chart (or focus it and use the arrow keys) to read a point; each chart has a Table view. Bars turn yellow and red when they reach their thresholds. Figures refresh every 30 seconds; history is kept 30 days.
86
+
80
87
  ### Container Terminal (`/containers/terminal/[id]`)
81
88
  Live web terminal into a Docker container. Run commands, inspect files, debug issues — like SSH but in the browser.
82
89
 
@@ -1,6 +1,6 @@
1
1
  # What PULSE Can Answer
2
2
 
3
- > **Last updated:** 2026-09-22
3
+ > **Last updated:** 2026-09-26
4
4
 
5
5
  PULSE is the in-dashboard AI concierge for Rev4a. This document defines what she can and cannot answer.
6
6
 
@@ -10,6 +10,7 @@ PULSE is the in-dashboard AI concierge for Rev4a. This document defines what she
10
10
 
11
11
  - "Where do I find the agents page?"
12
12
  - "How do I get to the container list?"
13
+ - "Where do I see CPU, RAM or disk usage of the server?" — The System page (`/system`), and the Machine card on the Dashboard.
13
14
  - "Where can I see the first-run wizard?"
14
15
  - "Is there a page for cron jobs?"
15
16
  - "Where can I see my AI providers?" — The Gateway page (`/gateway`) shows providers and the model catalogue. Which model an agent runs is on the Agents page, in that agent's detail panel.
@@ -95,13 +96,13 @@ PULSE is the in-dashboard AI concierge for Rev4a. This document defines what she
95
96
  - "Why are costs missing for some sessions?"
96
97
  - "Why can't I connect to the container terminal?"
97
98
  - "Why is the gateway sync not working?"
98
- - "Why am I seeing a setup banner?"
99
99
  - "How do I check if the daemon is running?"
100
100
  - "Why are credentials not showing in my agent?"
101
101
  - "What does the yellow diamond next to an agent mean?"
102
102
  - "Why is the agent base image banner showing?"
103
103
  - "How do I fix 'image is outdated'?"
104
- - "How do agent image updates work?" — The banner on the Agents page runs a `docker pull`. It is a download, not a build: there is no modal, no live log, and no way to abort from the UI.
104
+ - "How do agent image updates work?" — The banner on the Agents page runs a `docker pull` and shows the percent of layers finished and docker's latest output line, with a Cancel button once the download is running; there is no modal or live log. The outcome is reported even if it happened while you were away, as long as it is less than 30 seconds old.
105
+ - "How big is model X / how many parameters does it have?" — Open the **Details** modal from the model row on the Gateway page: parameter counts, architecture and benchmark numbers come from the model's card on Hugging Face, with the vendor's own claim where it makes one. Models whose weights are not published (OpenAI, Anthropic, Google, Qwen's Plus/Max line) show no parameter count: the size is not public and Rev4a does not estimate it.
105
106
  - "How do I update Rev4a itself?" — The "Update available" banner starts `rev4a update` in the background (the same command as from a shell). It installs the new version and Rev4a restarts itself, so the dashboard drops for a short while; the banner waits for the new version and then confirms it, or says the update did not complete. The update's output is written to `update.log` in the Rev4a data directory (`~/.config/rev4a/data/update.log`) — the only place that says why an update failed.
106
107
  - "Why is my Telegram bot not connecting?"
107
108
  - "Why can't I approve a pairing code?"