@flame0510/project-aether 1.7.0 → 1.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -72,7 +72,7 @@
72
72
 
73
73
  **Credentials vault:** The `/credentials` page stores third-party service tokens in `credentials.db` and installs them into containers on explicit sync. Provider API keys live separately in `provider-keys.json`, managed from the Gateway UI.
74
74
 
75
- **Container management UI:** The `/containers` page lists all running Docker containers with resource usage and quick links, and opens a web terminal into any of them. Fully functional.
75
+ **Container management UI:** The `/containers` page lists all Docker containers, running or stopped, with status, image, IP and ports, and opens a web terminal into any of them. Fully functional. (Machine CPU, RAM and disk are on `/system`.)
76
76
 
77
77
  **Workspace API:** Lazy-loaded file tree explorer with real-time reads (no caching), supports both host and container workspaces via `docker exec`.
78
78
 
@@ -81,7 +81,7 @@
81
81
  | Feature | Status | Notes |
82
82
  |---|---|---|
83
83
  | Credentials (UI + API) | 🟢 Implemented | SQLite, no encryption yet |
84
- | Container page | 🟢 Functional | Lists all containers, resources, web terminal |
84
+ | Container page | 🟢 Functional | Lists all containers, web terminal |
85
85
  | Agent creation wizard | 🟢 Implemented | One-click create with model/template selection |
86
86
  | Provider proxy | 🟢 Implemented | Agents route through Rev4a provider gateway |
87
87
  | Shared volumes | 🟢 Implemented | Skills + repos mounted on all agents |
@@ -575,15 +575,16 @@ The Rev4a daemon (`daemon.js`) is a standalone Node.js process that bridges the
575
575
  1. Poll `openclaw sessions --json --all-agents` every 30 s
576
576
  2. Upsert session rows into `events.db`
577
577
  3. Emit `spawn` / `complete` / `fail` / `spawn_timeout` events into the `events` table
578
- 4. Collect system metrics (CPU, RAM, disk) after every completed poll
579
- 5. Detect anomalies (CPU > 85%, RAM > 90%) and record them
580
- 6. Manage DB lifecycle (WAL mode, checkpoint after each cycle)
578
+ 4. Sample the host machine (CPU, RAM, swap, storage) every 30 s into `metrics.db`, on a
579
+ timer of its own
580
+ 5. Detect anomalies (CPU > 85%, RAM > 90%, a disk > 90%) and record them as events
581
+ 6. Manage DB lifecycle (WAL mode; events.db checkpointed after each poll cycle)
581
582
 
582
583
  ### Poll Interval
583
584
 
584
585
  A fixed 30 s timer, whether or not sessions are working. When the OpenClaw CLI is not on
585
- the host, the daemon logs it once and stays idle — and then records no system metrics
586
- either, since they are collected at the end of a poll.
586
+ the host, the session poll logs it once and stays idle; the machine metrics keep being
587
+ sampled, since they have their own timer.
587
588
 
588
589
  ### Cost Estimation
589
590
 
@@ -625,29 +626,41 @@ The daemon uses `INSERT … ON CONFLICT DO UPDATE` with these rules:
625
626
 
626
627
  ### System Metrics
627
628
 
628
- After every completed poll the daemon records a `system_metrics` row. Data sources:
629
+ Every 30 s (`METRICS_INTERVAL_MS`) the daemon samples the machine it runs on and writes a
630
+ `system_metrics` row plus one `system_disks` row per filesystem into `metrics.db` — its
631
+ own database next to `events.db` (schema in `docs/dev/DATABASE.md`). Machine-wide sources:
632
+
633
+ - **CPU**: busy share of all cores since the previous sample, from host ticks
634
+ (`os.cpus()`, i.e. `/proc/stat` on Linux) — the whole machine, not Rev4a's cgroup
635
+ - **RAM / swap**: `/proc/meminfo` on Linux, used = `MemTotal − MemAvailable` (what `free`
636
+ shows); on macOS (development) active + wired + compressed pages from `vm_stat` and
637
+ `vm.swapusage`; elsewhere `os.totalmem()` / `os.freemem()` and no swap
638
+ - **Storage**: `fs.statfsSync()` on `/`, Docker's data root (`docker info`, asked once and
639
+ retried every 10 min while Docker is down) and Rev4a's data directory, each device once;
640
+ a path that cannot be read (Docker Desktop's root inside its VM) is skipped
641
+ - **Load**: `os.loadavg()` 1 / 5 / 15 min
642
+
643
+ The first sample comes about 5 s after start (the CPU figure needs a delta from the baseline
644
+ taken at load), then one every 30 s. Samples older than 30 days are pruned at every sample.
645
+ `metrics.db` is opened in a try: if it cannot be (corrupt, unwritable), sampling is off —
646
+ `[METRICS] disabled` in the log — and the session poll keeps running. The API side reads
647
+ it through `lib/metrics-db.ts` (read-only, `null` when there is nothing to read);
648
+ `GET /api/metrics` serves the System page (`/system`) and the dashboard's Machine card.
629
649
 
630
- - **CPU**: cgroup v2 usage delta (`/sys/fs/cgroup/cpu.stat`) when available, falls back to host ticks from `os.cpus()`
631
- - **RAM**: `os.totalmem()` / `os.freemem()`
632
- - **Disk**: `fs.statfsSync()` on the directory holding the events database — the
633
- filesystem Rev4a's own data lives on; works the same on Linux and macOS
634
- - **Load**: `os.loadavg()[0]`
650
+ ### Anomaly Detection
635
651
 
636
- Metrics older than 30 days are pruned automatically each cycle.
652
+ After each sample the daemon inserts a `system_anomaly` event into `events.db`
653
+ (`{ metric, values, threshold, message }`) and logs `[ANOMALY] <message>` when:
637
654
 
638
- ### Anomaly Detection
655
+ | Metric | Condition | Repeats |
656
+ |---|---|---|
657
+ | CPU | > 85% on two consecutive samples | at most once per 5 min |
658
+ | RAM | > 90% on two consecutive samples | at most once per 5 min |
659
+ | Disk (each filesystem) | ≥ 90% | once when it reaches it, no cooldown — again only after it drops back under, or after the daemon restarts |
639
660
 
640
- The daemon compares the last two metric samples. When both exceed a threshold and the
641
- cooldown (5 min per metric) has passed, it inserts a `system_anomaly` event
642
- (`{ metric, values, threshold, message }`) and logs `[ANOMALY] CPU high: N%` or
643
- `[ANOMALY] RAM high: N%`. No feature consumes `system_anomaly` events yet; they only pass through the generic
661
+ No feature consumes `system_anomaly` events yet; they only pass through the generic
644
662
  event feeds.
645
663
 
646
- | Metric | Threshold (two consecutive samples) |
647
- |---|---|
648
- | CPU | > 85% |
649
- | RAM | > 90% |
650
-
651
664
  ### DB Safety
652
665
 
653
666
  - WAL mode + `PRAGMA synchronous = NORMAL` for concurrent read safety
@@ -82,6 +82,7 @@ then use it — a one-off copy of markup is the thing this section exists to pre
82
82
  5. **CSS vars over hardcoded colors.** Never inline `#22c55e`, `#ef4444`, `#888` — use the CSS token.
83
83
  6. **Loading feedback.** Use `loading` prop on `Button` (shows inline spinner). Do not render separate loader elements.
84
84
  7. **Minimal animation.** Transitions on hover/active (0.1s–0.12s), spin on loading spinner. No bounce, shake, or decorative animations.
85
+ 8. **Data display: violet for the data, status colours for levels.** A chart series is `var(--violet)` (2px line, 10% wash), grid and axis text are `var(--border)` / `var(--text-dim)`, never the series colour. A `Meter` fill is violet when fine and `var(--yellow)` / `var(--red)` at its warning / critical level, with the level also written out — status never rests on colour alone. Focus on a chart is a 1px `outline` in `var(--violet-dim)`, not a shadow. See `TimeSeriesChart` and `Meter` in `docs/FRONTEND-ARCHITECTURE.md`.
85
86
 
86
87
  ### React hook
87
88
 
@@ -38,6 +38,8 @@ All shared UI primitives live in `app/components/ui/` and are exported from `app
38
38
  | `ModalityIcons` | `ModalityIcons.tsx` | Capability badges parsed from a model's `modality` string (`"text+image->text"`). Input types render as icons; a non-text **output** is called out separately, since reading an image and generating one are different capabilities. Pass `dynamic` for router models, which advertise the union of everything they might route to and so show "varies" instead. |
39
39
  | `Metric` | `Metric.tsx` | Metric card with title, value, subtitle, tone, and a `size` (`md` default, `sm` for a grid inside a dialog: smaller value, tighter spacing). |
40
40
  | `StatusCard` | `StatusCard.tsx` | Health/status report card. |
41
+ | `Meter` / `meterLevel` | `Meter.tsx` | Horizontal 0-100 % fill whose colour carries the level (`ok` accent, `warning` yellow, `critical` red); the level is also written out beside it, never colour alone. `meterLevel(percent, warning, critical)` picks the level. ARIA `role="meter"`. |
42
+ | `TimeSeriesChart` | `TimeSeriesChart.tsx` | One series over time in plain SVG measured to its container (text never scaled): a 2px accent line for the value, a 10 % wash up to the bucket's peak, a hairline grid at 0 / 50 / 100 %, the latest value labelled at the line's end. Points further apart than 1.5 buckets are not joined (a gap means no samples). Hover or keyboard focus (← → Esc) shows a crosshair snapped to the nearest point with value and peak; a **Table** disclosure holds the same data. `stale` dims the previous render while a new range loads. Used by the System page. |
41
43
  | `Surface` | `Surface.tsx` | Shared panel/card surface, variant prop. |
42
44
  | `Page` / `PageHeader` | `Page.tsx` | Full-page layout shell. |
43
45
  | `Icons` | `Icons.tsx` | SVG icons (`EyeIcon`, `EyeOffIcon`), 16/20px shared. |
@@ -45,6 +47,8 @@ All shared UI primitives live in `app/components/ui/` and are exported from `app
45
47
  | `UpdateSection` | `app/agents/UpdateSection.tsx` | OPENCLAW VERSION section of the agent detail panel: the version the agent runs, **Update to <version>** when a newer supported version is downloaded (confirm modal, disabled while the agent is stopped), the running update's steps with backup progress, the outcome, and **Roll back to <version>** after an update. Polls `/api/agents/[id]/update` every 2 s while an update or rollback runs; one action at a time (keyed busy state). |
46
48
  | `BackupSection` | inline in `app/agents/PageClient.tsx` | BACKUP section of the agent detail panel, on the cold backup and the restore. **Backup Now** starts `POST /api/agents/[id]/cold-backup`; while the job runs a banner shows the file and its live percent with **Cancel** (`DELETE /cold-backup`). **Restore** (after a confirm) starts `POST /restore` and a banner shows `Restoring <file>…` (no percent: the extract is a single `tar xzf`, and there is no Cancel). Both sections poll their `GET` every 2 s while running, and on mount pick up a job that is already running — a backup lives in a Docker helper, a restore in `agent_restores`, so reloading the page or navigating away never loses them nor allows a second one (the server answers 409 anyway). Delete per row; all actions disabled while one runs; keyed busy state `{ kind, file }` so only the row in action shows the spinner. On the agent list, an activity Badge (fed by `/api/agents/activity-summary`, polled at 2 s only while something runs, otherwise riding the 15 s list poll) reads `BACKUP nn%`, `RESTORING`, `RECREATING`, `EDITING` or `UPDATING`. |
47
49
  | `RecreateSection` | inline in `app/agents/PageClient.tsx` | RECREATE section of the agent detail panel. **Recreate Container** starts `POST /api/agents/[id]/recreate` (202) after a confirm; a banner then shows the phase — *Backing up … nn%* while the cold backup runs, *Recreating container…* while the container is rebuilt and the gateway starts. The section polls `GET /recreate` every 2 s, and on mount picks up a recreate that is already running, so a reload or navigation never loses it; it refetches the agent once the job reports `done`. |
50
+ | System page | `app/system/PageClient.tsx` | `/system`: CPU, memory and storage of the host, from `GET /api/metrics`. Three cards (CPU with cores and load; memory with available and swap; one `Meter` per filesystem with its roles) — thresholds 85/95 % for CPU and memory, 80/90 % for storage — then a range switch (`Tabs`: 1h · 24h · 7d · 30d) scoping the charts below it: CPU and memory average with peak, one chart per filesystem. Polls every 30 s with an `AbortController` ref (a range change aborts the previous fetch). Times in the configured Rev4a timezone (`useRev4aTimezone`). Before the first answer the cards and charts are skeletons shaped like them (`app/system/SystemSkeleton.tsx`, also the route's `loading.tsx`), so nothing jumps; the hostname line is cut with an ellipsis and never widens the page. Says *No samples yet* before the daemon's first sample and *Not collecting* when the latest sample is older than three intervals. Two columns of charts from lg (992px) up, one below. |
51
+ | Machine card | `app/components/SystemCockpit.tsx` | On the dashboard, next to Active sessions: CPU, RAM and the fullest disk now, each figure coloured only when past its threshold (same thresholds as the System page), the issues named in words, and a link to `/system`. It says *No samples yet* before the first sample, *Not collecting* when the newest is stale, and *Metrics unavailable* when a refresh fails (keeping the last figures). A `Surface`, not a `Metric`: a danger `Metric` paints its whole value red. |
48
52
  | `ImageDownloadBanner` | `app/agents/ImageDownloadBanner.tsx` | Agent image banner on the Agents page. Polls `/api/agents/image-status` every 2 s; offers **Download Image** when no supported version is downloaded, **Download <version>** when the registry publishes a newer one, and while a download runs shows the percent of layers finished and the latest line of docker's output (the bar is indeterminate until a percent can be computed). **Cancel** (`DELETE /api/agents/download-image`, own loading state) appears once `image-status` reports the download running; before that the button reads *Starting…* and is disabled. Every outcome — ready and cancelled for 3 s, a failure until the next action — comes from the server's `lastResult`, so one that ended while the page was closed or reloading is still reported if it is less than 30 s old (aged with `serverTime`). Each outcome is announced once per page load, although the Agents page mounts the banner in three places (mobile list, mobile detail, desktop). Downloading changes no agent. |
49
53
  | `VersionBanner` | `app/components/VersionBanner.tsx` | "Update available" banner for Rev4a itself, when `GET /api/update-check?check=1` reports a newer published version (dismissable per version, remembered in `localStorage`). **Update now** starts `POST /api/update-check`; because that update restarts the server, the banner cannot be told the outcome by the response: it records what it asked for in `sessionStorage`, polls `/api/update-check` until the installed version moves (two minutes at most) and reloads, then on the next mount either confirms "Updated to vX" or reports that the update did not complete and points at `update.log` (the update's own output). It never reloads blindly onto the same version. |
50
54
  | `BrowserAccessSection` / `OpenControlUiButton` | `app/agents/BrowserAccessSection.tsx` | Browser access to one agent's Control UI, in its detail panel: requests waiting for approval (Approve / Reject) and approved browsers (Rename / Revoke), refreshed every 5 s while mounted. A successful approve, reject, rename or revoke updates the list at once, since the refresh behind it runs the OpenClaw CLI and takes seconds; a read started before the mutation is discarded. On agents that require approval, "Invite link" fetches `/api/agents/[id]/invite-link` and shows the link in a read-only field with Copy, which uses the Clipboard API in a secure context and the field's selection over plain HTTP, plus a warning when the link uses localhost. `OpenControlUiButton` opens `/api/agents/[id]/open-control-ui` in a new tab inside the click; that route redirects to a one-time link that pairs the browser with no approval, or to the plain token link when none can be issued. Used on the agent cards and in the panel. |
@@ -58,7 +62,7 @@ All shared UI primitives live in `app/components/ui/` and are exported from `app
58
62
  | Wizard icons | `app/wizard/icons.tsx` | Shared SVG icons (flyweight pattern): `ArrowRightIcon`, `CheckIcon`, `CheckCircleIcon`, `DockerIcon`, `GatewayIcon`, `AgentIcon`, `TemplateIcon`, `LinkIcon`, `ConfigIcon`, `InformationIcon`. |
59
63
  | `tokens` | `tokens.ts` | TypeScript types for `Tone` and related token values. |
60
64
  | `Badge` | `Badge.tsx` | Inline status tag with tone variants (success, danger, warning, neutral). Used for channel chips, pairing labels, error/success messages. |
61
- | `Skeleton` + shape helpers | `app/components/Skeleton.tsx` | Shimmer placeholders. `Skeleton` is the primitive; the rest mirror one specific layout each: `TreeSkeleton`, `CodeSkeleton`, `CronJobsSkeleton`, `CronRunsSkeleton`, `CredentialCardsSkeleton`, `AgentSyncRowsSkeleton`, `SkeletonLines`, `SkeletonMetric`, `CardRowSkeleton`. A shape helper must match the real markup it stands in for — same row structure, same paddings, same element count where the count is known — so nothing reflows when data replaces it. |
65
+ | `Skeleton` + shape helpers | `app/components/Skeleton.tsx` | Shimmer placeholders. `Skeleton` is the primitive; the rest mirror one specific layout each: `TreeSkeleton`, `CodeSkeleton`, `CronJobsSkeleton`, `CronRunsSkeleton`, `CredentialCardsSkeleton`, `AgentSyncRowsSkeleton`, `SkeletonLines`, `SkeletonMetric`, `CardRowSkeleton`. A shape helper must match the real markup it stands in for — same row structure, same paddings, same element count where the count is known — so nothing reflows when data replaces it. Inside text (a `<p>`, a `<span>`) pass `as="span"`: a `<div>` there is invalid HTML, and the parser splits the paragraph in the server-rendered page. The System page keeps its own shapes in `app/system/SystemSkeleton.tsx`, each shimmer inside the class of the text it replaces so the line boxes are the real ones — measured to 0 px of movement on desktop and mobile. |
62
66
 
63
67
  ### Rules
64
68
 
@@ -153,7 +157,7 @@ For interactive pages that can change view/file/tab quickly:
153
157
  For every new/changed page:
154
158
 
155
159
  1. No new one-off card styles unless justified.
156
- 2. Use `app/components/ui/*` primitives first — `Button`, `Input`, `Select`, `Modal`, `Toast`, `Surface`, `Metric`, `StatusCard`, `Pill`, `Page`, `Tabs`, `ItemList`, `FilterBar`, `PropertyList`, `TemplateOption`, `LoadingSpinner`.
160
+ 2. Use `app/components/ui/*` primitives first — `Button`, `Input`, `Select`, `Modal`, `Toast`, `Surface`, `Metric`, `StatusCard`, `Meter`, `TimeSeriesChart`, `Pill`, `Page`, `Tabs`, `ItemList`, `FilterBar`, `PropertyList`, `TemplateOption`, `LoadingSpinner`.
157
161
  3. Async data must render skeleton, not fake zero values.
158
162
  4. Business logic stays in `lib`/API routes; UI consumes typed payloads.
159
163
  5. If a pattern is introduced, name it and keep it in `lib/patterns` or `app/components/ui`.
package/docs/REV4A.md CHANGED
@@ -97,9 +97,10 @@ When environment variables are saved via the Config page (Save & Restart):
97
97
 
98
98
  | Route | Description |
99
99
  |---|---|
100
- | `/` | System overview — containers, resources, quick links |
100
+ | `/` | System overview — health, sessions, the machine's CPU / RAM / disk (links to `/system`), cost, live feed |
101
101
  | `/agents` | Running agents (Docker containers with `AGENT_ID`), gateway token management, agent creation wizard, channel manager (Telegram pairing) |
102
102
  | `/containers` | All Docker containers on the host |
103
+ | `/system` | The host machine: CPU, memory and swap, storage per filesystem — current values and 1h / 24h / 7d / 30d history (from `metrics.db`, sampled by the daemon every 30 s) |
103
104
  | `/wizard` | First-run setup wizard (Welcome → Providers → Agent → Ready) |
104
105
  | `/workspace` | File explorer with tree view + editor — VPS host or container workspaces |
105
106
  | `/lineage` | Agent lineage / orchestration tree |
@@ -1906,25 +1906,50 @@ Register a parent→child relationship between sessions.
1906
1906
  ## Metrics
1907
1907
 
1908
1908
  ### `GET /api/metrics`
1909
- System metrics from the daemon's `system_metrics` table: the latest sample, the most
1910
- recent samples, and 24 h CPU aggregates.
1909
+ Machine metrics of the host Rev4a runs on — CPU, RAM, swap, storage — read from
1910
+ `metrics.db` (written by the daemon every 30 s, see `docs/dev/DATABASE.md`): the latest
1911
+ sample, live host facts, and a bucketed history for one range. The System page
1912
+ (`/system`) and the dashboard's Machine card use it.
1911
1913
 
1912
1914
  **Auth:** Bearer or browser cookie
1913
1915
 
1914
- **Query:** `limit` — how many `history` rows (default 60, max 1000): the most recent ones, returned oldest first.
1916
+ **Query:** `range` — `1h` (default), `24h`, `7d` or `30d`; anything else answers `400`.
1917
+ The range is cut into buckets of `bucket_s` seconds (about 120 points, never shorter than
1918
+ 60 s — two samples, so a timer that drifts a second leaves no false gaps), each with its
1919
+ average and peak: 60 s for `1h`, 720 s for `24h`, 5040 s for `7d`, 21 600 s for `30d`.
1915
1920
 
1916
1921
  **Response:**
1917
1922
  ```json
1918
1923
  {
1919
- "latest": { "cpu": 12.5, "ram_used_mb": 1820, "ram_total_mb": 4096, "disk_used_gb": 38, "disk_total_gb": 100, "load_avg": 0.42, "ts": 1749201000 },
1920
- "history": [ { "cpu": 12.5, "ram_used_mb": 1820, "ram_total_mb": 4096, "ts": 1749201000 } ],
1921
- "stats_24h": { "cpu_avg": 14, "cpu_max": 68 }
1924
+ "latest": {
1925
+ "ts": 1790419256, "age_s": 12, "cpu": 21.5,
1926
+ "ram_used_mb": 4779, "ram_total_mb": 7851, "swap_used_mb": 2048, "swap_total_mb": 4096,
1927
+ "load_avg_1m": 1.35, "load_avg_5m": 1.56, "load_avg_15m": 1.21,
1928
+ "disks": [
1929
+ { "mount": "/", "roles": ["root", "docker", "data"], "used_mb": 61440, "avail_mb": 38912, "total_mb": 102400, "percent": 61 }
1930
+ ]
1931
+ },
1932
+ "host": { "hostname": "vps", "platform": "linux", "cores": 4, "uptime_s": 290000 },
1933
+ "interval_s": 30,
1934
+ "range": "1h",
1935
+ "bucket_s": 60,
1936
+ "history": [ { "ts": 1790419230, "cpu_avg": 21.5, "cpu_max": 29.1, "ram_avg": 60.9, "ram_max": 61.2 } ],
1937
+ "disk_history": [ { "mount": "/", "points": [ { "ts": 1790419230, "percent": 61 } ] } ]
1922
1938
  }
1923
1939
  ```
1924
1940
 
1925
- `ts` is in **seconds**. The daemon records one sample per completed poll (every 30 s) and keeps
1926
- 30 days. A host without the OpenClaw CLI records none: the daemon stays idle there,
1927
- and `latest` is `null`.
1941
+ - `ts` values are in **seconds**; a `history` point's `ts` is the start of its bucket.
1942
+ Buckets with no samples are absent (the daemon was not running), not zero.
1943
+ - `ram_used_mb` is total − available, the figure `free` shows (Linux `/proc/meminfo`); on
1944
+ macOS active + wired + compressed pages. `swap_*` are `null` where there is no swap
1945
+ figure. `ram_avg`/`ram_max` are percentages.
1946
+ - `disks`: each filesystem once — `/`, Docker's data root (`docker info`), and Rev4a's
1947
+ data directory, merged when they are the same device; `roles` says which of the three
1948
+ it is. `percent` is used / (used + available), as `df` computes it.
1949
+ - `latest` is `null` before the daemon's first sample (the database is created when the
1950
+ daemon starts, the first sample about 5 s later), and when `metrics.db` has no tables. `age_s` older than 3 × `interval_s`
1951
+ means the daemon is not sampling; the page says so.
1952
+ - `host` is read live by the API process, which runs on the same host.
1928
1953
 
1929
1954
  ---
1930
1955
 
@@ -1932,7 +1957,11 @@ and `latest` is `null`.
1932
1957
 
1933
1958
  ### `GET /api/system-health`
1934
1959
  Aggregated checks of the runtime, cron jobs and usage cost, with recommendations. It
1935
- does not report CPU, RAM, disk or the Docker daemon (those are in `GET /api/metrics`).
1960
+ does not report CPU, RAM or disk (those are in `GET /api/metrics`). Its `runtime.heartbeat`
1961
+ check ("Daemon heartbeat") reads the age of the newest machine metrics sample in
1962
+ `metrics.db`: it shows the daemon is alive, not that the OpenClaw session poll ingests —
1963
+ the two run on separate timers. A `metrics.db` that cannot be read makes this one check
1964
+ say so; it does not fail the others.
1936
1965
 
1937
1966
  **Auth:** browser cookie or bearer token
1938
1967
 
@@ -1950,7 +1979,7 @@ does not report CPU, RAM, disk or the Docker daemon (those are in `GET /api/metr
1950
1979
  }
1951
1980
  ```
1952
1981
 
1953
- Checks: `runtime.sessions`, `runtime.ingestion`, `runtime.errors24h`,
1982
+ Checks: `runtime.sessions`, `runtime.heartbeat`, `runtime.errors24h`,
1954
1983
  `cost.usageBasedToday`, `cron.jobs`. `health` (overall and per check) is `"ok"` /
1955
1984
  `"warning"` / `"error"`; a recommendation's `severity` is `"info"` / `"warning"` /
1956
1985
  `"critical"`. `generatedAt` and `createdAt` are ISO strings.
@@ -2,13 +2,16 @@
2
2
 
3
3
  > **Last updated:** 2026-09-26
4
4
 
5
- Rev4a uses a single SQLite file (`events.db`) in WAL mode.
5
+ Rev4a keeps three SQLite files in WAL mode, by default all in the data directory's `data/`:
6
+ `events.db` (sessions and events, below), `metrics.db` (machine metrics, see
7
+ [Metrics Database](#metrics-database-metricsdb)) and `credentials.db` (see
8
+ [Credentials Database](#credentials-database-credentialsdb)).
6
9
 
7
10
  ## Location
8
11
 
9
12
  | Environment | Path |
10
13
  |---|---|
11
- | VPS (default) | `/path/to/rev4a/data/events.db` |
14
+ | Default | `~/.config/rev4a/data/events.db` (`$REV4A_DATA_DIR/data/events.db`) |
12
15
  | Custom | Set `REV4A_DB` env var |
13
16
 
14
17
  ## WAL Mode
@@ -19,8 +22,9 @@ PRAGMA synchronous = NORMAL;
19
22
  ```
20
23
 
21
24
  - The daemon writes; the Next.js API routes open read-only connections
22
- - `PRAGMA wal_checkpoint(PASSIVE)` runs after each daemon poll cycle
23
- - WAL files: `events.db-shm`, `events.db-wal` — do not delete while daemon is running
25
+ - `events.db`: `PRAGMA wal_checkpoint(PASSIVE)` after each daemon poll cycle, `FULL` every 10
26
+ - `metrics.db`: SQLite's automatic checkpoint (every 1000 pages) — no explicit one
27
+ - WAL files (`*.db-shm`, `*.db-wal`) — do not delete while the daemon is running
24
28
 
25
29
  ## Tables
26
30
 
@@ -88,27 +92,9 @@ CREATE TABLE cost_override (
88
92
 
89
93
  When a `cost_override` row exists for the current month, the UI displays the override value instead of the computed sum.
90
94
 
91
- ### `system_metrics` — infrastructure metrics
92
-
93
- ```sql
94
- CREATE TABLE system_metrics (
95
- id INTEGER PRIMARY KEY AUTOINCREMENT,
96
- ts INTEGER NOT NULL,
97
- cpu_percent REAL,
98
- ram_used_mb INTEGER,
99
- ram_total_mb INTEGER,
100
- disk_used_gb REAL,
101
- disk_total_gb REAL,
102
- load_avg_1m REAL
103
- );
104
-
105
- CREATE INDEX idx_metrics_ts ON system_metrics(ts);
106
- ```
107
-
108
- Rows older than 30 days are pruned automatically by the daemon (one row per completed 30 s poll;
109
- none on a host without the OpenClaw CLI, where the daemon stays idle).
110
- No feature consumes the `system_anomaly` events the daemon writes to `events` yet; they
111
- only pass through the generic event feeds.
95
+ The daemon also writes a `system_anomaly` event (session `system`) when the machine
96
+ crosses a threshold — see [Metrics Database](#metrics-database-metricsdb). No feature
97
+ consumes those events yet; they only pass through the generic event feeds.
112
98
 
113
99
  ### `agent_upgrades` — agent updates and rollbacks
114
100
 
@@ -255,6 +241,60 @@ off the machine. `restore.sh` refuses an archive with links in it and a running
255
241
  SQLite `-wal` cannot mix into the restored state), then resets the permissions. Agent volumes are not included —
256
242
  the dashboard's cold backups cover them.
257
243
 
244
+ ## Metrics Database (`metrics.db`)
245
+
246
+ Machine metrics of the host — CPU, RAM, swap, storage — kept apart from `events.db`, so
247
+ the event log holds sessions and events only and the history can be dropped on its own.
248
+
249
+ **Location:** `data/metrics.db`, next to `events.db` (it follows `REV4A_DB`'s directory; both
250
+ the daemon and the API resolve the data directory the same way, `REV4A_DATA_DIR` included).
251
+ **Created by** the daemon when it starts (`CREATE TABLE IF NOT EXISTS`), which is also its
252
+ only writer; `GET /api/metrics` and `GET /api/system-health` read it
253
+ (`lib/metrics-db.ts`). Until the daemon has run once the file does not exist, and the
254
+ API answers with `latest: null`. If the daemon cannot open it (corrupt, unwritable), it
255
+ logs `[METRICS] disabled` and keeps polling sessions without sampling.
256
+ **Sampling:** the first sample about 5 s after the daemon starts, then every 30 s, on a timer
257
+ of its own — independent of the OpenClaw session poll.
258
+ **Retention:** 30 days, pruned at every sample: ~86 400 `system_metrics` rows, and one
259
+ `system_disks` row per filesystem per sample (up to ~259 200 for three).
260
+
261
+ ```sql
262
+ CREATE TABLE system_metrics (
263
+ id INTEGER PRIMARY KEY AUTOINCREMENT,
264
+ ts INTEGER NOT NULL, -- Unix seconds
265
+ cpu_percent REAL, -- all cores, busy share since the previous sample
266
+ ram_used_mb INTEGER, -- total − available (Linux), as `free` shows it
267
+ ram_total_mb INTEGER,
268
+ swap_used_mb INTEGER, -- NULL where there is no swap figure
269
+ swap_total_mb INTEGER,
270
+ load_avg_1m REAL,
271
+ load_avg_5m REAL,
272
+ load_avg_15m REAL
273
+ );
274
+ CREATE INDEX idx_metrics_ts ON system_metrics(ts);
275
+
276
+ -- One row per filesystem per sample: `/`, Docker's data root and Rev4a's data directory,
277
+ -- each device once; roles lists which of them it is (e.g. 'root,docker,data').
278
+ CREATE TABLE system_disks (
279
+ ts INTEGER NOT NULL, -- the same ts as the system_metrics row
280
+ mount TEXT NOT NULL,
281
+ roles TEXT,
282
+ used_mb INTEGER, -- blocks − free
283
+ avail_mb INTEGER, -- what a normal user can still write
284
+ total_mb INTEGER
285
+ );
286
+ CREATE INDEX idx_disks_ts ON system_disks(ts);
287
+ CREATE INDEX idx_disks_mount_ts ON system_disks(mount, ts);
288
+ ```
289
+
290
+ **Anomalies** go to `events.db` as `system_anomaly` events: CPU above 85 % or RAM above
291
+ 90 % on two consecutive samples (at most once per 5 minutes per metric), and a filesystem
292
+ at 90 % or more once when it reaches the threshold, with no cooldown (again only after it
293
+ has dropped back under, or once after each daemon start: the state is kept in memory).
294
+
295
+ Installs from before this database may still have a `system_metrics` table in
296
+ `events.db`: nothing reads or writes it any more, and it can be dropped.
297
+
258
298
  ## Credentials Database (`credentials.db`)
259
299
 
260
300
  The credential hub uses a **separate SQLite file** (`credentials.db`) alongside
@@ -35,20 +35,24 @@ sends exactly three things:
35
35
  Lineage is **not** pushed over the stream; the Lineage page fetches it itself.
36
36
 
37
37
  The event feed is as close to real-time as Rev4a gets. However, events are
38
- generated by the daemon's poll cycle, so an event that just happened may take up
38
+ generated by the daemon every 30 seconds — session events by its poll of OpenClaw,
39
+ system anomalies by its machine sampling — so an event that just happened may take up
39
40
  to 30 seconds to appear, plus up to 5 more to reach the browser.
40
41
 
41
42
  ---
42
43
 
43
44
  ## System metrics (CPU, RAM, disk)
44
45
 
45
- **Update frequency:** every completed daemon poll cycle (30 seconds). A host without the
46
- OpenClaw CLI records no metrics at all: the daemon stays idle there.
46
+ **Update frequency:** every 30 seconds, sampled by the Rev4a daemon on its own timer
47
+ (it does not depend on OpenClaw being installed on the host). The System page refreshes
48
+ every 30 seconds too, so a figure is at most about a minute old.
47
49
 
48
- **Retention:** 30 days (older rows are pruned automatically each cycle)
50
+ **Retention:** 30 days, in their own database (`metrics.db`).
49
51
 
50
- These metrics are collected and stored, and served by `GET /api/metrics`. They
51
- are **not** shown on the dashboard — see the System Health note below.
52
+ Shown on the **System** page (`/system`) — current CPU, memory, swap and storage, and
53
+ their history over 1 hour, 24 hours, 7 days or 30 days — and, as three figures, on the
54
+ Dashboard's Machine card. If the newest sample is older than 90 seconds the page says
55
+ the collector is not running instead of showing stale numbers as current.
52
56
 
53
57
  ---
54
58
 
@@ -138,8 +142,9 @@ The tracked set is `USER.md`, `MEMORY.md`, `AGENTS.md`, `SOUL.md`, `HEARTBEAT.md
138
142
  ## How to check if data is fresh
139
143
 
140
144
  System health is a card on the **Dashboard**, not a page of its own. Its checks
141
- cover session runtime, ingestion freshness, errors in the last 24 hours, today's
142
- usage-based cost, and cron jobs. It reports no CPU, RAM, or disk figures.
145
+ cover session runtime, the daemon heartbeat, errors in the last 24 hours, today's
146
+ usage-based cost, and cron jobs. CPU, RAM and disk are on the System page (`/system`)
147
+ and the Dashboard's Machine card instead.
143
148
 
144
149
  To check directly on the server:
145
150
  ```bash
@@ -43,7 +43,7 @@ A third-party service token (GitHub PAT, Trello API key, Vercel token, Supabase
43
43
  A scheduled task that runs an agent automatically at a fixed time (e.g. every night at 3:15 AM). Configured with standard cron syntax. Jobs can be enabled/disabled per entry.
44
44
 
45
45
  ## Dashboard
46
- The main page of Rev4a (`/`). Shows live sessions, cost summary, health cards for Rev4a's runtime, cron and lineage, and a real-time event feed. It shows no CPU, RAM, disk or load average, and there is no setup banner: an incomplete wizard redirects to `/wizard`.
46
+ The main page of Rev4a (`/`). Shows live sessions, cost summary, health cards for Rev4a's runtime, cron and lineage, and a real-time event feed. Its Machine card shows the server's CPU, RAM and fullest disk right now, with a link to the System page. There is no setup banner: an incomplete wizard redirects to `/wizard`.
47
47
 
48
48
  ## Event
49
49
  A lifecycle occurrence for a session: spawned, completed, failed, or missing too long (spawn timeout); the daemon also records system anomalies (high CPU or RAM) as events. Events appear in the live feed on the Dashboard and are pushed via SSE every 5 seconds.
@@ -130,7 +130,10 @@ Rev4a's store for third-party service tokens (GitHub, Trello, Vercel, Supabase,
130
130
  ## Workspace
131
131
  The directory where an agent's operational files live (config, memory files, skills, plugins, scripts). The Workspace page lets you browse, view, and edit files across the host workspace, which appears as `Local`, and all agent containers.
132
132
 
133
+ ## System (page)
134
+ The page with the machine Rev4a runs on: CPU (with cores and load), memory and swap, and each disk's used and free space, now and over 1 hour, 24 hours, 7 days or 30 days. A bar turns yellow and then red when it reaches a threshold (CPU and memory 85 % / 95 %, disk 80 % / 90 %), and the level is written next to it.
135
+
133
136
  ## System Health
134
- A card on the Dashboard reporting whether Rev4a itself is working, not the hardware it runs on. `/api/system-health` returns checks for session runtime, feed ingestion freshness, errors in the last 24 hours, today's usage-based cost, and cron jobs, each with a health of `ok`, `warning` or `error`, plus a list of recommendations.
137
+ A card on the Dashboard reporting whether Rev4a itself is working, not the hardware it runs on. `/api/system-health` returns checks for session runtime, the daemon heartbeat (how old its latest machine sample is), errors in the last 24 hours, today's usage-based cost, and cron jobs, each with a health of `ok`, `warning` or `error`, plus a list of recommendations.
135
138
 
136
- It reports no CPU, RAM, disk or load average. Those metrics are collected by the daemon and stored for 30 days, and are served by `/api/metrics`, but no page displays them.
139
+ It reports no CPU, RAM, disk or load average: those are on the System page (`/system`), sampled every 30 seconds by the daemon and kept for 30 days.
@@ -9,7 +9,7 @@ Rev4a is the control panel for your AI agent infrastructure. It shows you everyt
9
9
  ### Dashboard (`/`)
10
10
  The main page. See live sessions (who's working right now), cost summary, an overall health indicator, and a real-time event feed. Click any session row to open the Session Drawer and see every tool call the agent made.
11
11
 
12
- The health cards cover Rev4a's own runtime, cron and watchdog, and lineage. They do **not** show CPU, RAM, disk, or load average: those metrics are collected and stored, but no page displays them. There is no wizard banner on the dashboard; an incomplete setup redirects to `/wizard` instead.
12
+ The health cards cover Rev4a's own runtime, cron and watchdog, and lineage. The Machine card shows the server's CPU, RAM and fullest disk right now, coloured when it reaches a threshold, and links to the System page for details. There is no wizard banner on the dashboard; an incomplete setup redirects to `/wizard` instead.
13
13
 
14
14
  ### First-run wizard (`/wizard`)
15
15
  First-run setup wizard that guides new users through configuration. Four steps:
@@ -81,6 +81,9 @@ Full list of all Docker containers on the server, including stopped ones. Each r
81
81
 
82
82
  The page shows no CPU or memory figures and has no start, stop, or restart buttons: it is read-only apart from the terminal link. Container lifecycle is managed from the Agents page.
83
83
 
84
+ ### System (`/system`)
85
+ The machine Rev4a runs on. Three cards — CPU (percent, cores, load average), memory (used, available, swap) and storage (each disk with its used and free space, and whether it is the system disk, where Docker keeps agents, or where Rev4a keeps its data) — then charts of CPU, memory and each disk over 1 hour, 24 hours, 7 days or 30 days (for CPU and memory the average line with the peak shaded). Hover a chart (or focus it and use the arrow keys) to read a point; each chart has a Table view. Bars turn yellow and red when they reach their thresholds. Figures refresh every 30 seconds; history is kept 30 days.
86
+
84
87
  ### Container Terminal (`/containers/terminal/[id]`)
85
88
  Live web terminal into a Docker container. Run commands, inspect files, debug issues — like SSH but in the browser.
86
89
 
@@ -10,6 +10,7 @@ PULSE is the in-dashboard AI concierge for Rev4a. This document defines what she
10
10
 
11
11
  - "Where do I find the agents page?"
12
12
  - "How do I get to the container list?"
13
+ - "Where do I see CPU, RAM or disk usage of the server?" — The System page (`/system`), and the Machine card on the Dashboard.
13
14
  - "Where can I see the first-run wizard?"
14
15
  - "Is there a page for cron jobs?"
15
16
  - "Where can I see my AI providers?" — The Gateway page (`/gateway`) shows providers and the model catalogue. Which model an agent runs is on the Agents page, in that agent's detail panel.
@@ -91,16 +91,6 @@ export function initializeRev4aDb(dbPath) {
91
91
  updated_at INTEGER
92
92
  );
93
93
 
94
- CREATE TABLE IF NOT EXISTS system_metrics (
95
- ts INTEGER PRIMARY KEY,
96
- cpu_percent REAL,
97
- ram_used_mb INTEGER,
98
- ram_total_mb INTEGER,
99
- disk_used_gb REAL,
100
- disk_total_gb REAL,
101
- load_avg_1m REAL
102
- );
103
-
104
94
  CREATE TABLE IF NOT EXISTS alert_state (
105
95
  alert_key TEXT PRIMARY KEY,
106
96
  kind TEXT NOT NULL,
@@ -114,7 +104,6 @@ export function initializeRev4aDb(dbPath) {
114
104
  CREATE INDEX IF NOT EXISTS idx_events_session ON events(session_id, ts);
115
105
  CREATE INDEX IF NOT EXISTS idx_sessions_updated ON sessions(updated_at);
116
106
  CREATE INDEX IF NOT EXISTS idx_sessions_started ON sessions(started_at);
117
- CREATE INDEX IF NOT EXISTS idx_metrics_ts ON system_metrics(ts);
118
107
  CREATE INDEX IF NOT EXISTS idx_tool_calls_session ON tool_calls(session_id, ts);
119
108
  CREATE INDEX IF NOT EXISTS idx_alert_state_updated ON alert_state(updated_at);
120
109
  `);
@@ -0,0 +1,48 @@
1
+ // Read access to the machine metrics database (daemon.js writes it).
2
+ import fs from 'fs';
3
+ import path from 'path';
4
+ import Database from 'better-sqlite3';
5
+ import { DB_PATH } from './db';
6
+
7
+ /** Next to events.db, as daemon.js places it; its own file so the event log stays clean. */
8
+ export const METRICS_DB_PATH = path.join(path.dirname(DB_PATH), 'metrics.db');
9
+
10
+ /**
11
+ * Open the metrics database read-only, or null when there is nothing to read yet: no file
12
+ * (first start, or a daemon that never ran) or a file without the daemon's tables (it
13
+ * failed right after creating it). The daemon owns the schema.
14
+ */
15
+ export function openMetricsDb(): Database.Database | null {
16
+ if (!fs.existsSync(METRICS_DB_PATH)) return null;
17
+ const db = new Database(METRICS_DB_PATH, { readonly: true, fileMustExist: true });
18
+ try {
19
+ const tables = db
20
+ .prepare("SELECT COUNT(*) AS n FROM sqlite_master WHERE type = 'table' AND name IN ('system_metrics', 'system_disks')")
21
+ .get() as { n: number };
22
+ if (tables.n === 2) return db;
23
+ } catch (e) {
24
+ db.close();
25
+ throw e;
26
+ }
27
+ db.close();
28
+ return null;
29
+ }
30
+
31
+ /**
32
+ * Seconds-since-epoch of the newest sample, or null when there is none. `available` is
33
+ * false when there is no metrics database to read, `error` when it could not be read.
34
+ * Never throws: a broken metrics.db must not fail a caller about something else.
35
+ */
36
+ export function latestMetricSample(): { available: boolean; ts: number | null; error?: string } {
37
+ let db: Database.Database | null = null;
38
+ try {
39
+ db = openMetricsDb();
40
+ if (!db) return { available: false, ts: null };
41
+ const row = db.prepare('SELECT MAX(ts) AS latest FROM system_metrics').get() as { latest: number | null };
42
+ return { available: true, ts: row.latest };
43
+ } catch (e) {
44
+ return { available: false, ts: null, error: (e as Error).message };
45
+ } finally {
46
+ db?.close();
47
+ }
48
+ }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@flame0510/project-aether",
3
- "version": "1.7.0",
3
+ "version": "1.8.0",
4
4
  "description": "Rev4a — Revolution for Agents. OpenClaw agent fleet orchestrator.",
5
5
  "keywords": [
6
6
  "openclaw",
package/scripts/backup.sh CHANGED
@@ -7,9 +7,9 @@
7
7
  #
8
8
  # What it archives: the whole Rev4a data directory — $REV4A_DATA_DIR when set,
9
9
  # otherwise ~/.config/rev4a (the same rule as lib/rev4a-paths.ts). That is `.env`
10
- # (secrets and configuration), everything under `data/` (events.db and credentials.db
11
- # with their -wal/-shm files, provider-keys.json, agents-token.json, model overrides,
12
- # the update log) and `shared/` (shared skills and rules). New files there are included
10
+ # (secrets and configuration), everything under `data/` (events.db, metrics.db and
11
+ # credentials.db with their -wal/-shm files, provider-keys.json, agents-token.json, model
12
+ # overrides, the update log) and `shared/` (shared skills and rules). New files there are included
13
13
  # automatically.
14
14
  #
15
15
  # Not included: agent volumes (the cold backups in the dashboard cover those), the