@flame0510/project-aether 1.10.0 → 1.11.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1 -1
- package/app/agents/CostSection.tsx +2 -8
- package/app/agents/PageClient.tsx +7 -0
- package/app/agents/PanelRow.tsx +9 -0
- package/app/agents/ResourceSection.tsx +105 -0
- package/app/api/assistant/route.ts +2 -2
- package/app/api/containers/route.ts +18 -2
- package/app/api/costs/agent/route.ts +2 -3
- package/app/api/metrics/alerts/route.ts +47 -0
- package/app/api/metrics/containers/route.ts +54 -0
- package/app/api/metrics/route.ts +4 -6
- package/app/components/ui/Accordion.tsx +45 -0
- package/app/components/ui/TimeSeriesChart.tsx +20 -2
- package/app/components/ui/index.ts +1 -0
- package/app/containers/ContainersClient.tsx +48 -6
- package/app/globals.css +36 -0
- package/app/system/AgentCharts.tsx +138 -0
- package/app/system/AgentsSection.tsx +281 -0
- package/app/system/PageClient.tsx +7 -7
- package/app/system/RecentAlerts.tsx +72 -0
- package/app/system/SystemSkeleton.tsx +53 -1
- package/app/system/loading.tsx +5 -1
- package/daemon.js +195 -4
- package/docs/ARCHITECTURE.md +36 -5
- package/docs/FRONTEND-ARCHITECTURE.md +12 -3
- package/docs/REV4A.md +4 -4
- package/docs/dev/API-REFERENCE.md +91 -17
- package/docs/dev/DATABASE.md +58 -2
- package/docs/rag/DATA-FRESHNESS.md +16 -1
- package/docs/rag/GLOSSARY.md +2 -2
- package/docs/rag/REV4A-OVERVIEW.md +7 -3
- package/docs/rag/WHAT-I-CAN-ANSWER.md +4 -1
- package/lib/container-metrics.ts +340 -0
- package/lib/docker-socket-path.js +133 -0
- package/lib/docker-socket.ts +10 -99
- package/lib/docker-stats.js +284 -0
- package/lib/metrics-db.ts +15 -6
- package/lib/utils/format.ts +29 -3
- package/package.json +1 -1
- package/scripts/test-docker-stats.mjs +270 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Data Freshness in Rev4a
|
|
2
2
|
|
|
3
|
-
> **Last updated:** 2026-
|
|
3
|
+
> **Last updated:** 2026-10-01
|
|
4
4
|
|
|
5
5
|
Understanding how current the data in Rev4a is — and what is truly real-time vs. periodically updated.
|
|
6
6
|
|
|
@@ -56,6 +56,21 @@ the collector is not running instead of showing stale numbers as current.
|
|
|
56
56
|
|
|
57
57
|
---
|
|
58
58
|
|
|
59
|
+
## Container consumption (System page, Containers page, agent panel)
|
|
60
|
+
|
|
61
|
+
**Update frequency:** every 60 seconds for CPU, memory, processes, network and block I/O of
|
|
62
|
+
each running container; every 10 minutes for the disk Docker holds (volumes, writable layers,
|
|
63
|
+
images, build cache). Both are sampled by the Rev4a daemon on timers of their own and kept 30
|
|
64
|
+
days in `metrics.db`. The first figures after the daemon starts take about 20 seconds (a CPU
|
|
65
|
+
figure needs two readings). A stopped container leaves a gap in its charts, not zeros.
|
|
66
|
+
|
|
67
|
+
CPU is a share of Docker's cores and memory a share of Docker's memory: on Docker Desktop that
|
|
68
|
+
is its virtual machine (for example 8 GB of a 16 GB Mac), on a server it is the machine. If the
|
|
69
|
+
newest sample is older than three minutes the System page says it is not collecting — the
|
|
70
|
+
figures still on show are the last known ones, not live. The agent panel stands its "now" rows
|
|
71
|
+
down, and the Containers page clears its usage row rather than leave stale figures. A container
|
|
72
|
+
still seen in Docker's list but without fresh statistics says "no data", not "stopped".
|
|
73
|
+
|
|
59
74
|
## Agent costs (Costs page)
|
|
60
75
|
|
|
61
76
|
**Update frequency:** every 5 minutes. The Rev4a daemon asks each running agent
|
package/docs/rag/GLOSSARY.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Rev4a Glossary
|
|
2
2
|
|
|
3
|
-
> **Last updated:** 2026-
|
|
3
|
+
> **Last updated:** 2026-10-01
|
|
4
4
|
|
|
5
5
|
Terms you'll encounter while using the Rev4a dashboard.
|
|
6
6
|
|
|
@@ -22,7 +22,7 @@ A compressed archive (`.tar.gz`) of an agent's persistent volume (`/root/`). Bac
|
|
|
22
22
|
Which browsers may open an agent's Control UI. On OpenClaw 9.x every new browser must be approved once; the approval is remembered per browser. Managed in the "BROWSER ACCESS" section of the agent detail panel: approve or reject waiting browsers, rename or revoke approved ones. The Open button uses a one-time link that skips the approval. The Invite link button gives a link for someone else: their browser still waits for approval.
|
|
23
23
|
|
|
24
24
|
## Container
|
|
25
|
-
A Docker container running on the server. Each agent runs in its own container. The Containers page shows all containers, including infrastructure ones such as databases and reverse proxies, with a web terminal link (a real shell in the browser; a dropped connection is resumed for 2 minutes).
|
|
25
|
+
A Docker container running on the server. Each agent runs in its own container. The Containers page shows all containers (a running one also with its CPU, memory and processes now), including infrastructure ones such as databases and reverse proxies, with a web terminal link (a real shell in the browser; a dropped connection is resumed for 2 minutes).
|
|
26
26
|
|
|
27
27
|
## Channel
|
|
28
28
|
A communication channel (Telegram) configured on an agent. Channels allow users to send DMs to the agent via messaging apps. The Channel Manager modal lets you connect/disconnect Telegram and manage pairings (approve/reject senders).
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# What is Rev4a?
|
|
2
2
|
|
|
3
|
-
> **Last updated:** 2026-
|
|
3
|
+
> **Last updated:** 2026-10-01
|
|
4
4
|
|
|
5
5
|
Rev4a is the control panel for your AI agent infrastructure. It shows you everything your agents are doing, how much they cost, and whether the system is healthy — all in one dashboard.
|
|
6
6
|
|
|
@@ -77,9 +77,9 @@ Interactive graph showing session family trees — which agent spawned which chi
|
|
|
77
77
|
Connects agents to AI providers. Add and remove provider API keys, enable or disable individual models in the catalogue read from `models.config.json`, and push the resulting model list to every agent container. It does not assign models to agents: choosing which model an agent runs is done per agent, in the Model section of the agent's detail panel on the Agents page. Every model row has a **Details** button: the modal shows the model's description, price and provider, its architecture and parameter count when the weights are public, reasoning modes, capabilities, benchmarks with sources, licence and how much disk the weights take — the numbers come from the model card on Hugging Face and from curated entries that cite a source, never from estimates.
|
|
78
78
|
|
|
79
79
|
### Containers (`/containers`)
|
|
80
|
-
Full list of all Docker containers on the server, including stopped ones. Each row shows name, status, image, IP, ports, and an agent pill where applicable. Click "Terminal" to open an interactive shell into that container.
|
|
80
|
+
Full list of all Docker containers on the server, including stopped ones. Each row shows name, status, image, IP, ports, and an agent pill where applicable. Click "Terminal" to open an interactive shell into that container. A running container also shows what it consumes now — CPU share, memory and processes — with a "Charts" link to its history on the System page.
|
|
81
81
|
|
|
82
|
-
The page
|
|
82
|
+
The page has no start, stop, or restart buttons: it is read-only apart from the terminal link. Container lifecycle is managed from the Agents page.
|
|
83
83
|
|
|
84
84
|
### Costs (`/costs`)
|
|
85
85
|
What each agent spent, as its own OpenClaw priced it with the prices Rev4a syncs. Pick today, 7 days or 30 days (UTC days): the total, where the money went (input, output, cache), the tokens and how much input came from cache, a chart of the spend per day, the spend by agent and by model, and the most expensive sessions. Calls that could not be priced are listed by model, and agents whose figures are not current are flagged. When a DeepSeek model is on offer, the header shows whether DeepSeek is billing peak or off-peak (half price) and until when. Figures are read from the agents every 5 minutes. *Against the bill* compares them with what DeepSeek and OpenRouter actually charged (their balance or usage, read every hour); a small gap is normal, since the same key may pay for the assistant too. Each agent's own spend is also in its panel on the Agents page (COSTS).
|
|
@@ -87,6 +87,10 @@ What each agent spent, as its own OpenClaw priced it with the prices Rev4a syncs
|
|
|
87
87
|
### System (`/system`)
|
|
88
88
|
The machine Rev4a runs on. Three cards — CPU (percent, cores, load average), memory (used, available, swap) and storage (each disk with its used and free space, and whether it is the system disk, where Docker keeps agents, or where Rev4a keeps its data) — then charts of CPU, memory and each disk over 1 hour, 24 hours, 7 days or 30 days (for CPU and memory the average line with the peak shaded). Hover a chart (or focus it and use the arrow keys) to read a point; each chart has a Table view. Bars turn yellow and red when they reach their thresholds. Figures refresh every 30 seconds; history is kept 30 days.
|
|
89
89
|
|
|
90
|
+
### Agents and containers (on the System page)
|
|
91
|
+
Under the machine's charts: what Docker holds on disk (images, volumes, writable layers, build cache, and how much of each no container uses), then one row per container sorted by CPU. A closed row shows its name, whether it is an agent, and its CPU share, memory, storage and processes now, with the average and peak over the chosen range (1 hour to 30 days); open it for the same three charts as the machine — CPU, Memory, Storage — for that container. "Everything else" is what is not a container (the host, Rev4a, other software). Each agent's panel on the Agents page has the same numbers in its RESOURCES section, with a link to its charts. Below, the machine's recent alerts: CPU, memory or a disk that crossed its limit, and when.
|
|
92
|
+
When Docker still sees a container but it has no current statistics, its row says "no data". If collection stops, the current figures are hidden and old rows say "not current" instead of implying a live reading.
|
|
93
|
+
|
|
90
94
|
### Container Terminal (`/containers/terminal/[id]`)
|
|
91
95
|
Live web terminal into a running Docker container — like SSH in the browser: Tab completion, command history, `top`, `vi`, resizing with the window. If the connection drops (network, laptop sleep) the shell keeps running for 2 minutes and reconnecting picks it up; reloading the page resumes it too. **Close** ends the shell. If the container is stopped or missing, the page says so and offers **Reconnect**. Only a logged-in user can open it.
|
|
92
96
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# What PULSE Can Answer
|
|
2
2
|
|
|
3
|
-
> **Last updated:** 2026-
|
|
3
|
+
> **Last updated:** 2026-10-01
|
|
4
4
|
|
|
5
5
|
PULSE is the in-dashboard AI concierge for Rev4a. This document defines what she can and cannot answer.
|
|
6
6
|
|
|
@@ -11,6 +11,9 @@ PULSE is the in-dashboard AI concierge for Rev4a. This document defines what she
|
|
|
11
11
|
- "Where do I find the agents page?"
|
|
12
12
|
- "How do I get to the container list?"
|
|
13
13
|
- "Where do I see CPU, RAM or disk usage of the server?" — The System page (`/system`), and the Machine card on the Dashboard.
|
|
14
|
+
- "Which agent is using the most CPU or memory?" — The "Agents and containers" section of the System page: one row per container, sorted by CPU, with memory, storage and processes; open a row for its charts. An agent's own panel has the same numbers under RESOURCES.
|
|
15
|
+
- "What is filling the disk?" — The "Docker storage" block on the System page: images, volumes (an unused volume is named), writable layers and build cache.
|
|
16
|
+
- "Why is there a red bar or an alert on the System page?" — "Recent alerts" at the bottom of the System page lists what crossed its limit and when.
|
|
14
17
|
- "Where can I see the first-run wizard?"
|
|
15
18
|
- "Is there a page for cron jobs?"
|
|
16
19
|
- "Where can I see my AI providers?" — The Gateway page (`/gateway`) shows providers and the model catalogue. Which model an agent runs is on the Agents page, in that agent's detail panel.
|
|
@@ -0,0 +1,340 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* What each container consumes, read from metrics.db (daemon.js writes it — see
|
|
3
|
+
* docs/dev/DATABASE.md, Container consumption). The queries behind the System page's
|
|
4
|
+
* *Agents* section, the agent panel and the Containers page; the routes stay thin.
|
|
5
|
+
*
|
|
6
|
+
* CPU is stored in cores and shown as a share of Docker's cores (`docker_host.ncpu`): on the
|
|
7
|
+
* VPS that is the machine, on Docker Desktop its VM. Memory is the working set in MB,
|
|
8
|
+
* with its share of Docker's memory next to it.
|
|
9
|
+
*/
|
|
10
|
+
import type Database from 'better-sqlite3';
|
|
11
|
+
|
|
12
|
+
/** How often daemon.js samples containers, and reads Docker's disk (seconds). */
|
|
13
|
+
export const CONTAINER_INTERVAL_S = 60;
|
|
14
|
+
export const STORAGE_INTERVAL_S = 600;
|
|
15
|
+
/** History points per chart. */
|
|
16
|
+
const POINTS = 120;
|
|
17
|
+
/** A container whose newest sample is older than this many intervals is not running. */
|
|
18
|
+
const RUNNING_WITHIN_INTERVALS = 3;
|
|
19
|
+
|
|
20
|
+
const TABLES = ['container_metrics', 'container_sources', 'docker_host', 'docker_storage'];
|
|
21
|
+
|
|
22
|
+
/** False until the daemon has created the tables (first start after an update). */
|
|
23
|
+
export function containerTablesReady(db: Database.Database): boolean {
|
|
24
|
+
const row = db
|
|
25
|
+
.prepare(`SELECT COUNT(*) AS n FROM sqlite_master WHERE type = 'table' AND name IN (${TABLES.map(() => '?').join(', ')})`)
|
|
26
|
+
.get(...TABLES) as { n: number };
|
|
27
|
+
return row.n === TABLES.length;
|
|
28
|
+
}
|
|
29
|
+
|
|
30
|
+
/** Seconds per chart point: at least two samples, so a drifting timer leaves no empty bucket. */
|
|
31
|
+
export function bucketSeconds(spanS: number, intervalS: number): number {
|
|
32
|
+
return Math.max(2 * intervalS, Math.round(spanS / POINTS));
|
|
33
|
+
}
|
|
34
|
+
|
|
35
|
+
export interface DockerHost { ncpu: number; mem_total_mb: number }
|
|
36
|
+
|
|
37
|
+
export function readDockerHost(db: Database.Database): DockerHost | null {
|
|
38
|
+
const row = db.prepare('SELECT ncpu, mem_total_mb FROM docker_host WHERE id = 1').get() as
|
|
39
|
+
{ ncpu: number | null; mem_total_mb: number | null } | undefined;
|
|
40
|
+
return row && row.ncpu && row.mem_total_mb ? { ncpu: row.ncpu, mem_total_mb: row.mem_total_mb } : null;
|
|
41
|
+
}
|
|
42
|
+
|
|
43
|
+
/** `value` as a percent of `whole`, to `digits` decimals — CPU needs two: an idle agent is a few hundredths of a percent of the machine. */
|
|
44
|
+
const pct = (value: number | null, whole: number | null | undefined, digits = 1): number | null =>
|
|
45
|
+
value === null || !whole ? null : Math.round((value / whole) * 100 * 10 ** digits) / 10 ** digits;
|
|
46
|
+
const round2 = (n: number | null): number | null => (n === null ? null : Math.round(n * 100) / 100);
|
|
47
|
+
|
|
48
|
+
/** The named volumes a container uses, as daemon.js recorded them. */
|
|
49
|
+
function volumesOf(list: string | null): string[] {
|
|
50
|
+
return list ? list.split(',').filter(Boolean) : [];
|
|
51
|
+
}
|
|
52
|
+
|
|
53
|
+
export interface ContainerNow {
|
|
54
|
+
cpu_cores: number | null;
|
|
55
|
+
cpu_percent: number | null;
|
|
56
|
+
mem_mb: number | null;
|
|
57
|
+
mem_percent: number | null;
|
|
58
|
+
pids: number | null;
|
|
59
|
+
net_rx_bps: number | null;
|
|
60
|
+
net_tx_bps: number | null;
|
|
61
|
+
blk_read_bps: number | null;
|
|
62
|
+
blk_write_bps: number | null;
|
|
63
|
+
}
|
|
64
|
+
|
|
65
|
+
export interface ContainerRow {
|
|
66
|
+
container: string;
|
|
67
|
+
name: string;
|
|
68
|
+
agent_id: string | null;
|
|
69
|
+
is_agent: boolean;
|
|
70
|
+
running: boolean;
|
|
71
|
+
last_seen: number;
|
|
72
|
+
now: ContainerNow | null;
|
|
73
|
+
/** Average and peak over the range, CPU as a share of Docker's cores. */
|
|
74
|
+
range: { cpu_avg_percent: number | null; cpu_max_percent: number | null; mem_avg_mb: number | null; mem_max_mb: number | null };
|
|
75
|
+
/** Named volumes plus the writable layer, from the last reading of Docker's disk. */
|
|
76
|
+
storage: { volume_mb: number | null; layer_mb: number | null; total_mb: number | null; ts: number | null };
|
|
77
|
+
}
|
|
78
|
+
|
|
79
|
+
export interface DockerStorage {
|
|
80
|
+
ts: number;
|
|
81
|
+
images_mb: number | null;
|
|
82
|
+
images_reclaimable_mb: number | null;
|
|
83
|
+
build_cache_mb: number | null;
|
|
84
|
+
build_cache_reclaimable_mb: number | null;
|
|
85
|
+
volumes_mb: number;
|
|
86
|
+
/** What no container uses, all of it — `unused_volumes` lists only the largest five. */
|
|
87
|
+
unused_volumes_mb: number;
|
|
88
|
+
layers_mb: number;
|
|
89
|
+
/** Volumes no container uses (cold backups, leftovers), largest first. */
|
|
90
|
+
unused_volumes: { name: string; size_mb: number }[];
|
|
91
|
+
}
|
|
92
|
+
|
|
93
|
+
export interface ContainerOverview {
|
|
94
|
+
host: DockerHost | null;
|
|
95
|
+
/** Unix seconds of the newest sample, and how old it is. */
|
|
96
|
+
sampled_at: number | null;
|
|
97
|
+
age_s: number | null;
|
|
98
|
+
containers: ContainerRow[];
|
|
99
|
+
/** The machine minus the containers: what is not one of them (the host, Rev4a, other software). */
|
|
100
|
+
rest: { cpu_percent: number | null; mem_mb: number | null } | null;
|
|
101
|
+
docker_storage: DockerStorage | null;
|
|
102
|
+
}
|
|
103
|
+
|
|
104
|
+
interface NowRow {
|
|
105
|
+
ts: number; cpu_cores: number | null; mem_mb: number | null; pids: number | null;
|
|
106
|
+
net_rx_bps: number | null; net_tx_bps: number | null; blk_read_bps: number | null; blk_write_bps: number | null;
|
|
107
|
+
}
|
|
108
|
+
|
|
109
|
+
/** The newest reading of each volume, layer and total, or null before the first. */
|
|
110
|
+
export function readDockerStorage(db: Database.Database): DockerStorage | null {
|
|
111
|
+
const ts = (db.prepare('SELECT MAX(ts) AS ts FROM docker_storage').get() as { ts: number | null }).ts;
|
|
112
|
+
if (ts === null) return null;
|
|
113
|
+
const rows = db.prepare('SELECT kind, name, size_mb, reclaimable_mb FROM docker_storage WHERE ts = ?').all(ts) as
|
|
114
|
+
{ kind: string; name: string; size_mb: number | null; reclaimable_mb: number | null }[];
|
|
115
|
+
const one = (kind: string) => rows.find((r) => r.kind === kind);
|
|
116
|
+
const sum = (kind: string) => rows.filter((r) => r.kind === kind).reduce((acc, r) => acc + (r.size_mb ?? 0), 0);
|
|
117
|
+
const unused = rows
|
|
118
|
+
.filter((r) => r.kind === 'volume' && (r.reclaimable_mb ?? 0) > 0)
|
|
119
|
+
.map((r) => ({ name: r.name, size_mb: r.reclaimable_mb as number }))
|
|
120
|
+
.sort((a, b) => b.size_mb - a.size_mb);
|
|
121
|
+
return {
|
|
122
|
+
ts,
|
|
123
|
+
images_mb: one('images')?.size_mb ?? null,
|
|
124
|
+
images_reclaimable_mb: one('images')?.reclaimable_mb ?? null,
|
|
125
|
+
build_cache_mb: one('build_cache')?.size_mb ?? null,
|
|
126
|
+
build_cache_reclaimable_mb: one('build_cache')?.reclaimable_mb ?? null,
|
|
127
|
+
volumes_mb: sum('volume'),
|
|
128
|
+
unused_volumes_mb: unused.reduce((a, v) => a + v.size_mb, 0),
|
|
129
|
+
layers_mb: sum('layer'),
|
|
130
|
+
unused_volumes: unused.slice(0, 5),
|
|
131
|
+
};
|
|
132
|
+
}
|
|
133
|
+
|
|
134
|
+
interface AggregateRow { container: string; cpu_avg: number | null; cpu_max: number | null; mem_avg: number | null; mem_max: number | null }
|
|
135
|
+
|
|
136
|
+
/**
|
|
137
|
+
* From 24 hours up, the average and peak scan every sample of the range: the page's heaviest
|
|
138
|
+
* query — synchronous, so it holds the server's event loop (about 100 ms for ten containers
|
|
139
|
+
* over 24 h, measured) on every 30 s poll — and its answer hardly moves in two minutes. So
|
|
140
|
+
* those ranges are kept that long; shorter ones are cheap and read fresh.
|
|
141
|
+
*/
|
|
142
|
+
const LONG_RANGE_S = 24 * 3_600;
|
|
143
|
+
const AGGREGATE_TTL_MS = 120_000;
|
|
144
|
+
const aggregateCache = new Map<string, { at: number; rows: Map<string, AggregateRow> }>();
|
|
145
|
+
|
|
146
|
+
function rangeAggregates(db: Database.Database, spanS: number, since: number): Map<string, AggregateRow> {
|
|
147
|
+
const cacheable = spanS >= LONG_RANGE_S;
|
|
148
|
+
const key = `${db.name}:${spanS}`;
|
|
149
|
+
if (cacheable) {
|
|
150
|
+
const hit = aggregateCache.get(key);
|
|
151
|
+
if (hit && Date.now() - hit.at < AGGREGATE_TTL_MS) return hit.rows;
|
|
152
|
+
}
|
|
153
|
+
const rows = new Map(
|
|
154
|
+
(db
|
|
155
|
+
.prepare(
|
|
156
|
+
`SELECT container, AVG(cpu_cores) AS cpu_avg, MAX(cpu_cores) AS cpu_max, AVG(mem_mb) AS mem_avg, MAX(mem_mb) AS mem_max
|
|
157
|
+
FROM container_metrics WHERE ts > ? GROUP BY container`,
|
|
158
|
+
)
|
|
159
|
+
.all(since) as AggregateRow[])
|
|
160
|
+
.map((r) => [r.container, r]),
|
|
161
|
+
);
|
|
162
|
+
if (cacheable) aggregateCache.set(key, { at: Date.now(), rows });
|
|
163
|
+
return rows;
|
|
164
|
+
}
|
|
165
|
+
|
|
166
|
+
/**
|
|
167
|
+
* Every container seen in the range: who it is, what it uses now, its average and peak over
|
|
168
|
+
* the range, and its disk — running ones first, by CPU.
|
|
169
|
+
*/
|
|
170
|
+
export function readContainerOverview(
|
|
171
|
+
db: Database.Database,
|
|
172
|
+
spanS: number,
|
|
173
|
+
nowS = Math.floor(Date.now() / 1000),
|
|
174
|
+
/** Cores of the machine the API runs on, to put the containers' CPU on the machine's scale in `rest`. */
|
|
175
|
+
hostCores?: number,
|
|
176
|
+
): ContainerOverview {
|
|
177
|
+
const host = readDockerHost(db);
|
|
178
|
+
const sampledAt = (db.prepare('SELECT MAX(ts) AS ts FROM container_metrics').get() as { ts: number | null }).ts;
|
|
179
|
+
const since = nowS - spanS;
|
|
180
|
+
|
|
181
|
+
const sources = db
|
|
182
|
+
.prepare('SELECT container, agent_id, name, is_agent, volumes, last_seen FROM container_sources WHERE last_seen > ?')
|
|
183
|
+
.all(since) as { container: string; agent_id: string | null; name: string | null; is_agent: number; volumes: string | null; last_seen: number }[];
|
|
184
|
+
|
|
185
|
+
const aggregates = rangeAggregates(db, spanS, since);
|
|
186
|
+
// Running is judged against the newest sample, not the wall clock: a daemon that stopped must
|
|
187
|
+
// not make every container look stopped (the page says "Not collecting" instead) — but never
|
|
188
|
+
// against a sample from the future, which a clock set back would leave behind.
|
|
189
|
+
const reference = sampledAt === null ? null : Math.min(sampledAt, nowS);
|
|
190
|
+
|
|
191
|
+
const storage = readDockerStorage(db);
|
|
192
|
+
const storageRows = storage
|
|
193
|
+
? (db.prepare('SELECT kind, name, size_mb FROM docker_storage WHERE ts = ?').all(storage.ts) as { kind: string; name: string; size_mb: number | null }[])
|
|
194
|
+
: [];
|
|
195
|
+
const sizeOf = (kind: string, name: string) => storageRows.find((r) => r.kind === kind && r.name === name)?.size_mb ?? null;
|
|
196
|
+
|
|
197
|
+
const newest = db.prepare(
|
|
198
|
+
`SELECT ts, cpu_cores, mem_mb, pids, net_rx_bps, net_tx_bps, blk_read_bps, blk_write_bps
|
|
199
|
+
FROM container_metrics WHERE container = ? ORDER BY ts DESC LIMIT 3`,
|
|
200
|
+
);
|
|
201
|
+
|
|
202
|
+
const containers: ContainerRow[] = sources.map((s) => {
|
|
203
|
+
const recent = newest.all(s.container) as NowRow[];
|
|
204
|
+
const latest = recent[0];
|
|
205
|
+
const running = Boolean(latest && reference !== null && latest.ts >= reference - RUNNING_WITHIN_INTERVALS * CONTAINER_INTERVAL_S);
|
|
206
|
+
// A rate is null on the first sample after a start; the previous reading stands in for it.
|
|
207
|
+
const rate = (key: keyof NowRow) => recent.find((r) => r[key] !== null)?.[key] ?? null;
|
|
208
|
+
const cores = running ? (rate('cpu_cores') as number | null) : null;
|
|
209
|
+
const now: ContainerNow | null = running
|
|
210
|
+
? {
|
|
211
|
+
cpu_cores: cores,
|
|
212
|
+
cpu_percent: pct(cores, host?.ncpu, 2),
|
|
213
|
+
mem_mb: latest.mem_mb,
|
|
214
|
+
mem_percent: pct(latest.mem_mb, host?.mem_total_mb),
|
|
215
|
+
pids: latest.pids,
|
|
216
|
+
net_rx_bps: rate('net_rx_bps') as number | null,
|
|
217
|
+
net_tx_bps: rate('net_tx_bps') as number | null,
|
|
218
|
+
blk_read_bps: rate('blk_read_bps') as number | null,
|
|
219
|
+
blk_write_bps: rate('blk_write_bps') as number | null,
|
|
220
|
+
}
|
|
221
|
+
: null;
|
|
222
|
+
const agg = aggregates.get(s.container);
|
|
223
|
+
const volumeSizes = volumesOf(s.volumes).map((v) => sizeOf('volume', v));
|
|
224
|
+
const volume = volumeSizes.length && volumeSizes.some((v) => v !== null) ? volumeSizes.reduce<number>((a, v) => a + (v ?? 0), 0) : null;
|
|
225
|
+
const layer = sizeOf('layer', s.container);
|
|
226
|
+
return {
|
|
227
|
+
container: s.container,
|
|
228
|
+
name: s.name || s.container,
|
|
229
|
+
agent_id: s.agent_id,
|
|
230
|
+
is_agent: s.is_agent === 1,
|
|
231
|
+
running,
|
|
232
|
+
last_seen: s.last_seen,
|
|
233
|
+
now,
|
|
234
|
+
range: {
|
|
235
|
+
cpu_avg_percent: pct(agg?.cpu_avg ?? null, host?.ncpu, 2),
|
|
236
|
+
cpu_max_percent: pct(agg?.cpu_max ?? null, host?.ncpu, 2),
|
|
237
|
+
mem_avg_mb: agg?.mem_avg == null ? null : Math.round(agg.mem_avg),
|
|
238
|
+
mem_max_mb: agg?.mem_max ?? null,
|
|
239
|
+
},
|
|
240
|
+
storage: {
|
|
241
|
+
volume_mb: volume,
|
|
242
|
+
layer_mb: layer,
|
|
243
|
+
total_mb: volume === null && layer === null ? null : (volume ?? 0) + (layer ?? 0),
|
|
244
|
+
ts: storage?.ts ?? null,
|
|
245
|
+
},
|
|
246
|
+
};
|
|
247
|
+
});
|
|
248
|
+
|
|
249
|
+
containers.sort((a, b) =>
|
|
250
|
+
Number(b.running) - Number(a.running)
|
|
251
|
+
|| (b.now?.cpu_percent ?? -1) - (a.now?.cpu_percent ?? -1)
|
|
252
|
+
|| a.name.localeCompare(b.name));
|
|
253
|
+
|
|
254
|
+
// Machine minus the containers, from the machine's newest sample.
|
|
255
|
+
const machine = db
|
|
256
|
+
.prepare('SELECT cpu_percent, ram_used_mb FROM system_metrics ORDER BY ts DESC LIMIT 1')
|
|
257
|
+
.get() as { cpu_percent: number | null; ram_used_mb: number | null } | undefined;
|
|
258
|
+
const running = containers.filter((c) => c.running && c.now);
|
|
259
|
+
// The machine's CPU is a share of the host's cores; a container's is of Docker's — the same on a
|
|
260
|
+
// server, not on Docker Desktop, whose VM may have fewer. Compare in cores over the host's.
|
|
261
|
+
const containersShare = hostCores && hostCores > 0
|
|
262
|
+
? (running.reduce((a, c) => a + (c.now?.cpu_cores ?? 0), 0) / hostCores) * 100
|
|
263
|
+
: running.reduce((a, c) => a + (c.now?.cpu_percent ?? 0), 0);
|
|
264
|
+
const rest = machine && host
|
|
265
|
+
? {
|
|
266
|
+
cpu_percent: machine.cpu_percent === null ? null : round2(Math.max(0, machine.cpu_percent - containersShare)),
|
|
267
|
+
mem_mb: machine.ram_used_mb === null ? null : Math.max(0, machine.ram_used_mb - running.reduce((a, c) => a + (c.now?.mem_mb ?? 0), 0)),
|
|
268
|
+
}
|
|
269
|
+
: null;
|
|
270
|
+
|
|
271
|
+
return {
|
|
272
|
+
host,
|
|
273
|
+
sampled_at: sampledAt,
|
|
274
|
+
age_s: sampledAt === null ? null : Math.max(0, nowS - sampledAt),
|
|
275
|
+
containers,
|
|
276
|
+
rest,
|
|
277
|
+
docker_storage: storage,
|
|
278
|
+
};
|
|
279
|
+
}
|
|
280
|
+
|
|
281
|
+
export interface ContainerHistory {
|
|
282
|
+
host: DockerHost | null;
|
|
283
|
+
container: string;
|
|
284
|
+
name: string | null;
|
|
285
|
+
bucket_s: number;
|
|
286
|
+
/** CPU as a share of Docker's cores (average and peak), memory in MB (average and peak). */
|
|
287
|
+
history: { ts: number; cpu_avg: number | null; cpu_max: number | null; mem_avg: number | null; mem_max: number | null }[];
|
|
288
|
+
storage_bucket_s: number;
|
|
289
|
+
/** Named volumes plus the writable layer, MB. */
|
|
290
|
+
storage_history: { ts: number; total_mb: number | null }[];
|
|
291
|
+
}
|
|
292
|
+
|
|
293
|
+
/** One container's history, bucketed like the machine's (`GET /api/metrics`). */
|
|
294
|
+
export function readContainerHistory(db: Database.Database, container: string, spanS: number, nowS = Math.floor(Date.now() / 1000)): ContainerHistory {
|
|
295
|
+
const host = readDockerHost(db);
|
|
296
|
+
const since = nowS - spanS;
|
|
297
|
+
const bucket = bucketSeconds(spanS, CONTAINER_INTERVAL_S);
|
|
298
|
+
const storageBucket = bucketSeconds(spanS, STORAGE_INTERVAL_S);
|
|
299
|
+
|
|
300
|
+
const source = db.prepare('SELECT name, volumes FROM container_sources WHERE container = ?').get(container) as
|
|
301
|
+
{ name: string | null; volumes: string | null } | undefined;
|
|
302
|
+
|
|
303
|
+
const rows = db
|
|
304
|
+
.prepare(
|
|
305
|
+
`SELECT (ts / CAST(@bucket AS INTEGER)) * CAST(@bucket AS INTEGER) AS ts,
|
|
306
|
+
AVG(cpu_cores) AS cpu_avg, MAX(cpu_cores) AS cpu_max, AVG(mem_mb) AS mem_avg, MAX(mem_mb) AS mem_max
|
|
307
|
+
FROM container_metrics WHERE container = @container AND ts > @since
|
|
308
|
+
GROUP BY ts / CAST(@bucket AS INTEGER) ORDER BY ts`,
|
|
309
|
+
)
|
|
310
|
+
.all({ bucket, container, since }) as { ts: number; cpu_avg: number | null; cpu_max: number | null; mem_avg: number | null; mem_max: number | null }[];
|
|
311
|
+
|
|
312
|
+
const volumes = volumesOf(source?.volumes ?? null);
|
|
313
|
+
const placeholders = volumes.map(() => '?').join(', ');
|
|
314
|
+
// Volumes and the writable layer, summed per reading, then averaged per bucket.
|
|
315
|
+
const storageRows = db
|
|
316
|
+
.prepare(
|
|
317
|
+
`SELECT (ts / CAST(? AS INTEGER)) * CAST(? AS INTEGER) AS ts, AVG(total) AS total_mb FROM (
|
|
318
|
+
SELECT ts, SUM(size_mb) AS total FROM docker_storage
|
|
319
|
+
WHERE ts > ? AND ((kind = 'layer' AND name = ?)${volumes.length ? ` OR (kind = 'volume' AND name IN (${placeholders}))` : ''})
|
|
320
|
+
GROUP BY ts
|
|
321
|
+
) GROUP BY ts / CAST(? AS INTEGER) ORDER BY ts`,
|
|
322
|
+
)
|
|
323
|
+
.all(storageBucket, storageBucket, since, container, ...volumes, storageBucket) as { ts: number; total_mb: number | null }[];
|
|
324
|
+
|
|
325
|
+
return {
|
|
326
|
+
host,
|
|
327
|
+
container,
|
|
328
|
+
name: source?.name ?? null,
|
|
329
|
+
bucket_s: bucket,
|
|
330
|
+
history: rows.map((r) => ({
|
|
331
|
+
ts: r.ts,
|
|
332
|
+
cpu_avg: pct(r.cpu_avg, host?.ncpu, 2),
|
|
333
|
+
cpu_max: pct(r.cpu_max, host?.ncpu, 2),
|
|
334
|
+
mem_avg: r.mem_avg === null ? null : Math.round(r.mem_avg),
|
|
335
|
+
mem_max: r.mem_max,
|
|
336
|
+
})),
|
|
337
|
+
storage_bucket_s: storageBucket,
|
|
338
|
+
storage_history: storageRows.map((r) => ({ ts: r.ts, total_mb: r.total_mb === null ? null : Math.round(r.total_mb) })),
|
|
339
|
+
};
|
|
340
|
+
}
|
|
@@ -0,0 +1,133 @@
|
|
|
1
|
+
'use strict';
|
|
2
|
+
|
|
3
|
+
/**
|
|
4
|
+
* Where the Docker Engine API socket is, and a JSON request against it.
|
|
5
|
+
*
|
|
6
|
+
* Plain CommonJS on purpose: the routes reach it through lib/docker-socket.ts, and
|
|
7
|
+
* daemon.js — run by node directly, with no bundler — requires it. One copy of the rules
|
|
8
|
+
* below, so the daemon cannot drift from the dashboard on where Docker lives.
|
|
9
|
+
*
|
|
10
|
+
* The socket path is NOT the same everywhere. `/var/run/docker.sock` is the Linux default
|
|
11
|
+
* (and what the VPS uses), but Docker Desktop on macOS puts it at `~/.docker/run/docker.sock`
|
|
12
|
+
* and does not create the /var/run symlink unless the user opts in. Hardcoding the Linux
|
|
13
|
+
* path made every read-only Docker route return an empty list on macOS — agents existed
|
|
14
|
+
* and ran, but the dashboard showed nothing, because creation shells out to the `docker`
|
|
15
|
+
* CLI (which reads the context) while listing went through this socket.
|
|
16
|
+
*/
|
|
17
|
+
const http = require('http');
|
|
18
|
+
const fs = require('fs');
|
|
19
|
+
const os = require('os');
|
|
20
|
+
const path = require('path');
|
|
21
|
+
const { execFileSync } = require('child_process');
|
|
22
|
+
|
|
23
|
+
let cached = null;
|
|
24
|
+
let lastMissAt = 0;
|
|
25
|
+
|
|
26
|
+
/**
|
|
27
|
+
* How long a failed resolution is remembered. Long enough that a Docker outage doesn't
|
|
28
|
+
* spawn a `docker context inspect` per request, short enough that a server which started
|
|
29
|
+
* before the Docker daemon picks it up on its own.
|
|
30
|
+
*/
|
|
31
|
+
const MISS_TTL_MS = 5000;
|
|
32
|
+
|
|
33
|
+
function fromDockerHostEnv() {
|
|
34
|
+
const raw = process.env.DOCKER_HOST;
|
|
35
|
+
if (!raw) return null;
|
|
36
|
+
// Only unix sockets are usable here; tcp:// would need a different client.
|
|
37
|
+
if (!raw.startsWith('unix://')) return null;
|
|
38
|
+
return raw.slice('unix://'.length);
|
|
39
|
+
}
|
|
40
|
+
|
|
41
|
+
function fromDockerContext() {
|
|
42
|
+
try {
|
|
43
|
+
const out = execFileSync(
|
|
44
|
+
'docker',
|
|
45
|
+
['context', 'inspect', '--format', '{{.Endpoints.docker.Host}}'],
|
|
46
|
+
{ encoding: 'utf-8', timeout: 3000, stdio: ['ignore', 'pipe', 'ignore'] },
|
|
47
|
+
).trim();
|
|
48
|
+
return out.startsWith('unix://') ? out.slice('unix://'.length) : null;
|
|
49
|
+
} catch {
|
|
50
|
+
return null;
|
|
51
|
+
}
|
|
52
|
+
}
|
|
53
|
+
|
|
54
|
+
function usable(candidate) {
|
|
55
|
+
if (!candidate) return null;
|
|
56
|
+
try {
|
|
57
|
+
fs.accessSync(candidate, fs.constants.R_OK | fs.constants.W_OK);
|
|
58
|
+
return candidate;
|
|
59
|
+
} catch {
|
|
60
|
+
return null;
|
|
61
|
+
}
|
|
62
|
+
}
|
|
63
|
+
|
|
64
|
+
/**
|
|
65
|
+
* Resolve the Docker socket, most authoritative source first. Returns null when Docker
|
|
66
|
+
* isn't reachable at all.
|
|
67
|
+
*
|
|
68
|
+
* A success is cached for the life of the process — the path doesn't move. A failure is
|
|
69
|
+
* only cached for MISS_TTL_MS: Rev4a can legitimately start before the Docker daemon is up
|
|
70
|
+
* (systemd ordering, Docker Desktop still booting), and caching that failure permanently
|
|
71
|
+
* would leave every Docker route returning an empty list until someone restarted the server.
|
|
72
|
+
*
|
|
73
|
+
* @returns {string | null}
|
|
74
|
+
*/
|
|
75
|
+
function resolveDockerSocket() {
|
|
76
|
+
if (cached) return cached;
|
|
77
|
+
if (Date.now() - lastMissAt < MISS_TTL_MS) return null;
|
|
78
|
+
|
|
79
|
+
const found =
|
|
80
|
+
usable(fromDockerHostEnv()) ??
|
|
81
|
+
usable(fromDockerContext()) ??
|
|
82
|
+
// Docker Desktop (macOS, and Windows with WSL integration)
|
|
83
|
+
usable(path.join(os.homedir(), '.docker', 'run', 'docker.sock')) ??
|
|
84
|
+
// Linux default — the VPS lands here
|
|
85
|
+
usable('/var/run/docker.sock');
|
|
86
|
+
|
|
87
|
+
if (found) cached = found;
|
|
88
|
+
else lastMissAt = Date.now();
|
|
89
|
+
|
|
90
|
+
return found;
|
|
91
|
+
}
|
|
92
|
+
|
|
93
|
+
/**
|
|
94
|
+
* A request against the Docker Engine API, resolved as JSON. Rejects when Docker is
|
|
95
|
+
* unreachable, the answer is an HTTP error (Docker's error bodies are JSON too — a status
|
|
96
|
+
* is never resolved as data), the payload isn't JSON, or `timeoutMs` (when given) passes.
|
|
97
|
+
*
|
|
98
|
+
* @template [T=any]
|
|
99
|
+
* @param {string} method
|
|
100
|
+
* @param {string} apiPath
|
|
101
|
+
* @param {number} [timeoutMs] 0 or omitted: no timeout
|
|
102
|
+
* @returns {Promise<T>}
|
|
103
|
+
*/
|
|
104
|
+
function dockerRequestJson(method, apiPath, timeoutMs = 0) {
|
|
105
|
+
return new Promise((resolve, reject) => {
|
|
106
|
+
const socketPath = resolveDockerSocket();
|
|
107
|
+
if (!socketPath) {
|
|
108
|
+
reject(new Error('Docker socket not found'));
|
|
109
|
+
return;
|
|
110
|
+
}
|
|
111
|
+
|
|
112
|
+
const req = http.request(
|
|
113
|
+
{ socketPath, path: apiPath, method, headers: { Host: 'localhost' } },
|
|
114
|
+
(res) => {
|
|
115
|
+
let data = '';
|
|
116
|
+
res.on('data', (chunk) => { data += chunk; });
|
|
117
|
+
res.on('end', () => {
|
|
118
|
+
if ((res.statusCode ?? 500) >= 400) {
|
|
119
|
+
reject(new Error(`Docker error ${res.statusCode}: ${data.slice(0, 200) || '(no body)'}`));
|
|
120
|
+
return;
|
|
121
|
+
}
|
|
122
|
+
try { resolve(JSON.parse(data)); }
|
|
123
|
+
catch { reject(new Error('Invalid JSON from Docker')); }
|
|
124
|
+
});
|
|
125
|
+
},
|
|
126
|
+
);
|
|
127
|
+
if (timeoutMs > 0) req.setTimeout(timeoutMs, () => req.destroy(new Error('Docker request timed out')));
|
|
128
|
+
req.on('error', reject);
|
|
129
|
+
req.end();
|
|
130
|
+
});
|
|
131
|
+
}
|
|
132
|
+
|
|
133
|
+
module.exports = { resolveDockerSocket, dockerRequestJson };
|