@andrian.yablonskyy/thub-coordinator 1.1.17 → 1.1.18
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +2 -2
- package/public/js/client-config.js +10 -1
- package/src/api/agent.js +15 -1
- package/src/api/resource.js +40 -0
- package/src/services/jobs.js +44 -1
- package/src/services/registry.js +15 -1
- package/test/scheduler.test.js +2 -2
- package/test/usb-power.test.js +129 -0
- package/test/views.test.js +25 -1
- package/views/help/_agent-cli.pug +13 -0
- package/views/help/_ci.pug +198 -9
- package/views/help/_client-machines.pug +3 -0
- package/views/help/_client-setup.pug +2 -0
- package/views/help/_docker.pug +29 -0
- package/views/help/_env.pug +3 -0
- package/views/help/_git.pug +24 -11
- package/views/help/_lab-devices.pug +154 -0
- package/views/help/_storage.pug +218 -0
- package/views/help/_troubleshooting.pug +8 -0
- package/views/help/_usb-power.pug +147 -0
- package/views/help/index.pug +7 -1
- package/views/mixins/client-config.pug +3 -0
|
@@ -16,6 +16,7 @@
|
|
|
16
16
|
li A physical machine (a mini-PC, NUC or Raspberry-class ARM64 board) next to the bench.
|
|
17
17
|
li USB for the DUT's debugger (ST-Link), USB-serial adapters (UART) and DUT USB. A #[strong powered USB hub] is recommended.
|
|
18
18
|
li #[code stlink-tools] and/or #[code openocd], #[code usbutils] (for the dashboard's USB scan).
|
|
19
|
+
li To power-cycle boards: a hub with per-port power switching and #[code uhubctl] (#[a(href="#usb-power") USB port power]).
|
|
19
20
|
li Docker only if jobs run containers in their #[code --command].
|
|
20
21
|
li One Client instance per board. One machine can drive several boards (up to 8 ST-Links, UARTs and USB devices per instance).
|
|
21
22
|
.col-md-6
|
|
@@ -85,3 +86,5 @@
|
|
|
85
86
|
li Whatever tools your jobs' commands use (git, docker, a flasher) are installed. Their credentials come with each job as #[code --env] (#[a(href="#git") Using git], #[a(href="#docker") Using Docker]), not from the host.
|
|
86
87
|
li Under systemd, the service can only write to its state directory. Jobs write to #[code $THUB_WORK_DIR], not to #[code ~].
|
|
87
88
|
li Optional: a nightly reboot from the runner card's #[strong Reboot] tab (cron, host-local time). It waits for running jobs.
|
|
89
|
+
li Optional (HW): boards on switchable USB ports, so jobs and their owners can power-cycle them (#[a(href="#usb-power") USB port power]).
|
|
90
|
+
li Optional: boards on a smart socket or PDU outlet that jobs switch themselves. Install the tools they call (#[code curl], #[code snmp], …) and give each instance its device (#[a(href="#lab-devices") Smart sockets, PDUs, devices]).
|
|
@@ -124,3 +124,5 @@
|
|
|
124
124
|
thub-client restart
|
|
125
125
|
thub-client --config ~/.config/thub/dut1.json status # a specific instance
|
|
126
126
|
thub-client udev --print # the udev rules this config generates
|
|
127
|
+
thub-client power status # USB port power (hw-devices.usbPower)
|
|
128
|
+
thub-client power reset --port 2 --delay 3 # power-cycle one board, 3 s off
|
package/views/help/_docker.pug
CHANGED
|
@@ -41,6 +41,35 @@
|
|
|
41
41
|
td: code --device /dev/thub/dut1-uart
|
|
42
42
|
td HW: gives the container the DUT's UART (#[code /dev/thub/dut1-uart]). Use #[code --privileged] only if you must.
|
|
43
43
|
|
|
44
|
+
h3.h6 Registry logins, per registry
|
|
45
|
+
p.small.
|
|
46
|
+
Always #[code docker login <registry> -u <user> --password-stdin], with the secret piped in from #[code --env]
|
|
47
|
+
(never #[code -p]). Only the user and password differ. For cloud registries, mint the short-lived token in the CI step
|
|
48
|
+
and pass only that.
|
|
49
|
+
.table-responsive
|
|
50
|
+
table.table.table-sm.small.align-middle
|
|
51
|
+
thead
|
|
52
|
+
tr
|
|
53
|
+
th Registry
|
|
54
|
+
th User
|
|
55
|
+
th Password / token
|
|
56
|
+
tbody
|
|
57
|
+
each row in [['Your own (registry:2, Harbor)', 'a robot or service account', 'its password or token'], ['JFrog Artifactory', 'a service user', 'an access token (see Artifact storage)'], ['GitHub (ghcr.io)', 'any GitHub user name', 'a PAT with read:packages, or the workflow\'s GITHUB_TOKEN'], ['GitLab', 'gitlab-ci-token', 'the job\'s CI_JOB_TOKEN (or a deploy token)'], ['Docker Hub', 'your user', 'a personal access token'], ['AWS ECR', 'AWS', 'aws ecr get-login-password (12 h)'], ['Google Artifact Registry', 'oauth2accesstoken', 'gcloud auth print-access-token (1 h)'], ['Azure Container Registry', 'a service principal or token name', 'its secret or token password']]
|
|
58
|
+
tr
|
|
59
|
+
td= row[0]
|
|
60
|
+
td: code= row[1]
|
|
61
|
+
td= row[2]
|
|
62
|
+
+code('AWS ECR: token minted in the CI step, never AWS credentials in the job').
|
|
63
|
+
# CI step with AWS access (e.g. OIDC → an IAM role that can pull from ECR)
|
|
64
|
+
export ECR_PASSWORD=$(aws ecr get-login-password --region eu-central-1)
|
|
65
|
+
thub run --type sw --env ECR_PASSWORD --env REG=123456789012.dkr.ecr.eu-central-1.amazonaws.com \
|
|
66
|
+
--command 'echo "$ECR_PASSWORD" | docker login "$REG" -u AWS --password-stdin &&
|
|
67
|
+
docker run --rm --user "$(id -u):$(id -g)" -v "$THUB_WORK_DIR/src:/work" -w /work "$REG/test-runner:1.4" ./run-tests.sh' \
|
|
68
|
+
--wait
|
|
69
|
+
p.small.
|
|
70
|
+
A registry with a private-CA certificate needs #[code /etc/docker/certs.d/<registry>[:port]/ca.crt] on the Client host
|
|
71
|
+
(plus #[code client.cert] / #[code client.key] there if it asks for a client certificate). Never use #[code insecure-registries].
|
|
72
|
+
|
|
44
73
|
h3.h6 Clone and test inside a container, with a deploy key and a registry login from the job
|
|
45
74
|
p.small.
|
|
46
75
|
Everything comes from the Agent as #[code --env]: the registry and its credentials, and the private key the
|
package/views/help/_env.pug
CHANGED
|
@@ -39,6 +39,9 @@
|
|
|
39
39
|
['JOB_ARG / JOB_ARG_<n>', '--arg', 'all, space-separated / each (also "$@")'],
|
|
40
40
|
['JOB_TIMEOUT', '--timeout', 'seconds'],
|
|
41
41
|
['JOB_PRIORITY', '--priority', '0–100'],
|
|
42
|
+
['JOB_POWER_ON_START', '--power-on-start', 'on, off or reset (HW, when given)'],
|
|
43
|
+
['JOB_POWER_ON_END', '--power-on-end', 'on, off or reset (HW, when given)'],
|
|
44
|
+
['JOB_POWER_RESET_DELAY', '--power-reset-delay', 'seconds (when given)'],
|
|
42
45
|
['JOB_META_<KEY>', '--meta key=value', 'the value'],
|
|
43
46
|
['<NAME>', '--env NAME=value', 'your own variables, under the names you chose']
|
|
44
47
|
]
|
package/views/help/_git.pug
CHANGED
|
@@ -18,19 +18,32 @@
|
|
|
18
18
|
The token appears only inside the command's environment. To keep it out of #[code .git/config] as well, pass it as a header:
|
|
19
19
|
#[code git -c http.extraHeader="Authorization: Bearer $GH_TOKEN" clone …].
|
|
20
20
|
|
|
21
|
-
h3.h6 SSH with a key
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
21
|
+
h3.h6 SSH with a deploy key and a pinned host key
|
|
22
|
+
p.small.
|
|
23
|
+
Create a key pair for TestHub only, without a passphrase (#[code ssh-keygen -t ed25519 -N '' -f thub_deploy]), and add the
|
|
24
|
+
#[strong public] key as a read-only deploy key: GitHub, Settings → Deploy keys; GitLab, Settings → Repository → Deploy keys;
|
|
25
|
+
Bitbucket, Repository settings → Access keys. The private key goes with the job as #[code --env]. Pin the server's
|
|
26
|
+
host key too: fetch it once with #[code ssh-keyscan], check it against the fingerprints your git host publishes,
|
|
27
|
+
and store both in your CI's secret store.
|
|
28
|
+
+code('Deploy key and known_hosts from --env, used for this job only').
|
|
29
|
+
# once, then store both in your CI's secret store:
|
|
30
|
+
ssh-keyscan github.com > known_hosts # GitLab: gitlab.com; Bitbucket: bitbucket.org; your server: -p <port> host
|
|
31
|
+
export GIT_KEY="$(cat thub_deploy)" GIT_KNOWN_HOSTS="$(cat known_hosts)"
|
|
32
|
+
|
|
33
|
+
thub run --type hw --label board:nucleo-f401re --env GIT_KEY --env GIT_KNOWN_HOSTS \
|
|
34
|
+
--command 'umask 077 &&
|
|
35
|
+
printf "%s\n" "$GIT_KEY" > "$THUB_WORK_DIR/id" && printf "%s\n" "$GIT_KNOWN_HOSTS" > "$THUB_WORK_DIR/known_hosts" &&
|
|
36
|
+
export GIT_SSH_COMMAND="ssh -i $THUB_WORK_DIR/id -o IdentitiesOnly=yes -o UserKnownHostsFile=$THUB_WORK_DIR/known_hosts -o StrictHostKeyChecking=yes" &&
|
|
37
|
+
git clone --depth 1 --recurse-submodules --shallow-submodules git@github.com:yourorg/firmware-tests.git src && cd src &&
|
|
29
38
|
./ci/test.sh' \
|
|
30
39
|
--wait
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
40
|
+
ul.small
|
|
41
|
+
li #[code GIT_SSH_COMMAND] is exported, so #[code git fetch] and submodules on the same host use it as well.
|
|
42
|
+
li The key and #[code known_hosts] are written with #[code umask 077] into the job's directory and deleted with it.
|
|
43
|
+
li #[code StrictHostKeyChecking=yes] refuses an unknown host. #[code accept-new] protects nothing here: every job starts with an empty #[code known_hosts].
|
|
44
|
+
li Another port (Bitbucket Server, Gitea): #[code ssh://git@git.example.com:7999/proj/repo.git] and #[code ssh-keyscan -p 7999 git.example.com].
|
|
45
|
+
li To keep the key off the disk, load it into #[code ssh-agent]: #[code eval "$(ssh-agent -s)" && printf "%s\n" "$GIT_KEY" | ssh-add -].
|
|
46
|
+
li A key in the service user's #[code ~/.ssh] needs no #[code --env], but every job on that host can then use it.
|
|
34
47
|
|
|
35
48
|
p.small.
|
|
36
49
|
To clone #[em inside a container] with a key passed the same way (held by #[code ssh-agent], never written to disk), see
|
|
@@ -0,0 +1,154 @@
|
|
|
1
|
+
+section('lab-devices', 'Smart sockets, PDUs and other lab devices', 'plug')
|
|
2
|
+
p.
|
|
3
|
+
TestHub has one power control built in: #[a(href="#usb-power") USB port power] with uhubctl. Anything else in the lab
|
|
4
|
+
is driven by #[strong the job itself]: a smart socket (Shelly, Tasmota, TP-Link Kasa, or any socket Home Assistant
|
|
5
|
+
controls), a PDU outlet (APC, Raritan, …), a USB relay board, a bench supply. The job's #[code --command], or a script
|
|
6
|
+
from its repository, calls the device with its own tools. The Client needs no configuration for it.
|
|
7
|
+
|
|
8
|
+
.table-responsive
|
|
9
|
+
table.table.table-sm.small.align-middle
|
|
10
|
+
thead
|
|
11
|
+
tr
|
|
12
|
+
th(style="width: 18%")
|
|
13
|
+
th USB port power (built in)
|
|
14
|
+
th Socket, PDU or other device
|
|
15
|
+
tbody
|
|
16
|
+
tr
|
|
17
|
+
th Configured in
|
|
18
|
+
td The Client's #[code hw-devices.usbPower]
|
|
19
|
+
td The job's command or script, plus a variable naming the bench's device
|
|
20
|
+
tr
|
|
21
|
+
th Start / end of job
|
|
22
|
+
td #[code --power-on-start] / #[code --power-on-end]. The end action runs even on cancel
|
|
23
|
+
td The script's own steps; on cancel, only as reliable as its trap (below)
|
|
24
|
+
tr
|
|
25
|
+
th While the job runs
|
|
26
|
+
td #[code thub power reset <jobId>] (owner)
|
|
27
|
+
td Whenever the script decides. It can't be triggered from outside
|
|
28
|
+
tr
|
|
29
|
+
th Credentials
|
|
30
|
+
td None
|
|
31
|
+
td #[code --env], or the Client host's environment
|
|
32
|
+
|
|
33
|
+
h3.h6 Before you start
|
|
34
|
+
ul.small
|
|
35
|
+
li #[strong The command runs on the Client host]: the device must be reachable from there (the lab network), never through the Coordinator.
|
|
36
|
+
li #[strong The tools must be on the Client host]: #[code curl], #[code snmpset] (#[code sudo apt install snmp]), #[code kasa] (#[code pipx install python-kasa]), …
|
|
37
|
+
li #[strong Credentials come with the job] as #[code --env], so they're masked and dropped when the job ends.
|
|
38
|
+
li
|
|
39
|
+
| #[strong Tell each Client which device is its bench's.] Jobs inherit the Client service's environment, so add a
|
|
40
|
+
| systemd drop-in per instance. Every job on that Client can read it, so put addresses there, not secrets.
|
|
41
|
+
| Alternatively, pin jobs with #[code --client] and map #[code $JOB_CLIENT] in your script.
|
|
42
|
+
+code('Client host').
|
|
43
|
+
sudo systemctl edit thub-client@dut1
|
|
44
|
+
# [Service]
|
|
45
|
+
# Environment=BENCH_POWER=shelly:10.0.20.11
|
|
46
|
+
sudo systemctl restart thub-client@dut1
|
|
47
|
+
systemctl show thub-client@dut1 -p Environment # check
|
|
48
|
+
|
|
49
|
+
h3.h6 One command per device
|
|
50
|
+
p.small.
|
|
51
|
+
Each example switches the device off, waits, switches it on, then tests. API details vary between models and firmware,
|
|
52
|
+
so check your device's manual.
|
|
53
|
+
+code('Shelly Gen2/Plus/Pro (RPC; digest auth, user admin)').
|
|
54
|
+
thub run --type hw --label board:nucleo-f401re --env SHELLY_PASSWORD \
|
|
55
|
+
--command 'S="http://$BENCH_IP/rpc/Switch.Set?id=0"
|
|
56
|
+
curl -fsS --digest -u "admin:$SHELLY_PASSWORD" "$S&on=false" && sleep 2 &&
|
|
57
|
+
curl -fsS --digest -u "admin:$SHELLY_PASSWORD" "$S&on=true" && sleep 3 && ./ci/test.sh' --wait
|
|
58
|
+
# Gen1: http://$BENCH_IP/relay/0?turn=off|on (basic auth: curl -u user:password)
|
|
59
|
+
+code('Tasmota (Power2 … for multi-relay devices)').
|
|
60
|
+
thub run --type hw … --env TASMOTA_PASSWORD \
|
|
61
|
+
--command 'T="http://$BENCH_IP/cm?user=admin&password=$TASMOTA_PASSWORD&cmnd=Power%20"
|
|
62
|
+
curl -fsS "${T}Off" && sleep 2 && curl -fsS "${T}On" && sleep 3 && ./ci/test.sh'
|
|
63
|
+
+code('Home Assistant (any socket it controls; a long-lived access token)').
|
|
64
|
+
thub run --type hw … --env HA_URL=http://ha.lab:8123 --env HA_TOKEN \
|
|
65
|
+
--command 'ha() { curl -fsS -X POST -H "Authorization: Bearer $HA_TOKEN" -H "Content-Type: application/json" \
|
|
66
|
+
-d "{\"entity_id\":\"$BENCH_SWITCH\"}" "$HA_URL/api/services/switch/turn_$1"; }
|
|
67
|
+
ha off && sleep 2 && ha on && sleep 3 && ./ci/test.sh'
|
|
68
|
+
+code('TP-Link Kasa / Tapo (python-kasa; newer firmware needs --username/--password)').
|
|
69
|
+
thub run --type hw … --command 'kasa --host "$BENCH_IP" off && sleep 2 && kasa --host "$BENCH_IP" on && sleep 3 && ./ci/test.sh'
|
|
70
|
+
+code('APC switched PDU (SNMP: 1 on, 2 off, 3 reboot; older AP79xx: sPDUOutletCtl …1.4.4.2.1.3.<outlet>)').
|
|
71
|
+
thub run --type hw … --env PDU_COMMUNITY \
|
|
72
|
+
--command 'snmpset -v1 -c "$PDU_COMMUNITY" "$PDU_HOST" 1.3.6.1.4.1.318.1.1.12.3.3.1.1.4.$PDU_OUTLET i 3 && sleep 15 && ./ci/test.sh'
|
|
73
|
+
# state: snmpget -v1 -c "$PDU_COMMUNITY" "$PDU_HOST" 1.3.6.1.4.1.318.1.1.12.3.5.1.1.4.$PDU_OUTLET (1 on, 2 off)
|
|
74
|
+
p.small.
|
|
75
|
+
Anything else with a command line or a network interface works the same way: an SSH-managed PDU
|
|
76
|
+
(#[code ssh pdu1 "olOff 5"]), a USB relay board (#[code usbrelay]), a bench supply over SCPI
|
|
77
|
+
(#[code echo 'OUTP OFF' | nc psu1.lab 5025]).
|
|
78
|
+
|
|
79
|
+
h3.h6 A power script in the test repository
|
|
80
|
+
p.small.
|
|
81
|
+
Once several jobs need it, keep the device logic in the repository the job clones. A single script can cover every
|
|
82
|
+
kind of device in the lab, keyed by #[code $BENCH_POWER]:
|
|
83
|
+
+code('ci/power.sh').
|
|
84
|
+
#!/bin/sh
|
|
85
|
+
# ci/power.sh on|off|reset [off-seconds] — switch this bench's socket or PDU outlet.
|
|
86
|
+
# BENCH_POWER: shelly:10.0.20.11 | tasmota:10.0.20.12 | ha:switch.bench1 | apc:pdu1.lab:5
|
|
87
|
+
# Credentials (--env): POWER_PASSWORD (Shelly, Tasmota), HA_URL + HA_TOKEN, PDU_COMMUNITY.
|
|
88
|
+
set -eu
|
|
89
|
+
: "${BENCH_POWER:?BENCH_POWER is not set for this Client}"
|
|
90
|
+
kind=${BENCH_POWER%%:*} target=${BENCH_POWER#*:}
|
|
91
|
+
|
|
92
|
+
switch() { # $1: on | off
|
|
93
|
+
case $kind in
|
|
94
|
+
shelly)
|
|
95
|
+
url="http://$target/rpc/Switch.Set?id=0&on=$([ "$1" = on ] && echo true || echo false)"
|
|
96
|
+
if [ -n "${POWER_PASSWORD:-}" ]; then curl -fsS --digest -u "admin:$POWER_PASSWORD" "$url"; else curl -fsS "$url"; fi ;;
|
|
97
|
+
tasmota)
|
|
98
|
+
curl -fsS "http://$target/cm?user=admin&password=${POWER_PASSWORD:-}&cmnd=Power%20$1" ;;
|
|
99
|
+
ha)
|
|
100
|
+
curl -fsS -X POST -H "Authorization: Bearer $HA_TOKEN" -H 'Content-Type: application/json' \
|
|
101
|
+
-d "{\"entity_id\":\"$target\"}" "$HA_URL/api/services/switch/turn_$1" ;;
|
|
102
|
+
apc)
|
|
103
|
+
snmpset -v1 -c "$PDU_COMMUNITY" "${target%:*}" "1.3.6.1.4.1.318.1.1.12.3.3.1.1.4.${target##*:}" \
|
|
104
|
+
i "$([ "$1" = on ] && echo 1 || echo 2)" ;;
|
|
105
|
+
*)
|
|
106
|
+
echo "power.sh: unknown device kind '$kind' in BENCH_POWER" >&2; exit 2 ;;
|
|
107
|
+
esac >/dev/null
|
|
108
|
+
echo "power: $BENCH_POWER $1"
|
|
109
|
+
}
|
|
110
|
+
|
|
111
|
+
case ${1:-} in
|
|
112
|
+
on|off) switch "$1" ;;
|
|
113
|
+
reset) switch off; sleep "${2:-1}"; switch on ;;
|
|
114
|
+
*) echo "usage: power.sh on|off|reset [off-seconds]" >&2; exit 2 ;;
|
|
115
|
+
esac
|
|
116
|
+
+code('Use it from the job').
|
|
117
|
+
thub run --type hw --label board:nucleo-f401re \
|
|
118
|
+
--download-file "$IMAGE_URL" --env GH_TOKEN --env POWER_PASSWORD \
|
|
119
|
+
--command 'git clone --depth 1 "https://x-access-token:$GH_TOKEN@github.com/yourorg/firmware-tests.git" src && cd src &&
|
|
120
|
+
ci/power.sh reset 2 && sleep 3 &&
|
|
121
|
+
st-flash --reset write "$THUB_DOWNLOAD_1" 0x08000000 && ./ci/test.sh' --wait
|
|
122
|
+
+code('Or download the script instead of cloning a repository').
|
|
123
|
+
thub run --type hw … --download-file https://artifactory.example.com/lab/tools/power.sh \
|
|
124
|
+
--command 'sh "$THUB_DOWNLOAD_1" reset 2 && ./run-tests.sh'
|
|
125
|
+
|
|
126
|
+
h3.h6 Switching it off again, also when the job is canceled
|
|
127
|
+
p.small.
|
|
128
|
+
When a job is canceled, times out or its Client is stopped, the Client sends #[code SIGTERM] to the job's
|
|
129
|
+
#[strong shell only], then #[code SIGKILL] 10 s later. To power the board off whatever happens:
|
|
130
|
+
ul.small
|
|
131
|
+
li Start the wrapper with #[code exec], so the wrapper is the shell that receives the signal.
|
|
132
|
+
li Run the tests in the background and #[code wait] for them. A shell runs a trap only after its current foreground command ends.
|
|
133
|
+
li Finish the cleanup within 10 s. A #[code SIGKILL] can't be trapped.
|
|
134
|
+
+code('ci/run.sh — cold-boot, flash, test, always power off').
|
|
135
|
+
#!/bin/sh
|
|
136
|
+
set -u
|
|
137
|
+
cd "$(dirname "$0")/.."
|
|
138
|
+
test_pid=
|
|
139
|
+
off() { ci/power.sh off || echo "power: could not switch $BENCH_POWER off" >&2; }
|
|
140
|
+
trap '[ -n "$test_pid" ] && kill "$test_pid" 2>/dev/null; off; exit 143' TERM INT
|
|
141
|
+
|
|
142
|
+
ci/power.sh reset 2 || exit 2
|
|
143
|
+
sleep 3
|
|
144
|
+
st-flash --reset write "$THUB_DOWNLOAD_1" 0x08000000 || { off; exit 2; }
|
|
145
|
+
./ci/test.sh "$@" & test_pid=$!
|
|
146
|
+
wait "$test_pid"; rc=$?
|
|
147
|
+
off
|
|
148
|
+
exit "$rc"
|
|
149
|
+
+code('Started with exec').
|
|
150
|
+
thub run --type hw … --command 'git clone … src && cd src && exec ci/run.sh' --wait
|
|
151
|
+
+note.
|
|
152
|
+
For a guarantee independent of the job, also set a timer on the device as a dead-man switch: Shelly's
|
|
153
|
+
#[code toggle_after] (#[code Switch.Set?id=0&on=true&toggle_after=3600] switches it off again after an hour) or Tasmota's
|
|
154
|
+
#[code PulseTime]. A failing #[code power.sh] step fails the job, since the command's exit code is the verdict.
|
|
@@ -0,0 +1,218 @@
|
|
|
1
|
+
+section('storage', 'Artifact storage and certificates: Artifactory, S3, Google Drive, (S)FTP', 'cloud-arrow-up')
|
|
2
|
+
p.
|
|
3
|
+
TestHub stores no files. A job's #[strong inputs] come from wherever their URLs point, and its #[strong outputs]
|
|
4
|
+
(reports, binaries, core dumps) go wherever its command uploads them. The patterns are the same for every kind of storage:
|
|
5
|
+
.table-responsive
|
|
6
|
+
table.table.table-sm.small.align-middle
|
|
7
|
+
thead
|
|
8
|
+
tr
|
|
9
|
+
th(style="width: 22%") Direction
|
|
10
|
+
th How
|
|
11
|
+
th(style="width: 30%") Credentials
|
|
12
|
+
tbody
|
|
13
|
+
tr
|
|
14
|
+
td Input, anonymous HTTP(S)
|
|
15
|
+
td #[code --download-file <url>]: fetched by the Client before the command (#[code $THUB_DOWNLOAD_1] …)
|
|
16
|
+
td None. A plain GET with no headers
|
|
17
|
+
tr
|
|
18
|
+
td Input, signed URL
|
|
19
|
+
td #[code --download-file "<pre-signed URL>"] (S3, GCS, Azure SAS, Artifactory signed URLs)
|
|
20
|
+
td In the URL. It shows on the job page, so keep its expiry short
|
|
21
|
+
tr
|
|
22
|
+
td Input, private
|
|
23
|
+
td Fetched in #[code --command] (#[code curl], #[code aws], #[code rclone], …)
|
|
24
|
+
td #[code --env NAME]: masked, dropped when the job ends
|
|
25
|
+
tr
|
|
26
|
+
td Input, client certificate
|
|
27
|
+
td Fetched in #[code --command] with #[code curl --cert] (#[code --download-file] can't present one)
|
|
28
|
+
td The certificate and key as #[code --env]
|
|
29
|
+
tr
|
|
30
|
+
td Output
|
|
31
|
+
td Uploaded in #[code --command], then listed in #[code $THUB_ARTIFACTS_FILE] so it shows on the job page
|
|
32
|
+
td #[code --env NAME]
|
|
33
|
+
ul.small
|
|
34
|
+
li The tools run on the #[strong Client host], which must reach the storage. Install #[code awscli], #[code rclone], … there, or run them in a container.
|
|
35
|
+
li Keep secrets off the command line (visible in #[code ps]): #[code curl -K -] reads them from standard input, and #[code aws] and #[code rclone] read them from the environment.
|
|
36
|
+
li Artifact links must be #[code http(s)]: an #[code s3://] or #[code sftp://] link is dropped from the list.
|
|
37
|
+
li Keep the verdict: run the tests, save #[code rc=$?], upload, then #[code exit $rc].
|
|
38
|
+
li Name outputs after the job (#[code …/$THUB_JOB_ID/…]) or a CI id, so a CI step can find them.
|
|
39
|
+
|
|
40
|
+
h3.h6 Helper: list what was published
|
|
41
|
+
p.small The examples below call #[code ci/artifacts.sh] from the tests repository, and assume the command has cloned it.
|
|
42
|
+
+code('ci/artifacts.sh').
|
|
43
|
+
#!/bin/sh
|
|
44
|
+
# ci/artifacts.sh name link [file] — append one entry to $THUB_ARTIFACTS_FILE (§7.3)
|
|
45
|
+
name=$1 link=$2 file=${3:-}
|
|
46
|
+
size=null; [ -n "$file" ] && size=$(wc -c < "$file" | tr -d ' ')
|
|
47
|
+
entry=$(printf '{"name":"%s","size":%s,"link":"%s","timestamp":%s}' "$name" "$size" "$link" "$(date +%s)")
|
|
48
|
+
if [ -s "$THUB_ARTIFACTS_FILE" ]; then # plain shell, not sed: signed URLs contain & and |
|
|
49
|
+
list=$(cat "$THUB_ARTIFACTS_FILE")
|
|
50
|
+
printf '%s,%s]\n' "${list%]}" "$entry" > "$THUB_ARTIFACTS_FILE"
|
|
51
|
+
else
|
|
52
|
+
printf '[%s]\n' "$entry" > "$THUB_ARTIFACTS_FILE"
|
|
53
|
+
fi
|
|
54
|
+
|
|
55
|
+
h3.h6 HTTPS certificates: private CAs and client certificates
|
|
56
|
+
p.small.
|
|
57
|
+
#[strong A private CA.] #[code --download-file] is fetched by the Client's Node.js process, which doesn't read the
|
|
58
|
+
system CA store, so the download fails with #[code SELF_SIGNED_CERT_IN_CHAIN]. Give the Client the CA with
|
|
59
|
+
#[code NODE_EXTRA_CA_CERTS], once per host, in a drop-in for the template unit, which applies to every instance.
|
|
60
|
+
#[code update-ca-certificates] covers #[code curl] and #[code git] in jobs, and Docker reads
|
|
61
|
+
#[code /etc/docker/certs.d/<registry>/ca.crt]. Never switch verification off in a job.
|
|
62
|
+
+code('Client host').
|
|
63
|
+
sudo cp lab-ca.crt /usr/local/share/ca-certificates/ && sudo update-ca-certificates # curl, git, wget, apt…
|
|
64
|
+
sudo systemctl edit thub-client@.service # every instance on this host
|
|
65
|
+
# [Service]
|
|
66
|
+
# Environment=NODE_EXTRA_CA_CERTS=/usr/local/share/ca-certificates/lab-ca.crt
|
|
67
|
+
sudo systemctl restart 'thub-client@*'
|
|
68
|
+
p.small.
|
|
69
|
+
#[strong A client certificate (mutual TLS).] #[code --download-file] can't present one. Fetch the file in
|
|
70
|
+
#[code --command] with #[code curl --cert], with the certificate and key passed as #[code --env] and written
|
|
71
|
+
(#[code umask 077]) to the job's directory, which is deleted with the job. For PKCS#12, use
|
|
72
|
+
#[code curl --cert-type P12 --cert file.p12:password]. One certificate for the whole lab can instead live on the host
|
|
73
|
+
(readable by the service user only), named in a drop-in: every job on that host can then use it.
|
|
74
|
+
+code('mTLS download in the command').
|
|
75
|
+
export CLIENT_CERT="$(cat thub-ci.crt)" CLIENT_KEY="$(cat thub-ci.key)" # PEM, from your CI's secret store
|
|
76
|
+
thub run --type hw --env CLIENT_CERT --env CLIENT_KEY \
|
|
77
|
+
--command 'umask 077 && printf "%s\n" "$CLIENT_CERT" > "$THUB_WORK_DIR/c.pem" && printf "%s\n" "$CLIENT_KEY" > "$THUB_WORK_DIR/k.pem" &&
|
|
78
|
+
curl -fsSL --cert "$THUB_WORK_DIR/c.pem" --key "$THUB_WORK_DIR/k.pem" -o app.bin https://artifactory.lab/fw-local/app/1.4.0/app.bin &&
|
|
79
|
+
st-flash --reset write app.bin 0x08000000 && ./ci/test.sh' --wait
|
|
80
|
+
|
|
81
|
+
h3.h6 JFrog Artifactory: service authentication
|
|
82
|
+
.table-responsive
|
|
83
|
+
table.table.table-sm.small.align-middle
|
|
84
|
+
thead
|
|
85
|
+
tr
|
|
86
|
+
th Credential
|
|
87
|
+
th Sent as
|
|
88
|
+
th Notes
|
|
89
|
+
tbody
|
|
90
|
+
tr
|
|
91
|
+
td #[strong Access token] (scoped, expiring)
|
|
92
|
+
td: code Authorization: Bearer <token>
|
|
93
|
+
td Recommended. A service user with read access to inputs and deploy access to outputs.
|
|
94
|
+
tr
|
|
95
|
+
td Reference token
|
|
96
|
+
td: code Authorization: Bearer <token>
|
|
97
|
+
td Short form of an access token, for CI variables with a length limit.
|
|
98
|
+
tr
|
|
99
|
+
td Identity token / password
|
|
100
|
+
td: code -u svc-thub:<token>
|
|
101
|
+
td For tools that only speak basic auth: #[code docker login], #[code pip].
|
|
102
|
+
tr
|
|
103
|
+
td API key
|
|
104
|
+
td: code X-JFrog-Art-Api
|
|
105
|
+
td Deprecated and turned off in current versions. Migrate.
|
|
106
|
+
tr
|
|
107
|
+
td Anonymous read
|
|
108
|
+
td —
|
|
109
|
+
td Public repositories only. Then #[code --download-file] is enough.
|
|
110
|
+
p.small.
|
|
111
|
+
Best: a #[strong short-lived token per run], minted in the CI step from a service token, so a leaked token expires on its own.
|
|
112
|
+
On GitHub Actions, JFrog's OIDC integration (#[code jfrog/setup-jfrog-cli] with an #[code oidc-provider-name]) needs no stored secret at all.
|
|
113
|
+
+code('CI step: a one-hour token for the job').
|
|
114
|
+
# CI step: ART_SERVICE_TOKEN is a CI secret for the service user svc-thub
|
|
115
|
+
export ART_TOKEN=$(curl -fsS -H "Authorization: Bearer $ART_SERVICE_TOKEN" -X POST \
|
|
116
|
+
-d scope=applied-permissions/user -d expires_in=3600 \
|
|
117
|
+
https://artifactory.example.com/access/api/v1/tokens | jq -r .access_token)
|
|
118
|
+
thub run --type hw --env ART_TOKEN --command '…' --wait
|
|
119
|
+
+code('curl: download, test, upload with checksum, list').
|
|
120
|
+
thub run --type hw --env ART_TOKEN \
|
|
121
|
+
--command 'ART=https://artifactory.example.com/artifactory
|
|
122
|
+
H() { printf "header = \"Authorization: Bearer %s\"\n" "$ART_TOKEN"; } # via curl -K -, so not visible in ps
|
|
123
|
+
H | curl -fsSL -K - -o app.bin "$ART/fw-local/app/1.4.0-42/app.bin" &&
|
|
124
|
+
st-flash --reset write app.bin 0x08000000 && ./ci/test.sh; rc=$?
|
|
125
|
+
url="$ART/qa-local/$THUB_JOB_ID/junit.xml"
|
|
126
|
+
H | curl -fsS -K - -H "X-Checksum-Sha256: $(sha256sum results/junit.xml | cut -d" " -f1)" -T results/junit.xml "$url" &&
|
|
127
|
+
sh ci/artifacts.sh junit.xml "$url" results/junit.xml
|
|
128
|
+
exit $rc' --wait
|
|
129
|
+
+code('JFrog CLI, configured from --env, nothing left in the home directory').
|
|
130
|
+
export JF_ACCESS_TOKEN="$ART_TOKEN"
|
|
131
|
+
thub run --type sw --env JF_URL=https://artifactory.example.com --env JF_ACCESS_TOKEN \
|
|
132
|
+
--command 'export JFROG_CLI_HOME_DIR="$THUB_WORK_DIR/.jfrog" CI=true
|
|
133
|
+
jf rt dl "fw-local/app/1.4.0-42/*.bin" downloads/ --flat && ./ci/test.sh; rc=$?
|
|
134
|
+
jf rt u "results/*.xml" "qa-local/$THUB_JOB_ID/" --flat
|
|
135
|
+
exit $rc' --wait
|
|
136
|
+
+code('Signed URL (Enterprise+): --download-file with nothing else').
|
|
137
|
+
IMAGE_URL=$(curl -fsS -H "Authorization: Bearer $ART_SERVICE_TOKEN" -H "Content-Type: application/json" -X POST \
|
|
138
|
+
-d '{"repo_path":"/fw-local/app/1.4.0-42/app.bin","valid_for_secs":3600}' \
|
|
139
|
+
https://artifactory.example.com/artifactory/api/signed/url)
|
|
140
|
+
thub run --type hw --download-file "$IMAGE_URL" --command 'st-flash --reset write "$THUB_DOWNLOAD_1" 0x08000000 && ./ci/test.sh' --wait
|
|
141
|
+
+code('The same token for Docker, PyPI and npm repositories').
|
|
142
|
+
# Docker registry (see §7.2): the token as the password
|
|
143
|
+
echo "$ART_TOKEN" | docker login artifactory.example.com -u svc-thub --password-stdin
|
|
144
|
+
docker pull artifactory.example.com/docker-local/team/test-runner:1.4
|
|
145
|
+
# PyPI remote/virtual repository
|
|
146
|
+
pip install -r requirements.txt --index-url "https://svc-thub:$ART_TOKEN@artifactory.example.com/artifactory/api/pypi/pypi/simple"
|
|
147
|
+
# npm: a project-local .npmrc in the work directory, not ~/.npmrc
|
|
148
|
+
printf '//artifactory.example.com/artifactory/api/npm/npm/:_authToken=%s\n' "$ART_TOKEN" > .npmrc
|
|
149
|
+
npm ci --registry https://artifactory.example.com/artifactory/api/npm/npm/
|
|
150
|
+
p.small.
|
|
151
|
+
Nexus raw repositories take #[code curl -u user:password --upload-file]. GitLab's generic packages take #[code curl -H "JOB-TOKEN: $CI_JOB_TOKEN" --upload-file].
|
|
152
|
+
|
|
153
|
+
h3.h6 AWS S3 (and MinIO, Ceph, Cloudflare R2, Wasabi)
|
|
154
|
+
p.small.
|
|
155
|
+
Pre-sign inputs in CI, so the job needs no AWS credentials. Pass short-lived session credentials from the CI step's role
|
|
156
|
+
(OIDC) for outputs. A pre-signed link lasts at most 7 days. For S3-compatible storage, add #[code --endpoint-url] to #[code aws].
|
|
157
|
+
+code('Inputs: pre-signed in the CI step').
|
|
158
|
+
# in the CI step, which has AWS access (e.g. GitHub/GitLab OIDC → an IAM role)
|
|
159
|
+
IMAGE_URL=$(aws s3 presign "s3://fw-builds/app/$SHA/app.bin" --expires-in 3600)
|
|
160
|
+
thub run --type hw --download-file "$IMAGE_URL" --command 'st-flash --reset write "$THUB_DOWNLOAD_1" 0x08000000 && ./ci/test.sh' --wait
|
|
161
|
+
+code('Outputs: aws s3 cp with session credentials').
|
|
162
|
+
thub run --type sw --env AWS_ACCESS_KEY_ID --env AWS_SECRET_ACCESS_KEY --env AWS_SESSION_TOKEN --env AWS_DEFAULT_REGION=eu-central-1 \
|
|
163
|
+
--command 'aws s3 cp "s3://fw-builds/app/1.4.0/app.elf" . && ./ci/test.sh; rc=$?
|
|
164
|
+
aws s3 cp --recursive results/ "s3://qa-results/$THUB_JOB_ID/" &&
|
|
165
|
+
sh ci/artifacts.sh report.html "$(aws s3 presign "s3://qa-results/$THUB_JOB_ID/report.html" --expires-in 604800)"
|
|
166
|
+
exit $rc' --wait
|
|
167
|
+
|
|
168
|
+
h3.h6 Google Drive
|
|
169
|
+
p.small.
|
|
170
|
+
Use #[a(href="https://rclone.org/drive/" target="_blank" rel="noopener") rclone] with a #[strong service account] that's a
|
|
171
|
+
member of a #[strong shared drive] (service accounts have no storage of their own). rclone takes its configuration from
|
|
172
|
+
environment variables. #[code rclone link] makes a file readable by anyone with the link; leave it out for private reports.
|
|
173
|
+
Public #[code uc?export=download] links work with #[code --download-file] only for small files.
|
|
174
|
+
+code('rclone with a service account').
|
|
175
|
+
# GDRIVE_SA: the service account's JSON key, as one line (a CI secret)
|
|
176
|
+
thub run --type sw --env GDRIVE_SA --env RCLONE_CONFIG_GD_TYPE=drive --env RCLONE_CONFIG_GD_SCOPE=drive \
|
|
177
|
+
--env RCLONE_CONFIG_GD_TEAM_DRIVE=0ABCdEfGhIjKlUk9PVA \
|
|
178
|
+
--command 'export RCLONE_CONFIG_GD_SERVICE_ACCOUNT_CREDENTIALS="$GDRIVE_SA"
|
|
179
|
+
rclone copyto "gd:firmware/1.4.0/app.bin" app.bin && ./ci/test.sh app.bin; rc=$?
|
|
180
|
+
rclone copy results/ "gd:qa/$THUB_JOB_ID/" &&
|
|
181
|
+
sh ci/artifacts.sh report.html "$(rclone link "gd:qa/$THUB_JOB_ID/report.html")"
|
|
182
|
+
exit $rc' --wait
|
|
183
|
+
+code('Or the Drive API with curl and an OAuth token from the CI step').
|
|
184
|
+
curl -fsSL -H "Authorization: Bearer $GDRIVE_TOKEN" -o app.bin "https://www.googleapis.com/drive/v3/files/$FILE_ID?alt=media&supportsAllDrives=true"
|
|
185
|
+
curl -fsS -H "Authorization: Bearer $GDRIVE_TOKEN" \
|
|
186
|
+
-F "metadata={\"name\":\"report.html\",\"parents\":[\"$FOLDER_ID\"]};type=application/json" -F "file=@results/report.html" \
|
|
187
|
+
"https://www.googleapis.com/upload/drive/v3/files?uploadType=multipart&supportsAllDrives=true"
|
|
188
|
+
|
|
189
|
+
h3.h6 FTP, FTPS and SFTP
|
|
190
|
+
p.small.
|
|
191
|
+
#[code curl] speaks all three. Prefer SFTP or FTPS (#[code --ssl-reqd]); plain FTP sends the password in clear text. Pin the
|
|
192
|
+
SFTP host key with #[code --hostpubsha256] (from #[code ssh-keyscan host | ssh-keygen -lf -], without #[code SHA256:]).
|
|
193
|
+
FTP and SFTP links aren't #[code http(s)], so list the server's HTTPS view, if it has one.
|
|
194
|
+
+code('SFTP download and upload (FTPS and key auth in the comments)').
|
|
195
|
+
thub run --type hw --env SFTP_USER --env SFTP_PASSWORD --env SFTP_HOSTKEY \
|
|
196
|
+
--command 'auth() { printf "user = \"%s:%s\"\n" "$SFTP_USER" "$SFTP_PASSWORD"; } # read by curl -K -, so not visible in ps
|
|
197
|
+
auth | curl -fsS -K - --hostpubsha256 "$SFTP_HOSTKEY" -o app.bin "sftp://files.lab/fw/1.4.0/app.bin" &&
|
|
198
|
+
st-flash --reset write app.bin 0x08000000 && ./ci/test.sh; rc=$?
|
|
199
|
+
auth | curl -fsS -K - --hostpubsha256 "$SFTP_HOSTKEY" --ftp-create-dirs -T results/junit.xml "sftp://files.lab/qa/$THUB_JOB_ID/"
|
|
200
|
+
exit $rc' --wait
|
|
201
|
+
# FTPS: auth | curl -fsS -K - --ssl-reqd --ftp-create-dirs -T results/junit.xml "ftp://files.lab/qa/$THUB_JOB_ID/"
|
|
202
|
+
# SSH key auth: printf '%s\n' "$SFTP_KEY" > "$THUB_WORK_DIR/.key" && chmod 600 "$THUB_WORK_DIR/.key" &&
|
|
203
|
+
# curl --key "$THUB_WORK_DIR/.key" -u "$SFTP_USER:" --hostpubsha256 "$SFTP_HOSTKEY" …
|
|
204
|
+
|
|
205
|
+
h3.h6 Any other HTTP storage: curl and your own authentication
|
|
206
|
+
+code('curl authentication patterns').
|
|
207
|
+
H() { printf 'header = "%s"\n' "$1"; } # headers via curl -K -, kept out of ps
|
|
208
|
+
H "Authorization: Bearer $API_TOKEN" | curl -fsSL -K - -o app.bin "$URL" # bearer token
|
|
209
|
+
H "X-API-Key: $API_KEY" | curl -fsSL -K - -o app.bin "$URL" # API-key header
|
|
210
|
+
printf 'user = "%s:%s"\n' "$USER_NAME" "$USER_PASS" | curl -fsSL -K - -o app.bin "$URL" # basic auth
|
|
211
|
+
curl -fsSL --netrc-file "$THUB_WORK_DIR/.netrc" -o app.bin "$URL" # a .netrc the command wrote from --env values
|
|
212
|
+
curl -fsSL --cert "$THUB_WORK_DIR/c.pem" --key "$THUB_WORK_DIR/k.pem" -o app.bin "$URL" # mutual TLS
|
|
213
|
+
# OAuth2 client credentials: get a token, then use it as a bearer token
|
|
214
|
+
TOKEN=$(curl -fsS -d grant_type=client_credentials -u "$CLIENT_ID:$CLIENT_SECRET" https://auth.example.com/oauth/token | jq -r .access_token)
|
|
215
|
+
# uploads: PUT (-T file), or a form POST (-F "file=@results/report.html")
|
|
216
|
+
p.small.mb-0.
|
|
217
|
+
Once the patterns settle, move them into #[code ci/fetch.sh] / #[code ci/publish.sh] in the tests repository, so every job's
|
|
218
|
+
#[code --command] stays one line. The CI examples in #[a(href="#ci") CI/CD integration] use this storage to bring JUnit reports back.
|
|
@@ -9,10 +9,18 @@
|
|
|
9
9
|
['The job is rejected with `422`', 'No registered Client can ever match its type, labels, group or client. Compare `thub resources` with your `--type` and `--label`s; your group (set on the dashboard, see `thub whoami`) must have a Client of that kind.'],
|
|
10
10
|
['The job stays `QUEUED`', 'Matching Clients are all `BUSY`, locked (`local`), in `MAINTENANCE` or offline, or the job is pinned with `--client`. It times out after `--timeout`.'],
|
|
11
11
|
['`ERROR` during prepare', 'A download failed; the job log names the URL and the reason.'],
|
|
12
|
+
['`Download failed … (SELF_SIGNED_CERT_IN_CHAIN)` or `UNABLE_TO_GET_ISSUER_CERT_LOCALLY`', 'The server\'s certificate is from a CA the Client doesn\'t trust. Node.js ignores the system CA store: add `Environment=NODE_EXTRA_CA_CERTS=/path/ca.crt` with `sudo systemctl edit thub-client@.service` and restart the Clients. A download that needs a client certificate can\'t use `--download-file` at all: fetch it with `curl --cert` in `--command` (Artifact storage and certificates).'],
|
|
13
|
+
['`Host key verification failed` on a git clone', 'The job\'s `known_hosts` has no key for that server (or a different one). Pass the output of `ssh-keyscan <host>` as `--env GIT_KNOWN_HOSTS` and use it with `UserKnownHostsFile` (Using git). `Permission denied (publickey)` means the host is fine, but the deploy key isn\'t on that repository.'],
|
|
14
|
+
['`docker login`: unauthorized', 'Check the user the registry expects: `AWS` for ECR, `oauth2accesstoken` for Google, `gitlab-ci-token` with `CI_JOB_TOKEN` for GitLab (Using Docker). ECR and Google tokens expire after 12 h and 1 h, so mint them in the CI step that submits the job.'],
|
|
12
15
|
['A git clone or docker pull in the command fails', 'They run in your `--command`, with the credentials you pass as `--env` (see Using git / Using Docker). `git` or `docker` must be installed on that Client host; the Client itself needs neither. Try the command with `--dry-run` to see its environment.'],
|
|
13
16
|
['`--git-repo` / `--docker-image`: unknown option, or the job is refused', 'They were removed: clone the repository and run containers in `--command`, passing tokens with `--env`. An older Agent that still sends them is refused with that message; update it (`thub self-update`).'],
|
|
14
17
|
['Permission denied on `/dev/ttyUSB*` or `docker.sock`', 'Docker or the device groups were added after the Client was installed. Re-run `sudo npm i -g @andrian.yablonskyy/thub-client` so the unit gets the `dialout`/`plugdev`/`docker` groups, then restart.'],
|
|
15
18
|
['A configured device is shown as missing', 'Check the `devpath` with `udevadm info -a -n <dev>`, then `thub-client udev --print`; restart the Client to regenerate the udev rules.'],
|
|
19
|
+
['`thub power` is refused', 'It says why: only the job\'s owner may switch power (`403`, even for an admin; `thub whoami` shows whose key you use); the job must be `PREPARING` or `RUNNING`; the Client must list ports in `hw-devices.usbPower.ports` and have `uhubctl` installed (restart it after installing, so it reports it). `--port` is a position in that list, not the hub\'s port number.'],
|
|
20
|
+
['USB power: `No compatible devices detected!`', 'Either the hub can\'t switch per-port power (`sudo uhubctl` doesn\'t list it), or the Client\'s user may not switch it: install its udev rules with `sudo thub-client --config <path> udev` and check its service has the `plugdev` group (reinstall the Client if `plugdev` was created after it).'],
|
|
21
|
+
['USB power is off, but the board stays on (or half-boots after a reset)', 'Something else powers it: a second USB cable (ST-Link and the board\'s own USB), a debugger or an external supply. Every cable that powers it needs a switched port in `usbPower.ports`. Some hubs also feed power back from upstream. A board that half-boots needs a longer delay: `--delay 3`.'],
|
|
22
|
+
['A job can\'t reach its smart socket or PDU', 'The job\'s command runs on the Client host, so the device must be reachable from there: `curl -sI http://<device>/` as the Client\'s user. Tools such as `curl` or `snmpset` must be installed on that host. `BENCH_POWER is not set`: add the instance\'s systemd drop-in and restart it (`systemctl show thub-client@<instance> -p Environment`).'],
|
|
23
|
+
['A board on a socket or PDU stays on after a canceled job', 'On cancel the Client sends `SIGTERM` to the job\'s shell only, then `SIGKILL` after 10 s. Start the wrapper with `exec`, run the tests in the background with `wait`, trap `TERM`, and finish within 10 s. Add a timer on the device (Shelly `toggle_after`, Tasmota `PulseTime`) as a backstop.'],
|
|
16
24
|
['Job `LOST`', 'The Client missed 3 heartbeats (network, reboot, crash). With `requeueOnLost` the job is retried once on another Client.'],
|
|
17
25
|
['A job can\'t write to `~` (read-only file system)', 'Under systemd the Client can only write to its state directory. Write to `$THUB_WORK_DIR`, and set `DOCKER_CONFIG` there for `docker login`.'],
|
|
18
26
|
['Live logs don\'t stream through the proxy', 'Turn off response buffering for the Coordinator (nginx: `proxy_buffering off;`, Caddy: `flush_interval -1`).'],
|