homebridge-roborock-matter 3.11.0 → 3.11.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +40 -14
- package/README.md +3 -2
- package/package.json +1 -1
- package/roborockLib/lib/localConnector.js +5 -0
- package/roborockLib/lib/messageQueueHandler.js +10 -0
- package/roborockLib/roborockAPI.js +160 -3
package/CHANGELOG.md
CHANGED
|
@@ -1,8 +1,34 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 3.11.2
|
|
4
|
+
|
|
5
|
+
**"Attempt 12 this run" counted every run since Homebridge started. So did "failed 10 times in a row", and the once-per-run explanation of why no room could be named was only ever printed on the first run of the process.**
|
|
6
|
+
|
|
7
|
+
Caught on my own a70. It ran a two-room clean from Apple Home, the cloud map channel was timing out that morning, and `get_map_v1` failed all ten times it was asked — ten guaranteed-to-fail cloud requests, 10 seconds of timeout each, spread across one ten-minute clean, and not one room named.
|
|
8
|
+
|
|
9
|
+
The clear that runs at every run boundary exists so nothing leaks into the next run. It dropped the cached room and nothing else. Every counter behind the log lines survived for the lifetime of the process, so three lines were telling you about a window that was not the one they named. A run that fails every attempt is exactly the run that leaves the counters high, and that run never had a cached room to clear in the first place — which is why the leak survived a release that went looking for the same class of thing. 3.11.0 stopped placeholder poses from inflating the miss count; it did not make the count per-run. It is per-run now.
|
|
10
|
+
|
|
11
|
+
**The failing fetch is also no longer retried at live-display cadence for the whole run.** After 2 failures in a row the gap doubles with each further failure, capped at 5 minutes, and drops straight back to the live cadence the moment one succeeds. The first two failures are deliberately not slowed: a single lost frame on a healthy channel must not make a working live room sluggish, the same rule 3.11.1's local-mute limit follows. A streak long enough to have been slowed now says so when it ends, because "failed N times in a row" at warn level had no counterpart and a channel that recovered left the log's last word saying it was broken.
|
|
12
|
+
|
|
13
|
+
Both protocol paths are covered — the classic `get_map_v1` fetch and the Q7/B01 SCMap fetch had the same two defects in the same shape. The test enumerates the rule across every live-room state the plugin keeps rather than the two call sites that happened to be found, so a third path added later fails the test instead of leaking quietly.
|
|
14
|
+
|
|
15
|
+
Not addressed, because there is no measurement to justify it: why the cloud map channel timed out at all. `get_map_v1` has the default 10-second timeout and had been resolving every position on this same robot nine days earlier, so a longer timeout would be a guess. The morning also carried an unrelated cloud `get_prop` timeout on the same robot, which points at the account's cloud rather than at the map request.
|
|
16
|
+
|
|
17
|
+
## 3.11.1
|
|
18
|
+
|
|
19
|
+
**The LAN port was open, the robot never answered on it, and the plugin kept asking for the life of the process — 10 seconds thrown away on every poll and every command.**
|
|
20
|
+
|
|
21
|
+
The reporter of [#8](https://github.com/mathiashornbek/homebridge-roborock-matter/issues/8) runs Homebridge on a NAS in one VLAN and his Saros 10 in another. Port 58867 is reachable across that boundary, so the local client completes its TCP handshake and records `Local connect state: true` — and then every single request dies of silence 10 seconds later. `get_prop` on both startups, `app_segment_clean`, `app_pause`, `app_start`: all of them, every time.
|
|
22
|
+
|
|
23
|
+
The plugin only ever gave up on the LAN when the _connect_ failed. A socket that connected and then answered nothing was retried forever, which is why he saw a `get_prop` timeout at every restart and why his commands were slow before they worked at all. A successful handshake proves the port is reachable. It does not prove the robot is listening.
|
|
24
|
+
|
|
25
|
+
3 local timeouts in a row on a socket that still reports itself connected now write the LAN off for that robot and use the cloud instead, with 1 log line that says the port is open and the robot is not replying — the distinction that decides whether there is any point rewriting a firewall rule. Any local reply resets the count, so a single lost frame on a healthy network changes nothing; permanently exiling a robot to the cloud over one dropped packet would be worse than the bug being fixed.
|
|
26
|
+
|
|
27
|
+
The diagnostics report names this case separately from a failed connect, because those two look identical in a log and lead to opposite conclusions.
|
|
28
|
+
|
|
3
29
|
## 3.11.0
|
|
4
30
|
|
|
5
|
-
**A Q7 said it was between rooms
|
|
31
|
+
**A Q7 said it was between rooms 226 times during one clean. It was in the bedroom the whole time.**
|
|
6
32
|
|
|
7
33
|
I caught a run on my own robots and pulled the numbers rather than the impressions. One Q7, 47 minutes, 227 live-room fetches. 226 of them placed it at cell 22280,22100 — the same cell every time — while the room outlines on that map span 38 to 293. The pose behind it was exactly (1100, 1100), which is the same constant two other people's Q7s reported back in August. The remaining fetches resolved Stue, then Gang, then Soveværelse, in the order the robot actually moved.
|
|
8
34
|
|
|
@@ -12,7 +38,7 @@ A position further outside the map than the map is wide is now recognised for wh
|
|
|
12
38
|
|
|
13
39
|
Two things fall out of it. The room still updates on a Q7, on the fetches that carry a true position; nothing about the tile changes. And a resolved room now prints the cell it resolved at, so a working position and a failing one can be compared in one log instead of across two field sessions — which is what this one cost.
|
|
14
40
|
|
|
15
|
-
**Also:
|
|
41
|
+
**Also: 50 log lines a minute, per robot, that could never mean anything.** With debug on, every known status attribute produced `Skipping known get_status attribute without a Homebridge state object` on every poll. That check dates from this library's ioBroker origins, where the object it looks for exists; under Homebridge it never does, so the branch fired for every attribute forever and reported only that the plugin is not ioBroker. On my own server the log ring had shrunk to 90 minutes — the window you need when something real goes wrong. It is gone. An attribute nobody has mapped yet is still named once with its value, which is the half that carries information.
|
|
16
42
|
|
|
17
43
|
## 3.10.2
|
|
18
44
|
|
|
@@ -44,9 +70,9 @@ The cloud timeouts underneath this are not the plugin's to fix and are not new;
|
|
|
44
70
|
|
|
45
71
|
Every robot gets a forced full Matter write every 60 seconds. Its evidence line is written only when the rendered line would read differently from the last one — deliberately, since 3.10.0, so an idle robot does not fill the log with identical minutes.
|
|
46
72
|
|
|
47
|
-
The reporter of #7 was asked to check whether those lines were still appearing while his Apple Home tile was dead. It is the cheapest way to separate "the plugin stopped" from "the Matter session died underneath a healthy plugin". His robot was docked at 100 %, so every line rendered identically, so the last one was written at startup and none followed. He looked, correctly reported
|
|
73
|
+
The reporter of #7 was asked to check whether those lines were still appearing while his Apple Home tile was dead. It is the cheapest way to separate "the plugin stopped" from "the Matter session died underneath a healthy plugin". His robot was docked at 100 %, so every line rendered identically, so the last one was written at startup and none followed. He looked, correctly reported 11 minutes of nothing, and the answer was worth nothing — absence was the deduplication working, not evidence about whether anything ran. The question was unanswerable from the log at any level, which for a liveness signal is the wrong outcome.
|
|
48
74
|
|
|
49
|
-
A suppressed publish is now recorded at debug, naming the robot, the values and what triggered it. The info log stays exactly as quiet as before, and with plugin debug on,
|
|
75
|
+
A suppressed publish is now recorded at debug, naming the robot, the values and what triggered it. The info log stays exactly as quiet as before, and with plugin debug on, 11 minutes of a docked robot leaves 11 traces — so a gap in them means something.
|
|
50
76
|
|
|
51
77
|
## 3.10.0
|
|
52
78
|
|
|
@@ -112,7 +138,7 @@ It was not the cloud. Measured over 30,224 log lines covering 49 restarts: 92 re
|
|
|
112
138
|
|
|
113
139
|
A flaky connection does not look like that. It gives a varying number of attempts at varying times. One attempt, every time, only at startup, only on the cloud-only protocol is a race — and it was an ordering mistake in the startup sequence. The dedicated Q7 status loop was started at the end of device creation, and it polls immediately; the sequence did not wait for the MQTT session until after device creation had returned. A Q7 request is cloud-only by construction, so that first poll was rejected before anything reached the wire. The wait was already there, with a comment explaining this exact hazard for the two calls after it. The loop start had simply slipped in front of it.
|
|
114
140
|
|
|
115
|
-
The same event explains the other half, which had been observed
|
|
141
|
+
The same event explains the other half, which had been observed 14 times and never connected to it: for about 27 seconds after every restart, a Q7's tile in Apple Home showed `battery=100%, operationalState=0, runMode=0` — the snapshot taken at registration rather than the robot. That window is not a separate phenomenon. A refused attempt still stamps the request throttle, so the 15-second tick that followed fell inside the 25-second idle gap and was dropped, and the robot's real status did not arrive until the tick at 30 seconds. Measured median: 31 seconds.
|
|
116
142
|
|
|
117
143
|
Two changes, because the ordering fix alone leaves the hazard reachable — the wait resolves on a 10-second timeout whether or not the broker came up:
|
|
118
144
|
|
|
@@ -222,7 +248,7 @@ Also on that page: saves report failure instead of looking like they worked, the
|
|
|
222
248
|
|
|
223
249
|
**The log.** Two lines were removed as duplicates: the poll-profile notice was keyed per robot while its text is per model, so two robots of one model printed the same sentence twice, naming neither; and every room change was announced by both the library and the Matter layer with the same prefix. `Service started` was printed on the failure path — the `getHomeDetail` catch falls through to the same callback — directly under the stack trace saying it had failed; it now says what actually happened. That stack trace is gone too: a Roborock outage or a DNS blip is a warning with a sentence, not an error with a Node stack. `Starting adapter. This might take a few minutes` (it takes one second) and `Lets go!!!!!!!` are gone with the rest of the ioBroker vocabulary, `Adapter not inited. Command not executed.` now names the robot and says to try again in a few seconds, and a robot going offline is a warning that says what to check — with the matching "back online" line uncommented after who knows how long.
|
|
224
250
|
|
|
225
|
-
**
|
|
251
|
+
**14 more log lines were printing a raw 22-character duid to users.** `log-lines-name-the-robot` only inspected template literals written inside the logging call, so anything built into a variable or an `Error` first was invisible — it was checking 39 of 59 calls in one file alone. It now follows the three laundering channels as well, and everything it found is fixed.
|
|
226
252
|
|
|
227
253
|
**One resource leak.** `localConnector.js` opened its UDP discovery socket at module load, so requiring the file bound a socket a cloud-only install never uses, a second discovery pass attached a second set of handlers to it, and the first pass's `close()` left it unbindable for the next. It is now created per run and closed once. That also removes the "A worker process has failed to exit gracefully" warning the suite has printed for months, which was masking any real leak.
|
|
228
254
|
|
|
@@ -244,7 +270,7 @@ Verified red against 3.5.3: 3 of 29 fail, exactly the disk-read and fallback rul
|
|
|
244
270
|
|
|
245
271
|
3.5.0 mentioned pairing in one sentence, and the sentence was wrong: it assumed the bridge needed pairing, not enabling.
|
|
246
272
|
|
|
247
|
-
The startup log now answers which of three situations you are in, once per start, and warns rather than informs when HAP is off — an info line about a feature that cannot work reads like the
|
|
273
|
+
The startup log now answers which of three situations you are in, once per start, and warns rather than informs when HAP is off — an info line about a feature that cannot work reads like the 90 other info lines a start produces. The settings page shows the steps under the toggle when the feature is on, and the README and the setting's own description carry the same three: **Plugins → homebridge-roborock-matter → ⋮ → Child Bridge Config**, check **Enable HAP**, restart, then **Connect to HomeKit** on that screen and scan that QR code. All four surfaces name the two codes that look right and are not — the main Homebridge code, and the robot's Matter code, which covers the vacuum only.
|
|
248
274
|
|
|
249
275
|
`__tests__/the-switches-say-which-qr-code-to-scan.test.js` enumerates the rule over the surfaces, because the original failure was that only one surface mentioned pairing at all. Matching ignores markup, so `<strong>` and `**bold**` count as the same instruction. Verified red against 3.5.2: 26 of 26 fail.
|
|
250
276
|
|
|
@@ -393,14 +419,14 @@ The prep sequence sends up to three commands one after another, each with a two-
|
|
|
393
419
|
|
|
394
420
|
**Nothing in this release changes what the plugin does. It removes things that were never doing anything, and two of them were actively lying.**
|
|
395
421
|
|
|
396
|
-
- **Ten of the
|
|
422
|
+
- **Ten of the 11 shipped languages could never load.** `this.language` is only ever set from `options.language`; the sole production construction site passes none, the UI server hardcodes `"en"`, and no setting exposes the choice. So de, es, fr, it, nl, pl, pt, ru, uk and zh-cn — 78 KB of translations — were installed on every user's disk and read by nobody. They are gone. A test now enumerates the rule rather than the ten filenames: **a locale that ships must be selectable**, so adding one back fails until there is actually a way to pick it.
|
|
397
423
|
- **The README claimed 463 automated tests in one paragraph and 263 in another.** Both were wrong. Two hand-written numbers describing one fact will drift apart and neither gets corrected, because nothing checks them. There is one number now, and a test checks it against what the suite actually declares. It deliberately does not pin an exact figure — `test.each` expands at runtime and no static reader can know by how much — it pins the two things that went wrong: state it once, and keep it in a defensible band.
|
|
398
424
|
- **The publish log line still rendered `fault=…` from an attribute withdrawn in 3.4.1.** The branch was unreachable, and worse, it read as evidence the feature still existed. Removing it settles a real contradiction: `matter-fault-reporting.test.js` pinned that `operationalError` is never published, while `matter-publish-line-logs-every-change.test.js` hand-built one and asserted it rendered. Two tests disagreeing about whether a feature exists is worse than either answer.
|
|
399
425
|
- **An orphaned ioBroker map viewer and a MITM sniffing script** (`roborockLib/lib/map/`, `roborockLib/lib/sniffing/`) were excluded from the npm package rather than deleted — which is exactly how they survived unreviewed for so long. Ignored by the package, invisible in review, referenced by nothing.
|
|
400
426
|
- **Ten functions whose definition was their only occurrence in the entire tree** are gone: `getHomeID`, `decodeSniffedMessage`, `getConnector`, `updateDataExtraData`, `setupBasicObjects`, `getCleanSummary`, `resolve102Message`, `resolve301Message`, `BytesToInt`, `getErrorCodeDescription`, plus the unused `B01_REQUEST_DPS`/`B01_RESPONSE_DPS` constants and three exports nothing imported. `resolveLiveRoomId` went too — a one-line wrapper over `describeLiveRoomResolution` with no production callers, kept alive only by tests. Two ways to ask the same question is how one of them drifts.
|
|
401
427
|
- **`errorCodes` was NOT removed**, though a first pass called it dead. `deviceFeatures.js` still uses the table for its `error_code` state mapping. Worth recording: the check that catches this is grepping the whole tree, not reasoning about one file.
|
|
402
428
|
- **Two user-facing claims were false.** The `preferCloudForMatterCommands` setting promised to keep "the legacy HomeKit accessories unchanged", and a startup log line told users "The existing HomeKit accessory will continue to work." This fork removed every HAP accessory by design — there is nothing to fall back to. The log line now says what to actually do: enable Matter for the bridge.
|
|
403
|
-
- **ROADMAP.md was eight releases stale**, still titled for a different package, pointing at an `AGENTS.md` that has never existed in the tree, and listing HomeKit controls as delivered features
|
|
429
|
+
- **ROADMAP.md was eight releases stale**, still titled for a different package, pointing at an `AGENTS.md` that has never existed in the tree, and listing HomeKit controls as delivered features 13 lines above its own note that all HAP accessories were removed. Rewritten, with the pre-Matter-only entries labelled rather than deleted so the history stays readable.
|
|
404
430
|
|
|
405
431
|
## 3.4.12
|
|
406
432
|
|
|
@@ -413,7 +439,7 @@ The fix that added the two fields was made one level down from the gatekeeper, w
|
|
|
413
439
|
|
|
414
440
|
## 3.4.11
|
|
415
441
|
|
|
416
|
-
**Two docked Q7s flipped their Apple Home clean mode to "Vacuum" and back every
|
|
442
|
+
**Two docked Q7s flipped their Apple Home clean mode to "Vacuum" and back every 90 seconds, and the plugin was reporting a level it had never measured.** Caught in a log from a plugin author's own robots on 3.4.10: every battery tick produced a pair of publishes about a second apart, the first saying `cleanMode=0` and the second saying `cleanMode=6` — ten pairs in 14 minutes, on both robots, at the same battery value.
|
|
417
443
|
|
|
418
444
|
Mode 6 is "Max Vacuum", the level the robot is actually set to. Mode 0 is plain "Vacuum", a level nobody selected. When suction-level clean modes are announced (`enableFanPowerCleanModes`), the reported mode is derived from the robot's live fan power — but the derivation had no answer for "the fan power cannot be read right now". It fell through to the last Matter selection, which defaults to plain Vacuum, so a momentary gap in the reading was published as a definite statement about the robot's suction level.
|
|
419
445
|
|
|
@@ -425,7 +451,7 @@ This is the same class of defect as 3.4.6 and 3.4.7: reporting a value derived f
|
|
|
425
451
|
|
|
426
452
|
## 3.4.10
|
|
427
453
|
|
|
428
|
-
**The Q7 position that never resolved to a room is not a position at all.** 3.4.9 asked the two Q7s to report the range their room outlines occupy, and they answered: Garage sat in a map spanning cells 52–171 by 43–187, 1. Sal in one spanning 38–293 by 90–227. Back-computing through each map's own origin and resolution gives the same coordinate for both — exactly (1100.0, 1100.0), on two robots, two maps, and
|
|
454
|
+
**The Q7 position that never resolved to a room is not a position at all.** 3.4.9 asked the two Q7s to report the range their room outlines occupy, and they answered: Garage sat in a map spanning cells 52–171 by 43–187, 1. Sal in one spanning 38–293 by 90–227. Back-computing through each map's own origin and resolution gives the same coordinate for both — exactly (1100.0, 1100.0), on two robots, two maps, and 12 minutes of active cleaning. A number that identical is arithmetic, not a place a robot stood, which means live-room tracking on these models has never worked from that field.
|
|
429
455
|
|
|
430
456
|
- **The miss line now surveys the payload rather than asserting anything about it.** It prints the size of every top-level field and every scalar inside the small ones, keyed by field path. Two consecutive lines are then a diff: the value that changed while the robot was driving is the position, and the submessage that grew is the trail behind it.
|
|
431
457
|
- **Varints are surveyed, not just floats.** The pose message carries an `update` flag alongside its coordinates, so a float-only dump would have printed two plausible-looking numbers and hidden the field saying they were stale.
|
|
@@ -488,7 +514,7 @@ On a v1 robot the difference between "Vacuum" and "Vacuum and mop" **is** the wa
|
|
|
488
514
|
- **A field nobody has seen before still gets through.** If a robot starts sending something new after hours of uptime, it is reported on its own — quietening the noise must not also hide the signal, since these lines are the raw material for a model profile.
|
|
489
515
|
- **The line is now usable in a model report.** It names the robot and model, lists every field with its value, and says plainly that control, battery, rooms and state do not depend on them. Object values are serialised instead of arriving as `[object Object]`, which is what `cleaning_info` looked like in #8 — the one field where the shape was the interesting part.
|
|
490
516
|
- **Two startup tests no longer assert on the clock.** Both checked that per-robot probes run concurrently by timing them against a 180 ms budget — and on a quiet machine they finish in ~65 ms, so the assertion could only ever fail for a reason it was not testing. A scheduling hiccup was enough to fail a build with the concurrency perfectly correct. The check was also redundant: serialized probes give a peak concurrency of 1, which the neighbouring assertion already catches exactly. They now assert the property directly — every probe started before the first one finished — which holds on any machine under any load. Same defect 3.4.2 removed from the B01 full-chain simulation.
|
|
491
|
-
- **The log-naming rule from 3.3.2 was itself only half enumerated.** It listed three files by hand, and the files it left out held
|
|
517
|
+
- **The log-naming rule from 3.3.2 was itself only half enumerated.** It listed three files by hand, and the files it left out held 12 log lines still printing a bare 22-character duid — including `Device <duid> is offline.`, which is exactly the line someone quotes when asking why a robot dropped out. A hand-written file list is the same mistake as a hand-written line list, one level up. The rule now discovers the file list from the source tree, so a new file is covered the moment it exists, and all 12 lines now name the robot.
|
|
492
518
|
|
|
493
519
|
## 3.4.2
|
|
494
520
|
|
|
@@ -521,7 +547,7 @@ So the feature cost people their tile and never once delivered the thing it prom
|
|
|
521
547
|
|
|
522
548
|
## 3.3.2
|
|
523
549
|
|
|
524
|
-
- **Finished a job 3.3.1 only half did.** That release converted the B01 log lines to use the robot's name and missed
|
|
550
|
+
- **Finished a job 3.3.1 only half did.** That release converted the B01 log lines to use the robot's name and missed 11 others, including the live-room line for classic S/Q-series robots and the battery resync line — which then appeared in a field log directly above a line that did use the name: `Battery resync for 3tELc5hUekaTlOJEW3YetI` followed by `Matter publish for Garage`. Every user-visible log line now names the robot, and a test enumerates the rule rather than the instances, so the next line written cannot reintroduce it.
|
|
525
551
|
|
|
526
552
|
## 3.3.1
|
|
527
553
|
|
|
@@ -543,7 +569,7 @@ A robot that has stopped because it is wedged under the sofa has always looked e
|
|
|
543
569
|
- **The Error state was never gated for a reason.** `ERROR` (3) is a member of even the basic advertised operational state list, so publishing it was always legal — it was being rewritten to `STOPPED` alongside the states that genuinely did need a gate. That is why no released version has ever shown a Roborock fault in Apple Home.
|
|
544
570
|
- **A detached water tank or mop pad is not a fault.** Both are the normal, correct configuration for a vacuum-only run, so reporting them would leave a permanent warning on every dry robot's tile. They are read, and deliberately ignored.
|
|
545
571
|
- **The fault detail can never cost you the tile.** `operationalError` travels in the same cluster payload as the operational state, so a Matter build that refuses the attribute would otherwise freeze Cleaning/Docked along with it — the reason the explicit write was removed back in 1.4.61. If the write is rejected, the plugin immediately re-publishes without it, logs a warning naming the reason, and stops sending it for the rest of the session. An endpoint that is merely still starting up keeps its normal retry and does not disable the feature.
|
|
546
|
-
- **Diagnostics no longer truncate away the answer.** A Roborock status payload runs to about
|
|
572
|
+
- **Diagnostics no longer truncate away the answer.** A Roborock status payload runs to about 50 fields and the export kept the first 30 — which are largely housekeeping, while the 20 it dropped included `dock_error_status`, the single field a question about the dock's water tanks turns on. The fields that matter for a fault report are now always kept, however far down the payload they sit, with the size cap otherwise unchanged and secret redaction untouched.
|
|
547
573
|
|
|
548
574
|
## 3.2.0
|
|
549
575
|
|
package/README.md
CHANGED
|
@@ -37,7 +37,7 @@ This is the most feature-packed, most thoroughly engineered Roborock plugin for
|
|
|
37
37
|
- 📍 **See where it's cleaning — live.** Apple Home shows _"Cleaning — Kitchen"_ with the room the robot is actually inside, updating as it moves from room to room. Works even for cleans started from the robot's button or the Roborock app. No other Homebridge plugin does this.
|
|
38
38
|
- 🧭 **One robot, one tile — and as many robots as you own.** Sign in once and your whole fleet comes along: every vacuum on your account appears as its own clean, native accessory in Apple Home. No clutter of fake fans and helper switches, and rooms appear with the names you gave them in the Roborock app.
|
|
39
39
|
- ⚡ **Fast and reliable.** Commands go directly to the robot over your own network whenever possible, with the Roborock cloud as automatic backup — and built-in diagnostics in the settings if you ever want to look under the hood.
|
|
40
|
-
- 🛡️ **Verified by Homebridge.** Reviewed and endorsed by the Homebridge team.
|
|
40
|
+
- 🛡️ **Verified by Homebridge.** Reviewed and endorsed by the Homebridge team. 1151 automated tests, zero known vulnerabilities, no analytics, and a startup designed to never crash your Homebridge — even when your Wi-Fi or the Roborock cloud has a bad day.
|
|
41
41
|
|
|
42
42
|
## Features
|
|
43
43
|
|
|
@@ -211,13 +211,14 @@ The complete path — robot → plugin → Homebridge → matter.js store — wa
|
|
|
211
211
|
|
|
212
212
|
- **Diagnostics first:** the plugin settings include per-device connection state, the last cloud/local transport used, a live **Test Local Connection** probe, and a **redacted diagnostics report** you can paste straight into a GitHub issue.
|
|
213
213
|
- **Robot shows "Updating…" in Apple Home:** remove the robot from Apple Home and pair it again — a stale controller cache from an earlier pairing is the usual cause (tracked upstream in homebridge/homebridge#3951).
|
|
214
|
+
- **Robot shows "No Response" on one Apple device while another shows it fine:** update to **iOS 26.6.1**. On earlier versions the controller could declare its own subscription invalid (Matter status `0x7D INVALID_SUBSCRIPTION`) and then never subscribe again, so one controller kept rendering the accessory while another had no live subscription to it — the same tile, dead on the phone and alive on the Mac at the same moment. The bridge does the right thing in that situation (it drops the dead subscription and re-announces over mDNS), but no Matter device can force a controller to subscribe, so there is nothing to fix on this side. Confirmed fixed on 26.6.1 by the reporter in [#7](https://github.com/mathiashornbek/homebridge-roborock-matter/issues/7). On an older version, "Add Accessory → More Options → Cancel" revives the tile for about a minute, because it forces the controller to re-resolve and briefly re-subscribe.
|
|
214
215
|
- **Rooms missing for a Q7/B01 robot:** wait for the `B01 rooms for ...` log line, then re-pair once so the Service Area cluster is announced with room data.
|
|
215
216
|
- **Debug logging needs two switches, not one:** the plugin's own **Debug Mode** only decides whether it _calls_ the debug logger — Homebridge decides whether anything is _printed_, and it suppresses plugin debug output unless Homebridge itself runs with `-D`. Turn on **Homebridge Settings → Homebridge Debug Mode** as well, or the log will look exactly the same as before.
|
|
216
217
|
- **Startup without network:** the plugin retries the Roborock cloud with increasing backoff (up to 10 attempts) and never crash-loops Homebridge; wrong credentials stop cleanly with a clear log message.
|
|
217
218
|
|
|
218
219
|
## Contributing
|
|
219
220
|
|
|
220
|
-
Model reports, diagnostics exports, and pull requests are very welcome. The codebase ships with
|
|
221
|
+
Model reports, diagnostics exports, and pull requests are very welcome. The codebase ships with 1151 tests (protocol fixtures verified against the [python-roborock](https://github.com/Python-roborock/python-roborock) reference), strict TypeScript checking, and CI across Node 22/24 × Homebridge 1.11/2.x — `npm test` before you push and you're set.
|
|
221
222
|
|
|
222
223
|
## Support the project
|
|
223
224
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "homebridge-roborock-matter",
|
|
3
|
-
"version": "3.11.
|
|
3
|
+
"version": "3.11.2",
|
|
4
4
|
"description": "The most complete Roborock plugin for Apple Home. Supports the entire Roborock lineup — from the classic S-series to the new 2025 Q7 series that no other plugin can control. Sign in with your Roborock account and get native start/stop, room cleaning, suction levels, battery, and live 'cleaning in the kitchen' room tracking. Verified by Homebridge.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"author": {
|
|
@@ -552,6 +552,11 @@ class localConnector {
|
|
|
552
552
|
const { resolve, timeout } = this.adapter.pendingRequests.get(id);
|
|
553
553
|
this.adapter.clearTimeout(timeout);
|
|
554
554
|
this.adapter.pendingRequests.delete(id);
|
|
555
|
+
// Proof that this socket is not mute, so any run of timeouts counted
|
|
556
|
+
// against it starts over.
|
|
557
|
+
if (this.adapter.noteLocalRequestSucceeded) {
|
|
558
|
+
this.adapter.noteLocalRequestSucceeded(duid);
|
|
559
|
+
}
|
|
555
560
|
resolve(result);
|
|
556
561
|
|
|
557
562
|
if (this.adapter.deviceNotify !== undefined) {
|
|
@@ -116,6 +116,7 @@ function getRequestTimeout(method, requestTimeoutMs) {
|
|
|
116
116
|
* @property {LoggerLike} log
|
|
117
117
|
* @property {(duid: string, update: TransportDiagnosticsUpdate) => Promise<void>} updateTransportDiagnostics
|
|
118
118
|
* @property {(duid: string) => Promise<boolean>} [ensureLocalConnection]
|
|
119
|
+
* @property {(duid: string, method?: string) => Promise<void>} [noteLocalRequestTimedOut]
|
|
119
120
|
* @property {(message: string, location: string, duid?: string) => void} catchError
|
|
120
121
|
* @property {(duid: string) => string} [describeDevice]
|
|
121
122
|
*/
|
|
@@ -347,6 +348,15 @@ class messageQueueHandler {
|
|
|
347
348
|
)
|
|
348
349
|
);
|
|
349
350
|
} else {
|
|
351
|
+
// A socket that keeps reporting itself connected while every
|
|
352
|
+
// request dies of silence is not a transport worth retrying
|
|
353
|
+
// forever. Fire-and-forget: the caller is owed its rejection now,
|
|
354
|
+
// not after the bookkeeping resolves.
|
|
355
|
+
if (this.adapter.noteLocalRequestTimedOut) {
|
|
356
|
+
Promise.resolve(
|
|
357
|
+
this.adapter.noteLocalRequestTimedOut(duid, method)
|
|
358
|
+
).catch(() => {});
|
|
359
|
+
}
|
|
350
360
|
reject(
|
|
351
361
|
new Error(
|
|
352
362
|
`Local request with id ${messageID} with method ${method} timed out after ${timeoutSeconds} seconds Local connect state: ${localConnectionState}`
|
|
@@ -66,6 +66,19 @@ const CLOUD_ONLY_TRANSPORT_MARKERS = Object.freeze({
|
|
|
66
66
|
// for robots the plugin never attempts a LAN connection to.
|
|
67
67
|
const UNEXPLAINED_REMOTE_REASON = "remote-device";
|
|
68
68
|
|
|
69
|
+
// A local socket that completed its TCP handshake and then answered nothing.
|
|
70
|
+
// This is a different failure from a connect that failed, and conflating the
|
|
71
|
+
// two is what made it invisible: the port is reachable, so nothing looks
|
|
72
|
+
// broken, while every request still dies of silence at its timeout.
|
|
73
|
+
const LOCAL_MUTE_REMOTE_REASON = "local-socket-connected-but-mute";
|
|
74
|
+
|
|
75
|
+
// Consecutive local timeouts tolerated before the LAN is written off for a
|
|
76
|
+
// robot. One is noise — a single lost frame on a healthy network is ordinary,
|
|
77
|
+
// and exiling that robot to the cloud for it would be worse than the bug this
|
|
78
|
+
// bound fixes. Three in a row on a socket that keeps reporting itself
|
|
79
|
+
// connected is not noise.
|
|
80
|
+
const LOCAL_MUTE_TIMEOUT_LIMIT = 3;
|
|
81
|
+
|
|
69
82
|
// Minimum gap between live-room map fetch attempts while cleaning. The map
|
|
70
83
|
// payload is an order of magnitude heavier than get_status, so it rides a
|
|
71
84
|
// slower cadence than the active status polls — but 20s meant a robot could
|
|
@@ -74,6 +87,57 @@ const UNEXPLAINED_REMOTE_REASON = "remote-device";
|
|
|
74
87
|
// modest while making the room track the robot closely enough to be useful.
|
|
75
88
|
const B01_LIVE_ROOM_MIN_FETCH_GAP_MS = 10000;
|
|
76
89
|
|
|
90
|
+
// A live-room fetch that keeps failing is a channel that is down for this
|
|
91
|
+
// robot, not a lost frame. The request is heavy, it always rides the cloud
|
|
92
|
+
// (get_map_v1 is a secure request), and every failure costs a full request
|
|
93
|
+
// timeout. Retrying it at live-display cadence for a whole run buys nothing:
|
|
94
|
+
// measured on an a70 whose map channel was timing out, one ten-minute clean
|
|
95
|
+
// spent ten guaranteed-to-fail cloud requests and never named a room. So widen
|
|
96
|
+
// the gap as failures pile up, and drop straight back to the live cadence the
|
|
97
|
+
// moment one answers. The first two failures are deliberately NOT slowed — a
|
|
98
|
+
// single lost frame on a healthy channel must not make a working live display
|
|
99
|
+
// sluggish, the same rule the local-mute limit follows.
|
|
100
|
+
const LIVE_ROOM_FAILURE_BACKOFF_AFTER = 2;
|
|
101
|
+
const LIVE_ROOM_FAILURE_BACKOFF_MAX_MS = 300000; // 5 min
|
|
102
|
+
|
|
103
|
+
/**
|
|
104
|
+
* Required gap before the next live-room fetch attempt, given how many
|
|
105
|
+
* attempts in a row have already failed.
|
|
106
|
+
* @param {number} [consecutiveFailures]
|
|
107
|
+
* @returns {number}
|
|
108
|
+
*/
|
|
109
|
+
function liveRoomFetchGapMs(consecutiveFailures) {
|
|
110
|
+
const over = Number(consecutiveFailures) - LIVE_ROOM_FAILURE_BACKOFF_AFTER;
|
|
111
|
+
if (!Number.isFinite(over) || over <= 0) {
|
|
112
|
+
return B01_LIVE_ROOM_MIN_FETCH_GAP_MS;
|
|
113
|
+
}
|
|
114
|
+
return Math.min(
|
|
115
|
+
B01_LIVE_ROOM_MIN_FETCH_GAP_MS * 2 ** over,
|
|
116
|
+
LIVE_ROOM_FAILURE_BACKOFF_MAX_MS
|
|
117
|
+
);
|
|
118
|
+
}
|
|
119
|
+
|
|
120
|
+
/**
|
|
121
|
+
* Zero the counters that describe THIS run. clearLiveRoomForDevice runs at
|
|
122
|
+
* every run boundary and its stated job is to stop state leaking into the next
|
|
123
|
+
* run, but it used to drop only the cached room. Everything else survived, so
|
|
124
|
+
* a line reading "attempt N this run" counted every run since Homebridge
|
|
125
|
+
* started, the placeholder explanation meant to be said once per run was only
|
|
126
|
+
* ever visible on the very first run of the process, and "failed N times in a
|
|
127
|
+
* row" could greet a new run's first failure with N already at 5. Resetting
|
|
128
|
+
* has to happen even when no room was ever resolved — a run that failed every
|
|
129
|
+
* attempt is precisely the run that left the counters high.
|
|
130
|
+
* @param {{consecutiveFailures?: number, unresolvedPoseCount?: number, placeholderReported?: boolean} | null | undefined} liveState
|
|
131
|
+
*/
|
|
132
|
+
function resetLiveRoomRunCounters(liveState) {
|
|
133
|
+
if (!liveState) {
|
|
134
|
+
return;
|
|
135
|
+
}
|
|
136
|
+
liveState.consecutiveFailures = 0;
|
|
137
|
+
liveState.unresolvedPoseCount = 0;
|
|
138
|
+
liveState.placeholderReported = false;
|
|
139
|
+
}
|
|
140
|
+
|
|
77
141
|
// B01/Q7 status cadence. These were literals in three places — the two gap
|
|
78
142
|
// values, the loop interval, and a hand-written startup line quoting all
|
|
79
143
|
// three. 3.2.0 changed the idle gap from 45s to 25s and the startup line kept
|
|
@@ -330,6 +394,12 @@ class Roborock {
|
|
|
330
394
|
// see markDeviceRemote.
|
|
331
395
|
/** @type {Map<string, string>} */
|
|
332
396
|
this.remoteDeviceReasons = new Map();
|
|
397
|
+
// Consecutive local request timeouts per robot, counted only while the
|
|
398
|
+
// local client still reports itself connected. Reset by any local reply:
|
|
399
|
+
// the count has to mean "this socket answers nothing", not "this socket has
|
|
400
|
+
// ever timed out".
|
|
401
|
+
/** @type {Map<string, number>} */
|
|
402
|
+
this.localMuteTimeouts = new Map();
|
|
333
403
|
|
|
334
404
|
this.name = "roborock";
|
|
335
405
|
this.deviceNotify = null;
|
|
@@ -1421,6 +1491,8 @@ class Roborock {
|
|
|
1421
1491
|
"udp-broadcast-discovery": "UDP broadcast discovery found the vacuum",
|
|
1422
1492
|
"marked-remote-after-connect-failure":
|
|
1423
1493
|
"local TCP connection failed and the vacuum was marked remote",
|
|
1494
|
+
[LOCAL_MUTE_REMOTE_REASON]:
|
|
1495
|
+
"the local TCP socket connected but the vacuum answered nothing on it, so the cloud is used instead",
|
|
1424
1496
|
[b01Q7Adapter.B01_CLOUD_ONLY_REMOTE_REASON]:
|
|
1425
1497
|
"this model speaks only Roborock's cloud protocol, which has no LAN control surface",
|
|
1426
1498
|
};
|
|
@@ -2959,6 +3031,63 @@ class Roborock {
|
|
|
2959
3031
|
});
|
|
2960
3032
|
}
|
|
2961
3033
|
|
|
3034
|
+
/**
|
|
3035
|
+
* A local reply arrived, so the socket is not mute. Called from the one place
|
|
3036
|
+
* a local response lands, so the count measures the socket's current silence
|
|
3037
|
+
* rather than its history.
|
|
3038
|
+
*
|
|
3039
|
+
* @param {string} duid
|
|
3040
|
+
* @returns {void}
|
|
3041
|
+
*/
|
|
3042
|
+
noteLocalRequestSucceeded(duid) {
|
|
3043
|
+
if (!duid) {
|
|
3044
|
+
return;
|
|
3045
|
+
}
|
|
3046
|
+
|
|
3047
|
+
this.localMuteTimeouts.delete(duid);
|
|
3048
|
+
}
|
|
3049
|
+
|
|
3050
|
+
/**
|
|
3051
|
+
* A local request timed out on a socket that reported itself connected. After
|
|
3052
|
+
* LOCAL_MUTE_TIMEOUT_LIMIT of those in a row the LAN is written off for this
|
|
3053
|
+
* robot and the cloud is used instead, because the alternative is paying the
|
|
3054
|
+
* full timeout on every poll and every command for the life of the process.
|
|
3055
|
+
*
|
|
3056
|
+
* @param {string} duid
|
|
3057
|
+
* @param {string} [method] the request that died, for the log line
|
|
3058
|
+
* @returns {Promise<void>}
|
|
3059
|
+
*/
|
|
3060
|
+
async noteLocalRequestTimedOut(duid, method) {
|
|
3061
|
+
if (!duid) {
|
|
3062
|
+
return;
|
|
3063
|
+
}
|
|
3064
|
+
|
|
3065
|
+
const failures = (this.localMuteTimeouts.get(duid) || 0) + 1;
|
|
3066
|
+
this.localMuteTimeouts.set(duid, failures);
|
|
3067
|
+
|
|
3068
|
+
if (failures < LOCAL_MUTE_TIMEOUT_LIMIT) {
|
|
3069
|
+
return;
|
|
3070
|
+
}
|
|
3071
|
+
|
|
3072
|
+
// Already written off: say it once, then stay quiet. The timeouts keep
|
|
3073
|
+
// coming until something upstream changes, and a warning per poll would
|
|
3074
|
+
// bury the one line that explains the switch.
|
|
3075
|
+
if (this.remoteDevices.has(duid)) {
|
|
3076
|
+
return;
|
|
3077
|
+
}
|
|
3078
|
+
|
|
3079
|
+
this.log.warn(
|
|
3080
|
+
`The local connection to ${this.describeDevice(duid)} connected but answered nothing: ` +
|
|
3081
|
+
`${failures} requests in a row timed out${method ? ` (last: ${method})` : ""}. ` +
|
|
3082
|
+
`Using the Roborock cloud for this robot instead. The LAN port is reachable, ` +
|
|
3083
|
+
`so this is not a blocked port — the robot is not replying on it. On a segmented ` +
|
|
3084
|
+
`network, check that the reply path back to Homebridge is open, not just the ` +
|
|
3085
|
+
`outbound one.`
|
|
3086
|
+
);
|
|
3087
|
+
|
|
3088
|
+
await this.markDeviceRemote(duid, LOCAL_MUTE_REMOTE_REASON);
|
|
3089
|
+
}
|
|
3090
|
+
|
|
2962
3091
|
/**
|
|
2963
3092
|
* Drops the remote marker and the reason that went with it.
|
|
2964
3093
|
*
|
|
@@ -4524,7 +4653,10 @@ class Roborock {
|
|
|
4524
4653
|
if (liveState.inflight) {
|
|
4525
4654
|
return liveState.inflight;
|
|
4526
4655
|
}
|
|
4527
|
-
if (
|
|
4656
|
+
if (
|
|
4657
|
+
Date.now() - liveState.lastAttemptAt <
|
|
4658
|
+
liveRoomFetchGapMs(liveState.consecutiveFailures)
|
|
4659
|
+
) {
|
|
4528
4660
|
return liveState.current;
|
|
4529
4661
|
}
|
|
4530
4662
|
liveState.lastAttemptAt = Date.now();
|
|
@@ -4566,7 +4698,7 @@ class Roborock {
|
|
|
4566
4698
|
|
|
4567
4699
|
const resolution2 = b01Q7Adapter.describeLiveRoomResolution(parsed);
|
|
4568
4700
|
const roomId = resolution2.roomId;
|
|
4569
|
-
|
|
4701
|
+
this.noteLiveRoomFetchRecovered(duid, liveState);
|
|
4570
4702
|
|
|
4571
4703
|
if (roomId === null) {
|
|
4572
4704
|
// Debug-only used to make this invisible, and it is the single most
|
|
@@ -4671,6 +4803,27 @@ class Roborock {
|
|
|
4671
4803
|
return this._b01LiveRoomState?.get(duid)?.current || null;
|
|
4672
4804
|
}
|
|
4673
4805
|
|
|
4806
|
+
/**
|
|
4807
|
+
* Close the loop on a failure streak. "Live-room map fetch has failed N
|
|
4808
|
+
* times in a row" had no counterpart, so a channel that came back left the
|
|
4809
|
+
* last word in the log saying it was broken. Said exactly when the backoff
|
|
4810
|
+
* had begun slowing the fetch down, so the message and the behaviour it
|
|
4811
|
+
* reports on cannot drift apart.
|
|
4812
|
+
* @param {string} duid
|
|
4813
|
+
* @param {{consecutiveFailures?: number} | null | undefined} liveState
|
|
4814
|
+
*/
|
|
4815
|
+
noteLiveRoomFetchRecovered(duid, liveState) {
|
|
4816
|
+
const failures = liveState?.consecutiveFailures || 0;
|
|
4817
|
+
if (failures > LIVE_ROOM_FAILURE_BACKOFF_AFTER) {
|
|
4818
|
+
this.log.info(
|
|
4819
|
+
`Live-room map fetch for ${this.describeDevice(duid)} recovered after ${failures} failed attempt(s).`
|
|
4820
|
+
);
|
|
4821
|
+
}
|
|
4822
|
+
if (liveState) {
|
|
4823
|
+
liveState.consecutiveFailures = 0;
|
|
4824
|
+
}
|
|
4825
|
+
}
|
|
4826
|
+
|
|
4674
4827
|
/** @param {string} duid */
|
|
4675
4828
|
clearB01LiveRoom(duid) {
|
|
4676
4829
|
const liveState = this._b01LiveRoomState?.get(duid);
|
|
@@ -4680,6 +4833,7 @@ class Roborock {
|
|
|
4680
4833
|
);
|
|
4681
4834
|
liveState.current = null;
|
|
4682
4835
|
}
|
|
4836
|
+
resetLiveRoomRunCounters(liveState);
|
|
4683
4837
|
}
|
|
4684
4838
|
|
|
4685
4839
|
/**
|
|
@@ -4727,6 +4881,7 @@ class Roborock {
|
|
|
4727
4881
|
);
|
|
4728
4882
|
liveState.current = null;
|
|
4729
4883
|
}
|
|
4884
|
+
resetLiveRoomRunCounters(liveState);
|
|
4730
4885
|
}
|
|
4731
4886
|
|
|
4732
4887
|
/**
|
|
@@ -4798,7 +4953,7 @@ class Roborock {
|
|
|
4798
4953
|
// microseconds).
|
|
4799
4954
|
const segmentId =
|
|
4800
4955
|
RRMapParser.resolveLiveSegmentFromMapBuffer(mapBuffer);
|
|
4801
|
-
|
|
4956
|
+
this.noteLiveRoomFetchRecovered(duid, liveState);
|
|
4802
4957
|
|
|
4803
4958
|
if (segmentId === null) {
|
|
4804
4959
|
this.log.debug(
|
|
@@ -4960,6 +5115,8 @@ module.exports = {
|
|
|
4960
5115
|
Roborock,
|
|
4961
5116
|
CLOUD_ONLY_TRANSPORT_MARKERS,
|
|
4962
5117
|
UNEXPLAINED_REMOTE_REASON,
|
|
5118
|
+
LOCAL_MUTE_REMOTE_REASON,
|
|
5119
|
+
LOCAL_MUTE_TIMEOUT_LIMIT,
|
|
4963
5120
|
};
|
|
4964
5121
|
|
|
4965
5122
|
////////////////////////////////////////////////////////////////////////////////////////////////////
|