homebridge-roborock-matter 3.30.0 → 3.32.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,146 @@
1
1
  # Changelog
2
2
 
3
+ ## 3.32.0
4
+
5
+ **The give-up register had never counted a single poll failure. Not one, in two releases that were built around it.**
6
+
7
+ ### What I shipped twice and never verified
8
+
9
+ 3.30.0 introduced a register that stops asking a robot a question it never answers. 3.31.0 routed three more polls into it and fixed how it attributed blame. Both releases described it — in the changelog, and in the log line you read — as the thing that ends the flood of repeated timeouts, naming the seven methods from #22 and #24 that motivated it.
10
+
11
+ It could not have counted any of them. `pollParameter` decided the outcome from its own `try/catch`:
12
+
13
+ ```js
14
+ try {
15
+ const answer = await vacuum.getParameter(duid, method);
16
+ this.noteMethodAnswered(duid, method); // ran on EVERY poll
17
+ return answer;
18
+ } catch (error) {
19
+ this.noteMethodUnanswered(duid, method, error); // unreachable
20
+ throw error;
21
+ }
22
+ ```
23
+
24
+ `vacuum.getParameter` swallows its own errors. Its catch calls `catchError`, which only logs, and the function resolves `undefined`. So the catch here never ran. Every timeout looked like a success, and `noteMethodAnswered` **reset the counter on each failed poll**.
25
+
26
+ Measured on the real classes before the fix: 20 consecutive timeouts produced **0 entries** in the register. And a robot that had somehow been left alone was told `answers get_room_mapping again; it is back in the normal poll cycle` on the retry that had just timed out.
27
+
28
+ The only methods it ever governed were the three map requests, because those go through `sendRequest` directly rather than `getParameter`. Everything else — `get_room_mapping`, `get_multi_maps_list`, `get_consumable`, `get_carpet_mode`, `get_carpet_clean_mode`, `get_water_box_custom_mode`, `get_timer` — was untouched. 3.31.0's "three polls were bypassing the register" fix wired them to nothing.
29
+
30
+ ### The register now watches the wire
31
+
32
+ A request is answered or it is not, and the only place that knows is the message layer. So that is where the register is fed from: `messageQueueHandler` reports a reply when one arrives and a timeout when one does not. `pollParameter` no longer guesses at the outcome — it just declares which requests it is entitled to skip.
33
+
34
+ That declaration matters, because the message layer sees **every** request, including `get_status` and every command. A robot/method pair is counted only after a caller that can skip it has claimed it, so the two things this register must never touch cannot be reached by construction. There are tests that fire a hundred `get_status` timeouts and a hundred command timeouts and assert nothing closes.
35
+
36
+ Measured after: 200 polls at a silent method send **6 requests**.
37
+
38
+ ### And the transport exclusion was dead code again
39
+
40
+ 3.30.0's version looked for `EAI_AGAIN`, `offline` and friends — words the timeout message never contains. 3.31.0 replaced it with a regex for `MQTT connection state: false`, and I wrote a checklist entry about building test fixtures from real strings. That regex cannot match either: a down link is rejected **earlier**, with a refusal that has no "timed out after" in it, so the timeout arm is only ever reached with the flag reading `true`. Two releases, two dead gates, two green tests built from strings the code cannot produce.
41
+
42
+ The timeout now carries what it knows as data — `unansweredRequest` and `transportWasUp`, the latter read at **rejection** time rather than copied before the send, because a link that dies mid-flight is the entire case the exclusion exists for. The register reads the fields. No prose is parsed.
43
+
44
+ ### The measurement I hid behind a log level
45
+
46
+ 3.31.0 added a line for every discarded protocol-301 map frame, to answer why `get_map_v1` times out forever on a robot that answers everything else. I put it at **debug**, which is off by default. On my own server it produced nothing; a user in #24 went looking and found nothing. We both read that as evidence. It was not evidence — it was a log level.
47
+
48
+ Discarded frames are now counted per robot, and the count rides along in the give-up message, at info level, where nobody has to know to go looking:
49
+
50
+ ```
51
+ Stueetage has not answered get_map_v1 6 times in a row, so the plugin stops
52
+ asking for about 6 hour(s)… No reply frames for this robot were received and
53
+ discarded, so the reply is not arriving at all rather than being lost on this
54
+ side.
55
+ ```
56
+
57
+ or, if it is our fault:
58
+
59
+ ```
60
+ …NOTE: 6 map reply frame(s) for this robot were received and then discarded by
61
+ the plugin (addressed-elsewhere: 6), so the robot IS answering and this side is
62
+ throwing it away. Please report this — it is a bug here, not on your robot.
63
+ ```
64
+
65
+ The same figure is in the diagnostic report. Either way the question gets answered by a log you were going to send anyway.
66
+
67
+ ### Field notes from 3.31.0
68
+
69
+ The give-up rule did fire correctly on my own a70: `get_map_v1`, 6 in a row, paused for 6 hours — where the same request used to fail 95 and 225 times. That part worked, because the map path was the one path the register could actually see.
70
+
71
+ 2045 tests, 14 of them new and all 14 red against 3.31.0.
72
+
73
+ ## 3.31.0
74
+
75
+ **Six bugs, three of them mine from last week. And the oldest unexplained failure in this project finally has an instrument pointed at it.**
76
+
77
+ ### The give-up rule tripped on network trouble — the exact thing it promised not to do
78
+
79
+ 3.30.0 added a register that stops asking a robot a request it never answers, and said, in the source and in the line the user reads: it never trips on a transport error, because those come back on their own.
80
+
81
+ It did. The exclusion was dead code.
82
+
83
+ `messageQueueHandler` builds its two timeout messages with the connection state interpolated as a **boolean**:
84
+
85
+ ```
86
+ … timed out after 10 seconds. MQTT connection state: false
87
+ … timed out after 10 seconds Local connect state: false
88
+ ```
89
+
90
+ The rule looked for `EAI_AGAIN`, `ENOTFOUND`, `ECONNREFUSED`, `ECONNRESET`, "not connected", "offline" — every one of which belongs to an error that never contains "timed out after" in the first place, and so had already been excluded a line earlier. The second gate had nothing left to exclude.
91
+
92
+ A four-minute network blip during a clean is 24 failed live-room polls at 10-second intervals. Six is all it takes. Live-room tracking then died for six hours, under a log line telling you this was not a connection failure.
93
+
94
+ It now reads the boolean, and also stands down when the cloud timeout reports that nothing is coming back over MQTT at all. The tests use the exact strings the plugin emits — 3.30.0's used hand-written messages the code cannot produce, which is precisely why they passed against a broken rule.
95
+
96
+ ### A robot was punished for a request it answered perfectly
97
+
98
+ The B01/Q7 live-room fetch is two requests: `get_map_list` (cheap) and `service.upload_by_mapid` (the heavy map payload, with its own 20-second timeout). 3.30.0 wrapped both in one try, acknowledged `get_map_list` only after the second had been fetched, decoded and cached, and recorded every failure — including the upload leg's own timeout — against `get_map_list`.
99
+
100
+ So a robot answering `get_map_list` in 200 ms every time, with a silent upload channel, had the wrong counter climb to six. The diagnostics then named `get_map_list`: the one channel that was working.
101
+
102
+ Each leg is now counted on its own.
103
+
104
+ ### The heartbeat stopped healing a write that never landed
105
+
106
+ Suppressing an unchanged `operationalState` is what stopped the tank notification repeating every 2 minutes. It also broke the safety net the whole publish-dedup design rests on.
107
+
108
+ A cluster write can be rejected by matter.js **after** `updateAccessoryState` has resolved — this codebase has documented that for months: `#assertCurrentPhase` throws, Homebridge swallows the throw, and the controller keeps what it last accepted. From 3.30.0, one such rejection made the plugin believe a value was published that never was, and nothing wrote that attribute again. A permanently stale tile, recoverable only by the robot reaching a different state, or a restart. Before 3.30.0 the heartbeat repaired it inside a minute.
109
+
110
+ A forced heartbeat now re-asserts `operationalState` once every 10 cycles — about once every 10 minutes. Measured: 6 re-assertions an hour instead of 60, the stale tile heals within 10 minutes, and the notification stays gone.
111
+
112
+ ### A map reply can be thrown away in silence — and now it says so
113
+
114
+ This is the one I care most about.
115
+
116
+ `get_map_v1` on a classic robot times out after 10 seconds, forever, while the same robot answers everything else. 95 times in a row on my own a70, 225 twelve days earlier, 40 on the a75 in #9. Nobody has ever been able to say why.
117
+
118
+ Map replies do not come back on the ordinary reply path — they arrive as protocol 301 frames. The 301 handler had several ways to drop one, every one of them a bare `return`. No log, no counter, nothing. A dropped 301 leaves the request to die on its timer, which looks **exactly** like a robot that never answered. There was nothing to diagnose it with.
119
+
120
+ One of those drops was also wrong:
121
+
122
+ ```js
123
+ if (!endpoint.startsWith(data2.endpoint)) return;
124
+ ```
125
+
126
+ `endpoint` is our own 8-character key. `data2.endpoint` is a 15-byte wire field with only _trailing_ nulls stripped. python-roborock, the reference implementation, compares it the other way round — `received.startswith(ours)`. As written, a robot that echoes our 8 characters followed by anything that is not a trailing null leaves a longer string, and an 8-character string can never `startsWith` a longer one. It failed closed, on a reply addressed to us, without a word.
127
+
128
+ The comparison now matches the reference, and **every** 301 drop explains itself and names the robot. I am not claiming this is the cause of the a70's timeouts. I am saying that from this release, if it is, the log says so.
129
+
130
+ ### One robot's unfinished photo could swallow every other robot's map
131
+
132
+ `photoGzipChunks` and `photoChunkID` were module-level variables shared by every robot on the account, cleared only when a photo transfer **completed**. A robot going offline between chunk 1 and chunk 2 left the id set forever — and from then on every 301 frame with `seq == 2`, from any robot, was swallowed into that stale buffer instead of being decoded as a map reply. A permanent, silent map outage on a multi-robot account, with no error anywhere.
133
+
134
+ The buffer is now per robot, and one whose request is no longer waiting is discarded.
135
+
136
+ ### Smaller things
137
+
138
+ - **A robot that comes back online starts with clean counters.** `forgetDevice` existed since 3.30.0 and was never called once. A robot offline for hours kept the counts it collected while unreachable, and because the rule deliberately keeps the counter when it lets one request through, a single probe timing out during the reconnect closed the method for another six hours.
139
+ - **Three polls were bypassing the register entirely** — `get_multi_maps_list` and `get_room_mapping`, two of the seven methods named in the 647 suppressed timeouts that motivated the rule. The test that was supposed to catch this asserted "exactly 3 call sites", a number that silently excluded them. It now asserts the rule instead: no optional poll reaches the robot except through the register.
140
+ - **`cleaning_info` is printed in full while a robot is cleaning.** Not a feature — a measurement. Every status poll already receives this object and throws it away, and it is documented as carrying `{target_segment_id, segment_id, …}`. If `segment_id` tracks the room, live-room tracking on classic robots needs no map at all, which would route around the timeout above entirely. One clean answers it.
141
+
142
+ 2031 tests, 33 of them new.
143
+
3
144
  ## 3.30.0
4
145
 
5
146
  **The water-tank notification that repeated every two minutes was us. Three people reported it, and for three releases this project told them it was Apple.**
package/README.md CHANGED
@@ -37,7 +37,7 @@ This is the most feature-packed, most thoroughly engineered Roborock plugin for
37
37
  - 📍 **See where it's cleaning — live.** Apple Home shows _"Cleaning — Kitchen"_ with the room the robot is actually inside, updating as it moves from room to room. Works even for cleans started from the robot's button or the Roborock app. No other Homebridge plugin does this.
38
38
  - 🧭 **One robot, one tile — and as many robots as you own.** Sign in once and your whole fleet comes along: every vacuum on your account appears as its own clean, native accessory in Apple Home. No clutter of fake fans and helper switches, and rooms appear with the names you gave them in the Roborock app.
39
39
  - ⚡ **Fast and reliable.** Commands go directly to the robot over your own network whenever possible, with the Roborock cloud as automatic backup — and built-in diagnostics in the settings if you ever want to look under the hood.
40
- - 🛡️ **Verified by Homebridge.** Reviewed and endorsed by the Homebridge team. 1998 automated tests, zero known vulnerabilities, no analytics, and a startup designed to never crash your Homebridge — even when your Wi-Fi or the Roborock cloud has a bad day.
40
+ - 🛡️ **Verified by Homebridge.** Reviewed and endorsed by the Homebridge team. 2045 automated tests, zero known vulnerabilities, no analytics, and a startup designed to never crash your Homebridge — even when your Wi-Fi or the Roborock cloud has a bad day.
41
41
 
42
42
  ## Features
43
43
 
@@ -264,7 +264,7 @@ The complete path — robot → plugin → Homebridge → matter.js store — wa
264
264
 
265
265
  ## Contributing
266
266
 
267
- Model reports, diagnostics exports, and pull requests are very welcome. The codebase ships with 1998 tests (protocol fixtures verified against the [python-roborock](https://github.com/Python-roborock/python-roborock) reference), strict TypeScript checking, and CI across Node 22/24 × Homebridge 1.11/2.x — `npm test` before you push and you're set.
267
+ Model reports, diagnostics exports, and pull requests are very welcome. The codebase ships with 2045 tests (protocol fixtures verified against the [python-roborock](https://github.com/Python-roborock/python-roborock) reference), strict TypeScript checking, and CI across Node 22/24 × Homebridge 1.11/2.x — `npm test` before you push and you're set.
268
268
 
269
269
  ## Support the project
270
270
 
@@ -219,6 +219,16 @@ const DOCK_ERROR_CLEAN_WATER_TANK_EMPTY = 38;
219
219
  * writeOperationalStateCluster().
220
220
  */
221
221
  const OPERATIONAL_STATE_CLUSTER = "rvcOperationalState";
222
+ /**
223
+ * How many forced publishes (heartbeats) may pass before `operationalState`
224
+ * is re-written even though it has not changed.
225
+ *
226
+ * 10 heartbeats is about 10 minutes. That restores the self-healing the
227
+ * heartbeat provided before 3.30.0, at a rate far too low for the
228
+ * wipe/raise pair matter.js 0.17.9 performs to become a notification anyone
229
+ * notices — the fault cycle needed one every single minute.
230
+ */
231
+ const RESYNC_OPERATIONAL_STATE_EVERY_FORCED_WRITES = 10;
222
232
  const RVC_OPERATIONAL_ERROR = {
223
233
  NO_ERROR: 0,
224
234
  UNABLE_TO_START_OR_RESUME: 1,
@@ -665,6 +675,10 @@ class RoborockMatterVacuumAccessory {
665
675
  // dropped on failure, and (c) the heartbeat performs a forced full publish
666
676
  // every cycle, self-healing any residual divergence within a minute.
667
677
  this.lastPublishedClusterJson = new Map();
678
+ // How many forced publishes in a row have skipped operationalState. See
679
+ // writeOperationalStateCluster(): without this, a write matter.js rejects
680
+ // silently can never come back.
681
+ this.forcedWritesSinceOperationalState = 0;
668
682
  this.serviceAreaProgress = [];
669
683
  this.selectedCleanMode = CLEAN_MODE_VACUUM;
670
684
  this.selectedCleanModeNeedsApply = false;
@@ -1218,6 +1232,7 @@ class RoborockMatterVacuumAccessory {
1218
1232
  this.lastPublishedClusterJson.clear();
1219
1233
  this.publishedOperationalState = undefined;
1220
1234
  this.publishedOperationalError = undefined;
1235
+ this.forcedWritesSinceOperationalState = 0;
1221
1236
  // …and nothing has been stated about it either, so the evidence line is
1222
1237
  // restated for the new node instead of being suppressed as unchanged.
1223
1238
  this.lastLoggedMatterPublishLine = null;
@@ -1770,13 +1785,41 @@ class RoborockMatterVacuumAccessory {
1770
1785
  * the store really was cleared); clearing and re-raising the tank still
1771
1786
  * produce exactly one event each.
1772
1787
  */
1773
- async writeOperationalStateCluster(matter, attributes) {
1788
+ async writeOperationalStateCluster(matter, attributes, options = {}) {
1774
1789
  const { operationalState, operationalError, ...rest } = attributes;
1775
1790
  const first = { ...rest };
1776
- const stateChanged = operationalState !== undefined &&
1791
+ let stateChanged = operationalState !== undefined &&
1777
1792
  operationalState !== this.publishedOperationalState;
1793
+ // THE HOLE 3.30.0 LEFT, AND WHY IT NEEDED CLOSING.
1794
+ //
1795
+ // Suppressing an unchanged `operationalState` is what stops the tank
1796
+ // fault being cleared and re-raised every minute. But the whole dedup
1797
+ // design rests on the heartbeat being a forced full write that self-heals
1798
+ // any divergence within a minute — and a cluster write can be rejected by
1799
+ // matter.js AFTER `updateAccessoryState` has already resolved. This file
1800
+ // documents that elsewhere: `OperationalStateServer.#assertCurrentPhase`
1801
+ // throws, and Homebridge swallows the throw, so the whole cluster write
1802
+ // is silently rejected and the controller keeps what it last accepted.
1803
+ //
1804
+ // Believing such a write landed and then never writing that attribute
1805
+ // again turns a one-off rejection into a permanently stale tile,
1806
+ // recoverable only by the robot reaching a different state or a restart.
1807
+ // Before 3.30.0 the heartbeat repaired it inside a minute.
1808
+ //
1809
+ // So: never on an ordinary publish, but a forced write re-asserts it once
1810
+ // every RESYNC_OPERATIONAL_STATE_EVERY_FORCED_WRITES heartbeats.
1811
+ if (!stateChanged &&
1812
+ options.force === true &&
1813
+ operationalState !== undefined) {
1814
+ this.forcedWritesSinceOperationalState += 1;
1815
+ if (this.forcedWritesSinceOperationalState >=
1816
+ RESYNC_OPERATIONAL_STATE_EVERY_FORCED_WRITES) {
1817
+ stateChanged = true;
1818
+ }
1819
+ }
1778
1820
  if (stateChanged) {
1779
1821
  first.operationalState = operationalState;
1822
+ this.forcedWritesSinceOperationalState = 0;
1780
1823
  }
1781
1824
  if (Object.keys(first).length > 0) {
1782
1825
  await matter.updateAccessoryState(this.accessory.UUID, OPERATIONAL_STATE_CLUSTER, first);
@@ -1802,7 +1845,7 @@ class RoborockMatterVacuumAccessory {
1802
1845
  await matter.updateAccessoryState(this.accessory.UUID, OPERATIONAL_STATE_CLUSTER, { operationalError });
1803
1846
  this.publishedOperationalError = errorStateId;
1804
1847
  }
1805
- async updateMatterState(partialClusters, reason = "state update") {
1848
+ async updateMatterState(partialClusters, reason = "state update", options = {}) {
1806
1849
  if (!this.registered) {
1807
1850
  return false;
1808
1851
  }
@@ -1829,7 +1872,9 @@ class RoborockMatterVacuumAccessory {
1829
1872
  await Promise.all(clusterEntries.map(async ([cluster, attributes]) => {
1830
1873
  try {
1831
1874
  if (cluster === OPERATIONAL_STATE_CLUSTER) {
1832
- await this.writeOperationalStateCluster(matter, attributes);
1875
+ await this.writeOperationalStateCluster(matter, attributes, {
1876
+ force: options.force === true,
1877
+ });
1833
1878
  }
1834
1879
  else {
1835
1880
  await matter.updateAccessoryState(this.accessory.UUID, cluster, attributes);
@@ -1846,6 +1891,7 @@ class RoborockMatterVacuumAccessory {
1846
1891
  // re-assert them.
1847
1892
  this.publishedOperationalState = undefined;
1848
1893
  this.publishedOperationalError = undefined;
1894
+ this.forcedWritesSinceOperationalState = 0;
1849
1895
  }
1850
1896
  failures.push(error);
1851
1897
  this.platform.log.debug(`Matter publish for cluster ${cluster} on ${this.accessory.UUID} failed: ${error instanceof Error ? error.message : String(error)}`);
@@ -1944,7 +1990,12 @@ class RoborockMatterVacuumAccessory {
1944
1990
  },
1945
1991
  }, "Battery resync nudge");
1946
1992
  }
1947
- const updated = await this.updateMatterState(clusters, reason);
1993
+ // `force` is carried down so writeOperationalStateCluster can tell a
1994
+ // heartbeat from an ordinary publish; it is the only cluster for which
1995
+ // the distinction still means anything.
1996
+ const updated = await this.updateMatterState(clusters, reason, {
1997
+ force: options.force === true,
1998
+ });
1948
1999
  if (updated) {
1949
2000
  this.logMatterPublishIfChanged(snapshot, reason);
1950
2001
  }