@observertc/observer-js 1.0.0-beta.18 → 1.0.0-beta.21
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +93 -3
- package/dist/index.d.mts +689 -71
- package/dist/index.d.ts +689 -71
- package/dist/index.js +129 -18
- package/dist/index.js.map +1 -1
- package/dist/index.mjs +127 -18
- package/dist/index.mjs.map +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -1114,9 +1114,10 @@ What each adds that no endpoint can know:
|
|
|
1114
1114
|
reasons "clients in unrelated calls share only the infrastructure, so it must be us" — right for
|
|
1115
1115
|
network symptoms, **wrong for endpoint ones**. `cpulimitation` across six unrelated calls is not an
|
|
1116
1116
|
SFU event; CPU is owned by the endpoint, so what those endpoints share is a browser version or a
|
|
1117
|
-
client release. Groups by `browser` / `engine` / `platform` / `operationSystem`, one
|
|
1118
|
-
instance. The gate is **relative risk**, not share: "30% of Chrome 141 is unhappy" means
|
|
1119
|
-
30% of everyone is, and a share-based rule simply indicts whichever browser is most
|
|
1117
|
+
client release. Groups by `browser` / `engine` / `platform` / `operationSystem` / `location`, one
|
|
1118
|
+
axis per instance. The gate is **relative risk**, not share: "30% of Chrome 141 is unhappy" means
|
|
1119
|
+
nothing if 30% of everyone is, and a share-based rule simply indicts whichever browser is most
|
|
1120
|
+
popular. See [the `location` axis](#grouping-by-place-the-location-axis) for the geographic form.
|
|
1120
1121
|
- **`SfuCongestionDetector`** — *is congestion spiking across the fleet right now?* Counts distinct
|
|
1121
1122
|
clients reporting congestion in fixed wall-clock buckets and compares each bucket against a
|
|
1122
1123
|
median+MAD baseline of the ones before it. Buckets rather than update ticks on purpose: the tick is
|
|
@@ -1153,6 +1154,45 @@ observer.addObserverDetector('observer-concurrent-issue-detector', {
|
|
|
1153
1154
|
});
|
|
1154
1155
|
```
|
|
1155
1156
|
|
|
1157
|
+
### Grouping by place: the `location` axis
|
|
1158
|
+
|
|
1159
|
+
If your clients report coordinates, `ClientPopulationIssueDetector` can group by **where they are**
|
|
1160
|
+
instead of what they run — which is the grouping network symptoms actually cluster by:
|
|
1161
|
+
|
|
1162
|
+
```ts
|
|
1163
|
+
observer.addObserverDetector('client-population-issue-detector', {
|
|
1164
|
+
issueTypes: [ 'congestion', 'ice-disconnected' ],
|
|
1165
|
+
groupBy: 'location',
|
|
1166
|
+
locationPrecision: 3, // geohash chars: 3 ~156 km, 4 ~39 km, 5 ~5 km
|
|
1167
|
+
resolveClientLocation: (client) => client.attachments?.geo as { latitude: number, longitude: number },
|
|
1168
|
+
});
|
|
1169
|
+
```
|
|
1170
|
+
|
|
1171
|
+
**The client still owns "RTT jumped".** client-monitor's `CongestionDetector` compares each peer
|
|
1172
|
+
connection's RTT against its own EWMA baseline and requires a bandwidth-limitation corroboration
|
|
1173
|
+
before raising `congestion`. Absolute RTT is not comparable between clients — someone 200 ms away is
|
|
1174
|
+
*always* 200 ms away, so the only signal is deviation from that client's own baseline, which is
|
|
1175
|
+
exactly what the client measures. The observer's contribution is the part no endpoint can see: that
|
|
1176
|
+
many of the affected clients are **in the same place at the same time**.
|
|
1177
|
+
|
|
1178
|
+
Three things to know:
|
|
1179
|
+
|
|
1180
|
+
- **Cells, not radii.** The population is a geohash prefix. "Within N km" is a clustering problem —
|
|
1181
|
+
order-dependent, no stable group name, pairwise cost — and a detector needs the *same* group key on
|
|
1182
|
+
every tick for its cooldown and control group to mean anything. The cost is that a cell boundary can
|
|
1183
|
+
split two adjacent clients, which biases towards missing a finding rather than inventing one.
|
|
1184
|
+
- **Only the cell key is reported.** `payload.population` is the geohash; coordinates never enter the
|
|
1185
|
+
issue. These payloads get archived into [call summaries](#call-summaries), so that matters.
|
|
1186
|
+
- **Geography is confounded with your topology.** The control group is "everyone outside this cell",
|
|
1187
|
+
which cannot separate *"the path into this region degraded"* from *"the SFU serving this region
|
|
1188
|
+
degraded"*. If a region maps largely onto one deployment, both hypotheses fit the same evidence —
|
|
1189
|
+
so the finding concludes `infrastructure` and points you at `SfuCongestionDetector` /
|
|
1190
|
+
`TurnServerHealthDetector`, which answer whether clients elsewhere on the same server also
|
|
1191
|
+
degraded. It does not claim an attribution it cannot support.
|
|
1192
|
+
|
|
1193
|
+
Coordinates are not in `ClientSample`, so `resolveClientLocation` is required; without it the detector
|
|
1194
|
+
logs a warning at construction and finds nothing, rather than quietly reporting no findings forever.
|
|
1195
|
+
|
|
1156
1196
|
### Validators — one-shot structural checks
|
|
1157
1197
|
|
|
1158
1198
|
Every detector above answers *"is something wrong right now?"* and runs on every tick, because the
|
|
@@ -1528,6 +1568,56 @@ const observer = new Observer({
|
|
|
1528
1568
|
For the mediasoup factory, the application puts `producerId` / `consumerId` (and optionally
|
|
1529
1569
|
`direction`, `label`) into the track `attachments`.
|
|
1530
1570
|
|
|
1571
|
+
### Resolution stands alone — you do not need router observation
|
|
1572
|
+
|
|
1573
|
+
Track resolution and [mediasoup router observation](#mediasoup-router-observation) are two
|
|
1574
|
+
**independent** opt-ins that happen to share a vendor name. Nothing in `RemoteTrackResolver` reads
|
|
1575
|
+
`ObservedMediasoupRouter`, and nothing in `ObservedMediasoupRouter` touches calls, clients or tracks.
|
|
1576
|
+
|
|
1577
|
+
So if your application already builds its own per-router report and you only want the detectors that
|
|
1578
|
+
need publisher↔subscriber links, set `createRemoteTrackResolver` and simply never call
|
|
1579
|
+
`observer.createObservedMediasoupRouter(...)`. No `MediasoupRouterSample` is created, nothing
|
|
1580
|
+
accumulates, and these keep working:
|
|
1581
|
+
|
|
1582
|
+
`UnconsumedTrackDetector`, `PublisherFaultCorroborationDetector`, `TrackDeliveryMismatchDetector`,
|
|
1583
|
+
`IssueFanOutDetector`, `RemoteTrackResolverValidator`, `SimulcastReceiverValidator`.
|
|
1584
|
+
|
|
1585
|
+
### Backing a strategy with your own mapping
|
|
1586
|
+
|
|
1587
|
+
The three resolvers are plain functions returning a link key, so they can read a table your own
|
|
1588
|
+
report already maintains rather than something the client attached. The only requirement is a key
|
|
1589
|
+
**visible on both sides**. SSRC is the useful one, because it needs no client cooperation at all —
|
|
1590
|
+
mediasoup knows each consumer's `rtpParameters.encodings[].ssrc` server-side, and the subscriber's
|
|
1591
|
+
inbound RTP reports the same value:
|
|
1592
|
+
|
|
1593
|
+
```ts
|
|
1594
|
+
// your own table, filled where you already create consumers
|
|
1595
|
+
const ssrcToProducerId = new Map<number, string>();
|
|
1596
|
+
|
|
1597
|
+
const observer = new Observer({
|
|
1598
|
+
createRemoteTrackResolver: (observedCall) => new RemoteTrackResolver(observedCall, {
|
|
1599
|
+
resolveOutboundTrackPublisherId: (out) => out.attachments?.producerId as string | undefined,
|
|
1600
|
+
resolveInboundTrackPublisherId: (inb) => {
|
|
1601
|
+
const ssrc = inb.getInboundRtp()?.ssrc;
|
|
1602
|
+
|
|
1603
|
+
return ssrc === undefined ? undefined : ssrcToProducerId.get(ssrc);
|
|
1604
|
+
},
|
|
1605
|
+
}),
|
|
1606
|
+
});
|
|
1607
|
+
```
|
|
1608
|
+
|
|
1609
|
+
> **A key that arrives late still links.** A track announces itself once, but its `attachments` are
|
|
1610
|
+
> replaced on every sample and a table like the one above is inherently racy against sample arrival.
|
|
1611
|
+
> Tracks whose key does not resolve at first sight are held and retried on their own
|
|
1612
|
+
> `*-track-updated`, i.e. exactly when new stats arrive for them — so a key that appears on the second
|
|
1613
|
+
> sample links then, rather than being lost for the track's lifetime. `resolver.pendingTrackCounts`
|
|
1614
|
+
> reports how many are still waiting; in a healthy setup it is `{ inbound: 0, outbound: 0 }`.
|
|
1615
|
+
|
|
1616
|
+
If a strategy resolves *nothing* the failure is quiet — "no subscribers" and "no links resolved" look
|
|
1617
|
+
identical from the outside. That is what
|
|
1618
|
+
[`RemoteTrackResolverValidator`](#validators--one-shot-structural-checks) is for; run it once in
|
|
1619
|
+
staging after wiring up a custom strategy.
|
|
1620
|
+
|
|
1531
1621
|
---
|
|
1532
1622
|
|
|
1533
1623
|
## Mediasoup router observation
|