@observertc/observer-js 1.0.0-beta.13 → 1.0.0-beta.15
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +643 -61
- package/dist/index.d.mts +2731 -74
- package/dist/index.d.ts +2731 -74
- package/dist/index.js +3471 -374
- package/dist/index.js.map +1 -1
- package/dist/index.mjs +3421 -373
- package/dist/index.mjs.map +1 -1
- package/package.json +14 -4
- package/llms-full.txt +0 -1632
package/README.md
CHANGED
|
@@ -40,11 +40,9 @@ and emits a single, unified stream of typed events the application can react to.
|
|
|
40
40
|
> works whether your project uses `import` (ESM) or `require()` (CommonJS). Everything — including
|
|
41
41
|
> the built-in file sink — is exported from the single `@observertc/observer-js` entry.
|
|
42
42
|
|
|
43
|
-
> **For AI agents:** [`llms.txt`](./llms.txt) is a curated map of these docs (
|
|
44
|
-
>
|
|
45
|
-
>
|
|
46
|
-
> ships in the npm package (so it's available wherever the library is installed); `llms.txt` and
|
|
47
|
-
> `AGENTS.md` live in the repo (and `llms.txt` belongs at the root of the docs site).
|
|
43
|
+
> **For AI agents:** [`llms.txt`](./llms.txt) is a curated map of these docs (it belongs at the root
|
|
44
|
+
> of the docs site); [`AGENTS.md`](./AGENTS.md) covers build/test commands and the conventions for
|
|
45
|
+
> working **in** this repository.
|
|
48
46
|
|
|
49
47
|
---
|
|
50
48
|
|
|
@@ -55,7 +53,7 @@ and emits a single, unified stream of typed events the application can react to.
|
|
|
55
53
|
3. [Data flow](#data-flow)
|
|
56
54
|
4. [Entity hierarchy](#entity-hierarchy)
|
|
57
55
|
5. [Ingestion: `accept()`, context & lifecycle](#ingestion-accept-context--lifecycle)
|
|
58
|
-
6. [
|
|
56
|
+
6. [When things update](#when-things-update)
|
|
59
57
|
7. [The event bus](#the-event-bus) ← the core of the API
|
|
60
58
|
8. [API reference](#api-reference)
|
|
61
59
|
9. [Schema types (`ClientSample`)](#schema-types-clientsample)
|
|
@@ -65,8 +63,9 @@ and emits a single, unified stream of typed events the application can react to.
|
|
|
65
63
|
13. [Sinks (per-client sample persistence)](#sinks-per-client-sample-persistence)
|
|
66
64
|
14. [Injecting data into a client](#injecting-data-into-a-client)
|
|
67
65
|
15. [Logging](#logging)
|
|
68
|
-
16. [
|
|
69
|
-
17. [
|
|
66
|
+
16. [Design notes](#design-notes)
|
|
67
|
+
17. [Error-handling philosophy](#error-handling-philosophy)
|
|
68
|
+
18. [Development & extension guide](#development--extension-guide)
|
|
70
69
|
|
|
71
70
|
---
|
|
72
71
|
|
|
@@ -105,10 +104,9 @@ import { Observer, ClientSample } from '@observertc/observer-js';
|
|
|
105
104
|
|
|
106
105
|
// 1. Create an observer.
|
|
107
106
|
const observer = new Observer({
|
|
108
|
-
// when the observer
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
defaultCallUpdatePolicy: 'update-on-any-client-updated',
|
|
107
|
+
// a call updates when any of its clients does, and the observer when any of its calls does —
|
|
108
|
+
// both default to true, so this line is only here to show the knob exists:
|
|
109
|
+
autoUpdateOnCallUpdate: true,
|
|
112
110
|
// optional auto-teardown:
|
|
113
111
|
closeCallIfEmptyForMs: 20_000,
|
|
114
112
|
closeClientIfIdleForMs: 60_000,
|
|
@@ -292,30 +290,38 @@ These return `undefined` (and warn) when the parent is closed; `createObservedCa
|
|
|
292
290
|
|
|
293
291
|
---
|
|
294
292
|
|
|
295
|
-
##
|
|
293
|
+
## When things update
|
|
296
294
|
|
|
297
|
-
"Update" means *recompute aggregated metrics and emit the `*-updated` event* at
|
|
298
|
-
|
|
299
|
-
is no built-in timer. An app that wants a fixed cadence can call `observer.update()` /
|
|
300
|
-
`call.update()` from its own `setInterval`. With `'none'`, **nothing auto-updates** — the level
|
|
301
|
-
updates only when the application calls the public `update()` itself.
|
|
295
|
+
"Update" means *recompute aggregated metrics, run the detectors, and emit the `*-updated` event* at
|
|
296
|
+
that level. Updates are **event-driven** — there is no built-in timer.
|
|
302
297
|
|
|
303
|
-
|
|
298
|
+
The rule is structural rather than configurable:
|
|
304
299
|
|
|
305
|
-
|
|
306
|
-
|
|
307
|
-
|
|
308
|
-
| `update-when-all-call-updated` | every call has updated since the last observer update |
|
|
309
|
-
| `none` | never automatically — only when the app calls `observer.update()` |
|
|
300
|
+
> **A call is updated when any of its clients is updated. The observer is updated when any of its
|
|
301
|
+
> calls is updated.** Composed, that means the observer is updated exactly when any client anywhere
|
|
302
|
+
> is updated.
|
|
310
303
|
|
|
311
|
-
|
|
312
|
-
`ObserverConfig.defaultCallUpdatePolicy`):
|
|
304
|
+
Two booleans, both defaulting to `true`, let you opt out of a link in that chain:
|
|
313
305
|
|
|
314
|
-
|
|
|
315
|
-
|
|
316
|
-
| `
|
|
317
|
-
| `
|
|
318
|
-
|
|
306
|
+
| Setting | Where | Effect when `false` |
|
|
307
|
+
|---------|-------|---------------------|
|
|
308
|
+
| `autoUpdateOnClientUpdate` | `ObservedCallSettings` | the call updates only when you call `call.update()` |
|
|
309
|
+
| `autoUpdateOnCallUpdate` | `ObserverConfig` | the observer updates only when you call `observer.update()` |
|
|
310
|
+
|
|
311
|
+
An app that wants a fixed cadence sets both to `false` and drives `observer.update()` from its own
|
|
312
|
+
`setInterval`. Note that **observer-scoped detectors and validators run nowhere else** — if the
|
|
313
|
+
observer never updates, they never run.
|
|
314
|
+
|
|
315
|
+
```ts
|
|
316
|
+
const observer = new Observer({ autoUpdateOnCallUpdate: false });
|
|
317
|
+
|
|
318
|
+
setInterval(() => observer.update(), 5_000);
|
|
319
|
+
```
|
|
320
|
+
|
|
321
|
+
> Earlier versions had an `updatePolicy` / `defaultCallUpdatePolicy` enum (`'update-on-any-…'`,
|
|
322
|
+
> `'update-when-all-…'`, `'none'`) and a pluggable `Updater` object. Both are gone. "When all clients
|
|
323
|
+
> have updated" sounds appealing and deadlocks on the first client that stops sending — one silent
|
|
324
|
+
> participant froze the whole call's aggregation until it timed out.
|
|
319
325
|
|
|
320
326
|
---
|
|
321
327
|
|
|
@@ -334,7 +340,7 @@ contains the ancestry from the observer down to the entity that raised it, plus
|
|
|
334
340
|
specific subject:
|
|
335
341
|
|
|
336
342
|
```ts
|
|
337
|
-
type ObserverEventBase = { observer: Observer };
|
|
343
|
+
type ObserverEventBase = { observer: Observer, context?: AcceptContext };
|
|
338
344
|
type ObservedCallScope = ObserverEventBase & { observedCall: ObservedCall };
|
|
339
345
|
type ObservedClientScope = ObservedCallScope & { observedClient: ObservedClient };
|
|
340
346
|
type ObservedPeerConnectionScope = ObservedClientScope & { observedPeerConnection: ObservedPeerConnection };
|
|
@@ -360,9 +366,11 @@ additional field(s) on top of that scope.
|
|
|
360
366
|
|
|
361
367
|
| Event | Extra payload | Fires when |
|
|
362
368
|
|-------|---------------|-----------|
|
|
363
|
-
| `observer-updated` | — | `observer.update()` ran (
|
|
369
|
+
| `observer-updated` | — | `observer.update()` ran (see [When things update](#when-things-update)) |
|
|
364
370
|
| `observer-closed` | — | `observer.close()` |
|
|
365
371
|
| `sample-rejected` | `{ reason: 'observer-closed' \| 'missing-callId' \| 'missing-clientId', sample: ClientSample }` | a sample was dropped by `accept()` |
|
|
372
|
+
| `observer-issue` | `{ issue: ObserverIssue }` | `observer.addIssue(...)` — a cross-call / SFU-wide finding (see [observer-level detectors](#observer-level-detectors-cross-call--sfu-wide)) |
|
|
373
|
+
| `validation-ready` | `{ validator: string, report: ValidationReport }` | a [validator](#validators--one-shot-structural-checks) settled — fires once per check, not per tick |
|
|
366
374
|
|
|
367
375
|
#### Mediasoup level — scope `{ observer, observedMediasoupRouter }`
|
|
368
376
|
|
|
@@ -383,7 +391,7 @@ See [Mediasoup router observation](#mediasoup-router-observation) for the full d
|
|
|
383
391
|
| `call-closed` | — | the call closed |
|
|
384
392
|
| `call-empty` | — | last client left the call |
|
|
385
393
|
| `call-not-empty` | — | first client joined a previously-empty call |
|
|
386
|
-
| `call-issue` | `{ issue:
|
|
394
|
+
| `call-issue` | `{ issue: ObserverIssue }` | `call.addIssue(...)` (server-side detector finding) |
|
|
387
395
|
|
|
388
396
|
#### Client level — scope `{ observer, observedCall, observedClient }`
|
|
389
397
|
|
|
@@ -396,7 +404,8 @@ See [Mediasoup router observation](#mediasoup-router-observation) for the full d
|
|
|
396
404
|
| `client-joined` | — | first `CLIENT_JOINED` event seen |
|
|
397
405
|
| `client-left` | — | `CLIENT_LEFT` seen (or inferred on close) |
|
|
398
406
|
| `client-rejoined` | `{ timestamp: number }` | a later `CLIENT_JOINED` after an earlier join |
|
|
399
|
-
| `client-issue` | `{ issue: ClientIssue }` | a client-reported issue arrived, or `client.addIssue(...)` |
|
|
407
|
+
| `client-issue` | `{ issue: ClientIssue }` | a client-reported issue arrived, or `client.addIssue(...)`. A keyed issue also opens an entry in `observedClient.activeIssues` |
|
|
408
|
+
| `client-issue-resolved` | `{ resolvedIssue: ResolvedActiveClientIssue }` | a stateful issue ended — the client sent its `<type>-resolved` companion, or the observer force-closed it. Carries the finished interval (`durationInMs`, `resolvedBy`) — see [client issues](#client-issues-the-lifecycle-and-the-division-of-labour) |
|
|
400
409
|
| `client-metadata` | `{ metaData: ClientMetaData }` | a client meta item arrived |
|
|
401
410
|
| `client-extension-stats` | `{ extensionStats: ExtensionStat }` | an app-defined extension stat arrived |
|
|
402
411
|
| `client-event` | `{ event: ClientEvent }` | any client event was processed |
|
|
@@ -449,8 +458,8 @@ listen to them, but prefer the bus equivalents above for application logic.
|
|
|
449
458
|
new Observer<AppData>(config?: ObserverConfig<AppData>)
|
|
450
459
|
|
|
451
460
|
type ObserverConfig<AppData = Record<string, unknown>> = {
|
|
452
|
-
|
|
453
|
-
|
|
461
|
+
// a call updates when any client does; the observer when any call does. Default true.
|
|
462
|
+
autoUpdateOnCallUpdate?: boolean;
|
|
454
463
|
appData?: AppData;
|
|
455
464
|
closeClientIfIdleForMs?: number;
|
|
456
465
|
closeCallIfEmptyForMs?: number;
|
|
@@ -487,7 +496,15 @@ Key members:
|
|
|
487
496
|
- `createObservedCall<T>(settings): ObservedCall<T> | undefined`
|
|
488
497
|
- `getOrCreateObservedCall<T>(settings): ObservedCall<T> | undefined`
|
|
489
498
|
- `update(): void` — force an aggregation/`observer-updated` tick
|
|
499
|
+
- `addObserverDetector(name, config?): this` — build a cross-call detector onto `observer.detectors`
|
|
500
|
+
- `addCallDetector(name, config?): this` / `removeCallDetector(name): this` — register a call-scoped
|
|
501
|
+
detector for every call created **from now on**
|
|
502
|
+
- `addValidator(name, config?): this` — start a one-shot structural check
|
|
490
503
|
- `close(): void`
|
|
504
|
+
- `readonly detectors: Detectors` — observer-scoped registry. **Starts empty**; nothing is implicit
|
|
505
|
+
- `readonly callDetectorConfigs: Map<name, config>` — what `addCallDetector` recorded
|
|
506
|
+
- `readonly validators: Set<RunningValidator>` — normally empty; each removes itself on finishing
|
|
507
|
+
- `readonly activeIssuesRegistry: ActiveIssuesRegistry` — the fleet's open client issues
|
|
491
508
|
- `readonly observedCalls: Map<string, ObservedCall>`
|
|
492
509
|
- `readonly observedTURN: ObservedTURN`
|
|
493
510
|
- `get appData()`, `get numberOfCalls()`
|
|
@@ -500,7 +517,8 @@ Key members:
|
|
|
500
517
|
|
|
501
518
|
```ts
|
|
502
519
|
type ObservedCallSettings<AppData = Record<string, unknown>> = {
|
|
503
|
-
|
|
520
|
+
// update this call whenever one of its clients accepts a sample. Default true.
|
|
521
|
+
autoUpdateOnClientUpdate?: boolean;
|
|
504
522
|
callId: string;
|
|
505
523
|
appData?: AppData;
|
|
506
524
|
closeCallIfEmptyForMs?: number;
|
|
@@ -512,8 +530,11 @@ Key members:
|
|
|
512
530
|
- `readonly callId: string`, `appData: AppData`
|
|
513
531
|
- `readonly observedClients: Map<string, ObservedClient>`, `get numberOfClients()`
|
|
514
532
|
- `getObservedClient<T>(clientId)`, `createObservedClient<T>(settings)`, `getOrCreateObservedClient<T>(settings)` (all `… | undefined`)
|
|
515
|
-
- `addIssue(issue:
|
|
533
|
+
- `addIssue(issue: ObserverIssue): void` — raise a **call-level** issue → emits `call-issue`
|
|
534
|
+
- `addDetector(name, config?): this` — build a call-scoped detector onto this call only
|
|
516
535
|
- `readonly detectors: Detectors` — server-side detector registry (empty by default; see [Detectors](#detectors-server-side-extension-point))
|
|
536
|
+
- `readonly activeIssuesRegistry: ActiveIssuesRegistry` — this call's open client issues, propagating into the observer's
|
|
537
|
+
- `readonly unconsumedOutboundTracks: Set<ObservedOutboundTrack>` — maintained by the resolver
|
|
517
538
|
- `scoreCalculator: ScoreCalculator`, `get score()`, `readonly calculatedScore`
|
|
518
539
|
- `remoteTrackResolver?: RemoteTrackResolver` — set from `ObserverConfig.createRemoteTrackResolver` at call creation (see [Remote track resolution](#remote-track-resolution-mediasoup--sfu))
|
|
519
540
|
- aggregates: `numberOfIssues`, `numberOfPeerConnections`, `numberOfInboundRtpStreams`,
|
|
@@ -562,11 +583,28 @@ Key members:
|
|
|
562
583
|
`iceCandidates`, `iceCandidatePairs`, `certificates`, `selectedIceCandidatePairs`,
|
|
563
584
|
`selectedIceCandiadtePairForTurn`
|
|
564
585
|
- State: `connectionState?`, `iceConnectionState?`, `iceGatheringState?`, `usingTURN`, `usingTCP`
|
|
565
|
-
- Metrics: `currentRttInMs?`, `
|
|
566
|
-
`availableOutgoingBitrate`, sending/receiving bitrates, packet rates,
|
|
567
|
-
byte/packet counters
|
|
586
|
+
- Metrics: `currentRttInMs?`, `iceRttInMs?`, `rtcpRttInMs?`, `sfuHopRttInMs?`, `currentJitter?`,
|
|
587
|
+
`availableIncomingBitrate`, `availableOutgoingBitrate`, sending/receiving bitrates, packet rates,
|
|
588
|
+
and `total*` / `delta*` byte/packet counters
|
|
568
589
|
- `accept(pcSample, context?)`, `close()`, `get score()`
|
|
569
590
|
|
|
591
|
+
**Two different round trips — don't mix them.** `iceRttInMs` comes from ICE/STUN consent checks and
|
|
592
|
+
measures the trip to *whatever terminates ICE*: **in an SFU topology that is the SFU**, so it is the
|
|
593
|
+
client↔SFU leg. `rtcpRttInMs` comes from RTCP receiver reports and is an **end-to-end** media-path
|
|
594
|
+
round trip. They are not interchangeable, and averaging them together produces a number that moves
|
|
595
|
+
as streams come and go for reasons unrelated to the network. `currentRttInMs` therefore *prefers*
|
|
596
|
+
RTCP and falls back to ICE — always one kind within a tick, never a blend. `sfuHopRttInMs`
|
|
597
|
+
(`rtcp − ice`) estimates everything past the SFU, which separates "this client's last mile is slow"
|
|
598
|
+
from "the path beyond the SFU is slow".
|
|
599
|
+
|
|
600
|
+
**Counter-reset boundaries.** Chrome resets an SSRC's cumulative counters when the codec switches
|
|
601
|
+
([crbug/webrtc/5361](https://bugs.chromium.org/p/webrtc/issues/detail?id=5361), open since 2015),
|
|
602
|
+
which otherwise shows up as a sawtooth spike or a negative bitrate. `ObservedInboundRtp` /
|
|
603
|
+
`ObservedOutboundRtp` therefore set **`counterResetBoundary`** on any tick where `codecId`,
|
|
604
|
+
`encoder`/`decoderImplementation` or `scalabilityMode` changed, and suppress every delta for that
|
|
605
|
+
tick. Without this, a room-wide codec rollout fires a synchronized fake-degradation alert across
|
|
606
|
+
every participant at once.
|
|
607
|
+
|
|
570
608
|
**Remote-RTP correlation (derived).** During `accept()`, receiver/sender reports are linked
|
|
571
609
|
to the local streams by `remoteId` (fallback SSRC) and surfaced as fields:
|
|
572
610
|
|
|
@@ -701,12 +739,28 @@ snapshot happens only once.
|
|
|
701
739
|
|
|
702
740
|
## Detectors (server-side extension point)
|
|
703
741
|
|
|
704
|
-
`observer-js`
|
|
705
|
-
|
|
706
|
-
|
|
707
|
-
|
|
742
|
+
`observer-js` ships [ten detectors](#built-in-detectors), each an opt-in extension you register
|
|
743
|
+
explicitly with `addObserverDetector` / `addCallDetector` / `addDetector` (see
|
|
744
|
+
[Registering detectors](#registering-detectors)) — none are created automatically. All of them
|
|
745
|
+
correlate **across** the clients of a call or the calls of a fleet, because that is the only thing a
|
|
746
|
+
server can do better than a browser: per-client signals — packet loss, jitter, RTT, freezes — are
|
|
747
|
+
already detected on the client and arrive on samples as `clientIssues` (surfaced via `client-issue`).
|
|
708
748
|
|
|
709
|
-
|
|
749
|
+
Findings are raised as **`ObserverIssue`** — `{ type, timestamp, payload? }` — and the payload is the
|
|
750
|
+
**object**, not a JSON string. A server-raised finding is delivered to an in-process handler, so
|
|
751
|
+
there is nothing to serialise for:
|
|
752
|
+
|
|
753
|
+
```ts
|
|
754
|
+
observer.on('call-issue', ({ observedCall, issue }) => {
|
|
755
|
+
issue.payload; // the object; no JSON.parse
|
|
756
|
+
issuePayloadOf(issue); // if you want to accept a string payload too
|
|
757
|
+
issuePayloadAsString(issue); // only at an edge that needs text (log, HTTP, queue)
|
|
758
|
+
});
|
|
759
|
+
```
|
|
760
|
+
|
|
761
|
+
(`ClientIssue`, the type on samples, keeps its string payload — that one really is a wire format.)
|
|
762
|
+
|
|
763
|
+
The registry is also an open extension point, on **`ObservedCall`**:
|
|
710
764
|
|
|
711
765
|
```ts
|
|
712
766
|
import { Observer, Detector } from '@observertc/observer-js';
|
|
@@ -730,21 +784,462 @@ observer.on('call-added', ({ observedCall }) => {
|
|
|
730
784
|
observer.on('call-issue', ({ observedCall, issue }) => { /* react */ });
|
|
731
785
|
```
|
|
732
786
|
|
|
733
|
-
|
|
787
|
+
### Client issues: the lifecycle, and the division of labour
|
|
788
|
+
|
|
789
|
+
The most important thing to understand about detection in this library is **what it deliberately
|
|
790
|
+
does not do**. A client running
|
|
791
|
+
[`client-monitor-js`](https://github.com/ObserveRTC/client-monitor-js) already ships ~20 detectors
|
|
792
|
+
that decide *what is wrong with that endpoint* — `congestion`, `cpulimitation`, `audio-concealment`,
|
|
793
|
+
`freezed-video-track`, `keyframe-storm`, `video-decoder-overloaded`, `stuck-decoder`,
|
|
794
|
+
`ice-disconnected`, and so on. Those verdicts are better than anything re-derived from raw counters
|
|
795
|
+
server-side, because they carry hysteresis and multi-signal confirmation: `audio-concealment`
|
|
796
|
+
subtracts silent concealment (raw `concealedSamples` rises during ordinary silence, so a naive
|
|
797
|
+
detector flags every quiet moment); `audio-jitter-buffer-stress` requires the buffer to be grown
|
|
798
|
+
**and** NetEQ to be time-stretching (a grown buffer alone means NetEQ is *succeeding*);
|
|
799
|
+
`ice-disconnected` only fires once `disconnected` has persisted, so the blips ICE heals on its own
|
|
800
|
+
never surface.
|
|
801
|
+
|
|
802
|
+
**observer-js does not repeat that work.** Its job is the question no browser can answer: *who else
|
|
803
|
+
is in this state right now, what do they have in common, and where in publisher → SFU → subscriber
|
|
804
|
+
does the fault begin?*
|
|
805
|
+
|
|
806
|
+
#### The wire format
|
|
807
|
+
|
|
808
|
+
From client-monitor-js **4.6.0** the whole issue lifecycle reaches the server. A stateful issue
|
|
809
|
+
arrives as two `clientIssues[]` entries sharing a `key`:
|
|
810
|
+
|
|
811
|
+
```
|
|
812
|
+
raise: { type: 'stuck-decoder', key, payload, timestamp: raisedAt }
|
|
813
|
+
resolution: { type: 'stuck-decoder-resolved', key, payload: { raisedAt, comment, …final }, timestamp: resolvedAt }
|
|
814
|
+
```
|
|
815
|
+
|
|
816
|
+
The observer opens an entry in `observedClient.activeIssues` on the raise and closes it on the
|
|
817
|
+
matching key, emitting **`client-issue-resolved`** with the finished interval. Handled for you:
|
|
818
|
+
|
|
819
|
+
- the `-resolved` **suffix** is stripped, so both entries share one logical `type`;
|
|
820
|
+
- a **re-raise** of a live key refreshes the payload without restarting `raisedAt`;
|
|
821
|
+
- **keyless** entries are one-shot — reported via `client-issue`, never tracked;
|
|
822
|
+
- issues still open when a client closes are **force-resolved** (`resolvedBy: 'client-closed'`), and
|
|
823
|
+
the registry additionally expires stale entries, so a crashed participant can't leave an issue
|
|
824
|
+
"active" forever.
|
|
825
|
+
|
|
826
|
+
#### Why intervals beat windows
|
|
827
|
+
|
|
828
|
+
This turns point-in-time symptom reports into **intervals**, and that is the whole game. *"Several
|
|
829
|
+
clients reported congestion in the last 10 seconds"* is a heuristic that has to guess whether the
|
|
830
|
+
symptoms are still happening. *"Several clients are congested **right now, simultaneously**"* is
|
|
831
|
+
ground truth, because the client says when the episode ends. Overlapping intervals are far stronger
|
|
832
|
+
evidence of a shared cause than near-in-time reports.
|
|
833
|
+
|
|
834
|
+
```ts
|
|
835
|
+
observer.on('client-issue', ({ observedClient, issue }) => { /* opened (or one-shot) */ });
|
|
836
|
+
observer.on('client-issue-resolved', ({ resolvedIssue }) => {
|
|
837
|
+
resolvedIssue.type; // 'stuck-decoder' — suffix stripped
|
|
838
|
+
resolvedIssue.durationInMs; // how long the episode lasted
|
|
839
|
+
resolvedIssue.resolvedBy; // 'client' | 'timeout' | 'client-closed'
|
|
840
|
+
});
|
|
841
|
+
|
|
842
|
+
// the live per-client mirror
|
|
843
|
+
observedClient.activeIssues; // ObservedClientIssueRegistry, keyed by issue.key
|
|
844
|
+
```
|
|
845
|
+
|
|
846
|
+
> **client-monitor-js >= 4.6.0 is required** for every issue-driven detector. There is no fallback
|
|
847
|
+
> path that infers these conditions from raw counters — the client decides better, and maintaining a
|
|
848
|
+
> worse second implementation to be polite to old clients is how both end up wrong. Issues without a
|
|
849
|
+
> `key` have no lifecycle (nothing could ever close them), so they stay one-shot: reported on
|
|
850
|
+
> `client-issue`, never registered.
|
|
851
|
+
|
|
852
|
+
#### `ActiveIssuesRegistry` — issues are **pushed**, not polled
|
|
853
|
+
|
|
854
|
+
A detector does not go looking for the issues it cares about. It implements `ActiveIssueTracker` and
|
|
855
|
+
registers for the types it consumes; the registry hands them over as they open and close.
|
|
856
|
+
|
|
857
|
+
```ts
|
|
858
|
+
observedCall.activeIssuesRegistry // this meeting
|
|
859
|
+
observer.activeIssuesRegistry // the fleet; every call's registry propagates into it
|
|
860
|
+
|
|
861
|
+
observer.activeIssuesRegistry.addIssueTracker('congestion', myDetector);
|
|
862
|
+
observer.activeIssuesRegistry.removeIssueTracker(myDetector);
|
|
863
|
+
|
|
864
|
+
registry.values(); // the open issues in this scope, oldest first
|
|
865
|
+
registry.size; // how many
|
|
866
|
+
```
|
|
867
|
+
|
|
868
|
+
The cost of a detector is then proportional to the issues it actually receives, not to the number of
|
|
869
|
+
participants: a healthy 500-client fleet does no per-tick work at all, because nothing was pushed.
|
|
870
|
+
|
|
871
|
+
**There is no wildcard.** A tracker names its types and sees nothing else. "Feed me everything and
|
|
872
|
+
I'll work out what matters" moves the decision from the application — which knows its client build
|
|
873
|
+
and its issue vocabulary — onto a detector that has to guess, and it makes the cost of a
|
|
874
|
+
subscription unbounded and invisible. If a detector should watch five types, list five types.
|
|
875
|
+
|
|
876
|
+
Onset spread is measured on the **observer clock**, never the client's. `raisedAt` comes from each
|
|
877
|
+
participant's own machine, and comparing those across clients makes clock skew look like a
|
|
878
|
+
synchronized infrastructure event.
|
|
879
|
+
|
|
880
|
+
### Observer-level detectors (cross-call / SFU-wide)
|
|
881
|
+
|
|
882
|
+
Some findings only exist **above** call scope — "many calls on the same SFU degraded at once" is far
|
|
883
|
+
more actionable than fifty individual client alerts. The same detector registry exists on the `Observer`,
|
|
884
|
+
runs on every `observer.update()`, and raises findings through `observer.addIssue(...)`, surfaced on
|
|
885
|
+
the bus as **`observer-issue`**:
|
|
886
|
+
|
|
887
|
+
```ts
|
|
888
|
+
observer.detectors.add({
|
|
889
|
+
name: 'sfu-wide-degradation',
|
|
890
|
+
update: () => {
|
|
891
|
+
const degradedCalls = [ ...observer.observedCalls.values() ].filter(isDegraded);
|
|
892
|
+
|
|
893
|
+
if (observer.numberOfCalls > 3 && degradedCalls.length / observer.numberOfCalls > 0.6) {
|
|
894
|
+
observer.addIssue({ type: 'SFU_WIDE_QUALITY_DEGRADATION', timestamp: Date.now() });
|
|
895
|
+
}
|
|
896
|
+
},
|
|
897
|
+
});
|
|
898
|
+
|
|
899
|
+
observer.on('observer-issue', ({ issue }) => alert(issue));
|
|
900
|
+
```
|
|
901
|
+
|
|
902
|
+
### Publisher → subscribers: the resolver links
|
|
903
|
+
|
|
904
|
+
The question a single browser can never answer is *"did **everyone** receiving Alice see the same
|
|
905
|
+
degradation?"*. The join is the publisher↔subscriber links maintained by a
|
|
906
|
+
[`RemoteTrackResolver`](#remote-track-resolution-mediasoup--sfu), and detectors walk them directly:
|
|
907
|
+
|
|
908
|
+
```ts
|
|
909
|
+
outboundTrack.remoteInboundTracks; // Set<ObservedInboundTrack> — every subscriber of this source
|
|
910
|
+
inboundTrack.remoteOutboundTrack; // the publisher, or undefined if unlinked
|
|
911
|
+
inboundTrack.getInboundRtp(); // that receiver's RTP stats
|
|
912
|
+
observedCall.unconsumedOutboundTracks; // published tracks with no subscriber at all
|
|
913
|
+
```
|
|
914
|
+
|
|
915
|
+
> A `TrackDistributionAggregator` class used to sit in front of these links and summarise every
|
|
916
|
+
> published track against all of its receivers, on every tick. It is gone. It scanned the majority
|
|
917
|
+
> (all published tracks) to find the interesting minority, which is the wrong axis — the detectors
|
|
918
|
+
> now start from the handful of *affected* tracks the issue registry pushed at them and resolve only
|
|
919
|
+
> those. The statistics helpers it used (`percentile`, `median`, `summarize`, `counterDelta`,
|
|
920
|
+
> `robustZScore`, `SlidingWindow`, `TrendTester`) are all still exported for building your own.
|
|
921
|
+
|
|
922
|
+
### Call health: `CallHealthAggregator`
|
|
923
|
+
|
|
924
|
+
The **client** axis. Where the resolver links answer "how was *this source* delivered?", this asks
|
|
925
|
+
"how is *each participant* doing, sending vs receiving?":
|
|
926
|
+
|
|
927
|
+
```ts
|
|
928
|
+
import { CallHealthAggregator } from '@observertc/observer-js';
|
|
929
|
+
|
|
930
|
+
const health = new CallHealthAggregator(observedCall).aggregate();
|
|
931
|
+
|
|
932
|
+
health.degradedRatio; // 0.82 — the number that distinguishes shared faults from individual ones
|
|
933
|
+
health.inboundDegradedRatio; // receiving side → egress/downstream suspicion
|
|
934
|
+
health.outboundDegradedRatio; // sending side → ingress suspicion
|
|
935
|
+
health.rttInMs?.median; // percentile rollups, never means
|
|
936
|
+
health.qualityLimitation; // { cpu, bandwidth, other } client counts
|
|
937
|
+
health.clients; // per-client entries with `reasons`, direction flags, TURN/TCP
|
|
938
|
+
```
|
|
939
|
+
|
|
940
|
+
### Registering detectors
|
|
941
|
+
|
|
942
|
+
**Nothing is created implicitly.** A `new Observer()` has zero detectors. There is no detector
|
|
943
|
+
configuration in `ObserverConfig` and no default set — an application says what it wants to watch, or
|
|
944
|
+
it watches nothing.
|
|
945
|
+
|
|
946
|
+
```ts
|
|
947
|
+
const observer = new Observer({
|
|
948
|
+
createRemoteTrackResolver: createDefaultMediasoupRemoteTrackResolverFactory(),
|
|
949
|
+
});
|
|
950
|
+
|
|
951
|
+
// observer-scoped (cross-call) — built immediately onto `observer.detectors`
|
|
952
|
+
observer.addObserverDetector('observer-concurrent-issue-detector', {
|
|
953
|
+
issueTypes: [ 'congestion', 'ice-disconnected', 'ice-connection-failed' ],
|
|
954
|
+
minAffectedCalls: 3,
|
|
955
|
+
});
|
|
956
|
+
observer.addObserverDetector('turn-server-outage-detector', { minClientsAtPeak: 10 });
|
|
957
|
+
|
|
958
|
+
// call-scoped — recorded in `observer.callDetectorConfigs`, applied to every call created AFTER this
|
|
959
|
+
observer.addCallDetector('call-concurrent-issue-detector', {
|
|
960
|
+
issueTypes: [ 'congestion', 'ice-disconnected' ],
|
|
961
|
+
});
|
|
962
|
+
observer.removeCallDetector('unconsumed-track-detector'); // undo, for future calls
|
|
963
|
+
|
|
964
|
+
// one specific call
|
|
965
|
+
observedCall.addDetector('issue-fan-out-detector', { issueTypes: [ 'freezed-video-track' ] });
|
|
966
|
+
```
|
|
967
|
+
|
|
968
|
+
Detectors are named by their kebab-case `NAME`, and the name types the config — an unknown name or a
|
|
969
|
+
key that belongs to a different detector will not compile. Each detector owns its defaults in its own
|
|
970
|
+
constructor, beside the doc explaining what each threshold means; there is no central table to keep
|
|
971
|
+
in sync.
|
|
972
|
+
|
|
973
|
+
> **Why no defaults?** A detector nobody asked for is a detector nobody will act on. It costs time on
|
|
974
|
+
> every tick and raises findings into a handler that was not written to expect them. Earlier versions
|
|
975
|
+
> auto-created everything from a three-state config slot; the result was applications receiving
|
|
976
|
+
> finding types they had never heard of.
|
|
977
|
+
|
|
978
|
+
Every issue-driven detector takes an explicit, non-empty `issueTypes` (or
|
|
979
|
+
`publisherIssueTypes`/`receiverIssueTypes`). There is no "watch everything" option — see
|
|
980
|
+
[the registry](#activeissuesregistry--issues-are-pushed-not-polled).
|
|
981
|
+
|
|
982
|
+
> **🔗 marks detectors that require a
|
|
983
|
+
> [`RemoteTrackResolver`](#remote-track-resolution-mediasoup--sfu).** They reason about a published
|
|
984
|
+
> track and its subscribers, so without the publisher↔subscriber links they see nothing and stay
|
|
985
|
+
> **silent forever** — which looks exactly like "no problems found". Configure
|
|
986
|
+
> `ObserverConfig.createRemoteTrackResolver`, and start the
|
|
987
|
+
> [`remote-track-resolver` validator](#validators--one-shot-structural-checks) to prove it is wired.
|
|
988
|
+
|
|
989
|
+
### Built-in detectors
|
|
990
|
+
|
|
991
|
+
They consume the verdicts `client-monitor-js` >= 4.6.0 already ships (raise + `<type>-resolved`) and
|
|
992
|
+
add only the cross-participant conclusion. **None re-derives a per-endpoint verdict from raw
|
|
993
|
+
counters** — that is the rule the whole design hangs on:
|
|
994
|
+
|
|
995
|
+
> **If a condition is detectable on the client, the client's issue is the source of truth.**
|
|
996
|
+
|
|
997
|
+
| Detector | 🔗 | Scope | Raises |
|
|
998
|
+
|----------|:--:|-------|--------|
|
|
999
|
+
| `CallConcurrentIssueDetector` | | call | `CONCURRENT_CLIENT_ISSUES`, `ISSUE_ONSET_BURST` |
|
|
1000
|
+
| `ObserverConcurrentIssueDetector` | | **observer** | `CROSS_CALL_CONCURRENT_ISSUES`, `CROSS_CALL_ISSUE_ONSET_BURST` |
|
|
1001
|
+
| `IssueFanOutDetector` | 🔗 | call | `PUBLISHED_TRACK_ISSUE_FAN_OUT`, `SINGLE_RECEIVER_ISSUE` |
|
|
1002
|
+
| `PublisherFaultCorroborationDetector` | 🔗 | call | `CORROBORATED_PUBLISHER_FAULT` |
|
|
1003
|
+
| `TrackDeliveryMismatchDetector` | 🔗 | call | `PUBLISHED_TRACK_NOT_DELIVERED`, `RECEIVER_TRACK_NOT_DELIVERED`, `PUBLISHER_TRACK_DRY` |
|
|
1004
|
+
| `UnconsumedTrackDetector` | 🔗 | call | `UNCONSUMED_PUBLISHED_TRACK` |
|
|
1005
|
+
| `ClientPopulationIssueDetector` | | **observer** | `CLIENT_POPULATION_ISSUE` |
|
|
1006
|
+
| `SfuCongestionDetector` | | **observer** | `sfu-congestion` |
|
|
1007
|
+
| `TurnServerHealthDetector` | | **observer** | `TURN_SERVER_DEGRADED` |
|
|
1008
|
+
| `TurnServerOutageDetector` | | **observer** | `TURN_SERVER_OUTAGE` |
|
|
1009
|
+
|
|
1010
|
+
What each adds that no endpoint can know:
|
|
1011
|
+
|
|
1012
|
+
- **`CallConcurrentIssueDetector`** — *who else in this meeting is in this state right now?* The
|
|
1013
|
+
difference between "one person's Wi-Fi" and "this room is broken".
|
|
1014
|
+
- **`ObserverConcurrentIssueDetector`** — *is our infrastructure in trouble?* A **separate class**,
|
|
1015
|
+
not the call one with a bigger denominator, because it is a different question with different
|
|
1016
|
+
gates. It requires the group to span at least `minAffectedCalls` **independent calls** (default
|
|
1017
|
+
`2`) and raises its own `CROSS_CALL_*` types. Without that gate, one thirty-person meeting where
|
|
1018
|
+
everyone is congested clears every client threshold and pages you for a single bad room the
|
|
1019
|
+
call-scoped detector already reported. Clients in different calls share no room, no publisher and
|
|
1020
|
+
no host — only the servers, which is what makes the finding conclusive. Note there is deliberately
|
|
1021
|
+
**no participant ratio** at this scope: six broken calls out of forty is a small share of all
|
|
1022
|
+
clients, and a ratio gate would hide exactly the event you want.
|
|
1023
|
+
- **`IssueFanOutDetector`** — *does this issue follow one published source, or one receiver?*
|
|
1024
|
+
- **`PublisherFaultCorroborationDetector`** — *do **both ends** of one track agree the source is at
|
|
1025
|
+
fault?* Fan-out sees one end and infers; this sees the publisher reporting `encoder-bottleneck`
|
|
1026
|
+
about its own send path *while* its subscribers report `freezed-video-track` about receiving it.
|
|
1027
|
+
Two independent parties, one conclusion, nothing left to deduce — hence the highest confidence in
|
|
1028
|
+
the library. Run both: fan-out is broader and catches the case where the publisher is fine and the
|
|
1029
|
+
SFU's forwarding is not.
|
|
1030
|
+
- **`ClientPopulationIssueDetector`** — *is this concentrated on one **kind of client**?* The one
|
|
1031
|
+
correlation here that is neither per-call nor per-server. Every other observer-scoped detector
|
|
1032
|
+
reasons "clients in unrelated calls share only the infrastructure, so it must be us" — right for
|
|
1033
|
+
network symptoms, **wrong for endpoint ones**. `cpulimitation` across six unrelated calls is not an
|
|
1034
|
+
SFU event; CPU is owned by the endpoint, so what those endpoints share is a browser version or a
|
|
1035
|
+
client release. Groups by `browser` / `engine` / `platform` / `operationSystem`, one axis per
|
|
1036
|
+
instance. The gate is **relative risk**, not share: "30% of Chrome 141 is unhappy" means nothing if
|
|
1037
|
+
30% of everyone is, and a share-based rule simply indicts whichever browser is most popular.
|
|
1038
|
+
- **`SfuCongestionDetector`** — *is congestion spiking across the fleet right now?* Counts distinct
|
|
1039
|
+
clients reporting congestion in fixed wall-clock buckets and compares each bucket against a
|
|
1040
|
+
median+MAD baseline of the ones before it. Buckets rather than update ticks on purpose: the tick is
|
|
1041
|
+
unevenly spaced and shorter than a client's sampling period, so counting on it compares windows of
|
|
1042
|
+
different lengths and calls the difference a signal. Only add it when the observer's calls all come
|
|
1043
|
+
from the **same SFU** — the finding's meaning is "these clients share only that server".
|
|
1044
|
+
- **`TrackDeliveryMismatchDetector`** — *are the two ends of a track disagreeing?*
|
|
1045
|
+
- **`UnconsumedTrackDetector`** — *is anyone actually subscribed?* (reads the resolver's silence)
|
|
1046
|
+
- **`TurnServerHealthDetector`** — *does trouble cluster on one relay?*
|
|
1047
|
+
- **`TurnServerOutageDetector`** — covers the case the health detector structurally cannot. The
|
|
1048
|
+
health detector groups clients by the server relaying them and asks how many report issues — it
|
|
1049
|
+
needs clients *on* the server to ask. When a TURN server dies, allocation fails: existing sessions
|
|
1050
|
+
drop and new clients never obtain a relay candidate through it, so they are never attributed to it
|
|
1051
|
+
at all. Its population goes to zero and the health detector falls silent for the worst possible
|
|
1052
|
+
reason. **Degradation makes clients unhappy; an outage makes them disappear.** Absence is a
|
|
1053
|
+
dangerous signal, so the **control group** is the heart of the design: a call ending, everyone
|
|
1054
|
+
leaving at 6pm, and a fleet-wide network event all look identical to an outage. It refuses to blame
|
|
1055
|
+
a server unless clients *not* relayed through it are demonstrably still connected
|
|
1056
|
+
(`requireControlGroup`, on by default).
|
|
1057
|
+
|
|
1058
|
+
#### There is no ICE detector
|
|
1059
|
+
|
|
1060
|
+
ICE trouble is reported by `client-monitor-js` >= 4.6.0 as the keyed issues `ice-disconnected`,
|
|
1061
|
+
`ice-connection-failed`, `ice-transport-stalled` and `unstable-ice-path`, each with hysteresis and
|
|
1062
|
+
multi-signal confirmation behind it. An `IceDisruptionDetector` used to re-derive that server-side
|
|
1063
|
+
from raw state transitions; it has been removed, because the server sees less and guesses more. The
|
|
1064
|
+
client knows whether `disconnected` persisted or healed in 200 ms; the observer does not.
|
|
1065
|
+
|
|
1066
|
+
Correlating ICE trouble is now configuration, not a class:
|
|
1067
|
+
|
|
1068
|
+
```ts
|
|
1069
|
+
observer.addObserverDetector('observer-concurrent-issue-detector', {
|
|
1070
|
+
issueTypes: [ 'ice-disconnected', 'ice-connection-failed', 'ice-transport-stalled' ],
|
|
1071
|
+
});
|
|
1072
|
+
```
|
|
1073
|
+
|
|
1074
|
+
### Validators — one-shot structural checks
|
|
1075
|
+
|
|
1076
|
+
Every detector above answers *"is something wrong right now?"* and runs on every tick, because the
|
|
1077
|
+
answer legitimately changes. A **validator** answers *"is this deployment built correctly?"* — which
|
|
1078
|
+
only changes when you deploy. So it is not configured on and left running: you **start** one, it runs
|
|
1079
|
+
until it can decide, reports once, and the observer drops it.
|
|
734
1080
|
|
|
735
1081
|
```ts
|
|
736
|
-
|
|
737
|
-
|
|
738
|
-
|
|
739
|
-
|
|
740
|
-
|
|
741
|
-
|
|
742
|
-
|
|
743
|
-
|
|
1082
|
+
observer.addValidator('simulcast-receivers', { minChecks: 5 });
|
|
1083
|
+
|
|
1084
|
+
observer.on('validation-ready', ({ validator, report }) => {
|
|
1085
|
+
if (!report.ready) return;
|
|
1086
|
+
if (report.verdict === 'layer-decided-lowest-common-denominator') page(validator, report);
|
|
1087
|
+
});
|
|
1088
|
+
|
|
1089
|
+
onDeploy(() => observer.addValidator('simulcast-receivers')); // check again
|
|
1090
|
+
```
|
|
1091
|
+
|
|
1092
|
+
`observer.validators` is the set currently running — normally empty, since each removes itself on
|
|
1093
|
+
finishing. There is no revalidation timer: a deploy, not elapsed time, is what makes a structural
|
|
1094
|
+
verdict stale, so re-checking means starting another.
|
|
1095
|
+
|
|
1096
|
+
| Validator | `addValidator` name | Question | Also raises |
|
|
1097
|
+
|-----------|---------------------|----------|-------------|
|
|
1098
|
+
| `SimulcastReceiverValidator` 🔗 | `simulcast-receivers` | Does the SFU pick layers per receiver, or drag the publisher down to the worst one? | `WORST_RECEIVER_CONTAGION` |
|
|
1099
|
+
| `RemoteTrackResolverValidator` | `remote-track-resolver` | Is the resolver actually linking anything? | `REMOTE_TRACK_LINKS_UNRESOLVED` |
|
|
1100
|
+
| `CodecConsistencyValidator` | `codec-consistency` | Is everyone on the same codec — and is it the one you think you negotiated? | `CODEC_INCONSISTENCY` |
|
|
1101
|
+
|
|
1102
|
+
**`SimulcastReceiverValidator`** — simulcast (or SVC) exists so one slow
|
|
1103
|
+
participant doesn't set everyone's quality: with several encodings the server hands the struggling
|
|
1104
|
+
receiver a lower layer and leaves the rest alone. Without it — or with a server that relays RTCP end
|
|
1105
|
+
to end, so the publisher's bandwidth estimate collapses to the slowest receiver — the only way to
|
|
1106
|
+
serve them is to make the *source* send less. Both causes look identical from outside; what the check
|
|
1107
|
+
establishes is whether per-receiver adaptation happens at all.
|
|
1108
|
+
|
|
1109
|
+
| `verdict` | meaning |
|
|
1110
|
+
|-----------|---------|
|
|
1111
|
+
| `layer-decided-per-receiver` | verified — a receiver fell far behind and the publisher carried on |
|
|
1112
|
+
| `layer-decided-lowest-common-denominator` | the publisher followed its worst receiver; everyone gets the slowest participant's quality |
|
|
1113
|
+
| `inconclusive` | cancelled, or the observer closed, before it could decide |
|
|
1114
|
+
|
|
1115
|
+
**Not finishing is not a pass.** The check only runs when a publisher has 3+ receivers and one is at
|
|
1116
|
+
most half the median; plenty of healthy deployments never present that. A validator that never sees it
|
|
1117
|
+
simply keeps running and never reports — it does not quietly succeed. `report.checks` counts the times
|
|
1118
|
+
the check genuinely ran, so an `inconclusive` with `checks: 0` says plainly that nothing was verified.
|
|
1119
|
+
|
|
1120
|
+
**`RemoteTrackResolverValidator`** exists because of a specific, nasty failure mode. Four things here
|
|
1121
|
+
are built on publisher↔subscriber links — `IssueFanOutDetector`,
|
|
1122
|
+
`PublisherFaultCorroborationDetector`, `TrackDeliveryMismatchDetector`, `UnconsumedTrackDetector`
|
|
1123
|
+
(and `SimulcastReceiverValidator`) — and every one of them correctly does *nothing* when the links
|
|
1124
|
+
are missing rather than guessing. So a resolver wired to the wrong id field leaves all of them
|
|
1125
|
+
permanently silent, and **silence is what a healthy deployment looks like too**: you would conclude
|
|
1126
|
+
your calls were clean when in fact nothing was ever examined. Verdicts: `links-resolved` /
|
|
1127
|
+
`no-links-resolved` / `inconclusive`. Run it at start-up and after changing the resolver or the SFU's
|
|
1128
|
+
id scheme.
|
|
1129
|
+
|
|
1130
|
+
**`CodecConsistencyValidator`** answers two things at once. A **split** — participants of one call on
|
|
1131
|
+
different codecs — is a real fault with a confusing symptom: an SFU that forwards without transcoding
|
|
1132
|
+
cannot serve them all, so some pairs see each other and some do not, with no error anywhere. Only
|
|
1133
|
+
something holding every participant at once can see it. The quieter half is the silent fallback: a
|
|
1134
|
+
deployment configured for VP9 or AV1 drops to VP8 whenever one endpoint cannot negotiate the
|
|
1135
|
+
preference, the call keeps working at a higher bitrate than budgeted, and the team believes it
|
|
1136
|
+
shipped AV1 months ago. Give it `expected` and it says so. Verdicts: `codec-consistent` /
|
|
1137
|
+
`codec-split` / `unexpected-codec` / `inconclusive`.
|
|
1138
|
+
|
|
1139
|
+
```ts
|
|
1140
|
+
observer.addValidator('remote-track-resolver');
|
|
1141
|
+
observer.addValidator('codec-consistency', { expected: { video: 'video/VP9', audio: 'audio/opus' } });
|
|
1142
|
+
```
|
|
1143
|
+
|
|
1144
|
+
#### Conclusions
|
|
1145
|
+
|
|
1146
|
+
Every issue-driven finding carries a `conclusion` in its payload — the interpretation step, so the
|
|
1147
|
+
person reading the alert doesn't have to perform it:
|
|
1148
|
+
|
|
1149
|
+
```jsonc
|
|
1150
|
+
{
|
|
1151
|
+
"type": "CROSS_CALL_ISSUE_ONSET_BURST",
|
|
1152
|
+
"issueType": "congestion",
|
|
1153
|
+
"calls": 40, "affectedCalls": 6,
|
|
1154
|
+
"perCall": [ { "callId": "…", "affectedClients": 4, "totalClients": 9 } ],
|
|
1155
|
+
"conclusion": {
|
|
1156
|
+
"faultDomain": "infrastructure",
|
|
1157
|
+
"summary": "network congestion is open across independent calls at the same time — 6 of 40 calls (11/300 clients)",
|
|
1158
|
+
"recommendation": "check SFU egress bandwidth and host network saturation before looking at any single participant",
|
|
1159
|
+
"confidence": 0.85
|
|
1160
|
+
}
|
|
744
1161
|
}
|
|
745
1162
|
```
|
|
746
1163
|
|
|
747
|
-
|
|
1164
|
+
`faultDomain` is one of `infrastructure`, `call`, `published-track`, `endpoint`, `client-population`
|
|
1165
|
+
or `unknown`, and it comes from the **spread**, not the issue type — congestion in one call is a
|
|
1166
|
+
meeting problem, congestion in six calls is a server problem, and the client reported the identical
|
|
1167
|
+
symptom in both.
|
|
1168
|
+
|
|
1169
|
+
One case is worth knowing about because it inverts the usual reading: **`cpu-limitation` spread
|
|
1170
|
+
across many independent calls concludes `client-population`, not `infrastructure`.** Endpoint CPU is
|
|
1171
|
+
owned by the endpoint, so breadth there points at what those endpoints share — a recent client
|
|
1172
|
+
release, a browser version, shared VDI hardware — and paging the SFU on-call would be wrong. The
|
|
1173
|
+
conclusion table encodes that so nobody has to rediscover it during an incident.
|
|
1174
|
+
|
|
1175
|
+
Unknown issue types (your own custom client detectors) still produce a structurally valid conclusion
|
|
1176
|
+
from the spread alone; they just get generic wording.
|
|
1177
|
+
|
|
1178
|
+
Two functions are exported, one per scope: `concludeCallIssue()` and `concludeObserverIssue()`. They
|
|
1179
|
+
are separate because a detector already knows its scope, and a single generic function forced every
|
|
1180
|
+
caller to pass the other scope's fields as placeholders — call-scoped detectors passing
|
|
1181
|
+
`affectedCalls: 1, totalCalls: 1` forever, observer-scoped ones passing a participant ratio that was
|
|
1182
|
+
deliberately never read. Placeholders like that invite being read as if they meant something.
|
|
1183
|
+
|
|
1184
|
+
#### Cost
|
|
1185
|
+
|
|
1186
|
+
Detectors run inside `call.update()`, on your event loop, so their cost matters. Two things keep it
|
|
1187
|
+
off the participant axis:
|
|
1188
|
+
|
|
1189
|
+
- **Issues are pushed, not polled.** A detector holds only what the registry handed it, so an
|
|
1190
|
+
`update()` that finds `size === 0` — the overwhelmingly common case — costs one comparison,
|
|
1191
|
+
whatever the participant count. Nothing iterates clients looking for trouble.
|
|
1192
|
+
- **So are unconsumed tracks.** `observedCall.unconsumedOutboundTracks` is maintained by the resolver
|
|
1193
|
+
as tracks gain and lose subscribers, so `UnconsumedTrackDetector` reads a set that is normally
|
|
1194
|
+
empty instead of walking every published track (529 µs → 65 µs per tick at 1 200 tracks).
|
|
1195
|
+
- **Track lookups start from the affected minority.** A detector resolving an issue to its published
|
|
1196
|
+
track searches the *reporting client's* peer connections (typically one or two), not the call.
|
|
1197
|
+
|
|
1198
|
+
At 20 calls × 12 participants (2 640 subscriptions) the whole detector pass costs ~1.3 ms per tick.
|
|
1199
|
+
`yarn bench` prints a per-detector breakdown for your own shape.
|
|
1200
|
+
|
|
1201
|
+
#### Worked examples
|
|
1202
|
+
|
|
1203
|
+
[`examples/detectors.ts`](./examples/detectors.ts) (`yarn example:detectors`) runs one scenario per
|
|
1204
|
+
detector — the question it answers, its full config, the synthetic traffic that makes it fire, and
|
|
1205
|
+
the finding with its conclusion. It asserts every expected finding is produced, so it doubles as a
|
|
1206
|
+
smoke test. [`examples/sfu-observer.ts`](./examples/sfu-observer.ts) (`yarn example`) is the end-to-end
|
|
1207
|
+
tour instead: ingest → correlate → react, with the mediasoup wiring alongside.
|
|
1208
|
+
|
|
1209
|
+
#### `TrackDeliveryMismatchDetector` — resolving an ambiguous symptom
|
|
1210
|
+
|
|
1211
|
+
A dry track ("no bytes are arriving") is the clearest symptom there is and, on its own, completely
|
|
1212
|
+
ambiguous. A receiver seeing silence cannot distinguish *the camera was switched off* from *the SFU
|
|
1213
|
+
stopped forwarding* from *my own consumer wedged* — all three look identical from the browser.
|
|
1214
|
+
|
|
1215
|
+
Joining the two ends of the published track resolves it:
|
|
1216
|
+
|
|
1217
|
+
| publisher | subscribers | verdict |
|
|
1218
|
+
|---|---|---|
|
|
1219
|
+
| sending | **all** dry | `PUBLISHED_TRACK_NOT_DELIVERED` — the forwarding path |
|
|
1220
|
+
| sending | **some** dry | `RECEIVER_TRACK_NOT_DELIVERED` — those consumers (in mediasoup: recreate them) |
|
|
1221
|
+
| dry | any dry | `PUBLISHER_TRACK_DRY` — the source stopped; **not** an SFU fault |
|
|
1222
|
+
|
|
1223
|
+
The publisher side is judged from both available signals: its own `dry-outbound-track` issue when the
|
|
1224
|
+
client reports one, and the observed outbound RTP (`deltaPacketsSent`) as fallback and corroboration.
|
|
1225
|
+
That combination is what makes the first row trustworthy — the server can state that packets
|
|
1226
|
+
demonstrably left the publisher during the same interval in which every receiver got nothing.
|
|
1227
|
+
|
|
1228
|
+
This is the "SFU forwarding mismatch" check, and it needs **no mediasoup instrumentation at all** —
|
|
1229
|
+
the clients' own dry-track verdicts plus the resolver links are sufficient.
|
|
1230
|
+
|
|
1231
|
+
#### `UnconsumedTrackDetector` — reading the resolver's silence
|
|
1232
|
+
|
|
1233
|
+
The one detector where the *absence* of links is the signal: a track still pushing packets whose
|
|
1234
|
+
`remoteInboundTracks` set is empty, i.e. uplink and SFU ingress spent on media nobody receives
|
|
1235
|
+
(everyone has the publisher hidden, a simulcast layer no viewer selects, or an app that forgot to
|
|
1236
|
+
stop a track). It waits `minUnconsumedDurationInMs` first, since a gap between publishing and the
|
|
1237
|
+
first subscription is normal at join time.
|
|
1238
|
+
|
|
1239
|
+
Note the trap this one has to guard against, and why it checks `call.remoteTrackResolver` at runtime
|
|
1240
|
+
rather than trusting the flag alone: **"no subscribers" and "no resolver configured" produce the
|
|
1241
|
+
identical observation.** Without a resolver it would report every published track in the call as
|
|
1242
|
+
unconsumed.
|
|
748
1243
|
|
|
749
1244
|
## Remote track resolution (mediasoup / SFU)
|
|
750
1245
|
|
|
@@ -836,6 +1331,80 @@ on your own cadence read `observedRouter.sample` (snapshot/serialize/persist wha
|
|
|
836
1331
|
you don't, and close routers you no longer track. (mediasoup also typically shards routers across
|
|
837
1332
|
workers/cores, which keeps any one router small.)
|
|
838
1333
|
|
|
1334
|
+
### Extending the sample, and building your own report
|
|
1335
|
+
|
|
1336
|
+
The sample is yours to annotate. Every entity — the router, each transport, producer, consumer, data
|
|
1337
|
+
producer and data consumer — has an `attachments?: Record<string, unknown>` slot, and there are three
|
|
1338
|
+
ways to fill it, from most declarative to most ad-hoc.
|
|
1339
|
+
|
|
1340
|
+
**1. `enrich` — mirror mediasoup's own `appData`.** The common case: your application already keeps
|
|
1341
|
+
`participantId`, `purpose` and similar on the mediasoup objects, and you want them on the sample.
|
|
1342
|
+
Runs once per entity at creation, before the corresponding event:
|
|
1343
|
+
|
|
1344
|
+
```ts
|
|
1345
|
+
observer.createObservedMediasoupRouter({
|
|
1346
|
+
router,
|
|
1347
|
+
enrich: {
|
|
1348
|
+
producer: (producer) => ({ participantId: producer.appData.participantId, purpose: producer.appData.purpose }),
|
|
1349
|
+
consumer: (consumer) => ({ subscriberId: consumer.appData.subscriberId }),
|
|
1350
|
+
transport: (transport) => ({ role: transport.appData.role }),
|
|
1351
|
+
},
|
|
1352
|
+
});
|
|
1353
|
+
```
|
|
1354
|
+
|
|
1355
|
+
A throwing enricher is caught and logged — it can't take the router's bookkeeping down with it.
|
|
1356
|
+
|
|
1357
|
+
**2. Lifecycle events — enrich on the fly.** Each entity announces itself as
|
|
1358
|
+
`<entity>-sample-added` and `<entity>-sample-closed`, carrying **the live sample object** (not a
|
|
1359
|
+
copy) plus the mediasoup object it came from. Mutating it in the handler is the intended pattern:
|
|
1360
|
+
|
|
1361
|
+
```ts
|
|
1362
|
+
observedRouter.on('producer-sample-added', ({ sample, producer, transport }) => {
|
|
1363
|
+
sample.attachments = { ...sample.attachments, participantId: lookup(producer.id) };
|
|
1364
|
+
});
|
|
1365
|
+
|
|
1366
|
+
observedRouter.on('producer-sample-closed', ({ sample }) => {
|
|
1367
|
+
archive(sample); // its `closedAt` is set
|
|
1368
|
+
});
|
|
1369
|
+
```
|
|
1370
|
+
|
|
1371
|
+
Events: `transport-sample-added` / `-closed`, `producer-sample-added` / `-closed`,
|
|
1372
|
+
`consumer-sample-added` / `-closed`, `data-producer-sample-added` / `-closed`,
|
|
1373
|
+
`data-consumer-sample-added` / `-closed`.
|
|
1374
|
+
|
|
1375
|
+
**3. `attachTo(id, attachments)` — annotate later, from anywhere.** When the knowledge arrives after
|
|
1376
|
+
the entity did (a signalling message, a database lookup that resolved):
|
|
1377
|
+
|
|
1378
|
+
```ts
|
|
1379
|
+
observedRouter.attachTo(producerId, { participantId, joinedFrom: 'mobile' }); // merges
|
|
1380
|
+
```
|
|
1381
|
+
|
|
1382
|
+
Ids are unique across mediasoup entity kinds, so one method covers all of them. It returns `false`
|
|
1383
|
+
for an unknown id rather than failing quietly — which matters when application events race the
|
|
1384
|
+
mediasoup ones. For direct access there are typed accessors: `getTransportSample(id)`,
|
|
1385
|
+
`getProducerSample(id)`, `getConsumerSample(id)`, `getDataProducerSample(id)`,
|
|
1386
|
+
`getDataConsumerSample(id)`. They index the *same* objects the arrays hold, so a lookup is O(1)
|
|
1387
|
+
instead of a `sample.producers.find(...)` scan.
|
|
1388
|
+
|
|
1389
|
+
#### Building your own report
|
|
1390
|
+
|
|
1391
|
+
`observedRouter.sample` is live — arrays grow and `history` entries are appended as the router runs,
|
|
1392
|
+
so a report built directly on it keeps changing after you think you're done. Use **`snapshot()`** for
|
|
1393
|
+
a detached deep copy:
|
|
1394
|
+
|
|
1395
|
+
```ts
|
|
1396
|
+
const report = {
|
|
1397
|
+
...observedRouter.snapshot(), // never moves again
|
|
1398
|
+
generatedAt: Date.now(),
|
|
1399
|
+
region: process.env.REGION,
|
|
1400
|
+
};
|
|
1401
|
+
```
|
|
1402
|
+
|
|
1403
|
+
> **Note on typing.** The sample types no longer carry a `Record<string, unknown>` index signature.
|
|
1404
|
+
> That signature allowed arbitrary top-level keys but also silently accepted typos on real fields and
|
|
1405
|
+
> weakened autocomplete. Custom data belongs in `attachments`, which is typed as such. If you were
|
|
1406
|
+
> assigning ad-hoc keys directly onto a sample object, move them into `attachments`.
|
|
1407
|
+
|
|
839
1408
|
### Matching peer connections — by **event**, not by storage
|
|
840
1409
|
|
|
841
1410
|
The observer correlates the SFU side with the client side **at the peer-connection level**: a
|
|
@@ -1151,6 +1720,16 @@ filtering, per-module routing, and full silencing.
|
|
|
1151
1720
|
|
|
1152
1721
|
---
|
|
1153
1722
|
|
|
1723
|
+
## Design notes
|
|
1724
|
+
|
|
1725
|
+
**[`docs/design-notes.md`](./docs/design-notes.md)** covers the reasoning behind the library rather
|
|
1726
|
+
than its API: why client-detectable conditions are never re-derived server-side, why each shipped
|
|
1727
|
+
detector exists, what was deliberately *not* built and why, the WebRTC domain facts that shaped the
|
|
1728
|
+
implementation (ICE-Lite disconnect waves, counter resets, why ICE and RTCP RTT must never be
|
|
1729
|
+
blended), and an operational threshold reference.
|
|
1730
|
+
|
|
1731
|
+
---
|
|
1732
|
+
|
|
1154
1733
|
## Error-handling philosophy
|
|
1155
1734
|
|
|
1156
1735
|
The library **warns and degrades; it does not throw** on operational problems:
|
|
@@ -1184,10 +1763,13 @@ lint + typecheck + **build** + test on every push/PR.
|
|
|
1184
1763
|
|
|
1185
1764
|
**Project layout** (`src/`): `Observer.ts`, `ObservedCall.ts`, `ObservedClient.ts`,
|
|
1186
1765
|
`ObservedPeerConnection.ts`, the `Observed*` sub-stat classes, `ObserverEvents.ts` (the typed
|
|
1187
|
-
event map + scope types), `detectors/` (`Detector`, `Detectors
|
|
1188
|
-
|
|
1189
|
-
`
|
|
1190
|
-
`
|
|
1766
|
+
event map + scope types), `detectors/` (`Detector`, `Detectors`, and one file per detector),
|
|
1767
|
+
`validators/` (`Validator`, `Validators`, one file per validator), `issues/` (`ActiveClientIssue`,
|
|
1768
|
+
`ActiveIssueTracker`, `ActiveIssuesRegistry`, `ObservedClientIssueRegistry`), `scores/`,
|
|
1769
|
+
`resolvers/` (remote-track resolvers), `utils/` (`stats`, `SlidingWindow`, `TrendTester`,
|
|
1770
|
+
`CallHealthAggregator`), `common/` (`logger`, `utils`, `Middleware`), `schema/`
|
|
1771
|
+
(sample/event/meta types), and `sinks/` (the `ClientSampleSink` base + `JsonlFileSink` /
|
|
1772
|
+
`InMemorySink`, re-exported from the package root).
|
|
1191
1773
|
|
|
1192
1774
|
**Conventions to follow when developing further:**
|
|
1193
1775
|
|