@observertc/observer-js 1.0.0-beta.14 → 1.0.0-beta.16

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -53,7 +53,7 @@ and emits a single, unified stream of typed events the application can react to.
53
53
  3. [Data flow](#data-flow)
54
54
  4. [Entity hierarchy](#entity-hierarchy)
55
55
  5. [Ingestion: `accept()`, context & lifecycle](#ingestion-accept-context--lifecycle)
56
- 6. [Update policies](#update-policies)
56
+ 6. [When things update](#when-things-update)
57
57
  7. [The event bus](#the-event-bus) ← the core of the API
58
58
  8. [API reference](#api-reference)
59
59
  9. [Schema types (`ClientSample`)](#schema-types-clientsample)
@@ -63,8 +63,9 @@ and emits a single, unified stream of typed events the application can react to.
63
63
  13. [Sinks (per-client sample persistence)](#sinks-per-client-sample-persistence)
64
64
  14. [Injecting data into a client](#injecting-data-into-a-client)
65
65
  15. [Logging](#logging)
66
- 16. [Error-handling philosophy](#error-handling-philosophy)
67
- 17. [Development & extension guide](#development--extension-guide)
66
+ 16. [Design notes](#design-notes)
67
+ 17. [Error-handling philosophy](#error-handling-philosophy)
68
+ 18. [Development & extension guide](#development--extension-guide)
68
69
 
69
70
  ---
70
71
 
@@ -103,10 +104,9 @@ import { Observer, ClientSample } from '@observertc/observer-js';
103
104
 
104
105
  // 1. Create an observer.
105
106
  const observer = new Observer({
106
- // when the observer aggregates call/client metrics:
107
- updatePolicy: 'update-when-all-call-updated',
108
- // default policy applied to calls created automatically by accept():
109
- defaultCallUpdatePolicy: 'update-on-any-client-updated',
107
+ // a call updates when any of its clients does, and the observer when any of its calls does —
108
+ // both default to true, so this line is only here to show the knob exists:
109
+ autoUpdateOnCallUpdate: true,
110
110
  // optional auto-teardown:
111
111
  closeCallIfEmptyForMs: 20_000,
112
112
  closeClientIfIdleForMs: 60_000,
@@ -290,30 +290,38 @@ These return `undefined` (and warn) when the parent is closed; `createObservedCa
290
290
 
291
291
  ---
292
292
 
293
- ## Update policies
293
+ ## When things update
294
294
 
295
- "Update" means *recompute aggregated metrics and emit the `*-updated` event* at that level.
296
- Both the observer and each call have a configurable trigger. Updates are **event-driven** — there
297
- is no built-in timer. An app that wants a fixed cadence can call `observer.update()` /
298
- `call.update()` from its own `setInterval`. With `'none'`, **nothing auto-updates** — the level
299
- updates only when the application calls the public `update()` itself.
295
+ "Update" means *recompute aggregated metrics, run the detectors, and emit the `*-updated` event* at
296
+ that level. Updates are **event-driven** — there is no built-in timer.
300
297
 
301
- **Observer-level** (`ObserverConfig.updatePolicy`, default `update-when-all-call-updated`):
298
+ The rule is structural rather than configurable:
302
299
 
303
- | Policy | Triggers `observer.update()` when… |
304
- |--------|-------------------------------------|
305
- | `update-on-any-call-updated` | any call updates |
306
- | `update-when-all-call-updated` | every call has updated since the last observer update |
307
- | `none` | never automatically — only when the app calls `observer.update()` |
300
+ > **A call is updated when any of its clients is updated. The observer is updated when any of its
301
+ > calls is updated.** Composed, that means the observer is updated exactly when any client anywhere
302
+ > is updated.
308
303
 
309
- **Call-level** (`ObservedCallSettings.updatePolicy`, defaulted from
310
- `ObserverConfig.defaultCallUpdatePolicy`):
304
+ Two booleans, both defaulting to `true`, let you opt out of a link in that chain:
311
305
 
312
- | Policy | Triggers `call.update()` when… |
313
- |--------|--------------------------------|
314
- | `update-on-any-client-updated` | any client in the call updates |
315
- | `update-when-all-client-updated` | every client has updated since the last call update |
316
- | `none` | never automatically — only when the app calls `call.update()` |
306
+ | Setting | Where | Effect when `false` |
307
+ |---------|-------|---------------------|
308
+ | `autoUpdateOnClientUpdate` | `ObservedCallSettings` | the call updates only when you call `call.update()` |
309
+ | `autoUpdateOnCallUpdate` | `ObserverConfig` | the observer updates only when you call `observer.update()` |
310
+
311
+ An app that wants a fixed cadence sets both to `false` and drives `observer.update()` from its own
312
+ `setInterval`. Note that **observer-scoped detectors and validators run nowhere else** — if the
313
+ observer never updates, they never run.
314
+
315
+ ```ts
316
+ const observer = new Observer({ autoUpdateOnCallUpdate: false });
317
+
318
+ setInterval(() => observer.update(), 5_000);
319
+ ```
320
+
321
+ > Earlier versions had an `updatePolicy` / `defaultCallUpdatePolicy` enum (`'update-on-any-…'`,
322
+ > `'update-when-all-…'`, `'none'`) and a pluggable `Updater` object. Both are gone. "When all clients
323
+ > have updated" sounds appealing and deadlocks on the first client that stops sending — one silent
324
+ > participant froze the whole call's aggregation until it timed out.
317
325
 
318
326
  ---
319
327
 
@@ -332,7 +340,7 @@ contains the ancestry from the observer down to the entity that raised it, plus
332
340
  specific subject:
333
341
 
334
342
  ```ts
335
- type ObserverEventBase = { observer: Observer };
343
+ type ObserverEventBase = { observer: Observer, context?: AcceptContext };
336
344
  type ObservedCallScope = ObserverEventBase & { observedCall: ObservedCall };
337
345
  type ObservedClientScope = ObservedCallScope & { observedClient: ObservedClient };
338
346
  type ObservedPeerConnectionScope = ObservedClientScope & { observedPeerConnection: ObservedPeerConnection };
@@ -358,10 +366,11 @@ additional field(s) on top of that scope.
358
366
 
359
367
  | Event | Extra payload | Fires when |
360
368
  |-------|---------------|-----------|
361
- | `observer-updated` | — | `observer.update()` ran (per the observer update policy) |
369
+ | `observer-updated` | — | `observer.update()` ran (see [When things update](#when-things-update)) |
362
370
  | `observer-closed` | — | `observer.close()` |
363
371
  | `sample-rejected` | `{ reason: 'observer-closed' \| 'missing-callId' \| 'missing-clientId', sample: ClientSample }` | a sample was dropped by `accept()` |
364
- | `observer-issue` | `{ issue: ClientIssue }` | `observer.addIssue(...)` — a cross-call / SFU-wide finding (see [observer-level detectors](#observer-level-detectors-cross-call--sfu-wide)) |
372
+ | `observer-issue` | `{ issue: ObserverIssue }` | `observer.addIssue(...)` — a cross-call / SFU-wide finding (see [observer-level detectors](#observer-level-detectors-cross-call--sfu-wide)) |
373
+ | `validation-ready` | `{ validator: string, report: ValidationReport }` | a [validator](#validators--one-shot-structural-checks) settled — fires once per check, not per tick |
365
374
 
366
375
  #### Mediasoup level — scope `{ observer, observedMediasoupRouter }`
367
376
 
@@ -382,7 +391,7 @@ See [Mediasoup router observation](#mediasoup-router-observation) for the full d
382
391
  | `call-closed` | — | the call closed |
383
392
  | `call-empty` | — | last client left the call |
384
393
  | `call-not-empty` | — | first client joined a previously-empty call |
385
- | `call-issue` | `{ issue: ClientIssue }` | `call.addIssue(...)` (server-side detector finding) |
394
+ | `call-issue` | `{ issue: ObserverIssue }` | `call.addIssue(...)` (server-side detector finding) |
386
395
 
387
396
  #### Client level — scope `{ observer, observedCall, observedClient }`
388
397
 
@@ -396,7 +405,7 @@ See [Mediasoup router observation](#mediasoup-router-observation) for the full d
396
405
  | `client-left` | — | `CLIENT_LEFT` seen (or inferred on close) |
397
406
  | `client-rejoined` | `{ timestamp: number }` | a later `CLIENT_JOINED` after an earlier join |
398
407
  | `client-issue` | `{ issue: ClientIssue }` | a client-reported issue arrived, or `client.addIssue(...)`. A keyed issue also opens an entry in `observedClient.activeIssues` |
399
- | `client-issue-resolved` | `{ resolvedIssue: ResolvedClientIssue }` | a stateful issue ended — the client sent its `<type>-resolved` companion, or the observer force-closed it. Carries the finished interval (`durationInMs`, `resolvedBy`) — see [client issues](#client-issues-the-lifecycle-and-the-division-of-labour) |
408
+ | `client-issue-resolved` | `{ resolvedIssue: ResolvedActiveClientIssue }` | a stateful issue ended — the client sent its `<type>-resolved` companion, or the observer force-closed it. Carries the finished interval (`durationInMs`, `resolvedBy`) — see [client issues](#client-issues-the-lifecycle-and-the-division-of-labour) |
400
409
  | `client-metadata` | `{ metaData: ClientMetaData }` | a client meta item arrived |
401
410
  | `client-extension-stats` | `{ extensionStats: ExtensionStat }` | an app-defined extension stat arrived |
402
411
  | `client-event` | `{ event: ClientEvent }` | any client event was processed |
@@ -449,8 +458,8 @@ listen to them, but prefer the bus equivalents above for application logic.
449
458
  new Observer<AppData>(config?: ObserverConfig<AppData>)
450
459
 
451
460
  type ObserverConfig<AppData = Record<string, unknown>> = {
452
- updatePolicy?: 'update-on-any-call-updated' | 'update-when-all-call-updated' | 'none';
453
- defaultCallUpdatePolicy?: ObservedCallSettings['updatePolicy'];
461
+ // a call updates when any client does; the observer when any call does. Default true.
462
+ autoUpdateOnCallUpdate?: boolean;
454
463
  appData?: AppData;
455
464
  closeClientIfIdleForMs?: number;
456
465
  closeCallIfEmptyForMs?: number;
@@ -487,7 +496,21 @@ Key members:
487
496
  - `createObservedCall<T>(settings): ObservedCall<T> | undefined`
488
497
  - `getOrCreateObservedCall<T>(settings): ObservedCall<T> | undefined`
489
498
  - `update(): void` — force an aggregation/`observer-updated` tick
499
+ - `addObserverDetector(name, config?): this` — build a cross-call detector onto `observer.detectors`
500
+ - `addCallDetector(name, config?): this` — register a call-scoped detector for every call created
501
+ from now on
502
+ - `removeCallDetector(name, { includeOpenCalls? }): number` — stop building it, and (by default) drop
503
+ it from calls already open. Returns how many live instances were removed
504
+ - `removeObserverDetector(name): number` — remove an observer-scoped detector. For one specific
505
+ instance use `observer.detectors.remove(detector)`
506
+ - `addValidator(name, config?): this` — start a one-shot structural check
507
+ - `cancelValidator(name | validator, reason?): number` — stop a running check; it finishes
508
+ `inconclusive` with the reason and emits `validation-ready`
490
509
  - `close(): void`
510
+ - `readonly detectors: Detectors` — observer-scoped registry. **Starts empty**; nothing is implicit
511
+ - `readonly callDetectorConfigs: Map<name, config>` — what `addCallDetector` recorded
512
+ - `readonly validators: Set<RunningValidator>` — normally empty; each removes itself on finishing
513
+ - `readonly activeIssuesRegistry: ActiveIssuesRegistry` — the fleet's open client issues
491
514
  - `readonly observedCalls: Map<string, ObservedCall>`
492
515
  - `readonly observedTURN: ObservedTURN`
493
516
  - `get appData()`, `get numberOfCalls()`
@@ -500,7 +523,8 @@ Key members:
500
523
 
501
524
  ```ts
502
525
  type ObservedCallSettings<AppData = Record<string, unknown>> = {
503
- updatePolicy?: 'update-on-any-client-updated' | 'update-when-all-client-updated' | 'none';
526
+ // update this call whenever one of its clients accepts a sample. Default true.
527
+ autoUpdateOnClientUpdate?: boolean;
504
528
  callId: string;
505
529
  appData?: AppData;
506
530
  closeCallIfEmptyForMs?: number;
@@ -512,8 +536,13 @@ Key members:
512
536
  - `readonly callId: string`, `appData: AppData`
513
537
  - `readonly observedClients: Map<string, ObservedClient>`, `get numberOfClients()`
514
538
  - `getObservedClient<T>(clientId)`, `createObservedClient<T>(settings)`, `getOrCreateObservedClient<T>(settings)` (all `… | undefined`)
515
- - `addIssue(issue: ClientIssue): void` — raise a **call-level** issue → emits `call-issue`
539
+ - `addIssue(issue: ObserverIssue): void` — raise a **call-level** issue → emits `call-issue`
540
+ - `addDetector(name, config?): this` — build a call-scoped detector onto this call only
541
+ - `removeDetector(name): number` — remove it from this call, `close()`ing it. For one specific
542
+ instance use `call.detectors.remove(detector)`
516
543
  - `readonly detectors: Detectors` — server-side detector registry (empty by default; see [Detectors](#detectors-server-side-extension-point))
544
+ - `readonly activeIssuesRegistry: ActiveIssuesRegistry` — this call's open client issues, propagating into the observer's
545
+ - `readonly unconsumedOutboundTracks: Set<ObservedOutboundTrack>` — maintained by the resolver
517
546
  - `scoreCalculator: ScoreCalculator`, `get score()`, `readonly calculatedScore`
518
547
  - `remoteTrackResolver?: RemoteTrackResolver` — set from `ObserverConfig.createRemoteTrackResolver` at call creation (see [Remote track resolution](#remote-track-resolution-mediasoup--sfu))
519
548
  - aggregates: `numberOfIssues`, `numberOfPeerConnections`, `numberOfInboundRtpStreams`,
@@ -718,12 +747,28 @@ snapshot happens only once.
718
747
 
719
748
  ## Detectors (server-side extension point)
720
749
 
721
- `observer-js` deliberately ships **no built-in detectors**. Per-client signals — packet loss,
722
- jitter, RTT, freezes, etc. — are already detectable on the client and arrive on samples as
723
- `clientIssues` (surfaced via `client-issue`). Server-side detection should focus on what only
724
- the server can see by **correlating data across the clients of a call**.
750
+ `observer-js` ships [ten detectors](#built-in-detectors), each an opt-in extension you register
751
+ explicitly with `addObserverDetector` / `addCallDetector` / `addDetector` (see
752
+ [Registering detectors](#registering-detectors)) — none are created automatically. All of them
753
+ correlate **across** the clients of a call or the calls of a fleet, because that is the only thing a
754
+ server can do better than a browser: per-client signals — packet loss, jitter, RTT, freezes — are
755
+ already detected on the client and arrive on samples as `clientIssues` (surfaced via `client-issue`).
756
+
757
+ Findings are raised as **`ObserverIssue`** — `{ type, timestamp, payload? }` — and the payload is the
758
+ **object**, not a JSON string. A server-raised finding is delivered to an in-process handler, so
759
+ there is nothing to serialise for:
760
+
761
+ ```ts
762
+ observer.on('call-issue', ({ observedCall, issue }) => {
763
+ issue.payload; // the object; no JSON.parse
764
+ issuePayloadOf(issue); // if you want to accept a string payload too
765
+ issuePayloadAsString(issue); // only at an edge that needs text (log, HTTP, queue)
766
+ });
767
+ ```
768
+
769
+ (`ClientIssue`, the type on samples, keeps its string payload — that one really is a wire format.)
725
770
 
726
- The hook lives on **`ObservedCall`**:
771
+ The registry is also an open extension point, on **`ObservedCall`**:
727
772
 
728
773
  ```ts
729
774
  import { Observer, Detector } from '@observertc/observer-js';
@@ -803,29 +848,47 @@ observer.on('client-issue-resolved', ({ resolvedIssue }) => {
803
848
  });
804
849
 
805
850
  // the live per-client mirror
806
- observedClient.activeIssues; // Map<key, ActiveClientIssue>
851
+ observedClient.activeIssues; // ObservedClientIssueRegistry, keyed by issue.key
807
852
  ```
808
853
 
809
- #### `IssueRegistry` — querying the active set
854
+ > **client-monitor-js >= 4.6.0 is required** for every issue-driven detector. There is no fallback
855
+ > path that infers these conditions from raw counters — the client decides better, and maintaining a
856
+ > worse second implementation to be polite to old clients is how both end up wrong. Issues without a
857
+ > `key` have no lifecycle (nothing could ever close them), so they stay one-shot: reported on
858
+ > `client-issue`, never registered.
859
+
860
+ #### `ActiveIssuesRegistry` — issues are **pushed**, not polled
861
+
862
+ A detector does not go looking for the issues it cares about. It implements `ActiveIssueTracker` and
863
+ registers for the types it consumes; the registry hands them over as they open and close.
810
864
 
811
865
  ```ts
812
- import { IssueRegistry } from '@observertc/observer-js';
866
+ observedCall.activeIssuesRegistry // this meeting
867
+ observer.activeIssuesRegistry // the fleet; every call's registry propagates into it
813
868
 
814
- const registry = new IssueRegistry(observedCall); // or `observer` for cross-call scope
869
+ observer.activeIssuesRegistry.addIssueTracker('congestion', myDetector);
870
+ observer.activeIssuesRegistry.removeIssueTracker(myDetector);
815
871
 
816
- registry.cohortOf('congestion'); // { clientIds, affectedRatio, onsetSpreadInMs, … }
817
- registry.cohorts(); // every shared issue type, largest cohort first
818
- registry.byTrackIds(trackIds); // issues attributed to a published track's receivers
872
+ registry.values(); // the open issues in this scope, oldest first
873
+ registry.size; // how many
819
874
  ```
820
875
 
821
- `onsetSpreadInMs` is measured on the **observer clock**, never the client's. `raisedAt` comes from
822
- each participant's own machine, and comparing those across clients makes clock skew look like a
876
+ The cost of a detector is then proportional to the issues it actually receives, not to the number of
877
+ participants: a healthy 500-client fleet does no per-tick work at all, because nothing was pushed.
878
+
879
+ **There is no wildcard.** A tracker names its types and sees nothing else. "Feed me everything and
880
+ I'll work out what matters" moves the decision from the application — which knows its client build
881
+ and its issue vocabulary — onto a detector that has to guess, and it makes the cost of a
882
+ subscription unbounded and invisible. If a detector should watch five types, list five types.
883
+
884
+ Onset spread is measured on the **observer clock**, never the client's. `raisedAt` comes from each
885
+ participant's own machine, and comparing those across clients makes clock skew look like a
823
886
  synchronized infrastructure event.
824
887
 
825
888
  ### Observer-level detectors (cross-call / SFU-wide)
826
889
 
827
890
  Some findings only exist **above** call scope — "many calls on the same SFU degraded at once" is far
828
- more actionable than fifty individual client alerts. The same registry exists on the `Observer`,
891
+ more actionable than fifty individual client alerts. The same detector registry exists on the `Observer`,
829
892
  runs on every `observer.update()`, and raises findings through `observer.addIssue(...)`, surfaced on
830
893
  the bus as **`observer-issue`**:
831
894
 
@@ -844,41 +907,30 @@ observer.detectors.add({
844
907
  observer.on('observer-issue', ({ issue }) => alert(issue));
845
908
  ```
846
909
 
847
- ### Publisher → subscribers: `TrackDistributionAggregator`
910
+ ### Publisher → subscribers: the resolver links
848
911
 
849
912
  The question a single browser can never answer is *"did **everyone** receiving Alice see the same
850
- degradation?"*. `TrackDistributionAggregator` answers it by walking the publisher→subscriber links
851
- maintained by a [`RemoteTrackResolver`](#remote-track-resolution-mediasoup--sfu) and summarizing one
852
- published track against all of its receivers:
913
+ degradation?"*. The join is the publisher↔subscriber links maintained by a
914
+ [`RemoteTrackResolver`](#remote-track-resolution-mediasoup--sfu), and detectors walk them directly:
853
915
 
854
916
  ```ts
855
- import { TrackDistributionAggregator } from '@observertc/observer-js';
856
-
857
- const aggregator = new TrackDistributionAggregator(observedCall);
858
-
859
- for (const d of aggregator.aggregate()) {
860
- d.trackId; // the published track
861
- d.publisher.healthy; // is the source itself fine? (+ .reasons)
862
- d.numberOfReceivers; // 17
863
- d.numberOfDegradedReceivers; // 14
864
- d.degradedRatio; // 0.82
865
- d.fractionLost?.p95; // percentile summaries across receivers
866
- d.freezes; // { affectedReceivers, total } — fan-out counters
867
- d.plis;
868
- d.receivers; // per-receiver entries with `degraded` + `reasons`
869
- }
917
+ outboundTrack.remoteInboundTracks; // Set<ObservedInboundTrack> — every subscriber of this source
918
+ inboundTrack.remoteOutboundTrack; // the publisher, or undefined if unlinked
919
+ inboundTrack.getInboundRtp(); // that receiver's RTP stats
920
+ observedCall.unconsumedOutboundTracks; // published tracks with no subscriber at all
870
921
  ```
871
922
 
872
- Each receiver is judged against `ReceiverHealthThresholds` (loss, freezes, dropped frames,
873
- concealment, jitter-buffer delay, RTT — override any of them). Summaries are **medians and
874
- percentiles, not means**, because one participant at 1500 ms RTT would otherwise hide nine healthy
875
- ones. The helpers behind it (`percentile`, `median`, `summarize`, `counterDelta`, `SlidingWindow`)
876
- are exported for building your own detectors.
923
+ > A `TrackDistributionAggregator` class used to sit in front of these links and summarise every
924
+ > published track against all of its receivers, on every tick. It is gone. It scanned the majority
925
+ > (all published tracks) to find the interesting minority, which is the wrong axis — the detectors
926
+ > now start from the handful of *affected* tracks the issue registry pushed at them and resolve only
927
+ > those. The statistics helpers it used (`percentile`, `median`, `summarize`, `counterDelta`,
928
+ > `robustZScore`, `SlidingWindow`, `TrendTester`) are all still exported for building your own.
877
929
 
878
930
  ### Call health: `CallHealthAggregator`
879
931
 
880
- The second aggregation axis — the **client** axis. Where `TrackDistributionAggregator` asks "how was
881
- *this source* delivered?", this asks "how is *each participant* doing, sending vs receiving?":
932
+ The **client** axis. Where the resolver links answer "how was *this source* delivered?", this asks
933
+ "how is *each participant* doing, sending vs receiving?":
882
934
 
883
935
  ```ts
884
936
  import { CallHealthAggregator } from '@observertc/observer-js';
@@ -893,43 +945,342 @@ health.qualityLimitation; // { cpu, bandwidth, other } client counts
893
945
  health.clients; // per-client entries with `reasons`, direction flags, TURN/TCP
894
946
  ```
895
947
 
896
- ### Built-in detectors
948
+ ### Registering detectors
949
+
950
+ **Nothing is created implicitly.** A `new Observer()` has zero detectors. There is no detector
951
+ configuration in `ObserverConfig` and no default set — an application says what it wants to watch, or
952
+ it watches nothing.
953
+
954
+ ```ts
955
+ const observer = new Observer({
956
+ createRemoteTrackResolver: createDefaultMediasoupRemoteTrackResolverFactory(),
957
+ });
958
+
959
+ // observer-scoped (cross-call) — built immediately onto `observer.detectors`
960
+ observer.addObserverDetector('observer-concurrent-issue-detector', {
961
+ issueTypes: [ 'congestion', 'ice-disconnected', 'ice-connection-failed' ],
962
+ minAffectedCalls: 3,
963
+ });
964
+ observer.addObserverDetector('turn-server-outage-detector', { minClientsAtPeak: 10 });
965
+
966
+ // call-scoped — recorded in `observer.callDetectorConfigs`, applied to every call created AFTER this
967
+ observer.addCallDetector('call-concurrent-issue-detector', {
968
+ issueTypes: [ 'congestion', 'ice-disconnected' ],
969
+ });
970
+ // one specific call
971
+ observedCall.addDetector('issue-fan-out-detector', { issueTypes: [ 'freezed-video-track' ] });
972
+ ```
973
+
974
+ Every `add*` is **chainable** — it returns the owning entity:
975
+
976
+ ```ts
977
+ observer
978
+ .addObserverDetector('turn-server-health-detector')
979
+ .addObserverDetector('turn-server-outage-detector', { minClientsAtPeak: 10 })
980
+ .addValidator('remote-track-resolver');
981
+ ```
982
+
983
+ #### Removing them
984
+
985
+ By **name**, on the entity — which removes *every* instance under that name:
986
+
987
+ ```ts
988
+ observer.removeObserverDetector('turn-server-outage-detector'); // → 1
989
+ observer.removeCallDetector('call-concurrent-issue-detector'); // stops it everywhere
990
+ observedCall.removeDetector('issue-fan-out-detector'); // this call only
991
+ ```
992
+
993
+ By **instance**, through the registry — which is where instances live, since `add*` returns the
994
+ entity rather than the detector:
897
995
 
898
- All of them are **opt-in** — the library enables none by default. Register call-scoped ones on
899
- `observedCall.detectors` and observer-scoped ones on `observer.detectors`:
996
+ ```ts
997
+ observer
998
+ .addObserverDetector('client-population-issue-detector', { issueTypes: [ 'cpulimitation' ], groupBy: 'browser' })
999
+ .addObserverDetector('client-population-issue-detector', { issueTypes: [ 'cpulimitation' ], groupBy: 'operationSystem' });
1000
+
1001
+ const [ byBrowser, byOs ] = observer.detectors.getAll('client-population-issue-detector');
1002
+
1003
+ observer.detectors.remove(byOs); // keeps the browser axis running
1004
+ ```
1005
+
1006
+ `Detectors` is a small collection: `instances` (a copy, in registration order), `listOfNames`,
1007
+ `size`, `get(name)`, `getAll(name)`, `has(name)`, `add(detector)`, `remove(detector)`,
1008
+ `removeByName(name)`, `clear()`, and it is iterable — `for (const detector of call.detectors)`.
1009
+ `instances` being a copy is deliberate: removing while iterating the live array would skip entries,
1010
+ and "drop the ones that look like X" is the most natural thing to want to write.
1011
+
1012
+ Two things worth knowing:
1013
+
1014
+ - **By name removes every instance under it**, not the first. A name can legitimately be registered
1015
+ more than once — `ClientPopulationIssueDetector` is meant to be added once per `groupBy` axis — and
1016
+ "remove whichever is first in the array" is not something a caller can predict from a name. Go via
1017
+ `detectors.getAll(name)` + `detectors.remove(instance)` when you mean one of them.
1018
+ - **`removeCallDetector` affects calls already open, by default.** Otherwise whether a detector runs
1019
+ would depend on when a call happened to join, which is not a state anyone can reason about. Pass
1020
+ `{ includeOpenCalls: false }` to change only what future calls are built with.
1021
+
1022
+ Every removal path calls the detector's `close()`, so it unsubscribes from the issue registry and
1023
+ drops any timers or bus listeners. A detector removed without closing would keep being fed matching
1024
+ issues for the life of the call — invisible, unbounded, and it would still look healthy if you
1025
+ inspected it.
1026
+
1027
+ Detectors are named by their kebab-case `NAME`, and the name types the config — an unknown name or a
1028
+ key that belongs to a different detector will not compile. Each detector owns its defaults in its own
1029
+ constructor, beside the doc explaining what each threshold means; there is no central table to keep
1030
+ in sync.
1031
+
1032
+ > **Why no defaults?** A detector nobody asked for is a detector nobody will act on. It costs time on
1033
+ > every tick and raises findings into a handler that was not written to expect them. Earlier versions
1034
+ > auto-created everything from a three-state config slot; the result was applications receiving
1035
+ > finding types they had never heard of.
1036
+
1037
+ Every issue-driven detector takes an explicit, non-empty `issueTypes` (or
1038
+ `publisherIssueTypes`/`receiverIssueTypes`). There is no "watch everything" option — see
1039
+ [the registry](#activeissuesregistry--issues-are-pushed-not-polled).
900
1040
 
901
1041
  > **🔗 marks detectors that require a
902
1042
  > [`RemoteTrackResolver`](#remote-track-resolution-mediasoup--sfu).** They reason about a published
903
1043
  > track and its subscribers, so without the publisher↔subscriber links they see nothing and stay
904
1044
  > **silent forever** — which looks exactly like "no problems found". Configure
905
- > `ObserverConfig.createRemoteTrackResolver` before registering them.
1045
+ > `ObserverConfig.createRemoteTrackResolver`, and start the
1046
+ > [`remote-track-resolver` validator](#validators--one-shot-structural-checks) to prove it is wired.
906
1047
 
907
- **Issue-driven** (preferred — they consume the client's own verdicts):
1048
+ ### Built-in detectors
908
1049
 
909
- | Detector | 🔗 | Scope | Raises |
910
- |----------|:--:|-------|--------|
911
- | `ConcurrentIssueDetector` | | call **or** observer | `CONCURRENT_CLIENT_ISSUES`, `ISSUE_ONSET_BURST` |
912
- | `IssueFanOutDetector` | 🔗 | call | `PUBLISHED_TRACK_ISSUE_FAN_OUT`, `SINGLE_RECEIVER_ISSUE` |
913
- | `TrackDeliveryMismatchDetector` | 🔗 | call | `PUBLISHED_TRACK_NOT_DELIVERED`, `RECEIVER_TRACK_NOT_DELIVERED`, `PUBLISHER_TRACK_DRY` |
1050
+ They consume the verdicts `client-monitor-js` >= 4.6.0 already ships (raise + `<type>-resolved`) and
1051
+ add only the cross-participant conclusion. **None re-derives a per-endpoint verdict from raw
1052
+ counters** — that is the rule the whole design hangs on:
914
1053
 
915
- **Metric-driven** (work without client issues — the fallback path, and the things no client issue
916
- can express):
1054
+ > **If a condition is detectable on the client, the client's issue is the source of truth.**
917
1055
 
918
1056
  | Detector | 🔗 | Scope | Raises |
919
1057
  |----------|:--:|-------|--------|
920
- | `WorstReceiverContagionDetector` | 🔗 | call | `WORST_RECEIVER_CONTAGION` |
921
- | `CommonSourceDegradationDetector` | 🔗 | call | `PUBLISHER_HEALTHY_SUBSCRIBERS_DEGRADED`, `PUBLISHER_DEGRADED_FOR_ALL_SUBSCRIBERS`, `SINGLE_SUBSCRIBER_DEGRADED`, `MULTIPLE_SUBSCRIBERS_DEGRADED` |
922
- | `PliAndFreezeFanOutDetector` | 🔗 | call | `PUBLISHER_PLI_STORM`, `PUBLISHED_VIDEO_FROZEN_FOR_MULTIPLE_RECEIVERS` |
923
- | `AudioImpairmentFanOutDetector` | 🔗 | call | `PUBLISHED_AUDIO_DEGRADED_FOR_MAJORITY`, `CALL_WIDE_AUDIO_JITTER_BUFFER_STRESS` |
1058
+ | `CallConcurrentIssueDetector` | | call | `CONCURRENT_CLIENT_ISSUES`, `ISSUE_ONSET_BURST` |
1059
+ | `ObserverConcurrentIssueDetector` | | **observer** | `CROSS_CALL_CONCURRENT_ISSUES`, `CROSS_CALL_ISSUE_ONSET_BURST` |
1060
+ | `IssueFanOutDetector` | 🔗 | call | `PUBLISHED_TRACK_ISSUE_FAN_OUT`, `SINGLE_RECEIVER_ISSUE` |
1061
+ | `PublisherFaultCorroborationDetector` | 🔗 | call | `CORROBORATED_PUBLISHER_FAULT` |
1062
+ | `TrackDeliveryMismatchDetector` | 🔗 | call | `PUBLISHED_TRACK_NOT_DELIVERED`, `RECEIVER_TRACK_NOT_DELIVERED`, `PUBLISHER_TRACK_DRY` |
924
1063
  | `UnconsumedTrackDetector` | 🔗 | call | `UNCONSUMED_PUBLISHED_TRACK` |
925
- | `CallWideDegradationDetector` | | call | `CALL_WIDE_QUALITY_DEGRADATION`, `CALL_WIDE_INBOUND_DEGRADATION`, `CALL_WIDE_OUTBOUND_DEGRADATION` |
926
- | `IceDisruptionDetector` | | call | `CALL_ICE_DISRUPTION` |
1064
+ | `ClientPopulationIssueDetector` | | **observer** | `CLIENT_POPULATION_ISSUE` |
1065
+ | `SfuCongestionDetector` | | **observer** | `sfu-congestion` |
927
1066
  | `TurnServerHealthDetector` | | **observer** | `TURN_SERVER_DEGRADED` |
1067
+ | `TurnServerOutageDetector` | | **observer** | `TURN_SERVER_OUTAGE` |
1068
+
1069
+ What each adds that no endpoint can know:
1070
+
1071
+ - **`CallConcurrentIssueDetector`** — *who else in this meeting is in this state right now?* The
1072
+ difference between "one person's Wi-Fi" and "this room is broken".
1073
+ - **`ObserverConcurrentIssueDetector`** — *is our infrastructure in trouble?* A **separate class**,
1074
+ not the call one with a bigger denominator, because it is a different question with different
1075
+ gates. It requires the group to span at least `minAffectedCalls` **independent calls** (default
1076
+ `2`) and raises its own `CROSS_CALL_*` types. Without that gate, one thirty-person meeting where
1077
+ everyone is congested clears every client threshold and pages you for a single bad room the
1078
+ call-scoped detector already reported. Clients in different calls share no room, no publisher and
1079
+ no host — only the servers, which is what makes the finding conclusive. Note there is deliberately
1080
+ **no participant ratio** at this scope: six broken calls out of forty is a small share of all
1081
+ clients, and a ratio gate would hide exactly the event you want.
1082
+ - **`IssueFanOutDetector`** — *does this issue follow one published source, or one receiver?*
1083
+ - **`PublisherFaultCorroborationDetector`** — *do **both ends** of one track agree the source is at
1084
+ fault?* Fan-out sees one end and infers; this sees the publisher reporting `encoder-bottleneck`
1085
+ about its own send path *while* its subscribers report `freezed-video-track` about receiving it.
1086
+ Two independent parties, one conclusion, nothing left to deduce — hence the highest confidence in
1087
+ the library. Run both: fan-out is broader and catches the case where the publisher is fine and the
1088
+ SFU's forwarding is not.
1089
+ - **`ClientPopulationIssueDetector`** — *is this concentrated on one **kind of client**?* The one
1090
+ correlation here that is neither per-call nor per-server. Every other observer-scoped detector
1091
+ reasons "clients in unrelated calls share only the infrastructure, so it must be us" — right for
1092
+ network symptoms, **wrong for endpoint ones**. `cpulimitation` across six unrelated calls is not an
1093
+ SFU event; CPU is owned by the endpoint, so what those endpoints share is a browser version or a
1094
+ client release. Groups by `browser` / `engine` / `platform` / `operationSystem`, one axis per
1095
+ instance. The gate is **relative risk**, not share: "30% of Chrome 141 is unhappy" means nothing if
1096
+ 30% of everyone is, and a share-based rule simply indicts whichever browser is most popular.
1097
+ - **`SfuCongestionDetector`** — *is congestion spiking across the fleet right now?* Counts distinct
1098
+ clients reporting congestion in fixed wall-clock buckets and compares each bucket against a
1099
+ median+MAD baseline of the ones before it. Buckets rather than update ticks on purpose: the tick is
1100
+ unevenly spaced and shorter than a client's sampling period, so counting on it compares windows of
1101
+ different lengths and calls the difference a signal. Only add it when the observer's calls all come
1102
+ from the **same SFU** — the finding's meaning is "these clients share only that server".
1103
+ - **`TrackDeliveryMismatchDetector`** — *are the two ends of a track disagreeing?*
1104
+ - **`UnconsumedTrackDetector`** — *is anyone actually subscribed?* (reads the resolver's silence)
1105
+ - **`TurnServerHealthDetector`** — *does trouble cluster on one relay?*
1106
+ - **`TurnServerOutageDetector`** — covers the case the health detector structurally cannot. The
1107
+ health detector groups clients by the server relaying them and asks how many report issues — it
1108
+ needs clients *on* the server to ask. When a TURN server dies, allocation fails: existing sessions
1109
+ drop and new clients never obtain a relay candidate through it, so they are never attributed to it
1110
+ at all. Its population goes to zero and the health detector falls silent for the worst possible
1111
+ reason. **Degradation makes clients unhappy; an outage makes them disappear.** Absence is a
1112
+ dangerous signal, so the **control group** is the heart of the design: a call ending, everyone
1113
+ leaving at 6pm, and a fleet-wide network event all look identical to an outage. It refuses to blame
1114
+ a server unless clients *not* relayed through it are demonstrably still connected
1115
+ (`requireControlGroup`, on by default).
1116
+
1117
+ #### There is no ICE detector
1118
+
1119
+ ICE trouble is reported by `client-monitor-js` >= 4.6.0 as the keyed issues `ice-disconnected`,
1120
+ `ice-connection-failed`, `ice-transport-stalled` and `unstable-ice-path`, each with hysteresis and
1121
+ multi-signal confirmation behind it. An `IceDisruptionDetector` used to re-derive that server-side
1122
+ from raw state transitions; it has been removed, because the server sees less and guesses more. The
1123
+ client knows whether `disconnected` persisted or healed in 200 ms; the observer does not.
1124
+
1125
+ Correlating ICE trouble is now configuration, not a class:
1126
+
1127
+ ```ts
1128
+ observer.addObserverDetector('observer-concurrent-issue-detector', {
1129
+ issueTypes: [ 'ice-disconnected', 'ice-connection-failed', 'ice-transport-stalled' ],
1130
+ });
1131
+ ```
928
1132
 
929
- If your clients report issues, the issue-driven pair **subsumes** the fan-out and ICE detectors —
930
- `IssueFanOutDetector` covers freeze/PLI/concealment fan-out generically, and `ConcurrentIssueDetector`
931
- with the ICE issue types replaces `IceDisruptionDetector` (and does it better: the client already
932
- suppresses the transient blips ICE heals by itself). Run both families only while migrating.
1133
+ ### Validators — one-shot structural checks
1134
+
1135
+ Every detector above answers *"is something wrong right now?"* and runs on every tick, because the
1136
+ answer legitimately changes. A **validator** answers *"is this deployment built correctly?"* — which
1137
+ only changes when you deploy. So it is not configured on and left running: you **start** one, it runs
1138
+ until it can decide, reports once, and the observer drops it.
1139
+
1140
+ ```ts
1141
+ observer.addValidator('simulcast-receivers', { minChecks: 5 });
1142
+
1143
+ observer.on('validation-ready', ({ validator, report }) => {
1144
+ if (!report.ready) return;
1145
+ if (report.verdict === 'layer-decided-lowest-common-denominator') page(validator, report);
1146
+ });
1147
+
1148
+ onDeploy(() => observer.addValidator('simulcast-receivers')); // check again
1149
+ ```
1150
+
1151
+ `observer.validators` is the set currently running — normally empty, since each removes itself on
1152
+ finishing. There is no revalidation timer: a deploy, not elapsed time, is what makes a structural
1153
+ verdict stale, so re-checking means starting another.
1154
+
1155
+ **Cancelling.** A check that has not decided can be stopped, by name or by instance:
1156
+
1157
+ ```ts
1158
+ observer.cancelValidator('simulcast-receivers', 'sfu redeployed');
1159
+
1160
+ // or one specific instance — `observer.validators` holds what is running
1161
+ for (const validator of observer.validators) observer.cancelValidator(validator, 'shutting down');
1162
+ ```
1163
+
1164
+ Cancelling is **not** silent discarding. The validator finishes `inconclusive` with your reason,
1165
+ emits `validation-ready` like any other completion, and removes itself. That matters twice over:
1166
+ anything waiting on the verdict would otherwise wait forever, and *"we stopped asking"* is a
1167
+ materially different outcome from *"we asked and learned nothing"* — which is exactly what an
1168
+ `inconclusive` carrying a reason records. Pass a real reason; the default tells the reader nothing
1169
+ they could not already infer. `observer.close()` cancels whatever is still running with
1170
+ `'observer closed'`.
1171
+
1172
+ | Validator | `addValidator` name | Question | Also raises |
1173
+ |-----------|---------------------|----------|-------------|
1174
+ | `SimulcastReceiverValidator` 🔗 | `simulcast-receivers` | Does the SFU pick layers per receiver, or drag the publisher down to the worst one? | `WORST_RECEIVER_CONTAGION` |
1175
+ | `RemoteTrackResolverValidator` | `remote-track-resolver` | Is the resolver actually linking anything? | `REMOTE_TRACK_LINKS_UNRESOLVED` |
1176
+ | `CodecConsistencyValidator` | `codec-consistency` | Is everyone on the same codec — and is it the one you think you negotiated? | `CODEC_INCONSISTENCY` |
1177
+
1178
+ **`SimulcastReceiverValidator`** — simulcast (or SVC) exists so one slow
1179
+ participant doesn't set everyone's quality: with several encodings the server hands the struggling
1180
+ receiver a lower layer and leaves the rest alone. Without it — or with a server that relays RTCP end
1181
+ to end, so the publisher's bandwidth estimate collapses to the slowest receiver — the only way to
1182
+ serve them is to make the *source* send less. Both causes look identical from outside; what the check
1183
+ establishes is whether per-receiver adaptation happens at all.
1184
+
1185
+ | `verdict` | meaning |
1186
+ |-----------|---------|
1187
+ | `layer-decided-per-receiver` | verified — a receiver fell far behind and the publisher carried on |
1188
+ | `layer-decided-lowest-common-denominator` | the publisher followed its worst receiver; everyone gets the slowest participant's quality |
1189
+ | `inconclusive` | cancelled, or the observer closed, before it could decide |
1190
+
1191
+ **Not finishing is not a pass.** The check only runs when a publisher has 3+ receivers and one is at
1192
+ most half the median; plenty of healthy deployments never present that. A validator that never sees it
1193
+ simply keeps running and never reports — it does not quietly succeed. `report.checks` counts the times
1194
+ the check genuinely ran, so an `inconclusive` with `checks: 0` says plainly that nothing was verified.
1195
+
1196
+ **`RemoteTrackResolverValidator`** exists because of a specific, nasty failure mode. Four things here
1197
+ are built on publisher↔subscriber links — `IssueFanOutDetector`,
1198
+ `PublisherFaultCorroborationDetector`, `TrackDeliveryMismatchDetector`, `UnconsumedTrackDetector`
1199
+ (and `SimulcastReceiverValidator`) — and every one of them correctly does *nothing* when the links
1200
+ are missing rather than guessing. So a resolver wired to the wrong id field leaves all of them
1201
+ permanently silent, and **silence is what a healthy deployment looks like too**: you would conclude
1202
+ your calls were clean when in fact nothing was ever examined. Verdicts: `links-resolved` /
1203
+ `no-links-resolved` / `inconclusive`. Run it at start-up and after changing the resolver or the SFU's
1204
+ id scheme.
1205
+
1206
+ **`CodecConsistencyValidator`** answers two things at once. A **split** — participants of one call on
1207
+ different codecs — is a real fault with a confusing symptom: an SFU that forwards without transcoding
1208
+ cannot serve them all, so some pairs see each other and some do not, with no error anywhere. Only
1209
+ something holding every participant at once can see it. The quieter half is the silent fallback: a
1210
+ deployment configured for VP9 or AV1 drops to VP8 whenever one endpoint cannot negotiate the
1211
+ preference, the call keeps working at a higher bitrate than budgeted, and the team believes it
1212
+ shipped AV1 months ago. Give it `expected` and it says so. Verdicts: `codec-consistent` /
1213
+ `codec-split` / `unexpected-codec` / `inconclusive`.
1214
+
1215
+ ```ts
1216
+ observer.addValidator('remote-track-resolver');
1217
+ observer.addValidator('codec-consistency', { expected: { video: 'video/VP9', audio: 'audio/opus' } });
1218
+ ```
1219
+
1220
+ #### Conclusions
1221
+
1222
+ Every issue-driven finding carries a `conclusion` in its payload — the interpretation step, so the
1223
+ person reading the alert doesn't have to perform it:
1224
+
1225
+ ```jsonc
1226
+ {
1227
+ "type": "CROSS_CALL_ISSUE_ONSET_BURST",
1228
+ "issueType": "congestion",
1229
+ "calls": 40, "affectedCalls": 6,
1230
+ "perCall": [ { "callId": "…", "affectedClients": 4, "totalClients": 9 } ],
1231
+ "conclusion": {
1232
+ "faultDomain": "infrastructure",
1233
+ "summary": "network congestion is open across independent calls at the same time — 6 of 40 calls (11/300 clients)",
1234
+ "recommendation": "check SFU egress bandwidth and host network saturation before looking at any single participant",
1235
+ "confidence": 0.85
1236
+ }
1237
+ }
1238
+ ```
1239
+
1240
+ `faultDomain` is one of `infrastructure`, `call`, `published-track`, `endpoint`, `client-population`
1241
+ or `unknown`, and it comes from the **spread**, not the issue type — congestion in one call is a
1242
+ meeting problem, congestion in six calls is a server problem, and the client reported the identical
1243
+ symptom in both.
1244
+
1245
+ One case is worth knowing about because it inverts the usual reading: **`cpu-limitation` spread
1246
+ across many independent calls concludes `client-population`, not `infrastructure`.** Endpoint CPU is
1247
+ owned by the endpoint, so breadth there points at what those endpoints share — a recent client
1248
+ release, a browser version, shared VDI hardware — and paging the SFU on-call would be wrong. The
1249
+ conclusion table encodes that so nobody has to rediscover it during an incident.
1250
+
1251
+ Unknown issue types (your own custom client detectors) still produce a structurally valid conclusion
1252
+ from the spread alone; they just get generic wording.
1253
+
1254
+ Two functions are exported, one per scope: `concludeCallIssue()` and `concludeObserverIssue()`. They
1255
+ are separate because a detector already knows its scope, and a single generic function forced every
1256
+ caller to pass the other scope's fields as placeholders — call-scoped detectors passing
1257
+ `affectedCalls: 1, totalCalls: 1` forever, observer-scoped ones passing a participant ratio that was
1258
+ deliberately never read. Placeholders like that invite being read as if they meant something.
1259
+
1260
+ #### Cost
1261
+
1262
+ Detectors run inside `call.update()`, on your event loop, so their cost matters. Two things keep it
1263
+ off the participant axis:
1264
+
1265
+ - **Issues are pushed, not polled.** A detector holds only what the registry handed it, so an
1266
+ `update()` that finds `size === 0` — the overwhelmingly common case — costs one comparison,
1267
+ whatever the participant count. Nothing iterates clients looking for trouble.
1268
+ - **So are unconsumed tracks.** `observedCall.unconsumedOutboundTracks` is maintained by the resolver
1269
+ as tracks gain and lose subscribers, so `UnconsumedTrackDetector` reads a set that is normally
1270
+ empty instead of walking every published track (529 µs → 65 µs per tick at 1 200 tracks).
1271
+ - **Track lookups start from the affected minority.** A detector resolving an issue to its published
1272
+ track searches the *reporting client's* peer connections (typically one or two), not the call.
1273
+
1274
+ At 20 calls × 12 participants (2 640 subscriptions) the whole detector pass costs ~1.3 ms per tick.
1275
+ `yarn bench` prints a per-detector breakdown for your own shape.
1276
+
1277
+ #### Worked examples
1278
+
1279
+ [`examples/detectors.ts`](./examples/detectors.ts) (`yarn example:detectors`) runs one scenario per
1280
+ detector — the question it answers, its full config, the synthetic traffic that makes it fire, and
1281
+ the finding with its conclusion. It asserts every expected finding is produced, so it doubles as a
1282
+ smoke test. [`examples/sfu-observer.ts`](./examples/sfu-observer.ts) (`yarn example`) is the end-to-end
1283
+ tour instead: ingest → correlate → react, with the mediasoup wiring alongside.
933
1284
 
934
1285
  #### `TrackDeliveryMismatchDetector` — resolving an ambiguous symptom
935
1286
 
@@ -966,99 +1317,6 @@ rather than trusting the flag alone: **"no subscribers" and "no resolver configu
966
1317
  identical observation.** Without a resolver it would report every published track in the call as
967
1318
  unconsumed.
968
1319
 
969
- #### `WorstReceiverContagionDetector` — the one worth reading about
970
-
971
- In a correctly built SFU the RTCP feedback loop is **terminated at the server**: each receiver's
972
- reports drive what *that* receiver is sent. When the loop is relayed end-to-end instead, the
973
- publisher's bandwidth estimate collapses to the minimum across all receivers — so one participant on
974
- a bad 3G link silently downgrades the stream **everyone** sees. This is the "lowest common
975
- denominator" failure simulcast exists to prevent.
976
-
977
- It's detected as a *correlation over a window*, not a threshold: the publisher's outbound bitrate
978
- moving in lockstep with the **worst** receiver's inbound bitrate, while the median receiver has
979
- headroom — and tracking the worst receiver more closely than the median (otherwise everyone is just
980
- moving together, which is ordinary adaptation). The damage is invisible from every endpoint: the
981
- publisher sees "my bitrate dropped", each healthy receiver sees "my video got worse", and only the
982
- server can see the causal link.
983
-
984
- ```ts
985
- import {
986
- ConcurrentIssueDetector, IssueFanOutDetector, WorstReceiverContagionDetector,
987
- CallWideDegradationDetector, TurnServerHealthDetector,
988
- } from '@observertc/observer-js';
989
-
990
- observer.on('call-added', ({ observedCall }) => {
991
- // issue-driven: correlate the verdicts the clients already reached
992
- observedCall.detectors.add(new IssueFanOutDetector(observedCall));
993
- observedCall.detectors.add(new ConcurrentIssueDetector(observedCall));
994
- // metric-driven: things no client issue can express
995
- observedCall.detectors.add(new WorstReceiverContagionDetector(observedCall));
996
- observedCall.detectors.add(new CallWideDegradationDetector(observedCall));
997
- });
998
-
999
- // cross-call, so these go on the observer and raise `observer-issue`
1000
- observer.detectors.add(new TurnServerHealthDetector(observer));
1001
- observer.detectors.add(new ConcurrentIssueDetector(observer, {
1002
- issueTypes: [ 'ice-disconnected', 'ice-connection-failed', 'congestion' ],
1003
- }));
1004
-
1005
- observer.on('call-issue', ({ observedCall, issue }) => handle(observedCall, issue));
1006
- observer.on('observer-issue', ({ issue }) => handle(undefined, issue));
1007
- ```
1008
-
1009
- Every detector takes an options object to tune `minReceivers`/`minClients`, the ratio thresholds, the
1010
- per-entity health thresholds, and the debounce (`consecutiveTicks`, or `windowMs` + `cooldownMs` for
1011
- the window-based ones) — so an alert needs a condition to *persist*, not just appear in one sample.
1012
- Findings that depend on publisher↔subscriber links need a
1013
- [`RemoteTrackResolver`](#remote-track-resolution-mediasoup--sfu); without one those detectors stay
1014
- silent. A detector may implement `close()` (called when it's removed or the call/observer closes) —
1015
- `IceDisruptionDetector` uses it to drop the bus listeners it subscribes with.
1016
-
1017
- ### `CommonSourceDegradationDetector`
1018
-
1019
- The first detector built on the aggregator. It classifies where a fault lies by comparing the source
1020
- against its receivers, and raises a `call-issue` whose `type` is one of:
1021
-
1022
- | Finding | Meaning |
1023
- |---------|---------|
1024
- | `PUBLISHER_HEALTHY_SUBSCRIBERS_DEGRADED` | source egress fine, most receivers degraded → **downstream / SFU** suspected |
1025
- | `PUBLISHER_DEGRADED_FOR_ALL_SUBSCRIBERS` | the source itself is impaired → **publisher-side** |
1026
- | `SINGLE_SUBSCRIBER_DEGRADED` | one receiver suffers while the rest are fine → **that receiver's network** |
1027
- | `MULTIPLE_SUBSCRIBERS_DEGRADED` | several (but not most) receivers on the same source |
1028
-
1029
- ```ts
1030
- import { CommonSourceDegradationDetector } from '@observertc/observer-js';
1031
-
1032
- observer.on('call-added', ({ observedCall }) => {
1033
- observedCall.detectors.add(new CommonSourceDegradationDetector(observedCall, {
1034
- minReceivers: 3, // ratios need a meaningful denominator
1035
- degradedRatioThreshold: 0.6, // "most receivers"
1036
- consecutiveTicks: 2, // must hold 2 ticks — avoids flapping on one bad sample
1037
- }));
1038
- });
1039
- ```
1040
-
1041
- The `call-issue` payload carries the evidence: publisher health and reasons, receiver/degraded
1042
- counts, `degradedRatio`, `affectedClientIds`, freeze/PLI fan-out and the loss/bitrate summaries.
1043
- It requires a configured `RemoteTrackResolver` — with no links there is nothing to compare and the
1044
- detector stays silent.
1045
-
1046
- `Detector` interface and the registry:
1047
-
1048
- ```ts
1049
- interface Detector { readonly name: string; update(): void; }
1050
-
1051
- class Detectors {
1052
- add(d: Detector): void;
1053
- remove(d: Detector): void;
1054
- clear(): void;
1055
- update(): void; // called by ObservedCall.update(); guards each detector in try/catch
1056
- get listOfNames(): string[];
1057
- }
1058
- ```
1059
-
1060
- ---
1061
-
1062
1320
  ## Remote track resolution (mediasoup / SFU)
1063
1321
 
1064
1322
  In an SFU, one participant's **outbound** track is delivered to other participants as **inbound**
@@ -1538,6 +1796,16 @@ filtering, per-module routing, and full silencing.
1538
1796
 
1539
1797
  ---
1540
1798
 
1799
+ ## Design notes
1800
+
1801
+ **[`docs/design-notes.md`](./docs/design-notes.md)** covers the reasoning behind the library rather
1802
+ than its API: why client-detectable conditions are never re-derived server-side, why each shipped
1803
+ detector exists, what was deliberately *not* built and why, the WebRTC domain facts that shaped the
1804
+ implementation (ICE-Lite disconnect waves, counter resets, why ICE and RTCP RTT must never be
1805
+ blended), and an operational threshold reference.
1806
+
1807
+ ---
1808
+
1541
1809
  ## Error-handling philosophy
1542
1810
 
1543
1811
  The library **warns and degrades; it does not throw** on operational problems:
@@ -1571,10 +1839,13 @@ lint + typecheck + **build** + test on every push/PR.
1571
1839
 
1572
1840
  **Project layout** (`src/`): `Observer.ts`, `ObservedCall.ts`, `ObservedClient.ts`,
1573
1841
  `ObservedPeerConnection.ts`, the `Observed*` sub-stat classes, `ObserverEvents.ts` (the typed
1574
- event map + scope types), `detectors/` (`Detector`, `Detectors`), `scores/`, `updaters/`
1575
- (update-policy strategies), `utils/` (remote-track resolvers), `common/` (`logger`, `utils`,
1576
- `Middleware`), `schema/` (sample/event/meta types), and `sinks/` (the `ClientSampleSink` base +
1577
- `JsonlFileSink` / `InMemorySink`, re-exported from the package root).
1842
+ event map + scope types), `detectors/` (`Detector`, `Detectors`, and one file per detector),
1843
+ `validators/` (`Validator`, `Validators`, one file per validator), `issues/` (`ActiveClientIssue`,
1844
+ `ActiveIssueTracker`, `ActiveIssuesRegistry`, `ObservedClientIssueRegistry`), `scores/`,
1845
+ `resolvers/` (remote-track resolvers), `utils/` (`stats`, `SlidingWindow`, `TrendTester`,
1846
+ `CallHealthAggregator`), `common/` (`logger`, `utils`, `Middleware`), `schema/`
1847
+ (sample/event/meta types), and `sinks/` (the `ClientSampleSink` base + `JsonlFileSink` /
1848
+ `InMemorySink`, re-exported from the package root).
1578
1849
 
1579
1850
  **Conventions to follow when developing further:**
1580
1851