@observertc/observer-js 1.0.0-beta.13 → 1.0.0-beta.14
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +396 -9
- package/dist/index.d.mts +1253 -8
- package/dist/index.d.ts +1253 -8
- package/dist/index.js +1952 -59
- package/dist/index.js.map +1 -1
- package/dist/index.mjs +1915 -58
- package/dist/index.mjs.map +1 -1
- package/package.json +3 -5
- package/llms-full.txt +0 -1632
package/README.md
CHANGED
|
@@ -40,11 +40,9 @@ and emits a single, unified stream of typed events the application can react to.
|
|
|
40
40
|
> works whether your project uses `import` (ESM) or `require()` (CommonJS). Everything — including
|
|
41
41
|
> the built-in file sink — is exported from the single `@observertc/observer-js` entry.
|
|
42
42
|
|
|
43
|
-
> **For AI agents:** [`llms.txt`](./llms.txt) is a curated map of these docs (
|
|
44
|
-
>
|
|
45
|
-
>
|
|
46
|
-
> ships in the npm package (so it's available wherever the library is installed); `llms.txt` and
|
|
47
|
-
> `AGENTS.md` live in the repo (and `llms.txt` belongs at the root of the docs site).
|
|
43
|
+
> **For AI agents:** [`llms.txt`](./llms.txt) is a curated map of these docs (it belongs at the root
|
|
44
|
+
> of the docs site); [`AGENTS.md`](./AGENTS.md) covers build/test commands and the conventions for
|
|
45
|
+
> working **in** this repository.
|
|
48
46
|
|
|
49
47
|
---
|
|
50
48
|
|
|
@@ -363,6 +361,7 @@ additional field(s) on top of that scope.
|
|
|
363
361
|
| `observer-updated` | — | `observer.update()` ran (per the observer update policy) |
|
|
364
362
|
| `observer-closed` | — | `observer.close()` |
|
|
365
363
|
| `sample-rejected` | `{ reason: 'observer-closed' \| 'missing-callId' \| 'missing-clientId', sample: ClientSample }` | a sample was dropped by `accept()` |
|
|
364
|
+
| `observer-issue` | `{ issue: ClientIssue }` | `observer.addIssue(...)` — a cross-call / SFU-wide finding (see [observer-level detectors](#observer-level-detectors-cross-call--sfu-wide)) |
|
|
366
365
|
|
|
367
366
|
#### Mediasoup level — scope `{ observer, observedMediasoupRouter }`
|
|
368
367
|
|
|
@@ -396,7 +395,8 @@ See [Mediasoup router observation](#mediasoup-router-observation) for the full d
|
|
|
396
395
|
| `client-joined` | — | first `CLIENT_JOINED` event seen |
|
|
397
396
|
| `client-left` | — | `CLIENT_LEFT` seen (or inferred on close) |
|
|
398
397
|
| `client-rejoined` | `{ timestamp: number }` | a later `CLIENT_JOINED` after an earlier join |
|
|
399
|
-
| `client-issue` | `{ issue: ClientIssue }` | a client-reported issue arrived, or `client.addIssue(...)` |
|
|
398
|
+
| `client-issue` | `{ issue: ClientIssue }` | a client-reported issue arrived, or `client.addIssue(...)`. A keyed issue also opens an entry in `observedClient.activeIssues` |
|
|
399
|
+
| `client-issue-resolved` | `{ resolvedIssue: ResolvedClientIssue }` | a stateful issue ended — the client sent its `<type>-resolved` companion, or the observer force-closed it. Carries the finished interval (`durationInMs`, `resolvedBy`) — see [client issues](#client-issues-the-lifecycle-and-the-division-of-labour) |
|
|
400
400
|
| `client-metadata` | `{ metaData: ClientMetaData }` | a client meta item arrived |
|
|
401
401
|
| `client-extension-stats` | `{ extensionStats: ExtensionStat }` | an app-defined extension stat arrived |
|
|
402
402
|
| `client-event` | `{ event: ClientEvent }` | any client event was processed |
|
|
@@ -562,11 +562,28 @@ Key members:
|
|
|
562
562
|
`iceCandidates`, `iceCandidatePairs`, `certificates`, `selectedIceCandidatePairs`,
|
|
563
563
|
`selectedIceCandiadtePairForTurn`
|
|
564
564
|
- State: `connectionState?`, `iceConnectionState?`, `iceGatheringState?`, `usingTURN`, `usingTCP`
|
|
565
|
-
- Metrics: `currentRttInMs?`, `
|
|
566
|
-
`availableOutgoingBitrate`, sending/receiving bitrates, packet rates,
|
|
567
|
-
byte/packet counters
|
|
565
|
+
- Metrics: `currentRttInMs?`, `iceRttInMs?`, `rtcpRttInMs?`, `sfuHopRttInMs?`, `currentJitter?`,
|
|
566
|
+
`availableIncomingBitrate`, `availableOutgoingBitrate`, sending/receiving bitrates, packet rates,
|
|
567
|
+
and `total*` / `delta*` byte/packet counters
|
|
568
568
|
- `accept(pcSample, context?)`, `close()`, `get score()`
|
|
569
569
|
|
|
570
|
+
**Two different round trips — don't mix them.** `iceRttInMs` comes from ICE/STUN consent checks and
|
|
571
|
+
measures the trip to *whatever terminates ICE*: **in an SFU topology that is the SFU**, so it is the
|
|
572
|
+
client↔SFU leg. `rtcpRttInMs` comes from RTCP receiver reports and is an **end-to-end** media-path
|
|
573
|
+
round trip. They are not interchangeable, and averaging them together produces a number that moves
|
|
574
|
+
as streams come and go for reasons unrelated to the network. `currentRttInMs` therefore *prefers*
|
|
575
|
+
RTCP and falls back to ICE — always one kind within a tick, never a blend. `sfuHopRttInMs`
|
|
576
|
+
(`rtcp − ice`) estimates everything past the SFU, which separates "this client's last mile is slow"
|
|
577
|
+
from "the path beyond the SFU is slow".
|
|
578
|
+
|
|
579
|
+
**Counter-reset boundaries.** Chrome resets an SSRC's cumulative counters when the codec switches
|
|
580
|
+
([crbug/webrtc/5361](https://bugs.chromium.org/p/webrtc/issues/detail?id=5361), open since 2015),
|
|
581
|
+
which otherwise shows up as a sawtooth spike or a negative bitrate. `ObservedInboundRtp` /
|
|
582
|
+
`ObservedOutboundRtp` therefore set **`counterResetBoundary`** on any tick where `codecId`,
|
|
583
|
+
`encoder`/`decoderImplementation` or `scalabilityMode` changed, and suppress every delta for that
|
|
584
|
+
tick. Without this, a room-wide codec rollout fires a synchronized fake-degradation alert across
|
|
585
|
+
every participant at once.
|
|
586
|
+
|
|
570
587
|
**Remote-RTP correlation (derived).** During `accept()`, receiver/sender reports are linked
|
|
571
588
|
to the local streams by `remoteId` (fallback SSRC) and surfaced as fields:
|
|
572
589
|
|
|
@@ -730,6 +747,302 @@ observer.on('call-added', ({ observedCall }) => {
|
|
|
730
747
|
observer.on('call-issue', ({ observedCall, issue }) => { /* react */ });
|
|
731
748
|
```
|
|
732
749
|
|
|
750
|
+
### Client issues: the lifecycle, and the division of labour
|
|
751
|
+
|
|
752
|
+
The most important thing to understand about detection in this library is **what it deliberately
|
|
753
|
+
does not do**. A client running
|
|
754
|
+
[`client-monitor-js`](https://github.com/ObserveRTC/client-monitor-js) already ships ~20 detectors
|
|
755
|
+
that decide *what is wrong with that endpoint* — `congestion`, `cpulimitation`, `audio-concealment`,
|
|
756
|
+
`freezed-video-track`, `keyframe-storm`, `video-decoder-overloaded`, `stuck-decoder`,
|
|
757
|
+
`ice-disconnected`, and so on. Those verdicts are better than anything re-derived from raw counters
|
|
758
|
+
server-side, because they carry hysteresis and multi-signal confirmation: `audio-concealment`
|
|
759
|
+
subtracts silent concealment (raw `concealedSamples` rises during ordinary silence, so a naive
|
|
760
|
+
detector flags every quiet moment); `audio-jitter-buffer-stress` requires the buffer to be grown
|
|
761
|
+
**and** NetEQ to be time-stretching (a grown buffer alone means NetEQ is *succeeding*);
|
|
762
|
+
`ice-disconnected` only fires once `disconnected` has persisted, so the blips ICE heals on its own
|
|
763
|
+
never surface.
|
|
764
|
+
|
|
765
|
+
**observer-js does not repeat that work.** Its job is the question no browser can answer: *who else
|
|
766
|
+
is in this state right now, what do they have in common, and where in publisher → SFU → subscriber
|
|
767
|
+
does the fault begin?*
|
|
768
|
+
|
|
769
|
+
#### The wire format
|
|
770
|
+
|
|
771
|
+
From client-monitor-js **4.6.0** the whole issue lifecycle reaches the server. A stateful issue
|
|
772
|
+
arrives as two `clientIssues[]` entries sharing a `key`:
|
|
773
|
+
|
|
774
|
+
```
|
|
775
|
+
raise: { type: 'stuck-decoder', key, payload, timestamp: raisedAt }
|
|
776
|
+
resolution: { type: 'stuck-decoder-resolved', key, payload: { raisedAt, comment, …final }, timestamp: resolvedAt }
|
|
777
|
+
```
|
|
778
|
+
|
|
779
|
+
The observer opens an entry in `observedClient.activeIssues` on the raise and closes it on the
|
|
780
|
+
matching key, emitting **`client-issue-resolved`** with the finished interval. Handled for you:
|
|
781
|
+
|
|
782
|
+
- the `-resolved` **suffix** is stripped, so both entries share one logical `type`;
|
|
783
|
+
- a **re-raise** of a live key refreshes the payload without restarting `raisedAt`;
|
|
784
|
+
- **keyless** entries are one-shot — reported via `client-issue`, never tracked;
|
|
785
|
+
- issues still open when a client closes are **force-resolved** (`resolvedBy: 'client-closed'`), and
|
|
786
|
+
the registry additionally expires stale entries, so a crashed participant can't leave an issue
|
|
787
|
+
"active" forever.
|
|
788
|
+
|
|
789
|
+
#### Why intervals beat windows
|
|
790
|
+
|
|
791
|
+
This turns point-in-time symptom reports into **intervals**, and that is the whole game. *"Several
|
|
792
|
+
clients reported congestion in the last 10 seconds"* is a heuristic that has to guess whether the
|
|
793
|
+
symptoms are still happening. *"Several clients are congested **right now, simultaneously**"* is
|
|
794
|
+
ground truth, because the client says when the episode ends. Overlapping intervals are far stronger
|
|
795
|
+
evidence of a shared cause than near-in-time reports.
|
|
796
|
+
|
|
797
|
+
```ts
|
|
798
|
+
observer.on('client-issue', ({ observedClient, issue }) => { /* opened (or one-shot) */ });
|
|
799
|
+
observer.on('client-issue-resolved', ({ resolvedIssue }) => {
|
|
800
|
+
resolvedIssue.type; // 'stuck-decoder' — suffix stripped
|
|
801
|
+
resolvedIssue.durationInMs; // how long the episode lasted
|
|
802
|
+
resolvedIssue.resolvedBy; // 'client' | 'timeout' | 'client-closed'
|
|
803
|
+
});
|
|
804
|
+
|
|
805
|
+
// the live per-client mirror
|
|
806
|
+
observedClient.activeIssues; // Map<key, ActiveClientIssue>
|
|
807
|
+
```
|
|
808
|
+
|
|
809
|
+
#### `IssueRegistry` — querying the active set
|
|
810
|
+
|
|
811
|
+
```ts
|
|
812
|
+
import { IssueRegistry } from '@observertc/observer-js';
|
|
813
|
+
|
|
814
|
+
const registry = new IssueRegistry(observedCall); // or `observer` for cross-call scope
|
|
815
|
+
|
|
816
|
+
registry.cohortOf('congestion'); // { clientIds, affectedRatio, onsetSpreadInMs, … }
|
|
817
|
+
registry.cohorts(); // every shared issue type, largest cohort first
|
|
818
|
+
registry.byTrackIds(trackIds); // issues attributed to a published track's receivers
|
|
819
|
+
```
|
|
820
|
+
|
|
821
|
+
`onsetSpreadInMs` is measured on the **observer clock**, never the client's. `raisedAt` comes from
|
|
822
|
+
each participant's own machine, and comparing those across clients makes clock skew look like a
|
|
823
|
+
synchronized infrastructure event.
|
|
824
|
+
|
|
825
|
+
### Observer-level detectors (cross-call / SFU-wide)
|
|
826
|
+
|
|
827
|
+
Some findings only exist **above** call scope — "many calls on the same SFU degraded at once" is far
|
|
828
|
+
more actionable than fifty individual client alerts. The same registry exists on the `Observer`,
|
|
829
|
+
runs on every `observer.update()`, and raises findings through `observer.addIssue(...)`, surfaced on
|
|
830
|
+
the bus as **`observer-issue`**:
|
|
831
|
+
|
|
832
|
+
```ts
|
|
833
|
+
observer.detectors.add({
|
|
834
|
+
name: 'sfu-wide-degradation',
|
|
835
|
+
update: () => {
|
|
836
|
+
const degradedCalls = [ ...observer.observedCalls.values() ].filter(isDegraded);
|
|
837
|
+
|
|
838
|
+
if (observer.numberOfCalls > 3 && degradedCalls.length / observer.numberOfCalls > 0.6) {
|
|
839
|
+
observer.addIssue({ type: 'SFU_WIDE_QUALITY_DEGRADATION', timestamp: Date.now() });
|
|
840
|
+
}
|
|
841
|
+
},
|
|
842
|
+
});
|
|
843
|
+
|
|
844
|
+
observer.on('observer-issue', ({ issue }) => alert(issue));
|
|
845
|
+
```
|
|
846
|
+
|
|
847
|
+
### Publisher → subscribers: `TrackDistributionAggregator`
|
|
848
|
+
|
|
849
|
+
The question a single browser can never answer is *"did **everyone** receiving Alice see the same
|
|
850
|
+
degradation?"*. `TrackDistributionAggregator` answers it by walking the publisher→subscriber links
|
|
851
|
+
maintained by a [`RemoteTrackResolver`](#remote-track-resolution-mediasoup--sfu) and summarizing one
|
|
852
|
+
published track against all of its receivers:
|
|
853
|
+
|
|
854
|
+
```ts
|
|
855
|
+
import { TrackDistributionAggregator } from '@observertc/observer-js';
|
|
856
|
+
|
|
857
|
+
const aggregator = new TrackDistributionAggregator(observedCall);
|
|
858
|
+
|
|
859
|
+
for (const d of aggregator.aggregate()) {
|
|
860
|
+
d.trackId; // the published track
|
|
861
|
+
d.publisher.healthy; // is the source itself fine? (+ .reasons)
|
|
862
|
+
d.numberOfReceivers; // 17
|
|
863
|
+
d.numberOfDegradedReceivers; // 14
|
|
864
|
+
d.degradedRatio; // 0.82
|
|
865
|
+
d.fractionLost?.p95; // percentile summaries across receivers
|
|
866
|
+
d.freezes; // { affectedReceivers, total } — fan-out counters
|
|
867
|
+
d.plis;
|
|
868
|
+
d.receivers; // per-receiver entries with `degraded` + `reasons`
|
|
869
|
+
}
|
|
870
|
+
```
|
|
871
|
+
|
|
872
|
+
Each receiver is judged against `ReceiverHealthThresholds` (loss, freezes, dropped frames,
|
|
873
|
+
concealment, jitter-buffer delay, RTT — override any of them). Summaries are **medians and
|
|
874
|
+
percentiles, not means**, because one participant at 1500 ms RTT would otherwise hide nine healthy
|
|
875
|
+
ones. The helpers behind it (`percentile`, `median`, `summarize`, `counterDelta`, `SlidingWindow`)
|
|
876
|
+
are exported for building your own detectors.
|
|
877
|
+
|
|
878
|
+
### Call health: `CallHealthAggregator`
|
|
879
|
+
|
|
880
|
+
The second aggregation axis — the **client** axis. Where `TrackDistributionAggregator` asks "how was
|
|
881
|
+
*this source* delivered?", this asks "how is *each participant* doing, sending vs receiving?":
|
|
882
|
+
|
|
883
|
+
```ts
|
|
884
|
+
import { CallHealthAggregator } from '@observertc/observer-js';
|
|
885
|
+
|
|
886
|
+
const health = new CallHealthAggregator(observedCall).aggregate();
|
|
887
|
+
|
|
888
|
+
health.degradedRatio; // 0.82 — the number that distinguishes shared faults from individual ones
|
|
889
|
+
health.inboundDegradedRatio; // receiving side → egress/downstream suspicion
|
|
890
|
+
health.outboundDegradedRatio; // sending side → ingress suspicion
|
|
891
|
+
health.rttInMs?.median; // percentile rollups, never means
|
|
892
|
+
health.qualityLimitation; // { cpu, bandwidth, other } client counts
|
|
893
|
+
health.clients; // per-client entries with `reasons`, direction flags, TURN/TCP
|
|
894
|
+
```
|
|
895
|
+
|
|
896
|
+
### Built-in detectors
|
|
897
|
+
|
|
898
|
+
All of them are **opt-in** — the library enables none by default. Register call-scoped ones on
|
|
899
|
+
`observedCall.detectors` and observer-scoped ones on `observer.detectors`:
|
|
900
|
+
|
|
901
|
+
> **🔗 marks detectors that require a
|
|
902
|
+
> [`RemoteTrackResolver`](#remote-track-resolution-mediasoup--sfu).** They reason about a published
|
|
903
|
+
> track and its subscribers, so without the publisher↔subscriber links they see nothing and stay
|
|
904
|
+
> **silent forever** — which looks exactly like "no problems found". Configure
|
|
905
|
+
> `ObserverConfig.createRemoteTrackResolver` before registering them.
|
|
906
|
+
|
|
907
|
+
**Issue-driven** (preferred — they consume the client's own verdicts):
|
|
908
|
+
|
|
909
|
+
| Detector | 🔗 | Scope | Raises |
|
|
910
|
+
|----------|:--:|-------|--------|
|
|
911
|
+
| `ConcurrentIssueDetector` | | call **or** observer | `CONCURRENT_CLIENT_ISSUES`, `ISSUE_ONSET_BURST` |
|
|
912
|
+
| `IssueFanOutDetector` | 🔗 | call | `PUBLISHED_TRACK_ISSUE_FAN_OUT`, `SINGLE_RECEIVER_ISSUE` |
|
|
913
|
+
| `TrackDeliveryMismatchDetector` | 🔗 | call | `PUBLISHED_TRACK_NOT_DELIVERED`, `RECEIVER_TRACK_NOT_DELIVERED`, `PUBLISHER_TRACK_DRY` |
|
|
914
|
+
|
|
915
|
+
**Metric-driven** (work without client issues — the fallback path, and the things no client issue
|
|
916
|
+
can express):
|
|
917
|
+
|
|
918
|
+
| Detector | 🔗 | Scope | Raises |
|
|
919
|
+
|----------|:--:|-------|--------|
|
|
920
|
+
| `WorstReceiverContagionDetector` | 🔗 | call | `WORST_RECEIVER_CONTAGION` |
|
|
921
|
+
| `CommonSourceDegradationDetector` | 🔗 | call | `PUBLISHER_HEALTHY_SUBSCRIBERS_DEGRADED`, `PUBLISHER_DEGRADED_FOR_ALL_SUBSCRIBERS`, `SINGLE_SUBSCRIBER_DEGRADED`, `MULTIPLE_SUBSCRIBERS_DEGRADED` |
|
|
922
|
+
| `PliAndFreezeFanOutDetector` | 🔗 | call | `PUBLISHER_PLI_STORM`, `PUBLISHED_VIDEO_FROZEN_FOR_MULTIPLE_RECEIVERS` |
|
|
923
|
+
| `AudioImpairmentFanOutDetector` | 🔗 | call | `PUBLISHED_AUDIO_DEGRADED_FOR_MAJORITY`, `CALL_WIDE_AUDIO_JITTER_BUFFER_STRESS` |
|
|
924
|
+
| `UnconsumedTrackDetector` | 🔗 | call | `UNCONSUMED_PUBLISHED_TRACK` |
|
|
925
|
+
| `CallWideDegradationDetector` | | call | `CALL_WIDE_QUALITY_DEGRADATION`, `CALL_WIDE_INBOUND_DEGRADATION`, `CALL_WIDE_OUTBOUND_DEGRADATION` |
|
|
926
|
+
| `IceDisruptionDetector` | | call | `CALL_ICE_DISRUPTION` |
|
|
927
|
+
| `TurnServerHealthDetector` | | **observer** | `TURN_SERVER_DEGRADED` |
|
|
928
|
+
|
|
929
|
+
If your clients report issues, the issue-driven pair **subsumes** the fan-out and ICE detectors —
|
|
930
|
+
`IssueFanOutDetector` covers freeze/PLI/concealment fan-out generically, and `ConcurrentIssueDetector`
|
|
931
|
+
with the ICE issue types replaces `IceDisruptionDetector` (and does it better: the client already
|
|
932
|
+
suppresses the transient blips ICE heals by itself). Run both families only while migrating.
|
|
933
|
+
|
|
934
|
+
#### `TrackDeliveryMismatchDetector` — resolving an ambiguous symptom
|
|
935
|
+
|
|
936
|
+
A dry track ("no bytes are arriving") is the clearest symptom there is and, on its own, completely
|
|
937
|
+
ambiguous. A receiver seeing silence cannot distinguish *the camera was switched off* from *the SFU
|
|
938
|
+
stopped forwarding* from *my own consumer wedged* — all three look identical from the browser.
|
|
939
|
+
|
|
940
|
+
Joining the two ends of the published track resolves it:
|
|
941
|
+
|
|
942
|
+
| publisher | subscribers | verdict |
|
|
943
|
+
|---|---|---|
|
|
944
|
+
| sending | **all** dry | `PUBLISHED_TRACK_NOT_DELIVERED` — the forwarding path |
|
|
945
|
+
| sending | **some** dry | `RECEIVER_TRACK_NOT_DELIVERED` — those consumers (in mediasoup: recreate them) |
|
|
946
|
+
| dry | any dry | `PUBLISHER_TRACK_DRY` — the source stopped; **not** an SFU fault |
|
|
947
|
+
|
|
948
|
+
The publisher side is judged from both available signals: its own `dry-outbound-track` issue when the
|
|
949
|
+
client reports one, and the observed outbound RTP (`deltaPacketsSent`) as fallback and corroboration.
|
|
950
|
+
That combination is what makes the first row trustworthy — the server can state that packets
|
|
951
|
+
demonstrably left the publisher during the same interval in which every receiver got nothing.
|
|
952
|
+
|
|
953
|
+
This is the "SFU forwarding mismatch" check, and it needs **no mediasoup instrumentation at all** —
|
|
954
|
+
the clients' own dry-track verdicts plus the resolver links are sufficient.
|
|
955
|
+
|
|
956
|
+
#### `UnconsumedTrackDetector` — reading the resolver's silence
|
|
957
|
+
|
|
958
|
+
The one detector where the *absence* of links is the signal: a track still pushing packets whose
|
|
959
|
+
`remoteInboundTracks` set is empty, i.e. uplink and SFU ingress spent on media nobody receives
|
|
960
|
+
(everyone has the publisher hidden, a simulcast layer no viewer selects, or an app that forgot to
|
|
961
|
+
stop a track). It waits `minUnconsumedDurationInMs` first, since a gap between publishing and the
|
|
962
|
+
first subscription is normal at join time.
|
|
963
|
+
|
|
964
|
+
Note the trap this one has to guard against, and why it checks `call.remoteTrackResolver` at runtime
|
|
965
|
+
rather than trusting the flag alone: **"no subscribers" and "no resolver configured" produce the
|
|
966
|
+
identical observation.** Without a resolver it would report every published track in the call as
|
|
967
|
+
unconsumed.
|
|
968
|
+
|
|
969
|
+
#### `WorstReceiverContagionDetector` — the one worth reading about
|
|
970
|
+
|
|
971
|
+
In a correctly built SFU the RTCP feedback loop is **terminated at the server**: each receiver's
|
|
972
|
+
reports drive what *that* receiver is sent. When the loop is relayed end-to-end instead, the
|
|
973
|
+
publisher's bandwidth estimate collapses to the minimum across all receivers — so one participant on
|
|
974
|
+
a bad 3G link silently downgrades the stream **everyone** sees. This is the "lowest common
|
|
975
|
+
denominator" failure simulcast exists to prevent.
|
|
976
|
+
|
|
977
|
+
It's detected as a *correlation over a window*, not a threshold: the publisher's outbound bitrate
|
|
978
|
+
moving in lockstep with the **worst** receiver's inbound bitrate, while the median receiver has
|
|
979
|
+
headroom — and tracking the worst receiver more closely than the median (otherwise everyone is just
|
|
980
|
+
moving together, which is ordinary adaptation). The damage is invisible from every endpoint: the
|
|
981
|
+
publisher sees "my bitrate dropped", each healthy receiver sees "my video got worse", and only the
|
|
982
|
+
server can see the causal link.
|
|
983
|
+
|
|
984
|
+
```ts
|
|
985
|
+
import {
|
|
986
|
+
ConcurrentIssueDetector, IssueFanOutDetector, WorstReceiverContagionDetector,
|
|
987
|
+
CallWideDegradationDetector, TurnServerHealthDetector,
|
|
988
|
+
} from '@observertc/observer-js';
|
|
989
|
+
|
|
990
|
+
observer.on('call-added', ({ observedCall }) => {
|
|
991
|
+
// issue-driven: correlate the verdicts the clients already reached
|
|
992
|
+
observedCall.detectors.add(new IssueFanOutDetector(observedCall));
|
|
993
|
+
observedCall.detectors.add(new ConcurrentIssueDetector(observedCall));
|
|
994
|
+
// metric-driven: things no client issue can express
|
|
995
|
+
observedCall.detectors.add(new WorstReceiverContagionDetector(observedCall));
|
|
996
|
+
observedCall.detectors.add(new CallWideDegradationDetector(observedCall));
|
|
997
|
+
});
|
|
998
|
+
|
|
999
|
+
// cross-call, so these go on the observer and raise `observer-issue`
|
|
1000
|
+
observer.detectors.add(new TurnServerHealthDetector(observer));
|
|
1001
|
+
observer.detectors.add(new ConcurrentIssueDetector(observer, {
|
|
1002
|
+
issueTypes: [ 'ice-disconnected', 'ice-connection-failed', 'congestion' ],
|
|
1003
|
+
}));
|
|
1004
|
+
|
|
1005
|
+
observer.on('call-issue', ({ observedCall, issue }) => handle(observedCall, issue));
|
|
1006
|
+
observer.on('observer-issue', ({ issue }) => handle(undefined, issue));
|
|
1007
|
+
```
|
|
1008
|
+
|
|
1009
|
+
Every detector takes an options object to tune `minReceivers`/`minClients`, the ratio thresholds, the
|
|
1010
|
+
per-entity health thresholds, and the debounce (`consecutiveTicks`, or `windowMs` + `cooldownMs` for
|
|
1011
|
+
the window-based ones) — so an alert needs a condition to *persist*, not just appear in one sample.
|
|
1012
|
+
Findings that depend on publisher↔subscriber links need a
|
|
1013
|
+
[`RemoteTrackResolver`](#remote-track-resolution-mediasoup--sfu); without one those detectors stay
|
|
1014
|
+
silent. A detector may implement `close()` (called when it's removed or the call/observer closes) —
|
|
1015
|
+
`IceDisruptionDetector` uses it to drop the bus listeners it subscribes with.
|
|
1016
|
+
|
|
1017
|
+
### `CommonSourceDegradationDetector`
|
|
1018
|
+
|
|
1019
|
+
The first detector built on the aggregator. It classifies where a fault lies by comparing the source
|
|
1020
|
+
against its receivers, and raises a `call-issue` whose `type` is one of:
|
|
1021
|
+
|
|
1022
|
+
| Finding | Meaning |
|
|
1023
|
+
|---------|---------|
|
|
1024
|
+
| `PUBLISHER_HEALTHY_SUBSCRIBERS_DEGRADED` | source egress fine, most receivers degraded → **downstream / SFU** suspected |
|
|
1025
|
+
| `PUBLISHER_DEGRADED_FOR_ALL_SUBSCRIBERS` | the source itself is impaired → **publisher-side** |
|
|
1026
|
+
| `SINGLE_SUBSCRIBER_DEGRADED` | one receiver suffers while the rest are fine → **that receiver's network** |
|
|
1027
|
+
| `MULTIPLE_SUBSCRIBERS_DEGRADED` | several (but not most) receivers on the same source |
|
|
1028
|
+
|
|
1029
|
+
```ts
|
|
1030
|
+
import { CommonSourceDegradationDetector } from '@observertc/observer-js';
|
|
1031
|
+
|
|
1032
|
+
observer.on('call-added', ({ observedCall }) => {
|
|
1033
|
+
observedCall.detectors.add(new CommonSourceDegradationDetector(observedCall, {
|
|
1034
|
+
minReceivers: 3, // ratios need a meaningful denominator
|
|
1035
|
+
degradedRatioThreshold: 0.6, // "most receivers"
|
|
1036
|
+
consecutiveTicks: 2, // must hold 2 ticks — avoids flapping on one bad sample
|
|
1037
|
+
}));
|
|
1038
|
+
});
|
|
1039
|
+
```
|
|
1040
|
+
|
|
1041
|
+
The `call-issue` payload carries the evidence: publisher health and reasons, receiver/degraded
|
|
1042
|
+
counts, `degradedRatio`, `affectedClientIds`, freeze/PLI fan-out and the loss/bitrate summaries.
|
|
1043
|
+
It requires a configured `RemoteTrackResolver` — with no links there is nothing to compare and the
|
|
1044
|
+
detector stays silent.
|
|
1045
|
+
|
|
733
1046
|
`Detector` interface and the registry:
|
|
734
1047
|
|
|
735
1048
|
```ts
|
|
@@ -836,6 +1149,80 @@ on your own cadence read `observedRouter.sample` (snapshot/serialize/persist wha
|
|
|
836
1149
|
you don't, and close routers you no longer track. (mediasoup also typically shards routers across
|
|
837
1150
|
workers/cores, which keeps any one router small.)
|
|
838
1151
|
|
|
1152
|
+
### Extending the sample, and building your own report
|
|
1153
|
+
|
|
1154
|
+
The sample is yours to annotate. Every entity — the router, each transport, producer, consumer, data
|
|
1155
|
+
producer and data consumer — has an `attachments?: Record<string, unknown>` slot, and there are three
|
|
1156
|
+
ways to fill it, from most declarative to most ad-hoc.
|
|
1157
|
+
|
|
1158
|
+
**1. `enrich` — mirror mediasoup's own `appData`.** The common case: your application already keeps
|
|
1159
|
+
`participantId`, `purpose` and similar on the mediasoup objects, and you want them on the sample.
|
|
1160
|
+
Runs once per entity at creation, before the corresponding event:
|
|
1161
|
+
|
|
1162
|
+
```ts
|
|
1163
|
+
observer.createObservedMediasoupRouter({
|
|
1164
|
+
router,
|
|
1165
|
+
enrich: {
|
|
1166
|
+
producer: (producer) => ({ participantId: producer.appData.participantId, purpose: producer.appData.purpose }),
|
|
1167
|
+
consumer: (consumer) => ({ subscriberId: consumer.appData.subscriberId }),
|
|
1168
|
+
transport: (transport) => ({ role: transport.appData.role }),
|
|
1169
|
+
},
|
|
1170
|
+
});
|
|
1171
|
+
```
|
|
1172
|
+
|
|
1173
|
+
A throwing enricher is caught and logged — it can't take the router's bookkeeping down with it.
|
|
1174
|
+
|
|
1175
|
+
**2. Lifecycle events — enrich on the fly.** Each entity announces itself as
|
|
1176
|
+
`<entity>-sample-added` and `<entity>-sample-closed`, carrying **the live sample object** (not a
|
|
1177
|
+
copy) plus the mediasoup object it came from. Mutating it in the handler is the intended pattern:
|
|
1178
|
+
|
|
1179
|
+
```ts
|
|
1180
|
+
observedRouter.on('producer-sample-added', ({ sample, producer, transport }) => {
|
|
1181
|
+
sample.attachments = { ...sample.attachments, participantId: lookup(producer.id) };
|
|
1182
|
+
});
|
|
1183
|
+
|
|
1184
|
+
observedRouter.on('producer-sample-closed', ({ sample }) => {
|
|
1185
|
+
archive(sample); // its `closedAt` is set
|
|
1186
|
+
});
|
|
1187
|
+
```
|
|
1188
|
+
|
|
1189
|
+
Events: `transport-sample-added` / `-closed`, `producer-sample-added` / `-closed`,
|
|
1190
|
+
`consumer-sample-added` / `-closed`, `data-producer-sample-added` / `-closed`,
|
|
1191
|
+
`data-consumer-sample-added` / `-closed`.
|
|
1192
|
+
|
|
1193
|
+
**3. `attachTo(id, attachments)` — annotate later, from anywhere.** When the knowledge arrives after
|
|
1194
|
+
the entity did (a signalling message, a database lookup that resolved):
|
|
1195
|
+
|
|
1196
|
+
```ts
|
|
1197
|
+
observedRouter.attachTo(producerId, { participantId, joinedFrom: 'mobile' }); // merges
|
|
1198
|
+
```
|
|
1199
|
+
|
|
1200
|
+
Ids are unique across mediasoup entity kinds, so one method covers all of them. It returns `false`
|
|
1201
|
+
for an unknown id rather than failing quietly — which matters when application events race the
|
|
1202
|
+
mediasoup ones. For direct access there are typed accessors: `getTransportSample(id)`,
|
|
1203
|
+
`getProducerSample(id)`, `getConsumerSample(id)`, `getDataProducerSample(id)`,
|
|
1204
|
+
`getDataConsumerSample(id)`. They index the *same* objects the arrays hold, so a lookup is O(1)
|
|
1205
|
+
instead of a `sample.producers.find(...)` scan.
|
|
1206
|
+
|
|
1207
|
+
#### Building your own report
|
|
1208
|
+
|
|
1209
|
+
`observedRouter.sample` is live — arrays grow and `history` entries are appended as the router runs,
|
|
1210
|
+
so a report built directly on it keeps changing after you think you're done. Use **`snapshot()`** for
|
|
1211
|
+
a detached deep copy:
|
|
1212
|
+
|
|
1213
|
+
```ts
|
|
1214
|
+
const report = {
|
|
1215
|
+
...observedRouter.snapshot(), // never moves again
|
|
1216
|
+
generatedAt: Date.now(),
|
|
1217
|
+
region: process.env.REGION,
|
|
1218
|
+
};
|
|
1219
|
+
```
|
|
1220
|
+
|
|
1221
|
+
> **Note on typing.** The sample types no longer carry a `Record<string, unknown>` index signature.
|
|
1222
|
+
> That signature allowed arbitrary top-level keys but also silently accepted typos on real fields and
|
|
1223
|
+
> weakened autocomplete. Custom data belongs in `attachments`, which is typed as such. If you were
|
|
1224
|
+
> assigning ad-hoc keys directly onto a sample object, move them into `attachments`.
|
|
1225
|
+
|
|
839
1226
|
### Matching peer connections — by **event**, not by storage
|
|
840
1227
|
|
|
841
1228
|
The observer correlates the SFU side with the client side **at the peer-connection level**: a
|