@observertc/observer-js 1.0.0-beta.9 → 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1174 -166
- package/dist/index.d.mts +4012 -208
- package/dist/index.d.ts +4012 -208
- package/dist/index.js +4149 -601
- package/dist/index.js.map +1 -1
- package/dist/index.mjs +4091 -600
- package/dist/index.mjs.map +1 -1
- package/package.json +15 -4
package/README.md
CHANGED
|
@@ -31,8 +31,9 @@ and emits a single, unified stream of typed events the application can react to.
|
|
|
31
31
|
- **Drop it in safely** — warn-don't-throw, a pluggable logger, dual **ESM + CommonJS**, and **no**
|
|
32
32
|
media-stack dependency in the core.
|
|
33
33
|
|
|
34
|
-
> **Status:** `1.0.0
|
|
35
|
-
> against directly
|
|
34
|
+
> **Status:** `1.0.0` — the first stable release. The API described here is current and intended to
|
|
35
|
+
> be implemented against directly; it follows semver from here, so a breaking change means a 2.0.0.
|
|
36
|
+
> [`CHANGELOG.md`](./CHANGELOG.md) is the full statement of what the 1.0.0 line contains. This document is written to be self-sufficient: an engineer (or an AI
|
|
36
37
|
> agent) should be able to integrate the library, or develop it further, from this file alone.
|
|
37
38
|
> A companion doc, [`docs/logging.md`](./docs/logging.md), covers logging integration in depth.
|
|
38
39
|
|
|
@@ -40,6 +41,10 @@ and emits a single, unified stream of typed events the application can react to.
|
|
|
40
41
|
> works whether your project uses `import` (ESM) or `require()` (CommonJS). Everything — including
|
|
41
42
|
> the built-in file sink — is exported from the single `@observertc/observer-js` entry.
|
|
42
43
|
|
|
44
|
+
> **For AI agents:** [`llms.txt`](./llms.txt) is a curated map of these docs (it belongs at the root
|
|
45
|
+
> of the docs site); [`AGENTS.md`](./AGENTS.md) covers build/test commands and the conventions for
|
|
46
|
+
> working **in** this repository.
|
|
47
|
+
|
|
43
48
|
---
|
|
44
49
|
|
|
45
50
|
## Table of contents
|
|
@@ -49,17 +54,20 @@ and emits a single, unified stream of typed events the application can react to.
|
|
|
49
54
|
3. [Data flow](#data-flow)
|
|
50
55
|
4. [Entity hierarchy](#entity-hierarchy)
|
|
51
56
|
5. [Ingestion: `accept()`, context & lifecycle](#ingestion-accept-context--lifecycle)
|
|
52
|
-
6. [
|
|
57
|
+
6. [When things update](#when-things-update)
|
|
53
58
|
7. [The event bus](#the-event-bus) ← the core of the API
|
|
54
59
|
8. [API reference](#api-reference)
|
|
55
60
|
9. [Schema types (`ClientSample`)](#schema-types-clientsample)
|
|
56
61
|
10. [Detectors (server-side extension point)](#detectors-server-side-extension-point)
|
|
57
|
-
11. [
|
|
58
|
-
12. [
|
|
59
|
-
13. [
|
|
60
|
-
14. [
|
|
61
|
-
15. [
|
|
62
|
-
16. [
|
|
62
|
+
11. [Call summaries](#call-summaries)
|
|
63
|
+
12. [Remote track resolution (mediasoup / SFU)](#remote-track-resolution-mediasoup--sfu)
|
|
64
|
+
13. [Mediasoup router observation](#mediasoup-router-observation)
|
|
65
|
+
14. [Sinks (per-client sample persistence)](#sinks-per-client-sample-persistence)
|
|
66
|
+
15. [Injecting data into a client](#injecting-data-into-a-client)
|
|
67
|
+
16. [Logging](#logging)
|
|
68
|
+
17. [Design notes](#design-notes)
|
|
69
|
+
18. [Error-handling philosophy](#error-handling-philosophy)
|
|
70
|
+
19. [Development & extension guide](#development--extension-guide)
|
|
63
71
|
|
|
64
72
|
---
|
|
65
73
|
|
|
@@ -98,10 +106,9 @@ import { Observer, ClientSample } from '@observertc/observer-js';
|
|
|
98
106
|
|
|
99
107
|
// 1. Create an observer.
|
|
100
108
|
const observer = new Observer({
|
|
101
|
-
// when the observer
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
defaultCallUpdatePolicy: 'update-on-any-client-updated',
|
|
109
|
+
// a call updates when any of its clients does, and the observer when any of its calls does —
|
|
110
|
+
// both default to true, so this line is only here to show the knob exists:
|
|
111
|
+
autoUpdateOnCallUpdate: true,
|
|
105
112
|
// optional auto-teardown:
|
|
106
113
|
closeCallIfEmptyForMs: 20_000,
|
|
107
114
|
closeClientIfIdleForMs: 60_000,
|
|
@@ -238,8 +245,7 @@ observer.addAcceptMiddleware(route, filter);
|
|
|
238
245
|
// observer.removeAcceptMiddleware(route);
|
|
239
246
|
```
|
|
240
247
|
|
|
241
|
-
This is a lightweight global injection point
|
|
242
|
-
`ClientSampleProcessor` pipeline in the roadmap. When no middleware is registered, `accept()`
|
|
248
|
+
This is a lightweight global injection point. When no middleware is registered, `accept()`
|
|
243
249
|
dispatches directly with no overhead.
|
|
244
250
|
|
|
245
251
|
### `context` (the `AcceptContext`)
|
|
@@ -257,7 +263,10 @@ deliberately distinct:
|
|
|
257
263
|
|
|
258
264
|
- **`appData`** — application-assigned extra info that identifies/decorates an entity, fixed at
|
|
259
265
|
creation (via `settings.appData` or the `createCallAppData` / `createClientAppData` factories),
|
|
260
|
-
or assigned by the app on the `*-added` events. The library never changes it.
|
|
266
|
+
or assigned by the app on the `*-added` events. The library never changes it. The factories
|
|
267
|
+
**receive the context of the `accept()` that triggered the creation**, so a fact carried on the
|
|
268
|
+
context can be baked into `appData` at birth — but it is copied by the factory, deliberately, not
|
|
269
|
+
written across by the library.
|
|
261
270
|
- **`context`** — passed per `accept()`, may differ on every call, and is carried straight
|
|
262
271
|
through to the `*-updated` events that the `accept()` triggers, then discarded.
|
|
263
272
|
|
|
@@ -286,30 +295,38 @@ These return `undefined` (and warn) when the parent is closed; `createObservedCa
|
|
|
286
295
|
|
|
287
296
|
---
|
|
288
297
|
|
|
289
|
-
##
|
|
298
|
+
## When things update
|
|
299
|
+
|
|
300
|
+
"Update" means *recompute aggregated metrics, run the detectors, and emit the `*-updated` event* at
|
|
301
|
+
that level. Updates are **event-driven** — there is no built-in timer.
|
|
290
302
|
|
|
291
|
-
|
|
292
|
-
Both the observer and each call have a configurable trigger. Updates are **event-driven** — there
|
|
293
|
-
is no built-in timer. An app that wants a fixed cadence can call `observer.update()` /
|
|
294
|
-
`call.update()` from its own `setInterval`. With `'none'`, **nothing auto-updates** — the level
|
|
295
|
-
updates only when the application calls the public `update()` itself.
|
|
303
|
+
The rule is structural rather than configurable:
|
|
296
304
|
|
|
297
|
-
**
|
|
305
|
+
> **A call is updated when any of its clients is updated. The observer is updated when any of its
|
|
306
|
+
> calls is updated.** Composed, that means the observer is updated exactly when any client anywhere
|
|
307
|
+
> is updated.
|
|
298
308
|
|
|
299
|
-
|
|
300
|
-
|--------|-------------------------------------|
|
|
301
|
-
| `update-on-any-call-updated` | any call updates |
|
|
302
|
-
| `update-when-all-call-updated` | every call has updated since the last observer update |
|
|
303
|
-
| `none` | never automatically — only when the app calls `observer.update()` |
|
|
309
|
+
Two booleans, both defaulting to `true`, let you opt out of a link in that chain:
|
|
304
310
|
|
|
305
|
-
|
|
306
|
-
|
|
311
|
+
| Setting | Where | Effect when `false` |
|
|
312
|
+
|---------|-------|---------------------|
|
|
313
|
+
| `autoUpdateOnClientUpdate` | `ObservedCallSettings` | the call updates only when you call `call.update()` |
|
|
314
|
+
| `autoUpdateOnCallUpdate` | `ObserverConfig` | the observer updates only when you call `observer.update()` |
|
|
307
315
|
|
|
308
|
-
|
|
309
|
-
|
|
310
|
-
|
|
311
|
-
|
|
312
|
-
|
|
316
|
+
An app that wants a fixed cadence sets both to `false` and drives `observer.update()` from its own
|
|
317
|
+
`setInterval`. Note that **observer-scoped detectors and validators run nowhere else** — if the
|
|
318
|
+
observer never updates, they never run.
|
|
319
|
+
|
|
320
|
+
```ts
|
|
321
|
+
const observer = new Observer({ autoUpdateOnCallUpdate: false });
|
|
322
|
+
|
|
323
|
+
setInterval(() => observer.update(), 5_000);
|
|
324
|
+
```
|
|
325
|
+
|
|
326
|
+
> Earlier versions had an `updatePolicy` / `defaultCallUpdatePolicy` enum (`'update-on-any-…'`,
|
|
327
|
+
> `'update-when-all-…'`, `'none'`) and a pluggable `Updater` object. Both are gone. "When all clients
|
|
328
|
+
> have updated" sounds appealing and deadlocks on the first client that stops sending — one silent
|
|
329
|
+
> participant froze the whole call's aggregation until it timed out.
|
|
313
330
|
|
|
314
331
|
---
|
|
315
332
|
|
|
@@ -328,7 +345,7 @@ contains the ancestry from the observer down to the entity that raised it, plus
|
|
|
328
345
|
specific subject:
|
|
329
346
|
|
|
330
347
|
```ts
|
|
331
|
-
type ObserverEventBase = { observer: Observer };
|
|
348
|
+
type ObserverEventBase = { observer: Observer, context?: AcceptContext };
|
|
332
349
|
type ObservedCallScope = ObserverEventBase & { observedCall: ObservedCall };
|
|
333
350
|
type ObservedClientScope = ObservedCallScope & { observedClient: ObservedClient };
|
|
334
351
|
type ObservedPeerConnectionScope = ObservedClientScope & { observedPeerConnection: ObservedPeerConnection };
|
|
@@ -354,16 +371,18 @@ additional field(s) on top of that scope.
|
|
|
354
371
|
|
|
355
372
|
| Event | Extra payload | Fires when |
|
|
356
373
|
|-------|---------------|-----------|
|
|
357
|
-
| `observer-updated` | — | `observer.update()` ran (
|
|
374
|
+
| `observer-updated` | — | `observer.update()` ran (see [When things update](#when-things-update)) |
|
|
358
375
|
| `observer-closed` | — | `observer.close()` |
|
|
359
376
|
| `sample-rejected` | `{ reason: 'observer-closed' \| 'missing-callId' \| 'missing-clientId', sample: ClientSample }` | a sample was dropped by `accept()` |
|
|
377
|
+
| `observer-issue` | `{ issue: ObserverIssue }` | `observer.addIssue(...)` — a cross-call / SFU-wide finding (see [observer-level detectors](#observer-level-detectors-cross-call--sfu-wide)) |
|
|
378
|
+
| `validation-ready` | `{ validator: string, report: ValidationReport }` | a [validator](#validators--one-shot-structural-checks) settled — fires once per check, not per tick |
|
|
360
379
|
|
|
361
380
|
#### Mediasoup level — scope `{ observer, observedMediasoupRouter }`
|
|
362
381
|
|
|
363
382
|
| Event | Extra | Fires when |
|
|
364
383
|
|-------|-------|-----------|
|
|
365
384
|
| `mediasoup-router-added` | — | `observer.createObservedMediasoupRouter(...)` registered a router |
|
|
366
|
-
| `mediasoup-router-matched-with-
|
|
385
|
+
| `mediasoup-router-matched-with-peer-connection` | `{ observedCall, observedClient, observedPeerConnection }` | a newly added peer connection's id matched one of the router's WebRTC transport ids. **Opt-in** via `matchPeerConnectionByWebRtcTransportId: true`. |
|
|
367
386
|
| `mediasoup-router-removed` | — | the underlying mediasoup router closed (its `router.observer` `close` fired) |
|
|
368
387
|
|
|
369
388
|
See [Mediasoup router observation](#mediasoup-router-observation) for the full design and examples.
|
|
@@ -377,7 +396,8 @@ See [Mediasoup router observation](#mediasoup-router-observation) for the full d
|
|
|
377
396
|
| `call-closed` | — | the call closed |
|
|
378
397
|
| `call-empty` | — | last client left the call |
|
|
379
398
|
| `call-not-empty` | — | first client joined a previously-empty call |
|
|
380
|
-
| `call-issue` | `{ issue:
|
|
399
|
+
| `call-issue` | `{ issue: CallIssue }` | `call.addIssue(...)` (server-side detector finding) |
|
|
400
|
+
| `call-summary` | `{ summary: CallSummary }` | the call is closing and a [summary](#call-summaries) was configured. Emitted **inside** `close()`, while the call is still reachable |
|
|
381
401
|
|
|
382
402
|
#### Client level — scope `{ observer, observedCall, observedClient }`
|
|
383
403
|
|
|
@@ -390,7 +410,8 @@ See [Mediasoup router observation](#mediasoup-router-observation) for the full d
|
|
|
390
410
|
| `client-joined` | — | first `CLIENT_JOINED` event seen |
|
|
391
411
|
| `client-left` | — | `CLIENT_LEFT` seen (or inferred on close) |
|
|
392
412
|
| `client-rejoined` | `{ timestamp: number }` | a later `CLIENT_JOINED` after an earlier join |
|
|
393
|
-
| `client-issue` | `{ issue: ClientIssue }` | a client-reported issue arrived, or `client.addIssue(...)` |
|
|
413
|
+
| `client-issue` | `{ issue: ClientIssue }` | a client-reported issue arrived, or `client.addIssue(...)`. A keyed issue also opens an entry in `observedClient.activeIssues` |
|
|
414
|
+
| `client-issue-resolved` | `{ resolvedIssue: ResolvedActiveClientIssue }` | a stateful issue ended — the client sent its `<type>-resolved` companion, or the observer force-closed it. Carries the finished interval (`durationInMs`, `resolvedBy`) — see [client issues](#client-issues-the-lifecycle-and-the-division-of-labour) |
|
|
394
415
|
| `client-metadata` | `{ metaData: ClientMetaData }` | a client meta item arrived |
|
|
395
416
|
| `client-extension-stats` | `{ extensionStats: ExtensionStat }` | an app-defined extension stat arrived |
|
|
396
417
|
| `client-event` | `{ event: ClientEvent }` | any client event was processed |
|
|
@@ -402,7 +423,6 @@ See [Mediasoup router observation](#mediasoup-router-observation) for the full d
|
|
|
402
423
|
| `peer-connection-added` / `peer-connection-closed` | — | lifecycle of the PC |
|
|
403
424
|
| `peer-connection-updated` | `{ context?: AcceptContext }` | the PC processed a sample |
|
|
404
425
|
| `ice-connection-state-changed` / `ice-gathering-state-changed` / `connection-state-changed` | `{ state: string }` | driven by client events |
|
|
405
|
-
| `selected-candidate-pair-changed` | — | *declared; not currently emitted* |
|
|
406
426
|
| `inbound-track-added` / `-updated` / `-removed` / `-muted` / `-unmuted` | `{ observedInboundTrack }` | |
|
|
407
427
|
| `outbound-track-added` / `-updated` / `-removed` / `-muted` / `-unmuted` | `{ observedOutboundTrack }` | |
|
|
408
428
|
| `inbound-rtp-added` / `-updated` / `-removed` | `{ observedInboundRtp }` | `-updated` fires every tick |
|
|
@@ -444,45 +464,76 @@ listen to them, but prefer the bus equivalents above for application logic.
|
|
|
444
464
|
new Observer<AppData>(config?: ObserverConfig<AppData>)
|
|
445
465
|
|
|
446
466
|
type ObserverConfig<AppData = Record<string, unknown>> = {
|
|
447
|
-
|
|
448
|
-
|
|
467
|
+
// a call updates when any client does; the observer when any call does. Default true.
|
|
468
|
+
autoUpdateOnCallUpdate?: boolean;
|
|
449
469
|
appData?: AppData;
|
|
450
470
|
closeClientIfIdleForMs?: number;
|
|
451
471
|
closeCallIfEmptyForMs?: number;
|
|
472
|
+
// accumulate a per-call summary (see Call summaries). Absent or null = off, and nothing
|
|
473
|
+
// subscribes to anything. `{}` is valid: a summary with no built-in sections.
|
|
474
|
+
callSummary?: Partial<CallSummaryConfig> | null;
|
|
452
475
|
// appData factories — run when an entity is created without explicit appData
|
|
453
476
|
// (incl. lazily by accept()). appData is application-owned; accept `context` never touches it.
|
|
454
|
-
createCallAppData?: (p: { callId: string; observer: Observer }) => Record<string, unknown>;
|
|
455
|
-
createClientAppData?: (p: { clientId: string; observedCall: ObservedCall }) => Record<string, unknown>;
|
|
477
|
+
createCallAppData?: (p: { callId: string; observer: Observer; acceptCtx?: AcceptContext }) => Record<string, unknown>;
|
|
478
|
+
createClientAppData?: (p: { clientId: string; observedCall: ObservedCall; acceptCtx?: AcceptContext }) => Record<string, unknown>;
|
|
456
479
|
// sink factory — produces a per-client sink that receives every accepted sample (see Sinks).
|
|
457
480
|
createClientSink?: (p: { clientId: string; observedCall: ObservedCall }) => ClientSampleSink | undefined;
|
|
458
|
-
// track-resolver factory — produces a call's RemoteTrackResolver (see Remote track resolution).
|
|
459
|
-
|
|
481
|
+
// remote-track-resolver factory — produces a call's RemoteTrackResolver (see Remote track resolution).
|
|
482
|
+
createRemoteTrackResolver?: (observedCall: ObservedCall) => RemoteTrackResolver | undefined;
|
|
460
483
|
};
|
|
461
484
|
```
|
|
462
485
|
|
|
463
486
|
**appData factories.** Instead of pre-creating a call/client (or assigning on `call-added` /
|
|
464
|
-
`client-added`) just to enrich its `appData`, register a factory once. It runs
|
|
465
|
-
|
|
466
|
-
|
|
467
|
-
|
|
468
|
-
|
|
487
|
+
`client-added`) just to enrich its `appData`, register a factory once. It runs whenever the entity
|
|
488
|
+
is created without an explicit `settings.appData` — including the lazy creation inside `accept()`.
|
|
489
|
+
The `client` factory receives the already-created parent `observedCall`, so it can derive fields
|
|
490
|
+
from it.
|
|
491
|
+
|
|
492
|
+
Both also receive **`acceptCtx`**: the [`AcceptContext`](#context-the-acceptcontext) of the
|
|
493
|
+
`accept()` that caused the creation, or `undefined` when you created the entity yourself. This is
|
|
494
|
+
what lets an [accept middleware](#accept-middlewares-global-pre-dispatch-hook) resolve something
|
|
495
|
+
once — a tenant, a trace id — and have it land in `appData` at birth, instead of every factory
|
|
496
|
+
re-deriving it from the sample.
|
|
469
497
|
|
|
470
498
|
```ts
|
|
471
499
|
const observer = new Observer({
|
|
472
|
-
createCallAppData: ({ callId })
|
|
473
|
-
createClientAppData: ({ clientId, observedCall }) => ({ clientId,
|
|
500
|
+
createCallAppData: ({ callId, acceptCtx }) => ({ callId, startedAt: Date.now(), tenant: acceptCtx?.tenant }),
|
|
501
|
+
createClientAppData: ({ clientId, observedCall }) => ({ clientId, tenant: observedCall.appData.tenant }),
|
|
474
502
|
});
|
|
503
|
+
|
|
504
|
+
observer.accept(sample, { tenant: 'acme' });
|
|
475
505
|
```
|
|
476
506
|
|
|
507
|
+
`appData` stays application-owned: the context is *offered* to the factory, never written across by
|
|
508
|
+
the library, and it is still not stored on any entity.
|
|
509
|
+
|
|
477
510
|
Key members:
|
|
478
511
|
|
|
479
512
|
- `accept(sample: ClientSample, context?: AcceptContext): void`
|
|
480
513
|
- `addAcceptMiddleware(...mw: AcceptMiddleware[]): this` / `removeAcceptMiddleware(...mw): this` — global pre-dispatch sample hooks (see [Accept middlewares](#accept-middlewares-global-pre-dispatch-hook))
|
|
481
514
|
- `getObservedCall<T>(callId): ObservedCall<T> | undefined`
|
|
482
|
-
- `createObservedCall<T>(settings): ObservedCall<T> | undefined`
|
|
483
|
-
- `getOrCreateObservedCall<T>(settings): ObservedCall<T> | undefined`
|
|
515
|
+
- `createObservedCall<T>(settings, acceptCtx?): ObservedCall<T> | undefined`
|
|
516
|
+
- `getOrCreateObservedCall<T>(settings, acceptCtx?): ObservedCall<T> | undefined`
|
|
517
|
+
- `addIssue(issue: Omit<ObserverIssue, 'scope'>): void` — raise an **observer-level** finding → emits `observer-issue`. `scope` is stamped for you
|
|
484
518
|
- `update(): void` — force an aggregation/`observer-updated` tick
|
|
519
|
+
- `addObserverDetector(name, config?): this` — build a cross-call detector onto `observer.detectors`
|
|
520
|
+
- `addCallDetector(name, config?): this` — register a call-scoped detector for every call created
|
|
521
|
+
from now on
|
|
522
|
+
- `removeCallDetector(name, { includeOpenCalls? }): number` — stop building it, and (by default) drop
|
|
523
|
+
it from calls already open. Returns how many live instances were removed
|
|
524
|
+
- `removeObserverDetector(name): number` — remove an observer-scoped detector. For one specific
|
|
525
|
+
instance use `observer.detectors.remove(detector)`
|
|
526
|
+
- `addValidator(name, config?): this` — start a one-shot structural check
|
|
527
|
+
- `cancelValidator(name | validator, reason?): number` — stop a running check; it finishes
|
|
528
|
+
`inconclusive` with the reason and emits `validation-ready`
|
|
485
529
|
- `close(): void`
|
|
530
|
+
- `readonly detectors: Detectors` — observer-scoped registry. **Starts empty**; nothing is implicit
|
|
531
|
+
- `readonly callDetectorConfigs: Map<name, config>` — what `addCallDetector` recorded
|
|
532
|
+
- `readonly callSummaryCollector?: CallSummaryCollector` — owns the resolved `config.callSummary`,
|
|
533
|
+
the summary subscriptions, and the summaries. `undefined` when summaries are off, which is the
|
|
534
|
+
only place that answer lives
|
|
535
|
+
- `readonly validators: Set<RunningValidator>` — normally empty; each removes itself on finishing
|
|
536
|
+
- `readonly activeIssuesRegistry: ActiveIssuesRegistry` — the fleet's open client issues
|
|
486
537
|
- `readonly observedCalls: Map<string, ObservedCall>`
|
|
487
538
|
- `readonly observedTURN: ObservedTURN`
|
|
488
539
|
- `get appData()`, `get numberOfCalls()`
|
|
@@ -495,7 +546,8 @@ Key members:
|
|
|
495
546
|
|
|
496
547
|
```ts
|
|
497
548
|
type ObservedCallSettings<AppData = Record<string, unknown>> = {
|
|
498
|
-
|
|
549
|
+
// update this call whenever one of its clients accepts a sample. Default true.
|
|
550
|
+
autoUpdateOnClientUpdate?: boolean;
|
|
499
551
|
callId: string;
|
|
500
552
|
appData?: AppData;
|
|
501
553
|
closeCallIfEmptyForMs?: number;
|
|
@@ -506,14 +558,20 @@ Key members:
|
|
|
506
558
|
|
|
507
559
|
- `readonly callId: string`, `appData: AppData`
|
|
508
560
|
- `readonly observedClients: Map<string, ObservedClient>`, `get numberOfClients()`
|
|
509
|
-
- `getObservedClient<T>(clientId)`, `createObservedClient<T>(settings)`, `getOrCreateObservedClient<T>(settings)` (all `… | undefined`)
|
|
510
|
-
- `addIssue(issue:
|
|
561
|
+
- `getObservedClient<T>(clientId)`, `createObservedClient<T>(settings, acceptCtx?)`, `getOrCreateObservedClient<T>(settings, acceptCtx?)` (all `… | undefined`)
|
|
562
|
+
- `addIssue(issue: Omit<CallIssue, 'scope'>): void` — raise a **call-level** finding → emits `call-issue`. `scope` is stamped for you
|
|
563
|
+
- `addDetector(name, config?): this` — build a call-scoped detector onto this call only
|
|
564
|
+
- `removeDetector(name): number` — remove it from this call, `close()`ing it. For one specific
|
|
565
|
+
instance use `call.detectors.remove(detector)`
|
|
511
566
|
- `readonly detectors: Detectors` — server-side detector registry (empty by default; see [Detectors](#detectors-server-side-extension-point))
|
|
567
|
+
- `readonly activeIssuesRegistry: ActiveIssuesRegistry` — this call's open client issues, propagating into the observer's
|
|
568
|
+
- `readonly unconsumedOutboundTracks: Set<ObservedOutboundTrack>` — maintained by the resolver
|
|
512
569
|
- `scoreCalculator: ScoreCalculator`, `get score()`, `readonly calculatedScore`
|
|
513
|
-
- `remoteTrackResolver?: RemoteTrackResolver` — set from `ObserverConfig.
|
|
570
|
+
- `remoteTrackResolver?: RemoteTrackResolver` — set from `ObserverConfig.createRemoteTrackResolver` at call creation (see [Remote track resolution](#remote-track-resolution-mediasoup--sfu))
|
|
514
571
|
- aggregates: `numberOfIssues`, `numberOfPeerConnections`, `numberOfInboundRtpStreams`,
|
|
515
572
|
`numberOfOutboundRtpStreams`, `numberOfDataChannels`, `maxNumberOfClients`,
|
|
516
573
|
`clientsUsedTurn: Set<string>`, `startedAt?`, `endedAt?`, `closedAt?`, `closed`
|
|
574
|
+
- `summary?: CallSummary` — the live record of this call, when [summaries](#call-summaries) are on
|
|
517
575
|
- `update()`, `close()`
|
|
518
576
|
|
|
519
577
|
### `ObservedClient`
|
|
@@ -533,7 +591,7 @@ Key members:
|
|
|
533
591
|
- `readonly sink?: ClientSampleSink` — the per-client sink (see [Sinks](#sinks-per-client-sample-persistence)), if `createClientSink` is configured; listen on it for `close`/`error`
|
|
534
592
|
- **Injection API** (queue app data to be merged into the next sample processing):
|
|
535
593
|
`injectEvent(ClientEvent)`, `injectIssue(ClientIssue)`, `injectMetaData(ClientMetaData)`,
|
|
536
|
-
`injectExtensionStat(ExtensionStat)`, `injectAttachment(
|
|
594
|
+
`injectExtensionStat(ExtensionStat)`, `injectAttachment(attachments: Record<string, unknown>)`
|
|
537
595
|
- **Direct add API** (process immediately): `addIssue(ClientIssue)`, `addMetadata(ClientMetaData)`,
|
|
538
596
|
`addExtensionStats(ExtensionStat)`
|
|
539
597
|
- Metrics (current/derived): `currentAvgRttInMs?`, `currentMinRttInMs?`, `currentMaxRttInMs?`,
|
|
@@ -557,11 +615,28 @@ Key members:
|
|
|
557
615
|
`iceCandidates`, `iceCandidatePairs`, `certificates`, `selectedIceCandidatePairs`,
|
|
558
616
|
`selectedIceCandiadtePairForTurn`
|
|
559
617
|
- State: `connectionState?`, `iceConnectionState?`, `iceGatheringState?`, `usingTURN`, `usingTCP`
|
|
560
|
-
- Metrics: `currentRttInMs?`, `
|
|
561
|
-
`availableOutgoingBitrate`, sending/receiving bitrates, packet rates,
|
|
562
|
-
byte/packet counters
|
|
618
|
+
- Metrics: `currentRttInMs?`, `iceRttInMs?`, `rtcpRttInMs?`, `sfuHopRttInMs?`, `currentJitter?`,
|
|
619
|
+
`availableIncomingBitrate`, `availableOutgoingBitrate`, sending/receiving bitrates, packet rates,
|
|
620
|
+
and `total*` / `delta*` byte/packet counters
|
|
563
621
|
- `accept(pcSample, context?)`, `close()`, `get score()`
|
|
564
622
|
|
|
623
|
+
**Two different round trips — don't mix them.** `iceRttInMs` comes from ICE/STUN consent checks and
|
|
624
|
+
measures the trip to *whatever terminates ICE*: **in an SFU topology that is the SFU**, so it is the
|
|
625
|
+
client↔SFU leg. `rtcpRttInMs` comes from RTCP receiver reports and is an **end-to-end** media-path
|
|
626
|
+
round trip. They are not interchangeable, and averaging them together produces a number that moves
|
|
627
|
+
as streams come and go for reasons unrelated to the network. `currentRttInMs` therefore *prefers*
|
|
628
|
+
RTCP and falls back to ICE — always one kind within a tick, never a blend. `sfuHopRttInMs`
|
|
629
|
+
(`rtcp − ice`) estimates everything past the SFU, which separates "this client's last mile is slow"
|
|
630
|
+
from "the path beyond the SFU is slow".
|
|
631
|
+
|
|
632
|
+
**Counter-reset boundaries.** Chrome resets an SSRC's cumulative counters when the codec switches
|
|
633
|
+
([crbug/webrtc/5361](https://bugs.chromium.org/p/webrtc/issues/detail?id=5361), open since 2015),
|
|
634
|
+
which otherwise shows up as a sawtooth spike or a negative bitrate. `ObservedInboundRtp` /
|
|
635
|
+
`ObservedOutboundRtp` therefore set **`counterResetBoundary`** on any tick where `codecId`,
|
|
636
|
+
`encoder`/`decoderImplementation` or `scalabilityMode` changed, and suppress every delta for that
|
|
637
|
+
tick. Without this, a room-wide codec rollout fires a synchronized fake-degradation alert across
|
|
638
|
+
every participant at once.
|
|
639
|
+
|
|
565
640
|
**Remote-RTP correlation (derived).** During `accept()`, receiver/sender reports are linked
|
|
566
641
|
to the local streams by `remoteId` (fallback SSRC) and surfaced as fields:
|
|
567
642
|
|
|
@@ -606,13 +681,17 @@ type PeerConnectionSample = {
|
|
|
606
681
|
certificates?;
|
|
607
682
|
};
|
|
608
683
|
|
|
609
|
-
type ClientEvent = { type: string; payload?: string
|
|
610
|
-
type ClientIssue = { type: string; payload?: string; timestamp?: number };
|
|
611
|
-
type ClientMetaData = { type: string; payload?: string
|
|
612
|
-
type ExtensionStat = { type: string; payload?: string };
|
|
684
|
+
type ClientEvent = { type: string; payload?: Record<string, unknown>; timestamp?: number; /* +ids */ };
|
|
685
|
+
type ClientIssue = { type: string; payload?: Record<string, unknown>; key?: string; timestamp?: number };
|
|
686
|
+
type ClientMetaData = { type: string; payload?: Record<string, unknown>; timestamp?: number; /* +ids */ };
|
|
687
|
+
type ExtensionStat = { type: string; payload?: Record<string, unknown> };
|
|
613
688
|
```
|
|
614
689
|
|
|
615
|
-
`payload` fields are JSON
|
|
690
|
+
`payload` fields are objects (schema **3.7.0**: free-form JSON, so a payload may nest objects and
|
|
691
|
+
arrays), and the library reads the ones it understands. **Older clients keep working**: a payload
|
|
692
|
+
from a pre-3.5.0 client arrives as a JSON string and is parsed on the way in, so a fleet running a
|
|
693
|
+
mix of client versions needs no coordination — every reader sees the same object either way. Values
|
|
694
|
+
inside a payload are `unknown`, so narrow before use (`typeof p.trackId === 'string'`).
|
|
616
695
|
|
|
617
696
|
**`ClientEventTypes`** (enum of known `event.type` values): `CLIENT_JOINED`, `CLIENT_LEFT`,
|
|
618
697
|
`PEER_CONNECTION_OPENED/CLOSED/STATE_CHANGED`, `MEDIA_TRACK_ADDED/REMOVED/MUTED/UNMUTED/RESUMED`,
|
|
@@ -654,8 +733,8 @@ then lean **steady-state ticks**.
|
|
|
654
733
|
],
|
|
655
734
|
|
|
656
735
|
"clientMetaItems": [ // environment & devices, one-off (10 in the real sample)
|
|
657
|
-
{ "type": "USER_AGENT_DATA", "payload": "
|
|
658
|
-
{ "type": "MEDIA_DEVICE", "payload":
|
|
736
|
+
{ "type": "USER_AGENT_DATA", "payload": { "…": "Chrome 148 / macOS" } },
|
|
737
|
+
{ "type": "MEDIA_DEVICE", "payload": { "label": "BRIO 4K Stream Edition" } }
|
|
659
738
|
// …mic / camera / speaker devices…
|
|
660
739
|
],
|
|
661
740
|
|
|
@@ -696,12 +775,29 @@ snapshot happens only once.
|
|
|
696
775
|
|
|
697
776
|
## Detectors (server-side extension point)
|
|
698
777
|
|
|
699
|
-
`observer-js`
|
|
700
|
-
|
|
701
|
-
|
|
702
|
-
|
|
778
|
+
`observer-js` ships [ten detectors](#built-in-detectors), each an opt-in extension you register
|
|
779
|
+
explicitly with `addObserverDetector` / `addCallDetector` / `addDetector` (see
|
|
780
|
+
[Registering detectors](#registering-detectors)) — none are created automatically. All of them
|
|
781
|
+
correlate **across** the clients of a call or the calls of a fleet, because that is the only thing a
|
|
782
|
+
server can do better than a browser: per-client signals — packet loss, jitter, RTT, freezes — are
|
|
783
|
+
already detected on the client and arrive on samples as `clientIssues` (surfaced via `client-issue`).
|
|
784
|
+
|
|
785
|
+
Findings are raised as **`CallIssue`** or **`ObserverIssue`** — both share `IssueBase`: `{ type,
|
|
786
|
+
timestamp, conclusion?, payload? }` — and the payload is the **object**, not a JSON string. A
|
|
787
|
+
server-raised finding is delivered to an in-process handler, so there is nothing to serialise for:
|
|
788
|
+
|
|
789
|
+
```ts
|
|
790
|
+
observer.on('call-issue', ({ observedCall, issue }) => {
|
|
791
|
+
issue.payload; // the object; no JSON.parse
|
|
792
|
+
issue.conclusion?.faultDomain; // a first-class field, not payload.conclusion
|
|
793
|
+
issuePayloadAsString(issue); // only at an edge that needs text (log, HTTP, queue)
|
|
794
|
+
});
|
|
795
|
+
```
|
|
796
|
+
|
|
797
|
+
(`ClientIssue`, the type on samples, is a wire format — its payload may arrive as a JSON string from
|
|
798
|
+
a pre-3.5.0 client, which is why the library parses it before you see it.)
|
|
703
799
|
|
|
704
|
-
The
|
|
800
|
+
The registry is also an open extension point, on **`ObservedCall`**:
|
|
705
801
|
|
|
706
802
|
```ts
|
|
707
803
|
import { Observer, Detector } from '@observertc/observer-js';
|
|
@@ -712,7 +808,7 @@ class MyCrossClientDetector implements Detector {
|
|
|
712
808
|
update() { // called on every call.update()
|
|
713
809
|
// …inspect this.call.observedClients across participants…
|
|
714
810
|
if (/* condition only visible server-side */ false) {
|
|
715
|
-
this.call.addIssue({ type: this.name, payload:
|
|
811
|
+
this.call.addIssue({ type: this.name, payload: { /* … */ }, timestamp: Date.now() });
|
|
716
812
|
// → emitted on the bus as 'call-issue'
|
|
717
813
|
}
|
|
718
814
|
}
|
|
@@ -725,27 +821,719 @@ observer.on('call-added', ({ observedCall }) => {
|
|
|
725
821
|
observer.on('call-issue', ({ observedCall, issue }) => { /* react */ });
|
|
726
822
|
```
|
|
727
823
|
|
|
728
|
-
|
|
824
|
+
### Client issues: the lifecycle, and the division of labour
|
|
825
|
+
|
|
826
|
+
The most important thing to understand about detection in this library is **what it deliberately
|
|
827
|
+
does not do**. A client running
|
|
828
|
+
[`client-monitor-js`](https://github.com/ObserveRTC/client-monitor-js) already ships ~20 detectors
|
|
829
|
+
that decide *what is wrong with that endpoint* — `congestion`, `cpulimitation`, `audio-concealment`,
|
|
830
|
+
`freezed-video-track`, `keyframe-storm`, `video-decoder-overloaded`, `stuck-decoder`,
|
|
831
|
+
`ice-disconnected`, and so on. Those verdicts are better than anything re-derived from raw counters
|
|
832
|
+
server-side, because they carry hysteresis and multi-signal confirmation: `audio-concealment`
|
|
833
|
+
subtracts silent concealment (raw `concealedSamples` rises during ordinary silence, so a naive
|
|
834
|
+
detector flags every quiet moment); `audio-jitter-buffer-stress` requires the buffer to be grown
|
|
835
|
+
**and** NetEQ to be time-stretching (a grown buffer alone means NetEQ is *succeeding*);
|
|
836
|
+
`ice-disconnected` only fires once `disconnected` has persisted, so the blips ICE heals on its own
|
|
837
|
+
never surface.
|
|
838
|
+
|
|
839
|
+
**observer-js does not repeat that work.** Its job is the question no browser can answer: *who else
|
|
840
|
+
is in this state right now, what do they have in common, and where in publisher → SFU → subscriber
|
|
841
|
+
does the fault begin?*
|
|
842
|
+
|
|
843
|
+
#### The wire format
|
|
844
|
+
|
|
845
|
+
From client-monitor-js **4.6.0** the whole issue lifecycle reaches the server. A stateful issue
|
|
846
|
+
arrives as two `clientIssues[]` entries sharing a `key`:
|
|
847
|
+
|
|
848
|
+
```
|
|
849
|
+
raise: { type: 'stuck-decoder', key, payload, timestamp: raisedAt }
|
|
850
|
+
resolution: { type: 'stuck-decoder-resolved', key, payload: { raisedAt, comment, …final }, timestamp: resolvedAt }
|
|
851
|
+
```
|
|
852
|
+
|
|
853
|
+
The observer opens an entry in `observedClient.activeIssues` on the raise and closes it on the
|
|
854
|
+
matching key, emitting **`client-issue-resolved`** with the finished interval. Handled for you:
|
|
855
|
+
|
|
856
|
+
- the `-resolved` **suffix** is stripped, so both entries share one logical `type`;
|
|
857
|
+
- a **re-raise** of a live key refreshes the payload without restarting `raisedAt`;
|
|
858
|
+
- **keyless** entries are one-shot — reported via `client-issue`, never tracked;
|
|
859
|
+
- issues still open when a client closes are **force-resolved** (`resolvedBy: 'client-closed'`), and
|
|
860
|
+
the registry additionally expires stale entries, so a crashed participant can't leave an issue
|
|
861
|
+
"active" forever.
|
|
862
|
+
|
|
863
|
+
#### Why intervals beat windows
|
|
864
|
+
|
|
865
|
+
This turns point-in-time symptom reports into **intervals**, and that is the whole game. *"Several
|
|
866
|
+
clients reported congestion in the last 10 seconds"* is a heuristic that has to guess whether the
|
|
867
|
+
symptoms are still happening. *"Several clients are congested **right now, simultaneously**"* is
|
|
868
|
+
ground truth, because the client says when the episode ends. Overlapping intervals are far stronger
|
|
869
|
+
evidence of a shared cause than near-in-time reports.
|
|
729
870
|
|
|
730
871
|
```ts
|
|
731
|
-
|
|
732
|
-
|
|
733
|
-
|
|
734
|
-
|
|
735
|
-
|
|
736
|
-
|
|
737
|
-
|
|
738
|
-
|
|
872
|
+
observer.on('client-issue', ({ observedClient, issue }) => { /* opened (or one-shot) */ });
|
|
873
|
+
observer.on('client-issue-resolved', ({ resolvedIssue }) => {
|
|
874
|
+
resolvedIssue.type; // 'stuck-decoder' — suffix stripped
|
|
875
|
+
resolvedIssue.durationInMs; // how long the episode lasted
|
|
876
|
+
resolvedIssue.resolvedBy; // 'client' | 'timeout' | 'client-closed'
|
|
877
|
+
});
|
|
878
|
+
|
|
879
|
+
// the live per-client mirror
|
|
880
|
+
observedClient.activeIssues; // ObservedClientIssueRegistry, keyed by issue.key
|
|
881
|
+
```
|
|
882
|
+
|
|
883
|
+
> **client-monitor-js >= 4.6.0 is required** for every issue-driven detector. There is no fallback
|
|
884
|
+
> path that infers these conditions from raw counters — the client decides better, and maintaining a
|
|
885
|
+
> worse second implementation to be polite to old clients is how both end up wrong. Issues without a
|
|
886
|
+
> `key` have no lifecycle (nothing could ever close them), so they stay one-shot: reported on
|
|
887
|
+
> `client-issue`, never registered.
|
|
888
|
+
|
|
889
|
+
#### `ActiveIssuesRegistry` — issues are **pushed**, not polled
|
|
890
|
+
|
|
891
|
+
A detector does not go looking for the issues it cares about. It implements `ActiveIssueTracker` and
|
|
892
|
+
registers for the types it consumes; the registry hands them over as they open and close.
|
|
893
|
+
|
|
894
|
+
```ts
|
|
895
|
+
observedCall.activeIssuesRegistry // this meeting
|
|
896
|
+
observer.activeIssuesRegistry // the fleet; every call's registry propagates into it
|
|
897
|
+
|
|
898
|
+
observer.activeIssuesRegistry.addIssueTracker('congestion', myDetector);
|
|
899
|
+
observer.activeIssuesRegistry.removeIssueTracker(myDetector);
|
|
900
|
+
|
|
901
|
+
registry.values(); // the open issues in this scope, oldest first
|
|
902
|
+
registry.size; // how many
|
|
903
|
+
```
|
|
904
|
+
|
|
905
|
+
The cost of a detector is then proportional to the issues it actually receives, not to the number of
|
|
906
|
+
participants: a healthy 500-client fleet does no per-tick work at all, because nothing was pushed.
|
|
907
|
+
|
|
908
|
+
**There is no wildcard.** A tracker names its types and sees nothing else. "Feed me everything and
|
|
909
|
+
I'll work out what matters" moves the decision from the application — which knows its client build
|
|
910
|
+
and its issue vocabulary — onto a detector that has to guess, and it makes the cost of a
|
|
911
|
+
subscription unbounded and invisible. If a detector should watch five types, list five types.
|
|
912
|
+
|
|
913
|
+
Onset spread is measured on the **observer clock**, never the client's. `raisedAt` comes from each
|
|
914
|
+
participant's own machine, and comparing those across clients makes clock skew look like a
|
|
915
|
+
synchronized infrastructure event.
|
|
916
|
+
|
|
917
|
+
### Observer-level detectors (cross-call / SFU-wide)
|
|
918
|
+
|
|
919
|
+
Some findings only exist **above** call scope — "many calls on the same SFU degraded at once" is far
|
|
920
|
+
more actionable than fifty individual client alerts. The same detector registry exists on the `Observer`,
|
|
921
|
+
runs on every `observer.update()`, and raises findings through `observer.addIssue(...)`, surfaced on
|
|
922
|
+
the bus as **`observer-issue`**:
|
|
923
|
+
|
|
924
|
+
```ts
|
|
925
|
+
observer.detectors.add({
|
|
926
|
+
name: 'sfu-wide-degradation',
|
|
927
|
+
update: () => {
|
|
928
|
+
const degradedCalls = [ ...observer.observedCalls.values() ].filter(isDegraded);
|
|
929
|
+
|
|
930
|
+
if (observer.numberOfCalls > 3 && degradedCalls.length / observer.numberOfCalls > 0.6) {
|
|
931
|
+
observer.addIssue({ type: 'SFU_WIDE_QUALITY_DEGRADATION', timestamp: Date.now() });
|
|
932
|
+
}
|
|
933
|
+
},
|
|
934
|
+
});
|
|
935
|
+
|
|
936
|
+
observer.on('observer-issue', ({ issue }) => alert(issue));
|
|
937
|
+
```
|
|
938
|
+
|
|
939
|
+
### Publisher → subscribers: the resolver links
|
|
940
|
+
|
|
941
|
+
The question a single browser can never answer is *"did **everyone** receiving Alice see the same
|
|
942
|
+
degradation?"*. The join is the publisher↔subscriber links maintained by a
|
|
943
|
+
[`RemoteTrackResolver`](#remote-track-resolution-mediasoup--sfu), and detectors walk them directly:
|
|
944
|
+
|
|
945
|
+
```ts
|
|
946
|
+
outboundTrack.remoteInboundTracks; // Set<ObservedInboundTrack> — every subscriber of this source
|
|
947
|
+
inboundTrack.remoteOutboundTrack; // the publisher, or undefined if unlinked
|
|
948
|
+
inboundTrack.getInboundRtp(); // that receiver's RTP stats
|
|
949
|
+
observedCall.unconsumedOutboundTracks; // published tracks with no subscriber at all
|
|
950
|
+
```
|
|
951
|
+
|
|
952
|
+
> A `TrackDistributionAggregator` class used to sit in front of these links and summarise every
|
|
953
|
+
> published track against all of its receivers, on every tick. It is gone. It scanned the majority
|
|
954
|
+
> (all published tracks) to find the interesting minority, which is the wrong axis — the detectors
|
|
955
|
+
> now start from the handful of *affected* tracks the issue registry pushed at them and resolve only
|
|
956
|
+
> those. The statistics helpers it used (`percentile`, `median`, `summarize`, `counterDelta`,
|
|
957
|
+
> `robustZScore`, `SlidingWindow`, `TrendTester`) are all still exported for building your own.
|
|
958
|
+
|
|
959
|
+
### Call health: `CallHealthAggregator`
|
|
960
|
+
|
|
961
|
+
The **client** axis. Where the resolver links answer "how was *this source* delivered?", this asks
|
|
962
|
+
"how is *each participant* doing, sending vs receiving?":
|
|
963
|
+
|
|
964
|
+
```ts
|
|
965
|
+
import { CallHealthAggregator } from '@observertc/observer-js';
|
|
966
|
+
|
|
967
|
+
const health = new CallHealthAggregator(observedCall).aggregate();
|
|
968
|
+
|
|
969
|
+
health.degradedRatio; // 0.82 — the number that distinguishes shared faults from individual ones
|
|
970
|
+
health.inboundDegradedRatio; // receiving side → egress/downstream suspicion
|
|
971
|
+
health.outboundDegradedRatio; // sending side → ingress suspicion
|
|
972
|
+
health.rttInMs?.median; // percentile rollups, never means
|
|
973
|
+
health.qualityLimitation; // { cpu, bandwidth, other } client counts
|
|
974
|
+
health.clients; // per-client entries with `reasons`, direction flags, TURN/TCP
|
|
975
|
+
```
|
|
976
|
+
|
|
977
|
+
### Registering detectors
|
|
978
|
+
|
|
979
|
+
**Nothing is created implicitly.** A `new Observer()` has zero detectors. There is no detector
|
|
980
|
+
configuration in `ObserverConfig` and no default set — an application says what it wants to watch, or
|
|
981
|
+
it watches nothing.
|
|
982
|
+
|
|
983
|
+
```ts
|
|
984
|
+
const observer = new Observer({
|
|
985
|
+
createRemoteTrackResolver: createDefaultMediasoupRemoteTrackResolverFactory(),
|
|
986
|
+
});
|
|
987
|
+
|
|
988
|
+
// observer-scoped (cross-call) — built immediately onto `observer.detectors`
|
|
989
|
+
observer.addObserverDetector('observer-concurrent-issue-detector', {
|
|
990
|
+
issueTypes: [ 'congestion', 'ice-disconnected', 'ice-connection-failed' ],
|
|
991
|
+
minAffectedCalls: 3,
|
|
992
|
+
});
|
|
993
|
+
observer.addObserverDetector('turn-server-outage-detector', { minClientsAtPeak: 10 });
|
|
994
|
+
|
|
995
|
+
// call-scoped — recorded in `observer.callDetectorConfigs`, applied to every call created AFTER this
|
|
996
|
+
observer.addCallDetector('call-concurrent-issue-detector', {
|
|
997
|
+
issueTypes: [ 'congestion', 'ice-disconnected' ],
|
|
998
|
+
});
|
|
999
|
+
// one specific call
|
|
1000
|
+
observedCall.addDetector('issue-fan-out-detector', { issueTypes: [ 'freezed-video-track' ] });
|
|
1001
|
+
```
|
|
1002
|
+
|
|
1003
|
+
Every `add*` is **chainable** — it returns the owning entity:
|
|
1004
|
+
|
|
1005
|
+
```ts
|
|
1006
|
+
observer
|
|
1007
|
+
.addObserverDetector('turn-server-health-detector')
|
|
1008
|
+
.addObserverDetector('turn-server-outage-detector', { minClientsAtPeak: 10 })
|
|
1009
|
+
.addValidator('remote-track-resolver');
|
|
1010
|
+
```
|
|
1011
|
+
|
|
1012
|
+
#### Removing them
|
|
1013
|
+
|
|
1014
|
+
By **name**, on the entity — which removes *every* instance under that name:
|
|
1015
|
+
|
|
1016
|
+
```ts
|
|
1017
|
+
observer.removeObserverDetector('turn-server-outage-detector'); // → 1
|
|
1018
|
+
observer.removeCallDetector('call-concurrent-issue-detector'); // stops it everywhere
|
|
1019
|
+
observedCall.removeDetector('issue-fan-out-detector'); // this call only
|
|
1020
|
+
```
|
|
1021
|
+
|
|
1022
|
+
By **instance**, through the registry — which is where instances live, since `add*` returns the
|
|
1023
|
+
entity rather than the detector:
|
|
1024
|
+
|
|
1025
|
+
```ts
|
|
1026
|
+
observer
|
|
1027
|
+
.addObserverDetector('client-population-issue-detector', { issueTypes: [ 'cpulimitation' ], groupBy: 'browser' })
|
|
1028
|
+
.addObserverDetector('client-population-issue-detector', { issueTypes: [ 'cpulimitation' ], groupBy: 'operationSystem' });
|
|
1029
|
+
|
|
1030
|
+
const [ byBrowser, byOs ] = observer.detectors.getAll('client-population-issue-detector');
|
|
1031
|
+
|
|
1032
|
+
observer.detectors.remove(byOs); // keeps the browser axis running
|
|
1033
|
+
```
|
|
1034
|
+
|
|
1035
|
+
`Detectors` is a small collection: `instances` (a copy, in registration order), `listOfNames`,
|
|
1036
|
+
`size`, `get(name)`, `getAll(name)`, `has(name)`, `add(detector)`, `remove(detector)`,
|
|
1037
|
+
`removeByName(name)`, `clear()`, and it is iterable — `for (const detector of call.detectors)`.
|
|
1038
|
+
`instances` being a copy is deliberate: removing while iterating the live array would skip entries,
|
|
1039
|
+
and "drop the ones that look like X" is the most natural thing to want to write.
|
|
1040
|
+
|
|
1041
|
+
Two things worth knowing:
|
|
1042
|
+
|
|
1043
|
+
- **By name removes every instance under it**, not the first. A name can legitimately be registered
|
|
1044
|
+
more than once — `ClientPopulationIssueDetector` is meant to be added once per `groupBy` axis — and
|
|
1045
|
+
"remove whichever is first in the array" is not something a caller can predict from a name. Go via
|
|
1046
|
+
`detectors.getAll(name)` + `detectors.remove(instance)` when you mean one of them.
|
|
1047
|
+
- **`removeCallDetector` affects calls already open, by default.** Otherwise whether a detector runs
|
|
1048
|
+
would depend on when a call happened to join, which is not a state anyone can reason about. Pass
|
|
1049
|
+
`{ includeOpenCalls: false }` to change only what future calls are built with.
|
|
1050
|
+
|
|
1051
|
+
Every removal path calls the detector's `close()`, so it unsubscribes from the issue registry and
|
|
1052
|
+
drops any timers or bus listeners. A detector removed without closing would keep being fed matching
|
|
1053
|
+
issues for the life of the call — invisible, unbounded, and it would still look healthy if you
|
|
1054
|
+
inspected it.
|
|
1055
|
+
|
|
1056
|
+
Detectors are named by their kebab-case `NAME`, and the name types the config — an unknown name or a
|
|
1057
|
+
key that belongs to a different detector will not compile. Each detector owns its defaults in its own
|
|
1058
|
+
constructor, beside the doc explaining what each threshold means; there is no central table to keep
|
|
1059
|
+
in sync.
|
|
1060
|
+
|
|
1061
|
+
> **Why no defaults?** A detector nobody asked for is a detector nobody will act on. It costs time on
|
|
1062
|
+
> every tick and raises findings into a handler that was not written to expect them. Earlier versions
|
|
1063
|
+
> auto-created everything from a three-state config slot; the result was applications receiving
|
|
1064
|
+
> finding types they had never heard of.
|
|
1065
|
+
|
|
1066
|
+
Every issue-driven detector takes an explicit, non-empty `issueTypes` (or
|
|
1067
|
+
`publisherIssueTypes`/`receiverIssueTypes`). There is no "watch everything" option — see
|
|
1068
|
+
[the registry](#activeissuesregistry--issues-are-pushed-not-polled).
|
|
1069
|
+
|
|
1070
|
+
> **🔗 marks detectors that require a
|
|
1071
|
+
> [`RemoteTrackResolver`](#remote-track-resolution-mediasoup--sfu).** They reason about a published
|
|
1072
|
+
> track and its subscribers, so without the publisher↔subscriber links they see nothing and stay
|
|
1073
|
+
> **silent forever** — which looks exactly like "no problems found". Configure
|
|
1074
|
+
> `ObserverConfig.createRemoteTrackResolver`, and start the
|
|
1075
|
+
> [`remote-track-resolver` validator](#validators--one-shot-structural-checks) to prove it is wired.
|
|
1076
|
+
|
|
1077
|
+
### Built-in detectors
|
|
1078
|
+
|
|
1079
|
+
They consume the verdicts `client-monitor-js` >= 4.6.0 already ships (raise + `<type>-resolved`) and
|
|
1080
|
+
add only the cross-participant conclusion. **None re-derives a per-endpoint verdict from raw
|
|
1081
|
+
counters** — that is the rule the whole design hangs on:
|
|
1082
|
+
|
|
1083
|
+
> **If a condition is detectable on the client, the client's issue is the source of truth.**
|
|
1084
|
+
|
|
1085
|
+
| Detector | 🔗 | Scope | Raises |
|
|
1086
|
+
|----------|:--:|-------|--------|
|
|
1087
|
+
| `CallConcurrentIssueDetector` | | call | `CONCURRENT_CLIENT_ISSUES`, `ISSUE_ONSET_BURST` |
|
|
1088
|
+
| `ObserverConcurrentIssueDetector` | | **observer** | `CROSS_CALL_CONCURRENT_ISSUES`, `CROSS_CALL_ISSUE_ONSET_BURST` |
|
|
1089
|
+
| `IssueFanOutDetector` | 🔗 | call | `PUBLISHED_TRACK_ISSUE_FAN_OUT`, `SINGLE_RECEIVER_ISSUE` |
|
|
1090
|
+
| `PublisherFaultCorroborationDetector` | 🔗 | call | `CORROBORATED_PUBLISHER_FAULT` |
|
|
1091
|
+
| `TrackDeliveryMismatchDetector` | 🔗 | call | `PUBLISHED_TRACK_NOT_DELIVERED`, `RECEIVER_TRACK_NOT_DELIVERED`, `PUBLISHER_TRACK_DRY` |
|
|
1092
|
+
| `UnconsumedTrackDetector` | 🔗 | call | `UNCONSUMED_PUBLISHED_TRACK` |
|
|
1093
|
+
| `ClientPopulationIssueDetector` | | **observer** | `CLIENT_POPULATION_ISSUE` |
|
|
1094
|
+
| `SfuCongestionDetector` | | **observer** | `sfu-congestion` |
|
|
1095
|
+
| `TurnServerHealthDetector` | | **observer** | `TURN_SERVER_DEGRADED` |
|
|
1096
|
+
| `TurnServerOutageDetector` | | **observer** | `TURN_SERVER_OUTAGE` |
|
|
1097
|
+
|
|
1098
|
+
What each adds that no endpoint can know:
|
|
1099
|
+
|
|
1100
|
+
- **`CallConcurrentIssueDetector`** — *who else in this meeting is in this state right now?* The
|
|
1101
|
+
difference between "one person's Wi-Fi" and "this room is broken".
|
|
1102
|
+
- **`ObserverConcurrentIssueDetector`** — *is our infrastructure in trouble?* A **separate class**,
|
|
1103
|
+
not the call one with a bigger denominator, because it is a different question with different
|
|
1104
|
+
gates. It requires the group to span at least `minAffectedCalls` **independent calls** (default
|
|
1105
|
+
`2`) and raises its own `CROSS_CALL_*` types. Without that gate, one thirty-person meeting where
|
|
1106
|
+
everyone is congested clears every client threshold and pages you for a single bad room the
|
|
1107
|
+
call-scoped detector already reported. Clients in different calls share no room, no publisher and
|
|
1108
|
+
no host — only the servers, which is what makes the finding conclusive. Note there is deliberately
|
|
1109
|
+
**no participant ratio** at this scope: six broken calls out of forty is a small share of all
|
|
1110
|
+
clients, and a ratio gate would hide exactly the event you want.
|
|
1111
|
+
- **`IssueFanOutDetector`** — *does this issue follow one published source, or one receiver?*
|
|
1112
|
+
- **`PublisherFaultCorroborationDetector`** — *do **both ends** of one track agree the source is at
|
|
1113
|
+
fault?* Fan-out sees one end and infers; this sees the publisher reporting `encoder-bottleneck`
|
|
1114
|
+
about its own send path *while* its subscribers report `freezed-video-track` about receiving it.
|
|
1115
|
+
Two independent parties, one conclusion, nothing left to deduce — hence the highest confidence in
|
|
1116
|
+
the library. Run both: fan-out is broader and catches the case where the publisher is fine and the
|
|
1117
|
+
SFU's forwarding is not.
|
|
1118
|
+
- **`ClientPopulationIssueDetector`** — *is this concentrated on one **kind of client**?* The one
|
|
1119
|
+
correlation here that is neither per-call nor per-server. Every other observer-scoped detector
|
|
1120
|
+
reasons "clients in unrelated calls share only the infrastructure, so it must be us" — right for
|
|
1121
|
+
network symptoms, **wrong for endpoint ones**. `cpulimitation` across six unrelated calls is not an
|
|
1122
|
+
SFU event; CPU is owned by the endpoint, so what those endpoints share is a browser version or a
|
|
1123
|
+
client release. Groups by `browser` / `engine` / `platform` / `operationSystem` / `location`, one
|
|
1124
|
+
axis per instance. The gate is **relative risk**, not share: "30% of Chrome 141 is unhappy" means
|
|
1125
|
+
nothing if 30% of everyone is, and a share-based rule simply indicts whichever browser is most
|
|
1126
|
+
popular. See [the `location` axis](#grouping-by-place-the-location-axis) for the geographic form.
|
|
1127
|
+
- **`SfuCongestionDetector`** — *is congestion spiking across the fleet right now?* Counts distinct
|
|
1128
|
+
clients reporting congestion in fixed wall-clock buckets and compares each bucket against a
|
|
1129
|
+
median+MAD baseline of the ones before it. Buckets rather than update ticks on purpose: the tick is
|
|
1130
|
+
unevenly spaced and shorter than a client's sampling period, so counting on it compares windows of
|
|
1131
|
+
different lengths and calls the difference a signal. Only add it when the observer's calls all come
|
|
1132
|
+
from the **same SFU** — the finding's meaning is "these clients share only that server".
|
|
1133
|
+
- **`TrackDeliveryMismatchDetector`** — *are the two ends of a track disagreeing?*
|
|
1134
|
+
- **`UnconsumedTrackDetector`** — *is anyone actually subscribed?* (reads the resolver's silence)
|
|
1135
|
+
- **`TurnServerHealthDetector`** — *does trouble cluster on one relay?*
|
|
1136
|
+
- **`TurnServerOutageDetector`** — covers the case the health detector structurally cannot. The
|
|
1137
|
+
health detector groups clients by the server relaying them and asks how many report issues — it
|
|
1138
|
+
needs clients *on* the server to ask. When a TURN server dies, allocation fails: existing sessions
|
|
1139
|
+
drop and new clients never obtain a relay candidate through it, so they are never attributed to it
|
|
1140
|
+
at all. Its population goes to zero and the health detector falls silent for the worst possible
|
|
1141
|
+
reason. **Degradation makes clients unhappy; an outage makes them disappear.** Absence is a
|
|
1142
|
+
dangerous signal, so the **control group** is the heart of the design: a call ending, everyone
|
|
1143
|
+
leaving at 6pm, and a fleet-wide network event all look identical to an outage. It refuses to blame
|
|
1144
|
+
a server unless clients *not* relayed through it are demonstrably still connected
|
|
1145
|
+
(`requireControlGroup`, on by default).
|
|
1146
|
+
|
|
1147
|
+
#### There is no ICE detector
|
|
1148
|
+
|
|
1149
|
+
ICE trouble is reported by `client-monitor-js` >= 4.6.0 as the keyed issues `ice-disconnected`,
|
|
1150
|
+
`ice-connection-failed`, `ice-transport-stalled` and `unstable-ice-path`, each with hysteresis and
|
|
1151
|
+
multi-signal confirmation behind it. An `IceDisruptionDetector` used to re-derive that server-side
|
|
1152
|
+
from raw state transitions; it has been removed, because the server sees less and guesses more. The
|
|
1153
|
+
client knows whether `disconnected` persisted or healed in 200 ms; the observer does not.
|
|
1154
|
+
|
|
1155
|
+
Correlating ICE trouble is now configuration, not a class:
|
|
1156
|
+
|
|
1157
|
+
```ts
|
|
1158
|
+
observer.addObserverDetector('observer-concurrent-issue-detector', {
|
|
1159
|
+
issueTypes: [ 'ice-disconnected', 'ice-connection-failed', 'ice-transport-stalled' ],
|
|
1160
|
+
});
|
|
1161
|
+
```
|
|
1162
|
+
|
|
1163
|
+
### Grouping by place: the `location` axis
|
|
1164
|
+
|
|
1165
|
+
If your clients report coordinates, `ClientPopulationIssueDetector` can group by **where they are**
|
|
1166
|
+
instead of what they run — which is the grouping network symptoms actually cluster by:
|
|
1167
|
+
|
|
1168
|
+
```ts
|
|
1169
|
+
observer.addObserverDetector('client-population-issue-detector', {
|
|
1170
|
+
issueTypes: [ 'congestion', 'ice-disconnected' ],
|
|
1171
|
+
groupBy: 'location',
|
|
1172
|
+
locationPrecision: 3, // geohash chars: 3 ~156 km, 4 ~39 km, 5 ~5 km
|
|
1173
|
+
resolveClientLocation: (client) => client.attachments?.geo as { latitude: number, longitude: number },
|
|
1174
|
+
});
|
|
1175
|
+
```
|
|
1176
|
+
|
|
1177
|
+
**The client still owns "RTT jumped".** client-monitor's `CongestionDetector` compares each peer
|
|
1178
|
+
connection's RTT against its own EWMA baseline and requires a bandwidth-limitation corroboration
|
|
1179
|
+
before raising `congestion`. Absolute RTT is not comparable between clients — someone 200 ms away is
|
|
1180
|
+
*always* 200 ms away, so the only signal is deviation from that client's own baseline, which is
|
|
1181
|
+
exactly what the client measures. The observer's contribution is the part no endpoint can see: that
|
|
1182
|
+
many of the affected clients are **in the same place at the same time**.
|
|
1183
|
+
|
|
1184
|
+
Three things to know:
|
|
1185
|
+
|
|
1186
|
+
- **Cells, not radii.** The population is a geohash prefix. "Within N km" is a clustering problem —
|
|
1187
|
+
order-dependent, no stable group name, pairwise cost — and a detector needs the *same* group key on
|
|
1188
|
+
every tick for its cooldown and control group to mean anything. The cost is that a cell boundary can
|
|
1189
|
+
split two adjacent clients, which biases towards missing a finding rather than inventing one.
|
|
1190
|
+
- **Only the cell key is reported.** `payload.population` is the geohash; coordinates never enter the
|
|
1191
|
+
issue. These payloads get archived into [call summaries](#call-summaries), so that matters.
|
|
1192
|
+
- **Geography is confounded with your topology.** The control group is "everyone outside this cell",
|
|
1193
|
+
which cannot separate *"the path into this region degraded"* from *"the SFU serving this region
|
|
1194
|
+
degraded"*. If a region maps largely onto one deployment, both hypotheses fit the same evidence —
|
|
1195
|
+
so the finding concludes `infrastructure` and points you at `SfuCongestionDetector` /
|
|
1196
|
+
`TurnServerHealthDetector`, which answer whether clients elsewhere on the same server also
|
|
1197
|
+
degraded. It does not claim an attribution it cannot support.
|
|
1198
|
+
|
|
1199
|
+
Coordinates are not in `ClientSample`, so `resolveClientLocation` is required; without it the detector
|
|
1200
|
+
logs a warning at construction and finds nothing, rather than quietly reporting no findings forever.
|
|
1201
|
+
|
|
1202
|
+
### Validators — one-shot structural checks
|
|
1203
|
+
|
|
1204
|
+
Every detector above answers *"is something wrong right now?"* and runs on every tick, because the
|
|
1205
|
+
answer legitimately changes. A **validator** answers *"is this deployment built correctly?"* — which
|
|
1206
|
+
only changes when you deploy. So it is not configured on and left running: you **start** one, it runs
|
|
1207
|
+
until it can decide, reports once, and the observer drops it.
|
|
1208
|
+
|
|
1209
|
+
```ts
|
|
1210
|
+
observer.addValidator('simulcast-receivers', { minChecks: 5 });
|
|
1211
|
+
|
|
1212
|
+
observer.on('validation-ready', ({ validator, report }) => {
|
|
1213
|
+
if (!report.ready) return;
|
|
1214
|
+
if (report.verdict === 'layer-decided-lowest-common-denominator') page(validator, report);
|
|
1215
|
+
});
|
|
1216
|
+
|
|
1217
|
+
onDeploy(() => observer.addValidator('simulcast-receivers')); // check again
|
|
1218
|
+
```
|
|
1219
|
+
|
|
1220
|
+
`observer.validators` is the set currently running — normally empty, since each removes itself on
|
|
1221
|
+
finishing. There is no revalidation timer: a deploy, not elapsed time, is what makes a structural
|
|
1222
|
+
verdict stale, so re-checking means starting another.
|
|
1223
|
+
|
|
1224
|
+
**Cancelling.** A check that has not decided can be stopped, by name or by instance:
|
|
1225
|
+
|
|
1226
|
+
```ts
|
|
1227
|
+
observer.cancelValidator('simulcast-receivers', 'sfu redeployed');
|
|
1228
|
+
|
|
1229
|
+
// or one specific instance — `observer.validators` holds what is running
|
|
1230
|
+
for (const validator of observer.validators) observer.cancelValidator(validator, 'shutting down');
|
|
1231
|
+
```
|
|
1232
|
+
|
|
1233
|
+
Cancelling is **not** silent discarding. The validator finishes `inconclusive` with your reason,
|
|
1234
|
+
emits `validation-ready` like any other completion, and removes itself. That matters twice over:
|
|
1235
|
+
anything waiting on the verdict would otherwise wait forever, and *"we stopped asking"* is a
|
|
1236
|
+
materially different outcome from *"we asked and learned nothing"* — which is exactly what an
|
|
1237
|
+
`inconclusive` carrying a reason records. Pass a real reason; the default tells the reader nothing
|
|
1238
|
+
they could not already infer. `observer.close()` cancels whatever is still running with
|
|
1239
|
+
`'observer closed'`.
|
|
1240
|
+
|
|
1241
|
+
| Validator | `addValidator` name | Question | Also raises |
|
|
1242
|
+
|-----------|---------------------|----------|-------------|
|
|
1243
|
+
| `SimulcastReceiverValidator` 🔗 | `simulcast-receivers` | Does the SFU pick layers per receiver, or drag the publisher down to the worst one? | `WORST_RECEIVER_CONTAGION` |
|
|
1244
|
+
| `RemoteTrackResolverValidator` | `remote-track-resolver` | Is the resolver actually linking anything? | `REMOTE_TRACK_LINKS_UNRESOLVED` |
|
|
1245
|
+
| `CodecConsistencyValidator` | `codec-consistency` | Is everyone on the same codec — and is it the one you think you negotiated? | `CODEC_INCONSISTENCY` |
|
|
1246
|
+
|
|
1247
|
+
**`SimulcastReceiverValidator`** — simulcast (or SVC) exists so one slow
|
|
1248
|
+
participant doesn't set everyone's quality: with several encodings the server hands the struggling
|
|
1249
|
+
receiver a lower layer and leaves the rest alone. Without it — or with a server that relays RTCP end
|
|
1250
|
+
to end, so the publisher's bandwidth estimate collapses to the slowest receiver — the only way to
|
|
1251
|
+
serve them is to make the *source* send less. Both causes look identical from outside; what the check
|
|
1252
|
+
establishes is whether per-receiver adaptation happens at all.
|
|
1253
|
+
|
|
1254
|
+
| `verdict` | meaning |
|
|
1255
|
+
|-----------|---------|
|
|
1256
|
+
| `layer-decided-per-receiver` | verified — a receiver fell far behind and the publisher carried on |
|
|
1257
|
+
| `layer-decided-lowest-common-denominator` | the publisher followed its worst receiver; everyone gets the slowest participant's quality |
|
|
1258
|
+
| `inconclusive` | cancelled, or the observer closed, before it could decide |
|
|
1259
|
+
|
|
1260
|
+
**Not finishing is not a pass.** The check only runs when a publisher has 3+ receivers and one is at
|
|
1261
|
+
most half the median; plenty of healthy deployments never present that. A validator that never sees it
|
|
1262
|
+
simply keeps running and never reports — it does not quietly succeed. `report.checks` counts the times
|
|
1263
|
+
the check genuinely ran, so an `inconclusive` with `checks: 0` says plainly that nothing was verified.
|
|
1264
|
+
|
|
1265
|
+
**`RemoteTrackResolverValidator`** exists because of a specific, nasty failure mode. Four things here
|
|
1266
|
+
are built on publisher↔subscriber links — `IssueFanOutDetector`,
|
|
1267
|
+
`PublisherFaultCorroborationDetector`, `TrackDeliveryMismatchDetector`, `UnconsumedTrackDetector`
|
|
1268
|
+
(and `SimulcastReceiverValidator`) — and every one of them correctly does *nothing* when the links
|
|
1269
|
+
are missing rather than guessing. So a resolver wired to the wrong id field leaves all of them
|
|
1270
|
+
permanently silent, and **silence is what a healthy deployment looks like too**: you would conclude
|
|
1271
|
+
your calls were clean when in fact nothing was ever examined. Verdicts: `links-resolved` /
|
|
1272
|
+
`no-links-resolved` / `inconclusive`. Run it at start-up and after changing the resolver or the SFU's
|
|
1273
|
+
id scheme.
|
|
1274
|
+
|
|
1275
|
+
**`CodecConsistencyValidator`** answers two things at once. A **split** — participants of one call on
|
|
1276
|
+
different codecs — is a real fault with a confusing symptom: an SFU that forwards without transcoding
|
|
1277
|
+
cannot serve them all, so some pairs see each other and some do not, with no error anywhere. Only
|
|
1278
|
+
something holding every participant at once can see it. The quieter half is the silent fallback: a
|
|
1279
|
+
deployment configured for VP9 or AV1 drops to VP8 whenever one endpoint cannot negotiate the
|
|
1280
|
+
preference, the call keeps working at a higher bitrate than budgeted, and the team believes it
|
|
1281
|
+
shipped AV1 months ago. Give it `expected` and it says so. Verdicts: `codec-consistent` /
|
|
1282
|
+
`codec-split` / `unexpected-codec` / `inconclusive`.
|
|
1283
|
+
|
|
1284
|
+
```ts
|
|
1285
|
+
observer.addValidator('remote-track-resolver');
|
|
1286
|
+
observer.addValidator('codec-consistency', { expected: { video: 'video/VP9', audio: 'audio/opus' } });
|
|
1287
|
+
```
|
|
1288
|
+
|
|
1289
|
+
#### `CallIssue` vs `ObserverIssue`
|
|
1290
|
+
|
|
1291
|
+
Server-raised findings come in two kinds, distinguished by the scope that raised them:
|
|
1292
|
+
|
|
1293
|
+
| | raised by | delivered as | `scope` |
|
|
1294
|
+
|---|---|---|---|
|
|
1295
|
+
| `CallIssue` | `observedCall.addIssue(...)` | `call-issue` | `'call'` |
|
|
1296
|
+
| `ObserverIssue` | `observer.addIssue(...)` | `observer-issue` | `'observer'` |
|
|
1297
|
+
|
|
1298
|
+
Both share `IssueBase` — `type`, `timestamp`, `conclusion?`, `payload?` — and `Issue` is the union,
|
|
1299
|
+
discriminated on `scope`.
|
|
1300
|
+
|
|
1301
|
+
```ts
|
|
1302
|
+
observer.on('call-issue', ({ observedCall, issue }) => {
|
|
1303
|
+
issue.scope; // 'call'
|
|
1304
|
+
observedCall.callId; // the call — NOT repeated in the payload
|
|
1305
|
+
issue.conclusion?.faultDomain;
|
|
1306
|
+
issue.payload; // evidence only
|
|
1307
|
+
});
|
|
1308
|
+
|
|
1309
|
+
observer.on('observer-issue', ({ issue }) => {
|
|
1310
|
+
issue.scope; // 'observer'
|
|
1311
|
+
});
|
|
1312
|
+
```
|
|
1313
|
+
|
|
1314
|
+
`scope` is stamped by `addIssue` rather than asked of the detector: it is a fact about *where the
|
|
1315
|
+
finding was raised*, which the entity knows and a detector should not have to restate. Having it on
|
|
1316
|
+
the issue — not merely implied by which event fired — keeps a finding self-describing once it leaves
|
|
1317
|
+
the bus, into a shared handler, a log line or a queue.
|
|
1318
|
+
|
|
1319
|
+
**The payload is evidence and nothing else.** It no longer repeats `type`, `scope`, or the `callId`
|
|
1320
|
+
already carried by the event, and `conclusion` was lifted out of it to a first-class field. A payload
|
|
1321
|
+
that restates its own envelope invites the two to disagree — and they did, because nothing kept them
|
|
1322
|
+
in step. `payload` is always an object (the `string` form is gone, along with `issuePayloadOf`); use
|
|
1323
|
+
`issuePayloadAsString(issue)` at a boundary that genuinely needs text.
|
|
1324
|
+
|
|
1325
|
+
#### Conclusions
|
|
1326
|
+
|
|
1327
|
+
Every issue-driven finding carries a `conclusion` — the interpretation step, so the person reading
|
|
1328
|
+
the alert doesn't have to perform it. It sits **beside** the evidence, not inside it:
|
|
1329
|
+
|
|
1330
|
+
```jsonc
|
|
1331
|
+
{
|
|
1332
|
+
"type": "CROSS_CALL_ISSUE_ONSET_BURST",
|
|
1333
|
+
"scope": "observer",
|
|
1334
|
+
"timestamp": 1739812345678,
|
|
1335
|
+
"conclusion": {
|
|
1336
|
+
"faultDomain": "infrastructure",
|
|
1337
|
+
"summary": "network congestion is open across independent calls at the same time — 6 of 40 calls (11/300 clients)",
|
|
1338
|
+
"recommendation": "check SFU egress bandwidth and host network saturation before looking at any single participant",
|
|
1339
|
+
"confidence": 0.85
|
|
1340
|
+
},
|
|
1341
|
+
"payload": {
|
|
1342
|
+
"issueType": "congestion",
|
|
1343
|
+
"calls": 40, "affectedCalls": 6,
|
|
1344
|
+
"perCall": [ { "callId": "…", "affectedClients": 4, "totalClients": 9 } ]
|
|
1345
|
+
}
|
|
739
1346
|
}
|
|
740
1347
|
```
|
|
741
1348
|
|
|
742
|
-
|
|
1349
|
+
`faultDomain` is one of `infrastructure`, `call`, `published-track`, `endpoint`, `client-population`
|
|
1350
|
+
or `unknown`, and it comes from the **spread**, not the issue type — congestion in one call is a
|
|
1351
|
+
meeting problem, congestion in six calls is a server problem, and the client reported the identical
|
|
1352
|
+
symptom in both.
|
|
1353
|
+
|
|
1354
|
+
One case is worth knowing about because it inverts the usual reading: **`cpu-limitation` spread
|
|
1355
|
+
across many independent calls concludes `client-population`, not `infrastructure`.** Endpoint CPU is
|
|
1356
|
+
owned by the endpoint, so breadth there points at what those endpoints share — a recent client
|
|
1357
|
+
release, a browser version, shared VDI hardware — and paging the SFU on-call would be wrong. The
|
|
1358
|
+
conclusion table encodes that so nobody has to rediscover it during an incident.
|
|
1359
|
+
|
|
1360
|
+
Unknown issue types (your own custom client detectors) still produce a structurally valid conclusion
|
|
1361
|
+
from the spread alone; they just get generic wording.
|
|
1362
|
+
|
|
1363
|
+
Two functions are exported, one per scope: `concludeCallIssue()` and `concludeObserverIssue()`. They
|
|
1364
|
+
are separate because a detector already knows its scope, and a single generic function forced every
|
|
1365
|
+
caller to pass the other scope's fields as placeholders — call-scoped detectors passing
|
|
1366
|
+
`affectedCalls: 1, totalCalls: 1` forever, observer-scoped ones passing a participant ratio that was
|
|
1367
|
+
deliberately never read. Placeholders like that invite being read as if they meant something.
|
|
1368
|
+
|
|
1369
|
+
#### Cost
|
|
1370
|
+
|
|
1371
|
+
Detectors run inside `call.update()`, on your event loop, so their cost matters. Two things keep it
|
|
1372
|
+
off the participant axis:
|
|
1373
|
+
|
|
1374
|
+
- **Issues are pushed, not polled.** A detector holds only what the registry handed it, so an
|
|
1375
|
+
`update()` that finds `size === 0` — the overwhelmingly common case — costs one comparison,
|
|
1376
|
+
whatever the participant count. Nothing iterates clients looking for trouble.
|
|
1377
|
+
- **So are unconsumed tracks.** `observedCall.unconsumedOutboundTracks` is maintained by the resolver
|
|
1378
|
+
as tracks gain and lose subscribers, so `UnconsumedTrackDetector` reads a set that is normally
|
|
1379
|
+
empty instead of walking every published track (529 µs → 65 µs per tick at 1 200 tracks).
|
|
1380
|
+
- **Track lookups start from the affected minority.** A detector resolving an issue to its published
|
|
1381
|
+
track searches the *reporting client's* peer connections (typically one or two), not the call.
|
|
1382
|
+
|
|
1383
|
+
At 20 calls × 12 participants (2 640 subscriptions) the whole detector pass costs ~1.3 ms per tick.
|
|
1384
|
+
`yarn bench` prints a per-detector breakdown for your own shape.
|
|
1385
|
+
|
|
1386
|
+
#### Worked examples
|
|
1387
|
+
|
|
1388
|
+
[`examples/detectors.ts`](./examples/detectors.ts) (`yarn example:detectors`) runs one scenario per
|
|
1389
|
+
detector — the question it answers, its full config, the synthetic traffic that makes it fire, and
|
|
1390
|
+
the finding with its conclusion. It asserts every expected finding is produced, so it doubles as a
|
|
1391
|
+
smoke test. [`examples/sfu-observer.ts`](./examples/sfu-observer.ts) (`yarn example`) is the end-to-end
|
|
1392
|
+
tour instead: ingest → correlate → react, with the mediasoup wiring alongside.
|
|
1393
|
+
|
|
1394
|
+
#### `TrackDeliveryMismatchDetector` — resolving an ambiguous symptom
|
|
1395
|
+
|
|
1396
|
+
A dry track ("no bytes are arriving") is the clearest symptom there is and, on its own, completely
|
|
1397
|
+
ambiguous. A receiver seeing silence cannot distinguish *the camera was switched off* from *the SFU
|
|
1398
|
+
stopped forwarding* from *my own consumer wedged* — all three look identical from the browser.
|
|
1399
|
+
|
|
1400
|
+
Joining the two ends of the published track resolves it:
|
|
1401
|
+
|
|
1402
|
+
| publisher | subscribers | verdict |
|
|
1403
|
+
|---|---|---|
|
|
1404
|
+
| sending | **all** dry | `PUBLISHED_TRACK_NOT_DELIVERED` — the forwarding path |
|
|
1405
|
+
| sending | **some** dry | `RECEIVER_TRACK_NOT_DELIVERED` — those consumers (in mediasoup: recreate them) |
|
|
1406
|
+
| dry | any dry | `PUBLISHER_TRACK_DRY` — the source stopped; **not** an SFU fault |
|
|
1407
|
+
|
|
1408
|
+
The publisher side is judged from both available signals: its own `dry-outbound-track` issue when the
|
|
1409
|
+
client reports one, and the observed outbound RTP (`deltaPacketsSent`) as fallback and corroboration.
|
|
1410
|
+
That combination is what makes the first row trustworthy — the server can state that packets
|
|
1411
|
+
demonstrably left the publisher during the same interval in which every receiver got nothing.
|
|
1412
|
+
|
|
1413
|
+
This is the "SFU forwarding mismatch" check, and it needs **no mediasoup instrumentation at all** —
|
|
1414
|
+
the clients' own dry-track verdicts plus the resolver links are sufficient.
|
|
1415
|
+
|
|
1416
|
+
#### `UnconsumedTrackDetector` — reading the resolver's silence
|
|
1417
|
+
|
|
1418
|
+
The one detector where the *absence* of links is the signal: a track still pushing packets whose
|
|
1419
|
+
`remoteInboundTracks` set is empty, i.e. uplink and SFU ingress spent on media nobody receives
|
|
1420
|
+
(everyone has the publisher hidden, a simulcast layer no viewer selects, or an app that forgot to
|
|
1421
|
+
stop a track). It waits `minUnconsumedDurationInMs` first, since a gap between publishing and the
|
|
1422
|
+
first subscription is normal at join time.
|
|
1423
|
+
|
|
1424
|
+
Note the trap this one has to guard against, and why it checks `call.remoteTrackResolver` at runtime
|
|
1425
|
+
rather than trusting the flag alone: **"no subscribers" and "no resolver configured" produce the
|
|
1426
|
+
identical observation.** Without a resolver it would report every published track in the call as
|
|
1427
|
+
unconsumed.
|
|
1428
|
+
|
|
1429
|
+
## Call summaries
|
|
1430
|
+
|
|
1431
|
+
Everything else in this library is about *now*. Detectors answer "is something wrong right now",
|
|
1432
|
+
validators answer a structural question once, and both read state the call throws away when it ends.
|
|
1433
|
+
A **call summary** is the one thing that outlives the call: who was in it, what was raised against
|
|
1434
|
+
it, how it scored — the questions asked *after* the meeting, by support, by billing, by whoever is
|
|
1435
|
+
writing the incident note.
|
|
1436
|
+
|
|
1437
|
+
It is configured on the observer, at construction:
|
|
1438
|
+
|
|
1439
|
+
```ts
|
|
1440
|
+
const observer = new Observer({
|
|
1441
|
+
callSummary: {
|
|
1442
|
+
include: [ 'clients', 'issues', 'turnServers', 'scores' ],
|
|
1443
|
+
},
|
|
1444
|
+
});
|
|
1445
|
+
|
|
1446
|
+
observer.on('call-summary', ({ summary }) => archive(summary));
|
|
1447
|
+
```
|
|
1448
|
+
|
|
1449
|
+
Omit `callSummary`, or set it to `null`, and there are no summaries and **not one extra bus
|
|
1450
|
+
subscription**. Pass an object — `{}` is valid — and every call this observer creates carries one.
|
|
1451
|
+
|
|
1452
|
+
> **Why construction-time, when detectors are added per call?** A summary is a record of what
|
|
1453
|
+
> happened, and a record you can switch on halfway through is a record with a hole in it. Calls that
|
|
1454
|
+
> started before the switch would carry different sections from calls that started after, with
|
|
1455
|
+
> nothing on either to say which. One shape for every call, or none.
|
|
1456
|
+
|
|
1457
|
+
### Sections are opt-in, and absence means "not collected"
|
|
1458
|
+
|
|
1459
|
+
`include` picks from four built-ins, and **the default is `[]`** — none of them:
|
|
1460
|
+
|
|
1461
|
+
| Section | Contains |
|
|
1462
|
+
|---------|----------|
|
|
1463
|
+
| `clients` | `clientIds` (join order), `peak`, `joined`, `left`. Identifiers and counts only |
|
|
1464
|
+
| `issues` | `CallIssue[]`, in the order raised, capped by `maxIssues` |
|
|
1465
|
+
| `turnServers` | `serverUrls` that carried media, and `clientsRelayed` |
|
|
1466
|
+
| `scores` | `min` / `max` / `median` of the call score, and `samples` |
|
|
1467
|
+
|
|
1468
|
+
**A missing section means it was never collected — never "nothing happened".** Reading
|
|
1469
|
+
`summary.issues === undefined` as "this call was clean" is the one misreading this type invites, so
|
|
1470
|
+
there is no default-empty section to make it easy. This is the same rule as `inconclusive` on a
|
|
1471
|
+
validator: silence is not success.
|
|
1472
|
+
|
|
1473
|
+
The `clients` section is deliberately identifiers and counts. Anything *about* a client — browser,
|
|
1474
|
+
platform, region — is already on `observedClient` while the call is live, and belongs in
|
|
1475
|
+
`attachments` via an enricher if you want it kept; see below.
|
|
1476
|
+
|
|
1477
|
+
### Enrichers: fold in anything, from any call-scoped event
|
|
1478
|
+
|
|
1479
|
+
```ts
|
|
1480
|
+
new Observer({
|
|
1481
|
+
callSummary: {
|
|
1482
|
+
include: [ 'issues' ],
|
|
1483
|
+
enrich: {
|
|
1484
|
+
'client-joined': (summary, { observedClient }) => {
|
|
1485
|
+
// serialisable facts only — the region string, never the live object it came from
|
|
1486
|
+
((summary.attachments.regions ??= []) as string[]).push(String(observedClient.appData.region));
|
|
1487
|
+
},
|
|
1488
|
+
},
|
|
1489
|
+
},
|
|
1490
|
+
});
|
|
1491
|
+
```
|
|
1492
|
+
|
|
1493
|
+
Each enricher is typed against its own event's payload. Only **call-scoped** events are accepted —
|
|
1494
|
+
the ones carrying an `observedCall`. An enricher on `observer-issue` or `validation-ready` will not
|
|
1495
|
+
compile, because there is no single call to attribute a fleet-wide fact to, and quietly writing it
|
|
1496
|
+
into every open summary would be worse than a type error.
|
|
1497
|
+
|
|
1498
|
+
The library never writes to `summary.attachments`, so nothing you put there can collide with a
|
|
1499
|
+
section added in a future version.
|
|
1500
|
+
|
|
1501
|
+
> **Why `attachments` and not `appData`.** `appData` is live working state hung off an entity for
|
|
1502
|
+
> that entity's lifetime, and it may hold things that cannot be serialised — a mediasoup router, a
|
|
1503
|
+
> socket. A summary is the opposite: it outlives the call so it can be **shipped**, and it reaches
|
|
1504
|
+
> you on `call-summary` while the call it describes is being torn down, so an unserialisable value
|
|
1505
|
+
> in it points at something already gone. Same contract as `attachments` on a `ClientSample`: read
|
|
1506
|
+
> the live object off `observedCall` / `observedClient` in the enricher, attach what serialises —
|
|
1507
|
+
> the router's `id`, not the router. An enricher that throws is logged and skipped — a summary is a
|
|
1508
|
+
side-channel, and nothing about a call should break because a field could not be recorded.
|
|
1509
|
+
|
|
1510
|
+
### Caps announce what they dropped
|
|
1511
|
+
|
|
1512
|
+
`maxIssues` (default `500`) and `maxClientIds` (default `10_000`) bound the two unbounded lists.
|
|
1513
|
+
When either bites, `summary.truncated` appears with the shortfall — present **only** when something
|
|
1514
|
+
was actually dropped. That is what makes dropping safe: the true count is recoverable as
|
|
1515
|
+
`issues.length + (truncated?.issues ?? 0)`. A silently truncated summary is worse than no summary,
|
|
1516
|
+
because someone will count `issues.length` and report it as the issue count.
|
|
1517
|
+
|
|
1518
|
+
`issues` is the plain array, with no derived tallies alongside it. A count is `issues.length` and a
|
|
1519
|
+
per-type count is one `filter` — both cheaper at the call site than kept correct here.
|
|
1520
|
+
|
|
1521
|
+
### Reading it
|
|
1522
|
+
|
|
1523
|
+
`observedCall.summary` is live: read it at any point during the call. It is also delivered once on
|
|
1524
|
+
`call-summary`, emitted **inside** `close()` while the call is still in `observer.observedCalls` —
|
|
1525
|
+
after that the call is gone and there is nothing left to ask. `observer.close()` closes its calls
|
|
1526
|
+
first and its collector afterwards, so every summary still makes it out.
|
|
1527
|
+
|
|
1528
|
+
Cost is **one bus listener per subscribed event type, for the whole observer** — not one per call. A
|
|
1529
|
+
per-call design would be quadratic in concurrent calls: at 500 calls and eight events, 4 000
|
|
1530
|
+
listeners each doing 500 no-op invocations per event. Percentiles are computed once, at close.
|
|
743
1531
|
|
|
744
1532
|
## Remote track resolution (mediasoup / SFU)
|
|
745
1533
|
|
|
746
1534
|
In an SFU, one participant's **outbound** track is delivered to other participants as **inbound**
|
|
747
1535
|
tracks (one **publisher** → many **subscribers**). Correlation is **opt-in** per observer: set
|
|
748
|
-
`ObserverConfig.
|
|
1536
|
+
`ObserverConfig.createRemoteTrackResolver`, a factory invoked when each call is created that returns the
|
|
749
1537
|
call's `RemoteTrackResolver` (or `undefined` for none).
|
|
750
1538
|
|
|
751
1539
|
`RemoteTrackResolver` is a generic, strategy-driven class. It subscribes to the bus (filtered to
|
|
@@ -756,7 +1544,7 @@ the tracks: `inboundTrack.remoteOutboundTrack` and `outboundTrack.remoteInboundT
|
|
|
756
1544
|
import { Observer, createDefaultMediasoupRemoteTrackResolverFactory } from '@observertc/observer-js';
|
|
757
1545
|
|
|
758
1546
|
const observer = new Observer({
|
|
759
|
-
|
|
1547
|
+
createRemoteTrackResolver: createDefaultMediasoupRemoteTrackResolverFactory(),
|
|
760
1548
|
});
|
|
761
1549
|
|
|
762
1550
|
// later, given tracks (links are kept up to date as tracks come and go):
|
|
@@ -775,7 +1563,7 @@ id is just whatever links a subscribed track to the published one:
|
|
|
775
1563
|
import { Observer, RemoteTrackResolver } from '@observertc/observer-js';
|
|
776
1564
|
|
|
777
1565
|
const observer = new Observer({
|
|
778
|
-
|
|
1566
|
+
createRemoteTrackResolver: (observedCall) => new RemoteTrackResolver(observedCall, {
|
|
779
1567
|
resolveOutboundTrackPublisherId: (out) => out.attachments?.mediaId as string | undefined,
|
|
780
1568
|
resolveInboundTrackPublisherId: (inb) => inb.attachments?.mediaId as string | undefined,
|
|
781
1569
|
resolveInboundTrackSubscriberId: (inb) => inb.attachments?.subId as string | undefined, // optional
|
|
@@ -786,6 +1574,56 @@ const observer = new Observer({
|
|
|
786
1574
|
For the mediasoup factory, the application puts `producerId` / `consumerId` (and optionally
|
|
787
1575
|
`direction`, `label`) into the track `attachments`.
|
|
788
1576
|
|
|
1577
|
+
### Resolution stands alone — you do not need router observation
|
|
1578
|
+
|
|
1579
|
+
Track resolution and [mediasoup router observation](#mediasoup-router-observation) are two
|
|
1580
|
+
**independent** opt-ins that happen to share a vendor name. Nothing in `RemoteTrackResolver` reads
|
|
1581
|
+
`ObservedMediasoupRouter`, and nothing in `ObservedMediasoupRouter` touches calls, clients or tracks.
|
|
1582
|
+
|
|
1583
|
+
So if your application already builds its own per-router report and you only want the detectors that
|
|
1584
|
+
need publisher↔subscriber links, set `createRemoteTrackResolver` and simply never call
|
|
1585
|
+
`observer.createObservedMediasoupRouter(...)`. No `MediasoupRouterSample` is created, nothing
|
|
1586
|
+
accumulates, and these keep working:
|
|
1587
|
+
|
|
1588
|
+
`UnconsumedTrackDetector`, `PublisherFaultCorroborationDetector`, `TrackDeliveryMismatchDetector`,
|
|
1589
|
+
`IssueFanOutDetector`, `RemoteTrackResolverValidator`, `SimulcastReceiverValidator`.
|
|
1590
|
+
|
|
1591
|
+
### Backing a strategy with your own mapping
|
|
1592
|
+
|
|
1593
|
+
The three resolvers are plain functions returning a link key, so they can read a table your own
|
|
1594
|
+
report already maintains rather than something the client attached. The only requirement is a key
|
|
1595
|
+
**visible on both sides**. SSRC is the useful one, because it needs no client cooperation at all —
|
|
1596
|
+
mediasoup knows each consumer's `rtpParameters.encodings[].ssrc` server-side, and the subscriber's
|
|
1597
|
+
inbound RTP reports the same value:
|
|
1598
|
+
|
|
1599
|
+
```ts
|
|
1600
|
+
// your own table, filled where you already create consumers
|
|
1601
|
+
const ssrcToProducerId = new Map<number, string>();
|
|
1602
|
+
|
|
1603
|
+
const observer = new Observer({
|
|
1604
|
+
createRemoteTrackResolver: (observedCall) => new RemoteTrackResolver(observedCall, {
|
|
1605
|
+
resolveOutboundTrackPublisherId: (out) => out.attachments?.producerId as string | undefined,
|
|
1606
|
+
resolveInboundTrackPublisherId: (inb) => {
|
|
1607
|
+
const ssrc = inb.getInboundRtp()?.ssrc;
|
|
1608
|
+
|
|
1609
|
+
return ssrc === undefined ? undefined : ssrcToProducerId.get(ssrc);
|
|
1610
|
+
},
|
|
1611
|
+
}),
|
|
1612
|
+
});
|
|
1613
|
+
```
|
|
1614
|
+
|
|
1615
|
+
> **A key that arrives late still links.** A track announces itself once, but its `attachments` are
|
|
1616
|
+
> replaced on every sample and a table like the one above is inherently racy against sample arrival.
|
|
1617
|
+
> Tracks whose key does not resolve at first sight are held and retried on their own
|
|
1618
|
+
> `*-track-updated`, i.e. exactly when new stats arrive for them — so a key that appears on the second
|
|
1619
|
+
> sample links then, rather than being lost for the track's lifetime. `resolver.pendingTrackCounts`
|
|
1620
|
+
> reports how many are still waiting; in a healthy setup it is `{ inbound: 0, outbound: 0 }`.
|
|
1621
|
+
|
|
1622
|
+
If a strategy resolves *nothing* the failure is quiet — "no subscribers" and "no links resolved" look
|
|
1623
|
+
identical from the outside. That is what
|
|
1624
|
+
[`RemoteTrackResolverValidator`](#validators--one-shot-structural-checks) is for; run it once in
|
|
1625
|
+
staging after wiring up a custom strategy.
|
|
1626
|
+
|
|
789
1627
|
---
|
|
790
1628
|
|
|
791
1629
|
## Mediasoup router observation
|
|
@@ -799,121 +1637,224 @@ transitions. `ObservedMediasoupRouter` captures that server-side view into a
|
|
|
799
1637
|
### The concept
|
|
800
1638
|
|
|
801
1639
|
You hand the observer a live mediasoup `Router`; it attaches to mediasoup's own `observer` API and,
|
|
802
|
-
from then on, **passively
|
|
1640
|
+
from then on, **passively tracks** the router's topology and lifecycle — with no polling and no
|
|
803
1641
|
changes to your media code:
|
|
804
1642
|
|
|
805
|
-
- new transports (`webrtc` / `plain` / `pipe` / `direct`), their selected `tuple`,
|
|
806
|
-
transitions
|
|
1643
|
+
- new transports (`webrtc` / `plain` / `pipe` / `direct`), their selected `tuple`, ICE/DTLS/SCTP
|
|
1644
|
+
state transitions and `connectedAt`;
|
|
807
1645
|
- producers (codec, SSRCs/RIDs, `pause`/`resume`) and consumers (`pause`/`resume`,
|
|
808
1646
|
`producerPaused`/`producerResumed`);
|
|
809
1647
|
- data producers and data consumers;
|
|
810
1648
|
- `createdAt` / `closedAt` for every entity above.
|
|
811
1649
|
|
|
812
|
-
|
|
813
|
-
[`src/schema/MediasoupRouter.ts`](./src/schema/MediasoupRouter.ts)
|
|
814
|
-
|
|
1650
|
+
It keeps all of this **in memory**, in a single `MediasoupRouterSample` exposed as
|
|
1651
|
+
`observedRouter.sample` — see [`src/schema/MediasoupRouter.ts`](./src/schema/MediasoupRouter.ts). The
|
|
1652
|
+
sample **accumulates for the life of the router**: closed transports/producers/consumers are kept
|
|
1653
|
+
(with their `closedAt` set), not removed. Read it whenever you like — it's a plain object you own.
|
|
1654
|
+
|
|
1655
|
+
### Memory & large meetings
|
|
1656
|
+
|
|
1657
|
+
This is intentionally the **simplest** approach — everything lives in memory and nothing is sampled
|
|
1658
|
+
or evicted for you. That's fine for typical rooms, but be aware of the cost at scale:
|
|
1659
|
+
|
|
1660
|
+
- **Consumers grow as O(N²)** on a single flat router: with `N` participants each producing audio +
|
|
1661
|
+
video and consuming everyone else, the sample holds roughly `2·N·(N−1)` consumer records (≈ 19,800
|
|
1662
|
+
for `N` = 100).
|
|
1663
|
+
- The sample is **cumulative** — closed entities and their `history` are retained — so it also grows
|
|
1664
|
+
with call duration and churn (renegotiation, simulcast layer changes, rejoins).
|
|
1665
|
+
|
|
1666
|
+
A 100-participant flat router can therefore reach tens of MB and keep growing. There is **no built-in
|
|
1667
|
+
sink, snapshotting, or eviction** — by design. **If you run large meetings, do your own sampling:**
|
|
1668
|
+
on your own cadence read `observedRouter.sample` (snapshot/serialize/persist what you need), drop what
|
|
1669
|
+
you don't, and close routers you no longer track. (mediasoup also typically shards routers across
|
|
1670
|
+
workers/cores, which keeps any one router small.)
|
|
1671
|
+
|
|
1672
|
+
### Extending the sample, and building your own report
|
|
1673
|
+
|
|
1674
|
+
The sample is yours to annotate. Every entity — the router, each transport, producer, consumer, data
|
|
1675
|
+
producer and data consumer — has an `attachments?: Record<string, unknown>` slot, and there are three
|
|
1676
|
+
ways to fill it, from most declarative to most ad-hoc.
|
|
1677
|
+
|
|
1678
|
+
**1. `enrich` — mirror mediasoup's own `appData`.** The common case: your application already keeps
|
|
1679
|
+
`participantId`, `purpose` and similar on the mediasoup objects, and you want them on the sample.
|
|
1680
|
+
Runs once per entity at creation, before the corresponding event:
|
|
1681
|
+
|
|
1682
|
+
```ts
|
|
1683
|
+
observer.createObservedMediasoupRouter({
|
|
1684
|
+
router,
|
|
1685
|
+
enrich: {
|
|
1686
|
+
producer: (producer) => ({ participantId: producer.appData.participantId, purpose: producer.appData.purpose }),
|
|
1687
|
+
consumer: (consumer) => ({ subscriberId: consumer.appData.subscriberId }),
|
|
1688
|
+
transport: (transport) => ({ role: transport.appData.role }),
|
|
1689
|
+
},
|
|
1690
|
+
});
|
|
1691
|
+
```
|
|
1692
|
+
|
|
1693
|
+
A throwing enricher is caught and logged — it can't take the router's bookkeeping down with it.
|
|
1694
|
+
|
|
1695
|
+
**2. Lifecycle events — enrich on the fly.** Each entity announces itself as
|
|
1696
|
+
`<entity>-sample-added` and `<entity>-sample-closed`, carrying **the live sample object** (not a
|
|
1697
|
+
copy) plus the mediasoup object it came from. Mutating it in the handler is the intended pattern:
|
|
1698
|
+
|
|
1699
|
+
```ts
|
|
1700
|
+
observedRouter.on('producer-sample-added', ({ sample, producer, transport }) => {
|
|
1701
|
+
sample.attachments = { ...sample.attachments, participantId: lookup(producer.id) };
|
|
1702
|
+
});
|
|
1703
|
+
|
|
1704
|
+
observedRouter.on('producer-sample-closed', ({ sample }) => {
|
|
1705
|
+
archive(sample); // its `closedAt` is set
|
|
1706
|
+
});
|
|
1707
|
+
```
|
|
1708
|
+
|
|
1709
|
+
Events: `transport-sample-added` / `-closed`, `producer-sample-added` / `-closed`,
|
|
1710
|
+
`consumer-sample-added` / `-closed`, `data-producer-sample-added` / `-closed`,
|
|
1711
|
+
`data-consumer-sample-added` / `-closed`.
|
|
1712
|
+
|
|
1713
|
+
**3. `attachTo(id, attachments)` — annotate later, from anywhere.** When the knowledge arrives after
|
|
1714
|
+
the entity did (a signalling message, a database lookup that resolved):
|
|
1715
|
+
|
|
1716
|
+
```ts
|
|
1717
|
+
observedRouter.attachTo(producerId, { participantId, joinedFrom: 'mobile' }); // merges
|
|
1718
|
+
```
|
|
1719
|
+
|
|
1720
|
+
Ids are unique across mediasoup entity kinds, so one method covers all of them. It returns `false`
|
|
1721
|
+
for an unknown id rather than failing quietly — which matters when application events race the
|
|
1722
|
+
mediasoup ones. For direct access there are typed accessors: `getTransportSample(id)`,
|
|
1723
|
+
`getProducerSample(id)`, `getConsumerSample(id)`, `getDataProducerSample(id)`,
|
|
1724
|
+
`getDataConsumerSample(id)`. They index the *same* objects the arrays hold, so a lookup is O(1)
|
|
1725
|
+
instead of a `sample.producers.find(...)` scan.
|
|
815
1726
|
|
|
816
|
-
|
|
1727
|
+
#### Building your own report
|
|
817
1728
|
|
|
818
|
-
|
|
819
|
-
on
|
|
820
|
-
|
|
821
|
-
pairing means. Stamp the `callId` into the router's `appData`, build your own index, attach the
|
|
822
|
-
sample to the call in your database — whatever fits your system. The library stays unopinionated and
|
|
823
|
-
loosely coupled.
|
|
1729
|
+
`observedRouter.sample` is live — arrays grow and `history` entries are appended as the router runs,
|
|
1730
|
+
so a report built directly on it keeps changing after you think you're done. Use **`snapshot()`** for
|
|
1731
|
+
a detached deep copy:
|
|
824
1732
|
|
|
825
|
-
|
|
826
|
-
|
|
1733
|
+
```ts
|
|
1734
|
+
const report = {
|
|
1735
|
+
...observedRouter.snapshot(), // never moves again
|
|
1736
|
+
generatedAt: Date.now(),
|
|
1737
|
+
region: process.env.REGION,
|
|
1738
|
+
};
|
|
1739
|
+
```
|
|
827
1740
|
|
|
828
|
-
|
|
829
|
-
|
|
830
|
-
|
|
831
|
-
|
|
832
|
-
|
|
833
|
-
|
|
834
|
-
|
|
835
|
-
|
|
1741
|
+
> **Note on typing.** The sample types no longer carry a `Record<string, unknown>` index signature.
|
|
1742
|
+
> That signature allowed arbitrary top-level keys but also silently accepted typos on real fields and
|
|
1743
|
+
> weakened autocomplete. Custom data belongs in `attachments`, which is typed as such. If you were
|
|
1744
|
+
> assigning ad-hoc keys directly onto a sample object, move them into `attachments`.
|
|
1745
|
+
|
|
1746
|
+
### Matching peer connections — by **event**, not by storage
|
|
1747
|
+
|
|
1748
|
+
The observer correlates the SFU side with the client side **at the peer-connection level**: a
|
|
1749
|
+
mediasoup WebRTC transport and a client's `RTCPeerConnection` share the same id, so whenever an
|
|
1750
|
+
observed peer connection's id matches one of the router's WebRTC transport ids, that's a match.
|
|
1751
|
+
|
|
1752
|
+
**The observer does not store the router (or its sample) on any entity.** Instead, for **every**
|
|
1753
|
+
matching peer connection it emits **`mediasoup-router-matched-with-peer-connection`** and steps
|
|
1754
|
+
back — *your application* decides what the pairing means. The payload carries the full peer-connection
|
|
1755
|
+
ancestry, so you get the router **and** the matched `observedPeerConnection`, `observedClient` and
|
|
1756
|
+
`observedCall` in one place. Stamp the `routerId` into the peer connection's / client's `appData`,
|
|
1757
|
+
build your own index, attach the server sample to the call in your database — whatever fits.
|
|
1758
|
+
|
|
1759
|
+
This matching is **opt-in**: pass `matchPeerConnectionByWebRtcTransportId: true` to
|
|
1760
|
+
`createObservedMediasoupRouter`. When enabled, as peer connections are observed
|
|
1761
|
+
(`peer-connection-added`) the observer checks whether the peer connection's id is one of the router's
|
|
1762
|
+
WebRTC transport ids; on a hit it emits — once per matching peer connection — and keeps watching, so a
|
|
1763
|
+
router serving many participants emits one match per participant's transport. When the flag is omitted
|
|
1764
|
+
or `false`, no matching is performed and the event never fires. The internal listener is removed
|
|
1765
|
+
automatically when the router closes or the observer closes.
|
|
1766
|
+
|
|
1767
|
+
### Ordering contract — observe the router first
|
|
1768
|
+
|
|
1769
|
+
Matching is **forward-only by design**, and that is sufficient because the lifecycle ordering is
|
|
1770
|
+
**guaranteed, not racy**:
|
|
1771
|
+
|
|
1772
|
+
- `ObservedMediasoupRouter` works purely by **subscribing to mediasoup's `observer` API**, so it can
|
|
1773
|
+
only see events that happen *after* it is created. You therefore create it the moment the router
|
|
1774
|
+
exists — **before** any transport is added to it — and it captures the rest going forward.
|
|
1775
|
+
- A mediasoup transport is always created **on the server first**; only then can the client connect
|
|
1776
|
+
to it, produce/consume, and begin shipping `ClientSample`s. So a peer connection — and the
|
|
1777
|
+
`peer-connection-added` event it triggers — can never appear before its server-side WebRTC
|
|
1778
|
+
transport already exists (and has been observed by the router).
|
|
1779
|
+
|
|
1780
|
+
Put together: by the time a `peer-connection-added` fires, the router has already recorded that
|
|
1781
|
+
transport's id in `webrtcTransportIds`, so a single forward-looking listener catches every match. No
|
|
1782
|
+
back-scan of existing peer connections and no re-check on transport creation are needed — the
|
|
1783
|
+
observer deliberately does **not** look backwards.
|
|
1784
|
+
|
|
1785
|
+
**Your responsibility:** call `createObservedMediasoupRouter(...)` as early as the router exists
|
|
1786
|
+
(before transports are added or samples are accepted). If you register the router *after* its
|
|
1787
|
+
transports are created or after the client's first sample, those events are already in the past and
|
|
1788
|
+
the corresponding matches are missed.
|
|
836
1789
|
|
|
837
1790
|
When the underlying mediasoup router closes, its `close` propagates to `ObservedMediasoupRouter`,
|
|
838
|
-
which emits **`mediasoup-router-removed
|
|
839
|
-
|
|
1791
|
+
which sets the sample's `closedAt` and emits **`mediasoup-router-removed`** — your cue to read /
|
|
1792
|
+
persist the final `observedRouter.sample` and drop your reference to it.
|
|
840
1793
|
|
|
841
1794
|
### Options — `observer.createObservedMediasoupRouter(settings)`
|
|
842
1795
|
|
|
843
1796
|
| Field | Type | Required | Meaning |
|
|
844
1797
|
|-------|------|----------|---------|
|
|
845
|
-
| `router` | `mediasoup.types.Router` | yes | the live router to observe; the observer attaches to `router.observer` |
|
|
846
|
-
| `
|
|
847
|
-
| `
|
|
848
|
-
| `
|
|
849
|
-
| `callId` | `string` | no | **explicit match**: emit `mediasoup-router-matched-with-call` now if this call exists |
|
|
850
|
-
| `bindCallByWebRtcTransportId` | `boolean` | no | **implicit match**: discover the call(s) by correlating WebRTC transport ids with peer-connection ids |
|
|
1798
|
+
| `router` | `mediasoup.types.Router` | yes | the live router to observe; the observer attaches to `router.observer`. `.id` and the sample's `routerId` come from `router.id` |
|
|
1799
|
+
| `appData` | `Record<string, unknown>` | no | application-owned bag on the `ObservedMediasoupRouter` |
|
|
1800
|
+
| `attachments` | `Record<string, unknown>` | no | free-form data; carried on `sample.attachments` |
|
|
1801
|
+
| `matchPeerConnectionByWebRtcTransportId` | `boolean` | no | opt in to peer-connection matching: emit `mediasoup-router-matched-with-peer-connection` for each peer connection whose id matches one of the router's WebRTC transport ids. Omitted / `false` → no matching, the event never fires |
|
|
851
1802
|
|
|
852
|
-
|
|
853
|
-
|
|
1803
|
+
Peer-connection matching is **off by default**; enable it with
|
|
1804
|
+
`matchPeerConnectionByWebRtcTransportId: true`. Returns the `ObservedMediasoupRouter`, or `undefined`
|
|
1805
|
+
if the observer is closed (a router with the same id returns the existing instance — both warn).
|
|
854
1806
|
|
|
855
|
-
Useful members on the returned object: `.sample` (the `MediasoupRouterSample
|
|
856
|
-
|
|
857
|
-
`.close()`.
|
|
1807
|
+
Useful members on the returned object: `.sample` (the in-memory `MediasoupRouterSample`, with
|
|
1808
|
+
`createdAt` / `closedAt?` on it), `.appData`, `.attachments`, `.webrtcTransportIds: Set<string>`,
|
|
1809
|
+
`.id`, `.close()`.
|
|
858
1810
|
|
|
859
1811
|
### Example
|
|
860
1812
|
|
|
861
1813
|
```ts
|
|
862
|
-
import { Observer } from '@observertc/observer-js';
|
|
863
|
-
import type { ObservedMediasoupRouterScope,
|
|
1814
|
+
import { Observer, InMemorySink } from '@observertc/observer-js';
|
|
1815
|
+
import type { ObservedMediasoupRouterScope, ObservedPeerConnectionScope } from '@observertc/observer-js';
|
|
864
1816
|
|
|
865
1817
|
const observer = new Observer();
|
|
866
1818
|
|
|
867
|
-
// 1) Feed client samples as usual so the observer knows about calls & peer connections.
|
|
1819
|
+
// 1) Feed client samples as usual so the observer knows about calls, clients & peer connections.
|
|
868
1820
|
// (e.g. transport-layer: observer.accept(clientSample, context))
|
|
869
1821
|
|
|
870
|
-
// 2) Observe the SFU side
|
|
1822
|
+
// 2) Observe the SFU side; opt in to peer-connection matching. State accumulates in `.sample`.
|
|
871
1823
|
const router = /* your mediasoup router */ undefined as any;
|
|
872
1824
|
const observedRouter = observer.createObservedMediasoupRouter({
|
|
873
1825
|
router,
|
|
874
|
-
|
|
875
|
-
bindCallByWebRtcTransportId: true, // discover the call by peer-connection correlation
|
|
876
|
-
appData: {}, // we'll record the matched callId here
|
|
1826
|
+
matchPeerConnectionByWebRtcTransportId: true,
|
|
877
1827
|
});
|
|
878
1828
|
|
|
879
|
-
//
|
|
880
|
-
|
|
881
|
-
|
|
882
|
-
|
|
883
|
-
|
|
884
|
-
|
|
885
|
-
|
|
1829
|
+
// For large meetings, sample it yourself on your own cadence (see "Memory & large meetings"):
|
|
1830
|
+
// setInterval(() => persist(observedRouter.sample), 10_000);
|
|
1831
|
+
|
|
1832
|
+
// 3) Every peer connection whose id matches one of the router's WebRTC transport ids fires this —
|
|
1833
|
+
// WE decide what to do with each pairing. The payload carries the full ancestry.
|
|
1834
|
+
observer.on('mediasoup-router-matched-with-peer-connection',
|
|
1835
|
+
({ observedMediasoupRouter, observedCall, observedPeerConnection }:
|
|
1836
|
+
ObservedMediasoupRouterScope & ObservedPeerConnectionScope) => {
|
|
1837
|
+
(observedPeerConnection.appData ??= {}).routerId = observedMediasoupRouter.id;
|
|
1838
|
+
myStore.linkRouterToCall(observedCall.callId, observedMediasoupRouter.id);
|
|
886
1839
|
},
|
|
887
1840
|
);
|
|
888
1841
|
|
|
889
|
-
// 4) The router closed —
|
|
1842
|
+
// 4) The router closed — read/persist the final state, then drop your reference.
|
|
890
1843
|
observer.on('mediasoup-router-removed', ({ observedMediasoupRouter }: ObservedMediasoupRouterScope) => {
|
|
891
|
-
|
|
892
|
-
myStore.saveRouterSample(callId, observedMediasoupRouter.sample);
|
|
893
|
-
});
|
|
894
|
-
|
|
895
|
-
// (optional) react to the router being registered at all:
|
|
896
|
-
observer.on('mediasoup-router-added', ({ observedMediasoupRouter }) => {
|
|
897
|
-
console.log('observing router', observedMediasoupRouter.id);
|
|
1844
|
+
persist(observedMediasoupRouter.sample); // its `closedAt` is set
|
|
898
1845
|
});
|
|
899
1846
|
```
|
|
900
1847
|
|
|
901
|
-
|
|
902
|
-
|
|
903
|
-
```ts
|
|
904
|
-
observer.createObservedMediasoupRouter({ router, routerId: router.id, callId });
|
|
905
|
-
// → `mediasoup-router-matched-with-call` fires immediately if `callId` is a known call
|
|
906
|
-
```
|
|
907
|
-
|
|
908
|
-
### Why event-driven instead of storing on the call
|
|
1848
|
+
### Why event-driven matching instead of storing on the call
|
|
909
1849
|
|
|
910
1850
|
- **Loose coupling.** The call model stays about client telemetry; the SFU view lives on its own
|
|
911
|
-
|
|
912
|
-
- **You own the association.** One router
|
|
913
|
-
|
|
914
|
-
|
|
915
|
-
- **
|
|
916
|
-
|
|
1851
|
+
`ObservedMediasoupRouter` and is associated only if and how *you* choose.
|
|
1852
|
+
- **You own the association.** One router serves many peer connections (across clients and calls),
|
|
1853
|
+
and the right place to keep that mapping is application-specific — so the observer hands you each
|
|
1854
|
+
peer-connection match and gets out of the way.
|
|
1855
|
+
- **You own the sampling.** The router sample is plain in-memory state you read on your own terms;
|
|
1856
|
+
for large meetings, sample/persist it yourself (see [Memory & large meetings](#memory--large-meetings))
|
|
1857
|
+
rather than relying on the library to evict — it deliberately doesn't.
|
|
917
1858
|
|
|
918
1859
|
---
|
|
919
1860
|
|
|
@@ -1039,6 +1980,60 @@ const observer = new Observer({ createClientSink });
|
|
|
1039
1980
|
the bus with full ancestry. `ClientSampleSinkFactory` is
|
|
1040
1981
|
`(p: { clientId: string; observedCall: ObservedCall }) => ClientSampleSink | undefined`.
|
|
1041
1982
|
|
|
1983
|
+
---
|
|
1984
|
+
|
|
1985
|
+
## Injecting data into a client
|
|
1986
|
+
|
|
1987
|
+
Sometimes the application holds data that belongs on a client's record but isn't part of the
|
|
1988
|
+
client-reported `ClientSample` — a room id or display name, an application-level event
|
|
1989
|
+
(*"recording started"*), a server-detected issue, an extension stat, or a device/meta item.
|
|
1990
|
+
`ObservedClient` exposes **injection** methods that merge such data into the client's sample stream,
|
|
1991
|
+
so it updates the live model **and** is persisted to the client's
|
|
1992
|
+
[sink](#sinks-per-client-sample-persistence) exactly like sampled data.
|
|
1993
|
+
|
|
1994
|
+
| Method | Adds to the sample's | Surfaces as |
|
|
1995
|
+
|--------|----------------------|-------------|
|
|
1996
|
+
| `injectAttachment(attachments)` | `attachments` (merged via `Object.assign`) | `observedClient.attachments` |
|
|
1997
|
+
| `injectEvent(event: ClientEvent)` | `clientEvents` | `client-event` (plus any state the event drives) |
|
|
1998
|
+
| `injectIssue(issue: ClientIssue)` | `clientIssues` | `client-issue` |
|
|
1999
|
+
| `injectMetaData(meta: ClientMetaData)` | `clientMetaItems` | `client-metadata` |
|
|
2000
|
+
| `injectExtensionStat(stat: ExtensionStat)` | `extensionStats` | `client-extension-stats` |
|
|
2001
|
+
|
|
2002
|
+
### When the injected data lands
|
|
2003
|
+
|
|
2004
|
+
Injection is timing-aware so nothing is dropped, regardless of *when* you call it:
|
|
2005
|
+
|
|
2006
|
+
- **During a sample's processing** — e.g. from inside a `client-updated` / `client-event` handler,
|
|
2007
|
+
which run within `accept()` — the data is applied to the **current** sample immediately: reflected
|
|
2008
|
+
in entity state and written to the sink as part of that sample.
|
|
2009
|
+
- **Between samples** — the data is buffered and merged into the **next** `accept()`'s sample.
|
|
2010
|
+
- **On `close()` with pending injections and no further sample** — the buffer is flushed as a final
|
|
2011
|
+
synthetic sample (applied to state and written to the sink) before the sink is ended, so a
|
|
2012
|
+
last-moment injection is never lost.
|
|
2013
|
+
|
|
2014
|
+
In every case the injected data both updates the live `ObservedClient` and reaches the per-client
|
|
2015
|
+
sink — the sink always receives the final, **injection-merged** sample (the sink write happens at the
|
|
2016
|
+
end of `accept()`, after the merge).
|
|
2017
|
+
|
|
2018
|
+
### Example
|
|
2019
|
+
|
|
2020
|
+
```ts
|
|
2021
|
+
// Enrich at creation from your app's knowledge of the participant. Injecting in `client-added`
|
|
2022
|
+
// (which runs just before the first accept) lands on the first sample.
|
|
2023
|
+
observer.on('client-added', ({ observedClient }) => {
|
|
2024
|
+
observedClient.injectAttachment({ roomId: lookupRoomId(observedClient.clientId) });
|
|
2025
|
+
});
|
|
2026
|
+
|
|
2027
|
+
// Application-level signals at any time:
|
|
2028
|
+
const client = observer.getObservedCall(callId)?.getObservedClient(clientId);
|
|
2029
|
+
client?.injectEvent({ type: 'RECORDING_STARTED', timestamp: Date.now() });
|
|
2030
|
+
client?.injectIssue({ type: 'app-kicked-participant', timestamp: Date.now() });
|
|
2031
|
+
```
|
|
2032
|
+
|
|
2033
|
+
`attachments` are latest-wins (like sampled `attachments`): injecting a key overwrites its previous
|
|
2034
|
+
value. `appData` is unaffected — injections flow into the sample/telemetry, not the app-owned
|
|
2035
|
+
`appData` bag (see [Ingestion](#ingestion-accept-context--lifecycle)).
|
|
2036
|
+
|
|
1042
2037
|
## Logging
|
|
1043
2038
|
|
|
1044
2039
|
`observer-js` logs through a single, swappable sink. Out of the box it writes `debug` and
|
|
@@ -1063,6 +2058,16 @@ filtering, per-module routing, and full silencing.
|
|
|
1063
2058
|
|
|
1064
2059
|
---
|
|
1065
2060
|
|
|
2061
|
+
## Design notes
|
|
2062
|
+
|
|
2063
|
+
**[`docs/design-notes.md`](./docs/design-notes.md)** covers the reasoning behind the library rather
|
|
2064
|
+
than its API: why client-detectable conditions are never re-derived server-side, why each shipped
|
|
2065
|
+
detector exists, what was deliberately *not* built and why, the WebRTC domain facts that shaped the
|
|
2066
|
+
implementation (ICE-Lite disconnect waves, counter resets, why ICE and RTCP RTT must never be
|
|
2067
|
+
blended), and an operational threshold reference.
|
|
2068
|
+
|
|
2069
|
+
---
|
|
2070
|
+
|
|
1066
2071
|
## Error-handling philosophy
|
|
1067
2072
|
|
|
1068
2073
|
The library **warns and degrades; it does not throw** on operational problems:
|
|
@@ -1096,10 +2101,13 @@ lint + typecheck + **build** + test on every push/PR.
|
|
|
1096
2101
|
|
|
1097
2102
|
**Project layout** (`src/`): `Observer.ts`, `ObservedCall.ts`, `ObservedClient.ts`,
|
|
1098
2103
|
`ObservedPeerConnection.ts`, the `Observed*` sub-stat classes, `ObserverEvents.ts` (the typed
|
|
1099
|
-
event map + scope types), `detectors/` (`Detector`, `Detectors
|
|
1100
|
-
|
|
1101
|
-
`
|
|
1102
|
-
`
|
|
2104
|
+
event map + scope types), `detectors/` (`Detector`, `Detectors`, and one file per detector),
|
|
2105
|
+
`validators/` (`Validator`, `Validators`, one file per validator), `issues/` (`ActiveClientIssue`,
|
|
2106
|
+
`ActiveIssueTracker`, `ActiveIssuesRegistry`, `ObservedClientIssueRegistry`), `scores/`,
|
|
2107
|
+
`resolvers/` (remote-track resolvers), `utils/` (`stats`, `SlidingWindow`, `TrendTester`,
|
|
2108
|
+
`CallHealthAggregator`), `common/` (`logger`, `utils`, `Middleware`), `schema/`
|
|
2109
|
+
(sample/event/meta types), and `sinks/` (the `ClientSampleSink` base + `JsonlFileSink` /
|
|
2110
|
+
`InMemorySink`, re-exported from the package root).
|
|
1103
2111
|
|
|
1104
2112
|
**Conventions to follow when developing further:**
|
|
1105
2113
|
|