playlist-data-engine 1.7.2 → 1.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. package/README.md +20 -3
  2. package/bin/cli.cjs +85 -0
  3. package/dist/core/parser/TrackExtras.d.ts +150 -1
  4. package/dist/core/parser/TrackExtras.d.ts.map +1 -1
  5. package/dist/gateway-CDMPqFEH.js +1320 -0
  6. package/dist/gateway-DKa45Uz6.cjs +6 -0
  7. package/dist/gateway.d.ts +3 -1
  8. package/dist/gateway.d.ts.map +1 -1
  9. package/dist/gateway.js +1 -1
  10. package/dist/gateway.mjs +22 -14
  11. package/dist/index.d.ts +4 -3
  12. package/dist/index.d.ts.map +1 -1
  13. package/dist/playlist-data-engine.js +34 -34
  14. package/dist/playlist-data-engine.mjs +964 -1240
  15. package/dist/utils/engineDocs.d.ts +33 -0
  16. package/dist/utils/engineDocs.d.ts.map +1 -0
  17. package/dist/utils/playlistUtils.d.ts +37 -0
  18. package/dist/utils/playlistUtils.d.ts.map +1 -1
  19. package/dist/utils/validators.d.ts +27 -0
  20. package/dist/utils/validators.d.ts.map +1 -1
  21. package/docs/DATA_ENGINE_REFERENCE.md +6660 -0
  22. package/docs/USAGE_IN_OTHER_PROJECTS.md +587 -0
  23. package/docs/features/AUDIO_ANALYSIS.md +610 -0
  24. package/docs/features/BEAT_DETECTION.md +5250 -0
  25. package/docs/features/COMBAT_SYSTEM.md +1632 -0
  26. package/docs/features/CONTENT_PACKS.md +464 -0
  27. package/docs/features/CUSTOM_CONTENT.md +603 -0
  28. package/docs/features/ENEMY_GENERATION.md +1711 -0
  29. package/docs/features/EQUIPMENT_SYSTEM.md +2279 -0
  30. package/docs/features/EXTENSIBILITY_GUIDE.md +1106 -0
  31. package/docs/features/GATEWAY_RESOLUTION.md +725 -0
  32. package/docs/features/IRL_SENSORS.md +360 -0
  33. package/docs/features/PLAYLIST_PARSING.md +446 -0
  34. package/docs/features/PREREQUISITES.md +571 -0
  35. package/docs/features/ROLLS_AND_SEEDS.md +687 -0
  36. package/docs/features/XP_AND_STATS.md +1221 -0
  37. package/llms.txt +33 -0
  38. package/package.json +9 -2
  39. package/skills/playlist-data-engine/SKILL.md +69 -0
  40. package/dist/gateway-DUk4nCao.cjs +0 -1
  41. package/dist/gateway-DyR4M-uH.js +0 -681
@@ -0,0 +1,610 @@
1
+ # Audio Analysis Documentation
2
+
3
+ The Playlist Data Engine provides audio analysis modes for extracting meaningful data from music files. Each mode serves different use cases, from quick character generation to timeline visualization.
4
+
5
+ ---
6
+
7
+ ## Overview
8
+
9
+ The engine's audio analysis is powered by the Web Audio API and provides two distinct modes for general audio analysis:
10
+
11
+ | Mode | Method | Purpose | Use Case |
12
+ |------|--------|---------|----------|
13
+ | **Triple Tap Real-Time** | `extractSonicFingerprint()` | Quick analysis at key positions | Character generation, quick profiling |
14
+ | **Full Song Timeline** | `analyzeTimeline()` | Complete track analysis | Waveform visualization, level generation |
15
+ | **Music Classification** | `analyze()` | Deep ML classification | Genre, mood, and vibe detection |
16
+ | **Pitch Analysis** | `analyze()` | Full-track pitch detection | Melody analysis, note detection |
17
+
18
+ > **Note**: For rhythm game features like beat detection, beat streaming, and chart creation, see [BEAT_DETECTION.md](BEAT_DETECTION.md).
19
+
20
+ ### Source Files
21
+
22
+ | Component | Location |
23
+ |-----------|----------|
24
+ | **AudioAnalyzer** (main class) | [src/core/analysis/AudioAnalyzer.ts](../src/core/analysis/AudioAnalyzer.ts) |
25
+ | **MusicClassifier** (ML classification) | [src/core/analysis/MusicClassifier.ts](../src/core/analysis/MusicClassifier.ts) |
26
+ | **PitchAnalyzer** (pitch detection) | [src/core/analysis/PitchAnalyzer.ts](../src/core/analysis/PitchAnalyzer.ts) |
27
+ | **SpectrumScanner** (frequency bands) | [src/core/analysis/SpectrumScanner.ts](../src/core/analysis/SpectrumScanner.ts) |
28
+ | **Audio Types** | [src/core/types/AudioProfile.ts](../src/core/types/AudioProfile.ts) |
29
+
30
+ ---
31
+
32
+ ## ⚠️ Import path — TensorFlow.js
33
+
34
+ The audio analysis surface (`AudioAnalyzer`, `MusicClassifier`, `EssentiaPitchDetector`,
35
+ `PitchAnalyzer`, and the level-generation classes) depends on **`@tensorflow/tfjs`**
36
+ (~14 MB). These symbols are exported from the **`playlist-data-engine/analysis`**
37
+ subpath, NOT the default entry. Importing them from the bare `playlist-data-engine`
38
+ will fail and/or pull TensorFlow into your main bundle.
39
+
40
+ ```ts
41
+ // ✅ correct — TF-bearing symbols come from /analysis
42
+ import { AudioAnalyzer, MusicClassifier, PitchAnalyzer } from 'playlist-data-engine/analysis';
43
+
44
+ // ❌ wrong — these are NOT on the default (TF-free) entry
45
+ import { AudioAnalyzer } from 'playlist-data-engine';
46
+ ```
47
+
48
+ The default `playlist-data-engine` entry is kept TensorFlow-free on purpose so
49
+ apps that only need parsing/generation/gateway utilities don't pay the TF cost.
50
+ For web apps, run analysis in a **Web Worker** so the TF runtime loads on a
51
+ worker thread. (See `src/features/directoryTools/audioAnalysis/analyzeWorker.ts`
52
+ in ApeTapes for a reference worker implementation.)
53
+
54
+ ---
55
+
56
+ ## 3-Tap Real-Time Analysis
57
+
58
+ The original `AudioAnalyzer` real-time analysis uses the "Triple Tap" strategy: analyzing three key positions (5%, 40%, 70%) in tracks longer than 3 seconds, or the full buffer for shorter clips.
59
+
60
+ ### Method
61
+
62
+ ```typescript
63
+ extractSonicFingerprint(audioUrl: string): Promise<AudioProfile>
64
+ ```
65
+
66
+ ### Usage
67
+
68
+ ```typescript
69
+ import { AudioAnalyzer } from 'playlist-data-engine/analysis';
70
+
71
+ const analyzer = new AudioAnalyzer({
72
+ includeAdvancedMetrics: true, // Include spectral centroid, rolloff, zero-crossing rate
73
+ trebleBoost: 0.6, // Reduce treble dominance
74
+ bassBoost: 1.2, // Increase bass presence
75
+ });
76
+
77
+ const profile = await analyzer.extractSonicFingerprint(track.audio_url);
78
+
79
+ console.log(`Bass: ${profile.bass_dominance}`);
80
+ console.log(`Mid: ${profile.mid_dominance}`);
81
+ console.log(`Treble: ${profile.treble_dominance}`);
82
+ console.log(`RMS Energy: ${profile.rms_energy}`);
83
+ console.log(`Dynamic Range: ${profile.dynamic_range}`);
84
+ ```
85
+
86
+ ### AudioProfile Output
87
+
88
+ | Property | Type | Description |
89
+ |----------|------|-------------|
90
+ | `bass_dominance` | `number` | Relative bass energy (0-1, normalized with others) |
91
+ | `mid_dominance` | `number` | Relative mid-range energy (0-1) |
92
+ | `treble_dominance` | `number` | Relative treble energy (0-1) |
93
+ | `average_amplitude` | `number` | Average amplitude across all samples |
94
+ | `rms_energy` | `number` | Root mean square energy (perceived loudness) |
95
+ | `dynamic_range` | `number` | Peak amplitude minus RMS energy |
96
+ | `spectral_centroid` | `number?` | Brightness indicator (with advanced metrics) |
97
+ | `spectral_rolloff` | `number?` | Frequency below which 85% of energy is contained |
98
+ | `zero_crossing_rate` | `number?` | Measure of noisiness/percussiveness |
99
+ | `analysis_metadata` | `object` | Duration, sample positions, timestamp |
100
+
101
+ ### Frequency Bands
102
+
103
+ The analyzer separates audio into three perceptual frequency bands:
104
+
105
+ | Band | Frequency Range | Typical Content |
106
+ |------|-----------------|-----------------|
107
+ | **Bass** | 20 - 400 Hz | Kick drums, bass guitar, sub-bass |
108
+ | **Mid** | 400 - 4000 Hz | Vocals, guitars, keyboards |
109
+ | **Treble** | 4000 - 14000 Hz | Hi-hats, cymbals, high harmonics |
110
+
111
+ ### Triple Tap Strategy
112
+
113
+ For tracks longer than 3 seconds, the analyzer samples at three positions:
114
+
115
+ | Position | Rationale |
116
+ |----------|-----------|
117
+ | **5%** | Capture intro, often distinctive |
118
+ | **40%** | Typically in the main body/chorus |
119
+ | **70%** | Late section, often bridge or climax |
120
+
121
+ This provides representative coverage while avoiding expensive full-track analysis.
122
+
123
+ ---
124
+
125
+ ## Full Song Analysis
126
+
127
+ For applications requiring complete track data (waveform visualization, timeline displays, level generation), use `analyzeTimeline()`.
128
+
129
+ ### Method
130
+
131
+ ```typescript
132
+ analyzeTimeline(audioUrl: string, strategy: SamplingStrategy): Promise<AudioTimelineEvent[]>
133
+ ```
134
+
135
+ ### Sampling Strategies
136
+
137
+ ```typescript
138
+ // Option A: Sample every N seconds
139
+ const timeline = await analyzer.analyzeTimeline(audioUrl, {
140
+ type: 'interval',
141
+ intervalSeconds: 2 // Sample every 2 seconds
142
+ });
143
+
144
+ // Option B: Generate exactly N data points
145
+ const timeline = await analyzer.analyzeTimeline(audioUrl, {
146
+ type: 'count',
147
+ count: 100 // Exactly 100 evenly-spaced samples
148
+ });
149
+ ```
150
+
151
+ ### AudioTimelineEvent Output
152
+
153
+ | Property | Type | Description |
154
+ |----------|------|-------------|
155
+ | `timestamp` | `number` | Position in the track (seconds) |
156
+ | `duration` | `number` | Length of analyzed segment |
157
+ | `bass` | `number` | Bass dominance (0-1, normalized) |
158
+ | `mid` | `number` | Mid dominance (0-1, normalized) |
159
+ | `treble` | `number` | Treble dominance (0-1, normalized) |
160
+ | `amplitude` | `number` | RMS energy for this segment |
161
+ | `rms_energy` | `number` | Root mean square energy |
162
+ | `peak` | `number` | Peak amplitude |
163
+ | `dynamic_range` | `number` | Peak minus RMS |
164
+ | `spectral_centroid` | `number` | Brightness indicator |
165
+ | `spectral_rolloff` | `number` | Energy distribution measure |
166
+ | `zero_crossing_rate` | `number` | Noisiness/percussiveness |
167
+
168
+ ### Usage Example
169
+
170
+ ```typescript
171
+ import { AudioAnalyzer } from 'playlist-data-engine/analysis';
172
+
173
+ const analyzer = new AudioAnalyzer();
174
+
175
+ // Generate 100 data points across the song for visualization
176
+ const timeline = await analyzer.analyzeTimeline(audioUrl, {
177
+ type: 'count',
178
+ count: 100
179
+ });
180
+
181
+ // Find the loudest moment
182
+ const peakMoment = timeline.reduce((max, event) =>
183
+ event.rms_energy > max.rms_energy ? event : max
184
+ );
185
+ console.log(`Peak at ${peakMoment.timestamp}s`);
186
+
187
+ // Build a simple waveform visualization
188
+ timeline.forEach(event => {
189
+ const barHeight = Math.round(event.rms_energy * 20);
190
+ console.log(`${'█'.repeat(barHeight)} [${event.timestamp.toFixed(1)}s]`);
191
+ });
192
+ ```
193
+
194
+
195
+ ---
196
+
197
+ ## Music Classification (Genre, Mood, Vibe)
198
+
199
+ For deep semantic analysis of music including mood, themes, and vibe metrics (like danceability), use the `MusicClassifier`. This uses multiple `essentia.js` models.
200
+
201
+ ### Method
202
+
203
+ ```typescript
204
+ analyze(audioUrl: string): Promise<MusicClassificationProfile>
205
+ ```
206
+
207
+ ### Usage Example
208
+
209
+ ```typescript
210
+ import { MusicClassifier } from 'playlist-data-engine/analysis';
211
+
212
+ const classifier = new MusicClassifier({
213
+ topN: 3, // Return top 3 genres/moods
214
+ threshold: 0.1, // 10% confidence threshold
215
+ analysisDurationSeconds: 30, // Analyze 30 seconds instead of full song
216
+ analysisStartPosition: 0.5 // Start from middle of song
217
+ });
218
+
219
+ const profile = await classifier.analyze('https://example.com/audio.mp3');
220
+
221
+ console.log(`Genre: ${profile.primary_genre}`);
222
+ console.log(`Moods: ${profile.mood_tags.join(', ')}`);
223
+ console.log(`Danceability: ${profile.vibe_metrics.danceability}`);
224
+ console.log(`Energy: ${profile.vibe_metrics.energy}`);
225
+ ```
226
+
227
+ ### MusicClassificationProfile Output
228
+
229
+ | Property | Type | Description |
230
+ |----------|------|-------------|
231
+ | `genres` | `ClassificationTag[]` | Top matched genres |
232
+ | `moods` | `ClassificationTag[]` | Top matched moods/themes |
233
+ | `primary_genre` | `string` | Highest confidence genre |
234
+ | `mood_tags` | `string[]` | Most relevant mood keywords |
235
+ | `vibe_metrics` | `VibeMetrics` | Danceability, energy, valence, etc. |
236
+
237
+ ### Partial Song Analysis
238
+
239
+ By default, `analyze()` processes the entire audio signal. For faster classification (genre, mood, vibe), you can analyze only a segment of the song by setting `analysisDurationSeconds` and `analysisStartPosition` in the constructor options.
240
+
241
+ | Option | Type | Default | Description |
242
+ |--------|------|---------|-------------|
243
+ | `analysisDurationSeconds` | number | — | Seconds of audio to analyze. Full song when not set. Recommended: 30s |
244
+ | `analysisStartPosition` | number | `0.5` | Position to start (0.0 = start, 0.5 = middle, 1.0 = end). Only applies when duration is set |
245
+
246
+ ```typescript
247
+ // Analyze middle 30 seconds of each song (faster, good accuracy for genre/mood)
248
+ const classifier = new MusicClassifier({
249
+ analysisDurationSeconds: 30,
250
+ analysisStartPosition: 0.5
251
+ });
252
+
253
+ // Analyze first 15 seconds (fastest, lower accuracy)
254
+ const quickClassifier = new MusicClassifier({
255
+ analysisDurationSeconds: 15,
256
+ analysisStartPosition: 0
257
+ });
258
+ ```
259
+
260
+ > **Note**: The audio file is still fully downloaded and decoded (required by the Web Audio API). Only the ML feature extraction and model inference operate on the segment, which is where the speed improvement comes from. A 30-second segment processes ~8x faster than a 4-minute song.
261
+
262
+ ---
263
+
264
+ ### Two-Step Model Architecture
265
+
266
+ The `MusicClassifier` supports both single-step (one model) and two-step (embedding + classifier) architectures. This enables using state-of-the-art models like Discogs-EffNet for embeddings combined with specialized classifier heads.
267
+
268
+ #### Architecture Compatibility Table
269
+
270
+ Different model architectures require different mel-band configurations for feature extraction:
271
+
272
+ | Architecture | Mel Bands | Extractor | Compatible Models |
273
+ |--------------|-----------|-----------|-------------------|
274
+ | `musicnn` | 96 | Essentia musicnn | MusiCNN classifiers, MSD models |
275
+ | `effnet` | 128 | Custom | Discogs-EffNet embeddings |
276
+ | `vggish` | 64 | Essentia vggish | VGGish classifiers, AudioSet |
277
+ | `tempocnn` | 40 | Essentia tempocnn | TempoCNN tempo estimation |
278
+
279
+ > **Important**: The engine automatically detects the architecture from the model URL and uses the correct mel-band configuration. No manual configuration needed!
280
+
281
+ #### Model Configuration Formats
282
+
283
+ Every model option (`genre`, `mood`, `danceability`, `voice`, `acoustic`) accepts two formats:
284
+
285
+ ##### Format 1: Single-Step Model Config (with explicit type)
286
+
287
+ For URLs where architecture cannot be detected (e.g., Arweave URLs), use `SingleStepModelConfig`:
288
+
289
+ ```typescript
290
+ import { MusicClassifier, type SingleStepModelConfig } from 'playlist-data-engine/analysis';
291
+
292
+ const config: SingleStepModelConfig = {
293
+ modelUrl: 'https://arweave.net/xxx/model.json',
294
+ modelType: 'musicnn', // Explicitly specify architecture
295
+ genreType: 'jamendo', // Explicitly specify genre list (for genre models)
296
+ labels: ['custom', 'labels'] // Optional custom labels
297
+ };
298
+ ```
299
+
300
+ | Property | Type | Required | Description |
301
+ |----------|------|----------|-------------|
302
+ | `modelUrl` | `string` | Yes | URL to the model file |
303
+ | `modelType` | `ModelArchitecture` | Yes | Explicit architecture: `'musicnn'` \| `'effnet'` \| `'vggish'` \| `'tempocnn'` |
304
+ | `genreType` | `GenreListType` | No* | Explicit genre list: `'jamendo'` \| `'discogs400'` \| `'tzanetakis'` \| `'mtt_musicnn'` (*required for genre models with Arweave URLs) |
305
+ | `labels` | `string[]` | No | Custom output labels |
306
+
307
+ ##### Format 2: Two-Step Model Config (with explicit types)
308
+
309
+ Separate embedding and classifier models with optional explicit type parameters:
310
+
311
+ ```typescript
312
+ import { MusicClassifier, type TwoStepModelConfig } from 'playlist-data-engine/analysis';
313
+
314
+ const config: TwoStepModelConfig = {
315
+ embedding: '/models/discogs-effnet-bs64-1.json',
316
+ classifier: '/models/mtg_jamendo_genre-discogs-effnet-1.json',
317
+ // Optional explicit types (override URL detection)
318
+ embeddingType: 'effnet', // ModelArchitecture
319
+ classifierType: 'discogs400' // GenreListType (for genre models)
320
+ };
321
+ ```
322
+
323
+ | Property | Type | Required | Description |
324
+ |----------|------|----------|-------------|
325
+ | `embedding` | `string` | Yes | URL to the embedding model |
326
+ | `classifier` | `string` | Yes | URL to the classifier model |
327
+ | `labels` | `string[]` | No | Custom output labels |
328
+ | `embeddingType` | `ModelArchitecture` | No | Explicit embedding type: `'musicnn'` \| `'effnet'` \| `'vggish'` \| `'tempocnn'` |
329
+ | `classifierType` | `GenreListType` | No | Explicit genre list type: `'jamendo'` \| `'discogs400'` \| `'tzanetakis'` \| `'mtt_musicnn'` |
330
+
331
+ ##### Why Use Explicit Type Parameters?
332
+
333
+ URL-based detection works for conventional file paths like `/models/effnet-classifier.json`, but fails for:
334
+ - **Arweave URLs**: `https://arweave.net/xxx/model.json` contains no architecture hints
335
+ - **Custom hosting**: URLs without descriptive filenames
336
+ - **Proxied URLs**: Gateway URLs that obscure the original filename
337
+
338
+ #### Signal Flow Diagrams
339
+
340
+ **Single-Step Flow:**
341
+ ```
342
+ ┌─────────────────┐ ┌──────────────────┐ ┌─────────────────┐
343
+ │ Audio Signal │ ──▶ │ Feature │ ──▶ │ Single Model │
344
+ │ (16kHz mono) │ │ Extractor │ │ (classifier) │
345
+ └─────────────────┘ │ (96 bands) │ └────────┬────────┘
346
+ └──────────────────┘ │
347
+ ▼
348
+ ┌─────────────────┐
349
+ │ Class Labels │
350
+ │ (genre, mood) │
351
+ └─────────────────┘
352
+ ```
353
+
354
+ **Two-Step Flow:**
355
+ ```
356
+ ┌─────────────────┐ ┌──────────────────┐ ┌─────────────────┐
357
+ │ Audio Signal │ ──▶ │ Feature │ ──▶ │ Embedding │
358
+ │ (16kHz mono) │ │ Extractor │ │ Model │
359
+ └─────────────────┘ │ (128 bands) │ │ (1280-dim) │
360
+ └──────────────────┘ └────────┬────────┘
361
+ │
362
+ Architecture-specific │
363
+ mel-band config ▼
364
+ ┌─────────────────┐
365
+ │ Classifier │
366
+ │ Model │
367
+ │ (class probs) │
368
+ └────────┬────────┘
369
+ │
370
+ ▼
371
+ ┌─────────────────┐
372
+ │ Class Labels │
373
+ │ (genre, mood) │
374
+ └─────────────────┘
375
+ ```
376
+
377
+ #### Configuration Examples
378
+
379
+ **Mixed Configuration (Single + Two-Step):**
380
+
381
+ ```typescript
382
+ const classifier = new MusicClassifier({
383
+ models: {
384
+ // Two-step: embedding + classifier (uses 128-band extractor)
385
+ genre: {
386
+ embedding: '/models/discogs-effnet-bs64-1.json',
387
+ classifier: '/models/mtg_jamendo_genre-discogs-effnet-1.json'
388
+ },
389
+ // Two-step: same embedding cached, different classifier
390
+ mood: {
391
+ embedding: '/models/discogs-effnet-bs64-1.json',
392
+ classifier: '/models/mtg_jamendo_moodtheme-discogs-effnet-1.json'
393
+ },
394
+ // Single-step: requires modelUrl + modelType (uses 64-band vggish extractor)
395
+ danceability: {
396
+ modelUrl: '/models/danceability-vggish-audioset-1.json',
397
+ modelType: 'vggish'
398
+ },
399
+ // Single-step: optional voice detection
400
+ voice: {
401
+ modelUrl: '/models/voice-detector.json',
402
+ modelType: 'musicnn'
403
+ },
404
+ // Two-step: optional acoustic detection
405
+ acoustic: {
406
+ embedding: '/models/discogs-effnet-bs64-1.json',
407
+ classifier: '/models/acoustic-classifier.json'
408
+ }
409
+ }
410
+ });
411
+ ```
412
+
413
+ #### Using Arweave-Hosted Models (Zero Setup)
414
+
415
+ The default configuration uses pre-configured Arweave-hosted models:
416
+
417
+ ```typescript
418
+ import { MusicClassifier } from 'playlist-data-engine/analysis';
419
+
420
+ // Zero setup - models load from Arweave automatically
421
+ const classifier = new MusicClassifier();
422
+
423
+ const profile = await classifier.analyze('https://example.com/track.mp3');
424
+ ```
425
+
426
+ | Model | Architecture | Labels | Source |
427
+ |-------|-------------|--------|--------|
428
+ | **Genre** | Two-step (effnet + classifier) | discogs400 (400+ subgenres) | Arweave |
429
+ | **Mood** | Two-step (effnet + classifier) | JAMENDO_MOODS (60 themes) | Arweave |
430
+ | **Danceability** | Single-step (musicnn) | Binary | Turbo Gateway |
431
+
432
+ #### Using Presets
433
+
434
+ Instead of raw URLs, use preset names to select pre-configured models. This is the simplest way to swap genre/mood/danceability models without managing URLs.
435
+
436
+ ```typescript
437
+ import { MusicClassifier } from 'playlist-data-engine/analysis';
438
+
439
+ // Use presets for genre and mood
440
+ const classifier = new MusicClassifier({
441
+ preset: { genre: 'jamendo', mood: 'jamendo' }
442
+ });
443
+
444
+ // Mix presets with custom URLs — explicit models take precedence
445
+ const classifier = new MusicClassifier({
446
+ preset: { genre: 'tzanetakis' },
447
+ models: {
448
+ mood: { modelUrl: '/models/custom-mood.json', modelType: 'musicnn' }
449
+ }
450
+ });
451
+ ```
452
+
453
+ **Available presets:**
454
+
455
+ | Category | Preset | Architecture | Labels |
456
+ |----------|--------|--------------|--------|
457
+ | Genre | `discogs400` | Two-step (effnet + discogs400) | 400+ subgenres |
458
+ | Genre | `jamendo` | Two-step (effnet + jamendo) | MTG Jamendo 87 |
459
+ | Genre | `tzanetakis` | Single-step (musicnn) | GTZAN 10 |
460
+ | Genre | `musicnn` | Single-step (musicnn) | MagnaTagATune 50 |
461
+ | Mood | `jamendo` | Two-step (effnet + jamendo) | 60 themes |
462
+ | Mood | `happyMusicnn` | Single-step (musicnn) | Binary (happy/not happy) |
463
+ | Danceability | `default` | Single-step (musicnn) | Binary |
464
+
465
+ To enumerate available presets at runtime:
466
+
467
+ ```typescript
468
+ import { AVAILABLE_PRESETS } from 'playlist-data-engine/analysis';
469
+ console.log(AVAILABLE_PRESETS.genre); // ['discogs400', 'jamendo', 'tzanetakis', 'musicnn']
470
+ console.log(AVAILABLE_PRESETS.mood); // ['jamendo', 'happyMusicnn']
471
+ console.log(AVAILABLE_PRESETS.danceability); // ['default']
472
+ ```
473
+
474
+ **Partial Override & Custom Arweave Models:**
475
+
476
+ ```typescript
477
+ import { MusicClassifier, DEFAULT_ARWEAVE_MODELS } from 'playlist-data-engine/analysis';
478
+
479
+ const classifier = new MusicClassifier({
480
+ models: {
481
+ // Use default Arweave model
482
+ genre: DEFAULT_ARWEAVE_MODELS.genre,
483
+ // Override with local model
484
+ mood: {
485
+ modelUrl: '/models/mood-musicnn-msd-1.json',
486
+ modelType: 'musicnn'
487
+ },
488
+ // Or use custom Arweave URLs with explicit types
489
+ danceability: {
490
+ modelUrl: 'https://arweave.net/YOUR_TXID/model.json',
491
+ modelType: 'musicnn' // Required for Arweave URLs
492
+ }
493
+ }
494
+ });
495
+ ```
496
+
497
+ > **Note:** Single-step mood models use binary labels (`'happy'`, `'not happy'`) while two-step mood models use the full JAMENDO_MOODS taxonomy (60 mood/theme labels).
498
+
499
+ #### Embedding Model Caching
500
+
501
+ When using two-step models with shared embeddings (e.g., same Discogs-EffNet for genre and mood), the embedding model is automatically cached and reused:
502
+
503
+ ```typescript
504
+ const classifier = new MusicClassifier({
505
+ cacheEmbeddings: true, // Default: true
506
+ models: {
507
+ genre: {
508
+ embedding: '/models/discogs-effnet-bs64-1.json', // Loaded once
509
+ classifier: '/models/mtg_jamendo_genre-discogs-effnet-1.json'
510
+ },
511
+ mood: {
512
+ embedding: '/models/discogs-effnet-bs64-1.json', // Reused from cache!
513
+ classifier: '/models/mtg_jamendo_moodtheme-discogs-effnet-1.json'
514
+ }
515
+ }
516
+ });
517
+
518
+ // Later, to free memory:
519
+ classifier.clearEmbeddingCache(); // Clear embedding models only
520
+ classifier.clearClassifierCache(); // Clear classifier models only
521
+ classifier.clearAllCaches(); // Clear everything
522
+ ```
523
+
524
+ #### Metadata Tracking
525
+
526
+ The `analysis_metadata.models_used` field tracks which models were used:
527
+
528
+ ```typescript
529
+ const profile = await classifier.analyze(audioUrl);
530
+
531
+ console.log(profile.analysis_metadata.models_used);
532
+ // Two-step: ['/models/discogs-effnet-bs64-1.json -> /models/mtg_jamendo_genre-discogs-effnet-1.json', ...]
533
+ // Single-step: ['/models/genre-musicnn-msd-1.json', ...]
534
+ ```
535
+
536
+ ---
537
+
538
+ ## Pitch Analysis
539
+
540
+ For full-track pitch detection without beat-map dependency, use the `PitchAnalyzer`. This provides per-frame pitch results, melody contour analysis, and summary statistics directly from raw audio.
541
+
542
+ ### Method
543
+
544
+ ```typescript
545
+ analyze(audioUrl: string): Promise<PitchAnalysisProfile>
546
+ ```
547
+
548
+ ### Usage Example
549
+
550
+ ```typescript
551
+ import { PitchAnalyzer } from 'playlist-data-engine/analysis';
552
+
553
+ const analyzer = new PitchAnalyzer({
554
+ algorithm: 'pitch_melodia',
555
+ includeContour: true
556
+ });
557
+
558
+ const profile = await analyzer.analyze('https://example.com/audio.mp3');
559
+
560
+ console.log(`Voicing ratio: ${(profile.voicingRatio * 100).toFixed(1)}%`);
561
+ console.log(`Range: ${profile.lowestNote} - ${profile.highestNote}`);
562
+ console.log(`Contour direction: ${profile.contour?.direction}`);
563
+ ```
564
+
565
+ ### Constructor Options
566
+
567
+ | Option | Type | Default | Description |
568
+ |--------|------|---------|-------------|
569
+ | `algorithm` | `PitchAlgorithm` | `'pitch_melodia'` | Pitch detection algorithm |
570
+ | `minFrequency` | `number` | `80` | Minimum frequency in Hz |
571
+ | `maxFrequency` | `number` | algorithm-dependent | Maximum frequency in Hz (1000 for pyin_legacy, 20000 for others) |
572
+ | `sampleRate` | `number` | `44100` | Target sample rate |
573
+ | `hopSize` | `number` | `1024` | Hop size in samples (~23ms at 44.1kHz, ~43 frames/sec). Larger = faster but lower time resolution |
574
+ | `crepeModelUrl` | `string` | — | CREPE model URL (for pitch_crepe algorithm) |
575
+ | `resolveUrl` | `(url) => Promise<string>` | — | URL resolver for Arweave URLs |
576
+ | `includeContour` | `boolean` | `true` | Include melody contour analysis |
577
+ | `onProgress` | `(phase, progress) => void` | — | Progress callback |
578
+
579
+ ### PitchAnalysisProfile Output
580
+
581
+ | Property | Type | Description |
582
+ |----------|------|-------------|
583
+ | `pitchResults` | `PitchResult[]` | Per-frame pitch detection results |
584
+ | `contour` | `PitchContour` | Melody contour (when `includeContour: true`) |
585
+ | `voicingRatio` | `number` | Ratio of voiced frames (0.0 - 1.0) |
586
+ | `averageFrequency` | `number` | Average frequency of voiced frames (Hz) |
587
+ | `medianFrequency` | `number` | Median frequency of voiced frames (Hz) |
588
+ | `minFrequency` | `number` | Minimum detected frequency (Hz) |
589
+ | `maxFrequency` | `number` | Maximum detected frequency (Hz) |
590
+ | `pitchRangeSemitones` | `number` | Pitch range in semitones |
591
+ | `lowestNote` | `string \| null` | Lowest detected note name |
592
+ | `highestNote` | `string \| null` | Highest detected note name |
593
+ | `noteDistribution` | `{ note, count, percentage }[]` | Note frequency distribution |
594
+ | `totalFrames` | `number` | Total frames analyzed |
595
+ | `voicedFrames` | `number` | Voiced frames count |
596
+ | `directionStats` | `DirectionStats` | Direction statistics (when contour enabled) |
597
+ | `intervalStats` | `IntervalStats` | Interval statistics (when contour enabled) |
598
+ | `analysis_metadata` | `{ algorithm_used, analyzed_at, duration_analyzed }` | Pipeline metadata |
599
+
600
+ ### PitchAnalyzer vs PitchBeatLinker
601
+
602
+ The `PitchAnalyzer` is a standalone analyzer that operates on raw audio with **no dependency on beat detection**. Use it for general pitch analysis tasks.
603
+
604
+ For rhythm game chart generation where pitch needs to be aligned to beats, use `PitchBeatLinker` combined with `MelodyContourAnalyzer` (see [BEAT_DETECTION.md](BEAT_DETECTION.md)).
605
+
606
+ ---
607
+
608
+ ## Related Documentation
609
+
610
+ - **[BEAT_DETECTION.md](BEAT_DETECTION.md)** - Beat detection, rhythm games, and chart creation features