playlist-data-engine 1.7.3 → 1.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +14 -0
- package/bin/cli.cjs +85 -0
- package/dist/gateway-CDMPqFEH.js +1320 -0
- package/dist/gateway-DKa45Uz6.cjs +6 -0
- package/dist/gateway.d.ts +1 -0
- package/dist/gateway.d.ts.map +1 -1
- package/dist/gateway.js +1 -1
- package/dist/gateway.mjs +22 -19
- package/dist/index.d.ts +1 -0
- package/dist/index.d.ts.map +1 -1
- package/dist/playlist-data-engine.js +4 -4
- package/dist/playlist-data-engine.mjs +30 -27
- package/dist/utils/engineDocs.d.ts +33 -0
- package/dist/utils/engineDocs.d.ts.map +1 -0
- package/docs/DATA_ENGINE_REFERENCE.md +6660 -0
- package/docs/USAGE_IN_OTHER_PROJECTS.md +587 -0
- package/docs/features/AUDIO_ANALYSIS.md +610 -0
- package/docs/features/BEAT_DETECTION.md +5250 -0
- package/docs/features/COMBAT_SYSTEM.md +1632 -0
- package/docs/features/CONTENT_PACKS.md +464 -0
- package/docs/features/CUSTOM_CONTENT.md +603 -0
- package/docs/features/ENEMY_GENERATION.md +1711 -0
- package/docs/features/EQUIPMENT_SYSTEM.md +2279 -0
- package/docs/features/EXTENSIBILITY_GUIDE.md +1106 -0
- package/docs/features/GATEWAY_RESOLUTION.md +725 -0
- package/docs/features/IRL_SENSORS.md +360 -0
- package/docs/features/PLAYLIST_PARSING.md +446 -0
- package/docs/features/PREREQUISITES.md +571 -0
- package/docs/features/ROLLS_AND_SEEDS.md +687 -0
- package/docs/features/XP_AND_STATS.md +1221 -0
- package/llms.txt +33 -0
- package/package.json +9 -2
- package/skills/playlist-data-engine/SKILL.md +69 -0
- package/dist/gateway-C_p9Ku3O.js +0 -1211
- package/dist/gateway-Ceg-5xug.cjs +0 -1
|
@@ -0,0 +1,610 @@
|
|
|
1
|
+
# Audio Analysis Documentation
|
|
2
|
+
|
|
3
|
+
The Playlist Data Engine provides audio analysis modes for extracting meaningful data from music files. Each mode serves different use cases, from quick character generation to timeline visualization.
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## Overview
|
|
8
|
+
|
|
9
|
+
The engine's audio analysis is powered by the Web Audio API and provides two distinct modes for general audio analysis:
|
|
10
|
+
|
|
11
|
+
| Mode | Method | Purpose | Use Case |
|
|
12
|
+
|------|--------|---------|----------|
|
|
13
|
+
| **Triple Tap Real-Time** | `extractSonicFingerprint()` | Quick analysis at key positions | Character generation, quick profiling |
|
|
14
|
+
| **Full Song Timeline** | `analyzeTimeline()` | Complete track analysis | Waveform visualization, level generation |
|
|
15
|
+
| **Music Classification** | `analyze()` | Deep ML classification | Genre, mood, and vibe detection |
|
|
16
|
+
| **Pitch Analysis** | `analyze()` | Full-track pitch detection | Melody analysis, note detection |
|
|
17
|
+
|
|
18
|
+
> **Note**: For rhythm game features like beat detection, beat streaming, and chart creation, see [BEAT_DETECTION.md](BEAT_DETECTION.md).
|
|
19
|
+
|
|
20
|
+
### Source Files
|
|
21
|
+
|
|
22
|
+
| Component | Location |
|
|
23
|
+
|-----------|----------|
|
|
24
|
+
| **AudioAnalyzer** (main class) | [src/core/analysis/AudioAnalyzer.ts](../src/core/analysis/AudioAnalyzer.ts) |
|
|
25
|
+
| **MusicClassifier** (ML classification) | [src/core/analysis/MusicClassifier.ts](../src/core/analysis/MusicClassifier.ts) |
|
|
26
|
+
| **PitchAnalyzer** (pitch detection) | [src/core/analysis/PitchAnalyzer.ts](../src/core/analysis/PitchAnalyzer.ts) |
|
|
27
|
+
| **SpectrumScanner** (frequency bands) | [src/core/analysis/SpectrumScanner.ts](../src/core/analysis/SpectrumScanner.ts) |
|
|
28
|
+
| **Audio Types** | [src/core/types/AudioProfile.ts](../src/core/types/AudioProfile.ts) |
|
|
29
|
+
|
|
30
|
+
---
|
|
31
|
+
|
|
32
|
+
## ⚠️ Import path — TensorFlow.js
|
|
33
|
+
|
|
34
|
+
The audio analysis surface (`AudioAnalyzer`, `MusicClassifier`, `EssentiaPitchDetector`,
|
|
35
|
+
`PitchAnalyzer`, and the level-generation classes) depends on **`@tensorflow/tfjs`**
|
|
36
|
+
(~14 MB). These symbols are exported from the **`playlist-data-engine/analysis`**
|
|
37
|
+
subpath, NOT the default entry. Importing them from the bare `playlist-data-engine`
|
|
38
|
+
will fail and/or pull TensorFlow into your main bundle.
|
|
39
|
+
|
|
40
|
+
```ts
|
|
41
|
+
// ✅ correct — TF-bearing symbols come from /analysis
|
|
42
|
+
import { AudioAnalyzer, MusicClassifier, PitchAnalyzer } from 'playlist-data-engine/analysis';
|
|
43
|
+
|
|
44
|
+
// ❌ wrong — these are NOT on the default (TF-free) entry
|
|
45
|
+
import { AudioAnalyzer } from 'playlist-data-engine';
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
The default `playlist-data-engine` entry is kept TensorFlow-free on purpose so
|
|
49
|
+
apps that only need parsing/generation/gateway utilities don't pay the TF cost.
|
|
50
|
+
For web apps, run analysis in a **Web Worker** so the TF runtime loads on a
|
|
51
|
+
worker thread. (See `src/features/directoryTools/audioAnalysis/analyzeWorker.ts`
|
|
52
|
+
in ApeTapes for a reference worker implementation.)
|
|
53
|
+
|
|
54
|
+
---
|
|
55
|
+
|
|
56
|
+
## 3-Tap Real-Time Analysis
|
|
57
|
+
|
|
58
|
+
The original `AudioAnalyzer` real-time analysis uses the "Triple Tap" strategy: analyzing three key positions (5%, 40%, 70%) in tracks longer than 3 seconds, or the full buffer for shorter clips.
|
|
59
|
+
|
|
60
|
+
### Method
|
|
61
|
+
|
|
62
|
+
```typescript
|
|
63
|
+
extractSonicFingerprint(audioUrl: string): Promise<AudioProfile>
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
### Usage
|
|
67
|
+
|
|
68
|
+
```typescript
|
|
69
|
+
import { AudioAnalyzer } from 'playlist-data-engine/analysis';
|
|
70
|
+
|
|
71
|
+
const analyzer = new AudioAnalyzer({
|
|
72
|
+
includeAdvancedMetrics: true, // Include spectral centroid, rolloff, zero-crossing rate
|
|
73
|
+
trebleBoost: 0.6, // Reduce treble dominance
|
|
74
|
+
bassBoost: 1.2, // Increase bass presence
|
|
75
|
+
});
|
|
76
|
+
|
|
77
|
+
const profile = await analyzer.extractSonicFingerprint(track.audio_url);
|
|
78
|
+
|
|
79
|
+
console.log(`Bass: ${profile.bass_dominance}`);
|
|
80
|
+
console.log(`Mid: ${profile.mid_dominance}`);
|
|
81
|
+
console.log(`Treble: ${profile.treble_dominance}`);
|
|
82
|
+
console.log(`RMS Energy: ${profile.rms_energy}`);
|
|
83
|
+
console.log(`Dynamic Range: ${profile.dynamic_range}`);
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
### AudioProfile Output
|
|
87
|
+
|
|
88
|
+
| Property | Type | Description |
|
|
89
|
+
|----------|------|-------------|
|
|
90
|
+
| `bass_dominance` | `number` | Relative bass energy (0-1, normalized with others) |
|
|
91
|
+
| `mid_dominance` | `number` | Relative mid-range energy (0-1) |
|
|
92
|
+
| `treble_dominance` | `number` | Relative treble energy (0-1) |
|
|
93
|
+
| `average_amplitude` | `number` | Average amplitude across all samples |
|
|
94
|
+
| `rms_energy` | `number` | Root mean square energy (perceived loudness) |
|
|
95
|
+
| `dynamic_range` | `number` | Peak amplitude minus RMS energy |
|
|
96
|
+
| `spectral_centroid` | `number?` | Brightness indicator (with advanced metrics) |
|
|
97
|
+
| `spectral_rolloff` | `number?` | Frequency below which 85% of energy is contained |
|
|
98
|
+
| `zero_crossing_rate` | `number?` | Measure of noisiness/percussiveness |
|
|
99
|
+
| `analysis_metadata` | `object` | Duration, sample positions, timestamp |
|
|
100
|
+
|
|
101
|
+
### Frequency Bands
|
|
102
|
+
|
|
103
|
+
The analyzer separates audio into three perceptual frequency bands:
|
|
104
|
+
|
|
105
|
+
| Band | Frequency Range | Typical Content |
|
|
106
|
+
|------|-----------------|-----------------|
|
|
107
|
+
| **Bass** | 20 - 400 Hz | Kick drums, bass guitar, sub-bass |
|
|
108
|
+
| **Mid** | 400 - 4000 Hz | Vocals, guitars, keyboards |
|
|
109
|
+
| **Treble** | 4000 - 14000 Hz | Hi-hats, cymbals, high harmonics |
|
|
110
|
+
|
|
111
|
+
### Triple Tap Strategy
|
|
112
|
+
|
|
113
|
+
For tracks longer than 3 seconds, the analyzer samples at three positions:
|
|
114
|
+
|
|
115
|
+
| Position | Rationale |
|
|
116
|
+
|----------|-----------|
|
|
117
|
+
| **5%** | Capture intro, often distinctive |
|
|
118
|
+
| **40%** | Typically in the main body/chorus |
|
|
119
|
+
| **70%** | Late section, often bridge or climax |
|
|
120
|
+
|
|
121
|
+
This provides representative coverage while avoiding expensive full-track analysis.
|
|
122
|
+
|
|
123
|
+
---
|
|
124
|
+
|
|
125
|
+
## Full Song Analysis
|
|
126
|
+
|
|
127
|
+
For applications requiring complete track data (waveform visualization, timeline displays, level generation), use `analyzeTimeline()`.
|
|
128
|
+
|
|
129
|
+
### Method
|
|
130
|
+
|
|
131
|
+
```typescript
|
|
132
|
+
analyzeTimeline(audioUrl: string, strategy: SamplingStrategy): Promise<AudioTimelineEvent[]>
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
### Sampling Strategies
|
|
136
|
+
|
|
137
|
+
```typescript
|
|
138
|
+
// Option A: Sample every N seconds
|
|
139
|
+
const timeline = await analyzer.analyzeTimeline(audioUrl, {
|
|
140
|
+
type: 'interval',
|
|
141
|
+
intervalSeconds: 2 // Sample every 2 seconds
|
|
142
|
+
});
|
|
143
|
+
|
|
144
|
+
// Option B: Generate exactly N data points
|
|
145
|
+
const timeline = await analyzer.analyzeTimeline(audioUrl, {
|
|
146
|
+
type: 'count',
|
|
147
|
+
count: 100 // Exactly 100 evenly-spaced samples
|
|
148
|
+
});
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
### AudioTimelineEvent Output
|
|
152
|
+
|
|
153
|
+
| Property | Type | Description |
|
|
154
|
+
|----------|------|-------------|
|
|
155
|
+
| `timestamp` | `number` | Position in the track (seconds) |
|
|
156
|
+
| `duration` | `number` | Length of analyzed segment |
|
|
157
|
+
| `bass` | `number` | Bass dominance (0-1, normalized) |
|
|
158
|
+
| `mid` | `number` | Mid dominance (0-1, normalized) |
|
|
159
|
+
| `treble` | `number` | Treble dominance (0-1, normalized) |
|
|
160
|
+
| `amplitude` | `number` | RMS energy for this segment |
|
|
161
|
+
| `rms_energy` | `number` | Root mean square energy |
|
|
162
|
+
| `peak` | `number` | Peak amplitude |
|
|
163
|
+
| `dynamic_range` | `number` | Peak minus RMS |
|
|
164
|
+
| `spectral_centroid` | `number` | Brightness indicator |
|
|
165
|
+
| `spectral_rolloff` | `number` | Energy distribution measure |
|
|
166
|
+
| `zero_crossing_rate` | `number` | Noisiness/percussiveness |
|
|
167
|
+
|
|
168
|
+
### Usage Example
|
|
169
|
+
|
|
170
|
+
```typescript
|
|
171
|
+
import { AudioAnalyzer } from 'playlist-data-engine/analysis';
|
|
172
|
+
|
|
173
|
+
const analyzer = new AudioAnalyzer();
|
|
174
|
+
|
|
175
|
+
// Generate 100 data points across the song for visualization
|
|
176
|
+
const timeline = await analyzer.analyzeTimeline(audioUrl, {
|
|
177
|
+
type: 'count',
|
|
178
|
+
count: 100
|
|
179
|
+
});
|
|
180
|
+
|
|
181
|
+
// Find the loudest moment
|
|
182
|
+
const peakMoment = timeline.reduce((max, event) =>
|
|
183
|
+
event.rms_energy > max.rms_energy ? event : max
|
|
184
|
+
);
|
|
185
|
+
console.log(`Peak at ${peakMoment.timestamp}s`);
|
|
186
|
+
|
|
187
|
+
// Build a simple waveform visualization
|
|
188
|
+
timeline.forEach(event => {
|
|
189
|
+
const barHeight = Math.round(event.rms_energy * 20);
|
|
190
|
+
console.log(`${'█'.repeat(barHeight)} [${event.timestamp.toFixed(1)}s]`);
|
|
191
|
+
});
|
|
192
|
+
```
|
|
193
|
+
|
|
194
|
+
|
|
195
|
+
---
|
|
196
|
+
|
|
197
|
+
## Music Classification (Genre, Mood, Vibe)
|
|
198
|
+
|
|
199
|
+
For deep semantic analysis of music including mood, themes, and vibe metrics (like danceability), use the `MusicClassifier`. This uses multiple `essentia.js` models.
|
|
200
|
+
|
|
201
|
+
### Method
|
|
202
|
+
|
|
203
|
+
```typescript
|
|
204
|
+
analyze(audioUrl: string): Promise<MusicClassificationProfile>
|
|
205
|
+
```
|
|
206
|
+
|
|
207
|
+
### Usage Example
|
|
208
|
+
|
|
209
|
+
```typescript
|
|
210
|
+
import { MusicClassifier } from 'playlist-data-engine/analysis';
|
|
211
|
+
|
|
212
|
+
const classifier = new MusicClassifier({
|
|
213
|
+
topN: 3, // Return top 3 genres/moods
|
|
214
|
+
threshold: 0.1, // 10% confidence threshold
|
|
215
|
+
analysisDurationSeconds: 30, // Analyze 30 seconds instead of full song
|
|
216
|
+
analysisStartPosition: 0.5 // Start from middle of song
|
|
217
|
+
});
|
|
218
|
+
|
|
219
|
+
const profile = await classifier.analyze('https://example.com/audio.mp3');
|
|
220
|
+
|
|
221
|
+
console.log(`Genre: ${profile.primary_genre}`);
|
|
222
|
+
console.log(`Moods: ${profile.mood_tags.join(', ')}`);
|
|
223
|
+
console.log(`Danceability: ${profile.vibe_metrics.danceability}`);
|
|
224
|
+
console.log(`Energy: ${profile.vibe_metrics.energy}`);
|
|
225
|
+
```
|
|
226
|
+
|
|
227
|
+
### MusicClassificationProfile Output
|
|
228
|
+
|
|
229
|
+
| Property | Type | Description |
|
|
230
|
+
|----------|------|-------------|
|
|
231
|
+
| `genres` | `ClassificationTag[]` | Top matched genres |
|
|
232
|
+
| `moods` | `ClassificationTag[]` | Top matched moods/themes |
|
|
233
|
+
| `primary_genre` | `string` | Highest confidence genre |
|
|
234
|
+
| `mood_tags` | `string[]` | Most relevant mood keywords |
|
|
235
|
+
| `vibe_metrics` | `VibeMetrics` | Danceability, energy, valence, etc. |
|
|
236
|
+
|
|
237
|
+
### Partial Song Analysis
|
|
238
|
+
|
|
239
|
+
By default, `analyze()` processes the entire audio signal. For faster classification (genre, mood, vibe), you can analyze only a segment of the song by setting `analysisDurationSeconds` and `analysisStartPosition` in the constructor options.
|
|
240
|
+
|
|
241
|
+
| Option | Type | Default | Description |
|
|
242
|
+
|--------|------|---------|-------------|
|
|
243
|
+
| `analysisDurationSeconds` | number | — | Seconds of audio to analyze. Full song when not set. Recommended: 30s |
|
|
244
|
+
| `analysisStartPosition` | number | `0.5` | Position to start (0.0 = start, 0.5 = middle, 1.0 = end). Only applies when duration is set |
|
|
245
|
+
|
|
246
|
+
```typescript
|
|
247
|
+
// Analyze middle 30 seconds of each song (faster, good accuracy for genre/mood)
|
|
248
|
+
const classifier = new MusicClassifier({
|
|
249
|
+
analysisDurationSeconds: 30,
|
|
250
|
+
analysisStartPosition: 0.5
|
|
251
|
+
});
|
|
252
|
+
|
|
253
|
+
// Analyze first 15 seconds (fastest, lower accuracy)
|
|
254
|
+
const quickClassifier = new MusicClassifier({
|
|
255
|
+
analysisDurationSeconds: 15,
|
|
256
|
+
analysisStartPosition: 0
|
|
257
|
+
});
|
|
258
|
+
```
|
|
259
|
+
|
|
260
|
+
> **Note**: The audio file is still fully downloaded and decoded (required by the Web Audio API). Only the ML feature extraction and model inference operate on the segment, which is where the speed improvement comes from. A 30-second segment processes ~8x faster than a 4-minute song.
|
|
261
|
+
|
|
262
|
+
---
|
|
263
|
+
|
|
264
|
+
### Two-Step Model Architecture
|
|
265
|
+
|
|
266
|
+
The `MusicClassifier` supports both single-step (one model) and two-step (embedding + classifier) architectures. This enables using state-of-the-art models like Discogs-EffNet for embeddings combined with specialized classifier heads.
|
|
267
|
+
|
|
268
|
+
#### Architecture Compatibility Table
|
|
269
|
+
|
|
270
|
+
Different model architectures require different mel-band configurations for feature extraction:
|
|
271
|
+
|
|
272
|
+
| Architecture | Mel Bands | Extractor | Compatible Models |
|
|
273
|
+
|--------------|-----------|-----------|-------------------|
|
|
274
|
+
| `musicnn` | 96 | Essentia musicnn | MusiCNN classifiers, MSD models |
|
|
275
|
+
| `effnet` | 128 | Custom | Discogs-EffNet embeddings |
|
|
276
|
+
| `vggish` | 64 | Essentia vggish | VGGish classifiers, AudioSet |
|
|
277
|
+
| `tempocnn` | 40 | Essentia tempocnn | TempoCNN tempo estimation |
|
|
278
|
+
|
|
279
|
+
> **Important**: The engine automatically detects the architecture from the model URL and uses the correct mel-band configuration. No manual configuration needed!
|
|
280
|
+
|
|
281
|
+
#### Model Configuration Formats
|
|
282
|
+
|
|
283
|
+
Every model option (`genre`, `mood`, `danceability`, `voice`, `acoustic`) accepts two formats:
|
|
284
|
+
|
|
285
|
+
##### Format 1: Single-Step Model Config (with explicit type)
|
|
286
|
+
|
|
287
|
+
For URLs where architecture cannot be detected (e.g., Arweave URLs), use `SingleStepModelConfig`:
|
|
288
|
+
|
|
289
|
+
```typescript
|
|
290
|
+
import { MusicClassifier, type SingleStepModelConfig } from 'playlist-data-engine/analysis';
|
|
291
|
+
|
|
292
|
+
const config: SingleStepModelConfig = {
|
|
293
|
+
modelUrl: 'https://arweave.net/xxx/model.json',
|
|
294
|
+
modelType: 'musicnn', // Explicitly specify architecture
|
|
295
|
+
genreType: 'jamendo', // Explicitly specify genre list (for genre models)
|
|
296
|
+
labels: ['custom', 'labels'] // Optional custom labels
|
|
297
|
+
};
|
|
298
|
+
```
|
|
299
|
+
|
|
300
|
+
| Property | Type | Required | Description |
|
|
301
|
+
|----------|------|----------|-------------|
|
|
302
|
+
| `modelUrl` | `string` | Yes | URL to the model file |
|
|
303
|
+
| `modelType` | `ModelArchitecture` | Yes | Explicit architecture: `'musicnn'` \| `'effnet'` \| `'vggish'` \| `'tempocnn'` |
|
|
304
|
+
| `genreType` | `GenreListType` | No* | Explicit genre list: `'jamendo'` \| `'discogs400'` \| `'tzanetakis'` \| `'mtt_musicnn'` (*required for genre models with Arweave URLs) |
|
|
305
|
+
| `labels` | `string[]` | No | Custom output labels |
|
|
306
|
+
|
|
307
|
+
##### Format 2: Two-Step Model Config (with explicit types)
|
|
308
|
+
|
|
309
|
+
Separate embedding and classifier models with optional explicit type parameters:
|
|
310
|
+
|
|
311
|
+
```typescript
|
|
312
|
+
import { MusicClassifier, type TwoStepModelConfig } from 'playlist-data-engine/analysis';
|
|
313
|
+
|
|
314
|
+
const config: TwoStepModelConfig = {
|
|
315
|
+
embedding: '/models/discogs-effnet-bs64-1.json',
|
|
316
|
+
classifier: '/models/mtg_jamendo_genre-discogs-effnet-1.json',
|
|
317
|
+
// Optional explicit types (override URL detection)
|
|
318
|
+
embeddingType: 'effnet', // ModelArchitecture
|
|
319
|
+
classifierType: 'discogs400' // GenreListType (for genre models)
|
|
320
|
+
};
|
|
321
|
+
```
|
|
322
|
+
|
|
323
|
+
| Property | Type | Required | Description |
|
|
324
|
+
|----------|------|----------|-------------|
|
|
325
|
+
| `embedding` | `string` | Yes | URL to the embedding model |
|
|
326
|
+
| `classifier` | `string` | Yes | URL to the classifier model |
|
|
327
|
+
| `labels` | `string[]` | No | Custom output labels |
|
|
328
|
+
| `embeddingType` | `ModelArchitecture` | No | Explicit embedding type: `'musicnn'` \| `'effnet'` \| `'vggish'` \| `'tempocnn'` |
|
|
329
|
+
| `classifierType` | `GenreListType` | No | Explicit genre list type: `'jamendo'` \| `'discogs400'` \| `'tzanetakis'` \| `'mtt_musicnn'` |
|
|
330
|
+
|
|
331
|
+
##### Why Use Explicit Type Parameters?
|
|
332
|
+
|
|
333
|
+
URL-based detection works for conventional file paths like `/models/effnet-classifier.json`, but fails for:
|
|
334
|
+
- **Arweave URLs**: `https://arweave.net/xxx/model.json` contains no architecture hints
|
|
335
|
+
- **Custom hosting**: URLs without descriptive filenames
|
|
336
|
+
- **Proxied URLs**: Gateway URLs that obscure the original filename
|
|
337
|
+
|
|
338
|
+
#### Signal Flow Diagrams
|
|
339
|
+
|
|
340
|
+
**Single-Step Flow:**
|
|
341
|
+
```
|
|
342
|
+
┌─────────────────┐ ┌──────────────────┐ ┌─────────────────┐
|
|
343
|
+
│ Audio Signal │ ──▶ │ Feature │ ──▶ │ Single Model │
|
|
344
|
+
│ (16kHz mono) │ │ Extractor │ │ (classifier) │
|
|
345
|
+
└─────────────────┘ │ (96 bands) │ └────────┬────────┘
|
|
346
|
+
└──────────────────┘ │
|
|
347
|
+
▼
|
|
348
|
+
┌─────────────────┐
|
|
349
|
+
│ Class Labels │
|
|
350
|
+
│ (genre, mood) │
|
|
351
|
+
└─────────────────┘
|
|
352
|
+
```
|
|
353
|
+
|
|
354
|
+
**Two-Step Flow:**
|
|
355
|
+
```
|
|
356
|
+
┌─────────────────┐ ┌──────────────────┐ ┌─────────────────┐
|
|
357
|
+
│ Audio Signal │ ──▶ │ Feature │ ──▶ │ Embedding │
|
|
358
|
+
│ (16kHz mono) │ │ Extractor │ │ Model │
|
|
359
|
+
└─────────────────┘ │ (128 bands) │ │ (1280-dim) │
|
|
360
|
+
└──────────────────┘ └────────┬────────┘
|
|
361
|
+
│
|
|
362
|
+
Architecture-specific │
|
|
363
|
+
mel-band config ▼
|
|
364
|
+
┌─────────────────┐
|
|
365
|
+
│ Classifier │
|
|
366
|
+
│ Model │
|
|
367
|
+
│ (class probs) │
|
|
368
|
+
└────────┬────────┘
|
|
369
|
+
│
|
|
370
|
+
▼
|
|
371
|
+
┌─────────────────┐
|
|
372
|
+
│ Class Labels │
|
|
373
|
+
│ (genre, mood) │
|
|
374
|
+
└─────────────────┘
|
|
375
|
+
```
|
|
376
|
+
|
|
377
|
+
#### Configuration Examples
|
|
378
|
+
|
|
379
|
+
**Mixed Configuration (Single + Two-Step):**
|
|
380
|
+
|
|
381
|
+
```typescript
|
|
382
|
+
const classifier = new MusicClassifier({
|
|
383
|
+
models: {
|
|
384
|
+
// Two-step: embedding + classifier (uses 128-band extractor)
|
|
385
|
+
genre: {
|
|
386
|
+
embedding: '/models/discogs-effnet-bs64-1.json',
|
|
387
|
+
classifier: '/models/mtg_jamendo_genre-discogs-effnet-1.json'
|
|
388
|
+
},
|
|
389
|
+
// Two-step: same embedding cached, different classifier
|
|
390
|
+
mood: {
|
|
391
|
+
embedding: '/models/discogs-effnet-bs64-1.json',
|
|
392
|
+
classifier: '/models/mtg_jamendo_moodtheme-discogs-effnet-1.json'
|
|
393
|
+
},
|
|
394
|
+
// Single-step: requires modelUrl + modelType (uses 64-band vggish extractor)
|
|
395
|
+
danceability: {
|
|
396
|
+
modelUrl: '/models/danceability-vggish-audioset-1.json',
|
|
397
|
+
modelType: 'vggish'
|
|
398
|
+
},
|
|
399
|
+
// Single-step: optional voice detection
|
|
400
|
+
voice: {
|
|
401
|
+
modelUrl: '/models/voice-detector.json',
|
|
402
|
+
modelType: 'musicnn'
|
|
403
|
+
},
|
|
404
|
+
// Two-step: optional acoustic detection
|
|
405
|
+
acoustic: {
|
|
406
|
+
embedding: '/models/discogs-effnet-bs64-1.json',
|
|
407
|
+
classifier: '/models/acoustic-classifier.json'
|
|
408
|
+
}
|
|
409
|
+
}
|
|
410
|
+
});
|
|
411
|
+
```
|
|
412
|
+
|
|
413
|
+
#### Using Arweave-Hosted Models (Zero Setup)
|
|
414
|
+
|
|
415
|
+
The default configuration uses pre-configured Arweave-hosted models:
|
|
416
|
+
|
|
417
|
+
```typescript
|
|
418
|
+
import { MusicClassifier } from 'playlist-data-engine/analysis';
|
|
419
|
+
|
|
420
|
+
// Zero setup - models load from Arweave automatically
|
|
421
|
+
const classifier = new MusicClassifier();
|
|
422
|
+
|
|
423
|
+
const profile = await classifier.analyze('https://example.com/track.mp3');
|
|
424
|
+
```
|
|
425
|
+
|
|
426
|
+
| Model | Architecture | Labels | Source |
|
|
427
|
+
|-------|-------------|--------|--------|
|
|
428
|
+
| **Genre** | Two-step (effnet + classifier) | discogs400 (400+ subgenres) | Arweave |
|
|
429
|
+
| **Mood** | Two-step (effnet + classifier) | JAMENDO_MOODS (60 themes) | Arweave |
|
|
430
|
+
| **Danceability** | Single-step (musicnn) | Binary | Turbo Gateway |
|
|
431
|
+
|
|
432
|
+
#### Using Presets
|
|
433
|
+
|
|
434
|
+
Instead of raw URLs, use preset names to select pre-configured models. This is the simplest way to swap genre/mood/danceability models without managing URLs.
|
|
435
|
+
|
|
436
|
+
```typescript
|
|
437
|
+
import { MusicClassifier } from 'playlist-data-engine/analysis';
|
|
438
|
+
|
|
439
|
+
// Use presets for genre and mood
|
|
440
|
+
const classifier = new MusicClassifier({
|
|
441
|
+
preset: { genre: 'jamendo', mood: 'jamendo' }
|
|
442
|
+
});
|
|
443
|
+
|
|
444
|
+
// Mix presets with custom URLs — explicit models take precedence
|
|
445
|
+
const classifier = new MusicClassifier({
|
|
446
|
+
preset: { genre: 'tzanetakis' },
|
|
447
|
+
models: {
|
|
448
|
+
mood: { modelUrl: '/models/custom-mood.json', modelType: 'musicnn' }
|
|
449
|
+
}
|
|
450
|
+
});
|
|
451
|
+
```
|
|
452
|
+
|
|
453
|
+
**Available presets:**
|
|
454
|
+
|
|
455
|
+
| Category | Preset | Architecture | Labels |
|
|
456
|
+
|----------|--------|--------------|--------|
|
|
457
|
+
| Genre | `discogs400` | Two-step (effnet + discogs400) | 400+ subgenres |
|
|
458
|
+
| Genre | `jamendo` | Two-step (effnet + jamendo) | MTG Jamendo 87 |
|
|
459
|
+
| Genre | `tzanetakis` | Single-step (musicnn) | GTZAN 10 |
|
|
460
|
+
| Genre | `musicnn` | Single-step (musicnn) | MagnaTagATune 50 |
|
|
461
|
+
| Mood | `jamendo` | Two-step (effnet + jamendo) | 60 themes |
|
|
462
|
+
| Mood | `happyMusicnn` | Single-step (musicnn) | Binary (happy/not happy) |
|
|
463
|
+
| Danceability | `default` | Single-step (musicnn) | Binary |
|
|
464
|
+
|
|
465
|
+
To enumerate available presets at runtime:
|
|
466
|
+
|
|
467
|
+
```typescript
|
|
468
|
+
import { AVAILABLE_PRESETS } from 'playlist-data-engine/analysis';
|
|
469
|
+
console.log(AVAILABLE_PRESETS.genre); // ['discogs400', 'jamendo', 'tzanetakis', 'musicnn']
|
|
470
|
+
console.log(AVAILABLE_PRESETS.mood); // ['jamendo', 'happyMusicnn']
|
|
471
|
+
console.log(AVAILABLE_PRESETS.danceability); // ['default']
|
|
472
|
+
```
|
|
473
|
+
|
|
474
|
+
**Partial Override & Custom Arweave Models:**
|
|
475
|
+
|
|
476
|
+
```typescript
|
|
477
|
+
import { MusicClassifier, DEFAULT_ARWEAVE_MODELS } from 'playlist-data-engine/analysis';
|
|
478
|
+
|
|
479
|
+
const classifier = new MusicClassifier({
|
|
480
|
+
models: {
|
|
481
|
+
// Use default Arweave model
|
|
482
|
+
genre: DEFAULT_ARWEAVE_MODELS.genre,
|
|
483
|
+
// Override with local model
|
|
484
|
+
mood: {
|
|
485
|
+
modelUrl: '/models/mood-musicnn-msd-1.json',
|
|
486
|
+
modelType: 'musicnn'
|
|
487
|
+
},
|
|
488
|
+
// Or use custom Arweave URLs with explicit types
|
|
489
|
+
danceability: {
|
|
490
|
+
modelUrl: 'https://arweave.net/YOUR_TXID/model.json',
|
|
491
|
+
modelType: 'musicnn' // Required for Arweave URLs
|
|
492
|
+
}
|
|
493
|
+
}
|
|
494
|
+
});
|
|
495
|
+
```
|
|
496
|
+
|
|
497
|
+
> **Note:** Single-step mood models use binary labels (`'happy'`, `'not happy'`) while two-step mood models use the full JAMENDO_MOODS taxonomy (60 mood/theme labels).
|
|
498
|
+
|
|
499
|
+
#### Embedding Model Caching
|
|
500
|
+
|
|
501
|
+
When using two-step models with shared embeddings (e.g., same Discogs-EffNet for genre and mood), the embedding model is automatically cached and reused:
|
|
502
|
+
|
|
503
|
+
```typescript
|
|
504
|
+
const classifier = new MusicClassifier({
|
|
505
|
+
cacheEmbeddings: true, // Default: true
|
|
506
|
+
models: {
|
|
507
|
+
genre: {
|
|
508
|
+
embedding: '/models/discogs-effnet-bs64-1.json', // Loaded once
|
|
509
|
+
classifier: '/models/mtg_jamendo_genre-discogs-effnet-1.json'
|
|
510
|
+
},
|
|
511
|
+
mood: {
|
|
512
|
+
embedding: '/models/discogs-effnet-bs64-1.json', // Reused from cache!
|
|
513
|
+
classifier: '/models/mtg_jamendo_moodtheme-discogs-effnet-1.json'
|
|
514
|
+
}
|
|
515
|
+
}
|
|
516
|
+
});
|
|
517
|
+
|
|
518
|
+
// Later, to free memory:
|
|
519
|
+
classifier.clearEmbeddingCache(); // Clear embedding models only
|
|
520
|
+
classifier.clearClassifierCache(); // Clear classifier models only
|
|
521
|
+
classifier.clearAllCaches(); // Clear everything
|
|
522
|
+
```
|
|
523
|
+
|
|
524
|
+
#### Metadata Tracking
|
|
525
|
+
|
|
526
|
+
The `analysis_metadata.models_used` field tracks which models were used:
|
|
527
|
+
|
|
528
|
+
```typescript
|
|
529
|
+
const profile = await classifier.analyze(audioUrl);
|
|
530
|
+
|
|
531
|
+
console.log(profile.analysis_metadata.models_used);
|
|
532
|
+
// Two-step: ['/models/discogs-effnet-bs64-1.json -> /models/mtg_jamendo_genre-discogs-effnet-1.json', ...]
|
|
533
|
+
// Single-step: ['/models/genre-musicnn-msd-1.json', ...]
|
|
534
|
+
```
|
|
535
|
+
|
|
536
|
+
---
|
|
537
|
+
|
|
538
|
+
## Pitch Analysis
|
|
539
|
+
|
|
540
|
+
For full-track pitch detection without beat-map dependency, use the `PitchAnalyzer`. This provides per-frame pitch results, melody contour analysis, and summary statistics directly from raw audio.
|
|
541
|
+
|
|
542
|
+
### Method
|
|
543
|
+
|
|
544
|
+
```typescript
|
|
545
|
+
analyze(audioUrl: string): Promise<PitchAnalysisProfile>
|
|
546
|
+
```
|
|
547
|
+
|
|
548
|
+
### Usage Example
|
|
549
|
+
|
|
550
|
+
```typescript
|
|
551
|
+
import { PitchAnalyzer } from 'playlist-data-engine/analysis';
|
|
552
|
+
|
|
553
|
+
const analyzer = new PitchAnalyzer({
|
|
554
|
+
algorithm: 'pitch_melodia',
|
|
555
|
+
includeContour: true
|
|
556
|
+
});
|
|
557
|
+
|
|
558
|
+
const profile = await analyzer.analyze('https://example.com/audio.mp3');
|
|
559
|
+
|
|
560
|
+
console.log(`Voicing ratio: ${(profile.voicingRatio * 100).toFixed(1)}%`);
|
|
561
|
+
console.log(`Range: ${profile.lowestNote} - ${profile.highestNote}`);
|
|
562
|
+
console.log(`Contour direction: ${profile.contour?.direction}`);
|
|
563
|
+
```
|
|
564
|
+
|
|
565
|
+
### Constructor Options
|
|
566
|
+
|
|
567
|
+
| Option | Type | Default | Description |
|
|
568
|
+
|--------|------|---------|-------------|
|
|
569
|
+
| `algorithm` | `PitchAlgorithm` | `'pitch_melodia'` | Pitch detection algorithm |
|
|
570
|
+
| `minFrequency` | `number` | `80` | Minimum frequency in Hz |
|
|
571
|
+
| `maxFrequency` | `number` | algorithm-dependent | Maximum frequency in Hz (1000 for pyin_legacy, 20000 for others) |
|
|
572
|
+
| `sampleRate` | `number` | `44100` | Target sample rate |
|
|
573
|
+
| `hopSize` | `number` | `1024` | Hop size in samples (~23ms at 44.1kHz, ~43 frames/sec). Larger = faster but lower time resolution |
|
|
574
|
+
| `crepeModelUrl` | `string` | — | CREPE model URL (for pitch_crepe algorithm) |
|
|
575
|
+
| `resolveUrl` | `(url) => Promise<string>` | — | URL resolver for Arweave URLs |
|
|
576
|
+
| `includeContour` | `boolean` | `true` | Include melody contour analysis |
|
|
577
|
+
| `onProgress` | `(phase, progress) => void` | — | Progress callback |
|
|
578
|
+
|
|
579
|
+
### PitchAnalysisProfile Output
|
|
580
|
+
|
|
581
|
+
| Property | Type | Description |
|
|
582
|
+
|----------|------|-------------|
|
|
583
|
+
| `pitchResults` | `PitchResult[]` | Per-frame pitch detection results |
|
|
584
|
+
| `contour` | `PitchContour` | Melody contour (when `includeContour: true`) |
|
|
585
|
+
| `voicingRatio` | `number` | Ratio of voiced frames (0.0 - 1.0) |
|
|
586
|
+
| `averageFrequency` | `number` | Average frequency of voiced frames (Hz) |
|
|
587
|
+
| `medianFrequency` | `number` | Median frequency of voiced frames (Hz) |
|
|
588
|
+
| `minFrequency` | `number` | Minimum detected frequency (Hz) |
|
|
589
|
+
| `maxFrequency` | `number` | Maximum detected frequency (Hz) |
|
|
590
|
+
| `pitchRangeSemitones` | `number` | Pitch range in semitones |
|
|
591
|
+
| `lowestNote` | `string \| null` | Lowest detected note name |
|
|
592
|
+
| `highestNote` | `string \| null` | Highest detected note name |
|
|
593
|
+
| `noteDistribution` | `{ note, count, percentage }[]` | Note frequency distribution |
|
|
594
|
+
| `totalFrames` | `number` | Total frames analyzed |
|
|
595
|
+
| `voicedFrames` | `number` | Voiced frames count |
|
|
596
|
+
| `directionStats` | `DirectionStats` | Direction statistics (when contour enabled) |
|
|
597
|
+
| `intervalStats` | `IntervalStats` | Interval statistics (when contour enabled) |
|
|
598
|
+
| `analysis_metadata` | `{ algorithm_used, analyzed_at, duration_analyzed }` | Pipeline metadata |
|
|
599
|
+
|
|
600
|
+
### PitchAnalyzer vs PitchBeatLinker
|
|
601
|
+
|
|
602
|
+
The `PitchAnalyzer` is a standalone analyzer that operates on raw audio with **no dependency on beat detection**. Use it for general pitch analysis tasks.
|
|
603
|
+
|
|
604
|
+
For rhythm game chart generation where pitch needs to be aligned to beats, use `PitchBeatLinker` combined with `MelodyContourAnalyzer` (see [BEAT_DETECTION.md](BEAT_DETECTION.md)).
|
|
605
|
+
|
|
606
|
+
---
|
|
607
|
+
|
|
608
|
+
## Related Documentation
|
|
609
|
+
|
|
610
|
+
- **[BEAT_DETECTION.md](BEAT_DETECTION.md)** - Beat detection, rhythm games, and chart creation features
|