vibes-plug 2.14.1 → 3.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (151) hide show
  1. package/.claude/rules/vibes-plug-core.md +5 -0
  2. package/.cursor/rules/vibes-plug-core.mdc +7 -2
  3. package/.cursorrules +8 -2
  4. package/AGENTS.md +23 -2
  5. package/CHANGELOG.md +114 -0
  6. package/CLAUDE.md +10 -3
  7. package/README.md +216 -611
  8. package/bin/vibes.mjs +1104 -0
  9. package/package.json +11 -3
  10. package/plugin.json +4 -3
  11. package/scripts/check-anti-slop.mjs +53 -0
  12. package/scripts/install.js +3 -1
  13. package/scripts/update_skills.js +1 -1
  14. package/scripts/update_skills.mjs +86 -0
  15. package/scripts/validate-skills.mjs +111 -0
  16. package/skills/accessibility-testing-expert/SKILL.md +117 -116
  17. package/skills/affective-computing-emotion-ai/SKILL.md +83 -0
  18. package/skills/agentic-coding-workflow-expert/SKILL.md +297 -0
  19. package/skills/agentic-memory-architect/SKILL.md +52 -0
  20. package/skills/agentic-micro-economy-architect/SKILL.md +92 -0
  21. package/skills/ai-llm-integration-expert/SKILL.md +330 -194
  22. package/skills/ai-media-generation-expert/SKILL.md +173 -172
  23. package/skills/ai-prompt-engineering-expert/SKILL.md +204 -134
  24. package/skills/ai-safety-governance-expert/SKILL.md +223 -0
  25. package/skills/angular-expert/SKILL.md +149 -148
  26. package/skills/anti-slop/SKILL.md +134 -133
  27. package/skills/api-design-expert/SKILL.md +4 -3
  28. package/skills/api-gateway-proxy-expert/SKILL.md +3 -2
  29. package/skills/app-analyzer-optimizer/SKILL.md +4 -3
  30. package/skills/apple-ecosystem-expert/SKILL.md +6 -5
  31. package/skills/astro-framework-expert/SKILL.md +201 -200
  32. package/skills/async-queue-temporal-expert/SKILL.md +218 -217
  33. package/skills/authentication-identity-expert/SKILL.md +174 -173
  34. package/skills/autonomous-red-teamer/SKILL.md +338 -203
  35. package/skills/autonomous-tdd-debugger/SKILL.md +6 -5
  36. package/skills/biome-linter-formatter-expert/SKILL.md +90 -89
  37. package/skills/blockchain-web3-expert/SKILL.md +116 -115
  38. package/skills/brainstorming/SKILL.md +392 -377
  39. package/skills/browser-automation-expert/SKILL.md +260 -222
  40. package/skills/bun-runtime-expert/SKILL.md +5 -4
  41. package/skills/chatbot-messaging-expert/SKILL.md +115 -114
  42. package/skills/ci-cd-devops-architect/SKILL.md +3 -2
  43. package/skills/cloud-hosting-expert/SKILL.md +5 -4
  44. package/skills/coderabbit/SKILL.md +5 -4
  45. package/skills/compliance-gdpr-privacy-expert/SKILL.md +3 -2
  46. package/skills/composable-mach-architect/SKILL.md +338 -0
  47. package/skills/cron-scheduler-expert/SKILL.md +5 -4
  48. package/skills/data-pipeline-etl-expert/SKILL.md +3 -2
  49. package/skills/data-telemetry-expert/SKILL.md +5 -4
  50. package/skills/data-visualization-expert/SKILL.md +155 -154
  51. package/skills/database-orm-expert/SKILL.md +166 -165
  52. package/skills/deep-research-analyst/SKILL.md +182 -136
  53. package/skills/dependency-upgrade-migrator/SKILL.md +11 -10
  54. package/skills/design-system-architect/SKILL.md +4 -3
  55. package/skills/desktop-electron-expert/SKILL.md +129 -128
  56. package/skills/documentation-site-expert/SKILL.md +60 -59
  57. package/skills/doku-mcp-server/SKILL.md +5 -4
  58. package/skills/doku-payment-gateway/SKILL.md +250 -232
  59. package/skills/domain-driven-design-expert/SKILL.md +3 -2
  60. package/skills/e2e-testing-expert/SKILL.md +5 -4
  61. package/skills/ecommerce-expert/SKILL.md +88 -87
  62. package/skills/email-notification-expert/SKILL.md +5 -4
  63. package/skills/ephemeral-generative-ui-architect/SKILL.md +88 -0
  64. package/skills/error-resilience-expert/SKILL.md +14 -13
  65. package/skills/event-driven-architect/SKILL.md +5 -4
  66. package/skills/feature-flag-analytics-expert/SKILL.md +3 -2
  67. package/skills/file-upload-media-expert/SKILL.md +5 -4
  68. package/skills/firebase-security-expert/SKILL.md +5 -4
  69. package/skills/form-validation-expert/SKILL.md +7 -6
  70. package/skills/frontier-ai-models-expert/SKILL.md +116 -0
  71. package/skills/fullstack-expert/SKILL.md +185 -184
  72. package/skills/gemini-agent-booster/SKILL.md +248 -172
  73. package/skills/geospatial-maps-expert/SKILL.md +81 -80
  74. package/skills/global-a11y-i18n-expert/SKILL.md +5 -4
  75. package/skills/glsl-shader-expert/SKILL.md +191 -190
  76. package/skills/go-programming-expert/SKILL.md +5 -4
  77. package/skills/graph-rag-knowledge-expert/SKILL.md +201 -200
  78. package/skills/graphql-apollo-expert/SKILL.md +5 -4
  79. package/skills/headless-cms-expert/SKILL.md +182 -181
  80. package/skills/hig/SKILL.md +5 -4
  81. package/skills/js-backend-expert/SKILL.md +219 -218
  82. package/skills/legacy-code-translator/SKILL.md +6 -5
  83. package/skills/llm-finops-router/SKILL.md +52 -0
  84. package/skills/local-slm-edge-ai-expert/SKILL.md +168 -167
  85. package/skills/logging-error-tracking-expert/SKILL.md +5 -4
  86. package/skills/mcp-server-architect/SKILL.md +315 -307
  87. package/skills/micro-frontend-architect/SKILL.md +5 -4
  88. package/skills/mobile-expo-expert/SKILL.md +5 -4
  89. package/skills/modern-css-native-expert/SKILL.md +190 -189
  90. package/skills/monorepo-architect/SKILL.md +5 -4
  91. package/skills/mpa-orchestrator/SKILL.md +41 -4
  92. package/skills/multi-agent-orchestration/SKILL.md +388 -254
  93. package/skills/mvc-expert/SKILL.md +5 -4
  94. package/skills/n8n-automation-expert/SKILL.md +90 -89
  95. package/skills/nextjs-app-router-expert/SKILL.md +3 -2
  96. package/skills/openapi-swagger-codegen-expert/SKILL.md +4 -3
  97. package/skills/payment-gateway-expert/SKILL.md +131 -128
  98. package/skills/pdf-document-generation-expert/SKILL.md +92 -91
  99. package/skills/performance-web-vitals/SKILL.md +5 -4
  100. package/skills/post-quantum-crypto-migrator/SKILL.md +3 -2
  101. package/skills/prd-architect/SKILL.md +183 -182
  102. package/skills/proactive-background-watcher/SKILL.md +5 -4
  103. package/skills/production-ready-hardener/SKILL.md +10 -9
  104. package/skills/pwa-offline-first-expert/SKILL.md +227 -226
  105. package/skills/pydantic-ai-expert/SKILL.md +162 -161
  106. package/skills/python-programming-expert/SKILL.md +5 -4
  107. package/skills/rate-limit-abuse-prevention/SKILL.md +5 -4
  108. package/skills/realtime-collaboration-expert/SKILL.md +3 -2
  109. package/skills/rich-text-editor-expert/SKILL.md +178 -177
  110. package/skills/rust-programming-expert/SKILL.md +5 -4
  111. package/skills/saas-architect/SKILL.md +155 -154
  112. package/skills/saas-billing/SKILL.md +394 -382
  113. package/skills/saas-multi-tenant/SKILL.md +7 -6
  114. package/skills/scalability-clean-code/SKILL.md +5 -4
  115. package/skills/search-engine-expert/SKILL.md +90 -89
  116. package/skills/self-healing-cloud-orchestrator/SKILL.md +3 -2
  117. package/skills/senior-frontend/SKILL.md +14 -9
  118. package/skills/seo/SKILL.md +4 -4
  119. package/skills/session-memory-manager/SKILL.md +129 -128
  120. package/skills/solidjs-expert/SKILL.md +81 -80
  121. package/skills/spa-orchestrator/SKILL.md +5 -4
  122. package/skills/sse-websocket-streaming-expert/SKILL.md +3 -2
  123. package/skills/state-management-expert/SKILL.md +5 -4
  124. package/skills/supabase-security-expert/SKILL.md +5 -4
  125. package/skills/svelte-sveltekit-expert/SKILL.md +92 -91
  126. package/skills/svg-animation-motion-expert/SKILL.md +3 -2
  127. package/skills/synthetic-data-finetuning-expert/SKILL.md +156 -155
  128. package/skills/tailwind-expert/SKILL.md +62 -5
  129. package/skills/tanstack-query-expert/SKILL.md +5 -4
  130. package/skills/tauri-expert/SKILL.md +5 -4
  131. package/skills/typescript-expert/SKILL.md +5 -4
  132. package/skills/ui-ux-pro-max/SKILL.md +7 -6
  133. package/skills/vector-db-rag-expert/SKILL.md +209 -208
  134. package/skills/vercel-ai-sdk-expert/SKILL.md +226 -181
  135. package/skills/voice-ai-realtime-agent/SKILL.md +243 -242
  136. package/skills/vue-frontend-expert/SKILL.md +5 -4
  137. package/skills/wasm-edge-computing-expert/SKILL.md +3 -2
  138. package/skills/web-3d-graphics-expert/SKILL.md +314 -313
  139. package/skills/web-game-engine-expert/SKILL.md +330 -329
  140. package/skills/web-scraper/SKILL.md +158 -157
  141. package/skills/website-design-cloner/SKILL.md +5 -4
  142. package/skills/webxr-ar-vr-expert/SKILL.md +163 -162
  143. package/skills/wordpress-headless-expert/SKILL.md +145 -144
  144. package/skills/zero-tech-debt-auditor/SKILL.md +115 -0
  145. package/skills/zero-to-prod-orchestrator/SKILL.md +281 -229
  146. package/skills/zero-trust-secret-vault/SKILL.md +3 -2
  147. package/BLUEPRINT.md +0 -319
  148. package/skills/bootstrap-to-modern/SKILL.md +0 -94
  149. package/skills/multiple-entry-points/SKILL.md +0 -91
  150. package/skills/secure-fuzz-testing/SKILL.md +0 -207
  151. package/skills/visual-qa-vision-agent/SKILL.md +0 -71
@@ -1,172 +1,173 @@
1
- ---
2
- name: ai-media-generation-expert
3
- description: "Expert guide for AI image generation (Flux, DALL-E, Stable Diffusion), video generation (Sora, Runway), voice synthesis (ElevenLabs TTS), and speech recognition (Whisper STT) integration / Panduan ahli integrasi AI generasi gambar, video, suara (TTS), dan pengenalan suara (STT)."
4
- author: "Roedy Rustam"
5
- ---
6
-
7
- # AI Media Generation Expert (2026 Edition)
8
-
9
- [English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
10
-
11
- ---
12
-
13
- <a name="english"></a>
14
- ## English
15
-
16
- ### Orchestration & Integration
17
- - **`ai-llm-integration-expert`**: Core LLM patterns and model selection.
18
- - **`file-upload-media-expert`**: Storage, CDN, and media pipeline for generated assets.
19
- - **`sse-websocket-streaming-expert`**: Real-time streaming for progressive image/audio generation.
20
- - **`async-queue-temporal-expert`**: Background job processing for long-running generation tasks.
21
- - **`zero-trust-secret-vault`**: Secure API key management for Replicate, OpenAI, ElevenLabs.
22
-
23
- ### Description
24
- Production-grade guide for integrating AI-powered media generation into web and mobile applications. Covers image generation (Flux 1.1 Pro, SDXL, DALL-E 3), video generation (Sora, Runway Gen-3), text-to-speech (ElevenLabs v3, OpenAI TTS), speech-to-text (Whisper large-v3, Deepgram Nova-3), and voice cloning. Includes async processing patterns, cost optimization, and content safety filtering.
25
-
26
- ### Trigger Conditions
27
- - Integrating AI image generation (Flux, DALL-E, Stable Diffusion, Midjourney API).
28
- - Building text-to-speech or speech-to-text features.
29
- - Implementing voice cloning or AI avatar generation.
30
- - Adding AI video generation or editing capabilities.
31
- - Building media processing pipelines with AI models.
32
-
33
- ---
34
-
35
- ### Core Architecture
36
-
37
- #### 1. Image Generation
38
-
39
- **Provider Selection Matrix:**
40
-
41
- | Provider | Model | Speed | Quality | Cost | Best For |
42
- |----------|-------|-------|---------|------|----------|
43
- | Replicate | Flux 1.1 Pro | ~3s | ★★★★★ | $0.04/img | Photorealism, text rendering |
44
- | Replicate | Flux Schnell | ~1s | ★★★★ | $0.003/img | High-volume, previews |
45
- | OpenAI | DALL-E 3 | ~5s | ★★★★ | $0.04/img | Creative, prompt following |
46
- | Stability | SDXL Turbo | ~2s | ★★★★ | $0.002/img | Self-hosted, customization |
47
-
48
- **Recommendation:** Use **Flux 1.1 Pro** via Replicate for production quality. Use **Flux Schnell** for previews/drafts.
49
-
50
- ```typescript
51
- // Image generation with Replicate (Flux)
52
- import Replicate from 'replicate';
53
-
54
- const replicate = new Replicate({ auth: process.env.REPLICATE_API_TOKEN });
55
-
56
- async function generateImage(prompt: string, options?: {
57
- width?: number;
58
- height?: number;
59
- model?: 'flux-1.1-pro' | 'flux-schnell';
60
- }) {
61
- const model = options?.model ?? 'flux-1.1-pro';
62
- const output = await replicate.run(
63
- `black-forest-labs/${model}`,
64
- {
65
- input: {
66
- prompt,
67
- width: options?.width ?? 1024,
68
- height: options?.height ?? 1024,
69
- num_inference_steps: model === 'flux-schnell' ? 4 : 28,
70
- },
71
- }
72
- );
73
- return output; // URL to generated image
74
- }
75
- ```
76
-
77
- #### 2. Text-to-Speech (TTS)
78
-
79
- **Provider Selection:**
80
-
81
- | Provider | Model | Latency | Expressiveness | Cost |
82
- |----------|-------|---------|----------------|------|
83
- | ElevenLabs | eleven_v3 | ~500ms | ★★★★★ | $0.30/1K chars |
84
- | ElevenLabs | eleven_flash_v2_5 | ~200ms | ★★★★ | $0.15/1K chars |
85
- | OpenAI | tts-1-hd | ~300ms | ★★★ | $0.030/1K chars |
86
-
87
- **Recommendation:** Use **ElevenLabs eleven_v3** for expressive content. Use **eleven_flash_v2_5** for real-time chatbots.
88
-
89
- ```typescript
90
- // ElevenLabs TTS streaming
91
- async function textToSpeech(text: string, voiceId: string): Promise<ReadableStream> {
92
- const response = await fetch(
93
- `https://api.elevenlabs.io/v1/text-to-speech/${voiceId}/stream`,
94
- {
95
- method: 'POST',
96
- headers: {
97
- 'xi-api-key': process.env.ELEVENLABS_API_KEY!,
98
- 'Content-Type': 'application/json',
99
- },
100
- body: JSON.stringify({
101
- text,
102
- model_id: 'eleven_v3',
103
- voice_settings: { stability: 0.5, similarity_boost: 0.75 },
104
- }),
105
- }
106
- );
107
- return response.body!; // Stream audio chunks
108
- }
109
- ```
110
-
111
- #### 3. Speech-to-Text (STT)
112
-
113
- ```typescript
114
- // OpenAI Whisper transcription
115
- import OpenAI from 'openai';
116
-
117
- const openai = new OpenAI();
118
-
119
- async function transcribeAudio(audioFile: File) {
120
- const transcription = await openai.audio.transcriptions.create({
121
- file: audioFile,
122
- model: 'whisper-1',
123
- response_format: 'verbose_json',
124
- timestamp_granularities: ['word', 'segment'],
125
- });
126
- return transcription;
127
- }
128
- ```
129
-
130
- #### 4. Video Generation
131
-
132
- ```typescript
133
- // Replicate video generation (async with webhook)
134
- async function generateVideo(prompt: string) {
135
- const prediction = await replicate.predictions.create({
136
- model: 'minimax/video-01',
137
- input: { prompt, duration: 5 },
138
- webhook: `${process.env.APP_URL}/api/webhooks/replicate`,
139
- webhook_events_filter: ['completed'],
140
- });
141
- return prediction.id; // Poll or wait for webhook
142
- }
143
- ```
144
-
145
- #### 5. Production Patterns
146
-
147
- - **Async Processing:** Always use background jobs (BullMQ, Inngest) for generation tasks >2s.
148
- - **Webhook Architecture:** Use webhooks for Replicate predictions instead of polling.
149
- - **Content Safety:** Implement NSFW filtering (OpenAI Moderation API, Replicate safety checker).
150
- - **Cost Control:** Set per-user daily limits. Cache identical prompts. Use cheaper models for drafts.
151
- - **Storage Pipeline:** Generate → upload to S3/R2 → serve via CDN → store URL in DB.
152
-
153
- ---
154
-
155
- <a name="bahasa-indonesia"></a>
156
- ## Bahasa Indonesia
157
-
158
- ### Integrasi Orkestrasi
159
- - **`ai-llm-integration-expert`**: Pola LLM inti dan pemilihan model.
160
- - **`file-upload-media-expert`**: Penyimpanan, CDN, dan pipeline media untuk aset yang dihasilkan.
161
- - **`sse-websocket-streaming-expert`**: Streaming real-time untuk generasi gambar/audio progresif.
162
- - **`async-queue-temporal-expert`**: Pemrosesan job latar belakang untuk tugas generasi yang lama.
163
-
164
- ### Deskripsi
165
- Panduan tingkat produksi untuk mengintegrasikan generasi media berbasis AI ke dalam aplikasi web dan mobile. Mencakup generasi gambar (Flux 1.1 Pro, SDXL, DALL-E 3), generasi video (Sora, Runway Gen-3), text-to-speech (ElevenLabs v3, OpenAI TTS), speech-to-text (Whisper large-v3, Deepgram Nova-3), dan kloning suara.
166
-
167
- ### Kondisi Pemicu
168
- - Mengintegrasikan generasi gambar AI (Flux, DALL-E, Stable Diffusion).
169
- - Membangun fitur text-to-speech atau speech-to-text.
170
- - Mengimplementasikan kloning suara atau generasi avatar AI.
171
- - Menambahkan kemampuan generasi atau editing video AI.
172
- - Membangun pipeline pemrosesan media dengan model AI.
1
+ ---
2
+ name: ai-media-generation-expert
3
+ description: "Expert guide for AI image generation (Flux, DALL-E, Stable Diffusion), video generation (Sora, Runway), voice synthesis (ElevenLabs TTS), and speech recognition (Whisper STT) integration / Panduan ahli integrasi AI generasi gambar, video, suara (TTS), dan pengenalan suara (STT)."
4
+ author: "Roedy Rustam"
5
+ version: "3.0.0"
6
+ ---
7
+
8
+ # AI Media Generation Expert (2026 Edition)
9
+
10
+ [English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
11
+
12
+ ---
13
+
14
+ <a name="english"></a>
15
+ ## English
16
+
17
+ ### Orchestration & Integration
18
+ - **`ai-llm-integration-expert`**: Core LLM patterns and model selection.
19
+ - **`file-upload-media-expert`**: Storage, CDN, and media pipeline for generated assets.
20
+ - **`sse-websocket-streaming-expert`**: Real-time streaming for progressive image/audio generation.
21
+ - **`async-queue-temporal-expert`**: Background job processing for long-running generation tasks.
22
+ - **`zero-trust-secret-vault`**: Secure API key management for Replicate, OpenAI, ElevenLabs.
23
+
24
+ ### Description
25
+ Production-grade guide for integrating AI-powered media generation into web and mobile applications. Covers image generation (Flux 1.1 Pro, SDXL, DALL-E 3), video generation (Sora, Runway Gen-3), text-to-speech (ElevenLabs v3, OpenAI TTS), speech-to-text (Whisper large-v3, Deepgram Nova-3), and voice cloning. Includes async processing patterns, cost optimization, and content safety filtering.
26
+
27
+ ### Trigger Conditions
28
+ - Integrating AI image generation (Flux, DALL-E, Stable Diffusion, Midjourney API).
29
+ - Building text-to-speech or speech-to-text features.
30
+ - Implementing voice cloning or AI avatar generation.
31
+ - Adding AI video generation or editing capabilities.
32
+ - Building media processing pipelines with AI models.
33
+
34
+ ---
35
+
36
+ ### Core Architecture
37
+
38
+ #### 1. Image Generation
39
+
40
+ **Provider Selection Matrix:**
41
+
42
+ | Provider | Model | Speed | Quality | Cost | Best For |
43
+ |----------|-------|-------|---------|------|----------|
44
+ | Replicate | Flux 1.1 Pro | ~3s | ★★★★★ | $0.04/img | Photorealism, text rendering |
45
+ | Replicate | Flux Schnell | ~1s | ★★★★ | $0.003/img | High-volume, previews |
46
+ | OpenAI | DALL-E 3 | ~5s | ★★★★ | $0.04/img | Creative, prompt following |
47
+ | Stability | SDXL Turbo | ~2s | ★★★★ | $0.002/img | Self-hosted, customization |
48
+
49
+ **Recommendation:** Use **Flux 1.1 Pro** via Replicate for production quality. Use **Flux Schnell** for previews/drafts.
50
+
51
+ ```typescript
52
+ // Image generation with Replicate (Flux)
53
+ import Replicate from 'replicate';
54
+
55
+ const replicate = new Replicate({ auth: process.env.REPLICATE_API_TOKEN });
56
+
57
+ async function generateImage(prompt: string, options?: {
58
+ width?: number;
59
+ height?: number;
60
+ model?: 'flux-1.1-pro' | 'flux-schnell';
61
+ }) {
62
+ const model = options?.model ?? 'flux-1.1-pro';
63
+ const output = await replicate.run(
64
+ `black-forest-labs/${model}`,
65
+ {
66
+ input: {
67
+ prompt,
68
+ width: options?.width ?? 1024,
69
+ height: options?.height ?? 1024,
70
+ num_inference_steps: model === 'flux-schnell' ? 4 : 28,
71
+ },
72
+ }
73
+ );
74
+ return output; // URL to generated image
75
+ }
76
+ ```
77
+
78
+ #### 2. Text-to-Speech (TTS)
79
+
80
+ **Provider Selection:**
81
+
82
+ | Provider | Model | Latency | Expressiveness | Cost |
83
+ |----------|-------|---------|----------------|------|
84
+ | ElevenLabs | eleven_v3 | ~500ms | ★★★★★ | $0.30/1K chars |
85
+ | ElevenLabs | eleven_flash_v2_5 | ~200ms | ★★★★ | $0.15/1K chars |
86
+ | OpenAI | tts-1-hd | ~300ms | ★★★ | $0.030/1K chars |
87
+
88
+ **Recommendation:** Use **ElevenLabs eleven_v3** for expressive content. Use **eleven_flash_v2_5** for real-time chatbots.
89
+
90
+ ```typescript
91
+ // ElevenLabs TTS streaming
92
+ async function textToSpeech(text: string, voiceId: string): Promise<ReadableStream> {
93
+ const response = await fetch(
94
+ `https://api.elevenlabs.io/v1/text-to-speech/${voiceId}/stream`,
95
+ {
96
+ method: 'POST',
97
+ headers: {
98
+ 'xi-api-key': process.env.ELEVENLABS_API_KEY!,
99
+ 'Content-Type': 'application/json',
100
+ },
101
+ body: JSON.stringify({
102
+ text,
103
+ model_id: 'eleven_v3',
104
+ voice_settings: { stability: 0.5, similarity_boost: 0.75 },
105
+ }),
106
+ }
107
+ );
108
+ return response.body!; // Stream audio chunks
109
+ }
110
+ ```
111
+
112
+ #### 3. Speech-to-Text (STT)
113
+
114
+ ```typescript
115
+ // OpenAI Whisper transcription
116
+ import OpenAI from 'openai';
117
+
118
+ const openai = new OpenAI();
119
+
120
+ async function transcribeAudio(audioFile: File) {
121
+ const transcription = await openai.audio.transcriptions.create({
122
+ file: audioFile,
123
+ model: 'whisper-1',
124
+ response_format: 'verbose_json',
125
+ timestamp_granularities: ['word', 'segment'],
126
+ });
127
+ return transcription;
128
+ }
129
+ ```
130
+
131
+ #### 4. Video Generation
132
+
133
+ ```typescript
134
+ // Replicate video generation (async with webhook)
135
+ async function generateVideo(prompt: string) {
136
+ const prediction = await replicate.predictions.create({
137
+ model: 'minimax/video-01',
138
+ input: { prompt, duration: 5 },
139
+ webhook: `${process.env.APP_URL}/api/webhooks/replicate`,
140
+ webhook_events_filter: ['completed'],
141
+ });
142
+ return prediction.id; // Poll or wait for webhook
143
+ }
144
+ ```
145
+
146
+ #### 5. Production Patterns
147
+
148
+ - **Async Processing:** Always use background jobs (BullMQ, Inngest) for generation tasks >2s.
149
+ - **Webhook Architecture:** Use webhooks for Replicate predictions instead of polling.
150
+ - **Content Safety:** Implement NSFW filtering (OpenAI Moderation API, Replicate safety checker).
151
+ - **Cost Control:** Set per-user daily limits. Cache identical prompts. Use cheaper models for drafts.
152
+ - **Storage Pipeline:** Generate → upload to S3/R2 → serve via CDN → store URL in DB.
153
+
154
+ ---
155
+
156
+ <a name="bahasa-indonesia"></a>
157
+ ## Bahasa Indonesia
158
+
159
+ ### Integrasi Orkestrasi
160
+ - **`ai-llm-integration-expert`**: Pola LLM inti dan pemilihan model.
161
+ - **`file-upload-media-expert`**: Penyimpanan, CDN, dan pipeline media untuk aset yang dihasilkan.
162
+ - **`sse-websocket-streaming-expert`**: Streaming real-time untuk generasi gambar/audio progresif.
163
+ - **`async-queue-temporal-expert`**: Pemrosesan job latar belakang untuk tugas generasi yang lama.
164
+
165
+ ### Deskripsi
166
+ Panduan tingkat produksi untuk mengintegrasikan generasi media berbasis AI ke dalam aplikasi web dan mobile. Mencakup generasi gambar (Flux 1.1 Pro, SDXL, DALL-E 3), generasi video (Sora, Runway Gen-3), text-to-speech (ElevenLabs v3, OpenAI TTS), speech-to-text (Whisper large-v3, Deepgram Nova-3), dan kloning suara.
167
+
168
+ ### Kondisi Pemicu
169
+ - Mengintegrasikan generasi gambar AI (Flux, DALL-E, Stable Diffusion).
170
+ - Membangun fitur text-to-speech atau speech-to-text.
171
+ - Mengimplementasikan kloning suara atau generasi avatar AI.
172
+ - Menambahkan kemampuan generasi atau editing video AI.
173
+ - Membangun pipeline pemrosesan media dengan model AI.