vibes-plug 2.14.1 → 3.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (151) hide show
  1. package/.claude/rules/vibes-plug-core.md +5 -0
  2. package/.cursor/rules/vibes-plug-core.mdc +7 -2
  3. package/.cursorrules +8 -2
  4. package/AGENTS.md +23 -2
  5. package/CHANGELOG.md +114 -0
  6. package/CLAUDE.md +10 -3
  7. package/README.md +216 -611
  8. package/bin/vibes.mjs +1104 -0
  9. package/package.json +11 -3
  10. package/plugin.json +4 -3
  11. package/scripts/check-anti-slop.mjs +53 -0
  12. package/scripts/install.js +3 -1
  13. package/scripts/update_skills.js +1 -1
  14. package/scripts/update_skills.mjs +86 -0
  15. package/scripts/validate-skills.mjs +111 -0
  16. package/skills/accessibility-testing-expert/SKILL.md +117 -116
  17. package/skills/affective-computing-emotion-ai/SKILL.md +83 -0
  18. package/skills/agentic-coding-workflow-expert/SKILL.md +297 -0
  19. package/skills/agentic-memory-architect/SKILL.md +52 -0
  20. package/skills/agentic-micro-economy-architect/SKILL.md +92 -0
  21. package/skills/ai-llm-integration-expert/SKILL.md +330 -194
  22. package/skills/ai-media-generation-expert/SKILL.md +173 -172
  23. package/skills/ai-prompt-engineering-expert/SKILL.md +204 -134
  24. package/skills/ai-safety-governance-expert/SKILL.md +223 -0
  25. package/skills/angular-expert/SKILL.md +149 -148
  26. package/skills/anti-slop/SKILL.md +134 -133
  27. package/skills/api-design-expert/SKILL.md +4 -3
  28. package/skills/api-gateway-proxy-expert/SKILL.md +3 -2
  29. package/skills/app-analyzer-optimizer/SKILL.md +4 -3
  30. package/skills/apple-ecosystem-expert/SKILL.md +6 -5
  31. package/skills/astro-framework-expert/SKILL.md +201 -200
  32. package/skills/async-queue-temporal-expert/SKILL.md +218 -217
  33. package/skills/authentication-identity-expert/SKILL.md +174 -173
  34. package/skills/autonomous-red-teamer/SKILL.md +338 -203
  35. package/skills/autonomous-tdd-debugger/SKILL.md +6 -5
  36. package/skills/biome-linter-formatter-expert/SKILL.md +90 -89
  37. package/skills/blockchain-web3-expert/SKILL.md +116 -115
  38. package/skills/brainstorming/SKILL.md +392 -377
  39. package/skills/browser-automation-expert/SKILL.md +260 -222
  40. package/skills/bun-runtime-expert/SKILL.md +5 -4
  41. package/skills/chatbot-messaging-expert/SKILL.md +115 -114
  42. package/skills/ci-cd-devops-architect/SKILL.md +3 -2
  43. package/skills/cloud-hosting-expert/SKILL.md +5 -4
  44. package/skills/coderabbit/SKILL.md +5 -4
  45. package/skills/compliance-gdpr-privacy-expert/SKILL.md +3 -2
  46. package/skills/composable-mach-architect/SKILL.md +338 -0
  47. package/skills/cron-scheduler-expert/SKILL.md +5 -4
  48. package/skills/data-pipeline-etl-expert/SKILL.md +3 -2
  49. package/skills/data-telemetry-expert/SKILL.md +5 -4
  50. package/skills/data-visualization-expert/SKILL.md +155 -154
  51. package/skills/database-orm-expert/SKILL.md +166 -165
  52. package/skills/deep-research-analyst/SKILL.md +182 -136
  53. package/skills/dependency-upgrade-migrator/SKILL.md +11 -10
  54. package/skills/design-system-architect/SKILL.md +4 -3
  55. package/skills/desktop-electron-expert/SKILL.md +129 -128
  56. package/skills/documentation-site-expert/SKILL.md +60 -59
  57. package/skills/doku-mcp-server/SKILL.md +5 -4
  58. package/skills/doku-payment-gateway/SKILL.md +250 -232
  59. package/skills/domain-driven-design-expert/SKILL.md +3 -2
  60. package/skills/e2e-testing-expert/SKILL.md +5 -4
  61. package/skills/ecommerce-expert/SKILL.md +88 -87
  62. package/skills/email-notification-expert/SKILL.md +5 -4
  63. package/skills/ephemeral-generative-ui-architect/SKILL.md +88 -0
  64. package/skills/error-resilience-expert/SKILL.md +14 -13
  65. package/skills/event-driven-architect/SKILL.md +5 -4
  66. package/skills/feature-flag-analytics-expert/SKILL.md +3 -2
  67. package/skills/file-upload-media-expert/SKILL.md +5 -4
  68. package/skills/firebase-security-expert/SKILL.md +5 -4
  69. package/skills/form-validation-expert/SKILL.md +7 -6
  70. package/skills/frontier-ai-models-expert/SKILL.md +116 -0
  71. package/skills/fullstack-expert/SKILL.md +185 -184
  72. package/skills/gemini-agent-booster/SKILL.md +248 -172
  73. package/skills/geospatial-maps-expert/SKILL.md +81 -80
  74. package/skills/global-a11y-i18n-expert/SKILL.md +5 -4
  75. package/skills/glsl-shader-expert/SKILL.md +191 -190
  76. package/skills/go-programming-expert/SKILL.md +5 -4
  77. package/skills/graph-rag-knowledge-expert/SKILL.md +201 -200
  78. package/skills/graphql-apollo-expert/SKILL.md +5 -4
  79. package/skills/headless-cms-expert/SKILL.md +182 -181
  80. package/skills/hig/SKILL.md +5 -4
  81. package/skills/js-backend-expert/SKILL.md +219 -218
  82. package/skills/legacy-code-translator/SKILL.md +6 -5
  83. package/skills/llm-finops-router/SKILL.md +52 -0
  84. package/skills/local-slm-edge-ai-expert/SKILL.md +168 -167
  85. package/skills/logging-error-tracking-expert/SKILL.md +5 -4
  86. package/skills/mcp-server-architect/SKILL.md +315 -307
  87. package/skills/micro-frontend-architect/SKILL.md +5 -4
  88. package/skills/mobile-expo-expert/SKILL.md +5 -4
  89. package/skills/modern-css-native-expert/SKILL.md +190 -189
  90. package/skills/monorepo-architect/SKILL.md +5 -4
  91. package/skills/mpa-orchestrator/SKILL.md +41 -4
  92. package/skills/multi-agent-orchestration/SKILL.md +388 -254
  93. package/skills/mvc-expert/SKILL.md +5 -4
  94. package/skills/n8n-automation-expert/SKILL.md +90 -89
  95. package/skills/nextjs-app-router-expert/SKILL.md +3 -2
  96. package/skills/openapi-swagger-codegen-expert/SKILL.md +4 -3
  97. package/skills/payment-gateway-expert/SKILL.md +131 -128
  98. package/skills/pdf-document-generation-expert/SKILL.md +92 -91
  99. package/skills/performance-web-vitals/SKILL.md +5 -4
  100. package/skills/post-quantum-crypto-migrator/SKILL.md +3 -2
  101. package/skills/prd-architect/SKILL.md +183 -182
  102. package/skills/proactive-background-watcher/SKILL.md +5 -4
  103. package/skills/production-ready-hardener/SKILL.md +10 -9
  104. package/skills/pwa-offline-first-expert/SKILL.md +227 -226
  105. package/skills/pydantic-ai-expert/SKILL.md +162 -161
  106. package/skills/python-programming-expert/SKILL.md +5 -4
  107. package/skills/rate-limit-abuse-prevention/SKILL.md +5 -4
  108. package/skills/realtime-collaboration-expert/SKILL.md +3 -2
  109. package/skills/rich-text-editor-expert/SKILL.md +178 -177
  110. package/skills/rust-programming-expert/SKILL.md +5 -4
  111. package/skills/saas-architect/SKILL.md +155 -154
  112. package/skills/saas-billing/SKILL.md +394 -382
  113. package/skills/saas-multi-tenant/SKILL.md +7 -6
  114. package/skills/scalability-clean-code/SKILL.md +5 -4
  115. package/skills/search-engine-expert/SKILL.md +90 -89
  116. package/skills/self-healing-cloud-orchestrator/SKILL.md +3 -2
  117. package/skills/senior-frontend/SKILL.md +14 -9
  118. package/skills/seo/SKILL.md +4 -4
  119. package/skills/session-memory-manager/SKILL.md +129 -128
  120. package/skills/solidjs-expert/SKILL.md +81 -80
  121. package/skills/spa-orchestrator/SKILL.md +5 -4
  122. package/skills/sse-websocket-streaming-expert/SKILL.md +3 -2
  123. package/skills/state-management-expert/SKILL.md +5 -4
  124. package/skills/supabase-security-expert/SKILL.md +5 -4
  125. package/skills/svelte-sveltekit-expert/SKILL.md +92 -91
  126. package/skills/svg-animation-motion-expert/SKILL.md +3 -2
  127. package/skills/synthetic-data-finetuning-expert/SKILL.md +156 -155
  128. package/skills/tailwind-expert/SKILL.md +62 -5
  129. package/skills/tanstack-query-expert/SKILL.md +5 -4
  130. package/skills/tauri-expert/SKILL.md +5 -4
  131. package/skills/typescript-expert/SKILL.md +5 -4
  132. package/skills/ui-ux-pro-max/SKILL.md +7 -6
  133. package/skills/vector-db-rag-expert/SKILL.md +209 -208
  134. package/skills/vercel-ai-sdk-expert/SKILL.md +226 -181
  135. package/skills/voice-ai-realtime-agent/SKILL.md +243 -242
  136. package/skills/vue-frontend-expert/SKILL.md +5 -4
  137. package/skills/wasm-edge-computing-expert/SKILL.md +3 -2
  138. package/skills/web-3d-graphics-expert/SKILL.md +314 -313
  139. package/skills/web-game-engine-expert/SKILL.md +330 -329
  140. package/skills/web-scraper/SKILL.md +158 -157
  141. package/skills/website-design-cloner/SKILL.md +5 -4
  142. package/skills/webxr-ar-vr-expert/SKILL.md +163 -162
  143. package/skills/wordpress-headless-expert/SKILL.md +145 -144
  144. package/skills/zero-tech-debt-auditor/SKILL.md +115 -0
  145. package/skills/zero-to-prod-orchestrator/SKILL.md +281 -229
  146. package/skills/zero-trust-secret-vault/SKILL.md +3 -2
  147. package/BLUEPRINT.md +0 -319
  148. package/skills/bootstrap-to-modern/SKILL.md +0 -94
  149. package/skills/multiple-entry-points/SKILL.md +0 -91
  150. package/skills/secure-fuzz-testing/SKILL.md +0 -207
  151. package/skills/visual-qa-vision-agent/SKILL.md +0 -71
@@ -1,242 +1,243 @@
1
- ---
2
- name: voice-ai-realtime-agent
3
- description: "Expert guide for Ultra-Low Latency Conversational Voice AI (<300ms), WebRTC bidirectional streaming, OpenAI Realtime API, Gemini Multimodal Live Audio, LiveKit Agents, and Semantic VAD / Panduan ahli AI suara percakapan real-time berlatensi ultra-rendah."
4
- author: "Roedy Rustam"
5
- ---
6
-
7
- # Voice AI Realtime Agent (2026 Edition)
8
-
9
- [English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
10
-
11
- ---
12
-
13
- <a name="english"></a>
14
- ## English
15
-
16
- ### Description
17
- Expert guide for building ultra-low-latency (<300ms), bi-directional conversational voice AI applications. Covers WebRTC, full-duplex WebSocket audio streaming (PCM 24kHz), OpenAI Realtime API, Gemini Multimodal Live API, LiveKit Agents SDK, and smart interruption (barge-in) handling.
18
-
19
- ### Trigger Conditions
20
- - Applications requiring sub-second, spoken conversation with an AI agent.
21
- - Voice customer service bots, verbal copilots, language tutors, and interactive voice assistants.
22
- - Implementation of WebRTC audio streaming, full-duplex WebSocket audio (PCM 24kHz), and Silero VAD.
23
- - Setting up OpenAI Realtime API (`gpt-4o-realtime-preview`) or Gemini Multimodal Live API.
24
-
25
- ---
26
-
27
- ## 1. Core Architecture: Full-Duplex Speech-to-Speech
28
-
29
- Traditional voice pipelines chain STT ➔ LLM ➔ TTS with cumulative latency exceeding 1,200ms–2,500ms. Modern 2026 voice agents use **native speech-to-speech** or **streamable full-duplex WebRTC pipelines** achieving natural, human-like reaction times (~250–350ms).
30
-
31
- ```
32
- User Mic ──► [WebRTC / WebSocket] ──► [VAD: Silero / WebRTC VAD]
33
- │
34
- ▼
35
- User Speaks <── [Audio Output] ◄── [Native Audio Stream / Cartesia] ◄── [OpenAI Realtime / Gemini Live]
36
- │
37
- └── User Interrupts (Barge-in) ──► Instant Buffer Flush & Cancel Audio Frame Emission
38
- ```
39
-
40
- ---
41
-
42
- ## 2. Production Recipe: LiveKit Agents + OpenAI Realtime (Python)
43
-
44
- ```python
45
- # agent.py - Production Voice Agent Worker with LiveKit & OpenAI Realtime
46
- import asyncio
47
- import os
48
- from livekit import rtc
49
- from livekit.agents import (
50
- AutoSubscribe,
51
- JobContext,
52
- JobProcess,
53
- WorkerOptions,
54
- cli,
55
- llm,
56
- )
57
- from livekit.agents.pipeline import VoicePipelineAgent
58
- from livekit.plugins import deepgram, openai, silero
59
-
60
- async def entrypoint(ctx: JobContext):
61
- # Connect to room with audio only to minimize bandwidth & latency
62
- await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY)
63
-
64
- # Wait for the user participant to join
65
- participant = await ctx.wait_for_participant()
66
-
67
- # Define agent instructions and tools
68
- initial_ctx = llm.ChatContext().append(
69
- role="system",
70
- text=(
71
- "You are a helpful, concise voice assistant. "
72
- "Respond naturally in 1-2 short sentences. Never output markdown, bullet points, or emojis."
73
- )
74
- )
75
-
76
- # Realtime Voice Pipeline: Deepgram (STT) + OpenAI (LLM) + Cartesia/OpenAI (TTS)
77
- # Or use native OpenAI Realtime Model: gpt-4o-realtime-preview
78
- agent = VoicePipelineAgent(
79
- vad=silero.VAD.load(
80
- min_speech_duration=0.1,
81
- min_silence_duration=0.3, # Snappy turn-taking
82
- prefix_padding_duration=0.2,
83
- ),
84
- stt=deepgram.STT(model="nova-2", language="id"), # Multi-language support
85
- llm=openai.LLM(model="gpt-4o-mini"),
86
- tts=openai.TTS(voice="alloy"),
87
- chat_ctx=initial_ctx,
88
- allow_interruptions=True, # Barge-in capability
89
- interrupt_speech_duration=0.3, # Immediate cutoff when user talks
90
- )
91
-
92
- agent.start(ctx.room, participant)
93
-
94
- # Greet user immediately
95
- await agent.say("Halo! Ada yang bisa saya bantu hari ini?", now=True)
96
-
97
- if __name__ == "__main__":
98
- cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))
99
- ```
100
-
101
- ---
102
-
103
- ## 3. Production Recipe: Gemini Multimodal Live Audio (TypeScript / Node.js)
104
-
105
- ```typescript
106
- // gemini-live-audio.ts - Bidirectional WebSocket PCM 24kHz
107
- import WebSocket from 'ws';
108
-
109
- interface GeminiAudioConfig {
110
- apiKey: string;
111
- model?: string;
112
- systemInstruction?: string;
113
- }
114
-
115
- export class GeminiVoiceAgent {
116
- private ws: WebSocket | null = null;
117
- private isConnected = false;
118
-
119
- constructor(private config: GeminiAudioConfig) {}
120
-
121
- public async connect(): Promise<void> {
122
- const url = `wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1alpha.GenerativeService.BidiGenerateContent?key=${this.config.apiKey}`;
123
-
124
- this.ws = new WebSocket(url);
125
-
126
- this.ws.on('open', () => {
127
- this.isConnected = true;
128
- this.sendInitialHandshake();
129
- });
130
-
131
- this.ws.on('message', (data: WebSocket.Data) => {
132
- this.handleIncomingAudio(data);
133
- });
134
- }
135
-
136
- private sendInitialHandshake(): void {
137
- const setupMessage = {
138
- setup: {
139
- model: `models/${this.config.model || 'gemini-2.0-flash-exp'}`,
140
- generationConfig: {
141
- responseModalities: ["AUDIO"],
142
- speechConfig: {
143
- voiceConfig: {
144
- prebuiltVoiceConfig: { voiceName: "Puck" }
145
- }
146
- }
147
- },
148
- systemInstruction: {
149
- parts: [{ text: this.config.systemInstruction || "You are a conversational voice agent. Keep answers brief." }]
150
- }
151
- }
152
- };
153
- this.ws?.send(JSON.stringify(setupMessage));
154
- }
155
-
156
- // Stream raw PCM 16-bit 24kHz mono audio from mic
157
- public sendAudioChunk(pcm16Chunk: Buffer): void {
158
- if (!this.isConnected || !this.ws) return;
159
-
160
- const base64Audio = pcm16Chunk.toString('base64');
161
- const msg = {
162
- realtimeInput: {
163
- mediaChunks: [
164
- {
165
- mimeType: "audio/pcm;rate=24000",
166
- data: base64Audio
167
- }
168
- ]
169
- }
170
- };
171
- this.ws.send(JSON.stringify(msg));
172
- }
173
-
174
- private handleIncomingAudio(data: WebSocket.Data): void {
175
- try {
176
- const response = JSON.parse(data.toString());
177
- const parts = response.serverContent?.modelTurn?.parts;
178
- if (parts) {
179
- for (const part of parts) {
180
- if (part.inlineData?.data) {
181
- const pcmBuffer = Buffer.from(part.inlineData.data, 'base64');
182
- this.playAudioSpeaker(pcmBuffer);
183
- }
184
- }
185
- }
186
- } catch {
187
- // Binary PCM frame handler
188
- }
189
- }
190
-
191
- private playAudioSpeaker(pcmChunk: Buffer): void {
192
- // Send to WebRTC audio track or audio output device
193
- }
194
- }
195
- ```
196
-
197
- ---
198
-
199
- ## 4. Key 2026 Performance Guardrails
200
-
201
- 1. **Barge-in Latency Budget (<150ms)**: When the user speaks while the bot is talking, cancel outgoing audio immediately. Do not wait for the LLM to finish streaming its chunk.
202
- 2. **Audio Sample Rates**:
203
- - Mic Input: 16kHz or 24kHz 16-bit Linear PCM Mono.
204
- - Bot Output: 24kHz PCM for crystal-clear natural prosody.
205
- 3. **Turn-Taking Jitter Prevention**: Use minimum silence thresholds between `300ms` and `450ms`. Lower thresholds cause the bot to interrupt users when they pause to think; higher thresholds make the conversation feel robotic.
206
-
207
- ---
208
-
209
- ## Orchestration & Integration
210
-
211
- - **`ai-llm-integration-expert`**: For base LLM prompt routing and function calling during conversation.
212
- - **`realtime-collaboration-expert`**: For syncing WebRTC tracks and room states with client applications.
213
- - **`gemini-agent-booster`**: Connects Gemini 3.x / 2.0 Flash thinking models to live voice agents.
214
- - **`mobile-expo-expert`**: Audio streaming implementation in React Native with `expo-av` and WebRTC shim.
215
-
216
- ---
217
-
218
- <a name="bahasa-indonesia"></a>
219
- ## Bahasa Indonesia
220
-
221
- ### Deskripsi
222
- Panduan ahli untuk membangun aplikasi AI suara percakapan dua arah berlatensi ultra-rendah (<300ms). Mencakup integrasi WebRTC, streaming audio WebSocket full-duplex (PCM 24kHz), OpenAI Realtime API, Gemini Multimodal Live API, LiveKit Agents SDK, dan penanganan interupsi cerdas (*barge-in*).
223
-
224
- ### Kondisi Pemicu
225
- - Kebutuhan interaksi percakapan verbal instan di bawah satu detik dengan agen AI.
226
- - Bot layanan pelanggan berbasis suara, asisten verbal, tutor bahasa interaktif.
227
- - Implementasi streaming audio WebRTC, WebSocket PCM 24kHz dua arah, dan Silero Voice Activity Detection (VAD).
228
- - Konfigurasi OpenAI Realtime API (`gpt-4o-realtime-preview`) atau Gemini Multimodal Live API.
229
-
230
- ### Panduan Inti Arsitektur Suara Real-time
231
- 1. **Full-Duplex Speech-to-Speech**: Mengganti pipeline sekuensial tradisional (STT ➔ LLM ➔ TTS) dengan pipeline streamable WebRTC atau model native speech-to-speech untuk memangkas latensi dari ~2000ms menjadi ~300ms.
232
- 2. **Penanganan Interupsi (Barge-In)**: Saat VAD mendeteksi suara pengguna baru saat bot sedang berbicara, buffer audio keluar harus di-flush dalam waktu <150ms tanpa menunggu LLM menyelesaikan kalimatnya.
233
- 3. **Standar Format Audio**: Input mikrofon PCM 16-bit 16kHz/24kHz Mono, dan output speaker 24kHz untuk intonasi yang alami dan jernih.
234
-
235
- ---
236
-
237
- ## Integrasi Orkestrasi
238
-
239
- - **`ai-llm-integration-expert`**: Routing instruksi sistem dasar dan pemanggilan tool fungsi selama percakapan suara.
240
- - **`realtime-collaboration-expert`**: Sinkronisasi track audio WebRTC dan status room pengguna.
241
- - **`gemini-agent-booster`**: Integrasi model multimodal Gemini 2.0/3.x Flash untuk live audio.
242
- - **`mobile-expo-expert`**: Implementasi audio streaming di React Native menggunakan `expo-av` dan WebRTC shim.
1
+ ---
2
+ name: voice-ai-realtime-agent
3
+ description: "Expert guide for Ultra-Low Latency Conversational Voice AI (<300ms), WebRTC bidirectional streaming, OpenAI Realtime API, Gemini Multimodal Live Audio, LiveKit Agents, and Semantic VAD / Panduan ahli AI suara percakapan real-time berlatensi ultra-rendah."
4
+ author: "Roedy Rustam"
5
+ version: "3.0.0"
6
+ ---
7
+
8
+ # Voice AI Realtime Agent (2026 Edition)
9
+
10
+ [English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
11
+
12
+ ---
13
+
14
+ <a name="english"></a>
15
+ ## English
16
+
17
+ ### Description
18
+ Expert guide for building ultra-low-latency (<300ms), bi-directional conversational voice AI applications. Covers WebRTC, full-duplex WebSocket audio streaming (PCM 24kHz), OpenAI Realtime API, Gemini Multimodal Live API, LiveKit Agents SDK, and smart interruption (barge-in) handling.
19
+
20
+ ### Trigger Conditions
21
+ - Applications requiring sub-second, spoken conversation with an AI agent.
22
+ - Voice customer service bots, verbal copilots, language tutors, and interactive voice assistants.
23
+ - Implementation of WebRTC audio streaming, full-duplex WebSocket audio (PCM 24kHz), and Silero VAD.
24
+ - Setting up OpenAI Realtime API (`gpt-4o-realtime-preview`) or Gemini Multimodal Live API.
25
+
26
+ ---
27
+
28
+ ## 1. Core Architecture: Full-Duplex Speech-to-Speech
29
+
30
+ Traditional voice pipelines chain STT ➔ LLM ➔ TTS with cumulative latency exceeding 1,200ms–2,500ms. Modern 2026 voice agents use **native speech-to-speech** or **streamable full-duplex WebRTC pipelines** achieving natural, human-like reaction times (~250–350ms).
31
+
32
+ ```
33
+ User Mic ──► [WebRTC / WebSocket] ──► [VAD: Silero / WebRTC VAD]
34
+ │
35
+ ▼
36
+ User Speaks <── [Audio Output] ◄── [Native Audio Stream / Cartesia] ◄── [OpenAI Realtime / Gemini Live]
37
+ │
38
+ └── User Interrupts (Barge-in) ──► Instant Buffer Flush & Cancel Audio Frame Emission
39
+ ```
40
+
41
+ ---
42
+
43
+ ## 2. Production Recipe: LiveKit Agents + OpenAI Realtime (Python)
44
+
45
+ ```python
46
+ # agent.py - Production Voice Agent Worker with LiveKit & OpenAI Realtime
47
+ import asyncio
48
+ import os
49
+ from livekit import rtc
50
+ from livekit.agents import (
51
+ AutoSubscribe,
52
+ JobContext,
53
+ JobProcess,
54
+ WorkerOptions,
55
+ cli,
56
+ llm,
57
+ )
58
+ from livekit.agents.pipeline import VoicePipelineAgent
59
+ from livekit.plugins import deepgram, openai, silero
60
+
61
+ async def entrypoint(ctx: JobContext):
62
+ # Connect to room with audio only to minimize bandwidth & latency
63
+ await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY)
64
+
65
+ # Wait for the user participant to join
66
+ participant = await ctx.wait_for_participant()
67
+
68
+ # Define agent instructions and tools
69
+ initial_ctx = llm.ChatContext().append(
70
+ role="system",
71
+ text=(
72
+ "You are a helpful, concise voice assistant. "
73
+ "Respond naturally in 1-2 short sentences. Never output markdown, bullet points, or emojis."
74
+ )
75
+ )
76
+
77
+ # Realtime Voice Pipeline: Deepgram (STT) + OpenAI (LLM) + Cartesia/OpenAI (TTS)
78
+ # Or use native OpenAI Realtime Model: gpt-4o-realtime-preview
79
+ agent = VoicePipelineAgent(
80
+ vad=silero.VAD.load(
81
+ min_speech_duration=0.1,
82
+ min_silence_duration=0.3, # Snappy turn-taking
83
+ prefix_padding_duration=0.2,
84
+ ),
85
+ stt=deepgram.STT(model="nova-2", language="id"), # Multi-language support
86
+ llm=openai.LLM(model="gpt-4o-mini"),
87
+ tts=openai.TTS(voice="alloy"),
88
+ chat_ctx=initial_ctx,
89
+ allow_interruptions=True, # Barge-in capability
90
+ interrupt_speech_duration=0.3, # Immediate cutoff when user talks
91
+ )
92
+
93
+ agent.start(ctx.room, participant)
94
+
95
+ # Greet user immediately
96
+ await agent.say("Halo! Ada yang bisa saya bantu hari ini?", now=True)
97
+
98
+ if __name__ == "__main__":
99
+ cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))
100
+ ```
101
+
102
+ ---
103
+
104
+ ## 3. Production Recipe: Gemini Multimodal Live Audio (TypeScript / Node.js)
105
+
106
+ ```typescript
107
+ // gemini-live-audio.ts - Bidirectional WebSocket PCM 24kHz
108
+ import WebSocket from 'ws';
109
+
110
+ interface GeminiAudioConfig {
111
+ apiKey: string;
112
+ model?: string;
113
+ systemInstruction?: string;
114
+ }
115
+
116
+ export class GeminiVoiceAgent {
117
+ private ws: WebSocket | null = null;
118
+ private isConnected = false;
119
+
120
+ constructor(private config: GeminiAudioConfig) {}
121
+
122
+ public async connect(): Promise<void> {
123
+ const url = `wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1alpha.GenerativeService.BidiGenerateContent?key=${this.config.apiKey}`;
124
+
125
+ this.ws = new WebSocket(url);
126
+
127
+ this.ws.on('open', () => {
128
+ this.isConnected = true;
129
+ this.sendInitialHandshake();
130
+ });
131
+
132
+ this.ws.on('message', (data: WebSocket.Data) => {
133
+ this.handleIncomingAudio(data);
134
+ });
135
+ }
136
+
137
+ private sendInitialHandshake(): void {
138
+ const setupMessage = {
139
+ setup: {
140
+ model: `models/${this.config.model || 'gemini-2.0-flash-exp'}`,
141
+ generationConfig: {
142
+ responseModalities: ["AUDIO"],
143
+ speechConfig: {
144
+ voiceConfig: {
145
+ prebuiltVoiceConfig: { voiceName: "Puck" }
146
+ }
147
+ }
148
+ },
149
+ systemInstruction: {
150
+ parts: [{ text: this.config.systemInstruction || "You are a conversational voice agent. Keep answers brief." }]
151
+ }
152
+ }
153
+ };
154
+ this.ws?.send(JSON.stringify(setupMessage));
155
+ }
156
+
157
+ // Stream raw PCM 16-bit 24kHz mono audio from mic
158
+ public sendAudioChunk(pcm16Chunk: Buffer): void {
159
+ if (!this.isConnected || !this.ws) return;
160
+
161
+ const base64Audio = pcm16Chunk.toString('base64');
162
+ const msg = {
163
+ realtimeInput: {
164
+ mediaChunks: [
165
+ {
166
+ mimeType: "audio/pcm;rate=24000",
167
+ data: base64Audio
168
+ }
169
+ ]
170
+ }
171
+ };
172
+ this.ws.send(JSON.stringify(msg));
173
+ }
174
+
175
+ private handleIncomingAudio(data: WebSocket.Data): void {
176
+ try {
177
+ const response = JSON.parse(data.toString());
178
+ const parts = response.serverContent?.modelTurn?.parts;
179
+ if (parts) {
180
+ for (const part of parts) {
181
+ if (part.inlineData?.data) {
182
+ const pcmBuffer = Buffer.from(part.inlineData.data, 'base64');
183
+ this.playAudioSpeaker(pcmBuffer);
184
+ }
185
+ }
186
+ }
187
+ } catch {
188
+ // Binary PCM frame handler
189
+ }
190
+ }
191
+
192
+ private playAudioSpeaker(pcmChunk: Buffer): void {
193
+ // Send to WebRTC audio track or audio output device
194
+ }
195
+ }
196
+ ```
197
+
198
+ ---
199
+
200
+ ## 4. Key 2026 Performance Guardrails
201
+
202
+ 1. **Barge-in Latency Budget (<150ms)**: When the user speaks while the bot is talking, cancel outgoing audio immediately. Do not wait for the LLM to finish streaming its chunk.
203
+ 2. **Audio Sample Rates**:
204
+ - Mic Input: 16kHz or 24kHz 16-bit Linear PCM Mono.
205
+ - Bot Output: 24kHz PCM for crystal-clear natural prosody.
206
+ 3. **Turn-Taking Jitter Prevention**: Use minimum silence thresholds between `300ms` and `450ms`. Lower thresholds cause the bot to interrupt users when they pause to think; higher thresholds make the conversation feel robotic.
207
+
208
+ ---
209
+
210
+ ## Orchestration & Integration
211
+
212
+ - **`ai-llm-integration-expert`**: For base LLM prompt routing and function calling during conversation.
213
+ - **`realtime-collaboration-expert`**: For syncing WebRTC tracks and room states with client applications.
214
+ - **`gemini-agent-booster`**: Connects Gemini 3.x / 2.0 Flash thinking models to live voice agents.
215
+ - **`mobile-expo-expert`**: Audio streaming implementation in React Native with `expo-av` and WebRTC shim.
216
+
217
+ ---
218
+
219
+ <a name="bahasa-indonesia"></a>
220
+ ## Bahasa Indonesia
221
+
222
+ ### Deskripsi
223
+ Panduan ahli untuk membangun aplikasi AI suara percakapan dua arah berlatensi ultra-rendah (<300ms). Mencakup integrasi WebRTC, streaming audio WebSocket full-duplex (PCM 24kHz), OpenAI Realtime API, Gemini Multimodal Live API, LiveKit Agents SDK, dan penanganan interupsi cerdas (*barge-in*).
224
+
225
+ ### Kondisi Pemicu
226
+ - Kebutuhan interaksi percakapan verbal instan di bawah satu detik dengan agen AI.
227
+ - Bot layanan pelanggan berbasis suara, asisten verbal, tutor bahasa interaktif.
228
+ - Implementasi streaming audio WebRTC, WebSocket PCM 24kHz dua arah, dan Silero Voice Activity Detection (VAD).
229
+ - Konfigurasi OpenAI Realtime API (`gpt-4o-realtime-preview`) atau Gemini Multimodal Live API.
230
+
231
+ ### Panduan Inti Arsitektur Suara Real-time
232
+ 1. **Full-Duplex Speech-to-Speech**: Mengganti pipeline sekuensial tradisional (STT ➔ LLM ➔ TTS) dengan pipeline streamable WebRTC atau model native speech-to-speech untuk memangkas latensi dari ~2000ms menjadi ~300ms.
233
+ 2. **Penanganan Interupsi (Barge-In)**: Saat VAD mendeteksi suara pengguna baru saat bot sedang berbicara, buffer audio keluar harus di-flush dalam waktu <150ms tanpa menunggu LLM menyelesaikan kalimatnya.
234
+ 3. **Standar Format Audio**: Input mikrofon PCM 16-bit 16kHz/24kHz Mono, dan output speaker 24kHz untuk intonasi yang alami dan jernih.
235
+
236
+ ---
237
+
238
+ ## Integrasi Orkestrasi
239
+
240
+ - **`ai-llm-integration-expert`**: Routing instruksi sistem dasar dan pemanggilan tool fungsi selama percakapan suara.
241
+ - **`realtime-collaboration-expert`**: Sinkronisasi track audio WebRTC dan status room pengguna.
242
+ - **`gemini-agent-booster`**: Integrasi model multimodal Gemini 2.0/3.x Flash untuk live audio.
243
+ - **`mobile-expo-expert`**: Implementasi audio streaming di React Native menggunakan `expo-av` dan WebRTC shim.
@@ -1,7 +1,8 @@
1
1
  ---
2
2
  name: vue-frontend-expert
3
3
  description: "Expert guide for Vue 3 (Composition API), Nuxt 3, and Pinia. Covers advanced reactive state management, `<script setup>` syntax, Vue Router, VueUse, and SPA/SSR architectural patterns in English and Indonesian."
4
- author: "Roedy Rustam"
4
+ author: "Roedy Rustam"
5
+ version: "3.0.0"
5
6
  ---
6
7
 
7
8
  # Vue Frontend Expert (Vue 3 / Nuxt 3)
@@ -14,7 +15,7 @@ author: "Roedy Rustam"
14
15
  ## English
15
16
 
16
17
  ### Orchestration & Integration
17
- Connects and orchestrates with relevant domain skills like `brainstorming`, `zero-to-prod-orchestrator`, and `project-context-mapper` to ensure cohesive execution.
18
+ Connects and orchestrates with relevant domain skills like `brainstorming`, `zero-to-prod-orchestrator`, and `session-memory-manager` to ensure cohesive execution.
18
19
 
19
20
  ### Description
20
21
  Production-grade guidance for building highly reactive and scalable frontend applications using **Vue 3 (Composition API)** and **Nuxt 3**. Covers the latest ecosystem tools including **Pinia** for state management, **VueUse** for composables, and **Tailwind CSS v4** integration.
@@ -111,7 +112,7 @@ This skill should be referenced by the following orchestrators:
111
112
  ## Bahasa Indonesia
112
113
 
113
114
  ### Integrasi Orkestrasi
114
- Terhubung dan mengorkestrasi skill domain yang relevan seperti `brainstorming`, `zero-to-prod-orchestrator`, dan `project-context-mapper` untuk memastikan eksekusi yang kohesif.
115
+ Terhubung dan mengorkestrasi skill domain yang relevan seperti `brainstorming`, `zero-to-prod-orchestrator`, dan `session-memory-manager` untuk memastikan eksekusi yang kohesif.
115
116
 
116
117
  ### Deskripsi
117
118
  Panduan tingkat produksi untuk membangun aplikasi frontend reaktif dan terukur menggunakan **Vue 3 (Composition API)** dan **Nuxt 3**. Mencakup pengelolaan state dengan **Pinia**, utilitas **VueUse**, dan integrasi **Tailwind CSS v4**.
@@ -129,4 +130,4 @@ Aktifkan skill ini ketika pengguna sedang:
129
130
  - **Pinia:** Gunakan Pinia dengan pola *Setup Store* (mirip Composition API) alih-alih pola Options (state, getters, actions).
130
131
  - **Nuxt 3 Fetching:** Gunakan `useFetch` atau `useAsyncData` di dalam komponen Nuxt untuk pengambilan data saat SSR, bukan `onMounted` dengan `fetch` biasa.
131
132
  - **Reaktivitas:** Gunakan `ref` untuk nilai primitif (string, number) dan `reactive` untuk objek bersarang yang kompleks. Utamakan `computed` daripada `watch` untuk state turunan.
132
- - **VueUse:** Jangan menulis fungsi utilitas dari nol jika sudah ada di *library* VueUse (contoh: `useIntersectionObserver`, `useLocalStorage`).
133
+ - **VueUse:** Jangan menulis fungsi utilitas dari nol jika sudah ada di *library* VueUse (contoh: `useIntersectionObserver`, `useLocalStorage`).
@@ -1,7 +1,8 @@
1
1
  ---
2
2
  name: wasm-edge-computing-expert
3
3
  description: "Expert guide for WebAssembly (WASM) and Edge Computing. Covers WASI preview 2, Spin/Fermyon, Cloudflare Workers WASM, and high-performance browser computing / Panduan ahli untuk WebAssembly (WASM) dan Edge Computing. Mencakup WASI preview 2, Spin/Fermyon, Cloudflare Workers WASM, dan komputasi performa tinggi di browser."
4
- author: "Roedy Rustam"
4
+ author: "Roedy Rustam"
5
+ version: "3.0.0"
5
6
  ---
6
7
 
7
8
  # WASM & Edge Computing Expert
@@ -94,4 +95,4 @@ Untuk performa multi-threading sejati di browser menggunakan WASM, gunakan `Shar
94
95
  ## Integrasi Orkestrasi
95
96
  - Memperkuat `api-gateway-proxy-expert` untuk logika kustom di edge.
96
97
  - Terintegrasi dengan `rust-programming-expert` untuk kompilasi modul.
97
- - Melengkapi `performance-web-vitals` dengan memindahkan komputasi berat dari *main thread* JavaScript.
98
+ - Melengkapi `performance-web-vitals` dengan memindahkan komputasi berat dari *main thread* JavaScript.