vibes-plug 2.11.0 → 3.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (181) hide show
  1. package/.claude/rules/vibes-plug-core.md +5 -0
  2. package/.cursor/rules/vibes-plug-core.mdc +8 -3
  3. package/.cursorrules +9 -3
  4. package/AGENTS.md +25 -4
  5. package/CHANGELOG.md +151 -0
  6. package/CLAUDE.md +15 -8
  7. package/README.md +216 -641
  8. package/bin/vibes.mjs +1104 -0
  9. package/index.js +1 -1
  10. package/package.json +11 -3
  11. package/plugin.json +4 -3
  12. package/scripts/check-anti-slop.js +53 -0
  13. package/scripts/check-anti-slop.mjs +53 -0
  14. package/scripts/generate_swarm_gif.py +2 -2
  15. package/scripts/install.js +3 -1
  16. package/scripts/update_skills.js +1 -1
  17. package/scripts/update_skills.mjs +86 -0
  18. package/scripts/validate-skills.mjs +111 -0
  19. package/skills/accessibility-testing-expert/SKILL.md +117 -116
  20. package/skills/affective-computing-emotion-ai/SKILL.md +83 -0
  21. package/skills/agentic-coding-workflow-expert/SKILL.md +297 -0
  22. package/skills/agentic-memory-architect/SKILL.md +52 -0
  23. package/skills/agentic-micro-economy-architect/SKILL.md +92 -0
  24. package/skills/ai-llm-integration-expert/SKILL.md +330 -187
  25. package/skills/ai-media-generation-expert/SKILL.md +173 -172
  26. package/skills/ai-prompt-engineering-expert/SKILL.md +170 -50
  27. package/skills/ai-safety-governance-expert/SKILL.md +223 -0
  28. package/skills/angular-expert/SKILL.md +149 -148
  29. package/skills/anti-slop/SKILL.md +134 -0
  30. package/skills/api-design-expert/SKILL.md +4 -3
  31. package/skills/api-gateway-proxy-expert/SKILL.md +3 -2
  32. package/skills/app-analyzer-optimizer/SKILL.md +4 -3
  33. package/skills/apple-ecosystem-expert/SKILL.md +6 -5
  34. package/skills/astro-framework-expert/SKILL.md +201 -200
  35. package/skills/async-queue-temporal-expert/SKILL.md +218 -240
  36. package/skills/authentication-identity-expert/SKILL.md +79 -184
  37. package/skills/autonomous-red-teamer/SKILL.md +338 -203
  38. package/skills/autonomous-tdd-debugger/SKILL.md +6 -5
  39. package/skills/biome-linter-formatter-expert/SKILL.md +90 -89
  40. package/skills/blockchain-web3-expert/SKILL.md +116 -115
  41. package/skills/brainstorming/SKILL.md +392 -377
  42. package/skills/browser-automation-expert/SKILL.md +260 -222
  43. package/skills/bun-runtime-expert/SKILL.md +5 -4
  44. package/skills/chatbot-messaging-expert/SKILL.md +115 -114
  45. package/skills/ci-cd-devops-architect/SKILL.md +3 -2
  46. package/skills/cloud-hosting-expert/SKILL.md +5 -4
  47. package/skills/coderabbit/SKILL.md +5 -4
  48. package/skills/compliance-gdpr-privacy-expert/SKILL.md +3 -2
  49. package/skills/composable-mach-architect/SKILL.md +338 -0
  50. package/skills/cron-scheduler-expert/SKILL.md +5 -4
  51. package/skills/data-pipeline-etl-expert/SKILL.md +3 -2
  52. package/skills/data-telemetry-expert/SKILL.md +5 -4
  53. package/skills/data-visualization-expert/SKILL.md +155 -154
  54. package/skills/database-orm-expert/SKILL.md +102 -240
  55. package/skills/deep-research-analyst/SKILL.md +182 -0
  56. package/skills/dependency-upgrade-migrator/SKILL.md +11 -10
  57. package/skills/design-system-architect/SKILL.md +34 -3
  58. package/skills/desktop-electron-expert/SKILL.md +129 -128
  59. package/skills/documentation-site-expert/SKILL.md +60 -59
  60. package/skills/doku-mcp-server/SKILL.md +5 -4
  61. package/skills/doku-payment-gateway/SKILL.md +250 -232
  62. package/skills/domain-driven-design-expert/SKILL.md +3 -2
  63. package/skills/e2e-testing-expert/SKILL.md +5 -4
  64. package/skills/ecommerce-expert/SKILL.md +88 -87
  65. package/skills/email-notification-expert/SKILL.md +35 -7
  66. package/skills/ephemeral-generative-ui-architect/SKILL.md +88 -0
  67. package/skills/error-resilience-expert/SKILL.md +26 -4
  68. package/skills/event-driven-architect/SKILL.md +5 -4
  69. package/skills/feature-flag-analytics-expert/SKILL.md +3 -2
  70. package/skills/file-upload-media-expert/SKILL.md +5 -4
  71. package/skills/firebase-security-expert/SKILL.md +5 -4
  72. package/skills/form-validation-expert/SKILL.md +7 -6
  73. package/skills/frontier-ai-models-expert/SKILL.md +116 -0
  74. package/skills/fullstack-expert/SKILL.md +68 -144
  75. package/skills/gemini-agent-booster/SKILL.md +248 -172
  76. package/skills/geospatial-maps-expert/SKILL.md +81 -80
  77. package/skills/global-a11y-i18n-expert/SKILL.md +5 -4
  78. package/skills/glsl-shader-expert/SKILL.md +155 -71
  79. package/skills/go-programming-expert/SKILL.md +5 -4
  80. package/skills/graph-rag-knowledge-expert/SKILL.md +201 -159
  81. package/skills/graphql-apollo-expert/SKILL.md +5 -4
  82. package/skills/headless-cms-expert/SKILL.md +182 -181
  83. package/skills/hig/SKILL.md +5 -4
  84. package/skills/js-backend-expert/SKILL.md +219 -218
  85. package/skills/legacy-code-translator/SKILL.md +6 -5
  86. package/skills/llm-finops-router/SKILL.md +52 -0
  87. package/skills/local-slm-edge-ai-expert/SKILL.md +168 -167
  88. package/skills/logging-error-tracking-expert/SKILL.md +5 -4
  89. package/skills/mcp-server-architect/SKILL.md +316 -294
  90. package/skills/micro-frontend-architect/SKILL.md +5 -4
  91. package/skills/mobile-expo-expert/SKILL.md +5 -4
  92. package/skills/modern-css-native-expert/SKILL.md +190 -189
  93. package/skills/monorepo-architect/SKILL.md +5 -4
  94. package/skills/mpa-orchestrator/SKILL.md +41 -4
  95. package/skills/multi-agent-orchestration/SKILL.md +388 -254
  96. package/skills/mvc-expert/SKILL.md +5 -4
  97. package/skills/n8n-automation-expert/SKILL.md +90 -89
  98. package/skills/nextjs-app-router-expert/SKILL.md +3 -2
  99. package/skills/openapi-swagger-codegen-expert/SKILL.md +4 -3
  100. package/skills/payment-gateway-expert/SKILL.md +131 -128
  101. package/skills/pdf-document-generation-expert/SKILL.md +92 -91
  102. package/skills/performance-web-vitals/SKILL.md +5 -4
  103. package/skills/post-quantum-crypto-migrator/SKILL.md +3 -2
  104. package/skills/prd-architect/SKILL.md +85 -109
  105. package/skills/proactive-background-watcher/SKILL.md +5 -4
  106. package/skills/production-ready-hardener/SKILL.md +25 -27
  107. package/skills/pwa-offline-first-expert/SKILL.md +227 -185
  108. package/skills/pydantic-ai-expert/SKILL.md +162 -0
  109. package/skills/python-programming-expert/SKILL.md +5 -4
  110. package/skills/rate-limit-abuse-prevention/SKILL.md +5 -4
  111. package/skills/realtime-collaboration-expert/SKILL.md +3 -2
  112. package/skills/rich-text-editor-expert/SKILL.md +178 -177
  113. package/skills/rust-programming-expert/SKILL.md +5 -4
  114. package/skills/saas-architect/SKILL.md +155 -0
  115. package/skills/saas-billing/SKILL.md +394 -382
  116. package/skills/saas-multi-tenant/SKILL.md +7 -6
  117. package/skills/scalability-clean-code/SKILL.md +5 -4
  118. package/skills/search-engine-expert/SKILL.md +90 -89
  119. package/skills/self-healing-cloud-orchestrator/SKILL.md +3 -2
  120. package/skills/senior-frontend/SKILL.md +21 -18
  121. package/skills/senior-frontend/scripts/frontend_scaffolder.py +1 -1
  122. package/skills/seo/SKILL.md +4 -4
  123. package/skills/session-memory-manager/SKILL.md +129 -0
  124. package/skills/solidjs-expert/SKILL.md +81 -80
  125. package/skills/spa-orchestrator/SKILL.md +5 -4
  126. package/skills/sse-websocket-streaming-expert/SKILL.md +3 -2
  127. package/skills/state-management-expert/SKILL.md +5 -4
  128. package/skills/supabase-security-expert/SKILL.md +5 -4
  129. package/skills/svelte-sveltekit-expert/SKILL.md +92 -91
  130. package/skills/svg-animation-motion-expert/SKILL.md +3 -2
  131. package/skills/synthetic-data-finetuning-expert/SKILL.md +156 -0
  132. package/skills/tailwind-expert/SKILL.md +62 -5
  133. package/skills/tanstack-query-expert/SKILL.md +5 -4
  134. package/skills/tauri-expert/SKILL.md +5 -4
  135. package/skills/typescript-expert/SKILL.md +5 -4
  136. package/skills/ui-ux-pro-max/SKILL.md +7 -4
  137. package/skills/vector-db-rag-expert/SKILL.md +209 -208
  138. package/skills/vercel-ai-sdk-expert/SKILL.md +226 -0
  139. package/skills/voice-ai-realtime-agent/SKILL.md +243 -202
  140. package/skills/vue-frontend-expert/SKILL.md +5 -4
  141. package/skills/wasm-edge-computing-expert/SKILL.md +3 -2
  142. package/skills/web-3d-graphics-expert/SKILL.md +259 -82
  143. package/skills/web-game-engine-expert/SKILL.md +278 -50
  144. package/skills/web-scraper/SKILL.md +158 -157
  145. package/skills/website-design-cloner/SKILL.md +5 -4
  146. package/skills/webxr-ar-vr-expert/SKILL.md +105 -65
  147. package/skills/wordpress-headless-expert/SKILL.md +145 -144
  148. package/skills/zero-tech-debt-auditor/SKILL.md +115 -0
  149. package/skills/zero-to-prod-orchestrator/SKILL.md +281 -227
  150. package/skills/zero-trust-secret-vault/SKILL.md +3 -2
  151. package/BLUEPRINT.md +0 -309
  152. package/skills/ai-cost-token-optimizer/SKILL.md +0 -82
  153. package/skills/ai-evals-benchmark-expert/SKILL.md +0 -188
  154. package/skills/asisten-ramah/SKILL.md +0 -47
  155. package/skills/auto-doc-updater/SKILL.md +0 -220
  156. package/skills/autonomous-chaos-monkey/SKILL.md +0 -63
  157. package/skills/background-jobs-queue-expert/SKILL.md +0 -235
  158. package/skills/bootstrap-to-modern/SKILL.md +0 -94
  159. package/skills/database-migration-versioning-expert/SKILL.md +0 -90
  160. package/skills/edge-serverless-db-expert/SKILL.md +0 -99
  161. package/skills/mcp-client-orchestrator/SKILL.md +0 -76
  162. package/skills/mobile-push-notification-expert/SKILL.md +0 -71
  163. package/skills/monday-design-aesthetic/SKILL.md +0 -73
  164. package/skills/multiple-entry-points/SKILL.md +0 -91
  165. package/skills/project-context-mapper/SKILL.md +0 -85
  166. package/skills/saas-mvp-launcher/SKILL.md +0 -260
  167. package/skills/saas-transformer/SKILL.md +0 -500
  168. package/skills/saas-transformer/references/billing_integration_guide.md +0 -401
  169. package/skills/secure-fuzz-testing/SKILL.md +0 -207
  170. package/skills/self-evolving-memory-graph/SKILL.md +0 -91
  171. package/skills/session-context-loader/SKILL.md +0 -83
  172. package/skills/session-handoff-resume/SKILL.md +0 -164
  173. package/skills/skill-baru/SKILL.md +0 -178
  174. package/skills/supabase-migration/SKILL.md +0 -91
  175. package/skills/token-saver/SKILL.md +0 -119
  176. package/skills/ui-components-expert/SKILL.md +0 -166
  177. package/skills/vibe-code-gardener/SKILL.md +0 -181
  178. package/skills/visual-qa-vision-agent/SKILL.md +0 -71
  179. /package/skills/{saas-transformer → saas-architect}/references/feature_gating_patterns.md +0 -0
  180. /package/skills/{saas-transformer → saas-architect}/references/saas_transformation_checklist.md +0 -0
  181. /package/skills/{saas-transformer → saas-architect}/scripts/saas_transformation_scanner.py +0 -0
@@ -1,202 +1,243 @@
1
- ---
2
- name: voice-ai-realtime-agent
3
- description: "Expert guide for Ultra-Low Latency Conversational Voice AI (<300ms), WebRTC bidirectional streaming, OpenAI Realtime API, Gemini Multimodal Live Audio, LiveKit Agents, and Semantic VAD / Panduan ahli AI suara percakapan real-time berlatensi ultra-rendah."
4
- author: "Roedy Rustam"
5
- ---
6
-
7
- # Voice AI Realtime Agent (2026 Edition)
8
-
9
- Expert guide for building ultra-low-latency (<300ms), bi-directional conversational voice AI applications. Covers WebRTC, full-duplex WebSocket audio streaming (PCM 24kHz), OpenAI Realtime API, Gemini Multimodal Live API, LiveKit Agents SDK, and smart interruption (barge-in) handling.
10
-
11
- *Panduan ahli untuk membangun aplikasi AI suara percakapan dua arah berlatensi ultra-rendah (<300ms) menggunakan WebRTC, WebSocket full-duplex, OpenAI Realtime API, Gemini Multimodal Live API, LiveKit Agents SDK, dan penanganan interupsi (barge-in).*
12
-
13
- ---
14
-
15
- ## 1. Core Architecture: Full-Duplex Speech-to-Speech
16
-
17
- Traditional voice pipelines chain STT ➔ LLM ➔ TTS with cumulative latency exceeding 1,200ms–2,500ms. Modern 2026 voice agents use **native speech-to-speech** or **streamable full-duplex WebRTC pipelines** achieving natural, human-like reaction times (~250–350ms).
18
-
19
- ```
20
- User Mic ──► [WebRTC / WebSocket] ──► [VAD: Silero / WebRTC VAD]
21
- │
22
- ▼
23
- User Speaks <── [Audio Output] ◄── [Native Audio Stream / Cartesia] ◄── [OpenAI Realtime / Gemini Live]
24
- │
25
- └── User Interrupts (Barge-in) ──► Instant Buffer Flush & Cancel Audio Frame Emission
26
- ```
27
-
28
- ---
29
-
30
- ## 2. Production Recipe: LiveKit Agents + OpenAI Realtime (Python)
31
-
32
- ```python
33
- # agent.py - Production Voice Agent Worker with LiveKit & OpenAI Realtime
34
- import asyncio
35
- import os
36
- from livekit import rtc
37
- from livekit.agents import (
38
- AutoSubscribe,
39
- JobContext,
40
- JobProcess,
41
- WorkerOptions,
42
- cli,
43
- llm,
44
- )
45
- from livekit.agents.pipeline import VoicePipelineAgent
46
- from livekit.plugins import deepgram, openai, silero
47
-
48
- async def entrypoint(ctx: JobContext):
49
- # Connect to room with audio only to minimize bandwidth & latency
50
- await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY)
51
-
52
- # Wait for the user participant to join
53
- participant = await ctx.wait_for_participant()
54
-
55
- # Define agent instructions and tools
56
- initial_ctx = llm.ChatContext().append(
57
- role="system",
58
- text=(
59
- "You are a helpful, concise voice assistant. "
60
- "Respond naturally in 1-2 short sentences. Never output markdown, bullet points, or emojis."
61
- )
62
- )
63
-
64
- # Realtime Voice Pipeline: Deepgram (STT) + OpenAI (LLM) + Cartesia/OpenAI (TTS)
65
- # Or use native OpenAI Realtime Model: gpt-4o-realtime-preview
66
- agent = VoicePipelineAgent(
67
- vad=silero.VAD.load(
68
- min_speech_duration=0.1,
69
- min_silence_duration=0.3, # Snappy turn-taking
70
- prefix_padding_duration=0.2,
71
- ),
72
- stt=deepgram.STT(model="nova-2", language="id"), # Multi-language support
73
- llm=openai.LLM(model="gpt-4o-mini"),
74
- tts=openai.TTS(voice="alloy"),
75
- chat_ctx=initial_ctx,
76
- allow_interruptions=True, # Barge-in capability
77
- interrupt_speech_duration=0.3, # Immediate cutoff when user talks
78
- )
79
-
80
- agent.start(ctx.room, participant)
81
-
82
- # Greet user immediately
83
- await agent.say("Halo! Ada yang bisa saya bantu hari ini?", now=True)
84
-
85
- if __name__ == "__main__":
86
- cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))
87
- ```
88
-
89
- ---
90
-
91
- ## 3. Production Recipe: Gemini Multimodal Live Audio (TypeScript / Node.js)
92
-
93
- ```typescript
94
- // gemini-live-audio.ts - Bidirectional WebSocket PCM 24kHz
95
- import WebSocket from 'ws';
96
-
97
- interface GeminiAudioConfig {
98
- apiKey: string;
99
- model?: string;
100
- systemInstruction?: string;
101
- }
102
-
103
- export class GeminiVoiceAgent {
104
- private ws: WebSocket | null = null;
105
- private isConnected = false;
106
-
107
- constructor(private config: GeminiAudioConfig) {}
108
-
109
- public async connect(): Promise<void> {
110
- const url = `wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1alpha.GenerativeService.BidiGenerateContent?key=${this.config.apiKey}`;
111
-
112
- this.ws = new WebSocket(url);
113
-
114
- this.ws.on('open', () => {
115
- this.isConnected = true;
116
- this.sendInitialHandshake();
117
- });
118
-
119
- this.ws.on('message', (data: WebSocket.Data) => {
120
- this.handleIncomingAudio(data);
121
- });
122
- }
123
-
124
- private sendInitialHandshake(): void {
125
- const setupMessage = {
126
- setup: {
127
- model: `models/${this.config.model || 'gemini-2.0-flash-exp'}`,
128
- generationConfig: {
129
- responseModalities: ["AUDIO"],
130
- speechConfig: {
131
- voiceConfig: {
132
- prebuiltVoiceConfig: { voiceName: "Puck" }
133
- }
134
- }
135
- },
136
- systemInstruction: {
137
- parts: [{ text: this.config.systemInstruction || "You are a conversational voice agent. Keep answers brief." }]
138
- }
139
- }
140
- };
141
- this.ws?.send(JSON.stringify(setupMessage));
142
- }
143
-
144
- // Stream raw PCM 16-bit 24kHz mono audio from mic
145
- public sendAudioChunk(pcm16Chunk: Buffer): void {
146
- if (!this.isConnected || !this.ws) return;
147
-
148
- const base64Audio = pcm16Chunk.toString('base64');
149
- const msg = {
150
- realtimeInput: {
151
- mediaChunks: [
152
- {
153
- mimeType: "audio/pcm;rate=24000",
154
- data: base64Audio
155
- }
156
- ]
157
- }
158
- };
159
- this.ws.send(JSON.stringify(msg));
160
- }
161
-
162
- private handleIncomingAudio(data: WebSocket.Data): void {
163
- try {
164
- const response = JSON.parse(data.toString());
165
- const parts = response.serverContent?.modelTurn?.parts;
166
- if (parts) {
167
- for (const part of parts) {
168
- if (part.inlineData?.data) {
169
- const pcmBuffer = Buffer.from(part.inlineData.data, 'base64');
170
- this.playAudioSpeaker(pcmBuffer);
171
- }
172
- }
173
- }
174
- } catch {
175
- // Binary PCM frame handler
176
- }
177
- }
178
-
179
- private playAudioSpeaker(pcmChunk: Buffer): void {
180
- // Send to WebRTC audio track or audio output device
181
- }
182
- }
183
- ```
184
-
185
- ---
186
-
187
- ## 4. Key 2026 Performance Guardrails
188
-
189
- 1. **Barge-in Latency Budget (<150ms)**: When the user speaks while the bot is talking, cancel outgoing audio immediately. Do not wait for the LLM to finish streaming its chunk.
190
- 2. **Audio Sample Rates**:
191
- - Mic Input: 16kHz or 24kHz 16-bit Linear PCM Mono.
192
- - Bot Output: 24kHz PCM for crystal-clear natural prosody.
193
- 3. **Turn-Taking Jitter Prevention**: Use minimum silence thresholds between `300ms` and `450ms`. Lower thresholds cause the bot to interrupt users when they pause to think; higher thresholds make the conversation feel robotic.
194
-
195
- ---
196
-
197
- ## Orchestration & Integration
198
-
199
- - **`ai-llm-integration-expert`**: For base LLM prompt routing and function calling during conversation.
200
- - **`realtime-collaboration-expert`**: For syncing WebRTC tracks and room states with client applications.
201
- - **`gemini-agent-booster`**: Connects Gemini 3.x / 2.0 Flash thinking models to live voice agents.
202
- - **`mobile-expo-expert`**: Audio streaming implementation in React Native with `expo-av` and WebRTC shim.
1
+ ---
2
+ name: voice-ai-realtime-agent
3
+ description: "Expert guide for Ultra-Low Latency Conversational Voice AI (<300ms), WebRTC bidirectional streaming, OpenAI Realtime API, Gemini Multimodal Live Audio, LiveKit Agents, and Semantic VAD / Panduan ahli AI suara percakapan real-time berlatensi ultra-rendah."
4
+ author: "Roedy Rustam"
5
+ version: "3.0.0"
6
+ ---
7
+
8
+ # Voice AI Realtime Agent (2026 Edition)
9
+
10
+ [English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
11
+
12
+ ---
13
+
14
+ <a name="english"></a>
15
+ ## English
16
+
17
+ ### Description
18
+ Expert guide for building ultra-low-latency (<300ms), bi-directional conversational voice AI applications. Covers WebRTC, full-duplex WebSocket audio streaming (PCM 24kHz), OpenAI Realtime API, Gemini Multimodal Live API, LiveKit Agents SDK, and smart interruption (barge-in) handling.
19
+
20
+ ### Trigger Conditions
21
+ - Applications requiring sub-second, spoken conversation with an AI agent.
22
+ - Voice customer service bots, verbal copilots, language tutors, and interactive voice assistants.
23
+ - Implementation of WebRTC audio streaming, full-duplex WebSocket audio (PCM 24kHz), and Silero VAD.
24
+ - Setting up OpenAI Realtime API (`gpt-4o-realtime-preview`) or Gemini Multimodal Live API.
25
+
26
+ ---
27
+
28
+ ## 1. Core Architecture: Full-Duplex Speech-to-Speech
29
+
30
+ Traditional voice pipelines chain STT ➔ LLM ➔ TTS with cumulative latency exceeding 1,200ms–2,500ms. Modern 2026 voice agents use **native speech-to-speech** or **streamable full-duplex WebRTC pipelines** achieving natural, human-like reaction times (~250–350ms).
31
+
32
+ ```
33
+ User Mic ──► [WebRTC / WebSocket] ──► [VAD: Silero / WebRTC VAD]
34
+ │
35
+ ▼
36
+ User Speaks <── [Audio Output] ◄── [Native Audio Stream / Cartesia] ◄── [OpenAI Realtime / Gemini Live]
37
+ │
38
+ └── User Interrupts (Barge-in) ──► Instant Buffer Flush & Cancel Audio Frame Emission
39
+ ```
40
+
41
+ ---
42
+
43
+ ## 2. Production Recipe: LiveKit Agents + OpenAI Realtime (Python)
44
+
45
+ ```python
46
+ # agent.py - Production Voice Agent Worker with LiveKit & OpenAI Realtime
47
+ import asyncio
48
+ import os
49
+ from livekit import rtc
50
+ from livekit.agents import (
51
+ AutoSubscribe,
52
+ JobContext,
53
+ JobProcess,
54
+ WorkerOptions,
55
+ cli,
56
+ llm,
57
+ )
58
+ from livekit.agents.pipeline import VoicePipelineAgent
59
+ from livekit.plugins import deepgram, openai, silero
60
+
61
+ async def entrypoint(ctx: JobContext):
62
+ # Connect to room with audio only to minimize bandwidth & latency
63
+ await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY)
64
+
65
+ # Wait for the user participant to join
66
+ participant = await ctx.wait_for_participant()
67
+
68
+ # Define agent instructions and tools
69
+ initial_ctx = llm.ChatContext().append(
70
+ role="system",
71
+ text=(
72
+ "You are a helpful, concise voice assistant. "
73
+ "Respond naturally in 1-2 short sentences. Never output markdown, bullet points, or emojis."
74
+ )
75
+ )
76
+
77
+ # Realtime Voice Pipeline: Deepgram (STT) + OpenAI (LLM) + Cartesia/OpenAI (TTS)
78
+ # Or use native OpenAI Realtime Model: gpt-4o-realtime-preview
79
+ agent = VoicePipelineAgent(
80
+ vad=silero.VAD.load(
81
+ min_speech_duration=0.1,
82
+ min_silence_duration=0.3, # Snappy turn-taking
83
+ prefix_padding_duration=0.2,
84
+ ),
85
+ stt=deepgram.STT(model="nova-2", language="id"), # Multi-language support
86
+ llm=openai.LLM(model="gpt-4o-mini"),
87
+ tts=openai.TTS(voice="alloy"),
88
+ chat_ctx=initial_ctx,
89
+ allow_interruptions=True, # Barge-in capability
90
+ interrupt_speech_duration=0.3, # Immediate cutoff when user talks
91
+ )
92
+
93
+ agent.start(ctx.room, participant)
94
+
95
+ # Greet user immediately
96
+ await agent.say("Halo! Ada yang bisa saya bantu hari ini?", now=True)
97
+
98
+ if __name__ == "__main__":
99
+ cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))
100
+ ```
101
+
102
+ ---
103
+
104
+ ## 3. Production Recipe: Gemini Multimodal Live Audio (TypeScript / Node.js)
105
+
106
+ ```typescript
107
+ // gemini-live-audio.ts - Bidirectional WebSocket PCM 24kHz
108
+ import WebSocket from 'ws';
109
+
110
+ interface GeminiAudioConfig {
111
+ apiKey: string;
112
+ model?: string;
113
+ systemInstruction?: string;
114
+ }
115
+
116
+ export class GeminiVoiceAgent {
117
+ private ws: WebSocket | null = null;
118
+ private isConnected = false;
119
+
120
+ constructor(private config: GeminiAudioConfig) {}
121
+
122
+ public async connect(): Promise<void> {
123
+ const url = `wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1alpha.GenerativeService.BidiGenerateContent?key=${this.config.apiKey}`;
124
+
125
+ this.ws = new WebSocket(url);
126
+
127
+ this.ws.on('open', () => {
128
+ this.isConnected = true;
129
+ this.sendInitialHandshake();
130
+ });
131
+
132
+ this.ws.on('message', (data: WebSocket.Data) => {
133
+ this.handleIncomingAudio(data);
134
+ });
135
+ }
136
+
137
+ private sendInitialHandshake(): void {
138
+ const setupMessage = {
139
+ setup: {
140
+ model: `models/${this.config.model || 'gemini-2.0-flash-exp'}`,
141
+ generationConfig: {
142
+ responseModalities: ["AUDIO"],
143
+ speechConfig: {
144
+ voiceConfig: {
145
+ prebuiltVoiceConfig: { voiceName: "Puck" }
146
+ }
147
+ }
148
+ },
149
+ systemInstruction: {
150
+ parts: [{ text: this.config.systemInstruction || "You are a conversational voice agent. Keep answers brief." }]
151
+ }
152
+ }
153
+ };
154
+ this.ws?.send(JSON.stringify(setupMessage));
155
+ }
156
+
157
+ // Stream raw PCM 16-bit 24kHz mono audio from mic
158
+ public sendAudioChunk(pcm16Chunk: Buffer): void {
159
+ if (!this.isConnected || !this.ws) return;
160
+
161
+ const base64Audio = pcm16Chunk.toString('base64');
162
+ const msg = {
163
+ realtimeInput: {
164
+ mediaChunks: [
165
+ {
166
+ mimeType: "audio/pcm;rate=24000",
167
+ data: base64Audio
168
+ }
169
+ ]
170
+ }
171
+ };
172
+ this.ws.send(JSON.stringify(msg));
173
+ }
174
+
175
+ private handleIncomingAudio(data: WebSocket.Data): void {
176
+ try {
177
+ const response = JSON.parse(data.toString());
178
+ const parts = response.serverContent?.modelTurn?.parts;
179
+ if (parts) {
180
+ for (const part of parts) {
181
+ if (part.inlineData?.data) {
182
+ const pcmBuffer = Buffer.from(part.inlineData.data, 'base64');
183
+ this.playAudioSpeaker(pcmBuffer);
184
+ }
185
+ }
186
+ }
187
+ } catch {
188
+ // Binary PCM frame handler
189
+ }
190
+ }
191
+
192
+ private playAudioSpeaker(pcmChunk: Buffer): void {
193
+ // Send to WebRTC audio track or audio output device
194
+ }
195
+ }
196
+ ```
197
+
198
+ ---
199
+
200
+ ## 4. Key 2026 Performance Guardrails
201
+
202
+ 1. **Barge-in Latency Budget (<150ms)**: When the user speaks while the bot is talking, cancel outgoing audio immediately. Do not wait for the LLM to finish streaming its chunk.
203
+ 2. **Audio Sample Rates**:
204
+ - Mic Input: 16kHz or 24kHz 16-bit Linear PCM Mono.
205
+ - Bot Output: 24kHz PCM for crystal-clear natural prosody.
206
+ 3. **Turn-Taking Jitter Prevention**: Use minimum silence thresholds between `300ms` and `450ms`. Lower thresholds cause the bot to interrupt users when they pause to think; higher thresholds make the conversation feel robotic.
207
+
208
+ ---
209
+
210
+ ## Orchestration & Integration
211
+
212
+ - **`ai-llm-integration-expert`**: For base LLM prompt routing and function calling during conversation.
213
+ - **`realtime-collaboration-expert`**: For syncing WebRTC tracks and room states with client applications.
214
+ - **`gemini-agent-booster`**: Connects Gemini 3.x / 2.0 Flash thinking models to live voice agents.
215
+ - **`mobile-expo-expert`**: Audio streaming implementation in React Native with `expo-av` and WebRTC shim.
216
+
217
+ ---
218
+
219
+ <a name="bahasa-indonesia"></a>
220
+ ## Bahasa Indonesia
221
+
222
+ ### Deskripsi
223
+ Panduan ahli untuk membangun aplikasi AI suara percakapan dua arah berlatensi ultra-rendah (<300ms). Mencakup integrasi WebRTC, streaming audio WebSocket full-duplex (PCM 24kHz), OpenAI Realtime API, Gemini Multimodal Live API, LiveKit Agents SDK, dan penanganan interupsi cerdas (*barge-in*).
224
+
225
+ ### Kondisi Pemicu
226
+ - Kebutuhan interaksi percakapan verbal instan di bawah satu detik dengan agen AI.
227
+ - Bot layanan pelanggan berbasis suara, asisten verbal, tutor bahasa interaktif.
228
+ - Implementasi streaming audio WebRTC, WebSocket PCM 24kHz dua arah, dan Silero Voice Activity Detection (VAD).
229
+ - Konfigurasi OpenAI Realtime API (`gpt-4o-realtime-preview`) atau Gemini Multimodal Live API.
230
+
231
+ ### Panduan Inti Arsitektur Suara Real-time
232
+ 1. **Full-Duplex Speech-to-Speech**: Mengganti pipeline sekuensial tradisional (STT ➔ LLM ➔ TTS) dengan pipeline streamable WebRTC atau model native speech-to-speech untuk memangkas latensi dari ~2000ms menjadi ~300ms.
233
+ 2. **Penanganan Interupsi (Barge-In)**: Saat VAD mendeteksi suara pengguna baru saat bot sedang berbicara, buffer audio keluar harus di-flush dalam waktu <150ms tanpa menunggu LLM menyelesaikan kalimatnya.
234
+ 3. **Standar Format Audio**: Input mikrofon PCM 16-bit 16kHz/24kHz Mono, dan output speaker 24kHz untuk intonasi yang alami dan jernih.
235
+
236
+ ---
237
+
238
+ ## Integrasi Orkestrasi
239
+
240
+ - **`ai-llm-integration-expert`**: Routing instruksi sistem dasar dan pemanggilan tool fungsi selama percakapan suara.
241
+ - **`realtime-collaboration-expert`**: Sinkronisasi track audio WebRTC dan status room pengguna.
242
+ - **`gemini-agent-booster`**: Integrasi model multimodal Gemini 2.0/3.x Flash untuk live audio.
243
+ - **`mobile-expo-expert`**: Implementasi audio streaming di React Native menggunakan `expo-av` dan WebRTC shim.
@@ -1,7 +1,8 @@
1
1
  ---
2
2
  name: vue-frontend-expert
3
3
  description: "Expert guide for Vue 3 (Composition API), Nuxt 3, and Pinia. Covers advanced reactive state management, `<script setup>` syntax, Vue Router, VueUse, and SPA/SSR architectural patterns in English and Indonesian."
4
- author: "Roedy Rustam"
4
+ author: "Roedy Rustam"
5
+ version: "3.0.0"
5
6
  ---
6
7
 
7
8
  # Vue Frontend Expert (Vue 3 / Nuxt 3)
@@ -14,7 +15,7 @@ author: "Roedy Rustam"
14
15
  ## English
15
16
 
16
17
  ### Orchestration & Integration
17
- Connects and orchestrates with relevant domain skills like `brainstorming`, `zero-to-prod-orchestrator`, and `project-context-mapper` to ensure cohesive execution.
18
+ Connects and orchestrates with relevant domain skills like `brainstorming`, `zero-to-prod-orchestrator`, and `session-memory-manager` to ensure cohesive execution.
18
19
 
19
20
  ### Description
20
21
  Production-grade guidance for building highly reactive and scalable frontend applications using **Vue 3 (Composition API)** and **Nuxt 3**. Covers the latest ecosystem tools including **Pinia** for state management, **VueUse** for composables, and **Tailwind CSS v4** integration.
@@ -111,7 +112,7 @@ This skill should be referenced by the following orchestrators:
111
112
  ## Bahasa Indonesia
112
113
 
113
114
  ### Integrasi Orkestrasi
114
- Terhubung dan mengorkestrasi skill domain yang relevan seperti `brainstorming`, `zero-to-prod-orchestrator`, dan `project-context-mapper` untuk memastikan eksekusi yang kohesif.
115
+ Terhubung dan mengorkestrasi skill domain yang relevan seperti `brainstorming`, `zero-to-prod-orchestrator`, dan `session-memory-manager` untuk memastikan eksekusi yang kohesif.
115
116
 
116
117
  ### Deskripsi
117
118
  Panduan tingkat produksi untuk membangun aplikasi frontend reaktif dan terukur menggunakan **Vue 3 (Composition API)** dan **Nuxt 3**. Mencakup pengelolaan state dengan **Pinia**, utilitas **VueUse**, dan integrasi **Tailwind CSS v4**.
@@ -129,4 +130,4 @@ Aktifkan skill ini ketika pengguna sedang:
129
130
  - **Pinia:** Gunakan Pinia dengan pola *Setup Store* (mirip Composition API) alih-alih pola Options (state, getters, actions).
130
131
  - **Nuxt 3 Fetching:** Gunakan `useFetch` atau `useAsyncData` di dalam komponen Nuxt untuk pengambilan data saat SSR, bukan `onMounted` dengan `fetch` biasa.
131
132
  - **Reaktivitas:** Gunakan `ref` untuk nilai primitif (string, number) dan `reactive` untuk objek bersarang yang kompleks. Utamakan `computed` daripada `watch` untuk state turunan.
132
- - **VueUse:** Jangan menulis fungsi utilitas dari nol jika sudah ada di *library* VueUse (contoh: `useIntersectionObserver`, `useLocalStorage`).
133
+ - **VueUse:** Jangan menulis fungsi utilitas dari nol jika sudah ada di *library* VueUse (contoh: `useIntersectionObserver`, `useLocalStorage`).
@@ -1,7 +1,8 @@
1
1
  ---
2
2
  name: wasm-edge-computing-expert
3
3
  description: "Expert guide for WebAssembly (WASM) and Edge Computing. Covers WASI preview 2, Spin/Fermyon, Cloudflare Workers WASM, and high-performance browser computing / Panduan ahli untuk WebAssembly (WASM) dan Edge Computing. Mencakup WASI preview 2, Spin/Fermyon, Cloudflare Workers WASM, dan komputasi performa tinggi di browser."
4
- author: "Roedy Rustam"
4
+ author: "Roedy Rustam"
5
+ version: "3.0.0"
5
6
  ---
6
7
 
7
8
  # WASM & Edge Computing Expert
@@ -94,4 +95,4 @@ Untuk performa multi-threading sejati di browser menggunakan WASM, gunakan `Shar
94
95
  ## Integrasi Orkestrasi
95
96
  - Memperkuat `api-gateway-proxy-expert` untuk logika kustom di edge.
96
97
  - Terintegrasi dengan `rust-programming-expert` untuk kompilasi modul.
97
- - Melengkapi `performance-web-vitals` dengan memindahkan komputasi berat dari *main thread* JavaScript.
98
+ - Melengkapi `performance-web-vitals` dengan memindahkan komputasi berat dari *main thread* JavaScript.