vibes-plug 2.11.0 → 2.14.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (69) hide show
  1. package/.cursor/rules/vibes-plug-core.mdc +3 -3
  2. package/.cursorrules +3 -3
  3. package/AGENTS.md +4 -4
  4. package/BLUEPRINT.md +16 -6
  5. package/CHANGELOG.md +37 -0
  6. package/CLAUDE.md +8 -8
  7. package/README.md +85 -115
  8. package/index.js +1 -1
  9. package/package.json +2 -2
  10. package/plugin.json +2 -2
  11. package/scripts/check-anti-slop.js +53 -0
  12. package/scripts/generate_swarm_gif.py +2 -2
  13. package/skills/ai-llm-integration-expert/SKILL.md +22 -15
  14. package/skills/ai-prompt-engineering-expert/SKILL.md +133 -83
  15. package/skills/anti-slop/SKILL.md +133 -0
  16. package/skills/async-queue-temporal-expert/SKILL.md +135 -158
  17. package/skills/authentication-identity-expert/SKILL.md +172 -278
  18. package/skills/brainstorming/SKILL.md +26 -26
  19. package/skills/database-orm-expert/SKILL.md +164 -303
  20. package/skills/deep-research-analyst/SKILL.md +136 -0
  21. package/skills/design-system-architect/SKILL.md +31 -1
  22. package/skills/email-notification-expert/SKILL.md +31 -4
  23. package/skills/error-resilience-expert/SKILL.md +21 -0
  24. package/skills/fullstack-expert/SKILL.md +183 -260
  25. package/skills/glsl-shader-expert/SKILL.md +190 -107
  26. package/skills/graph-rag-knowledge-expert/SKILL.md +42 -1
  27. package/skills/mcp-server-architect/SKILL.md +15 -1
  28. package/skills/prd-architect/SKILL.md +181 -206
  29. package/skills/production-ready-hardener/SKILL.md +16 -19
  30. package/skills/pwa-offline-first-expert/SKILL.md +42 -1
  31. package/skills/pydantic-ai-expert/SKILL.md +161 -0
  32. package/skills/saas-architect/SKILL.md +154 -0
  33. package/skills/senior-frontend/SKILL.md +9 -11
  34. package/skills/senior-frontend/scripts/frontend_scaffolder.py +1 -1
  35. package/skills/session-memory-manager/SKILL.md +128 -0
  36. package/skills/synthetic-data-finetuning-expert/SKILL.md +155 -0
  37. package/skills/ui-ux-pro-max/SKILL.md +4 -2
  38. package/skills/vercel-ai-sdk-expert/SKILL.md +181 -0
  39. package/skills/voice-ai-realtime-agent/SKILL.md +41 -1
  40. package/skills/web-3d-graphics-expert/SKILL.md +313 -137
  41. package/skills/web-game-engine-expert/SKILL.md +329 -102
  42. package/skills/webxr-ar-vr-expert/SKILL.md +162 -123
  43. package/skills/zero-to-prod-orchestrator/SKILL.md +26 -24
  44. package/skills/ai-cost-token-optimizer/SKILL.md +0 -82
  45. package/skills/ai-evals-benchmark-expert/SKILL.md +0 -188
  46. package/skills/asisten-ramah/SKILL.md +0 -47
  47. package/skills/auto-doc-updater/SKILL.md +0 -220
  48. package/skills/autonomous-chaos-monkey/SKILL.md +0 -63
  49. package/skills/background-jobs-queue-expert/SKILL.md +0 -235
  50. package/skills/database-migration-versioning-expert/SKILL.md +0 -90
  51. package/skills/edge-serverless-db-expert/SKILL.md +0 -99
  52. package/skills/mcp-client-orchestrator/SKILL.md +0 -76
  53. package/skills/mobile-push-notification-expert/SKILL.md +0 -71
  54. package/skills/monday-design-aesthetic/SKILL.md +0 -73
  55. package/skills/project-context-mapper/SKILL.md +0 -85
  56. package/skills/saas-mvp-launcher/SKILL.md +0 -260
  57. package/skills/saas-transformer/SKILL.md +0 -500
  58. package/skills/saas-transformer/references/billing_integration_guide.md +0 -401
  59. package/skills/self-evolving-memory-graph/SKILL.md +0 -91
  60. package/skills/session-context-loader/SKILL.md +0 -83
  61. package/skills/session-handoff-resume/SKILL.md +0 -164
  62. package/skills/skill-baru/SKILL.md +0 -178
  63. package/skills/supabase-migration/SKILL.md +0 -91
  64. package/skills/token-saver/SKILL.md +0 -119
  65. package/skills/ui-components-expert/SKILL.md +0 -166
  66. package/skills/vibe-code-gardener/SKILL.md +0 -181
  67. /package/skills/{saas-transformer → saas-architect}/references/feature_gating_patterns.md +0 -0
  68. /package/skills/{saas-transformer → saas-architect}/references/saas_transformation_checklist.md +0 -0
  69. /package/skills/{saas-transformer → saas-architect}/scripts/saas_transformation_scanner.py +0 -0
@@ -0,0 +1,155 @@
1
+ ---
2
+ name: synthetic-data-finetuning-expert
3
+ description: "Expert guide for synthetic dataset generation, LLM-as-a-judge filtering, QLoRA fine-tuning (Unsloth), DPO alignment, and GGUF/Ollama export for local SLMs / Panduan ahli generasi data sintetis, fine-tuning QLoRA, DPO, dan ekspor GGUF/Ollama."
4
+ author: vibes-plug-swarm
5
+ ---
6
+
7
+ # Synthetic Data & Fine-Tuning Expert (Custom Domain SLMs)
8
+
9
+ [English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
10
+
11
+ ---
12
+
13
+ <a name="english"></a>
14
+ ## English
15
+
16
+ ### Orchestration & Integration
17
+ Connects and orchestrates with domain skills like `local-slm-edge-ai-expert`, `ai-prompt-engineering-expert`, `ai-evals-benchmark-expert`, `python-programming-expert`, and `ai-cost-token-optimizer` to build high-performance, cost-effective domain models.
18
+
19
+ ### Description
20
+ Production-grade guide for generating synthetic training datasets, curating high-signal instruction pairs, executing Parameter-Efficient Fine-Tuning (QLoRA / LoRA) with **Unsloth** and Hugging Face TRL, performing Direct Preference Optimization (DPO), and quantizing custom Small Language Models (SLMs) to GGUF for edge or on-premise execution.
21
+
22
+ **Swarm Synergy:** Within the **AI Engineering Swarm**, this skill serves as the Model Specialization Lead. When frontier API costs or latency become prohibitive, it trains, aligns, and deploys hyper-efficient domain SLMs (1B–8B parameters) in Phase 4.
23
+
24
+ ### Trigger Conditions
25
+ - Generating domain-specific synthetic training data from seed documents, codebases, or APIs.
26
+ - Filtering low-quality or hallucinated synthetic data using LLM-as-a-judge curation pipelines.
27
+ - Fine-tuning open-weights models (Llama 3.3, Qwen 2.5, Mistral) on custom tasks using 4-bit QLoRA.
28
+ - Aligning model outputs using Direct Preference Optimization (DPO) to enforce specific response styles.
29
+ - Quantizing fine-tuned models to GGUF (q4_k_m, q8_0) for zero-latency local inference with Ollama or llama.cpp.
30
+
31
+ ### Synthetic Data & Fine-Tuning Lifecycle
32
+
33
+ ```
34
+ 1. SEED EXTRACTION & SYNTHESIS
35
+ [Raw Docs / Codebase] ──► [Frontier LLM / Distilabel] ──► Raw Instruction Pairs (10k+)
36
+
37
+ 2. QUALITY FILTERING (LLM-AS-A-JUDGE)
38
+ Raw Instruction Pairs ──► [Rubric Scorer / De-duplication] ──► Curated Gold Dataset (2k-5k)
39
+
40
+ 3. 4-BIT QLORA FINE-TUNING (UNSLOTH)
41
+ Base Model (e.g. Qwen 2.5-Coder) + LoRA Adapters ──► SFT / DPO Training Loop
42
+
43
+ 4. QUANTIZATION & LOCAL DEPLOYMENT
44
+ Merged 16-bit Weights ──► [llama.cpp GGUF Export] ──► Local Ollama Service (<50ms latency)
45
+ ```
46
+
47
+ ### Core Implementation Guidelines
48
+
49
+ #### 1. Synthetic Instruction Generation & Rejection Sampling
50
+ Use frontier models to generate input-output pairs with strict rejection criteria:
51
+ ```python
52
+ from pydantic import BaseModel, Field
53
+ from openai import OpenAI
54
+ import json
55
+
56
+ client = OpenAI()
57
+
58
+ class SyntheticInstructionPair(BaseModel):
59
+ user_instruction: str = Field(description="Realistic developer query")
60
+ input_context: str = Field(description="Code snippet or API contract")
61
+ ground_truth_response: str = Field(description="Flawless, production-grade output")
62
+ quality_score: int = Field(ge=1, le=5, description="Self-evaluated quality score")
63
+
64
+ def generate_synthetic_samples(seed_text: str) -> list[SyntheticInstructionPair]:
65
+ response = client.beta.chat.completions.parse(
66
+ model="gpt-4o-mini",
67
+ messages=[
68
+ {"role": "system", "content": "Generate 5 diverse, hard edge-case instruction pairs based on the seed code."},
69
+ {"role": "user", "content": seed_text}
70
+ ],
71
+ response_format=SyntheticInstructionPair,
72
+ )
73
+ # Filter out anything below score 4 (Rejection Sampling)
74
+ return [sample for sample in [response.choices[0].message.parsed] if sample.quality_score >= 4]
75
+ ```
76
+
77
+ #### 2. High-Speed 4-Bit QLoRA with Unsloth
78
+ Fine-tune on consumer GPUs (e.g., RTX 3090/4090 or single A10G) with 5x faster throughput:
79
+ ```python
80
+ from unsloth import FastLanguageModel
81
+ import torch
82
+ from trl import SFTTrainer
83
+ from transformers import TrainingArguments
84
+
85
+ # 1. Load Model & Tokenizer in 4-bit
86
+ max_seq_length = 2048
87
+ model, tokenizer = FastLanguageModel.from_pretrained(
88
+ model_name="unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit",
89
+ max_seq_length=max_seq_length,
90
+ load_in_4bit=True,
91
+ )
92
+
93
+ # 2. Add LoRA Adapters
94
+ model = FastLanguageModel.get_peft_model(
95
+ model,
96
+ r=16,
97
+ target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],
98
+ lora_alpha=16,
99
+ lora_dropout=0, # Optimized 0 dropout for Unsloth
100
+ bias="none",
101
+ )
102
+
103
+ # 3. Supervised Fine-Tuning (SFT)
104
+ trainer = SFTTrainer(
105
+ model=model,
106
+ tokenizer=tokenizer,
107
+ train_dataset=dataset,
108
+ dataset_text_field="text",
109
+ max_seq_length=max_seq_length,
110
+ args=TrainingArguments(
111
+ per_device_train_batch_size=2,
112
+ gradient_accumulation_steps=4,
113
+ warmup_steps=10,
114
+ max_steps=100,
115
+ learning_rate=2e-4,
116
+ fp16=not torch.cuda.is_bf16_supported(),
117
+ bf16=torch.cuda.is_bf16_supported(),
118
+ logging_steps=1,
119
+ output_dir="outputs",
120
+ ),
121
+ )
122
+ trainer.train()
123
+ ```
124
+
125
+ #### 3. GGUF Export for Local Ollama Deployment
126
+ Export merged model to 4-bit or 8-bit quantized GGUF format:
127
+ ```python
128
+ # Save to 16bit or GGUF directly
129
+ model.save_pretrained_gguf("custom-domain-coder", tokenizer, quantization_method="q4_k_m")
130
+
131
+ # Generate Modelfile for Ollama:
132
+ # FROM ./custom-domain-coder-q4_k_m.gguf
133
+ # PARAMETER temperature 0.2
134
+ # SYSTEM You are an expert domain coder.
135
+ ```
136
+
137
+ ---
138
+
139
+ <a name="bahasa-indonesia"></a>
140
+ ## Bahasa Indonesia
141
+
142
+ ### Integrasi Orkestrasi
143
+ Terhubung dan mengorkestrasi skill domain yang relevan seperti `local-slm-edge-ai-expert`, `ai-prompt-engineering-expert`, `ai-evals-benchmark-expert`, `python-programming-expert`, dan `ai-cost-token-optimizer` untuk membangun model spesialis domain dengan performa tinggi dan biaya hemat.
144
+
145
+ ### Deskripsi
146
+ Panduan produksi untuk menghasilkan dataset pelatihan sintetis, mengurasi pasangan instruksi bernilai tinggi, mengeksekusi Parameter-Efficient Fine-Tuning (QLoRA / LoRA) dengan **Unsloth** dan Hugging Face TRL, menerapkan Direct Preference Optimization (DPO), dan mengkuantisasi Small Language Models (SLM) kustom ke format GGUF untuk inferensi lokal atau on-premise berlatensi ultra-rendah.
147
+
148
+ **Sinergi Swarm:** Di dalam **AI Engineering Swarm**, skill ini memegang peranan sebagai Pemimpin Spesialisasi Model. Ketika biaya API atau latensi model cloud frontier terlalu tinggi, skill ini melatih, menyelaraskan (*align*), dan men-deploy SLM domain yang sangat efisien (1B–8B parameter) pada Fase 4.
149
+
150
+ ### Kondisi Pemicu
151
+ - Menghasilkan data pelatihan sintetis spesifik domain dari dokumen panduan, codebase, atau skema API.
152
+ - Menyaring data sintetis berkualitas rendah menggunakan pipeline penilaian otomatis *LLM-as-a-judge*.
153
+ - Melakukan fine-tuning model berbobot terbuka (Llama 3.3, Qwen 2.5, Mistral) dengan QLoRA 4-bit secara hemat memori VRAM.
154
+ - Menyelaraskan respon model dengan Direct Preference Optimization (DPO) agar mematuhi aturan format dan gaya tertentu.
155
+ - Mengkuantisasi model hasil fine-tuning ke format GGUF (q4_k_m, q8_0) untuk inferensi lokal instan di Ollama atau llama.cpp.
@@ -58,7 +58,8 @@ python scripts/search.py "<query>" --domain <domain> --max-results 3
58
58
  - `icons`: Icon usage guidance, SVG libraries (Lucide, Heroicons), and code imports
59
59
  - `react`: React & Next.js performance optimizations, re-render fixes & dynamic imports
60
60
  - `web`: Web interface guidelines (ARIA, focus traps, virtual list, form inputs)
61
- - `m3`: Material Design 3 specific design tokens, color roles, elevation, and component specs
61
+ - `m3`: Material Design 3 specific design tokens, color roles, elevation, and component specs
62
+ - `preset-monday`: Monday.com spacious SaaS aesthetic (Vibrant Blue `#0073ea`, Clean White `#FFFFFF` + Light Grays `#F9F9F9`/`#F5F6F8`, Dark Footer `#111111`, Figtree/Inter fonts, 12-16px card radius, 80-120px vertical spacing)
62
63
 
63
64
  *Example:* `python scripts/search.py "fintech dark theme" --domain color`
64
65
 
@@ -159,7 +160,8 @@ python scripts/search.py "<kueri>" --domain <domain> --max-results 3
159
160
  - `icons`: Panduan penggunaan ikon, pustaka SVG (Lucide, Heroicons), & import kode
160
161
  - `react`: Optimasi performa React & Next.js, perbaikan re-render & dynamic import
161
162
  - `web`: Pedoman antarmuka web (ARIA, focus trap, virtual list, input form)
162
- - `m3`: Token desain spesifik Material Design 3, peran warna, elevasi, dan spesifikasi komponen
163
+ - `m3`: Token desain spesifik Material Design 3, peran warna, elevasi, dan spesifikasi komponen
164
+ - `preset-monday`: Estetika SaaS lapang ala Monday.com (Biru Cerah `#0073ea`, Putih Bersih `#FFFFFF` + Abu-abu Terang `#F9F9F9`/`#F5F6F8`, Footer Gelap `#111111`, font Figtree/Inter, radius kartu 12-16px, jarak seksi vertikal 80-120px)
163
165
 
164
166
  *Contoh:* `python scripts/search.py "fintech dark theme" --domain color`
165
167
 
@@ -0,0 +1,181 @@
1
+ ---
2
+ name: vercel-ai-sdk-expert
3
+ description: "Expert guide for Vercel AI SDK (Core, UI, RSC), streaming structured data, multi-provider model switching, tool calling loops, and React 19/Next.js 15 AI engineering / Panduan ahli Vercel AI SDK, streaming data terstruktur, dan integrasi AI pada React 19/Next.js 15."
4
+ author: vibes-plug-swarm
5
+ ---
6
+
7
+ # Vercel AI SDK Expert (Core, UI & Fullstack AI Engineering)
8
+
9
+ [English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
10
+
11
+ ---
12
+
13
+ <a name="english"></a>
14
+ ## English
15
+
16
+ ### Orchestration & Integration
17
+ Connects and orchestrates with domain skills like `senior-frontend`, `nextjs-app-router-expert`, `ai-llm-integration-expert`, `ui-components-expert`, and `multi-agent-orchestration` to deliver reactive, streaming AI interfaces.
18
+
19
+ ### Description
20
+ Production-grade guide for building AI applications using the **Vercel AI SDK (Core & UI)**. Covers unified model provider abstraction (`@ai-sdk/anthropic`, `@ai-sdk/openai`, `@ai-sdk/google`), streaming text and structured objects (`streamText`, `streamObject`), dynamic multi-step tool execution loops with `maxSteps`, client-side React 19 hooks (`useChat`, `useCompletion`), streaming data attachments (`createDataStreamResponse`), and generative UI rendering.
21
+
22
+ **Swarm Synergy:** Within the **Frontend & UI Swarm**, this skill serves as the AI UI Presentation Lead. It translates complex backend multi-agent outputs and streaming tokens into accessible, beautiful web components in Phase 4 & Phase 5.
23
+
24
+ ### Trigger Conditions
25
+ - Integrating conversational chat, streaming completions, or generative UI in React 19 / Next.js 15.
26
+ - Implementing structured data extraction using `generateObject` or `streamObject` with Zod schemas.
27
+ - Building autonomous multi-step tool-calling loops on Next.js Route Handlers or Server Actions.
28
+ - Switching seamlessly across frontier providers (Claude 3.7 Sonnet, Gemini 3.8 Flash, OpenAI o3/GPT-4.5, Ollama).
29
+ - Building streaming data channels with custom metadata, tool status indicators, and citations.
30
+
31
+ ### Vercel AI SDK Architecture (Core vs UI)
32
+
33
+ ```
34
+ ┌─────────────────────────────────────────────────────────────┐
35
+ │ CLIENT LAYER │
36
+ │ useChat / useCompletion / Generative UI React Components │
37
+ │ • Optimistic updates • Stream reader • Tool invocation │
38
+ └──────────────────────────────▲──────────────────────────────┘
39
+ │ HTTP SSE / Data Stream Protocol
40
+ ┌──────────────────────────────▼──────────────────────────────┐
41
+ │ SERVER ROUTE / ACTION │
42
+ │ streamText({ │
43
+ │ model: anthropic('claude-3-7-sonnet-20250219'), │
44
+ │ tools: { weatherTool, dbQueryTool }, │
45
+ │ maxSteps: 5, │
46
+ │ }).toDataStreamResponse() │
47
+ └─────────────────────────────────────────────────────────────┘
48
+ ```
49
+
50
+ ### Core Implementation Guidelines
51
+
52
+ #### 1. Next.js 15 Route Handler with Multi-Step Tool Calling Loop
53
+ Use `streamText` with `maxSteps` to enable the model to autonomously call tools, review results, and continue reasoning:
54
+ ```typescript
55
+ // app/api/chat/route.ts
56
+ import { anthropic } from '@ai-sdk/anthropic';
57
+ import { streamText, tool } from 'ai';
58
+ import { z } from 'zod';
59
+
60
+ export const maxDuration = 60; // Allow long-running agentic reasoning
61
+
62
+ export async function POST(req: Request) {
63
+ const { messages } = await req.json();
64
+
65
+ const result = streamText({
66
+ model: anthropic('claude-3-7-sonnet-20250219'),
67
+ messages,
68
+ maxSteps: 5, // Enables iterative tool calling loop
69
+ tools: {
70
+ calculateMetrics: tool({
71
+ description: 'Computes analytical metrics from raw time series data',
72
+ parameters: z.object({
73
+ datasetId: z.string(),
74
+ metricType: z.enum(['p95_latency', 'error_rate', 'throughput']),
75
+ }),
76
+ execute: async ({ datasetId, metricType }) => {
77
+ const data = await fetchDatasetMetrics(datasetId, metricType);
78
+ return { datasetId, metricType, value: data.result };
79
+ },
80
+ }),
81
+ },
82
+ system: 'You are an elite software performance auditor. Always back up your conclusions with data tool outputs.',
83
+ });
84
+
85
+ return result.toDataStreamResponse();
86
+ }
87
+ ```
88
+
89
+ #### 2. Streaming Type-Safe Structured Objects (`streamObject`)
90
+ Stream structured JSON objects directly into the UI while generating:
91
+ ```typescript
92
+ import { google } from '@ai-sdk/google';
93
+ import { streamObject } from 'ai';
94
+ import { z } from 'zod';
95
+
96
+ export async function POST(req: Request) {
97
+ const { codeDiff } = await req.json();
98
+
99
+ const result = streamObject({
100
+ model: google('gemini-2.5-flash'),
101
+ schema: z.object({
102
+ securityVulnerabilities: z.array(z.object({
103
+ severity: z.enum(['low', 'medium', 'high', 'critical']),
104
+ cwe: z.string(),
105
+ explanation: z.string(),
106
+ suggestedFix: z.string(),
107
+ })),
108
+ overallRiskScore: z.number().min(0).max(100),
109
+ passesReview: z.boolean(),
110
+ }),
111
+ prompt: `Audit the following git diff for security regressions:\n${codeDiff}`,
112
+ });
113
+
114
+ return result.toTextStreamResponse();
115
+ }
116
+ ```
117
+
118
+ #### 3. Client Hook Integration (`useChat` with Tool Invocations)
119
+ Render real-time streaming tokens, loading skeletons, and interactive tool call results:
120
+ ```tsx
121
+ 'use client';
122
+
123
+ import { useChat } from '@ai-sdk/react';
124
+
125
+ export function AgenticChat() {
126
+ const { messages, input, handleInputChange, handleSubmit, isLoading } = useChat({
127
+ maxSteps: 5,
128
+ });
129
+
130
+ return (
131
+ <div className="flex flex-col h-[600px] w-full max-w-2xl mx-auto border rounded-xl p-4 bg-background">
132
+ <div className="flex-1 overflow-y-auto space-y-4 pr-2">
133
+ {messages.map((m) => (
134
+ <div key={m.id} className={`flex ${m.role === 'user' ? 'justify-end' : 'justify-start'}`}>
135
+ <div className={`p-3 rounded-lg max-w-[80%] ${m.role === 'user' ? 'bg-primary text-primary-foreground' : 'bg-muted'}`}>
136
+ <div className="whitespace-pre-wrap">{m.content}</div>
137
+ {m.toolInvocations?.map((toolInvocation) => (
138
+ <div key={toolInvocation.toolCallId} className="mt-2 text-xs p-2 bg-black/10 rounded">
139
+ <span className="font-semibold">Tool [{toolInvocation.toolName}]:</span>{' '}
140
+ {'result' in toolInvocation ? JSON.stringify(toolInvocation.result) : 'Executing...'}
141
+ </div>
142
+ ))}
143
+ </div>
144
+ </div>
145
+ ))}
146
+ </div>
147
+ <form onSubmit={handleSubmit} className="flex gap-2 pt-3 border-t">
148
+ <input
149
+ value={input}
150
+ onChange={handleInputChange}
151
+ placeholder="Ask the agent..."
152
+ className="flex-1 px-3 py-2 border rounded-md"
153
+ />
154
+ <button type="submit" disabled={isLoading} className="px-4 py-2 bg-primary text-primary-foreground rounded-md">
155
+ Send
156
+ </button>
157
+ </form>
158
+ </div>
159
+ );
160
+ }
161
+ ```
162
+
163
+ ---
164
+
165
+ <a name="bahasa-indonesia"></a>
166
+ ## Bahasa Indonesia
167
+
168
+ ### Integrasi Orkestrasi
169
+ Terhubung dan mengorkestrasi skill domain yang relevan seperti `senior-frontend`, `nextjs-app-router-expert`, `ai-llm-integration-expert`, `ui-components-expert`, dan `multi-agent-orchestration` untuk menghadirkan antarmuka AI yang reaktif dan berlatensi rendah.
170
+
171
+ ### Deskripsi
172
+ Panduan produksi untuk membangun aplikasi AI menggunakan **Vercel AI SDK (Core & UI)**. Mencakup abstraksi penyedia model terpadu (`@ai-sdk/anthropic`, `@ai-sdk/openai`, `@ai-sdk/google`), streaming teks dan objek terstruktur (`streamText`, `streamObject`), siklus eksekusi tool multi-langkah otonom dengan `maxSteps`, hook klien React 19 (`useChat`, `useCompletion`), streaming respons saluran data (`createDataStreamResponse`), dan rendering Generative UI.
173
+
174
+ **Sinergi Swarm:** Di dalam **Frontend & UI Swarm**, skill ini berperan sebagai Pemimpin Presentasi UI AI. Skill ini bertugas mentransformasikan keluaran multi-agen backend dan token streaming menjadi komponen web yang interaktif, aksesibel, dan elegan pada Fase 4 & Fase 5.
175
+
176
+ ### Kondisi Pemicu
177
+ - Mengintegrasikan chat percakapan, streaming respons, atau generative UI di React 19 / Next.js 15.
178
+ - Menerapkan ekstraksi data terstruktur dengan validasi skema Zod via `generateObject` atau `streamObject`.
179
+ - Membangun loop pemanggilan tool (*tool-calling loops*) multi-langkah di Route Handler atau Server Actions.
180
+ - Beralih fleksibel antar penyedia model frontier (Claude 3.7 Sonnet, Gemini 3.8 Flash, OpenAI o3/GPT-4.5, Ollama).
181
+ - Mengelola status eksekusi tool, indikator loading, dan rendering komponen UI secara dinamis saat streaming berlangsung.
@@ -6,9 +6,21 @@ author: "Roedy Rustam"
6
6
 
7
7
  # Voice AI Realtime Agent (2026 Edition)
8
8
 
9
+ [English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
10
+
11
+ ---
12
+
13
+ <a name="english"></a>
14
+ ## English
15
+
16
+ ### Description
9
17
  Expert guide for building ultra-low-latency (<300ms), bi-directional conversational voice AI applications. Covers WebRTC, full-duplex WebSocket audio streaming (PCM 24kHz), OpenAI Realtime API, Gemini Multimodal Live API, LiveKit Agents SDK, and smart interruption (barge-in) handling.
10
18
 
11
- *Panduan ahli untuk membangun aplikasi AI suara percakapan dua arah berlatensi ultra-rendah (<300ms) menggunakan WebRTC, WebSocket full-duplex, OpenAI Realtime API, Gemini Multimodal Live API, LiveKit Agents SDK, dan penanganan interupsi (barge-in).*
19
+ ### Trigger Conditions
20
+ - Applications requiring sub-second, spoken conversation with an AI agent.
21
+ - Voice customer service bots, verbal copilots, language tutors, and interactive voice assistants.
22
+ - Implementation of WebRTC audio streaming, full-duplex WebSocket audio (PCM 24kHz), and Silero VAD.
23
+ - Setting up OpenAI Realtime API (`gpt-4o-realtime-preview`) or Gemini Multimodal Live API.
12
24
 
13
25
  ---
14
26
 
@@ -200,3 +212,31 @@ export class GeminiVoiceAgent {
200
212
  - **`realtime-collaboration-expert`**: For syncing WebRTC tracks and room states with client applications.
201
213
  - **`gemini-agent-booster`**: Connects Gemini 3.x / 2.0 Flash thinking models to live voice agents.
202
214
  - **`mobile-expo-expert`**: Audio streaming implementation in React Native with `expo-av` and WebRTC shim.
215
+
216
+ ---
217
+
218
+ <a name="bahasa-indonesia"></a>
219
+ ## Bahasa Indonesia
220
+
221
+ ### Deskripsi
222
+ Panduan ahli untuk membangun aplikasi AI suara percakapan dua arah berlatensi ultra-rendah (<300ms). Mencakup integrasi WebRTC, streaming audio WebSocket full-duplex (PCM 24kHz), OpenAI Realtime API, Gemini Multimodal Live API, LiveKit Agents SDK, dan penanganan interupsi cerdas (*barge-in*).
223
+
224
+ ### Kondisi Pemicu
225
+ - Kebutuhan interaksi percakapan verbal instan di bawah satu detik dengan agen AI.
226
+ - Bot layanan pelanggan berbasis suara, asisten verbal, tutor bahasa interaktif.
227
+ - Implementasi streaming audio WebRTC, WebSocket PCM 24kHz dua arah, dan Silero Voice Activity Detection (VAD).
228
+ - Konfigurasi OpenAI Realtime API (`gpt-4o-realtime-preview`) atau Gemini Multimodal Live API.
229
+
230
+ ### Panduan Inti Arsitektur Suara Real-time
231
+ 1. **Full-Duplex Speech-to-Speech**: Mengganti pipeline sekuensial tradisional (STT ➔ LLM ➔ TTS) dengan pipeline streamable WebRTC atau model native speech-to-speech untuk memangkas latensi dari ~2000ms menjadi ~300ms.
232
+ 2. **Penanganan Interupsi (Barge-In)**: Saat VAD mendeteksi suara pengguna baru saat bot sedang berbicara, buffer audio keluar harus di-flush dalam waktu <150ms tanpa menunggu LLM menyelesaikan kalimatnya.
233
+ 3. **Standar Format Audio**: Input mikrofon PCM 16-bit 16kHz/24kHz Mono, dan output speaker 24kHz untuk intonasi yang alami dan jernih.
234
+
235
+ ---
236
+
237
+ ## Integrasi Orkestrasi
238
+
239
+ - **`ai-llm-integration-expert`**: Routing instruksi sistem dasar dan pemanggilan tool fungsi selama percakapan suara.
240
+ - **`realtime-collaboration-expert`**: Sinkronisasi track audio WebRTC dan status room pengguna.
241
+ - **`gemini-agent-booster`**: Integrasi model multimodal Gemini 2.0/3.x Flash untuk live audio.
242
+ - **`mobile-expo-expert`**: Implementasi audio streaming di React Native menggunakan `expo-av` dan WebRTC shim.