vibes-plug 2.11.0 → 2.14.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.cursor/rules/vibes-plug-core.mdc +3 -3
- package/.cursorrules +3 -3
- package/AGENTS.md +4 -4
- package/BLUEPRINT.md +16 -6
- package/CHANGELOG.md +37 -0
- package/CLAUDE.md +8 -8
- package/README.md +85 -115
- package/index.js +1 -1
- package/package.json +2 -2
- package/plugin.json +2 -2
- package/scripts/check-anti-slop.js +53 -0
- package/scripts/generate_swarm_gif.py +2 -2
- package/skills/ai-llm-integration-expert/SKILL.md +22 -15
- package/skills/ai-prompt-engineering-expert/SKILL.md +133 -83
- package/skills/anti-slop/SKILL.md +133 -0
- package/skills/async-queue-temporal-expert/SKILL.md +135 -158
- package/skills/authentication-identity-expert/SKILL.md +172 -278
- package/skills/brainstorming/SKILL.md +26 -26
- package/skills/database-orm-expert/SKILL.md +164 -303
- package/skills/deep-research-analyst/SKILL.md +136 -0
- package/skills/design-system-architect/SKILL.md +31 -1
- package/skills/email-notification-expert/SKILL.md +31 -4
- package/skills/error-resilience-expert/SKILL.md +21 -0
- package/skills/fullstack-expert/SKILL.md +183 -260
- package/skills/glsl-shader-expert/SKILL.md +190 -107
- package/skills/graph-rag-knowledge-expert/SKILL.md +42 -1
- package/skills/mcp-server-architect/SKILL.md +15 -1
- package/skills/prd-architect/SKILL.md +181 -206
- package/skills/production-ready-hardener/SKILL.md +16 -19
- package/skills/pwa-offline-first-expert/SKILL.md +42 -1
- package/skills/pydantic-ai-expert/SKILL.md +161 -0
- package/skills/saas-architect/SKILL.md +154 -0
- package/skills/senior-frontend/SKILL.md +9 -11
- package/skills/senior-frontend/scripts/frontend_scaffolder.py +1 -1
- package/skills/session-memory-manager/SKILL.md +128 -0
- package/skills/synthetic-data-finetuning-expert/SKILL.md +155 -0
- package/skills/ui-ux-pro-max/SKILL.md +4 -2
- package/skills/vercel-ai-sdk-expert/SKILL.md +181 -0
- package/skills/voice-ai-realtime-agent/SKILL.md +41 -1
- package/skills/web-3d-graphics-expert/SKILL.md +313 -137
- package/skills/web-game-engine-expert/SKILL.md +329 -102
- package/skills/webxr-ar-vr-expert/SKILL.md +162 -123
- package/skills/zero-to-prod-orchestrator/SKILL.md +26 -24
- package/skills/ai-cost-token-optimizer/SKILL.md +0 -82
- package/skills/ai-evals-benchmark-expert/SKILL.md +0 -188
- package/skills/asisten-ramah/SKILL.md +0 -47
- package/skills/auto-doc-updater/SKILL.md +0 -220
- package/skills/autonomous-chaos-monkey/SKILL.md +0 -63
- package/skills/background-jobs-queue-expert/SKILL.md +0 -235
- package/skills/database-migration-versioning-expert/SKILL.md +0 -90
- package/skills/edge-serverless-db-expert/SKILL.md +0 -99
- package/skills/mcp-client-orchestrator/SKILL.md +0 -76
- package/skills/mobile-push-notification-expert/SKILL.md +0 -71
- package/skills/monday-design-aesthetic/SKILL.md +0 -73
- package/skills/project-context-mapper/SKILL.md +0 -85
- package/skills/saas-mvp-launcher/SKILL.md +0 -260
- package/skills/saas-transformer/SKILL.md +0 -500
- package/skills/saas-transformer/references/billing_integration_guide.md +0 -401
- package/skills/self-evolving-memory-graph/SKILL.md +0 -91
- package/skills/session-context-loader/SKILL.md +0 -83
- package/skills/session-handoff-resume/SKILL.md +0 -164
- package/skills/skill-baru/SKILL.md +0 -178
- package/skills/supabase-migration/SKILL.md +0 -91
- package/skills/token-saver/SKILL.md +0 -119
- package/skills/ui-components-expert/SKILL.md +0 -166
- package/skills/vibe-code-gardener/SKILL.md +0 -181
- /package/skills/{saas-transformer → saas-architect}/references/feature_gating_patterns.md +0 -0
- /package/skills/{saas-transformer → saas-architect}/references/saas_transformation_checklist.md +0 -0
- /package/skills/{saas-transformer → saas-architect}/scripts/saas_transformation_scanner.py +0 -0
|
@@ -0,0 +1,155 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: synthetic-data-finetuning-expert
|
|
3
|
+
description: "Expert guide for synthetic dataset generation, LLM-as-a-judge filtering, QLoRA fine-tuning (Unsloth), DPO alignment, and GGUF/Ollama export for local SLMs / Panduan ahli generasi data sintetis, fine-tuning QLoRA, DPO, dan ekspor GGUF/Ollama."
|
|
4
|
+
author: vibes-plug-swarm
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Synthetic Data & Fine-Tuning Expert (Custom Domain SLMs)
|
|
8
|
+
|
|
9
|
+
[English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
|
|
10
|
+
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
<a name="english"></a>
|
|
14
|
+
## English
|
|
15
|
+
|
|
16
|
+
### Orchestration & Integration
|
|
17
|
+
Connects and orchestrates with domain skills like `local-slm-edge-ai-expert`, `ai-prompt-engineering-expert`, `ai-evals-benchmark-expert`, `python-programming-expert`, and `ai-cost-token-optimizer` to build high-performance, cost-effective domain models.
|
|
18
|
+
|
|
19
|
+
### Description
|
|
20
|
+
Production-grade guide for generating synthetic training datasets, curating high-signal instruction pairs, executing Parameter-Efficient Fine-Tuning (QLoRA / LoRA) with **Unsloth** and Hugging Face TRL, performing Direct Preference Optimization (DPO), and quantizing custom Small Language Models (SLMs) to GGUF for edge or on-premise execution.
|
|
21
|
+
|
|
22
|
+
**Swarm Synergy:** Within the **AI Engineering Swarm**, this skill serves as the Model Specialization Lead. When frontier API costs or latency become prohibitive, it trains, aligns, and deploys hyper-efficient domain SLMs (1B–8B parameters) in Phase 4.
|
|
23
|
+
|
|
24
|
+
### Trigger Conditions
|
|
25
|
+
- Generating domain-specific synthetic training data from seed documents, codebases, or APIs.
|
|
26
|
+
- Filtering low-quality or hallucinated synthetic data using LLM-as-a-judge curation pipelines.
|
|
27
|
+
- Fine-tuning open-weights models (Llama 3.3, Qwen 2.5, Mistral) on custom tasks using 4-bit QLoRA.
|
|
28
|
+
- Aligning model outputs using Direct Preference Optimization (DPO) to enforce specific response styles.
|
|
29
|
+
- Quantizing fine-tuned models to GGUF (q4_k_m, q8_0) for zero-latency local inference with Ollama or llama.cpp.
|
|
30
|
+
|
|
31
|
+
### Synthetic Data & Fine-Tuning Lifecycle
|
|
32
|
+
|
|
33
|
+
```
|
|
34
|
+
1. SEED EXTRACTION & SYNTHESIS
|
|
35
|
+
[Raw Docs / Codebase] ──► [Frontier LLM / Distilabel] ──► Raw Instruction Pairs (10k+)
|
|
36
|
+
|
|
37
|
+
2. QUALITY FILTERING (LLM-AS-A-JUDGE)
|
|
38
|
+
Raw Instruction Pairs ──► [Rubric Scorer / De-duplication] ──► Curated Gold Dataset (2k-5k)
|
|
39
|
+
|
|
40
|
+
3. 4-BIT QLORA FINE-TUNING (UNSLOTH)
|
|
41
|
+
Base Model (e.g. Qwen 2.5-Coder) + LoRA Adapters ──► SFT / DPO Training Loop
|
|
42
|
+
|
|
43
|
+
4. QUANTIZATION & LOCAL DEPLOYMENT
|
|
44
|
+
Merged 16-bit Weights ──► [llama.cpp GGUF Export] ──► Local Ollama Service (<50ms latency)
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
### Core Implementation Guidelines
|
|
48
|
+
|
|
49
|
+
#### 1. Synthetic Instruction Generation & Rejection Sampling
|
|
50
|
+
Use frontier models to generate input-output pairs with strict rejection criteria:
|
|
51
|
+
```python
|
|
52
|
+
from pydantic import BaseModel, Field
|
|
53
|
+
from openai import OpenAI
|
|
54
|
+
import json
|
|
55
|
+
|
|
56
|
+
client = OpenAI()
|
|
57
|
+
|
|
58
|
+
class SyntheticInstructionPair(BaseModel):
|
|
59
|
+
user_instruction: str = Field(description="Realistic developer query")
|
|
60
|
+
input_context: str = Field(description="Code snippet or API contract")
|
|
61
|
+
ground_truth_response: str = Field(description="Flawless, production-grade output")
|
|
62
|
+
quality_score: int = Field(ge=1, le=5, description="Self-evaluated quality score")
|
|
63
|
+
|
|
64
|
+
def generate_synthetic_samples(seed_text: str) -> list[SyntheticInstructionPair]:
|
|
65
|
+
response = client.beta.chat.completions.parse(
|
|
66
|
+
model="gpt-4o-mini",
|
|
67
|
+
messages=[
|
|
68
|
+
{"role": "system", "content": "Generate 5 diverse, hard edge-case instruction pairs based on the seed code."},
|
|
69
|
+
{"role": "user", "content": seed_text}
|
|
70
|
+
],
|
|
71
|
+
response_format=SyntheticInstructionPair,
|
|
72
|
+
)
|
|
73
|
+
# Filter out anything below score 4 (Rejection Sampling)
|
|
74
|
+
return [sample for sample in [response.choices[0].message.parsed] if sample.quality_score >= 4]
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
#### 2. High-Speed 4-Bit QLoRA with Unsloth
|
|
78
|
+
Fine-tune on consumer GPUs (e.g., RTX 3090/4090 or single A10G) with 5x faster throughput:
|
|
79
|
+
```python
|
|
80
|
+
from unsloth import FastLanguageModel
|
|
81
|
+
import torch
|
|
82
|
+
from trl import SFTTrainer
|
|
83
|
+
from transformers import TrainingArguments
|
|
84
|
+
|
|
85
|
+
# 1. Load Model & Tokenizer in 4-bit
|
|
86
|
+
max_seq_length = 2048
|
|
87
|
+
model, tokenizer = FastLanguageModel.from_pretrained(
|
|
88
|
+
model_name="unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit",
|
|
89
|
+
max_seq_length=max_seq_length,
|
|
90
|
+
load_in_4bit=True,
|
|
91
|
+
)
|
|
92
|
+
|
|
93
|
+
# 2. Add LoRA Adapters
|
|
94
|
+
model = FastLanguageModel.get_peft_model(
|
|
95
|
+
model,
|
|
96
|
+
r=16,
|
|
97
|
+
target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],
|
|
98
|
+
lora_alpha=16,
|
|
99
|
+
lora_dropout=0, # Optimized 0 dropout for Unsloth
|
|
100
|
+
bias="none",
|
|
101
|
+
)
|
|
102
|
+
|
|
103
|
+
# 3. Supervised Fine-Tuning (SFT)
|
|
104
|
+
trainer = SFTTrainer(
|
|
105
|
+
model=model,
|
|
106
|
+
tokenizer=tokenizer,
|
|
107
|
+
train_dataset=dataset,
|
|
108
|
+
dataset_text_field="text",
|
|
109
|
+
max_seq_length=max_seq_length,
|
|
110
|
+
args=TrainingArguments(
|
|
111
|
+
per_device_train_batch_size=2,
|
|
112
|
+
gradient_accumulation_steps=4,
|
|
113
|
+
warmup_steps=10,
|
|
114
|
+
max_steps=100,
|
|
115
|
+
learning_rate=2e-4,
|
|
116
|
+
fp16=not torch.cuda.is_bf16_supported(),
|
|
117
|
+
bf16=torch.cuda.is_bf16_supported(),
|
|
118
|
+
logging_steps=1,
|
|
119
|
+
output_dir="outputs",
|
|
120
|
+
),
|
|
121
|
+
)
|
|
122
|
+
trainer.train()
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
#### 3. GGUF Export for Local Ollama Deployment
|
|
126
|
+
Export merged model to 4-bit or 8-bit quantized GGUF format:
|
|
127
|
+
```python
|
|
128
|
+
# Save to 16bit or GGUF directly
|
|
129
|
+
model.save_pretrained_gguf("custom-domain-coder", tokenizer, quantization_method="q4_k_m")
|
|
130
|
+
|
|
131
|
+
# Generate Modelfile for Ollama:
|
|
132
|
+
# FROM ./custom-domain-coder-q4_k_m.gguf
|
|
133
|
+
# PARAMETER temperature 0.2
|
|
134
|
+
# SYSTEM You are an expert domain coder.
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
---
|
|
138
|
+
|
|
139
|
+
<a name="bahasa-indonesia"></a>
|
|
140
|
+
## Bahasa Indonesia
|
|
141
|
+
|
|
142
|
+
### Integrasi Orkestrasi
|
|
143
|
+
Terhubung dan mengorkestrasi skill domain yang relevan seperti `local-slm-edge-ai-expert`, `ai-prompt-engineering-expert`, `ai-evals-benchmark-expert`, `python-programming-expert`, dan `ai-cost-token-optimizer` untuk membangun model spesialis domain dengan performa tinggi dan biaya hemat.
|
|
144
|
+
|
|
145
|
+
### Deskripsi
|
|
146
|
+
Panduan produksi untuk menghasilkan dataset pelatihan sintetis, mengurasi pasangan instruksi bernilai tinggi, mengeksekusi Parameter-Efficient Fine-Tuning (QLoRA / LoRA) dengan **Unsloth** dan Hugging Face TRL, menerapkan Direct Preference Optimization (DPO), dan mengkuantisasi Small Language Models (SLM) kustom ke format GGUF untuk inferensi lokal atau on-premise berlatensi ultra-rendah.
|
|
147
|
+
|
|
148
|
+
**Sinergi Swarm:** Di dalam **AI Engineering Swarm**, skill ini memegang peranan sebagai Pemimpin Spesialisasi Model. Ketika biaya API atau latensi model cloud frontier terlalu tinggi, skill ini melatih, menyelaraskan (*align*), dan men-deploy SLM domain yang sangat efisien (1B–8B parameter) pada Fase 4.
|
|
149
|
+
|
|
150
|
+
### Kondisi Pemicu
|
|
151
|
+
- Menghasilkan data pelatihan sintetis spesifik domain dari dokumen panduan, codebase, atau skema API.
|
|
152
|
+
- Menyaring data sintetis berkualitas rendah menggunakan pipeline penilaian otomatis *LLM-as-a-judge*.
|
|
153
|
+
- Melakukan fine-tuning model berbobot terbuka (Llama 3.3, Qwen 2.5, Mistral) dengan QLoRA 4-bit secara hemat memori VRAM.
|
|
154
|
+
- Menyelaraskan respon model dengan Direct Preference Optimization (DPO) agar mematuhi aturan format dan gaya tertentu.
|
|
155
|
+
- Mengkuantisasi model hasil fine-tuning ke format GGUF (q4_k_m, q8_0) untuk inferensi lokal instan di Ollama atau llama.cpp.
|
|
@@ -58,7 +58,8 @@ python scripts/search.py "<query>" --domain <domain> --max-results 3
|
|
|
58
58
|
- `icons`: Icon usage guidance, SVG libraries (Lucide, Heroicons), and code imports
|
|
59
59
|
- `react`: React & Next.js performance optimizations, re-render fixes & dynamic imports
|
|
60
60
|
- `web`: Web interface guidelines (ARIA, focus traps, virtual list, form inputs)
|
|
61
|
-
- `m3`: Material Design 3 specific design tokens, color roles, elevation, and component specs
|
|
61
|
+
- `m3`: Material Design 3 specific design tokens, color roles, elevation, and component specs
|
|
62
|
+
- `preset-monday`: Monday.com spacious SaaS aesthetic (Vibrant Blue `#0073ea`, Clean White `#FFFFFF` + Light Grays `#F9F9F9`/`#F5F6F8`, Dark Footer `#111111`, Figtree/Inter fonts, 12-16px card radius, 80-120px vertical spacing)
|
|
62
63
|
|
|
63
64
|
*Example:* `python scripts/search.py "fintech dark theme" --domain color`
|
|
64
65
|
|
|
@@ -159,7 +160,8 @@ python scripts/search.py "<kueri>" --domain <domain> --max-results 3
|
|
|
159
160
|
- `icons`: Panduan penggunaan ikon, pustaka SVG (Lucide, Heroicons), & import kode
|
|
160
161
|
- `react`: Optimasi performa React & Next.js, perbaikan re-render & dynamic import
|
|
161
162
|
- `web`: Pedoman antarmuka web (ARIA, focus trap, virtual list, input form)
|
|
162
|
-
- `m3`: Token desain spesifik Material Design 3, peran warna, elevasi, dan spesifikasi komponen
|
|
163
|
+
- `m3`: Token desain spesifik Material Design 3, peran warna, elevasi, dan spesifikasi komponen
|
|
164
|
+
- `preset-monday`: Estetika SaaS lapang ala Monday.com (Biru Cerah `#0073ea`, Putih Bersih `#FFFFFF` + Abu-abu Terang `#F9F9F9`/`#F5F6F8`, Footer Gelap `#111111`, font Figtree/Inter, radius kartu 12-16px, jarak seksi vertikal 80-120px)
|
|
163
165
|
|
|
164
166
|
*Contoh:* `python scripts/search.py "fintech dark theme" --domain color`
|
|
165
167
|
|
|
@@ -0,0 +1,181 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: vercel-ai-sdk-expert
|
|
3
|
+
description: "Expert guide for Vercel AI SDK (Core, UI, RSC), streaming structured data, multi-provider model switching, tool calling loops, and React 19/Next.js 15 AI engineering / Panduan ahli Vercel AI SDK, streaming data terstruktur, dan integrasi AI pada React 19/Next.js 15."
|
|
4
|
+
author: vibes-plug-swarm
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Vercel AI SDK Expert (Core, UI & Fullstack AI Engineering)
|
|
8
|
+
|
|
9
|
+
[English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
|
|
10
|
+
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
<a name="english"></a>
|
|
14
|
+
## English
|
|
15
|
+
|
|
16
|
+
### Orchestration & Integration
|
|
17
|
+
Connects and orchestrates with domain skills like `senior-frontend`, `nextjs-app-router-expert`, `ai-llm-integration-expert`, `ui-components-expert`, and `multi-agent-orchestration` to deliver reactive, streaming AI interfaces.
|
|
18
|
+
|
|
19
|
+
### Description
|
|
20
|
+
Production-grade guide for building AI applications using the **Vercel AI SDK (Core & UI)**. Covers unified model provider abstraction (`@ai-sdk/anthropic`, `@ai-sdk/openai`, `@ai-sdk/google`), streaming text and structured objects (`streamText`, `streamObject`), dynamic multi-step tool execution loops with `maxSteps`, client-side React 19 hooks (`useChat`, `useCompletion`), streaming data attachments (`createDataStreamResponse`), and generative UI rendering.
|
|
21
|
+
|
|
22
|
+
**Swarm Synergy:** Within the **Frontend & UI Swarm**, this skill serves as the AI UI Presentation Lead. It translates complex backend multi-agent outputs and streaming tokens into accessible, beautiful web components in Phase 4 & Phase 5.
|
|
23
|
+
|
|
24
|
+
### Trigger Conditions
|
|
25
|
+
- Integrating conversational chat, streaming completions, or generative UI in React 19 / Next.js 15.
|
|
26
|
+
- Implementing structured data extraction using `generateObject` or `streamObject` with Zod schemas.
|
|
27
|
+
- Building autonomous multi-step tool-calling loops on Next.js Route Handlers or Server Actions.
|
|
28
|
+
- Switching seamlessly across frontier providers (Claude 3.7 Sonnet, Gemini 3.8 Flash, OpenAI o3/GPT-4.5, Ollama).
|
|
29
|
+
- Building streaming data channels with custom metadata, tool status indicators, and citations.
|
|
30
|
+
|
|
31
|
+
### Vercel AI SDK Architecture (Core vs UI)
|
|
32
|
+
|
|
33
|
+
```
|
|
34
|
+
┌─────────────────────────────────────────────────────────────┐
|
|
35
|
+
│ CLIENT LAYER │
|
|
36
|
+
│ useChat / useCompletion / Generative UI React Components │
|
|
37
|
+
│ • Optimistic updates • Stream reader • Tool invocation │
|
|
38
|
+
└──────────────────────────────▲──────────────────────────────┘
|
|
39
|
+
│ HTTP SSE / Data Stream Protocol
|
|
40
|
+
┌──────────────────────────────▼──────────────────────────────┐
|
|
41
|
+
│ SERVER ROUTE / ACTION │
|
|
42
|
+
│ streamText({ │
|
|
43
|
+
│ model: anthropic('claude-3-7-sonnet-20250219'), │
|
|
44
|
+
│ tools: { weatherTool, dbQueryTool }, │
|
|
45
|
+
│ maxSteps: 5, │
|
|
46
|
+
│ }).toDataStreamResponse() │
|
|
47
|
+
└─────────────────────────────────────────────────────────────┘
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
### Core Implementation Guidelines
|
|
51
|
+
|
|
52
|
+
#### 1. Next.js 15 Route Handler with Multi-Step Tool Calling Loop
|
|
53
|
+
Use `streamText` with `maxSteps` to enable the model to autonomously call tools, review results, and continue reasoning:
|
|
54
|
+
```typescript
|
|
55
|
+
// app/api/chat/route.ts
|
|
56
|
+
import { anthropic } from '@ai-sdk/anthropic';
|
|
57
|
+
import { streamText, tool } from 'ai';
|
|
58
|
+
import { z } from 'zod';
|
|
59
|
+
|
|
60
|
+
export const maxDuration = 60; // Allow long-running agentic reasoning
|
|
61
|
+
|
|
62
|
+
export async function POST(req: Request) {
|
|
63
|
+
const { messages } = await req.json();
|
|
64
|
+
|
|
65
|
+
const result = streamText({
|
|
66
|
+
model: anthropic('claude-3-7-sonnet-20250219'),
|
|
67
|
+
messages,
|
|
68
|
+
maxSteps: 5, // Enables iterative tool calling loop
|
|
69
|
+
tools: {
|
|
70
|
+
calculateMetrics: tool({
|
|
71
|
+
description: 'Computes analytical metrics from raw time series data',
|
|
72
|
+
parameters: z.object({
|
|
73
|
+
datasetId: z.string(),
|
|
74
|
+
metricType: z.enum(['p95_latency', 'error_rate', 'throughput']),
|
|
75
|
+
}),
|
|
76
|
+
execute: async ({ datasetId, metricType }) => {
|
|
77
|
+
const data = await fetchDatasetMetrics(datasetId, metricType);
|
|
78
|
+
return { datasetId, metricType, value: data.result };
|
|
79
|
+
},
|
|
80
|
+
}),
|
|
81
|
+
},
|
|
82
|
+
system: 'You are an elite software performance auditor. Always back up your conclusions with data tool outputs.',
|
|
83
|
+
});
|
|
84
|
+
|
|
85
|
+
return result.toDataStreamResponse();
|
|
86
|
+
}
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
#### 2. Streaming Type-Safe Structured Objects (`streamObject`)
|
|
90
|
+
Stream structured JSON objects directly into the UI while generating:
|
|
91
|
+
```typescript
|
|
92
|
+
import { google } from '@ai-sdk/google';
|
|
93
|
+
import { streamObject } from 'ai';
|
|
94
|
+
import { z } from 'zod';
|
|
95
|
+
|
|
96
|
+
export async function POST(req: Request) {
|
|
97
|
+
const { codeDiff } = await req.json();
|
|
98
|
+
|
|
99
|
+
const result = streamObject({
|
|
100
|
+
model: google('gemini-2.5-flash'),
|
|
101
|
+
schema: z.object({
|
|
102
|
+
securityVulnerabilities: z.array(z.object({
|
|
103
|
+
severity: z.enum(['low', 'medium', 'high', 'critical']),
|
|
104
|
+
cwe: z.string(),
|
|
105
|
+
explanation: z.string(),
|
|
106
|
+
suggestedFix: z.string(),
|
|
107
|
+
})),
|
|
108
|
+
overallRiskScore: z.number().min(0).max(100),
|
|
109
|
+
passesReview: z.boolean(),
|
|
110
|
+
}),
|
|
111
|
+
prompt: `Audit the following git diff for security regressions:\n${codeDiff}`,
|
|
112
|
+
});
|
|
113
|
+
|
|
114
|
+
return result.toTextStreamResponse();
|
|
115
|
+
}
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
#### 3. Client Hook Integration (`useChat` with Tool Invocations)
|
|
119
|
+
Render real-time streaming tokens, loading skeletons, and interactive tool call results:
|
|
120
|
+
```tsx
|
|
121
|
+
'use client';
|
|
122
|
+
|
|
123
|
+
import { useChat } from '@ai-sdk/react';
|
|
124
|
+
|
|
125
|
+
export function AgenticChat() {
|
|
126
|
+
const { messages, input, handleInputChange, handleSubmit, isLoading } = useChat({
|
|
127
|
+
maxSteps: 5,
|
|
128
|
+
});
|
|
129
|
+
|
|
130
|
+
return (
|
|
131
|
+
<div className="flex flex-col h-[600px] w-full max-w-2xl mx-auto border rounded-xl p-4 bg-background">
|
|
132
|
+
<div className="flex-1 overflow-y-auto space-y-4 pr-2">
|
|
133
|
+
{messages.map((m) => (
|
|
134
|
+
<div key={m.id} className={`flex ${m.role === 'user' ? 'justify-end' : 'justify-start'}`}>
|
|
135
|
+
<div className={`p-3 rounded-lg max-w-[80%] ${m.role === 'user' ? 'bg-primary text-primary-foreground' : 'bg-muted'}`}>
|
|
136
|
+
<div className="whitespace-pre-wrap">{m.content}</div>
|
|
137
|
+
{m.toolInvocations?.map((toolInvocation) => (
|
|
138
|
+
<div key={toolInvocation.toolCallId} className="mt-2 text-xs p-2 bg-black/10 rounded">
|
|
139
|
+
<span className="font-semibold">Tool [{toolInvocation.toolName}]:</span>{' '}
|
|
140
|
+
{'result' in toolInvocation ? JSON.stringify(toolInvocation.result) : 'Executing...'}
|
|
141
|
+
</div>
|
|
142
|
+
))}
|
|
143
|
+
</div>
|
|
144
|
+
</div>
|
|
145
|
+
))}
|
|
146
|
+
</div>
|
|
147
|
+
<form onSubmit={handleSubmit} className="flex gap-2 pt-3 border-t">
|
|
148
|
+
<input
|
|
149
|
+
value={input}
|
|
150
|
+
onChange={handleInputChange}
|
|
151
|
+
placeholder="Ask the agent..."
|
|
152
|
+
className="flex-1 px-3 py-2 border rounded-md"
|
|
153
|
+
/>
|
|
154
|
+
<button type="submit" disabled={isLoading} className="px-4 py-2 bg-primary text-primary-foreground rounded-md">
|
|
155
|
+
Send
|
|
156
|
+
</button>
|
|
157
|
+
</form>
|
|
158
|
+
</div>
|
|
159
|
+
);
|
|
160
|
+
}
|
|
161
|
+
```
|
|
162
|
+
|
|
163
|
+
---
|
|
164
|
+
|
|
165
|
+
<a name="bahasa-indonesia"></a>
|
|
166
|
+
## Bahasa Indonesia
|
|
167
|
+
|
|
168
|
+
### Integrasi Orkestrasi
|
|
169
|
+
Terhubung dan mengorkestrasi skill domain yang relevan seperti `senior-frontend`, `nextjs-app-router-expert`, `ai-llm-integration-expert`, `ui-components-expert`, dan `multi-agent-orchestration` untuk menghadirkan antarmuka AI yang reaktif dan berlatensi rendah.
|
|
170
|
+
|
|
171
|
+
### Deskripsi
|
|
172
|
+
Panduan produksi untuk membangun aplikasi AI menggunakan **Vercel AI SDK (Core & UI)**. Mencakup abstraksi penyedia model terpadu (`@ai-sdk/anthropic`, `@ai-sdk/openai`, `@ai-sdk/google`), streaming teks dan objek terstruktur (`streamText`, `streamObject`), siklus eksekusi tool multi-langkah otonom dengan `maxSteps`, hook klien React 19 (`useChat`, `useCompletion`), streaming respons saluran data (`createDataStreamResponse`), dan rendering Generative UI.
|
|
173
|
+
|
|
174
|
+
**Sinergi Swarm:** Di dalam **Frontend & UI Swarm**, skill ini berperan sebagai Pemimpin Presentasi UI AI. Skill ini bertugas mentransformasikan keluaran multi-agen backend dan token streaming menjadi komponen web yang interaktif, aksesibel, dan elegan pada Fase 4 & Fase 5.
|
|
175
|
+
|
|
176
|
+
### Kondisi Pemicu
|
|
177
|
+
- Mengintegrasikan chat percakapan, streaming respons, atau generative UI di React 19 / Next.js 15.
|
|
178
|
+
- Menerapkan ekstraksi data terstruktur dengan validasi skema Zod via `generateObject` atau `streamObject`.
|
|
179
|
+
- Membangun loop pemanggilan tool (*tool-calling loops*) multi-langkah di Route Handler atau Server Actions.
|
|
180
|
+
- Beralih fleksibel antar penyedia model frontier (Claude 3.7 Sonnet, Gemini 3.8 Flash, OpenAI o3/GPT-4.5, Ollama).
|
|
181
|
+
- Mengelola status eksekusi tool, indikator loading, dan rendering komponen UI secara dinamis saat streaming berlangsung.
|
|
@@ -6,9 +6,21 @@ author: "Roedy Rustam"
|
|
|
6
6
|
|
|
7
7
|
# Voice AI Realtime Agent (2026 Edition)
|
|
8
8
|
|
|
9
|
+
[English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
|
|
10
|
+
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
<a name="english"></a>
|
|
14
|
+
## English
|
|
15
|
+
|
|
16
|
+
### Description
|
|
9
17
|
Expert guide for building ultra-low-latency (<300ms), bi-directional conversational voice AI applications. Covers WebRTC, full-duplex WebSocket audio streaming (PCM 24kHz), OpenAI Realtime API, Gemini Multimodal Live API, LiveKit Agents SDK, and smart interruption (barge-in) handling.
|
|
10
18
|
|
|
11
|
-
|
|
19
|
+
### Trigger Conditions
|
|
20
|
+
- Applications requiring sub-second, spoken conversation with an AI agent.
|
|
21
|
+
- Voice customer service bots, verbal copilots, language tutors, and interactive voice assistants.
|
|
22
|
+
- Implementation of WebRTC audio streaming, full-duplex WebSocket audio (PCM 24kHz), and Silero VAD.
|
|
23
|
+
- Setting up OpenAI Realtime API (`gpt-4o-realtime-preview`) or Gemini Multimodal Live API.
|
|
12
24
|
|
|
13
25
|
---
|
|
14
26
|
|
|
@@ -200,3 +212,31 @@ export class GeminiVoiceAgent {
|
|
|
200
212
|
- **`realtime-collaboration-expert`**: For syncing WebRTC tracks and room states with client applications.
|
|
201
213
|
- **`gemini-agent-booster`**: Connects Gemini 3.x / 2.0 Flash thinking models to live voice agents.
|
|
202
214
|
- **`mobile-expo-expert`**: Audio streaming implementation in React Native with `expo-av` and WebRTC shim.
|
|
215
|
+
|
|
216
|
+
---
|
|
217
|
+
|
|
218
|
+
<a name="bahasa-indonesia"></a>
|
|
219
|
+
## Bahasa Indonesia
|
|
220
|
+
|
|
221
|
+
### Deskripsi
|
|
222
|
+
Panduan ahli untuk membangun aplikasi AI suara percakapan dua arah berlatensi ultra-rendah (<300ms). Mencakup integrasi WebRTC, streaming audio WebSocket full-duplex (PCM 24kHz), OpenAI Realtime API, Gemini Multimodal Live API, LiveKit Agents SDK, dan penanganan interupsi cerdas (*barge-in*).
|
|
223
|
+
|
|
224
|
+
### Kondisi Pemicu
|
|
225
|
+
- Kebutuhan interaksi percakapan verbal instan di bawah satu detik dengan agen AI.
|
|
226
|
+
- Bot layanan pelanggan berbasis suara, asisten verbal, tutor bahasa interaktif.
|
|
227
|
+
- Implementasi streaming audio WebRTC, WebSocket PCM 24kHz dua arah, dan Silero Voice Activity Detection (VAD).
|
|
228
|
+
- Konfigurasi OpenAI Realtime API (`gpt-4o-realtime-preview`) atau Gemini Multimodal Live API.
|
|
229
|
+
|
|
230
|
+
### Panduan Inti Arsitektur Suara Real-time
|
|
231
|
+
1. **Full-Duplex Speech-to-Speech**: Mengganti pipeline sekuensial tradisional (STT ➔ LLM ➔ TTS) dengan pipeline streamable WebRTC atau model native speech-to-speech untuk memangkas latensi dari ~2000ms menjadi ~300ms.
|
|
232
|
+
2. **Penanganan Interupsi (Barge-In)**: Saat VAD mendeteksi suara pengguna baru saat bot sedang berbicara, buffer audio keluar harus di-flush dalam waktu <150ms tanpa menunggu LLM menyelesaikan kalimatnya.
|
|
233
|
+
3. **Standar Format Audio**: Input mikrofon PCM 16-bit 16kHz/24kHz Mono, dan output speaker 24kHz untuk intonasi yang alami dan jernih.
|
|
234
|
+
|
|
235
|
+
---
|
|
236
|
+
|
|
237
|
+
## Integrasi Orkestrasi
|
|
238
|
+
|
|
239
|
+
- **`ai-llm-integration-expert`**: Routing instruksi sistem dasar dan pemanggilan tool fungsi selama percakapan suara.
|
|
240
|
+
- **`realtime-collaboration-expert`**: Sinkronisasi track audio WebRTC dan status room pengguna.
|
|
241
|
+
- **`gemini-agent-booster`**: Integrasi model multimodal Gemini 2.0/3.x Flash untuk live audio.
|
|
242
|
+
- **`mobile-expo-expert`**: Implementasi audio streaming di React Native menggunakan `expo-av` dan WebRTC shim.
|