agentar 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Patrick Robinson
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,101 @@
1
+ # Agentar
2
+
3
+ *Agentar (agent + avatar): give your AI agent a body.*
4
+
5
+ agentar is a customizable, realistic 3D avatar for the agent you already use: Claude Code, Codex, or anything that can run a command or call an HTTP endpoint. When your agent talks, agentar speaks the words with text-to-speech and moves the avatar's face in sync with the audio. It lip-syncs, blinks, glances around, nods, and shows moods.
6
+
7
+ It runs locally in your browser with [three.js](https://threejs.org) and WebGL. The default voice is your operating system's own speech engine, so you need no API keys or cloud account.
8
+
9
+ ## Quick start
10
+
11
+ ```bash
12
+ npm install -g agentar
13
+ agentar start # starts the bridge and opens http://localhost:7777
14
+ ```
15
+
16
+ The first start downloads the default CC0 avatar (~37 MB) into `~/.agentar/builtin-models`. You can also run it once with `npx agentar start`.
17
+
18
+ From a checkout of this repo:
19
+
20
+ ```bash
21
+ npm install
22
+ npm run build
23
+ npm start # same as `agentar start`; the avatar goes into assets/models
24
+ ```
25
+
26
+ Click the page once so the browser allows sound. Then press **Speak** on the Talk tab.
27
+
28
+ For development with hot reload:
29
+
30
+ ```bash
31
+ npm run dev # bridge on :7777 and Vite on http://localhost:5173
32
+ ```
33
+
34
+ ## Connect your agent
35
+
36
+ The **Connect** tab in the app shows these commands with the correct paths already filled in. From a checkout, use `node <repo>/packages/cli/dist/index.js` in place of `agentar`.
37
+
38
+ | Harness | Command | What you get |
39
+ | --- | --- | --- |
40
+ | Claude Code (tools) | `claude mcp add agentar -- agentar mcp` | Claude can call `speak`, `set_mood`, `gesture` and `stop_speaking` |
41
+ | Claude Code (every reply) | `agentar install claude-code` | A Stop hook reads each final reply aloud |
42
+ | Codex CLI | `agentar install codex` | `notify` hook and the MCP server in `~/.codex/config.toml` |
43
+ | Anything | `curl -X POST localhost:7777/api/say -d '{"text":"Hi!"}' -H 'Content-Type: application/json'` | Plain HTTP |
44
+
45
+ To get a global `agentar` command, run `npm link -w @agentar/cli`.
46
+
47
+ Before speaking, agentar rewrites markdown for the ear. It removes code blocks, links, and paths, and it trims long replies.
48
+
49
+ ## Customize
50
+
51
+ - **Look**: body model (built-in or your own `.glb`/`.vrm` upload), and the colors of skin, hair, eyes, top, bottom, and shoes. Colors are re-tinted in the shader, so the texture detail stays. Also glasses, hats, and height.
52
+ - **Voice**: engine and voice, speed, pitch, and volume.
53
+ - `system`: offline, uses the macOS `say`, Linux `espeak-ng`, or Windows SAPI voices.
54
+ - `browser`: the Web Speech API.
55
+ - `edge`: Microsoft Edge's online neural voices. They sound much more natural than `say`. Install the free [`edge-tts`](https://github.com/rany2/edge-tts) CLI (`pipx install edge-tts` or `uv tool install edge-tts`). It needs internet, but no API key. If `edge-tts` is not on your `PATH`, set `AGENTAR_EDGE_TTS` to its full path.
56
+ - `openai`: needs `OPENAI_API_KEY`.
57
+ - `elevenlabs`: needs `ELEVENLABS_API_KEY`.
58
+ - `xai`: Grok voices. Needs `XAI_API_KEY` from [console.x.ai](https://console.x.ai). API usage is billed separately from a Grok app subscription.
59
+ - **Behavior**: resting mood, expressiveness, idle motion, and eye contact.
60
+ - **Scene**: framing (head, bust, or full body), lighting, background, and captions.
61
+
62
+ The bridge saves your settings to `~/.agentar/config.json`. It applies them live to every open view.
63
+
64
+ ## How lip-sync works
65
+
66
+ 1. The bridge renders speech audio (WAV or MP3) and sends every connected view a `speak` event.
67
+ 2. The browser decodes the audio. It computes a loudness envelope and finds the regions where speech is voiced.
68
+ 3. The text is converted to **visemes** (mouth shapes) with English letter-to-sound rules. The visemes are laid out over the voiced regions, so pauses in the audio are pauses in the mouth.
69
+ 4. Each frame samples the timeline with cross-fades (coarticulation). The audio loudness scales how wide the mouth opens. The result drives the model's `viseme_*` morph targets, with ARKit and VRM fallbacks.
70
+ 5. The browser voice engine gives no access to its audio. In that case the timeline follows the clock and re-aligns on each word-boundary event.
71
+
72
+ ## Repository layout
73
+
74
+ ```
75
+ packages/core Shared types, config schema, bridge protocol, text→viseme + audio analysis (no dependencies)
76
+ packages/avatar three.js renderer: model loading, pose, idle life, moods, recoloring, accessories, speech player
77
+ apps/bridge Local HTTP + WebSocket server: config storage, TTS engines, speech queue, static hosting
78
+ apps/web Vite app: the avatar view plus the customization panel
79
+ packages/mcp MCP server (stdio) exposing avatar tools to agents
80
+ packages/cli `agentar` command: start, say, mcp, hooks, installers
81
+ docs/ Architecture, roadmap, video-call setup
82
+ ```
83
+
84
+ ## Scripts
85
+
86
+ | Command | Does |
87
+ | --- | --- |
88
+ | `npm run build` | Builds every workspace in dependency order |
89
+ | `npm test` | Runs the Vitest suites (core, bridge, CLI) |
90
+ | `npm run typecheck` | Type-checks every workspace |
91
+ | `npm run dev` | Bridge (watch mode) + Vite dev server |
92
+ | `npm run fetch:models [-- --all]` | Downloads the avatar models |
93
+ | `npm run release:build` | Assembles the publishable `agentar` package in `release/agentar` (see [docs/RELEASING.md](docs/RELEASING.md)) |
94
+
95
+ ## Video calls
96
+
97
+ Put `http://localhost:7777/?stage=1` in an OBS Browser Source, then start OBS's Virtual Camera. The avatar can then appear in Zoom, Teams, or Meet. See [docs/meetings.md](docs/meetings.md), and see [docs/ROADMAP.md](docs/ROADMAP.md) for plans to have the agent join calls itself.
98
+
99
+ ## Credits
100
+
101
+ The lip-sync rules and the rest pose are adapted from [TalkingHead](https://github.com/met4citizen/TalkingHead) (MIT). See [THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md).
@@ -0,0 +1,34 @@
1
+ # Third-party notices
2
+
3
+ ## TalkingHead (MIT)
4
+
5
+ `packages/core/src/lipsync/english.ts` (letter-to-sound rules), `packages/avatar/src/pose.ts` (standing pose), and the arm and hand poses in `packages/avatar/src/gestures.ts` are adapted from TalkingHead by Mika Suominen.
6
+ The English rules derive from NRL Report 7948, "Automatic Translation of English Text to Phonetics by Means of Letter-to-Sound Rules" (Elovitz, Johnson, McHugh & Shore, 1976).
7
+
8
+ ```
9
+ MIT License
10
+
11
+ Copyright (c) 2023-2024 Mika Suominen
12
+
13
+ Permission is hereby granted, free of charge, to any person obtaining a copy
14
+ of this software and associated documentation files (the "Software"), to deal
15
+ in the Software without restriction, including without limitation the rights
16
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
17
+ copies of the Software, and to permit persons to whom the Software is
18
+ furnished to do so, subject to the following conditions:
19
+
20
+ The above copyright notice and this permission notice shall be included in all
21
+ copies or substantial portions of the Software.
22
+
23
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
24
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
25
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
26
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
27
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
28
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
29
+ SOFTWARE.
30
+ ```
31
+
32
+ ## Avatar models
33
+
34
+ The models are downloaded separately and are not part of this repository. See `assets/models/README.md` for their licenses. The default `mpfb.glb` is CC0. The other samples are for non-commercial use only.