# CLAUDE Source: https://docs.mascot.bot/CLAUDE # CLAUDE.md This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. ## Project Overview This is the documentation site for Mascotbot SDK, built with Mintlify - a modern documentation framework. The site documents the Mascotbot AI-powered mascots SDK with real-time lip-sync technology for React and React Native applications. ## Development Commands ```bash theme={null} # Install Mintlify CLI globally (required for development) npm i -g mintlify # Run development server locally mintlify dev # If dev server fails to start mintlify install ``` ## Architecture & Structure ### Technology Stack * **Documentation Framework**: Mintlify with MDX support * **Theme**: Maple theme with light/dark mode * **API Documentation**: OpenAPI 3.1.0 specification * **Content Format**: MDX (Markdown + JSX components) ### Directory Structure * `/api-reference/` - API documentation and OpenAPI spec * `openapi.yaml` - Viseme Prediction API specification * `endpoint/` - Individual endpoint documentation * `lipsync-integration.mdx` - Integration examples * `/libraries/` - SDK documentation * `react-sdk.mdx` - React SDK guide with WebGL2 renderer * `react-native-sdk.mdx` - React Native SDK guide * `/mascots/` - Pre-built mascot gallery * `/images/` - Documentation assets and screenshots * `/logo/` - Brand assets (light/dark variants) ### Navigation Configuration The site has two main dropdown sections configured in `docs.json`: 1. **Documentation** - Getting started, SDK libraries, ready-made mascots 2. **API Reference** - API introduction, endpoints, integration examples ## Key Development Tasks ### Adding Documentation Pages 1. Create MDX file in appropriate directory 2. Add frontmatter with title, description, and icon 3. Update navigation in `docs.json` under appropriate group 4. Preview locally with `mintlify dev` ### Updating API Documentation 1. Modify `api-reference/openapi.yaml` for API spec changes 2. Update endpoint MDX files in `api-reference/endpoint/` 3. Ensure examples match the OpenAPI specification ### Working with Code Examples * Use proper language identifiers for syntax highlighting * Include complete, runnable examples * Test code examples for accuracy * Show both TypeScript and JavaScript variants where applicable ### Testing Changes 1. Always preview with `mintlify dev` before committing 2. Check both light and dark theme appearance 3. Verify all links and images load correctly 4. Test interactive MDX components render properly 5. Ensure responsive design works on mobile viewports ## Content Guidelines ### MDX Best Practices * Use semantic headings (start with ##, not #) * Include descriptive frontmatter for SEO * Use Mintlify components: Note, Warning, Tip, Info callouts * Add code examples with proper syntax highlighting * Include screenshots and diagrams where helpful ### API Documentation Standards * Keep OpenAPI spec synchronized with actual API * Document all parameters, request/response bodies * Include realistic example values * Document error responses and status codes * Show streaming (SSE) examples for real-time endpoints ### SDK Documentation Focus * Provide complete setup instructions * Show common use cases with code examples * Document all hooks and components * Include troubleshooting sections * Explain performance considerations (e.g., WebGL2 renderer) ## Important Notes * The main branch auto-deploys via GitHub App integration * Mintlify provides built-in search functionality * External links configured: support email, Twitter/X, main website * The site uses contextual menu with copy, view, ChatGPT, and Claude options * Focus on accuracy - this is the primary resource for developers using Mascotbot SDK # Changelog Source: https://docs.mascot.bot/changelog Product updates, new features, and improvements ## SDK 0.2.3 — interactive avatar SDK, offline timelines, unified streaming The SDK is now two packages — `@mascotbot/core` and `@mascotbot/react` — a single composable surface for real-time interactive avatars. Continuously-updated models and assets are delivered from Mascotbot; you ship the avatar. The documentation has been rewritten end to end for the new surface. * **Real-time pipeline** — audio in, animated avatar out; one composable surface. * **Serializable offline timeline** — `processAudio()` returns a versioned `VisemeTimeline`. Persist it as JSON and replay later with zero reprocessing via `parseTimeline` + `setTimeline`. * **One streaming hook** — `useLipsyncStream` handles live audio with `source: mic | mediaStream | manual`. * **Realtime, your way** — wire OpenAI Realtime, Gemini Live, or ElevenLabs with their own SDKs and tap the audio; `createPCMStreamPlayer` plays raw-PCM providers and exposes the tap. * **Focused Rive surface** — `useMascotRive` + `useMascotInputs` (with `has()`); the SDK writes only mouth, `is_speaking`, and `stress` — everything else stays yours. * **Private registry install** — `npm.mascot.bot` with an `.npmrc` token; `mascot_dev_…` for localhost, `mascot_pub_…` for production. Upgrading an existing integration? The [migration guide](https://docs.mascot.bot/reference/migration) maps every change. [Read the docs →](https://docs.mascot.bot/overview) ## Ship your agent on your own domain Attach a subdomain or apex to any hosted agent, then set a per-agent page title, description, OG preview, and favicon so shared links look like yours — not ours. A cream editorial card titled "Bring your own domain" floating on a warm off-white canvas with faint concentric rings radiating behind — the card holds a globe icon pill at the top, the headline, a short subtitle, a labeled "Your domain" input with an https:// prefix and the placeholder chat.yourcompany.com, and a dark "Attach domain" button at the bottom * Custom domain attach via the Vercel Domains API — paste the hostname, copy one DNS record, we auto-poll and verify the moment propagation lands. * Per-agent SEO: ``, description, OG image, favicon — with a live Slack / Twitter / browser-tab preview right next to the inputs. * Works for both standalone and widget-hosted agents, subdomain or apex. [Try it →](https://app.mascot.bot/agents) </Update> <Update label="2026-04-22"> ## A lipsync playground against real voice streams The new [`/lipsync-test`](https://www.mascot.bot/lipsync-test) tool lets you upload any Rive mascot, wire a Gemini Live session or plug in your own ElevenLabs agent, and dial every SDK knob — critical-viseme holds, min intervals, transition speeds — while the mascot responds in real time. <Frame> <img alt="The /lipsync-test page rendered as a tilted editorial product shot on a warm cream canvas with a subtle dot grid — a friendly cartoon blue cat mascot sits centered on a sandy beach under a blue sky, flanked on the right by a vertical stack of cream settings panels for Rive File, Display Settings, Voice Provider (Gemini Live / ElevenLabs), Voice, Lip Sync Model, and Lip Sync Settings" /> </Frame> * Switch providers on the fly: **Gemini Live** (preset) or paste your own ElevenLabs API key + agent ID — credentials stay in your browser. * Every `NaturalLipSyncConfig` field is a live slider; changes apply mid-call without resetting playback. * Chip selector for the critical viseme set (r, l, f/v, p/b/m by default; toggle any of 21 to taste). * Replay captured responses against tweaked settings to A/B feel without re-speaking. * Upload any `.riv` file to test your own mascot; HD-zoom + pan the canvas to inspect mouth shapes. [Try it →](https://www.mascot.bot/lipsync-test) </Update> <Update label="2026-04-22"> ## Tune your mascot's lipsync mid-call, without reset Mascotbot SDK `0.2.0` adds critical-viseme holds, a configurable critical set, and a live-update pipeline — every phoneme knob now takes effect on the next viseme chunk, mid-call, without breaking the stream. <Frame> <img alt="A two-column release-card composition on a dark canvas — left, a real cream-colored Lip Sync Settings panel with the Natural Lip Sync toggle on, the Natural Lip Sync preset active, and five orange-tracked sliders for Min Viseme Interval, Merge Window, Key Viseme Preference, Similarity Threshold, and Critical Viseme Min Duration; right, a huge metallic-white wordmark reading SDK 0.2.0 on two lines" /> </Frame> * `criticalVisemeMinDuration` holds u/o/r/l/f/v/p/b/m phonemes long enough to register, dropping any non-critical viseme that would interrupt the hold window. * `criticalVisemeIds` exposes the critical set — drop vowels, add sibilants, whatever fits your character. Exports `DEFAULT_CRITICAL_VISEME_IDS` for consumers to derive custom lists. * `desktopTransitionSpeed` / `mobileTransitionSpeed` set the mouth-blend rate; default bumped from `11` → `22` for snappier blends. * Config edits now flow through the active `MascotPlayback` in place; slider tweaks and preset changes don't tear down the stream. * Rive runtime bumped to `@rive-app/react-webgl2@4.28.1` / `@rive-app/webgl2@2.37.2` — latest upstream fixes and perf work on the WebGL2 renderer. * Perf tuning validated on [`/lipsync-test`](https://www.mascot.bot/lipsync-test): memoized side panels, ref-based viewport pan/zoom, identity short-circuit on per-chunk config sync, NoiseOverlay visibility gate. [Download →](https://app.mascot.bot/sdk-access) </Update> <Update label="2026-04-18"> ## Clone a persistent voice AI agent for React and deploy in 30 minutes The new `react-website-demo` template is a Next.js 16 starter with an ElevenLabs voice widget that survives every page navigation — no re-init, no lost context — plus a companion tutorial that walks through the Context + Router pattern behind it. <Frame> <img alt="Two layered windows on a dark canvas — behind, a VS Code–style editor showing the react-website-demo template's VoiceProvider.tsx with its tab strip, sidebar file tree, and status bar; in front, a Safari-style browser window rendering the live MovingCo "Moving Made Simple" page with the voice-chat mascot in the bottom-right — code and live preview side-by-side" /> </Frame> * **Persistent across every route** — voice session lives in a React Context singleton, so conversations don't die when users click to a new page * **Three client tools wired out of the box** — the agent reads/writes your form, routes between pages, and triggers CTAs * **Full stack** — Next.js 16, React 19, TypeScript, ElevenLabs Conversational AI 0.5, MascotBot SDK 0.1.9 * **\<300ms voice latency · 30–45 min setup** — clone, set three env vars, one-click deploy to Vercel * **Companion tutorial** walks through the Context + Router pattern so you can adapt it to any React app [Read more →](https://templates.mascot.bot/voice-ai-agent-react-tutorial) </Update> <Update label="2026-04-18"> ## Drop a voice avatar onto any website with one line of HTML Design your widget visually in the dashboard — size, paddings, mobile overrides, custom button label — then paste a single `<script>` tag to ship it. Resize or restyle anytime; the live widget updates instantly without re-pasting the embed. <Frame> <img alt="Live widget hosting demo on mascot.bot — embedded mascot with Voice Chat button on a mock customer site" /> </Frame> * Visual builder with live desktop + mobile previews — drag sliders or type exact pixel values * Optional mobile-specific size + paddings with a configurable breakpoint * Bring your own Rive mascot or pick from the preset library * Pick any provider under the hood — ElevenLabs, Gemini Live, or OpenAI Realtime * Custom idle-state label on the Voice Chat button, plus an optional first-load reveal animation * Embed script auto-syncs from saved settings — no re-paste required after a redesign [Read more →](https://app.mascot.bot/agents/new) </Update> <Update label="2026-04-15"> ## More hours per dollar on every tier, with new Growth and Scale plans Pro now includes 150 hours of lipsync at the same \$149/month (2× the old allowance), Business drops overage rates by 75%, and two brand-new tiers (Growth \$499, Scale \$999) cover production workloads up to 5,000 hours per month. Existing subscribers are grandfathered on their current pricing. * **Starter** — \$49/mo · 20 h · \$2.99/h overage (unchanged) * **Pro** — \$149/mo · **150 h** (was 75 h) · **\$0.90/h** overage (was \$2.48/h) * **Business** — \$299/mo · 600 h · \$0.60/h overage * **Growth (new)** — \$499/mo · 1,500 h · \$0.45/h overage * **Scale (new)** — \$999/mo · 5,000 h · \$0.25/h overage * Yearly plans still get 20% off * Existing customers stay on their current pricing indefinitely — no migration [Read more →](https://app.mascot.bot/subscription) </Update> <Update label="2026-03-10"> ## Host voice agents on Gemini Live or OpenAI Realtime, not just ElevenLabs Pick your preferred real-time model when creating a hosted agent — ElevenLabs, Gemini Live, or OpenAI Realtime — and get the same lip-sync, gestures, and public share link across all three. <Frame> <img alt="Provider selector showing ElevenLabs, Gemini Live, and OpenAI Realtime cards in the agent creation wizard" /> </Frame> * Provider selector in the agent creation wizard; each agent is locked to one provider for its lifetime * **Gemini Live** — 30 Google voices, custom system instructions, configurable model, optional initial greeting * **OpenAI Realtime** — 10 voices (marin, cedar, alloy, …), `gpt-realtime` model, tunable voice-activity detection * **Video input** — Gemini agents accept live camera frames for vision-aware conversations * API keys stay encrypted server-side (AES-256-GCM); clients only ever see 9-minute ephemeral tokens * Works in both the standalone share link and the embeddable widget [Try it →](https://app.mascot.bot/agents/new) </Update> # Licensing & API Keys - Mascotbot Lipsync SDK Session Model Source: https://docs.mascot.bot/concepts/licensing-and-keys How Mascotbot lipsync SDK licensing works: development vs production keys, origin enforcement, the init / background-refresh session lifecycle, usage metering, and the key security model. Authentication uses a Mascotbot license key (`mascot_…`). Every session is authorized by the Mascotbot edge service, which also delivers the licensed model and the mascot / voice assets. Lipsync runs in realtime with no audio roundtrip — only licensing, model/asset delivery, and usage metering use the network. ## Key environments | Prefix | Valid origins | Metering | | -------------- | -------------------------------------------------------------- | ------------------------------- | | `mascot_dev_…` | `localhost`, `*.localhost`, `127.0.0.1`, private networks only | Development meters — no billing | | `mascot_pub_…` | Your registered, allow-listed public domains | Production meters | The pairing is enforced server-side and is intentionally strict: * A `mascot_dev_…` key from a public origin → rejected (`dev_key_on_public_domain`). * A `mascot_pub_…` key from `localhost` → rejected (`prod_key_on_localhost`). * A `mascot_pub_…` key from an origin not on its allow-list → rejected (`origin_not_allowed`). `devMode` is auto-detected for `localhost`, `*.localhost`, `127.0.0.1`, and private IPs; it skips the Origin allow-list and routes events to the dev meters. Get and manage keys at [app.mascot.bot/api-keys](https://app.mascot.bot/api-keys). ## Session lifecycle <Steps> <Step title="Init"> On mount, `MascotProvider` / `LipsyncClient.init` posts your key to the edge worker. On success the worker returns a short-lived license and the WASM runtime; `status` moves `initializing → ready`. </Step> <Step title="Background refresh"> While the client is active, the session auto-refreshes ahead of expiry. You do nothing — the `"refresh"` event fires on each successful cycle if you want to observe it. </Step> <Step title="Release"> Call `client.stop()` (or unmount the provider) to release resources. A backgrounded tab whose refresh ticks were throttled long enough can expire its session; the next call throws `RefusedError` with code `session_expired` and the user should reload. </Step> </Steps> `status` is one of `idle`, `initializing`, `ready`, `running`, `degraded`, `refused`, `error`. Read it from `useMascot()` (React) or `client.status` (vanilla), and gate audio work on `status === "ready"`. ## Usage metering Usage is attributed to your account through one of two meters, depending on plan: * **Speech seconds** — the entry plan bills by processed speech time. `processAudio()` and streaming sessions report `speechMs` (non-silent milliseconds); a persisted [`VisemeTimeline`](/concepts/visemes-and-timeline) carries `speechMs`, so replaying a stored timeline does not re-meter. * **Monthly active users (MAU)** — higher plans bill by MAU, attributed automatically per session. ## What your plan includes The subscription is the avatar platform around the SDK — you ship the avatar, Mascotbot keeps it running and improving: * **Continuously-updated models** delivered to the SDK (no rebuild on your side). * **The mascot & voice asset library** — ready-made avatars, or bring your own Rive ([Ready-made Mascots](/mascots/ready-to-use-mascots)). * **The commercial license** to ship those models and assets in your product. * **SDK + API access**, usage analytics, and support. Plans scale from a speech-seconds entry tier up to MAU-based tiers; current tiers and limits are at [mascot.bot pricing](https://mascot.bot/#pricing). ## Key security model * **`mascot_pub_…` publishable keys are safe in the browser bundle.** They are scoped to your allow-listed origins and to the runtime surface — a scraped production key cannot be used from another origin and cannot publish packages. This is why the key can live in `NEXT_PUBLIC_…` / client config. * **`.npmrc` registry tokens are not.** They grant install access to the private registry — keep them out of version control and inject from a CI secret. * **Standing third-party keys never touch the browser.** When you add a realtime AI provider or TTS, your OpenAI / Gemini / ElevenLabs *standing* key stays on the server; a route handler mints a short-lived client secret, ephemeral token, or signed URL per session. See [Realtime providers](/realtime/overview). ## When authorization fails License failures surface as typed errors with an actionable `.code` and `.message` so you can route the user to the right fix (re-subscribe, update card, swap dev/prod key). Branch on `error.code`, not the subclass. The full HTTP-status → code matrix and recommended UI per code is in [Error codes](/reference/error-codes). ## Next <Columns> <Card title="Installation" icon="download" href="/installation"> Registry + `.npmrc` setup. </Card> <Card title="Error codes" icon="triangle-exclamation" href="/reference/error-codes"> Every refusal code and its fix. </Card> <Card title="Realtime providers" icon="bolt" href="/realtime/overview"> Server-minted provider tokens. </Card> </Columns> # Rive Co-existence - The SDK Does Not Own Your Rive Instance Source: https://docs.mascot.bot/concepts/rive-coexistence The Mascotbot SDK writes only mouth visemes, is_speaking, and stress. Every other Rive input, ViewModel, event, and listener stays yours on the raw rive instance. ## The contract <Note> The SDK writes exactly three input families: mouth visemes (`100..118`), `is_speaking`, and `stress`. Everything else on the Rive instance — other state-machine inputs, data binding / ViewModels, events, listeners — is owned by the consumer and accessed directly on the raw `rive` object. The SDK never wraps, gates, proxies, or constrains that. `rive` is always fully exposed. </Note> This is a hard design rule, not a guideline. ## Why Avatars do far more than lip sync: gender / skin / outfit inputs, gesture triggers, click events, data-bound ViewModels, scene state. If the SDK owned the Rive instance, every one of those would have to be re-exposed through SDK API forever, and the SDK would become a bottleneck on the Rive runtime's own evolution. Keeping the SDK to a three-input writer keeps integration "bring your own Rive, we animate the mouth" — composable, and future-proof against Rive API changes. ## Get the raw instance ```tsx theme={null} import { useMascotRive } from "@mascotbot/react/rive"; function Costume() { const { rive } = useMascotRive(); // inside <Mascot> // rive is the unmodified @rive-app/* instance. // Set any input, fire any trigger, attach EventType.RiveEvent listeners, // bind a ViewModel — none of it involves the SDK. } ``` Framework-agnostic, the same is true of `getRiveInputs(rive)` from `@mascotbot/core/rive` — it reads inputs off a Rive instance you constructed and own. ## Presence checks — use `has()`, not raw introspection A missing input handle resolves to a silent no-op shim (`DEFAULT_SM_INPUT`) so the SDK's own mouth writes never throw on an artboard that lacks a viseme. That shim is **structurally identical** to a real input — you cannot tell them apart by inspection. To know whether an input actually exists, ask: ```tsx theme={null} import { useMascotInputs } from "@mascotbot/react/rive"; function Wave() { const { has, custom } = useMascotInputs<"wave">(); if (has("wave")) custom.wave.fire(); // consumer-owned, SDK-untouched return null; } ``` `has(name)` is the authoritative check. The framework-agnostic equivalents are `getRiveInputs(rive).has(name)` and `hasRiveInput(rive, name)` in `@mascotbot/core/rive`. `custom` from `useMascotInputs` is **never `undefined`**, which removes the optional-chaining tax from consumer code. Drive the input itself however you like — that part is entirely yours. ## Rive file requirements | Element | Requirement | | ------------- | ---------------------------------------------------------------------------------------- | | Artboard | `Character` | | State machine | `mascotStateMachine` (the SDK's input lookup also accepts the alternate `InLesson` name) | | Mouth inputs | Number inputs `100`–`118` (viseme ids) | | Optional | `is_speaking`, `eyes_smile`, `stress`, plus any consumer inputs (e.g. `gesture`) | Pass **only** `mascotStateMachine` in the `stateMachines` array to `new Rive(...)` (or rely on `STATE_MACHINE_NAMES[0]`). Rive 2.37+ throws on any unknown state-machine name in that array; the throw propagates through `initStateMachines`, fires `LoadError`, suppresses `Load`, and leaves the canvas blank. ## Next <Columns> <Card title="React hooks" icon="react" href="/libraries/react-hooks"> `useMascotRive`, `useMascotInputs`. </Card> <Card title="React SDK" icon="cube" href="/libraries/react-sdk"> Provider and client components. </Card> <Card title="Migration" icon="arrow-right-arrow-left" href="/reference/migration"> The 0.2.x symbol map. </Card> </Columns> # Visemes & the Viseme Timeline - Mascotbot Lip Sync Data Model Source: https://docs.mascot.bot/concepts/visemes-and-timeline How the Mascotbot SDK represents lip sync: viseme ids, the serializable run-length VisemeTimeline, and the framesToTimeline / timelineToCues / parseTimeline helpers. A **viseme** is the visual shape of the mouth for a sound — the visual counterpart of a phoneme. The SDK emits one of 22 viseme ids (`0`–`21`). Internally, ready-made Rive mascots map those onto number inputs `100`–`118` via the exported `VISEMES_MAP`; you rarely touch the raw ids directly. ## The model output Inference produces one viseme id per **10 ms frame**. `client.processAudio()` returns that as a `VisemeTimeline`, not a raw array: ```ts theme={null} const { timeline, durationMs, speechMs } = await client.processAudio(audio16kMono); ``` `processAudio()` returns the timeline directly. It is run-length-encoded, \~10× smaller than a per-frame array, and is exactly the change-event model the playback engine consumes, so there is no second representation to keep in sync. ## The `VisemeTimeline` shape ```ts theme={null} interface VisemeCue { t: number; // start time in ms v: number; // viseme id 0..21 } interface VisemeTimeline { version: number; // VISEME_TIMELINE_VERSION (currently 1) durationMs: number; // total audio duration speechMs: number; // non-silent ms detected (metering, preserved across persist/replay) frameMs: number; // engine frame interval the cues align to (10) cues: VisemeCue[]; // run-length; strictly increasing t; first cue is t: 0 } ``` It is plain JSON. Persist it anywhere — `localStorage`, your database, a CDN, a file — and replay it later without touching the model, the network, or a license refresh. ```json theme={null} { "version": 1, "durationMs": 1840, "speechMs": 1610, "frameMs": 10, "cues": [ { "t": 0, "v": 0 }, { "t": 120, "v": 7 }, { "t": 260, "v": 19 }, { "t": 410, "v": 0 } ] } ``` ## Helpers All three are pure functions on the package root. | Helper | Signature | Purpose | | ------------------ | -------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- | | `framesToTimeline` | `(argmax: readonly number[], opts: { speechMs: number; frameMs?: number }) → VisemeTimeline` | Build a timeline from a per-frame viseme array (if you assemble visemes yourself). | | `timelineToCues` | `(tl: VisemeTimeline) → { offset: number; visemeId: number }[]` | Expand a timeline into the cue list `MascotPlayback` consumes. The inverse of the change-event encoding. | | `parseTimeline` | `(input: unknown) → VisemeTimeline` | Validate untrusted/persisted JSON and return a typed timeline, or throw. | ```ts theme={null} import { framesToTimeline, timelineToCues, parseTimeline } from "@mascotbot/core"; ``` ## `parseTimeline` is the trust boundary Persisted JSON outlives SDK versions. `parseTimeline` is the single gate for loading a timeline back: it validates `version`, `frameMs`, monotonic cue offsets, the leading `t: 0`, and viseme-id ranges. On any mismatch it throws a `LipsyncError` whose `.code === "bad_timeline"`: ```ts theme={null} import { parseTimeline, LipsyncError } from "@mascotbot/core"; try { const tl = parseTimeline(JSON.parse(stored)); playback.setTimeline(tl); } catch (err) { if (err instanceof LipsyncError && err.code === "bad_timeline") { // stale or corrupt — regenerate via client.processAudio() } } ``` `VISEME_TIMELINE_VERSION` is bumped on any breaking shape or semantics change, so an old stored timeline fails loudly instead of animating garbage. Treat `bad_timeline` as "regenerate", never as a license or network condition. It is documented in the [error-code reference](/reference/error-codes). ## Next <Columns> <Card title="Offline lip sync" icon="box-archive" href="/libraries/offline-lipsync"> Generate → persist → replay in practice. </Card> <Card title="Rive co-existence" icon="puzzle-piece" href="/concepts/rive-coexistence"> How visemes reach the avatar. </Card> <Card title="Core client" icon="cube" href="/core/client"> `processAudio`, streaming sessions. </Card> </Columns> # LipsyncClient - Vanilla JavaScript Lip Sync Core API Source: https://docs.mascot.bot/core/client Use @mascotbot/core without React: LipsyncClient.init, processAudio, createStreamingSession, resample, diagnostics, events, and the vanilla Rive engine. `@mascotbot/core` is the framework-agnostic engine — no React, no Rive on the root entry. Use it in vanilla apps, Web Workers, or Node-adjacent environments. React apps get all of this re-exported through [`@mascotbot/react`](/libraries/react-sdk). ## Initialize ```ts theme={null} import { LipsyncClient } from "@mascotbot/core"; const client = await LipsyncClient.init({ apiKey: "mascot_pub_…", userId: "user_42", // optional — stable id for MAU attribution // licenseEndpoint, devMode, logger, fingerprintHash also accepted }); ``` `init` exchanges your key with the edge worker and loads the WASM runtime. Configuration mirrors [`<MascotProvider>`](/installation#4-configure-the-client). ## `processAudio(samples, opts?)` One-shot inference over a recorded buffer. `samples` is **16 kHz mono Float32 in `[-1, 1]`**; pass `opts.sampleRate` if your buffer is at another rate and let the client resample. ```ts theme={null} const { timeline, durationMs, speechMs } = await client.processAudio(audio16kMono); // timeline — VisemeTimeline (the serializable artifact) // durationMs — total audio duration // speechMs — non-silent ms detected ``` `timeline` is the offline artifact — persist it as JSON and replay later with zero reprocessing ([Offline lip sync](/libraries/offline-lipsync)). ### `resample(samples, fromSampleRate, toSampleRate?)` Linear resampler to prep audio for `processAudio` / streaming. `toSampleRate` defaults to `16000`. ```ts theme={null} const at16k = client.resample(decoded.getChannelData(0), decoded.sampleRate, 16000); ``` ## `createStreamingSession()` For live audio, push 25 ms (400-sample) windows of 16 kHz mono one at a time. ```ts theme={null} const session = client.createStreamingSession(); const frame = await session.pushWindow(audioWindow); // Promise — must await console.log(frame.visemeId, frame.silenceDetected, frame.frameIndex); session.close(); ``` `pushWindow` is async by design (inference is off the main thread). End-of- utterance phantom visemes are suppressed by the SDK's internal −50 dBFS silence gate. Full streaming guide, including the React wrapper: [Streaming sessions](/core/streaming). ## Events `client.on(event, listener)` returns an unsubscribe function. | Event | Fires when | | ----------- | ------------------------------------------------------ | | `"ready"` | Init finished; inference available | | `"refresh"` | A background license refresh succeeded (`{ epoch }`) | | `"refused"` | Authorization refused — listener gets a `RefusedError` | | `"error"` | A runtime error — listener gets an `Error` | ```ts theme={null} const off = client.on("refused", (err) => showBilling(err.code, err.message)); // later: off(); ``` `client.status` is the current `LipsyncStatus` (`idle | initializing | ready | running | degraded | refused | error`). ## `diagnostics()` Async — resolves to an opaque health snapshot. Useful for a debug panel. Always `await` it (it is computed off-thread): ```ts theme={null} const d = await client.diagnostics(); // { pendingSpeechMs, classifierEpoch, riskBand, installId, variantId } ``` ## Teardown ```ts theme={null} client.stop(); // stop active work, keep the instance reusable client.close(); // release everything; the instance is done ``` Call one of these when you are finished so resources and the audio graph are released. ## Vanilla Rive playback The Rive engine lives on the `/rive` subpath. Construct Rive yourself, read its inputs, and drive `MascotPlayback`: ```ts theme={null} import { Rive, EventType } from "@rive-app/webgl2"; import { MascotPlayback, getRiveInputs, STATE_MACHINE_NAMES } from "@mascotbot/core/rive"; const rive = new Rive({ src: "/mascot-fox.riv", canvas: document.querySelector("canvas")!, autoplay: true, stateMachines: STATE_MACHINE_NAMES[0], // "mascotStateMachine" — pass only this }); rive.on(EventType.Load, () => { const riveInputs = getRiveInputs(rive); const playback = new MascotPlayback({ riveInputs, stream: true, enableNaturalLipSync: true }); // Offline: replay a whole timeline (from processAudio or parseTimeline). playback.setTimeline(timeline); // Streaming alternative: playback.pushVisemes([{ offset: 0, visemeId: 7 }]); playback.play(); }); ``` `MascotPlayback` methods: `setTimeline(tl)`, `pushVisemes(cues)`, `play()`, `pause()`, `seek(ms)`, `reset()`, `setTransitionSpeed(n)`. `getRiveInputs(rive)` returns a bundle with `.has(name)`; `hasRiveInput(rive, name)` is the standalone presence check. The SDK only writes mouth / `is_speaking` / `stress` — see [Rive co-existence](/concepts/rive-coexistence). ## Next <Columns> <Card title="Streaming sessions" icon="tower-broadcast" href="/core/streaming"> `createStreamingSession` in depth. </Card> <Card title="PCM stream player" icon="waveform" href="/core/pcm-stream-player"> Play + tap raw provider PCM. </Card> <Card title="Offline lip sync" icon="box-archive" href="/libraries/offline-lipsync"> Generate → persist → replay. </Card> </Columns> # createPCMStreamPlayer - Play & Tap Raw PCM for Realtime Lip Sync Source: https://docs.mascot.bot/core/pcm-stream-player createPCMStreamPlayer plays streamed PCM16 gap-tolerantly and exposes a parallel MediaStream tap — the bridge for realtime AI providers that hand you raw audio. `createPCMStreamPlayer` does one thing: **play streamed PCM16 gap-tolerantly, and expose exactly what is playing as a `MediaStream`** you can feed to lip sync. It is the bridge for realtime AI providers that hand you raw audio chunks and do not play them (Gemini Live, OpenAI Realtime over WebSocket) and for server TTS that returns audio only. It knows nothing about any provider — provider transport parsing is your glue and stays in your app. ```ts theme={null} import { createPCMStreamPlayer } from "@mascotbot/core"; const player = createPCMStreamPlayer({ sampleRate: 24000 }); ``` ## API ```ts theme={null} const player = createPCMStreamPlayer(options); ``` `options: PCMStreamPlayerOptions` | Option | Type | Notes | | ----------------- | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `sampleRate` | `number` | Required. The PCM sample rate (e.g. `24000` for Gemini Live and OpenAI Realtime). | | `initialBufferMs` | `number?` | Pre-roll before playback starts (jitter cushion). | | `scheduleAheadMs` | `number?` | How far ahead chunks are scheduled. | | `onIdle` | `(() => void)?` | Called once each time the player drains **naturally** (queue empty + every scheduled buffer finished). Not called by `stop()`/`close()`; may re-fire after a later push. See [Knowing when playback finished](#knowing-when-playback-finished). | Returned `PCMStreamPlayer`: | Member | Type | Purpose | | ---------------------- | ------------------------ | ------------------------------------------------------------------------------------------------------------ | | `pushBase64PCM16(b64)` | `(string) => void` | Enqueue a base64 PCM16 chunk (Gemini `inlineData.data`, server TTS). | | `pushPCM16(bytes)` | `(Uint8Array) => void` | Enqueue a PCM16 byte chunk (OpenAI Realtime WS `ArrayBuffer`). | | `outputStream` | `MediaStream` (readonly) | The tap — feed this to `useLipsyncStream`. Runs parallel to the speakers. | | `isPlaying` | `boolean` (readonly) | `true` while audio is queued or scheduled; flips `false` on natural drain (when `onIdle` fires) or `stop()`. | | `stop()` | `() => void` | Drop queued audio immediately (barge-in / interruption). | | `resume()` | `() => void` | Resume the underlying `AudioContext` (call inside a user gesture). | | `close()` | `() => Promise<void>` | Tear down the player and its audio graph. | ## Pattern Create the player **inside the user gesture, before any `await`** — an `AudioContext` created in a post-fetch microtask starts suspended and cannot resume without another gesture. ```ts theme={null} import { createPCMStreamPlayer } from "@mascotbot/core"; import { useLipsyncStream } from "@mascotbot/react/rive"; const player = createPCMStreamPlayer({ sampleRate: 24000 }); // The tap drives the avatar; the player drives the speakers. useLipsyncStream({ client, playback, source: { kind: "mediaStream", stream: player.outputStream }, }); // Gemini Live (@google/genai): assistant audio is base64 PCM16 session.onmessage = (m) => { const b64 = m?.serverContent?.modelTurn?.parts?.[0]?.inlineData?.data; if (typeof b64 === "string") player.pushBase64PCM16(b64); if (m?.serverContent?.interrupted) player.stop(); }; // OpenAI Realtime (WebSocket): assistant audio is a PCM16 ArrayBuffer session.on("audio", (e) => player.pushPCM16(new Uint8Array(e.data))); session.on("audio_interrupted", () => player.stop()); ``` ## Knowing when playback finished There is no audio element to listen to, so the player surfaces a natural end through the `onIdle` option (with an `isPlaying` getter for polling). `onIdle` fires the moment the queue empties **and** every scheduled buffer has finished — i.e. all pushed audio has actually been heard. * It does **not** fire on `stop()` or `close()` — those are explicit interruptions, not a natural end. * It can fire more than once per player: a push after a drain restarts playback and a later drain fires it again. This is the signal to drive a **queue / sequential** consumer (streamed-TTS playlists, multi-utterance agents): advance to the next item, reset the avatar, or release resources — without estimating audio duration from byte counts. ```ts theme={null} const player = createPCMStreamPlayer({ sampleRate: 24000, onIdle: () => { // All queued audio has played out. If more is still being fetched, // ignore — the next push restarts the player and onIdle fires // again when that drains. if (queue.length > 0 || fetching) return; playback.reset(); // return the avatar to neutral void player.close(); // release the AudioContext (or keep for reuse) }, }); ``` ## When NOT to use it <Warning> Never route a **self-playing** provider through `createPCMStreamPlayer` — you would hear the voice twice (double audio). ElevenLabs Conversational AI and OpenAI Realtime over **WebRTC** play the audio themselves. For those, do not use the player: tap their existing playback with the SDK's cross-browser [`createElementTap()`](/realtime/overview#tap-a-playing-element) and feed that to `useLipsyncStream({ source: { kind: "mediaStream", stream } })`. </Warning> Decision rule: **does the provider play the audio for you?** Yes → tap its playback, no player. No (it hands you raw PCM) → `createPCMStreamPlayer`. ## Server TTS The same primitive powers "server returns audio only, the SDK does the lip sync": your route synthesizes speech and returns base64 PCM16; the client plays it through the player and the tap drives the mouth. No server-side visemes, no SSE protocol. See [Realtime providers](/realtime/overview). ## Next <Columns> <Card title="Realtime providers" icon="bolt" href="/realtime/overview"> The per-provider recipe. </Card> <Card title="Streaming & mic" icon="microphone" href="/libraries/streaming-and-mic"> Feeding the tap to lip sync. </Card> <Card title="Core client" icon="cube" href="/core/client"> The vanilla engine. </Card> </Columns> # Streaming Lip Sync Sessions - createStreamingSession & pushWindow Source: https://docs.mascot.bot/core/streaming Drive lip sync from live audio with the vanilla createStreamingSession API: push 25 ms windows, read per-frame results, handle silence, and barge-in. `createStreamingSession()` is the low-level streaming primitive on `LipsyncClient`. You feed it fixed-size audio windows and it returns one viseme result per window. Most React apps should use [`useLipsyncStream`](/libraries/streaming-and-mic) instead — it owns the audio graph for you. Reach for the raw session when you control the audio source yourself (a custom worklet, a Node pipeline, a non-React app). ## The window contract ```ts theme={null} const session = client.createStreamingSession(); // audioWindow: 16 kHz mono Float32 in [-1, 1], 400 samples (25 ms) const frame = await session.pushWindow(audioWindow); session.close(); ``` | Rule | Detail | | ----------- | --------------------------------------------------------------------- | | Sample rate | 16 kHz mono. Use `client.resample(buf, fromRate, 16000)` upstream. | | Window size | 400 samples — exactly 25 ms. | | Cadence | One window at a time, in order. `pushWindow` is async — `await` each. | | Cleanup | `session.close()` when the utterance/stream ends. | `pushWindow` resolves to a `LipsyncStreamingFrameResult`: ```ts theme={null} interface LipsyncStreamingFrameResult { visemeId: number; // viseme id for this window silenceDetected: boolean; // input was below the silence floor frameIndex: number; // monotonically increasing window counter } ``` ## Drive the avatar Append each result to a streaming `MascotPlayback`. `offset` is the audio-position in ms; the playback's animation-frame clock fires the viseme when its time arrives. ```ts theme={null} import { MascotPlayback, getRiveInputs } from "@mascotbot/core/rive"; const playback = new MascotPlayback({ riveInputs: getRiveInputs(rive), stream: true, enableNaturalLipSync: true }); playback.play(); let ms = 0; for await (const window of windows /* your 25 ms Float32 chunks */) { const frame = await session.pushWindow(window); if (!frame.silenceDetected) playback.pushVisemes([{ offset: ms, visemeId: frame.visemeId }]); ms += 25; } session.close(); ``` ## Silence handling `silenceDetected` reflects the SDK's internal **−50 dBFS input-amplitude silence gate**. It suppresses the phantom mouth shapes a naive pipeline would emit during the inference tail after speech stops — you do not implement your own gate. Treat `silenceDetected: true` windows as "mouth at rest". ## Barge-in and interruption To cut a response short (the user interrupts), stop feeding windows and reset playback: ```ts theme={null} playback.reset(); // clear queued cues; mouth returns to rest // open a fresh session for the next utterance if needed const next = client.createStreamingSession(); ``` If your audio arrives as raw PCM from a network source, pair [`createPCMStreamPlayer`](/core/pcm-stream-player) (which plays it gap-tolerantly and exposes a `MediaStream`) with `useLipsyncStream` rather than hand-windowing — that handles buffering and the tap for you. ## React equivalent `useLipsyncStream({ source: { kind: "manual" } })` wraps a streaming session and exposes `pushAudio` / `pushBase64PCM16` / `reset`, while keeping the audio graph stable across renders. Prefer it in React; use the raw session only when you need full control of the window loop. ## Next <Columns> <Card title="Core client" icon="cube" href="/core/client"> `init`, `processAudio`, events. </Card> <Card title="PCM stream player" icon="waveform" href="/core/pcm-stream-player"> Play + tap raw PCM. </Card> <Card title="Streaming & mic (React)" icon="microphone" href="/libraries/streaming-and-mic"> The React wrapper. </Card> </Columns> # Introduction Source: https://docs.mascot.bot/index Create engaging AI mascots with advanced lip-sync technology. Build interactive experiences for your applications. <div> <div> <img /> <img /> </div> <div> <div> Mascotbot Documentation </div> <p> Build interactive AI avatars — 120 FPS, real-time, voice-ready. The AI Avatar SDK for developers, brands, and voice agents. </p> <div> <HeroCard title="React SDK" description="Provider, hooks, and the Rive avatar layer" href="/libraries/react-sdk" /> <HeroCard title="Quickstart" description="Install, mount the provider, and animate an avatar in minutes" href="/quickstart" /> <HeroCard title="Realtime Providers" description="Connect OpenAI Realtime, Gemini Live, or ElevenLabs to a real-time avatar" href="/realtime/overview" /> <HeroCard title="Ready-made Mascots" description="Browse our collection of pre-built mascots with unique personalities" href="/mascots/ready-to-use-mascots" /> </div> <p> New in 0.2 — one SDK, continuously-updated models, and the offline timeline. Coming from an earlier version? See the <a href="/reference/migration">migration guide</a>, or start with the <a href="/overview">overview</a>. </p> </div> </div> # Install the Mascotbot Lipsync SDK - Private Registry & API Keys Source: https://docs.mascot.bot/installation Install @mascotbot/react and lipsync-core from the private npm registry, configure your .npmrc auth token, and pick the right development or production API key. The SDK is published to `npm.mascot.bot`, a **private registry** — it is not on npmjs.com. You need a Mascotbot API key to install and to run. ## 1. Get an API key Create a key at [app.mascot.bot/api-keys](https://app.mascot.bot/api-keys). Keys are prefixed by environment: | Prefix | Where it works | Metering | | -------------- | -------------------------------------------------------------- | ------------------------------------------------ | | `mascot_dev_…` | `localhost`, `*.localhost`, `127.0.0.1`, private networks only | Development meters — no billing, tamper-tolerant | | `mascot_pub_…` | Your registered public domains | Production meters — tamper detection active | A `mascot_dev_…` key sent from a public origin is **rejected by design**, and a `mascot_pub_…` key from `localhost` is rejected too. Use the matching key for the environment. See [Licensing & keys](/concepts/licensing-and-keys) for the full model. ## 2. Point npm at the private registry Add a `.npmrc` at the root of your project: ```ini theme={null} @mascotbot:registry=https://npm.mascot.bot/ //npm.mascot.bot/:_authToken=mascot_xxx ``` <Warning> Never commit `.npmrc` with the auth token. Add it to `.gitignore` and inject the token from a CI secret / environment variable instead. </Warning> ## 3. Install the package <CodeGroup> ```bash React theme={null} pnpm add @mascotbot/react ``` ```bash Vanilla / Node theme={null} pnpm add @mascotbot/core ``` ```bash Rive avatar (peer deps) theme={null} pnpm add @rive-app/react-webgl2 @rive-app/webgl2 ``` </CodeGroup> `@rive-app/webgl2` (and `@rive-app/react-webgl2` for React) is an **optional peer dependency** of the `/rive` subpaths. Install it only if you render an avatar; the audio pipeline alone does not need it. `npm` and `yarn` work the same way once the registry line is in `.npmrc`. ## 4. Configure the client Pass configuration to `<MascotProvider>` (React) or `LipsyncClient.init` (vanilla): | Option | Type | Default | Notes | | ----------------- | --------- | ---------------------------- | -------------------------------------------------------------------------------------------------------------------------------- | | `apiKey` | `string` | required | `mascot_dev_…` (localhost) or `mascot_pub_…` (production) | | `userId` | `string` | random | Stable per-user id for MAU billing attribution | | `licenseEndpoint` | `string` | `https://license.mascot.bot` | Override only if Mascotbot points you elsewhere | | `devMode` | `boolean` | auto-detect | Forced on for `localhost`, `*.localhost`, `127.0.0.1`, private IPs. Skips the Origin allow-list and routes events to dev meters. | | `fingerprintHash` | `string` | hashed UA + hardware | Override to control session attribution | ```tsx theme={null} "use client"; import { MascotProvider } from "@mascotbot/react"; export default function Layout({ children }: { children: React.ReactNode }) { return <MascotProvider apiKey={process.env.NEXT_PUBLIC_MASCOT_KEY!}>{children}</MascotProvider>; } ``` Production publishable keys (`mascot_pub_…`) are safe in client bundles — they are scoped to your allow-listed origins. Keep standing third-party keys (TTS, OpenAI, ElevenLabs) on the server. ## TypeScript All packages ship `.d.ts`. Public types are re-exported from each root entry point — you do not need to install `lipsync-core` separately just for types if you depend on `lipsync-react`. ## Next <Columns> <Card title="Quickstart" icon="rocket" href="/quickstart"> A working avatar in a few lines. </Card> <Card title="Licensing & keys" icon="key" href="/concepts/licensing-and-keys"> Dev vs production, session lifecycle, error codes. </Card> </Columns> # Mascotbot Avatars API - Ready-Made Mascot .riv Distribution Source: https://docs.mascot.bot/libraries/avatars A small public HTTP API that serves ready-made Mascotbot avatars as Rive .riv files plus per-version metadata. Fetch once, self-host, drive Rive from the metadata. A small public HTTP API that lists ready-made Mascotbot avatars (Rive `.riv` files) and serves their bytes plus per-version metadata. It is **not** an npm registry — just a manifest + binaries. Read access is **public**; no key required. The intended flow: fetch the `.riv` **once**, self-host it from your own CDN/origin, and drive Rive using the per-version metadata. * **Base URL:** `https://license.mascot.bot` ## Concepts * **Avatar** — a mascot identified by an `id` (e.g. `notion-guy`), with one or more **versions**. * **Version** — an immutable `(id, version)` pair where `version` is semver `x.y.z`. The bytes for a published version never change; publishers ship updates as new versions. Old versions are retained. * **`latest`** — the highest semver of an avatar. What `download` serves when no version is requested. * **Metadata** — a freeform JSON object attached to each version, returned verbatim. There is **no fixed schema** — different `.riv` files expose different artboards, state machines, and inputs, so the shape is whatever the publisher set. Mascotbot's own mascots mirror an `artboard` / `stateMachine` / `customization` shape; treat any field as optional. ## List avatars ```http theme={null} GET /v1/avatars ``` Returns the manifest. Response is `Cache-Control: public, max-age=60`; an empty pool is a valid `200` with `"avatars": []`. ```jsonc theme={null} { "schema": 2, "updatedAt": 1747000000000, "avatars": [ { "id": "notion-guy", "name": "Notion Guy", "description": null, "latest": "1.0.0", "metadata": { /* latest version's metadata, or null */ }, "downloadUrl": "https://license.mascot.bot/v1/avatars/notion-guy/download", "versions": [ { "version": "1.0.0", "sha256": "651ab97684ea…", "fileSize": 218080, "createdAt": 1747000000000, "metadata": { /* freeform, publisher-defined */ }, "downloadUrl": "https://license.mascot.bot/v1/avatars/notion-guy/download?version=1.0.0" } ] } ] } ``` ## Download a `.riv` ```http theme={null} GET /v1/avatars/<id>/download # serves `latest` GET /v1/avatars/<id>/download?version=1.0.0 # serves a pinned version ``` * `200` → the raw `.riv` (`application/octet-stream`, `Content-Disposition: attachment`). * `X-Avatar-Version: <served version>` is set on every response. * A **pinned** `?version=` response is `Cache-Control: public, max-age=31536000, immutable` — safe to cache forever. The unversioned `latest` response is `max-age=300` (it can move when a new version ships). * `HEAD` is supported. ## Recommended integration pattern <Steps> <Step title="Fetch the manifest once"> Poll on your own cadence (it changes only when a new version ships). </Step> <Step title="Pick the avatar + version"> Pin a `version` for reproducibility, or follow `latest` for auto-updates. </Step> <Step title="Download the .riv and self-host it"> Store it on your own CDN/origin keyed by `id@version` (or by `sha256`). Pinned versions are immutable — you never need to re-fetch. Do **not** serve `license.mascot.bot` to end users on every page load. </Step> <Step title="Drive Rive from the metadata"> The freeform `metadata` tells your renderer how to wire the file — `artboard`, `stateMachine`, available inputs, customization controls. </Step> </Steps> ```ts theme={null} const API = "https://license.mascot.bot"; const manifest = await fetch(`${API}/v1/avatars`).then(r => r.json()); const avatar = manifest.avatars.find((a: any) => a.id === "notion-guy"); const version = avatar.latest; // or pin "1.0.0" // Download once → store on your own hosting. const riv = await fetch( `${API}/v1/avatars/notion-guy/download?version=${version}`, ).then(r => r.arrayBuffer()); await myCdn.put(`mascots/notion-guy@${version}.riv`, riv); // Wire Rive from metadata (shape is publisher-defined). const md = avatar.metadata ?? {}; new Rive({ src: `https://my-cdn/mascots/notion-guy@${version}.riv`, artboard: md.artboard, stateMachines: md.stateMachine, // md.customization → build your UI controls (type / min / max / default) }); ``` In the React SDK, the same flow ends with passing your self-hosted URL to `<Mascot src={…}>` and reading `md.artboard` / `md.stateMachine` from the manifest — see [Rive co-existence](/concepts/rive-coexistence) for the input-ownership contract. ### Variants — same character, different `.riv` Some characters ship as **multiple avatars** because the standalone and embeddable-widget builds are genuinely different files with different Rive wiring. They are cross-linked in metadata and you pick the one matching how you render it: | `id` | use | artboard | stateMachine | | ------------------- | ------------------ | ----------- | -------------------- | | `notion-guy` | standalone hosting | `Character` | `InLesson` | | `notion-guy-widget` | embeddable widget | `Widget` | `mascotStateMachine` | The metadata on each one carries the full customization schema for that variant; the API doesn't merge them. ## Versioning A published `(id, version)` is **immutable** — to change an avatar, the publisher ships a **new version**, and clients that pinned the old one keep working unchanged. `latest` is the highest semver; clients that follow it get updates automatically, clients that pin a `version` are stable. Pinned download URLs are immutable and long-cached, so they are safe to bake into your own CDN keyed by `id@version` or `sha256`. ## Errors JSON `{ "code": "...", "message": "..." }`: | Status | `code` | Meaning | | ------ | ----------------------- | -------------------------------------------------------- | | 404 | `avatar_not_found` | no such `id` | | 404 | `version_not_found` | `id` exists, that `version` doesn't | | 404 | `avatar_binary_missing` | manifest entry exists but the binary is gone — report it | | 404 | `not_found` | unknown route | | 405 | `method_not_allowed` | non-`GET`/`HEAD` on a read route | | 500 | `manifest_corrupt` | server-side; surface a friendly retry | ## Quick reference | | | | --------------- | --------------------------------------------------------- | | List | `GET /v1/avatars` — public, 60 s cache | | Download latest | `GET /v1/avatars/<id>/download` | | Download pinned | `GET /v1/avatars/<id>/download?version=x.y.z` — immutable | | Base URL | `https://license.mascot.bot` | ## Next <Columns> <Card title="Ready-made mascots" icon="face-smile" href="/mascots/ready-to-use-mascots"> The catalog of characters available through this API. </Card> <Card title="Rive co-existence" icon="puzzle-piece" href="/concepts/rive-coexistence"> The input-ownership contract once a `.riv` is rendered. </Card> <Card title="React SDK intro" icon="react" href="/libraries/react-sdk"> Wiring a downloaded `.riv` into `<Mascot>`. </Card> </Columns> # ElevenLabs Avatar Integration - Add Real-time Visual Avatars to Your Voice AI Source: https://docs.mascot.bot/libraries/elevenlabs-avatar Complete guide to integrating animated avatars with ElevenLabs conversational AI. Real-time lip sync, WebSocket support, and production-ready React components. Start in 5 minutes → # ElevenLabs Avatar Integration - Real-time Visual Avatars for Your Voice AI Transform your ElevenLabs voice agents into engaging visual experiences with Mascot Bot SDK. Get perfect lip sync, seamless integration, and production-ready React components that work alongside the **official ElevenLabs SDK** — unchanged. The SDK works alongside your ElevenLabs setup and animates a real-time avatar from the audio. <video /> <Columns> <Card title="Quick Start" icon="rocket" href="#quick-start"> Add avatars in 5 minutes </Card> <Card title="Live Demo" icon="play" href="https://mascot.bot/elevenlabs-demo"> See voice avatars in action </Card> <Card title="GitHub Repo" icon="github" href="https://github.com/mascotbot-templates/elevenlabs-avatar"> Complete example code </Card> <Card title="Features" icon="sparkles" href="#features"> Real-time lip sync & more </Card> <Card title="API Reference" icon="code" href="#api-reference"> Complete hook documentation </Card> <Card title="Deploy" icon="triangle" href="https://vercel.com/new/clone?repository-url=https%3A%2F%2Fgithub.com%2Fmascotbot-templates%2Felevenlabs-avatar"> One-click deployment </Card> </Columns> ## Why Add Avatars to Your ElevenLabs Conversational AI? Voice-only AI can feel impersonal. By adding a **conversational AI avatar** with real-time lip sync, you create more engaging, human-like interactions. The Mascot Bot **voice to avatar SDK** works alongside your existing ElevenLabs setup — you keep using `@elevenlabs/client` exactly as you do today, and the SDK lip-syncs whatever the agent says in real time. ## Features ### <Icon icon="bullseye" />  Real-time Lip Sync The SDK turns ElevenLabs' audio output into a real-time talking avatar with no perceptible lag. There is no server round-trip for visemes — only ElevenLabs' own audio stream. ### <Icon icon="bolt" />  120fps Animation Performance Smooth, natural **voice-driven facial animation** powered by WebGL2 and the Rive runtime. ### <Icon icon="plug" />  Native ElevenLabs Support Works alongside `@elevenlabs/client` with zero conflicts and zero modifications to your ElevenLabs code. The SDK never proxies or intercepts the ElevenLabs connection — it only taps the audio it already plays. ### <Icon icon="palette" />  Customizable Avatars Choose from [ready-made mascots](/mascots/ready-to-use-mascots) or bring your own Rive file. The SDK only writes the mouth, `is_speaking`, and `stress` — every other input, outfit, gesture, and ViewModel stays yours ([Rive co-existence](/concepts/rive-coexistence)). ### <Icon icon="rotate" />  Streaming Avatar Audio ElevenLabs plays the assistant's audio itself; the SDK captures that exact playback as a `MediaStream` and lip-syncs it. The capture point is the playback point, so the mouth never drifts ahead of speech. ### <Icon icon="masks-theater" />  Natural Lip Sync Processing An optional post-processor merges rapid visemes and preserves the distinctive shapes for natural, non-robotic motion — [Natural lip sync](/libraries/natural-lip-sync). ## Quick Start ### Installation The SDK installs from the private registry `npm.mascot.bot`. Add an `.npmrc`, then install alongside the official ElevenLabs client: ```ini .npmrc theme={null} @mascotbot:registry=https://npm.mascot.bot/ //npm.mascot.bot/:_authToken=mascot_xxx ``` ```bash theme={null} pnpm add @mascotbot/react @rive-app/react-webgl2 @rive-app/webgl2 @elevenlabs/client ``` <Note> Get your Mascot Bot key at [app.mascot.bot/api-keys](https://app.mascot.bot/api-keys) (`mascot_dev_…` for localhost, `mascot_pub_…` for production). The SDK works alongside the official ElevenLabs SDK without any modifications. Full setup: [Installation](/installation). </Note> <Info> **Want a complete working example?** See the [open-source demo repository](https://github.com/mascotbot-templates/elevenlabs-avatar), or deploy it to Vercel with one click. </Info> ### Basic Integration Three pieces: a server route that mints an ElevenLabs signed URL, the ElevenLabs `Conversation`, and the SDK tapping its audio. ```tsx theme={null} "use client"; import { useState } from "react"; import { MascotProvider, useMascot } from "@mascotbot/react"; import { Mascot, MascotRive, useMascotPlayback, useLipsyncStream } from "@mascotbot/react/rive"; function App() { return ( <MascotProvider apiKey="mascot_pub_…"> <MascotProvider> <Mascot src="/mascot.riv"> <MascotRive /> <ElevenLabsAvatar /> </Mascot> </MascotProvider> </MascotProvider> ); } ``` The avatar is wired in [Step 2](#step-2-create-your-avatar-component). Your **ElevenLabs voice with avatar** is then ready — the SDK handles synchronization automatically. ## Complete Implementation Guide ### Step 1: Mint an ElevenLabs Signed URL (Server-Side) ElevenLabs needs a signed URL for the WebSocket. Mint it on the server so the standing `xi-api-key` never reaches the browser. This is the standard ElevenLabs signed-URL endpoint. ```typescript theme={null} // app/api/elevenlabs/signed-url/route.ts export const runtime = "nodejs"; export async function POST() { const key = process.env.ELEVENLABS_API_KEY; const agentId = process.env.ELEVENLABS_AGENT_ID; if (!key || !agentId) { return Response.json({ error: "ElevenLabs env not set" }, { status: 400 }); } const url = new URL("https://api.elevenlabs.io/v1/convai/conversation/get-signed-url"); url.searchParams.set("agent_id", agentId); const res = await fetch(url, { headers: { "xi-api-key": key }, cache: "no-store" }); if (!res.ok) return Response.json({ error: `ElevenLabs ${res.status}` }, { status: 502 }); const json = (await res.json()) as { signed_url?: string }; return Response.json({ signedUrl: json.signed_url }); } ``` <Info> Required environment variables (server-side only): * `ELEVENLABS_API_KEY` — your ElevenLabs API key * `ELEVENLABS_AGENT_ID` — your ElevenLabs Conversational AI agent id Your Mascot Bot key (`mascot_pub_…`) is a separate, browser-safe publishable key passed to `<MascotProvider>`. </Info> ### Step 2: Create Your Avatar Component ElevenLabs plays the assistant audio internally through an `<audio>` element. Capture that element, expose it as a `MediaStream`, and feed it to `useLipsyncStream`. Leave playback with ElevenLabs. ```tsx theme={null} "use client"; import { useEffect, useRef, useState } from "react"; import { useMascot } from "@mascotbot/react"; import { useMascotPlayback, useLipsyncStream } from "@mascotbot/react/rive"; // Stable module constant — see the natural-lip-sync warning below. const LIP_SYNC = { minVisemeInterval: 60, mergeWindow: 80 } as const; export function ElevenLabsAvatar() { const { client, status } = useMascot(); const playback = useMascotPlayback({ stream: true, enableNaturalLipSync: true, naturalLipSyncConfig: LIP_SYNC }); const [stream, setStream] = useState<MediaStream | null>(null); const teardownRef = useRef<null | (() => void)>(null); // The SDK lip-syncs whatever this MediaStream carries. const { error, attached } = useLipsyncStream({ client, playback, source: { kind: "mediaStream", stream }, }); useEffect(() => () => teardownRef.current?.(), []); const start = async () => { if (status !== "ready") return; // createElementTap from "@mascotbot/react" — create in // this click gesture so its AudioContext isn't suspended (Safari). const tap = createElementTap(); setStream(tap.stream); // Capture the hidden <audio> ElevenLabs creates (its srcObject is a // MediaStream of the worklet output we'll tap). const w = window as unknown as { Audio: typeof Audio; __el?: HTMLAudioElement | null }; const OrigAudio = w.Audio; w.Audio = function (...args: unknown[]) { const el = new OrigAudio(...(args as [])); w.__el = el; return el; } as unknown as typeof Audio; const { signedUrl } = await (await fetch("/api/elevenlabs/signed-url", { method: "POST" })).json(); const { Conversation } = await import("@elevenlabs/client"); const convo = await Conversation.startSession({ signedUrl }); // Live-tracks guard: on a 2nd "end → start" cycle, `w.__el` will // briefly still point at the previous call's <audio> element (its // srcObject is a MediaStream whose tracks are 'ended'). Attaching // to it produces silence for the entire new call. Require at // least one 'live' track before attaching. const isLive = (el: HTMLMediaElement | null | undefined) => !!el && el.srcObject instanceof MediaStream && el.srcObject.getAudioTracks().some((t) => t.readyState === "live"); let tries = 0; const iv = window.setInterval(() => { const el = w.__el; // tap.attach(el): cross-browser tap (Safari has no captureStream). // See /realtime/overview#tap-a-playing-element if (isLive(el)) { tap.attach(el as HTMLMediaElement); window.clearInterval(iv); } else if (++tries > 100) { window.clearInterval(iv); } }, 100); teardownRef.current = () => { window.clearInterval(iv); w.Audio = OrigAudio; // Null the stash so the *next* call's poll doesn't latch onto // this (now-dead) element before EL has constructed its new one. w.__el = null; // Close the tap's AudioContext + MediaStream — otherwise each // restart leaks a worklet graph. tap.close(); void convo.endSession(); }; }; return ( <div> <button onClick={start} disabled={status !== "ready"}>Start conversation</button> <span>{stream ? (attached ? "lip-sync attached" : "attaching…") : "idle"}</span> {error ? <p>{error.message}</p> : null} </div> ); } ``` <Warning> Do **not** route ElevenLabs through `createPCMStreamPlayer` — ElevenLabs plays the audio itself, so the player would make the voice play twice. Tap the audio it already renders, as above. </Warning> ### Step 3: Advanced Features #### Natural Lip Sync Configuration Tune the post-processor by passing a **stable** `naturalLipSyncConfig` to `useMascotPlayback`: ```tsx theme={null} // Module scope — a NEW object every render reinitializes playback and // breaks lip sync after the first chunk. const CONVERSATION = { minVisemeInterval: 60, mergeWindow: 80, keyVisemePreference: 0.7, preserveSilence: true, similarityThreshold: 0.6, preserveCriticalVisemes: true, } as const; const playback = useMascotPlayback({ enableNaturalLipSync: true, naturalLipSyncConfig: CONVERSATION }); ``` See [Natural lip sync](/libraries/natural-lip-sync) for every field, the defaults, and conversation / fast-speech / educational presets. #### Embedded Avatar Widget Mount the avatar small and fixed for an embeddable **AI agent with face**. The SDK only animates the mouth — your own widget chrome, click handlers, and Rive inputs are untouched: ```tsx theme={null} <div className="fixed bottom-4 right-4 w-64 h-64"> <Mascot src="/widget-mascot.riv"> <MascotRive /> <ElevenLabsAvatar /> </Mascot> </div> ``` Need a non-mouth input (a wave, a reveal)? Use `useMascotInputs().has(name)` then drive it yourself — [Rive co-existence](/concepts/rive-coexistence). #### Gestures on Every Agent Turn The legacy SDK auto-fired a `gesture` trigger at the start of every agent utterance (the old `gesture: true` flag on `useMascotElevenlabs`). 0.2.x removed the auto-fire — consumers wire it themselves. ElevenLabs makes this a one-liner: `Conversation.startSession` exposes an `onModeChange` callback that flips to `"speaking"` the moment the first audio chunk of a new turn lands. ```tsx theme={null} import { useRef } from "react"; import { useMascotInputs } from "@mascotbot/react/rive"; function ElevenLabsAvatarWithGestures() { // useMascotInputs() returns a fresh object every render — capture in a // ref so the long-lived EL callback always reads the current handle. const { custom } = useMascotInputs(); const customRef = useRef(custom); customRef.current = custom; const start = async () => { const { signedUrl } = await (await fetch("/api/elevenlabs/signed-url", { method: "POST" })).json(); const { Conversation } = await import("@elevenlabs/client"); await Conversation.startSession({ signedUrl, onModeChange: ({ mode }) => { // Fires exactly once per agent turn start; stays silent on the // "listening" back-edge so the gesture only plays as the agent // begins speaking. if (mode !== "speaking") return; customRef.current?.gesture?.fire?.(); }, // ... your existing onConnect / onDisconnect / onError }); }; return <button onClick={start}>Start conversation</button>; } ``` Declare `gesture` on the parent `<Mascot inputs={["gesture", ...]}>` so the SDK exposes a real trigger handle. On `.riv` files without a `gesture` input the SDK returns a no-op shim, so the optional-chain `fire?.()` stays safe. For a provider-agnostic approach driven by the speech envelope (works identically for OpenAI / Gemini), see [Stress emphasis and gestures](/realtime/overview#stress-emphasis-and-gestures). ### Step 4: Dynamic Variables ElevenLabs **dynamic variables** personalize conversations at runtime. They are an ElevenLabs feature and are completely independent of the SDK — pass them straight to `Conversation.startSession`: ```tsx theme={null} const dynamicVariables = { name: userName ?? "Guest", role: userRole ?? "user" }; const { signedUrl } = await (await fetch("/api/elevenlabs/signed-url", { method: "POST" })).json(); const { Conversation } = await import("@elevenlabs/client"); const convo = await Conversation.startSession({ signedUrl, dynamicVariables }); ``` Configure your agent prompt to use `{{name}}` / `{{role}}` placeholders. The SDK does not see or touch these — it only lip-syncs the resulting audio. ## API Reference This integration uses the standard SDK surface plus the official ElevenLabs client. | Surface | Role | | --------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- | | `<MascotProvider apiKey>` | Initializes the licensed avatar client. [Config](/installation#4-configure-the-client). | | `<MascotProvider>` / `<Mascot src>` / `<MascotRive>` | Load and render the Rive avatar. | | `useMascotPlayback({ stream: true, enableNaturalLipSync, naturalLipSyncConfig })` | The mouth playback engine. | | `useLipsyncStream({ client, playback, source: { kind: "mediaStream", stream } })` | Lip-syncs the tapped ElevenLabs audio. [Reference](/libraries/streaming-and-mic). | | `@elevenlabs/client` `Conversation` | The official ElevenLabs SDK — unchanged. | Server route: the plain ElevenLabs `GET /v1/convai/conversation/get-signed-url` with your `xi-api-key`. No Mascot Bot endpoint sits in the path. ## Use Cases ### AI Customer Service Avatar A visible **virtual assistant with face** for support — visual feedback during voice conversations, on-brand appearance, expressions driven by your own Rive inputs. ### Educational AI Tutor Avatar Clear articulation for learning; pair with the `educational` natural-lip-sync preset for crisper mouth shapes. ### Voice AI Virtual Receptionist A welcoming visual presence with natural conversation flow and a brand-customizable mascot. ## Technical Details ### Voice-to-Animation Pipeline 1. ElevenLabs streams and **plays** the assistant's audio (its own WebSocket, untouched). 2. The SDK captures that playback as a `MediaStream`. 3. The SDK infers a viseme per 10 ms frame from the audio. 4. The Rive runtime renders the mouth at up to 120fps. 5. Optional natural lip sync smooths the motion. No audio or viseme data is sent to or stored on Mascotbot servers. ### Performance * Low audio-to-visual delay (the capture point is the playback point). * WebGL2-accelerated rendering. * End-of-utterance phantom mouth shapes are suppressed by the SDK's internal silence gate — you do not implement one. ## Troubleshooting ### Avatar Not Moving? <Warning> Confirm `status === "ready"` from `useMascot()`, that the tapped `MediaStream` is non-null (the `window.Audio` patch must be installed **before** `Conversation.startSession`, and `el.srcObject` must be a `MediaStream`), and that the Rive file uses artboard `Character` + state machine `mascotStateMachine` with inputs `100`–`118`. </Warning> ### Only the First Second of Speech Animates? A new `naturalLipSyncConfig` object on every render reinitializes playback. Use a stable module constant or `useState`/`useMemo` — see the example above and [Troubleshooting](/libraries/react-troubleshooting). ### Hearing the Voice Twice? You routed ElevenLabs through `createPCMStreamPlayer`. ElevenLabs self-plays — tap its audio instead (Step 2), never the PCM player. ### Dynamic Variables Not Applied? They are an ElevenLabs concern. Ensure the agent prompt has the `{{placeholders}}` and that you pass `dynamicVariables` to `Conversation.startSession`. The SDK is not involved. ## FAQ ### Can You Add an Avatar to ElevenLabs? Yes. The SDK works alongside the official `@elevenlabs/client` with no modifications. You connect ElevenLabs as usual; the SDK lip-syncs its audio in real time. ### Does It Work With My Existing ElevenLabs Setup? Yes. Keep your `@elevenlabs/client` code exactly as it is — the SDK lip-syncs the audio ElevenLabs plays, in real time. ### Do I Modify My ElevenLabs Code? No. Keep your `Conversation` setup. You only add a `MediaStream` tap of the audio it plays. ### How Is the Lip Sync Synchronized? The audio is tapped at its playback point with a Web-Audio `MediaStreamDestination`, so visemes are derived from exactly what the user hears — the mouth cannot run ahead of the voice. ### Is Audio Sent to Mascot Bot? Your users' speech is processed by the SDK in their browser and isn't sent to or stored on Mascotbot servers. ### What Is the Voice Avatar SDK? A React/JavaScript library that adds a real-time, lip-synced avatar to any voice AI — including ElevenLabs Conversational AI. ## Start with ElevenLabs Avatar Today Ready to transform your voice AI? The **open-source avatar for ElevenLabs** makes it simple: <Columns> <Card title="Try Voice Avatar Demo" icon="play" href="https://mascot.bot/elevenlabs-demo"> Experience it yourself </Card> <Card title="Demo Repository" icon="github" href="https://github.com/mascotbot-templates/elevenlabs-avatar"> Complete working example </Card> </Columns> Unlike pre-rendered solutions, this is a **real-time alternative** — dynamic, responsive avatars that connect with users. ## Next Steps 1. Get a key at [app.mascot.bot/api-keys](https://app.mascot.bot/api-keys) and install from the [private registry](/installation). 2. Add `<MascotProvider>` + `<MascotProvider>`/`<Mascot>` and the avatar component above. 3. Choose a [ready-made mascot](/mascots/ready-to-use-mascots) or your own Rive file. 4. Tune motion with [natural lip sync](/libraries/natural-lip-sync); review the [realtime overview](/realtime/overview) for the general pattern. Transform your ElevenLabs implementation today with the most **developer-friendly avatar SDK for voice AI**. # Gemini Live API Avatar Integration - Build Interactive AI Avatars with Lip Sync Source: https://docs.mascot.bot/libraries/gemini-live-api-avatar Complete guide to building interactive AI avatars with Gemini Live API and Mascot Bot SDK. Real-time lip sync, ephemeral tokens, WebSocket streaming, and production-ready React components. TypeScript/JavaScript tutorial → # Gemini Live API Avatar — Build an Interactive AI Avatar with Real-time Lip Sync Add a **lip-synced animated avatar** to your Gemini Live API application in minutes. Mascot Bot SDK works alongside the official Google AI SDK (`@google/genai`) — your existing Gemini code stays untouched. The SDK plays Gemini's audio output and animates a real-time avatar from it. <img alt="Gemini Live API Avatar with webcam video — interactive AI mascot with real-time lip sync" /> <Columns> <Card title="Quick Start" icon="rocket" href="#quick-start"> Add avatars in 5 minutes </Card> <Card title="Live Demo" icon="play" href="https://mascot.bot/gemini-liveapi"> See Gemini Live avatar in action </Card> <Card title="GitHub Repo" icon="github" href="https://github.com/mascotbot-templates/gemini-live-api-avatar"> Complete example code </Card> <Card title="Features" icon="sparkles" href="#features"> Real-time lip sync & more </Card> <Card title="API Reference" icon="code" href="#api-reference"> Complete hook documentation </Card> <Card title="Deploy" icon="triangle" href="https://vercel.com/new/clone?repository-url=https%3A%2F%2Fgithub.com%2Fmascotbot-templates%2Fgemini-live-api-avatar"> One-click deployment </Card> </Columns> ## Why Add an Avatar to Your Gemini Live API App? Voice-only Gemini feels disembodied. A **lip-synced avatar** makes the assistant feel present. The Mascot Bot SDK adds that without changing how you use Gemini Live: you keep `@google/genai`, and the SDK lip-syncs Gemini's audio output in real time. ### How It Works: Real-Time Lip Sync Gemini Live streams the assistant's voice as raw base64 PCM16 (it does **not** play the audio for you). The pattern: 1. Your server mints a short-lived **ephemeral token** so the standing Gemini key never reaches the browser. 2. The browser connects to Gemini Live with `@google/genai` using that token. 3. [`createPCMStreamPlayer`](/core/pcm-stream-player) plays Gemini's PCM gap-tolerantly and exposes it as a `MediaStream`. 4. `useLipsyncStream` taps that stream; the SDK infers visemes and drives the avatar. No Mascot Bot endpoint sits in the audio path. See [Realtime overview](/realtime/overview) for the provider-agnostic version. ## Features ### <Icon icon="bullseye" />  Real-time Lip Sync for Gemini Live API Real-time viseme inference from Gemini's audio output — no server round-trip for visemes, no perceptible lag. ### <Icon icon="bolt" />  120fps Avatar Animation WebGL2 + Rive runtime for smooth, natural facial motion. ### <Icon icon="plug" />  Native Google AI SDK Compatibility Use `@google/genai` exactly as documented by Google. The SDK never proxies or wraps the Gemini connection — it only plays and taps the audio. ### <Icon icon="shield-halved" />  Ephemeral Token Security Mint single-use ephemeral tokens server-side with `ai.authTokens.create(...)`. The standing Gemini API key never reaches the client. ### <Icon icon="rotate" />  Streaming Avatar Audio `createPCMStreamPlayer` plays Gemini's PCM gap-tolerantly and exposes a parallel `MediaStream` tap so the avatar stays locked to what is heard. ### <Icon icon="masks-theater" />  Natural Lip Sync Processing Optional viseme post-processing for natural, non-robotic motion — [Natural lip sync](/libraries/natural-lip-sync). ### <Icon icon="video" />  Webcam Video Streaming Gemini Live can accept webcam frames (`session.sendRealtimeInput({ video })`). That is a Gemini capability you use directly through `@google/genai` — it is independent of lip sync and the SDK does not gate it. ### <Icon icon="clock" />  Session Management Gemini Live sessions are time-limited. Re-mint a token and reconnect when a session ends; the SDK's `player.stop()` handles barge-in/interruption. ## Quick Start ### Installation ```ini .npmrc theme={null} @mascotbot:registry=https://npm.mascot.bot/ //npm.mascot.bot/:_authToken=mascot_xxx ``` ```bash theme={null} pnpm add @mascotbot/react @rive-app/react-webgl2 @rive-app/webgl2 @google/genai ``` <Note> Get your Mascot Bot key at [app.mascot.bot/api-keys](https://app.mascot.bot/api-keys). Full registry/key setup: [Installation](/installation). `@google/genai` is Google's official SDK, used unchanged. </Note> ### Basic Integration ```tsx theme={null} "use client"; import { MascotProvider } from "@mascotbot/react"; import { Mascot, MascotRive } from "@mascotbot/react/rive"; export default function App() { return ( <MascotProvider apiKey="mascot_pub_…"> <MascotProvider> <Mascot src="/mascot.riv"> <MascotRive /> <GeminiAvatar /> </Mascot> </MascotProvider> </MascotProvider> ); } ``` `GeminiAvatar` is built in [Step 2](#step-2-create-your-avatar-component). ## Complete Implementation Guide ### Step 1: Set Up Ephemeral Token Generation (Server-Side) Mint a single-use ephemeral token so the standing `GEMINI_API_KEY` stays on the server. ```typescript theme={null} // app/api/gemini/token/route.ts export const runtime = "nodejs"; export async function POST() { const key = process.env.GEMINI_API_KEY; if (!key) return Response.json({ error: "GEMINI_API_KEY not set" }, { status: 400 }); const model = "models/gemini-3.1-flash-live-preview"; const { GoogleGenAI, Modality } = await import("@google/genai"); const ai = new GoogleGenAI({ apiKey: key, httpOptions: { apiVersion: "v1alpha" } }); const token = await ai.authTokens.create({ config: { uses: 1, newSessionExpireTime: new Date(Date.now() + 10 * 60 * 1000).toISOString(), liveConnectConstraints: { model, config: { responseModalities: [Modality.AUDIO] }, }, }, }); return Response.json({ ephemeralToken: token.name, model }); } ``` ### Step 2: Create Your Avatar Component Gemini Live does not play audio — `createPCMStreamPlayer` plays it and exposes the tap. The microphone is sent to Gemini via `session.sendRealtimeInput`. ```tsx theme={null} "use client"; import { useEffect, useRef, useState } from "react"; import { useMascot, createPCMStreamPlayer, type PCMStreamPlayer } from "@mascotbot/react"; import { useMascotPlayback, useLipsyncStream } from "@mascotbot/react/rive"; const LIP_SYNC = { minVisemeInterval: 60, mergeWindow: 80 } as const; export function GeminiAvatar() { const { client, status } = useMascot(); const playback = useMascotPlayback({ stream: true, enableNaturalLipSync: true, naturalLipSyncConfig: LIP_SYNC }); const playerRef = useRef<PCMStreamPlayer | null>(null); const [stream, setStream] = useState<MediaStream | null>(null); const teardownRef = useRef<null | (() => void)>(null); const { error, attached } = useLipsyncStream({ client, playback, source: { kind: "mediaStream", stream } }); useEffect(() => () => { teardownRef.current?.(); void playerRef.current?.close(); }, []); const connect = async () => { if (status !== "ready") return; // Create the player inside the click, before any await. const player = createPCMStreamPlayer({ sampleRate: 24000 }); playerRef.current = player; setStream(player.outputStream); const { ephemeralToken, model } = await (await fetch("/api/gemini/token", { method: "POST" })).json(); const { GoogleGenAI, Modality } = await import("@google/genai"); const ai = new GoogleGenAI({ apiKey: ephemeralToken, httpOptions: { apiVersion: "v1alpha" } }); // Liveness flag: the mic processor fires continuously — never send to a closed socket. let live = true; const session = await ai.live.connect({ model, // "models/gemini-3.1-flash-live-preview" config: { responseModalities: [Modality.AUDIO] }, callbacks: { onmessage: (msg: any) => { const b64 = msg?.serverContent?.modelTurn?.parts?.[0]?.inlineData?.data; if (typeof b64 === "string") player.pushBase64PCM16(b64); if (msg?.serverContent?.interrupted) player.stop(); }, onerror: () => { live = false; }, onclose: () => { live = false; }, }, }); // Mic: 16 kHz mono → PCM16 → sendRealtimeInput const mic = await navigator.mediaDevices.getUserMedia({ audio: { channelCount: 1, sampleRate: 16000 } }); const Ctor = (window as any).AudioContext || (window as any).webkitAudioContext; const ctx = new Ctor({ sampleRate: 16000 }); const src = ctx.createMediaStreamSource(mic); const proc = ctx.createScriptProcessor(4096, 1, 1); proc.onaudioprocess = (ev: AudioProcessingEvent) => { if (!live) return; const f32 = ev.inputBuffer.getChannelData(0); const pcm = new Int16Array(f32.length); for (let i = 0; i < f32.length; i++) { const s = Math.max(-1, Math.min(1, f32[i])); pcm[i] = s < 0 ? s * 0x8000 : s * 0x7fff; } let bin = ""; const bytes = new Uint8Array(pcm.buffer); for (let i = 0; i < bytes.length; i++) bin += String.fromCharCode(bytes[i]); try { session.sendRealtimeInput({ audio: { data: btoa(bin), mimeType: "audio/pcm;rate=16000" } }); } catch { live = false; } }; src.connect(proc); proc.connect(ctx.destination); session.sendClientContent({ turns: "Say a short friendly hello.", turnComplete: true }); teardownRef.current = () => { live = false; proc.onaudioprocess = null; proc.disconnect(); src.disconnect(); mic.getTracks().forEach((t) => t.stop()); void ctx.close(); session.close(); }; }; return ( <div> <button onClick={connect} disabled={status !== "ready"}>Connect</button> <span>{stream ? (attached ? "lip-sync attached" : "attaching…") : "idle"}</span> {error ? <p>{error.message}</p> : null} </div> ); } ``` ### Step 3: Advanced Features * **Natural lip sync** — pass a stable `naturalLipSyncConfig`; full reference and presets in [Natural lip sync](/libraries/natural-lip-sync). * **Barge-in** — `player.stop()` on `serverContent.interrupted` (shown above) drops queued audio instantly. * **Webcam video** — send frames to Gemini via `session.sendRealtimeInput({ video: … })`. This is a Gemini Live feature, used directly through `@google/genai`; it does not involve the lip sync SDK. * **Custom Rive inputs** — the SDK only writes the mouth. Drive gestures/outfits yourself; detect them with `useMascotInputs().has(name)` ([Rive co-existence](/concepts/rive-coexistence)). ## API Reference The integration uses the standard SDK surface plus Google's official SDK: | Surface | Role | | --------------------------------------------------------------- | -------------------------------------------------------------------------------- | | `<MascotProvider apiKey>` | Licensed avatar client. [Config](/installation#4-configure-the-client). | | `<MascotProvider>` / `<Mascot src>` / `<MascotRive>` | Load and render the avatar. | | `useMascotPlayback({ stream: true, enableNaturalLipSync })` | Mouth playback engine. | | `createPCMStreamPlayer({ sampleRate: 24000 })` | Plays Gemini PCM + exposes the tap. [Reference](/core/pcm-stream-player). | | `useLipsyncStream({ source: { kind: "mediaStream", stream } })` | Lip-syncs the tapped audio. [Reference](/libraries/streaming-and-mic). | | `@google/genai` `ai.live.connect` / `ai.authTokens.create` | Google's official SDK — unchanged. Model `models/gemini-3.1-flash-live-preview`. | ## Gemini Live API Pricing & Free Tier Gemini Live API usage is billed by Google per their pricing; the ephemeral-token model adds no Mascot Bot cost. Mascot Bot meters by your plan's speech/MAU allowance — replaying a persisted [timeline](/libraries/offline-lipsync) does not re-meter. Check current Gemini pricing in the Google AI documentation. ## Use Cases ### AI Customer Service Avatar A visible assistant for support — visual presence during Gemini voice conversations. ### Educational AI Tutor Pair with the `educational` natural-lip-sync preset for crisp articulation in language/learning apps. ### Voice AI Virtual Receptionist A branded, welcoming front desk powered by Gemini Live. ### AI Mascot for Streaming & Content A reactive on-screen character; drive non-mouth animation yourself via raw Rive inputs. ## Troubleshooting ### Avatar Not Moving? Confirm `status === "ready"`, that `player.outputStream` is set as the `mediaStream` source, and that the Rive file uses artboard `Character` + state machine `mascotStateMachine` with inputs `100`–`118`. ### Only First Second of Speech Animated? A non-stable `naturalLipSyncConfig` reinitializes playback. Use a module constant — [Troubleshooting](/libraries/react-troubleshooting). ### Connection Fails on Second Call? Ephemeral tokens are single-use (`uses: 1`). Mint a fresh token per session/reconnect. ### No Audio Playing? `createPCMStreamPlayer` must be created inside the user-gesture click before any `await`, or its `AudioContext` starts suspended. Also confirm you are calling `player.pushBase64PCM16` on `modelTurn` audio parts. ### Session Disconnects After \~10 Minutes? Gemini Live sessions are time-limited. Detect `onclose`, mint a new token, and reconnect. ## FAQ ### How Does Mascot Bot Work with the Google AI SDK? It runs alongside it. You use `@google/genai` as documented; the SDK plays Gemini's PCM and lip-syncs it in real time. ### Does It Work With My Existing Gemini Code? Yes. Use `@google/genai` exactly as documented; the SDK turns Gemini's audio output into a real-time avatar. ### Do I Modify My Gemini Code? No. Add a PCM player + a `useLipsyncStream` tap; the Gemini connection is unchanged. ### Can I Use My Own Ephemeral Token Setup? Yes. Any server route returning a valid `ai.authTokens.create` token name works. ### What Gemini Models Support the Live API? Use a Live-API model such as `models/gemini-3.1-flash-live-preview` with `apiVersion: "v1alpha"`. ### Is Audio Sent to Mascot Bot? Your users' speech is processed by the SDK in their browser and isn't sent to or stored on Mascotbot servers. ### Is This an Open-Source Alternative to Pre-rendered Interactive Avatars? Yes — a real-time alternative to server-rendered talking-head services. ## Start Building with Gemini Live API Avatar <Columns> <Card title="Live Demo" icon="play" href="https://mascot.bot/gemini-liveapi"> See it in action </Card> <Card title="Demo Repository" icon="github" href="https://github.com/mascotbot-templates/gemini-live-api-avatar"> Complete working example </Card> </Columns> ## Next Steps 1. Get a key at [app.mascot.bot/api-keys](https://app.mascot.bot/api-keys) and install from the [private registry](/installation). 2. Add the server token route and the avatar component above. 3. Tune motion with [natural lip sync](/libraries/natural-lip-sync). 4. Review the [realtime overview](/realtime/overview) and [PCM stream player](/core/pcm-stream-player) for the underlying pattern. # Natural Lip Sync Source: https://docs.mascot.bot/libraries/natural-lip-sync Create more realistic mouth movements with intelligent viseme processing # Natural Lip Sync Natural lip sync post-processes the viseme stream to produce more natural speech animation. Instead of snapping to every phoneme, it merges similar adjacent shapes and preserves the distinctive ones — the way a real mouth blends sounds during fast speech. ## Why Hitting every phoneme precisely looks robotic. In natural speech the mouth does not have time to fully form each shape; shapes blend, sometimes to the point of not changing at all. The processor mimics that blending while protecting the visually critical shapes (w/u, o, r, l, f/v, and bilabials p/b/m) so articulation still reads. ## Enable it It is one option on `useMascotPlayback`. This works for every path — offline, mic, and realtime: ```tsx theme={null} import { useMascotPlayback } from "@mascotbot/react/rive"; const playback = useMascotPlayback({ enableNaturalLipSync: true }); ``` That uses `DEFAULT_NATURAL_LIPSYNC_CONFIG`. To tune it, pass a `naturalLipSyncConfig`: ```tsx theme={null} const playback = useMascotPlayback({ enableNaturalLipSync: true, naturalLipSyncConfig: CONVERSATION, // a STABLE reference — see warning below }); ``` <Warning> `naturalLipSyncConfig` must be a **stable reference** — a module constant, or memoized with `useState` / `useMemo`. A new object literal on every render reinitializes playback and lip sync breaks after the first audio chunk. This is the single most common integration bug; see [Troubleshooting](/libraries/react-troubleshooting). </Warning> ## Configuration `NaturalLipSyncConfig` (all fields optional — unspecified fields fall back to `DEFAULT_NATURAL_LIPSYNC_CONFIG`): | Field | Type | Default | Meaning | | ------------------------------- | ------------------- | ----------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | | `minVisemeInterval` | `number` (ms) | `60` | Minimum time between visemes; closer ones merge (\~16 visemes/s max). | | `mergeWindow` | `number` (ms) | `80` | Look-ahead window for finding similar visemes to merge. | | `keyVisemePreference` | `number` 0–1 | `0.7` | Strength of preference for distinctive shapes. Higher keeps more. | | `preserveSilence` | `boolean` | `true` | Keep all silence visemes (recommended). | | `similarityThreshold` | `number` 0–1 | `0.6` | How similar two visemes must be to merge. Higher merges less. | | `preserveCriticalVisemes` | `boolean` | `true` | Never skip critical shapes (u/o/l/v/p/b/m). | | `criticalVisemeMinDuration` | `number` (ms) | `0` | Hold critical visemes at least this long (opt-in; `0` disables). | | `criticalVisemeAbsorbThreshold` | `number` (ms) | `30` | If holding a critical viseme shrinks the next one below this, drop the successor instead of flashing it. Only active when `criticalVisemeMinDuration > 0`. | | `criticalVisemeIds` | `readonly number[]` | `DEFAULT_CRITICAL_VISEME_IDS` | Which viseme ids are "critical". | `DEFAULT_CRITICAL_VISEME_IDS` is `[7, 8, 13, 14, 18, 21]` — w/u, o, r, l, f/v, and bilabials. Both `DEFAULT_NATURAL_LIPSYNC_CONFIG` and `DEFAULT_CRITICAL_VISEME_IDS` are exported so you can derive from them: ```ts theme={null} import { DEFAULT_NATURAL_LIPSYNC_CONFIG, DEFAULT_CRITICAL_VISEME_IDS } from "@mascotbot/react/rive"; // Hold s/z too, otherwise defaults export const WITH_SIBILANTS = { ...DEFAULT_NATURAL_LIPSYNC_CONFIG, criticalVisemeIds: [...DEFAULT_CRITICAL_VISEME_IDS, 15], }; ``` ## Presets Define presets as **module-level constants** so the reference is stable: ```ts theme={null} // Natural conversation — a good default for voice AI export const CONVERSATION = { minVisemeInterval: 60, mergeWindow: 80, keyVisemePreference: 0.7, preserveSilence: true, similarityThreshold: 0.6, preserveCriticalVisemes: true, } as const; // Fast / excited speech — coarser merging export const FAST_SPEECH = { minVisemeInterval: 90, mergeWindow: 120, keyVisemePreference: 0.6, preserveSilence: true, similarityThreshold: 0.4, preserveCriticalVisemes: true, criticalVisemeMinDuration: 70, } as const; // Clear articulation — education / language learning export const EDUCATIONAL = { minVisemeInterval: 40, mergeWindow: 50, keyVisemePreference: 0.9, preserveSilence: true, similarityThreshold: 0.8, preserveCriticalVisemes: true, } as const; ``` ```tsx theme={null} const playback = useMascotPlayback({ enableNaturalLipSync: true, naturalLipSyncConfig: CONVERSATION }); ``` Start from `CONVERSATION`. Raise `minVisemeInterval` / `mergeWindow` for smoother (lazier) motion; lower them for crisper articulation. ## Without React The processor is exported from `@mascotbot/core/rive` as a class and a function: ```ts theme={null} import { processNaturalLipSync, NaturalLipSyncProcessor } from "@mascotbot/core/rive"; ``` When you enable `enableNaturalLipSync` on `MascotPlayback` / `useMascotPlayback`, this runs internally — you only call it directly if you post-process visemes outside the playback engine. ## Next <Columns> <Card title="React hooks" icon="react" href="/libraries/react-hooks"> `useMascotPlayback` options. </Card> <Card title="Visemes & the timeline" icon="waveform-lines" href="/concepts/visemes-and-timeline"> What is being processed. </Card> <Card title="Troubleshooting" icon="wrench" href="/libraries/react-troubleshooting"> The stable-reference bug. </Card> </Columns> # Offline Lip Sync - Generate, Persist & Replay Viseme Timelines Source: https://docs.mascot.bot/libraries/offline-lipsync Run Mascotbot inference once, persist the VisemeTimeline as JSON, and replay it forever with zero reprocessing — ideal for prefetching, queues, and video export. The offline path is the SDK's most powerful pattern: **run inference once, persist the result as JSON, and replay it forever** without touching the model, the network, or a license refresh. It is the right tool for prefetching, queued playback, deterministic video export, and any case where the same audio is animated more than once. The artifact is the [`VisemeTimeline`](/concepts/visemes-and-timeline) — plain, versioned JSON. ## Generate → persist → replay <Steps> <Step title="Generate once"> Run `processAudio` (vanilla) or `useProcessAudio` (React). `result.timeline` is the artifact. </Step> <Step title="Persist anywhere"> It is plain JSON — `localStorage`, your database, a CDN object, a file. </Step> <Step title="Replay with zero reprocessing"> `parseTimeline(JSON.parse(stored))` → `playback.setTimeline(...)` → `playback.play()`. No inference, no streaming session, no network. </Step> </Steps> ### React ```tsx theme={null} "use client"; import { useMascot, useProcessAudio, parseTimeline } from "@mascotbot/react"; import { useMascotPlayback } from "@mascotbot/react/rive"; const KEY = "greeting.vtl"; function Greeting() { const { status } = useMascot(); const cached = typeof window !== "undefined" && localStorage.getItem(KEY); // Only run inference when there is no cached timeline. const { result } = useProcessAudio(cached ? null : "/audio/greeting.wav"); const playback = useMascotPlayback({ enableNaturalLipSync: true }); function play() { if (status !== "ready") return; let timeline; if (cached) { timeline = parseTimeline(JSON.parse(cached)); // zero reprocessing } else if (result) { timeline = result.timeline; localStorage.setItem(KEY, JSON.stringify(timeline)); // persist for next time } else return; new Audio("/audio/greeting.wav").play().catch(() => {}); playback.setTimeline(timeline); playback.play(); } return <button onClick={play} disabled={status !== "ready"}>Play</button>; } ``` ### Vanilla ```ts theme={null} import { LipsyncClient, parseTimeline } from "@mascotbot/core"; const client = await LipsyncClient.init({ apiKey: "mascot_pub_…" }); // 1. Generate once (16 kHz mono Float32 in [-1, 1]) const { timeline } = await client.processAudio(audio16kMono); // 2. Persist localStorage.setItem("greeting.vtl", JSON.stringify(timeline)); // 3. Later — replay, zero reprocessing const restored = parseTimeline(JSON.parse(localStorage.getItem("greeting.vtl")!)); playback.setTimeline(restored); playback.play(); ``` ## Assembling a timeline yourself If you already have per-frame viseme ids (e.g. from your own batch job), build a timeline with the pure converters instead of running inference: ```ts theme={null} import { framesToTimeline, timelineToCues } from "@mascotbot/core"; const timeline = framesToTimeline(visemeIdsPer10ms, { speechMs }); playback.setTimeline(timeline); // timelineToCues(timeline) → the { offset, visemeId }[] the engine consumes ``` ## Prefetching & queues Because a timeline is detached from the model, you can compute many ahead of time and play them instantly later: * **Prefetch on idle** — generate timelines for likely-next utterances during idle time; play from cache the moment they are needed (no inference latency at play time). * **Queue playback** — store a list of `{ audioUrl, timeline }` pairs; for each, start the audio and `playback.setTimeline(timeline)` in lockstep. * **Server-side precompute** — generate timelines in a build step or backend job, ship the JSON with your assets, and the client never runs inference for that content at all. `speechMs` rides inside the timeline, so cached replay never re-meters and never re-infers. ## Deterministic video export For frame-accurate rendering (recording the avatar to video), a timeline gives you a fixed, inspectable script: the same JSON produces the same mouth frames every run. Drive `playback.seek(ms)` to a render clock instead of wall-clock playback, capture the canvas per frame, and mux against the original audio. Because there is no live inference in the loop, export is reproducible and as fast as your renderer. ## Versioning & the trust boundary `parseTimeline` validates untrusted/persisted JSON and throws a `LipsyncError` with `.code === "bad_timeline"` on a version or shape mismatch — so a stale stored timeline fails loudly instead of animating garbage: ```ts theme={null} import { parseTimeline, LipsyncError } from "@mascotbot/core"; try { playback.setTimeline(parseTimeline(JSON.parse(stored))); } catch (err) { if (err instanceof LipsyncError && err.code === "bad_timeline") { // regenerate via client.processAudio(); do not treat as license/network } } ``` Always load persisted timelines through `parseTimeline`, never `JSON.parse` alone. `VISEME_TIMELINE_VERSION` bumps on breaking changes; old artifacts are rejected deterministically. ## Next <Columns> <Card title="Visemes & the timeline" icon="waveform-lines" href="/concepts/visemes-and-timeline"> The timeline format in detail. </Card> <Card title="Core client" icon="cube" href="/core/client"> `processAudio` and helpers. </Card> <Card title="Error codes" icon="triangle-exclamation" href="/reference/error-codes"> `bad_timeline` and the rest. </Card> </Columns> # OpenAI Realtime API Avatar — Build an Interactive Lip-Synced ChatGPT Avatar Source: https://docs.mascot.bot/libraries/openai-realtime-api-avatar Build an interactive lip-synced avatar powered by OpenAI Realtime API. Real-time lip sync, ephemeral token security, WebSocket streaming, natural viseme processing, and production-ready React components. Open-source alternative to HeyGen, D-ID, Synthesia. TypeScript/React tutorial → # OpenAI Realtime API Avatar — Build an Interactive Lip-Synced AI Avatar Add a **lip-synced animated avatar** to your OpenAI Realtime API application in minutes. Mascot Bot SDK works alongside the official OpenAI Agents Realtime SDK (`@openai/agents-realtime`) — your existing OpenAI code stays untouched. The SDK works alongside your OpenAI Realtime setup and animates a real-time avatar from the audio. <img alt="OpenAI Realtime API Avatar with real-time lip sync — interactive AI mascot powered by GPT" /> <Columns> <Card title="Quick Start" icon="rocket" href="#quick-start"> Add avatars in 5 minutes </Card> <Card title="Live Demo" icon="play" href="https://mascot.bot/openai-realtime"> See OpenAI Realtime avatar in action </Card> <Card title="GitHub Repo" icon="github" href="https://github.com/mascotbot-templates/openai-realtime-avatar"> Complete example code </Card> <Card title="Features" icon="sparkles" href="#features"> Real-time lip sync & more </Card> <Card title="API Reference" icon="code" href="#api-reference"> Complete hook documentation </Card> <Card title="Deploy" icon="triangle" href="https://vercel.com/new/clone?repository-url=https%3A%2F%2Fgithub.com%2Fmascotbot-templates%2Fopenai-realtime-avatar"> One-click deployment </Card> </Columns> ## Why Add an Avatar to Your OpenAI Realtime API App? A voice-only ChatGPT-style assistant is invisible. A **lip-synced avatar** gives it a face and makes interactions feel human. The Mascot Bot SDK adds that without changing how you use OpenAI Realtime: keep `@openai/agents-realtime`, and the SDK lip-syncs the assistant's audio in real time. ### How It Works: Tap the Assistant Audio OpenAI Realtime has two transports, and the integration differs only in how you obtain the assistant's audio as a `MediaStream`: * **WebRTC (recommended, cleanest).** The session plays the assistant audio into an `<audio>` element you supply; you tap it with the SDK's cross-browser [`createElementTap()`](/realtime/overview#tap-a-playing-element). No extra SDK audio piece. * **WebSocket.** The session hands you raw PCM16 chunks and does not play them; [`createPCMStreamPlayer`](/core/pcm-stream-player) plays them and exposes the tap. Either way: the browser connects to OpenAI with an **ephemeral client secret** minted server-side, and the SDK infers visemes from the assistant audio in real time. No Mascot Bot endpoint sits in the path. See [Realtime overview](/realtime/overview). ## Features ### <Icon icon="bullseye" />  Real-time Lip Sync for OpenAI Realtime API Real-time viseme inference from the assistant's audio — no server round-trip for visemes. ### <Icon icon="bolt" />  120fps Avatar Animation WebGL2 + Rive runtime for smooth facial motion. ### <Icon icon="plug" />  Native OpenAI Agents Realtime SDK Compatibility Use `@openai/agents-realtime` exactly as documented. The SDK never proxies the OpenAI connection. ### <Icon icon="shield-halved" />  Ephemeral Token Security Mint a short-lived `client_secret` server-side via `POST /v1/realtime/client_secrets`. The standing `OPENAI_API_KEY` never reaches the browser. ### <Icon icon="rotate" />  WebRTC or WebSocket Streaming WebRTC self-plays (tap via the SDK's [`createElementTap()`](/realtime/overview#tap-a-playing-element)); WebSocket hands you PCM (play + tap via `createPCMStreamPlayer`). Both end at one `useLipsyncStream` call. ### <Icon icon="masks-theater" />  Natural Lip Sync Processing Optional viseme post-processing for natural, non-robotic motion — [Natural lip sync](/libraries/natural-lip-sync). ### <Icon icon="microphone" />  Voice Activity Detection (VAD) Configure server-side VAD (`turn_detection: { type: "server_vad" }`) in the token route — an OpenAI feature, unchanged by the SDK. ### <Icon icon="clock" />  Session Management Re-mint a client secret per session. `player.stop()` / `session.on("audio_interrupted")` handle barge-in. ## Quick Start ### Installation ```ini .npmrc theme={null} @mascotbot:registry=https://npm.mascot.bot/ //npm.mascot.bot/:_authToken=mascot_xxx ``` ```bash theme={null} pnpm add @mascotbot/react @rive-app/react-webgl2 @rive-app/webgl2 @openai/agents-realtime ``` <Note> Get your Mascot Bot key at [app.mascot.bot/api-keys](https://app.mascot.bot/api-keys). Full setup: [Installation](/installation). `@openai/agents-realtime` is OpenAI's official SDK, used unchanged. </Note> ### Basic Integration ```tsx theme={null} "use client"; import { MascotProvider } from "@mascotbot/react"; import { Mascot, MascotRive } from "@mascotbot/react/rive"; export default function App() { return ( <MascotProvider apiKey="mascot_pub_…"> <MascotProvider> <Mascot src="/mascot.riv"> <MascotRive /> <OpenAIAvatar /> </Mascot> </MascotProvider> </MascotProvider> ); } ``` `OpenAIAvatar` is built in [Step 2](#step-2-create-your-avatar-component). ## Complete Implementation Guide ### Step 1: Set Up Ephemeral Token Generation (Server-Side) ```typescript theme={null} // app/api/openai/token/route.ts export const runtime = "nodejs"; export async function POST() { const key = process.env.OPENAI_API_KEY; if (!key) return Response.json({ error: "OPENAI_API_KEY not set" }, { status: 400 }); const model = "gpt-realtime"; const res = await fetch("https://api.openai.com/v1/realtime/client_secrets", { method: "POST", headers: { Authorization: `Bearer ${key}`, "Content-Type": "application/json" }, body: JSON.stringify({ session: { type: "realtime", model, output_modalities: ["audio"], audio: { input: { turn_detection: { type: "server_vad", threshold: 0.5, silence_duration_ms: 500 } }, output: { voice: "marin" }, }, }, }), }); if (!res.ok) return Response.json({ error: `OpenAI ${res.status}` }, { status: 502 }); const json = (await res.json()) as { value: string }; return Response.json({ clientSecret: json.value, model }); } ``` ### Step 2: Create Your Avatar Component **Recommended — WebRTC** (the session self-plays; tap its `<audio>`): ```tsx theme={null} "use client"; import { useEffect, useRef, useState } from "react"; import { useMascot } from "@mascotbot/react"; import { useMascotPlayback, useLipsyncStream } from "@mascotbot/react/rive"; const LIP_SYNC = { minVisemeInterval: 60, mergeWindow: 80 } as const; export function OpenAIAvatar() { const { client, status } = useMascot(); const playback = useMascotPlayback({ stream: true, enableNaturalLipSync: true, naturalLipSyncConfig: LIP_SYNC }); const [stream, setStream] = useState<MediaStream | null>(null); const teardownRef = useRef<null | (() => void)>(null); const { error, attached } = useLipsyncStream({ client, playback, source: { kind: "mediaStream", stream } }); useEffect(() => () => teardownRef.current?.(), []); const connect = async () => { if (status !== "ready") return; const { clientSecret } = await (await fetch("/api/openai/token", { method: "POST" })).json(); const { RealtimeAgent, RealtimeSession } = await import("@openai/agents-realtime"); const audioEl = new Audio(); // supply our own so we can tap it const agent = new RealtimeAgent({ name: "Assistant", instructions: "Keep replies short." }); const session = new RealtimeSession(agent, { transport: "webrtc" }); await session.connect({ apiKey: clientSecret, audioElement: audioEl }); // createElementTap from "@mascotbot/react" — cross-browser // tap (Safari has no captureStream). /realtime/overview#tap-a-playing-element const tap = createElementTap(); setStream(tap.stream); tap.attach(audioEl); teardownRef.current = () => { tap.close(); void session.close(); }; }; return ( <div> <button onClick={connect} disabled={status !== "ready"}>Connect</button> <span>{stream ? (attached ? "lip-sync attached" : "attaching…") : "idle"}</span> {error ? <p>{error.message}</p> : null} </div> ); } ``` **Alternative — WebSocket** (the session hands you raw PCM; play + tap it): ```tsx theme={null} import { createPCMStreamPlayer } from "@mascotbot/react"; import { WavRecorder } from "wavtools"; // inside connect(): const player = createPCMStreamPlayer({ sampleRate: 24000 }); // create in the click, before await setStream(player.outputStream); const session = new RealtimeSession(agent, { transport: "websocket", model }); session.on("audio", (e: { data: ArrayBuffer }) => player.pushPCM16(new Uint8Array(e.data))); session.on("audio_interrupted", () => player.stop()); await session.connect({ apiKey: clientSecret }); const recorder = new WavRecorder({ sampleRate: 24000 }); await recorder.begin(); // Pass the typed-array view itself (cast) — `.buffer` is over-long and yields "empty bytes". await recorder.record((d: { mono: Int16Array }) => session.sendAudio(d.mono as unknown as ArrayBuffer)); ``` <Warning> On WebSocket, never feed a self-playing element through the player and never use both paths at once — that double-plays the voice. The PCM player is for the WebSocket transport only; WebRTC self-plays and is tapped directly. </Warning> ### Step 3: Advanced Features * **Natural lip sync** — stable `naturalLipSyncConfig`; presets in [Natural lip sync](/libraries/natural-lip-sync). * **Barge-in** — `session.on("audio_interrupted", () => player.stop())` (WebSocket) or stop playback on the WebRTC element. * **VAD** — tune `turn_detection` in the token route (OpenAI feature). * **Custom Rive inputs** — the SDK only writes the mouth; drive gestures/outfits yourself, detect via `useMascotInputs().has(name)` ([Rive co-existence](/concepts/rive-coexistence)). ## API Reference The standard SDK surface plus OpenAI's official SDK: | Surface | Role | | --------------------------------------------------------------- | --------------------------------------------------------------------------------------------- | | `<MascotProvider apiKey>` | Licensed avatar client. [Config](/installation#4-configure-the-client). | | `<MascotProvider>` / `<Mascot src>` / `<MascotRive>` | Load and render the avatar. | | `useMascotPlayback({ stream: true, enableNaturalLipSync })` | Mouth playback engine. | | `useLipsyncStream({ source: { kind: "mediaStream", stream } })` | Lip-syncs the tapped audio. [Reference](/libraries/streaming-and-mic). | | `createPCMStreamPlayer({ sampleRate: 24000 })` | WebSocket transport only — plays PCM + exposes the tap. [Reference](/core/pcm-stream-player). | | `@openai/agents-realtime` `RealtimeSession` | OpenAI's official SDK — unchanged. Model `gpt-realtime`. | ## OpenAI Realtime API Pricing for Voice Avatars ### All-in Cost Per Hour (OpenAI + Mascot Bot) OpenAI Realtime usage is billed by OpenAI per their pricing. Mascot Bot meters by your plan's speech-seconds or MAU allowance and adds no per-minute audio cost. Replaying a persisted [timeline](/libraries/offline-lipsync) does not re-meter. Check current OpenAI Realtime pricing in the OpenAI documentation. ## MascotBot vs HeyGen vs D-ID vs Synthesia for Interactive Avatars Pre-rendered talking-head services (HeyGen, D-ID, Synthesia) generate video server-side and stream it back — higher latency, per-minute video cost, and no real-time control. Mascot Bot is a **real-time** alternative: a lightweight vector avatar lip-synced at up to 120fps, no video pipeline, and full control of every non-mouth animation through raw Rive. ## Use Cases ### AI Customer Service Avatar A visible support assistant with on-brand appearance. ### ChatGPT Avatar for Your Product Give your GPT-powered assistant a face that talks in real time. ### Educational AI Tutor Crisp articulation with the `educational` natural-lip-sync preset. ### Voice AI Virtual Receptionist A welcoming branded front desk. ### AI Mascot for Streaming & Content A reactive on-screen character; non-mouth animation is yours via raw Rive inputs. ## Troubleshooting ### Avatar Not Moving? Confirm `status === "ready"`, the tap stream is set (WebRTC: `createElementTap()`; WebSocket: `player.outputStream`), and the Rive file uses artboard `Character` + state machine `mascotStateMachine` with inputs `100`–`118`. ### Only First Second of Speech Animated? Non-stable `naturalLipSyncConfig` reinitializes playback — use a module constant ([Troubleshooting](/libraries/react-troubleshooting)). ### Connection Fails on Second Call? Client secrets are short-lived/single-use. Mint a fresh one per session. ### No Audio Playing? On WebSocket, `createPCMStreamPlayer` must be created inside the user-gesture click before any `await`. On WebRTC, ensure the supplied `<audio>` element is allowed to play (the click satisfies autoplay). ### "Invalid audio — empty bytes" Errors in Console? On the WebSocket mic path, pass the `Int16Array` view itself to `session.sendAudio` (`d.mono as unknown as ArrayBuffer`), not `d.mono.buffer` — the backing buffer is pooled/over-long and serializes as empty bytes. ### sendAudio Not Working? Confirm the recorder sample rate matches the session (24 kHz in the example) and that you started recording **after** `session.connect`. ## FAQ ### How Does Mascot Bot Work with the OpenAI Agents Realtime SDK? Alongside it. You use `@openai/agents-realtime` as documented; the SDK lip-syncs the assistant's audio in real time. ### Does It Work With My Existing OpenAI Code? Yes. Use `@openai/agents-realtime` as documented; the SDK turns the assistant's audio into a real-time avatar. ### Do I Modify My OpenAI Realtime Code? No. Add an audio tap (WebRTC) or a PCM player (WebSocket) plus one `useLipsyncStream` call. ### What OpenAI Models Support the Realtime API? A Realtime model such as `gpt-realtime`, configured in the server token route. ### How Does the SDK Connect to OpenAI? You connect directly to OpenAI with an ephemeral client secret minted by your server; the SDK only taps the resulting audio. ### Is Audio Sent to Mascot Bot? Your users' speech is processed by the SDK in their browser and isn't sent to or stored on Mascotbot servers. ### Does Mascot Bot Support Both OpenAI and Gemini? Yes — see the [Gemini Live guide](/libraries/gemini-live-api-avatar) and the [Realtime overview](/realtime/overview). ### Is This an Open-Source Alternative to HeyGen Interactive Avatar? Yes — a real-time alternative to server-rendered talking-head avatars. ## Start Building with OpenAI Realtime API Avatar <Columns> <Card title="Live Demo" icon="play" href="https://mascot.bot/openai-realtime"> See it in action </Card> <Card title="Demo Repository" icon="github" href="https://github.com/mascotbot-templates/openai-realtime-avatar"> Complete working example </Card> </Columns> ## Next Steps 1. Get a key at [app.mascot.bot/api-keys](https://app.mascot.bot/api-keys) and install from the [private registry](/installation). 2. Add the server token route and the avatar component above (WebRTC recommended). 3. Tune motion with [natural lip sync](/libraries/natural-lip-sync). 4. Review the [realtime overview](/realtime/overview) and [PCM stream player](/core/pcm-stream-player) for the underlying pattern. # Mascotbot React Hooks Reference - useMascot, useMascotPlayback & More Source: https://docs.mascot.bot/libraries/react-hooks Complete reference for the Mascotbot lipsync React hooks: useMascot, useProcessAudio, useMascotRive, useMascotInputs, useMascotPlayback, useLipsyncStream, useLoadRive, and useRiveAsset. Every hook in the React SDK, grouped by subpath. Audio-pipeline hooks come from `@mascotbot/react`; Rive hooks from `@mascotbot/react/rive`. ## Audio pipeline — `@mascotbot/react` ### `useMascot()` ```ts theme={null} const { client, status, error, reload } = useMascot(); ``` The licensed inference client from the enclosing `<MascotProvider>`. | Field | Type | Notes | | -------- | ----------------------- | -------------------------------------------------------------------------- | | `client` | `LipsyncClient \| null` | `null` until init resolves | | `status` | `LipsyncStatus` | `idle \| initializing \| ready \| running \| degraded \| refused \| error` | | `error` | `Error \| null` | A typed [error](/reference/error-codes) on `refused` / `error` | | `reload` | `() => void` | Re-runs init (e.g. after the user fixes a key) | Gate audio work on `status === "ready"`. ### `useProcessAudio(audioUrl)` ```ts theme={null} const { result, loading, error } = useProcessAudio("/audio/greeting.wav"); ``` Fetches the URL, decodes, resamples to 16 kHz, and runs inference **once**. Pass `null` to skip. `result` is a `ProcessAudioResult`: ```ts theme={null} result.timeline; // VisemeTimeline — the serializable artifact result.durationMs; // total audio duration result.speechMs; // non-silent ms detected ``` Hand `result.timeline` to `useMascotPlayback().setTimeline()`, or `JSON.stringify` it to persist and replay later with zero reprocessing — [Offline lip sync](/libraries/offline-lipsync). ## Rive layer — `@mascotbot/react/rive` ### `useMascotRive()` ```ts theme={null} const { rive, isRiveLoaded, RiveComponent, setImageAsset } = useMascotRive(); ``` The Rive instance + canvas for the enclosing `<Mascot>`. `rive` is the **raw, unmodified** `@rive-app/*` instance — yours for data binding, custom inputs, events, and ViewModels. The SDK never wraps it. `setImageAsset` swaps a runtime image asset (e.g. a custom face texture). ### `useMascotInputs<T>()` ```ts theme={null} const { riveInputs, custom, has } = useMascotInputs<"wave" | "reveal">(); if (has("wave")) custom.wave.fire(); ``` | Field | Notes | | ------------ | ------------------------------------------------------------------------------ | | `riveInputs` | SDK-driven input handles (mouth / `is_speaking` / `stress`) | | `custom` | Your declared inputs, typed by `T`. **Never `undefined`** | | `has(name)` | Authoritative presence check — use this, never raw `rive.stateMachineInputs()` | This is the supported way to detect and drive non-mouth inputs. See [Rive co-existence](/concepts/rive-coexistence). ### `useMascotPlayback(options?)` ```ts theme={null} const playback = useMascotPlayback({ stream: true, enableNaturalLipSync: true }); playback.setTimeline(result.timeline); // offline replay playback.pushVisemes(cues); // streaming append playback.stress([{ offset: 0, stress: 1 }]); // SDK-driven emphasis cues playback.play(); playback.pause(); playback.seek(0); playback.reset(); ``` Wraps the framework-agnostic `MascotPlayback`. `stress([{ offset, stress }])` schedules emphasis cues — the SDK animates the Rive `stress` input from them on the playback clock (see [Stress emphasis](/realtime/overview#stress-emphasis-and-gestures)). Options: | Option | Type | Purpose | | -------------------------------------------------- | ------------------------------- | -------------------------------------------------------------- | | `stream` | `boolean` | Streaming mode (mic / realtime); pair with `useLipsyncStream` | | `enableNaturalLipSync` | `boolean` | Smoother, less robotic merging | | `naturalLipSyncConfig` | `Partial<NaturalLipSyncConfig>` | Tune merging — [Natural lip sync](/libraries/natural-lip-sync) | | `desktopTransitionSpeed` / `mobileTransitionSpeed` | `number` | Mouth blend rate | | `setSpeakingState` | `boolean` | Auto-drive `is_speaking` (default on) | | `manualSpeakingStateControl` | `boolean` | Take manual control of `is_speaking` | <Warning> Pass a **stable** `naturalLipSyncConfig` reference (a module constant or `useState`/`useMemo`). A fresh object literal every render reinitializes playback and breaks lip sync after the first chunk — [Troubleshooting](/libraries/react-troubleshooting). </Warning> ### `useLipsyncStream(args)` ```ts theme={null} const { error, attached, pushAudio, pushBase64PCM16, reset } = useLipsyncStream({ client, playback, source: { kind: "mic" }, // | { kind: "mediaStream", stream } | { kind: "manual" } enabled: isLive, onFrame: (f) => {}, // optional per-window telemetry }); ``` The unified audio→viseme stream for live audio: microphone, a tapped `MediaStream` (realtime providers, played audio), or manual chunk pushing. Full guide: [Streaming & microphone](/libraries/streaming-and-mic). ### Low-level Rive loaders | Hook | Returns | Use when | | ------------------------------------------------ | ------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------- | | `useLoadRive(options?)` | `{ rive, isRiveLoaded, RiveComponent, setImageAsset }` | You need to load Rive yourself; pass an options object (`{ stateMachineName, ... }`). `<Mascot>` uses this internally. | | `useRiveAsset(riveParams?, opts?)` | `RiveState & { setImageAsset }` | Lowest-level escape hatch: `useRive` + a runtime image-asset swapper. | | `useSafeStateMachineInput(rive, _, name, init?)` | input or `DEFAULT_SM_INPUT` | Single-input lookup with the `mascotStateMachine` / `InLesson` fallback shim. | | `useIsMobileDevice()` | `boolean` | Touch + UA + screen-size heuristic (e.g. to pick a transition speed). | ## Next <Columns> <Card title="Streaming & mic" icon="microphone" href="/libraries/streaming-and-mic"> `useLipsyncStream` in depth. </Card> <Card title="Offline lip sync" icon="box-archive" href="/libraries/offline-lipsync"> `useProcessAudio` → persist → replay. </Card> <Card title="Rive co-existence" icon="puzzle-piece" href="/concepts/rive-coexistence"> Driving non-mouth inputs. </Card> </Columns> # Mascotbot React SDK Documentation Source: https://docs.mascot.bot/libraries/react-sdk Learn how to install and use the Mascotbot SDK with React. # Mascotbot React SDK `@mascotbot/react` is the React layer of the Mascotbot avatar SDK. It gives you a provider, a few hooks, and a thin set of Rive components. Speech goes in; the avatar speaks in real time. This page is the map. Each area links to its dedicated guide. ## Install ```ini .npmrc theme={null} @mascotbot:registry=https://npm.mascot.bot/ //npm.mascot.bot/:_authToken=mascot_xxx ``` <CodeGroup> ```bash pnpm theme={null} pnpm add @mascotbot/react pnpm add @rive-app/react-webgl2 @rive-app/webgl2 ``` ```bash npm theme={null} npm i @mascotbot/react npm i @rive-app/react-webgl2 @rive-app/webgl2 ``` ```bash yarn theme={null} yarn add @mascotbot/react yarn add @rive-app/react-webgl2 @rive-app/webgl2 ``` </CodeGroup> `@rive-app/react-webgl2` + `@rive-app/webgl2` are optional peer dependencies of the `/rive` subpath — install them only when you render an avatar. Full registry and key setup is in [Installation](/installation). ## The two subpaths | Import | Provides | | ----------------------- | --------------------------------------------------------------------------------------------------------------------- | | `@mascotbot/react` | `MascotProvider`, `useMascot`, `useProcessAudio` — the audio pipeline + the top-level provider. | | `@mascotbot/react/rive` | `Mascot`, `MascotRive`, `useMascotRive`, `useMascotInputs`, `useMascotPlayback`, `useLipsyncStream` — the Rive layer. | ## Composition One provider, one component. `<MascotProvider>` owns the licensed inference engine; `<Mascot>` mounts the Rive avatar inside it. ```tsx theme={null} "use client"; import { MascotProvider } from "@mascotbot/react"; import { Mascot, Fit, Alignment } from "@mascotbot/react/rive"; export default function App() { return ( <MascotProvider apiKey="mascot_pub_…"> <Mascot src="/mascot-fox.riv" layout={{ fit: Fit.Contain, alignment: Alignment.Center }} /> </MascotProvider> ); } ``` | Component | Purpose | | ----------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `<MascotProvider apiKey="…">` | Initializes one `LipsyncClient` and exposes it to descendants. Props extend the [client config](/installation#4-configure-the-client). | | `<Mascot src=… \| rive=…>` | Loads Rive and exposes its inputs via context. Pass a `.riv` URL **or** a pre-loaded Rive instance. Renders the canvas by default; children opt out and take control (you can still use `<MascotRive />` inside). | | `<MascotRive>` | The canvas component. Use it inside `<Mascot>` children when you need a custom loading slot or layout wrapper; otherwise `<Mascot>` mounts it for you. | <Tip> `<Mascot>` accepts a pre-loaded `rive` instance you constructed yourself. The SDK only writes the mouth, `is_speaking`, and `stress` — everything else stays yours. See [Rive co-existence](/concepts/rive-coexistence). </Tip> ## Where to go next <Columns> <Card title="Hooks reference" icon="react" href="/libraries/react-hooks"> `useMascot`, `useProcessAudio`, `useMascotRive`, `useMascotInputs`, `useMascotPlayback`, `useLipsyncStream`, and more. </Card> <Card title="Offline lip sync" icon="box-archive" href="/libraries/offline-lipsync"> Generate a timeline once, persist it, replay with zero reprocessing. </Card> <Card title="Streaming & microphone" icon="microphone" href="/libraries/streaming-and-mic"> `useLipsyncStream` for mic, tapped `MediaStream`, and manual audio. </Card> <Card title="Natural lip sync" icon="masks-theater" href="/libraries/natural-lip-sync"> Smoother, less robotic mouth movement. </Card> <Card title="Realtime providers" icon="bolt" href="/realtime/overview"> OpenAI Realtime, Gemini Live, ElevenLabs. </Card> <Card title="Troubleshooting" icon="wrench" href="/libraries/react-troubleshooting"> Common integration issues and fixes. </Card> </Columns> ## WebGL2 rendering The Rive avatar renders through `@rive-app/*-webgl2`: hardware-accelerated, smoother animation, advanced effects, efficient memory, and good scaling with multiple mascots. WebGL2 is required for the renderer (the audio pipeline does not need it). # Mascotbot React SDK Troubleshooting Source: https://docs.mascot.bot/libraries/react-troubleshooting # Troubleshooting Common integration issues with `@mascotbot/react` 0.2.x and their fixes. If your symptom is a license refusal, match the `error.code` against the [error-code reference](/reference/error-codes) first. ## Install fails with 401 / 403 from npm.mascot.bot The private registry needs a valid token in `.npmrc`: ```ini theme={null} @mascotbot:registry=https://npm.mascot.bot/ //npm.mascot.bot/:_authToken=mascot_xxx ``` A `403 wrong_key_scope` means the token is the wrong kind for the registry — mint one at [app.mascot.bot/api-keys](https://app.mascot.bot/api-keys). Make sure the `.npmrc` is at the project root and the token has no trailing newline. ## Status never reaches `ready` Read `status` and `error` from `useMascot()`: * `status === "refused"` → an authorization problem. Branch on `error.code` (`dev_key_on_public_domain`, `prod_key_on_localhost`, `origin_not_allowed`, `key_disabled`, …) and show the matching fix. See [Licensing & keys](/concepts/licensing-and-keys). * `status === "error"` with a `NetworkError` → the device cannot reach `license.mascot.bot`. Check connectivity, ad blockers, and corporate proxies. * Stuck on `initializing` with no error → WebAssembly or `crypto.subtle` is unavailable (old browser, insecure context). The SDK requires a secure context (HTTPS or `localhost`). ## Blank canvas — the avatar never renders Almost always the Rive state machine name. Pass **only** `mascotStateMachine` (or `STATE_MACHINE_NAMES[0]`) to Rive. Rive 2.37+ throws on any unknown state-machine name in the array; the throw fires `LoadError`, suppresses `Load`, and leaves the canvas blank. Also verify the artboard is named `Character` and the file exposes mouth inputs `100`–`118`. ## Mouth frozen during an active call The SDK does not freeze on parent re-renders — Rive input handles are referentially stable and playback is carried across any internal recreate. A frozen mouth during a call is almost always one of: * **Unstable `naturalLipSyncConfig`** — a new object literal every render reinitializes the natural-lipsync processor. Pass a module constant or a `useState`/`useMemo` reference. * **Audio is not reaching the tap** — pass `onFrame` to `useLipsyncStream` and log `silenceDetected` / `emittedVisemeId`. Rising emitted IDs with a dead mouth means audio is reaching the engine but the Rive handle isn't being written; a flat zero / `silenceDetected: true` means the tap is on a silent corpse (typical of a self-playing realtime provider torn down and not re-attached — see [ElevenLabs 2nd-call diagnostic](#elevenlabs-2nd-call-has-no-lip-sync-1st-call-worked)). * **Wrong `source` shape** — `{ kind: "mediaStream", stream }` where `stream` is `null` or has only ended tracks. ## Call disconnects immediately after `onConnect` (`reason: 'user'`) Symptom: `Conversation.startSession({ ... })` resolves, your `onConnect` fires, the agent's `first_message` may even reach `onMessage`, then `onStatusChange` flips to `disconnecting` → `disconnected` and `onDisconnect` runs with `details === { reason: 'user' }`. There's no network error, no `onError`, no server-side disconnect — **your own code called `endSession()`**. The usual cause is an unmount-cleanup effect whose dep array contains a `teardown` callback whose **identity flips on every render**: ```tsx theme={null} // BUG — re-runs the cleanup on EVERY render, not just unmount. const teardown = useCallback(() => { void convoRef.current?.endSession().catch(() => {}); // …other resource releases }, [setSomeMascotInput]); // ← unstable dep useEffect(() => () => teardown(), [teardown]); ``` `setSomeMascotInput` typically traces back to handles returned by `useMascotInputs()` (which intentionally return a **fresh `{ custom, has }` object per render** — see [Rive co-existence](/concepts/rive-coexistence)). Each fresh handle → new `useCallback` chain → new `teardown` identity → the cleanup runs, which calls `endSession()`, which surfaces as a `'user'` disconnect right after `onConnect`. Fix — stabilise the unmount cleanup with a ref so it runs **once on unmount only**, while always invoking the latest `teardown` closure: ```tsx theme={null} const teardownRef = useRef(teardown); teardownRef.current = teardown; // refresh every render, no deps useEffect(() => () => teardownRef.current?.(), []); // [] — true unmount ``` This pattern is correct regardless of how often `teardown`'s identity changes; the ref always points at the latest closure when the component finally unmounts. Apply the same shape to any other long-lived cleanup that depends on hook handles which are fresh-per-render (`useMascotInputs`, `useMascotRive` — see also the [ElevenLabs onModeChange recipe](/libraries/elevenlabs-avatar#gestures-on-every-agent-turn) which captures `custom` in a ref for the same reason). How to diagnose: temporarily log the disconnect detail and any frames so you can distinguish a self-end from an agent/server end: ```tsx theme={null} await Conversation.startSession({ signedUrl, onMessage: (m) => console.log("[debug] msg:", m), onStatusChange: (s) => console.log("[debug] status:", s.status), onDisconnect: (d) => console.log("[debug] disconnect reason:", d), // … }); ``` `reason: 'user'` is your code; `reason: 'agent'` is the agent stopping the call; any `onError` first means a server-side problem. ## Mouth flickers when speech stops This is handled by the SDK's internal **−50 dBFS silence gate** — do **not** add your own gate. If you still see phantom shapes at end of utterance, you are likely feeding a self-playing realtime provider through `createPCMStreamPlayer` (double audio / doubled inference). Tap the provider's own output instead — see [Realtime providers](/realtime/overview). ## Both the SDK and the provider play audio (double voice) `createPCMStreamPlayer` is **only** for providers that hand you raw PCM and do not play it (Gemini Live, OpenAI Realtime over WebSocket). For self-playing providers (ElevenLabs, OpenAI Realtime over WebRTC), do not use the player — tap their existing playback with the SDK's cross-browser [`createElementTap()`](/realtime/overview#tap-a-playing-element) and feed that to `useLipsyncStream({ source: { kind: "mediaStream", stream } })`. ## No audio in a realtime/TTS demo The `AudioContext` (and `createPCMStreamPlayer`) must be created inside the user-gesture handler, before any `await`. A context created in a post-fetch microtask starts suspended and cannot resume without another gesture. Create or `resume()` the player synchronously at the top of the click handler. ## ElevenLabs 2nd call has no lip sync (1st call worked) Symptom: an ElevenLabs widget animates the mouth on the very first call, you end it cleanly, then start a new call — voice plays, console is clean, but the mouth is frozen for the entire second call. You're using the `window.Audio` patch + `<audio>` poll pattern from [the ElevenLabs avatar guide](/libraries/elevenlabs-avatar) (the cross-browser tap approach). The class: * The patch stashes a reference to the `<audio>` element ElevenLabs constructs (e.g. `w.__el = el`) so a 100 ms poll can `tap.attach()` it once it's wired up. * On call-end, `endSession()` stops the conversation but the stashed reference and the `srcObject` MediaStream both *remain* on `window`. The MediaStream's audio tracks transition to `readyState: 'ended'`, but `el.srcObject instanceof MediaStream` is still `true`. * On call #2, the poll runs almost immediately — typically *before* ElevenLabs has called `new Audio()` again. The naive check `el && el.srcObject instanceof MediaStream` accepts the stale reference, `tap.attach()` lands on a silent corpse, and zero audio reaches the new tap. Two corrections, both required (one defends against the other failing): ```tsx theme={null} // 1. Reject any candidate whose audio tracks are no longer 'live'. const isLive = (el: HTMLMediaElement | null | undefined) => !!el && el.srcObject instanceof MediaStream && el.srcObject.getAudioTracks().some((t) => t.readyState === "live"); const iv = window.setInterval(() => { const el = w.__el; if (isLive(el)) { tap.attach(el as HTMLMediaElement); window.clearInterval(iv); } else if (++tries > 100) { window.clearInterval(iv); } }, 100); // 2. In teardown, null the stash AND close the tap. Otherwise the // next call's poll latches onto the stale ref before the next // `new Audio()` lands, and every restart leaks a worklet graph. teardownRef.current = () => { window.clearInterval(iv); w.Audio = OrigAudio; w.__el = null; // ← without this, next call's first poll // sees the still-set MediaStream and // attaches to the dead element tap.close(); // ← releases the tap's AudioContext + stream void convo.endSession(); }; ``` The `isLive` check is the real defense — even if you forget the null in teardown, no element with `readyState !== 'live'` will ever be attached. The null-out is belt-and-suspenders that also avoids one wasteful poll iteration. ## One widget's lip sync is fast/garbled *after another widget ran* Symptom: widget A (e.g. a Gemini call) works; you end it and start widget B (e.g. an ElevenLabs widget) on the same page, and B's mouth runs at \~2× / flickers. B alone, or B-then-A, is fine. The whole page shares **one `<MascotProvider>` → one `LipsyncClient`**. `useLipsyncStream`'s `mediaStream` pipeline is keyed on the **stream's identity** and tears down (closes its `AudioContext` + worklet + streaming session) only when that stream *reference* changes. If, on call-end, you only `player.stop()` but keep the same `player.outputStream` in state, the pipeline never tears down — it lingers on the shared client. Widget B then opens a *second* inference pipeline on the same client and the two corrupt each other's pacing. Fully release the pipeline on **every** call-end path (`stop`, `onclose`, error), symmetric with however you created it: ```tsx theme={null} player.stop(); player.close(); // releases the AudioContext, not just the queue playerRef.current = null; setVoiceStream(null); // ← the key line: changes the stream identity // so useLipsyncStream runs its teardown ``` `createPCMStreamPlayer().stop()` only drops queued audio (barge-in); `.close()` releases the context. Self-playing taps must likewise stop polling and `setStream(null)`. (Switching the avatar by unmounting the `<Mascot>` subtree tears down implicitly — this bug only surfaces when a call *ends without unmounting*.) ## Avatar customizations (gender, colors, outline) don't apply Symptom: you set `useMascotInputs().custom.gender.value = …` once (e.g. in a mount effect) and the avatar still shows defaults. Custom inputs are no-op shims until Rive has bound the real state-machine handles, which happens **asynchronously after load**. A single early write lands on a shim and is lost; the state machine then settles into its default pose and never re-evaluates. Consume **raw** `useMascotInputs()` (its `custom`/`has` are a fresh object every render — do **not** freeze them in a memo for this), gate the write on `has(...)`, and re-assert every render. The re-application is idempotent and load-bearing — a one-shot write is the bug: ```tsx theme={null} const { custom, has } = useMascotInputs(); useEffect(() => { if (!has("gender")) return; // real input bound yet? custom.gender.value = female ? 2 : 1; custom.colourful.value = true; }); // no dep array → re-asserts until (and after) Rive binds ``` ## Avatar is hidden behind a section background `<MascotRive>` renders a `position: relative` element with **no z-index**. A positioned background sibling (`z-5`, an absolute image, etc.) will paint over it. Wrap **only** `<MascotRive>` in a low positive z — never a wrapper that also contains your call controls, or that wrapper becomes a stacking context and traps the controls under a sibling gradient: ```tsx theme={null} <div className="relative h-full w-full z-[6]"> <MascotRive /> </div> ``` ## Next.js Pages Router: "Named export not found" Pages Router has stricter module resolution. Transpile the package: ```js theme={null} /** @type {import("next").NextConfig} */ module.exports = { transpilePackages: ["@mascotbot/react"] }; ``` Then clear the cache: `rm -rf .next`. ## CSP blocks the audio worklet The worklet is served from a Blob URL by default. If your Content Security Policy forbids `worker-src blob:`, either allow it or host the worklet yourself and pass its URL via `workletUrl` on `useLipsyncStream`. ## Still stuck? Compare against the reference integration in [`apps/lipsync-demo`](/overview) (single file, no design-system deps), or email [support@mascot.bot](mailto:support@mascot.bot). See also the [migration guide](/reference/migration). # Mascotbot Session Lifecycle - Refresh, Refused & Recovery Source: https://docs.mascot.bot/libraries/session-lifecycle How a Mascotbot lipsync session moves through idle → initializing → ready → refused, what triggers the refusal, and the canonical recovery pattern using reload() — pre-check at the top of every action plus a RefusedError catch mid-flight. After `<MascotProvider>` (or `LipsyncClient.init()`) the SDK runs an authenticated refresh loop in the background — it keeps the inference engine licensed without you doing anything. Almost always you ignore this and treat `status === "ready"` as the only check you need. The case worth a page is the rare **terminal cutoff**. If too many refresh requests fail in a tight burst (laptop sleep, a long-backgrounded mobile tab, a real network outage), the session moves to `"refused"` and **stays there** — no automatic retry. Recovery is one call: `reload()`. ## The status state machine `useMascot()` exposes one value, `status`: | Value | Means | Audio work allowed? | | -------------- | --------------------------------------------------------------------------------------------------- | ------------------- | | `idle` | Provider hasn't started initialization yet (e.g. `lazy`) | No | | `initializing` | License handshake in flight | No | | `ready` | Licensed; engine warm | Yes | | `running` | Currently inferring on a frame | Yes | | `degraded` | Soft signal from the engine; still usable | Yes | | `refused` | **Terminal cutoff** — recover with `reload()` | No | | `error` | Init itself failed (bad key, network down at boot) — recover with `reload()` after fixing the cause | No | `"refused"` and `"error"` are the two states that need active recovery. The others either advance on their own or simply mean "don't push audio yet." ## What triggers `"refused"` The refresh loop poisons the session after a small burst of consecutive refresh failures in a short window. The common real-world triggers are all **network-suppression-shaped** events: * A laptop closed mid-session and reopened minutes later. * A mobile browser tab backgrounded long enough for the OS to throttle timers and pause fetches. * A device losing connectivity entirely (subway, elevator, plane). * An aggressive ad-blocker / privacy extension that started intercepting the refresh endpoint. These are normal user behavior — the SDK has no way to distinguish them from a hostile network shim, so the fail-closed move is the same: stop and surface. A typed [`RefusedError`](/reference/error-codes) is fired on the client's `"refused"` event. The React Provider listens for it and flips `status → "refused"`, and any subsequent `processAudio()` / `pushWindow()` throws the same `RefusedError`. ## The canonical recovery pattern (React) Use **both** legs together. The pre-check covers "the cutoff already happened, the user is clicking again"; the catch covers "the cutoff fires **while** this call is in flight." ```tsx theme={null} "use client"; import { useCallback } from "react"; import { useMascot, RefusedError } from "@mascotbot/react"; function SpeakButton() { const { client, status, reload } = useMascot(); const speak = useCallback(async (audio: Float32Array) => { // (1) Pre-check: the Provider is already "refused" from a prior cutoff. // Kick off re-init and bail. The user's NEXT click runs on a fresh // client — no separate "Reconnect" button needed. if (status === "refused") { reload(); return; } if (!client || status !== "ready") return; try { const result = await client.processAudio(audio); // …setTimeline / play… } catch (err) { // (2) Mid-flight cutoff: the refresh budget burned WHILE this call was // in flight. Same recovery — reload() and let the next click run // on the fresh client. if (err instanceof RefusedError) { reload(); return; } throw err; } }, [client, status, reload]); return <button onClick={() => speak(/* … */)}>Speak</button>; } ``` That's the whole pattern. Copy it into every action that touches `client.*`. <Tip> **Don't block-and-wait inside the click.** `reload()` returns immediately and re-initializes asynchronously; the user's next click runs on the new client with a normal spinner. A `await initReady` inside the handler makes the first post-cutoff tap feel laggy without adding anything. </Tip> ### Why `reload()` instead of "just retry" `reload()` tears the client down and runs `LipsyncClient.init()` again — a fresh license, a fresh refresh chain, fresh in-memory state. A naive retry of `processAudio()` would just hit the same poisoned session. There is no "un-refuse" API; re-init is the recovery, and that's what `reload()` does. `reload()` is also the right move for `status === "error"` once you've fixed the underlying cause (e.g. the user pasted a valid key after seeing `invalid_api_key`). ## Vanilla equivalent Outside React, listen on the client and recreate it the same way: ```ts theme={null} import { LipsyncClient, RefusedError } from "@mascotbot/core"; let client = await LipsyncClient.init({ apiKey: "mascot_pub_…" }); function wire(c: LipsyncClient) { c.on("refused", async (_err) => { await c.close(); client = await LipsyncClient.init({ apiKey: "mascot_pub_…" }); wire(client); }); } wire(client); async function speak(audio: Float32Array) { try { return await client.processAudio(audio); } catch (err) { if (err instanceof RefusedError) { // The "refused" listener above is already re-initing; the next call // will land on the fresh client. Decide whether to retry now or on // the next user action. return null; } throw err; } } ``` The available events are `"ready" | "refused" | "error" | "refresh"` ([API conventions](/reference/api-conventions#1-events-vs-callbacks)). For recovery you only need `"refused"`. ## Tell the user what happened Reading `useMascot().error` gives you the typed cause when `status` is `"refused"` or `"error"`. Branch on `error.code`, not the class — that's the stable surface and the full code matrix lives in [Error codes](/reference/error-codes). For `RefusedError` in particular, the codes that map to "show a clear next step" UI are listed in that page's [branching example](/reference/error-codes#authorization-codes). For a plain network/sleep cutoff the user only needs to see "session expired — tap to continue"; the pre-check inside your action handler already does the rest on the next tap. ## Background tabs and mobile sleep — the practical caveat Most "the avatar stopped working when I came back" reports are not bugs — they're the device pausing the refresh loop long enough for the cutoff to fire. Two things make this a non-issue: 1. The recovery pattern above turns the next tap into a re-init. Users experience one slightly slower click, not a broken page. 2. If you can tell when the page becomes visible again (e.g. `document.visibilityState`), call `reload()` proactively so the session is warm before the user interacts. This is optional — the recovery pattern is already correct without it. ## Next <Columns> <Card title="Error codes" icon="triangle-exclamation" href="/reference/error-codes"> The full `RefusedError.code` matrix and recommended UI per code. </Card> <Card title="React hooks" icon="react" href="/libraries/react-hooks"> `useMascot` — the `status` / `error` / `reload` surface. </Card> <Card title="API conventions" icon="ruler-combined" href="/reference/api-conventions"> Events vs callbacks, the error taxonomy, module boundaries. </Card> </Columns> # Streaming & Microphone Lip Sync - useLipsyncStream Guide Source: https://docs.mascot.bot/libraries/streaming-and-mic Drive a Mascotbot avatar from live audio with useLipsyncStream: microphone input, tapping a played MediaStream, or pushing audio manually. One hook, three sources. `useLipsyncStream` is the single hook for **live** lip sync. It owns the audio graph, runs inference per window, and feeds visemes into a `MascotPlayback`. One hook, three input sources. ```ts theme={null} import { useLipsyncStream } from "@mascotbot/react/rive"; const { error, attached, pushAudio, pushBase64PCM16, reset } = useLipsyncStream({ client, // from useMascot() playback, // from useMascotPlayback({ stream: true }) source: { kind: "mic" }, // see sources below enabled: isLive, // optional gate onFrame: (f) => {}, // optional per-window telemetry debug: false, // optional verbose logging }); ``` Always create the playback with `stream: true` for live sources: ```ts theme={null} const playback = useMascotPlayback({ stream: true, enableNaturalLipSync: true }); ``` ## The three sources `source` is a discriminated union. ### `{ kind: "mic" }` The user's microphone. Optional `constraints?: MediaTrackConstraints`. ```tsx theme={null} "use client"; import { useEffect, useState } from "react"; import { useMascot } from "@mascotbot/react"; import { useMascotPlayback, useLipsyncStream } from "@mascotbot/react/rive"; function MicAvatar() { const { client, status } = useMascot(); const playback = useMascotPlayback({ stream: true, enableNaturalLipSync: true }); const [active, setActive] = useState(false); const isLive = active && status === "ready" && !!client; const { error } = useLipsyncStream({ client, playback, source: { kind: "mic" }, enabled: isLive, // gates getUserMedia + the audio graph without unmounting }); useEffect(() => { if (status !== "ready") setActive(false); }, [status]); return ( <> <button onClick={() => setActive((v) => !v)} disabled={status !== "ready"}> {active ? "Stop mic" : "Start mic"} </button> {error && <p>{error.message}</p>} </> ); } ``` The worklet outputs at zero gain, so the user's speakers do not echo the mic. ### `{ kind: "mediaStream", stream }` Tap any `MediaStream` — a played `<audio>`/`<video>` element, or a realtime AI provider's voice. The capture point is the playback point, so the mouth cannot drift ahead of the speech. Pass `null` to detach. ```tsx theme={null} import { createElementTap, type ElementTap } from "@mascotbot/react"; const audioRef = useRef<HTMLAudioElement>(null); const tapRef = useRef<ElementTap | null>(null); const [stream, setStream] = useState<MediaStream | null>(null); function play() { const el = audioRef.current; if (!el) return; if (!tapRef.current) { // create inside the click gesture tapRef.current = createElementTap(); setStream(tapRef.current.stream); } tapRef.current.attach(el); // idempotent el.play(); } useLipsyncStream({ client, playback, source: { kind: "mediaStream", stream } }); ``` `createElementTap()` is an SDK export (`@mascotbot/react`, re-exported from `lipsync-core`) — the cross-browser tap detailed in [Realtime providers → Tap a playing element](/realtime/overview#tap-a-playing-element) (it replaces `captureStream()`, which Safari does not implement). The same helper drives self-playing realtime providers (ElevenLabs, OpenAI Realtime over WebRTC) — see [Realtime providers](/realtime/overview). ### `{ kind: "manual" }` You push audio yourself with the returned functions — for providers that hand you raw PCM, or any custom source: ```ts theme={null} const { pushAudio, pushBase64PCM16, reset } = useLipsyncStream({ client, playback, source: { kind: "manual" }, }); await pushAudio(float32Samples, 24000); // Float32 [-1,1] + its sample rate await pushBase64PCM16(base64Chunk, 24000); // base64 PCM16 + its sample rate reset(); // drop buffered state (barge-in) ``` For raw-PCM realtime providers, pairing [`createPCMStreamPlayer`](/core/pcm-stream-player) with a `{ kind: "mediaStream" }` source is usually cleaner than manual pushing, because the player also handles gap-tolerant playback. ## Return value | Field | Type | Notes | | ----------------- | ----------------------------------------- | ----------------------------------- | | `error` | `Error \| null` | Stream/permission/inference error | | `attached` | `boolean` | The audio graph is live and tapping | | `pushAudio` | `(Float32Array, number) => Promise<void>` | Manual source only | | `pushBase64PCM16` | `(string, number) => Promise<void>` | Manual source only | | `reset` | `() => void` | Clears buffered stream state | ## Telemetry with `onFrame` `onFrame` fires once per inference window with per-frame detail — wire it to a debug HUD or metrics, not to your render path: ```ts theme={null} useLipsyncStream({ client, playback, source: { kind: "mic" }, onFrame: (f) => { // f.frameIndex, f.visemeId, f.silenceDetected, f.inferenceMs, // f.emittedVisemeId, f.overruns, f.audioContextRate }, }); ``` ## Lifecycle and stability `useLipsyncStream` keeps the audio graph alive across re-renders by routing `client`/`playback` through refs. Rive input handles are referentially stable and playback is carried across any internal recreate, so an unrelated parent re-render does not freeze the mouth. Only `enabled: false` (or a `null` `mediaStream`) tears the pipeline down — flip a button without unmounting the component. ## End-of-utterance silence is handled for you When speech stops, naive pipelines emit phantom mouth shapes during the inference tail. The SDK suppresses this with an internal **−50 dBFS input-amplitude silence gate** — you do **not** implement your own gate. The mouth simply settles to rest when the audio goes quiet, on every source. ## The audio worklet The worklet is embedded in the SDK and served from a Blob URL by default — there is no file to copy or host. Pass `workletUrl` only if your Content Security Policy forbids `worker-src blob:` or you want CDN caching. ## Next <Columns> <Card title="Realtime providers" icon="bolt" href="/realtime/overview"> Tap OpenAI / Gemini / ElevenLabs voices. </Card> <Card title="PCM stream player" icon="waveform" href="/core/pcm-stream-player"> Play + tap raw provider PCM. </Card> <Card title="Hooks reference" icon="react" href="/libraries/react-hooks"> Every hook at a glance. </Card> </Columns> # Meet Our Ready-to-Use Mascots: A Quick Start Guide Source: https://docs.mascot.bot/mascots/ready-to-use-mascots Below is a selection of pre-made mascots available to all subscribers. Feel free to use them for testing, early integration, or commercial projects while we craft your custom mascot. Each one offers unique personality traits, different visual styles, and multiple potential use cases. ## 1. Cat **Description**:\ A sleek, black feline with bright yellow eyes that radiates playful curiosity and cunning. **Potential Niches/Roles**: * Gaming or streaming channels seeking a witty sidekick * Tech-oriented brands wanting a clever problem-solver * Creative agencies looking for a charming but mischievous persona <iframe title="Cat Mascot Demo" /> *** ## 2. Panda **Description**:\ A friendly, futuristic panda with a sleek, robotic edge—ideal for conveying warmth and innovation. **Potential Niches/Roles**: * Family-friendly platforms or wellness apps in need of a kind, comforting figure * AI-focused businesses aiming for a blend of approachability and high-tech aesthetics * Educational tools requiring a gentle, nurturing presence <iframe title="Panda Mascot Demo" /> *** ## 3. Girl **Description**:\ A stylish, cyberpunk-inspired young woman who brings an edgy, confident energy to any project. **Potential Niches/Roles**: * Fashion or lifestyle brands emphasizing individuality and bold self-expression * Music or entertainment platforms seeking a vibrant persona with a rebellious streak * Futuristic or urban-themed products wanting a trendsetting spokesperson <iframe title="Girl Mascot Demo" /> *** ## 4. Robot **Description**:\ A nostalgic-yet-modern floating robot with an expressive screen, merging retro charm with futuristic flair. **Potential Niches/Roles**: * Tech startups wanting a friendly, mechanical guide * Gaming or sci-fi platforms seeking a quirky, robotic companion * E-learning or customer support applications needing a helpful digital assistant <iframe title="Cat Mascot Demo" /> *** ## Using a ready-made mascot Every mascot above ships as a Rive file with the artboard `Character` and the `mascotStateMachine` state machine the SDK expects — drop one into `<Mascot src>` and it lip-syncs with no extra setup: ```tsx theme={null} import { MascotProvider } from "@mascotbot/react"; import { Mascot, MascotRive } from "@mascotbot/react/rive"; <MascotProvider apiKey="mascot_pub_…"> <MascotProvider> <Mascot src="/cat.riv"> <MascotRive /> </Mascot> </MascotProvider> </MascotProvider> ``` The SDK only animates the mouth, `is_speaking`, and `stress` — every other input on these files (gestures, expressions, scene state) stays yours to drive directly. New here? Start with the [Quickstart](/quickstart), then the [React SDK](/libraries/react-sdk). # Mascotbot Avatar SDK Overview - Build Interactive AI Avatars Source: https://docs.mascot.bot/overview What the Mascotbot avatar SDK is: real-time interactive AI avatars for React and JavaScript — offline, microphone, and realtime voice-agent paths. The Mascotbot avatar SDK turns speech into a real-time talking avatar. It is a small, composable, low-level surface — **audio in → a serializable viseme timeline → a thin Rive playback layer** — backed by the licensed model and asset delivery. It does not ship a call UI, a TTS engine, or provider glue; those are recipes you compose, not framework you adopt. ## How it works <Steps> <Step title="Authorize"> `MascotProvider` (or `LipsyncClient.init`) exchanges your API key with the edge worker, which returns a short-lived license and the WASM runtime. Sessions auto-refresh in the background. </Step> <Step title="Process speech"> You hand the SDK 16 kHz mono audio — a recorded buffer, microphone windows, or a tapped `MediaStream`. The SDK produces a viseme id per 10 ms frame. </Step> <Step title="Animate"> Visemes drive the mouth inputs of a Rive avatar through the playback engine. The SDK writes only the mouth, `is_speaking`, and `stress` — everything else on the Rive file stays yours. </Step> </Steps> ## Packages The SDK is two packages, each with a root and a `/rive` subpath. Import the narrowest one for your use case. | Import | What it is | | ----------------------- | ---------------------------------------------------------------------------------------------------------- | | `@mascotbot/core` | Engine + offline `VisemeTimeline` + `createPCMStreamPlayer`. Framework-agnostic, no Rive, no React. | | `@mascotbot/core/rive` | Framework-agnostic Rive playback (`MascotPlayback`, `getRiveInputs`, `hasRiveInput`). | | `@mascotbot/react` | React provider + `useMascot` / `useProcessAudio`. | | `@mascotbot/react/rive` | React Rive layer: `<Mascot>`, `useMascotRive`, `useMascotInputs`, `useMascotPlayback`, `useLipsyncStream`. | Most React apps use `@mascotbot/react` + `@mascotbot/react/rive`. The `/rive` subpaths take `@rive-app/webgl2` (and `@rive-app/react-webgl2` for React) as an **optional peer dependency** — install it only if you render an avatar. ## Three integration paths <Columns> <Card title="Offline" icon="box-archive" href="/libraries/offline-lipsync"> Run inference once, persist the timeline as JSON, replay forever with zero reprocessing. </Card> <Card title="Microphone & streaming" icon="microphone" href="/libraries/streaming-and-mic"> Drive the avatar live from the user's mic, a tapped `MediaStream`, or manually pushed audio. </Card> <Card title="Realtime AI" icon="bolt" href="/realtime/overview"> Connect OpenAI Realtime, Gemini Live, or ElevenLabs by tapping the assistant's voice in real time. </Card> </Columns> All three end at the same place: a `MascotPlayback` instance driven by either a `VisemeTimeline` (offline) or a live audio source. There is no separate API to learn per path. ## What the SDK does and does not do The SDK writes exactly three Rive input families: mouth visemes (`100..118`), `is_speaking`, and `stress`. Every other state-machine input, data-binding ViewModel, event, and listener on the Rive instance is yours, accessed directly on the raw `rive` object. The SDK never wraps, gates, or proxies it. This contract is detailed in [Rive co-existence](/concepts/rive-coexistence). The SDK is intentionally minimal — audio in, animation out. Upgrading an existing integration? The [migration guide](/reference/migration) maps every change. ## Browser support * **Chrome / Edge** — full. * **Safari** (desktop 17+, iOS 17+) — full; WebAssembly + WebGL2 required for the Rive avatar. * **Firefox** — audio pipeline supported; the Rive renderer requires WebGL2. The SDK refuses to run if WebAssembly or `crypto.subtle` are unavailable. ## Next <Columns> <Card title="Installation" icon="download" href="/installation"> Private registry, keys, peer deps. </Card> <Card title="Quickstart" icon="rocket" href="/quickstart"> A working avatar in a few lines. </Card> <Card title="Visemes & the timeline" icon="waveform-lines" href="/concepts/visemes-and-timeline"> The core data model. </Card> </Columns> # Mascotbot Lipsync SDK Quickstart - Animate a Rive Avatar in Minutes Source: https://docs.mascot.bot/quickstart Wire the Mascotbot lipsync SDK into a React app: provider setup, drive a Rive avatar from an audio file, add live microphone input, or use the vanilla core without React. This gets you from an installed package to a talking avatar. If you have not installed yet, do [Installation](/installation) first. ## 1. Mount the provider `<MascotProvider>` initializes a single `LipsyncClient` for your app and exposes it through context. ```tsx theme={null} // app/layout.tsx (or wherever you mount providers) "use client"; import { MascotProvider } from "@mascotbot/react"; export default function Layout({ children }: { children: React.ReactNode }) { return <MascotProvider apiKey="mascot_pub_…">{children}</MascotProvider>; } ``` Anywhere inside it, `useMascot()` gives you the client and status, and `useProcessAudio()` fetches a URL, decodes, resamples to 16 kHz, and runs inference once: ```tsx theme={null} "use client"; import { useMascot, useProcessAudio } from "@mascotbot/react"; export function Demo() { const { status, error } = useMascot(); const { result, loading } = useProcessAudio("/audio/greeting.wav"); if (status === "initializing") return <p>Loading SDK…</p>; if (status === "error") return <p>{error?.message}</p>; if (loading || !result) return <p>Processing audio…</p>; // result.timeline — serializable viseme timeline (the offline artifact) // result.durationMs — total audio duration // result.speechMs — non-silent ms detected return <pre>{JSON.stringify(result.timeline, null, 2)}</pre>; } ``` `result.timeline` is a [`VisemeTimeline`](/concepts/visemes-and-timeline) — hand it to playback below, or `JSON.stringify` it to persist and replay later with zero reprocessing. ## 2. Drive a Rive avatar Wire the timeline into a Rive avatar's mouth state machine with the `/rive` subpath: ```tsx theme={null} "use client"; import { MascotProvider, useMascot, useProcessAudio } from "@mascotbot/react"; import { MascotProvider, Mascot, MascotRive, useMascotPlayback, Fit, Alignment, } from "@mascotbot/react/rive"; function App() { return ( <MascotProvider apiKey="mascot_pub_…"> <MascotProvider> <Mascot src="/mascot-fox.riv" layout={{ fit: Fit.Contain, alignment: Alignment.Center }} > <MascotRive /> <SamplePlayer /> </Mascot> </MascotProvider> </MascotProvider> ); } function SamplePlayer() { const { status } = useMascot(); const { result } = useProcessAudio("/audio/greeting.wav"); const playback = useMascotPlayback({ enableNaturalLipSync: true }); function play() { if (status !== "ready" || !result) return; new Audio("/audio/greeting.wav").play().catch(() => {}); playback.setTimeline(result.timeline); // offline timeline → mouth playback.play(); } return ( <button onClick={play} disabled={status !== "ready" || !result}> Play </button> ); } ``` <Tip> `useProcessAudio` runs inference once. Persist `result.timeline` (it is plain JSON) and on later loads skip decode + inference entirely: `playback.setTimeline(parseTimeline(JSON.parse(stored)))`. See [Offline lip sync](/libraries/offline-lipsync). </Tip> ### Rive avatar requirements If you author your own `.riv` file: | Element | Requirement | | ------------------------- | -------------------------------------- | | Artboard name | `Character` | | State machine name | `mascotStateMachine` | | Mouth inputs | Number inputs `100`–`118` (viseme ids) | | Emotion inputs (optional) | `is_speaking`, `eyes_smile` | | Stress input (optional) | `stress` (number) | The SDK writes only those inputs. Any other input, data binding, or event on the file is yours to drive on the raw `rive` instance (`useMascotRive().rive`) — see [Rive co-existence](/concepts/rive-coexistence). Or skip authoring entirely and use a [ready-made mascot](/mascots/ready-to-use-mascots). ## 3. Live microphone input Drive the avatar from the user's microphone in real time with `useLipsyncStream`: ```tsx theme={null} "use client"; import { useEffect, useState } from "react"; import { useMascot } from "@mascotbot/react"; import { useMascotPlayback, useLipsyncStream } from "@mascotbot/react/rive"; function MicAvatar() { const { client, status } = useMascot(); const playback = useMascotPlayback({ stream: true, enableNaturalLipSync: true }); const [active, setActive] = useState(false); const isLive = active && status === "ready" && !!client; const { error } = useLipsyncStream({ client, playback, source: { kind: "mic" }, enabled: isLive, // gates getUserMedia + the audio graph without unmount }); useEffect(() => { if (status !== "ready") setActive(false); }, [status]); return ( <> <button onClick={() => setActive((v) => !v)} disabled={status !== "ready"}> {active ? "Stop mic" : "Start mic"} </button> {error && <p>{error.message}</p>} </> ); } ``` The audio worklet is embedded in the SDK and served from a Blob URL by default — there is no file to copy. Pass `workletUrl` only if your CSP forbids `worker-src blob:`. The same hook handles realtime AI providers via `source: { kind: "mediaStream", stream }` — see [Realtime providers](/realtime/overview). ## 4. Vanilla JavaScript Not using React? Use `@mascotbot/core` directly: ```ts theme={null} import { LipsyncClient, parseTimeline } from "@mascotbot/core"; const client = await LipsyncClient.init({ apiKey: "mascot_pub_…", userId: "user_42", // optional, for accurate usage attribution }); // Pre-recorded audio (16 kHz mono Float32 in [-1, 1]) const { timeline } = await client.processAudio(audioBuffer); localStorage.setItem("greeting.vtl", JSON.stringify(timeline)); // persist // later: parseTimeline(JSON.parse(localStorage.getItem("greeting.vtl")!)) // Live streaming: push 25 ms (400-sample) windows one at a time const session = client.createStreamingSession(); const frame = await session.pushWindow(audioWindow); console.log(frame.visemeId, frame.silenceDetected); await client.stop(); // release resources when done ``` For the Rive engine without React, import from `@mascotbot/core/rive` — see [Core SDK](/core/client). ## Error handling `useMascot()` exposes `status` and `error`. Branch on `error.code`, not the subclass: ```tsx theme={null} import { RefusedError, NetworkError, EngineError } from "@mascotbot/react"; function StatusUI() { const { status, error } = useMascot(); if (status !== "error" && status !== "refused") return null; if (error instanceof RefusedError) { if (error.code === "key_disabled") return <ReSubscribe />; if (error.code === "dev_key_on_public_domain") return <UseProdKey />; } if (error instanceof NetworkError) return <p>Network issue — retry</p>; if (error instanceof EngineError) return <p>Inference error — refresh</p>; return <p>{error?.message}</p>; } ``` Full matrix: [Error codes](/reference/error-codes). ## Next <Columns> <Card title="Offline lip sync" icon="box-archive" href="/libraries/offline-lipsync"> Generate → persist → replay. </Card> <Card title="React hooks" icon="react" href="/libraries/react-hooks"> The full hook reference. </Card> <Card title="Realtime providers" icon="bolt" href="/realtime/overview"> OpenAI, Gemini, ElevenLabs. </Card> </Columns> # Realtime AI Voice Avatars - Lip Sync for OpenAI, Gemini & ElevenLabs Source: https://docs.mascot.bot/realtime/overview Add a real-time avatar to any voice AI: tap the assistant's audio into useLipsyncStream. OpenAI Realtime, Gemini Live, ElevenLabs. Adding a lip-synced avatar to a realtime voice assistant is one idea: > Give the SDK a `MediaStream` of the assistant's voice. It turns that > audio into a talking avatar in real time. ```ts theme={null} useLipsyncStream({ client, playback, source: { kind: "mediaStream", stream } }); ``` You wire the provider with *its own* official SDK and the SDK lip-syncs the audio in real time. The only question is how you obtain that `MediaStream`, which depends on whether the provider plays the audio for you. ## Pick the path by provider | Provider / transport | Plays audio itself? | How you get the stream | SDK piece | | -------------------------------- | ---------------------------------- | ---------------------------------------------- | ------------------------------------- | | **OpenAI Realtime — WebRTC** | Yes (into an `<audio>` you supply) | [`createElementTap()`](#tap-a-playing-element) | none | | **Gemini Live** | No (raw base64 PCM16 @ 24 kHz) | `createPCMStreamPlayer().outputStream` | [PCM player](/core/pcm-stream-player) | | **OpenAI Realtime — WebSocket** | No (raw PCM16 `ArrayBuffer`) | `createPCMStreamPlayer().outputStream` | [PCM player](/core/pcm-stream-player) | | **ElevenLabs Conversational AI** | Yes (internal worklet → `<audio>`) | [`createElementTap()`](#tap-a-playing-element) | none | <Warning> Never route a self-playing provider (ElevenLabs, OpenAI-WebRTC) through `createPCMStreamPlayer` — the voice would play twice. The player is **only** for providers that hand you raw PCM and do not play it. </Warning> ## Tap a playing element When a provider plays the audio itself (OpenAI WebRTC, ElevenLabs), tap the element it plays through with **`createElementTap()`** — an SDK helper that works in Chrome, Safari and Firefox (`HTMLMediaElement.captureStream()` is not implemented in Safari/WebKit, so the SDK does not use it): ```ts theme={null} import { createElementTap } from "@mascotbot/react"; // Create inside the click that starts the call (so its AudioContext isn't // born suspended). `stream` is usable immediately — silent until attach(). const tap = createElementTap(); useLipsyncStream({ client, playback, source: { kind: "mediaStream", stream: tap.stream }, }); tap.attach(audioEl); // now, or later once the provider's <audio> exists // teardown: tap.close(); ``` `tap.attach(el)` handles both element kinds: a file/URL `<audio>` (e.g. OpenAI WebRTC) is kept audible **and** tapped; an element whose `srcObject` is a `MediaStream` (e.g. ElevenLabs' worklet output) is tapped **only**, so the provider's own playback isn't doubled. `attach()` is idempotent and may run after creation. Also exported from `@mascotbot/core`. ## Provider tokens stay on the server Never ship a standing provider key to the browser. A server route mints a short-lived credential per session; the client connects with that. The [demo](/overview) ships reference route handlers for all three: | Provider | Server mints | Notes | | ---------- | ---------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- | | OpenAI | `POST https://api.openai.com/v1/realtime/client_secrets` → `clientSecret` | model `gpt-realtime` | | Gemini | `@google/genai` `ai.authTokens.create(...)` → ephemeral `token.name` | model `models/gemini-3.1-flash-live-preview`, `apiVersion: "v1alpha"` | | ElevenLabs | `GET https://api.elevenlabs.io/v1/convai/conversation/get-signed-url?agent_id=…` → `signedUrl` | `xi-api-key` header, server-side only | ## Path 1 — OpenAI Realtime (WebRTC), cleanest The provider auto-plays into an `<audio>` element. Supply your own so you can tap it; no SDK audio piece needed. ```ts theme={null} import { RealtimeAgent, RealtimeSession } from "@openai/agents-realtime"; const audioEl = new Audio(); const session = new RealtimeSession(new RealtimeAgent({ name: "Assistant" }), { transport: "webrtc" }); await session.connect({ apiKey: clientSecret, audioElement: audioEl }); // clientSecret from your server route const tap = createElementTap(); // create in the click — see "Tap a playing element" useLipsyncStream({ client, playback, source: { kind: "mediaStream", stream: tap.stream } }); tap.attach(audioEl); ``` ## Path 2 — Gemini Live / OpenAI Realtime (WebSocket) The provider streams raw PCM and does not play it. [`createPCMStreamPlayer`](/core/pcm-stream-player) plays it gap-tolerantly and exposes the same audio as a tappable `MediaStream`. ```ts theme={null} import { createPCMStreamPlayer } from "@mascotbot/core"; const player = createPCMStreamPlayer({ sampleRate: 24000 }); // both emit 24 kHz useLipsyncStream({ client, playback, source: { kind: "mediaStream", stream: player.outputStream } }); // Gemini Live (@google/genai): modelTurn audio part session.onmessage = (m) => { const b64 = m?.serverContent?.modelTurn?.parts?.[0]?.inlineData?.data; if (typeof b64 === "string") player.pushBase64PCM16(b64); if (m?.serverContent?.interrupted) player.stop(); }; // OpenAI Realtime (WebSocket): PCM16 ArrayBuffer session.on("audio", (e) => player.pushPCM16(new Uint8Array(e.data))); session.on("audio_interrupted", () => player.stop()); ``` The transport parsing above is provider glue and stays in your app — it must not enter the SDK. Only the play-and-tap primitive is shared. ## Path 3 — ElevenLabs Conversational AI ElevenLabs plays its assistant audio through a hidden `<audio>` whose `srcObject` is a `MediaStream` (an internal worklet → `MediaStreamDestination`). Create [`createElementTap()`](#tap-a-playing-element) in the click, patch `window.Audio` **before** `Conversation.startSession` to capture the element, then `tap.attach(el)` once its `srcObject` is set — the `srcObject` branch taps without re-outputting, so ElevenLabs' own playback is not doubled: ```ts theme={null} const tap = createElementTap(); // in the click, before startSession setStream(tap.stream); // → useLipsyncStream source: { kind: "mediaStream", stream } const w = window as unknown as { Audio: typeof Audio; __el?: HTMLAudioElement }; const Orig = w.Audio; w.Audio = function (...a: unknown[]) { const el = new Orig(...(a as [])); w.__el = el; return el; } as unknown as typeof Audio; const { Conversation } = await import("@elevenlabs/client"); await Conversation.startSession({ signedUrl }); const iv = setInterval(() => { const el = w.__el; if (el?.srcObject instanceof MediaStream) { clearInterval(iv); w.Audio = Orig; tap.attach(el); // srcObject branch → tap only; stays audible } }, 100); // teardown: tap.close(); ``` ## Server TTS For plain TTS, the server route returns **audio only** (base64 PCM16). The client plays it through `createPCMStreamPlayer` and the tap drives the mouth. The server only synthesizes speech; it never computes or streams visemes. ## End-of-utterance silence The SDK's internal **−50 dBFS silence gate** suppresses the phantom mouth shapes that appear when the assistant stops talking. You do not implement your own gate — this is handled for every realtime path. ## Stress emphasis and gestures `stress` and `gesture` add body to a talking avatar. There is no flag to "enable" them — you drive them, and each needs the matching input declared on the `.riv` (the [ready-made mascots](/mascots/ready-to-use-mascots) include `stress`). They work the same for every realtime provider; the natural trigger is **speech onset**, which `useLipsyncStream`'s `onFrame` gives you. ### `stress` — built-in emphasis `stress` is one of the three input families the SDK drives (mouth, `is_speaking`, `stress`). `useMascotPlayback()` returns a **`stress()`** method: you push emphasis cues `{ offset, stress }` and the SDK eases the Rive `stress` input toward each target. `offset` is ms on the playback clock; cues are applied in order, and a cue whose `offset` has already passed is applied on the next frame — so **`offset: 0` means "apply now"**. That makes the realtime pattern trivial: raise stress while the assistant speaks, drop it when it stops. ### `gesture` — your own one-shot trigger `gesture` is a consumer-owned input — the SDK never touches it. If your `.riv` declares one, fire it yourself with `useMascotInputs()`. `has("gesture")` confirms the input exists; `custom.gesture.fire?.()` triggers it (the optional-call form also tolerates a numeric `gesture` input). ### Wiring both for ElevenLabs (or any provider) `playback` must be created with `stream: true` for realtime. This drives `stress` on the speech envelope and fires `gesture` once per utterance — identical for OpenAI / Gemini, only the `stream` source differs: ```tsx theme={null} import { useRef } from "react"; import { useMascot } from "@mascotbot/react"; import { useMascotPlayback, useMascotInputs, useLipsyncStream } from "@mascotbot/react/rive"; function AvatarReactions({ stream }: { stream: MediaStream | null }) { const { client } = useMascot(); const playback = useMascotPlayback({ stream: true, enableNaturalLipSync: true }); const { has, custom } = useMascotInputs(); const speaking = useRef(false); useLipsyncStream({ client, playback, source: { kind: "mediaStream", stream }, // createElementTap() for ElevenLabs, or player.outputStream onFrame: (f) => { const isSpeech = !f.silenceDetected; if (isSpeech && !speaking.current) { speaking.current = true; playback.stress([{ offset: 0, stress: 1 }]); // emphasize while speaking if (has("gesture")) custom.gesture.fire?.(); // one-shot reaction } else if (!isSpeech && speaking.current) { speaking.current = false; playback.stress([{ offset: 0, stress: 0 }]); // ease back to neutral } }, }); return null; } ``` For a single emphasis bump instead of a sustained one, push `stress: 1` then release after a hold: `setTimeout(() => playback.stress([{ offset: 0, stress: 0 }]), 350)`. For offline playback the same `playback.stress([...])` works with real timeline offsets (e.g. `{ offset: 0, stress: 1 }`, `{ offset: 400, stress: 0 }`). `reset()` clears scheduled stress with the rest of playback. The timeline JSON does not carry stress — you always schedule it separately. See [Rive co-existence](/concepts/rive-coexistence) for why `gesture` is yours and `stress` is SDK-driven. ## Provider guides <Columns> <Card title="ElevenLabs avatar" icon="microphone-lines" href="/libraries/elevenlabs-avatar"> Conversational AI avatar. </Card> <Card title="Gemini Live avatar" icon="google" href="/libraries/gemini-live-api-avatar"> Gemini Live API avatar. </Card> <Card title="OpenAI Realtime avatar" icon="robot" href="/libraries/openai-realtime-api-avatar"> ChatGPT Realtime avatar. </Card> </Columns> ## Next <Columns> <Card title="PCM stream player" icon="waveform" href="/core/pcm-stream-player"> Play + tap raw PCM. </Card> <Card title="Streaming & mic" icon="microphone" href="/libraries/streaming-and-mic"> `useLipsyncStream` in depth. </Card> </Columns> # Mascotbot SDK API Conventions - Predictable, Hard-to-Misuse Design Source: https://docs.mascot.bot/reference/api-conventions The rules the Mascotbot lipsync SDK follows so it stays predictable: events vs callbacks, options objects, honest async, the error taxonomy, format versioning, and module boundaries. The SDK surface follows a small set of rules so it stays predictable and hard to misuse. Knowing them makes the whole API guessable. ## 1. Events vs. callbacks * **Lifecycle / multi-fire → an emitter.** `client.on("ready" | "refused" | "error" | "refresh", fn)` returns an unsubscribe function. Use it for things that happen repeatedly or that multiple listeners care about. * **One-shot wiring / per-item telemetry → an option callback.** e.g. `useLipsyncStream({ onFrame })`. Use it to configure one thing at construction. There is never a second mechanism for the same concern (no `onReady` prop when `client.on("ready")` exists). ## 2. Options object, not positional New hooks and functions take a **single options object** — e.g. `useLoadRive({ stateMachineName, ... })`. No public function takes multiple positional arguments where an options object would do. ## 3. Async is honest If a value is computed off-thread (worker, network), the method is `async` and stays `async` — `client.diagnostics()`, `session.pushWindow()`. Nothing hides a Promise behind a sync-looking API, so you always know where to `await`. ```ts theme={null} const d = await client.diagnostics(); // off-thread → async const frame = await session.pushWindow(window); // inference → async ``` ## 4. Error taxonomy Five classes — `LipsyncError` (base) plus `License` / `Network` / `Engine` / `RefusedError` — mapped to failure **domains**. You branch on `.code`, not the subclass. A new failure that is not a new domain reuses `LipsyncError` with a new `.code` (e.g. `bad_timeline`) rather than adding a subclass per code. Every code is registered in [Error codes](/reference/error-codes). ## 5. Serialized-format versioning Anything you can persist and feed back later carries an explicit `version` and a validating parser that rejects mismatches loudly. `VisemeTimeline` has `version: VISEME_TIMELINE_VERSION` plus `frameMs`, and `parseTimeline` throws `LipsyncError("bad_timeline", …)` on any incompatibility. The version is bumped on any breaking shape or semantics change; an old shape is never silently accepted. Always load persisted data through the validating parser, never `JSON.parse` alone. See [the timeline model](/concepts/visemes-and-timeline). ## 6. Module boundaries Import the narrowest entry point for the job: | Entry | Contains | Excludes | | ----------------------- | ----------------------------------------------------------- | ------------------------------------------------ | | `@mascotbot/core` | Engine, `VisemeTimeline` + helpers, `createPCMStreamPlayer` | No Rive, no React | | `@mascotbot/core/rive` | `MascotPlayback`, `getRiveInputs`, `hasRiveInput` | No React; `@rive-app/webgl2` is an optional peer | | `@mascotbot/react` | `MascotProvider`, `useMascot`, `useProcessAudio` | No Rive | | `@mascotbot/react/rive` | The React Rive layer | — | A Rive type never appears on the core root entry, and core never depends on React. The split keeps bundles minimal and the dependency graph honest. ## 7. The SDK stays out of your Rive instance The SDK writes exactly three input families — mouth visemes, `is_speaking`, `stress` — and nothing else. Every other Rive capability is reached directly on the raw instance. This is a hard contract, documented in [Rive co-existence](/concepts/rive-coexistence): the SDK never wraps, gates, proxies, or constrains Rive. ## Next <Columns> <Card title="Error codes" icon="triangle-exclamation" href="/reference/error-codes"> The full code matrix. </Card> <Card title="Rive co-existence" icon="puzzle-piece" href="/concepts/rive-coexistence"> The Rive ownership contract. </Card> <Card title="Migration" icon="arrow-right-arrow-left" href="/reference/migration"> The 0.2.x symbol map. </Card> </Columns> # Mascotbot SDK Error Codes - License, Network & Timeline Failures Source: https://docs.mascot.bot/reference/error-codes Every Mascotbot lipsync SDK error code, the HTTP status it maps from, the customer-readable message, and the recommended UI. Branch on error.code, not the subclass. Every authorization failure carries a semantic `code` and an actionable `message`. The SDK surfaces them as typed errors so your app can show the right next step ("re-subscribe", "update card", "use a dev key") instead of a flat 401. ## The taxonomy Five classes, mapped to failure **domains**: | Class | Domain | | -------------- | --------------------------------------------------------------------- | | `LipsyncError` | Base. Also used directly for client-side codes (e.g. `bad_timeline`). | | `LicenseError` | Authorization at `init`. | | `RefusedError` | Hard refusal (init or refresh) — not retryable. | | `NetworkError` | Transport failure reaching the edge service. | | `EngineError` | Inference/runtime failure. | <Tip> **Branch on `error.code`, not the subclass.** Codes are stable; the class is just the domain bucket. </Tip> ```tsx theme={null} import { RefusedError, NetworkError, EngineError } from "@mascotbot/react"; function StatusUI() { const { status, error } = useMascot(); if (status !== "error" && status !== "refused") return null; if (error instanceof RefusedError) { switch (error.code) { case "key_disabled": case "subscription_canceled": return <ReSubscribe />; case "subscription_past_due_expired": return <UpdateCard />; case "key_expired": return <RotateKey />; case "invalid_api_key": return <CheckKey />; case "prod_key_on_localhost": return <UseDevKey />; case "dev_key_on_public_domain": return <UseProdKey />; case "origin_not_allowed": return <CheckAllowlist />; case "session_expired": return <ReloadPrompt />; } } if (error instanceof NetworkError) return <p>Network issue — retry</p>; if (error instanceof EngineError) return <p>Inference error — refresh</p>; return <p>{error?.message}</p>; } ``` ## Authorization codes The edge worker verifies the key and maps the rejection to one envelope. | HTTP | `code` | Fires when | Recommended UI | | ---- | ------------------------------- | ---------------------------------------------------------------- | -------------------------------------------- | | 401 | `missing_bearer` | No `Authorization` header. | Fix the integration (key not sent). | | 401 | `invalid_api_key` | Key not recognized — typo'd or never existed. | "Check the key at app.mascot.bot/api-keys." | | 401 | `key_expired` | Dev keys auto-expire 30 days after creation. | "Mint a fresh key." | | 402 | `key_disabled` | Key disabled — canonical canceled/revoked path. | "Re-subscribe, or mint a new key." | | 402 | `subscription_past_due_expired` | Payment failed and the grace period ran out. | "Update your card." | | 402 | `subscription_canceled` | Subscription canceled (identity preserved). | "Re-subscribe." | | 403 | `wrong_key_scope` | Key used on a surface it is not authorized for. | "Issue the matching key type." | | 403 | `prod_key_on_localhost` | Production key from `localhost`. | "Use a development key locally." | | 403 | `dev_key_on_public_domain` | Development key from a public origin. | "Use a production publishable key." | | 403 | `origin_not_allowed` | Production key + origin not on the allow-list. | "Add the origin at app.mascot.bot/security." | | 403 | `origin_allowlist_empty` | Production key with no origins configured. | "Configure the allow-list." | | 403 | `missing_origin` | Production key with no `Origin` header. | Fix the request. | | 429 | `rate_limited` | Key verification throttled. Transient. | "Try again in a few seconds." | | 404 | `session_expired` | The referenced session is gone (refresh only). **Hard refusal.** | "Reload the page to start a fresh session." | `session_expired` realistically happens when a tab is backgrounded long enough that throttled refresh ticks let the session lapse. It is never recoverable by retry — only a fresh init (a reload) recovers, so prompt the user immediately. ## Client-side codes (no network) Some `LipsyncError`s are thrown entirely client-side and still carry `.code`: | Class | `code` | Fires when | Treat as | | -------------- | -------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------- | | `LipsyncError` | `bad_timeline` | `parseTimeline(input)` rejected a persisted/untrusted timeline — version mismatch, non-positive `frameMs`, non-monotonic cue offsets, missing leading `t: 0`, out-of-range viseme id, or wrong field types. | "Regenerate via `client.processAudio()`" — **not** a license/network condition. | ```ts theme={null} import { parseTimeline, LipsyncError } from "@mascotbot/core"; try { playback.setTimeline(parseTimeline(JSON.parse(stored))); } catch (err) { if (err instanceof LipsyncError && err.code === "bad_timeline") { // re-run processAudio() and re-persist } } ``` See [Offline lip sync](/libraries/offline-lipsync) and [the timeline model](/concepts/visemes-and-timeline). ## Wire format Every error response is JSON with `Content-Type: application/json`; the HTTP status is the envelope status: ```json theme={null} { "code": "key_disabled", "message": "This API key has been disabled. …" } ``` The SDK parses this into the typed error's `.code` / `.message`. When a body is not JSON (5xx, network blip) the SDK synthesizes a status-derived envelope, so `.code` is always present. ## Next <Columns> <Card title="Licensing & keys" icon="key" href="/concepts/licensing-and-keys"> Why these refusals happen. </Card> <Card title="API conventions" icon="ruler-combined" href="/reference/api-conventions"> Why the taxonomy is shaped this way. </Card> <Card title="Troubleshooting" icon="wrench" href="/libraries/react-troubleshooting"> Fixing them in practice. </Card> </Columns> # Migrating the Mascotbot Lipsync SDK Source: https://docs.mascot.bot/reference/migration How to move existing integrations onto the current Mascotbot lipsync SDK — the 0.2.x → 0.3.0 unified-naming cutover, plus the legacy → 0.2.x architectural migration for older integrations. Two migrations live here: 1. **0.2.x → 0.3.0** — current pre-release cutover. Package scope rename * symbol consolidation (single Provider, single mount component). Most readers want this. 2. **Legacy → 0.2.x** — older architectural migration (server-side visemes → on-device WebAssembly). Kept for integrators on the pre-0.2 SDK. *** # 0.2.x → 0.3.0 The 0.3.0 release unifies SDK naming across the package scope and the React component surface. **Hard cutover** — no deprecated aliases. The bytes computing visemes are unchanged; only names and the mount shape moved. ## Package scope ```ini .npmrc theme={null} # 0.2.x @mascotbot-sdk:registry=https://npm.mascot.bot/ # 0.3.0 @mascotbot:registry=https://npm.mascot.bot/ ``` ```diff theme={null} # package.json - "@mascotbot-sdk/lipsync-core": "^0.2.9", - "@mascotbot-sdk/lipsync-react": "^0.2.9", + "@mascotbot/core": "^0.3.0", + "@mascotbot/react": "^0.3.0", ``` The same renaming applies to `lipsync-native` → `@mascotbot/native` and `lipsync-react-native` → `@mascotbot/react-native` for RN integrators. ## React mount — one Provider, one component ```diff theme={null} - import { LipsyncProvider, useLipsync } from "@mascotbot-sdk/lipsync-react"; - import { - MascotProvider, MascotClient, MascotRive, - useMascotPlayback, useMascotInputs, - } from "@mascotbot-sdk/lipsync-react/rive"; + import { MascotProvider, useMascot } from "@mascotbot/react"; + import { + Mascot, MascotRive, + useMascotPlayback, useMascotInputs, + } from "@mascotbot/react/rive"; - <LipsyncProvider apiKey="mascot_pub_…"> - <MascotProvider> - <MascotClient src="/avatar.riv"> - <MascotRive /> - </MascotClient> - </MascotProvider> - </LipsyncProvider> + <MascotProvider apiKey="mascot_pub_…"> + <Mascot src="/avatar.riv" /> + </MascotProvider> ``` If you previously wrapped `<MascotRive />` (or a loading slot) inside `<MascotClient>`, pass those as children of `<Mascot>` instead — they take over from the default canvas: ```tsx theme={null} <MascotProvider apiKey="mascot_pub_…"> <Mascot src="/avatar.riv" inputs={["wave"]}> <div className="my-layout-wrapper"> <MascotRive /> <MyCustomOverlay /> </div> </Mascot> </MascotProvider> ``` ## Symbol rename — quick reference | 0.2.x | 0.3.0 | Notes | | ----------------------------------------------------------------------- | ----------------------------------------------------- | -------------------------------------------------------------------------------------- | | `<LipsyncProvider>` | `<MascotProvider>` | Top-level provider; takes `apiKey` | | `<MascotProvider>` (empty Rive marker) | **gone** | Merged into the top-level | | `<MascotClient>` + `<MascotRive />` | `<Mascot>` | Single component; renders canvas by default. `<MascotRive />` stays as an escape hatch | | `useLipsync()` | `useMascot()` | Same return shape | | `MascotLipsyncClient` | `LipsyncClient` | Class — package scope provides brand | | `MascotLipsyncConfig` | `LipsyncConfig` | Type | | `LipsyncContextValue` | `MascotContextValue` | Type returned by `useMascot()` | | `MascotClientContext` / `MascotClientContextType` / `MascotClientProps` | `MascotContext` / `MascotContextType` / `MascotProps` | Component renamed → matching types | **Unchanged on purpose** (feature-descriptive — `lipsync` here names what the thing is, not who ships it): * `useProcessAudio(url)` * `useLipsyncStream({ source })` * `useMascotRive()` / `useMascotInputs()` / `useMascotPlayback()` * `LipsyncError` / `LipsyncStatus` / `LipsyncLogger` * `LipsyncStreamSource` / `LipsyncStreamFrame` * `LicenseError` / `NetworkError` / `EngineError` / `RefusedError` * `VisemeTimeline` / `VisemeCue` / `MascotPlayback` (the class) / `NaturalLipSync*` ## Migration steps (0.2.x → 0.3.0) <Steps> <Step title="Update .npmrc + package.json"> Change `@mascotbot-sdk:registry=…` to `@mascotbot:registry=…`, then rewrite the four dep names. `pnpm install --force` after — pnpm caches tarballs by hash, the new scope needs a fresh resolve. </Step> <Step title="Run the rename grep"> ```bash theme={null} git grep -nE '\b(LipsyncProvider|MascotClient|useLipsync|MascotLipsyncClient|MascotLipsyncConfig|LipsyncProviderProps|LipsyncContextValue|MascotClientContext|MascotClientContextType|MascotClientProps)\b' ``` Every hit needs the substitution in the table above. `useLipsyncStream`, `LipsyncError`, etc. are intentionally preserved — they survive because the word `lipsync` precisely describes what they are. </Step> <Step title="Collapse the Provider + Mascot mount"> `<LipsyncProvider><MascotProvider>…` → single `<MascotProvider apiKey>`. `<MascotClient src><MascotRive /></MascotClient>` → `<Mascot src />` if you don't need a custom layout slot. If you do, change the tag name from `MascotClient` to `Mascot` and keep the children. </Step> <Step title="Typecheck + smoke"> `npx tsc --noEmit` is the gate. Runtime behavior is identical — if types pass, runtime almost always passes too. </Step> </Steps> *** # Legacy → 0.2.x The 0.2.x SDK was a new architecture, not a renamed release. If you integrated a pre-0.2 build, the changes below still apply (cascade 0.2.x → 0.3.0 from the section above on top). ## What changed, conceptually | Then (legacy) | Now (0.2.x+) | | --------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | One legacy `@mascotbot/react` package (pre-0.2) | `@mascotbot/core` + `@mascotbot/react`, each with a `/rive` subpath | | Visemes computed **server-side** and streamed to the client (REST `/v1/visemes`, SSE, a Mascot Bot WebSocket proxy, "viseme injection") | Visemes computed by the SDK on-device. No proxy, no REST viseme API, no SSE viseme protocol | | Provider-specific hooks (`useMascotElevenlabs`, `useMascotLiveAPI`, `useMascotOpenAI`) and a TTS hook (`useMascotSpeech`) | One realtime contract: tap the assistant's audio into `useLipsyncStream({ source: { kind: "mediaStream", stream } })`. TTS server returns audio only; the SDK does the lip sync | | `getSignedUrl` proxy that injected visemes | Provider wired with its **own** official SDK; server mints only a short-lived provider credential | | `result.argmax` per-frame array | `result.timeline` — a serializable, versioned `VisemeTimeline` | | Private-registry `.tgz` install | Private npm registry `npm.mascot.bot` with an `.npmrc` token | The dedicated provider guides ([ElevenLabs](/libraries/elevenlabs-avatar), [Gemini Live](/libraries/gemini-live-api-avatar), [OpenAI Realtime](/libraries/openai-realtime-api-avatar)) and [Realtime overview](/realtime/overview) show the new wiring end to end. ## Symbol map (legacy → 0.2.x — then add the 0.3.0 column above) | Legacy | 0.2.x replacement | | --------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Legacy `@mascotbot/react` (pre-0.2) | `@mascotbot/react` + `@mascotbot/react/rive` (0.3.0 names) | | `useMascot()` | `useMascotRive()` + `useMascotInputs()` (split the grab-bag — `rive`/render from the first, `custom`/`has`/`riveInputs` from the second) | | `useMascotClient()` | `useMascotRive()` + `useMascotInputs()` | | `useRiveInputs()` | `useMascotInputs()` | | `useMascotSpeech()` | Server route returns audio → `createPCMStreamPlayer` + `useLipsyncStream`. See [Realtime overview](/realtime/overview) | | `useMascotElevenlabs()` | [ElevenLabs guide](/libraries/elevenlabs-avatar) — own SDK + audio tap | | `useMascotLiveAPI()` | [Gemini Live guide](/libraries/gemini-live-api-avatar) | | `useMascotOpenAI()` | [OpenAI Realtime guide](/libraries/openai-realtime-api-avatar) | | `gesture: true` (auto-fire at utterance start) | The auto-fire is gone — the SDK never drives the `gesture` input. The built-in `stress` emphasis is `useMascotPlayback().stress([{ offset, stress }])`. A separate custom `gesture` trigger input is consumer-fired via `useMascotInputs().custom.gesture.fire()` — declare it on `<Mascot inputs={["gesture", ...]}>` first | | `useMicLipsync()` | `useLipsyncStream({ source: { kind: "mic" } })` | | `useStreamingLipsync()` | `useLipsyncStream({ source: { kind: "manual" } })` | | `useMediaStreamLipsync()` | `useLipsyncStream({ source: { kind: "mediaStream", stream } })` | | `playback.add(visemes)` | `playback.pushVisemes(cues)` (streaming) or `playback.setTimeline(tl)` (offline) | | `loadPrefetchedData(...)` | Persist a `VisemeTimeline` JSON; replay via `parseTimeline` + `setTimeline`. See [Offline lip sync](/libraries/offline-lipsync) | | `result.argmax` | `result.timeline` (a `VisemeTimeline`) | | `rive.stateMachineInputs(...)` introspection for presence | `useMascotInputs().has(name)` / `hasRiveInput(rive, name)` | | Mascot Bot proxy / `get-signed-url` / `/v1/visemes` | Removed — hybrid architecture; license init + refresh against `license.mascot.bot`. | ## Hooks are not reference-stable Legacy `useMascot()` returned a **memoised, stable** object; `useMascotRive()` / `useMascotInputs()` return a **fresh wrapper every render**. * Putting their return (`custom`, `has`, the whole handle) in a `useEffect` / `useCallback` **dependency array** re-runs it every render — e.g. a Rive event listener that re-binds \~60×/s. Capture what you need in a `useRef` and depend only on the **stable `rive` instance** (stable once loaded). `riveInputs` identity is stable, but the wrapper object and `has` are fresh per render — so this rule applies to them. <Warning> The one consumer responsibility that is real is **full per-call teardown on a shared client** — `stop()` + `close()` + null the stream state on every call-end path; see the [Troubleshooting guide](/libraries/react-troubleshooting). A lingering pipeline on the shared client corrupts other widgets that use the same `<MascotProvider>`. </Warning> ## Next <Columns> <Card title="Quickstart" icon="rocket" href="/quickstart"> The current happy path. </Card> <Card title="Realtime overview" icon="bolt" href="/realtime/overview"> Provider wiring. </Card> <Card title="Troubleshooting" icon="wrench" href="/libraries/react-troubleshooting"> Post-migration issues. </Card> </Columns>