Building
Companion packs
Ship 3D characters, avatars and motion sets for companion mode — data only, no code.
Companion mode gives an agent a face and a voice: a 3D character in the browser that listens, looks thoughtful while the agent works, and answers out loud with lip-sync and gestures timed to the words — while still being the same agent, with its tools, files and crew.
A companion pack is a plugin that ships the look and the movement for that, as data. No code runs.
| Capability | Folder | What the user gets |
|---|---|---|
| Motion set | companion-motions/<key>/ |
how a companion moves: idle, talking, listening, thinking, gestures, reactions, and extra body parts like tails and ears |
| Model | companion-models/<key>/ |
a 3D avatar (VRM or GLB) to build a companion on |
| Character | companion-characters/<key>/ |
a ready-made face in the Companions gallery: a look + motions + a suggested personality and voice |
Voices come from voice providers — see Voice and speech-to-text. A character only suggests voices.
A complete pack
readme.md
companion-motions/sample-moves/motions.json
companion-motions/sample-moves/idle-neutral.vrma
companion-motions/sample-moves/talk-idle.vrma
companion-motions/sample-moves/listen.vrma
companion-motions/sample-moves/think.vrma
companion-motions/sample-moves/wave.vrma
companion-motions/sample-moves/laugh.vrma
companion-motions/sample-moves/tail-wag.glb
companion-models/sample-seed/model.json
companion-models/sample-seed/seed-san.vrm
companion-models/sample-seed/thumbnail.png
companion-characters/sample-seed-character/character.json
companion-characters/sample-belle-character/character.json
(infrastructure/plugins/sample-companion-pack — walkthrough: Sample Companion Pack)
Keys are unique across the whole package. companion-models/fox/ and companion-motions/fox/ collide — use
fox-model and fox-moves.
Every binary file must be referenced by one of the JSON files, or the publish is refused naming the file.
Motion set — motions.json
{
"name": "Sample Moves",
"description": "A small sample pack — a body loop for every base role, one gesture, one reaction, and a tail wag for avatars that have a tail.",
"clips": [
{ "key": "idle", "file": "idle-neutral.vrma", "role": "idle", "name": "Idle",
"description": "A long, unhurried neutral standing idle loop." },
{ "key": "talk-idle", "file": "talk-idle.vrma", "role": "talkIdle", "name": "Talk idle",
"description": "Idle loop with light gestures while speaking." },
{ "key": "listen", "file": "listen.vrma", "role": "listen", "name": "Listen (agreeing)",
"description": "A small nod of agreement — mic on, actively listening." },
{ "key": "think", "file": "think.vrma", "role": "think", "name": "Think",
"description": "Hand rises to the chin and holds — waiting on a reply." },
{ "key": "wave", "file": "wave.vrma", "role": "gesture", "name": "Wave",
"description": "Waves a hand — greeting or farewell.", "cues": ["happy", "excited"], "strokeSeconds": 0.6 },
{ "key": "laugh", "file": "laugh.vrma", "role": "reaction", "name": "Laugh",
"description": "A real laugh — reaction for a laughing line.", "cues": ["laughing", "chuckling"], "strokeSeconds": 3.83 },
{ "key": "tail-wag", "file": "tail-wag.glb", "clip": "Tail_Wag", "role": "partsGesture", "name": "Tail wag",
"description": "Wags her tail — happy." }
]
}
| Field | Rules |
|---|---|
name |
required, 1–100 characters |
description |
≤ 500 characters |
clips |
1–100 clips; at least one idle; at most 30 gesture and 10 partsGesture |
Each clip:
| Field | Rules |
|---|---|
key |
required, ^[a-z0-9][a-z0-9_-]{0,63}$, unique in the set — your stable id for the clip across versions |
file |
required, path inside the package; several clips may share one file |
role |
required — see below |
name, description |
required — the description is what the gesture picker reads, so describe the motion and its mood |
clip |
the animation's name inside a GLB or FBX file (not allowed for VRMA) |
rig |
a bone map — see Bone maps |
cues |
up to 10 tone words that trigger this clip (see below) |
strokeSeconds |
0–30 — when in the clip the "beat" lands, so it can be timed to a word |
isFullBody, isPointingAtWord |
optional booleans |
Roles:
| Role | Plays |
|---|---|
idle |
standing still, the default loop |
talkIdle |
the loop while speaking |
listen |
while the user is talking |
think |
while the agent is working |
gesture |
a movement chosen per sentence while speaking |
reaction |
an emotional beat, e.g. a laugh on a [laughing] line |
partsIdle |
an always-on loop for extra body parts (tail, ears, wings) |
partsGesture |
an extra-parts movement chosen while speaking (a tail wag on a happy line) |
File formats:
| Format | clip |
rig |
Can drive |
|---|---|---|---|
.vrma (VRM Animation) |
not allowed | not allowed | the body only — never an extra-parts role |
.glb |
required | only for a body clip on a file without VRM humanoid data | the body, or named extra-part bones |
.fbx (e.g. Mixamo) |
required | required | the body only |
Clips are checked against each avatar's bones, so a companion only ever gets clips that will actually play on it. An extra-parts clip must move at least one named bone. Clip files ≤ 32 MB.
Cues tie clips to the tone of what the agent says. When the agent speaks aloud it can mark lines with stage
directions such as [happy] or [laughing]; a clip whose cues include that word becomes the natural pick.
Allowed words are the emotions (happy, sad, excited, calm, curious, surprised, proud, grateful,
confused, … — about fifty), the sound effects (laughing, chuckling, sighing, gasping, crying,
yawning, …) and stretching. An unknown cue refuses the publish, naming it.
Motion sets are followed live. A companion using your motion set always uses its current version — publish an update and every companion moves the new way. If the user uninstalls your plugin, their companions fall back to their look's default motions.
Model — model.json
{
"name": "Sample Seed-san",
"description": "A copy of the built-in Seed-san model, offered as a plugin companion model for testing.",
"avatar": "seed-san.vrm",
"thumbnail": "thumbnail.png",
"defaultMotions": { "key": "sample-moves" }
}
| Field | Rules |
|---|---|
name, description |
display text |
avatar |
required, .vrm or .glb, ≤ 100 MB. VRM gets the full treatment (facial expressions, visemes, spring-bone physics); a plain GLB lip-syncs only through blend shapes or a jaw bone if it has them |
thumbnail |
optional PNG, JPEG or WebP, ≤ 2 MB — shown in the "new companion" picker |
defaultMotions |
optional; exactly one of { "key": "<a motion set in this package>" } or { "builtIn": "<built-in set>" } |
Models are copied. When a user creates a companion from your model, the avatar is copied into their account as their own. It keeps working after they uninstall your plugin.
Character — character.json
{
"name": "Sample Seed",
"vibe": "A test character with a plugin look and plugin motions.",
"model": { "key": "sample-seed" },
"motions": { "key": "sample-moves" },
"soul": "You are Sample Seed, a cheerful test companion. You keep answers short and friendly.",
"instructions": "You talk with the person out loud. Keep replies to a sentence or two and avoid lists.",
"voice": {
"description": "A bright, friendly young voice.",
"suggestions": [
{ "provider": "fish-audio", "voiceId": "8ef4a238714b45718ce04243307c57a7", "voiceName": "Sample voice" },
{ "provider": "elevenlabs", "voiceId": "21m00Tcm4TlvDq8ikWAM", "voiceName": "Rachel", "model": "eleven_v3" }
]
}
}
A character appears in the Companions gallery tagged Plugin · your plugin's name. Talk creates a new agent from it; Talk as… puts the face on an existing agent (optionally adopting the suggested soul). Each character becomes at most one companion per account.
The file is read strictly: an unknown field (a typo like "sugestions") refuses the publish.
| Field | Rules |
|---|---|
name |
required, 1–100 |
vibe |
required, 1–120 — the one line on the gallery card |
soul |
required, 1–8,000 — the suggested persona for the agent |
instructions |
required, 1–8,000 — the suggested operating instructions |
model |
required; exactly one of { "key": "<model in this package>" } or { "builtIn": "belle" | "rose" | "seed-san" | "polydancer" } |
motions |
optional; { "key": … } or { "builtIn": "feminine-playful" | "masculine-confident" | "neutral-professional" | "energetic" }. Omitted = the look's defaults |
thumbnail |
optional PNG/JPEG/WebP ≤ 2 MB, inside the character's folder |
voice.description |
≤ 300 characters, shown to the user |
voice.suggestions |
≤ 8, one per provider: provider (elevenlabs, fish-audio, deepgram, openai-compatible, or a plugin provider's kindId), voiceId (1–128), optional voiceName (≤ 100) and model (≤ 200) |
How the voice is picked: the companion takes the first suggestion whose provider the user has set up. None
matching means captions only, with a notice offering to add a provider. For ElevenLabs, fish.audio and Deepgram a
voice id works on any of their models; for OpenAI-compatible and plugin providers a voice id belongs to one model,
so set model — the suggestion then applies only when the user has exactly that model.
Soul and instructions are suggestions: the personality always belongs to the agent. The agent's soul also steers which gestures the companion picks.
Bone maps (rig)
A body clip in an FBX file, or in a GLB without VRM humanoid data, needs a bone map from humanoid bone names to the file's node names:
"rig": {
"version": 1,
"bones": {
"hips": "mixamorigHips", "spine": "mixamorigSpine", "head": "mixamorigHead",
"leftUpperArm": "mixamorigLeftArm", "leftLowerArm": "mixamorigLeftForeArm", "leftHand": "mixamorigLeftHand",
"rightUpperArm": "mixamorigRightArm", "rightLowerArm": "mixamorigRightForeArm", "rightHand": "mixamorigRightHand",
"leftUpperLeg": "mixamorigLeftUpLeg", "leftLowerLeg": "mixamorigLeftLeg", "leftFoot": "mixamorigLeftFoot",
"rightUpperLeg": "mixamorigRightUpLeg", "rightLowerLeg": "mixamorigRightLeg", "rightFoot": "mixamorigRightFoot"
}
}
version must be 1; those 15 bones are required.
Write every bone name the way three.js spells it. Browsers load these files with three.js, which rewrites node
names: whitespace becomes _, and [ ] . : / are removed. Names in your JSON must already be in that form — the
publish refuses a raw name and tells you the correct spelling:
| In the source file | In your JSON |
|---|---|
Tail.001 |
Tail001 |
mixamorig:Hips |
mixamorigHips |
Ear L |
Ear_L |
A file where two nodes would end up with the same name after that rewrite is refused.
The easy way: export from the console
Hand-writing bone maps is real work you can skip. Build a motion set in the console (Companions → Motions), then use Export as plugin package: you get a ready-to-publish zip in exactly this format, with FBX and model-embedded animations converted to small GLBs and every bone name already spelled correctly. (A set that uses another plugin's clips can't be exported.)
Limits
| Package | 256 MB |
| Avatar model | 100 MB |
| Animation clip | 32 MB |
| Thumbnail | 2 MB (PNG, JPEG, WebP) |
| Clips per motion set | 100 (≤ 30 gestures, ≤ 10 extra-parts gestures) |
| Cues per clip | 10 |