Building

Companion packs

Ship 3D characters, avatars and motion sets for companion mode — data only, no code.

Companion mode gives an agent a face and a voice: a 3D character in the browser that listens, looks thoughtful while the agent works, and answers out loud with lip-sync and gestures timed to the words — while still being the same agent, with its tools, files and crew.

A companion pack is a plugin that ships the look and the movement for that, as data. No code runs.

Capability Folder What the user gets
Motion set companion-motions/<key>/ how a companion moves: idle, talking, listening, thinking, gestures, reactions, and extra body parts like tails and ears
Model companion-models/<key>/ a 3D avatar (VRM or GLB) to build a companion on
Character companion-characters/<key>/ a ready-made face in the Companions gallery: a look + motions + a suggested personality and voice

Voices come from voice providers — see Voice and speech-to-text. A character only suggests voices.

A complete pack

readme.md
companion-motions/sample-moves/motions.json
companion-motions/sample-moves/idle-neutral.vrma
companion-motions/sample-moves/talk-idle.vrma
companion-motions/sample-moves/listen.vrma
companion-motions/sample-moves/think.vrma
companion-motions/sample-moves/wave.vrma
companion-motions/sample-moves/laugh.vrma
companion-motions/sample-moves/tail-wag.glb
companion-models/sample-seed/model.json
companion-models/sample-seed/seed-san.vrm
companion-models/sample-seed/thumbnail.png
companion-characters/sample-seed-character/character.json
companion-characters/sample-belle-character/character.json

(infrastructure/plugins/sample-companion-pack — walkthrough: Sample Companion Pack)

Keys are unique across the whole package. companion-models/fox/ and companion-motions/fox/ collide — use fox-model and fox-moves.

Every binary file must be referenced by one of the JSON files, or the publish is refused naming the file.

Motion set — motions.json

{
  "name": "Sample Moves",
  "description": "A small sample pack — a body loop for every base role, one gesture, one reaction, and a tail wag for avatars that have a tail.",
  "clips": [
    { "key": "idle", "file": "idle-neutral.vrma", "role": "idle", "name": "Idle",
      "description": "A long, unhurried neutral standing idle loop." },
    { "key": "talk-idle", "file": "talk-idle.vrma", "role": "talkIdle", "name": "Talk idle",
      "description": "Idle loop with light gestures while speaking." },
    { "key": "listen", "file": "listen.vrma", "role": "listen", "name": "Listen (agreeing)",
      "description": "A small nod of agreement — mic on, actively listening." },
    { "key": "think", "file": "think.vrma", "role": "think", "name": "Think",
      "description": "Hand rises to the chin and holds — waiting on a reply." },
    { "key": "wave", "file": "wave.vrma", "role": "gesture", "name": "Wave",
      "description": "Waves a hand — greeting or farewell.", "cues": ["happy", "excited"], "strokeSeconds": 0.6 },
    { "key": "laugh", "file": "laugh.vrma", "role": "reaction", "name": "Laugh",
      "description": "A real laugh — reaction for a laughing line.", "cues": ["laughing", "chuckling"], "strokeSeconds": 3.83 },
    { "key": "tail-wag", "file": "tail-wag.glb", "clip": "Tail_Wag", "role": "partsGesture", "name": "Tail wag",
      "description": "Wags her tail — happy." }
  ]
}
Field Rules
name required, 1–100 characters
description ≤ 500 characters
clips 1–100 clips; at least one idle; at most 30 gesture and 10 partsGesture

Each clip:

Field Rules
key required, ^[a-z0-9][a-z0-9_-]{0,63}$, unique in the set — your stable id for the clip across versions
file required, path inside the package; several clips may share one file
role required — see below
name, description required — the description is what the gesture picker reads, so describe the motion and its mood
clip the animation's name inside a GLB or FBX file (not allowed for VRMA)
rig a bone map — see Bone maps
cues up to 10 tone words that trigger this clip (see below)
strokeSeconds 0–30 — when in the clip the "beat" lands, so it can be timed to a word
isFullBody, isPointingAtWord optional booleans

Roles:

Role Plays
idle standing still, the default loop
talkIdle the loop while speaking
listen while the user is talking
think while the agent is working
gesture a movement chosen per sentence while speaking
reaction an emotional beat, e.g. a laugh on a [laughing] line
partsIdle an always-on loop for extra body parts (tail, ears, wings)
partsGesture an extra-parts movement chosen while speaking (a tail wag on a happy line)

File formats:

Format clip rig Can drive
.vrma (VRM Animation) not allowed not allowed the body only — never an extra-parts role
.glb required only for a body clip on a file without VRM humanoid data the body, or named extra-part bones
.fbx (e.g. Mixamo) required required the body only

Clips are checked against each avatar's bones, so a companion only ever gets clips that will actually play on it. An extra-parts clip must move at least one named bone. Clip files ≤ 32 MB.

Cues tie clips to the tone of what the agent says. When the agent speaks aloud it can mark lines with stage directions such as [happy] or [laughing]; a clip whose cues include that word becomes the natural pick. Allowed words are the emotions (happy, sad, excited, calm, curious, surprised, proud, grateful, confused, … — about fifty), the sound effects (laughing, chuckling, sighing, gasping, crying, yawning, …) and stretching. An unknown cue refuses the publish, naming it.

Motion sets are followed live. A companion using your motion set always uses its current version — publish an update and every companion moves the new way. If the user uninstalls your plugin, their companions fall back to their look's default motions.

Model — model.json

{
  "name": "Sample Seed-san",
  "description": "A copy of the built-in Seed-san model, offered as a plugin companion model for testing.",
  "avatar": "seed-san.vrm",
  "thumbnail": "thumbnail.png",
  "defaultMotions": { "key": "sample-moves" }
}
Field Rules
name, description display text
avatar required, .vrm or .glb, ≤ 100 MB. VRM gets the full treatment (facial expressions, visemes, spring-bone physics); a plain GLB lip-syncs only through blend shapes or a jaw bone if it has them
thumbnail optional PNG, JPEG or WebP, ≤ 2 MB — shown in the "new companion" picker
defaultMotions optional; exactly one of { "key": "<a motion set in this package>" } or { "builtIn": "<built-in set>" }

Models are copied. When a user creates a companion from your model, the avatar is copied into their account as their own. It keeps working after they uninstall your plugin.

Character — character.json

{
  "name": "Sample Seed",
  "vibe": "A test character with a plugin look and plugin motions.",
  "model": { "key": "sample-seed" },
  "motions": { "key": "sample-moves" },
  "soul": "You are Sample Seed, a cheerful test companion. You keep answers short and friendly.",
  "instructions": "You talk with the person out loud. Keep replies to a sentence or two and avoid lists.",
  "voice": {
    "description": "A bright, friendly young voice.",
    "suggestions": [
      { "provider": "fish-audio", "voiceId": "8ef4a238714b45718ce04243307c57a7", "voiceName": "Sample voice" },
      { "provider": "elevenlabs", "voiceId": "21m00Tcm4TlvDq8ikWAM", "voiceName": "Rachel", "model": "eleven_v3" }
    ]
  }
}

A character appears in the Companions gallery tagged Plugin · your plugin's name. Talk creates a new agent from it; Talk as… puts the face on an existing agent (optionally adopting the suggested soul). Each character becomes at most one companion per account.

The file is read strictly: an unknown field (a typo like "sugestions") refuses the publish.

Field Rules
name required, 1–100
vibe required, 1–120 — the one line on the gallery card
soul required, 1–8,000 — the suggested persona for the agent
instructions required, 1–8,000 — the suggested operating instructions
model required; exactly one of { "key": "<model in this package>" } or { "builtIn": "belle" | "rose" | "seed-san" | "polydancer" }
motions optional; { "key": … } or { "builtIn": "feminine-playful" | "masculine-confident" | "neutral-professional" | "energetic" }. Omitted = the look's defaults
thumbnail optional PNG/JPEG/WebP ≤ 2 MB, inside the character's folder
voice.description ≤ 300 characters, shown to the user
voice.suggestions ≤ 8, one per provider: provider (elevenlabs, fish-audio, deepgram, openai-compatible, or a plugin provider's kindId), voiceId (1–128), optional voiceName (≤ 100) and model (≤ 200)

How the voice is picked: the companion takes the first suggestion whose provider the user has set up. None matching means captions only, with a notice offering to add a provider. For ElevenLabs, fish.audio and Deepgram a voice id works on any of their models; for OpenAI-compatible and plugin providers a voice id belongs to one model, so set model — the suggestion then applies only when the user has exactly that model.

Soul and instructions are suggestions: the personality always belongs to the agent. The agent's soul also steers which gestures the companion picks.

Bone maps (rig)

A body clip in an FBX file, or in a GLB without VRM humanoid data, needs a bone map from humanoid bone names to the file's node names:

"rig": {
  "version": 1,
  "bones": {
    "hips": "mixamorigHips", "spine": "mixamorigSpine", "head": "mixamorigHead",
    "leftUpperArm": "mixamorigLeftArm", "leftLowerArm": "mixamorigLeftForeArm", "leftHand": "mixamorigLeftHand",
    "rightUpperArm": "mixamorigRightArm", "rightLowerArm": "mixamorigRightForeArm", "rightHand": "mixamorigRightHand",
    "leftUpperLeg": "mixamorigLeftUpLeg", "leftLowerLeg": "mixamorigLeftLeg", "leftFoot": "mixamorigLeftFoot",
    "rightUpperLeg": "mixamorigRightUpLeg", "rightLowerLeg": "mixamorigRightLeg", "rightFoot": "mixamorigRightFoot"
  }
}

version must be 1; those 15 bones are required.

Write every bone name the way three.js spells it. Browsers load these files with three.js, which rewrites node names: whitespace becomes _, and [ ] . : / are removed. Names in your JSON must already be in that form — the publish refuses a raw name and tells you the correct spelling:

In the source file In your JSON
Tail.001 Tail001
mixamorig:Hips mixamorigHips
Ear L Ear_L

A file where two nodes would end up with the same name after that rewrite is refused.

The easy way: export from the console

Hand-writing bone maps is real work you can skip. Build a motion set in the console (Companions → Motions), then use Export as plugin package: you get a ready-to-publish zip in exactly this format, with FBX and model-embedded animations converted to small GLBs and every bone name already spelled correctly. (A set that uses another plugin's clips can't be exported.)

Limits

Package 256 MB
Avatar model 100 MB
Animation clip 32 MB
Thumbnail 2 MB (PNG, JPEG, WebP)
Clips per motion set 100 (≤ 30 gestures, ≤ 10 extra-parts gestures)
Cues per clip 10