Building

Voice and speech-to-text providers

Give companions a new voice vendor, or a new way to hear the user.

In companion mode an agent speaks its replies out loud and can listen to the user's voice. AgentParley ships fish.audio, ElevenLabs, Deepgram and OpenAI-compatible voices; a plugin can add any other vendor — for speaking, for listening, or both. Users add it under AI Providers → Add provider like a native vendor, then pick its voices on a companion.

These are model providers: the same manifest.modelProvider, the same tools.fail(kind, message), the same "no agent powers" rule. Two handler groups:

Group Handlers Rule
Voice (text-to-speech) speak + voices both or neither
Speech-to-text transcribe on its own, or alongside anything else

The reference implementation is OpenAI Voice — see the walkthrough.

// infrastructure/plugins/openai-voice/index.ts
globalThis.manifest = {
  egress: ["api.openai.com"],
  modelProvider: { kindId: "openai-voice", displayName: "OpenAI Voice" },
  settings: [
    {
      key: "apiKey",
      title: "API key",
      description: "An OpenAI API key with access to /v1/audio/speech.",
      type: "secret",
      required: true,
    },
  ],
};

speak — say one group of sentences

request: {
  voiceId: string;    // the voice the user picked (one of yours, from `voices`)
  text: string;       // one short group of sentences, ≤ ~300 characters
  model: string;      // the text-to-speech model id the user saved
}
returns: {
  audioBase64: string;                                  // raw PCM: signed 16-bit little-endian, 24 kHz, mono, no header
  words?: { text: string; start: number; end: number }[]; // optional per-word timing, seconds from the first sample
} | ModelProviderFailure
  • The text can contain cues — bracketed stage directions the agent writes for spoken replies, such as [happy], [whispering], [laughing], [pause]. If your vendor understands them, translate them; if not, strip them:

    // infrastructure/plugins/openai-voice/index.ts
    function stripCues(text: string): string {
      return text
        .replace(/[ \t]*\[\p{L}[\p{L} '\-]{0,38}\](?!\()/gu, "")
        .replace(/[ \t]+([.,!?…])/g, "$1")
        .trim();
    }
    
  • Audio format is fixed: raw 24 kHz mono 16-bit PCM, base64. Fetch binary from your vendor with host.fetch({ …, responseEncoding: "base64" }):

    // infrastructure/plugins/openai-voice/index.ts
    async speak(request, tools) {
      const apiKey = await host.settings.getSecret("apiKey");
      const response = await host.fetch({
        method: "POST",
        url: "https://api.openai.com/v1/audio/speech",
        headers: { Authorization: `Bearer ${apiKey}`, "Content-Type": "application/json" },
        body: JSON.stringify({
          model: "tts-1",
          voice: request.voiceId,
          input: stripCues(request.text),
          response_format: "pcm",
        }),
        responseEncoding: "base64",
      });
      if (response.status !== 200) return mapFailure(response.status, tools);
      return { audioBase64: response.body };
    },
    
  • Word timings make the face land on the word. If your vendor returns per-word timestamps, pass them as words; lip-sync and gestures then line up exactly. Without them the platform estimates from the audio itself. At most 2,000 entries; each needs 0 ≤ start ≤ end (invalid entries are dropped, the audio still plays).

  • Buffered, not streamed. The user hears nothing until your handler returns the whole group, so keep it quick: about 15 s per call, and the response must fit the transport (~4 MB, about 30 s of audio).

voices — list voices to pick from

request: {
  search: string | null;      // what the user typed in the picker
  mine: boolean;              // only voices this account owns on the vendor (vs. the public library)
  cursor: string | null;      // null for page 1; otherwise your own nextCursor from last time
  voiceIds: string[] | null;  // look up specific ids (e.g. to show an already-picked voice's name)
  model: string;              // the TTS model the voices are for
}
returns: {
  voices: { id: string; name: string; description?: string; languages?: string[];
            sampleUrl?: string; coverImageUrl?: string }[];
  nextCursor: string | null;  // null = no more pages
} | ModelProviderFailure

nextCursor is opaque to the platform — a page number, an offset, a vendor token. sampleUrl and coverImageUrl are loaded by the user's browser directly. About 10 s per call. There is no manual voice-id entry: a voice can only be picked from what voices returns.

// infrastructure/plugins/openai-voice/index.ts
voices(request) {
  const search = request.search?.trim().toLowerCase();
  const voices = search ? VOICES.filter((voice) => voice.name.toLowerCase().includes(search)) : VOICES;
  return { voices, nextCursor: null };
},

transcribe — turn the user's speech into text

request: {
  audioBase64: string;       // the whole recorded clip, base64
  mediaType: string;         // its MIME type — today always "audio/wav" (16 kHz, mono, 16-bit)
  model: string;             // the speech-to-text model id the user saved
  language: string | null;   // the user's language code, or null for auto-detect
}
returns: { text: string; durationSeconds?: number } | ModelProviderFailure
  • Recorded, not live. The platform waits until the user stops talking, then calls you once with the whole clip. Plugin speech-to-text can't show words while the user is still speaking (the built-in Deepgram path can).
  • durationSeconds is what the user is billed for (rounded up); omit it and a flat 1 second is billed.
  • About 20 s per call.
  • Vendors that want a file upload need multipart/form-data. The sandbox has no Buffer or TextEncoder, so build the body as bytes and send it with bodyEncoding: "base64" — OpenAI Voice shows how:
// infrastructure/plugins/openai-voice/index.ts
async transcribe(request, tools) {
  const apiKey = await host.settings.getSecret("apiKey");
  const boundary = "agentparley-openai-voice-boundary";
  const extension = extensionFor(request.mediaType);
  const preamble = asciiBytes(
    `--${boundary}\r\n` +
      `Content-Disposition: form-data; name="model"\r\n\r\n` +
      `whisper-1\r\n` +
      `--${boundary}\r\n` +
      `Content-Disposition: form-data; name="file"; filename="audio.${extension}"\r\n` +
      `Content-Type: ${request.mediaType}\r\n\r\n`,
  );
  const postamble = asciiBytes(`\r\n--${boundary}--\r\n`);
  const body = base64Encode([...preamble, ...base64Decode(request.audioBase64), ...postamble]);

  const response = await host.fetch({
    method: "POST",
    url: "https://api.openai.com/v1/audio/transcriptions",
    headers: { Authorization: `Bearer ${apiKey}`, "Content-Type": `multipart/form-data; boundary=${boundary}` },
    body,
    bodyEncoding: "base64",
  });
  if (response.status !== 200) return mapFailure(response.status, tools);

  const parsed = JSON.parse(response.body);
  return { text: typeof parsed.text === "string" ? parsed.text : "" };
},

Failures

Return tools.fail(kind, message):

  • keyRejected, noCredit, modelNotFound, requestRejected → the account's red provider banner.
  • rateLimited, unavailable → no banner; treated as a transient failure.

A thrown error or timeout counts as unavailable.

How a user sets it up

  1. Install your plugin.
  2. AI Providers → Add provider → "Your name from plugin", enter the key.
  3. Add a text-to-speech model and/or a speech-to-text model under it, with the model identifier your vendor uses — it reaches you as request.model.
  4. On a companion, open the voice picker, choose your provider, and pick a voice. For listening, choose your speech-to-text model as the account default or for that companion. The mic button appears once a speech-to-text model is available.

Companion characters can suggest your voices by your kindId — see Companion packs. For plugin providers, include model in the suggestion.