Examples

Example: OpenAI Voice

A voice and speech-to-text provider — binary audio in and out through host.fetch, typed vendor failures.

Example plugin · modelProvider.speak + voices + transcribe · one secret · egress api.openai.com

This plugin adds OpenAI's text-to-speech voices and Whisper transcription as a provider for companion mode. It's the reference for the two hard parts of an audio provider: binary audio through the sandbox, which has no Buffer, and mapping vendor errors to the platform's failure kinds.

Manifest

// infrastructure/plugins/openai-voice/index.ts
globalThis.manifest = {
  egress: ["api.openai.com"],
  modelProvider: { kindId: "openai-voice", displayName: "OpenAI Voice" },
  settings: [
    {
      key: "apiKey",
      title: "API key",
      description: "An OpenAI API key with access to /v1/audio/speech.",
      type: "secret",
      required: true,
    },
  ],
};

modelProvider.kindId makes "OpenAI Voice from plugin" appear under AI Providers → Add provider. The apiKey belongs to the provider the user adds there.

Voices

const VOICES: VoiceSummary[] = [
  { id: "alloy", name: "Alloy" },
  { id: "echo", name: "Echo" },
  { id: "fable", name: "Fable" },
  { id: "onyx", name: "Onyx" },
  { id: "nova", name: "Nova" },
  { id: "shimmer", name: "Shimmer" },
];

// in globalThis.modelProvider:
voices(request) {
  const search = request.search?.trim().toLowerCase();
  const voices = search ? VOICES.filter((voice) => voice.name.toLowerCase().includes(search)) : VOICES;
  return { voices, nextCursor: null };
},

A static list, filtered by the picker's search box; nextCursor: null means one page. A vendor with a large library would page with cursor, honour mine, and return sampleUrls so users can preview.

Speaking

async speak(request, tools) {
  const apiKey = await host.settings.getSecret("apiKey");
  const response = await host.fetch({
    method: "POST",
    url: "https://api.openai.com/v1/audio/speech",
    headers: { Authorization: `Bearer ${apiKey}`, "Content-Type": "application/json" },
    body: JSON.stringify({
      model: "tts-1",
      voice: request.voiceId,
      input: stripCues(request.text),
      response_format: "pcm",
    }),
    responseEncoding: "base64",
  });
  if (response.status !== 200) return mapFailure(response.status, tools);
  return { audioBase64: response.body };
},
  1. Cues are stripped. The agent's spoken text can contain [happy], [pause], [laughing]. OpenAI's TTS doesn't understand them, so stripCues removes any letters-only bracket (leaving [1] and markdown links).
  2. The format matches exactly. response_format: "pcm" returns 24 kHz, 16-bit little-endian mono — precisely what speak must return — so there's no conversion.
  3. Binary comes back as base64. responseEncoding: "base64" makes host.fetch return the raw bytes base64 encoded, which is the audioBase64 the platform wants.
  4. No words. OpenAI doesn't return word timings, so the platform estimates them from the audio for lip-sync and gestures.

(Note: this example always uses tts-1; a production plugin would send request.model, the TTS model the user added.)

Listening

transcribe receives the recorded clip (today a 16 kHz mono WAV) as base64 and must upload it as multipart/form-data. The sandbox has no Buffer/TextEncoder/atob, so the plugin builds the body as a byte array with a tiny base64 codec of its own, then sends it with bodyEncoding: "base64":

async transcribe(request, tools) {
  const apiKey = await host.settings.getSecret("apiKey");
  const boundary = "agentparley-openai-voice-boundary";
  const extension = extensionFor(request.mediaType);
  const preamble = asciiBytes(
    `--${boundary}\r\n` +
      `Content-Disposition: form-data; name="model"\r\n\r\n` +
      `whisper-1\r\n` +
      `--${boundary}\r\n` +
      `Content-Disposition: form-data; name="file"; filename="audio.${extension}"\r\n` +
      `Content-Type: ${request.mediaType}\r\n\r\n`,
  );
  const postamble = asciiBytes(`\r\n--${boundary}--\r\n`);
  const body = base64Encode([...preamble, ...base64Decode(request.audioBase64), ...postamble]);

  const response = await host.fetch({
    method: "POST",
    url: "https://api.openai.com/v1/audio/transcriptions",
    headers: { Authorization: `Bearer ${apiKey}`, "Content-Type": `multipart/form-data; boundary=${boundary}` },
    body,
    bodyEncoding: "base64",
  });
  if (response.status !== 200) return mapFailure(response.status, tools);

  const parsed = JSON.parse(response.body);
  return { text: typeof parsed.text === "string" ? parsed.text : "" };
},

Whisper chooses its decoder from the file name, so extensionFor(mediaType) maps the MIME type to an extension. The plugin doesn't return durationSeconds, so each transcription is billed as a flat 1 second; returning the vendor's duration would bill accurately.

Failures

function mapFailure(status: number, tools: ModelProviderTools): ModelProviderFailure {
  if (status === 401 || status === 403) return tools.fail("keyRejected", "OpenAI rejected the API key");
  if (status === 402) return tools.fail("noCredit", "OpenAI account is out of credit");
  if (status === 429) return tools.fail("rateLimited", "OpenAI rate limit exceeded");
  if (status >= 500) return tools.fail("unavailable", `OpenAI returned ${status}`);
  return tools.fail("requestRejected", `OpenAI returned ${status}`);
}

A bad key or empty balance shows the account's red provider banner, exactly like a native vendor; rate limits and outages are transient and show no banner.

Using it

  1. Install; AI Providers → Add provider → OpenAI Voice from plugin; paste the key.
  2. Add a text-to-speech model and a speech-to-text model under it.
  3. On a companion, pick one of the six voices; set the speech-to-text model for listening. The mic button appears.

Permissions: none. A provider plugin may not act as an agent, and this one doesn't need to.