Examples
Example: OpenAI Voice
A voice and speech-to-text provider — binary audio in and out through host.fetch, typed vendor failures.
Example plugin · modelProvider.speak + voices + transcribe · one secret · egress api.openai.com
This plugin adds OpenAI's text-to-speech voices and Whisper transcription as a provider for companion mode. It's
the reference for the two hard parts of an audio provider: binary audio through the sandbox, which has no
Buffer, and mapping vendor errors to the platform's failure kinds.
Manifest
// infrastructure/plugins/openai-voice/index.ts
globalThis.manifest = {
egress: ["api.openai.com"],
modelProvider: { kindId: "openai-voice", displayName: "OpenAI Voice" },
settings: [
{
key: "apiKey",
title: "API key",
description: "An OpenAI API key with access to /v1/audio/speech.",
type: "secret",
required: true,
},
],
};
modelProvider.kindId makes "OpenAI Voice from plugin" appear under AI Providers → Add provider. The apiKey
belongs to the provider the user adds there.
Voices
const VOICES: VoiceSummary[] = [
{ id: "alloy", name: "Alloy" },
{ id: "echo", name: "Echo" },
{ id: "fable", name: "Fable" },
{ id: "onyx", name: "Onyx" },
{ id: "nova", name: "Nova" },
{ id: "shimmer", name: "Shimmer" },
];
// in globalThis.modelProvider:
voices(request) {
const search = request.search?.trim().toLowerCase();
const voices = search ? VOICES.filter((voice) => voice.name.toLowerCase().includes(search)) : VOICES;
return { voices, nextCursor: null };
},
A static list, filtered by the picker's search box; nextCursor: null means one page. A vendor with a large
library would page with cursor, honour mine, and return sampleUrls so users can preview.
Speaking
async speak(request, tools) {
const apiKey = await host.settings.getSecret("apiKey");
const response = await host.fetch({
method: "POST",
url: "https://api.openai.com/v1/audio/speech",
headers: { Authorization: `Bearer ${apiKey}`, "Content-Type": "application/json" },
body: JSON.stringify({
model: "tts-1",
voice: request.voiceId,
input: stripCues(request.text),
response_format: "pcm",
}),
responseEncoding: "base64",
});
if (response.status !== 200) return mapFailure(response.status, tools);
return { audioBase64: response.body };
},
- Cues are stripped. The agent's spoken text can contain
[happy],[pause],[laughing]. OpenAI's TTS doesn't understand them, sostripCuesremoves any letters-only bracket (leaving[1]and markdown links). - The format matches exactly.
response_format: "pcm"returns 24 kHz, 16-bit little-endian mono — precisely whatspeakmust return — so there's no conversion. - Binary comes back as base64.
responseEncoding: "base64"makeshost.fetchreturn the raw bytes base64 encoded, which is theaudioBase64the platform wants. - No
words. OpenAI doesn't return word timings, so the platform estimates them from the audio for lip-sync and gestures.
(Note: this example always uses tts-1; a production plugin would send request.model, the TTS model the user
added.)
Listening
transcribe receives the recorded clip (today a 16 kHz mono WAV) as base64 and must upload it as
multipart/form-data. The sandbox has no Buffer/TextEncoder/atob, so the plugin builds the body as a byte
array with a tiny base64 codec of its own, then sends it with bodyEncoding: "base64":
async transcribe(request, tools) {
const apiKey = await host.settings.getSecret("apiKey");
const boundary = "agentparley-openai-voice-boundary";
const extension = extensionFor(request.mediaType);
const preamble = asciiBytes(
`--${boundary}\r\n` +
`Content-Disposition: form-data; name="model"\r\n\r\n` +
`whisper-1\r\n` +
`--${boundary}\r\n` +
`Content-Disposition: form-data; name="file"; filename="audio.${extension}"\r\n` +
`Content-Type: ${request.mediaType}\r\n\r\n`,
);
const postamble = asciiBytes(`\r\n--${boundary}--\r\n`);
const body = base64Encode([...preamble, ...base64Decode(request.audioBase64), ...postamble]);
const response = await host.fetch({
method: "POST",
url: "https://api.openai.com/v1/audio/transcriptions",
headers: { Authorization: `Bearer ${apiKey}`, "Content-Type": `multipart/form-data; boundary=${boundary}` },
body,
bodyEncoding: "base64",
});
if (response.status !== 200) return mapFailure(response.status, tools);
const parsed = JSON.parse(response.body);
return { text: typeof parsed.text === "string" ? parsed.text : "" };
},
Whisper chooses its decoder from the file name, so extensionFor(mediaType) maps the MIME type to an extension.
The plugin doesn't return durationSeconds, so each transcription is billed as a flat 1 second; returning the
vendor's duration would bill accurately.
Failures
function mapFailure(status: number, tools: ModelProviderTools): ModelProviderFailure {
if (status === 401 || status === 403) return tools.fail("keyRejected", "OpenAI rejected the API key");
if (status === 402) return tools.fail("noCredit", "OpenAI account is out of credit");
if (status === 429) return tools.fail("rateLimited", "OpenAI rate limit exceeded");
if (status >= 500) return tools.fail("unavailable", `OpenAI returned ${status}`);
return tools.fail("requestRejected", `OpenAI returned ${status}`);
}
A bad key or empty balance shows the account's red provider banner, exactly like a native vendor; rate limits and outages are transient and show no banner.
Using it
- Install; AI Providers → Add provider → OpenAI Voice from plugin; paste the key.
- Add a text-to-speech model and a speech-to-text model under it.
- On a companion, pick one of the six voices; set the speech-to-text model for listening. The mic button appears.
Permissions: none. A provider plugin may not act as an agent, and this one doesn't need to.