Building
Voice and speech-to-text providers
Give companions a new voice vendor, or a new way to hear the user.
In companion mode an agent speaks its replies out loud and can listen to the user's voice. AgentParley ships fish.audio, ElevenLabs, Deepgram and OpenAI-compatible voices; a plugin can add any other vendor — for speaking, for listening, or both. Users add it under AI Providers → Add provider like a native vendor, then pick its voices on a companion.
These are model providers: the same manifest.modelProvider, the same
tools.fail(kind, message), the same "no agent powers" rule. Two handler groups:
| Group | Handlers | Rule |
|---|---|---|
| Voice (text-to-speech) | speak + voices |
both or neither |
| Speech-to-text | transcribe |
on its own, or alongside anything else |
The reference implementation is OpenAI Voice — see the walkthrough.
// infrastructure/plugins/openai-voice/index.ts
globalThis.manifest = {
egress: ["api.openai.com"],
modelProvider: { kindId: "openai-voice", displayName: "OpenAI Voice" },
settings: [
{
key: "apiKey",
title: "API key",
description: "An OpenAI API key with access to /v1/audio/speech.",
type: "secret",
required: true,
},
],
};
speak — say one group of sentences
request: {
voiceId: string; // the voice the user picked (one of yours, from `voices`)
text: string; // one short group of sentences, ≤ ~300 characters
model: string; // the text-to-speech model id the user saved
}
returns: {
audioBase64: string; // raw PCM: signed 16-bit little-endian, 24 kHz, mono, no header
words?: { text: string; start: number; end: number }[]; // optional per-word timing, seconds from the first sample
} | ModelProviderFailure
The text can contain cues — bracketed stage directions the agent writes for spoken replies, such as
[happy],[whispering],[laughing],[pause]. If your vendor understands them, translate them; if not, strip them:// infrastructure/plugins/openai-voice/index.ts function stripCues(text: string): string { return text .replace(/[ \t]*\[\p{L}[\p{L} '\-]{0,38}\](?!\()/gu, "") .replace(/[ \t]+([.,!?…])/g, "$1") .trim(); }Audio format is fixed: raw 24 kHz mono 16-bit PCM, base64. Fetch binary from your vendor with
host.fetch({ …, responseEncoding: "base64" }):// infrastructure/plugins/openai-voice/index.ts async speak(request, tools) { const apiKey = await host.settings.getSecret("apiKey"); const response = await host.fetch({ method: "POST", url: "https://api.openai.com/v1/audio/speech", headers: { Authorization: `Bearer ${apiKey}`, "Content-Type": "application/json" }, body: JSON.stringify({ model: "tts-1", voice: request.voiceId, input: stripCues(request.text), response_format: "pcm", }), responseEncoding: "base64", }); if (response.status !== 200) return mapFailure(response.status, tools); return { audioBase64: response.body }; },Word timings make the face land on the word. If your vendor returns per-word timestamps, pass them as
words; lip-sync and gestures then line up exactly. Without them the platform estimates from the audio itself. At most 2,000 entries; each needs0 ≤ start ≤ end(invalid entries are dropped, the audio still plays).Buffered, not streamed. The user hears nothing until your handler returns the whole group, so keep it quick: about 15 s per call, and the response must fit the transport (~4 MB, about 30 s of audio).
voices — list voices to pick from
request: {
search: string | null; // what the user typed in the picker
mine: boolean; // only voices this account owns on the vendor (vs. the public library)
cursor: string | null; // null for page 1; otherwise your own nextCursor from last time
voiceIds: string[] | null; // look up specific ids (e.g. to show an already-picked voice's name)
model: string; // the TTS model the voices are for
}
returns: {
voices: { id: string; name: string; description?: string; languages?: string[];
sampleUrl?: string; coverImageUrl?: string }[];
nextCursor: string | null; // null = no more pages
} | ModelProviderFailure
nextCursor is opaque to the platform — a page number, an offset, a vendor token. sampleUrl and coverImageUrl
are loaded by the user's browser directly. About 10 s per call. There is no manual voice-id entry: a voice can only
be picked from what voices returns.
// infrastructure/plugins/openai-voice/index.ts
voices(request) {
const search = request.search?.trim().toLowerCase();
const voices = search ? VOICES.filter((voice) => voice.name.toLowerCase().includes(search)) : VOICES;
return { voices, nextCursor: null };
},
transcribe — turn the user's speech into text
request: {
audioBase64: string; // the whole recorded clip, base64
mediaType: string; // its MIME type — today always "audio/wav" (16 kHz, mono, 16-bit)
model: string; // the speech-to-text model id the user saved
language: string | null; // the user's language code, or null for auto-detect
}
returns: { text: string; durationSeconds?: number } | ModelProviderFailure
- Recorded, not live. The platform waits until the user stops talking, then calls you once with the whole clip. Plugin speech-to-text can't show words while the user is still speaking (the built-in Deepgram path can).
durationSecondsis what the user is billed for (rounded up); omit it and a flat 1 second is billed.- About 20 s per call.
- Vendors that want a file upload need
multipart/form-data. The sandbox has noBufferorTextEncoder, so build the body as bytes and send it withbodyEncoding: "base64"— OpenAI Voice shows how:
// infrastructure/plugins/openai-voice/index.ts
async transcribe(request, tools) {
const apiKey = await host.settings.getSecret("apiKey");
const boundary = "agentparley-openai-voice-boundary";
const extension = extensionFor(request.mediaType);
const preamble = asciiBytes(
`--${boundary}\r\n` +
`Content-Disposition: form-data; name="model"\r\n\r\n` +
`whisper-1\r\n` +
`--${boundary}\r\n` +
`Content-Disposition: form-data; name="file"; filename="audio.${extension}"\r\n` +
`Content-Type: ${request.mediaType}\r\n\r\n`,
);
const postamble = asciiBytes(`\r\n--${boundary}--\r\n`);
const body = base64Encode([...preamble, ...base64Decode(request.audioBase64), ...postamble]);
const response = await host.fetch({
method: "POST",
url: "https://api.openai.com/v1/audio/transcriptions",
headers: { Authorization: `Bearer ${apiKey}`, "Content-Type": `multipart/form-data; boundary=${boundary}` },
body,
bodyEncoding: "base64",
});
if (response.status !== 200) return mapFailure(response.status, tools);
const parsed = JSON.parse(response.body);
return { text: typeof parsed.text === "string" ? parsed.text : "" };
},
Failures
Return tools.fail(kind, message):
keyRejected,noCredit,modelNotFound,requestRejected→ the account's red provider banner.rateLimited,unavailable→ no banner; treated as a transient failure.
A thrown error or timeout counts as unavailable.
How a user sets it up
- Install your plugin.
- AI Providers → Add provider → "Your name from plugin", enter the key.
- Add a text-to-speech model and/or a speech-to-text model under it, with the model identifier your
vendor uses — it reaches you as
request.model. - On a companion, open the voice picker, choose your provider, and pick a voice. For listening, choose your speech-to-text model as the account default or for that companion. The mic button appears once a speech-to-text model is available.
Companion characters can suggest your voices by your kindId — see
Companion packs. For plugin providers, include model in the
suggestion.