Overview #
Express Voice Blend lets a person design a voice of their own and then speak with it. You choose one or more of the catalogue voices and how much of each, you may add a few short recordings of the voice you want it to lean towards — your own, or that of someone who has agreed — and you may set a gender and an approximate age. It gives back a voice code, a short piece of text that is the voice; send the code with any later line of text and the line comes back spoken in that voice. It is built for augmentative and alternative communication, where the voice a device speaks in stands for its user, and it was released on 2026-10-01: designed voices are spoken by a fine-tune that was measured against the unmodified base before release.
Express Voice Blend is a voice designer in front of a text-to-speech model. A voice, to SpeechT5, is a single 512-value speaker embedding, an x-vector; Express Voice stores seven of them and offers nothing in between. Express Voice Blend computes a new one for each design — a weighted blend of catalogue voices, mixed with the x-vectors of the caller’s clips, then moved along a gender axis and set to an age — and returns it, with a pitch and a rate, as a portable voice code of 1,387 characters. /v1/speak reads the code in place of a catalogue voice id and returns 16 kHz mono WAV. Because a designed voice can sit anywhere among the reference voices, it is spoken by its own fine-tune of SpeechT5, trained on several hundred audiobook readers so that it stays intelligible between and beyond the voices it has heard.
No voice is stored. A clip is read in memory, reduced to an embedding and dropped; it is not written to disk, logged or used for training, and neither is the embedding or the code. The one thing recorded is the consent attestation: when a design carries clips, a log line notes the account, the time and the number of clips, and nothing of the audio. The voice code is the only thing kept, and the customer keeps it: Falcon holds no list of designed voices and cannot return a code that has been lost.
It is the second member of the Voice family, beside Express Voice. Express Voice speaks in seven fixed voices chosen by id; Express Voice Blend speaks in a voice the user designed. Both answer on /v1/speak, and a voice code and a voice id cannot be sent together.
Intended use #
- A speech-generating app or device whose user wants a voice that feels like their own rather than one of a few stock voices: a blend of voices they like, adjusted by gender and age until it fits.
- Voice banking in a small way: a person who expects to lose their speech, or whose speech has changed, records up to five short clips, and the designed voice leans towards how they sound.
- A donor voice given with consent: a family member or a friend records the clips for a user who cannot, and attests to it.
- Giving each user of a shared communication app a distinct, stable voice that the app stores as one string beside their settings.
Out of scope #
- Imitating a person without their consent. A clip is accepted only with the attestation that the speaker is the user or has agreed; using the route to copy the voice of someone who has not agreed is a misuse of it.
- A faithful copy of a voice. The designed voice resembles the clips in pitch range and general colour; an x-vector does not carry an accent or a manner of speaking. How close the result sounds to the speaker has not been measured with listeners.
- Children’s voices: the age control covers young adult to older adult, and nothing in the training or reference data is a child.
- Languages other than English, singing, emotion or emphasis on request, and streaming. The limits of Express Voice apply to the speech itself.
Choose Express Voice Blend when #
- The voice belongs to one person and should be theirs: they will design it once, keep the code, and speak with it every day.
- A catalogue voice is close but not right, and a blend of two, a different age or a small change of pitch or pace would make it so.
- Otherwise choose Express Voice when one of its seven fixed voices is enough: it is released, measured, and needs no code to be stored.
Specification #
| Parameters | 144,431,684 — the SpeechT5 TTS base count; the fine-tune keeps its size |
|---|---|
| Base model | microsoft/speecht5_tts — SpeechT5 in its text-to-speech configuration, MIT licence |
| Vocoder | microsoft/speecht5_hifigan — HiFi-GAN, MIT licence; log-mel spectrogram to 16 kHz waveform |
| Speaker encoder | speechbrain/spkrec-xvect-voxceleb (Apache-2.0): one 512-value x-vector per clip, computed at request time and discarded with the clip |
| A voice | One unit-length 512-value x-vector: a weighted blend of catalogue voices, mixed with the clips’ x-vectors, then moved along a gender axis and an age setting |
| Attribute axes | Gender: the difference between the mean female and mean male reference voices. Age: approximate; a weak learned direction fitted on age-labelled speakers, applied together with a pitch and rate offset |
| Voice code | evb1. + 1,382 URL-safe characters (1,036 bytes): the embedding as 512 half-precision values, pitch, rate, three flags, CRC32 |
| Clips | 1 to 5 per design; WAV, FLAC or OGG; 3 to 60 s each; 10,000,000 bytes in all; read in windows of at most 12 s; never stored |
| Prosody controls | Speaking rate 0.75 to 1.33 and pitch -2 to +2 semitones, applied to the spectrogram at synthesis |
| Training corpus | LibriTTS-R — restored LibriVox audiobook speech, CC BY 4.0, 24 kHz resampled to 16 kHz; 494 readers, at most 30 clips each |
| Recipe | lr 2e-5, batch 16 × 2, bf16, speaker-balanced sampler, one x-vector per clip, at most 12,000 steps, checkpoint chosen on held-out loss |
| Output | 16 kHz mono WAV |
| Text limit | 600 characters per request on /v1/speak; 200 for a design preview |
| Runtime | The Express Voice GPU service (NVIDIA L4), beside the catalogue voices; it shuts down after four hours without calls |
Try it #
- You send
- A blend of the voices female-1 and male-2, one short recording of the user reading aloud, sent with their consent, and the age set to older.
- You get back
- A voice code of 1,387 characters. Sent with any later line of text, it returns that line spoken in the designed voice.
The same exchange as the API sees it:
{
"voices": [
{"id": "female-1", "weight": 0.6},
{"id": "male-2", "weight": 0.4}
],
"clips": ["UklGRiQAAABXQVZFZm10…"],
"clip_weight": 0.5,
"consent": true,
"gender": "neutral",
"age": "older",
"preview_text": "This is how I will sound."
}{
"ok": true,
"engine": "express-voice-blend",
"voice_code": "evb1.RVZCAQ…",
"summary": {
"voices": [
{"id": "female-1", "weight": 0.6},
{"id": "male-2", "weight": 0.4}
],
"clips": 1,
"clip_seconds": [12.4],
"gender": "neutral",
"age": "older"
},
"similar": [
{"id": "female-1", "similarity": 0.62},
{"id": "male-2", "similarity": 0.41},
{"id": "female-3", "similarity": 0.18}
],
"settings": {
"gender": 0,
"age": 68,
"age_mode": "embedding+prosody",
"pitch": -0.5,
"rate": 0.9,
"clip_weight": 0.5,
"attribute_strength": 1,
"gender_score": 0.03
},
"warnings": ["age_is_approximate"],
"preview_wav_base64": "UklGR…",
"preview_duration_seconds": 2.1
}Limits & safety #
Express Voice Blend hears only the clips it is sent, for the length of one request, and reads only the text it is asked to speak. It does not know who the speaker in a clip is, whether they agreed, or who will hear the result: the attestation is the caller’s statement, and the model cannot verify it.
- It does not verify consent.
consent: trueis required with clips and is an attestation by the caller, who is responsible for having it; the route is keyed so that every design belongs to an account. - It does not store audio, embeddings or voice codes, and it cannot recover a lost code or list the voices an account has designed.
- It does not copy a voice faithfully. An x-vector carries pitch range and general timbre; the result resembles the clips. How close it sounds to the speaker has not been measured with listeners; the output follows the requested embedding closely, so it must not be relied on to be distinguishable from the speaker, and it may match automated speaker verification.
- It does not make children’s voices: age covers young adult to older adult, and age is approximate throughout.
- It does not sign voice codes. The checksum shows a code is intact, not that the design route produced it; a holder of a key who can compute an x-vector elsewhere can construct a code without the attestation. The speaking route is keyed for that reason.
- It does not promise that every designed voice is clear. Voices far from the reference speakers — from noisy clips, or pushed to the end of an axis — are the weaker case, are flagged with a warning, and are measured only for audiobook readers in this version.
- It does not speak languages other than English, read more than 600 characters per request, or take emotion or emphasis instructions.
- It does not know what it is saying: any text within the limit is spoken, in the designed voice, so the host is responsible for the content it sends.
Related models #
Express Voice
Turns a short line of English text into speech, in a choice of clear male and female voices.
Text to speech in seven English voices, male and female: two SpeechT5 fine-tunes
Express Cue
Describe a sound in words and get a short audio clip back.
Renders a short sound cue from a text description through a deterministic synthesis engine
Latest versions #
| Version | Date | Status | Note |
|---|---|---|---|
| 1.0.0 | Released | First release: the voice designer, the voice code and the routes; a SpeechT5 fine-tune on 494 LibriTTS-R readers speaks designed voices (word error rate 0.45 %, base 1.94 %). |
Read the full documentation
Nine chapters: architecture, inputs and outputs, training, evaluation, API, runtime, limits and versions.