Express Voice Blend

Design a personal voice from catalogue voices, your own clips, a gender and an age

Released Released 2026-10-01; age is approximate Voice v1.0.0

Version 1.0.0 · Updated 2026-10-01 · Express · Model 03 / 03

← Overview page

Four to six sentences from Cognitio, written on request. Generated text: the page is the reference.

On this page

Overview #

Express Voice Blend lets a person design a voice of their own and then speak with it. You choose one or more of the catalogue voices and how much of each, you may add a few short recordings of the voice you want it to lean towards — your own, or that of someone who has agreed — and you may set a gender and an approximate age. It gives back a voice code, a short piece of text that is the voice; send the code with any later line of text and the line comes back spoken in that voice. It is built for augmentative and alternative communication, where the voice a device speaks in stands for its user, and it was released on 2026-10-01: designed voices are spoken by a fine-tune that was measured against the unmodified base before release.

Express Voice Blend is a voice designer in front of a text-to-speech model. A voice, to SpeechT5, is a single 512-value speaker embedding, an x-vector; Express Voice stores seven of them and offers nothing in between. Express Voice Blend computes a new one for each design — a weighted blend of catalogue voices, mixed with the x-vectors of the caller’s clips, then moved along a gender axis and set to an age — and returns it, with a pitch and a rate, as a portable voice code of 1,387 characters. /v1/speak reads the code in place of a catalogue voice id and returns 16 kHz mono WAV. Because a designed voice can sit anywhere among the reference voices, it is spoken by its own fine-tune of SpeechT5, trained on several hundred audiobook readers so that it stays intelligible between and beyond the voices it has heard.

No voice is stored. A clip is read in memory, reduced to an embedding and dropped; it is not written to disk, logged or used for training, and neither is the embedding or the code. The one thing recorded is the consent attestation: when a design carries clips, a log line notes the account, the time and the number of clips, and nothing of the audio. The voice code is the only thing kept, and the customer keeps it: Falcon holds no list of designed voices and cannot return a code that has been lost.

It is the second member of the Voice family, beside Express Voice. Express Voice speaks in seven fixed voices chosen by id; Express Voice Blend speaks in a voice the user designed. Both answer on /v1/speak, and a voice code and a voice id cannot be sent together.

Intended use #

  • A speech-generating app or device whose user wants a voice that feels like their own rather than one of a few stock voices: a blend of voices they like, adjusted by gender and age until it fits.
  • Voice banking in a small way: a person who expects to lose their speech, or whose speech has changed, records up to five short clips, and the designed voice leans towards how they sound.
  • A donor voice given with consent: a family member or a friend records the clips for a user who cannot, and attests to it.
  • Giving each user of a shared communication app a distinct, stable voice that the app stores as one string beside their settings.

Out of scope #

  • Imitating a person without their consent. A clip is accepted only with the attestation that the speaker is the user or has agreed; using the route to copy the voice of someone who has not agreed is a misuse of it.
  • A faithful copy of a voice. The designed voice resembles the clips in pitch range and general colour; an x-vector does not carry an accent or a manner of speaking. How close the result sounds to the speaker has not been measured with listeners.
  • Children’s voices: the age control covers young adult to older adult, and nothing in the training or reference data is a child.
  • Languages other than English, singing, emotion or emphasis on request, and streaming. The limits of Express Voice apply to the speech itself.

Choose Express Voice Blend when #

  • The voice belongs to one person and should be theirs: they will design it once, keep the code, and speak with it every day.
  • A catalogue voice is close but not right, and a blend of two, a different age or a small change of pitch or pace would make it so.
  • Otherwise choose Express Voice when one of its seven fixed voices is enough: it is released, measured, and needs no code to be stored.

Architecture #

Parameters144,431,684 — the SpeechT5 TTS base count; the fine-tune keeps its size
Base modelmicrosoft/speecht5_tts — SpeechT5 in its text-to-speech configuration, MIT licence
Vocodermicrosoft/speecht5_hifigan — HiFi-GAN, MIT licence; log-mel spectrogram to 16 kHz waveform
Speaker encoderspeechbrain/spkrec-xvect-voxceleb (Apache-2.0): one 512-value x-vector per clip, computed at request time and discarded with the clip
A voiceOne unit-length 512-value x-vector: a weighted blend of catalogue voices, mixed with the clips’ x-vectors, then moved along a gender axis and an age setting
Attribute axesGender: the difference between the mean female and mean male reference voices. Age: approximate; a weak learned direction fitted on age-labelled speakers, applied together with a pitch and rate offset
Voice codeevb1. + 1,382 URL-safe characters (1,036 bytes): the embedding as 512 half-precision values, pitch, rate, three flags, CRC32
Clips1 to 5 per design; WAV, FLAC or OGG; 3 to 60 s each; 10,000,000 bytes in all; read in windows of at most 12 s; never stored
Prosody controlsSpeaking rate 0.75 to 1.33 and pitch -2 to +2 semitones, applied to the spectrogram at synthesis
Training corpusLibriTTS-R — restored LibriVox audiobook speech, CC BY 4.0, 24 kHz resampled to 16 kHz; 494 readers, at most 30 clips each
Recipelr 2e-5, batch 16 × 2, bf16, speaker-balanced sampler, one x-vector per clip, at most 12,000 steps, checkpoint chosen on held-out loss
Output16 kHz mono WAV
Text limit600 characters per request on /v1/speak; 200 for a design preview
RuntimeThe Express Voice GPU service (NVIDIA L4), beside the catalogue voices; it shuts down after four hours without calls

Express Voice Blend has three parts: a speaker encoder that reads clips, a designer that does arithmetic on speaker embeddings, and a text-to-speech model that speaks with the result.

The speaker encoder is the x-vector network speechbrain/spkrec-xvect-voxceleb, licensed Apache-2.0: the same network that produced the stored voices of Express Voice. Each clip is decoded, mixed to one channel, resampled to 16 kHz and read in windows of at most 12 seconds; each window gives a 512-value vector of unit length, and the windows of a clip, and then the clips, are averaged. Express Voice runs this network once per voice before training; here it runs at request time, and its output is the only thing that survives the clip.

The designer works on directions, because SpeechT5 normalises the speaker embedding it is given and only the direction matters. A blend is the weighted sum of the chosen catalogue voices, renormalised. The clips’ embedding is mixed in by clip_weight. Gender is an axis through the reference voices, the difference between the mean of the female and the mean of the male speakers: gender names an absolute position on it, and the designer moves the voice there. Age is approximate by design. The voice is moved along an age direction fitted on age-labelled speakers in the same way, and, because that direction separates older from younger speakers only weakly, the age is also applied as a speaking-rate change with a pitch offset of at most half a semitone — slower and slightly lower for older (rate 0.90, half a semitone down), slightly quicker and higher for young adult (rate 1.04, half a semitone up). settings.age_mode in the response reports embedding+prosody, and every design that sets an age carries the warning age_is_approximate. A gender or age step that would carry the voice away from every real speaker the designer knows is shortened, and the response says so.

The voice code is the result made portable: the embedding as 512 half-precision numbers, the pitch in tenths of a semitone, the rate in per cent, three flags (built from clips, attributes applied, an out-of-distribution warning) and a CRC32 checksum, 1,036 bytes written as evb1. and 1,382 URL-safe characters. The checksum catches a code that was cut short or altered in storage. It is not a signature: it shows that a code is intact, not where it came from.

The text-to-speech model is SpeechT5 in its text-to-speech configuration with the HiFi-GAN vocoder published beside it, as in Express Voice: characters in, log-mel spectrogram frames out one step at a time, conditioned on the speaker embedding at the decoder pre-net. The released SpeechT5 checkpoint and the two Express Voice fine-tunes have each seen a limited set of voices; this version fine-tunes the base on a broad set of readers so that an embedding between two speakers, or one computed from a clip of someone it never heard, is still read clearly. Rate and pitch are applied to the spectrogram before the vocoder: the frames are stretched in time for rate, and stretched and then resampled for pitch, which moves the resonances of the voice along with its pitch; that is why pitch is limited to two semitones either way and is best kept small.

What is deliberately absent: no store of voices or clips on the server; no speaker identification — the encoder’s output is used to speak, not to say who is speaking; no emotion or style control; no language beyond English; no streaming decoder.

Inputs & outputs #

Input #

FieldTypeRequiredDescriptionLimit
voicesarray of {id, weight}No/v1/voice/design. Catalogue voices to blend, by the ids that GET /v1/voices publishes, each id once. weight is 0 or more and defaults to 1; the weights are normalised, so only their proportions matter. A bare id string is accepted for {id, weight: 1}.1 to 8 voices
clipsarray of stringNo/v1/voice/design. Recordings of the voice to move towards, each a base64 string or a data: URL of a WAV, FLAC or OGG file. The speaker must be the user or a donor who has agreed. Each clip is read in memory, reduced to a speaker embedding and dropped.1 to 5 clips, 3 to 60 s each, 10,000,000 bytes in all once decoded
consentbooleanNo/v1/voice/design. Required, as the JSON literal true, whenever clips is sent: the attestation that the speaker in every clip is the user, or has agreed to this use of their voice. Without it the clips are not read and the call answers 400 consent_required.must be true with clips
clip_weightnumberNo/v1/voice/design. The share of the clips against the blended catalogue voices when both are sent: 0 keeps the blend, 1 keeps the clips.0 to 1; default 0.5
genderstring or numberNo/v1/voice/design. Where the voice should sit between the female and male reference voices: female, neutral or male, or a number from -1 (female) through 0 (neutral) to 1 (male). An absolute target, not a nudge; omitted, the voice is left as the blend and clips made it.female, neutral, male, or -1 to 1
agestring or numberNo/v1/voice/design. An approximate age for the voice: young-adult (25), adult (35), middle-aged (50) or older (68), or a number of years. Adults only: there are no children’s voices in this version.one of four names, or 18 to 80
pitchnumberNo/v1/voice/design and /v1/speak. A small pitch shift in semitones applied when the voice speaks. It moves the colour of the voice with the pitch, so a lower setting also sounds somewhat more male and a higher one more female; use gender for a larger change. Set at design time it is stored in the voice code; sent with a voice code on /v1/speak it overrides the stored value for that call.-2 to 2; default 0
ratenumberNo/v1/voice/design and /v1/speak. The speaking rate, where 1 is unchanged, 0.75 is slower and 1.33 is faster. Stored in the voice code and overridable per call, like pitch.0.75 to 1.33; default 1
preview_textstringNo/v1/voice/design. A short line to hear the designed voice at once: when present, the response carries it spoken as a WAV.200 characters
voice_codestringNo/v1/speak. The voice code a design returned, unchanged, in place of a catalogue voice; sending both answers 400 voice_conflict. The code holds the whole voice; nothing is looked up on the server.1,387 characters
textstringNo/v1/speak. The English text to speak in the designed voice, under the same rules as Express Voice: whitespace is collapsed, and the service speaks it in phrases of at most 12 words.600 characters

A design request names what to build the voice from. Every field is optional, but at least one of voices, clips, gender and age must be present; with only a gender or an age, the design starts from the centre of the reference voices. A request that blends two catalogue voices, leans towards one clip of the user and asks for an older voice, with a line to hear at once:

json
{
  "voices": [
    {"id": "female-1", "weight": 0.6},
    {"id": "male-2", "weight": 0.4}
  ],
  "clips": ["UklGRiQAAABXQVZFZm10…"],
  "clip_weight": 0.5,
  "consent": true,
  "gender": "neutral",
  "age": "older",
  "preview_text": "This is how I will sound."
}

Limits:

  • voices holds 1 to 8 entries, each id once, from the catalogue that GET /v1/voices publishes. Weights are normalised; at least one must be above 0.
  • clips holds 1 to 5 recordings as base64 or data: URLs, WAV, FLAC or OGG, mono or stereo, at a standard sample rate from 8 kHz to 96 kHz (8, 11.025, 12, 16, 22.05, 24, 32, 44.1, 48, 88.2 or 96 kHz), each 3 to 60 seconds long, 10,000,000 bytes in all once decoded; the whole body is at most 16,000,000 bytes. A browser recording in WebM or Opus is not read: record or convert to one of the three formats. A clip that is silent, shorter than 3 seconds or longer than 60 is refused with 400 bad_clip, and detail says why. Speech recorded close to the microphone in a quiet room works best; music, several speakers or a distant microphone give an embedding unlike any reference voice, and the response then carries the warning clip_out_of_distribution.
  • consent must be the JSON literal true whenever clips is present. It is checked first: without it no clip is read.
  • gender is female, neutral or male, or a number from -1 to 1. age is young-adult, adult, middle-aged or older, or a number of years from 18 to 80.
  • pitch is -2 to 2 semitones and rate is 0.75 to 1.33; both are stored in the voice code.
  • preview_text is at most 200 characters.

Speaking with the result is a /v1/speak request that carries the code in place of a voice id:

json
{
  "text": "I would like a glass of water, please.",
  "voice_code": "evb1.RVZCAQ…"
}

Output #

FieldTypeDescription
voice_codestring/v1/voice/design. The designed voice: evb1. followed by 1,382 URL-safe characters that carry the 512-value speaker embedding, the pitch and rate, three flags and a checksum. The customer stores it; it is the only thing kept, and it is not kept by Falcon.
summaryobject/v1/voice/design. What the design was made from, as the request stated it: voices with their weights, clips (the count), clip_seconds (the length read from each clip, when clips were sent), gender and age (null when not sent).
similararray of {id, similarity}/v1/voice/design. Up to three catalogue voices nearest the result — nearest the clips, when clips were sent — best first, with a similarity from -1 to 1 measured around the centre of the reference voices.
settingsobject/v1/voice/design. What the designer applied: gender and age as numbers (or null), age_mode (embedding+prosody, or prosody when no age direction is loaded), the pitch and rate stored in the code, clip_weight, attribute_strength (below 1 when a gender or age step was shortened) and gender_score (where the result sits, -1 female to 1 male).
warningsarray of string/v1/voice/design. Empty, or any of clip_out_of_distribution, attributes_limited and age_is_approximate. A warning never withholds the voice code.
preview_wav_base64string/v1/voice/design, only when preview_text was sent. The preview line as a base64-encoded RIFF/WAVE file, 16 kHz, one channel, with its length in preview_duration_seconds.
wav_base64string/v1/speak. The text spoken in the designed voice, in the Express Voice envelope: engine is express-voice, voice is custom, and duration_seconds is the length of the audio.

A design answers with the voice code and a description of what was built. An illustrative response — the shape is fixed; the code is abbreviated here and is always 1,387 characters:

json
{
  "ok": true,
  "engine": "express-voice-blend",
  "voice_code": "evb1.RVZCAQ…",
  "summary": {
    "voices": [
      {"id": "female-1", "weight": 0.6},
      {"id": "male-2", "weight": 0.4}
    ],
    "clips": 1,
    "clip_seconds": [12.4],
    "gender": "neutral",
    "age": "older"
  },
  "similar": [
    {"id": "female-1", "similarity": 0.62},
    {"id": "male-2", "similarity": 0.41},
    {"id": "female-3", "similarity": 0.18}
  ],
  "settings": {
    "gender": 0,
    "age": 68,
    "age_mode": "embedding+prosody",
    "pitch": -0.5,
    "rate": 0.9,
    "clip_weight": 0.5,
    "attribute_strength": 1,
    "gender_score": 0.03
  },
  "warnings": ["age_is_approximate"],
  "preview_wav_base64": "UklGR…",
  "preview_duration_seconds": 2.1
}

voice_code is the voice. Store it exactly as it came, in a JSON field or a settings file, and send it back unchanged in a request body; a code that has been trimmed or edited fails its checksum and answers 400 bad_voice_code. Its characters are URL-safe, but keep it out of URLs, logs and analytics: when clips were used, the code is a voiceprint of the speaker, which could be matched against other recordings of that person, and it should be stored as biometric personal data. A code is not tied to the account that designed it: anyone who holds a code and a key can speak in that voice.

similar helps a user understand the result: the catalogue voices it most resembles, which is also a way to find the catalogue voices nearest to a person’s own recording before blending. settings reports what was applied, which can differ from what was asked: an age given as years comes back as that number, a pitch that the age offset moved comes back moved, and attribute_strength below 1 means a gender or age step was shortened to keep the voice among real speakers. The warnings are clip_out_of_distribution (the clips gave a voice unlike the reference speakers: heavy noise, music or tones), attributes_limited (a gender or age step was shortened) and age_is_approximate (the age was applied as a pitch and rate offset). None of them withholds the code.

/v1/speak with a voice code answers in the Express Voice envelope: engine is express-voice, voice is custom, and wav_base64 decodes to a RIFF/WAVE file at 16 kHz with one channel.

Training #

Two things are fitted: the weights that speak designed voices, and the axes the designer moves a voice along. No recordings of users are used for either, and no clip sent to /v1/voice/design is ever used for training: clips are not kept.

The weights are a full fine-tune of SpeechT5 TTS on LibriTTS-R (Koizumi et al., Interspeech 2023), a restored version of the LibriTTS corpus of LibriVox audiobook speech, released under CC BY 4.0, recorded at 24 kHz and resampled to 16 kHz for the base. The selection is deliberately broad and long: 494 readers (242 female, 252 male) from its clean training sets, 13,138 clips in all, at most 30 clips per reader, each clip 3 to 14 seconds and 10 to 40 words, with a speaker-balanced sampler so that no reader dominates a batch and one x-vector per clip rather than one per reader, so the model sees the natural spread of each voice. The recipe is the one that produced the six Hi-Fi voices of Express Voice: learning rate 2e-5, batch 16 with two-step gradient accumulation, bf16 precision, at most 12,000 steps on one NVIDIA A100 40 GB on Vertex AI, with the checkpoint chosen on a held-out split. The readers of the corpus’s development set are held out entirely and serve as unseen voices in evaluation. The job ran on 2026-10-01 for the full 12,000 steps in 86 minutes; the checkpoint that serves is the one with the lowest held-out loss, 0.394 on 133 clips, reached at step 5,000. The 37 held-out readers took no part in training.

The length of the clips is the lesson of an earlier attempt. A sibling of Express Voice fine-tuned on 109 speakers with short clips sounded natural and lost its place on sentences longer than about ten words, and was withdrawn; this fine-tune uses many speakers and long clips together.

The gender axis is fitted, in seconds and without a GPU, from speaker averages that already exist: the mean x-vectors of the 494 readers the fine-tune trained on, of the 109 speakers of the CSTR VCTK Corpus (CC BY 4.0), with their recorded gender, and of the seven catalogue voices, 610 reference speakers in all. Only the direction, the two class means and summary statistics are kept. The age direction is fitted on 1,725 speakers of the GLOBE corpus (CC0) who state an age band, from its test and validation sets; only the direction and the target of each age band are kept, never those speakers’ vectors. It is weak: on 340 speakers held out of the fit it separates older from younger speakers with a score of 0.61, where 0.5 is chance and 1 is perfect. That is why age is also applied as a rate and pitch offset and is always reported as approximate.

What it is not trained on: any recording of a user or of a clip donor, any child’s voice, any language other than English, or conversational and expressive speech. LibriTTS-R is read audiobook speech, mostly North American English.

Evaluation #

MetricValueSource
Word error rate, all designed voices (200 utterances, open recogniser)0.45 %blend_eval.models.finetune.overall.wer in the run's eval/results.json (Whisper small.en transcripts vs the input text), job express-voice-blend-20261001-143700
Word error rate by group — readers trained on / blends of readers / readers never heard / gender-shifted0.0 % / 0.09 % / 2.05 % / 0.0 %blend_eval.models.finetune.groups (40 / 80 / 40 / 40 utterances) in the same file
Utterances more than half wrong1 of 200 (a reader never heard)utterances_over_0.5 in the same file
Unmodified SpeechT5 base on the same set1.94 % overall, 6.53 % on readers never heard, 3 of 200 more than half wrongblend_eval.models.base in the same file
Faithfulness — cosine between the x-vector of the rendered speech and the embedding asked for0.968 mean (0.948 lowest)speaker_cosine_mean / speaker_cosine_min in the same file
Gender control — shifted voices that land on the requested side20 of 20gender_flip_rate / gender_flip_n in the same file
Held-out spectrogram loss (133 clips, best checkpoint)0.394eval_loss in the export's meta.json
Age direction — older against younger speakers, held-out (0.5 is chance)0.61heldout_age_auc in the fitted design file's summary (340 held-out GLOBE speakers)

The evaluation was fixed before the run. It renders a fixed set of sentences — short requests of the kind a communication app speaks, and long sentences of twenty words and more — in four groups of voices: readers the fine-tune trained on, blends of two and three readers, readers it never heard, and voices moved along the gender axis. Intelligibility is the word error rate of an open speech recogniser (Whisper small.en) against the input text, reported overall and per group, with a count of utterances that loop or stop early. Faithfulness is the cosine between the x-vector of the rendered speech and the embedding that was asked for. Control is whether the measured pitch and the gender-axis position of the output move as the gender setting moves. The same set is rendered with the unmodified SpeechT5 base, so that the fine-tune is released only if it is the better speaker of designed voices.

Result (2026-10-01, 200 utterances, 40 voices): the fine-tune misheard 0.45 % of words against 1.94 % for the base. By group: readers it trained on 0.0 %, blends of two and three readers 0.09 %, readers it never heard 2.05 % (base 6.53 %), gender-shifted voices 0.0 %. One utterance of 200 came out more than half wrong, in the voice of a reader it never heard; the base had three. The rendered speech matched the requested embedding with a mean cosine of 0.968 (base 0.965), and all 20 gender-shifted voices landed on the requested side. On that evidence the fine-tune speaks designed voices. One result goes the other way: on long sentences spoken in a single piece, which the service does not do (it speaks in phrases of at most 12 words), the fine-tune misheard 6.0 % of words and the base none.

Known gaps #

Every voice used in that evaluation is an audiobook reader or a blend of readers. Clips recorded by users of communication devices — on a telephone or tablet microphone, in a noisy room, or by a speaker whose speech is weak or dysarthric — have not been tested, and such clips may give embeddings unlike the reference voices; the clip_out_of_distribution warning is specified for that case and is itself unmeasured. The age direction is weak (0.61 on held-out speakers, where 0.5 is chance), and what it does to the rendered voice has not been measured. Voices unlike audiobook readers are the weaker case: a small local check on speakers of another corpus was cut short and showed no advantage over the base. No listening study with people who use augmentative communication has been run.

API #

EndpointAuthBody limitDescription
POST /v1/voice/designBearer preview key (fln_) or session token (fls_)16,000,000 bytes; clips ≤ 10,000,000 bytes decodedCatalogue voices, clips (with consent: true), gender, age, pitch and rate → a voice code, the nearest catalogue voices and, when asked, a spoken preview. Clips are processed in memory and not kept.
POST /v1/speakBearer preview key (fln_) or session token (fls_)text ≤ 600 charactersWith voice_code in place of voice: speak one text in the designed voice and return the WAV. Optional rate and pitch override the values stored in the code.
GET /v1/voicesBearer preview key (fln_) or session token (fls_)—The catalogue of voices a design can blend, and design: whether the designer is available on the voice service.

Route /v1/voice/design · quota bucket audio · body limit 16 MB.

Express Voice Blend is public on the Falcon API at POST /v1/voice/design, and its voices are spoken through POST /v1/speak with a voice_code. Both count against the audio preview quota bucket: one unit for each design that returns a code, with or without a preview, and one unit for each line spoken. The design body is limited to 16 MB. The design field of GET /v1/voices says whether the designer is available on the voice service. The base URL, envelope, authentication, rate limits and retry guidance are documented in API conventions.

The routes are keyed. There is no keyless route for this model: the public playground does not offer voice design, and it does not speak a voice code, because recordings of a voice are not accepted from anonymous visitors and a voice built from them is not spoken for one. The portal’s sandbox, for signed-in organisation users, runs the design route.

A worked example #

1. Design. Send the voices, the clips with the attestation, and the attributes:

http
POST /v1/voice/design HTTP/1.1
Authorization: Bearer fln_…
Content-Type: application/json
json
{
  "voices": [
    {"id": "female-1", "weight": 0.6},
    {"id": "male-2", "weight": 0.4}
  ],
  "clips": ["UklGRiQAAABXQVZFZm10…"],
  "consent": true,
  "age": "older",
  "preview_text": "This is how I will sound."
}
json
{
  "ok": true,
  "engine": "express-voice-blend",
  "voice_code": "evb1.RVZCAQ…",
  "summary": {
    "voices": [{"id": "female-1", "weight": 0.6}, {"id": "male-2", "weight": 0.4}],
    "clips": 1,
    "clip_seconds": [12.4],
    "gender": null,
    "age": "older"
  },
  "similar": [{"id": "female-1", "similarity": 0.62}, {"id": "male-2", "similarity": 0.41}],
  "settings": {"gender": null, "age": 68, "age_mode": "embedding+prosody", "pitch": -0.5, "rate": 0.9, "clip_weight": 0.5, "attribute_strength": 1, "gender_score": -0.21},
  "warnings": ["age_is_approximate"],
  "preview_wav_base64": "UklGR…",
  "preview_duration_seconds": 2.1
}

2. Store the code. Play the preview to the user; when they are happy with it, save voice_code with their profile. It is the only copy: Falcon does not keep it, and a design cannot be fetched again. To change the voice, design again and replace the code.

3. Speak. Send the code with each line, in place of voice:

json
{
  "text": "I would like a glass of water, please.",
  "voice_code": "evb1.RVZCAQ…"
}
json
{
  "ok": true,
  "engine": "express-voice",
  "voice": "custom",
  "duration_seconds": 2.6,
  "wav_base64": "UklGR…"
}

rate and pitch may be sent beside a voice code on /v1/speak to override the stored values for one call — a slower rate for a listener who needs it, say — within the same ranges as at design time. They apply to voice codes only.

Metering and errors #

A design counts one audio unit once it has returned a code; a spoken line counts one, as on Express Voice. Validation errors, 401, 402, 429 and every 503 do not count.

StatusCodeMeaning
400bad_jsonThe body is not valid JSON or is not an object.
400nothing_to_blendA design was sent with none of voices, clips, gender and age.
400bad_voicevoices is not a list of at most eight {id, weight} entries, an id is not in the catalogue or appears twice, or no weight is above 0; detail says which.
400consent_requiredclips was sent without consent: true. The clips were not read.
400bad_clipA clip is not base64, is not a WAV, FLAC or OGG file that can be decoded, has more than two channels or an unusual sample rate, is shorter than 3 s, longer than 60 s or silent, there are more than five, or clip_weight is not a number from 0 to 1; detail says which.
413clips_too_largeThe clips are over 10,000,000 bytes in all once decoded, or the body is over 16,000,000 bytes.
400bad_gendergender is not female, neutral, male or a number from -1 to 1.
400bad_ageage is not one of the four names or a number from 18 to 80.
400bad_pitchpitch is not a number from -2 to 2.
400bad_raterate is not a number from 0.75 to 1.33.
400bad_preview_textpreview_text is not a string of at most 200 characters, or cannot be spoken.
400bad_voice_code/v1/speak: voice_code is not a voice code, or it is cut short, altered or from another version.
400voice_conflict/v1/speak: voice and voice_code were both sent.
400text_required/v1/speak: text is missing or empty after trimming.
413text_too_long/v1/speak: text exceeds 600 characters.
401invalid_credentialsMissing, malformed or revoked bearer credential.
402payment_requiredThe credential’s owner is not in good standing with the preview.
403key_requiredA voice code was sent to the keyless playground route; designed voices are spoken with a key.
429quota_exceededThe audio bucket for the current UTC calendar month is exhausted; kind is audio.
502design_failedThe designer failed after a valid request; safe to retry once.
503design_warmingThe voice service is loading after a quiet period, about half a minute; Retry-After: 30 (10 while its one instance is busy). On /v1/speak the same state answers speak_warming.
503design_unavailableThe voice service has no designer: it is not configured, its designer did not load, or it predates this model. On /v1/speak with a voice code the same answer is given rather than speech in another voice.

GET /v1/voices carries the designer’s state beside the catalogue: "design": {"available": true, "model": "express-voice-blend"}. An application can read it once to decide whether to offer voice design.

Runtime & deployment #

KindGPU service
ResidentOn the Express Voice GPU service (one NVIDIA L4): the SpeechT5 weights that speak designed voices, the x-vector speaker encoder that reads clips, the attribute axes and the HiFi-GAN vocoder
ServingBehind the Falcon API routes /v1/voice/design and /v1/speak (with voice_code); /v1/voices reports whether the designer is available
Cold startShares the Express Voice service, which shuts down after four hours without calls; the call that wakes it is answered 503 design_warming or speak_warming with Retry-After: 30 while the models load
ConcurrencyOne design or one utterance per request
Timeout45 s in the reference client, then one retry

Express Voice Blend runs inside the Express Voice GPU service: one container on Cloud Run with one NVIDIA L4, the same one that holds the catalogue voices, fronted by the Falcon API. No second GPU is used. The service loads the weights that speak designed voices, the speaker encoder and the attribute axes when it starts, beside the two Express Voice fine-tunes, and a failure to load the designer leaves the catalogue voices serving.

  • Service contract. The service takes POST /v1/voice/design and POST /v1/speak with a voice_code, and reports the designer on GET /v1/voices and GET /health. The Falcon API checks the request — the attestation first — forwards only the fields the route knows, and passes the answer on.
  • Loading. The service shuts down after four hours without calls. The call that wakes it is answered 503 design_warming (or speak_warming on /v1/speak) with Retry-After: 30 while the models load; the reference client waits 2.5 s and retries once, and gives up on a request after 45 s.
  • What a design costs. Decoding the clips and computing their embeddings is a matter of a second or so; the arithmetic is negligible; a preview adds one synthesis. None of this has been timed on the serving hardware for this version.
  • Privacy. Clips exist only inside the request that carries them: decoded in memory, reduced to an embedding, dropped. Neither the API nor the service writes a clip, an embedding, a voice code or the text to disk or to its logs; the log line for a design carries counts and durations only.
  • Selection. A voice_code on /v1/speak selects the designed voice and this model’s weights; a voice id selects a catalogue voice and Express Voice’s. The two cannot be combined in one request.
  • What the host provides. The recording step and the consent that goes with it; a place to keep each user’s voice code; the text; and a player for the WAV. The host owns the decision to offer voice design at all, and to whom.

Limits & safety #

Express Voice Blend hears only the clips it is sent, for the length of one request, and reads only the text it is asked to speak. It does not know who the speaker in a clip is, whether they agreed, or who will hear the result: the attestation is the caller’s statement, and the model cannot verify it.

  • It does not verify consent. consent: true is required with clips and is an attestation by the caller, who is responsible for having it; the route is keyed so that every design belongs to an account.
  • It does not store audio, embeddings or voice codes, and it cannot recover a lost code or list the voices an account has designed.
  • It does not copy a voice faithfully. An x-vector carries pitch range and general timbre; the result resembles the clips. How close it sounds to the speaker has not been measured with listeners; the output follows the requested embedding closely, so it must not be relied on to be distinguishable from the speaker, and it may match automated speaker verification.
  • It does not make children’s voices: age covers young adult to older adult, and age is approximate throughout.
  • It does not sign voice codes. The checksum shows a code is intact, not that the design route produced it; a holder of a key who can compute an x-vector elsewhere can construct a code without the attestation. The speaking route is keyed for that reason.
  • It does not promise that every designed voice is clear. Voices far from the reference speakers — from noisy clips, or pushed to the end of an axis — are the weaker case, are flagged with a warning, and are measured only for audiobook readers in this version.
  • It does not speak languages other than English, read more than 600 characters per request, or take emotion or emphasis instructions.
  • It does not know what it is saying: any text within the limit is spoken, in the designed voice, so the host is responsible for the content it sends.

Out of scope. Express Voice Blend is not for impersonation, fraud, or any use where a listener would be led to believe a particular real person is speaking. It is not a medical device, and a designed voice is not a clinical voice-banking product: it has not been evaluated with people who use augmentative communication. The readers of LibriTTS-R and the speakers of VCTK did not record for this purpose beyond the terms of their corpora’s licences; designed voices are not to be presented as theirs.

Fixed weights per version; the model does not learn from requests.

Versions #

VersionDateStatusNote
1.0.0ReleasedFirst release: the voice designer, the voice code and the routes; a SpeechT5 fine-tune on 494 LibriTTS-R readers speaks designed voices (word error rate 0.45 %, base 1.94 %).

Version numbers follow the series convention: a major bump changes the contract — the request fields, the response shape, the sample rate or the voice code, so that codes from an earlier major version stop being read; a minor bump is a retrain or a new attribute axis with the same contract, after which a stored code still speaks but may sound somewhat different, because the code holds the embedding and the weights turn it into sound; a patch bump changes metadata or runtime only.

Current version: 1.0.0, released 2026-10-01: the designer, the voice code (version 1, prefix evb1.), the routes, and the fine-tune that speaks designed voices, with the evaluation above. Later versions will be recorded here and in the changelog will record it.

Base SpeechT5 TTS and HiFi-GAN vocoder under MIT; speaker encoder under Apache-2.0; LibriTTS-R, the fine-tuning corpus, and the CSTR VCTK Corpus, whose speaker averages fit the gender axis, are CC BY 4.0, attributed here and in the Express series notes; weights are not distributed during the private preview.