Humaneness Voice Small · S10 · reference-conditioned CFG

Can guidance amplify S10 voice acting without breaking the words?

A paired prompt-and-seed test of classifier-free guidance (CFG) at strengths 1, 2, 3, 4 and 5. Every take uses the same final S10 checkpoint and one of the two provided voice references. Strength 1 is the ordinary conditional model; stronger values extrapolate away from an emotion-neutral prompt.

Loading status…

Why these prompts?

The completed S3 probe scored 5,544 takes. Across its 40 isolated emotion scenes, the revised GENERAL/SCRIPT C01 directions had the largest mean target-emotion uplift (+0.24 to +0.34 raw scorer units) but mean WER around 0.23–0.25. Compact CAPTION/TRANSCRIPT C03 directions gave a smaller uplift (+0.06 to +0.10) with WER around 0.11–0.15; C02 was between them in WER. An earlier verbose C01 canary also exposed instruction readout, though the revised full C01 prompts did not trip the conservative lexical screen. This S10 follow-up prioritizes caption variants with better text fidelity. Selection was exploratory—only three S3 seeds per exact cell—and no S10 outcome is inferred from S3.

The eight tested targets are Affection, Amusement, Anger, Fear, Elation, Awe, Disgust and Pleasure/Ecstasy. Each has two selected prompt recipes, two reference voices, three paired seeds and five guidance strengths: 480 planned takes. The original full 40-emotion, style and VoiceNet sweep is here; this page is a focused CFG follow-up, not a rerun of every S3 condition.

Selected acting recipes

EmotionRecipe ARecipe B
AffectionC03 extremeC03 direct
AmusementC03 directC02 extreme
AngerC02 directC03 direct
FearC02 extremeC03 direct
ElationC03 directC03 embodied
AweC03 embodiedC03 extreme
DisgustC03 embodiedC02 embodied
Pleasure/EcstasyC02 embodiedC03 extreme (low-WER boundary control)

C02 and C03 are newly authored captions in those training-style surfaces, not outputs from BUD‑E Whisper on these invented scripts. The C03 Pleasure/Ecstasy comparison had low WER on S3 but did not improve its emotion head; it is retained as an explicit boundary control, not called an S3 steering winner.

What CFG changes—and what stays fixed

For each 80-ms acoustic frame, the semantic transformer is evaluated once with the requested emotion and once with the same transcript in a neutral caption. The local Talker then predicts its 12 codec channels autoregressively in both branches. At each channel we combine logits as neutral + g × (emotional − neutral), sample one token, and feed that same token to both branches before the next channel. We do not guide the structural end/continue decision; it comes from the emotional branch. This matches the earlier SFT‑3 CFG method, adapted to the current 600M-backbone/large-Talker S10 model.

Both branches keep the exact spoken words, reference codec tokens, and token budget; only the acting caption changes. Guidance is an inference-time method and does not alter the checkpoint. We compare matched prompt, reference and seed at g=1. The two user-supplied references are chris-ref-enhanced.wav (7.51 s, stereo 48 kHz) and Fairy-2.wav (25.68 s, mono 44.1 kHz), codec-encoded at the model’s required 12-codebook depth. Strong g may increase emotion, but it may also raise WER, hurt naturalness, or trigger longer/unstable speech; results are shown without hiding those trade-offs.

Search, compare and listen

Players lazy-load individual MP3s from an immutable dataset revision. Choose an emotion and reference, then sort by guidance or a metric. The exact conditional and neutral instructions appear with each take.