Humaneness Voice Small · S10 · reference-conditioned CFG
Can guidance amplify S10 voice acting without breaking the words?
A paired prompt-and-seed test of classifier-free guidance (CFG) at strengths 1, 2, 3, 4 and 5. Every take uses the same final S10 checkpoint and one of the two provided voice references. Strength 1 is the ordinary conditional model; stronger values extrapolate away from an emotion-neutral prompt.
Why these prompts?
The completed S3 probe scored 5,544 takes. Across its 40 isolated emotion scenes, the revised GENERAL/SCRIPT C01 directions had the largest mean target-emotion uplift (+0.24 to +0.34 raw scorer units) but mean WER around 0.23–0.25. Compact CAPTION/TRANSCRIPT C03 directions gave a smaller uplift (+0.06 to +0.10) with WER around 0.11–0.15; C02 was between them in WER. An earlier verbose C01 canary also exposed instruction readout, though the revised full C01 prompts did not trip the conservative lexical screen. This S10 follow-up prioritizes caption variants with better text fidelity. Selection was exploratory—only three S3 seeds per exact cell—and no S10 outcome is inferred from S3.
The eight tested targets are Affection, Amusement, Anger, Fear, Elation, Awe, Disgust and Pleasure/Ecstasy. Each has two selected prompt recipes, two reference voices, three paired seeds and five guidance strengths: 480 planned takes. The original full 40-emotion, style and VoiceNet sweep is here; this page is a focused CFG follow-up, not a rerun of every S3 condition.
Selected acting recipes
| Emotion | Recipe A | Recipe B |
|---|---|---|
| Affection | C03 extreme | C03 direct |
| Amusement | C03 direct | C02 extreme |
| Anger | C02 direct | C03 direct |
| Fear | C02 extreme | C03 direct |
| Elation | C03 direct | C03 embodied |
| Awe | C03 embodied | C03 extreme |
| Disgust | C03 embodied | C02 embodied |
| Pleasure/Ecstasy | C02 embodied | C03 extreme (low-WER boundary control) |
C02 and C03 are newly authored captions in those training-style surfaces, not outputs from BUD‑E Whisper on these invented scripts. The C03 Pleasure/Ecstasy comparison had low WER on S3 but did not improve its emotion head; it is retained as an explicit boundary control, not called an S3 steering winner.
What CFG changes—and what stays fixed
For each 80-ms acoustic frame, the semantic transformer is evaluated once with the requested emotion and once with the same transcript in a neutral caption. The local Talker then predicts its 12 codec channels autoregressively in both branches. At each channel we combine logits as neutral + g × (emotional − neutral), sample one token, and feed that same token to both branches before the next channel. We do not guide the structural end/continue decision; it comes from the emotional branch. This matches the earlier SFT‑3 CFG method, adapted to the current 600M-backbone/large-Talker S10 model.
Both branches keep the exact spoken words, reference codec tokens, and token budget; only the acting caption changes. Guidance is an inference-time method and does not alter the checkpoint. We compare matched prompt, reference and seed at g=1. The two user-supplied references are chris-ref-enhanced.wav (7.51 s, stereo 48 kHz) and Fairy-2.wav (25.68 s, mono 44.1 kHz), codec-encoded at the model’s required 12-codebook depth. Strong g may increase emotion, but it may also raise WER, hurt naturalness, or trigger longer/unstable speech; results are shown without hiding those trade-offs.
Matched outcome by guidance
The target column is the requested Empathic Insight emotion head (raw regression units, not a probability). WER is from Parakeet v3. Genuineness and vocal-burst blend are model estimates. Every delta uses the same emotion, prompt recipe, reference and seed at g=1. Model estimates are screening signals; listen to the audio before choosing a production strength.
What the 480 completed takes show
CFG steers the emotion scorer, but quality does not improve monotonically. Against the exactly matched g=1 takes, g=2 raises the requested emotion head by +0.049 raw units, with WER 0.159→0.166 and blend 3.558→3.038. g=5 has the largest average steering effect, +0.400, but WER rises to 0.208 and blend falls to 2.101; the count of WER-above-0.2 takes rises from 17/96 to 27/96. g=3 and g=4 fall between these points. Genuineness alone does not capture the audible/intelligibility trade-off.
There is no universal best strength. The automatic conservative guardrail below selects g=1 because every tested g>1 reduces mean blend by more than its allowed 0.5 units. g=2 is the most plausible exploratory setting if a listener accepts the moderate blend loss; g=4–5 are useful as high-intensity, per-emotion experiments, not a global default. Amusement and Elation show strong g=5 target-head gains, whereas Fear and Disgust do not show a reliable gain in this small eight-emotion panel. The Chris and Fairy reference splits also differ substantially. Each emotion/strength cell has only 12 takes; these are screening results, not proof of perceived acting quality. Listen to matched takes before adopting any setting.
Search, compare and listen
Players lazy-load individual MP3s from an immutable dataset revision. Choose an emotion and reference, then sort by guidance or a metric. The exact conditional and neutral instructions appear with each take.