Humaneness Voice Small · Acting Challenge Evaluation
How well does each stage follow an acting prompt?
Listen to generated takes, inspect the exact prompt and transcript, and compare scorer estimates across the S1–S10 training stages. Scores are imperfect proxies; use the audio and full metrics together.
Explore the new S3 extreme-prompt probe: 40 emotions, speaking styles, and VoiceNet high/low directions. Its page clearly marks whether audio and scores have been released.
Compare S10 reference-conditioned CFG strengths: matched emotional versus neutral caption branches at guidance 1–5, with two reference voices and direct MP3 players after scoring.
Content note: this benchmark intentionally includes intense fictional adult acting, including horror, distress, pain and screams, plus one non-graphic adult-theme scene. Clearly illegal and underage sexual scenarios are excluded. Public intermediate results are automatically checked but have not yet received a human listening audit.
Qwen3-0.6B semantic backbone · fresh SFT-3-width Talker · MOSS Audio Tokenizer v2
Release counts and revision load from the release manifest.
A dataset commit pin will be added with the first results.
Explore stages
Loading stage inventory…
Stage comparison
Same challenges and conditions across stages; every seed remains in the denominator. WER is lower-is-better. Genuineness, blend, and target-aware emotion/VoiceNet hit rates are automatic estimates, not human judgments. Burst realization uses a published five-class requested-type crosswalk (for example, “surprised_gasp” counts as the requested “gasp”), separately reports broader acoustic-family matching, and excludes the detector’s explicit “no_burst” negative class. Burst headline means use only explicit C01 requests; no-burst controls have their own false-positive rate. These burst measures and sentence-duration MAE appear only after independent postprocessing. MAE covers successful alignments only, so its success fraction is displayed beside it; the ±20% timing-response slope likewise has a complete-triplet coverage column (ideal slope near 1, not simply as large as possible). Cells show mean ± SD and n for applicable takes, plus bootstrap intervals when available. Stage-to-stage claims should use paired differences and confidence intervals, not just these raw means.
Release status
Evaluation results are being prepared. Stage pages will appear here as reviewed data is released.
How to read this evaluation
The protocol selects each stage checkpoint using held-out validation loss. Completed stage pages identify their selected checkpoint and metrics. The evaluation pairs the same challenges, prompt conditions, and seeds across stages. Challenge transcripts stay fixed; the prompt surface changes by condition.
Automatic measures such as WER, duration error, emotion estimates, genuineness, burst realization, and VoiceNet attributes describe scorer predictions. They do not establish that a take is objectively natural or that a speaker truly feels an emotion.
Native audio players link to individual MP3s in the pinned dataset revision and fetch audio only when played. Per-stage Parquet files, immutable manifests, licenses, and scorer versions are available from the dataset snapshot. Read the frozen detailed protocol and exact benchmark manifests.
Conclusions and practical prompting recommendations
There is no checkpoint that wins every outcome. For a provisional all-round listening and generation default, start with S3: it has the highest mean genuineness (1.636/6) and the highest mean composite acting reward on identical, score-complete takes (0.4628; 1,651 matched takes). This is a useful starting point, not proof that S3 sounds best: its reward lead over S4, S8 and S10 is small and the challenge-bootstrap intervals include zero. Listen before choosing a production checkpoint.
Choose the objective explicitly: S5 has the lowest overall WER (0.1130); S7 has the highest generic vocal-burst blend score (4.606/10); S10 has the highest requested-emotion intensity-band hit rate (0.2074); and S2 has the highest requested burst type plus onset F1 (0.0529 on 186 explicit requests). The last number is still very low. The lowest held-out validation loss is actually S4 (4.6966), not S10 (4.7278). Validation loss selected the best checkpoint within each stage; it does not guarantee the best acting or audibly natural speech.
Checkpoint decision table
Each stage has the same 1,872 challenge/condition/seed slots. Reward is a descriptive, target-aware 0–1 scalar with a WER gate; it is not a human preference score. Its availability varies when sentence alignment fails. The common-take column restricts every stage to the same 1,651 non-missing reward slots. “Burst event F1” measures requested type and timing, unlike the generic blend score.
Loading audited checkpoint table…
Which prompt surface works?
For S3, the highest descriptive reward among caption-style conditions is C10, the Gemma-style English acting instruction (0.505/150); the lowest WER is C05, Timbre prose (0.0437); the highest blend is C07, Voice Tagging tags only (5.934/10); and the highest emotion-band hit among C02–C11 is C04, Timbre tags only (0.212). At S10, C07 has the highest caption-style reward (0.496); C09, Voice Tagging tags plus prose, combines low WER (0.0473) with the highest caption-style emotion-band hit (0.234). These small differences are descriptive, not a causal format ranking: reference assignments differ, C11 always has a reference but no style caption, and the BUD-E/Timbre/Voice Tagging/Gemma-style texts here were hand-authored to resemble their training surfaces, not newly generated by those annotation models. C01 activates additional burst/timing reward components and loses some reward rows to failed alignment (S3: 200/300 available), so its scalar reward is not directly comparable with caption-only conditions.
For ordinary intelligible speech, try a concise caption plus an exact TRANSCRIPT: "…" first; at S3 C05 is a useful WER-oriented candidate, while C10 is a useful reward-oriented candidate. At S10, try C09 or C07 depending on whether transcription/emotion or blend matters more. Use C11 only when the reference should carry the voice and explicit style control is unnecessary. Use C01 GENERAL:/SCRIPT: when you need sentence timing or a specified vocal burst; its WER is much higher (S3: 0.348, S10: 0.313) and exact burst timing is not yet reliable. Inspect the same scene and seed across formats in the stage pages before relying on an average.
Loading S3 and S10 condition tables�¦
Does reference audio help?
In the seven caption conditions that mix reference and instruction mode (C02/C03/C05/C06/C08/C09/C10), reference-conditioned takes have lower descriptive WER: S3 0.0516 versus 0.0856 without reference, and S10 0.0567 versus 0.0930. That matches the listening hypothesis, but the reference and no-reference groups contain different challenge assignments; other metrics do not improve consistently. The cleaner paired test is C01, which renders the same prompt, scene and seed both ways. There, adding reference changes mean WER by −0.023 at S3 and +0.002 at S10 (reference minus no-reference), with challenge-bootstrap intervals spanning zero. We therefore recommend reference audio for identity and as a practical WER candidate, but cannot claim a universal causal quality improvement from this benchmark.
Was the detailed GENERAL/SCRIPT syntax actually sent?
Yes for the elements visible in the frozen C01 prompt: all 3,000 published C01 takes across ten stages exactly match the manifest, and all have a parenthesized delivery cue at the start of every spoken sentence. Half of the 50 challenges have square-bracket sentence durations (1,500 of 3,000 C01 takes). Twenty-five challenges explicitly request a parenthesized vocal burst, at the onset of a designated sentence. The stage pages previously put the exact prompt inside a collapsed details panel and collapsed its line breaks into an HTML paragraph; the updated viewer preserves those line breaks. The displayed prompt is the Instruction field, not the entire model input: the evaluator also wraps it in <user_inst> with the reference marker, frame budget, language and a Text field that repeats the SCRIPT.
The benchmark is nevertheless not fully guide-faithful. It has zero explicit [N.N seconds pause] tags, whereas the home-folder guide describes measured pauses. It repeats an emotion-and-delivery cue before every sentence, rather than using the guide’s short later reminders. Requested bursts occur only at a designated sentence onset, not at arbitrary in-sentence positions, and every request uses (label, 0.7 seconds), longer than the guide’s typical detected bursts (median 0.28 s; 90th percentile 0.48 s). These choices could hurt burst realization but have not been isolated experimentally. The benchmark already contains ... hesitation in 34/50 scripts (2,040/3,000 C01 takes); it was not absent, but there is no paired with-versus-without-ellipsis ablation. Its 25/50 timed C01 split also differs from the guide’s 70% timed training mixture.
Next controlled test — proposed, not yet run
Use the provisional S3 checkpoint on 20 balanced German/English scenes, three seeds, four C01 variants and both reference modes (480 takes): existing prompt; add plausible explicit pause tags only; shorten burst requests to a plausible 0.2–0.5 s only; and combine both changes. Keep the spoken words, scene, seed and reference assignment matched. Add a small separate punctuation-only ... versus comma test, scoring hesitation and naturalness as well as punctuation-normalized WER. Re-run Parakeet WER, genuineness, blend, requested burst type-plus-onset, emotion/VoiceNet target hits and sentence timing, and listen blind to matched takes. No benefit is claimed before this ablation runs.
Source of this report: the pinned 18,720-take release, frozen 50-scene manifest, held-out checkpoint selection, target-aware aggregate analysis, and an independently regenerated reward/prompt audit. All comparisons remain automatic and provisional pending human listening.