Humaneness Voice Small · S3 · controlled prompt probe

How far can S3 be pushed by prompting alone?

Explore extreme acting direction for all 40 emotion heads, eight speaking styles, four focus-emotion/style interactions, and high-versus-low directions for 56 vocal VoiceNet axes. The 57th axis, content appropriateness, is reported as non-vocal and intentionally not probed with unsafe content.

Content note: some fictional adult scripts convey fear, anger, pain or non-explicit adult attraction. The perceived-young-age voice test uses an unrelated neutral script and never intersects with adult themes.

Loading release status…

Experiment contract

All audio, when present, comes from the same S3 step 32,337 checkpoint, the earlier provisional all-round choice. This is a new prompt probe, not a training run. Each comparison keeps the exact three spoken sentences and seeds (17, 29, 43) fixed. The three surfaces are C01 GENERAL/SCRIPT, C02 BUD-E Whisper 1.0-style CAPTION/TRANSCRIPT, and C03 BUD-E Whisper 1.1-style CAPTION/TRANSCRIPT. Each gets direct, embodied-actor and extreme-direction variants. Emotion-only cases also get a neutral acting-control prompt with the same transcript. These are newly authored, style-matched evaluation captions, not fresh BUD-E/Whisper annotations.

Prompted VoiceNet high/low pairs use the taxonomy’s direction, not a generic notion of “better”: higher AGEV means older; higher AROU means more aroused; higher BKGN means cleaner background. Age tests use a neutral, non-sexual three-sentence script. Technical axes such as recording quality, noise and esthetics are exploratory and should not be interpreted as actor controllability. EXPL measures content appropriateness, not voice, and is excluded from prompt generation while still identified in the 57-axis inventory. No reference audio is used, to isolate text-prompt effects.

The emotion×style grid is an intentional interference test: a few instructions conflict (for example, furious loud ranting plus whispering). Do not count those conflicting pairs as failures of a single-axis emotion or style instruction. The primary controllability tables below use only the isolated emotion and style cases.

Canary finding: an earlier C01 draft put long prose inside every sentence cue, and the model sometimes read that prose aloud (5/26 C01 takes flagged by a conservative ASR prompt-readout screen). The frozen prompt matrix now keeps detailed direction in GENERAL: and uses short inline cues such as (Anger). The final page reports prompt-readout rates separately by format and strength; this warning is not a claim that the revised prompts have already solved the problem.

Outcomes are model-estimated 40-head emotion and 57-axis VoiceNet values, Parakeet WER, genuineness and blend. The emotion comparison uses the same script’s neutral-prompt control; VoiceNet compares paired high and low instructions on the same text/seed. These are screening results, not proof of human-perceived emotion or style. Failed/empty takes remain visible. Scores must be weighed against intelligibility and listening.

Search and listen

Filter the exact prompts below by emotion, style, VoiceNet axis, surface, strategy, or seed. Sort by target scorer, WER, genuineness or blend. The full prompt is shown for every take, with its spoken transcript immediately below. Players fetch individual MP3s only when played.