Acting Challenge Evaluation · Stage
S4
Each take now shows its named prompt format and exact Instruction prompt above the audio player. Use Sort → Prompt format to group takes by format, or the Prompt format filter to show just one. Each player loads its MP3 when you press Play; the default order is challenge, condition, and seed.
The visible prompt preserves its original line breaks, parentheses, square-bracket durations, and ellipses. The model also received an outer <user_inst> wrapper with the reference mode, length budget, language, and a Text field; the displayed prompt is the Instruction part, not that full wrapper.
Content note: some takes involve fictional adult horror, distress, pain or screams; one challenge has a non-graphic adult theme. Use the content-warning filter if you prefer to avoid them. These provisional results have not yet received a human listening audit.
What this stage page measures
Each acting scene has a fixed three- or four-sentence transcript and three random seeds. Conditions C01–C11 vary the trained prompt surface: procedural GENERAL/SCRIPT, BUD-E-like captions, Timbre Whisper-style tags/prose/both, Voice Tagging Whisper-style tags/prose/both, Gemma-style English acting direction, and transcript-only with a reference. Newly authored acting prompts imitate these formats; they are not new Whisper or Gemma model annotations. Instruction-only and reference-conditioned takes are identified separately. References are MOSS-v2 codec reconstructions of pinned, licensed source audio, not pristine source waveforms.
WER uses Parakeet ASR; genuineness, vocal-burst blend, 40 emotion heads, and 57 VoiceNet attributes are model estimates. They should be compared with listening, not interpreted as ground truth. Empty ASR, near-silent audio, and failed alignments remain in the denominator. Whole-clip duration error is only a diagnostic; forced-aligned sentence MAE is computed only on successfully aligned sentences, so alignment-success and complete-timing-triplet fractions are reported alongside it. The separate ±20% duration-response control is applicable only to the 36 predeclared baseline takes with both controls. Requested burst types use a published five-class canonical crosswalk from fine detector subtypes. The dashboard separately reports requested-type F1, requested-type plus onset F1, and broader family plus onset F1, only on explicit C01 burst prompts. Unmapped detector labels remain false positives. No-burst C01 controls have their own false-positive rate. Historical substring burst F1 is diagnostic only. No composite reward is used to select checkpoints.
Target-aware results by prompting condition
Means include all seeds, failures, and both languages. Confidence intervals are challenge-cluster bootstrap intervals; they are very wide in a five-scene preview. Emotion hit rates use predeclared targets and a frozen corpus calibration; VoiceNet hit rates use predeclared attributes and fixed tolerances. Neither is a human rating.
All scorer dimensions
Raw stage-wide distributions below show what the scorers predicted across every take, including unrequested attributes. A high raw emotion or VoiceNet mean is not evidence that the prompt was followed; target-aware comparisons require the authored target and calibrated metric.
Filter results
Downloads and provenance
S4 per-take Parquet · JSONL · release manifest. Targets authored to emulate training surfaces are labeled as targets, not observed annotations.