Voice profiles

Voice profiles. Each chain is a single cloned voice by construction, so there is no speaker identity to verify and no conversion to apply. They are here to show the same rules on material where identity is not in question.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken. Underlined descriptors are the ones that change across the chain; brackets inside the words are a real non-speech sound. The perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
BKGN — background noise levelcloned voice   strict_113 · #1

This is a VoiceNet dimension, not an emotion: background noise level (BKGN) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with background noise level (BKGN) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it at the very top of the range at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 1.00.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

5 clips · 45 s · de · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/100/100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.1 dB
per-clip
per-clip buttons play:
rule VN1k 5qmax 1.000cmax 0.250d_a 1.000d_b 1.000dataset vprof_vclang detotal 44.8schain gain +4.1 dBseam step 2.1 dBcrossfades 150/100/100/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult feminine voice · fairly smooth, wide pitch range
(affection, bitterness, sourness · normal-paced, highly aroused, neutral tension, ranting) Ich werde die Härte in dir nehmen, diesen steinigen Kern, und ihn durch ein lebendiges Herz ersetzen. (contented sigh) (shriek) Ein Herz, in dem Christus wirklich in dir herrscht. (contented sigh)
full caption & clip details
An adult feminine voice; delivery is highly aroused, normal-paced, neutral tension, variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, wide pitch range, normal breath; affect is deeply negative, very dominant, vulnerable; reads as affection, bitterness, sourness; style: ranting; average recording, quiet background; mildly explicit content; genuineness 1.2/6; vocal-burst blend 3.7/10; 16.6s, DE.
emolia_c0362__V__BKGN__extremely_low__de.c004.k0 · in -25.9 dBFS · gain +4.1 dB · vprof_vc-00038
(doubt, helplessness, disgust · brisk, energised, neutral tension, ranting) I just don't know what concrete plans these two clubs have for a joint team yet. (fast breathing) It's clear though, they both desperately want to strengthen their player lists.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, little disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as doubt, helplessness, disgust; style: ranting, playful; average recording, quiet background; mildly explicit content; genuineness 1.4/6; vocal-burst blend 2.4/10; 6.4s, EN.
emolia_c0362__E__Helplessness__B__en.c008.k2 · in -23.9 dBFS · gain +4.1 dB · vprof_vc-00038
(pride, contempt, triumph · fast, energised, slightly relaxed, dramatic) Endlich sind die schwarzen Pixel von meinen Miniaturbildern weg. (ahem) (guffaw) Ich kann meine Bilder jetzt richtig gut sehen.
full caption & clip details
A young adult feminine voice; delivery is energised, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as pride, contempt, triumph; style: dramatic, playful; good recording, quiet background; genuineness 1.2/6; vocal-burst blend 2.3/10; 5.6s, DE.
emolia_c0362__E__Relief__D__de.c034.k2 · in -22.7 dBFS · gain +4.1 dB · vprof_vc-00038
(shame, distress, disgust · brisk, energised, slightly tense, ranting) I feel so ashamed seeing how the vibrant green jungle just stops against that bleak, dusty expanse west of the peaks. It’s such a harsh, unforgivable difference.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, fairly guarded; reads as shame, distress, disgust; style: ranting, dramatic; good recording, no background noise; mildly explicit content; genuineness 0.4/6; vocal-burst blend 1.6/10; 8.1s, EN.
emolia_c0362__E__Shame__D__en.c045.k1 · in -22.6 dBFS · gain +4.1 dB · vprof_vc-00038
(awe, affection, thankfulness gratitude · brisk, normally alert, slightly relaxed, casual) For those of you new to the deep end of this digital life, nem was designed as a welcoming doorway into the whole world of blockchain. (deep breath) Think of it as the perfect little place to begin exploring all the curves and depths.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as awe, affection, thankfulness gratitude; style: casual, ranting; good recording, no background noise; mildly explicit content; genuineness 0.4/6; vocal-burst blend 0.6/10; 8.6s, EN.
emolia_c0362__P__explicit__en.c037.k3 · in -24.2 dBFS · gain +4.1 dB · vprof_vc-00038
BKGN — background noise levelcloned voice   strict_114 · #2

This is a VoiceNet dimension, not an emotion: background noise level (BKGN) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with background noise level (BKGN) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it at the very top of the range at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 1.00.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

5 clips · 46 s · de · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150/100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 4.3 dB
per-clip
per-clip buttons play:
rule VN1k 5qmax 1.000cmax 0.250d_a 1.000d_b 1.000dataset vprof_vclang detotal 45.7schain gain +4.3 dBseam step 4.3 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, wide pitch range
(disappointment, disgust, fatigue exhaustion · brisk, highly aroused, tense, dramatic) (contented sigh) Diese neuen haben unten am Gehäuse eine kleine Kerbe, damit du den Stecker genau an die Leiste schieben kannst. (sharp (contented sigh) whistle) Das macht die ganze Einrichtung so viel ordentlicher.
full caption & clip details
A young adult feminine voice; delivery is highly aroused, brisk, tense, variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is deeply negative, very dominant, vulnerable; reads as disappointment, disgust, fatigue exhaustion; style: dramatic, casual; below-average recording, quiet background; genuineness 1.9/6; vocal-burst blend 3.1/10; 12.2s, DE.
emolia_c0362__V__AROU__very_high__de.c011.k1 · in -25.4 dBFS · gain +4.3 dB · vprof_vc-00038
(disgust, shame, anger · fast, energised, neutral tension, ranting) Nachdem ich den ganzen Tag Ausrüstung transportiert habe, (slow breathing) wenn ich einen professionellen Trockner für Gebäude brauche, schaue ich bei Bosch, DeWalt, Makita, Metabo oder Steinel, weil ich mir kein weiteres Stück Schrott leisten kann, das mich fertig macht.
full caption & clip details
A young adult feminine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, no disfluency, wide pitch range, light breath; affect is negative, slightly dominant, guarded; reads as disgust, shame, anger; style: ranting, dramatic; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 5.8/10; 10.9s, DE.
emolia_c0362__E__Fatigue_Exhaustion__D__de.c008.k1 · in -25.6 dBFS · gain +4.3 dB · vprof_vc-00038
(bitterness, longing, distress · brisk, energised, neutral tension, ranting) Die Art, wie er meine Energie nutzt, sie einfach länger brennen lässt, lässt mich nach mehr dieser treibenden Hitze sehnen. (chuckle) Ich sehne mich nach dem reinen Umfang des Wertes, den mein Einsatz ihm bringt.
full caption & clip details
An adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, wide pitch range, light breath; affect is negative, slightly dominant, guarded; reads as bitterness, longing, distress; style: ranting, cartoonish; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 4.4/10; 6.5s, DE.
emolia_c0362__E__Sexual_Lust__C__de.c007.k3 · in -25.7 dBFS · gain +4.3 dB · vprof_vc-00038
(shame, distress, disgust · brisk, energised, slightly tense, dramatic) I feel so ashamed seeing how the vibrant green jungle just stops against that bleak, dusty expanse west of the peaks. It’s such a harsh, unforgivable difference.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, fairly guarded; reads as shame, distress, disgust; style: dramatic, ranting; good recording, no background noise; mildly explicit content; genuineness 0.4/6; vocal-burst blend 1.4/10; 8.1s, EN.
emolia_c0362__E__Shame__D__en.c045.k0 · in -21.5 dBFS · gain +4.3 dB · vprof_vc-00038
(brisk, normally alert, slightly relaxed, dramatic) Studies suggest these things are common factors in how someone develops a same-sex attraction. The research points to a few key areas that seem to play a role.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: dramatic, playful; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.7/10; 8.5s, EN.
emolia_c0362__V__BKGN__extremely_low__en.c019.k1 · in -24.5 dBFS · gain +4.3 dB · vprof_vc-00038
CHNK — chunking / phrasing densitycloned voice   strict_115 · #3

This is a VoiceNet dimension, not an emotion: chunking / phrasing density (CHNK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with chunking / phrasing density (CHNK) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it at the very top of the range at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 1.00.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

5 clips · 43 s · de · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.1 dB
per-clip
per-clip buttons play:
rule VN1k 5qmax 1.000cmax 0.250d_a 1.000d_b 1.000dataset vprof_vclang detotal 42.6schain gain +3.3 dBseam step 3.1 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an elderly masculine voice · fairly smooth, no background noise
(confusion, intoxication altered states of consciousness, pain · slow, lethargic, fully relaxed, ASMR) Dieser Essiggeruch, der trifft dich, oder? Als ob er die Konturen verschwimmen lässt, das Metall... instabil wirken lässt.
full caption & clip details
An elderly masculine voice; delivery is lethargic, slow, fully relaxed, steady; timbre is slightly cool, dark, fairly smooth, thin; slurred, frequent disfluency, narrow pitch range, breathless; affect is neutral, submissive, vulnerable; reads as confusion, intoxication altered states of consciousness, pain; style: ASMR, whispered; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.0/10; 12.5s, DE.
emolia_c2356__E__Intoxication_Altered_States_of_Consciousness__C__de.c027.k1 · in -23.3 dBFS · gain +3.3 dB · vprof_vc-00142
(doubt, confusion, affection · measured, energised, slightly relaxed, dramatic) I need to talk to the victim's best friend about this. (spitting) Maybe they saw something important.
full caption & clip details
An adult masculine voice; delivery is energised, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as doubt, confusion, affection; style: dramatic, playful; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.6/10; 5.2s, EN.
emolia_c2356__V__ATCK__extremely_low__en.c039.k3 · in -24.1 dBFS · gain +3.3 dB · vprof_vc-00142
(awe, distress, helplessness · normal-paced, normally alert, slightly relaxed, monologue) Diese endlose Reihe von Eichen verspottet uns und verschwimmt die Grenze, wo der Komfort endet und meine wache Anwesenheit beginnt. Diese riesige Glasfläche, eine Farce der Offenheit, lässt mich jeden einzelnen Schatten sehen.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, little disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, distress, helplessness; style: monologue, formal; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 3.3/10; 14.9s, DE.
emolia_c2356__E__Malevolence_Malice__C__de.c021.k0 · in -25.7 dBFS · gain +3.3 dB · vprof_vc-00142
(jealousy and envy, longing, affection · normal-paced, normally alert, slightly relaxed, casual) Seize this one day, friend, because soon you'll hear another's late-night woes. (surprised gasp) You'll see their blemishes, and endure their nightly coughing fit.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as jealousy and envy, longing, affection; style: casual, storytelling; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 3.4/10; 3.9s, EN.
emolia_c2356__V__AGEV__moderately_high__en.c011.k1 · in -23.4 dBFS · gain +3.3 dB · vprof_vc-00142
(triumph, elation, relief · brisk, normally alert, slightly relaxed, conversational) My body is finally achieving perfect equilibrium now that the systems are aligning and the vital energies are stabilizing. (cackle) This profound metabolic shift is the triumph I have worked toward.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as triumph, elation, relief; style: conversational, casual; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 2.3/10; 6.6s, EN.
emolia_c2356__E__Triumph__B__en.c007.k3 · in -20.2 dBFS · gain +3.3 dB · vprof_vc-00142
Interestcloned voice   strict_125 · #4

This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Interest essentially absent — 0.00, virtually no clip in this corpus scores lower — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 1.00.

Nothing was asked of the other axis, and in fact Amusement drifts down from 1.00 to 0.73 (-0.27), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

5 clips · 42 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.7 dB
per-clip
per-clip buttons play:
rule B1k 5qmax 0.999cmax 0.250d_a -0.267d_b 0.999dataset vprof_vclang entotal 41.6schain gain +5.0 dBseam step 1.7 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · fairly smooth, wide pitch range
(amusement, astonishment surprise, fatigue exhaustion · brisk, energised, neutral tension, playful) For instance, an unmarried mother gets no elder support or compensation for career setbacks related to the temporary search for a husband in Frankfurt am Main concerning the child.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, no audible breath; affect is positive, slightly submissive, neutral openness; reads as amusement, astonishment surprise, fatigue exhaustion; style: playful, casual; below-average recording, some background noise; contains vocal bursts: Breathy Giggle; genuineness 2.3/6; vocal-burst blend 4.2/10; 0.7s, EN.
emolia_c0645__V__WARM__extremely_low__en.c037.k1 · in -27.7 dBFS · gain +5.0 dB · vprof_vc-00050
(relief, astonishment surprise, disgust · measured, highly aroused, neutral tension, playful) It was just a shove, nothing like a punch, not from the side or behind. No fancy attack like that happened at all. (low mumble)
full caption & clip details
A young adult somewhat masculine voice; delivery is highly aroused, measured, neutral tension, variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; slurred, some disfluency, wide pitch range, normal breath; affect is negative, slightly dominant, vulnerable; reads as relief, astonishment surprise, disgust; style: playful, casual; below-average recording, quiet background; genuineness 0.8/6; vocal-burst blend 0.0/10; 11.9s, EN.
emolia_c0645__V__VALN__extremely_low__en.c030.k1 · in -26.0 dBFS · gain +5.0 dB · vprof_vc-00050
(teasing, disgust, impatience and irritability · brisk, energised, neutral tension, playful) Seriously, you can just trash specific post revisions, keeping only the ones you actually want. (wolf whistle) It's super lightweight, dead simple to use, and plugs right into the WordPress backend. (contented sigh)
full caption & clip details
An adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, slightly guarded; reads as teasing, disgust, impatience and irritability; style: playful, casual; good recording, no background noise; mildly explicit content; genuineness 0.6/6; vocal-burst blend 0.1/10; 10.2s, EN.
emolia_c0645__V__S_RANT__moderately_high__en.c006.k0 · in -25.2 dBFS · gain +5.0 dB · vprof_vc-00050
(astonishment surprise, teasing, contempt · brisk, normally alert, slightly relaxed, ranting) So, Riseup, another one of those United States email providers, they were actually compelled to hand over data to the authorities. It was just another example of the pressure they were under.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as astonishment surprise, teasing, contempt; style: ranting, playful; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.5/10; 8.7s, EN.
emolia_c0645__V__S_FORM__moderately_high__en.c019.k2 · in -24.4 dBFS · gain +5.0 dB · vprof_vc-00050
(interest, awe, hope enthusiasm optimism · normal-paced, energised, slightly relaxed, narration) We're diving deep into the gear that defines some of the biggest guitar and bass players out there. (relief sigh) This series showcases the hand tools that have shaped legendary sounds across music.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, awe, hope enthusiasm optimism; style: narration, playful; good recording, no background noise; mildly explicit content; genuineness 0.3/6; vocal-burst blend 0.0/10; 10.7s, EN.
emolia_c0645__V__S_NEWS__very_high__en.c006.k1 · in -24.4 dBFS · gain +5.0 dB · vprof_vc-00050
Interestcloned voice   strict_126 · #5

This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Interest essentially absent — 0.00, virtually no clip in this corpus scores lower — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 1.00.

Nothing was asked of the other axis, and in fact Amusement drifts down from 1.00 to 0.75 (-0.25), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

5 clips · 32 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/100/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.5 dB
per-clip
per-clip buttons play:
rule B1k 5qmax 0.999cmax 0.250d_a -0.250d_b 0.999dataset vprof_vclang entotal 31.9schain gain +4.2 dBseam step 2.5 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · fairly smooth
(amusement, teasing, pleasure ecstasy · fast, highly aroused, very tense, casual) Oh, so dog pneumonia, you say? The most obvious sign is just this ridiculous, hacking cough they keep doing. Seriously, it sounds like a tiny, furry opera singer who's just had too much kibble.
full caption & clip details
A young adult masculine voice; delivery is highly aroused, fast, very tense, volatile; timbre is slightly cool, slightly bright, fairly smooth, thin; very slurred, no disfluency, extreme pitch range, heavy breath; affect is elated, slightly dominant, guarded; reads as amusement, teasing, pleasure ecstasy; style: casual, playful; average recording, quiet background; contains vocal bursts: Childlike Giggle; genuineness 2.6/6; vocal-burst blend 4.6/10; 1.0s, EN.
emolia_c0645__X__ga_amuse_laugh__en.c027.k0 · in -26.8 dBFS · gain +4.2 dB · vprof_vc-00048
(sourness, impatience and irritability · brisk, energised, slightly relaxed, casual) We build the tech for platform linking and product handling, but the underwriter always holds the final call on claims and risk. (slow breathing) It's just... all procedure, you know.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as sourness, impatience and irritability; style: casual, playful; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.1/10; 6.2s, EN.
emolia_c0645__E__Emotional_Numbness__B__en.c036.k2 · in -24.2 dBFS · gain +4.2 dB · vprof_vc-00048
(elation, astonishment surprise, pleasure ecstasy · brisk, energised, slightly tense, ranting) Can you believe this? The Essen premiere of that ridiculous cooperative project with the Opéra national du Rhin Strasbourg is set for Saturday, March fourteenth, at seven in the evening. I just can't stomach it.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as elation, astonishment surprise, pleasure ecstasy; style: ranting, casual; good recording, no background noise; mildly explicit content; genuineness 0.6/6; vocal-burst blend 3.2/10; 7.9s, EN.
emolia_c0645__E__Disgust__D__en.c015.k0 · in -24.3 dBFS · gain +4.2 dB · vprof_vc-00048
(astonishment surprise, awe, pleasure ecstasy · brisk, energised, slightly relaxed, storytelling) Es ist unglaublich, wirklich. Jahre später tauchte der Eisaufstrich auf, eine brillante Abwandlung von gefrorenem Saft, völlig anders als bloßes Eis.
full caption & clip details
An adult feminine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as astonishment surprise, awe, pleasure ecstasy; style: storytelling, ranting; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.2/10; 8.1s, DE.
emolia_c0645__E__Awe__C__de.c018.k3 · in -23.6 dBFS · gain +4.2 dB · vprof_vc-00048
(interest, astonishment surprise, awe · brisk, energised, neutral tension, playful) (breathy giggle) Wait, we actually found this massive international site to search and book hotels? (contented sigh) (snort) I can pick out the perfect places and see all the prices saved!
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, vulnerable; reads as interest, astonishment surprise, awe; style: playful, dramatic; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 2.0/10; 9.2s, EN.
emolia_c0645__E__Astonishment_Surprise__B__en.c006.k0 · in -24.5 dBFS · gain +4.2 dB · vprof_vc-00048
Interestcloned voice   strict_127 · #6

This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Interest essentially absent — 0.00, virtually no clip in this corpus scores lower — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 1.00.

Nothing was asked of the other axis, and in fact Fatigue Exhaustion drifts down from 1.00 to 0.69 (-0.31), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

5 clips · 42 s · de · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150/150/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 4.0 dB
per-clip
per-clip buttons play:
rule B1k 5qmax 0.999cmax 0.250d_a -0.311d_b 0.999dataset vprof_vclang detotal 42.4schain gain +3.1 dBseam step 4.0 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, good recording
(fatigue exhaustion, relief, distress · slow, very low-energy, relaxed, monologue) (contented sigh) Da diese Verstopfung schon so lange anhält, (deep breath) musst du wirklich abschwellende Tropfen und Sprays benutzen.
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, relaxed, variable; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, vulnerable; reads as fatigue exhaustion, relief, distress; style: monologue; good recording, quiet background; genuineness 2.1/6; vocal-burst blend 7.1/10; 10.1s, DE.
emolia_c2249__V__ARSH__extremely_low__de.c009.k3 · in -25.3 dBFS · gain +3.1 dB · vprof_vc-00134
(intoxication altered states of consciousness, fatigue exhaustion, helplessness · measured, normally alert, slightly relaxed, storytelling) Westösterreich sieht mehr vom A H drei N zwei Virus. (growl) Ostösterreich hat auch die A H eins N eins Pandemie null neun, mit weniger B-Fällen.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, fatigue exhaustion, helplessness; style: storytelling, narration; good recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.2/10; 9.2s, DE.
emolia_c2249__V__ATCK__very_high__de.c005.k0 · in -21.3 dBFS · gain +3.1 dB · vprof_vc-00134
(astonishment surprise, awe, longing · normal-paced, highly aroused, slightly relaxed, monologue) The bonds we forged were deeper than I ever imagined. Saying goodbye to such fierce friendships really hurt.
full caption & clip details
An adult masculine voice; delivery is highly aroused, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, very thin; clear, no disfluency, wide pitch range, light breath; affect is negative, slightly dominant, fairly guarded; reads as astonishment surprise, awe, longing; style: monologue, storytelling; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 3.3/10; 3.1s, EN.
emolia_c2249__E__Pain__A__en.c020.k3 · in -21.8 dBFS · gain +3.1 dB · vprof_vc-00134
(disappointment, relief, doubt · normal-paced, normally alert, slightly relaxed, ranting) Kannst du glauben, dass die Umweltgruppen es endlich geschafft haben, das Berufungsgericht dazu zu bringen, dass die Emissionen nach dem Export wichtig sind? Es ist fast zum Verrücktwerden, wie lange es gedauert hat, bis sie das anerkannt haben, was wir schon immer wussten.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, relief, doubt; style: ranting, playful; good recording, no background noise; mildly explicit content; genuineness 0.9/6; vocal-burst blend 2.3/10; 14.0s, DE.
emolia_c2249__E__Jealousy_and_Envy__C__de.c023.k3 · in -23.9 dBFS · gain +3.1 dB · vprof_vc-00134
(interest, hope enthusiasm optimism, jealousy and envy · brisk, energised, neutral tension, conversational) I'm so impressed by how Chris Brogan and Pat Flynn use videos to genuinely get people onto their lists; it's just brilliant stuff. (contented sigh) Their strategies are truly inspiring.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, hope enthusiasm optimism, jealousy and envy; style: conversational, casual; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.6/10; 6.6s, EN.
emolia_c2249__E__Thankfulness_Gratitude__C__en.c002.k0 · in -22.7 dBFS · gain +3.1 dB · vprof_vc-00134
Contemplation ↓  /  Embarrassmentcloned voice   strict_089 · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Embarrassment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Embarrassment essentially absent — 0.04, lower than 96 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.96.

At the same time Contemplation goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.09 (lower than 91 % of clips in this corpus), a change of -0.90. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.25, then +0.25, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

5 clips · 42 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 100/100/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.8 dB
per-clip
per-clip buttons play:
rule PXRk 5qmax 0.898cmax 0.250d_a -0.898d_b 0.960dataset vprof_vclang entotal 42.4schain gain +5.7 dBseam step 0.8 dBcrossfades 100/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, balanced body, good recording, no background noise, light breath
(contemplation, fear, interest · brisk, normally alert, neutral tension, casual) What if architecture isn't just stone, but the way we interact with it, the actions that give it meaning? I fear that if we lose that human relationship, the whole cultural structure collapses.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as contemplation, fear, interest; style: casual, ranting; good recording, no background noise; mildly explicit content; genuineness 1.9/6; vocal-burst blend 2.2/10; 8.1s, EN.
emolia_c1651__E__Fear__D__en.c026.k0 · in -25.3 dBFS · gain +5.7 dB · vprof_vc-00099
(awe, contempt, fear · normal-paced, normally alert, slightly relaxed, monologue) Irans wilde Kreaturen, über seine vielfältigen Länder hinweg, wecken in mir etwas Urzeitliches. (contented sigh) Diese üppige, vielfältige Natur entfacht einen tiefen, ungezähmten Hunger in meiner Seele.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as awe, contempt, fear; style: monologue, narration; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.1/10; 11.4s, DE.
emolia_c1651__E__Sexual_Lust__A__de.c038.k2 · in -26.1 dBFS · gain +5.7 dB · vprof_vc-00099
(elation, hope enthusiasm optimism, interest · brisk, energised, neutral tension, casual) The possibilities in optimal control problems are incredible! We're so excited to explore distributed observation with constrained control, and seeing how constrained observation pairs with distributed control opens up such exciting avenues.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as elation, hope enthusiasm optimism, interest; style: casual, authoritative; good recording, no background noise; mildly explicit content; genuineness 0.4/6; vocal-burst blend 1.4/10; 11.4s, EN.
emolia_c1651__E__Hope_Enthusiasm_Optimism__B__en.c027.k2 · in -25.7 dBFS · gain +5.7 dB · vprof_vc-00099
(contempt, teasing, relief · brisk, energised, neutral tension, casual) Just add a contrasting color, for heaven's sake; ladies actually like red or white. (displeased grunt) Don't make this harder than it needs to be with something so basic.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as contempt, teasing, relief; style: casual, ranting; good recording, no background noise; mildly explicit content; genuineness 1.9/6; vocal-burst blend 3.7/10; 6.3s, EN.
emolia_c1651__E__Impatience_and_Irritability__B__en.c011.k0 · in -25.9 dBFS · gain +5.7 dB · vprof_vc-00099
(embarrassment, shame, helplessness · brisk, normally alert, neutral tension, casual) My little list of things I should have done is on the phone, on the kitchen table, even tucked beside my pillow. It just feels so painfully obvious to everyone.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, little disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as embarrassment, shame, helplessness; style: casual, ranting; good recording, no background noise; mildly explicit content; genuineness 2.9/6; vocal-burst blend 3.1/10; 5.7s, EN.
emolia_c1651__E__Shame__C__en.c029.k0 · in -25.7 dBFS · gain +5.7 dB · vprof_vc-00099
Fatigue Exhaustion ↓  /  Interestcloned voice   strict_090 · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest essentially absent — 0.00, virtually no clip in this corpus scores lower — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.99.

At the same time Fatigue Exhaustion goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.11 (lower than 89 % of clips in this corpus), a change of -0.89. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

5 clips · 45 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 100/150/150/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.4 dB
per-clip
per-clip buttons play:
rule PXRk 5qmax 0.888cmax 0.250d_a -0.888d_b 0.991dataset vprof_vclang entotal 45.1schain gain +6.1 dBseam step 1.4 dBcrossfades 100/150/150/100 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth
(fatigue exhaustion, sadness, disappointment · brisk, very low-energy, neutral tension, casual) Trying to balance all those different social needs is proving impossible. (contented sigh) (relief sigh) This single task is just tearing everything apart for us. (contented sigh)
full caption & clip details
A young adult feminine voice; delivery is very low-energy, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly negative, neutral stance, vulnerable; reads as fatigue exhaustion, sadness, disappointment; style: casual, playful; below-average recording, quiet background; mildly explicit content; genuineness 1.2/6; vocal-burst blend 1.1/10; 12.8s, EN.
emolia_c2675__V__S_DRAM__very_high__en.c033.k1 · in -27.5 dBFS · gain +6.1 dB · vprof_vc-00160
(emotional numbness, fear, disgust · brisk, normally alert, slightly relaxed, narration) Acne on the pope's skin often suggests he hasn't changed his undergarments or bedding. (deep breathing) It points to some serious oversight in his personal care.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as emotional numbness, fear, disgust; style: narration, authoritative; good recording, no background noise; mildly explicit content; genuineness 0.1/6; vocal-burst blend 0.4/10; 8.2s, EN.
emolia_c2675__V__TEMP__moderately_low__en.c003.k1 · in -26.1 dBFS · gain +6.1 dB · vprof_vc-00160
(astonishment surprise, sourness, impatience and irritability · brisk, normally alert, neutral tension, ranting) Diese modernen Geräte blasen nicht nur im Sommer kühle Luft; sie heizen dein Zuhause auch auf, wenn der Winter zuschlägt! (hiccup) Das ist unglaubliche Vielseitigkeit in einem einzigen tollen Gerät.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as astonishment surprise, sourness, impatience and irritability; style: ranting, dramatic; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.7/10; 7.7s, DE.
emolia_c2675__V__VOLT__very_high__de.c006.k0 · in -25.2 dBFS · gain +6.1 dB · vprof_vc-00160
(awe, pride, contentment · brisk, energised, slightly relaxed, authoritative) You can see everything: those stunning Hollywood Hills, the Sunset Strip, and Downtown la. Plus, Catalina Island and Century City offer incredible views.
full caption & clip details
An adult feminine voice; delivery is energised, brisk, slightly relaxed, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, full; clear, no disfluency, wide pitch range, minimal breath; affect is mildly negative, slightly dominant, slightly guarded; reads as awe, pride, contentment; style: authoritative; good recording, no background noise; mildly explicit content; genuineness 0.1/6; vocal-burst blend 0.9/10; 9.7s, EN.
emolia_c2675__V__S_FORM__moderately_high__en.c020.k2 · in -25.6 dBFS · gain +6.1 dB · vprof_vc-00160
(interest, awe, contemplation · normal-paced, normally alert, slightly relaxed, formal) But when you really look into the core of this creature, you'll see a different side to the mushroom's whole nature. It reveals a deeper truth about who it actually is.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, awe, contemplation; style: formal, narration; good recording, no background noise; mildly explicit content; genuineness 0.1/6; vocal-burst blend 0.5/10; 7.3s, EN.
emolia_c2675__V__S_WHIS__very_high__en.c044.k3 · in -25.4 dBFS · gain +6.1 dB · vprof_vc-00160
Fatigue Exhaustion ↓  /  Interestcloned voice   strict_091 · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest barely there — 0.11, lower than 89 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.89.

At the same time Fatigue Exhaustion goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.11 (lower than 89 % of clips in this corpus), a change of -0.89. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25, then +0.14 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

5 clips · 45 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 100/150/100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.6 dB
per-clip
per-clip buttons play:
rule PXRk 5qmax 0.888cmax 0.250d_a -0.888d_b 0.888dataset vprof_vclang entotal 45.3schain gain +5.1 dBseam step 2.6 dBcrossfades 100/150/100/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(fatigue exhaustion, relief, sexual lust · normal-paced, moderately variable, some disfluency, casual) (low mumble) Ugh, I need to power this through. (mournful wail) A quick shower and a gentle massage on my chest might finally get the pump going.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as fatigue exhaustion, relief, sexual lust; style: casual, conversational; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 0.2/10; 6.0s, EN.
emolia_c2508__E__Fatigue_Exhaustion__C__en.c032.k2 · in -24.4 dBFS · gain +5.1 dB · vprof_vc-00148
(disappointment, fear, emotional numbness · brisk, fairly steady, some disfluency, conversational) Ehrlich gesagt, fühlte sich das Schloss Hotel Velden wie ein Glücksspiel an, das einfach nicht aufgegangen ist. (guffaw) Aber weißt du, die ganze Wirtschaft schien alles zu ersticken, und das hat wahrscheinlich auch einen großen Teil dazu beigetragen.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, fear, emotional numbness; style: conversational, dramatic; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 2.8/10; 8.3s, DE.
emolia_c2508__E__Infatuation__C__de.c005.k1 · in -25.4 dBFS · gain +5.1 dB · vprof_vc-00148
(disgust · normal-paced, fairly steady, no disfluency, formal) Gib einfach exit in der neuen Shell ein, um die primäre Gruppe wieder in ihren Ausgangszustand zu versetzen. (sniff) Das sollte alles ordnungsgemäß rückgängig machen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust; style: formal, monologue; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 2.2/10; 10.5s, DE.
emolia_c2508__V__BKGN__moderately_high__de.c007.k2 · in -26.3 dBFS · gain +5.1 dB · vprof_vc-00148
(affection, distress, awe · normal-paced, fairly steady, no disfluency, formal) Dieser tiefe Dröhnton ist der Grund, warum wir streben, nicht nur für uns, sondern für den starken Rausch, den wir spüren, wenn wir das ganze Sein eines anderen beanspruchen. Dieses geteilte Feuer, dieser absolute Hunger, treibt alles an, was wir tun.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as affection, distress, awe; style: formal, didactic; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.3/10; 12.3s, DE.
emolia_c2508__E__Sexual_Lust__B__de.c030.k1 · in -26.0 dBFS · gain +5.1 dB · vprof_vc-00148
(interest, hope enthusiasm optimism, elation · normal-paced, fairly steady, little disfluency, casual) When we talk about biology, "trophy" opens up such amazing ideas—it could mean a whole location's nutrient boost in ecology, or the very life-giving nourishment within plants in botany!
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, hope enthusiasm optimism, elation; style: casual, conversational; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 2.4/10; 8.8s, EN.
emolia_c2508__E__Hope_Enthusiasm_Optimism__B__en.c019.k0 · in -23.4 dBFS · gain +5.1 dB · vprof_vc-00148
Astonishment Surprise ↓  /  Concentrationcloned voice   strict_101 · #10

This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Concentration up — by at least 0.25 each.

The chain starts with Concentration essentially absent — 0.02, lower than 98 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.98.

At the same time Astonishment Surprise goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.14 (lower than 86 % of clips in this corpus), a change of -0.86. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

5 clips · 44 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150/100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.5 dB
per-clip
per-clip buttons play:
rule AB2k 5qmax 0.865cmax 0.249d_a -0.865d_b 0.975dataset vprof_vclang entotal 43.6schain gain +4.8 dBseam step 1.5 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, fairly steady
(astonishment surprise, pleasure ecstasy, awe · normal-paced, neutral tension, some disfluency, conversational) She designed the house, I had no idea. It was for the whole family.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as astonishment surprise, pleasure ecstasy, awe; style: conversational, casual; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 6.7/10; 3.5s, EN.
emolia_c0693__V__R_MIXD__extremely_low__en.c020.k0 · in -23.5 dBFS · gain +4.8 dB · vprof_vc-00053
(sourness, contempt, teasing · normal-paced, slightly relaxed, little disfluency, casual) (contented sigh) Tell me honestly, do you really think you know who the true power players are in this whole mess? (drinking noises) I mean, who do you actually reckon holds the reins of control around here?
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as sourness, contempt, teasing; style: casual; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.7/10; 7.4s, EN.
emolia_c0693__V__R_MIXD__very_high__en.c013.k0 · in -24.5 dBFS · gain +4.8 dB · vprof_vc-00053
(sadness, longing, helplessness · normal-paced, slightly relaxed, no disfluency, narration) Linus drifted away from importance, becoming just a voice reading aloud now. (snorting giggle) He serves only to gently remind us of who he once was.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as sadness, longing, helplessness; style: narration, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.4/10; 7.7s, EN.
emolia_c0693__V__R_MIXD__moderately_low__en.c012.k1 · in -23.5 dBFS · gain +4.8 dB · vprof_vc-00053
(longing, distress, sadness · normal-paced, slightly relaxed, no disfluency, conversational) Mein Herz will eine Sache, aber die Logik besteht auf eine andere. Dennoch kann ich mich nicht entscheiden, weil beide Seiten genauso schwer und wackelig wirken.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as longing, distress, sadness; style: conversational, storytelling; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 3.0/10; 8.4s, DE.
emolia_c0693__V__REGS__moderately_high__de.c028.k3 · in -25.0 dBFS · gain +4.8 dB · vprof_vc-00053
(concentration, disgust, emotional numbness · brisk, slightly relaxed, some disfluency, casual) Die Blut- und Gewebeproben richtig zu entnehmen und ordnungsgemäß zu lagern, ist wirklich wichtig. Das ermöglicht uns die Analyse dieser Plasma-Kaskadenmarker, wie Komplement- und Gerinnungsproteine.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, disgust, emotional numbness; style: casual, monologue; good recording, no background noise; genuineness 2.4/6; vocal-burst blend 3.3/10; 17.1s, DE.
emolia_c0693__V__R_MIXD__moderately_low__de.c026.k3 · in -26.2 dBFS · gain +4.8 dB · vprof_vc-00053
Astonishment Surprise ↓  /  Concentrationcloned voice   strict_102 · #11

This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Concentration up — by at least 0.25 each.

The chain starts with Concentration essentially absent — 0.02, lower than 98 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.98.

At the same time Astonishment Surprise goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.14 (lower than 86 % of clips in this corpus), a change of -0.86. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

5 clips · 44 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 4.3 dB
per-clip
per-clip buttons play:
rule AB2k 5qmax 0.864cmax 0.249d_a -0.864d_b 0.977dataset vprof_vclang entotal 43.6schain gain +4.0 dBseam step 4.3 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body
(astonishment surprise, amusement, teasing · very slow, energised, fully relaxed, casual) A new study involving nearly ten thousand people suggests that forty-one percent of respiratory deaths within fifteen years might be linked to not enough vitamin D. (resonant hum) This research comes from the German Cancer Research Center.
full caption & clip details
An adult masculine voice; delivery is energised, very slow, fully relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, wide pitch range, breathless; affect is mildly positive, neutral stance, slightly guarded; reads as astonishment surprise, amusement, teasing; style: casual, playful; below-average recording, quiet background; contains vocal bursts: Ahem; genuineness 2.7/6; vocal-burst blend 1.8/10; 0.9s, EN.
emolia_c2435__V__ARSH__extremely_low__en.c037.k2 · in -23.4 dBFS · gain +4.0 dB · vprof_vc-00145
(fatigue exhaustion, distress · measured, subdued, neutral tension) Im Ernst, wie viel kann ein Kaffee dort kosten? Wer kauft das in Südafrika, wo Eskom ständig Strom abschaltet und saa kaum über die Runden kommt? (contented sigh)
full caption & clip details
An adult somewhat feminine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly negative, neutral stance, vulnerable; reads as fatigue exhaustion, distress; good recording, quiet background; explicit content; genuineness 1.6/6; vocal-burst blend 2.3/10; 12.6s, DE.
emolia_c2435__V__AGEV__moderately_high__de.c045.k1 · in -26.9 dBFS · gain +4.0 dB · vprof_vc-00145
(measured, normally alert, slightly relaxed, narration) The Thames near Kensington beckoned a visitor from Bordeaux. He admired the architecture of the Leadenhall Market structure.
full caption & clip details
An adult somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; mildly explicit content; genuineness 0.4/6; vocal-burst blend 1.5/10; 6.4s, EN.
emolia_c2435__E__Teasing__D__en.c008.k3 · in -22.5 dBFS · gain +4.0 dB · vprof_vc-00145
(fear · normal-paced, normally alert, slightly relaxed, narration) It can happen if the nail is chipped, bent, or if something is pressing on it from the outside while it's growing. (person whistling to get attention) That kind of pressure can really cause problems.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as fear; style: narration, monologue; good recording, no background noise; mildly explicit content; genuineness 0.8/6; vocal-burst blend 0.1/10; 8.2s, EN.
emolia_c2435__V__ATCK__moderately_low__en.c004.k1 · in -24.9 dBFS · gain +4.0 dB · vprof_vc-00145
(concentration, contemplation, anger · normal-paced, normally alert, slightly relaxed, formal) Durch diesen Akt, durch unser Hinfassen und Gebet, übergeben wir diese Person der Verkündigung des Evangeliums. Das bedeutet, sie trägt eine tiefe Verantwortung für die Einheit unserer Kirche.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, contemplation, anger; style: formal, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.7/10; 16.1s, DE.
emolia_c2435__V__AROU__extremely_low__de.c020.k3 · in -22.9 dBFS · gain +4.0 dB · vprof_vc-00145
Astonishment Surprise ↓  /  Contemptcloned voice   strict_103 · #12

This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Contempt up — by at least 0.25 each.

The chain starts with Contempt barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.86.

At the same time Astonishment Surprise goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.14 (lower than 86 % of clips in this corpus), a change of -0.86. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.18, then +0.25, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

5 clips · 30 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 100/150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 4.2 dB
per-clip
per-clip buttons play:
rule AB2k 5qmax 0.864cmax 0.250d_a -0.864d_b 0.865dataset vprof_vclang entotal 29.6schain gain +4.4 dBseam step 4.2 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an infant strongly feminine voice · neutral-bright, fairly smooth
(astonishment surprise, fear, disgust · very slow, energised, slightly relaxed, playful) You really need to ease into the new healthy diet, don't just switch everything at once. (surprised gasp) Tell me, am I missing any special steps for my rabbits when I make this food change?
full caption & clip details
An infant strongly feminine voice; delivery is energised, very slow, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, some disfluency, moderate pitch range, heavy breath; affect is mildly positive, slightly dominant, very vulnerable; reads as astonishment surprise, fear, disgust; style: playful, dramatic; below-average recording, quiet background; contains vocal bursts: Surprised Gasp; genuineness 1.7/6; vocal-burst blend 2.0/10; 0.5s, EN.
emolia_c1047__V__S_PLAY__very_high__en.c001.k0 · in -19.9 dBFS · gain +4.4 dB · vprof_vc-00076
(fatigue exhaustion, confusion · measured, very low-energy, neutral tension, monologue) Also, so ein paar Tage bevor deine Periode kommt, sinkt deine Temperatur einfach wieder ab. (surprised gasp) Das ist so ein kleines Signal, weißt du?
full caption & clip details
An adult somewhat feminine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, vulnerable; reads as fatigue exhaustion, confusion; style: monologue, ranting; average recording, no background noise; mildly explicit content; genuineness 1.7/6; vocal-burst blend 1.8/10; 6.9s, DE.
emolia_c1047__V__S_PLAY__very_high__de.c011.k1 · in -24.2 dBFS · gain +4.4 dB · vprof_vc-00076
(normal-paced, normally alert, slightly relaxed, authoritative) The pictures and other external content in this chapter follow the Creative Commons license too. (quiet sob) Always check the caption, though, for any exceptions to that rule.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: authoritative, playful; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.2/10; 7.5s, EN.
emolia_c1047__V__VFLX__very_high__en.c015.k2 · in -24.8 dBFS · gain +4.4 dB · vprof_vc-00076
(disgust, infatuation · normal-paced, normally alert, slightly relaxed, conversational) They're split up so that mucus can just flow right back into you through your mouth. (trembling whimper) That's how things are designed to happen naturally.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, wide pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as disgust, infatuation; style: conversational, playful; good recording, no background noise; mildly explicit content; genuineness 1.0/6; vocal-burst blend 0.5/10; 7.6s, EN.
emolia_c1047__V__S_MONO__very_high__en.c037.k1 · in -25.0 dBFS · gain +4.4 dB · vprof_vc-00076
(contempt, sourness, disgust · normal-paced, normally alert, slightly relaxed, narration) idiopathic non-specific lymphocytic infiltration manifesting as a chronic serositis of low-frequency amplitude
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as contempt, sourness, disgust; style: narration, authoritative; good recording, no background noise; mildly explicit content; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.7s, EN.
emolia_c1047__V__S_FORM__moderately_low__en.c012.k3 · in -24.3 dBFS · gain +4.4 dB · vprof_vc-00076
FULL — fullness of tonecloned voice   strict_110 · #13

This is a VoiceNet dimension, not an emotion: fullness of tone (FULL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with fullness of tone (FULL) at the very bottom of the range — 0.02, lower than 98 % of clips in this corpus — and ends with it high at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.75.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

4 clips · 45 s · de · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 120/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.7 dB
per-clip
per-clip buttons play:
rule VN1k 4qmax 0.750cmax 0.250d_a 0.750d_b 0.750dataset vprof_vclang detotal 45.0schain gain +3.2 dBseam step 0.7 dBcrossfades 120/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice
(fatigue exhaustion, infatuation, sexual lust · slow, lethargic, fully relaxed) Diese großen Wege, alle, sind mit Metalltieren verstopft. (person whistling to get attention) Solche Staus sind ein Übel auf jeder Hauptstraße dieser fliegenden Festung.
full caption & clip details
A young adult masculine voice; delivery is lethargic, slow, fully relaxed, fairly steady; timbre is slightly cool, dark, fairly smooth, thin; very slurred, no disfluency, moderate pitch range, breathless; affect is neutral, neutral stance, slightly guarded; below-average recording, no background noise; contains vocal bursts: Ahem; reads as fatigue exhaustion, infatuation, sexual lust; genuineness 1.2/6; vocal-burst blend 2.8/10; 0.5s, DE.
emolia_c0436__C__gravelly-orc-warlord__de.c007.k1 · in -23.6 dBFS · gain +3.2 dB · vprof_vc-00044
(teasing, amusement, pleasure ecstasy · normal-paced, energised, relaxed, casual) (chuckle) Honestly, this whole calculation and the idea of a fair market value, they solve so many headaches for crypto traders. It's like finally having instructions for building a really complicated digital Lego set. (ahem) (childlike giggle) (chuckle)
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as teasing, amusement, pleasure ecstasy; style: casual, conversational; good recording, quiet background; mildly explicit content; genuineness 2.3/6; vocal-burst blend 0.7/10; 32.0s, EN.
emolia_c0436__X__ga_amuse_laugh__en.c030.k2 · in -23.2 dBFS · gain +3.2 dB · vprof_vc-00044
(shame, sadness, doubt · normal-paced, normally alert, slightly relaxed, narration) It's just so disheartening because her whole case for damages rests on a contract breach. She says the defendant terminated things without any real, justifiable reason at all.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as shame, sadness, doubt; style: narration, authoritative; good recording, no background noise; mildly explicit content; genuineness 0.4/6; vocal-burst blend 0.5/10; 7.6s, EN.
emolia_c0436__E__Disappointment__A__en.c015.k1 · in -23.6 dBFS · gain +3.2 dB · vprof_vc-00044
(normal-paced, normally alert, slightly relaxed, narration) On the third of April, the event is scheduled for three o'clock in the afternoon. Please arrive no later than the thirty-first minute of that hour.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, authoritative; good recording, no background noise; explicit content; genuineness 1.2/6; vocal-burst blend 1.1/10; 5.3s, EN.
emolia_c0436__E__Distress__B__en.c008.k2 · in -22.9 dBFS · gain +3.2 dB · vprof_vc-00044
FULL — fullness of tonecloned voice   strict_111 · #14

This is a VoiceNet dimension, not an emotion: fullness of tone (FULL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with fullness of tone (FULL) at the very bottom of the range — 0.02, lower than 98 % of clips in this corpus — and ends with it high at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.75.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

4 clips · 18 s · de · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 11.9 dB
per-clip
per-clip buttons play:
rule VN1k 4qmax 0.750cmax 0.250d_a 0.750d_b 0.750dataset vprof_vclang detotal 17.8schain gain +3.3 dBseam step 11.9 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice
(relief, astonishment surprise, fatigue exhaustion · very slow, lethargic, fully relaxed) Man sagt, der Kanzler spinnt mit diesen großen europäischen Energieversorgern. (heavy breathing) Ehrlich gesagt, ich fühle mich so machtlos, das mitanzusehen.
full caption & clip details
An adult masculine voice; delivery is lethargic, very slow, fully relaxed, moderately variable; timbre is slightly cool, very dark, gravelly, thin; very slurred, frequent disfluency, moderate pitch range, breathless; affect is neutral, submissive, vulnerable; reads as relief, astonishment surprise, fatigue exhaustion; below-average recording, some background noise; contains vocal bursts: Contented Sigh; genuineness 1.7/6; vocal-burst blend 5.5/10; 1.1s, DE.
emolia_c2356__E__Helplessness__A__de.c008.k2 · in -34.4 dBFS · gain +3.3 dB · vprof_vc-00142
(teasing · brisk, energised, neutral tension, storytelling) Over six years, the biologists observed twenty-seven bottlenose dolphins off Bimini. (affirmative grunt) They meticulously tracked over seven hundred distinct hunting patterns.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as teasing; style: storytelling, playful; good recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.4/10; 5.1s, EN.
emolia_c2356__V__AGEV__extremely_low__en.c039.k1 · in -22.5 dBFS · gain +3.3 dB · vprof_vc-00142
(awe, longing, interest · brisk, energised, neutral tension, casual) Against Dortmund, de Jong just slipped it right in, honestly, I was watching that closely. (exasperated sigh) Such a delicate touch, sending it smoothly into the back of the net.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as awe, longing, interest; style: casual, storytelling; good recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.9/10; 5.6s, EN.
emolia_c2356__E__Interest__D__en.c035.k1 · in -23.6 dBFS · gain +3.3 dB · vprof_vc-00142
(pleasure ecstasy, affection, longing · normal-paced, normally alert, slightly relaxed, narration) Finding that empty space inside, I pour love into it. It blooms, a pure, breathtaking ecstasy that sweeps me away.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as pleasure ecstasy, affection, longing; style: narration; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.5/10; 6.4s, EN.
emolia_c2356__E__Pleasure_Ecstasy__B__en.c029.k1 · in -23.3 dBFS · gain +3.3 dB · vprof_vc-00142
FULL — fullness of tonecloned voice   strict_112 · #15

This is a VoiceNet dimension, not an emotion: fullness of tone (FULL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with fullness of tone (FULL) at the very bottom of the range — 0.02, lower than 98 % of clips in this corpus — and ends with it high at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.75.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

4 clips · 36 s · de · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 100/100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.3 dB
per-clip
per-clip buttons play:
rule VN1k 4qmax 0.750cmax 0.250d_a 0.750d_b 0.750dataset vprof_vclang detotal 35.5schain gain +7.0 dBseam step 3.3 dBcrossfades 100/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · fairly smooth, wide pitch range
(amusement, pleasure ecstasy, teasing · slow, energised, fully relaxed, casual) In den späteren Tagen des Zweiten Zeitalters schmiedete der große dunkle Herr Sauron neun Ringe, je einen für jeden auserwählten Narren. (snicker) Sie werden seinem Willen gehorchen.
full caption & clip details
A young adult masculine voice; delivery is energised, slow, fully relaxed, moderately variable; timbre is neutral-toned, very dark, fairly smooth, thin; very slurred, no disfluency, wide pitch range, breathless; affect is positive, slightly submissive, guarded; reads as amusement, pleasure ecstasy, teasing; style: casual, playful; average recording, quiet background; contains vocal bursts: Chuckle; genuineness 1.6/6; vocal-burst blend 4.5/10; 1.1s, DE.
emolia_c2154__C__gravelly-orc-warlord__de.c034.k2 · in -29.7 dBFS · gain +7.0 dB · vprof_vc-00124
(helplessness, distress, fear · slow, very low-energy, relaxed, casual) Die wohltätigen Zusagen von Bezos geben mir einfach einen Schauer. (snorting giggle) Zumindest zeigt irgendein Riese ein bisschen Menschlichkeit.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, slow, relaxed, variable; timbre is slightly cool, slightly dark, fairly smooth, slightly thin; very slurred, frequent disfluency, wide pitch range, breathless; affect is negative, submissive, vulnerable; reads as helplessness, distress, fear; style: casual, dramatic; below-average recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.8/10; 16.4s, DE.
emolia_c2154__X__cold_shiver__de.c008.k1 · in -29.6 dBFS · gain +7.0 dB · vprof_vc-00124
(awe, triumph · measured, energised, neutral tension, playful) And at Patriot Hills, the wind always rushes down from the mountain, hitting the aircraft right on the side. Whoa.
full caption & clip details
An adult masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, slightly guarded; reads as awe, triumph; style: playful, casual; below-average recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.1/10; 12.0s, EN.
emolia_c2154__B__slurping_noises__en.c007.k3 · in -26.3 dBFS · gain +7.0 dB · vprof_vc-00124
(fear, distress, disgust · normal-paced, energised, slightly relaxed, narration) Stell dir diese schreckliche Anämie vor, sie kann ein Leben ganz rauben. In den schlimmsten Fällen ist sie absolut tödlich.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, distress, disgust; style: narration, dramatic; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 2.0/10; 6.4s, DE.
emolia_c2154__E__Awe__C__de.c021.k1 · in -24.0 dBFS · gain +7.0 dB · vprof_vc-00124
Jealousy and Envycloned voice   strict_122 · #16

This chain comes from the one-sided rule: only Jealousy and Envy had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Jealousy and Envy barely there — 0.25, lower than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.75.

Nothing was asked of the other axis, and in fact Disappointment drifts down from 1.00 to 0.85 (-0.15), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

4 clips · 46 s · de · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.8 dB
per-clip
per-clip buttons play:
rule B1k 4qmax 0.750cmax 0.250d_a -0.148d_b 0.750dataset vprof_vclang detotal 45.8schain gain +4.3 dBseam step 0.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, wide pitch range
(disappointment, disgust, fatigue exhaustion · normal-paced, energised, neutral tension, cartoonish) Mann, ich bin gerade total fertig, echt mega niedergeschlagen. Meine ganze Stimmung ist gerade komplett im Eimer, das ist eine totale Katastrophe. (surprised gasp)
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as disappointment, disgust, fatigue exhaustion; style: cartoonish, ranting; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 0.8/10; 11.8s, DE.
emolia_c1969__V__ROUG__moderately_high__de.c007.k0 · in -24.1 dBFS · gain +4.3 dB · vprof_vc-00114
(malevolence malice, sourness, anger · normal-paced, normally alert, neutral tension, authoritative) Hört dies, denn der Herr spricht: Ich werde dieses Holz in Flammen setzen und jedes Blatt und jeden trockenen Zweig in euch verzehren. (exasperated sigh) Mein Feuer wird durchziehen und nichts Grünes oder Brüchiges unberührt lassen.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as malevolence malice, sourness, anger; style: authoritative, cartoonish; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 3.9/10; 15.0s, DE.
emolia_c1969__V__ESTH__very_high__de.c002.k3 · in -24.4 dBFS · gain +4.3 dB · vprof_vc-00114
(disgust, teasing, bitterness · normal-paced, energised, neutral tension, playful) Dieser Traum vom Collier für die ledige Frau? (wolf whistle) Das bedeutet wirklich, dass sie an einen zukünftigen Ehemann denkt. Es ist ein klassisches Symbol für Verpflichtung, weißt du.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as disgust, teasing, bitterness; style: playful, conversational; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 0.8/10; 12.2s, DE.
emolia_c1969__V__RESP__moderately_high__de.c002.k2 · in -24.2 dBFS · gain +4.3 dB · vprof_vc-00114
(jealousy and envy, doubt, sexual lust · brisk, energised, slightly relaxed, conversational) If you click it and buy something, it won't cost you anything extra. I'll just get a few pennies for advertising, you know?
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as jealousy and envy, doubt, sexual lust; style: conversational, playful; good recording, no background noise; mildly explicit content; genuineness 1.4/6; vocal-burst blend 1.4/10; 7.2s, EN.
emolia_c1969__V__RESP__moderately_high__en.c046.k0 · in -25.0 dBFS · gain +4.3 dB · vprof_vc-00114
Jealousy and Envycloned voice   strict_123 · #17

This chain comes from the one-sided rule: only Jealousy and Envy had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Jealousy and Envy barely there — 0.25, lower than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.75.

Nothing was asked of the other axis, and in fact Sadness drifts down from 1.00 to 0.94 (-0.06), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

4 clips · 42 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.7 dB
per-clip
per-clip buttons play:
rule B1k 4qmax 0.750cmax 0.250d_a -0.057d_b 0.750dataset vprof_vclang entotal 42.4schain gain +3.4 dBseam step 2.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, balanced body
(sadness, distress, fear · measured, energised, slightly relaxed, narration) When Malice finally crumbles, when all her meager health is gone and she fades to a mere whisper of a ghost, that's when the real end comes. A bitter curtain call for her.
full caption & clip details
An adult masculine voice; delivery is energised, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as sadness, distress, fear; style: narration, storytelling; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.7/10; 9.8s, EN.
emolia_c0760__E__Bitterness__C__en.c046.k0 · in -22.5 dBFS · gain +3.4 dB · vprof_vc-00058
(sourness, anger, contempt · normal-paced, energised, slightly relaxed, newsreading) Als wir uns diese Rollen genauer anschahen, wurde klar: je enger die Finanzen einer Nation mit dem größeren System verwoben sind, desto eher fühlt sich ihre Führung gezwungen, stärkere Unterstützung von der Europäischen Union zu benötigen.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sourness, anger, contempt; style: newsreading, storytelling; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.5/10; 13.8s, DE.
emolia_c0760__E__Concentration__B__de.c030.k2 · in -23.6 dBFS · gain +3.4 dB · vprof_vc-00058
(awe, longing, sadness · slow, very low-energy, relaxed, monologue) Es ist immer so viel schlimmer darunter, weißt du? (exhausted groan) Diese Schwellung sieht immer kleiner aus, bis du siehst, wie tief sie wirklich ist.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, slightly dark, rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is deeply negative, submissive, vulnerable; reads as awe, longing, sadness; style: monologue, dramatic; below-average recording, quiet background; genuineness 1.3/6; vocal-burst blend 1.2/10; 12.5s, DE.
emolia_c0760__X__whimpering__de.c007.k1 · in -25.0 dBFS · gain +3.4 dB · vprof_vc-00058
(jealousy and envy, contempt, bitterness · brisk, normally alert, slightly relaxed, authoritative) Diese nutzlosen Leute, zu dumm für Landwirtschaft oder Handel, wurden hierher aus England gezogen, weil sie eine Last für ihre Familien waren, und jetzt glauben sie, sie könnten uns befehlen?
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, contempt, bitterness; style: authoritative, casual; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 5.1/10; 6.7s, DE.
emolia_c0760__E__Bitterness__B__de.c017.k2 · in -22.3 dBFS · gain +3.4 dB · vprof_vc-00058
Contemplationcloned voice   strict_124 · #18

This chain comes from the one-sided rule: only Contemplation had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Contemplation barely there — 0.24, lower than 76 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.75.

Nothing was asked of the other axis, and in fact Anger drifts down from 1.00 to 0.80 (-0.20), which the rule did not require.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

4 clips · 36 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.5 dB
per-clip
per-clip buttons play:
rule B1k 4qmax 0.750cmax 0.250d_a -0.199d_b 0.750dataset vprof_vclang entotal 36.3schain gain +5.0 dBseam step 1.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice
(anger, intoxication altered states of consciousness, malevolence malice · slow, frantic, tense, casual) Look, this guide covers everything: the good parts, the problems it fixes, the main things to check out, and a whole lot more. Just follow along already.
full caption & clip details
An adult masculine voice; delivery is frantic, slow, tense, volatile; timbre is cool, dark, gravelly, thin; slurred, heavy disfluency, very wide pitch range, heavy breath; affect is deeply negative, very dominant, guarded; reads as anger, intoxication altered states of consciousness, malevolence malice; style: casual, cartoonish; poor recording, quiet background; mildly explicit content; genuineness 1.8/6; vocal-burst blend 0.9/10; 12.3s, EN.
emolia_c1950__E__Impatience_and_Irritability__C__en.c024.k3 · in -26.2 dBFS · gain +5.0 dB · vprof_vc-00113
(jealousy and envy, bitterness, contempt · brisk, energised, slightly relaxed, ranting) Ankara is happy because Azerbaijan's link holds, and Russia is busy clinging to its new allies. At least Turkey backs Abkhazia's freedom while Moscow worries about its mandates.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as jealousy and envy, bitterness, contempt; style: ranting, conversational; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.2/10; 7.8s, EN.
emolia_c1950__E__Sourness__B__en.c027.k1 · in -24.7 dBFS · gain +5.0 dB · vprof_vc-00113
(impatience and irritability, anger, disappointment · brisk, energised, slightly tense, casual) It's infuriating how that doctor just slapped simple cases into the average bin, or mislabeled them entirely, ignoring what the actual fees require. (effort grunt) I just can't believe they got away with such a blatant disregard for proper billing.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, anger, disappointment; style: casual, dramatic; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 6.6/10; 8.1s, EN.
emolia_c1950__E__Jealousy_and_Envy__A__en.c014.k1 · in -24.5 dBFS · gain +5.0 dB · vprof_vc-00113
(contemplation, bitterness, sadness · brisk, energised, neutral tension, dramatic) Sometimes it's a heavy sadness, but other times, it's just quiet. That nothingness... it's a strange kind of peace, like the volume on everything has just been turned down.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contemplation, bitterness, sadness; style: dramatic, ranting; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.4/10; 8.5s, EN.
emolia_c1950__E__Relief__C__en.c019.k1 · in -24.4 dBFS · gain +5.0 dB · vprof_vc-00113
Anger ↓  /  Concentrationcloned voice   strict_086 · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration below average — 0.26, lower than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.74.

At the same time Anger goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.28 (lower than 72 % of clips in this corpus), a change of -0.72. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

4 clips · 46 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.3 dB
per-clip
per-clip buttons play:
rule PXRk 4qmax 0.720cmax 0.249d_a -0.720d_b 0.736dataset vprof_vclang entotal 46.2schain gain +0.2 dBseam step 0.4 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult feminine voice · neutral-bright, fairly smooth, balanced body, no background noise, very low-energy, clear
(anger, relief, malevolence malice · slow, relaxed, moderately variable, whispered) All that poison—hatred, anger, violence—it'll never let them see it. Only when they finally just accept it, that’s when peace and love can truly arrive.
full caption & clip details
An adult feminine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly warm, neutral-bright, fairly smooth, balanced body; clear, no disfluency, very wide pitch range, minimal breath; affect is negative, slightly submissive, neutral openness; reads as anger, relief, malevolence malice; style: whispered, narration; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.9/10; 10.0s, EN.
anime_111__E__Sourness__C__en.c029.k0 · in -20.4 dBFS · gain +0.2 dB · vprof_vc-00009
(longing, interest, fatigue exhaustion · measured, slightly relaxed, fairly steady, narration) Ich bin so müde von diesen alten Bedarfsplanungs-Tools; zu versuchen, sie mit etwas Neuem zum Sprechen zu bringen, ist einfach erschöpfend. Wir brauchen Software, die uns wirklich ermöglicht, von jedem Team bessere Eingaben für das Gesamtbild zu bekommen.
full caption & clip details
An elderly feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as longing, interest, fatigue exhaustion; style: narration, monologue; very good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.8/10; 15.7s, DE.
anime_111__E__Fatigue_Exhaustion__D__de.c042.k0 · in -20.0 dBFS · gain +0.2 dB · vprof_vc-00009
(normal-paced, slightly relaxed, fairly steady, whispered) Knowing that hsbc customer data was exposed through the Swiss Leaks research network just feels... (mournful wail) devastating. It's such a profound breach of trust.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is slightly warm, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly negative, slightly submissive, slightly vulnerable; no dominant emotion; style: whispered, narration; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 4.8/10; 5.5s, EN.
anime_111__E__Sadness__B__en.c003.k1 · in -20.1 dBFS · gain +0.2 dB · vprof_vc-00009
(concentration, emotional numbness, sexual lust · normal-paced, slightly relaxed, fairly steady, whispered) Paroxysmale supraventrikuläre Tachykardie, eine rhythmische Beschleunigung über hundert Schläge pro Minute, stellt eine komplexe elektrophysiologische Herausforderung dar. Berücksichtigen Sie das verlängerte Intervall des qrs-Komplexes in der Elektrokardiogramm-Analyse.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is slightly warm, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, neutral openness; reads as concentration, emotional numbness, sexual lust; style: whispered, ASMR; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.0/10; 15.4s, DE.
anime_111__E__Interest__C__de.c034.k1 · in -20.4 dBFS · gain +0.2 dB · vprof_vc-00009
Interest ↓  /  Angercloned voice   strict_087 · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Anger is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Anger below average — 0.28, lower than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.72.

At the same time Interest goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.29 (lower than 71 % of clips in this corpus), a change of -0.71. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.25, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

4 clips · 38 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.6 dB
per-clip
per-clip buttons play:
rule PXRk 4qmax 0.710cmax 0.249d_a -0.710d_b 0.718dataset vprof_vclang entotal 38.0schain gain +3.9 dBseam step 0.6 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-bright, fairly smooth, balanced body, brisk
(interest, awe, hope enthusiasm optimism · normally alert, slightly relaxed, fairly steady, formal) So, in the fifth song of this epic, Homer tells us about Calypso, this beautiful, curly-haired nymph. She falls for the shipwrecked Odysseus and keeps him on her island of Ogygia for seven long years.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as interest, awe, hope enthusiasm optimism; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.7/10; 10.1s, EN.
emolia_c1699__V__S_STRY__very_high__en.c024.k2 · in -23.4 dBFS · gain +3.9 dB · vprof_vc-00104
(teasing, pleasure ecstasy, elation · energised, slightly relaxed, volatile, playful) Stell dir vor: du bist endlich in dieser Traumrolle, entwirfst exquisite Damenmode für die allerführenden Modehäuser. (trembling whimper) Es ist alles, wofür du dich je eingesetzt hast.
full caption & clip details
An adult feminine voice; delivery is energised, brisk, slightly relaxed, volatile; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as teasing, pleasure ecstasy, elation; style: playful, cartoonish; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 1.5/10; 11.5s, DE.
emolia_c1699__V__S_DRAM__moderately_high__de.c010.k1 · in -24.1 dBFS · gain +3.9 dB · vprof_vc-00104
(relief, sourness, amusement · very low-energy, neutral tension, moderately variable, casual) (breathy giggle) Even so, this rule didn't stop us from issuing the basic ruling. (convulsive sob) The defendant doesn't actually argue that the plaintiff was harmed by the unfair termination.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as relief, sourness, amusement; style: casual, playful; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 0.9/10; 9.9s, EN.
emolia_c1699__V__VALN__moderately_high__en.c008.k0 · in -24.0 dBFS · gain +3.9 dB · vprof_vc-00104
(anger, impatience and irritability, teasing · normally alert, slightly relaxed, fairly steady, formal) Es traf mich; Hintergrund, Akzent oder wie du dich präsentierst spielt wirklich keine Rolle. Wenn du das Problem mit deinen Fähigkeiten lösen kannst, bist du dabei, egal wie du bist.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, impatience and irritability, teasing; style: formal, dramatic; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 3.7/10; 7.0s, DE.
emolia_c1699__V__S_NARR__very_high__de.c020.k3 · in -24.0 dBFS · gain +3.9 dB · vprof_vc-00104
Impatience and Irritability ↓  /  Thankfulness Gratitudecloned voice   strict_098 · #21

This chain comes from the two-sided rule: it only counts if both emotions move — Impatience and Irritability down and Thankfulness Gratitude up — by at least 0.25 each.

The chain starts with Thankfulness Gratitude below average — 0.27, lower than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.73.

At the same time Impatience and Irritability goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.29 (lower than 71 % of clips in this corpus), a change of -0.71. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

4 clips · 33 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 4.8 dB
per-clip
per-clip buttons play:
rule AB2k 4qmax 0.709cmax 0.248d_a -0.709d_b 0.732dataset vprof_vclang entotal 33.3schain gain +3.1 dBseam step 4.8 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, good recording, no background noise, clear, light breath
(impatience and irritability, contempt, anger · brisk, energised, neutral tension, storytelling) Seriously, they just won't listen to any reason, it's infuriating. They ought to just talk to actual women in real life, for goodness sake.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as impatience and irritability, contempt, anger; style: storytelling, conversational; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.5/10; 7.2s, EN.
emolia_c0137__E__Impatience_and_Irritability__B__en.c030.k0 · in -19.7 dBFS · gain +3.1 dB · vprof_vc-00020
(infatuation, longing, fear · normal-paced, normally alert, slightly relaxed, narration) I need this to figure out when our bodies are bumping in this intimate little scene we're creating. (swallows) It's important for the physics of our closeness in this two-dimensional world.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation, longing, fear; style: narration, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.2/10; 8.7s, EN.
emolia_c0137__P__explicit__en.c010.k0 · in -24.5 dBFS · gain +3.1 dB · vprof_vc-00020
(relief, fear, disappointment · normal-paced, normally alert, slightly relaxed, ranting) Als die Historiker ihre Vorlesungen begannen und diese riesigen Themen behandelten, spürte ich, wie sich ein Knoten in meinem Magen zusammenzog. (fearful gasp) Dann am nächsten Tag, mit all diesen Gruppenprojektpräsentationen, machte ich mir einfach Sorgen, was sie sagen würden.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, fear, disappointment; style: ranting, narration; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 4.2/10; 12.9s, DE.
emolia_c0137__E__Fear__C__de.c006.k0 · in -25.5 dBFS · gain +3.1 dB · vprof_vc-00020
(thankfulness gratitude, contentment, pain · normal-paced, normally alert, slightly relaxed, narration) Ich bin so dankbar für die Masse, die in dieses kleine Volumenelement während dieser winzigen Zeit fließt. Es bedeutet mir so viel, sie auch wieder austreten zu sehen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as thankfulness gratitude, contentment, pain; style: narration, ranting; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 3.5/10; 4.9s, DE.
emolia_c0137__E__Thankfulness_Gratitude__B__de.c020.k3 · in -24.2 dBFS · gain +3.1 dB · vprof_vc-00020
Intoxication Altered States of Consciousness ↓  /  Bitternesscloned voice   strict_088 · #22

This chain comes from the proxy rule: the same two-sided test as above, but because Bitterness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Bitterness below average — 0.28, lower than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.72.

At the same time Intoxication Altered States of Consciousness goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.29 (lower than 71 % of clips in this corpus), a change of -0.71. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.25, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

4 clips · 33 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.9 dB
per-clip
per-clip buttons play:
rule PXRk 4qmax 0.706cmax 0.248d_a -0.706d_b 0.715dataset vprof_vclang entotal 32.7schain gain +4.4 dBseam step 1.9 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult somewhat masculine voice
(intoxication altered states of consciousness, pain, fatigue exhaustion · brisk, highly aroused, tense, casual) Martin, I find this demand completely ridiculous, especially considering the brother doesn't have to register his Lord's Table with you. What is the actual logic behind this contradictory nonsense?
full caption & clip details
An adult somewhat masculine voice; delivery is highly aroused, brisk, tense, volatile; timbre is slightly cool, dark, very rough, thin; slurred, almost no disfluency, very wide pitch range, audible breath; affect is deeply negative, slightly dominant, guarded; reads as intoxication altered states of consciousness, pain, fatigue exhaustion; style: casual, storytelling; below-average recording, quiet background; mildly explicit content; genuineness 1.8/6; vocal-burst blend 0.5/10; 7.0s, EN.
emolia_c0323__E__Anger__A__en.c043.k2 · in -24.4 dBFS · gain +4.4 dB · vprof_vc-00031
(normal-paced, normally alert, slightly relaxed, monologue) Some Muslim grooms insist on such proof, a profound requirement, to be absolutely certain of their bride's virginity. (nervous giggle) It is a practice that truly makes one stop and marvel at the depth of some traditions.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: monologue; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.7/10; 6.5s, EN.
emolia_c0323__E__Awe__A__en.c038.k0 · in -24.4 dBFS · gain +4.4 dB · vprof_vc-00031
(interest, concentration · normal-paced, normally alert, slightly relaxed, casual) It really strikes me how employer representation in the automotive industry mirrors the unique structures of each country's labor system, much like what we see with trade unions. It's a pattern that's deeply ingrained in the way each national system operates.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, concentration; style: casual, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.0/10; 10.9s, EN.
emolia_c0323__E__Contemplation__C__en.c026.k2 · in -23.8 dBFS · gain +4.4 dB · vprof_vc-00031
(bitterness, malevolence malice, disgust · brisk, energised, neutral tension, cartoonish) Mehr Kontrolle über deine Schwäche, oder die erbärmliche Freiheit, einfach nachzugeben? (hiccup) Wähle, welches Leid du ertragen möchtest.
full caption & clip details
An adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, guarded; reads as bitterness, malevolence malice, disgust; style: cartoonish, ranting; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 4.5/10; 8.8s, DE.
emolia_c0323__C__abyssal-tyrant__de.c033.k2 · in -25.6 dBFS · gain +4.4 dB · vprof_vc-00031
Sexual Lust ↓  /  Triumphcloned voice   strict_099 · #23

This chain comes from the two-sided rule: it only counts if both emotions move — Sexual Lust down and Triumph up — by at least 0.25 each.

The chain starts with Triumph below average — 0.32, lower than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.68.

At the same time Sexual Lust goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.31 (lower than 69 % of clips in this corpus), a change of -0.68. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.25, then +0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

4 clips · 29 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.9 dB
per-clip
per-clip buttons play:
rule AB2k 4qmax 0.680cmax 0.249d_a -0.680d_b 0.680dataset vprof_vclang entotal 28.5schain gain +4.1 dBseam step 1.9 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normal-paced, slightly relaxed, fairly steady
(sexual lust, disgust · normally alert, little disfluency, average clarity, conversational) Because you get to pick what you eat, (deep breath) when you eat it, and how it fits your life, based on your own tastes and what your body needs for energy.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as sexual lust, disgust; style: conversational, playful; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.0/10; 6.5s, EN.
emolia_c1576__V__TEMP__moderately_high__en.c009.k1 · in -24.6 dBFS · gain +4.1 dB · vprof_vc-00096
(normally alert, some disfluency, average clarity, casual) When we do this, we often take out things that are actually good for your body too. Those important nutrients get swept away with the rest.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 0.5/10; 6.9s, EN.
emolia_c1576__V__S_ASMR__extremely_low__en.c046.k0 · in -22.9 dBFS · gain +4.1 dB · vprof_vc-00096
(sourness, emotional numbness, sadness · energised, no disfluency, clear, storytelling) Zusätzlich zu ihrer Haftstrafe erhöhte das Obergericht auch den Schaden, den die sechs Täter ihren zwei Opfern schulden. (drinking noises) Das war eine erhebliche Erhöhung.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, emotional numbness, sadness; style: storytelling, narration; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.3/10; 8.9s, DE.
emolia_c1576__V__STNC__extremely_low__de.c010.k1 · in -24.8 dBFS · gain +4.1 dB · vprof_vc-00096
(triumph, anger, bitterness · energised, almost no disfluency, very clear, narration) The earth just bucked the whole city. We tumbled free, finding ourselves surrounded by broken concrete.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; very clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as triumph, anger, bitterness; style: narration, storytelling; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.2/10; 6.5s, EN.
emolia_c1576__V__S_MONO__extremely_low__en.c044.k2 · in -24.4 dBFS · gain +4.1 dB · vprof_vc-00096
DFLU — disfluencycloned voice   strict_107 · #24

This is a VoiceNet dimension, not an emotion: disfluency (DFLU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with disfluency (DFLU) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 20 s · de · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.2 dB
per-clip
per-clip buttons play:
rule VN1k 3qmax 0.500cmax 0.250d_a 0.500d_b 0.500dataset vprof_vclang detotal 19.6schain gain +3.0 dBseam step 1.2 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · balanced body, no background noise
(emotional numbness, pain, fatigue exhaustion · slow, normally alert, slightly relaxed, formal) Es fühlt sich einfach... leer an. Das zu nehmen, was andere verdient haben, dieser Austausch bedeutet nichts mehr. Nur eine stumpfe, flache Anerkennung der Transaktion.
full caption & clip details
A young adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, pain, fatigue exhaustion; style: formal, monologue; very good recording, no background noise; genuineness 0.7/6; vocal-burst blend 3.5/10; 1.3s, DE.
emolia_c2343__E__Emotional_Numbness__C__de.c041.k1 · in -22.9 dBFS · gain +3.0 dB · vprof_vc-00141
(fear, distress, helplessness · measured, normally alert, slightly relaxed, didactic) Wenn ich mir die Tabellen des Schadenskontrollhandbuchs anschaue, erinnere ich mich, wie viel dieses Schiff uns bedeutet; es zeigt, dass das Überfluten der zentralen Abteile die größte Gefahr darstellt.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, distress, helplessness; style: didactic, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.5/10; 10.3s, DE.
emolia_c2343__E__Affection__C__de.c018.k0 · in -22.5 dBFS · gain +3.0 dB · vprof_vc-00141
(relief · slow, very low-energy, relaxed, monologue) And at Patriot Hills the wind really whips off the mountain, right into the side of this aircraft. Whoa. (resonant hum)
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly warm, slightly dark, rough, balanced body; slurred, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as relief; style: monologue, conversational; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.1/10; 8.4s, EN.
emolia_c2343__B__slurping_noises__en.c007.k1 · in -23.8 dBFS · gain +3.0 dB · vprof_vc-00141
DFLU — disfluencycloned voice   strict_108 · #25

This is a VoiceNet dimension, not an emotion: disfluency (DFLU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with disfluency (DFLU) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 19 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.6 dB
per-clip
per-clip buttons play:
rule VN1k 3qmax 0.500cmax 0.250d_a 0.500d_b 0.500dataset vprof_vclang entotal 19.5schain gain +2.5 dBseam step 3.6 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(emotional numbness, longing, pain · slow, no disfluency, clear, formal) Not seeking Him, not knowing Him through His creation or the law deep inside us, it feels like we've left the door wide open to such sadness. It is a terrible thing to ignore what is truly good.
full caption & clip details
A young adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, longing, pain; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 4.5/10; 1.0s, EN.
emolia_c0436__E__Sadness__C__en.c017.k3 · in -22.4 dBFS · gain +2.5 dB · vprof_vc-00045
(astonishment surprise, awe, impatience and irritability · normal-paced, almost no disfluency, clear, narration) Frankly, I'm surprised by this massive free capacity and the willingness to cooperate here in Kassel. (effort grunt) It's almost overwhelming how much flexibility we have.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as astonishment surprise, awe, impatience and irritability; style: narration, formal; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.5/10; 7.2s, EN.
emolia_c0436__E__Impatience_and_Irritability__D__en.c008.k2 · in -20.6 dBFS · gain +2.5 dB · vprof_vc-00045
(pride, thankfulness gratitude, disgust · normal-paced, some disfluency, average clarity, monologue) Ich bin so stolz, dass unsere Anzeigenlinks direkt zu dieser sorgfältig gestalteten Zielseite führen und nicht nur zur Hauptseite. (humming) Das ist es, was unser Publikum wirklich mit dem richtigen Erlebnis verbindet.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as pride, thankfulness gratitude, disgust; style: monologue, casual; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 4.3/10; 11.7s, DE.
emolia_c0436__E__Pride__D__de.c033.k0 · in -24.2 dBFS · gain +2.5 dB · vprof_vc-00045
DFLU — disfluencycloned voice   strict_109 · #26

This is a VoiceNet dimension, not an emotion: disfluency (DFLU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with disfluency (DFLU) at the very bottom of the range — 0.00, virtually no clip in this corpus scores lower — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.50.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 21 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.6 dB
per-clip
per-clip buttons play:
rule VN1k 3qmax 0.500cmax 0.250d_a 0.500d_b 0.500dataset vprof_vclang entotal 20.9schain gain +4.7 dBseam step 0.6 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, good recording, no background noise, normal-paced, normally alert, slightly relaxed, fairly steady
(no disfluency, very clear, wide pitch range, narration) Liang Shanbo and Zhu Yingtai; their tragic love will stain everything with their needless suffering. They deserve nothing but the worst of it.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, very rough, very full; very clear, no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; no dominant emotion; style: narration, authoritative; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.8/10; 2.2s, EN.
emolia_c1056__E__Malevolence_Malice__C__en.c024.k1 · in -24.1 dBFS · gain +4.7 dB · vprof_vc-00077
(fear, longing, distress · little disfluency, clear, moderate pitch range, monologue) Ich sehe jetzt, wenn ich zurückblicke, diese ersten kleinen Beben. (fearful gasp) Sie waren Flüstern, kleine Warnungen, die mir sagten, dass etwas Schreckliches kommen würde.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, longing, distress; style: monologue, casual; good recording, no background noise; explicit content; genuineness 1.2/6; vocal-burst blend 4.6/10; 10.2s, DE.
emolia_c1056__E__Fear__A__de.c009.k0 · in -24.6 dBFS · gain +4.7 dB · vprof_vc-00077
(some disfluency, clear, moderate pitch range, casual) It covers everything from figuring out the math and designing the buildings to actually building and then keeping them up. That also means restoring and renovating existing structures.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.7/10; 8.8s, EN.
emolia_c1056__V__AGEV__extremely_low__en.c046.k0 · in -25.0 dBFS · gain +4.7 dB · vprof_vc-00077
Jealousy and Envycloned voice   strict_119 · #27

This chain comes from the one-sided rule: only Jealousy and Envy had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Jealousy and Envy around average — 0.50, right about the corpus median — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.50.

Nothing was asked of the other axis, and in fact Astonishment Surprise drifts down from 1.00 to 0.75 (-0.25), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 36 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.7 dB
per-clip
per-clip buttons play:
rule B1k 3qmax 0.500cmax 0.250d_a -0.245d_b 0.500dataset vprof_vclang entotal 35.9schain gain +3.7 dBseam step 0.7 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-bright, neutral tension, average clarity, wide pitch range
(astonishment surprise, awe, embarrassment · brisk, very low-energy, moderately variable, playful) I mean, I was totally expecting one picture, but then this entirely different one just popped up. (wolf whistle) It was quite a shock, honestly. (contented sigh)
full caption & clip details
A young adult masculine voice; delivery is very low-energy, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, submissive, vulnerable; reads as astonishment surprise, awe, embarrassment; style: playful, casual; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.7/10; 12.0s, EN.
emolia_c0241__V__VALS__moderately_high__en.c011.k2 · in -24.3 dBFS · gain +3.7 dB · vprof_vc-00027
(amusement, pleasure ecstasy, affection · normal-paced, energised, variable, storytelling) Deine Bauchmuskeln und dein Rumpf werden nicht nur durch endlose Crunches aufgebaut, weißt du. Du musst wirklich deinen ganzen Körper einbeziehen, dich nicht nur auf eine kleine Bewegung konzentrieren.
full caption & clip details
A child masculine voice; delivery is energised, normal-paced, neutral tension, variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as amusement, pleasure ecstasy, affection; style: storytelling, casual; good recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.5/10; 12.0s, DE.
emolia_c0241__V__VALN__moderately_high__de.c026.k3 · in -23.8 dBFS · gain +3.7 dB · vprof_vc-00027
(jealousy and envy, sadness, bitterness · normal-paced, energised, moderately variable, conversational) Manche Leute haben ganze Karrieren auf dem aufgebaut, was sie entdeckt haben. Außerdem wollten die Universitäten Anning als Frau einfach nicht um sich haben.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, neutral stance, slightly guarded; reads as jealousy and envy, sadness, bitterness; style: conversational, storytelling; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.2/10; 12.2s, DE.
emolia_c0241__V__S_CONV__moderately_high__de.c042.k2 · in -23.1 dBFS · gain +3.7 dB · vprof_vc-00027
Contentmentcloned voice   strict_120 · #28

This chain comes from the one-sided rule: only Contentment had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Contentment around average — 0.50, right about the corpus median — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.50.

Nothing was asked of the other axis, and in fact Sexual Lust drifts down from 1.00 to 0.85 (-0.15), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 27 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.8 dB
per-clip
per-clip buttons play:
rule B1k 3qmax 0.500cmax 0.250d_a -0.146d_b 0.500dataset vprof_vclang entotal 26.7schain gain +0.2 dBseam step 1.8 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-bright, balanced body
(sexual lust, awe, affection · brisk, normally alert, slightly relaxed, conversational) She sat at that enormous piano, so regal, so utterly breathtaking. I just stood there, completely lost in the sheer beauty of her waiting for my cue.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as sexual lust, awe, affection; style: conversational, casual; good recording, no background noise; genuineness 4.0/6; vocal-burst blend 5.7/10; 6.2s, EN.
emolia_c2322__E__Awe__B__en.c047.k1 · in -18.6 dBFS · gain +0.2 dB · vprof_vc-00140
(intoxication altered states of consciousness, confusion, affection · normal-paced, energised, relaxed, casual) H T T P colon slash slash example dot com slash images slash (ahem) picture dash one. This is a (ahem) test file path for voice (low mumble) actor (low mumble) practice today.
full caption & clip details
A child masculine voice; delivery is energised, normal-paced, relaxed, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; slurred, frequent disfluency, very wide pitch range, minimal breath; affect is positive, slightly dominant, fairly guarded; reads as intoxication altered states of consciousness, confusion, affection; style: casual, playful; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 2.3/10; 13.2s, EN.
emolia_c2322__E__Confusion__A__en.c008.k2 · in -20.4 dBFS · gain +0.2 dB · vprof_vc-00140
(contentment, relief, longing · brisk, normally alert, slightly relaxed, casual) Even though long walks bring a bit of lower back ache, finding a place to rest always brings such a deep sense of peace. (guffaw) It feels wonderful to finally let all that tension just melt away.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contentment, relief, longing; style: casual, conversational; good recording, quiet background; genuineness 2.0/6; vocal-burst blend 1.5/10; 7.6s, EN.
emolia_c2322__E__Contentment__A__en.c005.k3 · in -21.6 dBFS · gain +0.2 dB · vprof_vc-00140
Jealousy and Envycloned voice   strict_121 · #29

This chain comes from the one-sided rule: only Jealousy and Envy had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Jealousy and Envy around average — 0.50, right about the corpus median — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.50.

Nothing was asked of the other axis, and in fact Amusement drifts down from 1.00 to 0.72 (-0.27), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 26 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.2 dB
per-clip
per-clip buttons play:
rule B1k 3qmax 0.500cmax 0.250d_a -0.271d_b 0.500dataset vprof_vclang entotal 25.8schain gain +4.8 dBseam step 0.2 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, some disfluency
(amusement, teasing, astonishment surprise · normal-paced, energised, slightly relaxed, playful) I went against what the casino director said. He insisted there were so many concerts in Basel people wouldn't care for one more.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as amusement, teasing, astonishment surprise; style: playful, casual; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.3/10; 9.7s, EN.
emolia_c2545__V__ESTH__moderately_high__en.c042.k0 · in -24.9 dBFS · gain +4.8 dB · vprof_vc-00152
(anger, impatience and irritability, astonishment surprise · measured, normally alert, slightly relaxed, monologue) Verdammt, das ist ein totales Chaos, echt jetzt, was zum Teufel ist hier gerade los? Das ist ein kompletter Mist, ich sag's dir.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, impatience and irritability, astonishment surprise; style: monologue, narration; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 2.9/10; 7.1s, DE.
emolia_c2545__V__RCQL__moderately_low__de.c040.k2 · in -24.8 dBFS · gain +4.8 dB · vprof_vc-00152
(jealousy and envy, fatigue exhaustion, helplessness · measured, subdued, neutral tension, monologue) (contented sigh) Hey, ich habe vor ein paar Tagen alle meine Fotos von meinem alten Handy auf meinen Laptop verschoben, der Windows Zehn läuft. (normal breathing) Jetzt versuche ich, sie alle wieder auf mein neues Handy zu bekommen. (exhausted groan)
full caption & clip details
An adult masculine voice; delivery is subdued, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as jealousy and envy, fatigue exhaustion, helplessness; style: monologue, casual; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 4.3/10; 9.3s, DE.
emolia_c2545__V__R_CHST__moderately_high__de.c007.k1 · in -24.8 dBFS · gain +4.8 dB · vprof_vc-00152
Jealousy and Envy ↓  /  Contentmentcloned voice   strict_095 · #30

This chain comes from the two-sided rule: it only counts if both emotions move — Jealousy and Envy down and Contentment up — by at least 0.25 each.

The chain starts with Contentment around average — 0.51, higher than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.49.

At the same time Jealousy and Envy goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.51 (right about the corpus median), a change of -0.49. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 45 s · de · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.5 dB
per-clip
per-clip buttons play:
rule AB2k 3qmax 0.490cmax 0.248d_a -0.491d_b 0.490dataset vprof_vclang detotal 44.8schain gain +6.1 dBseam step 3.5 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-bright
(jealousy and envy, sourness, disgust · normal-paced, normally alert, slightly relaxed, narration) Diese neue Anlage hat so eine unglaubliche Gestaltungsfreiheit, weißt du? Sie hat einfach diese Effizienz, die alle wollen, und das macht mich echt neidisch.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, sourness, disgust; style: narration, storytelling; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.7/10; 8.0s, DE.
emolia_c0955__E__Jealousy_and_Envy__A__de.c019.k0 · in -24.3 dBFS · gain +6.1 dB · vprof_vc-00064
(relief, teasing, triumph · measured, highly aroused, neutral tension, dramatic) Leaving Vienna again, Ulrich sought refuge from fresh attacks. (exasperated sigh) He brought his letters to the imperial court in Korneuburg, then on to Neustadt. (contented sigh)
full caption & clip details
A young adult masculine voice; delivery is highly aroused, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, almost no disfluency, wide pitch range, normal breath; affect is deeply negative, slightly dominant, vulnerable; reads as relief, teasing, triumph; style: dramatic, cartoonish; below-average recording, quiet background; mildly explicit content; genuineness 0.4/6; vocal-burst blend 0.4/10; 26.1s, EN.
emolia_c0955__V__AGEV__moderately_low__en.c006.k2 · in -27.8 dBFS · gain +6.1 dB · vprof_vc-00064
(contentment, awe, elation · normal-paced, normally alert, slightly relaxed, casual) When you see these moments as a chance to transcend this life, I truly feel a surge of pride knowing you reach acceptance, love, and profound peace. (humming) It is a magnificent realization of the human spirit.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, very thin; clear, some disfluency, moderate pitch range, minimal breath; affect is mildly negative, neutral stance, fairly guarded; reads as contentment, awe, elation; style: casual, conversational; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.9/10; 11.0s, EN.
emolia_c0955__E__Pride__B__en.c039.k0 · in -24.6 dBFS · gain +6.1 dB · vprof_vc-00064
Hope Enthusiasm Optimism ↓  /  Longingcloned voice   strict_096 · #31

This chain comes from the two-sided rule: it only counts if both emotions move — Hope Enthusiasm Optimism down and Longing up — by at least 0.25 each.

The chain starts with Longing around average — 0.51, right about the corpus median — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.49.

At the same time Hope Enthusiasm Optimism goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.51 (higher than 51 % of clips in this corpus), a change of -0.49. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 20 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.1 dB
per-clip
per-clip buttons play:
rule AB2k 3qmax 0.489cmax 0.248d_a -0.489d_b 0.494dataset vprof_vclang entotal 20.0schain gain +3.5 dBseam step 3.1 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult somewhat feminine voice · neutral-toned, fairly smooth, good recording, no background noise
(hope enthusiasm optimism, elation, pleasure ecstasy · normal-paced, normally alert, slightly relaxed, whispered) It's genuinely thrilling that we can finally head into the action. We're going to be right in the thick of it.
full caption & clip details
An adult somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, elation, pleasure ecstasy; style: whispered, narration; good recording, no background noise; mildly explicit content; genuineness 0.6/6; vocal-burst blend 1.2/10; 5.8s, EN.
emolia_c1044__V__ROUG__moderately_low__en.c031.k1 · in -22.0 dBFS · gain +3.5 dB · vprof_vc-00075
(sexual lust, affection, disgust · normal-paced, normally alert, slightly relaxed, casual) Those mushroom garlands look lovely, but you should really keep them in the shade if you put them out on the street. (sharp inhale) The direct sun will just ruin the beauty of them too quickly.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as sexual lust, affection, disgust; style: casual, narration; good recording, no background noise; mildly explicit content; genuineness 1.3/6; vocal-burst blend 1.7/10; 6.6s, EN.
emolia_c1044__V__BRGT__extremely_low__en.c035.k3 · in -22.9 dBFS · gain +3.5 dB · vprof_vc-00075
(longing, affection, infatuation · brisk, very low-energy, slightly tense, monologue) I cannot love you truly until you let me feel your love first, my savior. (quiet sob) That is what my heart demands from you.
full caption & clip details
An adult feminine voice; delivery is very low-energy, brisk, slightly tense, variable; timbre is neutral-toned, slightly bright, fairly smooth, very thin; somewhat unclear, little disfluency, wide pitch range, audible breath; affect is negative, submissive, vulnerable; reads as longing, affection, infatuation; style: monologue, dramatic; good recording, no background noise; mildly explicit content; genuineness 1.8/6; vocal-burst blend 4.0/10; 7.8s, EN.
emolia_c1044__V__FOCS__very_high__en.c002.k0 · in -26.0 dBFS · gain +3.5 dB · vprof_vc-00075
Interest ↓  /  Malevolence Malicecloned voice   strict_083 · #32

This chain comes from the proxy rule: the same two-sided test as above, but because Malevolence Malice is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Malevolence Malice around average — 0.50, right about the corpus median — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.49.

At the same time Interest goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.51 (right about the corpus median), a change of -0.49. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 38 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.8 dB
per-clip
per-clip buttons play:
rule PXRk 3qmax 0.487cmax 0.249d_a -0.487d_b 0.494dataset vprof_vclang entotal 38.3schain gain +5.5 dBseam step 1.8 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, fairly smooth, good recording, no background noise, slightly relaxed, average clarity
(interest, hope enthusiasm optimism, awe · brisk, very low-energy, moderately variable, casual) So today, I want to show you how to make a vegetable curry that's super fast, really simple, and doesn't cost much, but tastes totally real.
full caption & clip details
An adult feminine voice; delivery is very low-energy, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, full; average clarity, little disfluency, wide pitch range, minimal breath; affect is positive, slightly submissive, neutral openness; reads as interest, hope enthusiasm optimism, awe; style: casual, playful; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.0/10; 11.3s, EN.
emolia_c0238__V__S_TECH__extremely_low__en.c044.k3 · in -24.1 dBFS · gain +5.5 dB · vprof_vc-00025
(sourness, disgust, fatigue exhaustion · normal-paced, normally alert, fairly steady, storytelling) Da Bitcoin fast vierundzwanzigtausend Dollar erreicht hat, steigen Polkadot und Chainlink zusammen mit Dogecoin. (coughing) Diese Altcoins mit großer Marktkapitalisierung steigen wirklich an, nachdem Bitcoin über dreieinundzwanzigtausend gebrochen ist.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, disgust, fatigue exhaustion; style: storytelling, formal; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.7/10; 14.4s, DE.
emolia_c0238__V__S_MONO__very_high__de.c014.k3 · in -25.9 dBFS · gain +5.5 dB · vprof_vc-00025
(malevolence malice, triumph, contempt · normal-paced, normally alert, volatile, casual) Vergiss das bloße Faulenzen an einem schönen Ort während deines Brandenburg-Urlaubs. (nervous giggle) Zieh stattdessen diese fantastische Alternative für deinen Gruppentrip in Betracht.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, volatile; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, no audible breath; affect is positive, slightly submissive, neutral openness; reads as malevolence malice, triumph, contempt; style: casual, conversational; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.4/10; 12.9s, DE.
emolia_c0238__V__TENS__moderately_high__de.c008.k0 · in -26.8 dBFS · gain +5.5 dB · vprof_vc-00025
Impatience and Irritability ↓  /  Contentmentcloned voice   strict_084 · #33

This chain comes from the proxy rule: the same two-sided test as above, but because Contentment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contentment around average — 0.51, higher than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.49.

At the same time Impatience and Irritability goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.51 (higher than 51 % of clips in this corpus), a change of -0.49. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 28 s · de · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 100/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 9.3 dB
per-clip
per-clip buttons play:
rule PXRk 3qmax 0.486cmax 0.250d_a -0.486d_b 0.488dataset vprof_vclang detotal 28.4schain gain -0.5 dBseam step 9.3 dBcrossfades 100/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice
(impatience and irritability, pleasure ecstasy, confusion · fast, highly aroused, slightly tense, cartoonish) Vor neunzig Millionen Jahren fanden Wissenschaftler in Yunnan, China, dieses unglaubliche Dinosaurierknochenbett. Es zeigt uns endlich genau, wie diese uralten Embryonen in ihren Eiern gewachsen sind!
full caption & clip details
A young adult feminine voice; delivery is highly aroused, fast, slightly tense, variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, no disfluency, very wide pitch range, normal breath; affect is elated, very dominant, guarded; reads as impatience and irritability, pleasure ecstasy, confusion; style: cartoonish, ranting; average recording, quiet background; mildly explicit content; genuineness 0.8/6; vocal-burst blend 2.1/10; 12.4s, DE.
emolia_c1905__E__Triumph__C__de.c021.k1 · in -24.2 dBFS · gain -0.5 dB · vprof_vc-00112
(shame, embarrassment, disappointment · brisk, normally alert, slightly relaxed, conversational) I feel so ashamed that the English version, the one it was first published in, is the only one that truly counts as official and legally sound. (gulps) It's just... embarrassing that there isn't one definitive, recognized text.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as shame, embarrassment, disappointment; style: conversational, casual; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 3.6/10; 8.6s, EN.
emolia_c1905__E__Shame__C__en.c005.k0 · in -24.3 dBFS · gain -0.5 dB · vprof_vc-00112
(contentment, pleasure ecstasy, amusement · very slow, very low-energy, relaxed, casual) We only run (low mumble) doubles in the summer hobby round, (soft hum) you know, but in the winter, we switch things up to mixed pairs for the actual matches.
full caption & clip details
An infant strongly feminine voice; delivery is very low-energy, very slow, relaxed, highly volatile; timbre is cool, bright, gravelly, thin; slurred, frequent disfluency, very wide pitch range, no audible breath; affect is elated, slightly submissive, very vulnerable; reads as contentment, pleasure ecstasy, amusement; style: casual, dramatic; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.2/10; 7.6s, EN.
emolia_c1905__V__ARSH__moderately_high__en.c011.k3 · in -14.9 dBFS · gain -0.5 dB · vprof_vc-00112
Infatuation ↓  /  Interestcloned voice   strict_085 · #34

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest around average — 0.51, higher than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.49.

At the same time Infatuation goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.51 (right about the corpus median), a change of -0.49. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 29 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.6 dB
per-clip
per-clip buttons play:
rule PXRk 3qmax 0.486cmax 0.250d_a -0.493d_b 0.486dataset vprof_vclang entotal 29.1schain gain +2.7 dBseam step 1.6 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed, fairly steady
(infatuation, affection, longing · normal-paced, normally alert, fairly steady) By sharing the broken bread and cup, Jesus showed God's love cannot be stopped. (contented sigh) Nothing can keep him from remaining forever within that boundless love.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as infatuation, affection, longing; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.2/10; 8.4s, EN.
emolia_c2496__E__Triumph__D__en.c035.k2 · in -22.2 dBFS · gain +2.7 dB · vprof_vc-00146
(fear, concentration, confusion · brisk, casual, formal) Seriously, the vestibular system reacts first, then that superior following and optokinetic system is slow, yet somehow precise. (effort grunt) And don't even start on cortical suppression of eye movements during fixation.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as fear, concentration, confusion; style: casual, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.0/10; 10.1s, EN.
emolia_c2496__E__Impatience_and_Irritability__D__en.c003.k0 · in -23.7 dBFS · gain +2.7 dB · vprof_vc-00146
(interest, elation, hope enthusiasm optimism · brisk, formal) To truly know how to date a Wiccan, I absolutely need to dive into the internet, exploring every wonderful detail of Wiccan dating. The sheer possibility of what awaits is just electrifying.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as interest, elation, hope enthusiasm optimism; style: formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.0/10; 10.9s, EN.
emolia_c2496__E__Pleasure_Ecstasy__A__en.c031.k0 · in -22.3 dBFS · gain +2.7 dB · vprof_vc-00146
Longing ↓  /  Triumphcloned voice   strict_097 · #35

This chain comes from the two-sided rule: it only counts if both emotions move — Longing down and Triumph up — by at least 0.25 each.

The chain starts with Triumph around average — 0.51, higher than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.49.

At the same time Longing goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.51 (higher than 51 % of clips in this corpus), a change of -0.49. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

3 clips · 37 s · de · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 5.1 dB
per-clip
per-clip buttons play:
rule AB2k 3qmax 0.485cmax 0.249d_a -0.485d_b 0.488dataset vprof_vclang detotal 36.6schain gain +2.9 dBseam step 5.1 dBcrossfades 100/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, good recording
(longing, jealousy and envy, infatuation · slow, very low-energy, relaxed, monologue) (contented sigh) Bleib doch bitte noch ein bisschen länger bei mir. (coughing) Ich will nicht, dass das so schnell vorbei ist.
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, relaxed, variable; timbre is neutral-toned, dark, fairly smooth, very full; somewhat unclear, some disfluency, fairly narrow pitch, audible breath; affect is negative, submissive, vulnerable; reads as longing, jealousy and envy, infatuation; style: monologue, whispered; good recording, quiet background; mildly explicit content; genuineness 1.8/6; vocal-burst blend 10.0/10; 6.3s, DE.
emolia_c1073__V__S_NEWS__moderately_low__de.c034.k2 · in -25.6 dBFS · gain +2.9 dB · vprof_vc-00080
(normal-paced, normally alert, slightly relaxed, didactic) Gerichte gingen auch davon aus, (snicker) besonders auf dem Markt für Sport- und Freizeitbekleidung, dass die Leute dort daran gewöhnt sind, sekundäre Marken neben etablierten Marken zu sehen, wie bei diesen Schuhen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 2.3/10; 12.4s, DE.
emolia_c1073__V__S_CONV__very_high__de.c013.k0 · in -20.5 dBFS · gain +2.9 dB · vprof_vc-00080
(triumph, pride, jealousy and envy · normal-paced, energised, slightly relaxed, storytelling) Der König, mächtigste und heiligste aller Könige, hielt den Stab seines Reiches. Er lebte in Hastinapura, der Stadt, die nach dem Elefanten benannt ist, und brachte großen Ruhm den Kuru-Herrschern.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as triumph, pride, jealousy and envy; style: storytelling, narration; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 1.7/10; 18.1s, DE.
emolia_c1073__V__S_CART__moderately_low__de.c047.k2 · in -24.6 dBFS · gain +2.9 dB · vprof_vc-00080
GEND — perceived gendercloned voice   strict_104 · #36

This is a VoiceNet dimension, not an emotion: perceived gender (GEND) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with perceived gender (GEND) below average — 0.35, lower than 65 % of clips in this corpus — and ends with it above average at 0.60, higher than 60 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

2 clips · 16 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.4 dB
per-clip
per-clip buttons play:
rule VN1k 2qmax 0.250cmax 0.250d_a 0.250d_b 0.250dataset vprof_vclang entotal 16.0schain gain +2.3 dBseam step 0.4 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(affection, infatuation, pleasure ecstasy · normal-paced, almost no disfluency, casual, conversational) Margherita, so deeply in love, felt a sudden, overwhelming awe for his art. (nervous giggle) She knew she would devote her life to that beautiful, infinite memory of him.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as affection, infatuation, pleasure ecstasy; style: casual, conversational; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 3.4/10; 7.8s, EN.
emolia_c2322__E__Awe__B__en.c001.k0 · in -22.1 dBFS · gain +2.3 dB · vprof_vc-00140
(elation · brisk, no disfluency, authoritative, newsreading) www Punkt example Punkt com Schrägstrich users Schrägstrich profile Unterstrich eins zwei drei Bitte besuchen Sie die Webseite jetzt, um alle Details dieses spannenden neuen Projekts zu sehen.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as elation; style: authoritative, newsreading; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.0/10; 8.3s, DE.
emolia_c2322__E__Doubt__B__de.c011.k3 · in -22.5 dBFS · gain +2.3 dB · vprof_vc-00140
FOCS — vocal focuscloned voice   strict_105 · #37

This is a VoiceNet dimension, not an emotion: vocal focus (FOCS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with vocal focus (FOCS) at the very bottom of the range — 0.02, lower than 98 % of clips in this corpus — and ends with it below average at 0.27, lower than 73 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

2 clips · 19 s · de · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.9 dB
per-clip
per-clip buttons play:
rule VN1k 2qmax 0.250cmax 0.250d_a 0.250d_b 0.250dataset vprof_vclang detotal 19.3schain gain +7.2 dBseam step 2.9 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult somewhat feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, slightly relaxed, fairly steady, average clarity
(longing, fatigue exhaustion, sadness · measured, subdued, frequent disfluency) Ich sehe genau, was sie getan haben. (relief sigh) Ich werde mich ihnen auf jeden Fall widersetzen.
full caption & clip details
An adult somewhat feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as longing, fatigue exhaustion, sadness; very good recording, no background noise; genuineness 1.5/6; vocal-burst blend 5.3/10; 4.3s, DE.
emolia_c0323__V__FOCS__extremely_low__de.c013.k2 · in -25.1 dBFS · gain +7.2 dB · vprof_vc-00033
(teasing, jealousy and envy, amusement · normal-paced, normally alert, some disfluency, playful) Manche Dinge leiten Wärme gut, andere nicht. (smack one s lips) Genau wie manche Materialien besser Strom leiten als andere, so tut das auch die Wärme. (surprised gasp)
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as teasing, jealousy and envy, amusement; style: playful, conversational; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 0.9/10; 15.1s, DE.
emolia_c0323__V__R_NASL__extremely_low__de.c003.k1 · in -28.1 dBFS · gain +7.2 dB · vprof_vc-00033
FOCS — vocal focuscloned voice   strict_106 · #38

This is a VoiceNet dimension, not an emotion: vocal focus (FOCS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with vocal focus (FOCS) at the very bottom of the range — 0.02, lower than 98 % of clips in this corpus — and ends with it below average at 0.27, lower than 73 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

2 clips · 17 s · de · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.6 dB
per-clip
per-clip buttons play:
rule VN1k 2qmax 0.250cmax 0.250d_a 0.250d_b 0.250dataset vprof_vclang detotal 17.1schain gain +2.7 dBseam step 3.6 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a child masculine voice · slightly cool, fairly smooth, no background noise, energised, moderately variable, very wide pitch range
(awe, intoxication altered states of consciousness, elation · slow, slightly relaxed, no disfluency, cartoonish) Bie bie see, für groß, kühn, kreativ. Rücken an Rücken, helles Leuchtfeuer ruft.
full caption & clip details
A child masculine voice; delivery is energised, slow, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, very thin; crisply articulate, no disfluency, very wide pitch range, audible breath; affect is positive, slightly dominant, vulnerable; reads as awe, intoxication altered states of consciousness, elation; style: cartoonish, dramatic; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 2.0/10; 6.3s, DE.
emolia_c2162__E__Fear__C__de.c009.k0 · in -20.8 dBFS · gain +2.7 dB · vprof_vc-00129
(intoxication altered states of consciousness, fatigue exhaustion, confusion · brisk, neutral tension, some disfluency, cartoonish) (surprised gasp) Nur die endgültige Montage der Continental-Motoren, wie die größten, erfolgt noch von Hand, und dieser ganze Prozess fühlt sich einfach... verschoben an. Ein Team baut es zusammen, und alles wirkt ein bisschen unscharf.
full caption & clip details
An adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; very clear, some disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, guarded; reads as intoxication altered states of consciousness, fatigue exhaustion, confusion; style: cartoonish, dramatic; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 2.4/10; 10.9s, DE.
emolia_c2162__E__Intoxication_Altered_States_of_Consciousness__C__de.c026.k1 · in -24.4 dBFS · gain +2.7 dB · vprof_vc-00129
Fearcloned voice   strict_116 · #39

This chain comes from the one-sided rule: only Fear had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Fear clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.25.

Nothing was asked of the other axis, and in fact Infatuation drifts down from 1.00 to 0.92 (-0.08), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

2 clips · 15 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.7 dB
per-clip
per-clip buttons play:
rule B1k 2qmax 0.250cmax 0.250d_a -0.077d_b 0.250dataset vprof_vclang entotal 15.0schain gain +2.6 dBseam step 2.7 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · fairly smooth, brisk
(infatuation, affection, astonishment surprise · energised, slightly relaxed, moderately variable, monologue) He found himself completely drawn to Lina. She felt the same way, and it was a connection that just deepened between them.
full caption & clip details
An adult feminine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as infatuation, affection, astonishment surprise; style: monologue, narration; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.4/10; 8.0s, EN.
emolia_c0645__V__S_PLAY__very_high__en.c043.k3 · in -24.1 dBFS · gain +2.6 dB · vprof_vc-00050
(astonishment surprise, fear, relief · very low-energy, very tense, variable, dramatic) (contented sigh) (contented sigh) Also, du hast diesen Furunkel, richtig? (sharp inhale) Das ist im Grunde eine gemeine, plötzliche Infektion, die tief drinnen eitert.
full caption & clip details
A young adult somewhat masculine voice; delivery is very low-energy, brisk, very tense, variable; timbre is slightly cool, very dark, fairly smooth, slightly thin; somewhat unclear, some disfluency, very wide pitch range, audible breath; affect is negative, very submissive, vulnerable; reads as astonishment surprise, fear, relief; style: dramatic, casual; below-average recording, quiet background; mildly explicit content; genuineness 2.6/6; vocal-burst blend 6.9/10; 7.2s, DE.
emolia_c0645__V__VALN__very_high__de.c006.k1 · in -21.4 dBFS · gain +2.6 dB · vprof_vc-00050
Jealousy and Envycloned voice   strict_117 · #40

This chain comes from the one-sided rule: only Jealousy and Envy had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Jealousy and Envy clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.25.

Nothing was asked of the other axis, and in fact Awe barely moves at all, sitting near 1.00 throughout.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

2 clips · 18 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.1 dB
per-clip
per-clip buttons play:
rule B1k 2qmax 0.250cmax 0.250d_a -0.009d_b 0.250dataset vprof_vclang entotal 17.6schain gain +2.9 dBseam step 0.1 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, good recording, slightly relaxed, light breath
(awe, pleasure ecstasy, affection · normal-paced, normally alert, fairly steady, narration) To behold that blessed leg, so richly adorned in dream, fills me with profound awe. It truly speaks of the clan's blossoming strength and the tribe's vast expansion.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, pleasure ecstasy, affection; style: narration, casual; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.7/10; 8.4s, EN.
emolia_c0760__E__Awe__A__en.c028.k1 · in -23.0 dBFS · gain +2.9 dB · vprof_vc-00058
(jealousy and envy, disgust, contempt · measured, energised, moderately variable, storytelling) Diese elenden Blumen sind so fad, völlig unscheinbar, doch sie leuchten mit diesem ekelhaften, lackartigen Glanz. (yawn) Es ist, als würde die ganze elende Schau zu sehr versuchen.
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as jealousy and envy, disgust, contempt; style: storytelling, cartoonish; good recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.5/10; 9.4s, DE.
emolia_c0760__E__Bitterness__C__de.c009.k0 · in -22.8 dBFS · gain +2.9 dB · vprof_vc-00058
Doubtcloned voice   strict_118 · #41

This chain comes from the one-sided rule: only Doubt had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Doubt clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.25.

Nothing was asked of the other axis, and in fact Fear drifts down from 1.00 to 0.84 (-0.16), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

2 clips · 44 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.0 dB
per-clip
per-clip buttons play:
rule B1k 2qmax 0.250cmax 0.250d_a -0.158d_b 0.250dataset vprof_vclang entotal 43.6schain gain +8.9 dBseam step 2.0 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · slightly cool, quiet background
(fear, disgust, helplessness · slow, very low-energy, fully relaxed, casual) This slushy ice is hitting me like tiny rocks! I can't breathe with this awful, freezing sleet!
full caption & clip details
A young adult masculine voice; delivery is very low-energy, slow, fully relaxed, variable; timbre is slightly cool, very dark, gravelly, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, heavy breath; affect is negative, submissive, vulnerable; reads as fear, disgust, helplessness; style: casual, playful; below-average recording, quiet background; mildly explicit content; genuineness 1.8/6; vocal-burst blend 1.8/10; 11.8s, EN.
emolia_c1682__X__pain_scream__en.c022.k3 · in -27.5 dBFS · gain +8.9 dB · vprof_vc-00100
(doubt, emotional numbness, teasing · fast, normally alert, slightly relaxed, casual) backslash slash backslash backslash backslash slash backslash backslash backslash slash backslash backslash backslash slash
full caption & clip details
A young adult somewhat masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, emotional numbness, teasing; style: casual, monologue; average recording, quiet background; explicit content; genuineness 1.6/6; vocal-burst blend 3.6/10; 32.0s, DE.
emolia_c1682__E__Concentration__C__de.c016.k0 · in -29.5 dBFS · gain +8.9 dB · vprof_vc-00100
Triumph ↓  /  Shamecloned voice   strict_092 · #42

This chain comes from the two-sided rule: it only counts if both emotions move — Triumph down and Shame up — by at least 0.25 each.

The chain starts with Shame clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.25.

At the same time Triumph goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.75 (higher than 75 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

2 clips · 32 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.3 dB
per-clip
per-clip buttons play:
rule AB2k 2qmax 0.250cmax 0.250d_a -0.250d_b 0.250dataset vprof_vclang entotal 31.9schain gain +5.3 dBseam step 0.3 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, balanced body, slightly relaxed
(triumph, pleasure ecstasy, teasing · measured, energised, volatile, casual) And with this adjustment, we've made those two triangles perfectly similar. (low mumble) This establishes the necessary geometric relationship we needed.
full caption & clip details
A young adult feminine voice; delivery is energised, measured, slightly relaxed, volatile; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; very slurred, frequent disfluency, very wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as triumph, pleasure ecstasy, teasing; style: casual, conversational; average recording, quiet background; genuineness 0.6/6; vocal-burst blend 0.0/10; 23.8s, EN.
emolia_c0323__V__S_DRAM__very_high__en.c002.k1 · in -25.4 dBFS · gain +5.3 dB · vprof_vc-00034
(shame, disappointment, sadness · normal-paced, normally alert, fairly steady, formal) Manchmal lässt dich der bittere Stich tiefer Enttäuschung jede Güte in Frage stellen, die du je gezeigt hast. Du fängst an zu fragen, ob all diese Güte wirklich etwas wert war.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, disappointment, sadness; style: formal, didactic; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.0/10; 8.1s, DE.
emolia_c0323__V__STRU__very_high__de.c027.k2 · in -25.1 dBFS · gain +5.3 dB · vprof_vc-00034
Sexual Lust ↓  /  Amusementcloned voice   strict_093 · #43

This chain comes from the two-sided rule: it only counts if both emotions move — Sexual Lust down and Amusement up — by at least 0.25 each.

The chain starts with Amusement clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.25.

At the same time Sexual Lust goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.75 (higher than 75 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

2 clips · 19 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 4.0 dB
per-clip
per-clip buttons play:
rule AB2k 2qmax 0.250cmax 0.250d_a -0.250d_b 0.250dataset vprof_vclang entotal 18.8schain gain +7.4 dBseam step 4.0 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · quiet background
(sexual lust, helplessness, pain · very slow, lethargic, fully relaxed, dramatic) Specifically, which regulatory frameworks govern the treatment of international site laborers and what protections do they possess? (pleasure moan) I need a precise overview of their codified entitlements.
full caption & clip details
An adult masculine voice; delivery is lethargic, very slow, fully relaxed, moderately variable; timbre is slightly warm, very dark, gravelly, very full; slurred, no disfluency, very wide pitch range, breathless; affect is deeply negative, dominant, guarded; reads as sexual lust, helplessness, pain; style: dramatic, cartoonish; below-average recording, quiet background; contains vocal bursts: Ahem; genuineness 1.2/6; vocal-burst blend 3.8/10; 2.8s, EN.
emolia_c0645__V__S_TECH__very_high__en.c012.k0 · in -24.2 dBFS · gain +7.4 dB · vprof_vc-00050
(amusement, pleasure ecstasy, teasing · normal-paced, energised, neutral tension, casual) Dieses hier hat viel mehr Chartbeispiele als alle zehn anderen Bücher zum Tageshandel, die ich habe. (swallows) Und das ohne Barrys zweites Chartbuch zu zählen, das auch großartig ist. (chuckle)
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, volatile; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as amusement, pleasure ecstasy, teasing; style: casual, conversational; good recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.5/10; 16.2s, DE.
emolia_c0645__V__VALS__very_high__de.c036.k0 · in -28.2 dBFS · gain +7.4 dB · vprof_vc-00050
Triumph ↓  /  Fearcloned voice   strict_080 · #44

This chain comes from the proxy rule: the same two-sided test as above, but because Fear is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fear clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.25.

At the same time Triumph goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.75 (higher than 75 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

2 clips · 18 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.3 dB
per-clip
per-clip buttons play:
rule PXRk 2qmax 0.250cmax 0.250d_a -0.250d_b 0.250dataset vprof_vclang entotal 17.6schain gain +3.4 dBseam step 2.3 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert, almost no disfluency, clear
(triumph, pride, relief · normal-paced, neutral tension, moderately variable, monologue) They finally approved it! (breathy giggle) The Norwegian parliament just greenlit the next shipment of those incredible F-thirty-five jets, a nearly four-billion-kroner deal. This is huge for our defense!
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, light breath; affect is mildly negative, neutral stance, vulnerable; reads as triumph, pride, relief; style: monologue, narration; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.0/10; 10.5s, EN.
emolia_c2046__E__Elation__C__en.c009.k0 · in -24.5 dBFS · gain +3.4 dB · vprof_vc-00115
(fear, awe, distress · brisk, slightly relaxed, fairly steady, casual) The sheer force of human blood rushing through those tiny vessels, capable of bursting into flame from ten meters away when breached... it's utterly terrifying.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as fear, awe, distress; style: casual, authoritative; good recording, no background noise; mildly explicit content; genuineness 0.3/6; vocal-burst blend 1.1/10; 7.3s, EN.
emolia_c2046__E__Awe__B__en.c019.k0 · in -22.2 dBFS · gain +3.4 dB · vprof_vc-00115
Triumph ↓  /  Contentmentcloned voice   strict_081 · #45

This chain comes from the proxy rule: the same two-sided test as above, but because Contentment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contentment clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.

At the same time Triumph goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.75 (higher than 75 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

2 clips · 24 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 6.2 dB
per-clip
per-clip buttons play:
rule PXRk 2qmax 0.250cmax 0.250d_a -0.250d_b 0.250dataset vprof_vclang entotal 24.1schain gain -0.4 dBseam step 6.2 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · slightly bright
(triumph, pleasure ecstasy, elation · fast, frantic, tense, ranting) It turns out that the very core of every ailment, every human frailty, traces back to the shadows of our emotions, even the most persistent bladder issues. We finally have the key to unlocking that connection.
full caption & clip details
A young adult masculine voice; delivery is frantic, fast, tense, volatile; timbre is slightly cool, slightly bright, very rough, thin; slurred, almost no disfluency, very wide pitch range, normal breath; affect is elated, very dominant, guarded; reads as triumph, pleasure ecstasy, elation; style: ranting, cartoonish; below-average recording, quiet background; genuineness 1.2/6; vocal-burst blend 1.5/10; 15.6s, EN.
emolia_c1950__E__Triumph__A__en.c017.k2 · in -18.3 dBFS · gain -0.4 dB · vprof_vc-00113
(contentment, thankfulness gratitude, affection · brisk, energised, slightly relaxed, casual) As the very Son of God who brought everything into being, Jesus celebrated creation without being bound by customs that warped its beauty. (contented sigh) He offered us a vision of perfect harmony, a source of endless hope for all of us.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contentment, thankfulness gratitude, affection; style: casual, dramatic; average recording, no background noise; genuineness 1.5/6; vocal-burst blend 4.2/10; 8.7s, EN.
emolia_c1950__E__Hope_Enthusiasm_Optimism__B__en.c034.k0 · in -24.4 dBFS · gain -0.4 dB · vprof_vc-00113
Impatience and Irritability ↓  /  Embarrassmentcloned voice   strict_094 · #46

This chain comes from the two-sided rule: it only counts if both emotions move — Impatience and Irritability down and Embarrassment up — by at least 0.25 each.

The chain starts with Embarrassment strongly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.25.

At the same time Impatience and Irritability goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.75 (higher than 75 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

2 clips · 17 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.3 dB
per-clip
per-clip buttons play:
rule AB2k 2qmax 0.249cmax 0.249d_a -0.249d_b 0.249dataset vprof_vclang entotal 16.6schain gain +4.9 dBseam step 0.3 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, neutral tension, moderately variable, average clarity
(impatience and irritability, jealousy and envy, anger · brisk, energised, some disfluency, casual) It's infuriating how competition law tries to stop businesses from cheating, but state aid law? It just scrutinizes how governments meddle, which is honestly so maddening.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as impatience and irritability, jealousy and envy, anger; style: casual, storytelling; good recording, no background noise; mildly explicit content; genuineness 1.3/6; vocal-burst blend 0.5/10; 9.3s, EN.
emolia_c0137__E__Jealousy_and_Envy__C__en.c031.k3 · in -25.1 dBFS · gain +4.9 dB · vprof_vc-00020
(embarrassment, shame, distress · normal-paced, normally alert, little disfluency, ranting) I feel so ashamed, I really do; I just used a snaffle, even though I knew it might not be quite right for the pull. (fast breathing) It's always a guessing game with the bit choice, honestly.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as embarrassment, shame, distress; style: ranting, storytelling; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 3.0/10; 7.4s, EN.
emolia_c0137__E__Shame__A__en.c033.k2 · in -24.7 dBFS · gain +4.9 dB · vprof_vc-00020
Concentration ↓  /  Interestcloned voice   strict_082 · #47

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.25.

At the same time Concentration goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.75 (higher than 75 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Guaranteed, without needing a check: one voice profile is one cloned voice by construction, so every clip here is the same synthetic speaker.

Voice consistency: not an issue here — every clip in this chain is the same cloned voice by construction, so there is no speaker mismatch to hear.

2 clips · 22 s · en · vprof_vc

Identity. a voice profile: one cloned voice per chain by construction, so there is no speaker identity to verify. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.7 dB
per-clip
per-clip buttons play:
rule PXRk 2qmax 0.249cmax 0.250d_a -0.250d_b 0.249dataset vprof_vclang entotal 21.9schain gain +2.4 dBseam step 0.7 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-bright, balanced body, good recording, slightly relaxed, clear
(concentration, pride, malevolence malice · brisk, normally alert, fairly steady, narration) This objective fits within the bachelor thesis introduction, outlining the precise rationale behind this investigation. (clears throat) We need to clearly establish the scope for our subsequent technical analysis.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as concentration, pride, malevolence malice; style: narration, dramatic; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.0/10; 9.8s, EN.
anime_062__V__S_TECH__moderately_high__en.c007.k2 · in -22.1 dBFS · gain +2.4 dB · vprof_vc-00004
(interest, awe, hope enthusiasm optimism · measured, very low-energy, moderately variable, storytelling) Diese intelligenten Technologien steigern wirklich, was jedes Gerät leisten kann. Sie eröffnen viel mehr Möglichkeiten, wie wir hier alles nutzen können. (contented sigh) (surprised gasp)
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as interest, awe, hope enthusiasm optimism; style: storytelling, narration; good recording, quiet background; genuineness 1.4/6; vocal-burst blend 1.0/10; 12.2s, DE.
anime_062__V__S_PLAY__moderately_low__de.c025.k1 · in -22.8 dBFS · gain +2.4 dB · vprof_vc-00004