What this rule requires. only the target emotion must be reached; the start axis is unconstrained — with a per-step cap of C = 0.25, and here only chains whose speaker similarity clears 0.80 on both conditions. No voice conversion has been applied.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken. Underlined descriptors are the ones that change across the chain; brackets inside the words are a real non-speech sound. The perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
Relief ↑timbre ≥0.80 strict_075 · #1
This chain comes from the one-sided rule: only Relief had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Relief essentially absent — 0.08, lower than 92 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.92.
Nothing was asked of the other axis, and in fact Confusion drifts down from 0.98 to 0.90 (-0.08), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.23, then +0.25, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 23 s · en · emolia
Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.836 and the worst against the first clip 0.823; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/100/150/150 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 4.2 dB
Unchanged across all 5 clips: a young adult masculine voice · quiet background, normally alert, some disfluency
(confusion, astonishment surprise, doubt · measured, fully relaxed, fairly steady, casual)Why are there two seemingly exit zones?
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, fully relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; slurred, some disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as confusion, astonishment surprise, doubt; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 1.0/10; 3.2s, EN.
EN_E1mqQZfqRuu_W000130 · in -22.5 dBFS · gain +3.3 dB · emolia-00780
(awe, astonishment surprise, confusion ·normal-paced, slightly relaxed, moderately variable, casual)I can't remember, does this guy get ma- wow. It even shows up in the overworld.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as awe, astonishment surprise, confusion; style: casual, conversational; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 1.9/10; 4.1s, EN.
EN_E1mqQZfqRuu_W000133 · in -21.3 dBFS · gain +3.3 dB · emolia-00780
(infatuation, affection, teasing· normal-paced, slightly relaxed, moderately variable, casual)(ahem) Heh, managed to dodge him. That was lucky. What happens if I go to the overworld with you?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, dark, slightly rough, thin; slurred, some disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as infatuation, affection, teasing; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 6.0/6; vocal-burst blend 1.1/10; 5.2s, EN.
EN_E1mqQZfqRuu_W000134 · in -25.5 dBFS · gain +3.3 dB · emolia-00780
(pain, embarrassment, helplessness· normal-paced, slightly relaxed, fairly steady, casual)That one just does damage to me like normal because of the Yoshi. Okay. Let's try the two left.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, balanced body; slurred, some disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as pain, embarrassment, helplessness; style: casual, whispered; below-average recording, quiet background; mildly explicit content; genuineness 4.0/6; vocal-burst blend 1.3/10; 6.3s, EN.
EN_E1mqQZfqRuu_W000137 · in -24.8 dBFS · gain +3.3 dB · emolia-00780
(relief, helplessness, fear·measured, slightly relaxed, fairly steady, casual)Let's go, Raft. I'm not trying to kill myself anymore. Whoops.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as relief, helplessness, fear; style: casual, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.5/10; 4.9s, EN.
EN_E1mqQZfqRuu_W000138 · in -23.0 dBFS · gain +3.3 dB · emolia-00780
Anger ↑timbre ≥0.80 strict_078 · #2
This chain comes from the one-sided rule: only Anger had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Anger essentially absent — 0.03, lower than 97 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.89.
Nothing was asked of the other axis, and in fact Pain drifts down from 0.95 to 0.35 (-0.60), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.20, then +0.25, then +0.20 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 36 s · zh · emolia
Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.866 and the worst against the first clip 0.898; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/100/100/150 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.5 dB
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, average recording, normally alert, slightly relaxed, moderate pitch range
(pain · measured, fairly steady, some disfluency, didactic)我相信你们一定会为公司当前面对的问题找到一条出路。解释过后呢,他就离开了。几小时后,他再一次回到会议室。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: didactic, monologue; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 0.6/10; 10.7s, ZH.
ZH_B00059_S03639_W000011 · in -18.4 dBFS · gain -1.1 dB · emolia-03871
(normal-paced, fairly steady, no disfluency, formal)发现这一群人已经想出了一个突破性的解决办法。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, didactic; average recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.4/10; 4.2s, ZH.
ZH_B00059_S03639_W000012 · in -18.3 dBFS · gain -1.1 dB · emolia-03871
(measured, fairly steady, some disfluency, didactic)这是他的高级研究人员和发展人员忽略了的方法。而他们呢早出来了,这件事情就证明了一点。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 2.2/10; 9.2s, ZH.
ZH_B00059_S03639_W000013 · in -19.0 dBFS · gain -1.1 dB · emolia-03871
(distress·fast, steady, no disfluency, narration)就是那些职位低微的人和那些高级行政人员的创造能力都是一样的。
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as distress; style: narration, monologue; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 1.7/10; 6.6s, ZH.
ZH_B00059_S03639_W000014 · in -19.2 dBFS · gain -1.1 dB · emolia-03871
(anger· fast, fairly steady, almost no disfluency, authoritative)只不过,那些处在低位的人,不相信自己,也不相信自己的想法罢了。
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as anger; style: authoritative, dramatic; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 3.1/10; 5.6s, ZH.
ZH_B00059_S03639_W000015 · in -20.7 dBFS · gain -1.1 dB · emolia-03871
This chain comes from the one-sided rule: only Thankfulness Gratitude had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Thankfulness Gratitude essentially absent — 0.06, lower than 94 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.87.
Nothing was asked of the other axis, and in fact Shame drifts down from 0.91 to 0.02 (-0.89), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.22, then +0.17, then +0.24 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 47 s · en · emolia
Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.826 and the worst against the first clip 0.843; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/100/150/150 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.1 dB
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, moderate pitch range
(shame · normal-paced, fairly steady, some disfluency, monologue)Which I can check in the dictionary, but I do know that it means to ruin a reputation.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as shame; style: monologue, authoritative; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 1.1/10; 5.3s, EN.
EN_B00030_S05057_W000102 · in -21.5 dBFS · gain +2.3 dB · emolia-00835
(measured, fairly steady, frequent disfluency, didactic)So I might just make a note there to remind myself to ruin a reputation. So here, great, I've got the word, I've got the grammar, I've got the collocations,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 0.8/10; 13.1s, EN.
EN_B00030_S05057_W000103 · in -22.4 dBFS · gain +2.3 dB · emolia-00835
(concentration, interest·normal-paced, moderately variable, frequent disfluency, conversational)Now, I know it's a lot of work, but what I'm doing is I'm not just practicing this word, right? I'm practicing different collocations to give me flexibility, essential for a band 9, uhm, (low mumble) and in my phrase I'm also
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, interest; style: conversational, casual; good recording, no background noise; genuineness 3.5/6; vocal-burst blend 3.6/10; 17.0s, EN.
EN_B00030_S05057_W000104 · in -21.9 dBFS · gain +2.3 dB · emolia-00835
(measured, fairly steady, some disfluency, whispered)Practicing different language, not just the word reputation. Practicing generous, practicing very,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: whispered, monologue; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 0.6/10; 6.3s, EN.
EN_B00030_S05057_W000105 · in -24.0 dBFS · gain +2.3 dB · emolia-00835
(thankfulness gratitude, pride·normal-paced, fairly steady, some disfluency, conversational)It's great, great practice. So that is a simple example of how I might record a word.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as thankfulness gratitude, pride; style: conversational, monologue; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.1/10; 6.1s, EN.
EN_B00030_S05057_W000106 · in -22.9 dBFS · gain +2.3 dB · emolia-00835
This chain comes from the one-sided rule: only Thankfulness Gratitude had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Thankfulness Gratitude barely there — 0.17, lower than 83 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.81.
Nothing was asked of the other axis, and in fact Longing drifts down from 0.99 to 0.91 (-0.08), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.09, then +0.25, then +0.24 — an uneven climb, but always in the same direction.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.80 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 27 s · en · podcast
Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.801 and the worst against the first clip 0.827; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150/100/150 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.0 dB
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, average recording, normal-paced, some disfluency, light breath
(longing, relief, infatuation · normally alert, neutral tension, fairly steady, casual)Like and I it was one of those days where I wasn't feeling it, so I was like, maybe if I dress like good, then I'll I'll wanna go work out and show it off.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as longing, relief, infatuation; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 8.4/10; 6.6s, EN.
51756_00456200 · in -29.7 dBFS · gain +8.8 dB · podcast-06012
(pleasure ecstasy, sexual lust, triumph· normally alert, neutral tension, moderately variable, casual)And (low mumble) uh once I wore 'em to the gym, that was it was it from there. Like I just fucking that's when I was like, fuck it. Like I would I'd go to like the club and those. I I just
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as pleasure ecstasy, sexual lust, triumph; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.9/6; vocal-burst blend 10.0/10; 7.0s, EN.
51756_00456912 · in -27.7 dBFS · gain +8.8 dB · podcast-06015
(impatience and irritability, intoxication altered states of consciousness, contempt·energised, neutral tension, moderately variable, casual)you've gone to the gym in those shoes, they're they're not fucked, but it's like they're kinda like
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as impatience and irritability, intoxication altered states of consciousness, contempt; style: casual, conversational; average recording, no background noise; explicit content; genuineness 4.8/6; vocal-burst blend 2.1/10; 3.6s, EN.
51756_00457800 · in -28.3 dBFS · gain +8.8 dB · podcast-06029
(affection, infatuation, embarrassment·normally alert, slightly relaxed, fairly steady, casual)at the same time, it's like but like every now and then I'll clean 'em up to put 'em with a nice outfit, you know, if I
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as affection, infatuation, embarrassment; style: casual, conversational; average recording, no background noise; genuineness 5.2/6; vocal-burst blend 7.3/10; 4.3s, EN.
51756_00458312 · in -29.1 dBFS · gain +8.8 dB · podcast-06002
(thankfulness gratitude, embarrassment, astonishment surprise· normally alert, neutral tension, moderately variable, casual)(low mumble) Um, but yeah, with these, I was like that for a while 'cause I was like, I spent like I that's probably the most money I've ever spent on shoes.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as thankfulness gratitude, embarrassment, astonishment surprise; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 7.6/10; 6.6s, EN.
51756_00458896 · in -29.5 dBFS · gain +8.8 dB · podcast-06007
Relief ↑timbre ≥0.80 strict_070 · #5
This chain comes from the one-sided rule: only Relief had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Relief barely there — 0.20, lower than 80 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.75.
Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism drifts down from 0.90 to 0.72 (-0.18), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 29 s · en · emolia
Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.865 and the worst against the first clip 0.882; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/100/150 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.3 dB
Unchanged across all 4 clips: a young adult feminine voice · fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, light breath
(hope enthusiasm optimism · fairly steady, little disfluency, clear, casual)In tutorial this week, we're gonna be discussing your experiences of self-directed learning.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as hope enthusiasm optimism; style: casual, monologue; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.4/10; 5.1s, EN.
EN_veU-QBiNbJy_W000073 · in -20.0 dBFS · gain -0.3 dB · emolia-02526
(doubt, concentration·steady, some disfluency, clear, whispered)Do you think that they need to be introverts or extroverts? What's the learning style of a self directed learner? Does level of education affect their ability to be self directed?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, some disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, neutral openness; reads as doubt, concentration; style: whispered, monologue; average recording, quiet background; genuineness 0.3/6; vocal-burst blend 0.6/10; 10.2s, EN.
EN_veU-QBiNbJy_W000074 · in -20.1 dBFS · gain -0.3 dB · emolia-02526
(doubt, contemplation·fairly steady, some disfluency, clear, whispered)And readiness, how do you know if you are or they are? Is this kind of autonomy innate or is it situational? Are you born with it or can you learn how to be more self-directed?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as doubt, contemplation; style: whispered, casual; good recording, quiet background; genuineness 1.0/6; vocal-burst blend 1.0/10; 10.8s, EN.
EN_veU-QBiNbJy_W000075 · in -20.1 dBFS · gain -0.3 dB · emolia-02526
(relief· fairly steady, some disfluency, average clarity, casual)And finally, does technology facilitate self-directed learning?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as relief; style: casual, storytelling; good recording, no background noise; genuineness 3.1/6; vocal-burst blend 2.9/10; 3.1s, EN.
EN_veU-QBiNbJy_W000076 · in -17.8 dBFS · gain -0.3 dB · emolia-02526
Contemplation ↑timbre ≥0.80 strict_074 · #6
This chain comes from the one-sided rule: only Contemplation had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Contemplation barely there — 0.18, lower than 82 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.73.
Nothing was asked of the other axis, and in fact Relief drifts down from 0.86 to 0.11 (-0.75), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.24, then +0.24 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 33 s · en · emolia
Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.820 and the worst against the first clip 0.820; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/100/150 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.4 dB
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, balanced body, normal-paced, normally alert, some disfluency, average clarity, moderate pitch range, light breath
(slightly relaxed, fairly steady, casual, conversational)Last couple of terms we have moving on a little bit in the song.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 2.6/10; 3.5s, EN.
EN_B00033_S08506_W000030 · in -19.5 dBFS · gain +0.8 dB · emolia-00889
(teasing·neutral tension, moderately variable, casual, conversational)Talk, you're talking, go viral. So (ahem) uhm, go viral. She's, she's talking probably about the media is getting a nice headline saying something scandalous about Taylor Swift and their videos going viral because people love to, haters gonna hate, right?
full caption & clip details
A young adult somewhat masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as teasing; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 3.8/6; vocal-burst blend 6.0/10; 13.9s, EN.
EN_B00033_S08506_W000031 · in -21.9 dBFS · gain +0.8 dB · emolia-00889
(infatuation, elation, interest·slightly relaxed, fairly steady, casual, monologue)And the last one, I just need this love spiral. Like spiral is literally something that does like this, right? (ahem) Uh, that makes a circular.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as infatuation, elation, interest; style: casual, monologue; good recording, quiet background; genuineness 3.0/6; vocal-burst blend 3.1/10; 8.5s, EN.
EN_B00033_S08506_W000032 · in -21.5 dBFS · gain +0.8 dB · emolia-00889
(interest, contemplation, confusion·relaxed, fairly steady, casual, conversational)Moving inward motion and when you talk about like someone's in a spiral, it's kind of like they're going round and round like a cycle. You could say as well.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as interest, contemplation, confusion; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 5.7/10; 7.8s, EN.
EN_B00033_S08506_W000033 · in -19.6 dBFS · gain +0.8 dB · emolia-00889
Anger ↑identity ≥0.80 strict_072 · #7
This chain comes from the one-sided rule: only Anger had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Anger below average — 0.28, lower than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.70.
Nothing was asked of the other axis, and in fact Thankfulness Gratitude drifts down from 0.99 to 0.57 (-0.42), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.23, then +0.22 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 47 s · mt · eurospeech
Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.901 and the worst against the first clip 0.901; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.5 dB
Unchanged across all 4 clips: a child feminine voice · slightly cool, fairly smooth, thin, average recording, quiet background, normally alert, slightly relaxed, moderately variable
(thankfulness gratitude · fast, frequent disfluency, average clarity, cartoonish)importanti li ttaxxa tibqa’ kompetittiva u attraenti sabiex ma jkunx hemm piż żejjed, speċjalment fuq il-kumpaniji ż-żgħar.
full caption & clip details
A child feminine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, frequent disfluency, very wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as thankfulness gratitude; style: cartoonish, playful; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.0/10; 10.1s, MT.
malta_13_099_11042018_10791680_10801769 · in -26.3 dBFS · gain +6.1 dB · eurospeech-02113
(concentration·normal-paced, some disfluency, clear, playful)Se nbiddel is-suġġett u se mmur għasseba’ punt tiegħi, li jittratta limportanza li jkun hemm mudell ħolistiku fuq l-edukazzjoni u t-taħriġ
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, some disfluency, very wide pitch range, normal breath; affect is neutral, slightly dominant, neutral openness; reads as concentration; style: playful, didactic; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 1.0/10; 12.0s, MT.
malta_13_099_11042018_10837712_10849696 · in -25.9 dBFS · gain +6.1 dB · eurospeech-02113
(measured, some disfluency, average clarity, didactic)(ahem) sabiex iż-żgħażagħ tagħna jkunu mħarrġa u preparati biex jaħdmu f’dan is-settur. Huwa fatt li hemm domanda kbira
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, very wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; no dominant emotion; style: didactic, dramatic; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.8/10; 10.8s, MT.
malta_13_099_11042018_10849696_10860511 · in -26.0 dBFS · gain +6.1 dB · eurospeech-02113
(anger, concentration, bitterness·normal-paced, frequent disfluency, average clarity, cartoonish)lid-data biex verament inkunu nafu l-livell ta’ dipendenza li pajjiżna għandu fuq dan is-settur u l-livell ta’ kontribut li dan is-settur għandu fl-ekonomija ta’ pajjiżna. Hekk biss
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, slightly dark, fairly smooth, thin; average clarity, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as anger, concentration, bitterness; style: cartoonish, playful; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 0.9/10; 14.7s, MT.
malta_13_099_11042018_11049872_11064608 · in -26.5 dBFS · gain +6.1 dB · eurospeech-02113
This chain comes from the one-sided rule: only Malevolence Malice had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Malevolence Malice barely there — 0.21, lower than 79 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.62.
Nothing was asked of the other axis, and in fact Concentration drifts down from 0.85 to 0.75 (-0.10), which the rule did not require.
It takes 4 clips to get there. Clip to clip the moves are +0.21, then +0.22, then +0.19 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 29 s · snippets
Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.831 and the worst against the first clip 0.892; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/150/100 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.7 dB
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, good recording, no background noise, normally alert, slightly relaxed, light breath
(measured, fairly steady, almost no disfluency, didactic)Philadelphia architect Thomas S. Stewart adapted Lefever's capital for the interior pilaster capitals of Richmond's St. Paul's Church.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: didactic, monologue; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.2/10; 10.3s.
batch60_part1_batch60_part1_chunk_1540_1_1877403 · in -26.0 dBFS · gain +6.4 dB · snippets-01196
(emotional numbness· measured, fairly steady, almost no disfluency, monologue)we see that the in-base have daylight or void between the columns.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as emotional numbness; style: monologue, formal; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.6/10; 5.5s.
batch60_part1_batch60_part1_chunk_1540_1_1877444 · in -25.1 dBFS · gain +6.4 dB · snippets-01196
(measured, fairly steady, some disfluency, monologue)And it was aimed not at the homeowners, but at the person doing the work of building.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: monologue, authoritative; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.2/10; 6.2s.
batch60_part1_batch60_part1_chunk_1540_1_1877534 · in -27.8 dBFS · gain +6.4 dB · snippets-01196
(slow, steady, almost no disfluency, narration)This is a design composed of antique specimens and reduced to accurate proportions.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, full; very clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.0/10; 7.2s.
batch60_part1_batch60_part1_chunk_1540_1_1877547 · in -26.9 dBFS · gain +6.4 dB · snippets-01196
Contemplation ↑identity ≥0.80 strict_077 · #9
This chain comes from the one-sided rule: only Contemplation had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Contemplation below average — 0.35, lower than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.61.
Nothing was asked of the other axis, and in fact Malevolence Malice drifts down from 0.92 to 0.66 (-0.26), which the rule did not require.
It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.22, then +0.01, then +0.21 — a plateau around step 3, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.80. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 31 s · snippets
Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.821 and the worst against the first clip 0.803; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/100/150/100 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.3 dB
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(malevolence malice · slow, fairly steady, no disfluency, formal)or even other Assyrians who had rebelled.
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice; style: formal, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 3.4/10; 3.3s.
batch232_part2_batch232_part2_chunk_555_1_554767 · in -26.4 dBFS · gain +7.8 dB · snippets-00689
(malevolence malice, awe, shame·measured, fairly steady, no disfluency, monologue)And killed that sea creature known as a Nahiru.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, awe, shame; style: monologue, narration; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 4.3/10; 3.6s.
batch232_part2_batch232_part2_chunk_555_1_554824 · in -28.7 dBFS · gain +7.8 dB · snippets-00689
(pride, fear, malevolence malice · measured, fairly steady, almost no disfluency, narration)and instead, set about destroying the cities with the same viciousness that the Assyrians had once reserved for the cities of Elam.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, fear, malevolence malice; style: narration, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.8/10; 9.1s.
batch232_part2_batch232_part2_chunk_555_1_554836 · in -27.8 dBFS · gain +7.8 dB · snippets-00689
(sadness, disappointment· measured, steady, almost no disfluency, narration)For centuries now, the powerful Elamites had been their rivals and kept their ambitions in check.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sadness, disappointment; style: narration, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.3/10; 7.2s.
batch232_part2_batch232_part2_chunk_555_1_554870 · in -28.1 dBFS · gain +7.8 dB · snippets-00689
(contemplation, emotional numbness· measured, steady, almost no disfluency, narration)and it shows that in the social upheaval of this period of chaos, some of the power of the nobility was being eroded.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, emotional numbness; style: narration, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.8/10; 8.7s.
batch232_part2_batch232_part2_chunk_555_1_554887 · in -27.6 dBFS · gain +7.8 dB · snippets-00689
This chain comes from the one-sided rule: only Malevolence Malice had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Malevolence Malice around average — 0.49, lower than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.50.
Nothing was asked of the other axis, and in fact Contemplation drifts down from 0.96 to 0.58 (-0.39), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 35 s · french · mls
Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.907 and the worst against the first clip 0.907; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.2 dB
Unchanged across all 3 clips: an elderly masculine voice · balanced body, measured, steady, fairly narrow pitch, light breath
(contemplation, shame, disappointment · very low-energy, relaxed, frequent disfluency, whispered)que l'on joue quelquefois cette actrice dit-il à martin me plaît beaucoup elle a un faux air de mademoiselle cunégonde je serais bien aise de la saluer
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, neutral openness; reads as contemplation, shame, disappointment; style: whispered, ASMR; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 3.5/10; 12.7s, FRENCH.
1840_1870_000034 · in -21.8 dBFS · gain +0.7 dB · mls-00062
(emotional numbness, sexual lust·normally alert, slightly relaxed, no disfluency, formal)l'abbé périgourdin s'offrit à l'introduire chez elle candide élevé en allemagne demanda quelle était l'étiquette et comment on traitait en france les reines d'angleterre
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, sexual lust; style: formal, monologue; average recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 12.3s, FRENCH.
1840_1870_000310 · in -20.9 dBFS · gain +0.7 dB · mls-00062
(malevolence malice, sourness, fear· normally alert, slightly relaxed, no disfluency, monologue)il faut distinguer dit l'abbé en province on les mène au cabaret à paris on les respecte quand elles sont belles et on les jette à la voirie quand elles sont mortes
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, sourness, fear; style: monologue, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 10.6s, FRENCH.
1840_1870_000283 · in -19.7 dBFS · gain +0.7 dB · mls-00062
Elation ↑timbre ≥0.80 strict_066 · #11
This chain comes from the one-sided rule: only Elation had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Elation around average — 0.46, lower than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.50.
Nothing was asked of the other axis, and in fact Contempt drifts down from 0.92 to 0.14 (-0.79), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 24 s · en · emolia
Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.815 and the worst against the first clip 0.815; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/100 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.0 dB
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, slightly relaxed, some disfluency, average clarity
(contempt, sourness · normal-paced, normally alert, fairly steady, monologue)(low mumble) Uhm, then the party is in deep doo-doo because (low mumble) uhm, they have not demonstrated unity.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, sourness; style: monologue, formal; good recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.2/10; 5.3s, EN.
EN_B00018_S07976_W000024 · in -17.9 dBFS · gain -1.9 dB · emolia-00605
(interest, concentration· normal-paced, energised, moderately variable, didactic)So that's a very important part of the coalition system, and that's something that we share between humans and chimpanzees. Now, how do you become an alpha male? First of all,
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as interest, concentration; style: didactic, authoritative; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.5/10; 11.7s, EN.
EN_B00018_S07976_W000025 · in -18.6 dBFS · gain -1.9 dB · emolia-00605
(elation, pride, triumph·brisk, normally alert, moderately variable, playful)You need to be impressive and intimidating and demonstrate your vigor on occasion and show that you're very strong and there's all sorts of ways of doing that.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as elation, pride, triumph; style: playful, casual; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 2.8/10; 7.7s, EN.
EN_B00018_S07976_W000026 · in -17.6 dBFS · gain -1.9 dB · emolia-00605
Contemplation ↑identity ≥0.80 strict_067 · #12
This chain comes from the one-sided rule: only Contemplation had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Contemplation around average — 0.47, lower than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.50.
Nothing was asked of the other axis, and in fact Pleasure Ecstasy drifts down from 0.93 to 0.30 (-0.63), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 41 s · nl · podcast
Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.900 and the worst against the first clip 0.895; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 4.9 dB
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, quiet background, normal-paced, some disfluency, average clarity, light breath
(pleasure ecstasy · normally alert, neutral tension, moderately variable, conversational)Metezakjes, hier heb je jezelf. Hier is je staggebegeleider, hier is je opleider, (low mumble) hier is de uitkering. (low mumble) Wie zit er nou op je nek? En met die Marokkaanse Meluksiejonges, kan je er van
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as pleasure ecstasy; style: conversational, playful; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 5.1/10; 12.6s, NL.
819784_00258120 · in -25.3 dBFS · gain +8.0 dB · podcast-04624
(contentment, thankfulness gratitude, relief·very low-energy, slightly relaxed, fairly steady, casual)(ahem) Groen heb ik twee edities gedraaid en met de Marokkaan en Melukers. Eentje. (low mumble) En allebei grote successen. (low mumble) En toen liepen ze af en toen was de vraag wat nu. En er was daar in de wijk geen draagkracht met de Molukse en Marokkaanse jongeren voor een nieuw project. Politie en gemeente wilde het heel graag, maar er was geen draagkracht in de wijk.
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, thankfulness gratitude, relief; style: casual, monologue; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 2.8/10; 20.5s, NL.
819784_00260623 · in -30.2 dBFS · gain +8.0 dB · podcast-04648
(contemplation, pain·normally alert, slightly relaxed, fairly steady, casual)bij Groen, (low mumble) was mijn draagkracht een beetje op. En toen heb ik dus zelf in mijn eigen interview een opstelling gedaan wat
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, pain; style: casual, conversational; good recording, quiet background; genuineness 3.9/6; vocal-burst blend 0.2/10; 8.2s, NL.
819784_00262768 · in -29.9 dBFS · gain +8.0 dB · podcast-04614
Shame ↑identity ≥0.80 strict_068 · #13
This chain comes from the one-sided rule: only Shame had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Shame around average — 0.47, lower than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.49.
Nothing was asked of the other axis, and in fact Thankfulness Gratitude drifts down from 0.97 to 0.70 (-0.27), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 43 s · hr · eurospeech
Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.930 and the worst against the first clip 0.946; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.7 dB
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, slightly dark, slightly rough, balanced body, average recording, quiet background, measured, fairly steady
(thankfulness gratitude · subdued, slightly relaxed, normal breath, monologue)što znači da imaju svu potrebnu dokumentaciju te 265 poduzetničkih potpornih institucija od kojih su 203 verificirane. Za sam postupak verifikacije
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: monologue, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.3/10; 13.1s, HR.
croatia_20211202090948-27559_1642368_1655488 · in -17.8 dBFS · gain -3.0 dB · eurospeech-01547
(disappointment, anger, triumph·normally alert, neutral tension, light breath, monologue)osnivači dostavljaju nadležnom ministarstvu propisanu dokumentaciju i u Registar upisuju tražene podatke iz čega proizlazi da se verifikacija temelji na administrativnom postupanju osnivača, dakle predlaganje donošenja odluka o osnivanju
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, anger, triumph; style: monologue, cartoonish; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.6/10; 17.2s, HR.
croatia_20211202090948-27559_1655488_1672704 · in -15.5 dBFS · gain -3.0 dB · eurospeech-01547
(shame, sadness, disgust· normally alert, slightly relaxed, light breath, monologue)na predstavničkom tijelu jedinice lokalne samouprave, zatim dostavi prikaz obuhvata zone i utvrđivanja strukture vlasništva na temelju zemljišnoknjižnih ili katastarskih
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, sadness, disgust; style: monologue, casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.9/10; 13.4s, HR.
croatia_20211202090948-27559_1672704_1686080 · in -19.2 dBFS · gain -3.0 dB · eurospeech-01547
Bitterness ↑identity ≥0.80 strict_069 · #14
This chain comes from the one-sided rule: only Bitterness had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Bitterness below average — 0.35, lower than 65 % of clips in this corpus — and ends with it strongly present at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.47.
Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.83 to 0.54 (-0.29), which the rule did not require.
It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.22 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 22 s · snippets
Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.863 and the worst against the first clip 0.863; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/100 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.1 dB
Unchanged across all 3 clips: an adult masculine voice · neutral-bright, fairly smooth, balanced body, good recording, no background noise, slightly relaxed, fairly steady, light breath
(measured, normally alert, almost no disfluency, monologue)and beyond that he comes up with a general procedure that can be used to find the sum of fifth powers or sixth powers.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.4/10; 7.6s.
batch73_part3_batch73_part3_chunk_1655_1_1508129 · in -22.0 dBFS · gain +2.0 dB · snippets-01267
(normal-paced, energised, almost no disfluency, authoritative)or the book of chapters on Hindu arithmetic.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.4/10; 3.5s.
batch73_part3_batch73_part3_chunk_1655_1_1508223 · in -21.9 dBFS · gain +2.0 dB · snippets-01267
(measured, energised, some disfluency, didactic)(ahem) Uh, he would actually publish 92 scientific works, many of which dealt with optics, many of which dealt with mathematics, other scientific topics.
full caption & clip details
An adult masculine voice; delivery is energised, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; very clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; no dominant emotion; style: didactic, monologue; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.0/10; 10.8s.
batch73_part3_batch73_part3_chunk_1655_1_1508233 · in -22.1 dBFS · gain +2.0 dB · snippets-01267
Pride ↑timbre ≥0.80 strict_060 · #15
This chain comes from the one-sided rule: only Pride had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Pride clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.
Nothing was asked of the other axis, and in fact Concentration drifts down from 0.97 to 0.21 (-0.76), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 16 s · en · emolia
Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.917 and the worst against the first clip 0.917; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.9 dB
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, quiet background, brisk, normally alert, slightly relaxed
(concentration · casual, conversational)Yeah, in this particular case, I agree. This particular case, in fact, it does not have any commission noise. The noise here is really the quantization noise. So I don't, I haven't even introduced a commission noise here. So this is basically just- Oh, okay.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration; style: casual, conversational; good recording, quiet background; genuineness 3.2/6; vocal-burst blend 2.8/10; 11.4s, EN.
EN_B00047_S02980_W000267 · in -22.2 dBFS · gain +0.7 dB · emolia-01146
(pride, impatience and irritability, triumph· casual, conversational)Yeah, just because I-I scale up the bits, so naturally my quantizing noise reduces, right? I'm not-
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as pride, impatience and irritability, triumph; style: casual, conversational; good recording, quiet background; genuineness 4.4/6; vocal-burst blend 5.0/10; 4.4s, EN.
EN_B00047_S02980_W000268 · in -18.3 dBFS · gain +0.7 dB · emolia-01146
This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Hope Enthusiasm Optimism clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.
Nothing was asked of the other axis, and in fact Amusement drifts down from 1.00 to 0.83 (-0.17), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.80 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.80. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 40 s · en · podcast
Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.804 and the worst against the first clip 0.804; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.1 dB
Unchanged across all 2 clips: a young adult feminine voice · slightly cool, brisk, energised, some disfluency
(amusement, embarrassment, intoxication altered states of consciousness · neutral tension, volatile, slurred, casual)(ahem) Um if it's personal related to community. However, you want to answer that question about keeping it real. (breathy giggle) Now she's got a modest.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, volatile; timbre is slightly cool, bright, gravelly, thin; slurred, some disfluency, very wide pitch range, heavy breath; affect is positive, slightly submissive, very vulnerable; reads as amusement, embarrassment, intoxication altered states of consciousness; style: casual, conversational; below-average recording, some background noise; genuineness 3.7/6; vocal-burst blend 1.4/10; 12.4s, EN.
618997_00264216 · in -14.6 dBFS · gain -3.5 dB · podcast-00160
(hope enthusiasm optimism, thankfulness gratitude, elation·slightly relaxed, moderately variable, average clarity, dramatic)hope you enjoyed the Mindsight Collective podcast and you have been uplifted. Be sure to subscribe to Stay Wise, follow if you want freedom tomorrow, rate and comment if you loved our content. Let's get the conversation going on Instagram at MindSight Collective and on Twitter, MindSight Tweets. Don't forget to go to our platform at mindsidecollective dot com for more provocative thought, art, music and inspiration.
full caption & clip details
A child feminine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, thankfulness gratitude, elation; style: dramatic, casual; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 3.1/10; 27.6s, EN.
618997_00309256 · in -17.8 dBFS · gain -3.5 dB · podcast-01143
Shame ↑identity ≥0.80 strict_062 · #17
This chain comes from the one-sided rule: only Shame had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Shame clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.25.
Nothing was asked of the other axis, and in fact Concentration drifts down from 0.98 to 0.90 (-0.08), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 35 s · da · eurospeech
Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.877 and the worst against the first clip 0.877; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 6.2 dB
Unchanged across all 2 clips: an adult feminine voice · neutral-bright, balanced body, average recording, quiet background, average clarity
(concentration, disappointment, fear · brisk, normally alert, slightly relaxed, formal)borgerforslag om eksempelvis afskaffelse af uddannelsesloftet og Det Konservative Folkeparti står på som forslagsstillere, men ikke bakker op om forslaget i sidste ende? Der har vi også et ansvar over for hinanden (ahem) i forhold til at respektere, at man kan være medforslagsstiller uden nødvendigvis at bakke op om selve (ahem) indholdet i det
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration, disappointment, fear; style: formal, authoritative; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.2/10; 19.8s, DA.
denmark_20171M078_2018-04-06_1000_7467584_7487360 · in -24.9 dBFS · gain +6.6 dB · eurospeech-00324
(shame, helplessness, longing·slow, very low-energy, neutral tension, monologue)i det enkelte borgerforslag. Jeg tror, at vi alle sammen skal starte med os selv, og det er der jo i virkeligheden rigtig mange, der har sagt meget, meget rigtigt og meget, meget
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, audible breath; affect is mildly positive, neutral stance, neutral openness; reads as shame, helplessness, longing; style: monologue, cartoonish; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.0/10; 15.6s, DA.
denmark_20171M078_2018-04-06_1000_7487360_7502912 · in -31.0 dBFS · gain +6.6 dB · eurospeech-00324
Confusion ↑identity ≥0.80 strict_063 · #18
This chain comes from the one-sided rule: only Confusion had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Confusion clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.
Nothing was asked of the other axis, and in fact Infatuation drifts down from 0.89 to 0.84 (-0.05), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 24 s · dutch · mls
Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.911 and the worst against the first clip 0.911; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.0 dB
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, slightly dark, slightly rough, balanced body, average recording, subdued, slightly relaxed, steady
(measured, some disfluency, normal breath, narration)toen zij de bogt van frankrijk door en in de spaansche zee gekomen waren werd het weêr ongestadiger en somtijds zelfs eenigzins stormachtig waarom nu en dan zich maurits meer beneden hield
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; average recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.0/10; 14.3s, DUTCH.
2450_10565_000440 · in -28.9 dBFS · gain +9.2 dB · mls-00087
(confusion, fatigue exhaustion·slow, frequent disfluency, audible breath, narration)dagelijks volgde hij de vermaning van zijne moeder en las een gedeelte uit den bijbel en wel bijzonder die gedeelten welke zij hem had aanbevolen
full caption & clip details
A middle-aged masculine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, fatigue exhaustion; style: narration, monologue; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.0/10; 10.0s, DUTCH.
2450_10565_001979 · in -29.9 dBFS · gain +9.2 dB · mls-00088
This chain comes from the one-sided rule: only Fatigue Exhaustion had to get where it was going, by at least 0.25. The other emotion was left completely free.
The chain starts with Fatigue Exhaustion clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.25.
Nothing was asked of the other axis, and in fact Doubt drifts down from 0.99 to 0.11 (-0.87), which the rule did not require.
It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
2 clips · 11 s · snippets
Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.842 and the worst against the first clip 0.842; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100 ms equal-power crossfades, and the chain is normalised as one signal.
joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.0 dB
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(doubt, sadness · normal-paced, almost no disfluency, narration, formal)his income must now support both his and Zahida's families and he's not sure he'll get work when he returns home.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt, sadness; style: narration, formal; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.8/10; 7.6s.
batch57_part2_batch57_part2_chunk_1513_1_1483938 · in -29.4 dBFS · gain +9.7 dB · snippets-01179
(measured, no disfluency, storytelling, casual)for week after week she made a four-hour bus journey to the courts.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, fairly guarded; no dominant emotion; style: storytelling, casual; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 3.8/10; 3.9s.
batch57_part2_batch57_part2_chunk_1513_1_1484150 · in -30.4 dBFS · gain +9.7 dB · snippets-01179