Proxy (PXR)

What this rule requires. the same two-sided test as Rule A, but the per-step smoothness cap is certified on a proxy axis for emotions that cannot ramp on their own — with a per-step cap of C = 0.25, and here only chains whose speaker similarity clears 0.80 on both conditions. No voice conversion has been applied.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken. Underlined descriptors are the ones that change across the chain; brackets inside the words are a real non-speech sound. The perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
Fear ↓  /  Astonishment Surprisetimbre ≥0.80   strict_015 · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Astonishment Surprise is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Astonishment Surprise barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.73.

At the same time Fear goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.25 (lower than 75 % of clips in this corpus), a change of -0.71. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.16, then +0.13, then +0.21 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 46 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.840 and the worst against the first clip 0.887; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/100/100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.4 dB
per-clip
per-clip buttons play:
rule PXRk 5qmax 0.713cmax 0.238d_a -0.713d_b 0.734dataset emolialang entotal 46.1schain gain -5.7 dBseam step 1.4 dBcrossfades 150/100/100/150 msmin_cos_consec 0.8398min_cos_anchor 0.8874
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fear, concentration · steady, almost no disfluency, formal, newsreading) This means that the bright hemisphere is visible from Earth when Iapetus is on the western side of Saturn, and that the dark hemisphere is visible when Iapetus is on the eastern side.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, concentration; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 9.2s, EN.
EN_c5BYOO0j3Fy_W000008 · in -14.9 dBFS · gain -5.7 dB · emolia-02588
(fairly steady, no disfluency, formal, monologue) The Dark Hemisphere was later named Cassini Regio in his honor
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.0/10; 3.5s, EN.
EN_c5BYOO0j3Fy_W000009 · in -13.4 dBFS · gain -5.7 dB · emolia-02588
(fairly steady, no disfluency, formal, monologue) Iopetus is named after the Titan Iopetus from Greek mythology
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.2/10; 3.6s, EN.
EN_c5BYOO0j3Fy_W000011 · in -14.4 dBFS · gain -5.7 dB · emolia-02588
(emotional numbness · fairly steady, no disfluency, newsreading, formal) The name was suggested by John Herschel – son of William Herschel, discoverer of Mimas and Enceladus – in his 1847 publication Results of Astronomical Observations Made at the Cape of Good Hope, in which he advocated naming the moons of Saturn after the Titans, brothers and sisters of the Titan Cronus – whom the Romans equated with their god Saturn.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 19.0s, EN.
EN_c5BYOO0j3Fy_W000012 · in -14.8 dBFS · gain -5.7 dB · emolia-02588
(fairly steady, no disfluency, authoritative, formal) When first discovered, Iapetus was among four Saturnian moons labeled the Sedera Lodoisia by their discoverer Giovanni Cassini after King Louis XIV. The other three were Tethys, Dione and Rhea
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.2s, EN.
EN_c5BYOO0j3Fy_W000013 · in -13.7 dBFS · gain -5.7 dB · emolia-02588
Thankfulness Gratitude ↓  /  Concentrationtimbre ≥0.80   strict_016 · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration barely there — 0.21, lower than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.72.

At the same time Thankfulness Gratitude goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.33 (lower than 67 % of clips in this corpus), a change of -0.67. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.12, then +0.19, then +0.20, then +0.21 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 34 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.882 and the worst against the first clip 0.883; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.5 dB
per-clip
per-clip buttons play:
rule PXRk 5qmax 0.671cmax 0.245d_a -0.671d_b 0.719dataset emolialang entotal 33.9schain gain -0.3 dBseam step 3.5 dBcrossfades 150/150/150/150 msmin_cos_consec 0.8820min_cos_anchor 0.8826
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, quiet background, light breath
(thankfulness gratitude, contentment, relief · normal-paced, energised, neutral tension, casual) Yeah, thank you. Thank you, Cindy. Thank you for sharing your perspective and I can sense and understand your frustration, (low mumble) uh, as well.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as thankfulness gratitude, contentment, relief; style: casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.9/10; 7.8s, EN.
EN_K0dRbVNBiF0_W000194 · in -20.8 dBFS · gain -0.3 dB · emolia-00786
(normal-paced, normally alert, slightly relaxed, casual) And so (low mumble) uhm, grandstanding, (ahem) some people call it performative allyship.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 2.4/10; 4.8s, EN.
EN_K0dRbVNBiF0_W000195 · in -17.3 dBFS · gain -0.3 dB · emolia-00786
(normal-paced, normally alert, neutral tension, casual) Uh, (low mumble) virtual signaling, I mean, we're starting to see all sorts of different, different ways to describe it. But, (low mumble) uhm,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 1.5/10; 6.7s, EN.
EN_K0dRbVNBiF0_W000196 · in -20.2 dBFS · gain -0.3 dB · emolia-00786
(normal-paced, normally alert, slightly relaxed, conversational) Here's, (ahem) uh, you know, here's, here's what we have to understand when it, when it comes to grandstanding, when it comes to any of these perspectives, to be honest with you.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: conversational, casual; good recording, quiet background; genuineness 4.5/6; vocal-burst blend 6.2/10; 7.7s, EN.
EN_K0dRbVNBiF0_W000197 · in -20.6 dBFS · gain -0.3 dB · emolia-00786
(concentration, interest · brisk, energised, neutral tension, authoritative) (low mumble) Uh, but I actually know that's, let's focus on grants, let's focus on the ability as to grandstanding. Cindy, you made a point that I think is really important.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, interest; style: authoritative, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.2/10; 7.5s, EN.
EN_K0dRbVNBiF0_W000198 · in -19.9 dBFS · gain -0.3 dB · emolia-00786
Thankfulness Gratitude ↓  /  Emotional Numbnessidentity ≥0.80   strict_010 · #3

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness below average — 0.32, lower than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.66.

At the same time Thankfulness Gratitude goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.27 (lower than 73 % of clips in this corpus), a change of -0.66. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.20, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.80 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.80. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 33 s · en · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.802 and the worst against the first clip 0.802; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.9 dB
per-clip
per-clip buttons play:
rule PXRk 4qmax 0.661cmax 0.248d_a -0.661d_b 0.663dataset podcastlang entotal 33.1schain gain -0.4 dBseam step 0.9 dBcrossfades 150/100/150 msmin_cos_consec 0.8024min_cos_anchor 0.8024
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, fairly steady, average clarity
(thankfulness gratitude, relief · normal-paced, neutral tension, some disfluency, conversational) they can get that help from, you know, CleanWright, or they can get that help from their distributor. Yeah. They can get that help from people in other businesses who (low mumble) uh have done similar things with restaurants, maybe. And I think that it
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as thankfulness gratitude, relief; style: conversational, casual; good recording, quiet background; genuineness 5.0/6; vocal-burst blend 9.9/10; 14.3s, EN.
65546_00066936 · in -19.8 dBFS · gain -0.4 dB · podcast-01153
(normal-paced, slightly relaxed, some disfluency, casual) the information is out there. You do have to do your homework,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 3.1/6; vocal-burst blend 2.2/10; 4.2s, EN.
65546_00068360 · in -18.9 dBFS · gain -0.4 dB · podcast-01118
(anger, contempt, distress · slow, slightly relaxed, frequent disfluency, casual) (ahem) but the car wash industry has so much runway in the self serving in bay because of these car washes that are not meeting their potential.
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as anger, contempt, distress; style: casual, conversational; good recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.9/10; 10.1s, EN.
65546_00068780 · in -19.7 dBFS · gain -0.4 dB · podcast-01120
(emotional numbness, helplessness · measured, slightly relaxed, some disfluency, casual) And the ROI is there. The operators just need the confidence
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as emotional numbness, helplessness; style: casual, conversational; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 1.5/10; 4.9s, EN.
65546_00069856 · in -19.5 dBFS · gain -0.4 dB · podcast-01129
Emotional Numbness ↓  /  Relieftimbre ≥0.80   strict_019 · #4

This chain comes from the proxy rule: the same two-sided test as above, but because Relief is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Relief essentially absent — 0.01, lower than 99 % of clips in this corpus — and ends with it clearly present at 0.69, higher than 69 % of clips in this corpus. That is a total rise of 0.68.

At the same time Emotional Numbness goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.23 (lower than 77 % of clips in this corpus), a change of -0.65. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.15, then +0.21, then +0.07, then +0.24 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 19 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.921 and the worst against the first clip 0.924; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.9 dB
per-clip
per-clip buttons play:
rule PXRk 5qmax 0.649cmax 0.240d_a -0.649d_b 0.672dataset emolialang entotal 19.0schain gain -5.2 dBseam step 1.9 dBcrossfades 100/150/150/150 msmin_cos_consec 0.9205min_cos_anchor 0.9244
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(formal, authoritative) == Institutes of national importance == 91 institutes
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 4.5s, EN.
EN_oaAekQKkoe4_W000043 · in -14.3 dBFS · gain -5.2 dB · emolia-02579
(formal, casual) All India Institutes of Medical Sciences' 7 functioning, 13 upcoming
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, casual; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.6/10; 4.8s, EN.
EN_oaAekQKkoe4_W000044 · in -14.2 dBFS · gain -5.2 dB · emolia-02579
(formal, authoritative) Indian Institutes of Technology – 23 functioning
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.7/10; 3.2s, EN.
EN_oaAekQKkoe4_W000045 · in -16.1 dBFS · gain -5.2 dB · emolia-02579
(formal, authoritative) National Institutes of Technology 31 functioning
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.5/10; 3.2s, EN.
EN_oaAekQKkoe4_W000046 · in -15.4 dBFS · gain -5.2 dB · emolia-02579
(formal, authoritative) Indian Institutes of Information Technology – 23 functioning
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.0/10; 3.9s, EN.
EN_oaAekQKkoe4_W000047 · in -15.7 dBFS · gain -5.2 dB · emolia-02579
Relief ↓  /  Emotional Numbnesstimbre ≥0.80   strict_011 · #5

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness below average — 0.32, lower than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.66.

At the same time Relief goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.34 (lower than 66 % of clips in this corpus), a change of -0.62. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.20, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 38 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.871 and the worst against the first clip 0.871; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/150/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.4 dB
per-clip
per-clip buttons play:
rule PXRk 4qmax 0.615cmax 0.246d_a -0.615d_b 0.659dataset emolialang entotal 38.5schain gain +2.6 dBseam step 2.4 dBcrossfades 100/150/100 msmin_cos_consec 0.8714min_cos_anchor 0.8714
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, balanced body, quiet background
(relief · slow, very low-energy, relaxed, casual) I think I translated that correctly. But anyway, nonviolence has got to be accepted by the council and they've got to come out against
full caption & clip details
A young adult masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as relief; style: casual, whispered; below-average recording, quiet background; genuineness 4.6/6; vocal-burst blend 1.3/10; 11.6s, EN.
EN_hDpxTjl-aU0_W000223 · in -23.1 dBFS · gain +2.6 dB · emolia-02412
(measured, normally alert, slightly relaxed, monologue) You might say that this was (low mumble) another asheq, another failure, but in fact in 1983
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 0.2/10; 6.2s, EN.
EN_hDpxTjl-aU0_W000225 · in -22.5 dBFS · gain +2.6 dB · emolia-02412
(bitterness, thankfulness gratitude, concentration · measured, subdued, slightly relaxed, monologue) The American Catholic bishops joined by French Catholic bishops and the German Catholic bishops all issued a statement called, God's Challenge and Our Response saying that nuclear weapons could not be
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as bitterness, thankfulness gratitude, concentration; style: monologue, didactic; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.5/10; 13.9s, EN.
EN_hDpxTjl-aU0_W000226 · in -21.7 dBFS · gain +2.6 dB · emolia-02412
(emotional numbness, malevolence malice, thankfulness gratitude · normal-paced, normally alert, slightly relaxed, monologue) Justified under just worth teaching as it has come down to us from Ambrose Augustine and all the rest of them.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, malevolence malice, thankfulness gratitude; style: monologue, casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.9/10; 7.2s, EN.
EN_hDpxTjl-aU0_W000227 · in -24.1 dBFS · gain +2.6 dB · emolia-02412
Disgust ↓  /  Paintimbre ≥0.80   strict_012 · #6

This chain comes from the proxy rule: the same two-sided test as above, but because Pain is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Pain below average — 0.35, lower than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.61.

At the same time Disgust goes the other way, from 0.76 (higher than 76 % of clips in this corpus) to 0.14 (lower than 86 % of clips in this corpus), a change of -0.62. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.17, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 21 s · zh · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.904 and the worst against the first clip 0.865; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.1 dB
per-clip
per-clip buttons play:
rule PXRk 4qmax 0.610cmax 0.241d_a -0.622d_b 0.610dataset emolialang zhtotal 21.1schain gain -2.8 dBseam step 2.1 dBcrossfades 100/100/150 msmin_cos_consec 0.9041min_cos_anchor 0.8650
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, no disfluency, clear, authoritative) 如果一些美容院自身的服务和环境太差,基本的环境卫生都不达标。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 2.8/10; 5.3s, ZH.
ZH_B00016_S07612_W000009 · in -17.6 dBFS · gain -2.8 dB · emolia-03434
(thankfulness gratitude · fast, some disfluency, clear, authoritative) 那么,你觉得消费者今后还会到你的美容院中来吗?这样一来,口碑丢失,想重新建立都非常难的。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: authoritative, monologue; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 4.8/10; 7.4s, ZH.
ZH_B00016_S07612_W000010 · in -18.4 dBFS · gain -2.8 dB · emolia-03434
(normal-paced, some disfluency, average clarity, authoritative) 第二,团购促销活动一定要和实际相符。然而,实际上。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 2.6/10; 4.6s, ZH.
ZH_B00016_S07612_W000011 · in -16.3 dBFS · gain -2.8 dB · emolia-03434
(pain · normal-paced, no disfluency, average clarity, authoritative) 有一些美容院经营管理者,希望达到吸引消费者的目的。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: authoritative, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 2.7/10; 4.2s, ZH.
ZH_B00016_S07612_W000012 · in -16.5 dBFS · gain -2.8 dB · emolia-03434
Disgust ↓  /  Contemplationtimbre ≥0.80   strict_014 · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contemplation below average — 0.28, lower than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.64.

At the same time Disgust goes the other way, from 0.74 (higher than 74 % of clips in this corpus) to 0.14 (lower than 86 % of clips in this corpus), a change of -0.60. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.23, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 29 s · ko · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.931 and the worst against the first clip 0.935; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.3 dB
per-clip
per-clip buttons play:
rule PXRk 4qmax 0.600cmax 0.231d_a -0.600d_b 0.642dataset emolialang kototal 28.8schain gain +1.8 dBseam step 3.3 dBcrossfades 150/150/150 msmin_cos_consec 0.9311min_cos_anchor 0.9347
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range
(brisk, little disfluency, clear, authoritative) 또한, hdmi-264라든지, 그런 비디오 코덱 또한 gpu를 통해서 바로 디코딩이 될 수가 있습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, dramatic; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 2.5/10; 5.1s, KO.
KO_W86C06i3VxY_W000027 · in -20.4 dBFS · gain +1.8 dB · emolia-03120
(normal-paced, some disfluency, average clarity, conversational) 하지만 이쪽까지는 이번 세션의 주요한 목적은 아니고요. 이쪽에 관해서 궁금하신 분들은 앞에서 설명드렸던
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.3/10; 7.7s, KO.
KO_W86C06i3VxY_W000028 · in -20.0 dBFS · gain +1.8 dB · emolia-03120
(brisk, some disfluency, average clarity, authoritative) 그 볼, 그 앞에서 적혀있던 그런 링크들을 통해서 여러 가지 메터리얼이 있으니까 그걸 보시면 될 것 같습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, monologue; good recording, no background noise; genuineness 2.4/6; vocal-burst blend 3.0/10; 5.8s, KO.
KO_W86C06i3VxY_W000029 · in -23.3 dBFS · gain +1.8 dB · emolia-03120
(contemplation · normal-paced, some disfluency, average clarity, conversational) 이렇게 엑셀트 컴포지팅이 다 좋은데 이게 여전히 문제가 있어요. 왜냐면 여전히 느리다는 거죠. 특히 기술이 발전하고 컨텐츠들이 좋아지면 좋아질수록
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contemplation; style: conversational, didactic; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 6.7/10; 10.6s, KO.
KO_W86C06i3VxY_W000030 · in -24.1 dBFS · gain +1.8 dB · emolia-03120
Infatuation ↓  /  Astonishment Surpriseidentity ≥0.80   strict_017 · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Astonishment Surprise is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Astonishment Surprise below average — 0.36, lower than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.61.

At the same time Infatuation goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.45 (lower than 55 % of clips in this corpus), a change of -0.54. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.16, then +0.13, then +0.18, then +0.13 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 40 s · en · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.830 and the worst against the first clip 0.833; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150/150/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 4.2 dB
per-clip
per-clip buttons play:
rule PXRk 5qmax 0.542cmax 0.219d_a -0.542d_b 0.607dataset podcastlang entotal 40.1schain gain +17.3 dBseam step 4.2 dBcrossfades 150/150/150/100 msmin_cos_consec 0.8304min_cos_anchor 0.8328
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, balanced body, normal-paced, normally alert, slightly relaxed
(infatuation, sexual lust, fear · steady, some disfluency, average clarity, casual) (contented sigh) She's also Imperial. It's this girl what's her name? Because she recently died, which is really upsetting because she's quite pretty still. But
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as infatuation, sexual lust, fear; style: casual, monologue; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 1.6/10; 6.7s, EN.
933965_00442952 · in -38.4 dBFS · gain +17.3 dB · podcast-05830
(infatuation, longing, affection · steady, frequent disfluency, somewhat unclear, casual) She was a Danish woman who like moved to Paris and met Jean-Luc Dardin. She was uh (low mumble) yeah, she was he kind of like the big star in all of his movies
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, neutral openness; reads as infatuation, longing, affection; style: casual, whispered; average recording, quiet background; mildly explicit content; genuineness 4.3/6; vocal-burst blend 2.7/10; 9.8s, EN.
933965_00443824 · in -39.8 dBFS · gain +17.3 dB · podcast-05849
(pain, embarrassment, pleasure ecstasy · fairly steady, some disfluency, average clarity, casual) of a sudden, but I was gonna say one of the biggest things I felt like most of the people in this movie just felt
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as pain, embarrassment, pleasure ecstasy; style: casual, conversational; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 2.9/10; 4.0s, EN.
933965_00446000 · in -35.6 dBFS · gain +17.3 dB · podcast-05818
(emotional numbness, interest, longing · fairly steady, some disfluency, average clarity, casual) younger than most of his films. Most of his films were like adults, and this one you really felt like it was just you were watching a bunch of kids, like just
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as emotional numbness, interest, longing; style: casual, monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 4.1/10; 7.2s, EN.
933965_00446400 · in -38.9 dBFS · gain +17.3 dB · podcast-05839
(astonishment surprise, amusement · fairly steady, some disfluency, average clarity, casual) Even though, well, I guess we're mid to late twenties now, but early watching early twenties people, I was like, wow, this is very youthful. like cast this cast very like you said was very clearly meant for the the
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as astonishment surprise, amusement; style: casual; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 3.6/10; 12.9s, EN.
933965_00447520 · in -35.6 dBFS · gain +17.3 dB · podcast-05817
Shame ↓  /  Contemplationidentity ≥0.80   strict_013 · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contemplation below average — 0.33, lower than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.60.

At the same time Shame goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.38 (lower than 62 % of clips in this corpus), a change of -0.52. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.12, then +0.24 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 18 s · de · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.868 and the worst against the first clip 0.892; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/150/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.8 dB
per-clip
per-clip buttons play:
rule PXRk 4qmax 0.524cmax 0.244d_a -0.524d_b 0.597dataset podcastlang detotal 18.2schain gain +0.5 dBseam step 1.8 dBcrossfades 100/150/100 msmin_cos_consec 0.8680min_cos_anchor 0.8920
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normal-paced, normally alert, slightly relaxed
(shame · formal, didactic) Zweitens, die Sorge vor einer Reduzierung sozialer Interaktionen. (ahem)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame; style: formal, didactic; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.0/10; 4.1s, DE.
316345_00062455 · in -18.6 dBFS · gain +0.5 dB · podcast-02186
(jealousy and envy, fatigue exhaustion · formal, didactic) Lernen Kinder noch ausreichend dem Austausch mit Mitschülern und Lehrkräften, wenn sie viel Zeit mit einer Lern-App verbringen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, fatigue exhaustion; style: formal, didactic; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.2/10; 7.0s, DE.
316345_00062860 · in -20.4 dBFS · gain +0.5 dB · podcast-02174
(conversational, formal) Das könnte soziale Kompetenzen beeinflussen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: conversational, formal; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.0/10; 3.3s, DE.
316345_00063560 · in -21.6 dBFS · gain +0.5 dB · podcast-02178
(contemplation · formal, didactic) Drittens die Frage nach der Abhängigkeit und dem kritischen Denken.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation; style: formal, didactic; very good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.8/10; 4.2s, DE.
316345_00063888 · in -22.2 dBFS · gain +0.5 dB · podcast-02180
Astonishment Surprise ↓  /  Sexual Lustidentity ≥0.80   strict_018 · #10

This chain comes from the proxy rule: the same two-sided test as above, but because Sexual Lust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Sexual Lust around average — 0.46, lower than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.51.

At the same time Astonishment Surprise goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.60. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.17, then +0.20, then -0.02 — not a clean run: step 4 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 44 s · en · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.827 and the worst against the first clip 0.814; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.5 dB
per-clip
per-clip buttons play:
rule PXRk 5qmax 0.514cmax 0.235d_a -0.599d_b 0.514dataset podcastlang entotal 43.5schain gain +3.7 dBseam step 3.5 dBcrossfades 100/150/150/150 msmin_cos_consec 0.8266min_cos_anchor 0.8142
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice
(astonishment surprise, confusion · slow, normally alert, relaxed, casual) Because Jesus didn't say if you fast, he says. He said, When?
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as astonishment surprise, confusion; style: casual, conversational; average recording, no background noise; genuineness 3.6/6; vocal-burst blend 0.9/10; 5.2s, EN.
451423_00093264 · in -23.9 dBFS · gain +3.7 dB · podcast-05491
(relief, awe, contemplation · measured, very low-energy, slightly relaxed, whispered) right? In Matthew chapter six, verse 16. This is what he said. When you fast, fasting was expected as a normal part of spiritual life. And in Matthew chapter 17 and verse 21, Jesus said, This kind does not go except by prayer
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as relief, awe, contemplation; style: whispered, casual; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 2.4/10; 22.1s, EN.
451423_00093896 · in -24.9 dBFS · gain +3.7 dB · podcast-05495
(emotional numbness, fatigue exhaustion · measured, normally alert, slightly relaxed, casual) Even one day a week or just one meal a day
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, dark, rough, thin; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as emotional numbness, fatigue exhaustion; style: casual, conversational; average recording, no background noise; genuineness 3.1/6; vocal-burst blend 3.5/10; 4.0s, EN.
451423_00097504 · in -21.4 dBFS · gain +3.7 dB · podcast-05510
(sexual lust, jealousy and envy, malevolence malice · slow, normally alert, slightly relaxed, casual) Fasting is you telling your body you're not in charge.
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is slightly warm, neutral-bright, rough, very full; average clarity, no disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as sexual lust, jealousy and envy, malevolence malice; style: casual, storytelling; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 3.1/10; 4.0s, EN.
451423_00098424 · in -22.9 dBFS · gain +3.7 dB · podcast-05510
(sexual lust, pain, fear · measured, subdued, slightly relaxed, monologue) It means you are temporarily denying your cravings, distractions, and comforts so that your spirit can
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is warm, slightly dark, slightly rough, very full; clear, almost no disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as sexual lust, pain, fear; style: monologue, whispered; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.6/10; 8.8s, EN.
451423_00098904 · in -22.9 dBFS · gain +3.7 dB · podcast-05501
Longing ↓  /  Impatience and Irritabilitytimbre ≥0.80   strict_005 · #11

This chain comes from the proxy rule: the same two-sided test as above, but because Impatience and Irritability is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Impatience and Irritability around average — 0.45, lower than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.48.

At the same time Longing goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.48 (lower than 52 % of clips in this corpus), a change of -0.50. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 20 s · zh · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.848 and the worst against the first clip 0.848; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.6 dB
per-clip
per-clip buttons play:
rule PXRk 3qmax 0.478cmax 0.249d_a -0.497d_b 0.478dataset emolialang zhtotal 20.3schain gain +0.3 dBseam step 0.6 dBcrossfades 150/100 msmin_cos_consec 0.8483min_cos_anchor 0.8483
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an elderly feminine voice · neutral-toned, slightly dark, fairly smooth, no background noise, measured, slightly relaxed, fairly steady, clear
(longing, infatuation, sexual lust · very low-energy, some disfluency, fairly narrow pitch, whispered) 周炳义为了郝冬梅拒绝了去当名副政委的秘书,为此都认为郝冬梅是个美女。所以好奇来看。
full caption & clip details
An elderly feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; clear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as longing, infatuation, sexual lust; style: whispered, monologue; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 1.7/10; 9.1s, ZH.
ZH_B00063_S02912_W000009 · in -20.2 dBFS · gain +0.3 dB · emolia-03904
(subdued, little disfluency, moderate pitch range, ASMR) 陶俊书特意叫了好多梅,过来,让人看几个女人。
full caption & clip details
A young adult feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; no dominant emotion; style: ASMR, whispered; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 2.9/10; 4.6s, ZH.
ZH_B00063_S02912_W000010 · in -20.1 dBFS · gain +0.3 dB · emolia-03904
(impatience and irritability · normally alert, no disfluency, moderate pitch range, whispered) 这次倒是直接表现的郝冬梅很一般的样子,甚至觉得周秉义重感情。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as impatience and irritability; style: whispered, ASMR; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 1.1/10; 6.8s, ZH.
ZH_B00063_S02912_W000011 · in -20.7 dBFS · gain +0.3 dB · emolia-03904
Contentment ↓  /  Fearidentity ≥0.80   strict_006 · #12

This chain comes from the proxy rule: the same two-sided test as above, but because Fear is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fear below average — 0.32, lower than 68 % of clips in this corpus — and ends with it strongly present at 0.78, higher than 78 % of clips in this corpus. That is a total rise of 0.46.

At the same time Contentment goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.47. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 28 s · en · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.854 and the worst against the first clip 0.854; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.7 dB
per-clip
per-clip buttons play:
rule PXRk 3qmax 0.463cmax 0.248d_a -0.466d_b 0.463dataset podcastlang entotal 27.5schain gain +8.6 dBseam step 1.7 dBcrossfades 150/150 msmin_cos_consec 0.8545min_cos_anchor 0.8545
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · slightly cool, fairly smooth, slightly thin, normal-paced, very low-energy
(contentment, affection, infatuation · neutral tension, moderately variable, some disfluency, casual) So it's like he's kinda like brushing it off, you know, and like his mom is asking were you in the same class together? Like was he a nice kid and he's like yeah we'd had math together he was alright and the mom is like just alright
full caption & clip details
A child feminine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as contentment, affection, infatuation; style: casual, playful; below-average recording, quiet background; genuineness 4.7/6; vocal-burst blend 5.5/10; 14.0s, EN.
508754_00347616 · in -28.2 dBFS · gain +8.6 dB · podcast-01146
(shame, relief, infatuation · neutral tension, moderately variable, frequent disfluency, casual) and he says like I don't want to be mean or anything but he was kind of full of himself, you know like he thought he was really cool.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as shame, relief, infatuation; style: casual, whispered; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 1.8/10; 8.5s, EN.
508754_00349016 · in -29.9 dBFS · gain +8.6 dB · podcast-04555
(slightly relaxed, fairly steady, almost no disfluency, whispered) And his father asks was he and he says to some people yeah girls were into
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: whispered, storytelling; average recording, no background noise; genuineness 2.2/6; vocal-burst blend 0.3/10; 5.4s, EN.
508754_00349860 · in -28.4 dBFS · gain +8.6 dB · podcast-04556
Fatigue Exhaustion ↓  /  Jealousy and Envyidentity ≥0.80   strict_007 · #13

This chain comes from the proxy rule: the same two-sided test as above, but because Jealousy and Envy is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Jealousy and Envy around average — 0.53, higher than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.46.

At the same time Fatigue Exhaustion goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.47 (lower than 53 % of clips in this corpus), a change of -0.48. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 40 s · mt · eurospeech

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.928 and the worst against the first clip 0.909; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.0 dB
per-clip
per-clip buttons play:
rule PXRk 3qmax 0.463cmax 0.245d_a -0.484d_b 0.463dataset eurospeechlang mttotal 39.8schain gain +2.5 dBseam step 3.0 dBcrossfades 150/150 msmin_cos_consec 0.9284min_cos_anchor 0.9093
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · balanced body, average recording, quiet background, normally alert
(fatigue exhaustion, impatience and irritability, longing · measured, neutral tension, moderately variable, monologue) niġu hawnhekk u ninjoraw totalment ir-realtà ta’ pajjiżna. Jew li niġu hawnhekk u rridu nibqgħu nagħfsu lil dawn in-nies
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is slightly cool, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as fatigue exhaustion, impatience and irritability, longing; style: monologue, didactic; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.0/10; 11.1s, MT.
malta_13_522_22112021_6594480_6605584 · in -21.0 dBFS · gain +2.5 dB · eurospeech-02148
(contempt, impatience and irritability, bitterness · measured, neutral tension, moderately variable, monologue) tad-dar tagħha għallużu personali u taħt regoli stretti. Aħjar ikollna soċjetà regolata milli jkollna l-ġungla li għandna bħalissa.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt, impatience and irritability, bitterness; style: monologue, didactic; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 3.7/10; 11.2s, MT.
malta_13_522_22112021_6620080_6631280 · in -24.1 dBFS · gain +2.5 dB · eurospeech-02148
(jealousy and envy, disappointment, pride · normal-paced, slightly relaxed, fairly steady, monologue) Dan hu l-ħsieb wara dan l-Abbozz ta’ Liġi, ċjoè li nirregolaw l-affarijiet. Ejja ma nħallux aktar lin-nies iħabbtu l-bieb tal-pusher li mhux kannabis biss se jkollu xi jbegħlek u jekk illum ma jkollux kannabis, se jipprova jbegħlek xi ħaġa oħra.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, disappointment, pride; style: monologue, didactic; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 2.8/10; 17.8s, MT.
malta_13_522_22112021_6631280_6649088 · in -22.7 dBFS · gain +2.5 dB · eurospeech-02148
Astonishment Surprise ↓  /  Infatuationidentity ≥0.80   strict_008 · #14

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation below average — 0.36, lower than 64 % of clips in this corpus — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.44.

At the same time Astonishment Surprise goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.44. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 13 s · snippets

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.842 and the worst against the first clip 0.842; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.9 dB
per-clip
per-clip buttons play:
rule PXRk 3qmax 0.441cmax 0.250d_a -0.442d_b 0.441dataset snippetslang undtotal 12.8schain gain -4.2 dBseam step 1.9 dBcrossfades 100/100 msmin_cos_consec 0.8419min_cos_anchor 0.8419
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(astonishment surprise · formal, narration) He realized that the coin had an image of a Greek or Roman face.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as astonishment surprise; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 3.2/10; 3.8s.
batch74_part1_batch74_part1_chunk_1666_1_1668411 · in -17.3 dBFS · gain -4.2 dB · snippets-01270
(longing · narration, formal) This is what happened to skipper and diver Trevor Small when he went sailing in Poole Bay.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing; style: narration, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.7/10; 5.2s.
batch74_part1_batch74_part1_chunk_1666_1_1668434 · in -16.2 dBFS · gain -4.2 dB · snippets-01270
(monologue, formal) They were frequent divers in the Monterey Bay National Marine Sanctuary.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.7/10; 4.0s.
batch74_part1_batch74_part1_chunk_1666_1_1668533 · in -14.3 dBFS · gain -4.2 dB · snippets-01270
Anger ↓  /  Contemplationidentity ≥0.80   strict_009 · #15

This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contemplation around average — 0.53, higher than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.42.

At the same time Anger goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.48 (lower than 52 % of clips in this corpus), a change of -0.48. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 36 s · dutch · mls

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.955 and the worst against the first clip 0.962; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.9 dB
per-clip
per-clip buttons play:
rule PXRk 3qmax 0.420cmax 0.240d_a -0.477d_b 0.420dataset mlslang nltotal 36.2schain gain +4.2 dBseam step 1.9 dBcrossfades 150/150 msmin_cos_consec 0.9548min_cos_anchor 0.9625
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, slightly relaxed, fairly steady
(anger, sexual lust, impatience and irritability · normal-paced, no disfluency, narration, monologue) ook meende ik ditmaal op den weg te zijn eener betere fortuin en geloofde reeds dat mijn goede engel over mijn kwaden geest getriumfeerd had toen ik vernam dat de kapitein in quaestie in spaansch braband gesneuveld was
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, sexual lust, impatience and irritability; style: narration, monologue; good recording, quiet background; genuineness 0.9/6; vocal-burst blend 0.0/10; 12.2s, DUTCH.
1724_2757_002343 · in -24.2 dBFS · gain +4.2 dB · mls-00104
(fear, disappointment, helplessness · measured, some disfluency, whispered, narration) ik had grond tot de hoop dat ik in zijne plaats zou worden aangesteld maar ik leerde t wel anders daar ik op eens dat is nu omstreeks twee maanden geleden zonder eenige fout begaan te hebben willekeurig uit mijn rang werd ontzet
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, disappointment, helplessness; style: whispered, narration; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.0/10; 13.6s, DUTCH.
1724_2757_000356 · in -23.5 dBFS · gain +4.2 dB · mls-00104
(contemplation, concentration, longing · measured, frequent disfluency, didactic, whispered) ik mag zeggen onverdiend en zonder dat ik tot hiertoe de oorzaak van deze cassatie heb kunnen raden die werd u dan niet aangezegd na gewezen vonnis
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, concentration, longing; style: didactic, whispered; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.0/10; 10.8s, DUTCH.
1724_2757_002431 · in -25.5 dBFS · gain +4.2 dB · mls-00104
Impatience and Irritability ↓  /  Intoxication Altered States of Consciousnesstimbre ≥0.80   strict_000 · #16

This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Intoxication Altered States of Consciousness clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.25.

At the same time Impatience and Irritability goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.74 (higher than 74 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 15 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.897 and the worst against the first clip 0.897; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.1 dB
per-clip
per-clip buttons play:
rule PXRk 2qmax 0.250cmax 0.250d_a -0.250d_b 0.250dataset emolialang entotal 14.6schain gain +3.7 dBseam step 1.1 dBcrossfades 150 msmin_cos_consec 0.8970min_cos_anchor 0.8970
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · balanced body, quiet background, neutral tension
(impatience and irritability, disgust, anger · measured, energised, moderately variable, didactic) And let's just make this more concrete for you. These people, these people, the four people who chose this, they made a mistake.
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is slightly cool, slightly dark, slightly rough, balanced body; very clear, frequent disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as impatience and irritability, disgust, anger; style: didactic, cartoonish; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.0/10; 9.5s, EN.
EN_B00033_S03995_W000113 · in -23.3 dBFS · gain +3.7 dB · emolia-00887
(normal-paced, normally alert, fairly steady, conversational) (low mumble) We're never going to get the make in, let's try and get the make in there. Come forward as far as you can and then really shout, yep.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: conversational, playful; good recording, quiet background; genuineness 4.3/6; vocal-burst blend 2.0/10; 5.2s, EN.
EN_B00033_S03995_W000114 · in -24.5 dBFS · gain +3.7 dB · emolia-00887
Relief ↓  /  Angertimbre ≥0.80   strict_004 · #17

This chain comes from the proxy rule: the same two-sided test as above, but because Anger is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Anger clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.25.

At the same time Relief goes the other way, from 0.77 (higher than 77 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 12 s · zh · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.856 and the worst against the first clip 0.856; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.3 dB
per-clip
per-clip buttons play:
rule PXRk 2qmax 0.250cmax 0.250d_a -0.250d_b 0.250dataset emolialang zhtotal 11.6schain gain -1.7 dBseam step 0.3 dBcrossfades 150 msmin_cos_consec 0.8562min_cos_anchor 0.8562
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(no disfluency, formal, authoritative) 原来几天前,继父打来电话,说村里准备重新画地。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.3/10; 3.9s, ZH.
ZH_B00063_S06302_W000030 · in -18.2 dBFS · gain -1.7 dB · emolia-03904
(almost no disfluency, monologue, authoritative) 离开家的十几年里,二毛家的宅基地被邻居侵占。继父说,再不回来,你的地就全没了,赶紧回来吧。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 3.1/10; 7.8s, ZH.
ZH_B00063_S06302_W000031 · in -18.5 dBFS · gain -1.7 dB · emolia-03904
Elation ↓  /  Disgustidentity ≥0.80   strict_001 · #18

This chain comes from the proxy rule: the same two-sided test as above, but because Disgust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Disgust clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.25.

At the same time Elation goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 34 s · en · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.937 and the worst against the first clip 0.937; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.5 dB
per-clip
per-clip buttons play:
rule PXRk 2qmax 0.250cmax 0.250d_a -0.250d_b 0.250dataset podcastlang entotal 34.2schain gain +6.5 dBseam step 0.5 dBcrossfades 150 msmin_cos_consec 0.9369min_cos_anchor 0.9369
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, slightly bright, fairly smooth, balanced body, no background noise, normally alert, neutral tension, moderately variable
(elation, contentment, amusement · brisk, casual, conversational) after three weeks, which obviously we did. Yeah. It was fine. But then when we came back, I think we had like maybe two weeks before a new crop of kids were coming because on the the kids only can really do like a six month contract. Like that's the longest that the kids do it, usually, unless they happen to be like very small
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as elation, contentment, amusement; style: casual, conversational; good recording, no background noise; genuineness 4.4/6; vocal-burst blend 6.6/10; 18.0s, EN.
148957_00291136 · in -26.2 dBFS · gain +6.5 dB · podcast-04408
(disgust, impatience and irritability, amusement · normal-paced, casual, conversational) like re are very petite, like 'cause when they get a little older, it's like hard. They have to be a certain height. They need to, you know, lots of restrictions for the kids. Plus you don't want kids like back to back on uh (contented sigh) you know, like doing a year. Like they need to only do (ahem) uh just to be kids. I think they
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as disgust, impatience and irritability, amusement; style: casual, conversational; average recording, no background noise; mildly explicit content; genuineness 5.4/6; vocal-burst blend 8.0/10; 16.4s, EN.
148957_00293000 · in -26.7 dBFS · gain +6.5 dB · podcast-01351
Thankfulness Gratitude ↓  /  Infatuationidentity ≥0.80   strict_002 · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.25.

At the same time Thankfulness Gratitude goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 26 s · no · eurospeech

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.928 and the worst against the first clip 0.928; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.1 dB
per-clip
per-clip buttons play:
rule PXRk 2qmax 0.249cmax 0.249d_a -0.249d_b 0.249dataset eurospeechlang nototal 26.1schain gain +5.4 dBseam step 0.1 dBcrossfades 150 msmin_cos_consec 0.9284min_cos_anchor 0.9284
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged somewhat feminine voice · neutral-toned, neutral-bright, fairly smooth, average recording, slightly relaxed, moderately variable, some disfluency, average clarity
(thankfulness gratitude, pride · measured, very low-energy, light breath, monologue) Å vera konkurransedyktig er viktig for arbeidsplassar og velferd, og ikkje minst at vi må vera eit inkluderande Norden. Noreg må vera ein pådrivar for å nå desse måla.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as thankfulness gratitude, pride; style: monologue, whispered; average recording, some background noise; genuineness 2.6/6; vocal-burst blend 0.2/10; 12.9s, NO.
norway_10303-1_1890880_1903744 · in -25.3 dBFS · gain +5.4 dB · eurospeech-02217
(normal-paced, normally alert, normal breath, monologue) Eg har stor tru på eit tett samarbeid i Norden, og at det framleis vil spela ei stor rolle framover. Denne stortingsmeldinga er eit godt grunnlag for å vidareutvikla dette samarbeidet.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, neutral stance, neutral openness; no dominant emotion; style: monologue, whispered; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.9/10; 13.4s, NO.
norway_10303-1_1903744_1917168 · in -25.5 dBFS · gain +5.4 dB · eurospeech-02217
Emotional Numbness ↓  /  Intoxication Altered States of Consciousnessidentity ≥0.80   strict_003 · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Intoxication Altered States of Consciousness clearly present — 0.58, higher than 58 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.25.

At the same time Emotional Numbness goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 15 s · snippets

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.832 and the worst against the first clip 0.832; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.4 dB
per-clip
per-clip buttons play:
rule PXRk 2qmax 0.248cmax 0.248d_a -0.248d_b 0.248dataset snippetslang undtotal 15.4schain gain +7.7 dBseam step 1.4 dBcrossfades 150 msmin_cos_consec 0.8317min_cos_anchor 0.8317
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, no disfluency, formal, monologue) conventional wisdom would say those are three knockouts.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 3.3/10; 3.8s.
batch77_part4_batch77_part4_chunk_1689_1_1990102 · in -26.7 dBFS · gain +7.7 dB · snippets-01287
(measured, almost no disfluency, monologue, newsreading) There was three principles that didn't come from us, but from a Swiss design engineering team. A small batch size that was financially viable, 100% recyclable.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, newsreading; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.9/10; 11.8s.
batch77_part4_batch77_part4_chunk_1689_1_1990129 · in -28.1 dBFS · gain +7.7 dB · snippets-01287