Rule A — two-sided

What this rule requires. both the start emotion and the target emotion must move by at least T — with a per-step cap of C = 0.25, and here only chains whose speaker similarity clears 0.80 on both conditions. No voice conversion has been applied.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken. Underlined descriptors are the ones that change across the chain; brackets inside the words are a real non-speech sound. The perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
Thankfulness Gratitude ↓  /  Relieftimbre ≥0.80   strict_035 · #1

This chain comes from the two-sided rule: it only counts if both emotions move — Thankfulness Gratitude down and Relief up — by at least 0.25 each.

The chain starts with Relief barely there — 0.21, lower than 79 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.65.

At the same time Thankfulness Gratitude goes the other way, from 0.78 (higher than 78 % of clips in this corpus) to 0.11 (lower than 89 % of clips in this corpus), a change of -0.66. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.09, then +0.12, then +0.20, then +0.23 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 35 s · zh · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.880 and the worst against the first clip 0.861; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/100/150/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.3 dB
per-clip
per-clip buttons play:
rule AB2k 5qmax 0.647cmax 0.235d_a -0.663d_b 0.647dataset emolialang zhtotal 35.5schain gain +0.4 dBseam step 1.3 dBcrossfades 100/100/150/100 msmin_cos_consec 0.8800min_cos_anchor 0.8612
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, normally alert, slightly relaxed, some disfluency, moderate pitch range
(fast, fairly steady, average clarity, casual) 我们是通过发现自己真正想要的东西,发现自己的目标而找到自我的。
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, authoritative; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.9/10; 5.7s, ZH.
ZH_B00004_S03103_W000007 · in -20.1 dBFS · gain +0.4 dB · emolia-03315
(measured, fairly steady, somewhat unclear, monologue) 那在有价值的人生道路上旅行的时候,我们如果能以他人为重,就能忘却自我。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 3.2/6; vocal-burst blend 3.0/10; 6.3s, ZH.
ZH_B00004_S03103_W000008 · in -21.1 dBFS · gain +0.4 dB · emolia-03315
(measured, fairly steady, somewhat unclear, monologue) 但是以他人为重呢,说的容易啊,实际上是非常难做到的。约翰呢讲了一个故事。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 2.6/10; 6.2s, ZH.
ZH_B00004_S03103_W000009 · in -19.9 dBFS · gain +0.4 dB · emolia-03315
(measured, steady, somewhat unclear, didactic) 他呢去印度跟他父亲和祖父在那边生活,他父亲祖父在那里做很多慈善项目。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.0/10; 7.1s, ZH.
ZH_B00004_S03103_W000010 · in -19.9 dBFS · gain +0.4 dB · emolia-03315
(normal-paced, fairly steady, somewhat unclear, monologue) 而且呢贫民窟啊和其他不幸地区的孩子们缺少基本的教育,他们所学的唯一语言就是当地方言,这会限制他们以后在生活中的机会。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 4.1/10; 10.6s, ZH.
ZH_B00004_S03103_W000011 · in -21.0 dBFS · gain +0.4 dB · emolia-03315
Astonishment Surprise ↓  /  Disgusttimbre ≥0.80   strict_036 · #2

This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Disgust up — by at least 0.25 each.

The chain starts with Disgust barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.76.

At the same time Astonishment Surprise goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.08 (lower than 92 % of clips in this corpus), a change of -0.64. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.20, then +0.23, then +0.14 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 47 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.867 and the worst against the first clip 0.842; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.6 dB
per-clip
per-clip buttons play:
rule AB2k 5qmax 0.640cmax 0.233d_a -0.640d_b 0.761dataset emolialang entotal 46.6schain gain -4.9 dBseam step 1.6 dBcrossfades 150/150/150/150 msmin_cos_consec 0.8668min_cos_anchor 0.8421
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fairly steady, light breath, formal, authoritative) There is evidence that at least two embassies were sent to the Roman Emperor Augustus by Pandya kings
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.6/10; 5.2s, EN.
EN_wby6d6Ua5Lu_W000046 · in -14.3 dBFS · gain -4.9 dB · emolia-01574
(fairly steady, light breath, formal, authoritative) Potsherds with Tamil writing have also been found in excavations on the Red Sea, suggesting the presence of Tamil merchants there
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.1/10; 7.0s, EN.
EN_wby6d6Ua5Lu_W000047 · in -15.5 dBFS · gain -4.9 dB · emolia-01574
(fairly steady, light breath, formal, newsreading) An anonymous 1st century traveller's account written in Greek, Periplus Maris Arithrae, describes the ports of the Pandya and Shara kingdoms in Damarica and their commercial activity in great detail
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; mildly explicit content; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.7s, EN.
EN_wby6d6Ua5Lu_W000048 · in -15.5 dBFS · gain -4.9 dB · emolia-01574
(emotional numbness · steady, minimal breath, newsreading, formal) Peri Plus also indicates that the chief exports of the ancient Tamils were pepper, malabathrum, pearls, ivory, silk, spikenard, diamonds, sapphires, and tortoise shell.The Classical period ended around the 4th century CE with invasions by the Calabra, referred to as the Calapyrur in Tamil literature and inscriptions.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 18.7s, EN.
EN_wby6d6Ua5Lu_W000049 · in -15.5 dBFS · gain -4.9 dB · emolia-01574
(disgust · fairly steady, light breath, formal, newsreading) These invaders are described as evil kings and barbarians coming from lands to the north of the Tamil country
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.2/10; 5.7s, EN.
EN_wby6d6Ua5Lu_W000050 · in -13.8 dBFS · gain -4.9 dB · emolia-01574
Doubt ↓  /  Disgusttimbre ≥0.80   strict_037 · #3

This chain comes from the two-sided rule: it only counts if both emotions move — Doubt down and Disgust up — by at least 0.25 each.

The chain starts with Disgust below average — 0.33, lower than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.64.

At the same time Doubt goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.22 (lower than 78 % of clips in this corpus), a change of -0.73. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.12, then +0.16, then +0.15 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 35 s · zh · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.846 and the worst against the first clip 0.824; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.8 dB
per-clip
per-clip buttons play:
rule AB2k 5qmax 0.637cmax 0.223d_a -0.728d_b 0.637dataset emolialang zhtotal 34.8schain gain -5.1 dBseam step 1.8 dBcrossfades 150/150/150/150 msmin_cos_consec 0.8464min_cos_anchor 0.8245
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, no background noise, normally alert, slightly relaxed
(doubt · measured, fairly steady, frequent disfluency, didactic) 最好是自己去查一查字典,他的示意有很多。比如说第一个。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: didactic, whispered; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 1.0/10; 7.4s, ZH.
ZH_B00038_S05687_W000007 · in -15.9 dBFS · gain -5.1 dB · emolia-03659
(normal-paced, moderately variable, some disfluency, storytelling) Young小朋友们说样子吗?是的样子,物体的形状。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: storytelling, dramatic; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 4.3/10; 5.8s, ZH.
ZH_B00038_S05687_W000008 · in -14.2 dBFS · gain -5.1 dB · emolia-03659
(measured, steady, frequent disfluency, didactic) 还有第二个人的神情模样人如果他说的。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; crisply articulate, frequent disfluency, narrow pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, ASMR; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.0/10; 7.3s, ZH.
ZH_B00038_S05687_W000009 · in -14.5 dBFS · gain -5.1 dB · emolia-03659
(jealousy and envy · measured, fairly steady, some disfluency, monologue) 想表达的想描述的是人的话,是人的神情模样、人的模样。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as jealousy and envy; style: monologue, didactic; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 1.7/10; 6.9s, ZH.
ZH_B00038_S05687_W000010 · in -14.8 dBFS · gain -5.1 dB · emolia-03659
(disgust, confusion · measured, moderately variable, frequent disfluency, whispered) 第三,用于做标准的东西用于做标准。他是榜样。他是。
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, slightly dark, fairly smooth, thin; clear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, neutral openness; reads as disgust, confusion; style: whispered, didactic; average recording, no background noise; genuineness 2.7/6; vocal-burst blend 1.7/10; 8.1s, ZH.
ZH_B00038_S05687_W000011 · in -15.5 dBFS · gain -5.1 dB · emolia-03659
Disgust ↓  /  Contentmenttimbre ≥0.80   strict_038 · #4

This chain comes from the two-sided rule: it only counts if both emotions move — Disgust down and Contentment up — by at least 0.25 each.

The chain starts with Contentment barely there — 0.24, lower than 76 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.63.

At the same time Disgust goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.14 (lower than 86 % of clips in this corpus), a change of -0.79. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.00, then +0.19, then +0.21 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 35 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.850 and the worst against the first clip 0.867; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/100/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 4.8 dB
per-clip
per-clip buttons play:
rule AB2k 5qmax 0.632cmax 0.236d_a -0.792d_b 0.632dataset emolialang entotal 34.8schain gain +0.2 dBseam step 4.8 dBcrossfades 100/100/150/150 msmin_cos_consec 0.8505min_cos_anchor 0.8671
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, normally alert, slightly relaxed, clear
(disgust, contempt · normal-paced, steady, little disfluency, formal) It's an adjective which means devoid of guile, innocent and without deception. For example,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, contempt; style: formal, whispered; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.6/10; 7.4s, EN.
EN_ttS0dbNDTfo_W000052 · in -19.0 dBFS · gain +0.2 dB · emolia-02615
(emotional numbness · measured, steady, no disfluency, formal) Every child is born as Skyless and Innocent Ping.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 2.1/10; 3.0s, EN.
EN_ttS0dbNDTfo_W000053 · in -20.3 dBFS · gain +0.2 dB · emolia-02615
(concentration · normal-paced, steady, some disfluency, whispered) Mnemonic for remembering the word guileless. So it can be broken up as guile, which is deception, and less, which means no. So, guileless means no deception or something that is innocent.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: whispered, monologue; good recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.4/10; 11.3s, EN.
EN_ttS0dbNDTfo_W000054 · in -22.0 dBFS · gain +0.2 dB · emolia-02615
(normal-paced, fairly steady, little disfluency, monologue) Some synonyms for the word guileless can be artless and genuous.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 2.0/10; 4.0s, EN.
EN_ttS0dbNDTfo_W000055 · in -17.2 dBFS · gain +0.2 dB · emolia-02615
(measured, steady, frequent disfluency, whispered) Antonyms could be scheming, crafty, devious. The next word is good, which is a verb meaning argon or
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: whispered, didactic; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.7/10; 9.7s, EN.
EN_ttS0dbNDTfo_W000056 · in -21.8 dBFS · gain +0.2 dB · emolia-02615
Embarrassment ↓  /  Emotional Numbnesstimbre ≥0.80   strict_039 · #5

This chain comes from the two-sided rule: it only counts if both emotions move — Embarrassment down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness essentially absent — 0.05, lower than 95 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.84.

At the same time Embarrassment goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.32 (lower than 68 % of clips in this corpus), a change of -0.62. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.25, then +0.22, then +0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 32 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.825 and the worst against the first clip 0.844; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/100/100/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.5 dB
per-clip
per-clip buttons play:
rule AB2k 5qmax 0.623cmax 0.246d_a -0.623d_b 0.840dataset emolialang entotal 32.0schain gain -2.8 dBseam step 2.5 dBcrossfades 150/100/100/100 msmin_cos_consec 0.8248min_cos_anchor 0.8440
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, light breath
(embarrassment, pain, relief · normal-paced, normally alert, neutral tension, conversational) (surprised gasp) No problem. I had (low mumble) uhm, chosen the topic of virtually connecting and the way in which I
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, pain, relief; style: conversational, casual; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 2.8/10; 6.8s, EN.
EN_TU4EGx5Tasg_W000126 · in -14.9 dBFS · gain -2.8 dB · emolia-01187
(normal-paced, normally alert, slightly relaxed, casual) To bring students, (ahem) uhm, to start talking amongst themselves basically in this.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.2/10; 5.0s, EN.
EN_TU4EGx5Tasg_W000128 · in -17.4 dBFS · gain -2.8 dB · emolia-01187
(contemplation, sadness, doubt · slow, very low-energy, relaxed, ASMR) I thought, (ahem) you know, one could balance this, that it could go either way and it depended on, (low mumble) um, you know, how much the students became involved, how, (ahem) um,
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, sadness, doubt; style: ASMR, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.3/10; 10.8s, EN.
EN_TU4EGx5Tasg_W000131 · in -18.6 dBFS · gain -2.8 dB · emolia-01187
(doubt · normal-paced, normally alert, slightly relaxed, monologue) In how much the instructor put it in print, so I use that as a kind of neutral.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as doubt; style: monologue, formal; average recording, no background noise; genuineness 2.4/6; vocal-burst blend 0.3/10; 5.9s, EN.
EN_TU4EGx5Tasg_W000132 · in -18.0 dBFS · gain -2.8 dB · emolia-01187
(normal-paced, normally alert, slightly relaxed, monologue) Response to that as far as teacher centric or learner centric.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 1.3/10; 3.9s, EN.
EN_TU4EGx5Tasg_W000133 · in -18.1 dBFS · gain -2.8 dB · emolia-01187
Thankfulness Gratitude ↓  /  Concentrationtimbre ≥0.80   strict_030 · #6

This chain comes from the two-sided rule: it only counts if both emotions move — Thankfulness Gratitude down and Concentration up — by at least 0.25 each.

The chain starts with Concentration below average — 0.26, lower than 74 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.59.

At the same time Thankfulness Gratitude goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.61. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.13, then +0.22 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 30 s · ko · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.911 and the worst against the first clip 0.940; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.2 dB
per-clip
per-clip buttons play:
rule AB2k 4qmax 0.595cmax 0.242d_a -0.611d_b 0.595dataset emolialang kototal 30.4schain gain -1.9 dBseam step 2.2 dBcrossfades 100/150/150 msmin_cos_consec 0.9107min_cos_anchor 0.9401
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, light breath
(thankfulness gratitude, confusion · normal-paced, fairly steady, some disfluency, authoritative) 왔다. 왔어. 진행이 왔답니다. 왔다 는 도올류 투산 1.6 터보 가수님 4wd 에 애프터 블록 기능에 대해서 알아보도록 하겠습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, confusion; style: authoritative, didactic; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 4.3/10; 8.4s, KO.
KO_LgVWkgLl-FA_W000000 · in -17.5 dBFS · gain -1.9 dB · emolia-03196
(measured, fairly steady, no disfluency, formal) after blow는 에어컨을 사용 후에 석기를 제거해주는 기능입니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 2.3/10; 4.8s, KO.
KO_LgVWkgLl-FA_W000001 · in -17.2 dBFS · gain -1.9 dB · emolia-03196
(measured, steady, some disfluency, authoritative) 차 안에의 곰팡이 냄새를 제게 해줘서 개적한 공기를 차 안에에 제공합니다.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, didactic; average recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.6/10; 6.8s, KO.
KO_LgVWkgLl-FA_W000002 · in -19.4 dBFS · gain -1.9 dB · emolia-03196
(measured, fairly steady, some disfluency, monologue) 그러나 과연, 그러한 기능이 들어가 있는 앱 더블로 기능이 정상적으로 작동하고 있는지 아닌지를 오늘 테스트해 보는 것이 이 영상의 목적입니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 2.8/10; 10.8s, KO.
KO_LgVWkgLl-FA_W000003 · in -18.6 dBFS · gain -1.9 dB · emolia-03196
Sexual Lust ↓  /  Concentrationtimbre ≥0.80   strict_031 · #7

This chain comes from the two-sided rule: it only counts if both emotions move — Sexual Lust down and Concentration up — by at least 0.25 each.

The chain starts with Concentration below average — 0.31, lower than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.61.

At the same time Sexual Lust goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.34 (lower than 66 % of clips in this corpus), a change of -0.59. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.21, then +0.15 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 32 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.926 and the worst against the first clip 0.881; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.8 dB
per-clip
per-clip buttons play:
rule AB2k 4qmax 0.590cmax 0.246d_a -0.590d_b 0.608dataset emolialang entotal 32.0schain gain +2.1 dBseam step 1.8 dBcrossfades 100/100/150 msmin_cos_consec 0.9256min_cos_anchor 0.8806
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · fairly smooth, slightly relaxed
(sexual lust, infatuation · slow, normally alert, fairly steady, casual) Click on open and then click on install now.
full caption & clip details
A young adult feminine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is slightly cool, dark, fairly smooth, thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, submissive, neutral openness; reads as sexual lust, infatuation; style: casual, whispered; average recording, no background noise; genuineness 2.9/6; vocal-burst blend 2.2/10; 4.2s, EN.
EN_B00081_S08593_W000006 · in -20.1 dBFS · gain +2.1 dB · emolia-01793
(interest, hope enthusiasm optimism · normal-paced, normally alert, fairly steady, casual) (ahem) Uhm, so once you've bought the plugin, you do need to install it into your blog and such. So I'm going to work through how we do all of that.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as interest, hope enthusiasm optimism; style: casual, monologue; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.0/10; 7.2s, EN.
EN_B00081_S08593_W000007 · in -21.9 dBFS · gain +2.1 dB · emolia-01793
(teasing · normal-paced, normally alert, fairly steady, casual) Once you've done your settings for Backup Buddy, scroll down again to the Backup Buddy section. (low mumble) Uhm, if you click on the backups menu,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as teasing; style: casual, playful; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 0.7/10; 8.9s, EN.
EN_B00081_S08593_W000008 · in -23.1 dBFS · gain +2.1 dB · emolia-01793
(concentration · slow, subdued, steady, whispered) Right here, (wistful sigh) uh, and (ahem) uhm, from there you can click on the add a new button. Now you could also click on the add new link in the menu under plugins here as well. (wistful sigh)
full caption & clip details
A young adult feminine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is slightly cool, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as concentration; style: whispered, ASMR; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 0.9/10; 12.0s, EN.
EN_B00081_S08593_W000009 · in -22.8 dBFS · gain +2.1 dB · emolia-01793
Astonishment Surprise ↓  /  Paintimbre ≥0.80   strict_033 · #8

This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Pain up — by at least 0.25 each.

The chain starts with Pain below average — 0.35, lower than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.61.

At the same time Astonishment Surprise goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.14 (lower than 86 % of clips in this corpus), a change of -0.59. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then +0.17, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 31 s · zh · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.834 and the worst against the first clip 0.834; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.3 dB
per-clip
per-clip buttons play:
rule AB2k 4qmax 0.588cmax 0.239d_a -0.588d_b 0.608dataset emolialang zhtotal 31.1schain gain -1.7 dBseam step 0.3 dBcrossfades 100/150/150 msmin_cos_consec 0.8341min_cos_anchor 0.8341
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, measured, normally alert
(formal, narration) 欧洲中心论似乎是欧美学者的不自觉的下意识。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 2.0/10; 4.8s, ZH.
ZH_B00081_S06905_W000007 · in -18.0 dBFS · gain -1.7 dB · emolia-04087
(confusion · monologue, didactic) 梁启超提出,中华民族概念极大增强了认同感和归属感,凝聚了一致对外的民族精神,激发了爱国情感,避免了国家四分五裂。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion; style: monologue, didactic; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.9/10; 14.1s, ZH.
ZH_B00081_S06905_W000008 · in -18.4 dBFS · gain -1.7 dB · emolia-04087
(dramatic, storytelling) 并逐渐赋予了政治、文化、经济等现代内涵。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: dramatic, storytelling; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 2.8/10; 5.1s, ZH.
ZH_B00081_S06905_W000009 · in -18.4 dBFS · gain -1.7 dB · emolia-04087
(pain · didactic, monologue) 一八九八年,严复出版天言论,阐述了民族生存竞争的群族理念论。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: didactic, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.7/10; 7.6s, ZH.
ZH_B00081_S06905_W000010 · in -18.4 dBFS · gain -1.7 dB · emolia-04087
Infatuation ↓  /  Thankfulness Gratitudetimbre ≥0.80   strict_034 · #9

This chain comes from the two-sided rule: it only counts if both emotions move — Infatuation down and Thankfulness Gratitude up — by at least 0.25 each.

The chain starts with Thankfulness Gratitude below average — 0.33, lower than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.59.

At the same time Infatuation goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.27 (lower than 73 % of clips in this corpus), a change of -0.61. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.12, then +0.23 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 39 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.876 and the worst against the first clip 0.917; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.8 dB
per-clip
per-clip buttons play:
rule AB2k 4qmax 0.583cmax 0.245d_a -0.583d_b 0.588dataset emolialang entotal 39.1schain gain -5.3 dBseam step 0.8 dBcrossfades 150/150/150 msmin_cos_consec 0.8765min_cos_anchor 0.9168
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(no disfluency, formal, authoritative) Saving faith is so-called because it has eternal life inseparably connected with it, and is a special operation of the Holy Spirit
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 6.9s, EN.
EN_ZHvQ24ZKUUU_W000096 · in -14.7 dBFS · gain -5.3 dB · emolia-02480
(awe, emotional numbness, relief · almost no disfluency, authoritative, formal) Paul writes in Ephesians 2 verses 8 to 9, "'For by grace you have been saved through faith, and this is not your own doing, it is the gift of God—not the result of works, so that no one may boast.'"
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, emotional numbness, relief; style: authoritative, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.4s, EN.
EN_ZHvQ24ZKUUU_W000098 · in -14.6 dBFS · gain -5.3 dB · emolia-02480
(doubt, concentration, contemplation · no disfluency, newsreading, authoritative) From this, some Protestants believe that faith itself is given as a gift of God, e.g. the Westminster Confession of Faith, although this interpretation is disputed by others who believe the Greek gender alignment indicates that the
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, concentration, contemplation; style: newsreading, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 14.8s, EN.
EN_ZHvQ24ZKUUU_W000099 · in -15.0 dBFS · gain -5.3 dB · emolia-02480
(thankfulness gratitude · no disfluency, formal, authoritative) According to Lutherans, saving faith is the knowledge of, acceptance of, and trust in the promise of the Gospel
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 5.5s, EN.
EN_ZHvQ24ZKUUU_W000101 · in -14.2 dBFS · gain -5.3 dB · emolia-02480
Jealousy and Envy ↓  /  Concentrationidentity ≥0.80   strict_032 · #10

This chain comes from the two-sided rule: it only counts if both emotions move — Jealousy and Envy down and Concentration up — by at least 0.25 each.

The chain starts with Concentration below average — 0.25, lower than 75 % of clips in this corpus — and ends with it strongly present at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.52.

At the same time Jealousy and Envy goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.28 (lower than 72 % of clips in this corpus), a change of -0.61. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.25, then +0.05 — a plateau around step 3, where it barely moves.

The largest step is 0.25, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 28 s · fr · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.889 and the worst against the first clip 0.832; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.1 dB
per-clip
per-clip buttons play:
rule AB2k 4qmax 0.520cmax 0.249d_a -0.610d_b 0.520dataset podcastlang frtotal 28.2schain gain -3.3 dBseam step 1.1 dBcrossfades 150/150/100 msmin_cos_consec 0.8894min_cos_anchor 0.8320
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(average clarity, conversational, casual) the people, for example, they travel in the agroaliment because they have always (ahem) (ahem) touched,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: conversational, casual; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.7/10; 5.4s, FR.
901679_00289144 · in -16.0 dBFS · gain -3.3 dB · podcast-05593
(disappointment, triumph, distress · average clarity, casual, conversational) because they are expert that in the agroaliment is that it's super fascinating to people who work as industrial who have
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, triumph, distress; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 8.0/10; 10.9s, FR.
901679_00289684 · in -17.1 dBFS · gain -3.3 dB · podcast-00837
(disappointment · somewhat unclear, monologue) (low mumble) the example of (low mumble) this. It's (low mumble) (low mumble) super good to do after 10 years of lay.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment; style: monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.6/10; 7.8s, FR.
901679_00290772 · in -16.6 dBFS · gain -3.3 dB · podcast-00837
(average clarity, casual, conversational) And that's what is. (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 2.6/10; 4.5s, FR.
901679_00291552 · in -17.5 dBFS · gain -3.3 dB · podcast-05591
Anger ↓  /  Elationtimbre ≥0.80   strict_025 · #11

This chain comes from the two-sided rule: it only counts if both emotions move — Anger down and Elation up — by at least 0.25 each.

The chain starts with Elation around average — 0.46, lower than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.49.

At the same time Anger goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.48 (lower than 52 % of clips in this corpus), a change of -0.48. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 35 s · de · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.907 and the worst against the first clip 0.876; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.2 dB
per-clip
per-clip buttons play:
rule AB2k 3qmax 0.476cmax 0.248d_a -0.476d_b 0.486dataset emolialang detotal 34.9schain gain +0.1 dBseam step 2.2 dBcrossfades 150/150 msmin_cos_consec 0.9074min_cos_anchor 0.8759
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, quiet background, normal-paced, normally alert
(anger, disappointment, doubt · conversational, casual) okay.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, disappointment, doubt; style: conversational, casual; good recording, quiet background; genuineness 4.9/6; vocal-burst blend 6.0/10; 10.1s, DE.
DE_mOQsc65izl8_W000011 · in -20.3 dBFS · gain +0.1 dB · emolia-00107
(disappointment, contemplation, longing · conversational, casual) so.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as disappointment, contemplation, longing; style: conversational, casual; good recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.5/10; 7.9s, DE.
DE_mOQsc65izl8_W000012 · in -18.7 dBFS · gain +0.1 dB · emolia-00107
(elation, interest, disgust · conversational, casual) so, (low mumble) so.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as elation, interest, disgust; style: conversational, casual; good recording, quiet background; genuineness 4.3/6; vocal-burst blend 3.5/10; 17.1s, DE.
DE_mOQsc65izl8_W000013 · in -20.9 dBFS · gain +0.1 dB · emolia-00107
Hope Enthusiasm Optimism ↓  /  Feartimbre ≥0.80   strict_029 · #12

This chain comes from the two-sided rule: it only counts if both emotions move — Hope Enthusiasm Optimism down and Fear up — by at least 0.25 each.

The chain starts with Fear around average — 0.49, right about the corpus median — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.47.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.68 (higher than 68 % of clips in this corpus) to 0.20 (lower than 80 % of clips in this corpus), a change of -0.48. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 14 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.849 and the worst against the first clip 0.807; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.0 dB
per-clip
per-clip buttons play:
rule AB2k 3qmax 0.469cmax 0.247d_a -0.480d_b 0.469dataset emolialang entotal 13.9schain gain -2.0 dBseam step 3.0 dBcrossfades 150/100 msmin_cos_consec 0.8485min_cos_anchor 0.8067
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(casual, monologue) We should also mention that they are engaged in international training with the Israeli police.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 1.8/10; 4.3s, EN.
EN_34RHb_kmwNQ_W000046 · in -15.8 dBFS · gain -2.0 dB · emolia-00519
(authoritative, dramatic) (low mumble) Uhm, and so we think this project really is the beginning of a militarized police base
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, dramatic; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 1.1/10; 5.8s, EN.
EN_34RHb_kmwNQ_W000047 · in -18.8 dBFS · gain -2.0 dB · emolia-00519
(fear, emotional numbness · authoritative, formal) By (ahem) a police state, which is out of control at this particular state.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as fear, emotional numbness; style: authoritative, formal; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 1.0/10; 4.0s, EN.
EN_34RHb_kmwNQ_W000050 · in -20.7 dBFS · gain -2.0 dB · emolia-00519
Triumph ↓  /  Contemptidentity ≥0.80   strict_026 · #13

This chain comes from the two-sided rule: it only counts if both emotions move — Triumph down and Contempt up — by at least 0.25 each.

The chain starts with Contempt around average — 0.53, higher than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.46.

At the same time Triumph goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.51 (higher than 51 % of clips in this corpus), a change of -0.49. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.97 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 44 s · da · eurospeech

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.967 and the worst against the first clip 0.957; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.0 dB
per-clip
per-clip buttons play:
rule AB2k 3qmax 0.456cmax 0.248d_a -0.486d_b 0.456dataset eurospeechlang datotal 44.1schain gain +3.0 dBseam step 1.0 dBcrossfades 150/150 msmin_cos_consec 0.9670min_cos_anchor 0.9575
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · slightly rough, balanced body, average recording, wide pitch range
(triumph, pride, thankfulness gratitude · measured, normally alert, neutral tension, cartoonish) med et mobilitetshensyn. Det har bl.a. hr. Ole Birk Olesen gjort rede for hele tiden er en afvejning, og vi skal selvfølgelig stå på mål for sikkerheden. (low mumble) Jeg har respekt for SF's holdning i den her sag, for den har jeg haft hele tiden,
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as triumph, pride, thankfulness gratitude; style: cartoonish, casual; average recording, some background noise; genuineness 4.5/6; vocal-burst blend 1.4/10; 14.2s, DA.
denmark_20201M079_2021-03-11_1000_17078672_17092848 · in -23.6 dBFS · gain +3.0 dB · eurospeech-00420
(concentration, doubt, anger · normal-paced, subdued, neutral tension, casual) og I gik til valg på, at I ikke ville have store knallerter tilladt fra 16 år, i modsætning til Socialdemokratiet. Men jeg vil gerne spørge SF, når det åbenbart er meget farligt at køre 45 km/t. på et motoriseret køretøj som 16-årig – det er jo ligesom det, der ligger i det
full caption & clip details
A middle-aged masculine voice; delivery is subdued, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as concentration, doubt, anger; style: casual, monologue; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 6.4/10; 17.5s, DA.
denmark_20201M079_2021-03-11_1000_17092848_17110384 · in -22.6 dBFS · gain +3.0 dB · eurospeech-00420
(contempt, bitterness, disgust · measured, normally alert, slightly relaxed, authoritative) agter at foreslå en afskaffelse af speedpedelec, som man jo gerne må køre på som 15-årig, bare man har et kørekort til en lille knallert.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; very clear, frequent disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as contempt, bitterness, disgust; style: authoritative, cartoonish; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.0/10; 12.7s, DA.
denmark_20201M079_2021-03-11_1000_17110384_17123120 · in -23.2 dBFS · gain +3.0 dB · eurospeech-00420
Longing ↓  /  Triumphidentity ≥0.80   strict_027 · #14

This chain comes from the two-sided rule: it only counts if both emotions move — Longing down and Triumph up — by at least 0.25 each.

The chain starts with Triumph around average — 0.51, higher than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.49.

At the same time Longing goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.47 (lower than 53 % of clips in this corpus), a change of -0.45. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 37 s · es · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.900 and the worst against the first clip 0.900; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.5 dB
per-clip
per-clip buttons play:
rule AB2k 3qmax 0.452cmax 0.244d_a -0.452d_b 0.488dataset podcastlang estotal 37.1schain gain +1.4 dBseam step 0.5 dBcrossfades 100/150 msmin_cos_consec 0.8999min_cos_anchor 0.8999
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normally alert, some disfluency
(longing, intoxication altered states of consciousness · fast, slightly relaxed, fairly steady, casual) (ahem) La Liga Mexicana de Football, siempre tan querida, siempre tan controversial. Obviamente
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 6.6/10; 5.2s, ES.
272933_00011708 · in -20.9 dBFS · gain +1.4 dB · podcast-05085
(intoxication altered states of consciousness, fatigue exhaustion, pleasure ecstasy · normal-paced, slightly relaxed, fairly steady, casual) estuvo cargada de polémicas, estuvo cargada de resultados sorprendentes, pero dio inicio el día viernes 4 de abril con el Querétaro León, empatando a unos, un juego un poco desabrido, pero que León, bueno,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, fatigue exhaustion, pleasure ecstasy; style: casual, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 8.6/10; 13.7s, ES.
272933_00012220 · in -21.2 dBFS · gain +1.4 dB · podcast-05085
(triumph, pleasure ecstasy, astonishment surprise · fast, neutral tension, moderately variable, casual) (low mumble) empezó muy fuerte, empezó con un paso muy firme, ha ido bajando, pierde contra Santos, que era uno de los últimos lugares de la tabla, y en esta ocasión empata con el Querétaro a unos. Eso fue el día viernes. Cerrando la actividad del día viernes, Tijuana cae, dos goles contra uno en contra del Necaxa.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as triumph, pleasure ecstasy, astonishment surprise; style: casual, monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 10.0/10; 18.5s, ES.
272933_00013591 · in -21.7 dBFS · gain +1.4 dB · podcast-06493
Intoxication Altered States of Consciousness ↓  /  Emotional Numbnessidentity ≥0.80   strict_028 · #15

This chain comes from the two-sided rule: it only counts if both emotions move — Intoxication Altered States of Consciousness down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness around average — 0.53, higher than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.40.

At the same time Intoxication Altered States of Consciousness goes the other way, from 0.85 (higher than 85 % of clips in this corpus) to 0.38 (lower than 62 % of clips in this corpus), a change of -0.47. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 17 s · snippets

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.874 and the worst against the first clip 0.880; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.5 dB
per-clip
per-clip buttons play:
rule AB2k 3qmax 0.404cmax 0.237d_a -0.468d_b 0.404dataset snippetslang undtotal 16.5schain gain -4.2 dBseam step 0.5 dBcrossfades 100/100 msmin_cos_consec 0.8738min_cos_anchor 0.8802
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(formal, monologue) from 1935 in all cinemas, these terrible propaganda films are shown.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.9/10; 4.9s.
batch41_part4_batch41_part4_chunk_1366_1_1233307 · in -15.3 dBFS · gain -4.2 dB · snippets-01106
(formal, monologue) in a bunker under the monumental chancellery, of the three main accomplices
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.3/10; 4.2s.
batch41_part4_batch41_part4_chunk_1366_1_1233383 · in -15.8 dBFS · gain -4.2 dB · snippets-01106
(emotional numbness · formal, narration) A large chemical plant is set up at the cutting edge of technology at the time. In particular, it manufactures rubber for tires.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.6s.
batch41_part4_batch41_part4_chunk_1366_1_1233397 · in -16.2 dBFS · gain -4.2 dB · snippets-01106
Contempt ↓  /  Concentrationtimbre ≥0.80   strict_020 · #16

This chain comes from the two-sided rule: it only counts if both emotions move — Contempt down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.25.

At the same time Contempt goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

The largest step is 0.25, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 19 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.915 and the worst against the first clip 0.915; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.7 dB
per-clip
per-clip buttons play:
rule AB2k 2qmax 0.250cmax 0.250d_a -0.250d_b 0.250dataset emolialang entotal 19.4schain gain -5.4 dBseam step 0.7 dBcrossfades 150 msmin_cos_consec 0.9153min_cos_anchor 0.9153
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(contempt, disgust · fairly steady, formal, newsreading) On occasion, the Zoroastrian clergy assisted Muslims in attacks against those who they deemed Zoroastrian heretics
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, disgust; style: formal, newsreading; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.0/10; 6.6s, EN.
EN_dxpk1X2S13Q_W000007 · in -14.2 dBFS · gain -5.4 dB · emolia-02173
(concentration · steady, newsreading, formal) Until the Arab invasion and subsequent Muslim conquest, in the mid-7th century Persia – modern-day Iran – was a politically independent state, spanning from the Mesopotamia to the Indus River and dominated by a Zoroastrian majority
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 13.0s, EN.
EN_dxpk1X2S13Q_W000011 · in -14.9 dBFS · gain -5.4 dB · emolia-02173
Concentration ↓  /  Shametimbre ≥0.80   strict_023 · #17

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Shame up — by at least 0.25 each.

The chain starts with Shame around average — 0.50, right about the corpus median — and ends with it clearly present at 0.75, higher than 75 % of clips in this corpus. That is a total rise of 0.25.

At the same time Concentration goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 23 s · zh · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.826 and the worst against the first clip 0.826; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.8 dB
per-clip
per-clip buttons play:
rule AB2k 2qmax 0.250cmax 0.250d_a -0.250d_b 0.250dataset emolialang zhtotal 23.3schain gain -0.8 dBseam step 1.8 dBcrossfades 150 msmin_cos_consec 0.8261min_cos_anchor 0.8261
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, slightly dark, fairly smooth, balanced body, average recording, no background noise, normally alert, slightly relaxed
(concentration · measured, fairly steady, some disfluency, monologue) 子贡问他的老师,孔子说啊,老师有一个可以一辈子都遵守不备的字吗?孔子说,就是庶自己,不希望别人在自己身上做的事儿,也就不要强加到别人身上。
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, didactic; average recording, no background noise; genuineness 2.2/6; vocal-burst blend 3.9/10; 16.2s, ZH.
ZH_B00001_S02597_W000004 · in -18.7 dBFS · gain -0.8 dB · emolia-03289
(slow, steady, frequent disfluency, didactic) 从儒家的解释来看呢,术的第一层含义就是内省,然后推己及人。
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.7/10; 7.2s, ZH.
ZH_B00001_S02597_W000005 · in -20.5 dBFS · gain -0.8 dB · emolia-03289
Fatigue Exhaustion ↓  /  Confusiontimbre ≥0.80   strict_024 · #18

This chain comes from the two-sided rule: it only counts if both emotions move — Fatigue Exhaustion down and Confusion up — by at least 0.25 each.

The chain starts with Confusion clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.25.

At the same time Fatigue Exhaustion goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 18 s · ko · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.917 and the worst against the first clip 0.917; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.0 dB
per-clip
per-clip buttons play:
rule AB2k 2qmax 0.250cmax 0.250d_a -0.250d_b 0.250dataset emolialang kototal 18.0schain gain -0.1 dBseam step 1.0 dBcrossfades 150 msmin_cos_consec 0.9167min_cos_anchor 0.9167
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, moderate pitch range, narration, formal) 있다, 그라모, 그걸로, 재산을 늘려봅시다 하네요.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 6.0/10; 4.0s, KO.
KO_O4l6V73sLxc_W000055 · in -19.1 dBFS · gain -0.1 dB · emolia-03027
(confusion · measured, fairly narrow pitch, formal, narration) 이 때 내가 와이프한테 내세운 조건이 일단 집을 팔고 귀금속 다 팔고 모든 대출 다 받아서 몰빵하자고 그 와중에 승부수를 던집니다. 결과는 하나도 실천하지 않았습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 4.0/10; 14.2s, KO.
KO_O4l6V73sLxc_W000056 · in -20.2 dBFS · gain -0.1 dB · emolia-03027
Bitterness ↓  /  Jealousy and Envyidentity ≥0.80   strict_021 · #19

This chain comes from the two-sided rule: it only counts if both emotions move — Bitterness down and Jealousy and Envy up — by at least 0.25 each.

The chain starts with Jealousy and Envy clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.25.

At the same time Bitterness goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 21 s · en · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.894 and the worst against the first clip 0.894; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.9 dB
per-clip
per-clip buttons play:
rule AB2k 2qmax 0.250cmax 0.250d_a -0.250d_b 0.250dataset podcastlang entotal 20.7schain gain -1.6 dBseam step 0.9 dBcrossfades 100 msmin_cos_consec 0.8938min_cos_anchor 0.8938
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(bitterness · fairly steady, moderate pitch range, casual, conversational) It stems from comfortability. I firmly believe in that because you have people, for example, who are overweight who want to lose weight, or people who are underweight, very skinny, don't eat enough, genetics are insane, and they burn way too many
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as bitterness; style: casual, conversational; good recording, no background noise; genuineness 4.3/6; vocal-burst blend 7.0/10; 13.0s, EN.
913266_00092728 · in -18.8 dBFS · gain -1.6 dB · podcast-03743
(jealousy and envy, infatuation, longing · moderately variable, wide pitch range, casual, conversational) And then you have someone who has the exact look. It could be, for example, my mom, Braden's mom. They're always like, I'm too skinny.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as jealousy and envy, infatuation, longing; style: casual, conversational; good recording, no background noise; genuineness 3.8/6; vocal-burst blend 6.6/10; 7.8s, EN.
913266_00094192 · in -17.8 dBFS · gain -1.6 dB · podcast-03732
Malevolence Malice ↓  /  Confusionidentity ≥0.80   strict_022 · #20

This chain comes from the two-sided rule: it only counts if both emotions move — Malevolence Malice down and Confusion up — by at least 0.25 each.

The chain starts with Confusion clearly present — 0.74, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.

At the same time Malevolence Malice goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 29 s · mt · eurospeech

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.936 and the worst against the first clip 0.936; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.8 dB
per-clip
per-clip buttons play:
rule AB2k 2qmax 0.247cmax 0.250d_a -0.250d_b 0.247dataset eurospeechlang mttotal 28.9schain gain +0.5 dBseam step 0.8 dBcrossfades 100 msmin_cos_consec 0.9355min_cos_anchor 0.9355
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · slightly cool, rough, thin, normal-paced, energised, neutral tension, moderately variable, wide pitch range
(malevolence malice, anger, bitterness · some disfluency, somewhat unclear, cartoonish, authoritative) u illi l-Kamra taħtar ukoll lil dawn is-sostituti tagħhom: L-Onor. Leo Brincat L-Onor. Dolores Cristina L-Onor. Joseph Falzon Huwa wkoll riżolut illi din innomina ta’ kull Rappreżentant u ta’ kull Sostitut tiegħu
full caption & clip details
A middle-aged masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly dark, rough, thin; somewhat unclear, some disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, fairly guarded; reads as malevolence malice, anger, bitterness; style: cartoonish, authoritative; poor recording, some background noise; genuineness 2.4/6; vocal-burst blend 4.5/10; 15.9s, MT.
malta_11_020_18062008_1750128_1766047 · in -20.2 dBFS · gain +0.5 dB · eurospeech-02066
(confusion, awe, thankfulness gratitude · almost no disfluency, slurred, cartoonish, authoritative) tibqa’ fis-seħħ sakemm dawn ir-Rappreżentanti u Sostituti jibqgħu Membri ta’ dan ilParlament jew sakemm din ilKamra ma tiddeċidix xort’ oħra. Illi l-Onorevoli Membri, hekk maħtura għandhom,
full caption & clip details
A middle-aged masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, dark, rough, thin; slurred, almost no disfluency, wide pitch range, audible breath; affect is negative, slightly dominant, fairly guarded; reads as confusion, awe, thankfulness gratitude; style: cartoonish, authoritative; below-average recording, quiet background; genuineness 1.6/6; vocal-burst blend 2.9/10; 13.0s, MT.
malta_11_020_18062008_1766047_1779088 · in -21.0 dBFS · gain +0.5 dB · eurospeech-02066