VoiceNet (VN1)

What this rule requires. one of 57 VoiceNet voice-descriptor dimensions sweeps by at least T; no proxies — with a per-step cap of C = 0.25, and here only chains whose speaker similarity clears 0.80 on both conditions. No voice conversion has been applied.
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken. Underlined descriptors are the ones that change across the chain; brackets inside the words are a real non-speech sound. The perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
S_MONO — style: monologuetimbre ≥0.80   strict_055 · #1

This is a VoiceNet dimension, not an emotion: style: monologue (S_MONO) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with style: monologue (S_MONO) low — 0.13, lower than 87 % of clips in this corpus — and ends with it at the very top of the range at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.85.

It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.16, then +0.25, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 27 s · ko · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.909 and the worst against the first clip 0.899; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/100/100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.7 dB
per-clip
per-clip buttons play:
rule VN1k 5qmax 0.850cmax 0.248d_a 0.850d_b 0.850dataset emolialang kototal 27.4schain gain -3.2 dBseam step 1.7 dBcrossfades 150/100/100/150 msmin_cos_consec 0.9092min_cos_anchor 0.8987
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, slightly relaxed, fairly steady, some disfluency
(helplessness, pain · fast, normally alert, average clarity, casual) 나중에 사이사이에 길이 만들어진다고 가정하면 이렇게 해야 될 것 같아요.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as helplessness, pain; style: casual, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.1/10; 3.6s, KO.
KO_aWOsiMf_bxu_W000103 · in -14.8 dBFS · gain -3.2 dB · emolia-03220
(measured, normally alert, slurred, casual) 일단 이거 받을게요. 운이라도 버려야 되니까. 이렇게 산대 받아주고.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 2.2/10; 4.4s, KO.
KO_aWOsiMf_bxu_W000104 · in -16.5 dBFS · gain -3.2 dB · emolia-03220
(confusion · measured, normally alert, slurred, casual) 하우스도 좀 더 확장을 할까요? (low mumble) 음, 2층으로 올려도 될 것 같긴 한데. 아, 그럴 필요 없구나. 자, 이거를.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as confusion; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 1.1/10; 7.4s, KO.
KO_aWOsiMf_bxu_W000105 · in -17.0 dBFS · gain -3.2 dB · emolia-03220
(doubt · measured, normally alert, slurred, monologue) 자, 연료가 부족해지니까 이쪽으로 이동해서 Fabricator Large를 만들고.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as doubt; style: monologue, ASMR; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 5.1/10; 5.2s, KO.
KO_aWOsiMf_bxu_W000106 · in -18.7 dBFS · gain -3.2 dB · emolia-03220
(contemplation · measured, subdued, somewhat unclear, monologue) 자, 자원이 모자르니까 좀 자원 많이 쥔 대로 좀, (low mumble) 음, 빨리 개간을 시켜야 될 것 같아요. 이런 데 있죠. 개간 시킵시다.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as contemplation; style: monologue, ASMR; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.6/10; 7.2s, KO.
KO_aWOsiMf_bxu_W000107 · in -17.3 dBFS · gain -3.2 dB · emolia-03220
S_CART — style: cartoonishtimbre ≥0.80   strict_056 · #2

This is a VoiceNet dimension, not an emotion: style: cartoonish (S_CART) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with style: cartoonish (S_CART) at the very bottom of the range — 0.04, lower than 96 % of clips in this corpus — and ends with it high at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.83.

It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.19, then +0.20, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 31 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.839 and the worst against the first clip 0.898; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/100/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.0 dB
per-clip
per-clip buttons play:
rule VN1k 5qmax 0.830cmax 0.233d_a 0.830d_b 0.830dataset emolialang entotal 30.7schain gain +5.6 dBseam step 3.0 dBcrossfades 150/100/150/150 msmin_cos_consec 0.8391min_cos_anchor 0.8982
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, light breath
(pleasure ecstasy, elation, embarrassment · normal-paced, normally alert, neutral tension, casual) I (low mumble) am a PlayStation Plus user, which you would think, oh, that's going to be DRM to hell, right? You know what happens when I play that in airplane mode on a plane? It works.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as pleasure ecstasy, elation, embarrassment; style: casual, conversational; good recording, no background noise; genuineness 3.8/6; vocal-burst blend 6.1/10; 8.3s, EN.
EN_B00039_S03178_W000277 · in -27.7 dBFS · gain +5.6 dB · emolia-01007
(sexual lust, pleasure ecstasy, infatuation · normal-paced, normally alert, slightly relaxed, casual) You know what happens when I don't have an internet connection? I'll play it. It works. Every bloody time.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sexual lust, pleasure ecstasy, infatuation; style: casual, conversational; good recording, no background noise; genuineness 2.4/6; vocal-burst blend 1.5/10; 5.3s, EN.
EN_B00039_S03178_W000278 · in -26.4 dBFS · gain +5.6 dB · emolia-01007
(bitterness, anger, impatience and irritability · normal-paced, energised, slightly relaxed, authoritative) Every time. Same with PSN and everything I download from that.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, fairly guarded; reads as bitterness, anger, impatience and irritability; style: authoritative, storytelling; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.4/10; 4.0s, EN.
EN_B00039_S03178_W000279 · in -27.4 dBFS · gain +5.6 dB · emolia-01007
(sourness, anger, teasing · normal-paced, energised, slightly relaxed, casual) Evidently, if you're being outdone on the DRM front by a company like Sony, then you are doing it wrong.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as sourness, anger, teasing; style: casual, conversational; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 1.3/10; 7.5s, EN.
EN_B00039_S03178_W000280 · in -25.9 dBFS · gain +5.6 dB · emolia-01007
(impatience and irritability, anger, contempt · brisk, energised, slightly relaxed, storytelling) And these are the guys that are patenting DRM to make discs not work on more than one machine.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, anger, contempt; style: storytelling, conversational; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.6/10; 6.3s, EN.
EN_B00039_S03178_W000281 · in -22.9 dBFS · gain +5.6 dB · emolia-01007
R_ORAL — resonance: oraltimbre ≥0.80   strict_059 · #3

This is a VoiceNet dimension, not an emotion: resonance: oral (R_ORAL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with resonance: oral (R_ORAL) at the very top of the range — 0.98, higher than 98 % of clips in this corpus — and works its way down to low at 0.15, lower than 85 % of clips in this corpus. That is a total fall of 0.83.

It takes 5 clips to get there. Clip to clip the moves are -0.25, then -0.12, then -0.24, then -0.22 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 29 s · zh · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.938 and the worst against the first clip 0.919; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.6 dB
per-clip
per-clip buttons play:
rule VN1k 5qmax 0.829cmax 0.247d_a -0.829d_b -0.829dataset emolialang zhtotal 29.1schain gain -1.6 dBseam step 2.6 dBcrossfades 150/150/150/150 msmin_cos_consec 0.9384min_cos_anchor 0.9195
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(pain, confusion · normal-paced, authoritative, storytelling) 而在教室门外,张建明忍不住看了看徐洛。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain, confusion; style: authoritative, storytelling; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 4.8/10; 3.2s, ZH.
ZH_B00013_S08715_W000013 · in -19.6 dBFS · gain -1.6 dB · emolia-03407
(normal-paced, formal, authoritative) 听说你入侵了向阳的神域,杀了他的哥布林。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 4.8/10; 3.4s, ZH.
ZH_B00013_S08715_W000014 · in -18.0 dBFS · gain -1.6 dB · emolia-03407
(normal-paced, formal, narration) 听到这个消息的时候,张建明还觉得自己是在梦中呢。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 5.7/10; 4.0s, ZH.
ZH_B00013_S08715_W000015 · in -19.9 dBFS · gain -1.6 dB · emolia-03407
(pride · measured, formal, authoritative) 张老师纠正一下,是他自己入侵,我的神谕,我只是反击而已,死这么说是真的了。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: formal, authoritative; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 4.7/10; 7.3s, ZH.
ZH_B00013_S08715_W000016 · in -17.3 dBFS · gain -1.6 dB · emolia-03407
(contentment · measured, narration, monologue) 张建明还是感觉到不可置信,徐洛的初始生物不是爬虫类吗?这可以说是最弱的物种了。虽然哥布林也不是太盛的存在,可是相比于爬虫来说也是天差地别啊。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment; style: narration, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 4.9/10; 11.9s, ZH.
ZH_B00013_S08715_W000017 · in -19.0 dBFS · gain -1.6 dB · emolia-03407
DFLU — disfluencyidentity ≥0.80   strict_057 · #4

This is a VoiceNet dimension, not an emotion: disfluency (DFLU) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with disfluency (DFLU) high — 0.88, higher than 88 % of clips in this corpus — and works its way down to low at 0.11, lower than 89 % of clips in this corpus. That is a total fall of 0.77.

It takes 5 clips to get there. Clip to clip the moves are -0.14, then -0.24, then -0.22, then -0.18 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 47 s · en · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.814 and the worst against the first clip 0.813; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/100/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.1 dB
per-clip
per-clip buttons play:
rule VN1k 5qmax 0.775cmax 0.239d_a -0.775d_b -0.775dataset podcastlang entotal 47.2schain gain +10.8 dBseam step 2.1 dBcrossfades 100/100/150/150 msmin_cos_consec 0.8142min_cos_anchor 0.8129
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, moderate pitch range, light breath
(infatuation, longing, jealousy and envy · normal-paced, normally alert, relaxed, casual) Yeah. Where they don't have the French yogurt. I mean they have whey, but that's like a different vibe. That's like the lesser version of the one I like, but that one's good too. I like little jars. And like, wait, like this is like six dollars for like a box of cheese it's like I could get
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as infatuation, longing, jealousy and envy; style: casual, monologue; average recording, no background noise; mildly explicit content; genuineness 5.2/6; vocal-burst blend 7.2/10; 16.4s, EN.
840040_00066336 · in -31.1 dBFS · gain +10.8 dB · podcast-04314
(jealousy and envy, infatuation, fatigue exhaustion · normal-paced, normally alert, neutral tension, casual) no, they're not cheap. Like I I just I hadn't been like I hadn't had access to such variety in a while. Like I was like really surprised it was so expensive. And I think it's gotten more expensive recently too.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as jealousy and envy, infatuation, fatigue exhaustion; style: casual, conversational; average recording, no background noise; genuineness 5.6/6; vocal-burst blend 9.0/10; 10.7s, EN.
840040_00068672 · in -29.2 dBFS · gain +10.8 dB · podcast-04326
(interest · normal-paced, normally alert, slightly relaxed, casual) (ahem) Um those sorts of items, (ahem) uh just because they go through like a lot of processes before they get to the store, so you're adding tariffs and cost of labor, which is gonna get more and more expensive (ahem) um because we're kicking out immigrants who want to do factory labor.
full caption & clip details
A young adult somewhat masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as interest; style: casual, whispered; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.4/10; 14.1s, EN.
840040_00069864 · in -31.4 dBFS · gain +10.8 dB · podcast-02397
(embarrassment · normal-paced, normally alert, slightly relaxed, casual) So yeah, I mean I I think people don't realize that you can
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as embarrassment; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 4.3/10; 3.0s, EN.
840040_00071272 · in -31.2 dBFS · gain +10.8 dB · podcast-04312
(fatigue exhaustion, contentment · measured, subdued, slightly relaxed, casual) it takes more time, of course, to prepare
full caption & clip details
A young adult feminine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as fatigue exhaustion, contentment; style: casual; good recording, no background noise; mildly explicit content; genuineness 2.3/6; vocal-burst blend 4.8/10; 3.4s, EN.
840040_00071576 · in -32.8 dBFS · gain +10.8 dB · podcast-04326
TENS — tensionidentity ≥0.80   strict_058 · #5

This is a VoiceNet dimension, not an emotion: tension (TENS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with tension (TENS) low — 0.11, lower than 89 % of clips in this corpus — and ends with it high at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.76.

It takes 5 clips to get there. Clip to clip the moves are +0.08, then +0.22, then +0.21, then +0.25 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 42 s · en · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.853 and the worst against the first clip 0.881; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/100/100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 2.5 dB
per-clip
per-clip buttons play:
rule VN1k 5qmax 0.759cmax 0.248d_a 0.759d_b 0.759dataset podcastlang entotal 42.5schain gain +8.4 dBseam step 2.5 dBcrossfades 150/100/100/150 msmin_cos_consec 0.8531min_cos_anchor 0.8813
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, normal-paced, normally alert, average clarity, light breath
(teasing, amusement, contempt · slightly relaxed, fairly steady, some disfluency, casual) you know you scroll, you you look you look over to your left or right or whatever, and you hover over the icon to to warp to that, and it'll you know, then the the good
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as teasing, amusement, contempt; style: casual, conversational; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 5.3/10; 7.9s, EN.
269133_00232236 · in -30.5 dBFS · gain +8.4 dB · podcast-02141
(embarrassment, sourness · relaxed, moderately variable, frequent disfluency, casual) around. (ahem) Um you do customize your character if you want, (low mumble) um, because it takes place like either on the the new republic or the empire. So you're playing as both both sides, which which is kind of cool. Oh,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as embarrassment, sourness; style: casual, playful; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 1.5/10; 13.3s, EN.
269133_00233280 · in -29.0 dBFS · gain +8.4 dB · podcast-01023
(fatigue exhaustion, infatuation · neutral tension, moderately variable, some disfluency, casual) yeah, it's it's pretty solid. (ahem) Um I'll probably end up depending. I think it's like four or five hours, so it's pretty short. So
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as fatigue exhaustion, infatuation; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 5.0/10; 5.5s, EN.
269133_00234840 · in -26.5 dBFS · gain +8.4 dB · podcast-01025
(embarrassment, confusion, disappointment · neutral tension, moderately variable, frequent disfluency, casual) the multiplayer at all? I didn't. I I really don't have any interest in the multiplayer. I I may check it out once just to say I did it, but (chuckle) (low mumble) um, yeah, I just yeah, I don't really have much interest in the multiplayer. Uh
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as embarrassment, confusion, disappointment; style: casual, conversational; average recording, quiet background; genuineness 5.6/6; vocal-burst blend 7.7/10; 12.8s, EN.
269133_00235736 · in -28.4 dBFS · gain +8.4 dB · podcast-01015
(pride, embarrassment, triumph · slightly relaxed, fairly steady, some disfluency, casual) and (ahem) uh I never I've played the original Halo Wars, but haven't played the second one,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as pride, embarrassment, triumph; style: casual, conversational; good recording, no background noise; genuineness 4.3/6; vocal-burst blend 5.1/10; 3.5s, EN.
269133_00237760 · in -27.0 dBFS · gain +8.4 dB · podcast-01017
ATCK — attack / onset sharpnesstimbre ≥0.80   strict_050 · #6

This is a VoiceNet dimension, not an emotion: attack / onset sharpness (ATCK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with attack / onset sharpness (ATCK) high — 0.83, higher than 83 % of clips in this corpus — and works its way down to low at 0.10, lower than 90 % of clips in this corpus. That is a total fall of 0.73.

It takes 4 clips to get there. Clip to clip the moves are -0.25, then -0.25, then -0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 24 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.912 and the worst against the first clip 0.868; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.2 dB
per-clip
per-clip buttons play:
rule VN1k 4qmax 0.732cmax 0.248d_a -0.732d_b -0.732dataset emolialang entotal 24.1schain gain +0.1 dBseam step 3.2 dBcrossfades 150/150/100 msmin_cos_consec 0.9125min_cos_anchor 0.8681
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, slightly relaxed, average clarity, moderate pitch range
(triumph · normal-paced, normally alert, fairly steady, casual) (ahem) Uhm, all throughout, and then here we're just calling them resultants. And so you've already taken
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as triumph; style: casual, authoritative; good recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.6/10; 4.1s, EN.
EN_B00076_S06764_W000022 · in -19.7 dBFS · gain +0.1 dB · emolia-01700
(malevolence malice, fear, teasing · normal-paced, normally alert, moderately variable, casual) If you cross any two vectors that are parallel, you get zero.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as malevolence malice, fear, teasing; style: casual, storytelling; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.7/10; 4.1s, EN.
EN_B00076_S06764_W000023 · in -17.8 dBFS · gain +0.1 dB · emolia-01700
(normal-paced, normally alert, fairly steady, casual) (ahem) If it helps simplify some of your components. Coordinate systems are tools to use. They're not
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 1.3/10; 4.9s, EN.
EN_B00076_S06764_W000024 · in -21.0 dBFS · gain +0.1 dB · emolia-01700
(concentration · measured, subdued, fairly steady, casual) The amount, right, not a vector, but the amount or length of one vector that is along another. And we use the COMP (low mumble) indicating this is the component.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: casual, didactic; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 2.4/10; 11.3s, EN.
EN_B00076_S06764_W000025 · in -21.2 dBFS · gain +0.1 dB · emolia-01700
VALS — valence stabilitytimbre ≥0.80   strict_051 · #7

This is a VoiceNet dimension, not an emotion: valence stability (VALS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with valence stability (VALS) at the very top of the range — 0.96, higher than 96 % of clips in this corpus — and works its way down to low at 0.23, lower than 77 % of clips in this corpus. That is a total fall of 0.73.

It takes 4 clips to get there. Clip to clip the moves are -0.23, then -0.25, then -0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 40 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.889 and the worst against the first clip 0.839; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 5.3 dB
per-clip
per-clip buttons play:
rule VN1k 4qmax 0.727cmax 0.248d_a -0.727d_b -0.727dataset emolialang entotal 40.1schain gain -1.4 dBseam step 5.3 dBcrossfades 150/150/150 msmin_cos_consec 0.8891min_cos_anchor 0.8391
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normally alert, fairly steady
(normal-paced, slightly relaxed, frequent disfluency, casual) Okay. Oh, I see many comments now. Let's see. Let me better cover.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, playful; good recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.4/10; 6.0s, EN.
EN_B00066_S08833_W000035 · in -22.1 dBFS · gain -1.4 dB · emolia-01507
(normal-paced, slightly relaxed, some disfluency, casual) Okay, so let me check some of the comments now. I think I'm reading some old comments, so let me scroll down a little bit.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.9/10; 7.7s, EN.
EN_B00066_S08833_W000036 · in -16.8 dBFS · gain -1.4 dB · emolia-01507
(concentration · measured, slightly relaxed, frequent disfluency, monologue) If daily is down and four hours bullish, then follow the daily timeframe. That means you wait until the lower timeframes point down to sell.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration; style: monologue, authoritative; good recording, quiet background; genuineness 0.9/6; vocal-burst blend 0.3/10; 10.5s, EN.
EN_B00066_S08833_W000037 · in -19.7 dBFS · gain -1.4 dB · emolia-01507
(bitterness, concentration · measured, neutral tension, frequent disfluency, monologue) So I recommend you to avoid this kind of a trace. So remember when it's range, the market goes up and down. There is no direction and you might have to endure the profits and losses for a couple of days.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as bitterness, concentration; style: monologue, didactic; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.8/10; 16.3s, EN.
EN_B00066_S08833_W000038 · in -18.1 dBFS · gain -1.4 dB · emolia-01507
COGL — cognitive loadtimbre ≥0.80   strict_054 · #8

This is a VoiceNet dimension, not an emotion: cognitive load (COGL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with cognitive load (COGL) high — 0.83, higher than 83 % of clips in this corpus — and works its way down to low at 0.11, lower than 89 % of clips in this corpus. That is a total fall of 0.72.

It takes 4 clips to get there. Clip to clip the moves are -0.24, then -0.25, then -0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.80 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.80. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 19 s · ja · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.803 and the worst against the first clip 0.803; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.3 dB
per-clip
per-clip buttons play:
rule VN1k 4qmax 0.723cmax 0.249d_a -0.723d_b -0.723dataset emolialang jatotal 19.2schain gain -2.1 dBseam step 1.3 dBcrossfades 150/150/150 msmin_cos_consec 0.8034min_cos_anchor 0.8034
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a child feminine voice · neutral-toned, fairly smooth, no background noise, light breath
(teasing, pain, confusion · measured, normally alert, slightly relaxed, storytelling) ほんとだ、ちょっと目立つわね。そっちのチェックのシャツも捨てるの?
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as teasing, pain, confusion; style: storytelling, conversational; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 6.8/10; 5.5s, JA.
JA_B00001_S09841_W000214 · in -18.1 dBFS · gain -2.1 dB · emolia-02947
(normal-paced, normally alert, slightly relaxed, storytelling) 旅先で男の人と女の人が話しています。
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: storytelling, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 3.4/10; 4.3s, JA.
JA_B00001_S09841_W000215 · in -19.4 dBFS · gain -2.1 dB · emolia-02947
(measured, normally alert, slightly relaxed, formal) 2人はどんな順番で動きますか?
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, neutral openness; no dominant emotion; style: formal, storytelling; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 2.4/10; 3.3s, JA.
JA_B00001_S09841_W000216 · in -18.4 dBFS · gain -2.1 dB · emolia-02947
(sexual lust, longing, confusion · fast, energised, neutral tension, storytelling) 行きたいところは美術館と有名な仏像のあるお寺と、あとはお土産の買い物かな。
full caption & clip details
A child feminine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, no disfluency, very wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as sexual lust, longing, confusion; style: storytelling, cartoonish; average recording, no background noise; genuineness 2.3/6; vocal-burst blend 4.7/10; 6.6s, JA.
JA_B00001_S09841_W000217 · in -17.1 dBFS · gain -2.1 dB · emolia-02947
FULL — fullness of toneidentity ≥0.80   strict_052 · #9

This is a VoiceNet dimension, not an emotion: fullness of tone (FULL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with fullness of tone (FULL) high — 0.82, higher than 82 % of clips in this corpus — and works its way down to low at 0.12, lower than 88 % of clips in this corpus. That is a total fall of 0.70.

It takes 4 clips to get there. Clip to clip the moves are -0.24, then -0.25, then -0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 27 s · en · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.824 and the worst against the first clip 0.869; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.1 dB
per-clip
per-clip buttons play:
rule VN1k 4qmax 0.703cmax 0.246d_a -0.703d_b -0.703dataset podcastlang entotal 27.1schain gain +6.6 dBseam step 1.1 dBcrossfades 150/100/150 msmin_cos_consec 0.8245min_cos_anchor 0.8692
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, some disfluency
(triumph, pride, thankfulness gratitude · normal-paced, conversational, casual) and not doing it, and a relationship, and then I'm easy and other things. So it was
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as triumph, pride, thankfulness gratitude; style: conversational, casual; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 9.0/10; 10.7s, EN.
60611_00121756 · in -26.1 dBFS · gain +6.6 dB · podcast-03844
(pride · normal-paced, casual, conversational) with (ahem) him to
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as pride; style: casual, conversational; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 3.7/10; 5.8s, EN.
60611_00122824 · in -26.8 dBFS · gain +6.6 dB · podcast-03842
(impatience and irritability · brisk, casual, conversational) But then the house and
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as impatience and irritability; style: casual, conversational; average recording, no background noise; genuineness 4.2/6; vocal-burst blend 8.6/10; 5.8s, EN.
60611_00123400 · in -27.4 dBFS · gain +6.6 dB · podcast-03839
(affection, pride, contentment · normal-paced, casual, conversational) I attribute much to (ahem) you, you were chicken, it was chicken. So
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as affection, pride, contentment; style: casual, conversational; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 3.2/10; 5.2s, EN.
60611_00124824 · in -26.3 dBFS · gain +6.6 dB · podcast-02420
COGL — cognitive loadidentity ≥0.80   strict_053 · #10

This is a VoiceNet dimension, not an emotion: cognitive load (COGL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with cognitive load (COGL) low — 0.23, lower than 77 % of clips in this corpus — and ends with it at the very top of the range at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.70.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.23, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 47 s · en · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.857 and the worst against the first clip 0.816; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/100/100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.1 dB
per-clip
per-clip buttons play:
rule VN1k 4qmax 0.702cmax 0.244d_a 0.702d_b 0.702dataset podcastlang entotal 47.2schain gain +6.2 dBseam step 1.1 dBcrossfades 100/100/100 msmin_cos_consec 0.8566min_cos_anchor 0.8158
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, slightly relaxed, fairly steady
(normal-paced, normally alert, frequent disfluency, casual) Those five countries are Australia, the US, which is mostly Alaska, (ahem) Brazil, Russia, and Canada.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 1.1/10; 7.9s, EN.
291478_00203412 · in -26.7 dBFS · gain +6.2 dB · podcast-04256
(normal-paced, normally alert, some disfluency, casual) 70% of the world's wilderness (ahem) in those five countries, the rest a lot of it is in Africa.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual; average recording, no background noise; genuineness 4.4/6; vocal-burst blend 1.5/10; 6.6s, EN.
291478_00204304 · in -26.0 dBFS · gain +6.2 dB · podcast-01740
(measured, normally alert, frequent disfluency, casual) 77% of land, not including Antarctica, and 87% of oceans had been modified by human intervention. So thoughts
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, whispered; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 0.4/10; 12.5s, EN.
291478_00205071 · in -26.8 dBFS · gain +6.2 dB · podcast-01750
(pride, awe, fear · slow, very low-energy, frequent disfluency, casual) modified by humans. So human beings, there's close to 7 billion right now on Earth. (ahem) The second most populous primate is the crab eating macaque from Southeast Asia at 2.5 million.
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, slightly guarded; reads as pride, awe, fear; style: casual, whispered; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 1.2/10; 20.5s, EN.
291478_00206600 · in -25.7 dBFS · gain +6.2 dB · podcast-01738
S_RANT — style: rantingtimbre ≥0.80   strict_045 · #11

This is a VoiceNet dimension, not an emotion: style: ranting (S_RANT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with style: ranting (S_RANT) above average — 0.62, higher than 62 % of clips in this corpus — and works its way down to low at 0.12, lower than 88 % of clips in this corpus. That is a total fall of 0.50.

It takes 3 clips to get there. Clip to clip the moves are -0.25, then -0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 18 s · zh · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.868 and the worst against the first clip 0.868; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 3.7 dB
per-clip
per-clip buttons play:
rule VN1k 3qmax 0.499cmax 0.250d_a -0.499d_b -0.499dataset emolialang zhtotal 18.1schain gain +8.3 dBseam step 3.7 dBcrossfades 100/150 msmin_cos_consec 0.8678min_cos_anchor 0.8678
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, normally alert, slightly relaxed, fairly steady, some disfluency
(thankfulness gratitude, confusion · fast, average clarity, casual, conversational) 放在手机半个小时已经过去了,但是意犹未尽根本就进入不了状态。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, confusion; style: casual, conversational; average recording, no background noise; genuineness 3.2/6; vocal-burst blend 4.8/10; 3.3s, ZH.
ZH_B00041_S05099_W000008 · in -26.5 dBFS · gain +8.3 dB · emolia-03691
(fast, slurred, casual, conversational) 但是现在我回完消息马上锁屏,根本不想浪费一点时间。所以过去的一周里,我每天只用了半小时手机。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 5.9/10; 7.3s, ZH.
ZH_B00041_S05099_W000009 · in -27.3 dBFS · gain +8.3 dB · emolia-03691
(contemplation, doubt · measured, slurred, whispered, monologue) 我睡眠更好了,出门更多了,情绪更积极了,做事更专注了。但当我真正做到了,才发现,放下手机。
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, doubt; style: whispered, monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 3.1/10; 7.7s, ZH.
ZH_B00041_S05099_W000010 · in -31.0 dBFS · gain +8.3 dB · emolia-03691
S_NARR — style: narrationtimbre ≥0.80   strict_048 · #12

This is a VoiceNet dimension, not an emotion: style: narration (S_NARR) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with style: narration (S_NARR) above average — 0.72, higher than 72 % of clips in this corpus — and works its way down to low at 0.23, lower than 77 % of clips in this corpus. That is a total fall of 0.50.

It takes 3 clips to get there. Clip to clip the moves are -0.25, then -0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 24 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.855 and the worst against the first clip 0.879; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.7 dB
per-clip
per-clip buttons play:
rule VN1k 3qmax 0.499cmax 0.250d_a -0.499d_b -0.499dataset emolialang entotal 23.9schain gain -2.4 dBseam step 0.7 dBcrossfades 150/150 msmin_cos_consec 0.8548min_cos_anchor 0.8787
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · fairly smooth, quiet background, normal-paced, normally alert, average clarity, light breath
(thankfulness gratitude, affection, contentment · slightly relaxed, fairly steady, little disfluency, monologue) To make sure that for those who do need an extra helping hand, we are giving them an extra helping hand.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, vulnerable; reads as thankfulness gratitude, affection, contentment; style: monologue, formal; good recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.5/10; 5.5s, EN.
EN_ZdH3r-KqWmA_W000134 · in -17.9 dBFS · gain -2.4 dB · emolia-00509
(sadness, disappointment, distress · slightly relaxed, fairly steady, little disfluency, authoritative) These have been difficult decisions to land. Some say it's too much, others say it's not enough. But I think you can see the genuineness with the approach that the Albanese government has taken
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, little disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as sadness, disappointment, distress; style: authoritative, formal; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.8/10; 11.3s, EN.
EN_ZdH3r-KqWmA_W000135 · in -17.8 dBFS · gain -2.4 dB · emolia-00509
(relief, triumph · neutral tension, moderately variable, some disfluency, formal) When we said we would assess payments, we would do what we can to adjust them in every budget, we've been doing that.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as relief, triumph; style: formal, authoritative; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.3/10; 7.5s, EN.
EN_ZdH3r-KqWmA_W000136 · in -17.2 dBFS · gain -2.4 dB · emolia-00509
S_TECH — style: technicaltimbre ≥0.80   strict_049 · #13

This is a VoiceNet dimension, not an emotion: style: technical (S_TECH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with style: technical (S_TECH) low — 0.23, lower than 77 % of clips in this corpus — and ends with it above average at 0.73, higher than 73 % of clips in this corpus. That is a total rise of 0.50.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 28 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.825 and the worst against the first clip 0.887; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.3 dB
per-clip
per-clip buttons play:
rule VN1k 3qmax 0.499cmax 0.249d_a 0.499d_b 0.499dataset emolialang entotal 28.3schain gain -0.3 dBseam step 1.3 dBcrossfades 150/150 msmin_cos_consec 0.8252min_cos_anchor 0.8874
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, brisk, energised, some disfluency, light breath
(fear, relief, doubt · neutral tension, moderately variable, average clarity, casual) Right, if they're a little more assertive and a little less fearful, you might not be able to tyrannize over them so easily. And so it's not necessarily the case at all that you would be happy about that.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as fear, relief, doubt; style: casual, storytelling; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 5.7/10; 8.1s, EN.
EN_B00032_S03853_W000238 · in -18.8 dBFS · gain -0.3 dB · emolia-00864
(interest, contempt, malevolence malice · neutral tension, moderately variable, average clarity, casual) So you gotta watch that sort of thing too, and maybe the person wouldn't even be that happy about it, because they're getting all sort of secondary benefits from being, you know, neurotic and martyred, because that's a vicious weapon. To be weak and useless, if you can wield that as a weapon, it's extraordinarily effective.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as interest, contempt, malevolence malice; style: casual, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 6.5/10; 16.1s, EN.
EN_B00032_S03853_W000239 · in -20.1 dBFS · gain -0.3 dB · emolia-00864
(sourness · slightly relaxed, fairly steady, clear, casual) So you gotta watch for that sort of thing to be working against your psychotherapeutic aims as well.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sourness; style: casual, conversational; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 2.3/10; 4.4s, EN.
EN_B00032_S03853_W000240 · in -20.8 dBFS · gain -0.3 dB · emolia-00864
FULL — fullness of toneidentity ≥0.80   strict_046 · #14

This is a VoiceNet dimension, not an emotion: fullness of tone (FULL) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with fullness of tone (FULL) around average — 0.54, higher than 54 % of clips in this corpus — and works its way down to at the very bottom of the range at 0.04, lower than 96 % of clips in this corpus. That is a total fall of 0.50.

It takes 3 clips to get there. Clip to clip the moves are -0.25, then -0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 26 s · es · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.829 and the worst against the first clip 0.862; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.6 dB
per-clip
per-clip buttons play:
rule VN1k 3qmax 0.497cmax 0.249d_a -0.497d_b -0.497dataset podcastlang estotal 26.1schain gain +6.3 dBseam step 0.6 dBcrossfades 100/150 msmin_cos_consec 0.8286min_cos_anchor 0.8621
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, average recording, quiet background, normally alert, fairly steady, somewhat unclear, moderate pitch range
(affection, embarrassment, hope enthusiasm optimism · measured, neutral tension, frequent disfluency, casual) contenido. Sí. (low mumble) A estos tipos de personajes deberías empezar a apuntar. O sea, a apuntar de quiero ser como tal o quiero hacer como tal. O bueno, en el tema gaming, por ejemplo, está el Capone.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as affection, embarrassment, hope enthusiasm optimism; style: casual, monologue; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 8.4/10; 14.1s, ES.
555590_00169912 · in -26.4 dBFS · gain +6.3 dB · podcast-05384
(jealousy and envy, intoxication altered states of consciousness, elation · normal-paced, relaxed, frequent disfluency, casual) O Dead, Juan Guarnizo. Todos estos son muy buenos para ir adaptando su contenido a lo que tú haces.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as jealousy and envy, intoxication altered states of consciousness, elation; style: casual, conversational; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 5.6/10; 8.2s, ES.
555590_00171536 · in -26.1 dBFS · gain +6.3 dB · podcast-05387
(doubt · normal-paced, fully relaxed, some disfluency, casual) si si el apoyo cualquier creador de contenido es muy bueno la verdad
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as doubt; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.8/6; vocal-burst blend 6.8/10; 4.1s, ES.
555590_00177680 · in -26.6 dBFS · gain +6.3 dB · podcast-05390
FOCS — vocal focusidentity ≥0.80   strict_047 · #15

This is a VoiceNet dimension, not an emotion: vocal focus (FOCS) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with vocal focus (FOCS) above average — 0.63, higher than 63 % of clips in this corpus — and works its way down to low at 0.15, lower than 85 % of clips in this corpus. That is a total fall of 0.48.

It takes 3 clips to get there. Clip to clip the moves are -0.23, then -0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 35 s · da · eurospeech

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.890 and the worst against the first clip 0.884; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150/150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 1.2 dB
per-clip
per-clip buttons play:
rule VN1k 3qmax 0.480cmax 0.248d_a -0.480d_b -0.480dataset eurospeechlang datotal 35.3schain gain +4.5 dBseam step 1.2 dBcrossfades 150/150 msmin_cos_consec 0.8901min_cos_anchor 0.8840
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a middle-aged somewhat feminine voice · slightly cool, neutral-bright, fairly smooth, average recording, quiet background, measured, moderately variable, frequent disfluency
(thankfulness gratitude, relief, pride · normally alert, neutral tension, wide pitch range, whispered) (low mumble) Jeg er ikke i tvivl om, at det var, fordi der var noget at kæmpe for, at de har ment så stærkt, at økonomisk uafhængighed var meget, meget
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as thankfulness gratitude, relief, pride; style: whispered, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 1.6/10; 10.8s, DA.
denmark_20111M062_2012-03-27_1300_14513344_14524096 · in -23.5 dBFS · gain +4.5 dB · eurospeech-00150
(pride, thankfulness gratitude, shame · normally alert, neutral tension, wide pitch range) (low mumble) Jeg vil også godt, sådan lidt inspireret af ordføreren, takke rødstrømperne for den kamp, de kæmpede i sin tid, og som ligesom har været fundamentet for, hvor langt vi er kommet i dag.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, audible breath; affect is mildly positive, neutral stance, neutral openness; reads as pride, thankfulness gratitude, shame; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.4/10; 13.8s, DA.
denmark_20111M062_2012-03-27_1300_14524096_14537936 · in -24.5 dBFS · gain +4.5 dB · eurospeech-00150
(sadness, shame · subdued, slightly relaxed, fairly narrow pitch, monologue) med den her debat om, om det virkelig kan være rigtigt, at vi som politikere skal vide bedre, hvad der er bedst for hvem,
full caption & clip details
A middle-aged somewhat feminine voice; delivery is subdued, measured, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as sadness, shame; style: monologue, authoritative; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.1/10; 11.0s, DA.
denmark_20111M062_2012-03-27_1300_14537936_14548944 · in -25.8 dBFS · gain +4.5 dB · eurospeech-00150
S_AUTH — style: authoritativeidentity ≥0.80   strict_040 · #16

This is a VoiceNet dimension, not an emotion: style: authoritative (S_AUTH) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with style: authoritative (S_AUTH) around average — 0.47, lower than 53 % of clips in this corpus — and works its way down to low at 0.22, lower than 78 % of clips in this corpus. That is a total fall of 0.25.

It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 26 s · en · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.893 and the worst against the first clip 0.893; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 100 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.4 dB
per-clip
per-clip buttons play:
rule VN1k 2qmax 0.250cmax 0.250d_a -0.250d_b -0.250dataset podcastlang entotal 25.6schain gain +12.2 dBseam step 0.4 dBcrossfades 100 msmin_cos_consec 0.8934min_cos_anchor 0.8934
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, moderately variable
(infatuation, relief · slightly relaxed, some disfluency, casual, conversational) Right. And that's what my friends were saying was like, hey, we're gonna get to the point where that is the most important. Like, you can't play if you don't have that stat.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as infatuation, relief; style: casual, conversational; good recording, quiet background; genuineness 4.0/6; vocal-burst blend 4.8/10; 8.6s, EN.
925276_00356104 · in -32.5 dBFS · gain +12.2 dB · podcast-01049
(intoxication altered states of consciousness, pride, embarrassment · relaxed, frequent disfluency, casual, conversational) I mean, I'm hitting I'm hitting everything. I have max level Whatever the fuck my you know arrows, right? And then I have seven point six percent, you know, more. But yeah, I I guess I see what you guys are saying. Is maybe when we start to play the elite dungeons,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as intoxication altered states of consciousness, pride, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 5.4/6; vocal-burst blend 8.4/10; 17.1s, EN.
925276_00358208 · in -32.1 dBFS · gain +12.2 dB · podcast-05718
R_THRT — resonance: throattimbre ≥0.80   strict_041 · #17

This is a VoiceNet dimension, not an emotion: resonance: throat (R_THRT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with resonance: throat (R_THRT) low — 0.25, lower than 75 % of clips in this corpus — and ends with it around average at 0.50, right about the corpus median. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

The largest step is 0.25, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 20 s · en · emolia

Identity. measured with Orange/Speaker-wavLM-tbr, a timbre embedding. Store-to-store on paired chains -tbr ≥ 0.788 is the equivalent of -id ≥ 0.80, so the 0.80 applied here is fractionally stricter than the calibrated point. The worst neighbour-to-neighbour similarity on this chain is 0.951 and the worst against the first clip 0.951; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.7 dB
per-clip
per-clip buttons play:
rule VN1k 2qmax 0.250cmax 0.250d_a 0.250d_b 0.250dataset emolialang entotal 20.5schain gain -2.0 dBseam step 0.7 dBcrossfades 150 msmin_cos_consec 0.9507min_cos_anchor 0.9507
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, average recording, quiet background, normally alert, neutral tension, some disfluency
(contemplation, concentration, doubt · brisk, fairly steady, conversational, casual) These are two different things, and I think we also have to understand the complexity. So I think there are two fundamental dilemmas that we have to come to terms with if we think about reforming the system.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as contemplation, concentration, doubt; style: conversational, casual; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 6.6/10; 8.3s, EN.
EN_a56AwJ5hgfc_W000223 · in -18.5 dBFS · gain -2.0 dB · emolia-00990
(contemplation, concentration, sourness · normal-paced, moderately variable, didactic, casual) The first is what I would call the decentralization dilemma. And I think that's very prevalent when you think again about the United States case where we want to have small banks, no big national banks, (ahem) uh, democracy, et cetera.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contemplation, concentration, sourness; style: didactic, casual; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 2.6/10; 12.3s, EN.
EN_a56AwJ5hgfc_W000224 · in -17.8 dBFS · gain -2.0 dB · emolia-00990
R_THRT — resonance: throatidentity ≥0.80   strict_044 · #18

This is a VoiceNet dimension, not an emotion: resonance: throat (R_THRT) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with resonance: throat (R_THRT) around average — 0.50, right about the corpus median — and works its way down to low at 0.25, lower than 75 % of clips in this corpus. That is a total fall of 0.25.

It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.

The largest step is 0.25, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 38 s · en · podcast

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.932 and the worst against the first clip 0.932; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.3 dB
per-clip
per-clip buttons play:
rule VN1k 2qmax 0.250cmax 0.250d_a -0.250d_b -0.250dataset podcastlang entotal 38.5schain gain -0.9 dBseam step 0.3 dBcrossfades 150 msmin_cos_consec 0.9324min_cos_anchor 0.9324
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, energised, moderately variable, some disfluency
(bitterness, interest, astonishment surprise · brisk, slightly relaxed, casual, dramatic) because they abandoned what was the world that they were going to use and started to divert and then kind of changed their mind. Anyway, there are a couple of of of things here. The the news was it's actually starting. So Andy Muschetti, who did It and It Chapter Two and did Mother and has been making his name in the horror genre, I love his direct role style.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as bitterness, interest, astonishment surprise; style: casual, dramatic; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 4.0/10; 21.1s, EN.
294291_00199544 · in -19.2 dBFS · gain -0.9 dB · podcast-06064
(pleasure ecstasy, affection, infatuation · normal-paced, neutral tension, casual, conversational) (low mumble) Um, I I really do. He's he's fantastic. I'm all in. (low mumble) Um, there were questions of Ezra Miller. He was like, look, if you're not gonna put a lot of heart in here and give me a lot to play with as an actor, and it's just zoom zoom, like forget it. I don't want to do this. And then can you let me do the character as I was making him be? All those things. So now with
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as pleasure ecstasy, affection, infatuation; style: casual, conversational; good recording, quiet background; genuineness 4.6/6; vocal-burst blend 10.0/10; 17.6s, EN.
294291_00201648 · in -18.9 dBFS · gain -0.9 dB · podcast-06053
S_STRY — style: storytellingidentity ≥0.80   strict_042 · #19

This is a VoiceNet dimension, not an emotion: style: storytelling (S_STRY) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with style: storytelling (S_STRY) at the very bottom of the range — 0.02, lower than 98 % of clips in this corpus — and ends with it below average at 0.27, lower than 73 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 35 s · lt · eurospeech

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.948 and the worst against the first clip 0.948; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.2 dB
per-clip
per-clip buttons play:
rule VN1k 2qmax 0.250cmax 0.250d_a 0.250d_b 0.250dataset eurospeechlang lttotal 35.2schain gain -3.5 dBseam step 0.2 dBcrossfades 150 msmin_cos_consec 0.9480min_cos_anchor 0.9480
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · slightly cool, neutral-bright, slightly rough, some background noise, brisk, energised, neutral tension, moderately variable
(anger, concentration, malevolence malice · light breath, authoritative, monologue) Taip pat labai svarbus projektas - tai Konkurencijos įstatymo projektas ir kartu visi lydintieji įstatymai. Šiame sektoriuje mes dar turime labai daug rimtų problemų ir artėjant prie Europos Sąjungos reikalavimų, ir mūsų pačių
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as anger, concentration, malevolence malice; style: authoritative, monologue; average recording, some background noise; genuineness 2.9/6; vocal-burst blend 3.2/10; 16.1s, LT.
lithuania_lithuania_11_15091998_5302224_5318320 · in -16.6 dBFS · gain -3.5 dB · eurospeech-01839
(triumph, elation, hope enthusiasm optimism · audible breath, cartoonish, dramatic) dar kartą grįšime prie klausimo, kaip veikia eksporto skatinimo programa bei įmonių gaivinimo programa, apie kurią buvo tiek daug kalbama, tiek daug (low mumble) vilčių į jas dedama. Tačiau susitikimuose su verslo atstovais aišku, kad, (ahem) jų nuomone, šios
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; somewhat unclear, some disfluency, wide pitch range, audible breath; affect is positive, slightly dominant, slightly guarded; reads as triumph, elation, hope enthusiasm optimism; style: cartoonish, dramatic; below-average recording, some background noise; genuineness 3.4/6; vocal-burst blend 8.2/10; 19.2s, LT.
lithuania_lithuania_11_15091998_5351696_5370943 · in -16.4 dBFS · gain -3.5 dB · eurospeech-01839
CHNK — chunking / phrasing densityidentity ≥0.80   strict_043 · #20

This is a VoiceNet dimension, not an emotion: chunking / phrasing density (CHNK) describes the voice or the recording itself — how it sounds — rather than what the speaker feels. The rule asked it to sweep by at least 0.25.

The chain starts with chunking / phrasing density (CHNK) around average — 0.53, higher than 53 % of clips in this corpus — and ends with it high at 0.78, higher than 78 % of clips in this corpus. That is a total rise of 0.25.

It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 30 s · italian · mls

Identity. measured with Orange/Speaker-wavLM-id, the verification model the 0.80 threshold belongs to. The worst neighbour-to-neighbour similarity on this chain is 0.912 and the worst against the first clip 0.912; both clear 0.80, which is the whole reason it is here and why no voice conversion was applied. The joins are 150 ms equal-power crossfades, and the chain is normalised as one signal.

joint
the same chain with each clip levelled to −20 dBFS separately (what the earlier grids did) — source level steps here reach 0.1 dB
per-clip
per-clip buttons play:
rule VN1k 2qmax 0.250cmax 0.250d_a 0.250d_b 0.250dataset mlslang ittotal 29.6schain gain +6.5 dBseam step 0.1 dBcrossfades 150 msmin_cos_consec 0.9124min_cos_anchor 0.9124
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(thankfulness gratitude, relief, shame · almost no disfluency, authoritative, monologue) ebbene un bel giorno si stancarono perdettero la pazienza alla fine chi sa da quanto tempo frenavano dentro le smanie della loro speranza frustrata di continuo e reprimevano i segni delle loro disillusioni
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, relief, shame; style: authoritative, monologue; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 2.6/10; 14.7s, ITALIAN.
7440_7720_000034 · in -26.5 dBFS · gain +6.5 dB · mls-00079
(shame, infatuation, pain · some disfluency, authoritative, monologue) il primo segno ch'io potei scorgere e che m'è rimasto impresso come in un dramma una frase che lasci intravedere la catastrofe fu quella mattina che dovevamo recarci alla vigna di ponte molle e giorgina si presentò al tranzi col capo chino
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, infatuation, pain; style: authoritative, monologue; good recording, quiet background; genuineness 1.7/6; vocal-burst blend 1.8/10; 15.0s, ITALIAN.
7440_7720_000012 · in -26.6 dBFS · gain +6.5 dB · mls-00079