Strict subset — hard identity cut, crossfaded, no voice conversion

126 trajectories rendered as a sample · speaker similarity ≥ 0.80 on both conditions · 150 ms equal-power crossfades · chain-level normalisation · nothing here has been voice-converted.

The proposal, in one paragraph. The voice-conversion pass measured well on emotion but badly on voice: it helps only the chains that were already beyond rescue, and on chains that clear a 0.80 identity cut it actively costs similarity (−0.117, improving only 2 % of them). So the recommendation is to stop converting and start filtering: take the chains that are already one speaker, join them with a short crossfade, and ship those. This page is that set. Under the strict filter there are 3,378,525 such chains in the mined corpus at T≥0.20 and 1,399,128 at T≥0.25, plus 949,212 voice-profile chains that need no speaker check at all — so the strict path is not a scarcity problem.

How many trajectories survive the strict criteria

The filter. A chain is counted only if min_cos_consec ≥ 0.80 and min_cos_anchor ≥ 0.80. Neighbour-only similarity does not chain — A can resemble B and B resemble C while A and C are plainly different people — so the anchored condition is the one that catches drift, and requiring both is the conservative reading. Chains with no similarity measurement at all are excluded, not assumed good. Across the whole mined set that leaves 4,130,756 of 9,744,940 chains (42.4 %). The per-step cap is C = 0.25 throughout, and membership at a given T is tested as qmax ≥ T and cmax ≤ C under each row's own rule — the only predicate that is exact for every rule.
rulewhat it requiresT ≥ 0.20T ≥ 0.25T ≥ 0.40T ≥ 0.50
Rule A — two-sided
AB2
both the start emotion and the target emotion must move by at least T619,202167,3045,434437
Rule B — one-sided
B1
only the target emotion must be reached; the start axis is unconstrained1,097,321457,63787,26320,934
Proxy (PXR)
PXR
the same two-sided test as Rule A, but the per-step smoothness cap is certified on a proxy axis for emotions that cannot ramp on their own629,944173,8435,637446
VoiceNet (VN1)
VN1
one of 57 VoiceNet voice-descriptor dimensions sweeps by at least T; no proxies1,032,058600,344162,56726,087
total, mined corpora3,378,5251,399,128260,90147,904
voice profiles vprof_vcone cloned voice per chain — no speaker filter applies1,307,695949,212676,132428,294

By chain length

A k=2 chain cannot exist above T=0.25 while the per-step cap is 0.25: with two clips the single step is the whole move, so cmax == qmax and nothing above the cap can qualify. That is a property of the rule, not a shortage of data.
ruleTk=2k=3k=4k=5
AB20.20203,215239,182120,15456,651
AB20.25090,04551,72725,532
AB20.4001,1822,6011,651
B10.20358,221282,975240,789215,336
B10.250166,115152,279139,243
B10.40018,66233,60934,992
PXR0.20203,481244,995123,31858,150
PXR0.25093,80253,64426,397
PXR0.4001,2342,6971,706
VN10.20352,805267,905220,972190,376
VN10.250242,810194,738162,796
VN10.40053,22161,09648,250

Which embedding backs each chain

The two similarity columns are different populations and must not be compared with each other. -id covers podcast and parts of eurospeech/mls/snippets (645,764 chains pass the strict filter); -tbr covers emolia (2,732,761). A higher pass-rate on one is mostly a different denominator and a different corpus mix, not a more lenient model.

Where -id exists it is preferred: it is the verification model and 0.80 is its threshold. Where only -tbr exists the same 0.80 is applied, which is fractionally stricter than the calibrated equivalent — store-to-store on paired chains, -tbr ≥ 0.788 matches -id ≥ 0.80 at 85.9 % agreement. Every card is labelled with which one backs it.

How these were joined

Crossfades. Every join is an equal-power crossfade (cos/sin, so the two sides sum to constant energy; a linear crossfade dips audibly at its midpoint on uncorrelated material). Nominal 150 ms, shortened to 100 ms where the incoming clip starts loud or the outgoing one ends loud, and never more than a quarter of either clip. Across 325 joins: 221 at the full 150 ms, 86 shortened for a hot onset, 15 for a hot tail, 3 by the duration cap. Shortest anywhere: 80 ms. This replaces the silent gap the earlier grids used.
Normalisation — and an honest caveat. The default render normalises each chain as one signal, so the loudness relationships between clips survive. On voice-converted audio that was clearly right, because SIDON returns every segment at a consistent level. On these original recordings it is less clean: the MOSS encoder levelled every clip independently and 59.5 % of clips hit its ±3 dB clamp, so some of the level differences preserved here are a pipeline artifact rather than performance.

Measured on this set: the largest step between adjacent clips is a median 1.80 dB, 5.14 dB at the 95th percentile, 53.14 dB at worst. Every card also carries the per-clip-levelled version one click away, so you can hear which you prefer. For a production corpus I would level per clip and keep joint normalisation for converted audio only.
A quality gate was applied on top of the rule. Two selected chains were dropped after rendering: one voice-profile chain containing a clip at −76.9 dBFS (silent — a broken generation) and one podcast chain whose clips all sat near −46 dBFS. The gate is: every clip at or above −40 dBFS and a source level spread of at most 12 dB. “Rather a bit too conservative” should include not shipping inaudible audio.

The sample

126 chains, chosen to span the conditions rather than to be the 126 best: every rule, every chain length k=2…5, both the mined corpora and the voice profiles, and within each cell the chains with the largest qualifying move first. Rules: AB2 31, B1 31, PXR 32, VN1 32. Lengths: k=2 32, k=3 32, k=4 30, k=5 32. Corpora: emolia 40, eurospeech 9, mls 4, podcast 19, snippets 7, vprof_vc 47.
pagewhat is on itchains
Rule A — two-sidedboth the start emotion and the target emotion must move by at least T20
Rule B — one-sidedonly the target emotion must be reached; the start axis is unconstrained19
Proxy (PXR)the same two-sided test as Rule A, but the per-step smoothness cap is certified on a proxy axis for emotions that cannot ramp on their own20
VoiceNet (VN1)one of 57 VoiceNet voice-descriptor dimensions sweeps by at least T; no proxies20
Voice profiles — one cloned voice per chainno speaker filter needed47