Strict subset — hard identity cut, crossfaded, no voice conversion
126 trajectories rendered as a sample · speaker similarity ≥ 0.80 on both conditions · 150 ms equal-power crossfades · chain-level normalisation · nothing here has been voice-converted.
The proposal, in one paragraph. The
voice-conversion pass measured well on emotion but badly on voice: it helps only the
chains that were already beyond rescue, and on chains that clear a 0.80 identity cut it
actively costs similarity (−0.117, improving only 2 % of them). So
the recommendation is to stop converting and start filtering: take the chains that
are already one speaker, join them with a short crossfade, and ship those. This page is
that set. Under the strict filter there are
3,378,525 such chains in the mined corpus at
T≥0.20 and 1,399,128 at
T≥0.25, plus 949,212 voice-profile
chains that need no speaker check at all — so the strict path is not a scarcity
problem.
How many trajectories survive the strict criteria
The filter. A chain is counted only if
min_cos_consec ≥ 0.80 and min_cos_anchor ≥ 0.80.
Neighbour-only similarity does not chain — A can resemble B and B resemble C while
A and C are plainly different people — so the anchored condition is the one that
catches drift, and requiring both is the conservative reading. Chains with no
similarity measurement at all are excluded, not assumed good. Across the whole mined
set that leaves 4,130,756 of 9,744,940 chains
(42.4 %). The per-step cap is
C = 0.25 throughout, and membership at a given T is
tested as qmax ≥ T and cmax ≤ C under each row's own rule — the
only predicate that is exact for every rule.
| rule | what it requires | T ≥ 0.20 | T ≥ 0.25 | T ≥ 0.40 | T ≥ 0.50 |
|---|
Rule A — two-sided
AB2 | both the start emotion and the target emotion must move by at least T | 619,202 | 167,304 | 5,434 | 437 |
Rule B — one-sided
B1 | only the target emotion must be reached; the start axis is unconstrained | 1,097,321 | 457,637 | 87,263 | 20,934 |
Proxy (PXR)
PXR | the same two-sided test as Rule A, but the per-step smoothness cap is certified on a proxy axis for emotions that cannot ramp on their own | 629,944 | 173,843 | 5,637 | 446 |
VoiceNet (VN1)
VN1 | one of 57 VoiceNet voice-descriptor dimensions sweeps by at least T; no proxies | 1,032,058 | 600,344 | 162,567 | 26,087 |
| total, mined corpora | | 3,378,525 | 1,399,128 | 260,901 | 47,904 |
voice profiles vprof_vc | one cloned voice per chain — no speaker filter applies | 1,307,695 | 949,212 | 676,132 | 428,294 |
By chain length
A k=2 chain cannot exist above
T=0.25 while the per-step cap is 0.25: with two clips the single
step is the whole move, so cmax == qmax and nothing above the cap
can qualify. That is a property of the rule, not a shortage of data.
| rule | T | k=2 | k=3 | k=4 | k=5 |
|---|
AB2 | 0.20 | 203,215 | 239,182 | 120,154 | 56,651 |
AB2 | 0.25 | 0 | 90,045 | 51,727 | 25,532 |
AB2 | 0.40 | 0 | 1,182 | 2,601 | 1,651 |
B1 | 0.20 | 358,221 | 282,975 | 240,789 | 215,336 |
B1 | 0.25 | 0 | 166,115 | 152,279 | 139,243 |
B1 | 0.40 | 0 | 18,662 | 33,609 | 34,992 |
PXR | 0.20 | 203,481 | 244,995 | 123,318 | 58,150 |
PXR | 0.25 | 0 | 93,802 | 53,644 | 26,397 |
PXR | 0.40 | 0 | 1,234 | 2,697 | 1,706 |
VN1 | 0.20 | 352,805 | 267,905 | 220,972 | 190,376 |
VN1 | 0.25 | 0 | 242,810 | 194,738 | 162,796 |
VN1 | 0.40 | 0 | 53,221 | 61,096 | 48,250 |
Which embedding backs each chain
The two similarity columns are different
populations and must not be compared with each other. -id covers podcast
and parts of eurospeech/mls/snippets (645,764 chains
pass the strict filter); -tbr covers emolia
(2,732,761). A higher pass-rate on one is mostly a
different denominator and a different corpus mix, not a more lenient model.
Where -id exists it is preferred: it is the verification model and
0.80 is its threshold. Where only -tbr exists the same 0.80 is applied,
which is fractionally stricter than the calibrated equivalent —
store-to-store on paired chains, -tbr ≥ 0.788 matches
-id ≥ 0.80 at 85.9 % agreement. Every card is labelled
with which one backs it.
How these were joined
Crossfades. Every join is an equal-power
crossfade (cos/sin, so the two sides sum to constant energy; a
linear crossfade dips audibly at its midpoint on uncorrelated material). Nominal
150 ms, shortened to 100 ms where the incoming clip starts loud or the outgoing
one ends loud, and never more than a quarter of either clip. Across
325 joins: 221 at the full 150 ms,
86 shortened for a hot onset, 15 for a hot
tail, 3 by the duration cap. Shortest anywhere:
80 ms. This replaces the silent gap the earlier grids
used.
Normalisation — and an honest caveat.
The default render normalises each chain as one signal, so the loudness
relationships between clips survive. On voice-converted audio that was clearly right,
because SIDON returns every segment at a consistent level. On these original
recordings it is less clean: the MOSS encoder levelled every clip independently and
59.5 % of clips hit its ±3 dB clamp, so some of the level differences
preserved here are a pipeline artifact rather than performance.
Measured on this set: the largest step between adjacent clips is a median
1.80 dB, 5.14 dB at the
95th percentile, 53.14 dB at worst. Every card also
carries the per-clip-levelled version one click away, so you can hear which you
prefer. For a production corpus I would level per clip and keep joint
normalisation for converted audio only.
A quality gate was applied on top of the rule.
Two selected chains were dropped after rendering: one voice-profile chain containing a
clip at −76.9 dBFS (silent — a broken generation) and one podcast chain
whose clips all sat near −46 dBFS. The gate is: every clip at or above
−40 dBFS and a source level spread of at most 12 dB. “Rather a bit
too conservative” should include not shipping inaudible audio.
The sample
126 chains, chosen to span the conditions rather than
to be the 126 best: every rule, every chain length k=2…5, both the mined corpora
and the voice profiles, and within each cell the chains with the largest
qualifying move first. Rules: AB2 31, B1 31, PXR 32, VN1 32.
Lengths: k=2 32, k=3 32, k=4 30, k=5 32.
Corpora: emolia 40, eurospeech 9, mls 4, podcast 19, snippets 7, vprof_vc 47.