NZBT
Non-zero boundary transitions · an unfinished audio experiment
NZBT is a proposal for making transitions part of the audio runtime. Computers can change a parameter, replace an algorithm, reassign a voice or alter a graph at a single sample boundary. The sound still has to get through that change. Usually the developer has to write the smoothing, fades and state handling into each instrument or effect. The question is whether more of that work could belong underneath, in the system running them.
Non-zero boundary transitions means giving the audible change a duration, even when the computational decision is discrete. It’s broader than “crossfade everything”. A continuous parameter might need a trajectory; an algorithm swap might need overlapping instances; a stateful process might need state transfer or preparation. The runtime would choose a suitable way through, according to what the process exposes and what the change requires.
I think that’s useful because it could make changing an instrument while it’s running less of a special case every developer has to solve again. You could explore its settings and structure while sound is happening, with reusable transition handling underneath. The aim is to reduce that burden, while leaving deliberate attacks, cuts and other musical transients alone.
The research challenged the idea as a universal guarantee for arbitrary black boxes. That led to the discussion about borrowing time: if immediate execution is the restriction, allow a bounded buffer and work towards what people can hear and feel. The later signal, copy and sample proposal adds three coordinated fallback paths, selected and blended mathematically around a change. It extends the runtime idea; it isn’t the whole thesis or a proof that hidden state never matters.
I’m a slopster. AI agents made the research and code, and I’ve argued with their conclusions along the way. The research notes and original documents are here, including the claims we challenged. The filter comparison below is one small working example, not the entire proposal and not a demonstration that it works for every process.
The numbers measure a chosen waveform target, not how good it sounds. Have a listen and make your own mind up.
One small demonstration
A driven 150 Hz low-pass resonant filter transitions to a 350 Hz filter. Signal continues the original filter; copy runs the replacement using a limited recorded input history; sample repeats the last 20 ms of the old output. All comparisons use the same 80 ms transition and the same copy pre-roll budget.
Solid: three paths · dashed: two paths · dotted: linear crossfade
Solid: signal · dashed: copy · dotted: sample
Numbers behind this run
One change at 1.50 seconds. Each button restarts the same tone and timing.
Ready to listen.
Playback is one 3.38-second sustained tone at 16 kHz, with a single 80 ms transition at 1.50 seconds. Short fades soften the start and end of each clip. A shared gain is calculated across all four methods to leave headroom; the methods are not individually normalized. Cold jump starts the replacement at the switch without pre-roll. This page calculates offline, not on an audio render thread.
The maths, if you want it
At each sample, choose nonnegative weights summing to one. Fix the first weights to signal alone and the last to copy alone. The reference r is a quintic transition between the available signal and copy outputs. It is an explicit design target, not an inaccessible ideal history or a validated model of hearing.
y[k] = Σᵢ wᵢ[k] pᵢ[k]
r[k] = (1 − s[k]) signal[k] + s[k] copy[k]
s = 10u³ − 15u⁴ + 6u⁵, u = k/N
J = mean((y − r)²) / P + 8 mean((Δy − Δr)²) / P
+ λ Σₖ ||Δw[k]||²
P = max(mean(r²), 10⁻⁶)For fixed path signals this is a convex quadratic objective on simplex constraints. Projected gradient descent with backtracking evaluates the objective and its gradient at every iteration. The code reports its iteration cap and projected-gradient residual; it does not claim convergence merely because it stopped. The sample weight may be zero. The three-path solver starts from the computed two-path solution and uses additional iterations, so this is not an equal-compute benchmark. The two-path optimizer uses the same objective and endpoints with sample disabled.
Amplitude and slope tracking are engineering proxies. The choice of target favors the quintic transition; a lower cost demonstrates improvement on this chosen objective, not perceptual superiority. A future controller would need an application-specific perceptual criterion and path feasibility rules.
What the research actually establishes
For a linear process driven by the same input during pre-roll, the initial-state difference evolves as ε(t₀) = exp(AH) ε(t₀ − H). This establishes remaining state mismatch when the initial mismatch is nonzero. It does not establish that the mixed output is discontinuous, that the difference is audible, or that a sample bridge cannot conceal it.
A memory time constant is not an exact finite memory cutoff. Finite pre-roll generally reduces a stable IIR filter’s state mismatch rather than making it exactly zero at H equal to a time constant. For the scalar decay case, ε(H) = ε(0) exp(−H/τ): there is no special impossibility boundary at H = τ. A zero initial mismatch also remains zero.
Smooth weights alone do not guarantee waveform or slope continuity. Path regularity, endpoint value and slope matching, and the way paths are activated also matter. Crossfade cancellation depends on the actual signals and weights; internal phase access is not universally required to inspect output phase relationships. Compute cost must be measured, not assumed to equal three times the original process.
Exact reproduction of an unknown alternate history remains distinct from delivering an acceptable transition. This demo supports neither a universal impossibility verdict nor a universal success verdict.
Limits and where this came from
The diagnostic history reference runs the replacement for one second plus 1.50 seconds before the switch using the same deterministic input. It is a finite longer-history benchmark, not an infinite-past ground truth. Its output discrepancy is measured separately and is never fed to the optimizer.
H here means accessible recorded history replayed into a new instance. The page does not implement a live output queue, emulate unknown future player actions, or measure end-to-end latency. Real replay requires sufficient compute before its deadline. Output buffering, recorded history length, transition duration and interaction delay are separate budgets.
The research and this implementation were made with AI agents. The first reports confused leftover state error with unavoidable audible failure, and described fixed blends as optimization. Those claims were challenged and the demo was rewritten. This version gives the controller an explicit objective; it is still a small numerical example, not a listening study or a finished audio framework.