loading

NZBT

A question, some research, an offering

I asked a question about how sound gets through a change, and ended up with this. There’s a runtime idea, some AI research I’ve argued with, and a small working experiment. I’m a slopster. I’m putting the useful bits here rather than making you read a report to find out what I’m offering.

If there’s something in it for you, have it. Build on it, pull it apart, or use it to ask a better question. I’ve left the original documents at the bottom for anyone who wants the detail.

I. The question

When an audio process changes, the computer can make the decision at one sample boundary. The sound has to get from what it was doing to what it’s doing next. Parameter smoothing, fades, state preparation and handoffs usually become things each developer has to build into their own instrument or effect.

Could more of that responsibility live in the runtime underneath? That’s the NZBT thesis: discrete computational changes given a non-zero interval in which to become audible. Parameters, algorithm swaps, voice reassignment, graph routing and model changes all belong to the question.

I think it’s useful because changing a running instrument shouldn’t always mean rebuilding the same transition machinery. A shared system could make more of those changes available while leaving the person using it free to explore.

II. A way through the change

This isn’t just “crossfade everything”. A parameter might have a meaningful path between values. Two different algorithms might need to run alongside each other. A process with memory might benefit from transferring its state, or being given recorded input to prepare it.

The original research considers an external wrapper, an interface for state transfer, and a pipeline that renders ahead. The runtime would need to know what a process permits and what the change requires. Part of the question is whether that contract genuinely reduces developers’ work, or just moves it somewhere else.

There are different things to preserve: waveform value, slope, phase, spectral behaviour, stored energy and the musical intention. They aren’t interchangeable. A drum attack or deliberate cut shouldn’t be softened just because a system has decided that every boundary is a problem.

III. Signal copy and sample

The later proposal adds three coordinated fallback paths. Signal continues or transforms what’s already available. Copy prepares the new behaviour in a parallel process. Sample uses captured audio to bridge a gap while another path becomes usable.

The maths chooses where each path is useful and how to blend them around a change. Feedforward can prepare for the requested change; feedback can adjust what comes out. You don’t have to use all three, and running a copy doesn’t automatically mean you can clone its hidden state.

Borrowing time is part of this. A bounded output buffer can give the system room to prepare before the listener hears the result. Recorded history, processing time, transition duration and delay felt by a player are different budgets. The aim is an acceptable audible result within those constraints, not necessarily an exact reconstruction of every possible internal history.

A little of the mathematical shape
y(t) = wₛ(t)yₛ(t) + w꜀(t)y꜀(t) + wₐ(t)yₐ(t)
wᵢ(t) ≥ 0; Σᵢ wᵢ(t) = 1

The candidate outputs must be aligned in time. Choose feasible weights over a bounded horizon using a defined error criterion, with penalties for abrupt weight changes and constraints on output level and processing deadlines. The browser demo uses amplitude and slope tracking; it isn’t a validated model of hearing.

IV. Things it sits beside

The research places the idea alongside control theory’s bumpless transfer, parameter dezippering, declarative graph replacement, time-varying filters and streaming neural audio. The offering is a way of bringing transition handling together, not a claim that smoothing or crossfading was invented here.

For a few starting points, the documents point to SuperCollider’s NodeProxy, scheduled automation in Web Audio, Wishnick’s time-varying filter work, and streamable neural audio synthesis. The full source lists and terminology maps are in the original files.

V. What’s still a question

The reports pushed back on a universal guarantee for arbitrary black boxes. Some of their arguments overreached: showing cancellation in one crossfade doesn’t rule out every external transition strategy; showing leftover state error doesn’t establish an unavoidable audible failure. A memory time constant also isn’t a hard cutoff where preparation suddenly succeeds or becomes impossible.

That leaves real constraints worth keeping: cancellation, feedback tails, nonlinear dynamics, startup history, CPU deadlines, interaction delay, sample bridges that sound frozen, and the difference between wanted and unwanted transients. Neural context and conditioning changes belong here too.

The proposed experiments include resonant-filter changes and feedback-delay changes, compared with raw replacement, crossfading and state-aware handling. Useful checks include added spectral energy, level dips, preserved tails, processing overruns, felt delay and developer effort. The reports’ numerical rejection thresholds are proposals, not universal perceptual laws or measured results.

VI. Have it

The audio experiment is one filter transition with an actual weight optimizer. It gives you something to hear and inspect. It doesn’t establish that the whole runtime proposal works, and a better score on its chosen target doesn’t prove it sounds better.

These are the complete, unchanged AI-generated documents, including the equations, architecture, prior-art comparisons, experiments and references. They also keep the claims we challenged. Read them as working material rather than settled findings.

There’s some slop here. If part of it is useful to what you’re doing, take that part.