The Deep Vocoder (private/src/DeepVocoder.hpp) is an experimental spectral processing system. The name “Deep” is a reference to the Deep Listening practices of Pauline Oliveros.
While a traditional vocoder imposes the amplitude envelope of an analysis signal (the modulator) onto a set of fixed bandpass filters acting on a synthesis signal (the carrier), the Deep Vocoder operates entirely differently. It analyzes an incoming audio signal using a Spectral Modeling Synthesis (SMS) approach and uses the extracted partials to “quantize” the fundamental pitches generated by the Nonagon sequencer.
The Deep Vocoder maintains a rolling buffer of incoming audio. Every hop (x_H samples), it performs a spectral analysis (m_spectralModel.ExtractAtoms).
Deep Vocoder uses the shared spectral model’s ordered one-to-one matcher with an explicit density radius of one semitone (1/12 octave) on each side of an old atom’s analysis frequency. Nearby pitch movement can continue the existing atom; peaks outside that window start new atoms while old ones decay. This keeps the new tracking model while giving the vocoder a wider continuation window than the shared one-cent default.
When a voice in the synthesizer is triggered by the Nonagon, its intended fundamental pitch (m_pitchCenter) is passed to the Deep Vocoder (TransformNote).
TransformNote receives the current input and refreshes that voice’s pitch,
ratios, and threshold parameters immediately. FFT-hop caching controls spectral
analysis, not which note is transformed. When disabled, the vocoder preserves
the trigger and passes the current pitch through with its post ratio; it never
substitutes the preceding note’s cached pitch.
The vocoder then searches the currently active spectral atoms to find the “best” match. It does this using a V-shaped thresholding function (MagnitudeThreshold):
m_gainThreshold).m_slopeUp) and below the target (m_slopeDown).The vocoder evaluates all current atoms:
magnitude / threshold.If no spectral atom meets the threshold criteria (e.g., if the input signal is silent, or if there are no partials remotely near the target pitch), the Deep Vocoder cancels the trigger (ahdControl->m_trig = false). The voice will not sound.
If a winning atom is found, the voice’s pitch is snapped (quantized) to the exact frequency of that partial, and the trigger proceeds normally.
This allows an external audio source to act as a dynamic, harmonically rich, real-time quantizer and gate for the sequenced melodies.