Ant64 Patches — The Unit of Voice Configuration

This document defines the patch as the unit of "what makes a voice sound like a thing" on Tempest — the FireStorm chipset audio engine, the FPGA voice array described in audio_more.md. It applies only to FireStorm chipset voices. Pulse's AMY synthesis and DeMon's SID/SAM each carry their own preset conventions and are out of scope here.

The principle, in one sentence: classic synthesisers and drum machines are reference patches against the general voice topology + a small pool of opt-in extras, not bespoke engines in their own right.

This document layers on a foundation that is design intent, not implemented silicon — the voice topology in audio_more.md is the target architecture and no FireStorm HDL exists yet. Where the patch model surfaces topology questions, those questions are flagged for the HDL pass rather than papered over.

The Principle

A FireStorm voice has a fixed general topology — 3 oscillator slots (each independently selecting VA / Wavetable / FM / Granular / PhysModel / Additive), 2 filters (Huovilainen ladder + SVF), 4 envelopes, 4 per-voice LFOs + 2 global, a 64-slot modulation matrix, and a VCA. Everything that sounds like a Juno, a Prophet, a DX7, a Minimoog, a 303, or a 909 is configuration of that voice — a patch. Where a particular classic device's character depends on stateful circuit behaviour that the general voice doesn't naturally produce, the patch engages one or more entries from the extras pool — a small, curated set of opt-in character circuits any voice can switch on.

Why the 303 forced the question

The TB-303 was the example that tested the principle hardest. Its character lives in stateful circuit details that don't fit the general voice's per-note envelope + clean Huovilainen ladder:

  • The accent-RC accumulator — a 47 kΩ × 33 nF RC network whose smoothed voltage drives both the filter cutoff and the VCA on accented notes, and that doesn't fully discharge between consecutive accents. Repeated accents stack, producing the distinctive "increasingly distressed animal cry" the 303 is known for.
  • The mismatched first pole in the diode ladder — a deliberately undersized capacitor (C18 ≈ half the others) gives the first stage roughly 2× the cutoff of the rest, preventing the feedback phase from reaching 360° and converting resonance-into-self-oscillation into resonance-into-overdrive. This is the 303 sound.

Earlier drafts of the audio documentation treated the 303 as a dedicated single-voice engine occupying one slot with its own DSP pipeline (see audio_more.md). The justification was real — those two behaviours genuinely don't fit a general voice — but it created a precedent the rest of the design didn't follow. Every other "classic sound" is already a patch, so why not the 303?

Three options considered

  1. Keep the 303 dedicated. Honest about how different it is. Costs: design inconsistency, and every future device-specific quirk (909-kick saturation, snare noise-into-resonance, and so on) becomes a candidate for its own bespoke engine.
  2. Make every voice's silicon rich enough to fully subsume the 303. Add the accent-RC accumulator and mismatched-ladder mode to every voice as standard. Cost: silicon area paid by all 128 voices for behaviour only acid-bass patches use.
  3. Optional per-voice extras pool. A small curated set of opt-in character circuits any patch can engage. Voice silicon stays lean, paying for extras only on voices whose patches engage them (via a per-voice extras-enable register and the accompanying state); 303-style behaviour becomes a configurable patch extension rather than bespoke silicon.

Option 3 wins. The 303 is a configuration of a general voice + two engaged extras (accent-RC accumulator + mismatched-ladder mode). The general topology pays no permanent silicon cost for it; the extras pool pays per-extra silicon once and is shared across whichever voices are currently running patches that engage them.

The same pattern extends cleanly to future device-specific character — a 909-kick saturation circuit, a snare noise-into-resonance cross-feedback, a sawtooth-to-square waveshaper — each becomes a new entry in the pool when (and only when) a reference patch needs it.

The extras pool

Initial pool:

Extra What it does Silicon cost (rough) First needed by
Accent-RC accumulator Stateful RC smoothing across accent-flagged notes; configurable RC time, accent amount, and target (cutoff, VCA, or modmatrix sink). Voltage decays between notes but doesn't fully discharge, so consecutive accents stack. Small — one S2.30 accumulator register + one decay multiply per voice that engages it TB-303 acid bass
Mismatched-ladder mode When the Huovilainen ladder is engaged, applies an asymmetric coefficient to the first pole (≈ 2× the cutoff of the other three). Converts resonance-into-self-oscillation into resonance-into-overdrive. Trivial — one extra coefficient register and a per-voice select bit TB-303 acid bass

Honest tradeoff. Every extra is silicon area, paid forever on every fabricated FireStorm whether or not any running patch engages it. The pool stays curated, not kitchen-sink. The bar for inclusion is that a reference patch needs it and the behaviour cannot be synthesised from existing voice features. Anything achievable by clever use of the general topology, the modulation matrix, the multi-segment ENV4, or sequencer-side logic on Pulse does not become an extra.

Candidate extras — flagged for evaluation, not yet committed:

  • Output saturation stage (post-VCA tanh with configurable drive). Useful for the 303 output character, and for 808/909 kicks. The filter's tanh saturation in the resonant feedback path may already cover this; the question is whether reference patches sound right routed through filter tanh alone.
  • Sawtooth-derived square waveshaper. The 303's square is sawtooth through a single-transistor waveshaper, subtly different from a clean square. This is probably best expressed as a wavetable preset rather than an extras-pool entry.
  • Noise-into-resonance cross-feedback for 909-snare character — where noise excites the tonal layer's resonance. May be expressible via the modulation matrix rather than a dedicated extra.
  • Fixed-time portamento (the 303's slide is fixed time, not fixed rate). This most naturally lives in Pulse's sequencer / note-dispatch logic rather than in FireStorm voice silicon.

These will be promoted into the pool only if reference-patch implementation surfaces them as genuinely missing.

Implementation notes (fixed-point)

These are the FPGA-level fixed-point details for the two extras the TB-303 reference patch engages.

Accent-RC accumulator. The trickiest part to get right in fixed-point:

// Per sample tick:
accent_voltage = accent_voltage * decay_coeff;     // S1.30 multiply
if (accent_triggered) {
    accent_voltage += meg_output * accent_amount;  // add MEG contribution
}
vcf_cutoff += accent_voltage >> accent_depth_shift;
vca_level  += (accent_voltage * vca_accent_scale) >> 16;

The RC time constant maps to a decay_coeff close to 1.0 — approximately 1 - (1 / (sample_rate * RC_time)). At 48 kHz with RC ≈ 47 kΩ × 33 nF = 1.55 ms, decay_coeff ≈ 0.99999 in S1.30 fixed-point. Needs sufficient precision — use S2.30 or S1.62 (64-bit accumulator) to avoid losing the accent voltage to rounding.

Mismatched-ladder mode is modelled by giving the first one-pole stage of the Huovilainen ladder a slightly different frequency coefficient — approximately 0.5× the capacitance translates to ~2× the cutoff frequency for that stage. This is what prevents the feedback phase from reaching 360° and suppresses self-oscillation (converting it into overdrive — the 303 character).

Data Model

A patch is a layered configuration of one or more voice slots. The same data model covers single-voice patches (the common case) and multi-slot patches (stacks, splits, drum kits with one slot per drum).

Top level

  • Metadata — name, author, format version, classification (bass / lead / pad / drum-kit / fx / etc.), tags, creation timestamp, source attribution where relevant (a reference patch like "TB-303 acid bass" credits the original device)
  • Voice slot list — 1 to N entries. The chipset's voice budget is the upper bound, but a single patch rarely uses more than a handful of slots; drum kits are the typical multi-slot case
  • Layering rules — for multi-slot patches, how the slots respond to note events: key range (split), velocity range (layer), MIDI channel, drum-note mapping (one slot per MIDI note), or unison stack (every slot plays every note with configurable detune)
  • Global FX sends — patch-level send levels into BBD chorus, delay, reverb, and any other global FX bus (the FX chain is not per-voice; sends are)

Per voice slot

  • Oscillators (× 3) — paradigm per oscillator slot (VA / Wavetable / FM 6-op / Granular / PhysModel / Additive) plus paradigm-specific parameters; detune, level, pan contribution, hard-sync routing, cross-modulation routing
  • Filter block — F1 mode (ladder LP-24 / LP-18 / LP-12), F2 mode (SVF with LP / HP / BP / Notch outputs), routing (serial / parallel / F1-only / F2-only / bypass), cutoff and resonance bases, keyboard tracking
  • Envelopes (× 4) — ENV1–3 are ADSR with optional hold; ENV4 is 8-segment loopable (Waldorf-style). Per envelope: rates, levels, key tracking, velocity-to-amount, and one or more modulation destinations
  • LFOs (× 4 per-voice + 2 global) — waveform, rate (or tempo-sync division), delay, fade-in, sync mode (free / note-triggered / key-sync)
  • Modulation matrix — up to 64 (source → destination × amount) bindings. Sources include every envelope and LFO, velocity, mono and poly aftertouch, mod wheel, pitch bend, key tracking, note number, random per-note, MIDI CCs, and audio-rate oscillator output. Destinations include essentially every continuous parameter the voice exposes
  • VCA / pan — velocity curve (linear / exponential / custom), base level, stereo pan with optional per-note random spread
  • Extras engaged — list of pool entries the slot has switched on, each with its own parameter set
  • FX sends (per slot) — slot-level contributions into each global FX send

A typical per-slot config is around 1–2 kB binary uncompressed, dominated by the modmatrix (64 slots × ~16 bytes each ≈ 1 kB). A drum kit with 8 slots is roughly 8–16 kB. DBFS comfortably stores tens of thousands.

Layering and voice allocation

The voice allocator treats a single-slot patch as one voice per note and an N-slot layered patch as N voices per note. The voice budget (128 VA / 256 sample / patch-dependent FM) is consumed accordingly. A monophonic patch (303-style bass with last-note priority and slide) reserves a single slot and uses last-note priority; layered patches reserve all their slots together so they trigger in lockstep.

When a patch engages extras the allocator prefers slots whose required extras hardware is currently idle — extras are physically present on every voice in the initial pool of two (the silicon cost is tiny), so contention is unlikely. If a future extras pool is implemented as a shared resource (a small number of extras instances multiplexed across the voice array), the allocator will fall back to either stealing an older voice's extras allocation (last-note priority) or to a degraded patch path that omits the extra. Patch metadata can declare whether degraded fallback is acceptable.

File Format

Two complementary representations:

Binary — for DBFS storage and fast load

  • 16-byte header: magic "Ant64Patch", format version (uint16), patch length (uint32), CRC-32 over body
  • Per-slot blocks, each tagged with slot-type byte and length so future fields can be added without invalidating older patches
  • Modmatrix block uses a compact (source, destination, amount) triple encoding, omitting zero-amount slots
  • Extras-engaged block is a list of (extra-id, param-blob) pairs, where param-blob length is determined by extra-id
  • Reference patches in the factory library ship as binary blobs either in the FireStorm bitstream's BRAM init or as a DBFS table loaded at boot

JSON — for editing, version control, and sharing

  • One-to-one round-trip with the binary form (every binary field has a JSON key, every JSON key produces deterministic binary)
  • Pretty-printed for readability; arrays for slot lists and modmatrix
  • The AntOS workstation app's patch editor reads and writes JSON natively; binary export happens on save or transfer to another device
  • Patches under version control live as JSON; the binary form is build artefact

Honest tradeoff. Binary is compact, fast to load, DBFS-friendly, and memory-mappable. JSON is verbose, slow to parse, and unfriendly to embedded loaders but vastly more useful for editing, diffing, and code review. Carrying both means writing the serialiser twice and keeping them in lockstep — the cost is real but the tooling benefit is larger.

Lifecycle

   AntOS loads patch from DBFS  ──► JSON or binary, parsed into in-memory struct
              │
              ▼
   Patch passed to Pulse         ──► Pulse holds the "active patch table" for each
                                     allocated voice slot (one entry per slot)
              │
              ▼
   Voice slot(s) bound            ──► Pulse writes voice config to FireStorm via OPI
                                     register window: oscillator paradigm bits, filter
                                     modes, envelope rates, modmatrix routings, extras
                                     enable bits, all of it
              │
              ▼
   Note-on event                  ──► Pulse triggers the slot(s); FireStorm's voice
                                     envelopes and LFOs start running autonomously
                                     and the voice generates audio into the mix bus
              │
              ▼
   During note hold               ──► Live parameter tweaks (jog dial, MIDI CC,
                                     modmatrix-bound sources) cause Pulse register
                                     writes into the running voice's config; the
                                     voice continues without re-binding
              │
              ▼
   Note-off                       ──► Envelopes enter release; voice eventually
                                     becomes idle and returns to the allocator
              │
              ▼
   Patch swap                     ──► If a different patch needs the slot, the
                                     allocator writes the new patch's config and
                                     the cycle starts again

For layered patches all N slots trigger in lockstep on note-on and all release on note-off. Patch metadata can declare per-slot offsets (e.g. layer-2 starts 20 ms after layer-1) when an authentic emulation needs them.

Reference Patch Library

Reference patches double as specification artefacts. Each one is a concrete configuration against the documented voice topology + extras pool; if a reference patch can't be expressed, the gap is real and surfaces below in Topology Validation Findings.

Patch Slots Oscillators Filter Key extras Notes
TB-303 acid bass 1 Osc 1 VA sawtooth (or wavetable-derived square) F1 ladder LP-18 accent-RC + mismatched-ladder Mono, last-note slide. Dual envelope: ENV1 → VCA (fixed long decay), ENV2 → cutoff (variable decay)
TR-909 kick 1 Osc 1 VA sine; Osc 2 noise burst gated by ENV4 segment F2 SVF HP for subsonic cleanup candidate output saturation ENV1 → pitch drop on Osc 1; short ENV2 → VCA
TR-909 snare 1 Osc 1 VA sine (~200 Hz tonal); Osc 2 noise via granular F2 SVF BP for noise body (candidate) noise-into-resonance Two envelopes for tonal vs noise decay
TR-909 hi-hat (closed / open) 2 Three VA squares per slot at non-musical ratios = six squares total F2 SVF BP, high cutoff — Topology finding — one slot only carries three oscillators; canonical hi-hat needs six. Two-slot layering matches the original directly
TR-909 clap 1 Osc 1 noise (granular or Additive with random partials) F2 SVF BP — ENV4 multi-segment (8 segments) generates the 4-pulse clap envelope
TR-808 kick 1 Osc 1 VA sine F1 ladder LP-12 (gentle warmth) candidate output saturation Longer decay than 909; pure tonal
M-86 Hoover 1 Osc 1, 2, 3 VA PWM-saw, slight detune F1 ladder LP-24 with envelope attack — Fast LFO 1 → all three PW modulations; ENV1 → pitch drop; BBD chorus in FX send
Juno-60 pad 1 Osc 1 sawtooth + sub-oct; Osc 2 sawtooth detuned F1 ladder LP-24 — Slow ENV1 → VCA, slow ENV2 → cutoff; BBD chorus in FX send
Prophet-5 brass 1 Osc 1, 2 VA sawtooth, detuned F1 ladder LP-24 with ENV2 → cutoff attack — The ENV2 attack on cutoff is the brass character
DX7 EP (algorithm 5) 1 Osc 1 FM 6-op bypass — Full DX7-compatible 4-stage operator envelopes
DX7 bass (algorithm 14) 1 Osc 1 FM 6-op bypass — As above
Minimoog lead 1 Osc 1, 2, 3 VA sawtooth, slight detune (unison) F1 ladder LP-24 with full resonance / self-osc — The 3-osc topology matches the original directly
Karplus-Strong pluck 1 Osc 1 PhysModel KS bypass or F1 LP-12 for damping — Built-in noise excite + delay-line damping
Wavetable sweep 1 Osc 1 Wavetable F1 ladder LP-18 with ENV sweep — LFO 1 or ENV → wavetable position
Granular cloud 1 Osc 1 Granular over live mic or sample F2 SVF BP for tonal selection — Live input granularisation; long grains, position scatter

Topology Validation Findings

The reference patches above surface a small set of voice-topology questions worth resolving before HDL freeze:

  1. The 909 hi-hat needs more simple oscillators than one voice slot provides. Three options exist: (a) the canonical patch uses two slots (recommended — matches the original's 6-oscillator topology directly), (b) the Additive paradigm's 64 partials per voice can place partials at the hi-hat ratios with a noise residual, (c) reduce to three squares and rely on filter ringing for additional content. The reference library should support all three and ship the 2-slot version as the canonical "909 closed hat" patch.
  2. Whether to add a dedicated output saturation stage or rely on filter tanh for character. Affects TB-303, 808/909 kicks. A prototype could compare both routes before committing extras-pool silicon.
  3. Fixed-time portamento for the 303 is a sequencer/control-plane concern (Pulse note dispatch), not voice DSP. Confirm Pulse's sequencer exposes it as a per-step option.
  4. Sawtooth-derived square — wavetable preset (essentially free) or per-VA-osc waveshaper toggle (a register bit and a mux). The wavetable route is leaner; reserve the waveshaper option for later if it proves necessary.

These are notes for the HDL design pass, not blockers for the patch model itself — the patch format and extras-pool concept stand regardless.

New Combinations — The Real Point

The principle isn't just emulation. The extras pool combined with the general voice topology yields configurations no historical device produced. The reference library establishes that classic sounds are reachable; the unprecedented sounds are what the architecture is actually for.

A few directions:

  • Wavetable acid bass — Osc 1 Wavetable + F1 ladder + accent-RC + mismatched- ladder. The 303's filter character driven by a wavetable that sweeps in addition to the filter sweep. Not a 303 emulation — something else entirely.
  • FM acid lead — Osc 1 FM 6-op + F1 ladder + accent-RC. The stacking accent behaviour on top of FM's harmonic richness; the cutoff stacking reshapes operator interactions rather than VA harmonics.
  • Granular Hoover — Osc 1 Granular over a long sample, with the M-86 LFO routing applied to grain duration instead of PWM. The Hoover topology applied to grains.
  • Karplus-303 — PhysModel KS as the oscillator, with mismatched-ladder mode on the filter and accent-RC on cutoff. Plucked-string excitation through a 303 filter character.
  • Additive 909-hat — 64 partials placed by drum-kit synthesis rules with the mismatched-ladder mode for character. Higher fidelity than the original could achieve with six squares.
  • Multi-paradigm voice — Osc 1 FM 6-op + Osc 2 VA sub + Osc 3 Granular pad texture from a sample buffer, all through one shared filter and VCA. This is the combination audio_more.md calls out specifically; no current hardware synth does all three at once in a single voice.

Each of these is a 30-line JSON patch, expressible against the same data model and extras pool the reference patches use. There is no separate engine for any of them.

Doc Revision List

Once patches.md is integrated, these specific sections in the existing docs need rewording (identified here, not edited — Anthony integrates after reading):

audio.md — high-level mentions:

  • Any reference to a "dedicated 303 engine" in the engine/paradigm summary
  • Any "engine" wording that conflates the paradigms (VA / Sample / FM, which are real engines) with device emulations (303, 909, etc., which are patches)

audio_design.md — marketing-tier copy:

  • The "Three engines, no compromises" section, if it lists the 303 alongside the paradigms — reframe so paradigms are engines and devices are reference patches
  • Any "dedicated 303 engine" callout

audio_more.md — technical reference:

  • ### Why the 303 is Hard to Clone (≈ line 1594) — done: kept the analysis, the reframing now happens via the new ### TB-303 Reference Patch section that follows it (no longer transitions into a dedicated engine architecture)
  • ### TB-303 Engine Architecture (FireStorm) (≈ line 1611) — done: retitled to ### TB-303 Reference Patch and restructured as a patch spec. The fixed-point implementation notes moved into this document under the extras pool
  • ### 303 vs General Voice Engine (≈ line 1698) — done: removed entirely; the new TB-303 Reference Patch section above makes the right argument and the separate "dedicated engine" framing dissolves
  • ### 303 Step Sequencer (≈ line 1686) — done: kept the section, reworded "drives the 303 engine" to "drives the TB-303 patch"
  • ### 303 Track (dedicated) (≈ line 2078, workstation UI) — done: renamed to ### 303 Acid Track and reworded the body to drop the engine framing
  • Any voice-budget / competitive / live-coder table that named "303 engine" or "303 acid engine" — done: those rows now name the TB-303 patch with a cross-link to the reference patch library

The extras pool should probably get a short anchor section in audio_more.md near the voice topology (just before or after ## Per-Voice Engine) pointing to patches.md for the full data model and reference library.

Open Questions

Items to settle before locking the patch format:

  • Voice-private vs shared extras pool. Voice-private means every voice in the array has its own copies (more silicon, simpler allocator); shared means a small number of extras-circuit instances multiplexed across voices (less silicon, more allocator complexity). Voice-private is recommended for the initial pool of two extras — the silicon cost is genuinely tiny — and shared deferred until the pool grows large enough to justify the multiplexing.
  • Patch-format versioning policy. Strict (refuse to load mismatched major version) or lenient (load with best-effort field mapping)? Recommend strict on major version, lenient on minor.
  • Missing extras at load time. A FireStorm bitstream variant might omit, say, the mismatched-ladder mode to save area. When a patch references an extra the current bitstream doesn't include, should it refuse to load, fall back to a degraded path that omits the extra, or warn? Patch metadata should declare its requirements; AntOS handles the mismatch policy from there.
  • Reference patch library distribution. Bundle in DBFS by default, or as optional content packs? Drum kits in particular take meaningful space.

Important: The Ant64 family of home computers are at early design/prototype stage, everything you see here is subject to change.