Ant64 Patches — The Unit of Voice Configuration
This document defines the patch as the unit of "what makes a voice sound like a thing" on Tempest — the FireStorm chipset audio engine, the FPGA voice array described in audio_more.md. It applies only to FireStorm chipset voices. Pulse's AMY synthesis and DeMon's SID/SAM each carry their own preset conventions and are out of scope here.
The principle, in one sentence: classic synthesisers and drum machines are reference patches against the general voice topology + a small pool of opt-in extras, not bespoke engines in their own right.
This document layers on a foundation that is design intent, not implemented silicon — the voice topology in audio_more.md is the target architecture and no FireStorm HDL exists yet. Where the patch model surfaces topology questions, those questions are flagged for the HDL pass rather than papered over.
The Principle
A FireStorm voice has a fixed general topology — 3 oscillator slots (each independently selecting VA / Wavetable / FM / Granular / PhysModel / Additive), 2 filters (Huovilainen ladder + SVF), 4 envelopes, 4 per-voice LFOs + 2 global, a 64-slot modulation matrix, and a VCA. Everything that sounds like a Juno, a Prophet, a DX7, a Minimoog, a 303, or a 909 is configuration of that voice — a patch. Where a particular classic device's character depends on stateful circuit behaviour that the general voice doesn't naturally produce, the patch engages one or more entries from the extras pool — a small, curated set of opt-in character circuits any voice can switch on.
Why the 303 forced the question
The TB-303 was the example that tested the principle hardest. Its character lives in stateful circuit details that don't fit the general voice's per-note envelope + clean Huovilainen ladder:
- The accent-RC accumulator — a 47 kΩ × 33 nF RC network whose smoothed voltage drives both the filter cutoff and the VCA on accented notes, and that doesn't fully discharge between consecutive accents. Repeated accents stack, producing the distinctive "increasingly distressed animal cry" the 303 is known for.
- The mismatched first pole in the diode ladder — a deliberately undersized capacitor (C18 ≈ half the others) gives the first stage roughly 2× the cutoff of the rest, preventing the feedback phase from reaching 360° and converting resonance-into-self-oscillation into resonance-into-overdrive. This is the 303 sound.
Earlier drafts of the audio documentation treated the 303 as a dedicated single-voice engine occupying one slot with its own DSP pipeline (see audio_more.md). The justification was real — those two behaviours genuinely don't fit a general voice — but it created a precedent the rest of the design didn't follow. Every other "classic sound" is already a patch, so why not the 303?
Three options considered
- Keep the 303 dedicated. Honest about how different it is. Costs: design inconsistency, and every future device-specific quirk (909-kick saturation, snare noise-into-resonance, and so on) becomes a candidate for its own bespoke engine.
- Make every voice's silicon rich enough to fully subsume the 303. Add the accent-RC accumulator and mismatched-ladder mode to every voice as standard. Cost: silicon area paid by all 128 voices for behaviour only acid-bass patches use.
- Optional per-voice extras pool. A small curated set of opt-in character circuits any patch can engage. Voice silicon stays lean, paying for extras only on voices whose patches engage them (via a per-voice extras-enable register and the accompanying state); 303-style behaviour becomes a configurable patch extension rather than bespoke silicon.
Option 3 wins. The 303 is a configuration of a general voice + two engaged extras (accent-RC accumulator + mismatched-ladder mode). The general topology pays no permanent silicon cost for it; the extras pool pays per-extra silicon once and is shared across whichever voices are currently running patches that engage them.
The same pattern extends cleanly to future device-specific character — a 909-kick saturation circuit, a snare noise-into-resonance cross-feedback, a sawtooth-to-square waveshaper — each becomes a new entry in the pool when (and only when) a reference patch needs it.
The extras pool
Initial pool:
| Extra | What it does | Silicon cost (rough) | First needed by |
|---|---|---|---|
| Accent-RC accumulator | Stateful RC smoothing across accent-flagged notes; configurable RC time, accent amount, and target (cutoff, VCA, or modmatrix sink). Voltage decays between notes but doesn't fully discharge, so consecutive accents stack. | Small — one S2.30 accumulator register + one decay multiply per voice that engages it | TB-303 acid bass |
| Mismatched-ladder mode | When the Huovilainen ladder is engaged, applies an asymmetric coefficient to the first pole (≈ 2× the cutoff of the other three). Converts resonance-into-self-oscillation into resonance-into-overdrive. | Trivial — one extra coefficient register and a per-voice select bit | TB-303 acid bass |
Honest tradeoff. Every extra is silicon area, paid forever on every fabricated FireStorm whether or not any running patch engages it. The pool stays curated, not kitchen-sink. The bar for inclusion is that a reference patch needs it and the behaviour cannot be synthesised from existing voice features. Anything achievable by clever use of the general topology, the modulation matrix, the multi-segment ENV4, or sequencer-side logic on Pulse does not become an extra.
Candidate extras — flagged for evaluation, not yet committed:
- Output saturation stage (post-VCA tanh with configurable drive). Useful for the 303 output character, and for 808/909 kicks. The filter's tanh saturation in the resonant feedback path may already cover this; the question is whether reference patches sound right routed through filter tanh alone.
- Sawtooth-derived square waveshaper. The 303's square is sawtooth through a single-transistor waveshaper, subtly different from a clean square. This is probably best expressed as a wavetable preset rather than an extras-pool entry.
- Noise-into-resonance cross-feedback for 909-snare character — where noise excites the tonal layer's resonance. May be expressible via the modulation matrix rather than a dedicated extra.
- Fixed-time portamento (the 303's slide is fixed time, not fixed rate). This most naturally lives in Pulse's sequencer / note-dispatch logic rather than in FireStorm voice silicon.
These will be promoted into the pool only if reference-patch implementation surfaces them as genuinely missing.
Implementation notes (fixed-point)
These are the FPGA-level fixed-point details for the two extras the TB-303 reference patch engages.
Accent-RC accumulator. The trickiest part to get right in fixed-point:
// Per sample tick:
accent_voltage = accent_voltage * decay_coeff; // S1.30 multiply
if (accent_triggered) {
accent_voltage += meg_output * accent_amount; // add MEG contribution
}
vcf_cutoff += accent_voltage >> accent_depth_shift;
vca_level += (accent_voltage * vca_accent_scale) >> 16;
The RC time constant maps to a decay_coeff close to 1.0 — approximately
1 - (1 / (sample_rate * RC_time)). At 48 kHz with RC ≈ 47 kΩ × 33 nF = 1.55 ms,
decay_coeff ≈ 0.99999 in S1.30 fixed-point. Needs sufficient precision — use
S2.30 or S1.62 (64-bit accumulator) to avoid losing the accent voltage to
rounding.
Mismatched-ladder mode is modelled by giving the first one-pole stage of the Huovilainen ladder a slightly different frequency coefficient — approximately 0.5× the capacitance translates to ~2× the cutoff frequency for that stage. This is what prevents the feedback phase from reaching 360° and suppresses self-oscillation (converting it into overdrive — the 303 character).
Data Model
A patch is a layered configuration of one or more voice slots. The same data model covers single-voice patches (the common case) and multi-slot patches (stacks, splits, drum kits with one slot per drum).
Top level
- Metadata — name, author, format version, classification (bass / lead / pad / drum-kit / fx / etc.), tags, creation timestamp, source attribution where relevant (a reference patch like "TB-303 acid bass" credits the original device)
- Voice slot list — 1 to N entries. The chipset's voice budget is the upper bound, but a single patch rarely uses more than a handful of slots; drum kits are the typical multi-slot case
- Layering rules — for multi-slot patches, how the slots respond to note events: key range (split), velocity range (layer), MIDI channel, drum-note mapping (one slot per MIDI note), or unison stack (every slot plays every note with configurable detune)
- Global FX sends — patch-level send levels into BBD chorus, delay, reverb, and any other global FX bus (the FX chain is not per-voice; sends are)
Per voice slot
- Oscillators (× 3) — paradigm per oscillator slot (VA / Wavetable / FM 6-op / Granular / PhysModel / Additive) plus paradigm-specific parameters; detune, level, pan contribution, hard-sync routing, cross-modulation routing
- Filter block — F1 mode (ladder LP-24 / LP-18 / LP-12), F2 mode (SVF with LP / HP / BP / Notch outputs), routing (serial / parallel / F1-only / F2-only / bypass), cutoff and resonance bases, keyboard tracking
- Envelopes (× 4) — ENV1–3 are ADSR with optional hold; ENV4 is 8-segment loopable (Waldorf-style). Per envelope: rates, levels, key tracking, velocity-to-amount, and one or more modulation destinations
- LFOs (× 4 per-voice + 2 global) — waveform, rate (or tempo-sync division), delay, fade-in, sync mode (free / note-triggered / key-sync)
- Modulation matrix — up to 64 (source → destination × amount) bindings. Sources include every envelope and LFO, velocity, mono and poly aftertouch, mod wheel, pitch bend, key tracking, note number, random per-note, MIDI CCs, and audio-rate oscillator output. Destinations include essentially every continuous parameter the voice exposes
- VCA / pan — velocity curve (linear / exponential / custom), base level, stereo pan with optional per-note random spread
- Extras engaged — list of pool entries the slot has switched on, each with its own parameter set
- FX sends (per slot) — slot-level contributions into each global FX send
A typical per-slot config is around 1–2 kB binary uncompressed, dominated by the modmatrix (64 slots × ~16 bytes each ≈ 1 kB). A drum kit with 8 slots is roughly 8–16 kB. DBFS comfortably stores tens of thousands.
Layering and voice allocation
The voice allocator treats a single-slot patch as one voice per note and an N-slot layered patch as N voices per note. The voice budget (128 VA / 256 sample / patch-dependent FM) is consumed accordingly. A monophonic patch (303-style bass with last-note priority and slide) reserves a single slot and uses last-note priority; layered patches reserve all their slots together so they trigger in lockstep.
When a patch engages extras the allocator prefers slots whose required extras hardware is currently idle — extras are physically present on every voice in the initial pool of two (the silicon cost is tiny), so contention is unlikely. If a future extras pool is implemented as a shared resource (a small number of extras instances multiplexed across the voice array), the allocator will fall back to either stealing an older voice's extras allocation (last-note priority) or to a degraded patch path that omits the extra. Patch metadata can declare whether degraded fallback is acceptable.
File Format
Two complementary representations:
Binary — for DBFS storage and fast load
- 16-byte header: magic
"Ant64Patch", format version (uint16), patch length (uint32), CRC-32 over body - Per-slot blocks, each tagged with slot-type byte and length so future fields can be added without invalidating older patches
- Modmatrix block uses a compact (source, destination, amount) triple encoding, omitting zero-amount slots
- Extras-engaged block is a list of (extra-id, param-blob) pairs, where param-blob length is determined by extra-id
- Reference patches in the factory library ship as binary blobs either in the FireStorm bitstream's BRAM init or as a DBFS table loaded at boot
JSON — for editing, version control, and sharing
- One-to-one round-trip with the binary form (every binary field has a JSON key, every JSON key produces deterministic binary)
- Pretty-printed for readability; arrays for slot lists and modmatrix
- The AntOS workstation app's patch editor reads and writes JSON natively; binary export happens on save or transfer to another device
- Patches under version control live as JSON; the binary form is build artefact
Honest tradeoff. Binary is compact, fast to load, DBFS-friendly, and memory-mappable. JSON is verbose, slow to parse, and unfriendly to embedded loaders but vastly more useful for editing, diffing, and code review. Carrying both means writing the serialiser twice and keeping them in lockstep — the cost is real but the tooling benefit is larger.
Lifecycle
AntOS loads patch from DBFS ──► JSON or binary, parsed into in-memory struct
│
▼
Patch passed to Pulse ──► Pulse holds the "active patch table" for each
allocated voice slot (one entry per slot)
│
▼
Voice slot(s) bound ──► Pulse writes voice config to FireStorm via OPI
register window: oscillator paradigm bits, filter
modes, envelope rates, modmatrix routings, extras
enable bits, all of it
│
▼
Note-on event ──► Pulse triggers the slot(s); FireStorm's voice
envelopes and LFOs start running autonomously
and the voice generates audio into the mix bus
│
▼
During note hold ──► Live parameter tweaks (jog dial, MIDI CC,
modmatrix-bound sources) cause Pulse register
writes into the running voice's config; the
voice continues without re-binding
│
▼
Note-off ──► Envelopes enter release; voice eventually
becomes idle and returns to the allocator
│
▼
Patch swap ──► If a different patch needs the slot, the
allocator writes the new patch's config and
the cycle starts again
For layered patches all N slots trigger in lockstep on note-on and all release on note-off. Patch metadata can declare per-slot offsets (e.g. layer-2 starts 20 ms after layer-1) when an authentic emulation needs them.
Reference Patch Library
Reference patches double as specification artefacts. Each one is a concrete configuration against the documented voice topology + extras pool; if a reference patch can't be expressed, the gap is real and surfaces below in Topology Validation Findings.
| Patch | Slots | Oscillators | Filter | Key extras | Notes |
|---|---|---|---|---|---|
| TB-303 acid bass | 1 | Osc 1 VA sawtooth (or wavetable-derived square) | F1 ladder LP-18 | accent-RC + mismatched-ladder | Mono, last-note slide. Dual envelope: ENV1 → VCA (fixed long decay), ENV2 → cutoff (variable decay) |
| TR-909 kick | 1 | Osc 1 VA sine; Osc 2 noise burst gated by ENV4 segment | F2 SVF HP for subsonic cleanup | candidate output saturation | ENV1 → pitch drop on Osc 1; short ENV2 → VCA |
| TR-909 snare | 1 | Osc 1 VA sine (~200 Hz tonal); Osc 2 noise via granular | F2 SVF BP for noise body | (candidate) noise-into-resonance | Two envelopes for tonal vs noise decay |
| TR-909 hi-hat (closed / open) | 2 | Three VA squares per slot at non-musical ratios = six squares total | F2 SVF BP, high cutoff | — | Topology finding — one slot only carries three oscillators; canonical hi-hat needs six. Two-slot layering matches the original directly |
| TR-909 clap | 1 | Osc 1 noise (granular or Additive with random partials) | F2 SVF BP | — | ENV4 multi-segment (8 segments) generates the 4-pulse clap envelope |
| TR-808 kick | 1 | Osc 1 VA sine | F1 ladder LP-12 (gentle warmth) | candidate output saturation | Longer decay than 909; pure tonal |
| M-86 Hoover | 1 | Osc 1, 2, 3 VA PWM-saw, slight detune | F1 ladder LP-24 with envelope attack | — | Fast LFO 1 → all three PW modulations; ENV1 → pitch drop; BBD chorus in FX send |
| Juno-60 pad | 1 | Osc 1 sawtooth + sub-oct; Osc 2 sawtooth detuned | F1 ladder LP-24 | — | Slow ENV1 → VCA, slow ENV2 → cutoff; BBD chorus in FX send |
| Prophet-5 brass | 1 | Osc 1, 2 VA sawtooth, detuned | F1 ladder LP-24 with ENV2 → cutoff attack | — | The ENV2 attack on cutoff is the brass character |
| DX7 EP (algorithm 5) | 1 | Osc 1 FM 6-op | bypass | — | Full DX7-compatible 4-stage operator envelopes |
| DX7 bass (algorithm 14) | 1 | Osc 1 FM 6-op | bypass | — | As above |
| Minimoog lead | 1 | Osc 1, 2, 3 VA sawtooth, slight detune (unison) | F1 ladder LP-24 with full resonance / self-osc | — | The 3-osc topology matches the original directly |
| Karplus-Strong pluck | 1 | Osc 1 PhysModel KS | bypass or F1 LP-12 for damping | — | Built-in noise excite + delay-line damping |
| Wavetable sweep | 1 | Osc 1 Wavetable | F1 ladder LP-18 with ENV sweep | — | LFO 1 or ENV → wavetable position |
| Granular cloud | 1 | Osc 1 Granular over live mic or sample | F2 SVF BP for tonal selection | — | Live input granularisation; long grains, position scatter |
Topology Validation Findings
The reference patches above surface a small set of voice-topology questions worth resolving before HDL freeze:
- The 909 hi-hat needs more simple oscillators than one voice slot provides. Three options exist: (a) the canonical patch uses two slots (recommended — matches the original's 6-oscillator topology directly), (b) the Additive paradigm's 64 partials per voice can place partials at the hi-hat ratios with a noise residual, (c) reduce to three squares and rely on filter ringing for additional content. The reference library should support all three and ship the 2-slot version as the canonical "909 closed hat" patch.
- Whether to add a dedicated output saturation stage or rely on filter tanh for character. Affects TB-303, 808/909 kicks. A prototype could compare both routes before committing extras-pool silicon.
- Fixed-time portamento for the 303 is a sequencer/control-plane concern (Pulse note dispatch), not voice DSP. Confirm Pulse's sequencer exposes it as a per-step option.
- Sawtooth-derived square — wavetable preset (essentially free) or per-VA-osc waveshaper toggle (a register bit and a mux). The wavetable route is leaner; reserve the waveshaper option for later if it proves necessary.
These are notes for the HDL design pass, not blockers for the patch model itself — the patch format and extras-pool concept stand regardless.
New Combinations — The Real Point
The principle isn't just emulation. The extras pool combined with the general voice topology yields configurations no historical device produced. The reference library establishes that classic sounds are reachable; the unprecedented sounds are what the architecture is actually for.
A few directions:
- Wavetable acid bass — Osc 1 Wavetable + F1 ladder + accent-RC + mismatched- ladder. The 303's filter character driven by a wavetable that sweeps in addition to the filter sweep. Not a 303 emulation — something else entirely.
- FM acid lead — Osc 1 FM 6-op + F1 ladder + accent-RC. The stacking accent behaviour on top of FM's harmonic richness; the cutoff stacking reshapes operator interactions rather than VA harmonics.
- Granular Hoover — Osc 1 Granular over a long sample, with the M-86 LFO routing applied to grain duration instead of PWM. The Hoover topology applied to grains.
- Karplus-303 — PhysModel KS as the oscillator, with mismatched-ladder mode on the filter and accent-RC on cutoff. Plucked-string excitation through a 303 filter character.
- Additive 909-hat — 64 partials placed by drum-kit synthesis rules with the mismatched-ladder mode for character. Higher fidelity than the original could achieve with six squares.
- Multi-paradigm voice — Osc 1 FM 6-op + Osc 2 VA sub + Osc 3 Granular pad texture from a sample buffer, all through one shared filter and VCA. This is the combination audio_more.md calls out specifically; no current hardware synth does all three at once in a single voice.
Each of these is a 30-line JSON patch, expressible against the same data model and extras pool the reference patches use. There is no separate engine for any of them.
Doc Revision List
Once patches.md is integrated, these specific sections in the existing docs need rewording (identified here, not edited — Anthony integrates after reading):
audio.md — high-level mentions:
- Any reference to a "dedicated 303 engine" in the engine/paradigm summary
- Any "engine" wording that conflates the paradigms (VA / Sample / FM, which are real engines) with device emulations (303, 909, etc., which are patches)
audio_design.md — marketing-tier copy:
- The "Three engines, no compromises" section, if it lists the 303 alongside the paradigms — reframe so paradigms are engines and devices are reference patches
- Any "dedicated 303 engine" callout
audio_more.md — technical reference:
### Why the 303 is Hard to Clone(≈ line 1594) — done: kept the analysis, the reframing now happens via the new### TB-303 Reference Patchsection that follows it (no longer transitions into a dedicated engine architecture)### TB-303 Engine Architecture (FireStorm)(≈ line 1611) — done: retitled to### TB-303 Reference Patchand restructured as a patch spec. The fixed-point implementation notes moved into this document under the extras pool### 303 vs General Voice Engine(≈ line 1698) — done: removed entirely; the new TB-303 Reference Patch section above makes the right argument and the separate "dedicated engine" framing dissolves### 303 Step Sequencer(≈ line 1686) — done: kept the section, reworded "drives the 303 engine" to "drives the TB-303 patch"### 303 Track (dedicated)(≈ line 2078, workstation UI) — done: renamed to### 303 Acid Trackand reworded the body to drop the engine framing- Any voice-budget / competitive / live-coder table that named "303 engine" or "303 acid engine" — done: those rows now name the TB-303 patch with a cross-link to the reference patch library
The extras pool should probably get a short anchor section in audio_more.md near
the voice topology (just before or after ## Per-Voice Engine) pointing to
patches.md for the full data model and reference library.
Open Questions
Items to settle before locking the patch format:
- Voice-private vs shared extras pool. Voice-private means every voice in the array has its own copies (more silicon, simpler allocator); shared means a small number of extras-circuit instances multiplexed across voices (less silicon, more allocator complexity). Voice-private is recommended for the initial pool of two extras — the silicon cost is genuinely tiny — and shared deferred until the pool grows large enough to justify the multiplexing.
- Patch-format versioning policy. Strict (refuse to load mismatched major version) or lenient (load with best-effort field mapping)? Recommend strict on major version, lenient on minor.
- Missing extras at load time. A FireStorm bitstream variant might omit, say, the mismatched-ladder mode to save area. When a patch references an extra the current bitstream doesn't include, should it refuse to load, fall back to a degraded path that omits the extra, or warn? Patch metadata should declare its requirements; AntOS handles the mismatch policy from there.
- Reference patch library distribution. Bundle in DBFS by default, or as optional content packs? Drum kits in particular take meaningful space.