Probe — tagged memory, and syncing to a running guest

SRAM and DDR3 on the Ant64 are 36 and 72 bits wide — nine bits per byte. Spending the ninth bit as a tag gives every byte in the machine its own breakpoint, checked for free in the cycle the data arrives. An AntOS program can then watch a simulated system's memory and act in sync with it — adding backdrops, sprites, audio, save-states — without altering a single byte of the guest's code.

Status: design. The mechanism is settled by the memory width; the modes, the event format and the substitution question are open. See Open questions.

Related: FireStorm · personality cartridges · emulation · develop


The mechanism

Wide memory carries a spare bit per byte, normally spent on parity or ECC. Spend it instead as a tag:

DDR3 ×72  =  64 data + 8 tag      →  1 tag bit per byte
SRAM ×36  =  32 data + 4 tag      →  1 tag bit per byte

The tag is fetched with the byte, in the same cycle, through the same wires. So the check is:

if (tag) → raise

There is no comparator, no CAM, no address decode, and no lookup on the fast path. The cost of having a breakpoint set is zero, and the cost of having ten million set is also zero. Only a hit costs anything.

That inversion is the whole idea. Conventional hardware watchpoints are scarce because each one is a comparator against the address bus — four of them, or eight, and you spend your debugging session deciding which to spend. Here the address bus is not consulted at all: the memory itself remembers which of its bytes are interesting.

On a hit, the address goes to a lookup — a table saying why this byte is tagged and what to do about it — and an event is emitted. The tag is the filter; the table is the meaning. Only tagged accesses ever reach the table, so it can be as rich as you like.


Instruction fetches carry more than one bit

The tag is one bit per byte, and a CPU fetching an instruction reads several bytes at once. So the width of the tag scales with the width of the fetch: a 4-byte instruction delivers four tag bits simultaneously, in the same cycle, through the same wires.

That turns the tag from a flag into an opcode. Sixteen values, read for free on every instruction fetch, with no table access at all.

This is why the instruction side and the data side are different features. A data access has whatever width the guest chose — one byte, sometimes four — so nothing can be assumed and the data side stays one bit plus a table lookup. An instruction fetch has a width the decoder knows, so the bits can be assembled into a code. The two live in different parts of the design: the code is formed in the fetch/decode path after instruction boundaries are known, not in the memory controller.

The rule, stated once, for any width

Instruction widths are not 2 and 4. They are 1 to 3 on a 6502, 1 to 4 on a Z80, 2 to 10 on a 68000, 1 to 15 on an x86 — and 3-byte fetches are common everywhere. One rule covers all of it:

The code is the instruction's tag bits, lowest address as the least significant bit, padded on the high side to the target's code width: with 1s if any bit is set, with 0s if none is.

Bit order matters and is easy to implement backwards: byte 0 of the instruction is bit 0, and the bytes a narrow instruction doesn't have are the high bits. That is what makes padding meaningful.

At a 4-bit code width, what each fetch width can reach:

Bytes fetched Reachable codes Count
1 0000, 1111 1 + nothing
2 + 1101, 1110 3 + nothing
3 + 1001, 1010, 1011, 1100 7 + nothing
4 + 0001–0111, 1000 15 + nothing

The property this buys: the sets nest

A code reachable at width w is reachable at every width above it. 1111 works everywhere; 1101/1110 work from two bytes up; 1001–1100 from three bytes up; the rest need four. Nothing is ambiguous, because a code can only have been produced by a fetch at least as wide as the narrowest that reaches it.

So codes are allocated by the narrowest instruction that must carry them, which is a real design rule rather than a table to memorise:

  • 1111 — reachable from a single byte, so it must be the universal meaning: consult the table.
  • 1101, 1110 — the scarcest useful codes. Whatever must work in compressed RISC-V or in 2-byte 68000 instructions goes here.
  • 1001–1100 — available to anything three bytes or wider. Comfortable on a 68000, where 2-byte instructions are common but 4-byte ones are not rare.
  • 0001–1000 — full-width only.

And it degrades to the one-bit design exactly. A single-byte fetch sets one bit, the rule fills the rest with ones, and the code is 1111 — consult the table, which is the behaviour specified everywhere else on this page. Clear the bit and it is 0000. The encoding is a strict superset of the one-bit scheme, not a parallel mechanism, and an 8-bit or variable-length guest needs no special case at all.

Settled: a 4-bit code and up to an 8-bit parameter

A 68000 move.l #imm,(addr).l is ten bytes; an x86 instruction can reach fifteen. Bytes 5 onward carry a parameter — a small immediate saying which counter to bump, which channel to raise, which bank to arm — capped at 8 bits, so tags above byte 12 take no part in the fetch and remain free for data-side use.

Instruction bytes Code bits Parameter bits
1–4 4 (padded) none
5 4 1
6 4 2
8 4 4
10 4 6
12 or more 4 8 (capped)

The code is one-padded; the parameter is zero-extended. These are opposite rules on purpose and the temptation to unify them should be resisted. Padding the code with 1s is what makes narrow fetches unambiguous. The parameter is a value, so an absent one must read as 0 — one-extending it would make every short instruction report parameter 255 and quietly select the wrong counter.

The parameter is what makes 16 codes feel like plenty. One code meaning "count", with the parameter selecting the counter, gives 256 counters from a single code; one code meaning "event" gives 256 channels. You need the other fifteen codes far less than the raw count suggests.

The honest limit: parameters need 5-byte instructions

Nothing below five bytes has room for one. So on RISC-V — maximum four bytes — there is never a parameter at all, and the same is true of the Z80, the 6502 and the 65816. It is a wide-CISC feature: the 68000 and x86 get it, and they are the targets whose instructions are long enough to spare the bits.

On targets with no parameter room the fallback is to spend codes instead: allocate 0001–1000 as "count into counter 0" through "counter 7" and you have eight counters at the cost of eight codes, or fall through to 1111 and the table and pay a memory access for unlimited ones. So code assignment is genuinely per-architecture — a 68000 profile and a RISC-V profile should not expect the same allocation, and the define_code table exists per target for exactly that reason.

The code width stays at 4 regardless, because a code indexes a register file — sixteen entries — and widening it turns the mechanism back into the memory lookup it exists to avoid.

What the codes are for

The value is that a code costs no memory access. The sixteen meanings live in a small register file — sixteen entries, configured once per target — so decoding one is a lookup into registers, not into DDR. That is the difference between a hit costing a memory round trip and costing nothing.

So spend the codes on the things that must be free, and let everything richer fall through:

  • 0000 — nothing.
  • 1111 — consult the table: the full per-address behaviour, all the modes, actions and substitutions described on this page. It has to be this one, because it is the only code a single-byte fetch can produce.
  • 1110, 1101 — the two most valuable inline behaviours, chosen carefully, because these are the narrowest codes after 1111 and so the only ones available in compressed RISC-V or 2-byte 68000 instructions.
  • 1001–1100 — four more, from three bytes up.
  • 0001–1000 — eight more, full-width only.

Good candidates for inline codes, roughly in order of how much they gain by avoiding the table: bump a counter (coverage and profiling over the whole guest with no event traffic at all — just counters), raise an event on a fixed channel, break, arm or disarm a tag bank, and region enter / exit markers for a hardware profiler.

The counter one deserves the emphasis. A code meaning "increment and continue" gives you execution counts for every tagged instruction with no queue, no drops and no watcher — which is a complete instruction profiler for a guest whose source you do not have.

Two consequences to design around

  • The tag map must know instruction boundaries. Placing a four-bit code requires knowing where the instruction starts and how wide it is. Tag the second byte of an instruction believing it a boundary and you produce a wrong code silently. So the instruction-side profile is a toolchain output, not a range — either a disassembly of a binary you did not build, or, far better, something your assembler emitted. See below.
  • Narrow-instruction code gets few inline codes. Compressed RISC-V, or a 68000 routine of 2-byte instructions, has 1101, 1110 and the table. Where the guest is built for size, most enhancement stays table-driven and the inline codes are a bonus rather than the mechanism. Worth knowing before assigning them.

The other direction: your own assembler emits the map

Everything above assumes you are working on somebody else's binary. Invert it. If you are developing for the target, the tag map is a build artifact — your assembler knows exactly where every instruction starts, how wide it is, what it is called and which source line it came from, so it can emit the map with certainty rather than leaving anyone to recover it.

That is a different tool for a different person, and a more valuable one. A homebrew author writing 68000 for an Amiga personality, or Z80 for a Spectrum core, has historically had a machine-language monitor and patience. This gives them something better than the original hardware ever offered:

  • Breakpoints that do not alter the instruction stream. This is the big one. A conventional breakpoint patches a trap instruction into your code, which changes timing and is therefore unusable in exactly the code that most needs debugging — raster chases, copper-synced effects, cycle-counted loops. A tag breakpoint changes nothing: the bytes the CPU fetches are byte-identical with the tag set or clear. You can debug timing-critical code without disturbing its timing, which on these machines was never previously possible.
  • Unlimited watchpoints, where the real hardware had none.
  • Zero-cost instrumentation. A count code next to an instruction gives execution counts with no event traffic at all — coverage and profiling on a machine that never had a profiler.
  • Source-level events. Because the map came from the assembler, an event resolves to a symbol and a line, not a bare address.

The natural authoring form is a directive in the source, so instrumentation lives beside the code it measures and disappears when you build without it:

        probe.count                 ; how often is this reached?
inner:  move.w  d0,(a0)+
        dbra    d1,inner
        probe.break  when=exec      ; stop here, without patching a trap

Practically, most retro toolchains emit a listing file rather than DWARF, so deriving the map from a listing matters more than supporting a debug format. Either is a build output, which is the point.

A stale map is the dangerous case. Relink the binary and every address moves; apply yesterday's map and you tag the middles of instructions, which is the silent-wrong-code failure this design is otherwise careful to avoid. The map must carry a hash of the binary it was generated for, and applying it to anything else must be refused rather than warned about.

This is deliberately not aimed at native Ant64 development. Code written for AntOS or FireStorm already has a real toolchain, GDB and source-level debugging — see develop. The value here is for software targeting a guest, where none of that exists and where, until now, the answer was a monitor and a hex dump.


Two modes, and they are not the same feature

The distinction matters more than it first appears, because the interesting use case needs the weaker one.

Trace — signal and continue Break — stall and wait
Guest runs on, undisturbed halts until the watcher releases it
Timing preserved exactly destroyed
For live enhancement, telemetry, sync debugging, inspection, stepping
Event delivery asynchronous, into a queue synchronous, handshaken

Enhancement needs trace, not break. The moment you stall a guest to draw a backdrop you have broken its frame timing, its audio, and any code that counts cycles — which, on the machines this is aimed at, is most of it. The watcher does not need the world frozen; it needs to know promptly that something happened. A few microseconds of latency is invisible. A stall is not.

Break is still worth having, because it turns the same hardware into a debugger with unlimited watchpoints — which is a serious tool on its own, for anyone bringing up a core or reverse-engineering a binary.


What an event has to carry

An address alone is too thin to act on. Each event should carry, at minimum:

Field Why
Address which byte
Access type — read / write / execute the same address means different things fetched, stored to, and run
Value on a write, the new value; on a read, what was served
Beam position — scanline and pixel the one that makes enhancement possible. "The guest wrote the player's X" is useful; "…on scanline 48" tells you whether you can still draw behind it this frame
Frame / cycle counter ordering and rate measurement
Guest PC, where available which routine did it, not just what it touched

Beam position is the field it would be easy to leave out and painful to add later. Everything about compositing a new layer into a running frame depends on knowing where in the frame you are.


Tag classes — one bit, many meanings

One bit cannot say why a byte matters, and it does not need to. It says "stop and look this up"; the table says the rest. A table entry is naturally:

  • Class — trace or break
  • Access mask — fire on read, write, execute, or a combination
  • Handler — an event id the watching program subscribes to
  • Rate limit — fire at most once per frame, per N hits, or on change of value only

That last one is not a nicety. A tag on a byte inside a copy loop fires tens of thousands of times a second; "only when the value changes" turns an unusable tag into a useful one, and it belongs in hardware where it costs nothing rather than in the watcher where it costs a round trip.


Profiles — and why this becomes a community artifact

Somebody has to decide which addresses matter. For a given piece of software that is a profile: a data file mapping addresses to meanings.

0x8A40  write  once-per-frame  "player.x"
0x8A42  write  once-per-frame  "player.y"
0x9100  exec   every           "draw_background"
0xC000  read   on-change       "level_number"

Which means the work of understanding a game is done once, by one person, and then published. Someone maps a game's routines; everyone else gets enhanced graphics by downloading a file. The profile is small, it is text, it is obviously shareable, and it pairs naturally with a personality cartridge or a game manifest.

And crucially, the original is untouched. No patched binary, no relocated code, no checksum failure — and many of these games do check themselves. The enhancement lives entirely outside the thing it enhances, which is also why it can be turned off, updated independently, or applied to a different revision of the same game by editing one table.

This works best exactly where the user identified: software built to run at absolute addresses, which is nearly everything from the era this machine is aimed at.


Finding the addresses

A profile has to start somewhere, and the same hardware is the tool for building one — because unlimited tags means you can afford to be indiscriminate.

  1. Tag broadly — a whole region, or all of RAM.
  2. Watch what fires, and when. The beam position alone separates "written during the vertical blank" from "written mid-frame", which is most of the way to knowing what a variable is for.
  3. Narrow. Do something in the game; see which tags moved.

That is the classic cheat-finder / trainer workflow, running at full speed with no emulator overhead, and it is worth exposing as a first-class discovery mode rather than leaving people to build it themselves.

Hunting a value, and the thing that makes it better than a software cheat finder

The concrete version: tag all of RAM in write mode, keep the addresses that decrement, play until you die, and intersect. Repeat after each death and the lives counter falls out. Increments instead of decrements, and you have the level.

Two refinements make it work on real software rather than in principle:

  • Classify the delta in hardware. The tag logic already compares values for rate = "change", so decremented by exactly 1 is free rather than a software scan. Much more selective than "went down".
  • Drop candidates that move while you are not dying. Timers, loop counters, falling sprites and DMA pointers all decrement. The step people skip is the negative one — play normally, discard everything that changed anyway — and it removes most of the set on its own.
  • Let disqualified addresses untag themselves, in hardware. A candidate that breaks the current rule clears its own tag bit on the very access that disproves it: no event queued, no software consulted, and that byte costs nothing ever again. This is the difference between whole-RAM tagging being a denial of service on the event queue and being the obvious opening move — the surviving set only shrinks, so the event rate falls away as the hunt proceeds rather than staying flat. The negative pass becomes the cheapest of all, since any write at all is disproof. The one cost is that pruning is destructive and hardware cannot undo it, so the candidate set is mirrored in software and can be re-tagged after an over-strict pass.

But the thing a software cheat finder cannot give you is the PC. Every candidate arrives with the instruction that wrote it, and that is a better discriminator than any value test: a lives counter is written by one instruction, a loop variable by many. It also changes what you do with the answer. Knowing the address lets you freeze a value; knowing the writer lets you stop the decrement itself:

p:suppress_writes(0x00C3, 0x00C3, { when = { pc = 0x8A17 } })

That is a complete trainer. It is also the one case where write suppression is clearly correct despite the warning above — the guest reads back the value it held before the write, which is exactly the intent. It is what a hand-written trainer has always done, NOP the DEC, except that nothing is patched, so it toggles at runtime and survives a self-check.

So a trainer becomes a short Luau script and a shared profile — no patched binary, no build step, and it applies to an unmodified original.


Actions — closing the latency loop

An event delivered to a watching program costs a PCIe crossing each way. That is fine for compositing a backdrop a frame later; it is useless for anything that has to land inside the guest's own timing.

So a tag can carry an action: a short, ordered sequence of writes applied the moment it fires, in hardware, with the watcher never consulted. Refill a counter, set a flag, swap a pointer, copy sixteen bytes. The trainer earlier stops being "suppress a write" and becomes "put the value back, plus whatever else should happen at the same instant".

Three constraints are structural rather than policy:

  • Action writes must be silent — they must not themselves trigger tags — or an action touching a tagged byte re-enters and you have a feedback loop running at memory speed. Chaining is opt-in and depth-limited.
  • Sequences must be short, because they execute in the memory path and every step costs the guest cycles. Bulk data belongs in a substitution buffer, which costs nothing per access. Actions are for a handful of writes, and the cap should be a documented number rather than an aspiration.
  • Commit timing matters. Applying mid-frame lets the guest observe a half-written sequence, so vblank is the safe default and "now" is right mainly when reacting to a write the guest has just made and will not re-read.

The step worth having beyond plain writes is arm / disarm. A step that enables or disables another tag turns actions into a small state machine that runs at guest speed with the watcher entirely outside it — boss-fight tags that stay disarmed until the guest's own code spawns the boss, armed by its execution rather than by polling. That is the mechanism that lets a large profile stay cheap, and it is why the step vocabulary includes tag control and not only memory.


Substitution — the bigger step, deliberately separate

Observation gets you addition: extra sprites, a parallax backdrop behind the guest's playfield, an extra audio layer, an overlay. It does not get you replacement, because the guest still draws its own version.

There is a natural extension. A tagged read could return a different byte to the guest than the one in memory — substituting graphics data, a table, or a constant, without altering the stored image at all.

This is genuinely more invasive and is a separate, explicit capability rather than a flag on a trace tag: a substitution is a lie told to the guest, and the failure mode when it is wrong is a crash rather than a missing sprite. It is off until switched on, per target. The API is in luau_probe; three findings from designing it are worth having here, because each one closed off an approach that looks obvious first.

Substitute reads, not writes. The instinct is to suppress the guest's drawing and draw your own. That fails, because the guest reads back what it wrote and you have desynchronised it from its own state. Replace the source data instead — swap the tiles and sprites it fetches, and the guest renders your artwork using its own code, at its own time, in the right place, with its logic untouched. It does the work; you only change what it is working from.

A substitution is a buffer, not a callback. A read cannot wait for a program on the other side of PCIe. So hardware serves from a pre-loaded buffer, and the watcher updates that buffer asynchronously whenever it likes. function(addr) return byte end is unimplementable at bus speed and is not on offer — which is a constraint worth knowing before designing around one.

when.pc is the answer to self-checking software, and there is a lot of it. Condition the substitution on the reading program counter: the renderer gets the new artwork, the checksum routine gets the original bytes, and the check still passes. One address telling two truths, distinguished by who asked. This is the main thing substitution adds to the tag hardware beyond the buffer, and it is what makes the feature viable rather than merely possible.

Changes commit at vertical blank by default, because altering what memory says mid-frame tears.


What it costs — the honest part

  • You give up ECC and parity on tagged memory. Nine bits per byte is nine bits; the ninth cannot be both a tag and a check bit. For a region hosting a simulated 8- or 16-bit system that is an easy trade. For memory holding AntOS or FireStorm's own state it may not be — so the tag capability probably wants to be per-region, with ECC kept where it earns its place.
  • Hit rate is the real budget. Tags are free; hits are not. A tag on an instruction inside a hot loop is a self-inflicted denial of service on the event queue. Rate limiting, on-change filtering and self-pruning are the mitigation, and all three belong in the hardware — a filter that costs a round trip to the watcher is not a filter.
  • The queue can overflow, and the policy must be explicit and visible: drop and count, or stall the guest. Silent loss would make an enhancement that works in testing and glitches in play — the worst possible failure.
  • Tag storage has to be initialised, and a stale tag map from a previous guest is a confusing bug. Clearing tags on region assignment should be the default, not something you remember.

API sketch

The full API is in luau_probe, with the design and native core in antos_probe. In outline:

local probe = require("probe")

local p = probe.attach("guest")            -- the running simulated system

p:tag(0x8A40, { on = "write", rate = "frame",  event = "player_moved" })
p:tag_range(0x9100, 0x9140, { on = "exec", event = "drawing" })
p:load_profile("D:/profiles/somegame.toml")   -- or just load a published one

p:on("player_moved", function(e)
  -- e.addr, e.value, e.scanline, e.frame, e.pc
  if e.scanline < 48 then backdrop:parallax(e.value) end
end)

The watcher is an ordinary AntOS program — C++ or Luau — so an enhancement layer is a normal piece of software with a normal API, not a special build of anything.


Open questions

  • Trace and break in one tag bit, or two capabilities? The class lives in the lookup table, so one bit is enough — but a break needs a handshake path back to the guest's clock enable that a trace does not.
  • Where does the lookup table live? BRAM is fast and small; DDR is large and adds latency to every hit. Likely a BRAM cache over a DDR table, sized by how many tags a realistic profile sets.
  • Per-region tag enable — which regions can be tagged, and how the ECC trade is expressed to the user.
  • Event queue depth and overflow policy, and how a dropped event is surfaced rather than hidden.
  • Does the watcher run on DeMon, on FireStorm, or either? Latency and the PCIe hop argue for FireStorm for anything frame-locked; DeMon is the easier place to write it. Possibly both, with a latency budget documented per side.
  • PC comparison width and cost in the tag hardware — when.pc is what makes substitution survive self-checking guests, so its precision (exact PC, range, or a coarse bucket) decides whether the feature works on real software.
  • Where substitution buffers live, and how much of DDR a realistic graphics replacement wants to reserve.
  • The action-sequence cap — how many steps can execute in the memory path without a visible stall, which sets what an action can usefully do and needs measuring rather than guessing.
  • Which behaviours get the two compression-surviving codes (1101, 1110). They are the scarcest resource in the whole design and the choice is hard to reverse once profiles exist in the wild.
  • How the code is assembled when a wide fetch spans two instructions — the bits arrive with the memory word, the boundary comes from the decoder, so the split has to happen after boundary determination and the timing of that wants checking against the fetch pipeline.

Related

Important: The Ant64 family of home computers are at early design/prototype stage, everything you see here is subject to change.