Luau Bindings — probe

The require("probe") script API — watch a simulated system's memory and act in sync with it, without altering a byte of the guest's code. Tag any address; get an event when it is read, written or executed. Everything here is [antos]. Design & native core: antos_probe. The mechanism: probe.

local probe = require("probe")
local p = probe.attach("guest")

p:tag(0x8A40, { on = "write", rate = "change", event = "player_x" })
p:on("player_x", function(e) backdrop:parallax(e.value, e.scanline) end)

All functions are AntOS additions ([antos]); failures return nil, errmsg. Addresses are in the guest's own address space.


Attaching

Function Behaviour
probe.targets() Taggable regions currently available: { name, base, size, tags_enabled, ecc }.
probe.attach(name) Open a target → probe handle.
p:detach() Release it. Tags set by this handle are cleared.
p:regions() The target's regions and whether each can be tagged.

Tagging a region costs its ECC — the ninth bit is either a check bit or a tag bit, not both. That is an easy trade for memory hosting a guest and a poor one for memory holding system state, which is why it is per-region and why targets() reports ecc alongside tags_enabled. Attaching clears any tags left by a previous guest.


Setting tags

Function Behaviour
p:tag(addr, opts) Tag one byte.
p:tag_range(from, to, opts) Tag a span — cheap, because tags cost nothing to hold.
p:untag(addr) · p:untag_range(from, to) Remove tags.
p:clear() Remove every tag on the target.
p:tags() Iterator over what is currently tagged, with each tag's options and hit count.

opts fields:

Field Values Meaning
on "read", "write", "exec", or an array Which accesses fire. Default "write".
event string The handler name delivered to p:on().
rate "every", "frame", "change", or a number How often it may fire — see below.
mode "trace" (default) or "break" Signal and continue, or halt the guest.
note string Carried into the event; what this address is.

rate is the field that decides whether a tag is usable

Tags are free to hold; hits are not. A tag on a byte inside a copy loop fires tens of thousands of times a second and will swamp the queue.

rate Fires
"every" on every matching access — for cold addresses only
"frame" at most once per frame
"change" only when the value differs from last time
"once" fires once, then removes its own tag — free thereafter
a number at most once per N matching accesses

The filtering happens in hardware, so a filtered tag costs nothing rather than costing a round trip. Reach for "change" or "frame" by default and drop to "every" only when you know the address is quiet.


Receiving events

Function Behaviour
p:on(event, fn) Subscribe. Returns a handle for p:off().
p:off(handle) Unsubscribe.
p:poll([max]) Drain events as an array instead of by callback, for a program with its own loop.
p:flush() Discard anything queued — useful after a guest reset.

Every event carries:

Field
addr the byte
access "read" / "write" / "exec"
value written, or served on a read
scanline, pixel where in the frame it happened
frame frame counter
pc the guest's program counter, where the target can supply it
note whatever you set on the tag
p:on("drawing", function(e)
  if e.scanline < 48 then          -- still time to composite behind it
    layer:draw_backdrop(level)
  end
end)

scanline is the field that makes enhancement possible rather than merely observable — knowing that the guest moved the player is useful; knowing you are still 200 lines from the beam reaching it is what lets you act on it this frame.


Break mode

Function Behaviour
p:pause() · p:resume() Halt and release the guest explicitly.
p:step([n]) Run n guest instructions and halt again.
p:halted() true while stopped, with the event that caused it.
p:read(addr, len) · p:write(addr, bytes) Guest memory, whether halted or running.
p:registers() Guest CPU state, where the target exposes it.

Break stops the world, and that is usually wrong. A halted guest loses its frame timing, its audio, and any code counting cycles — which on the machines this is aimed at is most of them. Break is for debugging and inspection. For enhancement, stay in "trace".


Profiles

A tag map is data, so the work of understanding a piece of software is done once and shared.

Function Behaviour
probe.load_profile(path) Read a profile file → a table of tag definitions.
p:apply_profile(profile) Set every tag it defines.
p:save_profile(path) Write the current tags out, notes included.
p:apply_profile(probe.load_profile("D:/profiles/somegame.toml"))

Because the guest is never modified, a profile can be swapped, updated or turned off at any time, and nothing checksums wrong — which matters, since a good deal of this software checks itself.


Discovery

Finding the addresses in the first place. Unlimited tags mean you can afford to be indiscriminate.

Function Behaviour
p:discover(from, to [, opts]) Tag a whole span and collect what actually fires.
p:candidates([filter]) What has fired, with hit counts, first/last frame and typical scanline.
p:narrow(predicate) Keep only candidates matching — the "do something in the game, see what moved" step.

Beam position does much of the work here on its own: an address written during the vertical blank behaves very differently from one written mid-frame, and that distinction alone usually separates state from rendering.


Hunting a value — lives, level, score

The classic cheat-finder loop: tag everything, watch how values move, narrow across repeated events. Ordinarily this runs inside an emulator at a fraction of speed; here it runs at full speed on real tagged memory, and every candidate arrives with the PC that wrote it, which turns out to matter more than the address.

Delta classes

Tag hardware already compares values for rate = "change", so it can classify the change for free. on = "write" accepts a delta:

delta Fires when the value
"changed" / "unchanged" moved, or didn't
"decremented" / "incremented" went down / up — with optional by = n
"zeroed" became zero
"set" became a specific to = value

delta = "decremented", by = 1 is far more selective than "went down", and it costs nothing because the comparison already exists.

The hunt

Function Behaviour
p:hunt(opts) Start a candidate set over a range → hunt handle. opts: from, to, delta, prune.
h:fence([name]) Mark "now" — subsequent tests are relative to this point.
h:keep(delta [, opts]) Keep only candidates that changed this way since the fence.
h:drop(delta [, opts]) Remove candidates that changed this way since the fence.
h:list() · h:count() Candidates: { addr, hits, pc, pcs, last_value, scanline }.
h:group() Merge adjacent addresses that always change together — multi-byte values.
h:snapshot([name]) · h:restore(name) Save and restore the candidate set — see over-pruning below.
h:reset() · h:stop() Start again; release the tags.

Self-pruning — why tagging all of RAM is cheap

prune = true (the default for a hunt) means a candidate that breaks the current rule clears its own tag bit, in hardware, on the access that disqualifies it. No event is queued, no software is consulted, and that byte never costs anything again.

This is what makes the opening move viable. Tagging every byte of RAM in write mode would otherwise generate an event on every write in the machine — thousands per frame, straight into an overflowing queue. With pruning, the surviving set only ever shrinks, so the event rate falls away as the hunt proceeds: the first pass sheds most of RAM within a second or two of play, and by the second fence the tagged set is small enough to be free.

It also makes the negative pass the cheapest one of all. h:drop("changed") disqualifies on any write, so a candidate dies on the first store to it — which is the fastest possible disproof.

h:count() therefore falls as you play, and p:stats() gains candidates_pruned so you can watch it happen.

Over-pruning is the one way to lose work. A rule that is too strict destroys candidates that hardware cannot give back — the tag bit is gone. So take h:snapshot() before a pass you are unsure of; the candidate list is held in software as well as in the tags, and h:restore() re-tags from it. The tags are the fast filter; the list is the durable record.

Outside a hunt the same mechanism is available as a rate: rate = "once" fires a tag exactly once and then removes it, which is the right setting for "tell me when this routine is first reached" and costs nothing thereafter.

local h = p:hunt{ from = 0x0000, to = 0x1FFF }

h:fence()                                   -- play until you die
h:keep("decremented", { by = 1 })           -- pass 1

h:fence()                                   -- play a while WITHOUT dying
h:drop("changed")                           -- kill timers and counters

h:fence()                                   -- die again
h:keep("decremented", { by = 1 })           -- pass 2

for _, c in ipairs(h:list()) do
  print(("%04X  hits=%d  written by PC %04X"):format(c.addr, c.hits, c.pc))
end

Four discriminators, in order of how much work they do

The naive set — "everything that decremented" — is large. Timers, loop counters, falling sprite Y positions, DMA pointers and frame counters all decrement. What separates them:

  1. h:drop("changed") between events. A lives counter does not move while you are playing normally. This single step usually removes most of the set, and it is the step people skip.
  2. Hit count. Lives change a handful of times per game; a timer changes sixty times a second. h:list() gives you hits — sort by it.
  3. PC stability. A lives counter is written by one instruction; a loop variable is written from many. c.pcs is the count of distinct writers, and pcs == 1 is a strong signal.
  4. Decrement by exactly 1. Cheap, and rules out pointers walking by 2 or 4.

Level counters answer to the same machinery with delta = "incremented" — though a level is often assigned rather than incremented, so "changed" with a low hit count is the more reliable hunt. Score is the awkward one: it changes constantly, spans several bytes and is frequently BCD, so start with h:group() and expect to work from the top byte down.

From candidate to trainer — and the exception to rule 1

Once you have the address and its writer, the trainer is one call:

p:suppress_writes(0x00C3, 0x00C3, { when = { pc = 0x8A17 } })

That is the whole thing: the instruction that decrements lives is stopped from taking effect, and nothing else that touches the address is disturbed.

This is the one case where write suppression is clearly right, against the warning in the substitution section. The usual danger is that the guest reads back what it wrote and finds something else — but here the value it reads back is exactly the value it had before the write, which is precisely the intent. It is also what a hand-written trainer has always done — NOP the DEC — except that nothing is patched, so it toggles at runtime and survives a checksum.

A trainer is therefore a short Luau script with no build step: hunt once, write the address and PC into a profile, and ship the profile.


Substitution — serving the guest different bytes

This is the one capability here that can crash the guest. Everything above observes; substitution lies. It is gated behind an explicit opt-in, and it is worth reading the three rules below before the function tables.

Where tagging lets you add to a guest, substitution lets you replace: a tagged read returns a byte you supply instead of the one in memory. The stored image is never modified, so nothing is patched, nothing is permanent, and turning it off restores the original exactly.

Three rules that shape the whole API

1. Substitute reads, not writes. The instinct is to suppress the guest's drawing and draw your own. Don't — the guest usually reads back what it wrote, and suppressing writes desynchronises it from its own state. Instead replace the data it reads: swap the tile and sprite graphics it fetches, and the guest draws your artwork using its code, at the right time, in the right place, with its own logic intact. Write suppression exists below for the cases that need it, and should be the second thing you try.

2. A substitution is a buffer, not a callback. The read path cannot wait for your program — a round trip is microseconds and the guest's bus is not going to pause for it. So you register a buffer that hardware serves from, and update that buffer whenever you like, asynchronously. There is no function(addr) return byte end, and there cannot be.

3. A self-checking guest will notice — a good deal of this software checksums its own data. The escape is when.pc: substitute only when the renderer reads the address, and serve the original when the checksum routine reads it. Both are true at the same address, distinguished by who is asking.

Enabling it

Function Behaviour
probe.substitution_available() Whether the platform build supports it at all.
p:enable_substitution(true) Explicit opt-in for this target. Off by default, always.

Defining substitutions

Function Behaviour
p:substitute(from, to, opts) Serve opts.data for reads in [from, to] → substitution id.
p:unsubstitute(id) Remove it; the guest sees real memory again from the next access.
p:substitutions() What is currently defined, with hit counts and enabled state.

opts fields:

Field Meaning
data The replacement bytes — a string or buffer. Must cover the range.
mask Optional per-byte mask: 1 = substitute, 0 = pass through. Patch one value inside a range without copying the rest.
when Restrict when the lie is told — see below. Omit to substitute for every read.
commit "vblank" (default) or "now".
enabled Start active. Default true.

when — the conditional half

Field Meaning
pc / pc_range Only when the read comes from this code. The answer to self-checking guests.
scanline Only during a scanline range — e.g. substitute for the playfield, not the status bar.
frame Only on frames matching a predicate, for alternation or fades.
p:enable_substitution(true)

-- New tile art, but only when the renderer fetches it.
-- The checksum routine at 0x2100 still reads the original and still passes.
local id = p:substitute(0x4000, 0x5FFF, {
  data = io.open("D:/mods/tiles.bin"):read("a"),
  when = { pc_range = { 0x8000, 0x81FF } },
})

Updating live

Function Behaviour
p:sub_update(id, offset, bytes) Replace part of a substitution's buffer.
p:sub_enable(id, on [, commit]) Turn one on or off.
p:sub_swap(id, data [, commit]) Exchange the whole buffer — for animation or a level change.

commit defaults to "vblank" throughout, because changing what memory says mid-frame tears exactly as you would expect. "now" exists for when the guest is halted, or when you know the region is not being read.

Write suppression — last resort

Function Behaviour
p:suppress_writes(from, to [, opts]) Writes in range are discarded rather than stored → id.
p:unsuppress(id) Stop.

The guest will read back what it thinks it wrote, and get something else. For anything holding state that is a corruption bug you have introduced deliberately. It is defensible for genuinely write-only regions — a hardware register window the guest never reads — and hard to justify anywhere else. If you are reaching for this to stop the guest drawing something, go back to rule 1: replacing the source data is almost always the better route.

Statistics

p:stats() gains substitutions_served and substitutions_passed — the second being reads that matched a range but failed their when. If a pc-conditional substitution is not working, that pair tells you immediately whether the condition is too narrow or the range is wrong.


Instruction codes — four bits instead of one

On an instruction fetch the CPU reads several bytes at once, so several tag bits arrive together. A 4-byte instruction therefore carries a 4-bit code rather than a flag — sixteen meanings, decoded from a sixteen-entry register file with no memory access at all.

Data accesses do not get this: the guest chooses their width, so the data side stays one bit plus the table. Instruction fetches do, because the decoder knows the width.

The codes

The rule, for any instruction width:

The code is the instruction's tag bits — lowest address as bit 0 — padded on the high side to the code width: with 1s if any bit is set, with 0s if none is.

So the reachable set grows with the fetch, and narrower widths reach a subset of wider ones. Nothing is ambiguous, and codes are allocated by the narrowest instruction that must carry them:

Code Needs at least Meaning
0000 1 byte nothing
1111 1 byte consult the table — the full per-address behaviour
1101, 1110 2 bytes the scarcest inline codes — compressed RISC-V, 2-byte 68000
1001–1100 3 bytes four more
0001–1000 4 bytes eight more

This subsumes the one-bit scheme rather than replacing it. A single-byte fetch — Z80, 6502, any narrow opcode — sets one bit, the rule fills the rest, and the code is 1111: consult the table, exactly as everywhere else in this API. Clear the bit and it is 0000. 8-bit and variable-length guests need no special handling and no separate mode.

The parameter — bytes 5 to 12

The code is always 4 bits. On instructions of five bytes or more, the bits above it are a parameter: up to 8 bits, capped there, with tags beyond byte 12 taking no part in the fetch.

Instruction bytes Parameter bits
1–4 none (reads 0)
5 / 6 / 8 / 10 1 / 2 / 4 / 6
12 or more 8

The code is one-padded; the parameter is zero-extended. Deliberately opposite. One-padding the code is what makes narrow fetches unambiguous; the parameter is a value, so an absent one must read 0 — one-extending it would make every short instruction report parameter 255 and select the wrong counter.

A code's spec takes channel = "param" or bank = "param" to read the value from the instruction instead of a fixed number, and the parameter is carried into events as e.param.

This is what makes 16 codes feel like plenty: one code meaning count with channel = "param" gives 256 counters; one meaning event gives 256 channels.

Where there is no parameter

Nothing under five bytes has room. RISC-V never has one — four bytes maximum — and neither do the Z80, 6502 or 65816. It is a wide-CISC feature that the 68000 and x86 get.

On those targets, spend codes instead: 0001–1000 as "count into counter 0…7" buys eight counters for eight codes, or fall through to 1111 and the table for unlimited ones at the cost of a memory access. Code assignment is per-architecture — a 68000 profile and a RISC-V profile should not expect the same allocation, which is why define_code is per target.

Function Behaviour
p:code_config() Report the code width, parameter cap and the widths this target actually produces.

Setting codes

Function Behaviour
p:code(addr, code [, opts]) Write an inline code at an instruction. opts.width = 2 or 4; omit and the target's disassembler supplies it.
p:code_at(addr) Read back the code at an address, with the width it was written for.
p:uncode(addr) Clear it.
p:code_map(list) Apply many at once — the output of a disassembly pass.
Function Behaviour
p:define_code(code, spec) Bind one of the sixteen codes to a behaviour for this target.
p:codes() The current sixteen definitions.

spec takes: { op = "none" | "count" | "event" | "break" | "arm" | "disarm" | "table", channel = n, bank = id }. channel and bank may be "param" to read the value from the instruction's parameter bits rather than a fixed number — see below.

-- A whole-guest instruction profiler with no event traffic at all.
p:define_code(0x1, { op = "count" })
p:code_map(disasm.every_branch("D:/roms/game.bin"))
-- ...play...
for addr, n in pairs(p:counters()) do print(("%04X %d"):format(addr, n)) end

op = "count" is the one that earns its place fastest: it increments a counter and continues, producing execution counts for every tagged instruction with no queue, no drops and no watcher — a complete profiler for a guest whose source you do not have.

Two things to get right

  • Codes come from a toolchain, not a range. Placing a 4-bit code needs the instruction's start and width. Tag the second byte of an instruction believing it a boundary and you write a wrong code, silently. p:code_map() takes disassembler or assembler output for exactly this reason.
  • Compressed-heavy code has two inline codes, not twelve. Assign 1101 and 1110 to whatever matters most there; everything else in compressed code goes through 1111 and the table.

Maps from your own toolchain

If you are developing for the guest rather than working on somebody else's binary, the map is a build artifact: your assembler knows every instruction's address, width, symbol and source line exactly.

Function Behaviour
probe.load_map(path) Load a toolchain-emitted map — addresses, widths, codes, symbols, source lines.
probe.map_from_listing(path [, opts]) Derive one from an assembler listing file, which is what most retro toolchains actually emit.
p:apply_map(map) Place every code and tag it defines. Refuses on binary-hash mismatch — see below.
p:resolve(addr) { symbol, offset, file, line } for an address, from the loaded map.
p:map() The map currently applied.
local m = probe.load_map("build/game.probemap")
p:apply_map(m)

p:on("hit", function(e)
  local s = p:resolve(e.addr)
  print(("%s+%d  (%s:%d)"):format(s.symbol, s.offset, s.file, s.line))
end)

apply_map refuses a map whose binary hash does not match what is loaded. Relink and every address moves; yesterday's map would tag the middles of instructions — the silent-wrong-code failure everything else here is built to avoid. This is a refusal, not a warning.

Breakpoints placed this way do not alter the instruction stream. A conventional breakpoint patches a trap into your code and changes its timing, which makes it useless in raster chases, copper-synced effects and cycle-counted loops — precisely the code that most needs debugging. A tag breakpoint leaves the fetched bytes byte-identical. That, plus unlimited watchpoints and free execution counters, is more debugging capability than the original hardware ever had.

(For native AntOS and FireStorm development you have a real toolchain and GDB already — see develop. This is for software targeting a guest.)


Actions — writes that fire with the tag

A tag can carry an action: an ordered sequence of writes applied the moment it fires, in hardware, with no round trip to your program.

That is the whole point of it. An event delivered to a watcher costs a PCIe crossing each way — fine for compositing a backdrop, useless for anything that must land inside the guest's own timing. An action is pre-loaded beside the tag, so the effect is immediate and deterministic, and the watcher can be told about it afterwards or not at all.

Defining one

Function Behaviour
p:action(steps [, opts]) Build an action from a list of steps → action id.
p:action_update(id, steps) Replace the steps.
p:action_enable(id, on) Arm or disarm without deleting.
p:action_run(id) Fire it now, by hand.
p:action_undo(id) Restore the bytes it overwrote — requires snapshot.
p:actions() What is defined, with fire counts.

Attach it to a tag with action = id:

local refill = p:action({
  { addr = 0x00C3, op = "set",  value = 3    },   -- lives
  { addr = 0x00C4, op = "or",   value = 0x80 },   -- invulnerable flag
  { addr = 0x0200, op = "copy", from = 0x3000, len = 16 },
}, { snapshot = true })

p:tag(0x8A17, { on = "exec", action = refill, event = "death_handler" })

Step operations

op Effect
"set" write value
"and" · "or" · "xor" bitwise against value
"add" · "sub" arithmetic, with wrap or saturate
"copy" copy len bytes from from
"fill" write value across len bytes
"restore" put back what snapshot captured
"arm" · "disarm" enable or disable another tag or action

Guards, timing and one-shot

opts on the action, or if on an individual step:

Field Meaning
if { addr, cmp = "eq"\|"ne"\|"lt"\|"gt", value } — apply only when this holds
at "now" (default), "vblank", or { scanline = n }
once fire once, then disarm
snapshot capture the original bytes first, so action_undo works
silent default true — see below

Three rules the hardware imposes

1. Action writes are silent by default. A write performed by an action does not itself trigger tags. Without that, an action touching a tagged byte re-enters, and you have built a feedback loop that runs at memory speed. Set silent = false deliberately, for chaining, and a depth limit still applies.

2. Sequences are short, and the cap is real. The steps execute in the memory path, so every step costs the guest cycles. This is for a handful of writes — refill a counter, set a flag, swap a pointer — not for moving a screen. Use a substitution buffer for bulk data; that costs nothing per access.

3. Prefer at = "vblank" when the guest may be reading. Applying mid-frame can let the guest observe a half-written sequence. "now" is right when you are reacting to a write the guest has just made and will not re-read this frame — the trainer case — and wrong for anything it is actively scanning.

arm / disarm — the useful part

Because a step can enable or disable another tag, actions compose into a small state machine that runs at guest speed with your program entirely out of the loop:

-- Arm the boss-fight tags only once the boss actually spawns,
-- so they cost nothing for the rest of the game.
local begin_boss = p:action{ { op = "arm", tag = boss_tags } }
p:tag(BOSS_SPAWN_PC, { on = "exec", rate = "once", action = begin_boss })

This is how an enhancement stays cheap: tags exist but stay disarmed until the moment they are relevant, armed by the guest's own execution rather than by polling from outside.

p:stats() gains actions_fired and actions_skipped — the second counting guard failures, which is the first thing to check when an action seems not to run.


Statistics — read these

Function Behaviour
p:stats() { hits, events_delivered, events_dropped, queue_depth, queue_high_water }

events_dropped is not decoration. A dropped event that nothing checks produces an enhancement that works on the bench and glitches in play. If it is non-zero, a tag needs a rate.


Related

Important: The Ant64 family of home computers are at early design/prototype stage, everything you see here is subject to change.