AntOS Library — probe (reference)

The probe library's design and native core — tagged-memory breakpoints on a running simulated system, and the event stream they produce. The script-facing API is luau_probe. The mechanism and its rationale are in probe.

What it is

probe is how an AntOS program watches a guest running on FireStorm and acts in sync with it. The Ant64's memories are 36 and 72 bits wide — nine bits per byte — and the spare bit is spent as a tag. Any tagged byte raises an event when it is read, written or executed; probe sets those tags, owns the table that says what each one means, and delivers the resulting events to the watching program.

The purpose it exists for is non-invasive enhancement: adding backdrops, sprites, audio or overlays to software whose code is never altered. It is also, with one flag changed, a debugger with an unlimited number of watchpoints.

Design notes

  • Trace is the default and break is the exception, because the primary use case cannot tolerate a stall. A trace event is emitted asynchronously and the guest runs on with its timing intact; a break halts the guest until released, which is correct for debugging and fatal for enhancement. The two share the tag bit and differ only in the table entry, but a break additionally needs a path back to the guest's clock enable — so a build without break support is a legitimate, smaller configuration.
  • Instruction fetches and data accesses are different mechanisms, and the reason is width. A CPU fetching a 4-byte instruction reads four tag bits at once, so the instruction side gets a 4-bit code — sixteen meanings decoded from a sixteen-entry register file, costing no memory access. A data access has whatever width the guest chose, so nothing can be assumed and the data side keeps one bit plus the table. The code is therefore assembled in the fetch/decode path, after instruction boundaries are known, not in the memory controller — the two sit in different parts of the design and should be built as such.
  • One padding rule covers every instruction width, and the reachable code sets nest. The code is the instruction's tag bits, lowest address as bit 0, padded on the high side to the target's code width — with 1s if any bit is set, 0s if none is. Real widths are 1–3 on a 6502, 1–4 on a Z80, 2–10 on a 68000 and 1–15 on x86, so this generality is not theoretical. The property that falls out: a code reachable at width w is reachable at every width above it, nothing is ambiguous, and codes are allocated by the narrowest instruction that must carry them — 1111 from one byte, 1101/1110 from two, 1001–1100 from three, the rest full-width. Two consequences to design to: 1111 must be consult the table, since it is the only code a single-byte fetch can produce, which makes the encoding a strict superset of the one-bit scheme requiring no special case for 8-bit or variable-length guests; and bit order is load-bearing — byte 0 as bit 0, with the bytes a narrow instruction lacks as the high bits — which is exactly the sort of thing implemented backwards once.
  • Fixed at a 4-bit code plus a parameter of up to 8 bits. Bytes 5 through 12 of a wide instruction carry the parameter — a small immediate selecting which counter, channel or bank the code acts on — and tags beyond byte 12 take no part in the fetch, staying available for data-side use. The two fields pad by opposite rules and must not be unified: the code is one-padded, because that is what disambiguates narrow fetches, while the parameter is zero-extended, because it is a value and an absent one has to read 0 — one-extending it would make every short instruction report 255 and address the wrong counter. The payoff is that sixteen codes go a long way: one code meaning count with a parameter-selected channel gives 256 counters, and the code width stays at 4 so it continues to index a register file rather than memory.
  • Parameters require five-byte instructions, so several targets never have one. RISC-V tops out at four bytes; so do the Z80, 6502 and 65816. It is a wide-CISC feature the 68000 and x86 get. On the others the fallback is to spend codes — 0001–1000 as counters 0–7 — or to fall through to the table and pay a memory access. Code assignment is therefore per-architecture, not a fixed platform convention, which is why the sixteen definitions are configured per target rather than baked in.
  • Instruction-side tag maps must come from a toolchain, since placing a code needs each instruction's start and width; and narrow-instruction code has few inline codes, so most enhancement there stays table-driven.
  • The tag is a filter, not a description. One bit cannot say why a byte matters and does not try to: it says "look this up". The table carries the class, the access mask, the handler id and the rate limit. Everything expensive is paid only on a hit, which is what makes tagging ten million bytes cost the same as tagging one.
  • A tag may carry an action — a short sequence of writes applied when it fires. This exists to close the latency loop: an event costs a PCIe crossing each way, which is fine for compositing and useless for anything that must land inside the guest's own timing. Pre-loading the writes beside the tag makes the effect immediate and deterministic. Three constraints follow and all three are hardware's, not policy:
    • Action writes are silent by default — they do not themselves trigger tags. Otherwise an action touching a tagged byte re-enters and builds a feedback loop running at memory speed. Non-silent chaining is opt-in and depth-limited.
    • Sequences are short and hard-capped, because the steps execute in the memory path and every one costs the guest cycles. Bulk data belongs in a substitution buffer, which costs nothing per access; actions are for a handful of writes.
    • Commit timing matters — applying mid-frame lets the guest observe a partly-written sequence, so "vblank" is available and "now" is correct mainly when reacting to a write the guest has just made.
  • arm / disarm steps make actions a state machine. A step that enables or disables another tag lets an enhancement keep its tags disarmed until the guest's own execution reaches the point where they matter — armed by the guest, not by polling. This is the main mechanism for keeping a large profile cheap, and it is why the step vocabulary includes tag control rather than only memory writes.
  • Tags can clear themselves, and that is what makes whole-RAM tagging affordable. A table entry may carry a rule; an access that breaks the rule writes the tag bit back as zero in the same path, with no event queued and no software consulted. The consequences are worth stating because they invert the usual cost model: the tagged set only ever shrinks, so the event rate during a search falls away instead of staying flat, and tagging every byte of a guest's RAM becomes a reasonable opening move rather than a denial of service on the queue. It also supplies the general rate = "once" behaviour for free. The cost is that pruning is destructive and hardware cannot undo it — hence the candidate set is mirrored in software so a set can be restored and re-tagged after an over-strict pass.
  • Value comparison earns its keep twice. The table already compares old against new to implement rate = "change", so classifying the direction of the change — decremented, incremented, zeroed, set-to-N, optionally by an exact amount — costs almost nothing on top. That is what makes hunting a value a hardware filter rather than a software scan over a large candidate set, and it is why the delta classes are in the tag options rather than in the library.
  • Rate limiting is hardware, deliberately. A tag inside a copy loop fires tens of thousands of times a second. Filtering to once-per-frame or on-change-of-value in the table costs nothing; doing it in the watcher costs a round trip per hit and would make an otherwise reasonable tag unusable. This is the single most important thing the table does beyond dispatch.
  • Events carry beam position. Scanline and pixel are in every event, not optional, because compositing into a live frame depends on knowing where in the frame the event happened. It is the field that would be cheap to add now and painful to retrofit.
  • Tag memory is per-region, and tagging costs ECC. The ninth bit cannot be both a tag and a check bit. Regions hosting a guest give up ECC willingly; regions holding AntOS or FireStorm state may not, so the capability is expressed per region rather than globally. Assigning a region clears its tags — a stale tag map from a previous guest is a confusing class of bug and the default should not permit it.
  • Overflow is reported, never silent. The event queue can fill. A dropped event that nothing counts produces an enhancement that works on the bench and glitches in play, which is the worst available failure, so drops are counted and surfaced through stats().
  • Read substitution is a separate capability, not a flag on a trace tag. Returning a different byte to the guest than the one in memory is a lie told to the guest, and its failure mode is a crash rather than a missing sprite. It is off by default and requires an explicit per-target opt-in. Three consequences shape its implementation:
    • A substitution is a buffer, not a callback. The read path cannot wait for a program on the other side of PCIe — microseconds against a guest bus that will not pause — so hardware serves from a pre-loaded buffer and the watcher updates that buffer asynchronously. A callback-shaped API would be unimplementable at speed and is not offered.
    • Substituting reads beats suppressing writes. Swapping the graphics data a guest fetches leaves its own logic, timing and state intact and lets its code draw the new artwork; suppressing its writes desynchronises it from state it will read back. Write suppression exists for genuinely write-only regions and is documented as a last resort.
    • when.pc is what survives a self-checking guest. Much of this software checksums its own data. Conditioning a substitution on the reading PC serves new artwork to the renderer and the original bytes to the verifier — the same address telling two truths, distinguished by who asked. That requires PC comparison in the tag hardware, which is the main cost substitution adds beyond the buffer itself.
    • Changes commit at vertical blank by default, since altering what memory says mid-frame tears.

Profiles and maps — two sources, one input

A tag map is data — address, width, code, access, rate, handler — loaded from a file rather than compiled in, and it arrives from one of two directions.

From the outside, a profile is the result of understanding somebody else's binary: hunted, disassembled, annotated. That makes it a shareable artifact — mapped once, published, applied by anyone — and it ships alongside a game the way a personality manifest does.

From the inside, a map is a build output of the guest's own toolchain. An assembler knows every instruction's address, width, symbol and source line with certainty, so it can emit the map rather than leaving anyone to recover it — and most retro toolchains emit a listing file, so deriving a map from a listing matters more in practice than supporting a debug format. This is the case that turns probe from an enhancement tool into a development tool for the guest, and its headline property is that a tag breakpoint leaves the fetched instruction stream byte-identical: unlike a patched trap, it can be used in timing-critical code without disturbing the timing.

Both land on the same input, so the library does not distinguish them beyond provenance. The one asymmetry that matters is safety: a map carries the hash of the binary it was generated for, and applying it to a different one is refused rather than warned about. Addresses move on every relink, and a stale map tags the middles of instructions — the silent-wrong-code failure the four-bit encoding is otherwise built to avoid.

Native core — libantos_probe

Plain C on DeMon, talking to the FPGA's tag and event hardware over PCIe: tag writes, table management, and draining the event FIFO into subscriber callbacks. The latency-sensitive path is the drain, so the core keeps the queue mapped rather than copying, and delivers in batches per frame rather than per event.

Where the watcher itself should run — DeMon, FireStorm, or either — is an open question in probe; the binding is the same on both sides.

Related

Important: The Ant64 family of home computers are at early design/prototype stage, everything you see here is subject to change.