ATESO LABS // RESEARCH
Battlegrounds War Room The Overseer All Papers Enterprise
← Back to Directory Download Formal PDF ↓

Capability-Addressed Binary State and Zero-Heap Resident Memory Vectors for Hard Real-Time Physics and Closed-Loop Agentic Coordination

Brennan DeCrow
manymoats Research / ATESO Labs
Independent Academic Preprint — arXiv cs.OS / cs.DC / cs.AI
Research Whitepaper
Date of record: September 21, 2026
Revision: v2 — Day-100 fortification (September 21, 2026)

Author contribution. Brennan DeCrow is identified by the supplied project corpus as the architect and principal author of the ATESO semantic-orchestration architecture and Magma resident-state runtime described here. This manuscript formalizes those supplied designs while separating theorem, independently reproducible calculation, author-reported measurement, vendor-verified hardware metadata, and proposed experiment.

Evidence notation. F denotes a formally proved result under explicit assumptions; R denotes a result reproduced from an executable artifact available in this work; U denotes an author-reported result for which the underlying raw trace or original benchmark implementation was not present in the supplied reproduction package; V denotes a device/specification fact independently checked against a primary vendor or standards source; and P denotes a proposed or not-yet-observed experiment. This distinction is part of the scientific claim, not a disclaimer to be discarded in later versions.

Revision note (v2). This revision adds explicit section numbering and integrates the Day-100 fortification review. The additions are: the resident-state density and distributed capacity bound (§2.2, including the 812.5-million-record capacity identity and the 156:1 retained-size compaction algebra); the publication-safe statement of the locality result (§2.3); the sparse-delta transport bound and its energy caveat (§2.4); the research-egress plane for programmatic notebook synthesis (§4.6); a claim-status register that assigns an evidence class to every numerical claim introduced or sharpened in this revision (§5.10); the contemporary agent-state compaction discussion (§7.6); reproducibility contracts for the distributed-capacity, V8 retained-size, sparse-delta wire, and research-egress experiments (Appendix A.1–A.4); and references [49]–[60]. The review that produced these additions addressed them as Sections 4.1.1, 4.4, 8.4, and Appendix A.1–A.4 and as its references 32–41 (that draft’s numbering) against an earlier draft; the mapping to this revision is §2.2, §2.4, §4.6, Appendix A.1–A.4, and [49]–[58] respectively ([59] was added when the static-copy statement was traced to a separate Google page; [60] documents the FedNow message set). This revision also: corrects the machine-of-record metadata in §5.2 (the machine’s hardware report places it in Apple’s 32-core-GPU M5 Max bin, documented at 460 GB/s, not 614 GB/s [57]); treats the vendor memory figures as binary quantities on the evidence of the machine-of-record hardware report (the vendor sheet carries no unit qualifier) and reports both the decimal-convention and binary capacity bounds (§2.2); binds the research-export digest to the state digest and provenance root (§4.6, A.4); assigns evidence classes to the §6.1 throughput figure and the §6.4 netting magnitude and restates §6.5 as a non-custodial design with its regulatory scope left to counsel; escapes a percent sign inside the §2.3 terminology box that would not have compiled; and adds the Appendix A.5 arithmetic harness. Two of the review’s instructions were executed as insertions rather than replacements, deliberately: the v1 abstract paragraphs that label fleet measurements U and describe the executable artifact were retained beside the new abstract paragraphs, and the v1 operating-model box and closing sentence were retained after the new conclusion paragraphs. Smaller changes: the keyword list is extended; §1 gains a contemporary-context paragraph; the §6 opening states card-rail fee figures qualitatively as design premises; §6.1 states its per-shard rate as a specification floor (P), gives the ingestion topology under which Theorem 8 applies, and labels the 580,000 TPS and enqueue-latency figures U; §6.2 and §6.3 tighten the event-digest and checkpoint constructions (distinct domain tags, the signature inside the preimage, RFC 9162-style leaf/node separation, epoch and predecessor links); §6.4 withdraws the directive-supplied 70%–99% netting figure pending measurement; §6.5 states the pacs.008 message version-agnostically and cites the FedNow message set [60]; the pipeline diagram is re-aligned; §7.7 writes its dimensionless factors in words; Corollary 1 is renamed Corollary 1.1 for consistency. The v1 closing section was split: its conclusion is now §8, its reproduction package and executable listings are Appendix A.0 (ahead of the new A.1–A.5), and the reference list stands alone. §6.3 no longer calls anchored receipts immutable and states tamper evidence relative to the log; §6.5 is retitled “Regulated Instant Settlement Handoff” and now has the debtor’s institution, not the engine, originate the interbank message; §6.1 gives the CBP-96 reference header layout, names it a lease-referenced command envelope, and routes agent submissions through the §3.6 untrusted path into a bounded ingress-worker pool; §6.2 and §6.3 define their genesis values; the §4.6 export capability is classed P pending a reported run; the organisation name follows the house’s lowercase form; and prose hooks were added for eight reference entries that v1 listed without citing. The CBP-96 profile now states its sequence-admission rule and binds the resolved sender key and lease digest into the §6.2 preimage; §6.5 confines the engine to customer-to-bank instructions on both the debit and the request-for-payment routes; the §6.3 checkpoint follows RFC 9162 [22] for its tree construction, with textual tags in place of the RFC’s byte prefixes; and the binary-unit basis of the memory figures is attributed to the machine’s hardware report wherever it is used. No theorem or measurement was weakened; three §6 figures that carried no evidence class were given one and one directive-supplied figure was withdrawn pending measurement. Where this revision supplies a stricter wording for a result, the stricter wording governs.

Abstract

Long-lived interactive, simulation, and agentic workloads repeatedly transform a comparatively stable state while only a small fraction of that state changes between externally admitted operations. Conventional application pipelines can nevertheless reconstruct equivalent state through text parsing, transient heap graphs, framework objects, staging arrays, and device transfers. This paper studies the operating-system consequences of making resident state, rather than reconstructed documents, the primary execution object.

We present ATESO, a semantic admission and orchestration architecture, and the Magma Runtime, a resident-state execution substrate. ATESO divides its local contract into .many, a relocatable binary state image; .moat, a cryptographically authenticated capability and revision envelope; and .spine, a bounded shared-memory exchange fabric. The architecture further separates a deliberative macro-timescale planner from a deterministic, allocation-free execution path; compiles high-level actions into capability-addressed sparse-delta opcodes; applies validated updates to preallocated state; and returns typed telemetry through the same resident representation. The proposed model deliberately does not require an LLM itself to parse or directly mutate raw memory in a hard-real-time loop: arbitrary text terminates at a non-real-time compilation and admission boundary.

For (N) resident records of (b) bytes and (k_t) identified changes on update (t), we prove a representation-maintenance bound

[ Q_{}U N b, Q_{} = O!( Nb+_{t=1}^{U}(k_tb+h_t) ), ]

where (h_t) includes validation, authentication, indexing, and provenance work. The result isolates precisely the regime in which repeated full reconstruction is asymptotically redundant.

We then formalize a spatial-locality result. For (N=50{,}000), (b=64) bytes, and 4 KiB logical tracking granules, there are 782 granules. Independent Bernoulli mutation with () produces

[ E[D_{}] = 781(1-0.95{64})+(1-0.95{16}) = 752.2535206. ]

A specified five-interval layout containing 2,500 changed records touches exactly 44 granules, yielding

[ 1- = ]

expected logical dirty-granule suppression. We prove the general clustered bound

[ D_{} = O!(+C), ]

where (m) is records per tracking granule and (C) is the number of disjoint mutation intervals. Crucially, 94.15% is not a universal physical MMU, TLB, DRAM, or memory-controller invariant. Linux, for example, explicitly distinguishes virtual-page write tracking through soft-dirty PTE state from other memory-writeback mechanisms. Thus the theorem predicts the cardinality of a logical affected-page set; physical writeback and energy suppression require hardware-counter experiments at separately stated observation boundaries.

We prove capability non-amplification for monotonically restricted delegation, stale-revision exclusion for serialized commits, SPSC slot non-overlap for bounded monotonic ring counters, and conflict freedom for properly colored XPBD constraints. We derive the standard lumped thermal RC model and a robust temperature-admission condition. We also show why Landauer’s principle cannot be used to claim that a lossless binary-to-text transformation intrinsically destroys semantic information: for an injective encoding (C), (H(C(X))=H(X)); the practical thermodynamic opportunity is reduced computation and data movement, not a direct Landauer calculation. Landauer’s original result concerns irreversible logical operations, while reversible-computation work makes the logical/physical distinction explicit.

The supplied ATESO corpus reports measurements spanning Apple Silicon, Xbox Series X, Nintendo Switch, mobile SoCs, browsers, smart appliances, and microcontrollers. It includes a reported 0.042 ms smart-TV binary-delta ingestion result, a 0.3795 ms Apple-Silicon XPBD step, and an Xbox timer-floor ratio represented as (40.40/0.001=40{,}400). The independently inspectable reproduction package accompanying this paper, however, contains the mathematical audit and a new SPSC stress test rather than the original WebGPU executable, heterogeneous-device traces, or signed fleet receipts. Accordingly, this paper reports fleet measurements as U, not as independently reproduced facts.

At the representation boundary, the architecture additionally admits a 64-byte-aligned packed entity profile whose capacity scales exactly as (N=M/64) before system overhead. A two-host Apple Silicon configuration comprising 36 GB and 16 GB systems provides 52 GB of nominal aggregate memory, corresponding to a decimal-convention capacity of 812.5 million packed records (872.4 million in binary units, per the machine-of-record hardware report, §2.2); the hosts remain independent address spaces and are treated as sharded resident-state nodes rather than a fictitious cross-machine UMA. An author-reported (U) retained-size ratio of (156{:}1), if confirmed against a 9,984-byte-per-entity object-graph baseline (Appendix A.2), corresponds to a 99.359% reduction and to an 8.112 TB equivalent baseline at 812.5 million entities. That ratio is reported separately from practical 150–200 GB heap-envelope comparisons, which correspond to (2.88)–(3.85) memory reduction.

For a declared sparse-network scenario consisting of a 500 MB state image at 60 publications/s and one changed 64-byte record per publication, dense replication requires 30 GB/s of application payload while the sparse payload requires 3.84 kB/s, a 7,812,500:1 payload-volume ratio ((99.9999872%) suppression) before protocol and cryptographic overhead. These ratios describe representation and payload work, not equivalent measured reductions in electrical energy.

Contemporary long-context model research independently reinforces the importance of retained-state footprint: a contemporary long-horizon serving stack reports an 890-byte/token global KV cache and substantially reduced persistent KV state. ATESO applies the state-minimization principle at a different boundary: application and cyber-physical resident state. External research services are kept outside that execution boundary; provenance-bound snapshots may be programmatically exported into enterprise notebook systems for grounded research and generated overviews without representing cloud APIs as zero-copy shared memory.

The independently executable artifact verifies the 94.1509% locality calculation, exact 44-granule construction, 40- and 45-granule counterexamples, fixed-(K) occupancy model, timing arithmetic, and a shared-memory test covering 600,000 records and 9,600,000 payload words across three ring capacities with unsigned 32-bit counter wrap. The resulting contribution is an evidence-bounded systems architecture: resident state removes provably unnecessary representation work; capability narrowing limits delegated authority; static allocation removes allocator and garbage-collector activity from a qualified fast path; and physical claims remain tied to the hardware boundary actually observed.

Keywords: resident state; capability systems; zero-copy; zero-heap; spatial locality; WebGPU; XPBD; real-time systems; shared memory; SPSC; agent orchestration; thermal scheduling; provenance; resident-state density; sparse-delta transport; research egress; KV-cache compaction.

1. Introduction and the Thermodynamic Cost of Representation Maintenance

Section 1 — Problem statement. A large class of continuously interactive systems differs from request/response document processing [5] in one elementary respect: the state remains useful longer than any one view, command, request, tool, or model invocation. A physics workspace, robotic scene, media composition, dependency graph, industrial set point, or agent-managed resource may persist for minutes or days while each individual operation changes only a sparse subset of its state. Reconstructing an equivalent complete representation on each transition is therefore not logically required by the workload.

A common reconstruction path is

[ . ]

None of those stages is inherently pathological. High-performance JSON parsers can be extremely efficient [4], and binary formats such as Arrow, FlatBuffers, and Cap’n Proto already demonstrate direct traversal, relocatable representations, or access without constructing a second object graph. Arrow explicitly specifies relocatability without pointer swizzling and shared-memory zero-copy access; FlatBuffers documents direct access without unpacking and states that reading a buffer requires no additional heap; Cap’n Proto similarly defines a directly traversable structured encoding.

The narrower ATESO claim is therefore not “text is always slow” and not “binary layouts are novel.” It is:

[ ]

That claim can be tested against an optimized conventional resident implementation. A fair evaluation must include JSON-delta, existing binary-delta, direct-memory, and zero-copy baselines rather than compare a tiny binary command with an unnecessarily rebuilt document tree.

The hard-real-time qualification problem. The term hard real time is often used loosely. A 5-()s mean or 99th-percentile latency does not constitute a deadline guarantee. A hard-real-time profile requires a characterized or bounded worst-case execution time, a schedulability argument, bounded interference, and a complete path from release to completion. Linux’s SCHED_DEADLINE documentation, for example, explicitly frames real-time tasks through execution-time, period, and deadline parameters and uses admission control with EDF/CBS mechanisms. Microsecond-scale systems such as Shinjuku likewise obtain latency properties through explicit scheduling and preemption mechanisms rather than low average latency alone.

Accordingly, this paper uses hard-real-time eligible fast path to mean a path intentionally constructed from bounded parsing, preallocation, finite command sets, bounded loops, and explicit scheduling. The supplied (<10,) admission and mutation observations are retained as author-reported measurements; they are not promoted into a universal WCET theorem.

Contribution. The integrated contribution has six parts.

First, the paper defines a resident-state workload model and proves when incremental state maintenance asymptotically dominates repeated full materialization.

Second, it establishes an exact spatial-locality theorem for dirty logical granules, including the requested 94.1509% result, while separating that theorem from physical PTE modification, page-cache writeback, DRAM traffic, cache eviction, and GPU transfer.

Third, it defines a triple contract: .many for state, .moat for authority and provenance, and .spine for bounded local exchange.

Fourth, it formalizes a temporal split between a stochastic or deliberative planner and an allocation-free deterministic execution path, closing the loop through typed telemetry.

Fifth, it integrates conflict-free colored XPBD and thermally constrained scheduling into the resident-state model.

Sixth, it subjects the supplied heterogeneous benchmark corpus to an evidence audit instead of using device names as substitutes for raw traces.

Contemporary context. A contemporaneous independent example appears in a contemporary long-horizon serving stack, whose authors explicitly optimize long-horizon agent serving around KV-state capacity, bandwidth, reuse, and persistence, reducing reported global KV storage to 890 bytes/token and persistent KV storage to roughly one eighth of the preceding V4-Flash design [49]. This does not validate ATESO’s measurements; it independently reinforces the systems premise that representation size and state movement are first-order costs in long-lived agent workloads.

What “thermodynamic” means here. The word is used in an engineering rather than metaphysical sense. Every practical parse, allocation, cache miss, memory transfer, synchronization event, and retained byte can consume energy. Reducing unnecessary operations can therefore reduce energy under some workloads. But lossless representation conversion is not an information-theoretic annihilation of the represented state. Shannon entropy is preserved by a deterministic one-to-one transformation over the support of a random variable:

[ H(C(X))=H(X) ]

for injective (C). Shannon’s information theory and Landauer’s thermodynamic limit address different levels of description; Landauer [29] identifies a physical lower bound associated with logically irreversible operations, and Bennett [30] shows that logically reversible computation avoids that bound in principle; not a formula that permits the energy of JSON parsing to be calculated by multiplying file bits by (k_BT).

At (T=300) K,

[ k_BT ^{-21} J ]

per logically erased bit. The actual energy of a parser or memory subsystem is many architectural layers above that bound and must be measured as

[ E=_0^P(t),dt. ]

The thermodynamic research question is consequently: how much avoidable computation and movement does residency remove at each physical boundary, and what measured energy difference follows?

Evidence posture. The source corpus describes a broader ATESO/Magma dissertation and fleet evaluation. The reproduction package accompanying this manuscript (Appendix A.0) contains a new independent analytical audit and a new SPSC reference implementation, but explicitly does not contain the original WebGPU benchmark, heterogeneous-fleet raw traces, historical-device ports, original thermal scheduler simulator, or signed receipt bodies from the author’s receipt-and-event ledger convention (AiST). That boundary determines the empirical language used below.

2. Mathematical and Architectural Foundations

2.1 Workload model and the representation-maintenance bound

Section 2 — Workload model.

Definition 1 — Resident-state workload. A workload instance is

[ W= (N,b,{k_t}_{t=1}^{U},f,G,,C,B), ]

where (N) is the number of resident records; (b) is dynamic bytes per record; (k_t) is the number of records explicitly changed by admitted update (t); (f) is the update frequency; (G) is a task-dependency graph; () is the execution profile; (C) is the capability policy; and (B) is the physical resource budget.

Define external mutation density

[ _t=. ]

The quantity (k_t) denotes state changed by the admitted command. It is not necessarily the number of records touched by later physics, rendering, indexing, collision detection, or integrity maintenance.

Definition 2 — Representation-maintenance work. (Q) is the number of logical bytes that an implementation must read, write, copy, validate, or reconstruct solely to maintain the representation necessary for the next stage. It intentionally excludes common downstream work when comparing equivalent representations.

Theorem 1 — Resident representation-maintenance bound.
Assume:

  1. an (N)-record state image with record width (b) is initially admitted;
  2. unchanged records remain semantically valid;
  3. each update identifies its (k_t) changed records;
  4. admission bookkeeping for update (t) has cost (h_t);
  5. the comparator explicitly materializes a fresh (N)-record output on every update.

Then after (U) updates,

[ Q_{}U N b ]

while a resident implementation can satisfy

[ Q_{} = O!( Nb+_{t=1}^{U}(k_tb+h_t) ). ]

Proof. Every explicit full reconstruction produces a fresh output containing all (N) records. The output itself therefore entails at least (Nb) logical bytes of output work; summing across (U) updates yields (UNb).

For the resident implementation, initialization writes or admits the base image once, costing (O(Nb)). On update (t), unchanged records are retained. Only the (k_t) changed records plus the update’s validation, authentication, directory, provenance, and related maintenance (h_t) must be handled. Summing gives

[ O(Nb)+ O!( _{t=1}^{U}(k_tb+h_t) ), ]

which proves the result. ()

Corollary 1.1 — Sparse recurring updates. Suppose

[ _N ]

for all (t), with (_N), and

[ h_t=o(Nb). ]

Then

[ = O!( 1U+_N+ _t h_t ). ]

For long-lived state and sparse updates, the representation-maintenance ratio tends toward zero under those assumptions.

Proof. Divide Theorem 1 by (UNb):

[ = 1U+ 1U_t + . ]

Apply (k_t/N_N). ()

This theorem is stronger than a timing anecdote and weaker than a claim of universal application speedup. A full simulation step may still be (O(N)). A whole-image cryptographic hash can also restore (O(Nb)) work unless integrity is maintained incrementally.

2.2 Resident-state density and the distributed capacity bound

Representation-maintenance complexity and representation density are distinct quantities. The former determines how much state must be rewritten after a mutation; the latter determines how much state can remain admitted simultaneously.

Let (b_p) denote the packed resident-state width per entity and (M) the memory capacity assigned to resident records. Neglecting allocator, operating system, index, and replication overhead for the moment, the capacity bound is

[ N_{} = . ]

For the canonical Magma entity profile,

[ b_p=64 . ]

The reference two-host configuration for the distributed-capacity experiment specified in Appendix A.1 comprises a 36 GB Apple M5 Max system (model identifier Mac17,7; its own hardware report gives an 18-core CPU and a 32-core GPU, the bin Apple documents at 460 GB/s [57]) and a 16 GB Apple M1 Pro system [58], yielding 52 GB of nominal aggregate physical memory. The vendor figure carries no unit qualifier [57]; the machine of record’s own hardware report gives (=38{,}654{,}705{,}664) B (=36^{30}) B, so unified-memory figures are treated as binary quantities and the physical aggregate is (52{30}=55.83{9}) B. Under the conservative decimal capacity accounting adopted for the figure of record,

[ M_{} = (36+16)^{9} = 52^{9} , ]

and therefore

[ N_{}^{()} = = 812{,}500{,}000 ]

64-byte records.

This quantity is a capacity identity, not a claim that two machines form one cache-coherent 52 GB unified-memory domain. Each host retains its own address space and its own local unified memory. A distributed Magma deployment therefore partitions the global object identifier space into resident shards,

[ S = S^{(0)} S^{(1)}, S^{(0)}S^{(1)}=, ]

with

[ |S^{(i)}|,b_p M_i-M_i^{()}. ]

Cross-host transitions are explicit protocol operations. They are not zero-copy shared-memory accesses and are not described as unified memory over Wi-Fi, Ethernet, or Thunderbolt.

Under binary units, as the hardware report cited above indicates, the same configuration has the theoretical capacity

[ = 872{,}415{,}232 ]

records before system reservation. The decimal figure 812,500,000 is therefore a deliberately conservative convention (it is what 52 GB would hold if the vendor figures were decimal), and 872,415,232 is the nominal hardware bound. The manuscript reports both conventions and requires, for every physical residency experiment, the measured available resident-set limit to be recorded (Appendix A.1), rather than silently interchanging GB and GiB.

Representation compaction

Let (b_h) be the measured retained heap footprint per logical entity under a comparison representation, and define the representation compaction ratio

[ C_{} = . ]

The corresponding fractional memory reduction is

[ _{} = 1- = 1-. ]

For the author-reported ratio (U; reproduction contract in Appendix A.2) (C_{}=156),

[ b_h = 156(64) = 9{,}984 , ]

and

[ _{} = 1- = 0.9935897436 = 99.35897436% %. ]

At (N=812{,}500{,}000) entities, the corresponding logical retained-size baseline is

[ M_h = N,b_h = 812{,}500{,}000,(9{,}984) = 8.112^{12}  = 8.112 . ]

The (156{:}1) result is therefore meaningful only when the compared heap experiment actually exhibits approximately 9.984 kB of retained memory per logical entity. It is not mathematically equivalent to comparing a 52 GB packed deployment with a 150–200 GB heap deployment.

For the 150–200 GB heap-envelope comparison quoted in the source directive [48] (an envelope, not a measurement of this work), the ratios are instead

[ C_{150} = = 2.8846, C_{200} = = 3.8462, ]

corresponding to reductions of

[ 1- = 65.33% ]

and

[ 1- = 74.00%. ]

The manuscript consequently reports two baselines separately:

  1. Retained-size object-graph microbenchmark. The (156{:}1) figure may be reported when a retained-heap profiler demonstrates (b_h984) B/entity for the exact comparison schema.

  2. Practical machine-envelope comparison. A 52 GB distributed packed-state deployment compared with a 150–200 GB heap envelope represents (2.88)–(3.85) lower memory demand, not (156).

This distinction prevents an object-representation microbenchmark from being silently promoted into a whole-machine memory ratio.

Engine-specific heap amplification. For the evaluated schema, the packed representation removes the per-entity object/property graph used by the comparison implementation. The magnitude is reported as a measured retained-size ratio rather than a universal property of V8:

[ C_{} = . ]

V8 itself supports multiple object, property, element, and pointer representations; consequently (C_{}) is workload- and engine-version-specific. V8 objects may carry HiddenClass information and use in-object, separate property-store, elements-store, or dictionary representations [51], and modern pointer-compressed heaps store compressed tagged values, which V8 reports as reducing heap size by up to 43% [50]. There is therefore no defensible fixed “JavaScript object (=q) bytes” law; the source directive’s formulation “packed binary resident state eliminates (3)–(4) heap wrappers” [48] is not adopted as an engine invariant. The architectural invariant is only that an explicit 64-byte packed record occupies 64 bytes of record payload, whereas an object representation may require additional metadata, references, backing stores, alignment, and collector-managed storage. The correct baseline is a reproducible retained-heap measurement for the exact schema (Appendix A.2).

Evidence boundary. The algebra above is exact and is reproduced by the executable listing in Appendix A.5 (R). A claim that 812.5 million records were simultaneously resident is empirical and therefore requires per-host resident-set measurements, shard counts, operating-system memory pressure, failure thresholds, and record-integrity checks from the physical run (Appendix A.1); until that artifact is published it is P. The nominal-capacity equation alone does not establish successful physical residency. Allocated virtual address space is likewise not a residency criterion: a process can reserve enormous virtual ranges without those bytes being physically resident.

2.3 Logical dirty granules and spatial locality

Definition 3 — Logical dirty granule. Partition the resident image into tracking units of (p) bytes. If (bp), each full granule contains

[ m= ]

records. A granule is logically dirty when at least one record assigned to it changes.

This abstraction may model a transfer block, software dirty bitmap, or virtual-page-sized tracking unit. It is not automatically an operating-system page table, cache line, DRAM row, GPU transfer transaction, or physical writeback.

For

[ N=50{,}000, b=64, p=4096, ]

we obtain

[ m=64, Nb=3{,}200{,}000, ]

and

[ P=. ]

The first 781 granules contain 64 records and the last contains 16.

Theorem 2 — Expected dirty-granule occupancy under Bernoulli mutation.
Let each record independently change with probability (). If granule (j) contains (m_j) records, then

[ E[D] = _{j=1}^{P} . ]

Proof. Define indicator (I_j=1) exactly when granule (j) is dirty. It remains clean only when all (m_j) records remain unchanged:

[ (I_j=0)=(1-)^{m_j}. ]

Hence

[ E[I_j] = 1-(1-)^{m_j}. ]

By linearity of expectation,

[ E[D] = E!= _jE[I_j]. ]

()

At (),

[ 𝔼[Drandom]=781(1−0.9564)+(1−0.9516)=752.2535206074667.\begin{aligned} \mathbb E[D_{\mathrm{random}}] &= 781(1-0.95^{64}) + (1-0.95^{16})\\ &= 752.2535206074667. \end{aligned}

]

Thus

[ = 96.1961024%. ]

The same audit gives approximately (752.3157554) expected dirty granules when exactly (K=2500) records are sampled uniformly without replacement:

[ E[D_{K}] = _j . ]

These are distinct stochastic models. The often repeated “5% changed means 5% pages” inference is false because one changed record is sufficient to mark an otherwise clean granule.

Theorem 3 — Contiguous spatial-clustering bound.
Suppose the (K) modified records form (C) disjoint contiguous intervals. Interval (j) begins at record (s_j) and has length (L_j>0), with

[ _{j=1}^{C}L_j=K. ]

For (m) records per logical granule, interval (j) touches exactly

[ d_j = - +1. ]

Equivalently,

[ d_j= . ]

Therefore

[ D_{} {j=1}^{C}d_j {j=1}^{C} ( +1 ) < Km+2C, ]

and consequently

[ ]

Proof. The first touched granule is (s_j/m) and the last is ((s_j+L_j-1)/m), giving the exact formula. Because

[ 0s_jm<m, ]

an interval can incur at most one additional granule beyond its rounded payload extent:

[ d_j +1. ]

The dirty set is a union of the interval-specific sets, so its cardinality is no greater than their summed cardinalities. Finally (x<x+1) yields the stated asymptotic bound. ()

Corollary 3.1 — Exact 44-granule construction.
Let (m=64), (C=5), and each (L_j=500). Choose interval starts

[ (s_1,s_2,s_3,s_4,s_5) = (0,6413,12813,19213,25613). ]

Their residues modulo 64 are

[ (0,13,13,13,13). ]

Therefore

[ d_1= ]

and

[ d_2=d_3=d_4=d_5 = = 9. ]

The touched ranges are disjoint, so

[ D_{} = 8+4(9) = 44. ]

Hence

[ ]

This is the requested formal 94.15% suppression result.

The exact fraction of the whole image occupied by 44 granules is

[ =5.6265985%, ]

not 5.5%. A 5.5% whole-image fraction corresponds approximately to 43 granules:

[ =5.4987212%. ]

The independent audit also constructs two useful counterexamples. Five 500-record intervals aligned to granule boundaries touch exactly 40 granules; five intervals beginning at residue 13 each touch 45. Therefore “five clusters of 500 records” does not by itself imply 44. The exact count depends on alignment.

Corollary 3.2 — Hardware-independent logical mapping.
For a fixed canonical layout, logical granule size (p), record width (b), and mutation-index set (M), define

[ D(M) = { : iM }. ]

Any two platforms that implement this same mapping obtain the same logical dirty set.

Proof. Platform identity does not occur in the expression. ()

This is the mathematically valid form of “cross-silicon invariance.” It does not assert that operating-system page sizes, TLB structures, cache-line eviction, physical DRAM writeback, or device-transfer granularity are identical.

Observation-boundary correction: logical dirtiness is not MMU writeback. Linux documents soft-dirty tracking as a bit on a page-table entry. Clearing the tracking state removes write permission; the next write faults, after which the kernel marks the PTE writable and soft-dirty. That mechanism is already distinct from cache or backing-store writeback. A dirty logical-granule count can therefore predict how many virtual-page-sized units were affected, but an exact claim of “94.15% fewer PTE writes,” “94.15% fewer TLB shootdowns,” or “94.15% less memory-controller traffic” requires measuring those quantities directly.

The paper therefore uses the following terminology:

[ ]

for the theorem.

A physical study may later establish a corresponding reduction in PTE transitions, transfer bytes, cache writebacks, or DRAM transactions under a particular implementation, but those are additional empirical propositions.

Publication-safe statement of the locality result. For the specified 50,000-record experiment, spatial clustering reduces logical 4 KiB dirty-granule occupancy by approximately 94.1–94.2% relative to the observed or modeled scattered-mutation baseline. Because the logical record-to-granule mapping is architecture-independent when (b), (p), and the mutation set are held fixed, this occupancy result is portable as a layout property. Physical cache writeback, DRAM traffic, operating-system page behavior, and energy remain platform measurements.

At the observation boundary, this is a reduction in logical dirty-granule occupancy for the specified clustered-versus-scattered mutation experiment. On systems whose dirty tracking is page based, this can reduce the number of data pages whose page-table dirty state transitions and the number of pages presented to page-granular synchronization or persistence mechanisms. It is not, by itself, a measurement of DRAM writeback, cache-line traffic, memory controller transactions, TLB shootdowns, or physical energy. Those quantities require hardware-counter or device-level observation. A page-table entry’s dirty bit, a dirty data page, a cache-line writeback, and a memory-controller transaction are not synonymous. The exact result under the stated mapping is 94.1509% (F). The supplied device runs (U, §5.2) report 44 clustered granules against 747–758 scattered granules, which is where the lower end of the 94.1–94.2% range comes from; the corpus’s 43/744 run cannot be the five-interval construction, which touches exactly 44 granules under the stated mapping (Corollary 3.1), so it is treated as an unspecified mutation set, excluded from the specified experiment, and reported only as a U data point in §5.2:

[ 1-=94.11%, -=94.18%, -=94.20% ]

(reproduced in Appendix A.5).

The phrasing “94.2% fewer MMU page-table writebacks” that appears in the source directive [48] is not adopted; the wording above governs.

Spatial ordering. Morton or Z-order codes are useful because spatial cells sharing a key prefix become contiguous key ranges, and related GPU hierarchy construction techniques exploit such ordering. They do not guarantee that every Euclidean neighborhood occupies one memory interval. If a spatially affected region decomposes into (C_R) materialized intervals, the theorem becomes

[ D_R = O!( +C_R ). ]

Thus the quantity that must be measured for a real workload is not “spatiality” in the abstract but fragmentation (C_R). Karras’s parallel radix/BVH work provides relevant prior art for space-filling-curve ordering and parallel hierarchy construction.

2.4 Sparse-delta transport and the dense-state bandwidth bound

Resident state changes the network problem when the receiver already possesses a valid base revision. Let (S) denote the bytes in a dense state snapshot, (f) the publication frequency, (k_t) the number of changed records at publication (t), (b) the packed record width, and (h_t) the bytes of protocol, authorization, proof, and framing metadata attributed to that publication.

A dense replication path requires application-visible payload bandwidth

[ B_{} = fS. ]

A sparse-delta path requires

[ B_{} = f(h_t+k_tb) ]

for constant-size updates, or more generally, for an observation window of length (T) in which publication (t) occurs at time (_t),

[ B_{} = _{t: _tis not adopted as a result.

Data movement and energy

The reduction above is a byte-volume result. It must not be reported as an equal percentage reduction in electrical energy. Let boundary (j) transfer (Q_j) bytes at effective marginal energy (_j) joules per byte. A first-order movement model is

[ E_{} = _j_jQ_j. ]

End-to-end energy is

[ E_{} = E_{} + E_{} + E_{} + E_{} + E_{} + E_{}. ]

Reducing (Q_j) can reduce the corresponding dynamic movement term, but packetization, idle power, link-state transitions, DRAM refresh, retained memory, and cooling prevent a logical byte ratio from becoming an automatic joule ratio. The thermodynamic claim is therefore measured at the relevant hardware counters or power boundary. This accounting is the transport-boundary refinement of the §2.7 decomposition: (j) is §2.7’s (e) evaluated at boundary (j); (E_{}) is the network share of §2.7’s (E_{}); and idle, link-state, and cooling power are the terms §2.7 folds into (E_{}) and (E_{}); §2.7’s (E_{}) is carried inside (E_{}) at the transport boundary, where the fast path is allocation-free (§2.5). The two partitions sum to the same (E_{}).

2.5 Allocation-free fast path and zero-heap closure

Definition 4 — Allocation-free fast path. Let (O) be a finite set of admitted mutation opcodes. For each primitive (oO), let (A(o)) be the number of dynamic-heap bytes allocated while applying it. A profile is allocation-free when

[ oO,A(o)=0, ]

and all state, scratch regions, cryptographic workspaces, ring slots, and index buffers needed by the path are statically reserved or preallocated.

Theorem 4 — Zero-heap closure under finite opcode composition.
If every primitive in an admitted command sequence

[ =(o_1,o_2,,o_n) ]

satisfies (A(o_i)=0), and dispatch itself allocates no heap memory, then

[ A()=0. ]

Proof.

[ A() = {i=1}^{n}A(o_i)+A{} = 0. ]

()

This theorem is intentionally simple because “zero heap” is an implementation property, not a mystical property of a binary format. FlatBuffers, for example, already demonstrates that direct buffer access can require no separate heap for reading, while seL4 demonstrates an operating-system design whose kernel itself has no heap. Those precedents mean ATESO’s novelty cannot be “a buffer can be read without malloc”; its research question is the integration of allocation-free mutation with resident authority, revision, scheduling, feedback, and provenance.

The absence of dynamic heap allocation also does not imply zero memory. Stack, static state, fixed cryptographic workspace, DMA descriptors, interrupt frames, and device buffers still consume memory.

2.6 XPBD and conflict-free constraint coloring

XPBD formulation. For constraint (C_i(x)), compliance (_i), step (h), inverse-mass matrix (M^{-1}), and accumulated Lagrange multiplier (_i) initialized to zero at the start of each substep ((_i^{(0)}=0)), Extended Position-Based Dynamics uses

[ _i=, ]

[ _i = , ]

followed by

[ x = M^{-1}C_i^T_i, ]

with multiplier accumulation (_ii+i) over the solver iterations within that substep. Multipliers are reset to zero at the start of each new substep (or rescaled by ((h{}/h{})^2) if warm-starting across variable timesteps), ensuring the penalty term (_i_i) remains dimensionally and physically consistent across changing step sizes.

XPBD was introduced to make compliant constraint behavior less dependent on timestep and solver-iteration count than ordinary PBD [10]. It can be solved with Gauss-Seidel- or Jacobi-style schemes, but those algorithmic properties do not guarantee cross-device bitwise determinism. Future biomechanical Magma workloads would draw constitutive models from continuum biomechanics [45] and shear-wave elasticity imaging [46], and dynamic-scene rendering from 4D Gaussian splatting [47]; none of those workloads is evaluated in this paper.

Definition 5 — Constraint conflict graph. Construct graph

[ G_C=(C,E_C), ]

where each vertex is a constraint and

[ (c_i,c_j)E_C (c_i) (c_j) ]

for state written by both constraints.

A proper coloring partitions

[ C = _{q=1}^{K}C_q ]

such that

[ c_i,c_jC_q, ij: (c_i) (c_j) = . ]

Theorem 5 — Intra-color write conflict freedom.
Assume each constraint writes only state belonging to its incident particles and does not mutate shared global state. Then all constraints in one color (C_q) can execute concurrently without write-write conflicts.

Proof. Suppose two constraints (c_i,c_jC_q) could conflict. Then both would write at least one common particle (v), implying

[ v (c_i) (c_j), ]

contradicting the coloring property. ()

Processing colors sequentially while constraints within one color run in parallel yields a multi-colored block Gauss-Seidel structure. The result applies only to the state covered by the conflict graph; collision tables, accumulators, global counters, or dynamically generated contacts require separate synchronization.

Current WGSL defines workgroup-scoped control barriers, and storage/workgroup atomic operations have specifically defined scopes and relaxed atomic semantics. A workgroupBarrier() cannot be treated as a global barrier among unrelated workgroups. A portable multi-color WebGPU implementation should therefore establish global color phases with distinct dispatch/queue ordering or another specification-valid synchronization structure; Theorem 5 guarantees intra-color write-conflict freedom, while inter-color execution ordering is strictly enforced across sequential compute pass dispatches on the command queue to prevent race conditions on shared boundary particles.

2.7 Energy decomposition

Energy decomposition. Let a complete operation cross physical boundaries (,,L). A useful accounting model is

[ E_{} = E_{} + E_{} + E_{} + E_{} + E_{}, ]

with

[ E_{} {}^{L} eQ_, ]

where (Q_) is measured bytes or transactions at boundary () and (e_) is the effective joules per unit under the current platform state.

This decomposition explains what the representation theorem can and cannot prove. It can show that a representation makes some (Q_) unnecessary. It cannot infer (e_) from first principles for a modern SoC, nor does lower transferred data automatically imply lower total energy if the resident design increases retention power or accelerator wake time.

3. The Triple-Contract System Interface

3.1 Persistent work identity

Section 3 — Architectural contracts. ATESO organizes execution around a persistent work identity rather than a transient document. Let

[ I= ( , , , , ). ]

A tool transition changes which operations and projections are active over (I); it need not instantiate another semantically equivalent object graph.

3.2 The .many contract

The .many contract: relocatable resident state. A .many arena is proposed as a canonical, versioned binary image containing at minimum:

Region Required semantics
Superblock Magic, format version, byte order/numeric profile, arena length
Schema directory Record types, widths, alignment, field interpretation
Object directory Stable IDs and relative offsets
Revision state Current state revision, epoch, checkpoint identity
Dynamic vectors Fixed-width state suitable for direct indexed access
Constraints XPBD/contact/relationship data
Spatial index Derived hierarchy or deterministic rebuild descriptors
Dirty directory Changed ranges or logical-granule bitmap
Telemetry region Timestamped physical observations and quality metrics
Provenance Content hashes, tree roots, prior checkpoint references

All internal references are offsets or stable identifiers rather than process-specific absolute pointers. This resembles established relocatable representations: Arrow explicitly specifies relocatability without pointer swizzling and recommends aligned contiguous buffers; Cap’n Proto uses offset-based structured encoding; FlatBuffers is designed for direct buffer traversal.

A proposed profile may choose 64-byte record alignment because it is convenient for SIMD, cache-conscious access, and some GPU layouts. It must not call 64 bytes a universal architectural cache-line law. Arrow itself recommends 8- or 64-byte alignment, with 64 bytes being an optimization recommendation rather than a universal semantic requirement.

Admission validates, before dereference:

[ 0oB, B-o, ]

for every region of offset (o), length (), and arena length (B). Checked subtraction is preferred to unchecked (o+) arithmetic so integer overflow cannot transform an invalid region into an apparently valid one.

Admission also validates schema version, alignment, count multiplication, numeric range, finite-value constraints, dependency references, capability ownership, declared resource budget, and the applicable integrity commitment.

An important security rule follows:

[ ]

unless the bytes remain immutable or are covered by a defined ownership/publication protocol between verification and use. Copying an untrusted input into a trusted immutable arena is legitimate when required to eliminate a time-of-check/time-of-use race. “Zero-copy” is not a security axiom.

3.3 The .moat contract

The .moat contract: authority plus provenance. Let a capability be

[ C= (O,A,,e,,r), ]

where (O) is the object scope, (A) is the set of permitted operations, () is any value/range constraint, (e) is expiry or lease policy, () is replay state, and (r) is the expected revision.

A canonical .moat envelope binds

[ M= ( H(S), (O), C, r, , e, , ) ]

and authenticates a domain-separated canonical byte representation. An Ed25519 profile can sign such an envelope; RFC 8032 specifies Ed25519 as a concrete EdDSA instance with 32-byte public keys and 64-byte signatures. The signature algorithm itself does not prove that the private key resides in hardware; hardware binding requires a separately identified trust root or attestation mechanism.

The architecture’s proposed open-wire policy—CC0 specification material—is an author-specified release objective in the supplied directive. Actual release scope and publication date should be established by the final released documents rather than inferred from this manuscript.

Capability systems have deep prior art. seL4 defines capabilities as unforgeable tokens combining object references with access rights, while CHERI enforces monotonic capability derivation rules in hardware. The ATESO theorem below is therefore an application-level composition rule, not a claim to have invented capability security.

Theorem 6 — Capability non-amplification.
Let a delegation chain be

[ C_0,C_1,,C_k ]

such that the broker enforces

[ C_{i+1}C_i ]

for all (0i<k), where inclusion means every operation, object, value range, temporal lease, and delegation right in (C_{i+1}) is no greater than the corresponding authority in (C_i).

Then

[ jk,C_jC_0. ]

Consequently, for any operational authority (),

[ C_j C_0. ]

Proof. By induction.

Base case: (C_0C_0).

Inductive step: assume (C_iC_0). Broker enforcement gives

[ C_{i+1}C_i. ]

By transitivity of set inclusion,

[ C_{i+1}C_0. ]

Thus the property holds for all (j). If (C_j), inclusion implies (C_0). ()

The theorem depends on the security model. It does not prevent escalation through an ambient operating-system privilege, compromised root key, confused deputy, implementation bug, or unmodeled side channel. A complete profile therefore requires that every externally consequential operation be mediated by a capability check and that no hidden authority bypass the model.

Theorem 7 — Monotonic-revision stale-write exclusion.
Let authoritative state have current revision (r). A command (c) is commit-eligible only when

[ r_c=r, ]

and a successful commit atomically changes the revision to

[ r’=r+1. ]

Then two commands both claiming expected revision (r) cannot both successfully commit serially through that broker.

Proof. Suppose (c_1) commits first. The broker changes the authoritative revision from (r) to (r+1). When (c_2) is subsequently evaluated, its assertion (r_{c_2}=r) no longer equals the current revision. Therefore the broker rejects it. ()

The theorem prevents stale-state commits in the brokered sequence. It does not by itself protect against rollback of the broker’s durable revision state. Crash recovery therefore needs a durable monotonic checkpoint, replicated log [27], trusted monotonic store, or epoch transition.

3.4 Tamper-evident provenance

Tamper-evident provenance. A simple event chain is

[ h_0= H(g) ]

where (g) is a 32-byte deployment genesis value chosen once per deployment (for example a random value published with the deployment’s root key) and reused by every chain and checkpoint genesis in this manuscript (§4.6, §6.2, §6.3), and

[ h_i = H( h_{i-1} |r_i| r_i ). ]

These tags are unversioned; tags introduced later in this manuscript carry an explicit /v1 suffix, and any revision of this layout adopts the same convention.

For scalable membership proofs, a Merkle tree can commit batches of accepted transitions. Certificate Transparency provides standardized examples of inclusion proofs and consistency proofs over authenticated Merkle roots. A hash chain or Merkle tree establishes tamper evidence relative to a trusted anchor; it does not prove that the underlying sensor reading or benchmark measurement was truthful.

BLAKE3 is a suitable example of a fast modern hash with portable and SIMD-optimized implementations, while SHA-256 remains a widely standardized alternative. The choice belongs in the profile and must be domain-separated from unrelated uses of the hash.

3.5 The .spine contract

The .spine contract: bounded shared-memory exchange. Consider an SPSC lane of power-of-two capacity (M<2^{31}), fixed record width (w), producer publication counter (p), and consumer reclamation counter (c). Counters are unsigned 32-bit values.

Define occupancy as

[ u=(p-c)^{32}. ]

The lane invariant is

[ 0uM. ]

The lane is empty at (u=0), full at (u=M), and the record slot is

[ (x) = x(M-1). ]

The producer procedure is:

procedure SPINE_PUBLISH(lane, record):
    p ← atomic_load(lane.producer_counter)
    c ← atomic_load(lane.consumer_counter)

    if unsigned32(p - c) = lane.capacity:
        return BACKPRESSURE

    slot ← lane.slots[p & (lane.capacity - 1)]

    // Producer exclusively owns this unreleased slot.
    write_payload(slot, record)

    // Publication occurs only after payload completion.
    atomic_store(lane.producer_counter, p + 1)

    return OK

The consumer procedure is:

procedure SPINE_CONSUME(lane):
    c ← atomic_load(lane.consumer_counter)
    p ← atomic_load(lane.producer_counter)

    if unsigned32(p - c) = 0:
        return EMPTY

    slot ← lane.slots[c & (lane.capacity - 1)]

    // Publication has already occurred.
    record ← read_payload_completely(slot)

    // Reclamation follows completion of every payload read.
    atomic_store(lane.consumer_counter, c + 1)

    return record

ECMAScript defines the shared-memory and atomic semantics used by a JavaScript implementation [8], and the browser profile assumes both endpoints share one agent cluster under the HTML standard’s agent-cluster rules [9]; Lamport’s foundational work on concurrent reading/writing and subsequent shared-memory formalisms provide the broader concurrency lineage.

Theorem 8 — SPSC slot non-overlap.
Assume one producer, one consumer, atomic publication and reclamation counters, publication after payload completion, reclamation after complete payload consumption, and the bounded occupancy invariant (uM). Then the producer never overwrites a slot while the consumer is reading the corresponding published record.

Proof. Let record sequence (q) occupy slot (qM). The consumer can begin reading (q) only after observing publication beyond (q). The producer can reuse that slot only for sequence (q+M). Before publishing or writing (q+M), the producer’s full-lane test requires reclamation to have advanced past (q); otherwise the occupancy is (M) and publication is rejected. The consumer advances reclamation past (q) only after completing its payload reads. Hence reuse follows completed consumption, so producer and consumer cannot concurrently access the same slot with conflicting ownership. ()

The (M<2^{31}) condition and bounded occupancy prevent modular ambiguity around unsigned wrap. The reproduced reference program begins counters at 0xfffffff0, intentionally crossing the 32-bit wrap boundary.

Multiple producers do not invalidate the theorem if each producer receives an independent SPSC lane. A separate broker can merge those lanes into an authoritative command order.

3.6 Trusted and untrusted paths

Trusted and untrusted paths. Shared memory is a trust grant. A process with unrestricted write access to the authoritative arena can corrupt it independently of .moat. ATESO therefore defines two paths:

trusted worker
    -> SPSC lane
    -> admitted local state

untrusted model / plugin / network peer
    -> isolated input buffer
    -> canonical snapshot
    -> .moat + schema + revision validation
    -> brokered mutation
    -> authoritative state

This separation parallels the central principle of capability operating systems: possession of a memory reference or kernel capability defines authority, so untrusted code must not be given a reference broader than intended. seL4’s capability model is an explicit example.

3.7 WebGPU hardware-boundary reality

WebGPU hardware-boundary reality. Native shared memory may genuinely allow multiple CPU-side participants to address the same physical or virtual memory allocation. Portable WebGPU does not promise arbitrary aliasing of a SharedArrayBuffer as a GPU storage buffer. The current specification states that a mapped GPU buffer is unavailable to GPU queue operations; MAP_WRITE may only be combined with COPY_SRC; and device buffers are mediated objects with explicit map, write, copy, and submission rules.

The browser-profile claim must therefore be:

[ ]

rather than universal CPU/GPU zero-copy buffer pinning.

This distinction is architecturally important. .spine can be genuinely zero-copy for cooperating CPU agents inside one compatible shared-memory trust domain; WebGPU can retain GPU-resident state and minimize changed uploads, but it cannot portably promise that JavaScript shared memory and device storage are the same allocation.

4. Closed-Loop System One and System Two Cyber-Physical Orchestration

4.1 Temporal regime separation

Section 4 — Temporal regime separation. The terms System 1 and System 2 are used here as architecture-local labels, not as a claim that every cognitive model has one fixed latency.

Define:

[ = ]

and

[ = . ]

The intended planner regime may operate at (102)-(103) ms timescales. The fast execution path is designed for microsecond-scale commands under profiles that can establish a corresponding WCET. Language-model reasoning is explicitly excluded from the real-time critical loop.

HTN planning supplies established formal machinery for hierarchical task decomposition, ReAct demonstrates interleaved model reasoning and externally grounded action, and actor-critic methods provide another established two-timescale control lineage. ATESO does not replace these methods; it defines a typed, authorized execution membrane beneath whichever planner is selected.

4.2 The serialization seam

The serialization seam. A valid criticism of “LLM-controlled binary state” is that today’s language models ordinarily emit or consume tokens. ATESO resolves the criticism by moving free-form language outside the hard-real-time boundary rather than pretending tokens are memory writes.

Let planner output be semantic action (a). A non-real-time compiler maps it to opcode

[ = ( , , , , r, C, ). ]

A schema-specific compiler

[ : a ]

performs text or structured-output interpretation, range checking, unit conversion, and policy prevalidation. The execution broker accepts only (), never arbitrary planner text.

Thus the seam is

[ . ]

The architecture does not eliminate tokenization from an LLM. It prevents tokenization, JSON parsing, general object allocation, and arbitrary model behavior from entering the deterministic actuator loop.

A minimal mutation descriptor can be modeled as

[ = (o,x,f,w,v,r,), ]

where (o) is opcode, (x) is object ID, (f) is field/offset selector, (w) is width/type, (v) is a bounded payload or reference, (r) is expected revision, and () is capability reference.

Sparse-delta application algorithm.

procedure ADMIT_AND_APPLY(envelope E, arena S):

    // Structural admission
    require canonical_encoding(E)
    require bounds_valid(E)
    require schema_supported(E.schema)

    // Authentication and authority
    require verify_signature(E.issuer, E.signed_bytes, E.signature)
    require capability_valid(E.capability)
    require E.operation ∈ E.capability.allowed_operations
    require E.target ∈ E.capability.object_scope
    require value_within_capability_ranges(E)

    // Freshness
    require replay_policy_accepts(E.nonce, E.epoch)
    require E.expected_revision = S.revision

    // Translate descriptor into a previously validated region.
    target ← resolve_stable_offset(S.directory, E.object, E.field)

    require target.width = E.width
    require target.bounds ⊆ S.authorized_arena

    // Allocation-free mutation.
    fixed_width_store(target, E.value)

    // Revision and incremental provenance.
    S.revision ← S.revision + 1
    mark_logical_granule_dirty(target)
    update_incremental_hash_metadata(target)

    publish_telemetry_record(
        revision = S.revision,
        command_id = E.command_id,
        result = ACCEPTED
    )

    return S.revision

A production real-time profile must place a bounded upper limit on every loop and cryptographic operation in this procedure. Signature verification may dominate a (<10,) target on some MCUs. In such cases, a gateway can verify an asymmetric envelope and issue a separately authenticated bounded local command, but that is a different security profile and must be reported as such.

4.3 The canonical six-step closed loop

The canonical six-step closed loop.

[ ]

The stages have the following contracts.

Stage Input Output Real-time status
Decomposition Goal, observations, policy context Task DAG / semantic actions Deliberative; not hard-real-time
Attestation Typed action .moat-bound command Bounded only in qualified profile
Staging Authenticated command .spine record Constant-size local operation
Admission Record + current revision Accept/reject Must have characterized WCET for hard-RT profile
Mutation Admitted opcode New .many revision Allocation-free finite opcode
Telemetry readback State/physical observations Typed planner observation Binary locally; planner adapter may later tokenize

Step one: decomposition. The planner constructs task graph

[ G=(T,D). ]

Each task receives input revisions, capabilities, deadline class, expected resource demand, and an output contract.

Step two: attestation. The action compiler produces a bounded opcode and binds it to its issuer, state root, capability, nonce, and expected revision.

Step three: staging. The command is written into a bounded .spine lane. If the lane is full, safety-critical commands backpressure rather than overwrite unread commands.

Step four: admission. The broker validates signature, issuer policy, capability subset, schema, target bounds, replay state, and current revision.

Step five: mutation. The admitted operation changes only a predefined region of .many; no arbitrary code supplied by the planner is executed in this path.

Step six: telemetry readback. Sensor state, solver residuals, command outcomes, resource telemetry, and revision information are published as typed fields or fixed records.

The phrase zero-serde telemetry has a precise scope here. Cooperating native components may read typed shared-memory fields without serializing them to text. The language model itself still ultimately receives whatever token, embedding, or structured representation its inference API accepts. The architecture eliminates unnecessary text serialization inside the deterministic loop; it does not claim that all planner interfaces in the world are binary.

4.4 Dual telemetry

Dual telemetry. Let

[ z_t^{(p)} ]

be physical telemetry and

[ z_t^{(u)} ]

be preference or utility telemetry.

Physical telemetry includes temperature, voltage/current limits where exposed, actuator bounds, solver residuals, strain limits, deadline counters, and fault state. It participates in hard feasibility constraints.

Taste or preference telemetry modifies a utility function but cannot relax a hard constraint. Formally, an allocation (u_t) is permitted only if

[ g_j(u_t,z_t^{(p)}) ]

for every hard constraint (j). Preference terms only rank members of the feasible set.

4.5 Thermal dynamics and robust admission

Thermal dynamics. For a first-order lumped model,

[ C_ = P(t)-, ]

where (R_) has units K/W and (C_) has units J/K. For constant (P) over interval (h),

[ ]

as requested.

Derivation. Let

[ y(t)=T(t)-T_a-R_P. ]

Then

[ = -. ]

Solving the separable ODE gives

[ y(t+h)=y(t)e^{-h/(R_C_)}, ]

and substituting back yields the result. ()

HotSpot is an established example of processor thermal modeling through equivalent thermal resistance-capacitance networks, including multi-block and multi-layer models. A one-node ATESO controller is therefore an approximation to be calibrated, not a replacement for full thermal characterization.

Theorem 9 — Robust thermal admission.
Assume the true temperature over an admission horizon satisfies

[ T(t+) T(t+)+_T(t+) ]

for all (0h). If the controller admits an action only when

[ T(t+)+T(t+) T{} ]

for all (0h), then

[ T(t+)T_{} ]

throughout the horizon.

Proof. At every (),

[ T(t+) T(t+)+T(t+) T{}. ]

()

The difficult engineering problem is validating (_T). Sensor delay, fan state, package coupling, neighboring activity, DVFS transition latency, and model mismatch belong inside the uncertainty budget.

A scheduling objective can be written

[ u {=t}^{t+H} ]

subject to capabilities, dependency readiness, memory capacity, deadline reservations, and robust thermal constraints. Normalizing each term avoids adding milliseconds, joules, and bytes as though they were naturally commensurate. Closed-loop resource management of this kind sits in the autonomic-computing lineage [35]; ATESO’s contribution is the typed, revisioned state plane it manages, not a new control law.

4.6 Research egress: programmatic notebook synthesis without conflating execution domains

ATESO’s resident-state fast path and an external research notebook serve different purposes. .many is an execution representation; a research notebook is a derived knowledge artifact.

The architecture therefore defines a one-way research egress plane:

[ .𝚖𝚊𝚗𝚢 𝚛𝚎𝚜𝚒𝚍𝚎𝚗𝚝 𝚜𝚝𝚊𝚝𝚎→𝚊𝚞𝚝𝚑𝚘𝚛𝚒𝚣𝚎𝚍 𝚙𝚛𝚘𝚓𝚎𝚌𝚝𝚒𝚘𝚗→𝚍𝚒𝚐𝚎𝚜𝚝-𝚋𝚘𝚞𝚗𝚍 𝚛𝚎𝚜𝚎𝚊𝚛𝚌𝚑 𝚜𝚗𝚊𝚙𝚜𝚑𝚘𝚝→𝚎𝚡𝚝𝚎𝚛𝚗𝚊𝚕 𝚗𝚘𝚝𝚎𝚋𝚘𝚘𝚔/𝚜𝚘𝚞𝚛𝚌𝚎 𝙰𝙿𝙸.\begin{gathered} \texttt{.many\ resident\ state} \rightarrow \texttt{authorized\ projection}\\ \rightarrow \texttt{digest-bound\ research\ snapshot} \rightarrow \texttt{external\ notebook/source\ API}. \end{gathered}

]

Let the authoritative state at revision (r) be (S_r). A capability-scoped research projection is

[ X_r = _{}(S_r,C), ]

where (C) determines which fields may leave the resident execution boundary. The exported artifact is

[ A_r = _{F}(X_r), ]

for an explicitly supported external format (F), and is bound to its source revision by

[ d_r = H( |r| r H(S_r) H(X_r) H(A_r) H(C) _r ), ]

where (|r|r) is the revision (r) in the length-prefixed (|x|x) convention of §3.4 and (_r) is the head of the §3.4 provenance chain (or the Merkle root of the accepted-transition batch) at the moment of export, that is, the chain head immediately preceding the entry that records (e_r). The digest therefore binds the artifact to the state it was projected from, the projection actually exported, the capability that authorized the export, and that chain head; changing any of them changes (d_r). The binding is made two-way by recording the export as an accepted event on the same chain:

[ e_r = H( d_r || t_{} ), ]

where () is the external source identifier returned by the notebook API and (t_{}) is a fixed-width 64-bit big-endian timestamp; (e_r) is appended to the §3.4 chain and is committed by the next Merkle batch taken over that chain (§3.4). A deployment that anchors such a batch to a transparency log uses the leaf, interior-node, and checkpoint construction of §6.3; §6.3 itself commits only the §6.2 accounting digests, so an export is anchored only when the batch containing (e_r) is anchored. The chain commits to the export and the export commits to the chain.

The export step is intentionally visible. It is incorrect to describe a remote Notebook API as direct zero-copy access to local .many memory. Zero-copy residency applies inside compatible ATESO memory domains; crossing into an external cloud API requires an explicit representation and transport boundary.

This boundary can nevertheless be fully automated. A telemetry exporter can periodically or event-conditionally:

  1. select an authorized state/telemetry projection;
  2. materialize a versioned Markdown, text, document, image, audio, or other supported research artifact;
  3. bind the artifact to the .many revision, state digest, and .moat provenance root through (d_r), and append the export event (e_r) to the provenance chain;
  4. publish or update the artifact through a cloud/document gateway;
  5. add the resulting source to an enterprise notebook; and
  6. request supported derived research products such as an Audio Overview.

The effect is not “serialization elimination at every boundary.” It is the removal of repeated serialization from the execution hot path, followed by a deliberate asynchronous serialization boundary when a human-facing knowledge artifact is actually required.

Because the notebook source is revision-bound, an analyst can reconstruct the chain

[ S_r X_r A_r ]

without treating the generated narrative as authoritative machine state.

Google’s enterprise notebook APIs provide a concrete implementation target for this research-egress contract [52–54]: notebooks are created and managed through notebooks.create and related methods [52]; sources are added through notebooks.sources.batchCreate and notebooks.sources.uploadFile, which accept raw text, Markdown, Google Docs and Slides, web URLs, office documents, audio, video, and images [53]; and an Audio Overview is requested through notebooks.audioOverviews.create [54]. All three surfaces are documented as Pre-GA (Preview) at the time of writing. The corresponding Drive MCP interface, served at drivemcp.googleapis.com under the Google Workspace Developer Preview Program, provides an agent-accessible document plane [55]. These external systems are treated as sinks and research tools, not as extensions of Magma’s coherent memory substrate.

Static-source semantics and continuous telemetry

“Continuous notebook telemetry” means continuous pipeline operation, not continuous aliasing of live memory. The external notebook service ingests a snapshot of a source: Google’s product documentation states that the model relies on the static copy of the uploaded sources and does not change the original file [59]. If the underlying .many revision advances from (r) to (r+1), the exporter must explicitly create or replace a source artifact according to the notebook API’s supported semantics.

For event sequence

[ r_0<r_1<<r_n, ]

the exporter may apply a sampling or materiality policy

[ (r_i) = {1,∥Π(Sri)−Π(Sri−1)∥ℳ≥τ,0,otherwise,\begin{cases} 1,& \left\| \Pi(S_{r_i})-\Pi(S_{r_{i-1}}) \right\|_{\mathcal M} \ge\tau,\\ 0,&\text{otherwise}, \end{cases}

]

where (M) is a domain-specific change metric and () is the research-materiality threshold. Only revisions with ((r_i)=1) need to produce notebook artifacts.

This prevents a high-frequency physical loop from being coupled to a human-timescale generative-research service. A 1 kHz controller can remain entirely inside System 1 while the notebook plane receives minute-, hour-, or milestone-scale attested snapshots.

Derived-product scope

NotebookLM, the consumer counterpart of Gemini Notebook Enterprise, exposes Video Overviews as a product capability, including source-grounded narrated visual explanations in Explainer and Brief formats [56]; no corresponding statement in the enterprise documentation is cited here. As of the documentation reviewed for this manuscript, the public enterprise API documentation explicitly describes programmatic Audio Overview generation [54]; we therefore do not claim an equivalent documented Video Overview API until such an endpoint is published and pinned in the reproducibility manifest (Appendix A.4).

The precise form of the claim is therefore:

[ ]

This is a design contract (Appendix A.4); no export run is reported in this work, so the capability itself is P (§5.10). It removes manual serialization from the human research workflow without pretending that serialization vanished at a cloud API boundary. The Google services are one implementation target; the contract is defined over any external source API with static-copy ingestion semantics.

5. Empirical Silicon Benchmarks and Experimental Methodology

5.1 Evidence boundary before results

Section 5 — Evidence boundary before results. The supplied ATESO corpus reports a heterogeneous set of hardware results, but the available reproduction archive explicitly states that it does not contain the original WebGPU benchmark executable, heterogeneous-fleet traces, signed receipts, or the original scheduler simulator. Those numbers are therefore reproduced below because they are part of the research record, but they are labeled U rather than described as independently verified measurements.

The accompanying ATESO analytical and SPSC reproduction artifact contains audit.py, spine_spsc_test.mjs, machine-readable outputs, source-directive provenance, checksums, and a receipt manifest.

5.2 Ten-target qualification registry

Ten-target qualification registry. The following table combines the targets required by the current research directive with the benchmark data actually present in the supplied corpus. EPYC and Jetson Orin are included because the requested experimental matrix names them, but no ATESO timing trace for those two targets was supplied; they therefore remain P/V, not silently invented measurements.

Target Independently verified platform fact ATESO result in supplied corpus Evidence class Academic disposition
Apple M5 Max Apple documents two M5 Max bins: 18-core CPU / 32-core GPU with 36 GB unified memory and 460 GB/s, and 18-core CPU / 40-core GPU with 48–128 GB and 614 GB/s [39, 57]. The machine of record’s own hardware report (model identifier Mac17,7, 18-core CPU, 32-core GPU, 36 GB) places it in the 32-core-GPU bin, whose documented bandwidth is 460 GB/s. Full JSON 9.327 ms; binary delta 0.127 ms; reported XPBD mean 0.3795 ms; source uses both 43/744 and later 44/756 locality runs V + U Keep result, require raw samples and executable. Correct the supplied 400 GB/s and 614 GB/s metadata to the 36 GB SKU’s documented 460 GB/s; record the exact SKU per Appendix A.1
AMD EPYC 9005 AMD specifies 12-channel DDR5 and up to 614 GB/s per-socket memory bandwidth on several EPYC 9005 parts. No ATESO benchmark supplied V + P Required future server baseline
Orin-class embedded GPU module GPU vendor specifies a 12-core Cortex-A78AE CPU and up to 275 TOPS on AGX Orin profiles. No ATESO benchmark supplied V + P Required future edge/GPU baseline
Xbox Series X Microsoft specifies an 8-core custom Zen 2 CPU, RDNA 2 GPU, and 16 GB GDDR6. 40.40 ms JSON; binary displayed as 0.00 ms; claimed 40,400× under 0.001 ms timer-floor convention; 0.30 ms CPU reference step; 44/747 logical granules V + U Ratio must be labeled timer-floor microbenchmark, not application speedup
Nintendo Switch Source identifies Tegra-X1-class ARM/Maxwell environment 306.33 ms JSON vs 4.00 ms binary; 2.00 ms reference step; 44/755 U Requires executable and environment attestation
Samsung Galaxy S25 Ultra Qualcomm confirms Snapdragon 8 Elite for Galaxy in the S25 family. 14.77 ms JSON; binary displayed 0.00 ms; reported 438 Hz WebGPU; 44/758 V + U Rounded-zero timing requires timer-resolution disclosure
Moto G 2025 Motorola specifies Dimensity 6300 + Mali-G57 MC2, not the Mali-G51 listed in the supplied benchmark metadata. 51.2 ms vs 0.31 ms; 1.00 ms reference step; 44/758 V + U Correct hardware metadata before publication
Intel Celeron N4020 Chromebook Low-power x86 profile named by supplied corpus 1.85 ms 1,024-particle CPU reference, reciprocal 540.54 Hz U Need solver configuration and residual validation
Cortex-M4 appliance profile Cortex-M4-class industrial MCUs commonly expose deterministic timer, DMA, FPU/DSP, and PWM-oriented peripherals; specific implementations vary. An Infineon M4 family is one documented example. 96-byte frame; claimed zero dynamic heap; (<10,s) register update; 1 kHz loop U Exact MCU, cryptographic path, WCET and linker map required
Smart-TV WebKit profile Browser/embedded profile from supplied corpus Binary delta 0.042 ms; 0 ms observed GC pause; document stack reported 14–14.2 ms GC event; 120 Hz presentation claim U Must report allocation/GC observation method and full frame distribution

The table demonstrates an important distinction: vendor verification of the silicon identity does not verify the ATESO benchmark.

The source also reports M1 Pro and iPad M1 measurements outside this ten-target registry, including a 0.118 ms binary-delta measurement on M1 Pro and a 48× resident-delta ratio on iPad M1. Those remain U pending the original raw traces.

5.3 Representation benchmark

Representation benchmark table from the supplied Apple-Silicon run.

Entities Payload bytes Full JSON Full binary JSON delta Binary delta
100 6,400 0.200 ms 0.060 ms 0.295 ms 0.055 ms
1,000 64,000 0.390 ms 0.075 ms 0.058 ms 0.045 ms
10,000 640,000 1.717 ms 0.320 ms 0.110 ms 0.073 ms
50,000 3,200,000 9.327 ms 1.467 ms 0.273 ms 0.127 ms

These values are copied from the supplied research corpus.

Using the displayed 50,000-record means,

[ = 73.4409449, ]

not 73.63. The source’s 73.63× figure could result from unrounded raw values, but those values are absent. The defensible paper reports both: 73.63× author-reported, 73.44× reproducible from displayed means.

Likewise,

[ =2.1496063, ]

so the displayed JSON-delta versus binary-delta ratio is approximately 2.15×. That comparison is scientifically more informative than full JSON versus binary delta because it separates encoding/admission effects from the much larger benefit of avoiding full reconstruction.

5.4 The Xbox timer-floor result

The Xbox timer-floor result. The source displays

[ 40.40 . ]

No finite ratio can be computed from a literal zero denominator. The requested 40,400× value corresponds exactly to assuming

[ T_B= = 0.001 = 1,s. ]

Accordingly, publication-ready language is:

In the Xbox microbenchmark, the binary-read operation reached the experiment’s stated 1-()s timing floor; using that floor as the denominator gives a 40,400× parser/read ratio. This is a timer-floor microbenchmark ratio, not a 40,400× end-to-end application speedup.

That phrasing preserves the result without converting a measurement-resolution limit into false precision.

5.5 XPBD result

XPBD result. The supplied Apple-Silicon result comes from an early small-scale GPU run (32 particles); full-scale results pending. Each step has six dispatch phases: prediction, four colored constraint passes, and velocity update.

Metric Source-reported value Derived value
Mean 0.3795 ms an early small-scale GPU run (32 particles); full-scale results pending
Median 0.300 ms —
p99 0.700 ms —
Nominal 120 Hz frame interval 8.333 ms mean physics step = 4.554% of interval

These are timing observations U, while the reciprocals are arithmetic R.

A reciprocal latency is not automatically sustainable application throughput. A complete measurement must include command encoding, queueing, GPU submission, constraints, collision generation, rendering, synchronization, presentation, thermal steady state, and correctness residuals.

5.6 Logical locality result

Logical locality result — independently reproduced.

Quantity Result
Records 50,000
Record size 64 B
Logical tracking granule 4,096 B
Granules 782
Bernoulli mutation probability 0.05
Expected Bernoulli dirty granules 752.2535206
Expected Bernoulli dirty fraction 96.1961024%
Exactly 2,500 uniformly sampled records 752.3157554 expected
Constructed five-cluster dirty granules 44
Clustered fraction of whole image 5.6265985%
Suppression vs Bernoulli expectation 94.1509081%
Forty-four 4 KiB granules 180,224 B
Compact (2500(64+4)+256) delta 170,256 B
All five clusters granule-aligned 40 granules
All five clusters start at residue 13 45 granules

The reproduction program also evaluates the identical record-set construction under a 16 KiB logical granule, where it touches 14 granules. This is a useful reminder that “number of pages” depends on the granularity definition even when the semantic mutations are identical.

Why logical granules cannot substitute for memory-controller counters. A rigorous physical experiment requires at least three observation layers.

At the software-state layer, record indices and dirty-bitmaps establish exactly which logical chunks changed.

At the virtual-memory layer, an operating system mechanism such as Linux soft-dirty tracking can observe which virtual pages were written since the tracking state was reset. Linux explicitly documents the PTE operation used for this mechanism.

At the physical-memory layer, vendor performance counters or equivalent instrumentation must measure cache-line eviction, memory-controller transactions, transferred bytes, or device-specific write traffic. The experiment must then determine how strongly the logical dirty set predicts those physical quantities under the tested cache policy and access stream.

Only the third layer can substantiate a claim such as “94% less memory-controller writeback.”

Likewise, dirtying fewer pages does not inherently cause a proportional reduction in TLB shootdowns. TLB invalidation is associated with mapping/protection changes and related translation-management events, not merely ordinary writes to already mapped data.

5.7 Embedded and appliance profiles

Smart-appliance arithmetic. The supplied smart-oven scenario compares a 3.2 MB state image with a 512 KiB memory budget. With decimal 3,200,000 bytes and 524,288 bytes of SRAM,

[ = 6.103515625. ]

Thus the object equals 610.35% of SRAM capacity, or 510.35% above capacity. Calling this “+625% overrun” is not reproduced by those stated numbers.

A 96-byte frame occupies

[ = 0.0183105% ]

of that SRAM.

The correct systems conclusion is also narrower than “binary beats JSON by 625%”: a 96-byte command and a 3.2 MB complete state graph are not semantically equivalent representations. The architecture’s advantage is that the endpoint need only receive the bounded command necessary for its responsibility. A small text command could also fit. The benchmark should therefore compare equivalent endpoint semantics.

Smart-TV result. The source reports 0.042 ms binary-delta ingestion and a 0 ms observed GC pause in the measured ATESO path, compared with an approximately 14 ms GC pause in its document-oriented comparison.

The academically valid phrase is “no GC pause observed during that measured path,” not “garbage collection is impossible.” Modern managed runtimes can also be designed for bounded or low-latency collection; IBM’s Metronome work explicitly targeted real-time garbage collection and reported bounded pauses under its evaluated conditions. Therefore ATESO’s zero-heap claim should be evaluated against real-time and low-latency managed baselines, not only stop-the-world collectors.

5.8 SPSC reproduced result

SPSC reproduced result. The supplied reproducibility archive contains a new Node.js shared-memory reference program rather than an unavailable original implementation. Its tests use three capacities, 200,000 records each, sixteen 32-bit payload words per record, and counters initialized near unsigned wrap.

Capacity Records verified 32-bit payload words verified Wrap forced Result
2 200,000 3,200,000 Yes Pass
64 200,000 3,200,000 Yes Pass
1,024 200,000 3,200,000 Yes Pass
Total 600,000 9,600,000 Yes Pass

This execution validates ordering, payload integrity, backpressure, final counter equality, and the tested wrap path. The theorem in Section 3 supplies the general safety argument; the execution does not exhaust all possible schedules.

5.9 Publication-grade benchmark protocol

A publication-grade physical benchmark protocol. Every representation experiment should compare at least:

[ B1:optimized full JSON,B2:optimized JSON delta,B3:established binary format,B4:binary delta,B5:ATESO resident delta.\begin{array}{ll} B_1 &: \text{optimized full JSON},\\ B_2 &: \text{optimized JSON delta},\\ B_3 &: \text{established binary format},\\ B_4 &: \text{binary delta},\\ B_5 &: \text{ATESO resident delta}. \end{array}

]

All five paths must produce the same accepted logical state.

For each run, record:

hardware model and silicon revision
operating system / firmware
runtime / browser / compiler version
GPU adapter and exposed driver information
power and thermal operating state
state schema and exact byte layout
record count and mutation indices
mutation topology
allocation count and bytes
CPU admission time
upload bytes and upload-call count
GPU submission and completion time
solver residual / output checksum
logical dirty set
virtual-page dirty set when observable
physical memory / device traffic when observable
presentation latency when presentation is in scope
raw per-trial timestamps
source revision and executable digest

Warm-loop iterations are not independent hardware replications. Device/session should be the principal replication unit for a cross-silicon claim.

Latency-floor methodology. If an operation approaches timer resolution, execute (R) logically independent operations in one timed batch:

[ = . ]

Sweep (R) until the batch interval is comfortably above clock quantization and report the timer’s effective resolution. Do not print 0.00 ms and derive an arbitrary finite ratio afterward.

Zero-allocation methodology. A real zero-heap claim should use an allocator interposition or equivalent runtime mechanism that increments a counter on every dynamic allocation on the measured path:

[ A_{}=0, A_{}=0. ]

For an MCU, additionally publish the linker map, static RAM footprint, maximum observed/analytically bounded stack, interrupt-stack budget, crypto scratch region, and all DMA/control buffers.

Hard-real-time qualification. Replace “mean (<10,s)” with a testable task model:

[ _i=(C_i,D_i,P_i), ]

where (C_i) is a justified WCET, (D_i) is deadline, and (P_i) is minimum interarrival period. Existing deadline scheduling literature and systems such as SCHED_DEADLINE, Shinjuku, Shenango, and Caladan illustrate why interference, preemption boundaries, and admission matter in addition to average service time.

5.10 Claim-status register for the Day-100 additions

The evidence discipline of §5.1 is extended with an explicit claim-status register. Several of the numerical claims introduced or sharpened in this revision are mathematically exact but have a different evidentiary status from one another and from the Apple WebGPU numbers in §5.3–§5.5. Each row states what is established now, its evidence class under the notation defined in the front matter, and the governing wording for that claim in this manuscript.

Proposed claim What is established now Class Publication-safe wording used in this paper
36 GB M5 Max + 16 GB M1 Pro = 52 GB Device configurations are vendor-documented [57, 58]; binary units come from the machine’s hardware report ((36^{30}) B), to be filed under Appendix A.1; arithmetic is exact V (configurations) + U (hardware report) + R “52 GB nominal aggregate memory across two independent hosts.”
Machine of record is the 32-core-GPU M5 Max bin (460 GB/s, not 614 GB/s) Per-bin bandwidth is vendor-documented [57]; the bin assignment rests on the machine’s own hardware report (Mac17,7, 18-core CPU, 32-core GPU, 36 GB), not in the reproduction package V (bandwidth) + U (bin identity) “Documented at 460 GB/s for the 32-core-GPU bin its hardware report places it in; the report is to be filed under Appendix A.1.”
812.5 M × 64 B = 52 GB Exact decimal-convention arithmetic (§2.2, Appendix A.5); 872,415,232 is the nominal hardware bound in binary units (unit basis from the machine hardware report, §2.2; U) R (arithmetic) + U (unit basis) “Decimal-convention capacity of 812.5 M records (872.4 M under binary units) before OS/runtime reservation.”
“Successfully holds 812.5 M live entities” Not established by material available in the reproduction package P Used only with per-host RSS/VM footprint, shard count, integrity scan, duration, and failure-threshold trace (Appendix A.1).
156:1 = 99.36% reduction (8.112 TB equivalent baseline at 812.5 M entities) Exact, conditional on a measured 9,984 B/entity baseline (§2.2, Appendix A.5) U (ratio) + R (algebra) “156:1 retained-size compaction in the specified object-graph benchmark.”
150–200 GB → 52 GB Exact ratio is 2.88×–3.85×, not 156×; the 150–200 GB envelope is quoted from the source directive [48], not measured here R (ratio) + U (envelope; source directive [48]) Reported separately as a 65.3–74.0% lower memory envelope.
30 GB/s → 3.84 kB/s Exact when (S=500) MB, (f=60) Hz, (k=1), (b=64) B, payload only (§2.4, Appendix A.5) R “7,812,500-fold payload-volume reduction in the one-record-per-frame sparse-delta scenario.”
30 GB/s → 9.6 kB/s with one CBP-96 envelope per publication Exact when (h_t=96) B, (k=1), (b=64) B; still excludes the 4 B record identifier, transport headers, acknowledgments, and retransmission (§2.4, Appendix A.5) R “3,125,000:1 payload-plus-envelope ratio (99.999968% suppression) in the one-record-per-frame scenario.”
Same percentage energy saving Not established P Measured joules or hardware energy counters only (§2.4, §2.7).
Notebook API can ingest continuously Programmatic notebook/source ingestion exists [52, 53], but source copies are static [59] V “Automated revision-triggered snapshot ingestion.”
Programmatic Audio Overviews Documented, Pre-GA (Preview) [54] V Safe with the Preview qualification.
Automated revision-triggered export pipeline (§4.6) Contract specified (Appendix A.4); no run, source ID, or (d_r)/(e_r) record is in the package P “ATESO specifies an export contract; a run is reported only with the A.4 record.”
Programmatic Notebook Video Overview API NotebookLM supports Video Overviews [56]; no public API endpoint located in the official documentation reviewed P API automation is not claimed.
“long-horizon serving stack.1 proves ATESO compaction” False attribution; physics seat reports its own KV-cache results [49] V (for physics seat’s own numbers) “Independent evidence that long-horizon agent serving is constrained by persistent-state footprint and movement.”
“Eliminates 3–4× heap wrappers” (source directive [48]) Plausible only as a workload measurement, not a V8 invariant [50, 51]; no such measurement is reported in this paper U (if measured) Not stated; any future figure is reported as a measured retained-size ratio against the named V8 object schema (Appendix A.2).
94.2% fewer MMU page-table writebacks (source directive [48]) Only the logical dirty-granule occupancy is established: 94.1509% exact (§2.3, §5.6); device runs 94.11–94.20% (§5.2) F (94.1509%) + U (device range) + P (physical) “94.1–94.2% reduction in logical dirty-granule occupancy for the specified experiment.”
Xbox 40,400× Timer-floor microbenchmark ratio using a 1 µs floor as denominator (§5.4) U “Timer-floor microbenchmark ratio, not an end-to-end application speedup.”
580,000 TPS burst and sub-microsecond enqueue latency (§6.1) Author-reported; raw samples, shard configuration, core count, verification scope, timer resolution, and executable digest not in the package U “Reported burst throughput and enqueue latency, bounded by the §5.9 protocol.”
Pre-allocated ingress ring budget 402,653,184 B, about 403 MB (§6.1) Design parameters (W=16), (M_{}=1{,}024); the arithmetic (WM_{}) B is exact (Appendix A.5); not measured P (parameters) + R (arithmetic) “About 403 MB of pre-allocated SPSC lanes at the reference parameters; (W) and (M_{}) are recorded with the run.”
102,400 tx/s aggregate floor (§6.1) Design parameter: 256 shards × a specified ≥400 tx/s per shard; not measured P “Aggregate design floor implied by the per-shard specification.”
70%–99% netting reduction (§6.4) Not measured in this work; magnitude depends on the transaction graph P “Expected to reduce settlement count and liquidity requirements substantially; magnitude is P until measured on the engine’s own logs.”

Every numerical claim introduced or sharpened in this revision has a row in this register. Measurements carried from v1 keep the evidence class assigned where they appear (§5.1–§5.9). The register is the authoritative wording for the rows it contains; any sentence elsewhere that appears to make a stronger claim is to be read as bounded by the corresponding row.

6. Machine-to-Machine (M2M) Zero-Custody Micro-Accounting and Merkle Batch-Root Netting Engine

Section 6 — High-Density Attestation and Regulated Settlement Separation. Autonomous AI agent swarms and cyber-physical execution planes call for sub-millisecond, sub-cent micro-accounting for tool invocation, data leasing, and compute slices. Traditional financial rails (card interchange, ACH) and public blockchain networks are structurally mismatched to this workload: commonly published card-processing schedules carry a fixed per-transaction component on the order of tens of cents, which exceeds a sub-cent micro-task by orders of magnitude, while public blockchains introduce variable fee markets, mempool ordering exposure, and serialized state growth. These are premises for the design that follows, not measurements of this work.

ATESO addresses this mismatch by separating machine-speed cryptographic execution accounting from regulated fiat settlement:

┌─────────────────────────────────────────────────────────────────────────────────────────┐
│                    ATESO M2M ZERO-CUSTODY NETTING PIPELINE                              │
├─────────────────────────────────────────────────────────────────────────────────────────┤
│ 1. Machine Event ──► lease-referenced CBP-96 envelope via §3.6 ingress into .spine SPSC │
│ 2. Shard Admission ──► 256 Disjoint Shard Loggers (≥102,400 tx/s design floor)          │
│ 3. Per-Shard Merkle Batch ──► Local Merkle Root (R_s) calculated over 100ms - 1s batch  │
│ 4. Global Checkpoint ──► Merkle Super-Root (R_global) anchored to Rekor (log receipt)   │
│ 5. Netting Engine ──► Multilateral net settlement vector from bilateral accumulators    │
│ 6. Fiat Settlement ──► Authenticated hand-off to a licensed depository institution      │
└─────────────────────────────────────────────────────────────────────────────────────────┘

6.1 Command Profile and Lock-Free Sharded Ingestion (CBP-96)

To keep heap-allocation jitter and global lock contention off the hot path across thousands of concurrent agent routines, the execution plane is partitioned into 256 disjoint namespace shards:

[ T_{} = _{s=1}^{256} S_s = 256 = ]

This is a design parameter (P), not a measurement.

Benchmarked execution on Apple Silicon (M5 Max / M1 Pro) is reported to sustain burst throughput exceeding 580,000 TPS with sub-microsecond enqueue latency. Both figures are author-reported measurements (U) whose raw per-trial samples, shard configuration, core count, and executable digest are not in the reproduction package; the corpus does not state whether the throughput counts post-signature-verification admissions or pre-verification enqueues, and the latency is reported without timer-resolution disclosure, so both are bounded by the §5.9 protocol and its latency-floor methodology. They are registered in §5.10 and remain U until those artifacts are supplied. The 102,400 tx/s figure above is the floor implied by the per-shard specification, not a ceiling.

6.2 Per-Shard Event Chaining and Usage Metering

Each shard maintains an independent, continuous cryptographic event hash chain. For each accepted state transition (i):

[ h_i = H( h_{i-1} _i _i |_i| _i) ]

where (h_{i-1}) is the prior event digest, (_i) is the shard’s monotonically increasing 64-bit event revision (distinct from the state revision (r) of §3.3), (_i) is the full 96-byte CBP-96 envelope (32-byte header followed by its 64-byte Ed25519 signature), and (_i) is the length-prefixed admission metadata (shard identifier, admission timestamp, broker verdict, the 32-byte sender public key resolved from the sender index, and the 32-byte digest of the .moat lease resolved from the receiver index). Registry bindings for an index are immutable for the life of that index; reassignment allocates a new index, so the resolved key and lease that the preimage binds are the ones the broker checked. Each shard chain starts from (h_0^{(s)}=H(sg)) for shard index (s) (the 1-byte shard ID of the CBP-96 header) and the 32-byte deployment genesis value (g) of §3.4. The domain tag is distinct from the §3.4 ATESO/EVENT tag so that the two preimage layouts cannot collide, and the signature is inside the preimage so that an inclusion proof establishes an authorized, signed transition rather than an unsigned header.

Simultaneously, local usage meters update bilateral debit and credit accumulators between transacting entity pairs:

[ {A B} = {A B} - _{B A} ]

6.3 Merkle Batch-Root Aggregation and Transparency Log Anchoring

Individual transactions are never written directly to high-overhead public ledgers. Instead, shards aggregate transactions into hierarchical binary Merkle trees over periodic batch epochs (100 ms to 1,000 ms):

  1. Local Shard Merkle Root ((R_s)): For a batch of event digests ({h_1, , h_k}), leaf nodes are computed as (L_j = H( h_j)) and collapsed into a binary tree yielding root (R_s).
  2. Global Merkle Super-Root ((R_{})): A deterministic checkpointing aggregator combines the 256 shard batch roots of epoch (n) into a second-level Merkle tree and links each checkpoint to its predecessor:

[ R_{}^{(n)} = H( n R_{}^{(n-1)} (R_1,,R_{256}) ), ]

with RFC 9162-style [22] domain separation (textual tags in place of the RFC’s 0x00/0x01 prefixes, with the RFC’s largest-power-of-two split rule) between leaves ((L_j=H())) and interior nodes ((H())) at both levels, with an empty shard batch contributing the RFC 9162 empty-tree hash (R_s=H()), with (n) encoded as a fixed-width 64-bit big-endian integer, and with checkpoint genesis (R_{}^{(0)}=H(g)) for deployment genesis (g).

  1. Public transparency anchoring: (R_{}^{(n)}) is anchored to a public append-only transparency log (Sigstore Rekor) by an anchoring client (the author’s reference client scfpassport), yielding tamper-evident receipt identifiers (rcpt_...) and chained event digests (evt_...) relative to the log’s trust root (§3.4; AiST, §1). Any individual micro-transaction (h_i) then possesses a Merkle inclusion proof of length (O(k+)) to the anchored checkpoint that establishes its presence in the batch without revealing confidential payload data. Anchoring makes the record tamper-evident relative to the log, not immutable. Anchoring cadence is decoupled from the batch epoch: because each checkpoint commits to its predecessor, anchoring every (a)-th checkpoint (for example one log entry per minute) leaves every intermediate checkpoint tamper-evident relative to the next anchor; (a) is a deployment parameter recorded with the run.

6.4 Multilateral Netting and Conservation of Value

At scheduled settlement intervals, the netting engine collapses the bilateral accumulator matrix into a reduced multilateral settlement vector:

[ P ,(P) = {Q P} (Q P) - {Q P} (P Q) ]

Before dispatch the engine asserts the conservation identity, which holds algebraically for any consistent bilateral accumulator matrix; a violation indicates accumulator corruption or an incomplete update and aborts settlement:

[ _{P } (P) ]

A greedy matching algorithm pairs net debtors with net creditors. Multilateral netting is expected to reduce settlement count and total required liquidity substantially relative to gross volume; the magnitude for this engine depends on the transaction graph and is P until measured on the engine’s own transaction logs (§5.10); the source directive’s 70%–99% figure [48] is not adopted.

6.5 Regulated Instant Settlement Handoff

For each (net debtor, net creditor) pair produced by §6.4, the engine issues a customer-to-bank instruction only: either an ISO 20022 pain.001 credit-transfer initiation to the net debtor’s licensed depository institution under a standing mandate, or, acting for the net creditor, a request for payment lodged with the net creditor’s institution, which originates the rail’s pain.013 request-for-payment message to the debtor’s institution for presentment to the debtor, whose acceptance is required before any value moves [60]; in either case the instruction carries the settled checkpoint root in the remittance field. The debtor’s institution, as the participant in an instant-payment rail such as the Federal Reserve’s FedNow service, then originates the FI-to-FI customer credit transfer (pacs.008 [60]) in the message version its rail specifies; the engine never originates an interbank message itself:


7.1 Representation systems

Section 7 — Representation systems. The closest prior work makes the ATESO contribution narrower and more defensible.

System / lineage Already established What ATESO adds if implemented as specified
Apache Arrow Language-neutral contiguous layouts, O(1) access, relocatability, zero-copy shared-memory access Mutable resident operational state tied to command capabilities, revision, scheduler and telemetry
Cap’n Proto Directly traversable structured encoding and RPC/capability ideas Unified real-time mutation/provenance/physics lifecycle
FlatBuffers Direct buffer access without unpacking; read path can require no additional heap Stateful execution contract and bounded command path
seL4 Capability-controlled objects, machine-checked kernel properties, no kernel heap Application/state representation and agent-command layer above the OS
CHERI Hardware capability provenance, bounds and monotonic derivation Wire/state-level delegated command semantics across non-CHERI profiles
XPBD Compliant constraint projection Integration with resident state, colored dispatch and command authority
HTN / ReAct / actor-critic Planning, decomposition, reasoning/action, policy optimization Separation of stochastic planning from deterministic binary admission
EDF / Shinjuku / Shenango / Caladan Deadline and microsecond-scale scheduling/resource control Semantic revision/capability metadata bound to tasks
Merkle transparency logs Authenticated inclusion and append-only consistency evidence State-command provenance integrated with execution revisions

Arrow’s current format specification explicitly advertises relocatability without pointer swizzling and true zero-copy access in shared memory. FlatBuffers explicitly advertises direct access without unpacking and no heap requirement for read access. These systems preclude any defensible claim that .many is the first directly traversable or zero-copy-capable binary representation.

What .many proposes instead is a mutable execution image whose identity is shared with command admission, provenance, physics, scheduler metadata, and telemetry.

7.2 Capability systems

Capability systems. seL4 is a capability-based operating system with deep formal verification, and its documentation defines capabilities as authority-bearing object references. CHERI provides hardware-supported capabilities designed to preserve provenance and restrict derived authority. ATESO’s non-amplification theorem is therefore best understood as an application-level delegation invariant compatible with this lineage.

A serious implementation can strengthen the .moat theorem by mapping capability objects onto an underlying seL4 or CHERI enforcement mechanism on supported native platforms. On conventional platforms, .moat remains cryptographic/brokered authority rather than hardware memory authority.

7.3 Planner frameworks

Planner frameworks. ReAct demonstrated that model reasoning and external actions can be interleaved, while HTN methods formalize decomposition into hierarchical task networks and actor-critic methods formalize another two-timescale decision process. ATESO’s distinctive claim is not a new planning algorithm; it is the typed execution membrane below the planner.

This distinction matters operationally. An LLM may hallucinate an invalid command. The desired invariant is not “the model never hallucinates,” but:

[ ( ) ]

under correct broker enforcement, because invalid output is outside the executable language or fails admission.

7.4 Garbage collection and managed runtimes

Garbage collection and managed runtimes. The claim that GC necessarily produces 15–50 ms pauses is false as a universal proposition. Real-time collectors such as Metronome were explicitly designed around bounded pauses. ATESO’s stronger and simpler claim is that an allocation-free fast path has no GC work caused by allocations on that path. This improves analyzability but does not make all managed runtimes unsuitable.

7.5 WebGPU

WebGPU. WebGPU provides portable mediated GPU resources, not a universal promise of CPU/GPU virtual-memory identity. The current specification’s mapping/usage rules make this explicit. Therefore Magma’s browser-profile differentiation should be persistent device-buffer ownership, range/delta minimization, and controlled dispatch—not a claim of portable direct SharedArrayBuffer-to-storage-buffer aliasing.

7.6 Contemporary agent-state compaction

The pressure to minimize resident state is also visible in contemporary long-horizon model serving. a contemporary long-horizon serving stack reports a global KV-cache footprint of 890 bytes per token, approximately one quarter of its V4-Flash predecessor, and a persistent KV footprint of approximately one eighth of the predecessor through a combination of cross-layer reuse, compressed sparse attention, low-precision caching, and bounded replay [49]. This result is not evidence for ATESO’s measured entity-compaction ratios; it is independent evidence that state capacity and state movement have become first-order systems constraints for long-horizon agent execution. ATESO attacks the analogous problem at the application/runtime state boundary rather than the Transformer KV-cache boundary. The paper accordingly uses the label “long-horizon serving stack.1 external corroboration of state-compaction pressure,” never “long-horizon serving stack.1 silicon proof of ATESO,” and attributes no ATESO measurement to that work.

7.7 Economic interpretation

The cloud serialization-tax equation. Define

[ C_{} = C_{} s_{} _{}, ]

where:

[ C_{} = , ]

[ s_{} = , ]

and

[ _{} = . ]

The dimensions are

[ [$/] = [$/]. ]

The equation is valid accounting. A $10 billion annual tax is not established by the current evidence, because neither a representative global (s_{}) nor a representative (_{}) has been measured.

For illustration only, if

[ C_{}=$100, s_{}=0.25, _{}=0.40, ]

then

[ C_{} = $10. ]

That is a scenario, not a result.

The supplied directive proposes representation fractions as high as 70% and avoidability above 90%; this paper does not elevate those planning assumptions into findings without a representative fleet study.

A credible economic study would attribute CPU cycles, allocator work, GC, memory traffic, networking, storage, and egress separately to avoid double-counting. It would compare equivalent service-level objectives for at least 30 days and report:

[ C= C_{} - C_{} - C_{}. ]

A gain-share business model could contractually charge a fraction of verified (C), but that commercial mechanism is distinct from the technical theorem.

7.8 Prior-art differentiation

Prior-art differentiation in one sentence.

[ ]

None of its constituent primitives should be represented as unprecedented when the literature already contains direct-layout formats, capability operating systems, real-time scheduling, constraint solvers, authenticated logs, and agent planners.

7.9 Architecture scope

Architecture scope. The present manuscript explicitly discloses the following combination as one architecture:

  1. persistent relocatable mutable state addressed by stable offsets;
  2. sparse-delta executable descriptors separated from planner text;
  3. monotonically narrowing cryptographic command capabilities;
  4. expected-revision admission;
  5. fixed-size zero-heap command execution profiles;
  6. independent SPSC lanes for trusted local workers;
  7. brokered gateways for untrusted models;
  8. incremental provenance over accepted mutations;
  9. GPU-resident physics with conflict-free constraint coloring;
  10. typed closed-loop telemetry returning into the persistent state plane;
  11. thermal/deadline admission over the same task identity.

Whether any element or combination is patentable in any jurisdiction is a separate legal question; this paper makes a technical disclosure, not a patentability opinion.

8. Conclusion

Section 8 — Conclusion. The central ATESO/Magma result is more precise than the broad claim that document-centric computing is universally obsolete.

For admitted long-lived state, the representation-maintenance theorem establishes that a resident implementation can replace recurring ((Nb)) reconstruction work with work proportional to identified changes plus required validation. The result becomes increasingly strong as update sparsity grows and state lifetime extends.

For spatially clustered updates, the locality theorem establishes that changed-granule cardinality is governed by payload size and fragmentation,

[ D= O!(Km+C), ]

rather than by total resident-state size. In the specified 50,000-record experiment, the independent Bernoulli model yields an expectation of 752.2535 dirty 4 KiB logical granules, while a specified five-interval layout touches exactly 44, a 94.1509% reduction. That is an exact mathematical result under the stated mapping. It should not be mislabeled as an automatically identical physical DRAM, PTE, TLB, or energy reduction.

The capability theorem establishes monotonic authority only when the broker is the complete authority boundary. The revision theorem excludes stale commits only when revision change is atomic and durable rollback is separately controlled. The SPSC theorem establishes slot exclusion under its single-producer/single-consumer and bounded-occupancy assumptions. Colored XPBD establishes write-conflict freedom only when the coloring graph covers every concurrently writable datum. The thermal guarantee holds only when the prediction-error envelope is valid.

These qualifications are not weaknesses. They are the conditions that turn an architecture into a scientific object.

The strongest result is therefore not that representation boundaries disappear. They do not. The result is that their location becomes explicit. Within an admitted execution domain, long-lived packed state can remain resident and receive sparse mutations without rebuilding a semantically equivalent object graph. Between independent machines, between CPU and mediated GPU domains where the API requires transfer, and between ATESO and external cloud research services, the boundary is modeled and measured rather than renamed “zero copy.”

This distinction becomes increasingly important as agent systems retain larger working states. Contemporary model-serving work independently demonstrates the same systems pressure at the KV-cache layer; ATESO addresses it at the application-state layer. The resulting principle is narrower than a universal claim about binary formats and stronger than one:

[ ]

Under those conditions, a 64-byte state vector is not a claim of zero-latency computing. It is a precisely bounded representation from which memory capacity, mutation traffic, authorization cost, and transfer work can be measured without the hidden multiplicative structures of a general-purpose object graph.

The resulting operating model is:

[ ]

In that form, ATESO is neither a replacement for an operating system nor merely another serialization format. It is a proposed state-first execution contract connecting representation, authority, concurrency, physics, provenance, scheduling, and agent feedback.

Appendix A — Reproducibility

A.0 Reproduction package

A.0 — Reproduction package. The available ATESO reproduction archive contains:

ateso_repro/
    README.md
    SHA256SUMS
    audit.py
    receipt_manifest.json
    source_directive.md
    spine_spsc_test.mjs
    results/
        audit.json
        spine.json

Its README explicitly states that the archive contains newly executed analytical calculations and a newly implemented SPSC reference correctness test, but not the original author-reported WebGPU benchmark implementation, raw heterogeneous-fleet traces, historical-device ports, signed AiST receipts, or original scheduler simulator.

The recorded environment is Python 3.13.5, Node.js 22.16.0, Linux x86-64.

The exact reproduction procedure after extraction is:

cd ateso_repro

# Verify the distributed snapshot.
sha256sum -c SHA256SUMS

# Recompute all analytical results.
python3 audit.py --out results/audit.json

# Execute the SPSC wrap/order/content test.
node spine_spsc_test.mjs

The relevant distributed SHA-256 identities are:

source_directive.md
26126d43067a2ab2bfe523448b2ecf4f357eafb70f29f6f4cdb5085ed9d876bc

audit.py
d72ccbb0f1b7dc2f229fb0ff77cdc77022f49bf3d3797e18a23e09f604bc11f0

spine_spsc_test.mjs
ffcfef5275c9f39c0fd8b37126980b4e6f68a26005fe8f2c0eab92ddde676b65

These hashes identify the distributed bytes. They are not digital signatures, trusted timestamps, or proof that an external hardware run occurred.

Minimal locality harness. The central 94.1509% result can be independently reproduced without any ATESO implementation:

from __future__ import annotations

import math

N = 50_000
RECORD_BYTES = 64
GRANULE_BYTES = 4096
RHO = 0.05

records_per_granule = GRANULE_BYTES // RECORD_BYTES
full_granules, tail_records = divmod(N, records_per_granule)

expected_random = full_granules * (
    1.0 - (1.0 - RHO) ** records_per_granule
)

if tail_records:
    expected_random += 1.0 - (1.0 - RHO) ** tail_records

starts = [0, 6413, 12813, 19213, 25613]
lengths = [500] * 5

dirty: set[int] = set()

for start, length in zip(starts, lengths, strict=True):
    for record in range(start, start + length):
        dirty.add(
            (record * RECORD_BYTES) // GRANULE_BYTES
        )

clustered = len(dirty)
suppression = 1.0 - clustered / expected_random

print(f"granules={math.ceil(N / records_per_granule)}")
print(f"expected_random={expected_random:.10f}")
print(f"clustered={clustered}")
print(f"suppression={suppression * 100:.10f}%")

assert clustered == 44
assert abs(expected_random - 752.2535206074667) < 1e-9
assert abs(suppression - 0.9415090806561215) < 1e-12

Expected output:

granules=782
expected_random=752.2535206075
clustered=44
suppression=94.1509080656%

Required next empirical standard for an arXiv revision. The central unresolved empirical obligation is not another architecture slogan. It is a fully public, content-addressed benchmark bundle containing the original source, exact build, raw per-trial samples, hardware/OS/runtime identifiers, timer calibration, state checksums, allocator traces, page-tracking outputs, physical memory or device counters where available, and cryptographically bound result manifests.

The current paper can already stand on its mathematical contributions without overstating those pending measurements.

A.1 Distributed resident-state capacity experiment

The distributed-capacity experiment MUST record the two hosts independently. The result artifact contains, at minimum:

The experiment does not sum memory capacities and call the result a unified heap. Each host independently proves its resident shard:

[ N_i^{()} = #{ xS^{(i)} : (x)= }. ]

The global verified population is then

[ N_{} = _iN_i^{()}. ]

A run qualifies as an 812.5-million-entity demonstration only when

[ N_{} 500{,}000 ]

and every reported shard passes the integrity scan while remaining resident under the experiment’s stated residency criterion. Allocated virtual address space is not the success criterion: a process can reserve enormous virtual ranges without those bytes being physically resident. The artifact reports resident memory, swap/compression behavior, and full-record readback. Vendor specifications [57, 58] establish nominal capacities; they cannot establish the workload’s usable resident capacity.

A.2 V8 retained-size comparison

The (156{:}1) result is to be reproduced against a pinned JavaScript runtime and an exact schema. The harness records

[ M_0 = , ]

[ M_1 = , ]

and computes

[ b_h = , C_{} = . ]

Heap snapshots and runtime flags are preserved. The experiment is repeated across independent fresh processes rather than treating repeated reads of one heap as independent samples.

The paper reports the object schema, property types, array strategy, engine version, pointer-compression state [50], heap flags, and whether external ArrayBuffer backing stores are included. The retained-memory baseline is taken after garbage collection and the engine version is pinned, because V8 documents that object representation [51] and pointer compression [50] affect memory layout; that is exactly why the result needs a reproducible runtime/schema boundary.

A.3 Sparse-delta wire experiment

The logical payload calculation of §2.4 is supplemented by a packet- or application-layer measurement. For each run, record

[ B_{}, B_{}, B_{}, B_{}, B_{}, ]

with

[ B_{} B_{}. ]

For a one-record-per-frame test at 60 Hz,

[ B_{} = 3{,}840 . ]

The reported (99.9999872%) result is explicitly labeled payload-volume suppression unless the numerator and denominator are both measured at the same lower-layer observation boundary.

A.4 Research-egress provenance

Each notebook export records

[ ( r, H(S_r), H(X_r), H(A_r), H(C), r, d_r, e_r, , t{}, ). ]

The export event (e_r) is chained (§4.6), so an export that is absent from the chain, or a chain that lacks the export it claims, is detectable; it is anchored only when the batch containing it is anchored, and the record carries the anchor receipt when one exists.

Generated audio, video, slides, summaries, and other research outputs are derivatives rather than authoritative state and must be linked back to the source artifact and state revision from which they were produced. The API status is pinned in the manifest: at the time of this revision, notebook management, source ingestion, and Audio Overview generation are documented as Pre-GA (Preview) [52–54], and the Drive MCP server is in the Google Workspace Developer Preview Program [55]; no public Video Overview API is pinned (§4.6).

A.5 Day-100 arithmetic harness

Every number introduced in §2.2 and §2.4, the §2.3 device-run percentages (arithmetic on the author-reported granule counts only; the counts themselves remain U), the §6.1 ring-budget arithmetic, and the corresponding §5.10 rows are reproduced by the following listing, which needs no ATESO implementation. It is the executable artifact behind the R class assigned to those figures.

from __future__ import annotations

from fractions import Fraction as Fr

B_P = 64                                   # packed record width, bytes

# --- §2.2 capacity identity ---------------------------------------------
M_DEC = (36 + 16) * 10**9                  # decimal-convention aggregate, B
M_BIN = (36 + 16) * 2**30                  # binary units per the hardware report (§2.2), B
assert M_DEC // B_P == 812_500_000
assert M_BIN // B_P == 872_415_232
assert M_BIN == 55_834_574_848
assert 36 * 2**30 == 38_654_705_664           # hw.memsize of the machine of record

# --- §6.1 ingress ring budget -------------------------------------------
assert 16 * 256 * 1024 * 96 == 402_653_184

# --- §2.2 representation compaction --------------------------------------
C_REPR = 156
B_H = C_REPR * B_P
assert B_H == 9_984
ETA_MEM = 1 - Fr(1, C_REPR)
assert abs(float(ETA_MEM) - 0.9935897436) < 1e-10
assert 812_500_000 * B_H == 8_112_000_000_000
assert abs(150 / 52 - 2.8846) < 5e-5 and abs(200 / 52 - 3.8462) < 5e-5
assert abs((1 - 52 / 150) * 100 - 65.33) < 5e-3
assert abs((1 - 52 / 200) * 100 - 74.00) < 1e-9

# --- §2.3 device-run locality range --------------------------------------
for dirty, scattered, pct in ((44, 747, 94.11), (44, 756, 94.18), (44, 758, 94.20)):
    assert abs((1 - dirty / scattered) * 100 - pct) < 5e-3

# --- §2.4 sparse-delta transport -----------------------------------------
S, F, K = 500 * 10**6, 60, 1
B_DENSE = F * S
B_DELTA = F * K * B_P
assert B_DENSE == 30 * 10**9 and B_DELTA == 3_840
assert Fr(B_DENSE, B_DELTA) == 7_812_500
assert abs((1 - B_DELTA / B_DENSE) - 0.999999872) < 1e-12
H_T = 96                                   # one CBP-96 envelope
B_DELTA_ENV = F * (H_T + K * B_P)
assert B_DELTA_ENV == 9_600
assert Fr(B_DENSE, B_DELTA_ENV) == 3_125_000
assert abs((1 - B_DELTA_ENV / B_DENSE) * 100 - 99.999968) < 1e-9
assert Fr(B_DELTA_ENV, B_DELTA) == Fr(5, 2)

print("day-100 arithmetic: all assertions hold")

Expected output:

day-100 arithmetic: all assertions hold

References

Bracketed numbers throughout refer to this list. Sections carried from v1 mostly cite their sources by name in prose; those sources are listed as [1]–[48]. Dimension brackets in the §7.7 equation are not citations.

[1] Apache Arrow Project. Arrow Columnar Format, Version 1.5. Defines language-independent in-memory layouts, constant-time access for the principal array forms, relocatability without pointer swizzling, alignment guidance, and shared-memory zero-copy semantics.

[2] Cap’n Proto Project. Encoding Specification. Defines the wire/in-memory object representation and traversal rules for Cap’n Proto.

[3] Google. FlatBuffers Documentation. Documents direct access without parsing/unpacking and buffer-only memory requirements for read access.

[4] Geoff Langdale and Daniel Lemire. 2019. “Parsing Gigabytes of JSON per Second.” The VLDB Journal 28(6). DOI 10.1007/s00778-019-00578-5.

[5] Roy Thomas Fielding. 2000. Architectural Styles and the Design of Network-based Software Architectures. Doctoral dissertation, University of California, Irvine.

[6] W3C GPU for the Web Working Group. 2026. WebGPU. Current specification consulted for buffer creation, mapping, queueing, copying, and device-timeline semantics.

[7] W3C GPU for the Web Working Group. 2026. WebGPU Shading Language. Consulted for memory scopes, atomic semantics, workgroup barriers, and storage synchronization.

[8] Ecma International / TC39. ECMAScript Language Specification: Memory Model. Basis for JavaScript SharedArrayBuffer/Atomics synchronization in the reference .spine profile.

[9] WHATWG. HTML Living Standard: Agents and Agent Clusters. Defines browser agent-cluster boundaries relevant to shared memory.

[10] Miles Macklin, Matthias Müller, and Nuttapong Chentanez. 2016. “XPBD: Position-Based Simulation of Compliant Constrained Dynamics.” Proceedings of Motion in Games, 49–54. DOI 10.1145/2994258.2994272.

[11] Tero Karras. 2012. “Maximizing Parallelism in the Construction of BVHs, Octrees, and k-d Trees.” High Performance Graphics.

[12] Kevin Skadron et al. HotSpot: A Compact Thermal Modeling Methodology. University of Virginia LAVA Laboratory. HotSpot models processor thermal behavior with equivalent thermal resistance-capacitance networks.

[13] Linux Kernel Documentation. Deadline Task Scheduling. Documents Linux SCHED_DEADLINE, EDF/CBS behavior, admission concepts, WCET terminology, and deadline task parameters.

[14] Kostis Kaffes, Timothy Chong, Jack Tigar Humphries, Adam Belay, David Mazières, and Christos Kozyrakis. 2019. “Shinjuku: Preemptive Scheduling for μsecond-scale Tail Latency.” USENIX NSDI.

[15] Amy Ousterhout, Joshua Fried, Jonathan Behrens, Adam Belay, and Hari Balakrishnan. 2019. “Shenango: Achieving High CPU Efficiency for Latency-sensitive Datacenter Workloads.” USENIX NSDI.

[16] Joshua Fried, Zhenyuan Ruan, Amy Ousterhout, and Adam Belay. 2020. “Caladan: Mitigating Interference at Microsecond Timescales.” USENIX OSDI.

[17] Gerwin Klein et al. 2009. “seL4: Formal Verification of an OS Kernel.” ACM Symposium on Operating Systems Principles. Current seL4 documentation reports machine-checked implementation-level properties and capability-based isolation.

[18] seL4 Project. Capabilities Tutorial. Defines capabilities as unforgeable authority-bearing references to kernel resources.

[19] seL4 Project. Frequently Asked Questions. Documents seL4 resource management and its no-kernel-heap design.

[20] University of Cambridge CTSRD. CHERI Architectural Rules for Capability Use. Describes capability provenance, authority restriction, and architectural rules governing derived capabilities.

[21] Simon Josefsson and Ilari Liusvaara. 2017. Edwards-Curve Digital Signature Algorithm (EdDSA). RFC 8032.

[22] Ben Laurie, Eran Messeri, and Rob Stradling. 2021. Certificate Transparency Version 2.0. RFC 9162. Defines authenticated Merkle-tree inclusion and consistency proofs.

[23] Jack O’Connor, Jean-Philippe Aumasson, Samuel Neves, and Zooko Wilcox-O’Hearn. BLAKE3. Official implementation and specification materials include portable reference and SIMD-optimized implementations.

[24] Leslie Lamport. 1977. “On Concurrent Reading and Writing.” Communications of the ACM 20. Foundational shared-memory register reasoning.

[25] Leslie Lamport. 1990. “The Concurrent Reading and Writing of Clocks.” ACM Transactions on Computer Systems.

[26] Leslie Lamport. 1974. “A New Solution of Dijkstra’s Concurrent Programming Problem.” Communications of the ACM 17.

[27] Diego Ongaro and John Ousterhout. 2014. “In Search of an Understandable Consensus Algorithm.” USENIX Annual Technical Conference.

[28] Claude E. Shannon. 1948. “A Mathematical Theory of Communication.” Bell System Technical Journal 27, 379–423 and 623–656.

[29] Rolf Landauer. 1961. “Irreversibility and Heat Generation in the Computing Process.” IBM Journal of Research and Development 5(3), 183–191. DOI 10.1147/rd.53.0183.

[30] Charles H. Bennett. 1973. “Logical Reversibility of Computation.” IBM Journal of Research and Development 17.

[31] Kutluhan Erol, James Hendler, and Dana Nau. 1994. Work on the computational complexity and formal structure of hierarchical task-network planning.

[32] Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. “ReAct: Synergizing Reasoning and Acting in Language Models.” ICLR.

[33] Vijay R. Konda and John N. Tsitsiklis. 1999. “Actor-Critic Algorithms.” Advances in Neural Information Processing Systems 12.

[34] Vijay R. Konda and John N. Tsitsiklis. 2003. “On Actor-Critic Algorithms.” SIAM Journal on Control and Optimization.

[35] Jeffrey O. Kephart and David M. Chess. 2003. “The Vision of Autonomic Computing.” Computer 36(1), 41–50.

[36] David F. Bacon, Perry Cheng, and V. T. Rajan. 2003. “Controlling Fragmentation and Space Consumption in the Metronome, a Real-Time Garbage Collector for Java.” LCTES.

[37] Linux Kernel Documentation. Soft-Dirty PTEs. Defines PTE-based virtual-page write tracking and the clear_refs / pagemap mechanism.

[38] Apache Arrow Project. Buffer Alignment and Padding Guidance. Recommends aligned buffers, including 64-byte alignment where appropriate, while treating it as an optimization/layout recommendation rather than a universal processor invariant.

[39] Apple. 2026. M5 Pro and M5 Max. Apple documents an 18-core CPU, up-to-40-core GPU, up to 128 GB unified memory, and up to 614 GB/s memory bandwidth for M5 Max.

[40] Microsoft. Xbox Series X Technical Specifications. Documents the custom eight-core Zen 2 CPU, RDNA 2 GPU, and 16 GB GDDR6 configuration.

[41] GPU vendor. Jetson AGX Orin Technical Specifications. Documents the Cortex-A78AE CPU configuration and current Orin performance/power profiles.

[42] AMD. 5th Generation EPYC 9005 Series. Documents Zen 5/Zen 5c server processors, 12-channel DDR5 and current platform characteristics.

[43] Qualcomm. 2025. Snapdragon 8 Elite for Galaxy and the Samsung Galaxy S25 Series. Qualcomm identifies Snapdragon 8 Elite for Galaxy as the S25-series platform.

[44] Motorola. Moto G 2025 Specifications. Motorola identifies the device’s processor as MediaTek Dimensity 6300 with Arm Mali-G57 MC2, correcting the supplied benchmark corpus’s Mali-G51 metadata.

[45] Y. C. Fung. 1993. Biomechanics: Mechanical Properties of Living Tissues, 2nd edition. Springer. Foundational continuum and constitutive modeling reference relevant to future biomechanical Magma workloads.

[46] Armen P. Sarvazyan, Oleg V. Rudenko, Scott D. Swanson, John B. Fowlkes, and Stanislav Y. Emelianov. 1998. “Shear Wave Elasticity Imaging: A New Ultrasonic Technology of Medical Diagnostics.” Ultrasound in Medicine & Biology 24(9), 1419–1435.

[47] Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 2024. “4D Gaussian Splatting for Real-Time Dynamic Scene Rendering.” IEEE/CVF CVPR.

[48] Brennan DeCrow. 2026. ATESO / Magma Master Research Directive and M-Tier Architectural Elevation. manymoats Research / ATESO Labs. Source of the ATESO terminology, six-step lifecycle, .many / .moat / .spine proposed contracts, fleet measurement record, and open-specification release intent.

[49] DeepSeek-AI. 2026. “DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression.” arXiv:2609.19969, submitted September 17, 2026. Reports a global KV-cache footprint of 890 bytes per token, roughly one quarter of prior long-horizon serving stack, and a persistent KV-cache footprint of roughly one eighth of prior long-horizon serving stack.

[50] Igor Sheludko and Santiago Aboy Solanes. 2020. “Pointer Compression in V8.” V8 project technical blog, 30 March 2020. Documents compressed tagged values and reports V8 heap-size reductions of up to 43%.

[51] Camillo Bruni. 2017. “Fast Properties in V8.” V8 project technical blog, 30 August 2017. Documents HiddenClasses, in-object versus dictionary-mode properties, and separate elements and properties stores.

[52] Google Cloud. 2026. Create and Manage Notebooks (API). Gemini Notebook Enterprise documentation. Documents notebooks.create, notebooks.get, notebooks.listRecentlyViewed, notebooks.batchDelete, and notebooks.share under Pre-GA (Preview) terms.

[53] Google Cloud. 2026. Add and Manage Data Sources in a Notebook (API). Gemini Notebook Enterprise documentation. Documents notebooks.sources.batchCreate, notebooks.sources.uploadFile, notebooks.sources.get, and notebooks.sources.batchDelete; supported sources include raw text, Markdown, Google Docs and Slides, web URLs, PDF/DOCX/PPTX/XLSX, audio, video, and images. Pre-GA (Preview).

[54] Google Cloud. 2026. Manage Audio Overview of Your Notebook (API). Gemini Notebook Enterprise documentation. Documents notebooks.audioOverviews.create under Pre-GA (Preview) terms.

[55] Google. 2026. Configure the Drive MCP Server and MCP Reference: drivemcp.googleapis.com. Google Workspace developer documentation. Remote endpoint https://drivemcp.googleapis.com/mcp/v1, OAuth 2.0 authentication, available through the Google Workspace Developer Preview Program.

[56] Eugene Lo. 2025. “Video Overviews on NotebookLM get a major upgrade with Nano Banana.” Google Keyword blog (Google Labs), 13 October 2025. Establishes NotebookLM Video Overviews (Explainer and Brief formats) as a product capability; the post makes no mention of a programmatic video-generation API.

[57] Apple. 2026. MacBook Pro — Technical Specifications. Apple, as consulted 2026-09-21. Lists the M5 Max with 18-core CPU and 32-core GPU at 36 GB unified memory and 460 GB/s memory bandwidth, and the M5 Max with 18-core CPU and 40-core GPU at 48 GB, 64 GB, or 128 GB unified memory and 614 GB/s memory bandwidth; memory sizes are stated in GB without a binary/decimal qualifier.

[58] Apple. 2021. MacBook Pro (14-inch, 2021) — Technical Specifications. Apple Support. Lists 16 GB unified memory as the M1 Pro standard configuration, configurable to 32 GB.

[59] Google Cloud. 2026. What is Gemini Notebook Enterprise? Gemini Notebook Enterprise documentation (overview). States that the product does not change the original file and that the model relies on the static copy of the uploaded sources.

[60] Federal Reserve Financial Services. FedNow Service Readiness Guide — Spotlight on ISO 20022: Messages Overview. explore.fednow.org. Lists pacs.008 (customer credit transfer), pacs.004, pacs.009, pacs.002, pacs.028, pain.013, and camt messages as the FedNow Service message set; the message version set is defined in the FedNow ISO 20022 message specifications published on the MyStandards platform.