ATESO LABS // EMPIRICAL BENCHMARKS
Papers Archive Claim Card 01 Falsification Ledger Home
Public Reproducibility Protocol · v1.4

Empirical Benchmarks & Replication Guide

Turning "formula, not measurement" into repeatable, verifiable commands on live silicon.
THE HONESTY CHARTER (THE FOUNDER’S RULE):
"The 4.14× figure is a formula, not a memory-chip measurement. The 65.9 µs line is a sum, not a measured round trip. Nothing on this page is a meter reading from a live chip, a water plant, or a grid."

This document separates mathematical derivations from hardware measurements. Below are the exact numbers captured on Apple Silicon and Chrome runtimes, the known measurement caveats, and the shell commands required to reproduce every result independently.

1. Summary of Measured vs. Derived Values

Formula Compaction
4.1445×
Closed-form equation $W = \beta / ((1 - \alpha)(1 + f))$ against V8 expansion $\beta = 3.50$.
Memory Bus Stride
11.1 µs
2,500 Float64Array writes on M5 Max unified memory vs 33.0 ms full JSON.stringify/parse.
Chrome Binary Delta
0.127 ms
Headless Chrome 50k entity delta vs 9.327 ms full JSON (73.6×). 2.16× vs JSON-delta.
Ring Buffer Throughput
553,274 TPS
Single-threaded SPSC ring buffer replay. 600,000 records pass with 0 dropped frames.

2. Detailed Suite Results & Caveats

Benchmark Suite What Was Measured Observed Result Scientific Scope & Caveat
Suite 1: Working Set Compaction V8 managed heap memory bloat vs packed binary slab. 4.14× Model Formula-derived from $\beta=3.50$, $\alpha=0.25$, $f=0.126$. Not a direct physical DRAM meter reading. Opaque blobs yield $W \approx 1.0\times$.
Suite 2: Unified Memory Stride 2,500 record updates via Float64Array vs JSON rebuild. 11.1 µs / frame Compares in-memory pointer stride to full stringify/parse. Stride does not perform cryptographic signature verification on every tick.
Suite 3: Headless Chrome Delta Chrome WebGPU writeBuffer vs full JSON transfer. 0.127 ms vs 9.327 ms Contiguous 160 KiB buffer write at offset 0. performance.now() timer floor is quantized at ~0.1 ms in hardened browsers.
Suite 4: XPBD WebGPU Physics Extended Position Based Dynamics shader round-trip. 2,635 Hz queue dispatch Shader clamped at 32 particles / 96 edges (MAX_P=32, MAX_E=96). Measures queue roundtrip latency, not massive N-body stress.
Suite 5: SPSC Ring Correctness 600,000 sequential inter-thread events. 31 / 31 tests pass Proves lock-free single-producer single-consumer correctness and partition recovery under simulated link drops.

3. How to Measure It Yourself (Replication Guide)

Every test is runnable directly from the repository using standard Node.js (v20+) on any POSIX or Apple Silicon workstation:

A. Run the Resident State Memory Benchmark

Measures memory access latency across 50,000 entities over 200 iterations:

# From the repo root
node ateso-resident-state-benchmark.mjs

# Expected output:
# [JSON baseline]   50,000 objects stringify+parse: ~33.0 ms
# [Resident slab]   5,000 Float64Array writes:      ~11.1 µs
# Ratio:            2,900× – 3,500× (unmerged, machine-dependent)

B. Run the 31-Test Maglev Spine Partition Suite

Tests socket timeouts, breaker stamps, partition recovery, and lock-free concurrency:

# Runs 31 unit and partition regression tests
node --test scripts/maglev-cluster-boot.test.mjs

# Expected output:
# ℹ tests 31
# ℹ pass 31
# ℹ fail 0
# ℹ duration_ms ~53,000ms

C. Verify the Working Set Multiplier Formula

Runs the closed-form equation tests across arbitrary heap payload sizes:

# Test the arena compaction model
node engines/arena-multiplier/test.mjs

# Validates:
# W = 4.1445 under direct mode (gamma = 1.0)
# W = 2.195 under fallback siege mode (gamma = 2.0)

4. Falsification Criteria

Our research claims are defined to be strictly falsifiable under Popperian criteria:

5. Provenance & Verification Stamps

Repository: manymoats-kpf
Verification Suite: scripts/maglev-cluster-boot.test.mjs (31/31 Pass)
Author & Architect: Brennan William DeCrow
Date: September 2026
License: MIT / Dual Commercial Enterprise License