๐Ÿ“„ Document ID: CP-ENG-WP-2026, revision 2 ๐Ÿ“ฆ Release: v0.10.0 ๐Ÿ›๏ธ Technical Whitepaper ๐Ÿ“… Revised September 2026

Virtual File System (VFS) Architecture Whitepaper

Zero-copy asset streaming from memory-mapped archives, with a Rust core, whole-archive sealing and embedded Lua 5.4

Target Audience:
Technical directors, engine and tools programmers, and build or pipeline engineers evaluating an asset file system for a game or a real-time application

1. Executive Summary (BLUF)

Bottom Line Up Front (BLUF)

Games and real-time applications read thousands of assets per level. A loader that opens and reads each file on its own pays a system call and a copy for every asset, and an archive shipped without authentication can be read, or edited, by anyone who has the file.

CyclePack VFS is a Rust library with a C ABI that addresses both, and says where it stops. It maps Format v3 archives and hands out a stored entry as a borrow of the mapping, with no copy and no lock on the read path; compressed and sealed entries are decoded, checked against their BLAKE3 digests and returned as a copy. A whole archive can be sealed with AES-256-GCM or ChaCha20-Poly1305. An embedded Lua 5.4 runtime scripts packing and pipeline steps; it is trusted by default, with an opt-in sandbox for scripts the host does not control. In figures:

What is hardware here, and what is a software stand-in

This paper names CXL.mem, NVMe ZNS, NVMe computational storage (TP 4091/4092), GPUDirect Storage, RoCE v2 RDMA and DirectStorage. In the shipping build each of them is served by a software stand-in; the one exception is CXL memory on a Linux host whose kernel serves a DAX device that can hold a region, which the library maps directly (never yet run on a DAX device). The library links no DirectStorage, CUDA, cuFile or verbs library. The hardware interfaces it drives are Linux io_uring, which has run on one kernel (Linux 6.12, aarch64, in a local container on an Apple-silicon Mac; not on bare metal, x86_64 or NVMe), and, for CXL, a DAX character device.

The formats, queue models and state machines are real, implemented and tested, which is what lets you write an application that is correct before the hardware arrives. What is not delivered today is anything that comes from the hardware itself: keys that never enter host RAM, transfers that never touch host memory, coherency across hosts, GPU-side decompression. A diagram below that shows a GPU, an SSD controller or a CXL switch says that it draws the target, not this build's data path.

Ask the binary rather than the brochure: cyclepack doctor prints a [Capability Backing] table for the machine you run it on, and cyclepack_query_capability_backing(CYCLEPACK_CAP_*) returns the same verdict to C callers: 2 Native, 1 Emulated, 0 Unavailable. CXL reports Native only for a DAX character device that Linux sysfs confirms (on the dax bus with device_dax bound, or in the dax class): the one CYCLEPACK_CXL_DAX_PATH names, or one listed under /sys/bus/dax or /sys/class/dax, whose alignment and size sysfs states, whose size holds a one-slot region, and that this process can open read-write and map in whole units of that alignment. Any other character device, /dev/null included, reads Emulated with the reason. The DAX link check has run on a Linux 6.12 kernel in an emulated x86_64 container against simulated DAX entries, the alignment and size reads only against fixtures, and none of it on a real DAX device. ZNS reports Emulated on every host in this build, because it has no zone-ioctl backend; on Linux, CYCLEPACK_ZNS_DEVICE only changes what the ZNS row says about that device.

0 copies
Stored Entries, Read in Place
A stored entry comes back as a borrow of the mapped archive: no read() per access once the archive is mapped, and no lock on the read path. Compressed or sealed entries are decoded into a copy. Source: src/vfs/mmap.rs, src/vfs/engine.rs.
166 exports
C ABI, Every Export Panic-Guarded
None without a panic guard, so a Rust panic becomes an error code instead of a dead host. Source: python3 scripts/generate_abi_manifest.py --check, run for this revision.
7 of 14
Capabilities Native, and the Rest Say So
cyclepack doctor on the recorded Apple-silicon Mac: 7 capabilities native, 7 emulated, 0 unavailable. For the seven software formats and algorithms, native means code in this library, not hardware. Run it on your machine.
80.2%
Line Coverage
Up from 71.5% when coverage was first measured, on 2026-09-20; scripts/coverage.sh enforces a 77% floor. Measured on 2026-09-24 and not re-run for this paper.

2. Introduction: The Asset Loading Problem

2.1 What Per-File Loading Costs

Most engines start with a loader that opens and reads one file per asset, and a packed archive whose format anyone can parse. Four costs follow, and each has an answer in CyclePack:

1
A system call and a copy per asset: each read() enters the kernel and copies from the page cache into a user buffer. A mapped archive lets the kernel fault pages in on first touch, and CyclePack hands out a stored entry where it lies.
2
Contention on a shared archive: a loader that serialises reads behind one lock stalls every thread that wants an asset. CyclePack's reads take &self and no lock; mounting takes &mut self.
3
Rebuilds for pipeline changes: when packing, validation or patch rules are compiled into the engine, changing them means a rebuild. CyclePack runs those steps as Lua 5.4 scripts, which a host can load again without one.
4
Archives anyone can read or edit: a plain archive gives up its paths and contents to whoever has the file. A sealed CyclePack archive hides paths, entry sizes and digests without the key, and a reader holding the key refuses an edited header or manifest at open, and an edited entry when it reads it. It is not DRM: a key shipped inside a game can be extracted from it.
Figure 2.1: Per-File Loader vs. CyclePack Mapped Archive
graph LR subgraph Legacy_Pipeline ["Per-file loader"] direction TB App1["Engine"] -->|"open() and read() per asset"| KernelVFS["Kernel VFS"] KernelVFS -->|"copy out of the page cache"| BufferCopy["User buffer, a copy"] end subgraph CyclePack_Pipeline ["CyclePack mapped archive"] direction TB App2["Engine"] -->|"read_slice(path): table lookup, no read() call"| CPKVfs["VfsEngine (Rust)"] CPKVfs -->|"stored entry: a borrow, no copy"| MapView["mmap of the .cpk archive"] MapView -.->|"first touch: page fault, the kernel reads the page"| BlockDev["Storage device"] end

3. Core Architecture: Rust & Lua

CyclePack keeps apart the part that must be fast and memory-safe, the Rust core, and the part that changes often, the pipeline scripts in Lua.

Figure 3.1: CyclePack Rust Core, I/O Backends & Embedded Lua Runtime
graph TD subgraph Host_Layer ["Host: C ABI, Rust, Node / Deno / Bun, WASM, Python"] HostApp["Engine, tool or service"] end subgraph Rust_VFS_Core ["Rust core (v0.10.0)"] VfsRouter["VfsEngine: path sanitiser, mounts tried by priority"] MmapStore["MmapStorage: Format v3 archives"] EntryCodec["Entry codec: AEAD open, LZ4 decode, BLAKE3 check"] LearnedMount["Learned-index mount: model and Bloom filter"] OverlayEng["LayeredOverlayEngine: copy-on-write overlays"] Formats["Formats: .gspx, .cdag, .svt, .cntc"] Scheduler["StreamingScheduler: priority bands, coalescing, budgets"] Backends["I/O backends: io_uring on Linux, mmap and pread elsewhere"] HwModels["Emulated models: DirectStorage, ZNS, CSD, GPUDirect, RDMA, CXL without DAX"] end subgraph Lua_Ecosystem ["Embedded Lua 5.4"] LuaVM["LuaRuntime::new (trusted) or LuaRuntime::sandboxed"] LuaBindings["Bindings: archive, compression, crypto, vfs, telemetry and more"] end HostApp -->|"C ABI: 166 exports, each panic-guarded"| VfsRouter VfsRouter --> MmapStore VfsRouter --> LearnedMount VfsRouter --> OverlayEng MmapStore --> EntryCodec VfsRouter --> Formats Scheduler --> Backends Scheduler --> HwModels LuaVM --> LuaBindings LuaBindings --> VfsRouter
๐Ÿฆ€ Rust Core

Memory Safety and a Lock-Free Read Path

The core is Rust. Other languages reach it through the C ABI, and every export runs behind a panic guard:

  • Zero-copy reads: a stored entry is a borrow of the mapping. Reads take &self and take no lock; mounting takes &mut self. VfsEngine and MmapStorage are Send + Sync.
  • 64 KiB where a GPU would want it: .gspx chunk payloads start on 64 KiB boundaries, SVT tiles are 64 KiB, and the DirectStorage-shaped container uses independent 64 KiB blocks, so a host can hand an extent onward without repacking it. Format v3 archive entries are not padded to 64 KiB. The hop into GPU memory belongs to the host and to hardware this library does not link.
  • In-Storage Compute (CSD) command model: The NVMe 2.0 TP 4091/4092 program model — codec, cipher and key slot attached to a read — is implemented, so an application issues the commands a device would accept. Today those programs execute on the host CPU: no controller is in the path and the key slots are host memory. The simulator's decrypt stage is a repeating-key XOR whatever cipher the command names, so it protects nothing; archive sealing is what protects content.
๐ŸŒ™ Embedded Lua 5.4

Pipeline Changes Without a Rebuild

The pipeline steps that change from project to project live in Lua scripts rather than in the host:

  • Scripted pipelines: packing, sealing, compression and patch steps are Lua scripts, run by cyclepack run or by a LuaRuntime your host embeds.
  • No rebuild to change them: a host loads a new script into its runtime (LuaRuntime::load_script) without recompiling. The library fetches no script over the network; where scripts come from is the host's decision.
  • Trusted by Default, Sandboxed on Demand: LuaRuntime::new(), which the cyclepack CLI uses for your own pack scripts, is a trusted runtime: scripts get the full standard library (os.execute, os.getenv, require, ...) and unrestricted host file access, like any program the user runs. Scripts the host does not control (mods, downloaded pipelines) belong in LuaRuntime::sandboxed(LuaSandbox { roots }), which removes os.execute, os.exit, os.remove, os.rename, os.tmpname, os.getenv, os.setlocale, dofile, loadfile, require and package, only loads text chunks, drops the telemetry mutators, and confines every CyclePack binding that touches host files to the configured roots. Paths are checked lexically and again after resolving symlinks; dangling symlinks are refused, and memory-mapped files cannot be written.
  • Caps: a Lua buffer and a file read through the Lua I/O binding are each capped at 512 MB, and decompression in Lua at 256 MiB.

3.3 Learned Index & Bloom Filter

A learned-index mount replaces the search structure over the entry table with a model. Paths are hashed with 64-bit FNV-1a and sorted by hash. A piecewise-linear model over the sorted hashes predicts where an entry sits, within a bounded error ε (clamped to 2–64): a lookup binary-searches the segments, scans the ±2ε entries around the prediction, and falls back to a binary search over the whole table when the window misses. Before any of that, a Bloom filter with three hash functions, sized from 4 bits per key and rounded up to a power of two, rejects most missing paths. (The code calls it RibbonFilter; it is a Bloom filter.) The index still keeps every entry — hash, path, offset, size — in memory: the model replaces the B-tree or hash-map search, not the table. No lookup latency has been measured for this paper.

Figure 3.2: Learned-Index Lookup with a Bloom Filter in Front
graph LR PathInput["Asset path: 'textures/hero.png'"] --> Hash64["64-bit FNV-1a hash"] Hash64 --> BloomFilter["Bloom filter: 3 hashes, 4 or more bits per key"] BloomFilter -->|"definitely absent"| MissOut["Not found, no search"] BloomFilter -->|"maybe present"| SegSearch["Binary search over segment start keys"] SegSearch --> PLR["Segment model: slope and intercept"] PLR --> EstPos["Predicted position"] EstPos --> Scan["Scan 2 epsilon either side"] Scan -->|"found"| ExactEntry["Entry: offset and size"] Scan -->|"window missed"| Fallback["Binary search over the whole table"] Fallback --> ExactEntry

3.4 CXL Memory: What This Build Maps

The pool API is written for CXL memory, where several hosts share one coherent pool. This build has no fabric. Entities live in one of three places, and the engine reports which (CxlTier): in this process's heap (CxlSharedMemoryPool, a versioned compare-and-swap under a read-write lock); in a MAP_SHARED region over a file that other processes on the same host map too (MappedEntityStore: 128-byte slots, each guarded by a seqlock whose version moves by compare_exchange); or, on Linux, in that same region over a DAX character device that sysfs confirms, mapped at offset 0 in whole units of the device's alignment and never past its size, the only tier cyclepack doctor reports as Native. That sysfs check has run on a Linux 6.12 kernel in an emulated x86_64 container against simulated DAX entries, the alignment and size reads only against fixtures, and none of it on a real DAX device. No tier adds coherency across hosts, and nothing in this build issues clwb or sfence or makes a write durable on persistent memory:

Figure 3.3: Where CXL Entities Live in This Build
graph TD subgraph Host_Compute_Nodes ["One host"] Host1["Process A: MappedEntityStore"] Host2["Process B: MappedEntityStore"] Host3["Process C: CxlSharedMemoryPool, heap, this process only"] end subgraph Region_Layout ["Mapped region, layout v1"] RegionHeader["4 KiB header: magic CXLR, CRC-32, epoch"] RegionIndex["Index: one AtomicU64 per slot"] RegionSlots["128-byte slots, each with a seqlock version"] RegionHeader --> RegionIndex RegionIndex --> RegionSlots end subgraph Region_Backing ["What backs the region"] ShmFile["File under /dev/shm or any path: Emulated"] DaxDevice["DAX device confirmed by Linux sysfs: Native"] end Host1 -->|"compare_exchange on the slot version"| RegionSlots Host2 -->|"compare_exchange on the slot version"| RegionSlots RegionSlots --> ShmFile RegionSlots --> DaxDevice DaxDevice -.->|"target, not in this build"| CxlSwitch["CXL switch sharing one pool across hosts"]

3.5 Scheduling & Backend Routing

Requests go through the StreamingScheduler: five priority bands (Realtime, High, Normal, Low, Background) with starvation aging, deadline-aware ordering, coalescing of adjacent or overlapping byte ranges on the same asset, cooperative cancellation tokens, and backpressure budgets on in-flight bytes, in-flight requests and queue depth. A request names the memory tier it is meant for (host RAM, staging RAM, GPU memory, the CXL pool); this build has no GPU residency path, so an emulated backend may not serve a GPU-destined request unless RuntimeConfig::allow_emulated_backends is set (it is off by default), and with it set the bytes land in host memory. A cost model picks the backend per request from EWMA latency, queue depth, deadline risk, read amplification and backend health, with a hysteresis margin against flapping and a fallback chain when a backend fails. No figure in this section is a measurement:

Figure 3.4: Streaming Scheduler & Cost-Model Backend Routing
graph LR AssetReq["AssetRequest: path, range, priority, deadline, target tier"] --> Bands["Priority bands: Realtime, High, Normal, Low, Background"] Bands --> Coalesce["Coalesce adjacent or overlapping ranges"] Coalesce --> Budget["Backpressure: in-flight bytes, requests, queue depth"] Budget --> CostModel["Cost model: EWMA latency, queue, deadline risk, health"] CostModel --> UringBackend["io_uring, Linux"] CostModel --> MmapBackend["mmap and pread"] CostModel -->|"only with allow_emulated_backends"| EmuBackend["Emulated DirectStorage and GPUDirect"]

3.6 Semantic Bitpacking Codecs

semantic_bitpack.rs holds small codecs for common game data: 2D and 3D Morton (Z-order) swizzling, a 6-byte drop-W quaternion, 2-byte octahedral normals, a 16-byte quantized transform (translation, rotation, scale), and an f16 delta codec for vertex positions. The quaternion, normal, transform and f16 codecs are lossy by design. They are building blocks for your own formats; nothing here migrates a schema from one layout to another:

Figure 3.5: Semantic Bitpacking Codecs and Their Output Sizes
graph LR QuatIn["Quaternion: 4 x f32, 16 bytes"] -->|"DropWQuaternion::pack"| QuatOut["6 bytes"] NormalIn["Unit normal: 3 x f32, 12 bytes"] -->|"OctahedralNormalCodec::encode"| NormalOut["2 bytes"] TrsIn["Translation, rotation, scale: 9 x f32, 36 bytes"] -->|"TransformBitpacker::pack"| TrsOut["16 bytes"] CoordIn["x, y, z: 3 x u16"] -->|"MortonSwizzler::encode_3d"| CoordOut["u64 Morton code"]

3.7 Generation Publishing & Rollback

Live updates use generations, in the manner of RCU. A Generation is an immutable snapshot of the asset tree. A reader takes a short read lock to clone an Arc<Generation> and then reads without holding it. A publisher prepares the next generation (hashing, verification, allocation) outside the lock, then swaps the active pointer, and refuses if the active generation changed after it staged. A retired generation lives until its last reader drops it, and the last eight stay in a history ring for rollback. This is reference counting, not epoch-based reclamation, and no swap time is quoted:

Figure 3.6: Generation Publishing, Reader Lifetimes & Rollback
graph TD Stage["prepare_patch: build and verify generation N+1 outside the lock"] --> Publish["publish_staged: swap the active pointer if it is still N"] Publish --> ActiveGen["Active generation: N+1"] Publish --> HistoryRing["History ring: the last 8 generations"] Reader1["Reader: acquire_active clones the Arc"] --> ActiveGen Reader2["Reader still holding generation N"] --> OldGen["Generation N lives until its last Arc drops"] HistoryRing -->|"rollback"| ActiveGen

3.8 Schema Compiler (cyclepack codegen)

Rather than maintain the same struct by hand in several languages, cyclepack codegen reads a .cpack schema and emits packed, layout-identical structs with their serialisers: C++ (#pragma pack(push, 1) and memcpy), C# (StructLayout Sequential, Pack = 1), Python (ctypes.LittleEndianStructure), Go (encoding/binary, little-endian), Rust (#[repr(C, packed)], read with read_unaligned) and Lua 5.4 (string.pack / string.unpack). The test suite compiles the generated code and round-trips it through every one of those toolchains present on the machine. There is no Zig emitter, and the generated code carries no static_assert on struct sizes:

Figure 3.7: One Schema, Six Languages
graph TD SchemaDef["Schema: entity.cpack"] --> CodegenEngine["cyclepack codegen (src/codegen/mod.rs)"] CodegenEngine --> CPPOut["C++: pragma pack(push, 1) struct"] CodegenEngine --> CSharpOut["C#: StructLayout Sequential, Pack = 1"] CodegenEngine --> PythonOut["Python: ctypes.LittleEndianStructure"] CodegenEngine --> GoOut["Go: encoding/binary, little-endian"] CodegenEngine --> RustOut["Rust: repr(C, packed), read_unaligned"] CodegenEngine --> LuaOut["Lua 5.4: string.pack and string.unpack"]

3.9 Hardware Probe & Fallbacks

HardwareProber::probe runs at initialisation. The CPU probes are real — is_x86_feature_detected! / is_aarch64_feature_detected! for AVX-512, AVX2, AES and NEON — but nothing in the crate branches on the result to select a SIMD path: the only readers of has_avx2, has_avx512, has_aes_ni and has_neon are the cyclepack doctor report and the Lua hardware table. The aes and blake3 crates do their own dispatch internally, and on AArch64 the aes crate’s ARMv8 backend sits behind the aes_armv8 cfg, which this build does not set, so CpuCryptoTier::backing() reports Emulated there even when the ARMv8 AES extension is detected. The CXL probe maps a DAX device only when sysfs confirms one, and otherwise uses a shared mapping or the in-process pool (section 3.4). The computational-storage probe always answers that no device is present. io_uring serves when the kernel sets up a ring for the process (run on one kernel, Linux 6.12 aarch64 in a local container), and the mmap and pread backends otherwise. To time the in-process CAS path on your host, run cargo bench --bench vfs_bench -- cxl_pool_versioned_cas.

Figure 3.8: Hardware Probe & the Fallback Each Branch Takes
graph TD Init["Start-up: HardwareProber::probe()"] --> CpuFeat["CPU features: AVX2, AVX-512, AES, NEON, reported only"] Init --> CheckCXL{"DAX device confirmed by Linux sysfs?"} CheckCXL -->|"yes"| DaxTier["CxlTier::DaxDevice: map it, Native"] CheckCXL -->|"no, region path set"| SharedTier["CxlTier::SharedMapping: MAP_SHARED file, Emulated"] CheckCXL -->|"no path"| InProcTier["CxlTier::InProcess: heap pool, Emulated"] Init --> CheckCSD["Computational storage: no device detected, programs run on the host CPU"] Init --> CheckUring{"Kernel sets up an io_uring ring?"} CheckUring -->|"yes, Linux"| UringTier["io_uring backend"] CheckUring -->|"no"| MmapTier["mmap and pread backends"]

3.10 Asynchronous Worker Pool (AsyncVfsWorkerPool)

A synchronous read on an event-loop thread (Node.js, Python asyncio, a game's main thread) stalls it. AsyncVfsWorkerPool runs reads on its own worker threads, fed through a channel, and completes them through a Rust closure, a C-ABI callback or a polling token. The Python SDK's read_async awaits it without blocking asyncio; the Node SDK's reads are synchronous in this release. No throughput figure has been measured for this paper:

Figure 3.9: Worker Pool & Completion Paths
graph LR EventLoop["Event loop or game thread"] -->|"submit a read with a callback"| TaskQueue["Task channel"] TaskQueue --> Worker1["Worker 1: VfsEngine read"] TaskQueue --> WorkerN["Worker N: VfsEngine read"] Worker1 -->|"completion"| CallbackTrampoline["C-ABI callback, Rust closure or polling token"] WorkerN -->|"completion"| CallbackTrampoline CallbackTrampoline --> EventLoop

3.11 WebGPU Preview & the Embedded Reader

  • WebGPU preview (@cyclechain/cyclepack-webgpu): 16-byte PackedBaseSplat records are unpacked in a WGSL vertex shader, with no CPU decode and no depth sort yet. It was not run for this paper: it needs a browser with WebGPU.
  • Embedded reader (EmbeddedIndexReader<MAX_FILES>): parses a legacy index and its data from byte slices into a fixed-size array, with no heap allocation. It is part of the standard-library crate and has not been built for or run on a microcontroller in this release; its entry table takes 16 bytes per slot, times MAX_FILES.

3.12 In-Process Arena & CycleRPC

SharedMemoryArena lets Python write through a memoryview into the engine's slab, so nothing is copied onto the CPython heap. The arena lives in this process; it is not shared with other processes. CycleRPC frames carry a fixed 24-byte header (magic CRPC, version, flags, request id, method id, payload size) and dispatch to handlers in process; the transport is yours to bring. No throughput or latency figure is quoted for either.

4. Deep-Dive: The Read Path & Streaming Formats

When a host asks for an asset, this is what the code does. No GPU and no SSD controller takes part. A borrow is not hashed, so verify an archive once when you install it (cyclepack verify runs every entry through the copying path).

Figure 4.1: What a Read Does in This Release
sequenceDiagram autonumber actor Host as Host engine participant Vfs as VfsEngine participant Store as MmapStorage, Format v3 participant Codec as Entry codec Host->>Vfs: read_slice("assets/meshes/hero.mesh") Vfs->>Vfs: sanitise the path, try the mounts under its prefix by priority Vfs->>Store: look the path up in the entry table built at mount alt stored entry of an unsealed archive Store-->>Host: a borrow of the mapping, no copy and no hash else LZ4 entry, or any entry of a sealed archive Store-->>Host: refused with E_ZERO_COPY_COMPRESSED or E_ZERO_COPY_ENCRYPTED Host->>Vfs: read("assets/meshes/hero.mesh") Vfs->>Codec: open the AEAD seal and decode LZ4 Codec->>Codec: compare with the entry's BLAKE3 digest Codec-->>Host: a checked copy of the bytes end

4.2 Gaussian Splats (.gspx) & the Meshlet Cluster DAG (.cdag)

The .gspx container stores Gaussian splats as 16-byte PackedBaseSplat records in chunks, each described by a 64-byte descriptor carrying a Morton code and a bounding box, and every chunk payload starts on a 64 KiB boundary. Chunk bounds and frustum culling run on the reader, on the CPU. The meshlet container holds clusters of up to 128 triangles and 64 vertices; a cluster's error is measured vertex displacement, accumulated up the DAG, one cut can mix levels, the same mesh always gives the same bytes, and culling runs on the CPU. Rasterising either is the engine's job:

Figure 4.2: What CyclePack Does on the CPU, and What the Engine Does
graph TD subgraph Container_Storage ["Containers this library writes and reads"] GspxArchive[".gspx: 16-byte splats, chunks on 64 KiB boundaries"] MeshletArchive[".cdag: clusters of up to 128 triangles and 64 vertices"] SvtArchive[".svt: 64 KiB tiles"] end subgraph Spatial_Streaming_Engine ["CyclePack, on the CPU"] CameraFrustum["Camera position and frustum, from the host"] ChunkCulling["Chunk bounds and frustum culling"] MeshletDAG["DAG cut from the measured error"] SvtResidency["Tile residency and page tables"] end subgraph Hardware_Execution ["The host engine"] EngineUpload["GPU upload, binding and drawing"] end CameraFrustum --> ChunkCulling CameraFrustum --> MeshletDAG GspxArchive --> ChunkCulling MeshletArchive --> MeshletDAG SvtArchive --> SvtResidency ChunkCulling --> EngineUpload MeshletDAG --> EngineUpload SvtResidency --> EngineUpload

4.3 Neural Texture Container (.cntc)

The .cntc container stores a texture as multi-resolution latent grids, quantized to INT8, plus the weights of a small MLP decoder, and can decode any UV on its own. Decoding runs on the CPU in this build. The bundled encoder downsamples the texture into the grids and trains no decoder, so its output is not what a trained neural codec would give, and no compression ratio is claimed:

Figure 4.3: The .cntc Container & Its CPU Decoder
graph LR subgraph VFS_Storage [".cntc container"] LatentGrid["Latent grids at several resolutions, INT8"] QuantWeights["Small MLP decoder weights"] end subgraph CPU_Decode ["Decode on the CPU, this build"] TexCoord["UV query"] --> SampleGrid["Sample and concatenate the latents"] SampleGrid --> MLPDecode["MLP forward pass"] MLPDecode --> FinalChannels["Texel channels"] end LatentGrid --> SampleGrid QuantWeights --> MLPDecode BundledEncoder["Bundled encoder: downsamples into the grids, trains no decoder"] -.-> LatentGrid

4.4 Zone-Graph & Markov Prefetching

MarkovPrefetchEngine keeps a graph of zones and their neighbours, and Markov transition counts learned from observed moves: second-order (where the player went next, given the previous and the current zone) weighted 0.7, and first-order weighted 0.3. Given a position and a velocity, it moves each candidate's score by 0.3 times the cosine of the angle between the velocity and the direction to that zone (up for zones ahead, down for zones behind), drops candidates below a confidence threshold, and hands their asset lists to the scheduler as speculative, cancellable requests. How often it guesses right depends on the game and has not been measured here; the telemetry counts prefetch hits and misses so you can measure it on yours:

Figure 4.4: Zone-Graph & Markov Prefetch Flow
graph LR PlayerTelemetry["Host: current zone, position, velocity"] --> MarkovTable["Markov counts: 2nd-order weighted 0.7, 1st-order 0.3"] ZoneGraph["Zone graph: neighbours"] --> MarkovTable MarkovTable --> VelocityWeight["Velocity weighting: plus 0.3 x cosine of the angle to the zone"] VelocityWeight --> Threshold["Keep zones above the confidence threshold"] Threshold --> PriorityQueue["Speculative, cancellable requests to the StreamingScheduler"] PriorityQueue --> PrefetchTelemetry["Telemetry: prefetch hits and misses"]

4.5 GPU Direct Decompression (DirectStorage & Console ASICs) — Planned

On a native DirectStorage or console stack, archive decompression moves off the host CPU into GPU compute queues (GDeflate) or dedicated silicon. CyclePack does not do this today. Its DirectStorage path writes a GDeflate-shaped container — GDEF magic, independent 64 KiB blocks, the layout a GPU decompression queue expects — but the bitstream inside is chunked LZ4, decompressed in parallel on the CPU. A GPU GDeflate decoder cannot read it, and a stream produced by one cannot be read here. The container shape is what makes the swap cheap: a native backend would replace the codec, not the format. A Windows DirectStorage backend is planned and needs Windows CI; no console is supported:

Figure 4.5: The DirectStorage Path Today, and the Planned One
graph LR subgraph This_Build ["This build"] direction TB Chunk1["GDEF container: independent 64 KiB blocks"] --> CpuLz4["Chunked LZ4, decoded in parallel on the CPU"] CpuLz4 --> HostMem["Host memory"] end subgraph Planned_Backend ["Planned: native DirectStorage backend"] direction TB Chunk2["Same container shape"] --> GpuQueue["DirectStorage queue, GPU GDeflate decode"] GpuQueue --> GpuVram["GPU memory"] end

4.6 Sparse Virtual Texturing (SVT)

The SVT model divides a texture into 64 KiB tiles and keeps residency and page tables for them. A sampler-feedback residency model records which tiles a frame wanted, and a tile that has not streamed in yet falls back to the nearest resident parent mip. Reading feedback back from the GPU and binding tiles in GPU memory are the engine's job:

Figure 4.6: SVT Residency Loop, Split Between the Engine and CyclePack
graph TD subgraph GPU_Raster_Pass ["Host engine and GPU"] SamplerFeedback["Sampler feedback: tiles a frame wanted"] TileBind["Bind tiles in GPU memory"] end subgraph CyclePack_SVT_Engine ["CyclePack SVT model"] ResidencyTable["Residency and page tables, 64 KiB tiles"] ParentMip["Tile not streamed yet: nearest resident parent mip"] TileRequests["Tile read requests"] end SamplerFeedback --> ResidencyTable ResidencyTable --> ParentMip ResidencyTable --> TileRequests TileRequests --> TileBind

4.7 Prebuilt BVH Container (.cblas)

streaming_bvh.rs stores a prebuilt bounding volume hierarchy of 32-byte nodes (MessagePack, .cblas), traverses it on the CPU for ray queries, and keeps a list of top-level instances. It does not build the hierarchy and does not touch a GPU acceleration structure: building one and refitting it is the engine's job, with DXR or Vulkan ray tracing. Its savings_metrics estimate assumes 0.5 ms of GPU rebuild per 10,000 triangles; that is an assumption written in the code, not a measurement.

5. Security Architecture & FFI Boundary Hardening

A host that runs scripts it did not write needs them kept away from its files and its process. For scripts run in LuaRuntime::sandboxed (the trusted LuaRuntime::new() has the same host access as the user), CyclePack enforces three boundaries:

Figure 5.1: Sandboxed Script, Error Path & the C ABI Edge
graph TD subgraph Script_Invocation ["1. Sandboxed script"] HostExec["Host runs a script in LuaRuntime::sandboxed"] --> VerifyBytecode["Text chunks only, bytecode refused"] VerifyBytecode --> ConfinedVM["Host files confined to the sandbox roots, re-checked after symlinks"] ConfinedVM --> FfiCall["Script calls a Rust binding"] end subgraph Error_Path ["2. Errors stay errors"] LuaFailure["Syntax error, nil index, stack overflow or a cap exceeded"] --> ErrReturn["mlua returns Err to the Rust caller"] ErrReturn --> NextScript["The runtime stays usable: the next script runs"] end subgraph FFI_Edge ["3. C ABI edge"] PanicGuard["Every export runs behind a panic guard"] --> ErrorCode["A panic becomes an error code, never an unwind into the host"] end FfiCall -->|"normal completion"| SuccessReturn["Result returned to the host"] FfiCall -->|"failure"| LuaFailure

Rust Destructors Run After Lua GC

A test hands 1,000 Rust userdata values to Lua, drops the table that holds them, runs a full collection twice and asserts that the Rust destructors ran for all 1,000 (tests/vfs_advanced_symlink_and_sandbox_tests.rs).

Host Protection & Panic Safety

A syntax error, a nil dereference or a stack overflow in a script comes back to Rust as an Err, and the next script on the same runtime still runs (tests/chaos_fault_injection_tests.rs). On the C side all 166 exports run behind a panic guard, and the release profile keeps panic = "unwind" so the guard can catch.

Bounded Decompression & Paths

Size and count fields are bounded before they size an allocation. Decompression stops at 256 MiB in Lua and WASM (configurable in WASM) and at 4 GiB natively. Copying reads check each entry's BLAKE3 digest; extract refuses paths that escape the output directory, and pack does not follow symlinks.

5.3 FastCDC Sub-File Patching & Copy-on-Write Overlays

Gear-hash content-defined chunking cuts a file at boundaries that depend on its content (2 KB minimum, 8 KB average, 64 KB maximum by default), so an edit changes only the chunks near it and a patch ships only those. A patch is written to a stage file, fsynced and renamed over the target, and the directory is then fsynced (not on Windows, nor where the file system refuses a directory sync); any other fsync error is returned. Copy-on-write overlays layer patched files over a base archive without rewriting it. Measured: for a 153,639-byte file with one changed chunk, the serialized patch was 17,553 bytes (cargo run --release --example cross_fastcdc_subfile_patch, macOS arm64, release build). Savings scale with how little of the file changed:

Figure 5.2: FastCDC Patch Build & Atomic Apply
graph LR subgraph Ingestion_Stage ["Patch build"] NewAsset["New version of an asset"] --> FastCDC["FastCDC gear-hash chunking, 2 to 64 KB"] FastCDC --> ChunkHash["BLAKE3 hash per chunk"] end subgraph VFS_Deduplication ["Chunk store"] ChunkHash --> ContentIndex["Content-addressed chunk index"] ContentIndex -->|"chunk already present"| ReusePtr["Reference the existing chunk"] ContentIndex -->|"new chunk"| PatchBlob["Ship it in the patch"] end subgraph Runtime_Overlay ["Apply"] ReusePtr --> Rebuild["Rebuild the file from its chunks"] PatchBlob --> Rebuild Rebuild --> AtomicRename["fsync, rename, fsync the directory"] end

5.4 Defense in Depth: Four Layers This Build Enforces

Each layer below is code in this build. None depends on a TPM, a secure enclave or a drive controller, and none of them is DRM:

Figure 5.3: Four Layers, From Archive Parsing to the C ABI
graph TD subgraph Layer_1 ["Layer 1: archive parsing"] BoundsCheck["Sizes and counts bounded before any allocation"] MagicCheck["Unknown file magic refused, its four bytes quoted"] BoundsCheck --> MagicCheck end subgraph Layer_2 ["Layer 2: sealing, Format v3"] AeadSeal["AES-256-GCM or ChaCha20-Poly1305, 32-byte keys only"] HkdfKeys["Per-archive keys: HKDF-SHA256 salted with the archive UUID"] GenFloor["Generation floor refuses an older archive"] AeadSeal --> HkdfKeys HkdfKeys --> GenFloor end subgraph Layer_3 ["Layer 3: opt-in Lua sandbox"] LuaVM["No os.execute, os.getenv or require, text chunks only"] MemCeiling["512 MB buffer and file caps, 256 MiB decompression"] LuaVM --> MemCeiling end subgraph Layer_4 ["Layer 4: C ABI"] GuardAll["166 exports, each behind a panic guard"] BufferContract["One buffer contract: a null buffer is the size query"] GuardAll --> BufferContract end MagicCheck --> AeadSeal GenFloor --> LuaVM MemCeiling --> GuardAll

Keys go in at open and nowhere else: cyclepack_create_unpacker_with_key and cyclepack_vfs_mount_mmap_with_key in C, Unpacker::open_with_key and MmapStorage::open_with_key in Rust, cyclepack.open(path, { key, minGeneration }) in Node, Deno and Bun, and CyclePackWasm.from_buffer_with_key in the browser. A wrong key fails at open, and a key that is not exactly 32 bytes is refused, never padded or truncated. Without the key, paths, entry sizes, digests and the exact entry count stay hidden; the payload size, the padded manifest size, the UUID, the generation and the cipher stay visible, and a reader's I/O pattern still shows where entries start and end.

5.5 Signed Patch Manifests

Over-the-air patch manifests are signed with RSA-2048 (PKCS#1 v1.5 over SHA-256), as are licence tokens. Parsing yields an UnverifiedPatchManifest whose fields stay unreachable until verify(public_key_pem) accepts the signature; skipping the check takes into_manifest_without_verifying, a name that is easy to grep for. The host supplies the public key, and the library downloads nothing:

Figure 5.4: Patch Manifests Are Verified by Construction
graph LR ManifestBytes["Received manifest bytes"] --> ParseStep["Parse: UnverifiedPatchManifest"] ParseStep -->|"verify(public_key_pem)"| SigCheck{"RSA-2048 signature valid?"} SigCheck -->|"yes"| VerifiedManifest["PatchManifest: fields readable"] SigCheck -->|"no"| Refused["Error: nothing released"] ParseStep -.->|"into_manifest_without_verifying, an opt-out by name"| VerifiedManifest

5.6 Testing, Fuzzing & What Has Not Been Done

Parser-invariant smoke tests feed random bytes to the binary and text parsers and require Ok or Err, never a panic (tests/fuzz_smoke_tests.rs). Eight cargo fuzz targets cover the CASC, Format v3, FastCDC metadata, .gspx, meshlet, peer-frame, patch-manifest and path-normalisation parsers; they were not run for this paper. scripts/verify_local.sh runs rustfmt, clippy with warnings denied, the test suite, the ABI manifest check and cargo deny. No Miri, Kani, Loom or LLVM sanitizer run is part of this release, release artefacts are not signed, no SBOM or provenance attestation ships, and there is no confidential-computing support: keys and decrypted bytes live in ordinary process memory.

6. Evidence & Benchmarks

Every figure below names its source. Commands were run on macOS 26.2 arm64 (Apple silicon) on 2026-09-25, on release/0.10.0 with the sealed manifest, unless the row says otherwise. Linux and Windows builds are not verified in this release. The same figures, with the same sources, are on the Evidence section of the platform page:

Measure Result How it was obtained Notes
C-ABI exports 166, none without a panic guard python3 scripts/generate_abi_manifest.py --check All 166 frozen in the 0.10.0 baseline (abi/baselines/0.10.0); 4 are new since 0.9.0-rc.1
Capability backing 7 of 14 native, 7 emulated, 0 unavailable cyclepack doctor Debug build; the answer differs by host
Contract suites 25 + 34 passed, 0 failed cargo test --test capability_truth_tests --test mmap_engine_contract_tests Debug build, rustc 1.90
Line coverage 80.2% (floor 77%) scripts/coverage.sh (cargo-llvm-cov, all features) Measured 2026-09-24, not re-run
FastCDC patch size 17,553 bytes for a 153,639-byte file cargo run --release --example cross_fastcdc_subfile_patch One changed chunk; release build
WASM reader size 132,677 bytes gzipped (358,811 raw) scripts/build_wasm.sh (wasm-pack 0.13) Inside the script's 250 KiB budget
Converter parity 25 of 25 matched node scripts/studio_parity_test.mjs Native and WASM outputs byte-identical, except a packed archive's random UUID
Defects fixed 135 over 5 review rounds docs/changelog.md Each finding checked twice before it was changed

No latency or throughput figure appears in this paper. The one criterion benchmark (benches/vfs_bench.rs) times a 64 KiB read from an in-memory mock, path normalisation, a 16-thread read-write lock loop in which one thread in four writes, the in-process CXL pool's versioned CAS with 16 threads, and Lua dispatch and SHA-256 across the FFI boundary; the suite's timing loops run on small temporary files already in the page cache. Neither measures a disk. Run cargo bench --bench vfs_bench for numbers from your own machine. When a benchmark reads a real archive from storage, its number, command and machine will be published together.

6.2 Telemetry: Prometheus Text & OTLP JSON

The telemetry module keeps atomic latency histograms with twelve fixed buckets from 1 µs to 50 ms, and counters for prefetch hits and misses, CXL CAS contention and Lua GC pauses. It renders them as Prometheus text (export_prometheus_text) or OTLP JSON (export_otlp_json). The library opens no socket: your host serves /metrics or ships the JSON. There are no eBPF probes:

Figure 6.1: Telemetry Recorded in Process, Exported by the Host
graph LR subgraph Storage_Execution_Layer ["Recorded in process"] ReadHist["Read latency histogram, 1 us to 50 ms"] PrefetchCount["Prefetch hits and misses"] CasCount["CXL CAS contention"] GcCount["Lua GC pauses"] end subgraph Export_Formats ["Rendered on request"] PromText["export_prometheus_text"] OtlpJson["export_otlp_json"] end subgraph Observability_Stack ["Your host"] HostExport["Serves /metrics or sends OTLP"] end ReadHist --> PromText PrefetchCount --> PromText CasCount --> OtlpJson GcCount --> OtlpJson PromText --> HostExport OtlpJson --> HostExport

6.3 Fault-Injection & Crash-Consistency Tests

Two suites inject faults rather than assume them away. tests/chaos_fault_injection_tests.rs flips bits in and truncates archive headers and payloads, feeds malformed RPC frames and out-of-bounds arena writes, runs scripts that fail, and flips bits under a Merkle proof; each time it requires an error, not a panic. tests/chaos_crash_consistency_tests.rs simulates a patch interrupted at each step (a chunk write before fsync, the manifest write, the rename, a staged generation, a rollback) and checks that the active generation is the old one or the new one, never a mix. Neither simulates device faults: no NVMe error, PCIe link drop or network partition, and there is no self-healing or failover. No performance-regression gate runs in this release:

Figure 6.2: A Patch Interrupted at Any Step Leaves the Old Generation Active
graph LR PatchStart["Patch transaction"] --> WriteChunks["Write chunks, fsync"] WriteChunks --> WriteManifest["Write the manifest"] WriteManifest --> RenameStep["Atomic rename"] RenameStep --> PublishStep["Publish the generation"] PublishStep --> NewActive["Active: the new generation"] WriteChunks -.->|"interrupted"| OldActive["Active: the old generation"] WriteManifest -.->|"interrupted"| OldActive RenameStep -.->|"interrupted"| OldActive

7. Licensing, Platforms & Deployment

๐Ÿ’ฐ

1. What It Saves

  • Copies: a stored entry is borrowed from the mapping, so its bytes are not copied into a second buffer.
  • Patch bytes: a FastCDC patch ships the changed chunks only — 17,553 bytes for a 153,639-byte file with one changed chunk. Savings scale with how little changed; no fleet-level cost figure is claimed.
โšก

2. Pipeline Changes

  • Scripts, not rebuilds: packing, sealing, compression and patch steps are Lua scripts a host can load again without recompiling.
  • Live updates: a new generation is published with one pointer swap, and a bad one rolls back from a ring of the last eight.
๐Ÿ›ก๏ธ

3. Compliance & Risk Mitigation

  • Sealed archives: AES-256-GCM or ChaCha20-Poly1305 seals a whole archive, so without its key a reader learns no path, entry size or digest, and a reader holding the key refuses an edited header or manifest at open, and an edited entry when it reads it. Decryption runs on the host CPU; no SSD controller is in the path in this build. It is not DRM: a key shipped inside a game can be extracted from it.
  • Source code: An Enterprise licence includes the source code of each licensed release, not the repository history.
๐ŸŽฎ

4. Engines & Platforms

  • Plugins: Unreal Engine 5 (compiled against engine mocks; no packaged engine has run it yet), Unity 6 (qualified on a standalone C# runner; Editor PlayMode pending), Godot 4.3+ (qualified headless on 4.7.2).
  • Platforms: verified this release on macOS arm64. Linux and Windows builds are not verified in this release, consoles are not supported, and mobile is not built yet. The WASM reader is built and smoke-tested.

7.2 Live Updates Without a Rebuild

When packing or validation rules are compiled into an engine, changing one means rebuilding and shipping the engine. With CyclePack, a content or script change ships as a FastCDC patch with a signed manifest; the host verifies the manifest, stages the next generation and publishes it with a pointer swap, and a bad patch rolls back to an earlier generation. Build times, patch sizes and swap times depend on your content and have not been measured for this paper:

Figure 7.1: Rules Compiled Into the Engine vs. a CyclePack Live Update
graph LR subgraph Traditional_Pipeline ["Rules compiled into the engine"] direction TB DevCommit1["Change to a packing or validation rule"] --> BinaryRebuild["Rebuild the engine"] BinaryRebuild --> GlobalPatch["Ship a new build"] end subgraph CyclePack_Pipeline ["CyclePack live update"] direction TB DevCommit2["Change a Lua script or an asset"] --> FastCDCPatch["FastCDC patch: changed chunks only"] FastCDCPatch --> VerifyManifest["Host verifies the RSA-signed manifest"] VerifyManifest --> LiveOverlay["Stage and publish a new generation"] LiveOverlay --> RollbackStep["Roll back from the history ring if needed"] end

7.3 Engines, Languages & Platforms

One C library and one header serve every integration. Three engine plugins are maintained; sixteen engine examples are starting points, not qualified like the plugins (Bevy, Phaser, three.js/Babylon and PICO-8/TIC-80 are incomplete in this release); and 33 language drivers call the C ABI, most of whose tests skip when their toolchain is absent:

Figure 7.2: One C Library, Its Plugins and Its Platforms in 0.10.0
graph TD subgraph Core_Runtime ["libcyclepack: cdylib and staticlib"] RustCore["Rust core"] LuaSandbox["Embedded Lua 5.4"] CABIFfi["C ABI: 166 exports"] LuaSandbox --> RustCore RustCore --> CABIFfi end subgraph Console_Desktop ["Platforms in 0.10.0"] MacArm["macOS arm64: verified"] LinuxHost["Linux: not verified this release"] WindowsHost["Windows: not verified this release"] WebWasm["WASM reader: built and smoke-tested"] NoConsole["Consoles: not supported"] end subgraph Engine_Ecosystem ["Engine plugins"] UE5["Unreal Engine 5: IPlatformFile layer and IoDispatcher backend"] Unity["Unity 6: UPM package"] Godot["Godot 4.3+: GDExtension"] end subgraph Other_Hosts ["Also over the C ABI"] EngineExamples["16 engine examples"] LanguageDrivers["33 language drivers"] end CABIFfi --> MacArm CABIFfi --> LinuxHost CABIFfi --> WindowsHost RustCore --> WebWasm CABIFfi --> UE5 CABIFfi --> Unity CABIFfi --> Godot CABIFfi --> EngineExamples CABIFfi --> LanguageDrivers

7.4 Remote Chunks: Cloud Sync & Peer Delivery

CloudSyncManager::compute_missing_chunks lists which chunks of a remote recipe the local set lacks; downloading them is the host's job. Remote reads go through an HttpRangeFetcher the host implements (the crate ships only a mock), behind two cache levels: an in-process map, which drops an arbitrary entry when it is over budget (not the least recently used), and one file per chunk on disk, checked against its BLAKE3 hash on read and never evicted in this release. The peer-swarm model checks every chunk's BLAKE3 hash before it enters the local store and can hedge an urgent request between peers and a CDN. RDMA over RoCE v2 is emulated in process: no verbs provider or NIC is used. No latency figure is claimed for any level:

Figure 7.3: Remote Chunks Through a Fetcher the Host Provides
graph TD subgraph Host_Side ["Your host"] RangeFetcher["HttpRangeFetcher: fetch_range and head_size"] end subgraph CyclePack_Side ["CyclePack"] MissingList["compute_missing_chunks: what the local set lacks"] MemCache["L1: in-process map"] DiskCache["L2: one file per chunk, BLAKE3-checked on read"] SwarmModel["Peer swarm model: each chunk hash-checked before use"] end MissingList --> RangeFetcher RangeFetcher --> DiskCache SwarmModel --> DiskCache DiskCache --> MemCache

7.5 Licence Terms, Source Access & Support

The terms below are the ones on the pricing section; where this paper and that section differ, that section is right.

  • Tiers: Indie is $0 for studios whose revenue plus funding over the last 12 months is under $200,000. Pro is $2,500 per title, one time, for a production budget up to $2,000,000. Enterprise starts at $15,000 per title, or a studio-wide agreement, for budgets over $2,000,000. No tier pays a runtime royalty. Studios above the Indie threshold can evaluate free for 30 days; no commercial release under an evaluation.
  • Enforcement in 0.10.0: none. A licence is a commercial term in this release: no command, tool or read path requires one, and no feature is locked without one.
  • Source code: An Enterprise licence includes the source code of each licensed release, not the repository history.
  • Support: community support on Indie, 48-hour support on Pro, 8-hour priority support on Enterprise.
  • No telemetry: the library opens no network socket. Telemetry stays in your process until your host exports it (Prometheus text or OTLP JSON).

8. Conclusion & Next Steps

CyclePack is a Rust virtual file system that reads stored entries without copying them, seals and verifies whole archives, scripts its pipeline in Lua 5.4, and tells you which of its hardware paths are native on your machine. The hardware models exist so that your code is correct before the hardware is; the formats, the read path and the sealing are what you ship today.

Evaluate CyclePack on Your Own Assets

Studios above the Indie threshold can evaluate free for 30 days, and Indie studios can apply for the free licence. Run cyclepack doctor on your target machines and cyclepack verify on your archives before you decide.

View Licensing Tiers