GPU architecture
Transport and visualization share one WebGPU device. React owns configuration and controls; the transport engine accepts a typed configuration and local datasets without React. Rendering consumes the engine's GPU field directly.
Scheduling decision
The engine uses persistent history state with bounded event batches: one invocation owns one live particle and advances up to the configured event count. Several compute dispatches share a submission, with scalar active-count readback between submissions. Histories persist across batches; they are never discarded for exceeding an event budget.
Full history kernels would minimize scheduling but stall the UI on long divergent histories. Event queues, compaction, and material sorting can improve occupancy at large populations but need scans, scatter passes, extra buffers, and scheduling decisions in WebGPU. This initial implementation chooses simpler bounded histories for small interactive populations and CSG models. It does not claim event sorting or persistent CUDA-style scheduling. Active occupancy is a diagnostic, not a compacted queue.
When timestamps are supported, the scheduler adjusts from one to eight dispatches per submission toward a 4–12 ms compute budget. This is a latency budget, not a hard GPU preemption guarantee. Rendering uses a separate animation-frame schedule; changing display settings does not restart transport. Geometry, material, source, or numerical edits invalidate fields and rebuild the affected configuration after a short debounce.
Resources and arithmetic
Particles use three aligned vec4 records: position/energy, direction/cumulative path length, and history/RNG/lifecycle/group state. The hybrid AoS layout is convenient for an invocation transporting one history; adjacent invocations access adjacent records. Geometry and material tables are immutable vec4 arrays. Cross-section searches operate on the resident processed energy grid.
Fission and particle buffers each allocate eight times the external population; device limits are checked before allocation. The engine uses seven storage bindings and one uniform binding, 64-thread workgroups, standard f32/u32 arithmetic, and supported integer atomics. No shader u64, floating-point atomics, CUDA features, or host pointers are required.
Philox4x32-10 matches Random123 known-answer vectors. Its counter partitions draw index, history index, generation, and branch epoch. Fission emission uses a separate high-bit draw domain and distinct sibling identities. Multiplication high words use exact 16-bit limb arithmetic. Uniforms are open-interval midpoints on 23 bits, avoiding f32 rounding to zero or one. The regression suite tests moments and distributions; it does not replace a full TestU01/PractRand certification.
Track lengths are stochastically rounded to scaled u32 scores. Stochastic rounding is unbiased under the RNG, adds a controlled quantization variance, and consumes independent draws. Compare-and-exchange checks each addition before overflow; the generation is rejected on overflow. Per-generation buffers reduce into f32 sums and sums of squares. This design trades atomic contention against low memory use; the saved benchmarks report its actual cost on the tested adapter.
Only a 64-byte scalar block is read between compute submissions. Source entropy adds a 512-bin histogram read at eigenvalue generation boundaries. Full field data stays resident until explicit export or validation. GPU timestamp queries are optional; unsupported devices display unavailable timing rather than a CPU duration labeled as GPU time.
Rendering and lifetime
Analytical geometry ray intersections and volume/slice sampling run in WGSL render shaders. The view reads accumulated moments directly. Geometry selection uses a small CPU analytical query for pointer picking; transport remains GPU-based. Canvas pixel ratio is capped at 1.5 to control fill cost. Scalar UI updates are throttled separately from rendering.
Pipeline compilation and dataset processing occur before a calculation begins. Compute pipelines and render pipelines are cached per device, so geometry edits reuse compiled shaders. Pause waits for the in-flight batch before buffers are reset or destroyed. Device loss stops calculations and gives a reload recovery path. GPU buffer allocations are bounded; imported datasets are validated before packing.
The backend boundary is the typed engine configuration, progress callback, run/pause/reset operations, and explicit field export. Future native backends can implement those operations without introducing an unused backend framework today.
Reference: Random123.