Play with it, right here.
A working demo with sample data. Nothing here touches a real system.
Follow any number down to 1
If n is even, halve it. If it is odd, triple it and add one. Repeat. Every number ever tried has reached 1, which is evidence, not proof: that gap is exactly what Collatz Lab is built to respect.
- Total steps
- 111
- Peak
- 9,232
- at step 77
- Stopping time
- 96
- first step below n
- Odd steps
- 41
- ×3 + 1 moves
n = 27 · peak 9,232 at step 77 · reaches 1 after 111 steps
Fast kernels, checked by a slow one
Collatz Lab runs on compute kernels I built for each kind of hardware: compiled loops for every CPU core in Numba and C, CUDA for NVIDIA cards and a hand-written Metal kernel for Apple GPUs. Each one has a written contract, and none of them is trusted until it agrees with a slow, exact Python reference.
SimulationOne 512-thread group · one odd seed per thread
cpu-direct · Reference
The yardstick
- CPU
- Python
- Exact integers
Every number, one at a time, all the way down to 1, in Python’s arbitrary-precision integers. Nothing can overflow and nothing is skipped. It is slow on purpose: exact references like it are what the fast kernels are checked against.
- Exact arithmetic, with no 64-bit limit to hit
- cpu-accelerated: the same exact maths, as one odd step plus its run of halvings
- The reference when full-orbit kernels are validated
Simulation: 1 core busy, 7 idle · every n, down to 1.
cpu-parallel · Every core
Every core, compiled
- CPU
- Numba JIT
- 64-bit
The same orbit, compiled to machine code by Numba and spread across every CPU core with prange. Fixed 64-bit integers keep it fast; a guard stops any seed before 3n + 1 could overflow and finishes it in exact arithmetic.
- Compiled once and cached between runs
- Thread count follows the CPU budget set in the dashboard
- cpu-parallel-odd: the same kernel on odd seeds only
Simulation: 8 cores · every n, down to 1.
cpu-sieve · Odd-only, early exit
The workhorse
- CPU
- Numba or C + OpenMP
- Canonical verifier
Even numbers are skipped, because n/2 is smaller and already checked. Each odd seed stops the moment it drops below itself, because everything smaller is already verified. Hundreds of steps shrink to a handful, and the range is still covered completely.
- Native C build with OpenMP dynamic scheduling, 4,096 seeds per chunk, so one long orbit never holds up the other cores
- Checkpoints every 250 million numbers by default; a stopped run resumes where it left off
- C and Numba paths share one Python finish: overflow patches and records
Measured250–400 M odd seeds/sAMD Ryzen 5 7600X, depending on batch size
Simulation: 8 cores · odd seeds only · stop below n.
gpu-sieve · Metal · Apple
Hand-written for Apple GPUs
- GPU
- Metal Shading Language
- Swift host
The cpu-sieve contract, written by hand in Metal. One GPU thread per odd seed, 512 threads per group. After each 3n + 1, a count-trailing-zeros instruction clears a whole run of halvings in one shift, unless that would skip past the exit.
- Each group reduces its records on the GPU to one 24-byte summary, so the CPU never reads every seed’s peak
- A persistent helper stays loaded between chunks; chunk size is set by benchmark, and 16,777,216 odd seeds won on the M1 Pro
- Refuses to start if the Swift and Metal struct layouts disagree
Measured≈ 500 M odd seeds/sApple M1 Pro laptop, lab benchmark log, 24 March 2026
Simulation: One 512-thread group · one odd seed per thread.
gpu-sieve · CUDA · NVIDIA
Built for NVIDIA cards
- GPU
- Numba CUDA
- NVIDIA
For NVIDIA GPUs, CUDA kernels compiled by Numba. Each thread works through 16 seeds in a row, so more of its time goes on arithmetic, and each 256-thread block reduces its records in shared memory, so only one summary per block travels back.
- Block size chosen so an RTX 4060 Ti keeps four blocks per multiprocessor
- gpu-collatz-accelerated: the full orbit to 1, with the same block reduction
- If any seed in a batch would overflow, the batch is redone on the CPU path with exact patching
Simulation: One 256-thread block · 16 odd seeds per thread.
cpu-barina · Experimental
Bigger jumps, still under audit
- CPU
- Numba
- Experimental
Based on David Bařina’s domain-switching method. Instead of testing odd or even at every step it hops: add one, strip the twos, multiply by a power of three from a table, subtract one, strip again. Many steps per loop iteration.
- Powers of three come from a 64-entry table
- It counts compressed steps, so its totals are not comparable with standard steps
- Kept out of standard validation until it has been audited
Simulation: 8 cores · odd seeds only · one hop, many steps.
Work to check every number up to 1,000,000
| Approach | Loop iterations |
|---|---|
| Every n, down to 1cpu-direct · naive reference | 131,434,424baseline |
| Odd n only, down to 1cpu-parallel-odd | 68,799,629÷ 1.9 |
| Odd n, stop below ncpu-sieve | 4,726,260÷ 27.8 |
| …with halvings cleared in one shiftgpu-sieve · Metal | 2,995,252÷ 43.9 |
Up to 1,000,000: the sieve needs 27.8× fewer loop iterations than the naive reference, and the Metal loop 43.9× fewer.
Exact counts computed for this page, not timings. One loop iteration is one Collatz step, except in the Metal loop, where one iteration can also clear a whole run of halvings.
Measured on real hardware
≈ 500Modd seeds checked per second
gpu-sieve on Metal, on an Apple M1 Pro laptop: every odd number up to 6.1 billion, 3.05 billion seeds, in a median 6.10 s over five runs, after a parity check against cpu-sieve.
- 418 million per second with chunks of 1M odd seeds
- 451 million per second with chunks of 2M odd seeds
- 469 million per second with chunks of 4M odd seeds
- 487 million per second with chunks of 8M odd seeds
- 503 million per second with chunks of 16M odd seeds
- cpu-sieve · AMD Ryzen 5 7600X
- 250–400M/s
- depending on batch size
- First GPU path · PyTorch MPS, same M1 Pro
- ≈ 19k/s
- the gap that led to the Metal kernel
How a result earns trust
Speed only counts if the numbers are right.
- 01A slow, exact reference
A pure-Python mirror of the sieve, with exact integers. A differential test holds the Numba, native C and GPU paths to it on the range 1 to 9,999; CI runs the suite on every push.
- 02No silent overflow
Fast paths use 64-bit integers and stop a seed before 3n + 1 could overflow. That seed is finished in exact arithmetic instead: never wrapped, never skipped. It happens: 10,994,465,585,663 climbs to 2.28 × 1024 before it drops below itself.
- 03Independent replay
Runs up to 10 million numbers are replayed in full by a second path. Larger runs: every record-setting seed is recomputed exactly, ten random windows of 10,000 numbers are re-run, and the verified range is checked for gaps.
- 04Evidence, not proof
Validated means an independent path agreed. It is trust in the numbers, never a proof of the conjecture, and the lab never claims one.
Sources: kernel code in services.py and native_sieve_kit; checks in CORRECTNESS_AND_VALIDATION.md; figures from PERFORMANCE_MACOS.md, STATUS_AND_NEXT_STEPS.md and the lab’s own Metal benchmark log (Apple M1 Pro, 24 March 2026).
Overview
- Role
- Creator
- Year
- 2025
- Platforms
- Desktop · local
- Status
- Research · in progress
Collatz Lab is my local-first research platform for the Collatz conjecture. It runs reproducible experiments on the CPU and GPU, validates them through independent code paths and links every claim to the runs behind it. It is explicitly not a proof attempt: validated computation is evidence, never proof.
The compute is my own. I built a kernel for each kind of hardware: Numba-compiled loops that use every CPU core, a native C version with OpenMP, CUDA kernels for NVIDIA cards and a hand-written Metal kernel, driven from Swift, for Apple GPUs. The main verifier skips even numbers and stops each odd seed as soon as it drops below itself, so most numbers are settled in a handful of steps. On an Apple M1 Pro, the Metal path checks about 500 million odd seeds a second.
Speed only counts if the numbers are right. Fast paths use 64-bit integers with a guard that hands any seed near overflow to exact arithmetic, a differential test holds every fast path to a slow pure-Python reference, and finished runs are replayed: in full when they are small, through record-setting seeds and random windows when they are large.
Around the kernels sits the lab itself: a FastAPI service, a command line, resumable checkpointed workers with a CPU and GPU budget, SQLite storage and a React dashboard with evidence, operations, live maths and paper views. Model assistance is limited to review and planning, never the source of truth.
8 things it does well.
- Hand-written Metal kernel, ≈500M odd seeds/s on an M1 Pro
- CUDA kernels for NVIDIA cards with shared-memory reductions
- Odd-only, early-exit CPU sieve in Numba and C with OpenMP
- Overflow guard that hands big seeds to exact arithmetic
- Every fast path tested against a slow, exact reference
- Runs replayed in full, or by record seeds and random windows
- Resumable, checkpointed workers with a CPU and GPU budget
- Claims linked to the runs behind them, with a live maths dashboard
Built with
- Python
- Numba
- C · OpenMP
- NVIDIA CUDA
- Apple Metal
- Swift
- PyTorch MPS
- NumPy
- FastAPI
- SQLite
- React
- Vite