Systems programmer. Go, C, C++ and assembly — concurrency, SIMD and the parts of
the runtime most people never have to think about. I make hot paths measurably faster,
and I write down how.
Numbers below are from the benchmarks in each repo, not estimates.
| Project | What it is | Measured |
|---|---|---|
BloomFilter Go AVX2 NEON |
Lock-free SIMD Bloom filter. Zero allocations on the hot path, 64-byte aligned, atomic CAS instead of locks. | 26 ns add · 23 ns contains 3–4× faster than willf/bloom |
SIMDCuckooFilter Go AVX2 NEON |
Cuckoo filter with hand-written assembly for bucket probing on x86-64 and ARM64. | 3–4× scalar (AVX2) 2–3× scalar (NEON) |
tributary C++20 header-only |
Many-producer → one-consumer fan-in over bounded SPSC rings. A producer never blocks; overload drops and counts instead. | 22–29 ns push, flat from 4 to 32 threads 3.87× throughput at 4 consumers |
CFD C11 OpenMP CUDA |
2D/3D incompressible Navier–Stokes solver. Multigrid, RANS turbulence, four backends. | Validated against Ghia lid-driven cavity, Taylor–Green, and channel flow at Reτ=395 |
Also building: thermolab and wavelab — interactive bilingual (EN/HE) university physics courses running entirely in the browser on JupyterLite.
I write up the optimisation work in long form — the wrong turns included.
- The Cleaning Robot Puzzle: A Lower Bound from a Linear Program
- The Cleaning Robot Puzzle: Five Squares at Every Corner
- The Cleaning Robot Puzzle: Cleaning a Box
- The Cleaning Robot Puzzle: A Floor with a Twist
- The Cleaning Robot Puzzle: A Room That Wraps Around
Level 1 — The systems check (C)
main(){int i=1801675112;puts(&i);}Reveal
hack — it prints the integer's bytes as ASCII: 0x6B636168 → h a c k.
Level 2 — The logic gate
If slow is smooth and smooth is fast, how much time is lost by rushing?
Reveal
All of it. Rushing creates mistakes, mistakes require fixing, fixing takes time.
Level 3 — The hidden flag
Somewhere on this page is a message that isn't rendered. Inspect the source, you must.
Reveal
Shell we play a game? — base64, in an HTML comment.
Level 4 — The byte stream (C)
#include <stdio.h>
int main(void){for(int i=0;i<15;i++)putchar(((char*)(int[]){2003790963,544434464,1869573491,682100})[i]);}Reveal
slow is smooth — Level 1's trick across four integers; the char* cast walks their bytes in memory order: 0x776F6C73 → s l o w, 0x20736920 → ␣ i s ␣, 0x6F6F6D73 → s m o o, 0x000A6874 → t h \n.




