Marco Haruni
I build efficient, reliable LLM systems—from model training and GPU kernels to inference engines and production serving.
My work includes CUDA and Triton kernels, paged KV caching, continuous batching, transformer training, JAX, PyTorch, and open-source ML systems.
I study Physiotherapy at MUHAS while independently specializing in machine learning systems and inference engineering.
Connect with me on GitHub, X, LinkedIn, or at mtaniharuni95@gmail.com.