Projects

ForgeEngine

Qwen3 inference engine with paged KV caching, continuous batching, CUDA/Triton paths, and streaming APIs.

LLM Serving Kernels

CUDA, Triton, CuTe, and vLLM implementations of core LLM inference operations.

Pallas MoE Kernels

JAX/Pallas routed-MoE kernels with numerical tests and GPU benchmarks.

Forge LLM

A readable PyTorch decoder-only model training stack.

Forge Position

Implementations and experiments for positional encoding in transformer models.