ForgeEngine
Qwen3 inference engine with paged KV caching, continuous batching, CUDA/Triton paths, and streaming APIs.
Qwen3 inference engine with paged KV caching, continuous batching, CUDA/Triton paths, and streaming APIs.
CUDA, Triton, CuTe, and vLLM implementations of core LLM inference operations.
JAX/Pallas routed-MoE kernels with numerical tests and GPU benchmarks.
An exact 10-million-parameter transformer and dense-versus-MoE experiments in JAX.
A readable PyTorch decoder-only model training stack.
Implementations and experiments for positional encoding in transformer models.
A byte-level BPE tokenizer and evaluation laboratory.