Writing Home

Systems Notes

Notes on AI infrastructure, compilers, and the hardware that runs models.

Current focus

Apple Silicon

Writing

14 articles · Newest first
GPU architecture

Apple GPU Evolution

Apple GPU evolution from Dynamic Caching to Neural Accelerators, with an MLX Metal GEMM source walkthrough.

GPU kernels

Hopper kernel for MXFP4

Notes on Triton's Hopper MXFP4 matmul path, from packed FP4 unpacking into BF16 to scale conversion with inline PTX.

NCCL

Build PyTorch with NCCL2

Covers NCCL2 installation, clean PyTorch source builds, multi-GPU all-reduce tests, and network interface troubleshooting.

Tensor internals

Figure out what b.add_(b) has done

Traces PyTorch tensor addition from Python bindings through ATen, TH tensor macros, BLAS paths, OpenMP splitting, and AVX2 vector kernels.