Sparse attention
Reducing KV-cache traffic for efficient LLM decoding at MatX.

Berkeley, California / 37.87° N
I work on sparse attention, efficient generation, and performance conscious machine learning systems at UC Berkeley.
GitHubalexander.tian [at] berkeley.edu
Current work
Reducing KV-cache traffic for efficient LLM decoding at MatX.
VAE optimization and sparse methods at BAIR, mentored by Haocheng Xi in the Pallas Group.
Hong Kong → New York → Berkeley
2026
A block-sparse attention method that uses compressed second-order statistics to estimate attention mass, approaching dense-attention retrieval quality with substantially less KV-cache read traffic.
2025
A formerly top-30, 3600+ rated, superhuman chess engine in C++ with neural network evaluation, competing in international championships.
Chess, piano, baking, algorithms, and random things like sci-fi lol