Alexander Tian seated at a restaurant

Berkeley, California / 37.87° N

Alexander Tian

I work on sparse attention, efficient generation, and performance conscious machine learning systems at UC Berkeley.

Current work

Sparse attention

Reducing KV-cache traffic for efficient LLM decoding at MatX.

Efficient generation

VAE optimization and sparse methods at BAIR, mentored by Haocheng Xi in the Pallas Group.

Hong Kong → New York → Berkeley

Selected projects

All projects

2026

COBS: Cumulant Order Block Sparse Attention

A block-sparse attention method that uses compressed second-order statistics to estimate attention mass, approaching dense-attention retrieval quality with substantially less KV-cache read traffic.

sparse attention · LLMs · KV cache · machine learning

2025

Altair Chess Engine

A formerly top-30, 3600+ rated, superhuman chess engine in C++ with neural network evaluation, competing in international championships.

C++ · chess · search · SIMD

Interests

All interests

Chess, piano, baking, algorithms, and random things like sci-fi lol