About Me
I am an M.S. student in Computer Science at UC San Diego, advised by Prof. Hao Zhang in the Hao AI Lab, and I hold a B.S. from ShanghaiTech University advised by Prof. Kewei Tu. My research focuses on LLM inference infrastructure, GPU/TPU kernel optimization, sparse attention, and efficient generative-model systems.
I am currently a Student Researcher at ByteDance Seed Infra, where I own an internal multi-agent TPU kernel optimization framework and optimize attention, MoE, and fused Pallas kernels for large-scale LLM inference. My work delivered a 25% attention end-to-end speedup, up to 2.24× post-attention kernel acceleration over XLA, and a 5.41× LQQ quantization kernel speedup.
I built the official SGLang serving backend for HiLS-Attention, enabling 512K-token inference with up to 13.5× prefill and 15.7× decode speedup. At Hao AI Lab, I am a core contributor and code owner of FastVideo (4.2K+ stars), working on distributed training, custom kernels, quantization, and inference optimization. Earlier, I led FlashMHF, backed by IO-aware Triton/CUDA kernels that reduce peak memory by 3–5×.
Publications
Hierarchical Sparse Attention Done Right
arXiv Preprint, Jul 2026
Built the official SGLang serving backend and reference inference system for learned chunk-wise sparse attention. Enables 512K-token serving and 64× context-length extrapolation, with up to 13.5× prefill and 15.7× decode speedup over dense attention.
Systems & Infrastructure
TPU Multi-Agent Kernel Optimization Framework
Internal Project, Jun 2026 - Present
Code owner of internal TPU kernel agent with multi-agent scheduling and self-evolving knowledge base. Accelerated attention end-to-end by 25% (186μs → 136μs) via DMA pipelining, emit pipeline, and instruction issue reduction; also optimized pre-attention and post-attention kernels in Pallas, achieving up to 2.24× speedup over XLA. Optimized MoE kernel: 5.41× LQQ quantization speedup, 5% MoE end-to-end improvement.
FastVideo (4.2K+ Stars)
Open-Source Project, Oct 2025 - Present
Large-scale training and inference infrastructure for generative models. Own custom kernels and inference optimization; optimized LTX 2.3 on a single NVIDIA GB300 to 3× the performance of a single H100. Co-led NVFP4/FP8 quantization-aware distillation achieving 1.7s end-to-end generation, and contributed the Blackwell-optimized DreamVerse runtime with FA4, torch.compile, and kernel fusion.
CUDA/C++ Parallel Image Rendering
Personal Project, 2023
Built a C++ path tracer supporting Lambertian, metal, dielectric, and emissive materials. Implemented motion blur, depth of field, and volumetric effects. Accelerated rendering via CUDA parallelization and importance sampling, achieving ~200× speedup vs. single-threaded CPU baseline.
Education
University of California, San Diego
Sep 2025 - Dec 2026 (Expected)Master of Science in Computer Science and Engineering
Teaching Assistant, CSE 291 / DSC 291: Machine Learning Systems
La Jolla, CA
University of California, Berkeley
Aug 2023 - Jan 2024Exchange Student, EECS Department
Berkeley, CA
ShanghaiTech University
Sep 2021 - Jun 2025Bachelor of Engineering in Computer Science and Technology
Shanghai, China