In the fall, I'll be joining Prime Intellect to scale their inference systems.

Previously, I was an intern on the Cloudflare Workers AI team, where I wrote some fast kernels and built secure systems. I'm also working on Tokenspeed, a speed-of-light LLM inference engine, with mentorship from Mingxing Zhang and Tsinghua MADSys group.

I spend my free time contributing to SGLang's Diffusion engine (I think diffusion is very cool!) while also learning about MLsys, distributed systems, and inference optimization."Once a word leaves your mouth, even four horses cannot chase it back." - 邓析

今年秋天,我将加入 Prime Intellect ,负责扩展他们的推理系统。

此前,我曾在 Cloudflare 的 Workers AI 团队实习,编写了一些高性能 kernel 并构建了安全系统。此外,我也在参与 Tokenspeed, 一个 speed-of-light LLM 推理引擎,并在 Mingxing Zhang 清华 MADSys 课题组的指导下工作。

我在空闲时间为 SGLang 的 Diffusion engine 做贡献(我觉得扩散模型非常酷!),同时学习机器学习系统、分布式系统和推理优化。一言既出,驷马难追

trees