I am an Applied Scientist at Amazon Web Services working on machine-learning systems for distributed training and inference. I received my Ph.D. in Computer and Information Science from the University of Pennsylvania in 2025, where I was advised by Vincent Liu. Before Penn, I completed my undergraduate studies at The Chinese University of Hong Kong and worked closely with James Cheng.

My research focuses on accelerator–network co-design for distributed training and inference. I am broadly interested in asynchronous runtimes, fine-grained accelerator scheduling, device-direct networking, collective communication, computation–communication overlap, compiler–runtime co-design, and data-center network simulation.

[Email] [CV] [LinkedIn] [GitHub]

Experience

Amazon Web Services — Applied Scientist
May 2025–Present

I design systems for accelerator execution and communication, including asynchronous graph-based runtimes, fine-grained on-device scheduling, and device-direct communication between AWS Trainium and Elastic Fabric Adapter (EFA).

ByteDance — Research Scientist Intern
May–August 2024

I worked on in-place gradient compression and topology-aware, adaptive collective communication for distributed training and inference.

University of Pennsylvania — Graduate Research Assistant
August 2019–May 2025

I developed systems for heterogeneous LLM serving, software-defined GPU scheduling, approximate data-center network simulation, and shared-parameter model serving.

Education

University of Pennsylvania — Ph.D. in Computer and Information Science, 2025
Thesis: Resource Sharing for Machine Learning Serving
Advisor: Vincent Liu

The Chinese University of Hong Kong — B.Sc. in Computer Science, 2019
Thesis: Efficient Similarity Search with Configurable Probabilistic Recall Guarantees

Publications

Multiplexed Heterogeneous LLM Serving via Stage-Aligned Parallelism
Tao Luo, Kelvin K.W. Ng, Zhen Ping Khor, Sidharth Sankhe, Boon Thau Loo, Vincent Liu
SoCC 2025 [PDF] [DOI]

Paella: Low-latency Model Serving with Software-defined GPU Scheduling
[Artifacts Available, Artifacts Functional, Results Reproduced]
Kelvin K.W. Ng, Henri Maxime Demoulin, Vincent Liu
SOSP 2023 [PDF]

MimicNet: Fast Performance Estimates for Data Center Networks with Machine Learning
[Artifacts Available, Artifacts Functional, Results Reproduced]
Qizhen Zhang, Kelvin K.W. Ng, Charles Kazer, Shen Yan, João Sedoc, Vincent Liu
SIGCOMM 2021 [PDF]

Norm-Explicit Quantization: Improving Vector Quantization for Maximum Inner Product Search
Xinyan Dai, Xiao Yan, Kelvin K.W. Ng, Jie Liu, James Cheng
AAAI 2020 [PDF]

Hyper-Sphere Quantization: Communication-Efficient SGD for Federated Learning
Xinyan Dai, Xiao Yan, Kaiwen Zhou, Han Yang, Kelvin K.W. Ng, James Cheng, Yu Fan
Preprint [PDF]

Pyramid: A General Framework for Distributed Similarity Search on Large-scale Datasets
Siyuan Deng, Xiao Yan, Kelvin K.W. Ng, Chenyu Jiang, James Cheng
BigData 2019 [PDF]

Fast Network Simulation Through Approximation or: How Blind Men Can Describe Elephants
Charles W. Kazer, João Sedoc, Kelvin K.W. Ng, Vincent Liu, Lyle H. Ungar
HotNets 2018 [PDF]

A General and Efficient Querying Method for Learning to Hash
Jinfeng Li, Xiao Yan, Jian Zhang, An Xu, James Cheng, Jie Liu, Kelvin K.W. Ng, Ti-chung Cheng
SIGMOD 2018 [PDF]

Guaranteed Sufficient Decrease for Stochastic Variance Reduced Gradient Optimization
Fanhua Shang, Yuanyuan Liu, Kaiwen Zhou, James Cheng, Kelvin K.W. Ng, Yuichi Yoshida
AISTATS 2018 [PDF]