Writing About

About

Everyone Should Enjoy AI

I’m Kaiyu Shi, an AI infrastructure engineer on ByteDance’s Doubao mobile assistant team. I believe everyone should enjoy AI—and strong, open infrastructure helps make that possible.

What I work on

From Models to Production

I work across model serving, training systems, and hardware optimization to make AI practical, efficient, and accessible.

  • Model serving: Deploy and optimize LLMs, VLMs, and GUI agents for Doubao mobile assistant, focusing on latency, throughput, and cost.
  • Training systems: Integrate distributed training stacks, including VeOmni and verl on Ascend 910B.
  • Hardware optimization: Improve model execution through AI compilers, NPU operators, and GPU kernels.

Selected work

Experience

ByteDance, Doubao Mobile Assistant Team AI Infrastructure Engineer · Oct 2024-Present

Large-model serving and training infrastructure

  • Deploy and optimize cloud LLM/VLM inference for Doubao mobile assistant, including GUI Agent services, with a focus on latency, throughput, and cost.
  • Build and integrate large-model training infrastructure with VeOmni and verl on Ascend 910B.
NIO, Shanghai AI Compiler System - Staff Engineer · Apr 2023-Oct 2024

AI compilers and NPU performance

  • Developed and tuned NPU kernels for in-house DSA hardware, including attention and NMS, plus scan, reduce, and sort primitives.
  • Translated model bottlenecks into compiler optimizations; contributed to operator fusion, code generation, and chip bring-up validation.
AISpeech, Shanghai Speech Chip Algorithm Engineer · 2020-Apr 2023

Edge inference and speech-chip systems

  • Built a C99-compatible inference engine and ONNX compiler pipeline for edge chips, covering graph optimization, code generation, and NPU operator and memory planning.
  • Developed compact speech models and mixed-precision QAT tooling (16/8/4/1-bit) for constrained devices, with deployment across multiple edge businesses.
AISpeech, Suzhou Algorithm Engineer · 2019-2020

Distributed speech training and GPU inference

  • Migrated speech training from Kaldi to a PyTorch distributed stack, delivering 2× training speed, one-fifth the memory use, and 5–10% lower relative WER.
  • Optimized GPU inference and CPU–GPU operators, cutting NNLM latency to one-fifth, improving end-to-end ASR RTF by 25×, and increasing concurrency by 2.5×.
Microsoft Research, Seattle Intern · Spring 2019

Memory-efficient pipeline parallelism

  • Researched pipeline model-parallel training for large DNNs, using activation exchange to reduce communication pressure.
  • Explored layered weight stashing to reduce memory use in PipeDream-style training while retaining pipeline utilization.
AISpeech, Suzhou Algorithm Intern · 2018-2019

Large-scale language model training

  • Built distributed NNLM training in PyTorch, using Block Momentum to reduce communication overhead and NCE to accelerate large-vocabulary training.
  • Trained on 80 GB of text in three days.

Background

Education and Toolbox

Education

Shanghai Jiao Tong University, SpeechLab

Master's student, 2019 - Computer Science

Shanghai Jiao Tong University

Bachelor's degree, 2016 - Materials Science and Engineering. Second degree in Computer Science and Technology

Skills

PyTorch LLM/VLM Serving Inference Optimization Training Infrastructure Mobile AI Python C/C++ Ascend 910B Heterogeneous Computing Distributed Systems DSA/NPU Design and Compile

Research

Publications

  1. Wang, Haoming, Haoyang Zou, Huatong Song, et al. UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning. arXiv:2509.02544, 2025.
  2. Yu, Kai, Rao Ma, Kaiyu Shi, and Qi Liu. Neural network language model compression with product quantization and soft binarization. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2020.
  3. Shi, Kaiyu, and Kai Yu. Structured Word Embedding for Low Memory Neural Network Language Model. Interspeech, 2018.
  4. Narayanan, D., Phanishayee, A., Shi, K., Chen, X., & Zaharia, M. Memory-efficient pipeline-parallel DNN training. International Conference on Machine Learning, 2021.
  5. Shi, Kaiyu, Xuan Liu, and Yanmin Qian. Speech Emotion Recognition Based on SVM and GMM-HMM Hybrid System. NCMMSC, 2017.

Research

Patents

  1. Shi, Kaiyu, Jinzhou Sun, Shuo Wang, and Feng Xue. Assembly methods, disassembly methods, devices and storage media. Chinese Patent CN116088864B, 2026.
  2. Sun, Jinzhou, Kaiyu Shi, Shuo Wang, and Feng Xue. Neural network model compilation method, device and storage medium. Chinese Patent Application Publication CN116225445A, 2023.
  3. Wang, Shuo, Kaiyu Shi, Jinzhou Sun, and Feng Xue. Performance simulation method, electronic device and storage medium. Chinese Patent Application Publication CN116432573A, 2023.
  4. Yu, Kai, Kaiyu Shi, Da Zheng, et al. Man-machine mixed interaction system and method based on audio. Chinese Patent CN106409283B, 2020.
  5. Yu, Kai, Xuan Liu, Di Cao, and Kaiyu Shi. Language model compression method and system. Chinese Patent Application Publication CN108874754A, 2018.
  6. Yu, Kai, and Kaiyu Shi. Compression method and system for neural network language model. Chinese Patent Application Publication CN108415888A, 2018.