AI Researcher · Incoming MSc AI Student
Building AI systems that are capable, efficient, and trustworthy.
Hello, I’m Chiwun (Christian) Yang (杨智桓; Yang Chi Wun in Cantonese). I work across machine learning theory and real-world AI systems, with a focus on large language models, efficient training and inference, AI security, and learning dynamics.
In September 2026, I will begin the MSc in Artificial Intelligence at City University of Hong Kong. I received my BEng in Artificial Intelligence from Sun Yat-sen University in 2025.
Research
I am interested in mechanism-first questions at the intersection of LLMs, optimization, and systems. I especially enjoy projects that connect mathematical structure to measurable model behavior and working implementations.
01
Efficient & Long-Context LLMs
Sparse attention, KV-cache compression, parameter-efficient learning, and systems for longer, faster inference.
02
Trustworthy & Secure AI
Copyright-aware optimization, privacy and data recovery, and mathematically auditable AI behavior.
03
Learning Dynamics & Theory
Scaling laws, optimization dynamics, generalization, and the theoretical foundations of modern architectures.
Publications & Selected Preprints
* denotes equal contribution where marked. For the complete and most current record, see Google Scholar.
- AAAI 2024 Timothy Chu*, Zhao Song*, and Chiwun Yang*. How to Protect Copyright Data in Optimization of Large Language Models?
- EMNLP 2025 Yingyu Liang*, Zhenmei Shi*, Zhao Song*, and Chiwun Yang*. Towards Infinite-Long Prefix in Transformer.
- ICML 2025 Jing Xiong, Jianghan Shen, Chuanyang Zheng, Zhongwei Wan, Chenyang Zhao, Chiwun Yang, Fanghua Ye, Hongxia Yang, Lingpeng Kong, and Ngai Wong. ParallelComp: Parallel Long-Context Compressor for Length Extrapolation.
- NeurIPS 2025 Yang Cao*, Xiaoyu Li*, Zhao Song*, and Chiwun Yang*. Efficient k-Sparse Band-Limited Interpolation with Improved Approximation Ratio.
- CPAL 2025 Majid Daliri*, Zhao Song*, and Chiwun Yang*. Unlock the Theory behind Scaling 1-bit Neural Networks.
- CPAL 2025 Yekun Ke*, Yingyu Liang*, Zhenmei Shi*, Zhao Song*, and Chiwun Yang*. Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond.
- ICLR 2025 SLLM Workshop Yichuan Deng, Zhao Song, Jing Xiong, and Chiwun Yang. How Sparse Attention Approximates Exact Attention? Your Attention is Naturally nC-Sparse.
- ICLR 2025 DeLTa Workshop Yang Cao*, Zhao Song*, and Chiwun Yang*. Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation.
- Preprint · 2023 Yichuan Deng*, Zhao Song*, Shenghao Xie*, and Chiwun Yang*. Unmasking Transformers: A Theoretical Approach to Data Recovery via Attention Weights.
- Preprint · 2025 Jiangxuan Long*, Zhao Song*, and Chiwun Yang*. Theoretical Foundation of Flow-Based Time Series Generation: Provable Approximation, Generalization, and Efficiency.
- Preprint · 2025 Chiwun Yang. Unifying Learning Dynamics and Generalization in Transformers Scaling Law.
Education & Experience
Academic Service
Reviewer for NeurIPS 2026, AAAI (2026–2027), ICLR (2025–2026), ICML (2025–2026), COLM (2025–2026), and AISTATS 2026, with additional reviewing service for an ICLR workshop and a CVPR workshop.
Let’s talk
Ideas, criticism, and collaborations are welcome.
I am always happy to discuss research on efficient and trustworthy AI, especially work that connects theory with systems.
christiannyang37 [at] gmail [dot] comLast updated: August 2026.