I am an undergraduate student in Mathematics & Applied Mathematics at the University of Nottingham Ningbo China. My work sits at the intersection of mathematical foundations and practical machine learning systems.
My primary research affiliation is with the Di² Lab at HKUST(GZ), where I work on text-to-motion generation and embodied AI under the supervision of Associate Professor Yutao YUE. I also work remotely with Assistant Professor Xiao Luo at the University of Wisconsin–Madison on efficient Vision-Language-Action models, including QuantVLA-related research. Previously, at the Shenzhen University of Advanced Technology, I worked on token-efficient gigapixel pathology reasoning and developed TC-SSA.
You can find my publications on Google Scholar. Total citations: -.
🔥 News
- 2026.08: 📦 Released the TC-SSA dataset on Hugging Face.
- 2026.08: 🌐 Joined Assistant Professor Xiao Luo at the University of Wisconsin–Madison as a Remote Research Assistant.
- 2026.06: 🚀 Joined Di² Lab at HKUST(GZ) as a Research Assistant, supervised by Associate Professor Yutao YUE.
- 2026.06: 🎉🎉 Our paper TC-SSA: Token Compression via Semantic Slot Aggregation was accepted to MICCAI 2026!
- 2026.06: 📄 Learnable Token Sparsification for Efficient Gigapixel Whole Slide Image Reasoning is now available on arXiv.
- 2026.03: 📄 TC-SSA: Token Compression via Semantic Slot Aggregation was released on arXiv.
- 2025.07: 🔬 Joined the Computer Vision and Recognition Center (AI觉-知研究中心) at Shenzhen University of Advanced Technology as a Research Assistant.
📝 Publications

TC-SSA: Token Compression via Semantic Slot Aggregation for Gigapixel Pathology Reasoning
Zhuo Chen, Xiaoyu Yang, Lijian Xu
📦 Dataset · Total downloads: -
- Aggregates all WSI patch features into a fixed budget of 32 semantic slots through sparse Top-2 routing.
- Retains global slide evidence with only 1.7% of the original visual tokens and 1.72T FLOPs.
- Achieves 78.34% overall accuracy and 77.14% diagnosis accuracy on SlideBench (TCGA).

Learnable Token Sparsification for Efficient Gigapixel Whole Slide Image Reasoning
Jingzhi Chen, Landi He, Zhuo Chen, Shawn Young, Lijian Xu
- Reframes WSI token pruning as an end-to-end learnable sparsification problem with decoupled training and inference.
- Combines a variance-preserving noise gate, differentiable Soft Top-K, and diagonal-attention denoising in SparseLearn.
- At inference, deterministic Hard Top-K retains just 32 tokens (0.78%) and reaches 73.32% overall accuracy on SlideBench (TCGA).
💻 Research Experience
B.Sc. in Mathematics & Applied Mathematics
University of Nottingham
- Relevant Coursework: Linear Algebra, Machine Learning, Deep Learning, Computer Vision.
Research Assistant
Shenzhen University of Advanced Technology — AI Center (Computer Vision and Recognition Center)
Under the supervision of Associate Professor Lijian Xu
- Developed TC-SSA, a semantic-slot token compression framework for efficient gigapixel whole-slide image reasoning.
- Compressed thousands of pathology patch tokens into 32 semantic slots while preserving global slide-level evidence.
Research Assistant
HKUST(GZ) — Di² Lab
Under the supervision of Associate Professor Yutao YUE
- Conducting research on text-to-motion generation and human motion synthesis from natural-language instructions.
- Exploring embodied AI systems that connect multimodal perception, language understanding, and action generation.
Remote Research Assistant
University of Wisconsin–Madison — Department of Statistics
Under the supervision of Assistant Professor Xiao Luo
- Working on QuantVLA-related research for efficient Vision-Language-Action models.
- Investigating post-training quantization and memory-efficient deployment for embodied AI systems.