I am an undergraduate student in Mathematics & Applied Mathematics at the University of Nottingham Ningbo China. My work sits at the intersection of mathematical foundations and practical machine learning systems.
I am currently a Research Assistant at the Di² Lab, HKUST(GZ), under the supervision of Associate Professor Yutao YUE. My research focuses on efficient multimodal reasoning, long-context visual understanding, and algorithms that make large vision-language models more scalable and reliable.
You can find my publications on Google Scholar. Total citations: -.
🔥 News
- 2026.06: 🚀 Joined Di² Lab at HKUST(GZ) as a Research Assistant, supervised by Associate Professor Yutao YUE.
- 2026.06: 🎉🎉 Our paper TC-SSA: Token Compression via Semantic Slot Aggregation was accepted to MICCAI 2026!
- 2026.06: 📄 Learnable Token Sparsification for Efficient Gigapixel Whole Slide Image Reasoning is now available on arXiv.
- 2026.03: 📄 TC-SSA: Token Compression via Semantic Slot Aggregation was released on arXiv.
- 2025.07: 🔬 Joined the Computer Vision and Recognition Center (AI觉-知研究中心) at Shenzhen University of Advanced Technology as a Research Assistant.
📝 Publications

TC-SSA: Token Compression via Semantic Slot Aggregation for Gigapixel Pathology Reasoning
Zhuo Chen, Xiaoyu Yang, Lijian Xu
- Aggregates all WSI patch features into a fixed budget of 32 semantic slots through sparse Top-2 routing.
- Retains global slide evidence with only 1.7% of the original visual tokens and 1.72T FLOPs.
- Achieves 78.34% overall accuracy and 77.14% diagnosis accuracy on SlideBench (TCGA).

Learnable Token Sparsification for Efficient Gigapixel Whole Slide Image Reasoning
Jingzhi Chen, Landi He, Zhuo Chen, Shawn Young, Lijian Xu
- Reframes WSI token pruning as an end-to-end learnable sparsification problem with decoupled training and inference.
- Combines a variance-preserving noise gate, differentiable Soft Top-K, and diagonal-attention denoising in SparseLearn.
- At inference, deterministic Hard Top-K retains just 32 tokens (0.78%) and reaches 73.32% overall accuracy on SlideBench (TCGA).
💻 Research Experience
B.Sc. in Mathematics & Applied Mathematics
University of Nottingham
- Relevant Coursework: Linear Algebra, Machine Learning, Deep Learning, Computer Vision.
Research Assistant
Shenzhen University of Advanced Technology — AI Center (Computer Vision and Recognition Center)
Under the supervision of Associate Professor Lijian Xu
- Conducting research on Large Visual Language Models (LVLMs), studying efficiency optimization.
- Investigating reasoning problems in long-context image understanding.
- Working with models such as Qwen2.5-VL to bridge sensory input with cognitive understanding.
Research Assistant
HKUST(GZ) — Di² Lab
Under the supervision of Associate Professor Yutao YUE
- Conducting research on efficient multimodal reasoning and long-context visual understanding.
- Exploring algorithmic approaches for robust and scalable Large Visual Language Models.