I am an undergraduate student in Mathematics & Applied Mathematics at the University of Nottingham Ningbo China. My work sits at the intersection of mathematical foundations and practical machine learning systems.

I am currently a Research Assistant at the Di² Lab, HKUST(GZ), under the supervision of Associate Professor Yutao YUE. My research focuses on efficient multimodal reasoning, long-context visual understanding, and algorithms that make large vision-language models more scalable and reliable.

You can find my publications on Google Scholar. Total citations: -.

🔥 News

  • 2026.06:  🚀 Joined Di² Lab at HKUST(GZ) as a Research Assistant, supervised by Associate Professor Yutao YUE.
  • 2026.06:  🎉🎉 Our paper TC-SSA: Token Compression via Semantic Slot Aggregation was accepted to MICCAI 2026!
  • 2026.06:  📄 Learnable Token Sparsification for Efficient Gigapixel Whole Slide Image Reasoning is now available on arXiv.
  • 2026.03:  📄 TC-SSA: Token Compression via Semantic Slot Aggregation was released on arXiv.
  • 2025.07:  🔬 Joined the Computer Vision and Recognition Center (AI觉-知研究中心) at Shenzhen University of Advanced Technology as a Research Assistant.

📝 Publications

MICCAI 2026
TC-SSA

TC-SSA: Token Compression via Semantic Slot Aggregation for Gigapixel Pathology Reasoning

Zhuo Chen, Xiaoyu Yang, Lijian Xu

  • Aggregates all WSI patch features into a fixed budget of 32 semantic slots through sparse Top-2 routing.
  • Retains global slide evidence with only 1.7% of the original visual tokens and 1.72T FLOPs.
  • Achieves 78.34% overall accuracy and 77.14% diagnosis accuracy on SlideBench (TCGA).
arXiv
SparseLearn

Learnable Token Sparsification for Efficient Gigapixel Whole Slide Image Reasoning

Jingzhi Chen, Landi He, Zhuo Chen, Shawn Young, Lijian Xu

  • Reframes WSI token pruning as an end-to-end learnable sparsification problem with decoupled training and inference.
  • Combines a variance-preserving noise gate, differentiable Soft Top-K, and diagonal-attention denoising in SparseLearn.
  • At inference, deterministic Hard Top-K retains just 32 tokens (0.78%) and reaches 73.32% overall accuracy on SlideBench (TCGA).

💻 Research Experience

2024.09 - 2028.06

B.Sc. in Mathematics & Applied Mathematics

University of Nottingham

  • Relevant Coursework: Linear Algebra, Machine Learning, Deep Learning, Computer Vision.
2025.07 - Present

Research Assistant

Shenzhen University of Advanced Technology — AI Center (Computer Vision and Recognition Center)

Under the supervision of Associate Professor Lijian Xu

  • Conducting research on Large Visual Language Models (LVLMs), studying efficiency optimization.
  • Investigating reasoning problems in long-context image understanding.
  • Working with models such as Qwen2.5-VL to bridge sensory input with cognitive understanding.
2026.06 - Present

Research Assistant

HKUST(GZ) — Di² Lab

Under the supervision of Associate Professor Yutao YUE

  • Conducting research on efficient multimodal reasoning and long-context visual understanding.
  • Exploring algorithmic approaches for robust and scalable Large Visual Language Models.