Publications
Publications
Research Monograph

Visual Object Tracking: An Evaluation Perspective
X. Zhao, Shiyu Hu, X. Yin
Springer, Part of the book series: Advances in Computer Vision and Pattern Recognition (ACVPR)
๐ Dynamic Vision ๐ Visual Intelligence Evaluation ๐ Task-Space Diagnosis
๐ Book
Peer-Reviewed Publications
Lead or Corresponding Author

Global Instance Tracking: Locating Target More Like Humans
Shiyu Hu, X. Zhao, L. Huang, K. Huang
IEEE Transactions on Pattern Analysis and Machine Intelligence (CCF-A Journal)
๐ Open-World Tracking ๐ Global Instance Localization ๐ Human-Referenced Evaluation
๐ Paper ๐ PDF ๐ชง Poster ๐ Platform ๐ง Toolkit ๐พ Dataset

SOTVerse: A User-defined Task Space of Single Object Tracking
Shiyu Hu, X. Zhao, K. Huang
International Journal of Computer Vision (CCF-A Journal)
๐ Open-World Tracking ๐ Task-Space Modeling ๐ Diagnostic Evaluation
๐ Paper ๐ PDF ๐ชง Poster ๐ Platform

BioDrone: A Bionic Drone-based Single Object Tracking Benchmark for Robust Vision
X. Zhao, Shiyu Huโ๏ธ, Y. Wang, J. Zhang, Y. Hu, R. Liu, H. Lin, Y. Li, R. Li, K. Liu, J. Li
International Journal of Computer Vision (CCF-A Journal)
๐ Robust Visual Tracking ๐ Open-Environment Vision ๐ UAV Benchmark
๐ Paper ๐ Platform ๐ PDF ๐ง Toolkit ๐พ Dataset

A Multi-modal Global Instance Tracking Benchmark (MGIT): Better Locating Target in Complex Spatio-temporal and causal Relationship
Shiyu Hu, D. Zhang, M. Wu, X. Feng, X. Li, X. Zhao, K. Huang
Conference on Neural Information Processing Systems (CCF-A Conference, Poster)
๐ Multimodal Video Understanding ๐ Global Instance Tracking ๐ Spatiotemporal-Causal Reasoning
๐ Paper ๐ PDF ๐ชง Poster ๐น Slides ๐ Platform ๐ง Toolkit ๐พ Dataset

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking
X. Feng*, Shiyu Hu*, X. Li, D. Zhang, M. Wu, J. Zhang, X. Chen, K. Huang (*Equal Contributions)
International Conference on Computer Vision (CCF-A Conference, Highlight)
๐ Vision-Language Tracking ๐ Dynamic State Alignment ๐ Multimodal Memory
๐ Paper ๐ PDF

MATrack: Efficient Multiscale Adaptive Tracker for Real-Time Nighttime UAV Operations
X. Li*, X. Li*, Shiyu Huโ๏ธ
International Conference on Robotics and Automation (CCF-B Conference)
๐ Robust Visual Tracking ๐ Nighttime UAV Vision ๐ Real-Time Adaptation
๐ Paper ๐ PDF

DARTer: Dynamic Adaptive Representation Tracker for Nighttime UAV Tracking
X. Li*, X. Li*, Shiyu Huโ๏ธ
International Conference on Multimedia Retrieval (CCF-B Conference)
๐ Robust Visual Tracking ๐ Nighttime UAV Vision ๐ Dynamic Representation Learning
๐ Paper ๐ PDF

Visual Intelligence Evaluation Techniques for Single Object Tracking: A Survey (ๅ็ฎๆ ่ท่ธชไธญ็่ง่งๆบ่ฝ่ฏไผฐๆๆฏ็ปผ่ฟฐ)
Shiyu Hu, X. Zhao, K. Huang
Journal of Images and Graphics (ใไธญๅฝๅพ่ฑกๅพๅฝขๅญฆๆฅใ, CCF-B Chinese Journal)
๐ Dynamic Vision ๐ Visual Intelligence Evaluation ๐ Capability Diagnosis
๐ Paper ๐ PDF

Improved SAR Aircraft Detection Algorithm Based on Visual State Space Models
Y. Wang, J. Zhang, Y. Wang, Shiyu Huโ๏ธ, B. Shen, Z. Hou, W. Zhou
IET Computer Vision (CCF-C Journal)
๐ Remote Sensing ๐ SAR Aircraft Detection ๐ Vision State-Space Models
Collaborative Work

Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
X. Li*, X. Li*, Shiyu Hu, K. Huang
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CCF-A Conference)
๐ Video-LLM Reasoning ๐ Evidence Purity ๐ Agentic Frame Selection
๐ Paper ๐ PDF

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos
X. Li*, X. Li*, Shiyu Hu, K. Huang, W. Zhang
Proceedings of the AAAI Conference on Artificial Intelligence (CCF-A Conference, Oral)
๐ Video-LLM Reasoning ๐ Stepwise Causal Reasoning ๐ Diagnostic Benchmark
๐ Paper ๐ PDF ๐น Slides

VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains
X. Li*, X. Li*, Shiyu Hu, Y. Guo, W. Zhang
Proceedings of the AAAI Conference on Artificial Intelligence (CCF-A Conference, Oral)
๐ LLM Reasoning ๐ Reasoning Verifiers ๐ RLVR Evaluation
๐ Paper ๐ PDF ๐น Slides

NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation Models
X. Feng, H. Yu, M. Wu, Shiyu Hu, J. Chen, C. Zhu, J. Wu, X. Chu, K. Huang
International Conference on Learning Representations (CCF-A Conference)
๐ Generative Video Models ๐ Narrative Coherence ๐ Narrative-Centric Evaluation
๐ Paper ๐ PDF

Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
X. Li*, X. Li*, J. Gao, R. Pi, Shiyu Hu, W. Zhang
Annual Meeting of the Association for Computational Linguistics (CCF-A Conference)
๐ Multimodal Reasoning ๐ Grounded Visual Evidence ๐ Adaptive Pixel-Space Reasoning
๐ Paper ๐ PDF

Tracking by Detection and Query: An Efficient End-to-End Framework for Multi-Object Tracking
S. Jia, Shiyu Hu, Y. Cao, F. Yang, X. Lu, X. Lu
Pattern Recognition (CCF-B Journal)
๐ Multi-Object Tracking ๐ Query-Based Association ๐ Efficient End-to-End Learning
๐ Paper ๐ PDF

Talk with Your Fingers: A Depth-Aware Benchmark for Air-Writing Recognition
M. Wu, Y. Zhao, X. Li, Shiyu Hu, Y. Cai, J. Wu, W. Wang, K. Huang
IEEE Transactions on Circuits and Systems for Video Technology (CCF-B Journal)
๐ Human-Centered Vision ๐ Depth-Aware Air-Writing ๐ Multimodal Benchmark
๐ Paper

COAL: Counterfactual and Observation-Enhanced Alignment Learning for Discriminative Referring Multi-Object Tracking
S. Jia, Shiyu Hu, Y. Wang, X. Cheng, Y. Cao, X. Lu
International Joint Conference on Artificial Intelligence (CCF-B Conference)
๐ Referring Multi-Object Tracking ๐ Robust Language Grounding ๐ Counterfactual Alignment Learning
DASTrack: Rethinking Temporal Modeling in Visual Object Tracking via Decoupled Auxiliary Supervision
D. Zhang, Shiyu Hu, H. Fu, X. Feng, Y. Wang, KH Cheong, K. Huang
European Conference on Computer Vision (CCF-B Conference)
๐ Visual Object Tracking ๐ Temporal Representation Learning ๐ Decoupled Auxiliary Supervision

Global-Local Semi-Supervised Modeling for Retinal Layer Boundary Estimation in OCT
K. Li, B. Parikh, H. Yue, Shiyu Hu, S. W. Tan, W. Y. Low, X. Su, KH Cheong
Annual International Conference of the IEEE Engineering in Medicine and Biology Society (CAAI-B Conference)
๐ Medical Imaging ๐ Retinal Boundary Estimation ๐ Semi-Supervised Structure Learning

Constraint-Driven Evolution of Multimodal Video Intelligence: A Network and System Perspective
X. Li*, X. Li*, Shiyu Hu, Z. Zhang, KH Cheong
IEEE Transactions on Network Science and Engineering
๐ Multimodal Video Intelligence ๐ Constraint-Aware AI ๐ System-Level Reliability
๐ Paper

CalcTutor: Multi-Agent LLM Grading of Handwritten Mathematics with RAG-Grounded Feedback for Adaptive Learning Support
L. Tan, B. Zhu, Shiyu Hu, A. Mishra, Darren J. Yeo, KH Cheong
Mathematics
๐ AI for Education ๐ Multi-Agent Assessment ๐ RAG-Grounded Feedback
๐ Paper

CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features
X. Feng, D. Zhang, Shiyu Hu, X. Li, M. Wu, J. Zhang, X. Chen, K. Huang
International Conference on Machine Learning (CCF-A Conference, Poster)
๐ Multimodal Tracking ๐ Spatiotemporal Representation Learning ๐ Efficient Sensor Fusion
๐ Paper ๐ PDF

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
X. Feng, D. Zhang, Shiyu Hu, X. Li, M. Wu, J. Zhang, X. Chen, K. Huang
IEEE International Conference on Acoustics, Speech, and Signal Processing (CCF-B Conference, Poster)
๐ Vision-Language Tracking ๐ Visual Grounding ๐ Foundation Model Transfer
๐ Paper ๐ PDF

Artificial Intelligence-Enabled Adaptive Learning Platforms: A Review
L. Tan, Shiyu Hu, Darren J. Yeo, KH Cheong
Computers & Education: Artificial Intelligence
๐ AI for Education ๐ Personalized Learning ๐ Adaptive Learning Systems
๐ Paper ๐ PDF

A Comprehensive Review on Automated Grading Systems in STEM Using AI Techniques
L. Tan, Shiyu Hu, Darren J. Yeo, KH Cheong
Mathematics
๐ AI for Education ๐ Automated STEM Assessment ๐ Learning Analytics
๐ Paper

Trustworthy AI in education: Framework, cases, and governance strategies
Y. Ma, X. Li, Shiyu Hu, S. Liu, KH Cheong
Innovation and Emerging Technologies
๐ Trustworthy AI ๐ AI for Education ๐ Governance and Fairness
๐ Paper

A Review of Intelligent Psychological Assessment Based on Interactive Environment (ๅบไบไบคไบ็ฏๅข็ๆบ่ฝๅๅฟ็ๆต่ฏ)
K. Huang, Y. Kang, C. Yan, Shiyu Hu, L. Wang, T. Tao, W. Gao
Chinese Mental Health Journal (ใไธญๅฝๅฟ็ๅซ็ๆๅฟใ, CSSCI Journal, Top Psychological Journal in China)
๐ Human-Centered AI ๐ Interactive Psychological Assessment ๐ Validity and Ethics

Beyond Accuracy: Tracking more like Human via Visual Search
D. Zhang, Shiyu Hu, X. Feng, X. Li, M. Wu, J. Zhang, K. Huang
Conference on Neural Information Processing Systems (CCF-A Conference, Poster)
๐ Human-Centered Vision ๐ Visual Search ๐ Process-Level Evaluation
๐ Paper ๐ PDF

MemVLT: Vision-Language Tracking with Adaptive Memory-based Prompts
X. Feng, X. Li, Shiyu Hu, D. Zhang, M. Wu, J. Zhang, X. Chen, K. Huang
Conference on Neural Information Processing Systems (CCF-A Conference, Poster)
๐ Vision-Language Tracking ๐ Long-Term Memory ๐ Adaptive Prompting
๐ Paper ๐ PDF

Robust Single-particle Cryo-EM Image Denoising and Restoration
J. Zhang, T. Zhao, Shiyu Hu, X. Zhao
IEEE International Conference on Acoustics, Speech, and Signal Processing (CCF-B Conference, Poster)
๐ AI for Science ๐ Cryo-EM Restoration ๐ Diffusion Models
๐ Paper ๐ PDF

Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-World
M. Wu, K. Huang, Y. Cai, Shiyu Hu, Y. Zhao, W. Wang
IEEE Transactions on Circuits and Systems for Video Technology (CCF-B Journal)
๐ Human-Computer Interaction ๐ Unconstrained Air-Writing ๐ Real-World Benchmark
๐ Paper ๐ PDF ๐ง Toolkit

VS-LLM: Visual-Semantic Depression Assessment based on LLM for Drawing Projection Test
M. Wu, Y. Kang, X. Li, Shiyu Hu, X. Chen, Y. kang, W. Wang, K. Huang
Chinese Conference on Pattern Recognition and Computer Vision (CCF-C Conference)
๐ Human-Centered AI ๐ Multimodal Mental Health Assessment ๐ LLM-Assisted Interpretation
๐ Paper ๐ PDF

A Hierarchical Theme Recognition Model for Sandplay Therapy
X. Feng, Shiyu Hu, X. Chen, K. Huang
Chinese Conference on Pattern Recognition and Computer Vision (CCF-C Conference, Poster)
๐ Human-Centered AI ๐ Computational Mental Health ๐ Knowledge-Guided Recognition
๐ Paper ๐ PDF ๐ Supplementary ๐ชง Poster

Rethinking Similar Object Interference in Single Object Tracking
Y. Wang, Shiyu Hu, X. Zhao
International Conference on Computer Science and Artificial Intelligence (EI Conference, Oral)
๐ Robust Visual Tracking ๐ Similar-Object Interference ๐ Failure Diagnosis
๐ Paper ๐ bibTex ๐ PDF

Revisiting Instance Search: A New Benchmark Using Cycle Self-training
Y. Zhang, C. Liu, W. Chen, X. Xu, F. Wang, H. Li, Shiyu Hu, X. Zhao
Neurocomputing (CCF-C Journal)
๐ Open-World Retrieval ๐ Cross-Camera Instance Search ๐ Cycle Self-Training
๐ Paper ๐ PDF ๐ Project

Visual Turing: The Next Development of Computer Vision in The View of Human-computer Gaming (่ง่งๅพ็ต๏ผไปไบบๆบๅฏนๆ็่ฎก็ฎๆบ่ง่งไธไธๆญฅๅๅฑ)
K. Huang, X. Zhao, Q. Li, Shiyu Hu
Journal of Graphics (ใๅพๅญฆๅญฆๆฅใ, CCF-C Chinese Journal)
๐ Human-Centered AI ๐ Visual Intelligence Evaluation ๐ Human-Machine Benchmarking
๐ Paper ๐ PDF
Workshop
AAAIW 2026Learning to Be Taught: A Structured SOEI Framework for Modeling and Evaluating Personality-Aligned Virtual Student Agents, Y. Ma*, Shiyu Hu*, X. Li, Y. Wang, Y. Chen, S. Liu, KH Cheong (*Equal Contributions), the AI for Education Workshop in the 40th Annual AAAI Conference on Artificial Intelligence (Workshop in CCF-A Conference), ๐น SlidesAAAIW 2026Redefining Educational Simulation: EduVerse as a User-Defined and Developmental Multi-Agent Simulation Space, Y. Ma*, Shiyu Hu*, B. Zhu, Y. Wang, Y. Kang, S. Liu, KH Cheong (*Equal Contributions), the AI for Education Workshop in the 40th Annual AAAI Conference on Artificial Intelligence (Workshop in CCF-A Conference), ๐น SlidesAAAIW 2026From Objective to Subjective: A Benchmark for Virtual Student Abilities, B. Zhu*, Shiyu Hu*, Y. Ma, Y. Zhang, KH Cheong (*Equal Contributions), the AI for Education Workshop in the 40th Annual AAAI Conference on Artificial Intelligence (Workshop in CCF-A Conference), ๐น SlidesCVPRW 2024Diverse Text Generation for Visual Language Tracking Based on LLM, X. Li, X. Feng, Shiyu Hu, M. Wu, D. Zhang, J. Zhang, K. Huang, the 3rd Workshop on Vision Datasets Understanding and DataCV Challenge in CVPR 2024 (Workshop in CCF-A Conference, Oral, Best Paper Honorable Mention), ๐ Paper ๐ PDF ๐ชง Poster ๐น Slides ๐ Platform ๐ง Toolkit ๐พ Dataset ๐ Award
Selected collaborative publications are shown for concise browsing.
Preprints
How to Respond, How to Memorize, How to Be Fast: A Survey of Streaming Video Understanding
X. Li*, Shiyu Hu*, X. Feng, J. Zhao, K. Huang (*Equal Contributions)
๐ Streaming Video Understanding ๐ Long-Term Memory ๐ Efficient Inference
๐ Paper

FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
Shiyu Hu*, X. Li*, X. Li, J. Zhang, Y. Wang, X. Zhao, KH Cheong (*Equal Contributions)
๐ Large Vision-Language Models ๐ Human-Aligned Video Understanding ๐ Multi-Annotator Evaluation
๐ Paper ๐ PDF ๐ Project

When LLMs Learn to be Students: The SOEI Framework for Modeling and Evaluating Virtual Student Agents in Educational Interaction
Y. Ma*, Shiyu Hu*, X. Li, Y. Wang, Y. Chen, S. Liu, KH Cheong (*Equal Contributions)
๐ AI for Education ๐ Virtual Student Agents ๐ Interaction-Centered Evaluation
๐ Paper ๐ PDF

EduVerse: A User-Defined Multi-Agent Simulation Space for Education Scenario
Y. Ma*, Shiyu Hu*, B. Zhu, Y. Wang, Y. Kang, S. Liu, KH Cheong (*Equal Contributions)
๐ AI for Education ๐ Multi-Agent Simulation ๐ User-Defined Classroom Space
๐ Paper ๐ PDF

EduPersona: Benchmarking Subjective Ability Boundaries of Virtual Student Agents
B. Zhu*, Shiyu Hu*, Y. Ma, Y. Zhang, KH Cheong (*Equal Contributions)
๐ AI for Education ๐ Persona-Aware Agents ๐ Subjective Ability Diagnosis
๐ Paper ๐ PDF

SOI is the Root of All Evil: Quantifying and Breaking Similar Object Interference in Single Object Tracking
Y. Wang*, Shiyu Hu*, S. Jia, P. Xu, H. Ma, Y. Ma, J. Zhang, X. Lu, X. Zhao (*Equal Contributions)
๐ Robust Visual Tracking ๐ Similar-Object Interference ๐ VLM-Guided Correction
๐ Paper ๐ PDF

STEMVerse: A Dual-Axis Diagnostic Framework for STEM Reasoning in Large Language Models
X. Li, X. Li, J. Zhao, Shiyu Huโ๏ธ
๐ AI for Education ๐ STEM Reasoning ๐ Cognitive Diagnosis
๐ Paper ๐ PDF

DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
X. Li, Shiyu Hu, X. Feng, D. Zhang, M. Wu, J. Zhang, K. Huang
๐ Vision-Language Tracking ๐ Data-Centric AI ๐ Diverse Text Benchmark
๐ Paper ๐ PDF ๐ Project

Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
X. Li, Shiyu Hu, X. Feng, D. Zhang, M. Wu, J. Zhang, K. Huang
๐ Vision-Language Tracking ๐ Multimodal Interaction ๐ Interactive Robustness Evaluation
๐ Paper ๐ PDF ๐ Project


Selected preprints are shown for concise browsing.
