About

Research Fellow · Nanyang Technological University

Shiyu Hu (胡世宇)

My research examines how the reliability of AI capabilities can be modeled, measured, and diagnosed in open and interactive contexts. I study whether AI systems can maintain object identity, world state, task-relevant evidence, and interaction context as observations unfold over time or humans participate. My current work spans open-world vision, multimodal reasoning, and human-centered agents.

News

2026.08 I will serve as a Publicity Chair for the 2026 CSIG Annual Conference on Video and Image Security, to be held on 21 November 2026 in Xiong’an, China. Further information will be shared as it becomes available.

2026.08 I will give a talk, From Object Tracking to Process Understanding: State Modeling in Dynamic Vision, at the International Conference on Image and Graphics (ICIG 2026) on 4 October 2026 in Singapore, as part of the forum Visual Intelligence in Transition: From Physical Perception to Psychological Cognition.

2026.08 Two papers (FNC and ViEBench) have been accepted by EMNLP 2026, including one main-conference paper and one Findings paper.

2026.06 One paper (DASTrack) has been accepted by the 2026 European Conference on Computer Vision (ECCV, CCF-B Conference).

2026.05 I am glad to have the opportunity to give a talk at the Chinese Congress on Image and Graphics (CCIG 2026) in Guangzhou, China. Many thanks to the special session Intelligent Evolution of Video and Image Security: Perception, Reasoning, and Adversarial Challenges. I sincerely look forward to exchanging ideas with everyone and hearing your valuable suggestions.

2026.05 I received a Reviewer Award from the 43rd International Conference on Machine Learning (ICML 2026).

2026.04 One paper (RGRL) has been accepted by the main conference of the 64th Annual Meeting of the Association for Computational Linguistics (ACL, CCF-A Conference).

2026.04 One research paper has been accepted by the 48th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC, CAAI-B Conference).

2026.04 One paper (COAL) has been accepted by the 35th International Joint Conference on Artificial Intelligence (IJCAI, CCF-B Conference).

2026.03 One paper (TBDQ) has been accepted by the Pattern Recognition (PR, CCF-B Journal).

2026.02 One paper (EARL) has been accepted by the 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR, CCF-A Conference).

2026.02 One paper (MATrack) has been accepted by the 2026 IEEE International Conference on Robotics & Automation (ICRA, CCF-B Conference).

2026.02 One paper (DAAWBench) has been accepted by the IEEE Transactions on Circuits and Systems for Video Technology (TCSVT, CCF-B Journal).

2026.02 One review paper has been accepted by the IEEE Transactions on Network Science and Engineering (TNSE).

2026.01 I am honored to have participated in Humanity’s Last Exam (HLE), published in Nature, by submitting expert-level questions related to AI as a member of the HLE Contributors Consortium.

2026.01 We presented two main-conference Oral papers (CausalStep Slides and VerifyBench Slides) and three workshop papers (SOEI Slides, EduVerse Slides, and EduPersona Slides) at AAAI 2026. Thanks to everyone who visited us at the Singapore EXPO for the discussions.

2026.01 One paper (NarrLV) has been accepted by the 14th International Conference on Learning Representations (ICLR, CCF-A Conference).

2025.12 We will conduct a Mini-Symposium (topic: Complex Network Systems and Large Language Models) on NODYCON 2026 (The Fifth International Nonlinear Dynamics Conference), more information will be released soon.

2025.11 Three papers have been accepted by the AI for Education Workshop in the 40th Annual AAAI Conference on Artificial Intelligence (AAAIW).

2025.11 Two papers (CausalStep and VerifyBench) have been accepted by the 40th Annual AAAI Conference on Artificial Intelligence (AAAI, CCF-A Conference, Oral).

2025.10 We have conducted a tutorial at 28th European Conference on Artificial Intelligence (ECAI) (26th October, 2025, Bologna, Italy).

2025.10 We have conducted a tutorial at 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC) (5th October, 2025, Vienna, Austria).

2025.08 We have conducted a tutorial at 34th International Joint Conference on Artificial Intelligence (IJCAI) (18th August, 2025, Montreal, Canada).

2025.07 Obtain IEEE SMCS TEAM Program Award.

2025.06 One paper (ATCTrack) has been accepted by International Conference on Computer Vision (ICCV, CCF-A conference, Highlight).

2025.05 One review paper has been accepted by Computers and Education: Artificial Intelligence.

2025.05 Our new work FIOVA is now online! We introduce a multi-annotator benchmark for human-aligned video captioning, supporting semantic diversity and cognitive-aware evaluation. Check out the project page and arXiv paper for more details.

2025.05 Our updated work SOEI now available! Building upon our previous framework, this version introduces interactive multi-turn simulation to model open-ended educational dialogues with cognitively plausible virtual students. We further validate the framework’s effectiveness through behavioral analysis, personality recognition, and teacher-student reflection. Read more in the arXiv paper.

2025.05 We will present our work (SOTVerse) at IJCV2024 during the VALSE2025 poster session (June 2025, Zhuhai, China).

2025.05 One paper (CSTrack) has been accepted by International Conference on Machine Learning (ICML, CCF-A conference).

2025.05 One paper (DARTer) has been accepted by International Conference on Multimedia Retrieval (ICMR, CCF-B conference).

2025.05 One paper (MSAD) has been accepted by IET Computer Vision (IET-CVI, CCF-C journal).

2025.04 We will conduct a tutorial at 28th European Conference on Artificial Intelligence (ECAI) (25th-30th October, 2025, Bologna, Italy).

2025.03 We will conduct a tutorial at 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC) (5th-8th October, 2025, Vienna, Austria).

2025.02 The book Visual Object Tracking: An Evaluation Perspective is online.

2025.01 One paper (CTVLT) has been accepted by IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP, CCF-B conference).

2025.01 A special issue (Techniques and Applications of Multimodal Data Fusion) in Electronics has been announced, all papers related to this topic are welcomed for submission!

2024.12 We have conducted a tutorial at Asian Conference on Computer Vision (ACCV) (Dec. 9th 2024, Hanoi, Vietnam).

2024.12 We have prepared a tutorial at International Conference on Pattern Recognition (ICPR) (Dec. 1st 2024, Kolkata, India).

2024.10 We have conducted a tutorial at IEEE International Conference on Image Processing (ICIP) (Oct. 27th 2024, Abu Dhabi, United Arab Emirates).

2024.09 Two papers (MemVLT and CPDTrack) have been accepted by Conference on Neural Information Processing Systems (NeurIPS, CCF-A Conference).

2024.08 One tutorial proposal has been accepted by Asian Conference on Computer Vision (ACCV), the tutorial will be conducted in Dec. 2024 (Hanoi, Vietnam).

2024.08 Start my work as a Research Fellow in Nanyang Technological University (NTU), Singapore.

2024.07 One tutorial proposal has been accepted by International Conference on Pattern Recognition (ICPR), the tutorial will be conducted in Dec. 2024 (Kolkata, India).

2024.06 One paper has been accepted by Chinese Conference on Pattern Recognition and Computer Vision (PRCV).

2024.06 One paper has been accepted by Chinese Mental Health Journal (《中国心理卫生杂志》).

2024.05 Obtain Beijing Outstanding Graduates (北京市优秀毕业生, top 5%).

2024.05 We have presented our work (Global Instance Tracking (GIT)) at TPAMI2023 during the VALSE2024 poster session (May 2024, Chongqing, China, see our Poster for more information).

2024.04 One paper has been accepted by the 3rd Workshop on Vision Datasets Understanding and DataCV Challenge in CVPR 2024 (CVPRW, Oral, Best Paper Honorable Mention).

2024.04 One paper has been accepted by IEEE Transactions on Circuits and Systems for Video Technology (TCSVT).

2024.01 One project about human-computer interaction in intelligent education has been funded by the 2023 Intelligent Education PhD Research Fund, supported by the Institute of AI Education Shanghai and East China Normal University.

2024.01 Got my Ph.D. degree at Institute of Automation, Chinese Academy of Sciences (CASIA) and University of Chinese Academy of Sciences (UCAS).

2023.12 One paper has been accepted by the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP, CCF-B conference).

2023.11 One paper has been accepted by International Conference on Computer Science and Artificial Intelligence (CSAI, Oral).

2023.10 Obtain China National Scholarship (国家奖学金, top 1%, only 8 Ph.D. students in main campus of UCAS win this scholarship).

2023.10 Obtain First Prize of Climbing Scholarship (攀登一等奖学金, only 6 students in CASIA win this scholarship).

2023.10 One paper has been accepted by International Journal of Computer Vision (IJCV, CCF-A journal).

2023.09 One survey has been accepted by Journal of Images and Graphics (《中国图像图形学报》).

2023.09 One paper has been accepted by Conference on Neural Information Processing Systems (NeurIPS, CCF-A conference).

2023.09 One paper has been accepted by International Journal of Computer Vision (IJCV, CCF-A journal).

2023.08 One paper has been accepted by Chinese Conference on Pattern Recognition and Computer Vision (PRCV, CCF-C conference).

2022.06 Obtain merit student of University of Chinese Academy of Sciences.

2022.06 One paper has been accepted by Neurocomputing (Neu).

2022.02 One paper has been accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI, CCF-A journal).

2021.06 One survey has been accepted by Journal of Graphics (《图学学报》).

Background

Work

Professional Experience

Research Intern

Institute of Electronics, Chinese Academy of Sciences (CASIE)

Education

Education

Ph.D.

Institute of Automation, Chinese Academy of Sciences (CASIA)

Field: Computer Application Technology

Supervisor: Prof. Kaiqi Huang

Co-supervisor: Prof. Xin Zhao

Thesis and degree details

Thesis: Research of Intelligence Evaluation Techniques for Single Object Tracking

Supervisor: Prof. Kaiqi Huang (IAPR Fellow, IEEE Senior Member, 10,000 Talents Program – Leading Talents)

Co-supervisor: Prof. Xin Zhao (IEEE Senior Member, Beijing Science Fund for Distinguished Young Scholars)

Thesis Committee: Prof. Jianbin Jiao; Prof. Yuxin Peng (National Science Fund for Distinguished Young Scholars); Prof. Yao Zhao (IEEE Fellow, IET Fellow, National Science Fund for Distinguished Young Scholars); Prof. Yunhong Wang (IEEE Fellow, IAPR Fellow, CCF Fellow); and Prof. Ming Tang

Thesis Defense Grade: Excellent

B.E. in the Elite Class

Beijing Institute of Technology (BIT)

School of Information and Electronics

Field: Information Engineering

Thesis and degree details

Undergraduate Thesis Supervisor: Prof. Senlin Luo

Thesis: Text Sentiment Analysis Based on Deep Neural Network

Thesis Defense Grade: Excellent

Research Interests

Research Trajectory

Research trajectory from visual perception tasks to human-grounded intelligence evaluation

My research began with visual object tracking and machine vision evaluation, spanning task modeling, evaluation environments, measurement techniques, and human-machine comparison. Inspired by the Turing Test, I proposed the Visual Turing Test to evaluate dynamic visual intelligence against human abilities.

The central question has remained consistent: how can AI capabilities be modeled, measured, and diagnosed in open, human-centered contexts? Rather than reporting benchmark scores alone, I study what a system can perceive, understand, and reliably maintain as tasks and environments become more open.

From visual localization to capability diagnosis

Visual Object Tracking provided a concrete task for studying dynamic visual ability. Global Instance Tracking (GIT) extends tracking to long-term target retrieval, while Multi-modal GIT (MGIT) introduces hierarchical semantics and spatiotemporal-causal reasoning.

From closed benchmarks to open environments

Human visual environments are continuous, open, and semantically rich. VideoCube organizes long videos through narrative structure, SOTVerse supports user-defined task spaces, and BioDrone examines reliable perception under physical disturbance.

From performance comparison to human-grounded evaluation

By placing people and models in comparable visual tasks, I study how their capabilities differ and where conventional metrics conceal those differences. This perspective now guides my work on human-grounded evaluation and reliable human-AI collaboration.

Current Research Directions

Open-World Vision

I study whether visual systems can maintain target identity and world state under occlusion, interference, and physical disturbance. Current topics include open-world tracking, vision-language grounding, visual memory, and reliable perception for UAVs and embodied systems.

Evidence-Grounded Multimodal Reasoning

I investigate whether multimodal models select and use the right evidence in images and long videos. My work covers spatiotemporal and causal reasoning, streaming memory, adaptive visual computation, and process-level verification.

Human-Centered Agents and AI for Education

I develop agents that model cognitive and learning states, capability boundaries, and social interaction. Education provides a practical setting for virtual students, personalized agents, multi-agent simulation, and reliable human-AI collaboration.

The 3E framework connecting environment, evaluation, and executors

The 3E framework connects Environment, Evaluation, and Executors in one evaluation loop. Across these directions, I build open task spaces, human-grounded protocols, and process-level diagnostics to reveal capability boundaries and improve model and system design.

Publications

Research Monograph

Springer 2025
sym

Visual Object Tracking: An Evaluation Perspective
X. Zhao, Shiyu Hu, X. Yin
Springer, Part of the book series: Advances in Computer Vision and Pattern Recognition (ACVPR)
Dynamic Vision · Visual Intelligence Evaluation · Task-Space Diagnosis
📘 Book

Peer-Reviewed Publications

Lead or Corresponding Author

TPAMI 2023
sym

Global Instance Tracking: Locating Target More Like Humans
Shiyu Hu, X. Zhao, L. Huang, K. Huang
IEEE Transactions on Pattern Analysis and Machine Intelligence (CCF-A Journal)
Open-World Tracking · Global Instance Localization · Human-Referenced Evaluation
📃 Paper 📑 PDF 🪧 Poster 🌐 Platform 🔧 Toolkit 💾 Dataset

IJCV 2024
sym

SOTVerse: A User-defined Task Space of Single Object Tracking
Shiyu Hu, X. Zhao, K. Huang
International Journal of Computer Vision (CCF-A Journal)
Open-World Tracking · Task-Space Modeling · Diagnostic Evaluation
📃 Paper 📑 PDF 🪧 Poster 🌐 Platform

IJCV 2024
sym

BioDrone: A Bionic Drone-based Single Object Tracking Benchmark for Robust Vision
X. Zhao, Shiyu Hu✉️, Y. Wang, J. Zhang, Y. Hu, R. Liu, H. Lin, Y. Li, R. Li, K. Liu, J. Li
International Journal of Computer Vision (CCF-A Journal)
Robust Visual Tracking · Open-Environment Vision · UAV Benchmark
📃 Paper 🌐 Platform 📑 PDF 🔧 Toolkit 💾 Dataset

NeurIPS 2023
sym

A Multi-modal Global Instance Tracking Benchmark (MGIT): Better Locating Target in Complex Spatio-temporal and causal Relationship
Shiyu Hu, D. Zhang, M. Wu, X. Feng, X. Li, X. Zhao, K. Huang
Conference on Neural Information Processing Systems (CCF-A Conference, Poster)
Multimodal Video Understanding · Global Instance Tracking · Spatiotemporal-Causal Reasoning
📃 Paper 📑 PDF 🪧 Poster 📹 Slides 🌐 Platform 🔧 Toolkit 💾 Dataset

ICCV 2025
sym

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking
X. Feng*, Shiyu Hu*, X. Li, D. Zhang, M. Wu, J. Zhang, X. Chen, K. Huang (*Equal Contributions)
International Conference on Computer Vision (CCF-A Conference, Highlight)
Vision-Language Tracking · Dynamic State Alignment · Multimodal Memory
📃 Paper 📑 PDF

ICRA 2026
sym

MATrack: Efficient Multiscale Adaptive Tracker for Real-Time Nighttime UAV Operations
X. Li*, X. Li*, Shiyu Hu✉️
International Conference on Robotics and Automation (CCF-B Conference)
Robust Visual Tracking · Nighttime UAV Vision · Real-Time Adaptation
📃 Paper 📑 PDF

ICMR 2025
sym

DARTer: Dynamic Adaptive Representation Tracker for Nighttime UAV Tracking
X. Li*, X. Li*, Shiyu Hu✉️
International Conference on Multimedia Retrieval (CCF-B Conference)
Robust Visual Tracking · Nighttime UAV Vision · Dynamic Representation Learning
📃 Paper 📑 PDF

中国图象图形学报 2024
sym

Visual Intelligence Evaluation Techniques for Single Object Tracking: A Survey (单目标跟踪中的视觉智能评估技术综述)
Shiyu Hu, X. Zhao, K. Huang
Journal of Images and Graphics (《中国图象图形学报》, CCF-B Chinese Journal)
Dynamic Vision · Visual Intelligence Evaluation · Capability Diagnosis
📃 Paper 📑 PDF

IET-CVI 2025
sym

Improved SAR Aircraft Detection Algorithm Based on Visual State Space Models
Y. Wang, J. Zhang, Y. Wang, Shiyu Hu✉️, B. Shen, Z. Hou, W. Zhou
IET Computer Vision (CCF-C Journal)
Remote Sensing · SAR Aircraft Detection · Vision State-Space Models

Selected lead or corresponding-author publications are shown for concise browsing.

Collaborative Work

CVPR 2026
sym

Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
X. Li*, X. Li*, Shiyu Hu, K. Huang
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CCF-A Conference)
Video-LLM Reasoning · Evidence Purity · Agentic Frame Selection
📃 Paper 📑 PDF

AAAI 2026
sym

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos
X. Li*, X. Li*, Shiyu Hu, K. Huang, W. Zhang
Proceedings of the AAAI Conference on Artificial Intelligence (CCF-A Conference, Oral)
Video-LLM Reasoning · Stepwise Causal Reasoning · Diagnostic Benchmark
📃 Paper 📑 PDF 📹 Slides

AAAI 2026
sym

VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains
X. Li*, X. Li*, Shiyu Hu, Y. Guo, W. Zhang
Proceedings of the AAAI Conference on Artificial Intelligence (CCF-A Conference, Oral)
LLM Reasoning · Reasoning Verifiers · RLVR Evaluation
📃 Paper 📑 PDF 📹 Slides

ICLR 2026
sym

NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation Models
X. Feng, H. Yu, M. Wu, Shiyu Hu, J. Chen, C. Zhu, J. Wu, X. Chu, K. Huang
International Conference on Learning Representations (CCF-A Conference)
Generative Video Models · Narrative Coherence · Narrative-Centric Evaluation
📃 Paper 📑 PDF

ACL 2026
sym

Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
X. Li*, X. Li*, J. Gao, R. Pi, Shiyu Hu, W. Zhang
Annual Meeting of the Association for Computational Linguistics (CCF-A Conference)
Multimodal Reasoning · Grounded Visual Evidence · Adaptive Pixel-Space Reasoning
📃 Paper 📑 PDF

PR 2026
sym

Tracking by Detection and Query: An Efficient End-to-End Framework for Multi-Object Tracking
S. Jia, Shiyu Hu, Y. Cao, F. Yang, X. Lu, X. Lu
Pattern Recognition (CCF-B Journal)
Multi-Object Tracking · Query-Based Association · Efficient End-to-End Learning
📃 Paper 📑 PDF

TCSVT 2026
sym

Talk with Your Fingers: A Depth-Aware Benchmark for Air-Writing Recognition
M. Wu, Y. Zhao, X. Li, Shiyu Hu, Y. Cai, J. Wu, W. Wang, K. Huang
IEEE Transactions on Circuits and Systems for Video Technology (CCF-B Journal)
Human-Centered Vision · Depth-Aware Air-Writing · Multimodal Benchmark
📃 Paper

IJCAI 2026
sym

COAL: Counterfactual and Observation-Enhanced Alignment Learning for Discriminative Referring Multi-Object Tracking
S. Jia, Shiyu Hu, Y. Wang, X. Cheng, Y. Cao, X. Lu
International Joint Conference on Artificial Intelligence (CCF-B Conference)
Referring Multi-Object Tracking · Robust Language Grounding · Counterfactual Alignment Learning

ECCV 2026
Complete DASTrack figure

DASTrack: Rethinking Temporal Modeling in Visual Object Tracking via Decoupled Auxiliary Supervision
D. Zhang, Shiyu Hu, H. Fu, X. Feng, Y. Wang, KH Cheong, K. Huang
European Conference on Computer Vision (CCF-B Conference)
Visual Object Tracking · Temporal Representation Learning · Decoupled Auxiliary Supervision

EMNLP 2026
Fake News Court multi-agent adversarial framework

Fake News Court: A Multi-Agent Adversarial Framework for Robust Detection of LLM-Generated Fake News
M. Peng, H. Gu, Y. Ma, Shiyu Hu, X. Huang
Conference on Empirical Methods in Natural Language Processing (CCF-B Conference)
LLM-Generated Fake News · Multi-Agent Adversarial Reasoning · Robust Detection

EMNLP Findings 2026
ViEBench grounded visual evidence evaluation framework

Beyond Accuracy: Evaluating Grounded Visual Evidence in Thinking with Images
X. Li*, X. Li*, R. Pi, Shiyu Hu, J. Zhao, J. Gao
Conference on Empirical Methods in Natural Language Processing (EMNLP Findings, CCF-B Conference)
Agentic Visual Reasoning · Grounded Visual Evidence · Dual-Axis Diagnosis
📃 Paper 📑 PDF

EMBC 2026
sym

Global-Local Semi-Supervised Modeling for Retinal Layer Boundary Estimation in OCT
K. Li, B. Parikh, H. Yue, Shiyu Hu, S. W. Tan, W. Y. Low, X. Su, KH Cheong
Annual International Conference of the IEEE Engineering in Medicine and Biology Society (CAAI-B Conference)
Medical Imaging · Retinal Boundary Estimation · Semi-Supervised Structure Learning

TNSE 2026
sym

Constraint-Driven Evolution of Multimodal Video Intelligence: A Network and System Perspective
X. Li*, X. Li*, Shiyu Hu, Z. Zhang, KH Cheong
IEEE Transactions on Network Science and Engineering
Multimodal Video Intelligence · Constraint-Aware AI · System-Level Reliability
📃 Paper

Mathematics 2026
sym

CalcTutor: Multi-Agent LLM Grading of Handwritten Mathematics with RAG-Grounded Feedback for Adaptive Learning Support
L. Tan, B. Zhu, Shiyu Hu, A. Mishra, Darren J. Yeo, KH Cheong
Mathematics
AI for Education · Multi-Agent Assessment · RAG-Grounded Feedback
📃 Paper

Nature 2026
Figure 1 from the HLE paper comparing frontier-model accuracy on HLE, GPQA, MATH, and MMLU

A benchmark of expert-level academic questions to assess AI capabilities
Center for AI Safety, Scale AI, and HLE Contributors Consortium
Nature, 649, 1139–1146 (2026)
Contribution: Shiyu Hu submitted expert-level questions related to AI to the HLE benchmark as a member of the HLE Contributors Consortium.
Frontier AI Evaluation · Expert-Level Benchmarking · Consortium Contribution
📃 Paper 🌐 Benchmark

ICML 2025
sym

CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features
X. Feng, D. Zhang, Shiyu Hu, X. Li, M. Wu, J. Zhang, X. Chen, K. Huang
International Conference on Machine Learning (CCF-A Conference, Poster)
Multimodal Tracking · Spatiotemporal Representation Learning · Efficient Sensor Fusion
📃 Paper 📑 PDF

ICASSP 2025
sym

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
X. Feng, D. Zhang, Shiyu Hu, X. Li, M. Wu, J. Zhang, X. Chen, K. Huang
IEEE International Conference on Acoustics, Speech, and Signal Processing (CCF-B Conference, Poster)
Vision-Language Tracking · Visual Grounding · Foundation Model Transfer
📃 Paper 📑 PDF

C&E:AI 2025
sym

Artificial Intelligence-Enabled Adaptive Learning Platforms: A Review
L. Tan, Shiyu Hu, Darren J. Yeo, KH Cheong
Computers & Education: Artificial Intelligence
AI for Education · Personalized Learning · Adaptive Learning Systems
📃 Paper 📑 PDF

Mathematics 2025
sym

A Comprehensive Review on Automated Grading Systems in STEM Using AI Techniques
L. Tan, Shiyu Hu, Darren J. Yeo, KH Cheong
Mathematics
AI for Education · Automated STEM Assessment · Learning Analytics
📃 Paper

Innovation and Emerging Technologies 2025
sym

Trustworthy AI in education: Framework, cases, and governance strategies
Y. Ma, X. Li, Shiyu Hu, S. Liu, KH Cheong
Innovation and Emerging Technologies
Trustworthy AI · AI for Education · Governance and Fairness
📃 Paper

中国心理卫生杂志 2025
sym

A Review of Intelligent Psychological Assessment Based on Interactive Environment (基于交互环境的智能化心理测评)
K. Huang, Y. Kang, C. Yan, Shiyu Hu, L. Wang, T. Tao, W. Gao
Chinese Mental Health Journal (《中国心理卫生杂志》, CSSCI Journal, Top Psychological Journal in China)
Human-Centered AI · Interactive Psychological Assessment · Validity and Ethics

NeurIPS 2024
sym

Beyond Accuracy: Tracking more like Human via Visual Search
D. Zhang, Shiyu Hu, X. Feng, X. Li, M. Wu, J. Zhang, K. Huang
Conference on Neural Information Processing Systems (CCF-A Conference, Poster)
Human-Centered Vision · Visual Search · Process-Level Evaluation
📃 Paper 📑 PDF

NeurIPS 2024
sym

MemVLT: Vision-Language Tracking with Adaptive Memory-based Prompts
X. Feng, X. Li, Shiyu Hu, D. Zhang, M. Wu, J. Zhang, X. Chen, K. Huang
Conference on Neural Information Processing Systems (CCF-A Conference, Poster)
Vision-Language Tracking · Long-Term Memory · Adaptive Prompting
📃 Paper 📑 PDF

ICASSP 2024
sym

Robust Single-particle Cryo-EM Image Denoising and Restoration
J. Zhang, T. Zhao, Shiyu Hu, X. Zhao
IEEE International Conference on Acoustics, Speech, and Signal Processing (CCF-B Conference, Poster)
AI for Science · Cryo-EM Restoration · Diffusion Models
📃 Paper 📑 PDF

TCSVT 2024
sym

Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-World
M. Wu, K. Huang, Y. Cai, Shiyu Hu, Y. Zhao, W. Wang
IEEE Transactions on Circuits and Systems for Video Technology (CCF-B Journal)
Human-Computer Interaction · Unconstrained Air-Writing · Real-World Benchmark
📃 Paper 📑 PDF 🔧 Toolkit

PRCV 2024
sym

VS-LLM: Visual-Semantic Depression Assessment based on LLM for Drawing Projection Test
M. Wu, Y. Kang, X. Li, Shiyu Hu, X. Chen, Y. kang, W. Wang, K. Huang
Chinese Conference on Pattern Recognition and Computer Vision (CCF-C Conference)
Human-Centered AI · Multimodal Mental Health Assessment · LLM-Assisted Interpretation
📃 Paper 📑 PDF

PRCV 2023
sym

A Hierarchical Theme Recognition Model for Sandplay Therapy
X. Feng, Shiyu Hu, X. Chen, K. Huang
Chinese Conference on Pattern Recognition and Computer Vision (CCF-C Conference, Poster)
Human-Centered AI · Computational Mental Health · Knowledge-Guided Recognition
📃 Paper 📑 PDF 📎 Supplementary 🪧 Poster

CSAI 2023
sym

Rethinking Similar Object Interference in Single Object Tracking
Y. Wang, Shiyu Hu, X. Zhao
International Conference on Computer Science and Artificial Intelligence (EI Conference, Oral)
Robust Visual Tracking · Similar-Object Interference · Failure Diagnosis
📃 Paper 🧾 BibTeX 📑 PDF

Neurocomputing 2022
sym

Revisiting Instance Search: A New Benchmark Using Cycle Self-training
Y. Zhang, C. Liu, W. Chen, X. Xu, F. Wang, H. Li, Shiyu Hu, X. Zhao
Neurocomputing (CCF-C Journal)
Open-World Retrieval · Cross-Camera Instance Search · Cycle Self-Training
📃 Paper 📑 PDF 🌐 Project

图学学报 2021
sym

Visual Turing: The Next Development of Computer Vision in The View of Human-computer Gaming (视觉图灵:从人机对抗看计算机视觉下一步发展)
K. Huang, X. Zhao, Q. Li, Shiyu Hu
Journal of Graphics (《图学学报》, CCF-C Chinese Journal)
Human-Centered AI · Visual Intelligence Evaluation · Human-Machine Benchmarking
📃 Paper 📑 PDF

Workshop

Selected collaborative publications are shown for concise browsing.

Preprints

Preprint
Complete figure for the survey of streaming video understanding

How to Respond, How to Memorize, How to Be Fast: A Survey of Streaming Video Understanding
X. Li*, Shiyu Hu*, X. Feng, J. Zhao, K. Huang (*Equal Contributions)
Streaming Video Understanding · Long-Term Memory · Efficient Inference
📃 Paper

Preprint
sym

FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
Shiyu Hu*, X. Li*, X. Li, J. Zhang, Y. Wang, X. Zhao, KH Cheong (*Equal Contributions)
Large Vision-Language Models · Human-Aligned Video Understanding · Multi-Annotator Evaluation
📃 Paper 📑 PDF 🌐 Project

Preprint
sym

When LLMs Learn to be Students: The SOEI Framework for Modeling and Evaluating Virtual Student Agents in Educational Interaction
Y. Ma*, Shiyu Hu*, X. Li, Y. Wang, Y. Chen, S. Liu, KH Cheong (*Equal Contributions)
AI for Education · Virtual Student Agents · Interaction-Centered Evaluation
📃 Paper 📑 PDF

Preprint
sym

EduVerse: A User-Defined Multi-Agent Simulation Space for Education Scenario
Y. Ma*, Shiyu Hu*, B. Zhu, Y. Wang, Y. Kang, S. Liu, KH Cheong (*Equal Contributions)
AI for Education · Multi-Agent Simulation · User-Defined Classroom Space
📃 Paper 📑 PDF

Preprint
sym

EduPersona: Benchmarking Subjective Ability Boundaries of Virtual Student Agents
B. Zhu*, Shiyu Hu*, Y. Ma, Y. Zhang, KH Cheong (*Equal Contributions)
AI for Education · Persona-Aware Agents · Subjective Ability Diagnosis
📃 Paper 📑 PDF

Preprint
sym

SOI is the Root of All Evil: Quantifying and Breaking Similar Object Interference in Single Object Tracking
Y. Wang*, Shiyu Hu*, S. Jia, P. Xu, H. Ma, Y. Ma, J. Zhang, X. Lu, X. Zhao (*Equal Contributions)
Robust Visual Tracking · Similar-Object Interference · VLM-Guided Correction
📃 Paper 📑 PDF

Preprint
sym

How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking
X. Li*, Shiyu Hu*, X. Feng, D. Zhang, M. Wu, J. Zhang, K. Huang (*Equal Contributions)
Vision-Language Tracking · Language Utility Diagnosis · Fine-Grained Evaluation
📃 Paper 📑 PDF

Preprint
Complete STEMVerse diagnostic framework for STEM reasoning

STEMVerse: A Dual-Axis Diagnostic Framework for STEM Reasoning in Large Language Models
X. Li, X. Li, J. Zhao, Shiyu Hu✉️
AI for Education · STEM Reasoning · Cognitive Diagnosis
📃 Paper 📑 PDF

Preprint
sym

DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
X. Li, Shiyu Hu, X. Feng, D. Zhang, M. Wu, J. Zhang, K. Huang
Vision-Language Tracking · Data-Centric AI · Diverse Text Benchmark
📃 Paper 📑 PDF 🌐 Project

Preprint
sym

Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
X. Li, Shiyu Hu, X. Feng, D. Zhang, M. Wu, J. Zhang, K. Huang
Vision-Language Tracking · Multimodal Interaction · Interactive Robustness Evaluation
📃 Paper 📑 PDF 🌐 Project

Preprint
sym

Nearing or Surpassing: Overall Evaluation of Human-Machine Dynamic Vision Ability
Shiyu Hu, X. Zhao, Y. Wang, Y. Shan, K. Huang
Human-Centered AI · Dynamic Vision Capability · Human-Machine Evaluation
📑 PDF

Selected preprints are shown for concise browsing.

Projects

This section documents research software, evaluation platforms, academic challenges, and funded projects. Associated research outputs are listed under Publications.

Research Software

Darknet-Cross

Lightweight Deep Learning Framework for Heterogeneous Computing

2018.03 – 2018.11

A cross-platform acceleration framework developed for Android and Ubuntu across mobile and desktop GPUs. This work formed the engineering component of my master's thesis at HKU.

GitHub repository →

Research Platforms

Data snapshot Platform usage figures below are reported as of July 2026.

VideoCube / MGIT

2019.11 – Present

Evaluation infrastructure for the Global Instance Tracking and Multi-modal Global Instance Tracking studies published at TPAMI 2023 and NeurIPS 2023. The platform supports dataset access, tracker registration, standardized submission, and reproducible evaluation.

1.66M+ visits1.9K+ users720+ trackers960+ submissions

Visit platform →

SOTVerse / VLTVerse

2021.07 – Present

A task-space evaluation platform for the SOTVerse and VLTVerse research program, including the study published in IJCV 2024. It provides structured evaluation across tracking scenarios, target categories, and linguistic specifications.

267K+ visitsTask-space evaluation

Visit platform →

BioDrone

2022.05 – Present

Benchmark and evaluation infrastructure for real-world drone tracking research, including the BioDrone study published in IJCV 2024. It supports dataset access, standardized benchmarking, and comparative evaluation.

541K+ visitsReal-world UAV tracking

Visit platform →

GOT-10k

2020.07 – Present

Long-term maintenance of the large-scale evaluation platform supporting the GOT-10k research published in TPAMI 2021. It provides benchmark access, tracker registration, result submission, and evaluation services.

7.56M+ visits11K+ users32.7K+ trackers406K+ submissions

Visit platform →

Challenges

Hislopvision Challenge

2023.05 – 2023.11

Organization of the Hislopvision track for the 3rd High-speed and Low-power Visual Understanding Challenge at PRCV 2023, with participating teams from Tsinghua University, Beijing Institute of Technology, and Jilin University.

Challenge platform →

Cell Tracking Challenge

2021.01 – 2021.04

Our method ranked second on Fluo-C2FL-MSC+ and third on Fluo-C2FL-Huh7, based on the challenge rankings recorded in October 2023.

Challenge website →

Funded Research Projects

MOE2024-TRF-004

Singapore MOE Tertiary Education Research Fund

Core ResearcherNTU Singapore2025 – Present

MOE2022-TRF-029

Singapore MOE Tertiary Education Research Fund

Core ResearcherNTU Singapore2023 – 2025

2023 Intelligent Education PhD Research Fund

Shanghai Institute of AI Education · East China Normal University

Research Contributor2024.01 – 2025.01

Honors and Awards

  • Award 2026 Reviewer Award, the 43rd International Conference on Machine Learning (ICML 2026)
  • Award 2025 IEEE SMCS TEAM Program Award by the IEEE Systems, Man, and Cybernetics Society
  • Award 2024 Best Paper Honorable Mention in the 3rd Workshop on Vision Datasets Understanding and DataCV Challenge in CVPR 2024 (CVPRW最佳论文提名)
  • Honor 2024 Beijing Outstanding Graduates (北京市优秀毕业生, top 5%)
  • Award 2023 China National Scholarship (国家奖学金, top 1%, only 8 Ph.D. students in main campus of University of Chinese Academy of Sciences win this scholarship)
  • Award 2023 First Prize of Climbing Scholarship in Institute of Automation, Chinese Academy of Sciences (攀登一等奖学金, only 6 students in Institute of Automation, Chinese Academy of Sciences win this scholarship)
  • Honor 2022 Merit Student of University of Chinese Academy of Sciences (中国科学院大学三好学生)

  • Honor 2017 Excellent Innovative Student of Beijing Institute of Technology (北京理工大学优秀创新学生)
  • Award 2016 College Scholarship of Chinese Academy of Sciences (中国科学院大学生奖学金)
  • Honor 2016 Excellent League Member on Youth Day Competition of Beijing Institute of Technology (北京理工大学优秀团员)
  • Award 2015 National First Prize in Contemporary Undergraduate Mathematical Contest in Modeling (CUMCM) (全国大学生数学建模竞赛国家一等奖, top 1%, only 1 team in Beijing Institute of Technology win this prize) [📑PDF] [Selected and Reviewed Outstanding Papers in CUMCM (2011-2015) (Chapter 9)]
  • Award 2015 First Prize of Mathematics Modeling Competition within Beijing Institute of Technology (北京理工大学数学建模校内选拔赛第一名)
  • Honor 2015 Outstanding Individual on Summer Social Practice of Beijing Institute of Technology (北京理工大学暑期社会实践优秀个人)
  • Award 2015 Second Prize on Summer Social Practice of Beijing Institute of Technology (北京理工大学暑期社会实践二等奖, team leader)
  • Honor 2015 Outstanding Student Cadre of Beijing Institute of Technology (北京理工大学优秀学生干部)
  • Honor 2015 Outstanding League Cadre on Youth Day Competition of Beijing Institute of Technology (北京理工大学优秀团干部)
  • Honor 2015 Outstanding Youth League Branch of Beijing Institute of Technology (北京理工大学优秀团支部, team leader)
  • Honor 2015 Top-10 Activities on Youth Day Competition of Beijing Institute of Technology (北京理工大学十佳团日活动, team leader)
  • Honor 2014 Outstanding Student of Beijing Institute of Technology (北京理工大学优秀学生)
  • Award 2014, 2015, 2016, 2017 Academic Scholarship of Beijing Institute of Technology (北京理工大学学业奖学金)

Activities and Service

Tutorial

34th International Joint Conference on Artificial Intelligence (IJCAI)

  • Title: Human-Centric and Multimodal Evaluation for Explainable AI: Moving Beyond Benchmarks
  • Date & Location: 14:00-15:30, 18th August, 2025, Montreal, Canada

28th European Conference on Artificial Intelligence (ECAI)

  • Title: From Benchmarking to Trustworthy AI: Rethinking Evaluation Methods Across Vision and Complex Systems
  • Date & Location: 26th October, 2025, Bologna, Italy
    Webpage

2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC)

  • Title: The Synergy of Large Language Models and Evolutionary Optimization on Complex Networks
  • Date & Location: 5th October, 2025, Vienna, Austria

31st IEEE International Conference on Image Processing (ICIP)

  • Title: An Evaluation Perspective in Visual Object Tracking: from Task Design to Benchmark Construction and Algorithm Analysis
  • Date & Location: 9:00-12:30, 27th October, 2024, Abu Dhabi, United Arab Emirates
  • Duration: Half-day
    Slides Webpage

27th International Conference on Pattern Recognition (ICPR)

  • Title: Visual Turing Test in Visual Object Tracking: A New Vision Intelligence Evaluation Technique based on Human-Machine Comparison
  • Date & Location: 14:30-18:00, 1st December, 2024, Kolkata, India
  • Duration: Half-day
    Slides

17th Asian Conference on Computer Vision (ACCV)

  • Title: From Machine-Machine Comparison to Human-Machine Comparison: Adapting Visual Turing Test in Visual Object Tracking
  • Date & Location: 9:00-12:00, 9th December, 2024, Hanoi, Vietnam
  • Duration: Half-day
    Slides Webpage

Mini-Symposium

The Fifth International Nonlinear Dynamics Conference (NODYCON 2026)

  • Title: Complex Network Systems and Large Language Models
  • Date & Location: 20th-23rd September, 2026, Sapienza University of Rome, Italy
    Webpage

Talk

International Conference on Image and Graphics (ICIG 2026)

  • Forum: Visual Intelligence in Transition: From Physical Perception to Psychological Cognition
  • Title: From Object Tracking to Process Understanding: State Modeling in Dynamic Vision
  • Date & Location: 4th October, 2026, Singapore

Chinese Congress on Image and Graphics (CCIG 2026)

  • Title: Visual Understanding Reliability in Open Environments: from Robust Perception to Semantic Consistency
  • Date & Location: 29th May, 2026, Guangzhou, China

Publicity Chair

2026 CSIG Annual Conference on Video and Image Security

  • Date & Location: 21st November, 2026, Xiong’an, China

TPC Member

Guest Editor

Associate Editor

Reviewer

  • Conferences: NeurIPS, ICML, ICLR, CVPR, ECCV, ICCV, ACL, AAAI, IJCAI, ACM MM, ICRA, and AISTATS.
  • Journals: ACM Computing Surveys, IEEE Transactions on Image Processing, SCIENCE CHINA Information Sciences, Pattern Recognition, Transactions on Machine Learning Research, IEEE Transactions on Network Science and Engineering, IEEE Transactions on Vehicular Technology, Information Fusion, Visual Intelligence, Engineering Applications of Artificial Intelligence, Expert Systems with Applications, Neurocomputing, and Knowledge-Based Systems.

Member

  • Committee: Technical Committee on Video and Image Security (CSIG-TCVIS).
  • Societies: Institute of Electrical and Electronics Engineers (IEEE), Chinese Association for Artificial Intelligence (CAAI), and China Computer Federation (CCF).

Contact

Visitor insights cumulative homepage page views View details
Estimated visitor distribution by region
Loading the estimated regional distribution.
Exact first-party statistics Recent page views View details
page views countries and regions Singapore Time (UTC+8)
Loading recent country and region statistics.