Research Profile

About关于我

Research Fellow · Nanyang Technological University研究员 · Nanyang Technological University

Shiyu Hu (胡世宇)

My research examines how the reliability of AI capabilities can be measured, diagnosed, and modeled in open and interactive contexts. I study whether AI systems can maintain object identity, world state, task-relevant evidence, and interaction context as observations unfold over time or humans participate. My current work spans open-world vision, multimodal reasoning, and human-centered agents.

我的研究关注开放与交互情境下 AI 能力可靠性的度量、诊断与建模。我研究的是,当观测随时间展开或人进入系统之后,AI 能否持续保持对象身份、世界状态、任务相关证据和交互上下文。当前工作涵盖开放世界视觉、多模态推理与人类中心智能体。

I am a Research Fellow in the School of Physical and Mathematical Sciences (SPMS), Nanyang Technological University (NTU), working with Assoc. Prof. Kang Hao Cheong.

我现任 School of Physical and Mathematical Sciences (SPMS), Nanyang Technological University (NTU) Research Fellow,与 Assoc. Prof. Kang Hao Cheong 合作开展研究。

News动态

2026.06 📝One paper (DASTrack) has been accepted by the 2026 European Conference on Computer Vision (ECCV, CCF-B Conference).

2026.05 📣I am glad to have the opportunity to give a talk at the Chinese Congress on Image and Graphics (CCIG 2026) in Guangzhou, China. Many thanks to the special session Intelligent Evolution of Video and Image Security: Perception, Reasoning, and Adversarial Challenges. I sincerely look forward to exchanging ideas with everyone and hearing your valuable suggestions😄

2026.05 🏆I received a Reviewer Award from the 43rd International Conference on Machine Learning (ICML 2026).

2026.04 📝One paper (RGRL) has been accepted by the main conference of the 64th Annual Meeting of the Association for Computational Linguistics (ACL, CCF-A Conference).

2026.04 📝One research paper has been accepted by the 48th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC, CAAI-B Conference).

2026.04 📝One paper (COAL) has been accepted by the 35th International Joint Conference on Artificial Intelligence (IJCAI, CCF-B Conference).

2026.03 📝One paper (TBDQ) has been accepted by the Pattern Recognition (PR, CCF-B Journal).

2026.02 📝One paper (EARL) has been accepted by the 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR, CCF-A Conference).

2026.02 📝One paper (MATrack) has been accepted by the 2026 IEEE International Conference on Robotics & Automation (ICRA, CCF-B Conference).

2026.02 📝One paper (DAAWBench) has been accepted by the IEEE Transactions on Circuits and Systems for Video Technology (TCSVT, CCF-B Journal).

2026.02 📝One review paper has been accepted by the IEEE Transactions on Network Science and Engineering (TNSE).

2026.01 📣I am honored to have participated in Humanity’s Last Exam (HLE), published in Nature, by submitting expert-level questions related to AI as a member of the HLE Contributors Consortium.

2026.01 📣We presented two main-conference Oral papers (📹 CausalStep Slides and 📹 VerifyBench Slides) and three workshop papers (📹 SOEI Slides, 📹 EduVerse Slides, and 📹 EduPersona Slides) at AAAI 2026. Thanks to everyone who visited us at the Singapore EXPO for the discussions.

2026.01 📝One paper (NarrLV) has been accepted by the 14th International Conference on Learning Representations (ICLR, CCF-A Conference).

2025.12 📣We will conduct a Mini-Symposium (topic: Complex Network Systems and Large Language Models) on NODYCON 2026 (The Fifth International Nonlinear Dynamics Conference), more information will be released soon.

2025.11 📝Three papers have been accepted by the AI for Education Workshop in the 40th Annual AAAI Conference on Artificial Intelligence (AAAIW).

2025.11 📝Two papers (CausalStep and VerifyBench) have been accepted by the 40th Annual AAAI Conference on Artificial Intelligence (AAAI, CCF-A Conference, Oral).

2025.10 📣We have conducted a tutorial at 28th European Conference on Artificial Intelligence (ECAI) (26th October, 2025, Bologna, Italy).

2025.10 📣We have conducted a tutorial at 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC) (5th October, 2025, Vienna, Austria).

2025.08 📣We have conducted a tutorial at 34th International Joint Conference on Artificial Intelligence (IJCAI) (18th August, 2025, Montreal, Canada).

2025.07 🏆Obtain IEEE SMCS TEAM Program Award.

2025.06 📝One paper (ATCTrack) has been accepted by International Conference on Computer Vision (ICCV, CCF-A conference, Highlight).

2025.05 📝One review paper has been accepted by Computers and Education: Artificial Intelligence.

2025.05 📣Our new work FIOVA is now online! We introduce a multi-annotator benchmark for human-aligned video captioning, supporting semantic diversity and cognitive-aware evaluation. Check out the project page and arXiv paper for more details.

2025.05 📣Our updated work SOEI now available! Building upon our previous framework, this version introduces interactive multi-turn simulation to model open-ended educational dialogues with cognitively plausible virtual students. We further validate the framework’s effectiveness through behavioral analysis, personality recognition, and teacher-student reflection. Read more in the arXiv paper.

2025.05 📣We will present our work (SOTVerse) at IJCV2024 during the VALSE2025 poster session (June 2025, Zhuhai, China).

2025.05 📝One paper (CSTrack) has been accepted by International Conference on Machine Learning (ICML, CCF-A conference).

2025.05 📝One paper (DARTer) has been accepted by International Conference on Multimedia Retrieval (ICMR, CCF-B conference).

2025.05 📝One paper (MSAD) has been accepted by IET Computer Vision (IET-CVI, CCF-C journal).

2025.04 📣We will conduct a tutorial at 28th European Conference on Artificial Intelligence (ECAI) (25th-30th October, 2025, Bologna, Italy).

2025.03 📣We will conduct a tutorial at 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC) (5th-8th October, 2025, Vienna, Austria).

2025.02 📖The book Visual Object Tracking: An Evaluation Perspective is online.

2025.01 📝One paper (CTVLT) has been accepted by IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP, CCF-B conference).

2025.01 📣A special issue (Techniques and Applications of Multimodal Data Fusion) in Electronics has been announced, all papers related to this topic are welcomed for submission!

2024.12 📣We have conducted a tutorial at Asian Conference on Computer Vision (ACCV) (Dec. 9th 2024, Hanoi, Vietnam).

2024.12 📣We have prepared a tutorial at International Conference on Pattern Recognition (ICPR) (Dec. 1st 2024, Kolkata, India).

2024.10 📣We have conducted a tutorial at IEEE International Conference on Image Processing (ICIP) (Oct. 27th 2024, Abu Dhabi, United Arab Emirates).

2024.09 📝Two papers (MemVLT and CPDTrack) have been accepted by Conference on Neural Information Processing Systems (NeurIPS, CCF-A Conference).

2024.08 📣One tutorial proposal has been accepted by Asian Conference on Computer Vision (ACCV), the tutorial will be conducted in Dec. 2024 (Hanoi, Vietnam).

2024.08 👩‍💻Start my work as a Research Fellow in Nanyang Technological University (NTU), Singapore.

2024.07 📣One tutorial proposal has been accepted by International Conference on Pattern Recognition (ICPR), the tutorial will be conducted in Dec. 2024 (Kolkata, India).

2024.06 📝One paper has been accepted by Chinese Conference on Pattern Recognition and Computer Vision (PRCV).

2024.06 📝One paper has been accepted by Chinese Mental Health Journal (《中国心理卫生杂志》).

2024.05 🏆Obtain Beijing Outstanding Graduates (北京市优秀毕业生, top 5%).

2024.05 📣We have presented our work (Global Instance Tracking (GIT)) at TPAMI2023 during the VALSE2024 poster session (May 2024, Chongqing, China, see our 🪧 Poster for more information).

2024.04 📝One paper has been accepted by the 3rd Workshop on Vision Datasets Understanding and DataCV Challenge in CVPR 2024 (CVPRW, Oral, Best Paper Honorable Mention).

2024.04 📝One paper has been accepted by IEEE Transactions on Circuits and Systems for Video Technology (TCSVT).

2024.01 🪙One project about human-computer interaction in intelligent education has been funded by the 2023 Intelligent Education PhD Research Fund, supported by the Institute of AI Education Shanghai and East China Normal University.

2024.01 👩‍🎓Got my Ph.D. degree at Institute of Automation, Chinese Academy of Sciences (CASIA) and University of Chinese Academy of Sciences (UCAS).

2023.12 📝One paper has been accepted by the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP, CCF-B conference).

2023.11 📝One paper has been accepted by International Conference on Computer Science and Artificial Intelligence (CSAI, Oral).

2023.10 🏆Obtain China National Scholarship (国家奖学金, top 1%, only 8 Ph.D. students in main campus of UCAS win this scholarship).

2023.10 🏆Obtain First Prize of Climbing Scholarship (攀登一等奖学金, only 6 students in CASIA win this scholarship).

2023.10 📝One paper has been accepted by International Journal of Computer Vision (IJCV, CCF-A journal).

2023.09 📝One survey has been accepted by Journal of Images and Graphics (《中国图像图形学报》).

2023.09 📝One paper has been accepted by Conference on Neural Information Processing Systems (NeurIPS, CCF-A conference).

2023.09 📝One paper has been accepted by International Journal of Computer Vision (IJCV, CCF-A journal).

2023.08 📝One paper has been accepted by Chinese Conference on Pattern Recognition and Computer Vision (PRCV, CCF-C conference).

2022.06 🏆Obtain merit student of University of Chinese Academy of Sciences.

2022.06 📝One paper has been accepted by Neurocomputing (Neu).

2022.02 📝One paper has been accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI, CCF-A journal).

2021.06 📝One survey has been accepted by Journal of Graphics (《图学学报》).

Background经历

Work工作

Professional Experience工作经历

Research Intern研究实习生

Institute of Electronics, Chinese Academy of Sciences (CASIE)

Education教育

Education教育经历

Ph.D.

Institute of Automation, Chinese Academy of Sciences (CASIA)

Field:专业: Computer Application Technology

Supervisor:导师: Prof. Kaiqi Huang

Co-supervisor:共同导师: Prof. Xin Zhao

Thesis and degree details论文与学位详情

Thesis:博士论文: Research of Intelligence Evaluation Techniques for Single Object Tracking

Supervisor:导师: Prof. Kaiqi Huang (IAPR Fellow, IEEE Senior Member, 10,000 Talents Program – Leading Talents)

Co-supervisor:共同导师: Prof. Xin Zhao (IEEE Senior Member, Beijing Science Fund for Distinguished Young Scholars)

Thesis Committee:答辩委员会: Prof. Jianbin Jiao; Prof. Yuxin Peng (National Science Fund for Distinguished Young Scholars); Prof. Yao Zhao (IEEE Fellow, IET Fellow, National Science Fund for Distinguished Young Scholars); Prof. Yunhong Wang (IEEE Fellow, IAPR Fellow, CCF Fellow); and Prof. Ming Tang

Thesis Defense Grade:答辩成绩: Excellent

M.Sc. in Computer Science

University of Hong Kong (HKU)

Supervisor:导师: Prof. Choli Wang

Thesis and degree details论文与学位详情

Thesis:硕士论文: NightRunner: Deep Learning for Autonomous Driving Cars after Dark [Project]

Thesis Defense Grade:答辩成绩: A+

B.E. in the Elite Class本科 · 徐特立英才班

Beijing Institute of Technology (BIT)

School of Information and Electronics

Field:专业: Information Engineering

Thesis and degree details论文与学位详情

Undergraduate Thesis Supervisor:本科论文导师: Prof. Senlin Luo

Thesis:本科论文: Text Sentiment Analysis Based on Deep Neural Network

Thesis Defense Grade:答辩成绩: Excellent

Research Interests研究方向

Research Trajectory研究脉络

Research trajectory from visual perception tasks to human-grounded intelligence evaluation

My research began with visual object tracking and machine vision evaluation, spanning task modeling, evaluation environments, measurement techniques, and human-machine comparison. Inspired by the Turing Test, I proposed the Visual Turing Test to evaluate dynamic visual intelligence against human abilities.

The central question has remained consistent: how can AI capabilities be measured and diagnosed in open, human-centered contexts? Rather than reporting benchmark scores alone, I study what a system can perceive, understand, and reliably maintain as tasks and environments become more open.

我的研究从视觉目标跟踪与机器视觉评价出发,逐步涉及任务建模、评测环境、度量方法和人机比较。受 Turing Test 启发,我提出 Visual Turing Test,尝试以人类能力为参照评价动态视觉智能。

这条研究路径始终围绕同一个问题展开:如何在开放且以人为中心的情境中度量和诊断 AI 能力? 相比只报告 benchmark 分数,我更关心当任务和环境逐渐开放时,系统究竟感知了什么、理解了什么,又能可靠保持什么。

From visual localization to capability diagnosis从视觉定位到能力诊断

Visual Object Tracking provided a concrete task for studying dynamic visual ability. Global Instance Tracking (GIT) extends tracking to long-term target retrieval, while Multi-modal GIT (MGIT) introduces hierarchical semantics and spatiotemporal-causal reasoning.

Visual Object Tracking 为研究动态视觉能力提供了一个具体入口。Global Instance Tracking (GIT) 将跟踪扩展到长时目标检索,Multi-modal GIT (MGIT) 则进一步引入层级语义与时空因果推理。

From closed benchmarks to open environments从封闭 benchmark 到开放环境

Human visual environments are continuous, open, and semantically rich. VideoCube organizes long videos through narrative structure, SOTVerse supports user-defined task spaces, and BioDrone examines reliable perception under physical disturbance.

人的视觉环境是连续、开放且具有丰富语义的。VideoCube 通过叙事结构组织长视频,SOTVerse 支持用户定义的任务空间,BioDrone 则考察物理扰动下的可靠感知。

From performance comparison to human-grounded evaluation从性能比较到人类参照评价

By placing people and models in comparable visual tasks, I study how their capabilities differ and where conventional metrics conceal those differences. This perspective now guides my work on human-grounded evaluation and reliable human-AI collaboration.

通过让人与模型完成可比较的视觉任务,我研究二者的能力差异,以及传统指标会遮蔽哪些差异。这一视角也延伸到我目前关于人类参照评价与可靠人机协作的研究。

Current Research Directions当前研究方向

Open-World Vision开放世界视觉

I study whether visual systems can maintain target identity and world state under occlusion, interference, and physical disturbance. Current topics include open-world tracking, vision-language grounding, visual memory, and reliable perception for UAVs and embodied systems.

我研究视觉系统在遮挡、干扰和物理扰动下能否保持目标身份与世界状态,当前关注 open-world tracking、vision-language grounding、visual memory,以及 UAV 和 embodied systems 中的可靠感知。

Evidence-Grounded Multimodal Reasoning证据驱动的多模态推理

I investigate whether multimodal models select and use the right evidence in images and long videos. My work covers spatiotemporal and causal reasoning, streaming memory, adaptive visual computation, and process-level verification.

我关注多模态模型是否真正从图像和长视频中选择并使用了正确证据,研究涉及时空与因果推理、streaming memory、自适应视觉计算和过程级验证。

Human-Centered Agents and AI for Education人类中心智能体与 AI for Education

I develop agents that model cognitive and learning states, capability boundaries, and social interaction. Education provides a practical setting for virtual students, personalized agents, multi-agent simulation, and reliable human-AI collaboration.

我研究能够建模认知与学习状态、能力边界和社会互动的智能体。教育场景为 virtual students、个性化智能体、多智能体仿真和可靠人机协作提供了具体而可检验的研究环境。

The 3E framework connecting environment, evaluation, and executors

The 3E framework connects Environment, Evaluation, and Executors in one evaluation loop. Across these directions, I build open task spaces, human-grounded protocols, and process-level diagnostics to reveal capability boundaries and improve model and system design.

3E framework 将 Environment、Evaluation 与 Executors 连接在同一个评价闭环中。沿着这些方向,我通过开放任务空间、人类参照协议和过程级诊断揭示能力边界,并据此改进模型与系统设计。

Publications论文发表

Research Monograph学术专著

Springer 2025
sym

Visual Object Tracking: An Evaluation Perspective
X. Zhao, Shiyu Hu, X. Yin
Springer, Part of the book series: Advances in Computer Vision and Pattern Recognition (ACVPR)
📌 Dynamic Vision 📌 Visual Intelligence Evaluation 📌 Task-Space Diagnosis
📃 Book

Peer-Reviewed Publications同行评审论文

Lead or Corresponding Author第一作者或通讯作者

TPAMI 2023
sym

Global Instance Tracking: Locating Target More Like Humans
Shiyu Hu, X. Zhao, L. Huang, K. Huang
IEEE Transactions on Pattern Analysis and Machine Intelligence (CCF-A Journal)
📌 Open-World Tracking 📌 Global Instance Localization 📌 Human-Referenced Evaluation
📃 Paper 📑 PDF 🪧 Poster 🌐 Platform 🔧 Toolkit 💾 Dataset

IJCV 2024
sym

SOTVerse: A User-defined Task Space of Single Object Tracking
Shiyu Hu, X. Zhao, K. Huang
International Journal of Computer Vision (CCF-A Journal)
📌 Open-World Tracking 📌 Task-Space Modeling 📌 Diagnostic Evaluation
📃 Paper 📑 PDF 🪧 Poster 🌐 Platform

IJCV 2024
sym

BioDrone: A Bionic Drone-based Single Object Tracking Benchmark for Robust Vision
X. Zhao, Shiyu Hu✉️, Y. Wang, J. Zhang, Y. Hu, R. Liu, H. Lin, Y. Li, R. Li, K. Liu, J. Li
International Journal of Computer Vision (CCF-A Journal)
📌 Robust Visual Tracking 📌 Open-Environment Vision 📌 UAV Benchmark
📃 Paper 🌐 Platform 📑 PDF 🔧 Toolkit 💾 Dataset

NeurIPS 2023
sym

A Multi-modal Global Instance Tracking Benchmark (MGIT): Better Locating Target in Complex Spatio-temporal and causal Relationship
Shiyu Hu, D. Zhang, M. Wu, X. Feng, X. Li, X. Zhao, K. Huang
Conference on Neural Information Processing Systems (CCF-A Conference, Poster)
📌 Multimodal Video Understanding 📌 Global Instance Tracking 📌 Spatiotemporal-Causal Reasoning
📃 Paper 📃 PDF 🪧 Poster 📹 Slides 🌐 Platform 🔧 Toolkit 💾 Dataset

ICCV 2025
sym

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking
X. Feng*, Shiyu Hu*, X. Li, D. Zhang, M. Wu, J. Zhang, X. Chen, K. Huang (*Equal Contributions)
International Conference on Computer Vision (CCF-A Conference, Highlight)
📌 Vision-Language Tracking 📌 Dynamic State Alignment 📌 Multimodal Memory
📃 Paper 📑 PDF

ICRA 2026
sym

MATrack: Efficient Multiscale Adaptive Tracker for Real-Time Nighttime UAV Operations
X. Li*, X. Li*, Shiyu Hu✉️
International Conference on Robotics and Automation (CCF-B Conference)
📌 Robust Visual Tracking 📌 Nighttime UAV Vision 📌 Real-Time Adaptation
📃 Paper 📑 PDF

ICMR 2025
sym

DARTer: Dynamic Adaptive Representation Tracker for Nighttime UAV Tracking
X. Li*, X. Li*, Shiyu Hu✉️
International Conference on Multimedia Retrieval (CCF-B Conference)
📌 Robust Visual Tracking 📌 Nighttime UAV Vision 📌 Dynamic Representation Learning
📃 Paper 📑 PDF

中国图象图形学报 2024
sym

Visual Intelligence Evaluation Techniques for Single Object Tracking: A Survey (单目标跟踪中的视觉智能评估技术综述)
Shiyu Hu, X. Zhao, K. Huang
Journal of Images and Graphics (《中国图象图形学报》, CCF-B Chinese Journal)
📌 Dynamic Vision 📌 Visual Intelligence Evaluation 📌 Capability Diagnosis
📃 Paper 📑 PDF

IET-CVI 2025
sym

Improved SAR Aircraft Detection Algorithm Based on Visual State Space Models
Y. Wang, J. Zhang, Y. Wang, Shiyu Hu✉️, B. Shen, Z. Hou, W. Zhou
IET Computer Vision (CCF-C Journal)
📌 Remote Sensing 📌 SAR Aircraft Detection 📌 Vision State-Space Models

Collaborative Work合作论文

CVPR 2026
sym

Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
X. Li*, X. Li*, Shiyu Hu, K. Huang
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CCF-A Conference)
📌 Video-LLM Reasoning 📌 Evidence Purity 📌 Agentic Frame Selection
📃 Paper 📑 PDF

AAAI 2026
sym

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos
X. Li*, X. Li*, Shiyu Hu, K. Huang, W. Zhang
Proceedings of the AAAI Conference on Artificial Intelligence (CCF-A Conference, Oral)
📌 Video-LLM Reasoning 📌 Stepwise Causal Reasoning 📌 Diagnostic Benchmark
📃 Paper 📑 PDF 📹 Slides

AAAI 2026
sym

VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains
X. Li*, X. Li*, Shiyu Hu, Y. Guo, W. Zhang
Proceedings of the AAAI Conference on Artificial Intelligence (CCF-A Conference, Oral)
📌 LLM Reasoning 📌 Reasoning Verifiers 📌 RLVR Evaluation
📃 Paper 📑 PDF 📹 Slides

ICLR 2026
sym

NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation Models
X. Feng, H. Yu, M. Wu, Shiyu Hu, J. Chen, C. Zhu, J. Wu, X. Chu, K. Huang
International Conference on Learning Representations (CCF-A Conference)
📌 Generative Video Models 📌 Narrative Coherence 📌 Narrative-Centric Evaluation
📃 Paper 📑 PDF

ACL 2026
sym

Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
X. Li*, X. Li*, J. Gao, R. Pi, Shiyu Hu, W. Zhang
Annual Meeting of the Association for Computational Linguistics (CCF-A Conference)
📌 Multimodal Reasoning 📌 Grounded Visual Evidence 📌 Adaptive Pixel-Space Reasoning
📃 Paper 📑 PDF

PR 2026
sym

Tracking by Detection and Query: An Efficient End-to-End Framework for Multi-Object Tracking
S. Jia, Shiyu Hu, Y. Cao, F. Yang, X. Lu, X. Lu
Pattern Recognition (CCF-B Journal)
📌 Multi-Object Tracking 📌 Query-Based Association 📌 Efficient End-to-End Learning
📃 Paper 📑 PDF

TCSVT 2026
sym

Talk with Your Fingers: A Depth-Aware Benchmark for Air-Writing Recognition
M. Wu, Y. Zhao, X. Li, Shiyu Hu, Y. Cai, J. Wu, W. Wang, K. Huang
IEEE Transactions on Circuits and Systems for Video Technology (CCF-B Journal)
📌 Human-Centered Vision 📌 Depth-Aware Air-Writing 📌 Multimodal Benchmark
📃 Paper

IJCAI 2026
sym

COAL: Counterfactual and Observation-Enhanced Alignment Learning for Discriminative Referring Multi-Object Tracking
S. Jia, Shiyu Hu, Y. Wang, X. Cheng, Y. Cao, X. Lu
International Joint Conference on Artificial Intelligence (CCF-B Conference)
📌 Referring Multi-Object Tracking 📌 Robust Language Grounding 📌 Counterfactual Alignment Learning

ECCV 2026
Complete DASTrack figure

DASTrack: Rethinking Temporal Modeling in Visual Object Tracking via Decoupled Auxiliary Supervision
D. Zhang, Shiyu Hu, H. Fu, X. Feng, Y. Wang, KH Cheong, K. Huang
European Conference on Computer Vision (CCF-B Conference)
📌 Visual Object Tracking 📌 Temporal Representation Learning 📌 Decoupled Auxiliary Supervision

EMBC 2026
sym

Global-Local Semi-Supervised Modeling for Retinal Layer Boundary Estimation in OCT
K. Li, B. Parikh, H. Yue, Shiyu Hu, S. W. Tan, W. Y. Low, X. Su, KH Cheong
Annual International Conference of the IEEE Engineering in Medicine and Biology Society (CAAI-B Conference)
📌 Medical Imaging 📌 Retinal Boundary Estimation 📌 Semi-Supervised Structure Learning

TNSE 2026
sym

Constraint-Driven Evolution of Multimodal Video Intelligence: A Network and System Perspective
X. Li*, X. Li*, Shiyu Hu, Z. Zhang, KH Cheong
IEEE Transactions on Network Science and Engineering
📌 Multimodal Video Intelligence 📌 Constraint-Aware AI 📌 System-Level Reliability
📃 Paper

Mathematics 2026
sym

CalcTutor: Multi-Agent LLM Grading of Handwritten Mathematics with RAG-Grounded Feedback for Adaptive Learning Support
L. Tan, B. Zhu, Shiyu Hu, A. Mishra, Darren J. Yeo, KH Cheong
Mathematics
📌 AI for Education 📌 Multi-Agent Assessment 📌 RAG-Grounded Feedback
📃 Paper

Nature 2026

A benchmark of expert-level academic questions to assess AI capabilities
Center for AI Safety, Scale AI, and HLE Contributors Consortium
Nature, 649, 1139–1146 (2026)
Contribution: Shiyu Hu submitted expert-level questions related to AI to the HLE benchmark as a member of the HLE Contributors Consortium.
📌 Frontier AI Evaluation 📌 Expert-Level Benchmarking 📌 Consortium Contribution
📃 Paper 🌐 Benchmark

ICML 2025
sym

CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features
X. Feng, D. Zhang, Shiyu Hu, X. Li, M. Wu, J. Zhang, X. Chen, K. Huang
International Conference on Machine Learning (CCF-A Conference, Poster)
📌 Multimodal Tracking 📌 Spatiotemporal Representation Learning 📌 Efficient Sensor Fusion
📃 Paper 📑 PDF

ICASSP 2025
sym

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
X. Feng, D. Zhang, Shiyu Hu, X. Li, M. Wu, J. Zhang, X. Chen, K. Huang
IEEE International Conference on Acoustics, Speech, and Signal Processing (CCF-B Conference, Poster)
📌 Vision-Language Tracking 📌 Visual Grounding 📌 Foundation Model Transfer
📃 Paper 📃 PDF

C&E:AI 2025
sym

Artificial Intelligence-Enabled Adaptive Learning Platforms: A Review
L. Tan, Shiyu Hu, Darren J. Yeo, KH Cheong
Computers & Education: Artificial Intelligence
📌 AI for Education 📌 Personalized Learning 📌 Adaptive Learning Systems
📃 Paper 📑 PDF

Mathematics 2025
sym

A Comprehensive Review on Automated Grading Systems in STEM Using AI Techniques
L. Tan, Shiyu Hu, Darren J. Yeo, KH Cheong
Mathematics
📌 AI for Education 📌 Automated STEM Assessment 📌 Learning Analytics
📃 Paper

Innovation and Emerging Technologies 2025
sym

Trustworthy AI in education: Framework, cases, and governance strategies
Y. Ma, X. Li, Shiyu Hu, S. Liu, KH Cheong
Innovation and Emerging Technologies
📌 Trustworthy AI 📌 AI for Education 📌 Governance and Fairness
📃 Paper

中国心理卫生杂志 2025
sym

A Review of Intelligent Psychological Assessment Based on Interactive Environment (基于交互环境的智能化心理测评)
K. Huang, Y. Kang, C. Yan, Shiyu Hu, L. Wang, T. Tao, W. Gao
Chinese Mental Health Journal (《中国心理卫生杂志》, CSSCI Journal, Top Psychological Journal in China)
📌 Human-Centered AI 📌 Interactive Psychological Assessment 📌 Validity and Ethics

NeurIPS 2024
sym

Beyond Accuracy: Tracking more like Human via Visual Search
D. Zhang, Shiyu Hu, X. Feng, X. Li, M. Wu, J. Zhang, K. Huang
Conference on Neural Information Processing Systems (CCF-A Conference, Poster)
📌 Human-Centered Vision 📌 Visual Search 📌 Process-Level Evaluation
📃 Paper 📑 PDF

NeurIPS 2024
sym

MemVLT: Vision-Language Tracking with Adaptive Memory-based Prompts
X. Feng, X. Li, Shiyu Hu, D. Zhang, M. Wu, J. Zhang, X. Chen, K. Huang
Conference on Neural Information Processing Systems (CCF-A Conference, Poster)
📌 Vision-Language Tracking 📌 Long-Term Memory 📌 Adaptive Prompting
📃 Paper 📑 PDF

ICASSP 2024
sym

Robust Single-particle Cryo-EM Image Denoising and Restoration
J. Zhang, T. Zhao, Shiyu Hu, X. Zhao
IEEE International Conference on Acoustics, Speech, and Signal Processing (CCF-B Conference, Poster)
📌 AI for Science 📌 Cryo-EM Restoration 📌 Diffusion Models
📃 Paper 📑 PDF

TCSVT 2024
sym

Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-World
M. Wu, K. Huang, Y. Cai, Shiyu Hu, Y. Zhao, W. Wang
IEEE Transactions on Circuits and Systems for Video Technology (CCF-B Journal)
📌 Human-Computer Interaction 📌 Unconstrained Air-Writing 📌 Real-World Benchmark
📃 Paper 📃 PDF 🔧 Toolkit

PRCV 2024
sym

VS-LLM: Visual-Semantic Depression Assessment based on LLM for Drawing Projection Test
M. Wu, Y. Kang, X. Li, Shiyu Hu, X. Chen, Y. kang, W. Wang, K. Huang
Chinese Conference on Pattern Recognition and Computer Vision (CCF-C Conference)
📌 Human-Centered AI 📌 Multimodal Mental Health Assessment 📌 LLM-Assisted Interpretation
📃 Paper 📃 PDF

PRCV 2023
sym

A Hierarchical Theme Recognition Model for Sandplay Therapy
X. Feng, Shiyu Hu, X. Chen, K. Huang
Chinese Conference on Pattern Recognition and Computer Vision (CCF-C Conference, Poster)
📌 Human-Centered AI 📌 Computational Mental Health 📌 Knowledge-Guided Recognition
📃 Paper 📑 PDF 🔖 Supplementary 🪧 Poster

CSAI 2023
sym

Rethinking Similar Object Interference in Single Object Tracking
Y. Wang, Shiyu Hu, X. Zhao
International Conference on Computer Science and Artificial Intelligence (EI Conference, Oral)
📌 Robust Visual Tracking 📌 Similar-Object Interference 📌 Failure Diagnosis
📃 Paper 🗒 bibTex 📑 PDF

Neurocomputing 2022
sym

Revisiting Instance Search: A New Benchmark Using Cycle Self-training
Y. Zhang, C. Liu, W. Chen, X. Xu, F. Wang, H. Li, Shiyu Hu, X. Zhao
Neurocomputing (CCF-C Journal)
📌 Open-World Retrieval 📌 Cross-Camera Instance Search 📌 Cycle Self-Training
📃 Paper 📑 PDF 🌐 Project

图学学报 2021
sym

Visual Turing: The Next Development of Computer Vision in The View of Human-computer Gaming (视觉图灵:从人机对抗看计算机视觉下一步发展)
K. Huang, X. Zhao, Q. Li, Shiyu Hu
Journal of Graphics (《图学学报》, CCF-C Chinese Journal)
📌 Human-Centered AI 📌 Visual Intelligence Evaluation 📌 Human-Machine Benchmarking
📃 Paper 📑 PDF

WorkshopWorkshop 论文

Selected collaborative publications are shown for concise browsing.默认展示部分合作论文,完整列表可按需展开。

Preprints预印本

Preprint
Complete figure for the survey of streaming video understanding

How to Respond, How to Memorize, How to Be Fast: A Survey of Streaming Video Understanding
X. Li*, Shiyu Hu*, X. Feng, J. Zhao, K. Huang (*Equal Contributions)
📌 Streaming Video Understanding 📌 Long-Term Memory 📌 Efficient Inference
📃 Paper

Preprint
sym

FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
Shiyu Hu*, X. Li*, X. Li, J. Zhang, Y. Wang, X. Zhao, KH Cheong (*Equal Contributions)
📌 Large Vision-Language Models 📌 Human-Aligned Video Understanding 📌 Multi-Annotator Evaluation
📃 Paper 📑 PDF 🌐 Project

Preprint
sym

When LLMs Learn to be Students: The SOEI Framework for Modeling and Evaluating Virtual Student Agents in Educational Interaction
Y. Ma*, Shiyu Hu*, X. Li, Y. Wang, Y. Chen, S. Liu, KH Cheong (*Equal Contributions)
📌 AI for Education 📌 Virtual Student Agents 📌 Interaction-Centered Evaluation
📃 Paper 📑 PDF

Preprint
sym

EduVerse: A User-Defined Multi-Agent Simulation Space for Education Scenario
Y. Ma*, Shiyu Hu*, B. Zhu, Y. Wang, Y. Kang, S. Liu, KH Cheong (*Equal Contributions)
📌 AI for Education 📌 Multi-Agent Simulation 📌 User-Defined Classroom Space
📃 Paper 📑 PDF

Preprint
sym

EduPersona: Benchmarking Subjective Ability Boundaries of Virtual Student Agents
B. Zhu*, Shiyu Hu*, Y. Ma, Y. Zhang, KH Cheong (*Equal Contributions)
📌 AI for Education 📌 Persona-Aware Agents 📌 Subjective Ability Diagnosis
📃 Paper 📑 PDF

Preprint
sym

SOI is the Root of All Evil: Quantifying and Breaking Similar Object Interference in Single Object Tracking
Y. Wang*, Shiyu Hu*, S. Jia, P. Xu, H. Ma, Y. Ma, J. Zhang, X. Lu, X. Zhao (*Equal Contributions)
📌 Robust Visual Tracking 📌 Similar-Object Interference 📌 VLM-Guided Correction
📃 Paper 📑 PDF

Preprint
sym

How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking
X. Li*, Shiyu Hu*, X. Feng, D. Zhang, M. Wu, J. Zhang, K. Huang (*Equal Contributions)
📌 Vision-Language Tracking 📌 Language Utility Diagnosis 📌 Fine-Grained Evaluation
📃 Paper 📑 PDF

Preprint
Complete STEMVerse diagnostic framework for STEM reasoning

STEMVerse: A Dual-Axis Diagnostic Framework for STEM Reasoning in Large Language Models
X. Li, X. Li, J. Zhao, Shiyu Hu✉️
📌 AI for Education 📌 STEM Reasoning 📌 Cognitive Diagnosis
📃 Paper 📑 PDF

Preprint
sym

DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
X. Li, Shiyu Hu, X. Feng, D. Zhang, M. Wu, J. Zhang, K. Huang
📌 Vision-Language Tracking 📌 Data-Centric AI 📌 Diverse Text Benchmark
📃 Paper 📑 PDF 🌐 Project

Preprint
sym

Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
X. Li, Shiyu Hu, X. Feng, D. Zhang, M. Wu, J. Zhang, K. Huang
📌 Vision-Language Tracking 📌 Multimodal Interaction 📌 Interactive Robustness Evaluation
📃 Paper 📑 PDF 🌐 Project

Preprint
sym

Beyond Accuracy: Evaluating Grounded Visual Evidence in Thinking with Images
X. Li*, X. Li*, R. Pi, Shiyu Hu, J. Zhao,J. Gao
📌 Agentic Visual Reasoning 📌 Grounded Visual Evidence 📌 Dual-Axis Diagnosis
📃 Paper 📑 PDF

Preprint
sym

Nearing or Surpassing: Overall Evaluation of Human-Machine Dynamic Vision Ability
Shiyu Hu, X. Zhao, Y. Wang, Y. Shan, K. Huang
📌 Human-Centered AI 📌 Dynamic Vision Capability 📌 Human-Machine Evaluation
📑 PDF

Selected preprints are shown for concise browsing.默认展示部分预印本,完整列表可按需展开。

Projects研究项目

This section documents research software, evaluation platforms, academic challenges, and funded projects. Associated research outputs are listed under Publications.

这里汇总研究软件、评测平台、学术竞赛与资助项目,相关研究成果统一列在 Publications 中。

Research Software研究软件

Darknet-Cross

Lightweight Deep Learning Framework for Heterogeneous Computing

2018.03 – 2018.11

A cross-platform acceleration framework developed for Android and Ubuntu across mobile and desktop GPUs. This work formed the engineering component of my master's thesis at HKU.

面向 Android 与 Ubuntu、覆盖移动端和桌面端 GPU 的跨平台加速框架。该工作构成我在 HKU 硕士论文的工程部分。

GitHub repository →GitHub 仓库 →

Research Platforms研究平台

Data snapshot数据快照 Platform usage figures below are reported as of July 2026.以下平台使用数据统计截至 2026 年 7 月。

VideoCube / MGIT

2019.11 – Present

Evaluation infrastructure for the Global Instance Tracking and Multi-modal Global Instance Tracking studies published at TPAMI 2023 and NeurIPS 2023. The platform supports dataset access, tracker registration, standardized submission, and reproducible evaluation.

服务于 TPAMI 2023 的 Global Instance Tracking 与 NeurIPS 2023 的 Multi-modal Global Instance Tracking 研究,支持数据集获取、tracker 注册、标准化提交和可复现评测。

1.66M+ visits1.9K+ users720+ trackers960+ submissions

Visit platform →访问平台 →

SOTVerse / VLTVerse

2021.07 – Present

A task-space evaluation platform for the SOTVerse and VLTVerse research program, including the study published in IJCV 2024. It provides structured evaluation across tracking scenarios, target categories, and linguistic specifications.

面向 SOTVerse 与 VLTVerse 研究的任务空间评测平台,包括发表于 IJCV 2024 的相关工作,可围绕跟踪场景、目标类别和语言描述开展结构化评价。

267K+ visitsTask-space evaluation

Visit platform →访问平台 →

BioDrone

2022.05 – Present

Benchmark and evaluation infrastructure for real-world drone tracking research, including the BioDrone study published in IJCV 2024. It supports dataset access, standardized benchmarking, and comparative evaluation.

面向真实无人机跟踪研究的 benchmark 与评测基础设施,包括发表于 IJCV 2024 的 BioDrone,支持数据集获取、标准化测试和比较评价。

541K+ visitsReal-world UAV tracking

Visit platform →访问平台 →

GOT-10k

2020.07 – Present

Long-term maintenance of the large-scale evaluation platform supporting the GOT-10k research published in TPAMI 2021. It provides benchmark access, tracker registration, result submission, and evaluation services.

长期维护服务于 TPAMI 2021 GOT-10k 研究的大规模评测平台,提供 benchmark 获取、tracker 注册、结果提交和在线评价服务。

7.56M+ visits11K+ users32.7K+ trackers406K+ submissions

Visit platform →访问平台 →

Challenges学术竞赛

Hislopvision Challenge

2023.05 – 2023.11

Organization of the Hislopvision track for the 3rd High-speed and Low-power Visual Understanding Challenge at PRCV 2023, with participating teams from Tsinghua University, Beijing Institute of Technology, and Jilin University.

组织 PRCV 2023 第三届高速低功耗视觉理解竞赛 Hislopvision 赛道,参赛团队来自 Tsinghua University、Beijing Institute of Technology 与 Jilin University。

Challenge platform →竞赛平台 →

Cell Tracking Challenge

2021.01 – 2021.04

Our method ranked second on Fluo-C2FL-MSC+ and third on Fluo-C2FL-Huh7, based on the challenge rankings recorded in October 2023.

截至 2023 年 10 月的竞赛榜单,我们的方法在 Fluo-C2FL-MSC+ 上排名第二,在 Fluo-C2FL-Huh7 上排名第三。

Challenge website →竞赛网站 →

Funded Research Project资助项目

Human-Computer Interaction in Intelligent Education

2024.01 – 2025.01

Proposal development and project delivery for a study funded by the 2023 Intelligent Education PhD Research Fund at the Shanghai Institute of AI Education, East China Normal University.

负责项目申请与实施,项目获 Shanghai Institute of AI Education, East China Normal University 2023 Intelligent Education PhD Research Fund 资助。

Honors and Awards奖励与荣誉

  • Award 2026 Reviewer Award, the 43rd International Conference on Machine Learning (ICML 2026)
  • Award 2025 IEEE SMCS TEAM Program Award by the IEEE Systems, Man, and Cybernetics Society
  • Award 2024 Best Paper Honorable Mention in the 3rd Workshop on Vision Datasets Understanding and DataCV Challenge in CVPR 2024 (CVPRW最佳论文提名)
  • Honor 2024 Beijing Outstanding Graduates (北京市优秀毕业生, top 5%)
  • Award 2023 China National Scholarship (国家奖学金, top 1%, only 8 Ph.D. students in main campus of University of Chinese Academy of Sciences win this scholarship)
  • Award 2023 First Prize of Climbing Scholarship in Institute of Automation, Chinese Academy of Sciences (攀登一等奖学金, only 6 students in Institute of Automation, Chinese Academy of Sciences win this scholarship)
  • Honor 2022 Merit Student of University of Chinese Academy of Sciences (中国科学院大学三好学生)

  • Honor 2017 Excellent Innovative Student of Beijing Institute of Technology (北京理工大学优秀创新学生)
  • Award 2016 College Scholarship of Chinese Academy of Sciences (中国科学院大学生奖学金)
  • Honor 2016 Excellent League Member on Youth Day Competition of Beijing Institute of Technology (北京理工大学优秀团员)
  • Award 2015 National First Prize in Contemporary Undergraduate Mathematical Contest in Modeling (CUMCM) (全国大学生数学建模竞赛国家一等奖, top 1%, only 1 team in Beijing Institute of Technology win this prize) [📑PDF] [📖Selected and Reviewed Outstanding Papers in CUMCM (2011-2015) (Chapter 9)]
  • Award 2015 First Prize of Mathematics Modeling Competition within Beijing Institute of Technology (北京理工大学数学建模校内选拔赛第一名)
  • Honor 2015 Outstanding Individual on Summer Social Practice of Beijing Institute of Technology (北京理工大学暑期社会实践优秀个人)
  • Award 2015 Second Prize on Summer Social Practice of Beijing Institute of Technology (北京理工大学暑期社会实践二等奖, team leader)
  • Honor 2015 Outstanding Student Cadre of Beijing Institute of Technology (北京理工大学优秀学生干部)
  • Honor 2015 Outstanding League Cadre on Youth Day Competition of Beijing Institute of Technology (北京理工大学优秀团干部)
  • Honor 2015 Outstanding Youth League Branch of Beijing Institute of Technology (北京理工大学优秀团支部, team leader)
  • Honor 2015 Top-10 Activities on Youth Day Competition of Beijing Institute of Technology (北京理工大学十佳团日活动, team leader)
  • Honor 2014 Outstanding Student of Beijing Institute of Technology (北京理工大学优秀学生)
  • Award 2014, 2015, 2016, 2017 Academic Scholarship of Beijing Institute of Technology (北京理工大学学业奖学金)

Activities and Service学术活动与服务

Tutorial教程报告

34th International Joint Conference on Artificial Intelligence (IJCAI)

  • Title: Human-Centric and Multimodal Evaluation for Explainable AI: Moving Beyond Benchmarks
  • Date & Location: 14:00-15:30, 18th August, 2025, Montreal, Canada

28th European Conference on Artificial Intelligence (ECAI)

  • Title: From Benchmarking to Trustworthy AI: Rethinking Evaluation Methods Across Vision and Complex Systems
  • Date & Location: 26th October, 2025, Bologna, Italy
    🌐 Webpage

2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC)

  • Title: The Synergy of Large Language Models and Evolutionary Optimization on Complex Networks
  • Date & Location: 5th October, 2025, Vienna, Austria

31st IEEE International Conference on Image Processing (ICIP)

  • Title: An Evaluation Perspective in Visual Object Tracking: from Task Design to Benchmark Construction and Algorithm Analysis
  • Date & Location: 9:00-12:30, 27th October, 2024, Abu Dhabi, United Arab Emirates
  • Duration: Half-day
    📹 Slides 🌐 Webpage

27th International Conference on Pattern Recognition (ICPR)

  • Title: Visual Turing Test in Visual Object Tracking: A New Vision Intelligence Evaluation Technique based on Human-Machine Comparison
  • Date & Location: 14:30-18:00, 1st December, 2024, Kolkata, India
  • Duration: Half-day
    📹 Slides

17th Asian Conference on Computer Vision (ACCV)

  • Title: From Machine-Machine Comparison to Human-Machine Comparison: Adapting Visual Turing Test in Visual Object Tracking
  • Date & Location: 9:00-12:00, 9th December, 2024, Hanoi, Vietnam
  • Duration: Half-day
    📹 Slides 🌐 Webpage

Mini-Symposium专题研讨会

The Fifth International Nonlinear Dynamics Conference (NODYCON 2026)

  • Title: Complex Network Systems and Large Language Models
  • Date & Location: 20th-23rd September, 2026, Sapienza University of Rome, Italy
    🌐 Webpage

Talk学术报告

Chinese Congress on Image and Graphics (CCIG 2026)

  • Title: Visual Understanding Reliability in Open Environments: from Robust Perception to Semantic Consistency
  • Date & Location: 29th May, 2026, Guangzhou, China

TPC Member程序委员会委员

Guest Editor客座编辑

Associate Editor副编辑

Reviewer审稿服务

  • Conferences: NeurIPS, ICML, ICLR, CVPR, ECCV, ICCV, ACL, AAAI, IJCAI, ACM MM, ICRA, and AISTATS.
  • Journals: ACM Computing Surveys, IEEE Transactions on Image Processing, SCIENCE CHINA Information Sciences, Pattern Recognition, Transactions on Machine Learning Research, IEEE Transactions on Network Science and Engineering, IEEE Transactions on Vehicular Technology, Information Fusion, Visual Intelligence, Engineering Applications of Artificial Intelligence, Expert Systems with Applications, Neurocomputing, and Knowledge-Based Systems.

Member学会会员

  • Societies: Institute of Electrical and Electronics Engineers (IEEE), China Society of Image and Graphics (CSIG), Chinese Association for Artificial Intelligence (CAAI), and China Computer Federation (CCF).

Contact联系方式

Visitor insights访问概览 cumulative homepage page views主页累计访问量
Estimated visitor distribution by region访客地区估算分布
Loading the estimated regional distribution.正在加载地区分布。
Exact first-party statistics近期精确统计 Recent page views近期访问 View details
page views访问量 countries and regions国家和地区 Singapore Time (UTC+8)
Loading recent country and region statistics.正在加载近期国家和地区统计。