Research Interests

Research Interests

Research Trajectory

Research trajectory from visual perception tasks to human-grounded intelligence evaluation

My research began with visual object tracking and machine vision evaluation, spanning task modeling, evaluation environments, measurement techniques, and human-machine comparison. Inspired by the Turing Test, I proposed the Visual Turing Test to evaluate dynamic visual intelligence against human abilities.

The central question has remained consistent: how can AI capabilities be measured and diagnosed in open, human-centered contexts? Rather than reporting benchmark scores alone, I study what a system can perceive, understand, and reliably maintain as tasks and environments become more open.

From visual localization to capability diagnosis

Visual Object Tracking provided a concrete task for studying dynamic visual ability. Global Instance Tracking (GIT) extends tracking to long-term target retrieval, while Multi-modal GIT (MGIT) introduces hierarchical semantics and spatiotemporal-causal reasoning.

From closed benchmarks to open environments

Human visual environments are continuous, open, and semantically rich. VideoCube organizes long videos through narrative structure, SOTVerse supports user-defined task spaces, and BioDrone examines reliable perception under physical disturbance.

From performance comparison to human-grounded evaluation

By placing people and models in comparable visual tasks, I study how their capabilities differ and where conventional metrics conceal those differences. This perspective now guides my work on human-grounded evaluation and reliable human-AI collaboration.

Current Research Directions

Open-World Vision

I study whether visual systems can maintain target identity and world state under occlusion, interference, and physical disturbance. Current topics include open-world tracking, vision-language grounding, visual memory, and reliable perception for UAVs and embodied systems.

Evidence-Grounded Multimodal Reasoning

I investigate whether multimodal models select and use the right evidence in images and long videos. My work covers spatiotemporal and causal reasoning, streaming memory, adaptive visual computation, and process-level verification.

Human-Centered Agents and AI for Education

I develop agents that model cognitive and learning states, capability boundaries, and social interaction. Education provides a practical setting for virtual students, personalized agents, multi-agent simulation, and reliable human-AI collaboration.

The 3E framework connecting environment, evaluation, and executors

The 3E framework connects Environment, Evaluation, and Executors in one evaluation loop. Across these directions, I build open task spaces, human-grounded protocols, and process-level diagnostics to reveal capability boundaries and improve model and system design.