๐๏ธ Research Interests
๐๏ธ Research Interests
Research Foundation

My research began with visual object tracking and machine vision evaluation. I have studied task modeling, evaluation environments, measurement techniques, and human-machine comparison. Inspired by the Turing Test, I proposed the Visual Turing Test to evaluate dynamic visual intelligence against human abilities.
The central idea has remained consistent: AI evaluation should reveal what a system can perceive, understand, and reliably maintain in real environments, rather than report a benchmark score alone.
1 ยท What abilities define human perception?
I used Visual Object Tracking as a representative task for studying dynamic visual ability. Traditional tracking assumes continuous motion and short-term observation. Global Instance Tracking (GIT) extends the task to long-term target retrieval, while Multi-modal GIT (MGIT) introduces hierarchical semantics and spatiotemporal reasoning. This work moved my research from perceptual localization toward cognitive visual tasks.
2 ยท What environments do humans perceive?
Human visual environments are continuous, open, and semantically rich. I developed VideoCube to organize long videos through narrative structure, and SOTVerse as an open task space for testing visual generalization. BioDrone further examines reliable perception under motion disturbance and real-world physical constraints.
3 ยท How large is the human-machine gap?
I constructed unified evaluation settings in which people and models perform comparable visual tasks. The results show different strengths: people use semantics and context more effectively, while machines often sustain precision and persistence. These comparisons shaped my interest in human-grounded evaluation and reliable human-AI collaboration.
Current Research Directions
Open-World Vision
I study dynamic visual perception and world-state modeling under occlusion, interference, and physical disturbance. Current topics include object tracking, vision-language grounding, long-horizon visual memory, and reliable perception for UAVs and embodied systems.
Multimodal Reasoning
I investigate whether multimodal foundation models select and use the right evidence in images and long videos. My work covers spatiotemporal and causal reasoning, streaming memory, adaptive visual computation, and process-level verification.
Human-Centered Agents
I develop agents that model cognitive and learning states, capability boundaries, and social interaction. Education provides a practical setting for personalized agents, virtual students, multi-agent simulation, and reliable human-AI collaboration.

The 3E framework connects Environment, Evaluation, and Executors in one evaluation loop. Across the three directions above, I build open task spaces, human-grounded protocols, and process-level diagnostics to identify capability boundaries and failure mechanisms, then use those findings to improve model and system design.
