๐Ÿ”๏ธ Research Interests

๐Ÿ”๏ธ Research Interests

Research Foundation

Research trajectory from visual perception tasks to human-grounded intelligence evaluation

My research began with visual object tracking and machine vision evaluation. I have studied task modeling, evaluation environments, measurement techniques, and human-machine comparison. Inspired by the Turing Test, I proposed the Visual Turing Test to evaluate dynamic visual intelligence against human abilities.

The central idea has remained consistent: AI evaluation should reveal what a system can perceive, understand, and reliably maintain in real environments, rather than report a benchmark score alone.


1 ยท What abilities define human perception?

I used Visual Object Tracking as a representative task for studying dynamic visual ability. Traditional tracking assumes continuous motion and short-term observation. Global Instance Tracking (GIT) extends the task to long-term target retrieval, while Multi-modal GIT (MGIT) introduces hierarchical semantics and spatiotemporal reasoning. This work moved my research from perceptual localization toward cognitive visual tasks.

2 ยท What environments do humans perceive?

Human visual environments are continuous, open, and semantically rich. I developed VideoCube to organize long videos through narrative structure, and SOTVerse as an open task space for testing visual generalization. BioDrone further examines reliable perception under motion disturbance and real-world physical constraints.

3 ยท How large is the human-machine gap?

I constructed unified evaluation settings in which people and models perform comparable visual tasks. The results show different strengths: people use semantics and context more effectively, while machines often sustain precision and persistence. These comparisons shaped my interest in human-grounded evaluation and reliable human-AI collaboration.

Current Research Directions

Open-World Vision

I study dynamic visual perception and world-state modeling under occlusion, interference, and physical disturbance. Current topics include object tracking, vision-language grounding, long-horizon visual memory, and reliable perception for UAVs and embodied systems.

Multimodal Reasoning

I investigate whether multimodal foundation models select and use the right evidence in images and long videos. My work covers spatiotemporal and causal reasoning, streaming memory, adaptive visual computation, and process-level verification.

Human-Centered Agents

I develop agents that model cognitive and learning states, capability boundaries, and social interaction. Education provides a practical setting for personalized agents, virtual students, multi-agent simulation, and reliable human-AI collaboration.

The 3E framework connecting environment, evaluation, and executors

The 3E framework connects Environment, Evaluation, and Executors in one evaluation loop. Across the three directions above, I build open task spaces, human-grounded protocols, and process-level diagnostics to identify capability boundaries and failure mechanisms, then use those findings to improve model and system design.