Accelerates agentic visual reasoning with a lightweight speculative planner and confidence-based gating, reducing costly tool-use loops while preserving accuracy.
Bonjour 👋 I am a Ph.D. candidate in the Department of Computer Science at the University of Rochester, advised by Prof. Jiebo Luo. Previously, I earned my master's degree from Peking University (PKU) under the supervision of Prof. Li Yuan and Prof. Jie Chen, and my bachelor's degree with honors from the University of Electronic Science and Technology of China.
My research focuses on vision-language alignment for agentic AI.
@San Francisco, Jul 2026
Accelerates agentic visual reasoning with a lightweight speculative planner and confidence-based gating, reducing costly tool-use loops while preserving accuracy.
Generates time-lapse videos of evolving objects and scenes from language, learning patterns of physical transformation as a step toward metamorphic world simulation.
Builds compact shared video-language representations with expectation-maximization, reducing semantic redundancy for cross-modal retrieval.
Accelerates agentic visual reasoning with a lightweight speculative planner and confidence-based gating, reducing costly tool-use loops while preserving accuracy.
Grounds long-video understanding in retrieved speech, on-screen text, and object cues, augmenting video-language models without additional training.
Improves hateful meme detection with Chain-of-Evolution prompting, using retrieved meme pairs and contextual cues to interpret evolving visual and textual meanings.
Learns fine-grained correspondences between video frames and text tokens through cooperative game theory, supporting video-text retrieval and video question answering.
Builds compact shared video-language representations with expectation-maximization, reducing semantic redundancy for cross-modal retrieval.
Examines multimodal social-media understanding across five tasks, identifying strengths in image-text interpretation and limitations in reasoning and factual reliability.
Evaluates whether generated videos preserve reference subjects, appear natural, and follow text instructions, with fine-grained metrics and a five-million-scale supporting dataset.
Measures meaningful transformations and temporal coherence in generated time-lapse videos, complementing visual quality and text alignment with evaluation of metamorphic change.
Benchmarks who is speaking, when to join a conversation, and how to respond, diagnosing gaps between audio-visual perception and socially appropriate interaction.
Generates time-lapse videos of evolving objects and scenes from language, learning patterns of physical transformation as a step toward metamorphic world simulation.
Preserves human identity in text-conditioned video generation using frequency-aware facial features, supporting controllable visual futures without per-person fine-tuning.
“Impression, Sunrise” (1872), by Claude Monet.