S3Mem accepted to EMNLP
Structured scene–event memory for long-horizon interactive question answering.
PhD Researcher · AI for Science · Agents
I am a PhD student in Artificial Intelligence at the University of Science and Technology of China, advised by Prof. Wanli Ouyang and Prof. Houqiang Li. I am currently a research intern at Shanghai AI Laboratory.
My research focuses on AI for Science and LLM agents. I build scientific reasoning benchmarks, evaluation tools, multimodal models, and long-horizon agent systems with structured memory. My broader goal is to make AI scientifically rigorous, dependable, and useful in real research workflows.
Questions that I am especially interested in:
Previously, I spent one year in a research internship at UCLA (Aug 2024–Aug 2025), working as a Research Assistant with Prof. Debiao Li on longitudinal breast MRI and AI for Medicine. I also worked as a Researcher at BAAI (Mar–Sep 2025), where I focused on AI4Medical. Earlier research included 3D vision at CUHK/HKCLR, AI for Science at Microsoft Research Asia, and work with Prof. Alois Knoll at the Technical University of Munich.
Newest first · Scroll down for more.
Structured scene–event memory for long-horizon interactive question answering.
Dual-anchored policy distillation for more stable model alignment.
Accurate, Interdisciplinary and Transparent Structure–Property Understanding with Deep Native Structural Reasoning.
Evaluating the clinical reality gap of MLLMs with multi-sequence cardiac MRI.
Visual-to-symbolic analytical solution inference from scientific field visualizations.
A benchmark for rigorous scientific instruction following.
An open-source evaluation toolkit for scientific general intelligence.
A foundation for scientific reasoning across multiple disciplines and representations.
A data-centric review spanning scientific data foundations, models, and agent frontiers.
Our undergraduate-level multimodal physics reasoning benchmark is now available.
A human-imperceptible physical attack for near-infrared face recognition models.
Unified spatio-temporal learning across ten tasks and four scientific disciplines.
Weakly supervised post-processing for medical binary segmentation with SAM.
A concise record of my academic and research path.
Advised by Prof. Wanli Ouyang and Prof. Houqiang Li; research on LLMs and agents.
Training in machine learning, robotics, and computer vision.
Scientific intelligence, model evaluation, and agent systems.
AI for Medicine (AI4Medical), medical multimodal models, and evaluation.
One-year research internship on longitudinal breast MRI, cancer-risk prediction, and AI for medical imaging, guided by Prof. Debiao Li.
3D vision, handheld laser scanning, SDK development, and point-cloud processing.
AI for Science.
Selected work grouped by authorship role.
SciIF evaluates whether models can satisfy the explicit constraints that make scientific answers rigorous—not only produce a plausible final answer. It provides structured verification for conditions, units, assumptions, and required solution processes.
S3Mem writes long interaction histories into structured scene-event units and retrieves compact evidence chains for question answering. The system targets accurate, token-efficient memory use in long-horizon agents.
* Equal contribution.
PhysUniBench evaluates conceptual, mathematical, and diagram-based reasoning across undergraduate physics. It is designed to expose failures that simpler answer-only benchmarks miss.
BiSeg-SAM introduces a weakly supervised post-processing framework that refines SAM outputs for medical binary segmentation, reducing dependence on expensive pixel-level annotations. The work was presented as an oral paper at BIBM 2024.
SciReasoner aligns language with heterogeneous scientific representations and trains deliberate reasoning across multiple scientific capability families and disciplines.
This data-centric survey organizes scientific LLMs, datasets, benchmarks, and domain-specific representations, then traces the field toward closed-loop autonomous research agents.
UniSTD is a unified Transformer framework that learns across ten spatio-temporal tasks in four disciplines, reducing dependence on task-specific architectures and training pipelines.
This work studies native structural reasoning for explaining structure-property relationships across biology, chemistry, and materials science with greater accuracy and transparency.
The paper demonstrates a stealthy black-box physical attack on near-infrared face recognition using infrared-absorbing ink patterns that remain difficult for people to perceive.
This work introduces visual-to-symbolic analytical inference: recovering executable symbolic solutions directly from scientific field visualizations and limited metadata.
DAPD addresses the “privilege illusion” in on-policy self-distillation by aligning teacher and student behavior along two matched-information paths.
CardioLens provides a leakage-resistant, multi-sequence cardiac MRI evaluation designed to reveal where strong benchmark results fail to translate into clinical usefulness.
SciEvalKit is an extensible open-source platform for reproducible evaluation across scientific disciplines, model families, datasets, and capability categories.
Peer review and community contributions.
SIG-Cardiac Workshop on Digital Heart in the MICCAI community.
Workshop page ↗Reviewer for BIBM, MICCAI, AAAI, ACL, and related venues in artificial intelligence, medical imaging, and scientific machine learning.