University of Pennsylvania
M.S. in Data Science
I'm a first-year M.S. student in Data Science at the University of Pennsylvania, after finishing my B.A. in Data Science at the University of Wisconsin–Madison. Since January 2025 I have been doing research at the LIME Lab at USC with Jieyu Zhao, and since August 2026 I have also been part of the KANG Lab at MBZUAI with Jian Kang. Also, I am fortunate to work closely with Linxin Song.
Questions behind my research
How can we measure what models and agents can really do, and explain why they fail?
When an agent acts on our behalf, how can it understand what we mean, and what its actions will cause?
How can we produce data at scale that stays verifiable, and turn it into better agents?
Benchmarks, failure taxonomies and verifiable judges that reveal where and why models break before they reach users.
Agents that operate real software on our behalf: how they fail, how to keep them safe, and how to make them reliable.
Better data as the lever: human-crafted benchmarks, verified task synthesis, and error discovery over massive knowledge bases.
Extending agents from digital worlds to the physical world: using them to control robots, and what safety and evaluation mean there.
M.S. in Data Science
B.A. in Data Science
Research Assistant·Advisor: Jian Kang
LLM & Agent Algorithm Intern
Research Assistant·Advisor: Jieyu Zhao
We introduce OS-Blind, a benchmark where every user instruction is benign and the harm comes from the environment or the execution outcome. Most computer-use agents exceed 90% attack success rate, and safety alignment mostly fires in the first few steps and rarely re-engages. Multi-agent systems are even more vulnerable: task decomposition strips the context and hides the user's intent from the agents that act.
Featured on Hugging Face Daily Papers (Apr 15, 2026)
Used for GUI-agent safety evaluation in the UI-Venus-2 Technical Report
Stochastic Error Ascent (SEA) finds knowledge errors in closed-weight LLMs under a strict query budget by iteratively retrieving candidates semantically close to observed failures, with hierarchical retrieval and a relation DAG to prune sources.
A domain-adversarial training plus Group-DRO framework that jointly enforces demographic-invariant representations and worst-group robustness for non-contact stroke diagnosis, built on a 12-subgroup multimodal dataset.