How can we deploy AI in everyday life, safely and efficiently?

I'm a first-year M.S. student in Data Science at the University of Pennsylvania, after finishing my B.A. in Data Science at the University of Wisconsin–Madison. Since January 2025 I have been doing research at the LIME Lab at USC with Jieyu Zhao, and since August 2026 I have also been part of the KANG Lab at MBZUAI with Jian Kang. Also, I am fortunate to work closely with Linxin Song.

Questions behind my research

  1. Q1

    How can we measure what models and agents can really do, and explain why they fail?

  2. Q2

    When an agent acts on our behalf, how can it understand what we mean, and what its actions will cause?

  3. Q3

    How can we produce data at scale that stays verifiable, and turn it into better agents?

LLM / VLM Evaluation

Benchmarks, failure taxonomies and verifiable judges that reveal where and why models break before they reach users.

Computer-Use Agents

Agents that operate real software on our behalf: how they fail, how to keep them safe, and how to make them reliable.

Data-Centric AI

Better data as the lever: human-crafted benchmarks, verified task synthesis, and error discovery over massive knowledge bases.

Robot Use · Exploring

Extending agents from digital worlds to the physical world: using them to control robots, and what safety and evaluation mean there.

01

Education

Expected May 2028

University of Pennsylvania

M.S. in Data Science

Graduated May 2026

University of Wisconsin–Madison

B.A. in Data Science

GPA 4.0 / 4.0Dean's List
02

Experience

KANG Lab, MBZUAI

Aug 2026 – Present

Research Assistant·Advisor: Jian Kang

  • Research focus: evaluation of computer-use agents.

Beijing Lanyue Intelligence

May 2026 – Aug 2026

LLM & Agent Algorithm Intern

  • Browser-use evaluation. Benchmarked 12 frontier LLMs on web automation in the browser-use framework and built a 3-class, 9-subclass failure taxonomy (reasoning, execution, environment).
  • Active-exploration task synthesis. LLM-written web tasks often invent controls and fields. Instead, a computer-use agent explores each site and records a capability card; only cross-page transitions it actually executed become edges of an information-flow DAG, and paths over the graph compose long-horizon, cross-site tasks. 331 sites explored, 75 tasks delivered.
  • Verifiable LLM-as-a-Judge. 384 rubrics, each anchored to a verbatim span of the task and scored only against screenshots, the action trajectory and visible page fields, so every failure points to a specific step. On GPT 5.6 Sol, 77.1% of rubrics pass.
30.8 steps / task on average 57.3% task success · GPT 5.6 Sol 97.0% judge–human agreement

LIME Lab, University of Southern California

Jan 2025 – Present

Research Assistant·Advisor: Jieyu Zhao

  • Co-first author of OS-Blind, a 300-task human-authored benchmark of unintended attacks on computer-use agents, spanning 12 harm categories and 8 desktop applications.
  • Second author of SEA (COLM 2025): budgeted discovery of knowledge deficiencies in closed-weight LLMs; benchmarked SEA across 8 LLMs against ACD and AutoBencher and analysed discovered errors with topic modeling.
03

Publications

* equal contribution

The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents

Xuwei Ding*, Skylar Zhai*, Linxin Song*, Jiate Li, Taiwei Shi, Nicholas Meade, Siva Reddy, Jian Kang, Jieyu Zhao

We introduce OS-Blind, a benchmark where every user instruction is benign and the harm comes from the environment or the execution outcome. Most computer-use agents exceed 90% attack success rate, and safety alignment mostly fires in the first few steps and rarely re-engages. Multi-agent systems are even more vulnerable: task decomposition strips the context and hides the user's intent from the agents that act.

300 tasks · 8 apps Claude 4.5 Sonnet: 73.0% ASR Multi-agent system: 92.7% ASR

Featured on Hugging Face Daily Papers (Apr 15, 2026)

Used for GUI-agent safety evaluation in the UI-Venus-2 Technical Report

Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base

Linxin Song, Xuwei Ding, Jieyu Zhang, Taiwei Shi, Ryotaro Shimizu, Rahul Gupta, Yang Liu, Jian Kang, Jieyu Zhao

Stochastic Error Ascent (SEA) finds knowledge errors in closed-weight LLMs under a strict query budget by iteratively retrieving candidates semantically close to observed failures, with hierarchical retrieval and a relation DAG to prune sources.

40.7× more errors than ACD 599× / 9× lower cost per error

FAST-CAD: A Fairness-Aware Framework for Non-Contact Stroke Diagnosis

T. Sha, Z. Chen, Z. Cheng, H. Zhai, Xuwei Ding, K. Wang

A domain-adversarial training plus Group-DRO framework that jointly enforces demographic-invariant representations and worst-group robustness for non-contact stroke diagnosis, built on a 12-subgroup multimodal dataset.