Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
Xuwei Ding's picture

Xuwei Ding

Xuwei04
1 4 1
abeQ213's profile picture
·

AI & ML interests

None yet

Recent Activity

updated a dataset 13 days ago
Xuwei04/gui-real-decks-zenodo
upvoted a paper 26 days ago
Codifying the Judge: Scalable Evaluation via Program Distillation
upvoted a paper about 1 month ago
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
View all activity

Organizations

Language, Intelligence, and Model Evaluation Lab's profile picture Lexmount's profile picture

upvoted a paper 26 days ago

Codifying the Judge: Scalable Evaluation via Program Distillation

Paper • 2607.22561 • Published May 29 • 8
upvoted a paper about 1 month ago

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

Paper • 2607.08964 • Published Jul 9 • 77
upvoted a paper 3 months ago

CocoaBench: Evaluating Unified Digital Agents in the Wild

Paper • 2604.11201 • Published Apr 13 • 37
upvoted a paper 4 months ago

The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents

Paper • 2604.10577 • Published Apr 12 • 27
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs