DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Paper • 2607.20465 • Published May 19 • 54
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents Paper • 2607.18754 • Published 12 days ago • 24
HippoCamp: Benchmarking Contextual Agents on Personal Computers Paper • 2604.01221 • Published Apr 1 • 30