Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published Jul 14 • 235
OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories Paper • 2608.08557 • Published 22 days ago • 2
Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging Paper • 2607.10428 • Published Jul 24 • 1
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning Paper • 2606.25325 • Published Jun 24
MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy Paper • 2606.27652 • Published Jun 26
EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams Paper • 2605.07299 • Published May 8
ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories Paper • 2606.11520 • Published Jun 9
From Pixels to Words -- Towards Native One-Vision Models at Scale Paper • 2605.28820 • Published May 27 • 76
GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning Paper • 2605.16371 • Published May 10
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Paper • 2605.12500 • Published May 12 • 198
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis Paper • 2604.15093 • Published Apr 16 • 30
ReinDriveGen: Reinforcement Post-Training for Out-of-Distribution Driving Scene Generation Paper • 2604.01129 • Published Apr 1 • 8
EVA: Efficient Reinforcement Learning for End-to-End Video Agent Paper • 2603.22918 • Published Mar 24 • 44
Delving into the Devils of Bird's-eye-view Perception: A Review, Evaluation and Recipe Paper • 2209.05324 • Published Sep 12, 2022