NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 3 days ago • 384
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published 16 days ago • 179
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory Paper • 2607.27919 • Published Jul 30 • 59
SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks Paper • 2606.15872 • Published Jun 14 • 13
Masking Stale Observations Helps Search Agents -- Until It Doesn't: A Regime Map and Its Mechanism Paper • 2606.00408 • Published May 29 • 66
UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering Paper • 2605.30076 • Published May 28 • 27
SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models Paper • 2605.23345 • Published May 22 • 18
From Web to Pixels: Bringing Agentic Search into Visual Perception Paper • 2605.12497 • Published May 12 • 14
Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers Paper • 2605.06169 • Published May 7 • 239
MedConclusion: A Benchmark for Biomedical Conclusion Generation from Structured Abstracts Paper • 2604.06505 • Published Apr 7 • 5
LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model Paper • 2604.20796 • Published Apr 22 • 244
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents Paper • 2604.07429 • Published Apr 8 • 123
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability Paper • 2604.06628 • Published Apr 8 • 330