Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents Paper • 2609.17653 • Published 6 days ago • 37 • 3
RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation Paper • 2609.16900 • Published 6 days ago • 42 • 4
Discovery Foundation Models: Toward Open-Ended Discovery Intelligence Paper • 2609.15973 • Published 7 days ago • 33 • 3
Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model Paper • 2609.18323 • Published 5 days ago • 115 • 5
JEPA-Anything: Learning Predictive Models across Different Worlds Paper • 2609.20800 • Published 4 days ago • 56 • 4
EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents Paper • 2609.17632 • Published 6 days ago • 34 • 3
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 4 days ago • 95 • 5
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 4 days ago • 133 • 4
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 7 days ago • 205 • 3
An Empirical Study of Harness Design for Coding Agents Paper • 2609.20804 • Published 4 days ago • 71 • 3
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 4 days ago • 94 • 3
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 14 days ago • 360 • 5
Studying Image Tokenizers as Visual Languages in Unified Multimodal Models Paper • 2609.09143 • Published 13 days ago • 29 • 3
Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning Paper • 2609.10445 • Published 12 days ago • 34 • 3
Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work Paper • 2609.11977 • Published 17 days ago • 114 • 4
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 7 days ago • 237 • 3
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 10 days ago • 256 • 8
ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs Paper • 2609.10895 • Published 12 days ago • 52 • 3