UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation Paper • 2609.12397 • Published 9 days ago • 43
EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents Paper • 2609.17632 • Published 11 days ago • 39
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction Paper • 2609.10715 • Published 17 days ago • 329
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published 23 days ago • 84
Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners Paper • 2608.19863 • Published Aug 20 • 6
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents Paper • 2608.18423 • Published Aug 19 • 21
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 265
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? Paper • 2608.10366 • Published Aug 11 • 11
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning Paper • 2608.06197 • Published Aug 6 • 47
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Paper • 2608.01851 • Published Aug 3 • 13
ChronoVision: Temporal Reasoning via Latent State Reconstruction Paper • 2608.05631 • Published Aug 6 • 40
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published Jul 31 • 114