ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation Paper • 2609.09076 • Published 10 days ago • 25
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 10 days ago • 94
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published 22 days ago • 157
Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization Paper • 2608.20281 • Published 29 days ago • 13
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence Paper • 2608.16590 • Published Aug 17 • 150
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Paper • 2608.13552 • Published Aug 13 • 47
SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models Paper • 2608.10538 • Published Aug 11 • 16
Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure Paper • 2608.08722 • Published Aug 9 • 8
Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval Paper • 2608.06060 • Published Aug 6 • 41
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation Paper • 2608.06374 • Published Aug 6 • 23
LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-Final-GGUF Image-Text-to-Text • 35B • Updated about 16 hours ago • 513k • 629