Positive Alignment: Artificial Intelligence for Human Flourishing Paper • 2605.10310 • Published Jun 19 • 1
A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms Paper • 2609.04170 • Published 7 days ago • 1
A theory of appropriateness with applications to generative artificial intelligence Paper • 2412.19010 • Published Dec 26, 2024 • 2
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models Paper • 2201.11903 • Published Jan 28, 2022 • 16
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published Aug 3 • 184
EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery Paper • 2606.13662 • Published Jun 11 • 33
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Paper • 2406.01574 • Published Jun 3, 2024 • 57
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration Paper • 2605.03042 • Published May 4 • 150
Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses? Paper • 2608.04828 • Published Aug 5 • 1