Enterprise Agents and Benchmarks Collection Enterprise agent ecosystem featuring AssetOpsBench (industrial) and ITBench (SRE, FinOps, CISO), CUGA to accelerate AI Automation • 22 items • Updated 4 days ago • 20
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training Paper • 2609.04094 • Published 6 days ago • 25
view article Article Real-Time Intelligence with IBM Time Series Models on Confluent ibm-research • 6 days ago • 48
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents Paper • 2607.08093 • Published Jul 9 • 5
view article Article ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration ibm-research • Jun 30 • 26
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Paper • 2606.22388 • Published Jun 21 • 97 • 3
SenTSR-Bench: Thinking with Injected Knowledge for Time-Series Reasoning Paper • 2602.19455 • Published Feb 23 • 1
Enterprise Agents and Benchmarks Collection Enterprise agent ecosystem featuring AssetOpsBench (industrial) and ITBench (SRE, FinOps, CISO), CUGA to accelerate AI Automation • 22 items • Updated 4 days ago • 20
Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows Paper • 2605.24219 • Published May 26 • 10
Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents Paper • 2606.12674 • Published Jun 10 • 6
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Paper • 2606.19704 • Published Jun 18 • 43
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Paper • 2606.19704 • Published Jun 18 • 43
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents Paper • 2606.19704 • Published Jun 18 • 43
Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents Paper • 2606.12674 • Published Jun 10 • 6