Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads Paper • 2509.16495 • Published Jan 26
Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents Paper • 2606.22953 • Published Jun 22 • 1
When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents Paper • 2606.22936 • Published Jun 22 • 2
FastKernels: Benchmarking GPU Kernel Generation in Production Paper • 2605.23215 • Published May 22 • 9
Dual-View Training for Instruction-Following Information Retrieval Paper • 2604.18845 • Published Apr 20 • 13