view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 1 day ago • 14
view article Article BenchMIRT: What are LLM benchmarks actually measuring? allenai • 3 days ago • 12
view article Article Measuring benchmark optimization in speech recognition +5 tlebryk02, bezzam, aliceebaird, dayllon, jpc, jens-hume-ai, tzirakis • 15 days ago • 65
view article Article What We Learned by Reproducing 2,200 papers from ICML abidlabs • 23 days ago • 109
view article Article State of Open Models: Summer 2026 Observations +1 AdinaY, multimodalart, irenesolaiman • 22 days ago • 186
Running 6 Transformers Model Architectures 📐 6 Browse and filter transformer model architecture diagrams
view article Article Introducing North Mini Code: Cohere’s First Model For Developers CohereLabs • Jun 9 • 85
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 Text Generation • 561B • Updated 11 days ago • 238k • • 338
view article Article Task-Seeded Synthetic Q&A Generation for Nemotron Pretraining nvidia • Jun 4 • 17