The 50/25/25 weighting is intentional. It was chosen through testing and by examining failure modes in other leaderboard designs: equal benchmark weighting would give commonsense 60% simply because three of the five tasks fall into that category, while harder near-chance tasks like ARC-C can introduce disproportionate noise for smaller models.
The main goal is to approximate balanced capability across commonsense, science, and grammar without giving a model specialized in only one domain an outsized advantage. Additionaly the range of benchmarks is to be expanded in the future, so naturally the weighting will need a revision.
That said, your point about weight sensitivity is fair. I may add a robustness check across a few predefined alternatives such as 40/30/30, equal domains, and equal benchmarks and report a separate rank range across those schemes alongside the bootstrap rank range.