Honesty has a price: How version 5 of the CrucibleMark benchmark rewrites the rankings
A failure in testing should count against a model. Until now, it did not always. Version 5 cleans up this quiet injustice, officially brings tool use into the overall score, and raises a question that goes far beyond numbers: What does it actually mean to be compared fairly?