Read the benchmark column before the score column.
GLM-5.3 at 88.2 and GPT-6 Astra at 57.7 are on Terminal-Bench 2.1 and Terminal-Bench 4.0 respectively. The 4.0 suite is substantially harder. Sorting a table like this by raw score produces a ranking that looks authoritative and means very little, which is exactly why we name the harness in the row rather than in a footnote.