The typical standard errors between pairs of models on this dataset as a function of the absolute accuracy.
The standard error of each model pair against their observed accuracy difference. Pairs below a reference line differ by more than the corresponding number of standard errors, i.e. they are statistically distinguishable at that level.