Screenshot of this question was making the rounds last week. But this article covers testing against all the well-known models out there.
Also includes outtakes on the ‘reasoning’ models.
Screenshot of this question was making the rounds last week. But this article covers testing against all the well-known models out there.
Also includes outtakes on the ‘reasoning’ models.
They didn’t take into account the “thinking mode” most model pass when thinking is activated
Sure they did. They even had a notation on the results table that grok passed expect when reasoning mode was off.
ETA: they even posted all the reasoning texts for the models they tested