13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS
13 points by ibragim_bad 3 hours ago | 4 comments

sathish316 9 minutes ago
What does it mean when Fable 5 is 1st place and Opus 5 is 3rd place, while Claude code is 7th place? Which model and effort is used for Claude in 7th place, compared to 1st and 3rd?
reply
dia80 49 minutes ago
Why test Fable high effort vs Sol medium? Especially when Sol comes out 4-5x cheaper in their tests at those effort levels.
reply
cbg0 37 minutes ago
I think you just answered your own question.

Edit: In DeepSWE Sol High scores the same as Fable High for ~1/3 of the cost.

reply
spullara 2 hours ago
They are all different problems for the different languages. I was hoping this was a benchmark that attempted to see which languages were more efficient to use with which models.
reply
rsyring 28 minutes ago
[dead]
reply