Hacker News
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
16 points by
matt_d
2 hours ago |
3 comments
akshay_akula
21 minutes ago
[ - ]
Evals on actual research workflows is the right direction, most agent benches are toy tasks.
reply
rubslopes
2 hours ago
[ - ]
I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.
reply
vatsachak
16 minutes ago
[ - ]
Damn. These things aren't AGI... but I don't care.
Luna is good enough for me to give a parser spec and have it write one.
reply