Headlong: A Microharness for Persistent Agents
31 points by lbw1215 3 hours ago | 10 comments

MikhailTal 3 hours ago
Very fascinating, super interesting engineering. Although i do find it very funny how they just bypass a massive vulnerability, basically zero data isolation (even between good actors, let alone bad ones) with 3 sentences. Only in the llm space you can slap a massive limitation like this in the middle of the article and continue like nothing happened

> Whatever anyone tells Audel becomes part of the single experience that every other conversation draws on. In practice, Audel is bad at keeping secrets. Ask it what it’s been working on with someone else and it will often just tell you, even though we’ve asked it not to. We also haven’t studied what happens when two people give conflicting instructions. For now, we assume anything you tell Audel is shared with everyone on the team.

reply
yewenjie 3 hours ago
Are there any objective metrics/ benchmarks that people test harnesses by?

There are just so many now that it's hard to personally test them all or just trust the vibes.

reply
andyk 42 minutes ago
andy here (headlong post author). terminal bench 3 is pretty popular for comparing different harnesses using the same underlying model (it's another laude project actually). artificial analysis has an index. you can look at the model cards of popular model releases- they tend to have the most popular current benchmarks on them. w/ headlong we decided to announce it before we've benchmarked it. we mostly wanted to informally share our experiences w/ it in this initial post. we plan to do some benchmarking coming up here soon tho
reply
airocker 2 hours ago
Sub Question : IS there a real successful agent product today that uses a library for harness(like langgraph etc)? Building our own worked for us. Works with our components(postgres, events ...) and scales naturally with our system.
reply
gexla 40 minutes ago
I don't know about real successful. Since you mentioned Langchain, you could look at https://www.langchain.com/dcode which is a CLI harness build off Langchain deep agents.
reply
0xbadcafebee 21 minutes ago
> Audel designed experiments to spawn recursive shellm sub-runs to work on subproblems. Most of the experiments failed, because shellm has a safety watchdog that kills any command that stays silent for 30 seconds. Audel fought the watchdog for about 40 minutes and mostly stopped using shellm sub-runs. Results from recursive sub-runs of shellm merged back into Audel’s mind 64 times in its first two days and 12 times in the twelve days since. We’ve since revamped the watchdog, and we’ll see if we can convince Audel to give recursion another shot.

This is why "I made it think in a loop" doesn't result in significant improvement in LLM performance. It's not learning. You need RLAIF, STAR, IDPO, etc to retrain the model to learn from its mistakes. And you need a human to review it so it's not compounding mistakes. It's expensive and time-consuming. Doing it wrong leads to bad outcomes. But not doing it leads to no significant improvement.

reply
JacobAsmuth 2 hours ago
The Googlers must be vague posting about something internal.
reply
jnwatson 2 hours ago
It buries the lede. Prime Agent sounds like a very cool project.
reply
ma2kx 53 minutes ago
I'm just exhausted. So I've today now learned about four new harness:

https://github.com/exoharness/exo/

https://github.com/laude-institute/headlong

https://github.com/microsoft/agent-lightning

and now https://github.com/PrimeIntellect-ai/prime-agent

Of course the don't have exactly the same scopes but they are in general all about persistent memory and / or continous agent loops. Like I miss those times where only once a week a new js framework was promoted.

reply
russellbeattie 3 hours ago
> "Headlong is a complete agent harness with a core of less than 10K lines of Bash..."

Wow. So, be nice or I'll replace you with a very large shell script?

reply
luciana1u 12 minutes ago
[dead]
reply