Docker Agent
137 points by saikatsg 5 hours ago | 57 comments

Olscore 41 minutes ago
Interesting seeing other people's approaches to orchestration.

I've been working on Pullboard and just open sourced it: https://github.com/pullboard-dev/pullboard

To me, orchestration is not really the core problem. The bigger challenge is handling agent coherency over long periods of time, (drift) so my solution is a living forum-like system where agents shout to one another, document work in items, and follow doctrine handed down by the developer and work against a known specification that's checkable. I have used this workflow in several large projects that MUST be correct, and it works for me. Currently cleaning it up for other users.

reply
pitched 50 seconds ago
Not saying this counts specifically but I love seeing devs who hate Jira build something like Jira to solve managing a team of agents. It’s a beautifully ironic thing to watch.
reply
rodolphoarruda 56 minutes ago
Last year (2025), I lost count of how many times I tried using an LLM to set up Docker. It was still the era of prompting, then copy-pasting the response and seeing if it worked... and it almost never did. It’s great to see this new Docker agent. For me, it would be even more appealing to see an agent and a model specialized in Docker working together, capable of proposing advanced configurations and getting them working on the first try.
reply
blakeashleyjr 4 hours ago
While I love new open source dev (especially in Go!), agent harnesses are turning into JS frameworks from yesteryear.

All the cool kids have one!

reply
yipinwong 2 hours ago
You know which one will win for sure. ones with the worst experience.

Vue and svelte are easy to use and learn/adopt. But look what won. React.js.

So whatever the ai agent with the vast user basae will win regardless.

---

I know it's a slippery slope, but given you brought up JS framwork analogy, I had to fall into it

reply
girvo 2 hours ago
I mean as someone who was there at the time, early React was easy to learn and adopt. It’s complexity came later after it had already won

Now Angular v1, man… if I ever see another digest loop error in my life, I will scream

reply
pjmlp 4 hours ago
Coupled with VSCode forks to manage them
reply
Terr_ 33 minutes ago
"Look, LLMs will free you from the tyranny of frameworks!"

"Uh, OK, so how do I get this thing working? Where are the docs?"

"Well, first you install its unique framework..."

reply
caskeycoding 4 hours ago
[flagged]
reply
verdverm 4 hours ago
Go doesn't have a good mod/plugin story for harness devs to provide to their users

I say this as a gopher who also has a custom harness written in Go, it's a real challenge compared to the TS setups, but TS plugins are also a security concern

reply
otterley 6 minutes ago
Why isn’t IPC good enough?
reply
gandreani 3 hours ago
I agree! Comes with the territory of a compiled language I suppose?

I haven't given this much thought (I also just started with golang last year) but I'm assuming the only real path here is to provide extension SDKs in Lua or other languages whose interpreters have been implemented in go.

I think that's what grafana's k6 [1] does with a JS interpreter

[1] https://github.com/grafana/k6

reply
ChaseRensberger 2 hours ago
curious your thoughts on my approach: https://docs.wingman.actor/extend/plugin-quickstart/
reply
ChaseRensberger 2 hours ago
none are as great as https://wingman.actor!

jk its trash

reply
CBLT 3 hours ago
I tried really hard to make this work for me a couple of weeks ago but it was the brittlest harness out of any that I've used. I'd come back when it's more mature.
reply
triyambakam 2 hours ago
> bittlest

Sorry, what do you mean there?

reply
shawabawa3 2 hours ago
Typo for brittlest, most brittle
reply
CBLT 2 hours ago
Correct. Updated the message to brittlest. Thank you.
reply
dgageot 3 hours ago
Would love to know what didn't work for you
reply
CBLT 2 hours ago
Sessions just broke all the time for me. The one time I dug into it, it ended up being a known issue where if codex returned over 10k characters for a turn it breaks the session. Because docker agent does not use protocols like ACP, instead trying to parse Codex's internal state in a brittle manner.

I tried not to do anything too crazy, but even use using an API key in the docker agent natively was getting me sessions that would hang and couldn't be recovered.

reply
aheritier 56 minutes ago
In your use case you wanted to control codex from docker agent thus you used our harness integration ( https://docker.github.io/docker-agent/features/harnesses/ ). It was not codex calling docker agent which should work with ACP ( https://docker.github.io/docker-agent/features/acp/ ). We could effectively look at using ACP from docker-agent to control others harnesses if it's useful.
reply
dgageot 2 hours ago
Combination with Codex is indeed not the most common use case around us. I'll see what can be improved. Sorry that it broke your workflow!
reply
nujabe 2 hours ago
How is not using it with the most popular coding agent not a common use case?
reply
girvo 2 hours ago
Not the most common use case around them: likely meaning that the Codex harness isn’t used that often internally at Docker, is what I would surmise, at least that is how I read it.

I work at another big tech company, and up until recently that would’ve been true for us as well (but it’s changed now, we partnered with OpenAI)

reply
McScrooge 4 hours ago
https://docker.github.io/docker-agent/configuration/sandbox/

If, like me, you couldn't find any security-related info on the linked page.

reply
gregwebs 2 hours ago
Docker Agent is a harness. There is a sandbox mode that can be used to run it in docker sandbox (a VM, not a container). If you don't use sandbox mode then I assume it is running in a container.

If you don't want to use their harness then you wouldn't use docker agent and instead use their `sbx` cli to run the harness of your choice (claude, codex, pi, etc).

I think it will be great if Docker can get people used to using secure VMs. I am developing a similar project (still a work in progress): https://github.com/gregwebs/agent-vm

reply
genghisjahn 2 hours ago
I've found sbx to be very helpful. really like the sentinel value wrapper they have going so you can add secrets but the model can't see them. Outbound calls get looked up by sentinel value and the real one goes out to whatever API you're auth'ing too. There's other features but just having claude run in a sandbox and easily see what it does and doesn't have access to has been great.
reply
ShinyLeftPad 3 hours ago
i remember going through the entire docs of docker sandbox and there was not one mention of attack vectors. did they fix that?
reply
kydanet 2 hours ago
[flagged]
reply
maxdo 4 hours ago
they hope they can annoy and email every person in the org asking for money , same way as they did it with docker.

looks like a weird abstractions. why do i need this if i have modal.com, e2b, cloudflare that use original docker + some toolings around + way to run + isolations on network level.

reply
beckford 4 hours ago
Docker Agent could be beneficial for research studies. For example, while writing a ML paper about agents I need to rerun experiments so that: - others can repeat my results - validate that variability from LLMs is not causing overconfidence in the results

But currently uv is good enough to create isolated, repeatable environments. And it isn't limited to Linux like standard Docker. So when I want to test small samples on my local PC's GPU (Windows or mac) before running on beefier hardware, I can do that easily.

reply
tamimio 4 hours ago
>What it is, what it isn’t

Yep, open ai sol model wrote that.

reply
lukemercado 4 hours ago
It frustrates me that there's a prisoners dilemma in communicating this to other humans. If I flag it and write it out (as you've done) I've let the author know that they're losing trust and let others know that they should be more skeptical. I've also created the perfect eval data pair for the ai labs. It's deeply frustrating.
reply
baby_souffle 3 hours ago
I think the cat is out of the bag.

There was a brief point in time at the beginning of the internet where if you saw an image online you could be kind of confident that it hadn't been manipulated significantly... But somewhere around the mid to late 2000s it became prudent to just assume any image you've seen online got at least a blemish filter pass through something like Photoshop.

Same thing now with written text. Sometimes you'll be a little bit more or less sure than an llm produced the output but never able to say 100% confidently.

reply
cyanydeez 3 hours ago
AI is going to eradicate, for good, all trust in the Internet.

That's a benified "not a drawback".

reply
IncreasePosts 4 hours ago
Your probably spare yourself some RSI by inverting this and only writing out a message when you find human generated content on hn
reply
mplewis 4 hours ago
What does this have to do with Docker?
reply
speedgoose 4 hours ago
I guess people at Docker wanted to make an agent.

Having it docker branded, I could understand. It’s confusing but the docker brand is strong.

But exposing it as a docker subcommand is very confusing to me.

reply
dgageot 3 hours ago
You don't really have to use the docker agent sub-command. docker-agent is a standalone binary. When it was created, though, 1.5 year ago, the idea was that it would be the docker compose for agents (original name was cagent, compose for agents).

But it evolved differently and although it runs very well in docker containers and docker sandboxes, it also runs very well anywhere else.

reply
binsquare 4 hours ago
It reads to me like chasing trends
reply
throwitaway222 4 hours ago
That's the economy we've always been in. X is popular, if your company doesn't have a solution for X you're irrelevant.
reply
sleepybrett 3 hours ago
that's all docker has been doing for years.
reply
verdverm 4 hours ago
the tools and features certainly indicate so
reply
redleader55 4 hours ago
The quest for staying relevant.
reply
bmacho 4 hours ago
Based on the repo short description and the first sentence of the readme:

It is called 'Docker' as the company, and it has nothing to do with the container technology.

reply
nwhnwh 2 hours ago
This made me laugh. Not sure why. But the whole situation is very absurd now.
reply
esafak 3 hours ago
It's the Docker you know and love... with AI!
reply
sneak 2 hours ago
Docker is rapidly fading into irrelevance now that they have missed their strategic acquisition interval via the Microsoft collab misstep, and now has to hype chase as fewer and fewer people believe in their long term viability as a company.

Great tech. Not a business.

reply
mococa 4 hours ago
Pretty vague
reply
smashah 2 hours ago
Why can't we set the harness? Seems like the missing piece
reply
sergiotapia 4 hours ago
very weird. why does this have dockers name on it? might as well use an agent harness from kohl's or steak n' shake. odd..
reply
dgageot 3 hours ago
It was created to be the docker compose for agents. 2 years ago, people, including ourselves, started to mix agentic loops and non agentic components in docker compose files. We created docker agent to make this more powerful.

At the end, the link with docker compose is just the yaml form. Which by the way can be replaced with hcl files or Go code.

reply
throwrioawfo 3 hours ago
gimmick
reply
esafak 3 hours ago
What's new or interesting here? Docker's AI story ought to be around sandboxing, yet the word appears nowhere. Oh, and it's YAML, which I despise.
reply
dgageot 3 hours ago
Docker Agent predates Docker sandboxes. It was created almost two years ago when not everybody had a AI harness. Nowadays, it does run very well in Docker Sandboxes but it runs also very well anywhere you like.
reply
benjy3379 3 hours ago
[flagged]
reply
oxcartctl 4 hours ago
[dead]
reply
rfgplk 4 hours ago
> Go

Nowadays there are only two justifiable public languages you should be using for anything a) C++ b) Rust

Rust itself is dubious due to abysmal compile times and the fact that the only guarantees it gives you are of security (aka a skill issue). Modern LLMs write C++ that is as safe if not safer than Rust. Any other language choice is objectively wrong.

* For those inquisitive enough the keyword "public" is doing all the heavy lifting here. Pretty much everyone should be developing and using an in-house DSL at this point.

reply
triyambakam 2 hours ago
This is an almost rage bait take, but in any case I agree that the models ability to write correct code and fuzz and test that code does make language choice a different tradeoff than it was before.
reply
furyofantares 3 hours ago
Yeah well I'm doing everything in AssemblyScript.
reply