My Agent Setup
77 points by carimura 5 hours ago | 39 comments

whazor 4 hours ago
I am now on the MCP route, where I have lots of personally hosted MCPs. Including managing calender/e-mail.

It does not cost too much effort to maintain MCP servers. No port forwarding or VPNs required thanks to OpenAI tunnels.

And security wise its quite nice, since you have to activate MCP or give permission sometimes. So each chat is kind of isolated from each-other.

reply
carimura 4 hours ago
interesting.. does that bloat the context window? I imagine Hermes already struggles with context window size.. been meaning to look into that.
reply
Atotalnoob 3 hours ago
Skills and MCPs CAN bloat the context window. Some harnesses do more progressive loading of skills/mcp depending on how many you have. Check your harnesses docs.

Personally, I feel skills+CLI is better, since only the description of the skill enters the context window until it’s required and CLIs should support help flags which will allow you do have progressive discovery.

And a CLI is also easy to use yourself.

reply
trollbridge 3 hours ago
Couple of notes on this

First, it matters how well the MCP is written; a well-written one won't be so massive

Secondly, different sessions should load different MCPs. Load only what is needed

Thirdly, if you're finding you need a dozen MCPs, you need to consolidate that into a single service that does the work and then presents a single, unified API (via MCP) to your agent

Finally, start using models with 1M context window limits. I really don't know how people can stand being stuck at a limit like 272K.

reply
carimura 19 minutes ago
ya all makes sense. hoping the harness can manage some of the MCP noise. I'm using 5.6-Sol so context window is 1M.
reply
whazor 4 hours ago
I don't have problems on ChatGPT with the amount of MCPs (five) I have. But each MCP can do quite a lot.
reply
richardkam0511 3 hours ago
I actually went down a different path and set it up so my friend and I can use the same agent with approval-gated turns. He's more technical than I am, but I believe I'm better at marketing. We can see both of our prompts and be on the same page. There is also a mode where we can branch off with our own agents. Watching his prompts made me better at producing my own and using the agent effectively.
reply
carimura 3 hours ago
That's neat. You could do this with Buzz by just having open room conversations with agents.
reply
orangebread 4 hours ago
Very cool setup. I went down a similar path and while it is definitely cool to see the extent of how agentic AI can be, I agree the ROI isn't _quite_ there yet.

I don't live in a world where I need to be constantly reading and replying to emails, so 95% of my inbox is just subscription spam. With development I need to be at the helm to design the planning requirements and actively make decisions before the agent goes off and executes the plan. But I don't need a personal agent for that, I work directly out of codex/claude-code.

reply
carimura 4 hours ago
same. I'd like to get to a place where I can have long-running tasks and goals for more over-night working, but it's not there yet.
reply
lw18511811620 2 hours ago
[flagged]
reply
lukasco 3 hours ago
I've been building a triage agent for my inbox and whatsapp (it's product shaped), which has ironically left me not building one of these. So even while productizing, I'm getting fomo on the full monty.

I've also been building a harness that maintains my apps which I'm hoping to open source.

Hard agree that these things don't have personal ROI, and are actually quite hard to build reliably.

But it's really fun! And having a bot fix a live error is pretty exciting.

reply
avereveard 2 hours ago
tried similar setups in the past, but messages is just not my jam, so built a html wrapper on top of claude code and open code that runs on a small-ish instance. looks like this https://i.imgur.com/lj9Fgco.png (screenshot anonymized with chatgpt) and allows multiple conversation per project, and has a few conveniences like scheduled tasks. one of the project in the list is the project itself, so I can add features whenever.
reply
tln 2 hours ago
That looks really useful

You haven't put this on github have you?

reply
aliasxneo 3 hours ago
Have you thought about browser/machine use at all? I like the idea of Grok Bot but would rather host/maintain it myself with my own agents.
reply
carimura 3 hours ago
Not yet. The agents do have access to the web but not really a proper browser. Browserbase or equivalent is next on my list.
reply
moribvndvs 4 hours ago
Not pointed at the author, but at the current state of affairs: this is fucking exhausting. We went and made a trillions-dollar market out of the bikeshedding maximization machine.
reply
carimura 4 hours ago
this author agrees. hence my attempt at steering part of that trillion dollars for my non-profit which maybe someday might do some good.
reply
morkalork 59 minutes ago
With Claude not only can you shave more yaks, you can build and run a yak farm and specialized yak-trimmer factory!
reply
Manfrednotfunny 4 hours ago
No we did not.

LLM was invented and its just clear that this is something someone needs to build.

Why?

Because it makes just sense. You don't want an agent running on a laptop you close. You want to keep context small, you want to split up work / parallize it etc.

I'm now waiting for a while until the open source agent platform emerges and im borderline motivated to build something but i'm not doing it. He did, which is not a crime.

reply
owencmcgrath 3 hours ago
The most interesting part of this to me is having an agent that reads logs and spins up fixes in real time.

Great write up!

reply
4lx87 3 hours ago
Have you experimented with using a single executive agent (instead of multiple agents)?
reply
carimura 3 hours ago
that was one of my FAQs at the bottom. I want separation of duties and least-privilege so agents can only access what they need for their particular duties.

That said, my vision is to eventually build manager agents to manage the minion agents. Like real people. I have no idea if "people" is the right analogy for all this work but my brain can't really wrap around a different analogy yet.

reply
tosh 4 hours ago
ty for the writeup!

being able to talk to each of the agents via dm (but also in group chats) sounds interesting

does that mean that you have 1 chat per domain specific agent? can you also start multiple sessions/threads or is that not part of the way you interact with them currently?

reply
carimura 4 hours ago
it's just like Slack really. I have a DM open for each agent where I can talk direct. But they are also part of individual rooms also where I just need to @ them and they open a thread. They can talk to each other also by @'ing each other (in public rooms or ones they both belong to)

It's literally just like humans.. except. not.

reply
tosh 4 hours ago
makes sense, ty
reply
gchp 3 hours ago
Cool! How much does this setup cost you, roughly?
reply
carimura 3 hours ago
$48/month for the droplet (could probably be cheaper on Hertzner) and $100/month for OpenAI plan. so ~$150/month but again this could be cheaper with a different VPS and using Terra/Luna.

I also have a claude max plan for coding but like I mentioned in the article coding is still separate from the agents.

reply
messh 3 hours ago
A box with 2vcpu, 4gb ram and 50gb hdd on https://shellbox.dev is $0.02/hr and you pay per minute only for what you use. Even using it 24/7 is like less than $15
reply
carimura 22 minutes ago
ya i'm sure there are lower cost options. My droplet is 8 gigs mem which I found necessary for 6 agents.
reply
JSR_FDED 4 hours ago
Appreciate the honesty and lack of hype.
reply
geooff_ 5 hours ago
It's always fun seeing how other people are using these tools.

> Has it been worth it? For the journey, yes, for the ROI, nope.

It's also nice seeing someone experiment without succumbing to AI psychosis.

reply
carimura 4 hours ago
i'd like to go into the psychosis but i'm not there yet.
reply
davidw 3 hours ago
What is the total monthly spend, approximately, on all this; if you don't mind saying, of course?
reply
carimura 3 hours ago
(copied from comment above)

$48/month for the droplet (could probably be cheaper on Hertzner) and $100/month for OpenAI plan. so ~$150/month but again this could be cheaper with a different VPS and using Terra/Luna.

I also have a claude max plan for coding but like I mentioned in the article coding is still separate from the agents.

reply
world2vec 2 hours ago
This is an ad
reply
carimura 21 minutes ago
that's news to me. for what?
reply
phplovesong 3 hours ago
I would never have the power to have AI just do everything. I like to create, not have some slop machine do whatever.
reply
sublinear 2 hours ago
I have gone back and forth on a similar sentiment a few times. I settled on this: the only responsible use of LLMs is to fill in gaps.

It's fine as a search tool or autocomplete. It can be okay to generate code if you didn't know how else to get started, or you've already limited the damage it would do by your own design.

People who overuse LLMs are usually trying to compensate for their lack of experience, structure in their work, or dysfunctional teams. Anyone trying to get hired should recognize it as a new red flag attempting to cover up the old red flags.

reply
carimura 3 hours ago
AI is far from doing everything in my setup. just looking for incremental wins to see if my slop machine can be useful.
reply
carimura 4 hours ago
[dead]
reply
vhbryhin 4 hours ago
[dead]
reply