I built a 500k-domain search engine for makers in a weekend for $10
88 points by dreamforever 5 hours ago | 51 comments

tpowell 31 minutes ago
It takes a bit of setup and a huge download, but every time I need a good domain I follow this old post from Derek Sivers. I have Claude de-dupe it and turn it into a searchable database (on my machine), then have it search genres and terms I'm looking for. It's a task Claude is very well-suited to, from the technical implementation to back-and-forth about selections. [link]: https://sive.rs/com
reply
iFire 4 hours ago
Here's my impressions of your algorithm:

1. read each site

2. rent a 4090 with https://vast.ai to run vllm

3. let llm model invent its own category and tag names freely

4. save 1KB of metadata each

  a. a small local language model that reads each one and writes a name, two or three sentences, a category, and a handful of tags.
5. `code is going up as open source` soon (TM)
reply
iambenm 2 hours ago
Code appears to already be up: https://github.com/alexmorleyfinch/marlin
reply
jeroenhd 4 hours ago
The technical details are on another page: https://alexmorleyfinch.github.io/marlin/history/v1/article/...

Your impressions seem about right, but there are a few control steps it seems.

reply
moffkalast 2 hours ago
They really needn't have specified "in a weekend" cause yeah we can tell.

Since when has low effort become a selling point anyhow?

reply
smokel 27 minutes ago
I typically interpret it as an excuse, not a selling point.
reply
dreamforever 4 minutes ago
it was by no means low effort
reply
headz 4 hours ago
TS;DR: Too Sloppy; Didn't Read.
reply
juleiie 4 hours ago
Yeah it’s ai generated obviously but as I said previously which some people didn’t like - don’t judge the tools, judge the content.

I don’t care if robot hand written this or black or white. It’s useful.

reply
order-matters 17 minutes ago
did you edit it? did you fully read everything yourself and take out anything extra that didnt need to be in there like unnecessary comparisons and typical AI idioms? tweak the language to be less dramatic or emphatic about things that do not need emphasis?

AI is a great tool, but it is somewhere between a 3D printer and a CNC machine. it can make something smooth and easy to handle and maybe good enough for personal use but you would want to run it through some Finishing steps before giving its output to someone else. In some cases thats using the output as a mold for a full recreation and other times maybe its just sanding it a bit, smoothing out some rough spots and putting paint on it.

youre right not to care about if something was made by a human or a tool, but the full presentation of the content including editing is part of the content when it comes to a write-up.

reply
abc3354 4 hours ago
In the expression "AI slop", "AI" is about the tools, "slop" is about the content
reply
uean 4 hours ago
I'm not going to spend an hour trying to distill the AI-slop to find out what potential golden nugget may lie in there.

It's impossible to judge the content if it's buried under a landfill. "If you won't take the time to write it, I won't take the time to read it."

reply
juleiie 4 hours ago
[flagged]
reply
vivzkestrel 3 hours ago
- you are going to regret offloading this much of your brain's critical thinking to an external source in the long run

- every night a good number of neuron connections in your brain are automatically severed because you did not use that ability

- over long periods of time, it ll take you to a point where you wont be able to write a simple factorial program on your own without gpt telling you

- i am not saying this to incite you or mock you or anything. i am just concerned about how much of your thinking you are offloading

reply
danggggg 3 hours ago
[dead]
reply
fg137 2 hours ago
> because that’s how things will look like from now on.

No it's not.

People are increasingly getting fed up with slops like this -- you can see comments in HN discussions.

Many people just completely skip those articles.

reply
prepend 48 minutes ago
More importantly, I’ve started my own little blacklist of people and site I just ignore.

There’s people at work whose messages get completely ignored after too many times of posting verbose, useless, ai slop.

reply
gist 49 minutes ago
> People are increasingly getting fed up with slops like this -- you can see comments in HN discussions.

Why claim 'people' (plenty of people get value from satisficing) and even try to use 'see comments in HN discussions' as if there is something that actually documents that ie 'we all think it'.

I skimmed it and even skimming it I got value from (what some others here) call 'slop'. (The value was basically it gave me an idea of something I could implement and again I didn't have to study it or even understand everything (or agree) just seeing someone did that made me think about something I might want to do.

reply
sporedro 2 hours ago
If I want to read AI output, I’ll just ask my own AI. I’m not going to bother reading someone else’s AI.
reply
uean 3 hours ago
Thanks for this. My point is proven.
reply
juleiie 3 hours ago
[flagged]
reply
flyingcoder 3 hours ago
How did web3 and crypto work out for you buddy?
reply
jorisw 3 hours ago
You’re confusing moving on with the times, and using new tools wrong
reply
prepend 49 minutes ago
It’s not my problem because I can safely skip ai slop without fomo.

It’s extremely rare that someone smart with something smart to say produces this crap.

Like someone said upthread if the author isn’t willing to spend time writing cogently then I don’t think it’s worth my time to try to parse it.

reply
yieldcrv 3 hours ago
yes but tell your slopywriter to be less verbose, get to the point
reply
juleiie 3 hours ago
[flagged]
reply
nxndjdkdksmsb 2 hours ago
[dead]
reply
fg137 2 hours ago
> judge the content

The content itself is slop.

reply
ImPostingOnHN 4 hours ago
"slop" is a perfectly valid judgement of content

surely if we're expected to read this ourselves, the author can write it themselves?

indeed, it's helpful to the author, too: writing helps you learn

reply
ModernMech 17 minutes ago
I dunno, you'll notice several people here saying that got value from it, despite others not engaging with it at all. If your filter is that you don't engage with anything you don't expect a priori to get value from (based on weak signals), you'll miss a lot.
reply
marginalia_nu 4 hours ago
Interesting project. Website discovery is indeed in a pretty dire spot, definitely a space that needs innovation. An auto-labeled website directory isn't that silly of an idea.

I have a 400 GB sqlite database with samples of rendered root document DOMs I use for ad detection in Marginalia Search I've been meaning to explore similar ideas using.

reply
alightsoul 2 hours ago
I have been wanting do do this. The biggest source of domains is certificate transparency logs. Also ICANN zone files. According to some scientific papers these cover 88% of all registered domains. You could crawl dns for CNAME records with all ipv4 IPs by distributing requests across dozens of DNS servers, the internet archive or the common crawl but doing it for the internet archive is a dick move without giving them money

There's about 200 million active domains currently. That's about 66% of all businesses worldwide of which there are around 300 million. Around 100 to 150 million have active webpages

reply
jeromechoo 2 hours ago
It feels like we've hit a point where search engines can become what "todo list apps" were for devs 10 years ago.

What a homebrewed solution lacks in coverage it excels in indexing and serving a small slice of the internet really really well.

reply
marginalia_nu 2 hours ago
To be fair they are a supremely interesting problem to hack away at, and one that will meet you where you are.

Almost anyone can put together a basic search engine in a few thousand lines of code, it's just not very hard to make a program that will index a few million documents better than Confluence.

Then, between that first ansatz and a working scalable internet search engine, you have a pile of interesting problems touching every aspect of computer science and computer hardware and networking, enough so that hundreds of people will have gotten PhDs in narrow sub-problems of those problems you'll be facing.

It's great because you can just tackle the stuff you feel comfortable approaching and leave the rest for later.

reply
lagrange77 32 minutes ago
Took me a few minutes to realise it's not a domain name search engine.
reply
fg137 2 hours ago
Sorry I have a lot of trouble understanding what this is useful for. Like, I am never going to replace it with Google, DuckDuckGo, ChatGPT or even Bing.
reply
dreamforever 20 minutes ago
It's not for that, sorry, I should have been more specific. It's for people who wanna put in the effort and steer their own crawl to surface their own slice of the web. The article is just a little story of the journey
reply
prepend 45 minutes ago
I was wondering the same thing. I’ve wished for just a big blob of the web to grep and regex through, but I don’t think this is that much easier than using duckduckgo or even google.
reply
eggbrain 4 hours ago
This is actually where I see software going in the short term -- cloud moving to local.

A few years ago, if you wanted translation, you'd use Google Translate. If you wanted to search the web, you'd use Google search.

But for a few gigabytes, you can now install nllb-200-distilled-600M, and get translations for almost any language locally. You can have your computer crawl the web, create abstracts and categorizations for websites, and build search exactly as you want it.

The main limiter now is hard drive space (and to an extent, local compute) -- but right now it feels like the 70s again where the terminal into a remote server turned into building applications locally.

reply
dylan604 3 hours ago
The number of times we've gone from cloud/server access via terminal to local compute back and forth is something that always makes me laugh a bit.
reply
an0malous 3 hours ago
It's more of a pipeline than a back-and-forth. New abilities happen in the cloud first because they require specialized, higher capacity resources and then move towards being local as the resource usage gets optimized.
reply
BaudouinVH 2 hours ago
How do you build a list of domains you want to index ? I see there is a fetcher and a spider in the code but so for I haven't found how to build that list.
reply
dreamforever 12 minutes ago
Ah, I forgot to mention that anywhere. You have to provide your own. You can start from a small set, like 10 websites you like that have a bit of character, and it will also add any domains it finds from those 10
reply
pimlottc 3 hours ago
Sometimes I think people forget how capable computers are. 500k is not much. You can just slap that in a Lucene instance. This is a solved problem.
reply
marginalia_nu 3 hours ago
Approaching search by just tossing the data in Lucene is how you end up with Confluence's search box though.
reply
pavel_lishin 2 hours ago
From the screenshot, it's very funny that one of the indexed sites is www.llresearch.org, which looks like it's run by a crackpot.
reply
dreamforever 10 minutes ago
There are all kinds of websites in here lol. There is some gold in here and I'm determined to surface it all. I had to wrap this up without full analysis cos it was dragging on
reply
orliesaurus 4 hours ago
Like a personal Google? How do you bypass all the captcha, ip bans, cloudflare turnstile antibot stuff etc?
reply
dreamforever 7 minutes ago
despite this limitation, there is still some good stuff out there, and with the priority steering, you can focus compute on what you actually want, fast and cheap.
reply
sandeepkd 3 hours ago
Thats the fun part, the user just went with happy path. Javascript, captchas, cloudflare protected content did not made to the catalogue. This sort of use case exists in LLM training data a lot which makes it easier. The data gathered by the user is not really practically useful cause there are way too many gotchas when it comes to web scraping and building a catalogue (source: I have done scraping for a particular domain data and had to do at least 10+ iterations to get it >90 right)
reply
voidUpdate 3 hours ago
They don't: "skips the model entirely if the page is empty, parked, or a bot-challenge wall"
reply
dewey 4 hours ago
I think Kagi Small Web filter would give you very similar results.
reply
dreamforever 7 minutes ago
I'll check them out!
reply
dreamforever 5 hours ago
Check out my latest project! You can fork it, tweak the policy manually or with AI, run the system and watch the data come in! It's engineered to keep a low data footprint, so 500k domains fits into 1GB on disk. If you have local models it's free! You just might not get the best throughput depending on your GPU. My production data is not exposed anywhere yet, and I may never expose it. The point is for you to fork and make your own policy, and thus your own personal search engine! The article covers basic analysis on my data, so it's worth a read if you're interested! A deeper analysis may arrive with V2 if I ever do it
reply
elorant 2 hours ago
Domains are way more than just 40M though.
reply
BaudouinVH 2 hours ago
From what I understand the aim was not to collect all the domains on the web but focus on personal website, etc. and avoid corporate web sites.
reply
whatistrending 2 hours ago
[dead]
reply
iFire 4 hours ago
[dead]
reply
nonewideas 4 hours ago
[dead]
reply
hns86vq0nb 4 hours ago
[dead]
reply