Most of the DB engines out there still seem to use a "n-threads" style parallelism with exchange operations and poor async I/O management.
DuckDB is improving on this front, but in some sense is catching up to R&D (and implementation!) that is now decades old.
A "rhetorical challenge" I like to give software developers working on systems like this is the following: If I gave you a computer with 1,024 cores and matching network and storage bandwidth -- but with significant latency -- could you keep a system like this 100% utilised with one query?
The answer for almost all software is "no".
For example, SQL Server tops out at 64 hardware threads for any one query: https://learn.microsoft.com/en-us/sql/database-engine/config...
GPU codes are starting to get there, but CPU codes are way behind on this frontier of computer science.
It's not just databases! Can you (de)compress a file in parallel? Verify its hash in parallel? Upload/download from storage with CPU and I/O task parallelism? Can you overlap all of these operation so nothing is ever waiting on anything else it doesn't have to?
This matters! I ran some tests with bioinformatics codes and found that most got stuck in tar pits. Many could not scale to modern SSDs with millions of IOPS or modern networking with hundreds of gigabits of throughput, no matter how many CPU cores were thrown at them.
PS: AMD's Zen 6 era EPYC 9006 processors will have 512 cores and 1,024 threads per two-socket system, so this is not hypothetical: https://www.amd.com/en/products/processors/server/epyc/9006-...
> One setting drives this,...
> The cost is now about the rows you actually touch, not rounds times table size.
Etc.
I get the brain scramblies [1] from trying to parse this writing style at work so I hate to see it elsewhere. Apologies if I'm wrong. But if I'm not then OP don't use an LLM to write for you. It's hazardous to your reader's health [2].
[1]: https://www.youtube.com/watch?v=ipUJq-odt5Q
[2]: https://discourse.haskell.org/t/how-to-keep-enjoying-program...
> Storing a lake as thousands of 1 MB Parquet files is a bad practice anyway, and 2.0 does not rescue it.
The "does not rescue it". No human would write like that.
> I'll explain what that means on a table you already know.
No I don't already know that table.
Also
> and claims 40x on graph reachability
Is really hard to parse.
The section on recursive CTEs wasn't well written and didn't explain how the optimisation was done. This article explains how the recursive CTEs were improved https://duckdb.org/2026/08/25/how-duckdb-runs-recursive-ctes...
This is what I don’t understand. Supposedly LLMs are trained on human text. Why do they come up with such unrealistic prose? Is it intentional because the companies want the tells to be obvious?
I'm starting to read it as a sign of low LLM effort not just low human effort. It seems most common when one few-sentence prompt leads it to generate 4+ paragraphs (and the longer the output, the worse the odds). Prompting to dig into each resulting paragraph one by one, to make them readable, makes it do higher-effort deep dives.
Just because some guy on some forum said that doesn't make it true. That's not how you establish facts regarding health claims.
I think you're right on your last point, btw, nobody will care and we'll be the poorer for it. Just like how McDonald's and Walmart have driven out and undercut localized taste and culture in the US, so it shall be with writing of all kinds. It's happening right now and readers are getting used to reading the intellectual equivalent of a Big Mac.
We invented algorithms and machines that ostensibly understand and interpret our words and can make things happen, which is a revolution of understanding. From a legibility perspective, these machines now can "understand" me just as well as if I were writing like Gene Wolfe, yet instead of embracing the idea of "wow I can talk to a computer" we have an overweight on "wow, I can have computer talk to everyone else on my behalf!".
Even if the transmission of information here was perfect I'd still question the utility of outsourcing your brain on writing a very short article compiled from release notes.
Apparently the readability of LLM prose seems to trigger stupid "tabs vs. spaces" style arguments. I encourage you to save your time and energy and do something productive with it instead.
More and more really important and useful information will arrive like this for us to consume. There's no way around. So this is a disservice for newcomers that could come and go unscathed but instead is crippled by these kinds of comments that brings nothing of substance to the table and has the potential to make them hate something they otherwise wouldn't even notice.
Also, from https://news.ycombinator.com/newsguidelines.html:
This sort of writing decreases readability. People say "just have your own LLM rewrite it" like you don't lose value when you go from prompt -> slop you didn't review enough to clean this crap up in -> someone else's prompt to change the style -> finally someone reads it.
Do you know what happens a double-digit percentage of the time when I ask Claude to rewrite some shit that it gives me like this? It says things like: "I overstated this, I rechecked and actually..." or "this claim doesn't hold up, actually [this other thing is true]..."
So it's a sign that the claims in the post likely weren't vetted very hard.
So if you aren't proofreading I'm gonna be skeptical. And saying "deal with it" doesn't rescue it. Does it?
And then there's the reflexive "you must just be an ideological hater." No. I'm someone who uses the tools in a domain where quality matters enough that I have to dig into the quality of the tool output and spot the tells for when it's low output, so that I can deliver shit that works reliably and consistently.
Also if you think LLMisms so bad it's spam, you also don't get a free pass. From the guidelines: