Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)
78 points by sebg 8 hours ago | 4 comments
miki123211 6 hours ago
Another great way to understand how vllm works is to read the code of nano-vllm[1]. It's basically "vllm but cut down to size. It's ~5kloc, supports just one model, disposes of some of the abstraction layers that vllm needs due to its codebase size, but contains all the major pieces that make an inference engine fast.
replyBinRoo 7 hours ago
Love that this goes beyond paged attention. Curious how this compares with Radix Attention [1]?
reply[1] https://sgl-project-sglang-93.mintlify.app/concepts/radix-at...
I wonder how much it would cost to vibe code the whole thing from scatch?
I wonder how much better models need to get before such a thing wouldn't look like code vomit?
https://github.com/mmastrac/diffgemma