Kubernetes on Oxide: How customer needs shaped our integrations
94 points by stevehipwell 4 hours ago | 32 comments

stevehipwell 3 hours ago
I'm interested to see how the `oxide-cloud-controller-manager` is being built for "modern" Kubernetes and if it leads to any signify difference compared to CCMs that originated in-tree. Given the way Oxide engineer their solutions this could be really interesting.

FYI I've got `karpenter-provider-oxide` on my bingo card...

reply
sudomateo 2 hours ago
There are few paths we can take here. The CCM's primary responsibility is to implement the cloudprovider.Interface[0] which has node, route, and service controllers. However, the CCM can run arbitrary named controllers as well, giving us a fun extension point for the future. The in-tree CCMs are being phased out in favor of out-of-tree CCMs. The AWS Cloud Provider[1] is a good example of what that's starting to look like.

My colleague demo'd Karpeneter internally. We haven't committed releasing it yet but we're discussing it.

[0] https://github.com/kubernetes/cloud-provider/blob/master/clo...

[1] https://github.com/kubernetes/cloud-provider-aws

reply
stevehipwell 2 hours ago
My point was that the major cloud CCMs are rooted (at least spiritually) in in-tree implementations. With Oxide having an API first strategy (similar in principle to a major cloud provider) the direction you take on your greenfield CCM could be interesting.

RE Karpenter, it always seemed like a natural fit for Oxide, even more so than some of the currently implementors. And now you have the expertise in the team, I'd be really interested in the reason if you don't go down that path.

reply
sudomateo 2 hours ago
Totally, the out-of-tree CCM gives us a good opportunity to differentiate.

Autoscaling is a requested feature, and Karpenter fits that shape naturally. The nuance is to decide where something like the Cluster API provider ends and Karpenter begins since there's a bit of overlap in concerns. Specifically, both want to manage Kubernetes nodes but for different reasons. We're discussing it though. We have a Kubernetes watercooler meeting today where we'll likely discuss the comments in this post!

reply
bmwagner10 60 minutes ago
Karpenter is definitely a natural fit for Oxide. There's some interesting boundaries that we're discussing internally, like @sudomateo mentioned. One of which is CAPI. There's a CAPI Karpenter provider that elmiko built (notably back in the very early days of Karpenter). I think there's room for both use-cases. Some may want to use CAPI for everything and others may want a platform that is based on CAPI but guest clusters are not.

One implementation specific detail which makes Karpenter interesting on Oxide is the tunable CPU and memory parameters rather than strict instance type shapes. The prototype I built for Karpenter on Oxide generates all possible "instance type" combinations, so you can create some really specific nodes to bin-pack pods.

Another interesting area is multiple providers. This is becoming pretty common across public clouds too. A multi-provider Karpenter is something I'm interested in and I know some folks have already been gluing together, but the Karpenter story isn't great on maintaining those since you basically need to compile them together today. CAPIs multi-provider story is a bit cleaner since it only relies on CRDs.

If you have ideas, let's chat in the Kubernetes slack #karpenter-dev.

reply
thegagne 2 hours ago
Yes this would be great addition! Also it would be awesome to see proper load balancer, ingress controller with `gateway api` and a `api gateway` (two separate things).

I’m not an oxide customer, just a fan. I have lots of ideas around how something like this should look, being disappointed with the complexity and also shortcomings of products on the market for this stuff, both in the cloud and on prem.

reply
sudomateo 2 hours ago
I'll need to record a video on what's possible today. My personal Kubernetes cluster running on Oxide uses the CCM for LoadBalancer services and NGINX Gateway Fabric for Gateway API things. I'm mostly using HTTPRoute resources today.

There's an opportunity to more tightly integrate at the network layer but we'll want to get our load balancer released first.

reply
lars_francke 47 minutes ago
Disclaimer: I'm totally biased here.

I talked with a colleague from Oxide in 2024 about your Kubernetes story and he said back then "not yet but soon-ish". Seems like soon-ish is now :)

We said we'd talk again when that happens but he's since left Oxide. If you (or well...your customers) are interested in a Kubernetes native data platform 100% open source we'd be very happy to talk about how that could work easily. As it's "just" Kubernetes it should be trivial but we'd be happy to test and add you to our list: https://docs.stackable.tech/home/stable/kubernetes/#supporte...

The offer stands. If you're interested, my mail is in my HN profile. https://stackable.tech

reply
overflowy 3 hours ago
I would absolutely kill for them to open source their documentation system.
reply
bcantrill 3 hours ago
reply
mixmastamyk 38 minutes ago
Why would a “doc system” use React?
reply
dcre 7 minutes ago
Because it’s a web site.
reply
overflowy 3 hours ago
OMG, you just made my day!
reply
ahl 2 hours ago
does this mean that you won't be killing anyone?
reply
bitlad 2 hours ago
I am wondering when would you use kubernetes on oxide vs running kubernetes with kubevirt on baremetal?

At first look, it feels like oxide is equivalent to proxmox or some virtualization tool, may be it uses qemu stuff underneath.

Just curious.

The reason I am asking this is that, we have lot onprem scenarios in our business. We are tightly coupled with k8s, to solve this we started building an internal project that is kubernetes API compatible [1] but runs containerd or WASM or our platform natively.

Just curious how oxides work in this scenario

[1] https://github.com/debarshibasak/superkube

reply
sudomateo 22 minutes ago
Good question. We use our own hypervisor[0] that's not KVM/QEMU. We don't have nested virtualization today so we don't follow the KubeVirt model, though we are discussing what CRDs like OxideInstance would look like for those that want to operate solely in Kubernetes manifests.

The core primitive on Oxide is the instance (virtual machine). We could support some container primitive, but that's a larger product direction discussion. Our host OS is Helios (Illumos) so there are details to iron out there regarding what abstractions we would build and expose to the users. Not impossible but not something we're currently pursuing either given that we have other immediate product asks.

If you're at a scale where compute density, power efficiency, security, and rack-level API management matters then that's where Oxide makes sense for you.

[0] https://github.com/oxidecomputer/propolis

reply
wmf 53 minutes ago
You would use Oxide if you hate Dell/HPE/Lenovo/Supermicro. Also most people won't run k8s on bare metal because they want to dynamically provision a bunch of clusters.
reply
bakies 56 minutes ago
Yeah me too. Like I kinda compare oxide to coreweave and coreweave is built on top of k8s afaik. I'm wondering what the underlying tech oxide leans on. If I was doing this I would be using talos and k8s on baremetal and building on top of that. Container first rather than VM first seems like a big advantage imo. Most workloads will be containers.
reply
bitlad 9 minutes ago
Coreweave is a cloud provider right? I dont think it is apple to apple comparison.
reply
e12e 59 minutes ago
When oxide provide your metal?
reply
pianoben 3 hours ago
I have seldom wanted anything as much as I want an Oxide rack at home. Maybe in 40 years we'll start to see the first ones show up in surplus auctions...
reply
sudomateo 40 minutes ago
You can run the control plane at home. I have a video on how to do it. Until we make a smaller footprint that's the only way you'd get Oxide at home.
reply
bakies 59 minutes ago
fwiw, I am very satisfied with talos and k8s at home.
reply
whazor 2 hours ago
Kubernetes is where Oxide, from my armchair, doesn't yet quite match the public cloud.

On AWS EKS Fargate, each pod runs in its own dedicated VM. With Oxide each k8s node is a VM, so you still need something like Talos.

On networking it looks like it is getting closer. Where you can have external subnet give pod routed IPs without overlay. But the gap, as marked by the article, is also the load balancing.

It would be nice if Kubernetes were a native feature out-of-the-box. Also integrated within the existing user/access control.

reply
sudomateo 2 hours ago
We're discussing what "native Kubernetes" looks like on Oxide in the limit. There's a bunch to build, some of which is blocked on product gaps. We'll get there though!

I don't necessarily want to match the public cloud experience if there's an opportunity for Oxide to exceed the public cloud experience. Eliminating the overlay is a good example of this. We have customers using external subnets to eliminate the overlay but we haven't integrated that into our controllers yet.

reply
moondev 2 hours ago
Love to see the CAPOx provider and buy in to Cluster API
reply
wolttam 2 hours ago
So you’d be attaching a new volume to the running worker VM for each PVC? This seems a little odd to me. Could you attach a single large volume, and do path-based provisioning on that?
reply
sudomateo 2 hours ago
> So you’d be attaching a new volume to the running worker VM for each PVC?

That's what we prototyped before local disk was released and before we started disk hot-plug work.

> Could you attach a single large volume, and do path-based provisioning on that?

Possibly. We'd still want disk hot-plug first. Otherwise, customers would have to create their cluster in a certain shape before using PVCs.

reply
itintheory 2 hours ago
Does seem like that could be an issue for RWX PVCs, unless those volumes can be mounted to multiple nodes. Depending on the workload, you might be better off using NFS.
reply
wsng 60 minutes ago
There is no one-size-fits-all, but I would say NFS is nowadays not a preferred option for most modern distributed datastores. They are fine with local storage only, and achieve coordination with higher-level protocols. Running them on top of NFS kills their performance.
reply
redwood 2 hours ago
How are folks managing stateful workloads like databases on Oxide?
reply
sudomateo 40 minutes ago
I cover this a bit in the post. Today they are running something like Longhorn backed by Oxide storage. When we release our CSI plugin they can use that instead.
reply