Run Qwen3 4B on 60% of a GPU with HAMi and KitOps
Jesse Williams · July 17, 2026

Run Qwen3 4B on 60% of a GPU with HAMi and KitOps

GPUs are expensive and scarce, and most inference workloads waste them. A 4-billion-parameter model with decent quantization needs only a fraction of a modern accelerator, yet a single A10 with 24GB of VRAM usually ends up dedicated to a workload that uses maybe 20% of the card. In this walkthrough we run a Qwen3 4B model on 60% of a single GPU on Kubernetes, with the model weights and the server shipped together as one signed, reproducible artifact. One kubectl apply schedules the pod onto a fractional slice of the GPU, and we get live streaming inference out of the other end.

Two open source projects do the heavy lifting: Project HAMi for the GPU virtualization, and KitOps for the model packaging. Along the way we explain why you would pull the model through your own registry instead of reaching out to a public hub at container start.

Two problems worth solving

The first problem is GPU utilization. Teams tend to pick one of two bad options. They give a whole card to one pod and burn an entire GPU on a workload that barely touches it, or they reach for manual MPS or time-slicing hacks that are brittle, offer no hard isolation, and let one noisy tenant starve the others.

The second problem is how you ship the model itself. The usual answer is to pull weights from a public model hub when the container starts. That is a runtime network call to a public service with no signing, no versioning you control, and no pinning you control. If the hub is down, your pod never comes up. And because the model in production is whatever sits at the other end of that download, you cannot prove it matches what your team approved. AI supply chain tampering happens exactly in that gap between approval and execution.

Project HAMi: slicing one GPU into many

HAMi (Heterogeneous AI Computing Virtualization Middleware) is a CNCF project that became an incubating project in July 2026. In one sentence: it lets you slice a single physical GPU into fractional virtual GPUs and schedule them on Kubernetes with hard limits on both memory and compute.

Using it is simple. Instead of asking for a whole NVIDIA GPU, you add two resource limits to your pod spec:

resources:
  limits:
    nvidia.com/gpumem-percentage: 60   # cap the container to 60% of VRAM
    nvidia.com/gpucores: 60            # cap the container to 60% of compute cores

HAMi installs a scheduler and a device plugin. The scheduler places your pod on a node with a free slice, and the device plugin enforces the cap inside the container. Here is the part that makes it real isolation: when your process runs nvidia-smi or asks CUDA how much memory exists, HAMi intercepts the call and reports the sliced number rather than the physical total. The application genuinely believes it is running on a smaller GPU and cannot overallocate even if it tries. That interception happens at the CUDA API layer, so no driver changes and no application changes are required.

KitOps: the model as a signed OCI artifact

KitOps is a CNCF project that solves the packaging problem. It introduces the ModelKit, an OCI-compliant package that bundles everything an AI project needs, including the model weights, configuration, tokenizers, datasets, code, and documentation, into a single versioned, signable artifact. Because ModelKits are OCI-compliant, they live in the same registries you already use for container images, and you push and pull them exactly like a Docker image.

You define a ModelKit with a Kitfile, a small YAML manifest that plays the role a Dockerfile plays for a container image:

manifestVersion: 1.0.0
package:
  name: qwen3-4b
  license: Apache-2.0
model:
  name: qwen3-4b
  path: ./models/qwen3-4b   # the Qwen3 model weights

For this demo the model lives in Jozu Hub, the registry built for ModelKits by the team behind KitOps. A ModelKit can be served in two shapes. The first is a plain ModelKit, which is just the data: you pull it and serve it with your own runtime. The second is a Rapid Inference Container (RIC), an image that has a vLLM server and the weights baked in, with an entry point that starts an OpenAI-compatible API on port 8000. You pull one image, run it, and you have a serving endpoint.

Why a registry instead of a public hub

Why not point vLLM straight at a public model repository? Four reasons.

  • Reproducibility and immutability. A public repository can change under you. Someone force-pushes new weights to main and your same deployment is now running something different. A ModelKit is content-addressed and immutable, so the digest pins the exact bytes.
  • No runtime dependency on a public hub. The artifact lives in your own registry next to your container images, behind your own access controls. Your cluster's image pull path is the only way to reach it.
  • Signed and verifiable provenance. ModelKits are signed, and their provenance can be verified, so you can attest exactly what is inside. This matters for the AI supply chain, which is the newest and softest attack surface in most organizations.
  • One self-contained artifact. With the server and the weights in a single RIC image, there is nothing to download at startup. You run kubectl apply and the pod is self-contained, which is also what lets it survive a demo cluster in a region with unreliable networking.

The architecture, end to finish

  1. The Qwen3 4B model is a ModelKit in Jozu Hub, packaged as a Rapid Inference Container. The vLLM server and the weights are signed together as one OCI image.
  2. A kubectl apply deploys it. The pod spec sets the scheduler name to the HAMi scheduler and adds the two HAMi limits: 60% memory and 60% cores.
  3. The HAMi scheduler finds a GPU slice and places the pod, and the device plugin carves out a 60% fractional vGPU on the physical GPU.
  4. The RIC entry point boots vLLM, loads the baked-in weights, and serves an OpenAI-compatible endpoint on port 8000.
  5. A normal chat completions request comes in, and tokens stream back.

The demo

After connecting to the cluster and confirming the nodes are up with kubectl get nodes, we prove three things.

HAMi placed the pod. The pod's scheduler name is the HAMi scheduler, not the default one, and the resource limits show 60 cores and 60% memory. Exec into the pod, run nvidia-smi, and the isolation is visible: the same tool that reports roughly 23,000 MiB on the bare-metal A10 reports only about 13,800 MiB inside the container, which is exactly 60% of the card. The application literally cannot see the rest of the GPU.

The endpoint is live. Port-forward the pod and ask it what model it is serving, and it answers Qwen3 4B.

Inference streams in real time. Send a normal chat completions request and watch the model stream tokens back over an OpenAI-compatible API, generated on a GPU slice capped at 60%. Tail the pod logs and the access line shows the request hitting the RIC and returning 200 OK, served inside the HAMi-virtualized pod by the vLLM server that came from Jozu Hub.

Why this combination matters

Put the two projects together and you get efficient use of scarce GPUs alongside a model artifact you can actually trust. HAMi gives you hard isolation so several small models can share one card without stepping on each other. KitOps gives you a signed, immutable, content-addressed artifact that ships the server and the weights as one unit, pulled through your own registry rather than a public hub. Efficient hardware and a verifiable AI supply chain are usually treated as separate concerns. Here they are one deployment.

Project HAMi is worth a look if you are fighting for GPU capacity, and KitOps is the on-ramp for packaging models as OCI artifacts. If you have questions, come find us in the community.

Share this post