Skip to content
All writing
5 min read

Time slicing or MIG: how to actually share a GPU

Two ways to run more than one workload on a GPU, and how to pick between them.

By default, Kubernetes gives a whole GPU to one pod. GPUs are an "extended resource" that can't be split, so even a notebook sitting at 4% utilisation holds the entire card, and the next job just waits. Nobody would accept that for CPU. For hardware that costs many times more per hour, most clusters accept it without a second thought.

NVIDIA gives you two ways to share a card: time slicing and MIG. They're not interchangeable, and picking one is really a question about who is sharing the card.

Time slicing

1 GPU

A
B
C
A
B
C

compute · taking turns →

memory · shared by A, B, C

Pods take turns. Memory is shared, so one greedy pod can crash another.

MIG

1 GPU · split into slices

A

own memory · own compute

B

own memory · own compute

C

own memory · own compute

Split in hardware. Each pod gets its own slice and can't touch the others.

same card, two ways to share it

Time slicing

Time slicing makes one physical GPU show up to the scheduler as several. Kubernetes places multiple pods on it, and the driver takes turns running their work. Each pod thinks it has a GPU. In reality it has a share of one.

The nice part is that it's just configuration, not a hardware feature. So it works on cards with no partitioning support at all, which is most of what you'll actually run outside the big datacentre GPUs.

The catch is isolation, and it helps to be specific about what you lose. Memory isn't split. Pods on the same card share its memory, so one greedy workload can push another into an out-of-memory crash it didn't cause. Compute isn't guaranteed either. If a neighbour keeps the card busy, your job just gets slower, and nothing tells you why.

MIG (Multi-Instance GPU)

MIG splits the card in hardware. Each slice gets its own memory, cache and compute, and a pod on one slice can't see or slow down the others. That's real isolation, not everyone agreeing to play nice.

The downsides: it only works on GPUs that support it, you choose the slice layout up front instead of per workload, and a job that needs the full card can't have it while the card is split. You give up flexibility to get a guarantee.

So which one?

I don't think utilisation is the deciding question. The real question is whether the people sharing the card can hurt each other, and whether that matters.

  • Dev, experiments and internal inference, run by people on the same team: time slicing. These workloads are bursty and mostly idle, and if someone hogs the card, you just talk to them.
  • Several teams with their own SLOs, or anything customer-facing: MIG, if your hardware has it. A hard limit is much easier to rely on than an agreement you have to keep checking.
  • Untrusted or externally submitted workloads: MIG, or separate nodes. Shared memory means one bad job can affect everyone on the card.

On the EKS platform I built, the nodes were g4dn, which come with T4 GPUs. T4s don't support MIG at all, so the hardware made the choice for me. That happens a lot. What still mattered was writing it down, so the next person reading the cluster config doesn't assume there's isolation that isn't there.

Don't skip the GPU metrics

Whichever you pick, you can't tell if it's working without GPU-level metrics. Normal Kubernetes metrics only tell you a pod holds a GPU. They don't tell you whether the GPU is actually busy.

On that platform, the DCGM exporter feeds GPU metrics into Prometheus. Two signals are worth alerting on: utilisation, to see whether sharing is actually being used or you're paying for idle cards, and memory pressure, which with time slicing is your only early warning before two pods run into each other.