Deploy Mistral 7B on Google Kubernetes Engine with full GPU observability using DCGM Exporter, Prometheus, and Grafana. This guide walks you through creating a GKE cluster with NVIDIA L4 GPUs, deploying an LLM inference workload with vLLM, and setting up a complete monitoring stack to observe GPU utilization, memory, temperature, power, and inference performance in real time.
| Duration | ~45 minutes |
| GPU | NVIDIA L4 (g2-standard-4) |
| Model | Mistral-7B-Instruct-v0.3 |
| Runtime | vLLM (OpenAI-compatible API) |
| Monitoring | Prometheus + Grafana + DCGM Exporter |
- Create a GKE cluster with a GPU node pool
- Deploy Mistral 7B using vLLM with chunked prefill for efficient inference
- Set up NVIDIA DCGM Exporter to collect 40+ GPU metrics
- Install Prometheus and Grafana via the
kube-prometheus-stackHelm chart - Configure PodMonitors to scrape both GPU and LLM inference metrics
- Visualize GPU health and inference performance in pre-built Grafana dashboards
- Query the model through its OpenAI-compatible API
Before you begin, make sure you have the following installed and configured:
- Google Cloud SDK (
gcloud) — Install guide - kubectl — Install guide
- Helm 3.x — Install guide
- A Google Cloud project with billing enabled
- A Hugging Face token — Create one here (requires accepting the Mistral 7B license)
Install the GKE auth plugin if you haven't already:
gcloud components install gke-gcloud-auth-pluginOpen Cloud Shell or your local terminal and set the following variables. Update the values to match your project:
export PROJECT_ID=<your-gcp-project-id>
export ZONE=us-central1-a
export CLUSTER_NAME=gpu-observability
export GPU_TYPE=nvidia-l4
export GPU_COUNT=1
export MACHINE_TYPE=g2-standard-4
export HF_TOKEN=<your-huggingface-token>Note: The
g2-standard-4machine type comes with 1 NVIDIA L4 GPU (24 GB VRAM), which is sufficient for Mistral 7B inference.
Create a GKE cluster with a GPU-enabled node pool. To deploy your own DCGM exporter with the full set of GPU metrics, pass --monitoring=SYSTEM to disable the default GKE-managed DCGM:
gcloud container clusters create ${CLUSTER_NAME} \
--project=${PROJECT_ID} \
--zone=${ZONE} \
--accelerator type=${GPU_TYPE},count=${GPU_COUNT},gpu-driver-version=latest \
--machine-type=${MACHINE_TYPE} \
--num-nodes=1 \
--monitoring=SYSTEMThis creates a single-node cluster with an NVIDIA L4 GPU and the latest GPU drivers automatically installed.
Fetch your cluster credentials:
gcloud container clusters get-credentials ${CLUSTER_NAME} \
--zone=${ZONE} \
--project=${PROJECT_ID}Verify the GPU node is ready:
kubectl get nodes -o wide
kubectl describe nodes -l cloud.google.com/gke-acceleratorYou should see nvidia.com/gpu: 1 in the node's allocatable resources.
Create the observability namespace, service account, and RBAC rules that allow Prometheus to scrape metrics across the cluster:
kubectl apply -f monitoring/namespace.yamlThis creates:
- Namespace:
observability - ServiceAccount:
observability-sa - ClusterRole: Permissions to read nodes, pods, endpoints, services, and metrics
- ClusterRoleBinding: Binds the role to the service account
Add the Prometheus community Helm repo and install the kube-prometheus-stack:
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo updatehelm install kube-prometheus-stack prometheus-community/kube-prometheus-stack \
--version 82.13.6 \
--namespace observability \
--set prometheus.prometheusSpec.serviceMonitorSelectorNilUsesHelmValues=false \
--set prometheus.prometheusSpec.podMonitorSelectorNilUsesHelmValues=falseImportant: The
--setflags ensure Prometheus picks up all PodMonitors and ServiceMonitors across the cluster, not just those created by the Helm chart.
Verify the stack is running:
kubectl get pods -n observabilityWait until all pods show Running or Completed.
NVIDIA Data Center GPU Manager (DCGM) Exporter collects detailed GPU telemetry. This project uses a two-DaemonSet pattern split across separate manifests for clarity:
| File | Component | Description |
|---|---|---|
dcgm.yaml |
DCGM Host Engine | Runs nv-hostengine on each GPU node (port 5555) |
exporter.yaml |
DCGM Exporter | Connects to the host engine and exposes Prometheus metrics (port 9400) |
configmap.yaml |
Metrics Config | Defines 40+ DCGM fields to export as a CSV |
podmonitor.yaml |
PodMonitor | Tells Prometheus to scrape the exporter every 30s |
First, deploy the ConfigMap that defines which GPU metrics to collect:
kubectl apply -f monitoring/dcgm/configmap.yamlThis configures 40+ metrics across these categories:
| Category | Metrics |
|---|---|
| Utilization | GPU utilization %, memory copy utilization |
| Temperature | GPU temp, memory temp |
| Power | Power draw (W), total energy consumption |
| Memory | Framebuffer used/free/total (MiB) |
| PCIe | TX/RX throughput, replay counter, link gen/width |
| Errors | XID errors, ECC errors, throttling violations |
| Profiling | Tensor/FP16/FP32/FP64 core activity, DRAM activity |
The host engine DaemonSet runs a privileged container on every GPU node to communicate with the NVIDIA driver:
kubectl apply -f monitoring/dcgm/dcgm.yamlThe exporter DaemonSet connects to the host engine and serves metrics at /metrics on port 9400:
kubectl apply -f monitoring/dcgm/exporter.yamlCreate the PodMonitor so Prometheus scrapes DCGM metrics every 30 seconds:
kubectl apply -f monitoring/dcgm/podmonitor.yamlTip: You can also deploy all four manifests at once:
kubectl apply -f monitoring/dcgm/
kubectl get pods -n observability -l app=nvidia-dcgm
kubectl get pods -n observability -l app.kubernetes.io/name=nvidia-dcgm-exporterSpot-check raw GPU metrics:
kubectl port-forward -n observability \
$(kubectl get pod -n observability -l app.kubernetes.io/name=nvidia-dcgm-exporter -o jsonpath='{.items[0].metadata.name}') \
9400:9400 &
curl -s localhost:9400/metrics | grep DCGM_FI_DEV_GPU_UTILkubectl create namespace mistralCreate a Kubernetes secret with your Hugging Face token so vLLM can download the model:
kubectl create secret generic hf-token-secret \
--from-literal=token=${HF_TOKEN} \
-n mistralThe model weights are cached on a 50 Gi persistent disk so subsequent pod restarts don't re-download:
kubectl apply -f workloads/vllm/mistral/pvc.yamlkubectl apply -f workloads/vllm/mistral/deployment.yamlThis creates a Deployment with the following configuration:
| Parameter | Value |
|---|---|
| Image | vllm/vllm-openai:latest |
| Model | mistralai/Mistral-7B-Instruct-v0.3 |
| GPU | 1x NVIDIA L4 |
| Memory | 20 GB limit |
| Flags | --enable-chunked-prefill --max_num_batched_tokens 1024 |
| API Port | 8000 (OpenAI-compatible) |
| Health | /health endpoint with startup/liveness/readiness probes |
Note: The startup probe allows up to 5 minutes for the model to download and load into GPU memory on first boot. Subsequent starts with a warm cache are much faster.
Watch the pod come up:
kubectl get pods -n mistral -wWait until the pod status shows Running and READY 1/1.
Create a PodMonitor so Prometheus scrapes vLLM's built-in /metrics endpoint every 15 seconds:
kubectl apply -f monitoring/vllm-podmonitor.yamlThis tells Prometheus to look for pods with label app: mistral-7b in the mistral namespace and scrape port http (8000) at /metrics.
Deploy all pre-built dashboards as ConfigMaps that Grafana auto-discovers:
chmod +x monitoring/deploy-dashboards.sh
./monitoring/deploy-dashboards.shThis deploys the following dashboards into a "GPU Observability" folder in Grafana:
| Dashboard | What it shows |
|---|---|
| GPU Infrastructure Overview | GPU utilization, temperature, power draw, memory usage (DCGM) |
| LLM Inference | Tokens/sec, P95/P99 latency, queue depth, error rate, RPS |
| GPU PCIe Health | PCIe bus throughput, replay counters, link state |
| GKE Cluster Overview | Node CPU/memory, pod counts, cluster-level resource usage |
Retrieve the Grafana admin password:
kubectl -n observability get secrets kube-prometheus-stack-grafana \
-o jsonpath="{.data.admin-password}" | base64 -d; echoStart a port-forward to access Grafana locally:
kubectl -n observability port-forward svc/kube-prometheus-stack-grafana 3000:80Open http://localhost:3000 in your browser and log in:
- Username:
admin - Password: (the output from the command above)
Navigate to Dashboards > GPU Observability to see your dashboards. The GPU Cluster Overview dashboard will immediately show GPU utilization, temperature, power draw, and memory usage from DCGM.
Port-forward the vLLM service to your local machine:
kubectl -n mistral port-forward \
$(kubectl get pod -n mistral -l app=mistral-7b -o jsonpath='{.items[0].metadata.name}') \
8000:8000 &Send a chat completion request:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "mistralai/Mistral-7B-Instruct-v0.3",
"messages": [
{"role": "user", "content": "Explain GPU observability in 3 sentences."}
],
"max_tokens": 256
}'You should see a JSON response with the model's completion. Check Grafana — you'll see the request reflected in the LLM Inference dashboard (tokens/sec, latency) and GPU metrics spike in the GPU Cluster Overview.
Send multiple concurrent requests to see the dashboards populate:
for i in $(seq 1 20); do
curl -s http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "mistralai/Mistral-7B-Instruct-v0.3",
"messages": [{"role": "user", "content": "Write a haiku about GPUs."}],
"max_tokens": 64
}' &
done
waitCheck all pods across namespaces:
kubectl get pods --all-namespacesView raw DCGM metrics:
kubectl port-forward -n observability \
$(kubectl get pod -n observability -l app.kubernetes.io/name=nvidia-dcgm-exporter -o name | head -1) \
9400:9400
curl localhost:9400/metricsView vLLM metrics:
curl localhost:8000/metricsRestart the vLLM deployment:
kubectl rollout restart deployment/mistral-7b -n mistralCheck GPU node resources:
kubectl describe nodes -l cloud.google.com/gke-acceleratorUse the teardown script to remove all deployed resources and optionally delete the cluster:
./scripts/teardown.shOr load configuration from a .env file to skip prompts:
./scripts/teardown.sh --env .envThe script removes workloads, monitoring resources, the Helm release, and RBAC — then optionally deletes the GKE cluster entirely. To skip the interactive prompt and always delete the cluster, set DELETE_CLUSTER=Y in your env file.
├── .env.example # Environment variable template
├── architecture.drawio # Architecture diagram source
├── image.png # Architecture diagram image
├── scripts/
│ ├── setup.sh # One-command interactive setup
│ └── teardown.sh # Remove all resources and optionally delete the cluster
├── workloads/
│ └── vllm/
│ └── mistral/
│ ├── deployment.yaml # vLLM Mistral 7B deployment
│ ├── pvc.yaml # 50Gi model cache
│ └── secret.yaml # HuggingFace token
└── monitoring/
├── namespace.yaml # Namespace + RBAC
├── resource-quota.yaml # Resource quota for observability namespace
├── servicemonitor.yaml # ServiceMonitor for NIM workloads
├── vllm-podmonitor.yaml # Prometheus scrape config for vLLM
├── dcgm/
│ ├── dcgm.yaml # DCGM host engine DaemonSet
│ ├── exporter.yaml # DCGM exporter DaemonSet
│ ├── configmap.yaml # GPU metrics field definitions (40+ metrics)
│ └── podmonitor.yaml # Prometheus scrape config for DCGM
├── grafana/
│ ├── dashboards.yaml # Grafana dashboard sidecar ConfigMap
│ └── dashboards/ # Individual dashboard JSON files
│ ├── gpu-infra-overview.json # GPU utilization, temp, power, memory
│ ├── llm-inference.json # vLLM inference metrics
│ ├── pcie.json # PCIe health
│ └── gke-cluster-overview.json # GKE cluster resources
└── deploy-dashboards.sh # Script to deploy all dashboards as ConfigMaps
If you prefer to skip the manual steps, the interactive setup script handles everything:
./scripts/setup.shOr load configuration from a .env file (copy .env.example to .env and fill in your values):
cp .env.example .env
# edit .env
./scripts/setup.sh --env .envThe script prompts for each decision interactively. Key env variables you can pre-set to skip prompts:
| Variable | Default | Description |
|---|---|---|
HF_TOKEN |
(prompt) | Hugging Face token for downloading Mistral 7B weights |
DEPLOY_MONITORING |
Y |
Deploy Prometheus + Grafana + DCGM |
DEPLOY_DASHBOARDS |
Y |
Deploy pre-built Grafana dashboards |
DISABLE_DCGM |
N |
Use --monitoring=SYSTEM to disable GKE-managed DCGM |
