Appearance
Quickstart
In this guide you'll pull the collector Docker image, point it at your OTel backend, and start seeing GPU and host metrics within minutes.
Prerequisites
- Linux host with NVIDIA, AMD, or Intel GPU (for GPU metrics)
- Docker installed
- An OpenTelemetry-compatible backend (ARMS, Grafana, Datadog, or any OTLP endpoint)
Don't have an OTel backend? Start ARMS locally
bash
docker run -d \
--name arms \
-p 3000:3000 \
-p 4318:4318 \
<gpu-collector-image>Then use http://localhost:4318 as your OTEL_EXPORTER_OTLP_ENDPOINT.
Pull the collector image
bash
docker pull <gpu-collector-image>Run the collector
NVIDIA GPU
bash
docker run -d \
--name otel-gpu-collector \
--gpus all \
--pid=host \
-e OTEL_SERVICE_NAME=my-app \
-e OTEL_RESOURCE_ATTRIBUTES='deployment.environment=production' \
-e OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318 \
<gpu-collector-image>Requires the NVIDIA Container Toolkit on the host. --pid=host is required for per-process GPU attribution (cmdline, PID, zombie state).
AMD GPU
bash
docker run -d \
--name otel-gpu-collector \
--device /dev/kfd:/dev/kfd \
--device /dev/dri:/dev/dri \
--pid=host \
-e OTEL_SERVICE_NAME=my-app \
-e OTEL_RESOURCE_ATTRIBUTES='deployment.environment=production' \
-e OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318 \
<gpu-collector-image>--pid=host is required for per-process GPU attribution (cmdline, PID, zombie state).
Intel GPU
bash
docker run -d \
--name otel-gpu-collector \
--device /dev/dri:/dev/dri \
--pid=host \
-e OTEL_SERVICE_NAME=my-app \
-e OTEL_RESOURCE_ATTRIBUTES='deployment.environment=production' \
-e OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318 \
<gpu-collector-image>Requires Linux kernel 5.10+ with the i915 or Xe driver. --pid=host is required for per-process GPU attribution.
Host metrics only
bash
docker run -d \
--name otel-gpu-collector \
-e OTEL_SERVICE_NAME=my-app \
-e OTEL_RESOURCE_ATTRIBUTES='deployment.environment=production' \
-e OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318 \
<gpu-collector-image>The collector will export host and process metrics even without GPU access.
Verify it's running
bash
docker logs otel-gpu-collectorYou should see output like:
time=2024-01-01T00:00:00Z level=INFO msg="starting opentelemetry-gpu-collector"
time=2024-01-01T00:00:00Z level=INFO msg="discovered GPU" address=0000:01:00.0 vendor=nvidia
time=2024-01-01T00:00:00Z level=INFO msg="system metrics collector initialized"
time=2024-01-01T00:00:00Z level=INFO msg="process metrics collector initialized"
time=2024-01-01T00:00:00Z level=INFO msg="collector running"View metrics in your backend
Open your OTel backend and look for metrics in the hw.gpu.*, system.*, and process.* namespaces.
If using ARMS, navigate to http://localhost:3000 and go to the Metrics section.
Docker Compose
Add the collector as a service alongside your existing stack:
yaml
services:
otel-gpu-collector:
image: <gpu-collector-image>
pid: host
environment:
OTEL_SERVICE_NAME: my-app
OTEL_RESOURCE_ATTRIBUTES: deployment.environment=production
OTEL_EXPORTER_OTLP_ENDPOINT: http://otel-collector:4318
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
depends_on:
- otel-collector
restart: alwayspid: host (Docker --pid=host) is required so the collector can see host workload PIDs under /proc for per-process GPU metrics, cmdline, and zombie detection.