Skip to content

ARMS uses OpenTelemetry to help you monitor NVIDIA GPUs. This includes tracking GPU metrics like utilization, temperature, memory usage and power consumption.

Using the SDK

Install ARMS

Open your command line or terminal and run:

shell
pip install --extra-index-url https://arms-packages.elsaifoundry.ai/root/elsai-arms/ elsai-arms==3.0.3

Initialize ARMS in your Application

Zero Code Instrumentation

Perfect for existing applications - no code modifications needed:

bash
# Configure via CLI arguments
elsai-arms-instrument \
  --service-name my-ai-app \
  --environment production \
  --otlp-endpoint YOUR_OTEL_ENDPOINT \
  python your_app.py
bash
# Configure via environment variables
export OTEL_SERVICE_NAME=my-ai-app
export OTEL_DEPLOYMENT_ENVIRONMENT=production
export OTEL_EXPORTER_OTLP_ENDPOINT=YOUR_OTEL_ENDPOINT

# Run with zero code changes
elsai-arms-instrument python your_app.py

INFO

Perfect for: Legacy applications, production systems where code changes need approval, quick testing, or when you want to add observability without touching existing code.

One-Line Instrumentation

Via Function Parameters
python
import elsai_arms

elsai_arms.init(otlp_endpoint="YOUR_OTEL_ENDPOINT")
Via Environment Variables

Add the following two lines to your application code:

python
import elsai_arms

elsai_arms.init()

Then, configure the your OTLP endpoint using environment variable:

shell
export OTEL_EXPORTER_OTLP_ENDPOINT=YOUR_OTEL_ENDPOINT

Replace: YOUR_OTEL_ENDPOINT with your OpenTelemetry backend. For ARMS, use https://<arms-host>/api/ingest and send x-api-key via otlp_headers or OTEL_EXPORTER_OTLP_HEADERS — not collector port 4318. For a generic OTel Collector, use http://127.0.0.1:4318.

To send metrics and traces to other Observability tools, refer to the supported destinations.

For more advanced configurations and application use cases, visit the ARMS Python repository.

Using the Collector

Pull otel-gpu-collector Docker Image

You can quickly start using the OTel GPU Collector by pulling the Docker image:

sh
docker pull <gpu-collector-image>

Run otel-gpu-collector Docker container

You can quickly start using the OTel GPU Collector by pulling the Docker image: Here's a quick example showing how to run the container with the required environment variables:

sh
docker run --gpus all --pid=host \
    -e OTEL_SERVICE_NAME='chatbot' \
    -e OTEL_RESOURCE_ATTRIBUTES='deployment.environment=staging' \
    -e OTEL_EXPORTER_OTLP_ENDPOINT="YOUR_OTEL_ENDPOINT" \
    -e OTEL_EXPORTER_OTLP_HEADERS="YOUR_OTEL_HEADERS" \
    <gpu-collector-image>

--pid=host is required for per-process GPU attribution (cmdline, PID, zombie state).

For more advanced configurations of the collector, visit the OTel GPU Collector repository.

Note: If ARMS and the collector run on the same Docker network, use the ARMS service hostname; otherwise use the host IP. You can also add the OTel GPU Collector under services in your own Compose file:

Docker Compose: Add the following config under services
yaml
otel-gpu-collector:
  image: <gpu-collector-image>
  pid: host
  environment:
    OTEL_SERVICE_NAME: 'chatbot'
    OTEL_RESOURCE_ATTRIBUTES: 'deployment.environment=staging'
    OTEL_EXPORTER_OTLP_ENDPOINT: "http://otel-collector:4318"
  device_requests:
  - driver: nvidia
    count: all
    capabilities: [gpu]
  depends_on:
  - otel-collector
  restart: always
Host IP: Use the Host IP to connect to OTel Collector
sh
OTEL_EXPORTER_OTLP_ENDPOINT="http://192.168.10.15:4318"

Environment Variables

OTel GPU Collector uses standard OpenTelemetry environment variables for configuration:

Environment VariableDescriptionDefault Value
OTEL_SERVICE_NAMEService/application name attached to all metricsdefault
OTEL_RESOURCE_ATTRIBUTESResource attributes, e.g. deployment.environment=productiondeployment.environment=default
OTEL_EXPORTER_OTLP_ENDPOINTOpenTelemetry OTLP endpoint URL(required)
OTEL_EXPORTER_OTLP_HEADERSHeaders for authenticating with the OTLP endpointIgnore if using ARMS
OTEL_METRIC_EXPORT_INTERVALMetric polling interval in milliseconds10000
  • Collected Metrics — Details on the types of metrics collected and their descriptions.

Copyright © 2026 elsai foundry.