Skip to content

ARMS uses OpenTelemetry Auto-Instrumentation to help you monitor LLM applications built using models from Amazon Bedrock. This includes tracking performance, token usage, costs, and how users interact with the application.

Auto-instrumentation means you don't have to set up monitoring manually for different LLMs, frameworks, or databases. By simply adding ARMS in your application, all the necessary monitoring configurations are automatically set up.

The integration is compatible with

  • Boto3 Python SDK client >= 1.34.138

Get started ​

Install ARMS ​

Open your command line or terminal and run:

shell
pip install --extra-index-url https://arms-packages.elsaifoundry.ai/root/elsai-arms/ elsai-arms==3.0.4

Initialize ARMS in your Application ​

Zero Code Instrumentation ​

Perfect for existing applications - no code modifications needed:

bash
# Configure via CLI arguments
elsai-arms-instrument \
  --service-name my-ai-app \
  --environment production \
  --otlp-endpoint YOUR_OTEL_ENDPOINT \
  python your_app.py
bash
# Configure via environment variables
export OTEL_SERVICE_NAME=my-ai-app
export OTEL_DEPLOYMENT_ENVIRONMENT=production
export OTEL_EXPORTER_OTLP_ENDPOINT=YOUR_OTEL_ENDPOINT

# Run with zero code changes
elsai-arms-instrument python your_app.py

INFO

Perfect for: Legacy applications, production systems where code changes need approval, quick testing, or when you want to add observability without touching existing code.

One-Line Instrumentation ​

Via Function Parameters ​
python
import elsai_arms

elsai_arms.init(otlp_endpoint="YOUR_OTEL_ENDPOINT")
Via Environment Variables ​

Add the following two lines to your application code:

python
import elsai_arms

elsai_arms.init()

Then, configure the your OTLP endpoint using environment variable:

shell
export OTEL_EXPORTER_OTLP_ENDPOINT=YOUR_OTEL_ENDPOINT

Replace: YOUR_OTEL_ENDPOINT with your OpenTelemetry backend. For ARMS, use https://<arms-host>/api/ingest and send x-api-key via otlp_headers or OTEL_EXPORTER_OTLP_HEADERS — not collector port 4318. For a generic OTel Collector, use http://127.0.0.1:4318.

To send metrics and traces to other Observability tools, refer to the supported destinations.

For more advanced configurations and application use cases, visit the ARMS Python repository.

Prompt cache tracking ​

When your Bedrock application uses prompt caching, ARMS tracks cache reads and writes on every span — including Converse, InvokeModel, and streaming paths.

Cache write tokens are priced separately by TTL. A five-minute write and a one-hour write carry different rates, so ARMS splits them on the span:

Span attributeMeaning
gen_ai.usage.cache_read_input_tokensTokens read from cache
gen_ai.usage.cache_write_input_tokensTotal tokens written to cache
gen_ai.usage.cache_creation.input_tokens.5mCache writes at five-minute TTL
gen_ai.usage.cache_creation.input_tokens.1hCache writes at one-hour TTL
gen_ai.usage.cache.requested_ttlTTL requested on the call (5m or 1h)

When you use elsai-agents (v0.3.3+), these attributes are stamped on model and agent spans automatically. Cache token counts also appear in ARMS trace and log detail views under the cache token fields.

To enable prompt caching in your agent, see Amazon Bedrock — Example: multi-agent, prompt cache, and ARMS.

Copyright © 2026 elsai foundry.