Skip to content

BedrockModel ​

Environment variables: AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_REGION, BEDROCK_MODEL_ID

python
import os
from elsai_model.bedrock import BedrockModel

model = BedrockModel(
    model_id=os.getenv("BEDROCK_MODEL_ID", "us.anthropic.claude-3-5-sonnet-20241022-v2:0"),
    region_name=os.getenv("AWS_REGION", "us-east-1"),
    max_tokens=256,
    temperature=0.2,
)
messages = [{"role": "user", "content": "Say hello in one short sentence."}]

response = model.invoke(messages)

for chunk in model.stream_text(messages):
    print(chunk, end="", flush=True)

Or via the factory (AWS credentials, not api_key / params):

python
from elsai_model import LLM, Provider

model = LLM(
    provider=Provider.BEDROCK,
    model=os.getenv("BEDROCK_MODEL_ID", "us.anthropic.claude-3-5-sonnet-20241022-v2:0"),
    region_name=os.getenv("AWS_REGION", "us-east-1"),
    aws_access_key=os.environ["AWS_ACCESS_KEY_ID"],
    aws_secret_key=os.environ["AWS_SECRET_ACCESS_KEY"],
    max_tokens=256,
    temperature=0.2,
)

Configuration ​

ParameterDescription
model_idBedrock foundation model ID (e.g. us.anthropic.claude-3-5-sonnet-20241022-v2:0)
region_nameAWS region where the model is enabled
max_tokensMaximum tokens to generate
temperatureSampling temperature
cache_configPrompt caching — see Prompt caching below

Credentials are read from the environment (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY) or the default AWS credential chain (IAM role, ~/.aws/credentials).

Prompt caching ​

Bedrock prompt caching stores static context (system instructions, documents) so later calls read from cache instead of re-processing the full prompt. This cuts cost and latency on repeated requests.

Requires agent-grade BedrockModel from elsai-model (Converse API). The legacy BedrockConnector InvokeModel path does not support caching.

Automatic caching with CacheConfig ​

Pass cache_config on BedrockModel — the SDK injects a cache point automatically. Use this with Agent(model=...) and plain string system prompts:

python
import os
from elsai_model.bedrock import BedrockModel
from elsai_model.models import CacheConfig

region = os.getenv("AWS_REGION", "us-west-2")

model = BedrockModel(
    model_id=os.getenv("BEDROCK_MODEL_ID", "global.anthropic.claude-sonnet-4-6"),
    region_name=region,
    max_tokens=512,
    temperature=0.2,
    cache_config=CacheConfig(strategy="auto", ttl="1h"),
)

Set ttl="1h" only on models that support one-hour cache. Omit ttl or use ttl="5m" for the default five-minute cache.

Standalone invoke:

python
messages = [{"role": "user", "content": "What is 2+2?"}]
response = model.invoke(messages)
print(response["metadata"]["usage"])

cache_prompt is deprecated — use cache_config instead.

Streaming with usage ​

Use stream_text_with_usage() when you need cache token counts during streaming. Text chunks omit usage; the final chunk includes it:

python
messages = [{"role": "user", "content": "Say hello in one short sentence."}]

for chunk in model.stream_text_with_usage(messages):
    print(chunk.get("text", ""), end="", flush=True)
    if "usage" in chunk:
        print("\nUsage:", chunk["usage"])

Usage fields ​

Cache-related fields in the normalized Usage dict:

FieldDescription
cacheReadInputTokensInput tokens read from cache on this call
cacheWriteInputTokensInput tokens written to cache on this call
cacheDetailsBreakdown of cache writes by TTL — each entry has ttl ("5m" or "1h") and inputTokens
requestedCacheTtlTTL the SDK requested on the outbound call ("5m" or "1h")

For agent usage with CacheConfig, see Amazon Bedrock — Prompt caching.


See also ​

Copyright © 2026 elsai foundry.