Appearance
BedrockModel
Environment variables: AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_REGION, BEDROCK_MODEL_ID
python
import os
from elsai_model.bedrock import BedrockModel
model = BedrockModel(
model_id=os.getenv("BEDROCK_MODEL_ID", "us.anthropic.claude-3-5-sonnet-20241022-v2:0"),
region_name=os.getenv("AWS_REGION", "us-east-1"),
max_tokens=256,
temperature=0.2,
)
messages = [{"role": "user", "content": "Say hello in one short sentence."}]
response = model.invoke(messages)
for chunk in model.stream_text(messages):
print(chunk, end="", flush=True)Or via the factory (AWS credentials, not api_key / params):
python
from elsai_model import LLM, Provider
model = LLM(
provider=Provider.BEDROCK,
model=os.getenv("BEDROCK_MODEL_ID", "us.anthropic.claude-3-5-sonnet-20241022-v2:0"),
region_name=os.getenv("AWS_REGION", "us-east-1"),
aws_access_key=os.environ["AWS_ACCESS_KEY_ID"],
aws_secret_key=os.environ["AWS_SECRET_ACCESS_KEY"],
max_tokens=256,
temperature=0.2,
)Configuration
| Parameter | Description |
|---|---|
model_id | Bedrock foundation model ID (e.g. us.anthropic.claude-3-5-sonnet-20241022-v2:0) |
region_name | AWS region where the model is enabled |
max_tokens | Maximum tokens to generate |
temperature | Sampling temperature |
cache_config | Prompt caching — see Prompt caching below |
Credentials are read from the environment (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY) or the default AWS credential chain (IAM role, ~/.aws/credentials).
Prompt caching
Bedrock prompt caching stores static context (system instructions, documents) so later calls read from cache instead of re-processing the full prompt. This cuts cost and latency on repeated requests.
Requires agent-grade BedrockModel from elsai-model (Converse API). The legacy BedrockConnector InvokeModel path does not support caching.
Automatic caching with CacheConfig
Pass cache_config on BedrockModel — the SDK injects a cache point automatically. Use this with Agent(model=...) and plain string system prompts:
python
import os
from elsai_model.bedrock import BedrockModel
from elsai_model.models import CacheConfig
region = os.getenv("AWS_REGION", "us-west-2")
model = BedrockModel(
model_id=os.getenv("BEDROCK_MODEL_ID", "global.anthropic.claude-sonnet-4-6"),
region_name=region,
max_tokens=512,
temperature=0.2,
cache_config=CacheConfig(strategy="auto", ttl="1h"),
)Set ttl="1h" only on models that support one-hour cache. Omit ttl or use ttl="5m" for the default five-minute cache.
Standalone invoke:
python
messages = [{"role": "user", "content": "What is 2+2?"}]
response = model.invoke(messages)
print(response["metadata"]["usage"])cache_prompt is deprecated — use cache_config instead.
Streaming with usage
Use stream_text_with_usage() when you need cache token counts during streaming. Text chunks omit usage; the final chunk includes it:
python
messages = [{"role": "user", "content": "Say hello in one short sentence."}]
for chunk in model.stream_text_with_usage(messages):
print(chunk.get("text", ""), end="", flush=True)
if "usage" in chunk:
print("\nUsage:", chunk["usage"])Usage fields
Cache-related fields in the normalized Usage dict:
| Field | Description |
|---|---|
cacheReadInputTokens | Input tokens read from cache on this call |
cacheWriteInputTokens | Input tokens written to cache on this call |
cacheDetails | Breakdown of cache writes by TTL — each entry has ttl ("5m" or "1h") and inputTokens |
requestedCacheTtl | TTL the SDK requested on the outbound call ("5m" or "1h") |
For agent usage with CacheConfig, see Amazon Bedrock — Prompt caching.
See also
- LLM Models overview — install, extras, factory
- LLM factory