Amazon Bedrock for Beginners: One AWS Service, Many AI Models

Imagine that you want to use several AI models in one application, such as Claude, Llama, Amazon Nova, Mistral, and Cohere. Normally, you might need a separate account, API key, SDK, and billing setup for each provider. Amazon Bedrock makes this much simpler. You use your AWS account and IAM permissions, choose a model, send your request, and Bedrock handles much of the connection work for you.

The problem Amazon Bedrock solves

Without Bedrock, each AI provider can become a separate setup

Suppose your application needs AI models from Anthropic, Meta, Amazon, Mistral, or Cohere. If you connect to every provider directly, you may need to do the following for each one:

  • Create a separate provider account.
  • Create and manage a separate API key.
  • Learn a different authentication method.
  • Manage a separate billing account.
  • Learn a different SDK or API endpoint.
  • Handle different security and networking requirements.

Amazon Bedrock reduces this work. Bedrock is a managed AWS service that gives you access to foundation models from different providers through your AWS account. Your application authenticates with AWS IAM, chooses a model, and sends the request to Bedrock. Bedrock then sends the request to the selected model and returns the answer.

Without Bedrock Your application
Anthropic account + API key Meta account + API key Mistral account + API key Cohere account + API key
With Bedrock Your application AWS IAM + Amazon Bedrock
Claude Llama Amazon Nova Mistral Cohere

One important point: Bedrock gives you a common AWS way to handle access, security, billing, monitoring, and infrastructure, but the models are still different from one another. Some models support different parameters or special features. AWS provides the Converse API to make common chat requests look more consistent across many models.

What does “fully managed” mean?

AWS manages the infrastructure; you focus on using the service

When Bedrock is described as a fully managed or serverless service, it means you do not have to create or manage the servers and GPU infrastructure that run the models. AWS operates and maintains that infrastructure. You still pay based on the Bedrock pricing option you use, just as different AWS services have different billing models.

Simple billing idea for different AWS services
AWS serviceSimple billing idea
AWS LambdaPay for requests and execution time
AWS FargatePay for CPU and memory while the task runs
Amazon S3Pay for storage and requests
DynamoDB On-DemandPay for read and write requests
Bedrock On-DemandPay mainly for input and output tokens
Bedrock Provisioned ThroughputPay for reserved model capacity over time
How Bedrock billing works

With On-Demand, you normally pay for the tokens you use

For common On-Demand model calls, you are usually charged based on tokens rather than how many seconds your request takes. There are two main types of tokens:

  • Input tokens — the text or prompt you send to the model.
  • Output tokens — the text the model generates for you.
A small example

Your prompt uses 15 input tokens. The model response uses 180 output tokens.

The cost is calculated using the price of the input tokens plus the price of the output tokens. Each model has its own pricing, so Claude, Nova, Llama, and Mistral may cost different amounts.

Your prompt
15 input tokens
Model Model response
180 output tokens

Cost = price of the input tokens + price of the output tokens

Bedrock also offers Provisioned Throughput. Think of this as reserving model capacity for your workload. It can be useful when you need more predictable performance or dedicated capacity. Because the capacity is reserved for you, you pay for that reserved capacity even when traffic is low.

Bedrock inference modes and how you pay
Inference modeHow you pay
On-DemandBased on input and output token usage
Provisioned ThroughputReserved capacity, typically billed over time
A common Bedrock misunderstanding

Bedrock gives you one AWS entry point, but every model is not identical

A common misunderstanding is that Bedrock gives exactly the same API request for every model. That is not always true.

Bedrock gives you common AWS authentication and a common service endpoint, but the request format depends on which Bedrock API you use:

  • InvokeModel — you send the request in the format expected by that specific model. Claude, Llama, Nova, and Mistral may use different JSON fields.
  • Converse — AWS gives you a more common chat-style request format. This is easier when your application may switch between supported models.

A simple analogy is a shopping mall. Bedrock is the main entrance to the mall. You enter once, but the individual stores inside are still different. In the same way, Bedrock gives you one AWS entry point, while each AI model can still have its own capabilities and special options.

Amazon Bedrock API flow: Application to SDK/CLI, then Converse or InvokeModel into Bedrock Runtime, then Claude, Llama, or Nova
Converse is the easier, more standard chat path. InvokeModel is the model-specific path. Both requests still go through Amazon Bedrock Runtime before the selected foundation model produces an answer.
Before Converse: using InvokeModel

The same question can require different JSON for different models

Imagine your application supports Claude Sonnet and Llama 3.3 70B. You want to ask both models the same question: “Explain Kubernetes.” With InvokeModel, the request body can look different for each model:

# Claude-style body
{
  "anthropic_version": "bedrock-2023-05-31",
  "messages": [{"role": "user", "content": "Explain Kubernetes."}],
  "max_tokens": 500,
  "temperature": 0.7
}

# Llama-style body
{
  "prompt": "Explain Kubernetes.",
  "max_gen_len": 500,
  "temperature": 0.7
}

The goal is the same, but the JSON fields are different. For example:

Claude and Llama request fields
ClaudeLlama
messagesprompt
max_tokensmax_gen_len
anthropic_versionnot required

InvokeModel uses a common endpoint pattern such as POST /model/{modelId}/invoke, but the JSON body can change depending on the model. In other words, Bedrock standardizes the AWS access and endpoint, but you may still need to understand each model’s native request format.

Converse API: a simpler chat interface

Most chat applications want to do the same basic thing: send a list of conversation messages to a model. AWS introduced the Converse API so that many supported models can use a common message structure, like this:

{
  "messages": [
    {
      "role": "user",
      "content": [{ "text": "Explain Kubernetes." }]
    }
  ]
}

With Converse, the basic request structure can stay the same even when you change the modelId from a Claude model to a Llama model. Bedrock handles the translation to the model’s native format behind the scenes. This makes your application code easier to maintain.

Why can models still be different?

Converse standardizes common chat features, but some model-specific features can still be different. For example, a model may have its own reasoning options, tool settings, safety controls, image or PDF support, or token limits. If your application depends on one of these special features, check that the new model supports it before switching model IDs.

What happens when your app sends a request?

Follow the Bedrock request one step at a time

If you already understand a basic AWS request flow such as client → load balancer → backend, Bedrock follows a similar idea: your application sends a request, AWS checks it and routes it, and the selected model does the AI work.

Amazon Bedrock architecture: client environment, AWS Bedrock service account front door and runtime, and model provider escrow account with hosted foundation model runtime
Your application does not directly call the public API of a model provider such as Anthropic. The request is handled through AWS-managed Bedrock infrastructure.

1. Your application sends the request

The request can come from a web application, AWS Lambda, the AWS CLI, or a program using an AWS SDK such as boto3. It is sent to the Bedrock Runtime endpoint. If your environment requires private networking, you can also use a VPC interface endpoint.

2. Amazon Bedrock checks the request

Think of Bedrock as the front desk. It first checks who you are using IAM, whether you are allowed to use the requested model, whether the request is valid, and whether you are within service limits. If something is wrong, the request can fail before it ever reaches the model.

3. Bedrock Runtime routes the request to the model

After the request passes the checks, Bedrock Runtime acts like a traffic controller. It sends the request to the correct model backend. If you used Converse, Bedrock may translate the common Converse format into the format expected by that model. It then waits for the model response and sends the result back to your application.

4. Prompt history in the AWS console

The Bedrock playground in the AWS console can keep a history of prompts you enter there. This does not mean that every request from your Python code, Java application, CLI, or REST API automatically appears in that console history.

5. The model runs inside AWS-managed infrastructure

Your application is not simply sending the prompt to a provider’s public website over the internet. The model is made available through AWS-managed Bedrock infrastructure. This is why AWS controls such as IAM, CloudTrail, and PrivateLink can be used around Bedrock access. The underlying model infrastructure may serve multiple customers, while customer requests remain isolated.

6. The model files are already deployed and ready

The model files are stored securely, but Bedrock does not download a huge model file from storage every time you ask a question. The model is already deployed on inference infrastructure, including GPUs, and is ready to process requests.

Do all these steps make Bedrock slow?

Most of the waiting time is usually the model generating the answer

A Bedrock request does pass through several AWS-managed steps, but those checks and routing steps are usually much faster than generating the AI response itself. The table below shows the general idea using typical example ranges:

Typical example time for each step of a Bedrock request
OperationExample time
Network to Bedrock5–20 ms
IAM authentication< 5 ms
Authorization and validation< 5 ms
Routing< 5 ms
Model begins inference50–200 ms
Model generates about 500 tokens1–5 seconds
Authentication and routing Milliseconds
Generating hundreds of tokens Seconds

The key idea is simple: authentication and routing usually take milliseconds, while generating hundreds of tokens can take seconds. In most cases, the model generation time is the larger part of the total response time.

What can you do with Amazon Bedrock?

Bedrock provides more than just access to an LLM

Amazon Bedrock is a platform for building and operating generative AI applications. You may use only one feature at first and add more features as your application grows.

Foundation models

Use models such as Claude, Llama, Nova, Mistral, Cohere, and others through your AWS account.

Guardrails

Add rules to help control harmful content, prompt injection, sensitive information, and restricted topics.

Knowledge Bases

Build managed RAG applications. Bedrock can retrieve useful information from your documents and give that context to the model so the answer is grounded in your data.

Flows

Connect prompts, models, knowledge bases, agents, and conditions into a visual workflow.

Imported or custom models

Run supported model architectures using Bedrock-managed inference. Not every model checkpoint can be imported.

Agents

Allow a model to use tools such as AWS Lambda functions, APIs, databases, and knowledge bases to complete tasks.

For imported models, Bedrock must support the model architecture. You cannot assume that every model from Hugging Face can be imported. Pricing can also be different: native foundation models are commonly priced by token usage, while imported models may require dedicated inference infrastructure.

Your first Python example

Use boto3 to talk to Bedrock Runtime

The goal is simple: your Python program sends a prompt and receives an AI response. Your code talks to Amazon Bedrock Runtime. Bedrock then sends the request to the foundation model you selected.

Python application (your code) boto3 (AWS SDK for Python) Amazon Bedrock Runtime Foundation model (Nova, Claude, Llama, etc.) AI response returns to your program

You will often see two Bedrock-related boto3 clients. They have different jobs:

The two Bedrock-related boto3 clients
boto3 clientWhat it is used for
bedrockManagement tasks, such as listing models and working with guardrails
bedrock-runtimeInference tasks: send prompts and receive model responses

Here is a small Converse example:

import boto3

client = boto3.client("bedrock-runtime", region_name="us-east-1")

response = client.converse(
    modelId="amazon.nova-lite-v1:0",
    messages=[{
        "role": "user",
        "content": [{"text": "Explain Amazon Bedrock in simple language."}]
    }]
)

print(response["output"]["message"]["content"][0]["text"])

What happens behind the scenes? boto3 creates an HTTPS request and signs it using AWS Signature Version 4. Bedrock Runtime checks your IAM permissions and service limits, sends the request to the selected model, and returns the generated text. If your AWS account has access to another supported model, you can change the modelId.

For a larger example, see the IdeaWeaver repository: bedrock_converse.py.

If you work in DevOps, SRE, or platform engineering

Monitor Bedrock like any other production service

In production, it is not enough to ask only, “Is Bedrock up?” You also want to understand traffic, response time, errors, limits, and cost. These are useful starting points:

1. Invocation count

How many model requests are being made? If the count suddenly drops to zero, first check your application and upstream traffic.

2. Invocation latency

How long do requests take? Look at averages, but also P95 and P99 so that slow requests are not hidden by a good average.

3. Errors

Watch for 4xx and 5xx errors such as AccessDenied, ValidationException, ResourceNotFound, ModelTimeout, or InternalServerException.

4. Time to first token (TTFT)

For streaming responses, users notice how quickly the first word appears. A good TTFT can make an application feel much faster.

5. Token usage and cost

Input and output tokens affect model cost. Track usage and compare it with AWS Cost Explorer so you can understand where your Bedrock spend is coming from.

Model IDs and inference profiles

An inference profile is a routing option, not a different AI model

For example, anthropic.claude-sonnet-4-6 and us.anthropic.claude-sonnet-4-6 do not represent two different Claude models. One identifies the base model, while the other is a US inference profile that can route requests across supported US deployments.

Base model ID compared with a US inference profile
IDMeaningWhen you might use it
anthropic.claude-sonnet-4-6Base model IDIdentify the foundation model itself
us.anthropic.claude-sonnet-4-6US inference profileLet AWS route to an available supported US deployment
Your application us.anthropic.claude-sonnet-4-6 (US inference profile)
Supported US Region Supported US Region Supported US Region
Same Claude Sonnet model

Think of an inference profile as a smart routing layer. When you use a US inference profile, Bedrock can choose an available supported US Region for the request instead of forcing your application to hard-code one deployment location. Similar routing profiles can exist for other geographies or global routing.

What should you remember?

Five beginner-friendly takeaways

  1. Amazon Bedrock gives you one AWS service for accessing many foundation models.
  2. With On-Demand inference, cost is mainly based on input and output tokens. Provisioned Throughput reserves capacity.
  3. InvokeModel can use model-specific request bodies, while Converse gives you a more common chat format.
  4. Your application sends requests through AWS-managed Bedrock infrastructure, and model generation usually takes much longer than authentication and routing.
  5. In production, monitor request count, latency, errors, time to first token, limits, and token cost.

Quick review

Try to explain each term in your own words before reading the answer:

The simplest way to remember Bedrock: one AWS entry point, many AI models. Use Converse for a common chat experience, understand token-based cost, and monitor it like any other production service.