Gartner changed its forecast for worldwide AI spending in 2026 three times in nine months. In January the number was $2.52 trillion. In May it became $2.59 trillion. By September 16 it had climbed to $2.7 trillion. Research firms rarely revise that often, or that far in one direction.
The headline figure hides a more useful detail. In a July 2026 release on AI models and platforms, a Gartner analyst said that as more models reach the market and usage-based bills get harder to predict, buyers will move towards platforms that help them pick tools, track performance, apply policy and control costs. Put simply, companies have plenty of AI models to choose from. What they lack is a sane way to run several of them without a separate contract, bill and security review for each one.
That is the job Amazon Bedrock was built for. This article explains how it works in plain language, what each part does, how it compares with other options, and, most usefully, how it behaves when the data is thin, the signals disagree, the clock is ticking or the traffic suddenly triples.
The short version
Bedrock is a service from Amazon Web Services (AWS) that lets a company use AI models from several makers through one account. Anthropic's Claude, OpenAI's GPT models, Meta's Llama, Mistral's models and Amazon's own Nova family are all available through it. AWS runs the computers that power the models. You send text in, get text back, and pay for what you use.
Why do large companies care? Mostly because of the controls around the models, not the models themselves. When OpenAI's models arrived on Bedrock in April 2026, AWS noted that they inherit the same protections as everything else there: IAM permissions, PrivateLink (which keeps traffic on Amazon's private network instead of the public internet), guardrails, encryption and CloudTrail logging, which records who called what and when. AWS also states that prompts and responses aren't used to train the models. Those points are what security and legal teams ask about first, and they are the real selling point of generative AI on AWS.
If your company already keeps its data and applications on AWS, Bedrock lets you add AI features without moving that data somewhere new or signing up with a string of AI vendors. If you don't use AWS, most of its advantages shrink, and a similar platform on your own cloud may suit you better.
How a single request travels through Bedrock
It helps to follow one request from start to finish before looking at the individual pieces. Say an employee types "What is our travel allowance for a client visit to Singapore?" into an internal assistant built on Bedrock. Here is roughly what happens in the second or two before an answer appears:
1) The assistant sends the question to Bedrock along with the ID of the model it wants to use.
2) AWS checks permissions. Is this application allowed to call this model at all? That check is handled by IAM (Identity and Access Management), the same system that controls access to everything else in the company's AWS account.
3) A guardrail screens the question for anything it shouldn't allow, such as an attempt to trick the model into ignoring its instructions.
4) A knowledge base searches the company's own documents and pulls out the passages most likely to answer the question, perhaps two paragraphs from the travel policy.
5) The model reads the question plus those passages and writes an answer.
6) The guardrail checks the answer on the way out. Is it supported by the passages? Does it contain anything that should be hidden, such as a personal phone number?
7) The answer goes back to the employee, and a record of the call is written to the company's logs.
Every step in that chain is a place where things can go right or wrong, and each one adds a little time and a little cost. Most of the practical advice later in this article comes back to one of these seven steps.
The building blocks, one at a time
The model catalog
At the centre of Bedrock are foundation models. These are very large AI models trained on huge amounts of general text, images and code, which gives them broad skills such as writing, summarising, translating and reasoning. "Foundation" signals that they are a base to build on. On their own they know nothing about your company.
How many are available depends on who you ask. AWS's Bedrock homepage speaks of hundreds, while an independent guide from ThinkMove Solutions counted nearly 100 in 2026. Different counting of versions and marketplace listings probably explains most of the gap. The more useful question is whether the specific models you need are offered in your AWS Region, because availability varies by location.
Bedrock also includes evaluation tools. You upload a set of sample questions, run them through several models, and compare the results side by side, either scored by a person or by another model acting as a grader. It is the closest thing to a fair taste test, and it saves teams from choosing a model because of a launch announcement.
The catalog keeps growing. On April 28, 2026, AWS announced OpenAI models, the Codex coding agent and a new Managed Agents product on Bedrock in limited preview. On September 8, 2026, OpenAI's GPT-6 Astra became generally available there, with a context window of up to one million input tokens. A context window is the amount of text a model can consider at once, and a token is a small chunk of text, roughly three-quarters of an English word.
The Bedrock API
Applications don't click buttons in a console. They talk to Bedrock through the Bedrock API, which is a defined set of commands one program can send to another. The part most teams use is the Converse API. It gives the same request and response format for many different models, so changing from one model to another often means editing a single line that names the model.
A newer endpoint called Mantle accepts requests written in OpenAI's format. That helps teams who already wrote their code for OpenAI. The consultancy Caylent cautioned in August 2026 that not every Bedrock feature available through the standard route is necessarily available through Mantle, so check the specific feature before relying on it.
Knowledge Bases
This is Bedrock's version of a technique called retrieval-augmented generation, usually shortened to RAG. The easiest way to understand it is to think of a student sitting an open-book exam. Instead of answering from memory, the model is handed the relevant pages from your own files first.
To make that possible, your documents are split into small pieces, and each piece is converted into a list of numbers (an "embedding") that represents its meaning. Those numbers are stored in a database that can find passages by meaning, so a search for "travel allowance" can also find a paragraph titled "per diem rates." At its 2026 New York Summit, AWS added a Managed Knowledge Base option aimed at companies building this kind of setup across large document collections.
Guardrails
Guardrails are safety and policy rules that sit outside the model. They can filter harmful content, block entire topics, hide personal details, spot attempts to manipulate the model, and run a "contextual grounding" check that scores whether an answer is backed by the source documents and actually addresses the question.
Keeping these rules outside the model has a practical benefit. If you switch models next year, your policies come with you unchanged. AWS also says guardrails can be applied to agents built with frameworks such as Strands Agents, including those running on AgentCore.
Agents and AgentCore
An agent is an AI system that takes actions instead of only answering. It might read a customer's email, look up their order, issue a refund within set limits and write a reply, deciding each next step as it goes.
AgentCore is the part of Bedrock designed to run agents reliably. In April 2026 AWS previewed a managed "harness" where developers list a model, instructions and tools, and each session runs in its own microVM, a small, sealed-off virtual computer. The harness reached general availability at the 2026 New York Summit. On September 18, 2026, AWS followed with a new AgentCore Runtime that frees up unused memory during a session and keeps start-up times consistent regardless of how large the agent's software is.
Customisation options
When a general model isn't quite right, Bedrock offers several ways to adjust it. Fine-tuning trains the model on your own examples. Reinforcement fine-tuning scores the model's practice answers against a reward rule you write, which cuts down the number of hand-labelled examples you need. Distillation trains a smaller, cheaper model to imitate a larger one, and AWS says distilled models can be up to 500% faster and 75% cheaper, with under 2% accuracy loss for tasks like RAG. Custom Model Import lets you bring in a model you trained elsewhere, and AWS doesn't charge for the import step.
In practice, most first projects do fine with clear instructions and a knowledge base. Customisation is worth the effort once you have proof that the standard model falls short on your task.
Pricing and service tiers
Bedrock mainly charges per token, counting both what you send and what comes back. A one-page document plus a short question might be around 800 tokens in, and a paragraph-long answer about 150 tokens out, with output tokens usually priced higher than input. On top of that, each request can be assigned a service tier, which trades speed for price:
These rates come from AWS's pricing page as of September 2026. Caylent's guide warns that support for each tier varies by model, so confirm on the model's own pricing card before building a budget around a discount.
Two other features cut costs in the right conditions. Prompt caching stores the unchanging part of a prompt, such as a long set of instructions, so it isn't reprocessed on every call. AWS says it can lower costs by up to 90% and response times by up to 85% on supported models. Intelligent Prompt Routing sends simple questions to a smaller model and harder ones to a larger model in the same family, which AWS says can save up to 30%.
Market statistics worth knowing
Bedrock compared with the other routes
A company has three broad ways to use foundation models in its products: a managed multi-model platform like Bedrock, a direct account with one model maker, or open-weight models (models whose files you can download) running on servers you manage.
The last row matters more than people expect. Model makers often ship new capabilities on their own services before those reach any cloud platform. If your product depends on having the newest feature on day one, a direct contract makes sense. If security review, audit trails and a single bill carry more weight, Bedrock tends to come out ahead.
Microsoft and Google run comparable platforms on their clouds, and the deciding factor is usually where your data already lives. For a company built on Amazon's cloud, generative AI on AWS means fewer data transfers, fewer contracts and one familiar permission system.
Stress-testing Bedrock: five situations demos skip
A demo uses a tidy question, a clean document and one user. Real use involves missing files, contradictions, impatient customers, strange inputs and sudden crowds. Here is how each one plays out on Amazon Bedrock, and what to do about it.
1. Data gaps
What goes wrong: The model can only work with what the knowledge base hands it. If the Singapore travel policy was never uploaded, the search still returns something, maybe the policy for domestic travel. The model then writes a fluent answer from the wrong document, and the employee has no reason to doubt it.
Gaps usually come from a few sources. Documents get updated but the knowledge base isn't re-synced, so answers lag behind reality. Tables get split across chunks, separating a city from its daily rate. Scanned PDFs contain pictures of words rather than text the system can read. And documents lack labels such as region or effective date, so the search can't narrow to the right version. With a "region: APAC" label on each travel document, a filter can stop the European policy from ever reaching the model.
What to do: Tell the model, in its instructions, to say plainly when the provided passages don't answer the question. Turn on the contextual grounding check and route low-scoring answers to a person. Re-sync the knowledge base automatically whenever a document is published. Use Bedrock's advanced parsing options for files full of tables or scans. Then keep a list of every question the system couldn't answer, because that list shows exactly which documents are missing.
WATCH OUT The grounding check confirms that an answer matches the documents it was given. It says nothing about whether those documents are current. An outdated policy produces a well-grounded but wrong answer.
2. Conflicting signals
What goes wrong: Two sources disagree and the system has to pick. The most common case is two versions of the same policy, one saying claims must be filed in 30 days and another saying 60. When both are retrieved, the model may split the difference or mention both without saying which applies.
Conflicts also appear between parts of the system. The model may produce a sound answer that a guardrail then blocks, for example a pharmacist's question about safe dosage limits caught by a harmful-content filter. On Bedrock, the guardrail's decision stands. Automated quality scores can clash with real users too, when a grading model rates replies well but customers keep rating them poorly for being long-winded. And Intelligent Prompt Routing can misjudge a short but tricky question and send it to the smaller model.
What to do: Delete superseded documents where possible. Where you must keep them, label each with an effective date and tell the model to prefer the newest while flagging any conflict. Write a clear, helpful message for moments when a guardrail blocks an answer, and review blocked requests weekly to catch filters set too tightly. Treat automated scores as one signal and keep a weekly human check. For high-stakes categories such as refunds, legal questions or medical content, skip routing and fix the model. And if you ask two models for a second opinion, agree on the tie-break rule before launch. Two voters can't produce a majority, so disagreements should go to a person.
3. Real-time decisions
What goes wrong: People judge speed by how long they wait for the first word. A reply that takes eight seconds to appear feels broken, even if the content is good. Some tasks also have limits that no large model can meet. Approving a card payment in a couple of hundred milliseconds is one example, and a generative model shouldn't make that live decision, though it can review flagged transactions shortly afterwards.
What to do: Stream answers with ConverseStream so text appears as it's written. Cache the fixed part of long prompts. Use a smaller model for the live step. Put the most time-sensitive feature on the Priority tier. Keep answers short with a cap on output length, and send the model fewer, better document passages. Set your own timeout and show a fallback, such as a handoff to a person, when it expires. With agents, cap the number of steps, because each tool call is another trip to the model.
4. Exceptions and edge cases
What goes wrong: Things that worked in testing fail on real inputs. These show up most often:
› Feature gaps between models. Caylent notes that tiers, batch processing, caching and cross-Region inference each have their own model-by-model support list, so switching models can quietly remove a feature you depended on.
› Very long inputs. A long email thread or contract can exceed the context window, and even when it fits, you pay for every token.
› Non-English text. Many models break languages such as Tamil or Hindi into more tokens than the same meaning in English, which raises cost and fills the context window sooner.
› Formatting slips. A model asked for strict JSON (a structured data format) sometimes returns something slightly off.
› Retired models. Older model versions reach end of life on published dates, and replacements don't always behave the same way, even when they score better on paper.
› Cache timing. The original prompt caching launch held cached content for five minutes after each use, with some newer models allowing an hour. Quiet apps may seldom benefit.
› Over-eager filters. Doctors, lawyers and security teams discuss topics that trip content filters for legitimate reasons.
What to do: Pin exact model versions and keep a saved test set to rerun before any switch. Validate structured output in code and retry once. Trim or summarise long inputs before sending. Test costs with real samples in every language you support. Tune filter strength per audience rather than using one setting for all.
5. Behaviour under pressure and at scale
What goes wrong: Each model has limits per AWS account and Region, usually counted in requests per minute and tokens per minute. Go over and Bedrock returns a throttling error instead of waiting in line. If your code simply retries at once, thousands of requests pile back in at the same moment and make the spike worse.
The tiers also behave differently when the whole system is busy. Priority requests get served first and Flex requests wait. Standard, Priority and Flex share one pool of on-demand quota, while Reserved has its own. When traffic goes past a reservation, the overflow is billed at on-demand rates. A September 2026 review by the pricing research site Flexprice pointed out that this keeps an app running but makes the bill least predictable exactly when a reservation was meant to fix it, and that on-demand usage has no built-in spending cap.
Agents add another layer. Caylent lists "counting user requests instead of model turns" as a common reason budgets go wrong. One question from a user can turn into five or ten model calls as an agent searches, checks and revises.
What to do:
A note on the "growing, randomised waits" in that table. The technique is called exponential backoff with jitter. After a rejection, wait a moment, then double the wait after each further rejection, adding a small random delay so requests don't all retry together. AWS's software kits can do this automatically when the "adaptive" retry mode is switched on. Quota increases can take time to approve, so request them well before launch, and run a load test at roughly twice your expected peak to see where the limits actually bite.
What it costs in practice
The cost-tracking firm CloudForecast published an example in August 2026: a support chatbot handling 10,000 conversations a day on Claude Sonnet 4.6 would cost around $3,150 a month at Standard rates, and roughly half that with caching and the Flex tier applied where suitable.
That estimate assumes about one model call per conversation. If the chatbot is really an agent making three calls behind every conversation, the Standard figure climbs towards $9,000, with no visible change for the customer. Build estimates around model calls, set AWS Budgets alerts at several levels of the monthly figure, and review your most expensive feature each month. In most products, one feature quietly accounts for most of the bill.
What this means for you
For developers in particular, the Bedrock API becomes the one place to measure speed, cost and quality across every model you try, which makes comparisons fair.
A sensible first month usually looks like this:
1) Choose one narrow task with a clear measure of success, such as drafting replies to password-reset tickets.
2) Collect 50 to 100 real examples, with good answers written by your own staff.
3) Run two or three foundation models against those examples in the playground and compare quality, speed and cost.
4) Add a knowledge base only if the task truly needs company documents.
5) Switch on a guardrail, budget alerts and timeouts before any real user sees the feature.
6) Release to a small group, read the logs weekly and adjust.
Bedrock is a poor fit if your data lives on another cloud, you need a model it doesn't carry, you must keep everything on your own premises, or you're building a small side project where one provider's key is simpler.
Key takeaways
Final thoughts
Getting an AI model to produce a decent paragraph is no longer the hard part. The hard part is making that paragraph accurate, approved, logged, affordable and quick when thousands of people ask at the same moment. Bedrock handles a large share of that work, especially for companies already running on AWS.
It won't make judgment calls for you. Someone still has to keep documents current, decide what happens when sources disagree, set limits and watch spending. Start with one narrow task, a set of 50 to 100 real test questions and a guardrail from the first day. Measure honestly, then expand. For teams on Amazon's cloud, that is the most direct route to generative AI on AWS that actually holds up in production.


