AWS Bedrock Model Lag: When Your AI Platform Leaves You a Generation Behind

AWS Bedrock Model Lag: When Your AI Platform Leaves You a Generation Behind

πŸ‘253views

AWS Bedrock lags because its model onboarding process is slower than the partnerships that specialized providers like Together AI build with model labs. By the time Bedrock adds a model, the lab has released a newer generation, leaving enterprises stuck behind.

Article Summary
  • 1.
    What it is
    You will learn how AWS Bedrock lags behind providers like Together AI in offering the latest open weight models from DeepSeek, Kimi and GLM, with specific version and cost comparisons.
  • 2.
    Why it matters
    The article argues that this lag is structural, not temporary, and that decoupling model choice from inference platform is the only way to avoid being permanently a generation behind.
  • 3.
    Key takeaway
    Platform choice should not dictate model estate; procure models independently from infrastructure.
~25 min read
Listen to this article2 plays

DeepSeek, Kimi and GLM show why model choice and cloud choice can no longer be the same decision

1. Bedrock is a very AWS answer to generative AI

Amazon Bedrock takes a chaotic market of model providers and wraps it in IAM, AWS networking, logging, governance, regional deployment and an existing enterprise commercial relationship. What comes out the other side is something a CIO can safely put on an architecture diagram, and for a bank or any large regulated enterprise that proposition is enormously attractive. I am not dismissing it.

There is one problem, and it is structural rather than a criticism of the engineering. For several important model families, the market is moving faster than Bedrock, and being six months behind in AI is nothing like being six months behind on Java or PostgreSQL. Six months can mean an entirely new model generation, a materially different capability envelope, a much larger context window and, quite often, a completely different inference cost. I call this Bedrock lag.

Before I go further, let me be clear about where this is going, because the obvious conclusion is the wrong one. The answer is not to replace Bedrock with Together AI. That simply exchanges one dependency for another. The architectural mistake is allowing your inference provider to define your model estate at all.

2. The problem is no longer catalogue size

It would be unfair to argue that AWS offers only a small handful of models, because the Bedrock catalogue has grown substantially. Together AI now advertises more than two hundred models across text, code, vision, video, image and audio, but raw counts are the least interesting part of the comparison, because a large catalogue full of last year’s models solves nothing.

The question worth asking is not how many models a provider has. It is how old the best model is in each family that matters to you. That distinction matters because the frontier of open weight AI is moving very quickly, and it is moving fastest inside the Chinese labs. DeepSeek, Moonshot AI and Z.ai are the clearest illustrations of the point.

3. Where the gap actually sits

As of 12 September 2026, the picture looks roughly like this. I have deliberately separated what Bedrock will sell you from where the family itself has actually reached, because those are now two different things.

FamilyAmazon BedrockTogether AI serverlessLatest family releaseDeepSeekV3.2V4 Flash 0731 and V4 Pro 0813, both live, with a V4.1 Flash listing already publishedV4.1 FlashKimiK2.5K3K3GLMGLM 5GLM 5.2, GLM 5.3 and GLM 5.3 Flash, all liveGLM 5.3

This is a meaningfully worse picture for Bedrock than the one I published in August. At that point two of Together’s newest endpoints, DeepSeek V4 Pro 0813 and GLM 5.3, carried live prices but were still marked as launching soon, which I flagged as evidence that even the fastest provider answers to the lab rather than the other way round. Both are now serving traffic, Together added a cheaper GLM 5.3 Flash endpoint on top, and Bedrock’s cards for these three families have not moved at all.

The DeepSeek row understates the gap if you read it as a single version jump. AWS’s DeepSeek card is still V3.2. Since then DeepSeek has shipped a V4 preview on 24 April, V4 Flash 0731 with open weights on 31 July, and the official V4 Pro 0813 GA release on 13 August, which supersedes the April preview with materially stronger agentic results. Bedrock has not moved through any of that sequence.

A reader could reasonably say a snapshot is persuasive but AWS just hasn’t caught up yet. What is harder to wave away is the onboarding latency itself, measured directly, and then measured again three weeks later while the gap kept widening.

ModelModel releasedBedrock availabilityLag as at 12 Sep 2026DeepSeek V3.21 Dec 202510 Feb 202671 daysKimi K2.527 Jan 202610 Feb 202614 daysGLM 511 Feb 202618 Mar 2026~35 daysGLM 5.216 Jun 2026not on Bedrock88+ daysKimi K327 Jul 2026not on Bedrock47+ daysDeepSeek V4 Flash 073131 Jul 2026not on Bedrock43+ daysDeepSeek V4 Pro 081313 Aug 2026not on Bedrock30+ daysGLM 5.314 Aug 2026not on Bedrock29+ days

That table changes the claim from “Bedrock is behind” to something more precise: Bedrock’s onboarding latency for these families is itself unpredictable, and the recent releases have simply not arrived. Fourteen days for Kimi K2.5, seventy one for DeepSeek V3.2, thirty five for GLM 5, and then five consecutive releases across three families with no Bedrock date at all, the oldest of them now approaching three months. You cannot plan an architecture around a dependency whose lead time swings between two weeks and an open ended wait.

Version numbers are also a weak proxy for capability on their own, so it is worth seeing what actually changed underneath them.

ModelContextWhat changedDeepSeek V3.2164KPrevious generationDeepSeek V4 Flash 07311M284B mixture of experts, 13B active per tokenDeepSeek V4 Pro 08131M1.6T mixture of experts, 49B active, compressed attention cutting per token compute and KV cacheKimi K2.5256KPrevious generationKimi K31M2.8T sparse mixture of experts, native visionGLM 5200KPrevious generationGLM 5.2256KRepository scale agentic coding, configurable thinking effortGLM 5.31MSame base model as 5.2, gains entirely from scaled post training

Those are not cosmetic differences. They are changes in architecture, context and modality, which is what makes “a generation behind” something a reader can see rather than infer from a version string.

AWS is aware of this problem. Bedrock launched Project Mantle in February specifically to simplify and accelerate model onboarding, and Mantle is how DeepSeek V3.2, Kimi K2.5 and the GLM 4.7 family reached Bedrock in the first place. That may well shrink the lag going forward. But an improved onboarding pipeline is still a pipeline, and it is a different thing from a direct contractual relationship with a lab that ships day zero access as a matter of course. A faster queue is not the same as skipping the queue, and the next section shows exactly what skipping the queue looks like when AWS has the relationship to do it.

4. The OpenAI counterexample shows that lag follows relationships

Everything in section 3 describes AWS working at arm’s length from a model creator, building a general onboarding pipeline and hoping it moves fast enough. OpenAI on Bedrock is the opposite case, and it is worth dwelling on because it shows that Bedrock itself is not structurally slow. The lag described above is specific to AWS’s relationship with the Chinese open weight labs, not to Bedrock as a platform.

Amazon announced a $50 billion investment in OpenAI in February 2026, completed it by the end of July, and AWS built a deep strategic go to market relationship on top of it, including becoming the exclusive third party cloud provider for OpenAI’s Frontier programme. The model availability followed. GPT-5.5 launched on 23 April with API access a day later, and reached Bedrock on 1 June, which is roughly a five week gap rather than a zero one, so I would not oversell the early part of this story. What happened next is the interesting part. Two specialised cybersecurity models, Daybreak Red and Daybreak Blue, arrived on Bedrock for eligible customers on 11 August. Then GPT-6 Astra, OpenAI’s largest and most capable model, went generally available on Bedrock on 8 September, the same day OpenAI itself launched it, alongside the OpenAI API and Microsoft Azure rather than behind them. By the time I checked the Bedrock OpenAI page on 12 September the catalogue listed GPT-6 Astra, the GPT-5.6 family of Terra, Sol and Luna, both Daybreak models, and GPT-5.5 and GPT-5.4 beneath them.

The engineering underneath is worth noting too, because it is the same platform people accuse of being slow. Bedrock runs these models on its own next generation inference engine, with an isolated request queue per customer, durable state capture so that a failed node does not force a request to restart from the beginning, and the full set of AWS governance controls, meaning IAM permissions, VPC and PrivateLink isolation, KMS encryption and CloudTrail audit logging, extended to the calls themselves. That is not a platform incapable of moving at frontier speed.

The Daybreak gating detail is worth noticing on its own, because it is the same pattern GLM 5.3 shows in section 7. A lab can ship a capable model and still choose to withhold general access to it on safety grounds, and that choice sits with the lab regardless of which cloud is hosting the weights.

Put the OpenAI timeline next to the DeepSeek, Kimi and GLM timelines and the real driver becomes clear. Where Amazon has invested $50 billion and AWS has built a deep strategic relationship, as it has with OpenAI, a new flagship model generation can reach Bedrock on the day it launches. Where AWS is running open weight models through a generic pipeline, for labs it has no comparable relationship with, onboarding takes anywhere from two weeks to, on current evidence, indefinitely. Together’s day zero access to Moonshot’s releases, described in section 6, is the same phenomenon seen from the other side.

So the lesson is not that Together is faster than AWS. The lesson is that model availability follows relationships, and no provider can own every relationship. That is why the enterprise has to own the control plane. AWS is extraordinarily close to OpenAI, so OpenAI models arrive immediately. Together is extraordinarily close to Moonshot, so Kimi arrives immediately. Neither of them can plausibly have the deepest relationship with every lab that matters, which means model lag is not fundamentally a cloud problem at all. It is a supply chain problem, and the architectural response to a supply chain problem is never to pick a single supplier.

5. DeepSeek makes the problem painfully obvious

Bedrock gives you DeepSeek V3.2 with a 164K context window. Together gives you DeepSeek V4 Flash 0731, live and callable, supporting a one million token context window and activating only 13 billion parameters out of a 284 billion parameter mixture of experts architecture, which is a different design rather than a version bump.

When I first wrote this section, DeepSeek’s flagship V4 Pro 0813 had a priced Together page marked launching soon. It is now live at $1.32 per million input tokens, $0.13 cached and $3.96 per million output tokens. That is a 1.6 trillion parameter model with 49 billion active parameters, a one million token context window, and compressed attention designs that DeepSeek says cut per token inference compute to a fraction of V3.2’s while shrinking KV cache usage substantially. Together’s catalogue has since added a V4.1 Flash listing on top of that. The gap between Together and the absolute frontier is measured in days. The gap between Bedrock and the frontier, for this family, is now measured in seasons.

The economics are where it gets uncomfortable. In the major US Bedrock regions, DeepSeek V3.2 currently costs $0.62 per million input tokens and $1.85 per million output tokens. Together charges $0.14 per million input tokens and $0.28 per million output tokens for DeepSeek V4 Flash 0731, with cached input falling to $0.03 per million tokens.

Take a meaningful workload consuming 100 million input tokens and generating 20 million output tokens.

ModelInput costOutput costTotalBedrock DeepSeek V3.2$62.00$37.00$99.00Together DeepSeek V4 Flash 0731$14.00$5.60$19.60Together DeepSeek V4 Pro 0813$132.00$79.20$211.20

That is roughly an eighty percent reduction in token cost on the Flash comparison while simultaneously moving forward an entire model generation, and a doubling of cost if you want the flagship instead. Obviously these are not the same model, and model selection should always be driven by evaluations rather than release dates, but that is precisely my point. If V4 Flash passes my evaluation then I should be able to choose it, and if V4 Pro justifies its premium on my own workload then I should be able to choose that, and my infrastructure platform should not be making either decision on my behalf simply because it has not finished onboarding the model.

6. Kimi shows why the gap may be structural rather than temporary

Kimi is the interesting case, because it cuts against the obvious narrative twice over. AWS has Kimi K2.5, launched on 27 January 2026, with a 256K context window at $0.60 per million input tokens and $3.00 per million output tokens in the main US regions. Moonshot released Kimi K3 on 27 July with 2.8 trillion parameters, native vision and a one million token context window, designed explicitly around long horizon coding, research and agentic workloads.

Together had it live almost immediately, and two days later Together and Moonshot announced a strategic partnership under which Together becomes a launch platform for Moonshot’s future open weight releases, with what Together describes as day zero access. That is the part worth dwelling on. A generic onboarding pipeline is structurally disadvantaged against a provider with a day zero integration agreement with the model creator itself, and Together’s relationship with Moonshot is exactly that kind of agreement rather than a queue AWS could simply work faster.

There is a sharper way to see the same point inside AWS’s own documentation. Bedrock the managed service does not have Kimi K3, six and a half weeks after release. But three days after Moonshot released the weights, AWS’s own machine learning team published a deployment guide for running K3 on SageMaker HyperPod and EKS, on p6-b300 instances with eight NVIDIA B300 GPUs. AWS’s infrastructure was ready for K3 within days. Bedrock, the managed product built on top of that infrastructure, still is not. That gap is the whole argument in one example: AWS can run the model. Bedrock will offer the model. Those are two different sentences, on two different timelines, and the article’s Bedrock lag is really a statement about the second one.

The second way Kimi cuts against the narrative is price. K3 is not cheaper. Together charges $3.00 per million input tokens and $15.00 per million output tokens, with cached input at $0.30, so the same workload looks like this.

Model100M input plus 20M outputBedrock Kimi K2.5$120Together Kimi K3$600

K3 costs five times as much, so you could reasonably ask why anyone would use it, and for many workloads the honest answer is that they would not. What K3 offers is a materially different capability envelope rather than a discount, and for a long running agentic task the million token context and native vision may well justify the multiple. I want my engineering team making that call on the evidence, and I do not want the absence of a model from my cloud provider’s catalogue making it for them.

7. GLM shows that the frontier can outrun everybody, and then catch up

Bedrock offers GLM 5, which launched on 11 February 2026 with a 200K context window, at $1.00 per million input tokens and $3.20 per million output tokens in the main US regions. Together is serving GLM 5.2, released on 16 June with a 256K context window and positioned around agentic software engineering, at $1.40 per million input tokens and $4.40 per million output tokens with cached input at $0.26.

On our example workload that is $164 for GLM 5 on Bedrock against $228 for GLM 5.2 on Together. Once again the newer model is the more expensive one, which is why this is not an argument that AWS is structurally overpriced. Together gives me access to a much larger portion of the current price performance frontier, and that frontier is sometimes cheaper and sometimes considerably more expensive while being considerably more capable. What I care about is being able to evaluate the trade.

Then Z.ai released GLM 5.3 on 14 August, and the GLM story stopped being a simple two column comparison. GLM 5.3 uses the same base model as 5.2, with the entire improvement coming from scaled post training, and it ships with a one million token context window and a claimed fifty percent gain on Z.ai’s own coding benchmark. Bedrock is now two releases behind on this family rather than one.

This is also the part of the article that has aged most usefully. When I first wrote it, nobody had GLM 5.3 on a normal serverless endpoint, because Z.ai withheld the open weights for roughly two weeks after launch, citing safety work following unexpectedly strong cyber capability. That is the same instinct behind OpenAI’s Daybreak gating in section 4: a lab deciding that raw capability and general availability are not the same decision. Once Z.ai released the weights, Together had GLM 5.3 live at the same $1.40 and $4.40 pricing as 5.2, was publishing production benchmark results against GLM 5.3 Flash by late August, and added the Flash endpoint itself at $0.15 per million input tokens and $0.50 per million output tokens with cached input at $0.03. Bedrock is still on GLM 5.

That sequence is the cleanest single illustration of the thesis. A lab gate delays everybody equally, including Together, and the moment the gate lifts, the provider with the relationship serves the model within weeks while the provider without one does not serve it at all. Provider independence never guaranteed immediate access to every frontier model, because nobody can serve weights that have not been released. What it guarantees is that your own provider’s catalogue timeline stops being an artificial extra constraint stacked on top of the lab’s release schedule.

8. The real cost of lag

Put the current pricing side by side and the shape of the problem becomes clear.

ProviderModelInput per 1MOutput per 1M100M in plus 20M outBedrockDeepSeek V3.2$0.62$1.85$99.00TogetherDeepSeek V4 Flash 0731$0.14$0.28$19.60TogetherDeepSeek V4 Pro 0813$1.32$3.96$211.20BedrockKimi K2.5$0.60$3.00$120.00TogetherKimi K3$3.00$15.00$600.00BedrockGLM 5$1.00$3.20$164.00TogetherGLM 5.2$1.40$4.40$228.00TogetherGLM 5.3$1.40$4.40$228.00TogetherGLM 5.3 Flash$0.15$0.50$25.00

Prices and availability checked 12 September 2026. Every row on the Together side of that table was callable at the time of writing, which was not true three weeks earlier, and that movement is itself the point. The Bedrock rows have not changed since March.

Cached input deserves more attention than it usually gets, particularly for coding agents. Together prices cached input at $0.03 per million tokens for DeepSeek V4 Flash and GLM 5.3 Flash, $0.13 for DeepSeek V4 Pro, $0.26 for the full GLM 5.2 and 5.3, and $0.30 for Kimi K3, and an agent does not behave anything like a chatbot. It repeatedly reads system prompts, repository context, tool definitions, instructions and prior conversation state, so caching can change the economics of an agentic workload far more than the headline token price suggests.

This is how model lag turns into an economic tax. You are not necessarily just running an older model, you may be paying a premium for the privilege of running an older model.

9. The reason enterprises still choose Bedrock

Model freshness is not my only concern when I am running a bank. I care where my data goes, whether I have private connectivity, how identity is enforced, where data resides and whether I can prove which infrastructure processed a given prompt. On all of that, Bedrock remains extremely strong.

AWS allows a workload to reach Bedrock through PrivateLink, so traffic can leave my VPC through an interface endpoint without an internet gateway, NAT, VPN or public IP address. It also has explicit regional and geographic inference semantics, and with geographic cross region inference AWS can route requests across regions while keeping processing inside a defined geography such as the US or the EU, with the documentation stating that traffic remains encrypted across the AWS private network throughout. The Chinese models I have discussed are actually more constrained than the rest of the catalogue, because the Bedrock cards for DeepSeek V3.2, Kimi K2.5 and GLM 5 expose in region endpoints only and currently show geographic and global inference as unsupported. That is restrictive, but it is also extremely clear, and for a regulated financial institution clarity is worth a great deal.

10. A Chinese model does not mean Chinese inference

This is where enterprise risk conversations most often go wrong, because there is a persistent tendency to collapse several distinct questions into a single objection about the country on the model card. There are at least five separate decisions hiding in there: model provenance, inference location, data residency, model licence and, as both GLM 5.3 and OpenAI’s Daybreak models have demonstrated, release gating. Each one deserves its own control.

Together states that third party models from companies such as DeepSeek, Qwen and Mistral run on Together’s own infrastructure, that requests are not passed back to the model author, and that its DeepSeek deployment runs in secure North American data centres. Its Moonshot arrangement is likewise described as US hosted. The architecture is therefore your application to Together infrastructure to DeepSeek weights, not your application to Together to DeepSeek servers in China.

The same principle applies to any open weight model. Weights can be deployed somewhere entirely separate from the organisation that trained them, and in financial services getting that distinction straight is worth more than any amount of policy language about model origin.

11. Where Together still needs to mature

Together is no longer the thin American inference API I once took it for. Its GPU infrastructure material describes workloads across more than twenty five cities, a US portfolio exceeding 2 GW, and more than 150 MW available in Europe across France, the Netherlands, Sweden and Romania, built on NVIDIA reference architectures with InfiniBand networking and managed Kubernetes or Slurm. It has also announced a multi year European programme with Hypertec and 5C targeting up to 2 GW and close to a hundred thousand Blackwell and successor GPUs, rolling out from late 2025 through 2028, so it would be wrong to treat all of that as live today. Alongside serverless inference it now offers provisioned throughput with reserved token capacity and an uptime SLA, dedicated endpoints with reserved GPUs, full GPU clusters and enterprise deployments with private networking. That is an infrastructure ladder rather than a SaaS product.

The gap is precision about locality on the commodity path. Enterprise residency options are documented, dedicated EU endpoints are available on higher tiers, and specific models are described as North American or US hosted. What I cannot do is look at the shared serverless API and reason about it the way I can reason about bedrock-runtime.eu-west-1.amazonaws.com. This is also why I would avoid describing the Together footprint as a collection of POPs. A CDN needs POPs. An AI cloud needs GPUs. What matters is where the accelerator performing the inference physically sits, where prompts and outputs are permitted to travel, and whether I can constrain that geography contractually. Together is improving here, but the AWS regional model remains easier for a regulated enterprise to defend to a regulator.

12. The answer is an internal model gateway

Do not allow either provider to become your architecture. If every application in your company calls Bedrock directly then Bedrock is your AI architecture, and if every application calls Together directly then Together is your AI architecture. Both outcomes are the same mistake wearing different logos.

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚ Application  β”‚
                         β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
                                β”‚
                                β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚ Internal AI Gateway  β”‚
                    β”‚                      β”‚
                    β”‚ Policy               β”‚
                    β”‚ Routing              β”‚
                    β”‚ Evaluation           β”‚
                    β”‚ Cost                 β”‚
                    β”‚ Data classification  β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”¬β”€β”€β”€β”˜
                            β”‚     β”‚    β”‚
                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β”‚    └──────────┐
                 β–Ό                β–Ό               β–Ό
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚ Amazon Bedrock  β”‚ β”‚ Together AIβ”‚ β”‚ Self hosted /   β”‚
        β”‚                 β”‚ β”‚            β”‚ β”‚ dedicated       β”‚
        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                 β”‚                β”‚                β”‚
                 └──────────┐     β”‚     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                            β–Ό     β–Ό     β–Ό
                     Open model estate

I have added a third branch to the diagram on purpose. A gateway that only ever chooses between Bedrock and Together is not provider independence, it is a dual vendor strategy with the same failure mode one level up, and the relationship argument in section 4 is exactly why. Each provider’s catalogue reflects the deals it has managed to sign, and those deals change. Self hosting, whether that is your own GPU estate or rented dedicated infrastructure, has to be a real option the gateway can route to, or you have just built a more elaborate way of picking sides.

With that in place the decisions become tractable. A sensitive workload requiring explicit AWS regional processing routes to Bedrock. A workload that wants GPT-6 Astra routes to Bedrock as well, because that is where AWS’s relationships have put the frontier. A coding agent where DeepSeek V4 Flash gives better economics routes to Together. A workload where Kimi K3 genuinely outperforms everything else uses Kimi K3. When a model becomes cheaper somewhere else next month, or a gated release finally opens, you move. The model decision and the infrastructure provider decision should be two separate decisions, and that is far more practical now that Together exposes OpenAI compatible APIs and Bedrock’s own Mantle engine does two distinct things at once. It runs OpenAI’s own frontier models directly, at OpenAI’s first party pricing, as described in section 4. Separately, it also exposes an OpenAI compatible interface, through the bedrock-mantle endpoint, for other models in the catalogue including DeepSeek V3.2. Those are two different features that happen to share a name, and treating them as the same thing is an easy mistake to make from outside AWS.

One caveat is important enough to state plainly: an OpenAI compatible endpoint is not the same thing as a portable model. The wire format lining up means your HTTP client does not need to change. It says nothing about whether the model behind it handles tool calling the same way, exposes the same reasoning effort controls, honours cached context the same way, accepts the same structured output schema, enforces the same token limits, or fails in the same manner under load. A gateway built only to the lowest common denominator of the API shape will work right up until an application depends on one of those details, and then it will fail in a way that is hard to diagnose because the API call still looks correct. The gateway needs a capability layer, not just a routing layer: something that knows, per model, whether it supports vision, what reasoning effort levers exist, whether cached input is priced and honoured, the real context ceiling, the approved data region and the approved data classification, and what the model actually scored in your own evaluations. That is a materially more interesting piece of infrastructure than a generic proxy, and it is also the part most vendors will not build for you.

13. The AI control plane belongs to the enterprise

The most dangerous question in this whole area is “what models does Bedrock support?” It sounds like a sensible architecture question, and it is the wrong one. The better question is what the best model for this workload is today, and where you are prepared to execute it, because those are two separate governance decisions. A model can be approved without every inference provider being approved. A provider can be approved for one data classification and refused for another. A European workload might require a European dedicated deployment while a low sensitivity coding workload runs happily on Together’s US infrastructure and a highly sensitive banking workload stays inside AWS behind PrivateLink. That is actual risk management. Saying “we use Bedrock” is not an AI strategy, it is a procurement decision wearing the costume of an architecture.

It helps to separate three layers that this article has been sliding between. AWS has the strongest cloud control plane: IAM, networking, residency and PrivateLink, built over two decades and genuinely hard to reproduce, plus the deepest commercial relationship with OpenAI. Together currently has the stronger open weight supply plane: direct lab relationships with Moonshot and Z.ai, fast onboarding, and a stack running from serverless inference through to dedicated clusters. Neither of those is your AI control plane. That third layer, the one that governs which evaluations you trust, which data classification maps to which provider, how cost is tracked and how routing decisions get made, belongs to you, and no vendor is going to build it on your behalf, because it is supposed to encode your risk appetite rather than theirs. No inference provider can be your AI control plane, because no inference provider can own every important model relationship.

So use Bedrock where Bedrock is best, use Together where Together is best, self host where that makes sense, and make sure the application never has to care which of those is true this quarter. The model that wins your evaluation in February may not make your shortlist in August. Your cloud provider should not get to decide which intelligence your company is allowed to evaluate.

14. References

Prices and availability checked 12 September 2026, except the historical Bedrock onboarding dates in section 3, which were established on 20 August 2026 and are unchanged. Given how quickly this catalogue is moving, treat anything below as a snapshot rather than a permanent fact.

  1. Amazon Bedrock adds support for six fully managed open weights models, covering DeepSeek V3.2, Kimi K2.5 and GLM 4.7 under Project Mantle, posted 10 Feb 2026.
  2. Amazon Bedrock expands support for AWS PrivateLink, covering the bedrock-mantle endpoint.
  3. Try the new console experience in Amazon Bedrock, on the Mantle inference engine and OpenAI compatible APIs.
  4. Minimax M2.5 and GLM 5 models now available on Amazon Bedrock, posted 18 Mar 2026.
  5. Deploying Kimi K3 on Amazon SageMaker HyperPod and Amazon EKS, AWS’s self host deployment guide published three days after Moonshot’s release.
  6. Together AI model catalogue, checked 12 Sep 2026, showing DeepSeek V4 Pro 0813, GLM 5.3 and GLM 5.3 Flash live with serverless pricing.
  7. Together AI pricing
  8. Together AI GLM 5.2 model page
  9. Together AI GLM 5.3 model page
  10. Together AI DeepSeek V4 Pro 0813 model page
  11. Together AI serverless inference
  12. Together AI DeepSeek V4 Pro announcement, covering the April preview release
  13. Together AI and Moonshot partnership announcement
  14. Together AI privacy, security and third party model hosting documentation.
  15. Together AI GPU infrastructure, European expansion and provisioned throughput material.
  16. Z.ai GLM 5.3 developer documentation
  17. Axios: China’s Z.ai holds GLM 5.3 release over hacking risks
  18. VentureBeat: GLM 5.3 is here with advanced cyber capabilities
  19. DeepSeek V3.2 model card, Amazon Bedrock, confirming the 1 Dec 2025 model launch date and a 164K context window, checked 12 Sep 2026.
  20. DeepSeek V4 Pro GA release note, DeepSeek’s own changelog dated 13 Aug 2026.
  21. DeepSeek V4 preview release note, covering the 24 Apr 2026 preview and open weights.
  22. OpenAI on Amazon Bedrock, overview page listing GPT-6 Astra, the GPT-5.6 family, GPT-5.5, GPT-5.4 and the gated Daybreak Red and Daybreak Blue cyber models, checked 12 Sep 2026.
  23. OpenAI models and Codex on Amazon Bedrock are now generally available, AWS blog published 1 Jun 2026, confirming GPT-5.5, GPT-5.4 and Codex GA at first party pricing on Bedrock’s inference engine.
  24. Take on your most ambitious work with GPT-6 Astra on Amazon Bedrock, AWS blog dated 8 Sep 2026.
  25. GPT-6 Astra model card, Amazon Bedrock, confirming a Bedrock model launch date of 8 Sep 2026.
  26. GPT-6 Astra from OpenAI is now available on Amazon Bedrock, Amazon’s rolling update page, also covering the 11 Aug 2026 Daybreak Red and Daybreak Blue availability.
  27. Introducing GPT-5.5, OpenAI, showing the 23 Apr 2026 launch and 24 Apr API availability.
  28. Amazon’s $50 billion investment in OpenAI: what to know, Amazon, Feb 2026.