AWS Bedrock Model Lag: When Your AI Platform Leaves You a Generation Behind
AWS Bedrock lags because its model onboarding process is slower than the partnerships that specialized providers like Together AI build with model labs. By the time Bedrock adds a model, the lab has released a newer generation, leaving enterprises stuck behind.
DeepSeek, Kimi and GLM show why model choice and cloud choice can no longer be the same decision
1. Bedrock is a very AWS answer to generative AI
Amazon Bedrock takes a chaotic market of model providers and wraps it in IAM, AWS networking, logging, governance, regional deployment and an existing enterprise commercial relationship. What comes out the other side is something a CIO can safely put on an architecture diagram, and for a bank or any large regulated enterprise that proposition is enormously attractive. I am not dismissing it.
There is one problem, and it is structural rather than a criticism of the engineering. The AI market is moving faster than Bedrock, and being six months behind in AI is nothing like being six months behind on Java or PostgreSQL. Six months can mean an entirely new model generation, a materially different capability envelope, a much larger context window and, quite often, a completely different inference cost. I call this Bedrock lag.
Before I go further, let me be clear about where this is going, because the obvious conclusion is the wrong one. The answer is not to replace Bedrock with Together AI. That simply exchanges one dependency for another. The architectural mistake is allowing your inference provider to define your model estate at all.
2. The problem is no longer catalogue size
It would be unfair to argue that AWS offers only a small handful of models, because the Bedrock catalogue has grown substantially. Together AI now advertises more than two hundred models across text, code, vision, video, image and audio, but raw counts are the least interesting part of the comparison, because a large catalogue full of last year’s models solves nothing.
The question worth asking is not how many models a provider has. It is how old the best model is in each family that matters to you. That distinction matters because the frontier of open weight AI is moving very quickly, and it is moving fastest inside the Chinese labs. DeepSeek, Moonshot AI and Z.ai are the clearest illustrations of the point.
3. Where the gap actually sits
As of 20 August 2026, the picture looks roughly like this. I have deliberately separated what Bedrock will sell you from where the family itself has actually reached, because those are now two different things.
| Family | Amazon Bedrock | Together AI serverless | Latest family release |
|---|---|---|---|
| DeepSeek | V3.2 | V4 Pro, V4 Flash 0731 | V4 Pro |
| Kimi | K2.5 | K3 | K3 |
| GLM | GLM 5 | GLM 5.2 | GLM 5.3 |
Version numbers are a weak proxy for capability, so it is worth seeing what actually changed underneath them.
| Model | Context | What changed |
|---|---|---|
| DeepSeek V3.2 | 128K | Previous generation |
| DeepSeek V4 Flash 0731 | 1M | 284B mixture of experts, 13B active per token |
| Kimi K2.5 | 256K | Previous generation |
| Kimi K3 | 1M | 2.8T sparse mixture of experts, native vision |
| GLM 5 | 200K | Previous generation |
| GLM 5.2 | 256K | Repository scale agentic coding, configurable thinking effort |
| GLM 5.3 | 1M | Same base model as 5.2, gains entirely from scaled post training |
Those are not cosmetic differences. They are changes in architecture, context and modality, which is what makes “a generation behind” something a reader can see rather than infer from a version string.
4. DeepSeek makes the problem painfully obvious
Bedrock gives you DeepSeek V3.2. Together gives you DeepSeek V4 Pro, released on 24 April 2026, alongside V4 Flash 0731, released on 30 July. The Together deployment of V4 Flash supports a one million token context window and activates only 13 billion parameters out of a 284 billion parameter mixture of experts architecture, which is a different design rather than a version bump.
The economics are where it gets uncomfortable. In the major US Bedrock regions, DeepSeek V3.2 currently costs $0.62 per million input tokens and $1.85 per million output tokens. Together charges $0.14 per million input tokens and $0.28 per million output tokens for DeepSeek V4 Flash 0731, with cached input falling to $0.03 per million tokens.
Take a meaningful workload consuming 100 million input tokens and generating 20 million output tokens.
| Model | Input cost | Output cost | Total |
|---|---|---|---|
| Bedrock DeepSeek V3.2 | $62.00 | $37.00 | $99.00 |
| Together DeepSeek V4 Flash | $14.00 | $5.60 | $19.60 |
That is roughly an eighty percent reduction in token cost while simultaneously moving forward an entire model generation. Obviously V4 Flash and V3.2 are not the same model, and model selection should always be driven by evaluations rather than release dates, but that is precisely my point. If V4 Flash passes my evaluation then I should be able to choose it, and my infrastructure platform should not be making that decision on my behalf simply because it has not finished onboarding the model.
5. Kimi shows why the gap may be structural rather than temporary
Kimi is the interesting case, because it cuts against the obvious narrative twice over. AWS has Kimi K2.5, launched on 27 January 2026, with a 256K context window at $0.60 per million input tokens and $3.00 per million output tokens in the main US regions. Moonshot released Kimi K3 on 27 July with 2.8 trillion parameters, native vision and a one million token context window, designed explicitly around long horizon coding, research and agentic workloads.
Together had it live almost immediately, and two days later Together and Moonshot announced a strategic partnership under which Together becomes a launch platform for Moonshot’s future open weight releases, with what Together describes as day zero access. That is the part worth dwelling on, because it explains why this gap is unlikely to close permanently. AWS has a model onboarding process. Together has been building direct relationships with the labs themselves. Of course AWS will eventually add K3, and the reflexive enterprise response is that the lag is temporary, but by the time Bedrock catches K3 Moonshot will have moved again. A procurement pipeline cannot outrun a partnership agreement.
The second way Kimi cuts against the narrative is price. K3 is not cheaper. Together charges $3.00 per million input tokens and $15.00 per million output tokens, with cached input at $0.30, so the same workload looks like this.
| Model | 100M input plus 20M output |
|---|---|
| Bedrock Kimi K2.5 | $120 |
| Together Kimi K3 | $600 |
K3 costs five times as much, so you could reasonably ask why anyone would use it, and for many workloads the honest answer is that they would not. What K3 offers is a materially different capability envelope rather than a discount, and for a long running agentic task the million token context and native vision may well justify the multiple. I want my engineering team making that call on the evidence, and I do not want the absence of a model from my cloud provider’s catalogue making it for them.
6. GLM shows that the frontier can outrun everybody, including Together
Bedrock offers GLM 5, which launched on 11 February 2026 with a 200K context window, at $1.00 per million input tokens and $3.20 per million output tokens in the main US regions. Together is serving GLM 5.2, released on 16 June with a 256K context window and positioned around agentic software engineering, at $1.40 per million input tokens and $4.40 per million output tokens with cached input at $0.26.
On our example workload that is $164 for GLM 5 on Bedrock against $228 for GLM 5.2 on Together. Once again the newer model is the more expensive one, which is why this is not an argument that AWS is structurally overpriced. Together gives me access to a much larger portion of the current price performance frontier, and that frontier is sometimes cheaper and sometimes considerably more expensive while being considerably more capable. What I care about is being able to evaluate the trade.
Then Z.ai released GLM 5.3 on 14 August, and the GLM story stopped being a simple two column comparison. GLM 5.3 uses the same base model as 5.2, with the entire improvement coming from scaled post training, and it ships with a one million token context window and a claimed fifty percent gain on Z.ai’s own coding benchmark. Bedrock is now two releases behind on this family rather than one.
The more interesting detail is that Together does not have it either. GLM 5.3 launched into a staged release, initially available only through Z.ai’s own coding plan and tooling, with open weights withheld for roughly two weeks while Z.ai completed safety work. The stated reason was that cyber capability improved faster than the lab expected as post training scaled, particularly on tasks moving from vulnerability discovery toward complete exploitation chains, and Z.ai reported a leading score on the CyberGym vulnerability discovery benchmark alongside a substantial disclosure ledger. Together’s own catalogue currently lists GLM 5.3 as launching soon rather than live.
That matters for two reasons. It demonstrates that provider speed has a ceiling, because no inference provider can serve weights a lab has not released. It also means the assumption underneath most enterprise open weight strategies, that anything with open weights will be available from somebody within days, is no longer reliable. Staged and gated releases are becoming a feature of the frontier, and dual use cyber capability is the reason.
7. The real cost of lag
Put the current pricing side by side and the shape of the problem becomes clear.
| Provider | Model | Input per 1M | Output per 1M | 100M in plus 20M out |
|---|---|---|---|---|
| Bedrock | DeepSeek V3.2 | $0.62 | $1.85 | $99.00 |
| Together | DeepSeek V4 Flash | $0.14 | $0.28 | $19.60 |
| Together | DeepSeek V4 Pro | $1.74 | $3.48 | $243.60 |
| Bedrock | Kimi K2.5 | $0.60 | $3.00 | $120.00 |
| Together | Kimi K3 | $3.00 | $15.00 | $600.00 |
| Bedrock | GLM 5 | $1.00 | $3.20 | $164.00 |
| Together | GLM 5.2 | $1.40 | $4.40 | $228.00 |
Cached input deserves more attention than it usually gets, particularly for coding agents. Together prices cached input at $0.03 per million tokens for DeepSeek V4 Flash, $0.30 for Kimi K3 and $0.26 for GLM 5.2, and an agent does not behave anything like a chatbot. It repeatedly reads system prompts, repository context, tool definitions, instructions and prior conversation state, so caching can change the economics of an agentic workload far more than the headline token price suggests.
This is how model lag turns into an economic tax. You are not necessarily just running an older model, you may be paying a premium for the privilege of running an older model.
8. The reason enterprises still choose Bedrock
Model freshness is not my only concern when I am running a bank. I care where my data goes, whether I have private connectivity, how identity is enforced, where data resides and whether I can prove which infrastructure processed a given prompt. On all of that, Bedrock remains extremely strong.
AWS allows a workload to reach Bedrock through PrivateLink, so traffic can leave my VPC through an interface endpoint without an internet gateway, NAT, VPN or public IP address. It also has explicit regional and geographic inference semantics, and with geographic cross region inference AWS can route requests across regions while keeping processing inside a defined geography such as the US or the EU, with the documentation stating that traffic remains encrypted across the AWS private network throughout. The Chinese models I have discussed are actually more constrained than the rest of the catalogue, because the Bedrock cards for DeepSeek V3.2, Kimi K2.5 and GLM 5 expose in region endpoints only and currently show geographic and global inference as unsupported. That is restrictive, but it is also extremely clear, and for a regulated financial institution clarity is worth a great deal.
9. A Chinese model does not mean Chinese inference
This is where enterprise risk conversations most often go wrong, because there is a persistent tendency to collapse several distinct questions into a single objection about the country on the model card. There are at least five separate decisions hiding in there: model provenance, inference location, data residency, model licence and, as GLM 5.3 has just demonstrated, release gating. Each one deserves its own control.
Together states that third party models from companies such as DeepSeek, Qwen and Mistral run on Together’s own infrastructure, that requests are not passed back to the model author, and that its DeepSeek deployment runs in secure North American data centres. Its Moonshot arrangement is likewise described as US hosted. The architecture is therefore your application to Together infrastructure to DeepSeek weights, not your application to Together to DeepSeek servers in China.
The same principle applies to any open weight model. Weights can be deployed somewhere entirely separate from the organisation that trained them, and in financial services getting that distinction straight is worth more than any amount of policy language about model origin.
10. Where Together still needs to mature
Together is no longer the thin American inference API I once took it for. Its GPU infrastructure material describes workloads across more than twenty five cities, a US portfolio exceeding 2 GW, and more than 150 MW available in Europe across France, the Netherlands, Sweden and Romania, built on NVIDIA reference architectures with InfiniBand networking and managed Kubernetes or Slurm. It has also announced a multi year European programme with Hypertec and 5C targeting up to 2 GW and close to a hundred thousand Blackwell and successor GPUs, rolling out from late 2025 through 2028, so it would be wrong to treat all of that as live today. Alongside serverless inference it now offers provisioned throughput with reserved token capacity and an uptime SLA, dedicated endpoints with reserved GPUs, full GPU clusters and enterprise deployments with private networking. That is an infrastructure ladder rather than a SaaS product.
The gap is precision about locality on the commodity path. Enterprise residency options are documented, dedicated EU endpoints are available on higher tiers, and specific models are described as North American or US hosted. What I cannot do is look at the shared serverless API and reason about it the way I can reason about bedrock-runtime.eu-west-1.amazonaws.com. This is also why I would avoid describing the Together footprint as a collection of POPs. A CDN needs POPs. An AI cloud needs GPUs. What matters is where the accelerator performing the inference physically sits, where prompts and outputs are permitted to travel, and whether I can constrain that geography contractually. Together is improving here, but the AWS regional model remains easier for a regulated enterprise to defend to a regulator.
11. The answer is an internal model gateway
Do not allow either provider to become your architecture. If every application in your company calls Bedrock directly then Bedrock is your AI architecture, and if every application calls Together directly then Together is your AI architecture. Both outcomes are the same mistake wearing different logos.
┌──────────────┐
│ Application │
└──────┬───────┘
│
▼
┌──────────────────────┐
│ Internal AI Gateway │
│ │
│ Policy │
│ Routing │
│ Evaluation │
│ Cost │
│ Data classification │
└───────┬─────┬────────┘
│ │
┌──────────┘ └──────────┐
▼ ▼
┌─────────────────┐ ┌─────────────────┐
│ Amazon Bedrock │ │ Together AI │
└─────────────────┘ └─────────────────┘
│ │
└──────────┐ ┌──────────┘
▼ ▼
Open model estateWith that in place the decisions become tractable. A sensitive workload requiring explicit AWS regional processing routes to Bedrock. A coding agent where DeepSeek V4 Flash gives better economics routes to Together. A workload where Kimi K3 genuinely outperforms everything else uses Kimi K3. When a model becomes cheaper somewhere else next month, or a gated release finally opens, you move. The model decision and the infrastructure provider decision should be two separate decisions, and that is far more practical now that Together exposes OpenAI compatible APIs and the newer Bedrock Mantle endpoint also supports an OpenAI compatible interface for supported models.
12. The most dangerous question is “what models does Bedrock support?”
That sounds like a sensible enterprise architecture question, and I have come to believe it is the wrong one. The better question is what the best model for this workload is today, and where we are prepared to execute it, because those are two separate governance decisions.
A model can be approved without every inference provider being approved. A provider can be approved for one data classification and refused for another. A European workload might require a European dedicated deployment while a low sensitivity coding workload runs perfectly happily on Together’s US infrastructure and a highly sensitive banking workload stays inside AWS behind PrivateLink. That is actual risk management. Saying “we use Bedrock” is not an AI strategy, it is a procurement decision wearing the costume of an architecture.
13. Bedrock’s greatest strength could become its weakness
The AWS value proposition is that it does the integration work on your behalf, which is immensely useful when the market moves at AWS speed. AI does not move at AWS speed. DeepSeek V4 Flash arrived on 30 July, Kimi K3 on 27 July, GLM 5.2 in June and GLM 5.3 in August, while Bedrock remains on DeepSeek V3.2, Kimi K2.5 and GLM 5.
The problem is not that the older models suddenly became bad, because they did not. The problem is that the abstraction designed to give you access to AI can quietly become the thing preventing you from accessing the most interesting AI, and the cost is measurable. In the DeepSeek example the newer model on Together costs roughly a fifth of the older model on Bedrock. In the Kimi example the newer model costs five times more and offers an entirely different capability envelope. Different models, different capabilities, different economics and real choice is exactly what a functioning technology market should look like.
14. The AI control plane now matters more than the cloud control plane
AWS still has the better cloud control plane, and its networking, IAM, regionality and integration remain exceptionally difficult to reproduce. Together increasingly looks like it may have the better open AI control plane, with direct relationships to the model builders, rapid access to new releases, a growing specialised GPU estate and a stack running from serverless inference through to dedicated clusters. Neither of them, though, has a control plane that reflects your risk appetite, your data classifications or your evaluation results, because that is not their job. It is yours.
So use Bedrock where Bedrock is best, use Together where Together is best, self host where that makes sense, and make sure the application never has to care which of those is true this quarter. In 2026 the model that wins your evaluation in February may not make your shortlist in August. Your cloud provider should not get to decide which intelligence your company is allowed to evaluate.
15. References
- Amazon Bedrock model documentation for DeepSeek V3.2, Kimi K2.5 and GLM 5.
- Amazon Bedrock pricing.
- Amazon Bedrock geographic inference and AWS PrivateLink documentation.
- Together AI model catalogue: https://www.together.ai/models
- Together AI pricing: https://www.together.ai/pricing
- Together AI GLM 5.2 model page: https://www.together.ai/models/glm-52
- Together AI GLM 5.3 model page, currently listed as launching soon: https://www.together.ai/models/glm-5-3
- Together AI serverless inference: https://www.together.ai/serverless-inference
- Together AI model information for DeepSeek V4 Flash, DeepSeek V4 Pro and Kimi K3.
- Together AI and Moonshot partnership announcement.
- Together AI privacy, security and third party model hosting documentation.
- Together AI GPU infrastructure, European expansion and provisioned throughput material.
- Z.ai GLM 5.3 documentation: https://docs.z.ai/guides/llm/glm-5.3
- Z.ai GLM 5.3 staged release and cyber capability coverage: https://www.axios.com/2026/08/14/china-open-source-ai-glm-53
- GLM 5.3 release analysis: https://venturebeat.com/technology/glm-5-3-is-here-with-advanced-cyber-capabilities-and-reportedly-already-found-a-serious-vulnerability-in-cursor