• Docs
  • Talk to an expert
Blog
Blog
BlogProductCase studiesNewsInsights
Blog

Your AI stack will change again. Stop rebuilding it.

AI EngineeringUpsun DispatchAgentic SDLCAI Agents
29 September 2026
Share

TL;DR 

  • The pattern: the model your team relies on today is unlikely to be the one you're relying on a year from now. 
  • The mistake: building your workflow around a specific model, a specific coding tool, or a specific vendor's agent, then rebuilding it every time one of those changes. 
  • The bet: build the layer that stays the same, the workflow, the gates, the audit trail, and let the model underneath it be swappable by design.

The model your team relies on today is unlikely to be the one you're relying on a year from now. If your team's process for shipping AI-assisted code is built around a specific model, coding assistant, or a vendor's take on an autonomous agent, you are not building infrastructure. You are building something you will tear out and rebuild the next time the leaderboard shifts.

This has already happened to almost every team using AI to write code, and it will happen again, because the thing that changes fastest in this stack is precisely the part most teams have built the least flexibility around.

Why the churn is structural, not incidental

Key takeaway: Model leadership rotates on a timescale of months, not years. Any workflow hardwired to a specific model inherits that instability directly.

Frontier model releases are not slowing down, and neither is the reshuffling of which one is actually best for a given task. A model that leads on raw reasoning benchmarks is not always the one that's cheapest for a routine PR review, and the model that's cheapest today is not guaranteed to still be cheapest once a competitor cuts pricing or ships a faster variant. Coding-specific fine-tunes appear, get adopted, and get superseded on a similar cadence.

None of this is a criticism of any specific vendor. It's a description of a market still finding its shape. The mistake is treating today's leaderboard position as a permanent architectural decision. It's a temporary condition your infrastructure should be built to absorb.

The "could I have built this" test

Key takeaway: If the honest answer to "could my team have built this" is yes, but you rationally chose not to, that's the correct reason to adopt a platform. If the answer is no, that's a red flag, not a selling point.

Any engineer evaluating a new layer in their stack should ask a specific question: could I have built this myself? The real question isn't whether it's magic. It's whether your team could stand this up given enough time, and if so, why you'd choose not to.

For a sandbox runtime, the honest answer is usually yes. You could build container isolation yourself, and several teams have. The reason not to is that it's commodity work: every reasonable implementation ends up "quite similar in capabilities," to borrow language from an engineer who actually prototyped across five different sandbox providers before picking one.

The isolation layer is commodity work. What's worth building is the evaluation framework itself. It's the thing that tells you which model earns its keep, for which task, at what cost. And because the market keeps shifting, it has to stay current too.

Buy the commodity, build the judgment

Key takeaway: The layer worth building yourself is the one that requires ongoing judgment. The layer worth buying is the one where every reasonable option behaves about the same.

This is the actual architectural principle, not "avoid vendor lock-in" as an abstract virtue, but a specific rule for deciding what belongs where in your stack. Sandbox runtimes, container isolation, raw compute: these are commodities. The differences between providers are real but marginal, and betting your architecture on any single one being permanently superior is a bet you will lose eventually, just not predictably.

Model access is the same category. Whichever frontier lab is ahead this quarter is a fact about this quarter, not a fact about your architecture. Building a proprietary model-access layer from scratch, rather than treating "bring your own key, swap it when you need to" as the default, means you've built your infrastructure's stability on top of the single most unstable variable in the stack.

What's actually worth building in-house, or worth demanding from whatever platform you adopt, is the layer that makes swapping cheap. This includes the workflow definition, the human approval gates, the audit trail, and the routing logic that decides which model handles which task based on cost and quality rather than habit. This is judgment. It should be yours, whether you build it or a platform builds it for you, because it's the part that has to keep adapting.

It's worth being honest about what that actually costs. Someone still has to evaluate models against real tasks, track when a cheaper or better option shows up, and keep that judgment current as the market shifts. That's ongoing work, not a one-time setup. The bet isn't that the work disappears. It's that a platform built for this can absorb it as a routing decision rather than a rebuild. That's a smaller, cheaper problem to keep solving than the alternative.

What this looks like in practice

Key takeaway: A workflow layer that is genuinely model-agnostic means a model change is a configuration update, not a migration project.

If your workflow, the sequence of steps, the gates, the logged record of what happened and why, is defined independently of which model executes each step, then a model change is a routing decision, not an architecture change. You update which model handles a given task, not the process the task runs through.

This is the practical test for whether a platform is actually solving the churn problem or just deferring it. If adopting a new model, or a better-priced one, or a faster one, requires touching your review process, your approval chain, or your audit tooling, then the platform has quietly recreated the lock-in it claimed to avoid. If it only requires a routing configuration change, it's actually built for the reality that the model underneath will keep shifting.

Upsun Dispatch is built on this specific bet: that workflows, not models, are the unit worth designing around. Bring whichever model your team already uses, provided once during setup. When something better ships, you're not touching your review process, your approval chain, or your audit tooling to take advantage of it. You're changing which model is connected. That's the whole point of decoupling the workflow from the model in the first place.

Read the Upsun Dispatch docs to see how the workflow layer stays stable while the model underneath it doesn't have to.


 

Frequently asked questions (FAQ)

Why does model choice change so often in AI-assisted development?  
Because coding-specific model performance and pricing both shift on a cadence measured in months, driven by new frontier releases, competitive pricing pressure, and fine-tunes optimized for specific coding tasks. There is no stable long-term leader to standardize around, only a current one.

What does "buy the commodity, build the judgment" actually mean?  
It means recognizing which parts of your AI stack behave similarly across any reasonable provider (sandbox isolation and raw compute are common examples) and which parts require ongoing, defensible reasoning, like which model to route a given task to. For that second category, you need to build or acquire your own capability.

Isn't building your own model-routing layer just adding complexity?  
Only if it's built as a one-time decision instead of an ongoing evaluation. A routing layer that's periodically re-evaluated against real cost and quality data is less complex over time than a hardcoded model choice that eventually forces a full workflow rebuild when it becomes outdated.

How do you evaluate whether a workflow is genuinely model-agnostic?  
Ask what has to change when you swap models. If the answer is "a configuration setting," the workflow is genuinely model-agnostic. If the answer involves touching review processes, approval logic, or audit tooling, the model choice was never actually decoupled from the architecture.

Stay updated

Subscribe to our monthly newsletter for the latest updates and news.

Deploy possibility.
Try Upsun for free.

Build with DispatchDeploy on Cloud