• Docs
  • Talk to an expert
Blog
Blog
BlogProductCase studiesNewsInsights
Blog

Best platforms for running AI agents in 2026

AIAI AgentsUpsun Dispatchdeveloper workflow
06 October 2026
Share

An AI agent that works on a laptop often struggles in production. There, it has to keep state between steps, survive restarts, run code safely, and leave a record of what it did. The platform underneath decides how much of that work your team does itself.

There is no single best platform for running AI agents. The right choice depends on what the agent does and where your systems already live:

  • Your stack runs on one hyperscaler: use its managed agent runtime, such as Amazon Bedrock AgentCore, Google's Gemini Enterprise Agent Platform, or Microsoft Foundry Agent Service.
  • You want the model provider to run the agent loop: use a model vendor's managed runtime, such as Claude Managed Agents.
  • You built the agent with LangGraph: use LangSmith Deployment.
  • The agent runs code it writes: add a sandbox such as E2B, Daytona, or Modal.
  • The agent is a service with its own database and API: use an application platform such as Upsun Cloud, Railway, or Render.
  • The agent must recover from failure mid-task: use a durable execution engine such as Trigger.dev or Temporal.
  • The agent works on your codebase: use a governed agent workflow such as Upsun Dispatch or GitHub Copilot cloud agent.

Most production setups combine two of these layers. This guide explains each layer, compares the leading platforms, and shows when each one is the right call.

AI agent platforms compared at a glance

Platform

Layer

Best for

Notable fact

Amazon Bedrock AgentCoreHyperscaler runtimeTeams standardized on AWSOne microVM per session, up to 8 hours
Microsoft Foundry Agent ServiceHyperscaler runtimeTeams standardized on AzureVM-isolated sessions that resume after idling
Gemini Enterprise Agent PlatformHyperscaler runtimeTeams standardized on Google CloudBuilt-in sessions and long-term memory
Claude Managed AgentsModel vendor runtimeTeams building on Claude$0.08 per session-hour plus tokens
LangSmith DeploymentFramework runtimeLangGraph agentsCloud, hybrid, or self-hosted
E2BSandboxShort, isolated code runsFirecracker microVMs; 24-hour cap on Pro
DaytonaSandboxFast, stateful sandboxesClaimed starts under 90 ms
ModalSandboxPython agents that need GPUsSandboxes run up to 24 hours
Upsun CloudApplication platformAgent APIs with their own dataData-complete preview environment per branch
Railway and RenderApplication platformSimple, small deploymentsLow setup effort
Trigger.devDurable executionMulti-step TypeScript tasksNo task timeouts
GitHub Copilot cloud agentGoverned workflowTeams inside GitHub59-minute limit per task
Upsun DispatchGoverned workflowIssue-to-PR agents with auditImmutable per-run log of plan, approver, and cost

What does it take to run an AI agent in production?

Running an agent in production means giving it a safe place to act, a way to remember, and a way to recover. A framework such as LangGraph or the Claude Agent SDK defines how the agent thinks. The platform decides where that logic runs and what happens when something goes wrong.

That difference explains why the market has split into layers. Some platforms run the agent loop. Some isolate the code an agent writes. Others keep long tasks alive through failures. A growing group controls what agents may do inside real systems.

The platforms in this guide were compared on six questions:

  • Isolation: How strongly is one agent session separated from another, and from your systems?
  • Duration and state: How long can a session run, and does its state survive a restart?
  • Data and services: Can the agent sit next to its database, cache, and vector store?
  • Deployment model: Is it fully managed, or can it run in your own cloud account?
  • Governance: Does each run have an identity, an approval step, and an audit record?
  • Lock-in: How tied is the agent to one cloud, one model, or one framework?

 

The best platforms for running AI agents, by layer

Each layer below solves a different part of the problem. The entries explain what each platform does well and where it stops.

1. Hyperscaler agent runtimes

These suit teams whose data, identity, and budgets already sit with one cloud provider. The trade-off is lock-in to that provider.

  • Amazon Bedrock AgentCore gives each session its own microVM, which can run for up to eight hours. A new runtime version, released in September 2026, reports cold starts of about two seconds at the 75th percentile.
  • Microsoft Foundry Agent Service runs hosted agents in VM-isolated sandboxes, one per session. Each session keeps its files between turns and resumes after going idle. Hosted agents support Python and C#.
  • Google's Gemini Enterprise Agent Platform, formerly Vertex AI, provides a managed agent runtime with built-in sessions and long-term memory. It fits teams already building on Gemini and Google Cloud.

2. Managed runtimes from model and framework vendors

These suit teams that want to hand over the agent loop itself. The trade-off is closer ties to one vendor's model or framework.

  • Claude Managed Agents launched in public beta in April 2026. It provides sandboxed execution, sessions that can run for hours, and credential vaults. Pricing adds $0.08 per session-hour to standard token rates.
  • LangSmith Deployment, formerly LangGraph Platform, hosts agents built with LangGraph. It adds persistent state and can run in LangChain's cloud, in a hybrid setup, or fully self-hosted.
  • OpenAI is closing Agent Builder after 30 November 2026 and points developers to its Agents SDK, which teams host themselves.

3. Code-execution sandboxes

A sandbox is the safe room where an agent runs code it has just written. Most teams pair one with an application platform or a hyperscaler runtime.

  • E2B runs each sandbox in a Firecracker microVM that starts in about 150 milliseconds. Sandboxes last up to one hour on the free plan and 24 hours on the Pro plan.
  • Daytona gives each sandbox its own kernel, filesystem, and network stack, with a claimed start time under 90 milliseconds. Snapshots let state carry across sessions. Upsun Dispatch uses Daytona for its agent sandboxes.
  • Modal sandboxes run for up to 24 hours and can use GPUs. It is the natural choice for Python teams whose agents also need GPU compute.

4. Application platforms for agent services

Many agents are long-lived services. They expose an API, hold conversations, and query their own data. These platforms run that service next to its databases.

  • Upsun Cloud runs Python and Node.js agents alongside managed PostgreSQL with pgvector, Redis, Chroma, and Qdrant. Every Git branch can get a preview environment that clones production apps, services, and data. This lets a team test an agent change against realistic data before release. Upsun does not offer GPUs, so agents call external model APIs for inference
  • Railway and Render offer a similar service-and-database model and suit smaller teams that want a simple setup. Northflank adds GPUs and the option to deploy into your own cloud account. 

5. Durable execution engines

These keep a multi-step agent task alive through crashes, retries, and waits for human input. They sit on top of a hosting platform rather than replacing one.

  • Trigger.dev runs TypeScript tasks with no timeouts, retries, and queues. Tasks can pause until a person approves them. It is open source and can be self-hosted.
  • Temporal offers the same durability guarantees with SDKs for several languages, including Go, Java, Python, and TypeScript.

6. Governed agent workflows for software delivery

This layer covers agents that change your codebase. Here the main question is control: who the agent acts as, who approves its work, and what record remains.

  • Upsun Dispatch is built around workflows rather than individual agents. A workflow is a fixed sequence of steps that agents and people complete the same way on every run. Typical workflows turn an issue into a pull request when someone comments @upsun-dispatch implement, or review an open pull request. Runs can also start from a GitHub event, a schedule, or an API call. Dispatch works with GitHub, GitLab, Jira, and Linear. Each agent step runs in its own disposable sandbox, with outbound network access blocked by default. The agent cannot push a branch or open a pull request itself; a separate system step does that. Human gates pause the workflow wherever a person must decide, and a person owns every merge. Every run is logged as immutable data: the issue, the agent's context, the plan, who approved it, and what it cost. Teams bring their own model provider, and Dispatch runs on its own or alongside Upsun Cloud.
  • GitHub Copilot cloud agent works from issues, pull request comments, Slack, Jira, and Linear. It runs in a GitHub Actions environment, with a hard limit of 59 minutes per task. It is the simplest choice for teams that already use GitHub and Copilot. 

Why does governance matter once agents act on real systems?

Once an agent can open pull requests, call APIs, or read customer data, the key question changes. It is no longer whether the agent can run. It is who the agent acted as, who approved its work, and what evidence remains.

Four controls answer those questions:

  1. Identity: each run acts under a named, verifiable identity rather than a shared token.
  2. Approval: a person approves consequential actions, such as merging code or changing data.
  3. Audit: each run leaves a record of what triggered it, what it planned, and what it cost.
  4. Scoped access: the agent receives only the secrets and environment configuration its task needs.

These controls get harder to keep as agent use grows. Teams usually adopt agents one project at a time. The result is sprawl: different agents on different providers, each with its own pipeline, runbook, and audit process. A review then means collecting evidence from several consoles by hand.

For that reason, governance deserves the same weight as cold-start time when choosing a platform. The goal is consistent audit evidence across providers without manually reconciling separate consoles.

Which platform is the right call for your team?

The right call follows from your main constraint. Start with the row that matches it, then add a second layer only when a real need appears.

Choose this when

Start with

Why it fits

Everything already runs on AWSAmazon Bedrock AgentCoreNative identity, billing, and per-session microVMs
Everything already runs on AzureMicrosoft Foundry Agent ServiceEntra ID per agent and stateful sessions
Your agents need GPUsModal or NorthflankBoth offer GPU compute; Upsun Cloud does not
Your agents run untrusted codeE2B or DaytonaFast, strongly isolated sandboxes
Your agent is an API with its own dataUpsun CloudManaged databases and data-complete preview environments per branch
You want the simplest hosting for a small projectRailway or RenderLow setup effort
Multi-step tasks must survive failuresTrigger.dev or TemporalDurable execution with retries
Your team lives in GitHub and uses CopilotGitHub Copilot cloud agentBuilt into the tools the team already uses
Agents change code across GitHub, Jira, and Linear, and you need per-run audit and costUpsun DispatchVerified identity, human-owned merges and an immutable run log


 

 Frequently asked questions

What is the difference between an AI agent framework and an AI agent platform?

A framework, such as LangGraph or the Claude Agent SDK, is a code library that defines how an agent reasons and uses tools. A platform is the infrastructure that runs that code, stores its state, and controls what it can reach.

Do AI agents need a sandbox?

An AI agent needs a sandbox when it runs code it generated itself. The sandbox keeps a faulty or manipulated command away from your systems and data. Agents that only call fixed APIs can often run without one.

Do AI agents need GPUs?

Most AI agents do not need GPUs, because they call hosted model APIs for inference. GPUs matter when a team runs its own models. Modal and Northflank offer GPUs; Upsun Cloud does not.

How long can an AI agent run?

Maximum run times differ widely. GitHub Copilot cloud agent stops at 59 minutes, AgentCore microVM sessions at eight hours, and Modal and E2B Pro sandboxes at 24 hours. Trigger.dev tasks have no timeout.

Can AI agents run in my own cloud account?

Yes. Hyperscaler runtimes run inside your cloud account by design. LangSmith Deployment and Trigger.dev can be self-hosted, and Northflank can deploy into your own cloud account.

What does it cost to run an AI agent?

Running an AI agent costs model tokens plus compute, and platforms charge for compute in different ways. Claude Managed Agents charges $0.08 per session-hour. Microsoft and AWS bill for the CPU and memory each session uses. E2B's Pro plan costs $150 per month. 

How can I audit what an AI agent did?

Choose a platform that records each run's trigger, plan, approver and cost, tied to a verified identity. Without that record, reviewing an agent's actions means rebuilding events from scattered logs.

Stay updated

Subscribe to our monthly newsletter for the latest updates and news.

Deploy possibility.
Try Upsun for free.

Build with DispatchDeploy on Cloud