
An AI agent that works on a laptop often struggles in production. There, it has to keep state between steps, survive restarts, run code safely, and leave a record of what it did. The platform underneath decides how much of that work your team does itself.
There is no single best platform for running AI agents. The right choice depends on what the agent does and where your systems already live:
Most production setups combine two of these layers. This guide explains each layer, compares the leading platforms, and shows when each one is the right call.
Platform | Layer | Best for | Notable fact |
| Amazon Bedrock AgentCore | Hyperscaler runtime | Teams standardized on AWS | One microVM per session, up to 8 hours |
| Microsoft Foundry Agent Service | Hyperscaler runtime | Teams standardized on Azure | VM-isolated sessions that resume after idling |
| Gemini Enterprise Agent Platform | Hyperscaler runtime | Teams standardized on Google Cloud | Built-in sessions and long-term memory |
| Claude Managed Agents | Model vendor runtime | Teams building on Claude | $0.08 per session-hour plus tokens |
| LangSmith Deployment | Framework runtime | LangGraph agents | Cloud, hybrid, or self-hosted |
| E2B | Sandbox | Short, isolated code runs | Firecracker microVMs; 24-hour cap on Pro |
| Daytona | Sandbox | Fast, stateful sandboxes | Claimed starts under 90 ms |
| Modal | Sandbox | Python agents that need GPUs | Sandboxes run up to 24 hours |
| Upsun Cloud | Application platform | Agent APIs with their own data | Data-complete preview environment per branch |
| Railway and Render | Application platform | Simple, small deployments | Low setup effort |
| Trigger.dev | Durable execution | Multi-step TypeScript tasks | No task timeouts |
| GitHub Copilot cloud agent | Governed workflow | Teams inside GitHub | 59-minute limit per task |
| Upsun Dispatch | Governed workflow | Issue-to-PR agents with audit | Immutable per-run log of plan, approver, and cost |
Running an agent in production means giving it a safe place to act, a way to remember, and a way to recover. A framework such as LangGraph or the Claude Agent SDK defines how the agent thinks. The platform decides where that logic runs and what happens when something goes wrong.
That difference explains why the market has split into layers. Some platforms run the agent loop. Some isolate the code an agent writes. Others keep long tasks alive through failures. A growing group controls what agents may do inside real systems.
The platforms in this guide were compared on six questions:
Each layer below solves a different part of the problem. The entries explain what each platform does well and where it stops.
These suit teams whose data, identity, and budgets already sit with one cloud provider. The trade-off is lock-in to that provider.
These suit teams that want to hand over the agent loop itself. The trade-off is closer ties to one vendor's model or framework.
A sandbox is the safe room where an agent runs code it has just written. Most teams pair one with an application platform or a hyperscaler runtime.
Many agents are long-lived services. They expose an API, hold conversations, and query their own data. These platforms run that service next to its databases.
These keep a multi-step agent task alive through crashes, retries, and waits for human input. They sit on top of a hosting platform rather than replacing one.
This layer covers agents that change your codebase. Here the main question is control: who the agent acts as, who approves its work, and what record remains.
Once an agent can open pull requests, call APIs, or read customer data, the key question changes. It is no longer whether the agent can run. It is who the agent acted as, who approved its work, and what evidence remains.
Four controls answer those questions:
These controls get harder to keep as agent use grows. Teams usually adopt agents one project at a time. The result is sprawl: different agents on different providers, each with its own pipeline, runbook, and audit process. A review then means collecting evidence from several consoles by hand.
For that reason, governance deserves the same weight as cold-start time when choosing a platform. The goal is consistent audit evidence across providers without manually reconciling separate consoles.
The right call follows from your main constraint. Start with the row that matches it, then add a second layer only when a real need appears.
| Choose this when | Start with | Why it fits |
| Everything already runs on AWS | Amazon Bedrock AgentCore | Native identity, billing, and per-session microVMs |
| Everything already runs on Azure | Microsoft Foundry Agent Service | Entra ID per agent and stateful sessions |
| Your agents need GPUs | Modal or Northflank | Both offer GPU compute; Upsun Cloud does not |
| Your agents run untrusted code | E2B or Daytona | Fast, strongly isolated sandboxes |
| Your agent is an API with its own data | Upsun Cloud | Managed databases and data-complete preview environments per branch |
| You want the simplest hosting for a small project | Railway or Render | Low setup effort |
| Multi-step tasks must survive failures | Trigger.dev or Temporal | Durable execution with retries |
| Your team lives in GitHub and uses Copilot | GitHub Copilot cloud agent | Built into the tools the team already uses |
| Agents change code across GitHub, Jira, and Linear, and you need per-run audit and cost | Upsun Dispatch | Verified identity, human-owned merges and an immutable run log |
What is the difference between an AI agent framework and an AI agent platform?
A framework, such as LangGraph or the Claude Agent SDK, is a code library that defines how an agent reasons and uses tools. A platform is the infrastructure that runs that code, stores its state, and controls what it can reach.
Do AI agents need a sandbox?
An AI agent needs a sandbox when it runs code it generated itself. The sandbox keeps a faulty or manipulated command away from your systems and data. Agents that only call fixed APIs can often run without one.
Do AI agents need GPUs?
Most AI agents do not need GPUs, because they call hosted model APIs for inference. GPUs matter when a team runs its own models. Modal and Northflank offer GPUs; Upsun Cloud does not.
How long can an AI agent run?
Maximum run times differ widely. GitHub Copilot cloud agent stops at 59 minutes, AgentCore microVM sessions at eight hours, and Modal and E2B Pro sandboxes at 24 hours. Trigger.dev tasks have no timeout.
Can AI agents run in my own cloud account?
Yes. Hyperscaler runtimes run inside your cloud account by design. LangSmith Deployment and Trigger.dev can be self-hosted, and Northflank can deploy into your own cloud account.
What does it cost to run an AI agent?
Running an AI agent costs model tokens plus compute, and platforms charge for compute in different ways. Claude Managed Agents charges $0.08 per session-hour. Microsoft and AWS bill for the CPU and memory each session uses. E2B's Pro plan costs $150 per month.
How can I audit what an AI agent did?
Choose a platform that records each run's trigger, plan, approver and cost, tied to a verified identity. Without that record, reviewing an agent's actions means rebuilding events from scattered logs.