Thursday, 01 October 2026 PDT | 01:56 AM
The 1 News Alt Logo Text Smart News for Global Indians

Build a multi

Entertainment October 01, 2026 12:00 PM
Build a multi

As organizations move from single-purpose agents to multi-agent systems, the infrastructure requirements change. A lone agent handling customer queries can run in a serverless environment with short-lived sessions. But when you need three agents collaborating on a creative workflow that spans several days, sharing context and building on each other’s output, serverless sessions that cap at a few hours don’t cut it.

In this post, we walk through deploying a music production pipeline: One agent runs a generative audio model on the instance’s own GPU. The other two open the .wav file it wrote, off a shared volume. By the end, you will have a track you can play. You will also have learned how to create capacity providers, deploy agents from different artifact types, orchestrate agent-to-agent collaboration using shared sessions, and persist workflows across multiple days.

Amazon Bedrock AgentCore offers two compute options for hosting agents. MicroVMs are the serverless option: fast cold starts, session isolation, and consumption-based pricing. Runtime Instances are the new option: AWS managed EC2 infrastructure for persistent, long-running agent workflows. Both use the same runtime APIs, but Instances add multi-day sessions, GPUs, persistent volumes, and the ability to colocate multiple agents on a single instance.

How Runtime Instances differs from MicroVM

Both options support custom frameworks (CrewAI, LangGraph, LlamaIndex, Strands Agents), work with your choice of foundation model, integrate with MCP and A2A, and share the same AgentCore runtime APIs. The difference is in the underlying compute model.

An agent is a workload running within a session. Unlike the MicroVM model, where one runtime hosts one agent, a single Instances session can host multiple agents. When two agent runtimes share the same capacity provider, you can invoke them with the same runtimeSessionId to land both agents on the same EC2 instance. There, they share a filesystem and can collaborate on the same task.

Figure 1: The three-agent pipeline (Compose, Deliver, Screen) sharing one GPU instance and session

We will build a music production system that uses three specialized agents:

The workflow: a producer starts a track. The composition agent writes a brief and renders real audio on the instance’s GPU. The delivery agent opens that file, measures it, applies a chain it derived from those measurements, and measures again to prove the result landed on target. The compliance agent then re-measures independently, checks the delivery targets, and screens the audio against the studio’s back catalog. If the screen flags a match, it calls back to the composition agent to generate an alternative. The producer ends up with a playable .wav and three reports explaining every decision.

What makes this possible on Runtime Instances:

From here on, this post is hands-on. You will prepare your AWS account, then run a three-agent pipeline that renders, generates, and clears a finished track, producing a .wav file you can play. Work through the steps in order. You will start by confirming the prerequisites, then define the three agents (Step 1), create a capacity provider that provisions the GPU instance and its persistent volumes (Step 2), deploy each agent as its own runtime (Step 3), and invoke them with a shared session ID so they can colocate on one instance and hand work to each other (Step 4). Finally, Step 5 shows how any one team can ship a new version of its agent without disturbing the others. The complete sample is in the AgentCore samples GitHub repository.

Before you begin, make sure you have:

For the complete working code, see the AgentCore samples repository on GitHub.

AgentCore Runtime Instances supports any agent framework. In this sample, each agent is a Python application built with Strands Agents.

Two details are commonly misconfigured, and both fail confusingly:

The SDK dispatches on the parameter name (it checks params[1] == "context"). That is the only way to read the session ID, which the agents need in order to find each other’s files and to call one another.

Second, build the Agent inside the handler, not at module scope:

A module-level Agent is shared across concurrent requests, and Strands rejects re-entrant invocation with Agent is already processing a request. History lives on the volume through FileSessionManager, which is how a session resumed days later remembers earlier decisions.

The composition agent turns the producer request (prompt) into a musical brief, using Claude Sonnet 4.6, then renders the actual audio with the ACE-Step foundation model:

The mode=prepare call builds the stack onto the volume, a virtualenv with CUDA PyTorch and ACE-Step. The render runs as a subprocess under that volume’s interpreter:

The delivery agent reads the rendered track off the shared filesystem and measures it. It uses Claude Sonnet 4.6 to choose EQ bands, compressor settings, and what to leave alone. The digital signal processing (DSP) is then applied. Then the output is measured again, so the plan is checked rather than trusted.

The compliance agent independently re-measures the finished delivery, checks it against the delivery targets the delivery agent claimed, and screens it for harmonic similarity against the studio’s own back catalog. When it finds a similarity, it raises a flag and calls back to the composition agent for remediation.

A capacity provider tells AgentCore what compute infrastructure to provision for your agents. You specify instance types and VPC placement. AgentCore handles provisioning and lifecycle management.

You need two IAM roles, and the distinction matters:

Both trust bedrock-agentcore.amazonaws.com.

Now create the capacity provider. Note that names must use underscores (hyphens are not allowed):

⚠️ GPU capacity. If you hit InsufficientInstanceCapacity across multiple Availability Zones (AZs) trying to allocate a GPU instance, you can switch to another GPU instance type, like g5.xlarge. Update allowedInstanceTypes accordingly.

Each agent here gets its own runtime. The runtimes are brought together at invoke time. When two runtimes share a capacity provider and you invoke them with the same runtimeSessionId, AgentCore places both agents on the same EC2 instance, where they share a filesystem and can collaborate on the same task. That’s how the delivery agent reads the .wav the composition agent wrote.

Next, point each runtime at the capacity provider and declare which volumes it mounts:

The compliance agent is the same call with a different artifact, a zip rather than an image, and only the workspace volume:

Now invoke the agents. This is where the three agents become a pipeline: you pass the same runtimeSessionId to each invocation. The first invocation is the slow one because it provisions the instance. Every call after that routes to the instance already running.

Here’s what that produces on a live g6.xlarge in us-east-2:

When you call StopRuntimeSession, the instance idles down automatically. No compute charges accrue while idle. When you invoke the session again, AgentCore resumes it, provided it lands in the same Availability Zone. Amazon Elastic Block Store (Amazon EBS) volumes are AZ-locked. If the original AZ is capacity-dry, volumes can’t reattach and persistence is lost. Use an AZ-pinned ODCR or MODELS_SNAPSHOT_ID so a re-placed resume can recover. Sessions can persist for up to 14 days.

One of the strengths of this architecture is independent deployment. When the Audio AI team ships a new version of the composition agent, they update only their runtime:

The delivery and compliance agents continue running unchanged. No coordination needed, no shared deploy pipeline, no risk of breaking another team’s agent with your update.

Delete the session first. It’s the most direct way to stop EC2 and Amazon EBS charges. Deleting a session deprovisions the EC2 resources: instance, network interface, and Amazon EBS volume. Deleting the capacity provider also stops and deletes its associated sessions and their persistent storage. However, you must first disassociate every runtime and runtime version from it, and that detachment is asynchronous. Session deletion is the fast path, and the one to reach for if you want to avoid ongoing charges.

Then delete the ECR repositories, S3 artifacts, IAM roles, and Amazon CloudWatch log groups.

In this post, we deployed a multi-agent music production system on Amazon Bedrock AgentCore Runtime Instances. Three agents, built by different teams using different packaging formats, collaborated within one session on a single GPU instance, and produced a track you can play.

The architecture demonstrates several patterns that apply beyond music production:

Music was a convenient vehicle, but nothing about the architecture is musical. Swap out the render step and the same three-agent shape fits the workloads AWS calls out for GPU Instances: 3D rendering, simulation, model inference, media processing. Or a long-running job where a pipeline produces a large artifact, hands it to a second agent to transform, and has a third check the result before it ships.

To learn more, see the Amazon Bedrock AgentCore documentation. For the complete working code from this post, explore the AgentCore samples repository. For advanced orchestration patterns like Graph, Swarm, or Workflow, see the Strands Agents documentation.