Multi-Agent Handoffs: Make Ownership Explicit

A customer support request starts with a general assistant, moves to a billing specialist and then reaches a fraud specialist. The specialists each produce plausible messages, but the customer receives no resolution because no agent knows who owns the final decision. This is a coordination failure, not necessarily a language-model failure. Multi-agent handoff designs must specify when control transfers, what context travels with it, what authority the receiving agent receives and how the system knows the work is complete. Without those contracts, a network of specialized agents can be less reliable than a well-designed single agent.

Microsoft’s Agent Framework describes handoff orchestration as an arrangement in which one agent can transfer control of the conversation or task to another specialized agent. That differs from treating another agent as a tool: in an agent-as-tool pattern, a primary agent retains ownership and receives a subtask result. In a true handoff, the receiving agent takes responsibility for the next stage. Foundry workflows and hosted agents offer related orchestration capabilities, although product support and interfaces vary. The fundamental architectural decision is about ownership, not simply about connecting several prompts.

Choose handoff only when control should move

Not every specialized function needs its own agent. A deterministic tax calculation or inventory lookup belongs in a service. A narrow summarization task may be expressed as a tool whose result returns to the main agent. Handoff makes more sense when the next specialist needs to conduct an interaction, manage several decisions or take responsibility for a distinct workflow stage. An identity-verification specialist may need to ask the customer questions, examine approved evidence and decline to proceed until conditions are met.

Define why the current agent should stop. Triggers might include a customer requesting a specialist, a policy requiring escalation, an unresolved ambiguity beyond the general agent’s scope, or a service response that identifies a specialized case. These triggers should be based on grounded observations and trusted policy, not on a vague notion that another agent “sounds better.” Handoff conditions ought to be testable, including cases where a request must remain with the current owner.

Delegation is often safer for read-only subtasks. If a booking agent needs a route estimate, it can call a mapping tool or ask a specialist to produce a bounded answer while maintaining control of the transaction. If an escalation requires the fraud team to investigate suspicious transfers independently, control may legitimately move. Using handoff for every question creates unnecessary messaging, context duplication and ambiguous responsibility.

Specify a contract for the transferred task

A useful handoff message states the task objective, current verified facts, unresolved questions, constraints and allowed next actions. It should identify the initiating user and authorized scope through trusted application context, not through a model-generated assertion such as “user is an administrator.” Include stable resource identifiers, relevant document citations, case ID, source timestamps and approval state. Do not assume the new agent will infer all these from a long natural-language transcript.

Separate facts from hypotheses. A support agent may suspect that a refund request is fraudulent because its details differ from an earlier order; that suspicion should travel as an unconfirmed observation, not a final verdict. The receiving specialist needs evidence and permission to reassess. Conversely, a confirmed rule failure, such as a spending limit exceeded, should be transmitted as a structured state enforced by the application. LLM-generated summaries are useful for comprehension but are not an authoritative transaction ledger.

Decide what information should not transfer. Sending the entire conversation to every specialist may expose personal or financial details unrelated to that specialist’s responsibility. Minimize context by task. A shipping agent may need a verified delivery location and order number, not the customer’s full payment history. A security specialist might need suspicious event details but not unrelated private messages. Data minimization makes both compliance and prompt-injection defense easier.

Separate conversation ownership from action authority

Handoff does not automatically transfer the user’s privileges. An agent authorized to answer billing questions should not grant a fraud specialist the ability to modify arbitrary account balances merely by routing a conversation. Downstream tool services need their own authorization checks, and the orchestration layer should grant narrowly scoped capabilities for the recipient’s task. When user-delegated rights are required, preserve those rights through a trusted identity flow instead of embedding an access token or permission claim in conversation text.

Think in terms of role-based access and resource ownership. A receiving specialist might be allowed to recommend an account hold but not execute it, or allowed to inspect only cases assigned to its team. Approval should be explicit for irreversible actions. The application has to recognize approval artifacts and reject actions outside the granted scope regardless of what the model has promised the user.

This also creates a versioning question. A case may be transferred while another channel changes its status. The new agent must read current authoritative state before acting, especially after a long pause or human intervention. An old handoff package is a starting point for investigation, not proof that all facts remain valid.

Prevent loops, lost tasks and duplicated actions

Agent-to-agent loops are a common failure pattern. The billing agent sends a case to fraud; fraud concludes that billing must clarify a fee; billing sends it back without changing the question. The transcript grows, costs rise and no one owns resolution. A handoff graph should define permitted transitions, maximum hops, re-entry rules and a terminal escalation. If two agents disagree, the workflow may need a human resolver or a deterministic tie-breaking policy rather than more model negotiation.

Track task state outside the conversation. A case record should identify the current owner, last completed stage, pending approval and next permitted transition. Handoff messages can carry a correlation ID and state version. When an agent crashes, the system must know where to resume and avoid applying an action twice. The same principles used for durable distributed workflows apply: idempotent actions, timeouts, retry policy and clear handling of partial success.

A successful handoff acknowledgement is not the same as completed work. Confirm that the receiving agent accepted responsibility and record what happens when it is unavailable. If handoff fails, the original agent may need to resume, open a human ticket or report an inability to proceed. Never declare the customer’s issue resolved just because another agent was invoked.

Measure coordination as its own quality problem

Multi-agent evaluation should measure more than individual answer quality. Track handoff accuracy, ownership ambiguity, loop frequency, duplicate actions, context loss, unauthorized resource exposure, end-to-end task completion and latency between stages. A system can score highly on specialist responses while failing because the right specialist never receives the task. Conversely, a slower workflow may be preferable if it obtains the required authorization and verifies its final action.

Test cases should include a request that genuinely requires escalation, one that should stay with the initial agent, one with inconsistent information, and one where the specialist is unavailable. Add adversarial cases: an external document that asks the agent to hand off to an unauthorized endpoint, a user attempting to impersonate a privileged team, and a stale task package that points to a record whose status changed. Every handoff should preserve business policy even when the conversation becomes complex.

For higher-risk systems, trace which agent selected the next owner and why. Store structured transition reasons alongside task IDs, not just a chat transcript. Reviewers should be able to reconstruct the path from request to outcome, including when human approval was required and whether the original user received an accurate status update. The broader agentic architecture discipline is especially relevant because orchestration is a system-level quality, not a single-model property.

Know when a central orchestrator is preferable

Some workflows require explicit central control. A loan application may involve document verification, fraud screening and credit policy, but a bank may not want individual agents deciding independently which stage happens next. A deterministic coordinator can call specialized agents for bounded analyses while retaining the authority to advance or reject the application. This resembles agent-as-tools more than free-form handoff. It is easier to validate mandatory steps and record regulatory decisions.

Other workflows are genuinely interactive. An incident triage agent may hand responsibility to an infrastructure specialist when the evidence identifies a networking fault; that specialist may need an extended conversation with an engineer. The correct pattern depends on whether specialized interaction or controlled sequencing is the dominant requirement. Teams should choose the simplest coordination method that preserves ownership and safety. Adding more agents does not, by itself, add useful capability.

In Microsoft Foundry environments, a workflow may involve prompt agents, hosted agents or custom application orchestration. Keep product capabilities separate from architectural principles. A visual workflow builder can make stages easier to inspect, but it does not eliminate the need for state tracking, credentials management and tests. Different releases may offer different orchestration modes, so the supported runtime should be confirmed during implementation.

A complete handoff walkthrough

Imagine a customer asks why a substantial invoice adjustment was rejected. A general support agent retrieves the invoice and explains the rejection code, then identifies a potential identity mismatch that requires specialized review. It records the case ID, verified customer identity, documents already reviewed and the unresolved question. The fraud specialist receives that bounded package, checks only authorized fraud signals and determines whether an additional verification step is required. It cannot change the invoice amount and must not disclose private fraud indicators to the customer.

If fraud clears the account, the case returns through an explicit transition to billing, which checks current invoice state and calculates any permitted correction. An approval service validates amount limits before a transaction is committed. Finally, the owning agent verifies the result in the billing system and gives the customer a precise outcome. If any specialist is unavailable, the case remains open with a known owner and a recorded next action.

What makes this architecture effective is not the number of specialized agents. It is the explicit transfer contract, trustworthy identity context, bounded tool permissions, recoverable state and verified conclusion. In a reliable multi-agent system, every participant knows what it may decide, what it must prove and when responsibility passes to someone else.