Building enterprise voice agents used to mean stitching together speech services, real-time models, turn detection, and tool orchestration—then owning every failure mode. Voice agents in Foundry Agent Service consolidate a multi-service voice solution into a versioned, managed agent. They reduce runtime engineering, but identity, authorization, data, networking, retention, and operations remain enterprise responsibilities.
Microsoft's launch announcement explains the product capabilities. This article asks a different question: should the feature be the starting point for an enterprise voice architecture? It applies the six criteria from my earlier architecture comparison and adds a 300-turn real-audio study across six configurations.
In this article, you will learn:
- what Foundry manages and what remains the solution team's responsibility;
- how to assess the data, network, identity, cost, and operational boundaries;
- what a 300-turn real-audio benchmark reveals about latency across six configurations.
As of 29 September 2026, voice-based agents remain in public preview. Product statements were checked on that date; measurements were collected on 26 September 2026.
Executive assessment
A voice agent now packages the model, instructions, audio configuration, turn detection, tools, interim responses, and storage policy behind a managed real-time endpoint. The solution team can focus less on assembling the runtime and more on governing the complete workload.
Foundry owns much more of the voice runtime, while the customer remains accountable for identity, authorization, data placement, network design, retention, and operational fitness.
| Criterion | Managed or provided by Microsoft | Customer responsibility |
|---|---|---|
| Features and effort | Managed real-time voice runtime, speech and model orchestration, turn detection, agent versioning, tool orchestration, channels, tracing, monitoring, and evaluation surfaces. | Client and channel integration, business workflows, tool implementation, caller controls, governance, regression testing, capacity planning, and incident management. |
| Data residency | Published service and model geography commitments, platform encryption, and documented storage and retention behavior. | Select deployment types and regions; map speech, model, storage, tools, telemetry, and channel processing; configure retention and document compliance. |
| Private networking | Agent Service network-isolation capabilities, private endpoints, and integration with customer-managed Storage, Azure AI Search, and Azure Cosmos DB. | Client connectivity, private DNS, routing, tool access, telemetry and telephony egress, firewall policy, and end-to-end reachability testing. |
| Authentication and authorization | Microsoft Entra authentication for supported access patterns, managed and agent identities, role-based access control, and delegated-identity primitives. | Caller sign-in and admission, per-user data authorization, tool permissions, delegated access, approval and step-up controls, and duplicate-action protection. |
| Cost | Service meters, usage signals, estimated-cost telemetry, and Azure Cost Management billing records. | Select model hosting and capacity, measure the complete workload, set budgets and forecasts, reconcile charges, and optimize usage. |
| Latency | Managed runtime plus service-side latency metrics, traces, and monitoring views. | Model, voice, turn-detection, tools, retrieval, network, and channel choices; real-audio and concurrency tests; user-facing latency objectives and startup experience. |
Connection startup: Establishing a new voice-agent connection took 3.72–12.37 seconds in this test. Connect before the conversation begins, or show the user that the agent is still connecting.
What is architecturally new?
The new API represents the experience as a VoiceAgentDefinition. Its kind is voice, and every create or update operation produces an immutable agent version behind a stable agent name. The definition selects:
- model hosting, model, instructions, and an optional greeting;
- input audio, transcription, enhancement, and turn detection;
- output audio, voice, language, speed, and optional avatar;
- tools, tool policy, interim responses, and conversation storage.
The runtime is reached over a persistent real-time connection. Foundry Agent Service owns the agent lifecycle and tool orchestration; Voice Live supplies the speech runtime.
A managed voice runtime reduces implementation work. It does not merge the identity, data, networking, and retention boundaries around it.
One feature, two model paths
Foundry derives the voice architecture from the selected model.
Native speech-to-speech
caller audio -> realtime speech model -> spoken response
This path is intended for natural, low-latency conversation. The model consumes and emits audio directly. Available voice and audio controls depend on the selected model.
Cascaded text model
caller audio -> speech recognition -> text model -> speech synthesis -> spoken response
This path provides broader text-model and Azure voice choices, plus explicit transcription, phrase-list, and speech-synthesis controls. It also places serial stages on the critical latency path.
Service-hosted (managed) versus customer-deployed (self_deployed) is a separate decision. It controls where the model is hosted and metered. The model itself determines whether the resulting path is native or cascaded.
Criterion 1 — Features and implementation effort
This is the feature's clearest architectural value. Foundry combines real-time sessions, turn taking, transcription, voices, agent versions, models, knowledge, tools, channels, traces, monitoring, evaluation, and optional avatars. The portal supports no-code testing, while client libraries and Azure Developer CLI make definitions deployable from source control.
What remains customer-owned
An enterprise still owns the client or approved phone channel, caller admission and abuse controls, user-specific retrieval and action authorization, approval and duplicate-action protection for consequential tools, human handoff, retention policy, regression testing, capacity, and incident management.
Channel support is narrower than generic agent publishing. The current voice-based Channels experience exposes a Preview web app and phone-number integrations through Teams Phone extensibility or Twilio. Standard Microsoft Teams or Microsoft 365 agent publishing is not a substitute for Teams Phone extensibility.
Criterion 2 — Data residency
Data residency cannot be answered with one region field. A voice turn crosses multiple processing and storage surfaces:
| Surface | Residency question |
|---|---|
| Speech and model | Where are turn detection, recognition, synthesis, and model inference processed? Is the model global, data-zone, or regional? |
| Agent state and conversations | Where is service state stored? Is store enabled for transcripts and raw audio? |
| Knowledge and tools | Where are files, indexes, retrieved passages, tool requests, and outputs processed? |
| Observability | Where does Application Insights store telemetry, and is sensitive-content capture enabled? |
| Channel | What additional processing is introduced by Azure Communication Services, Teams Phone, or Twilio? |
The Voice Live data-privacy documentation says that Voice Live itself does not retain customer data by default, but connected features can. If a customer opts into support logging, Microsoft can retain the relevant speech data in the resource region for up to 30 days. Agent Service documentation places data stored by stateful service features at rest in the Azure OpenAI resource geography; model inference and tools follow their own configuration and hosting location.
For a self_deployed configuration, model inference follows the selected Global, Data Zone, or regional deployment type; that choice does not determine the location of speech, storage, tools, or channels. A managed configuration uses a service-hosted model. The project region alone does not answer the full residency question; verify the current processing and storage commitments for the selected speech, model, storage, telemetry, and channel components in Microsoft documentation.
Criterion 3 — Private networking
Foundry Agent Service supports a network-secured Standard setup with a delegated subnet, private endpoints, customer-provided Azure Storage, Azure AI Search, and Azure Cosmos DB, and public network access disabled. The project managed identity receives data-plane access to those dependencies. That is the starting point, not proof that a voice solution is private end to end. Review these paths independently:
| Connection | What must be demonstrated |
|---|---|
| Client to voice-agent endpoint | Private DNS, successful real-time connection setup, reachability, and idle-timeout behavior from the deployed client environment, or an authenticated public edge |
| Voice runtime to selected model | The exact managed or self-deployed route used by the chosen voice configuration |
| Agent to data and tools | Private endpoints, DNS, role-based access control, route, credentials, and successful runtime access |
| Telemetry and telephony | Approved Application Insights egress plus provider-specific media and signaling paths |
The Agent Service private-networking guide documents the general private setup. It does not make a private endpoint on one resource proof of every voice-specific managed hop.
Criterion 4 — Authentication and authorization
Voice-agent security is easier to reason about as four identity layers:
- Caller identity. Access to the hosted web experience can be assigned to organizational users and groups. A custom client still needs sign-in, session admission, rate limiting, and business entitlements.
- Session invocation identity. Public SDK and portal access patterns use Microsoft Entra ID for access to the Foundry project and agent endpoint. Phone-channel integrations add provider-specific authentication and connection setup: Teams Phone extensibility uses Azure Communication Services, Event Grid, and app-registration setup, while Twilio uses a project connection backed by Twilio credentials. Applications should use an appropriate workload identity and must not put long-lived credentials in browser code.
- Agent identity. The Foundry agent-identity model also applies to voice agents and authenticates downstream tools. Publishing an agent as a general Agent Application creates a distinct identity, so project permissions do not automatically transfer.
- Delegated user identity. For attended scenarios, OAuth on-behalf-of flows let a downstream service evaluate both the agent identity and the user's delegated permissions.
A valid token proves identity; it does not prove that a requested action is safe. An agent might be allowed to query an index while the caller can see only some documents, or it might be able to call a tool that the caller cannot authorize. Apply the caller's document permissions before content reaches the model. For consequential actions, require application policy, step-up verification where needed, spoken confirmation, and duplicate-action protection so a reconnect cannot repeat a side effect. Do not treat a recognized voice or a phone number alone as proof of identity.
Criterion 5 — Cost
There is now official pricing guidance for voice-based agents, but it is more useful as billing anatomy than as a durable price table.
Five documented factors drive cost:
- Audio input and output tokens.
- Managed versus self-deployed model hosting.
- Connected session duration.
- Tool calls and the services behind them.
- Optional features such as avatars and stored conversations.
The complete architecture can also include speech recognition and synthesis, custom voice hosting, Application Insights, Azure Storage, Azure AI Search, Azure Cosmos DB, telephony, application ingress, identity, and network infrastructure.
With model_type: managed, usage is billed as part of the voice agent. With model_type: self_deployed, model usage is billed against the customer's own Foundry model deployment under that deployment's standard or provisioned capacity pricing model.
Voice traces can expose token usage and estimated-cost attributes while a system is being tuned. Microsoft explicitly states that these are estimates, not invoice records; Azure Cost Management remains the billing source of truth.
Why this article has no dollar estimate
A single per-conversation estimate would create false precision because region, agreement, model, audio volume, session duration, tools, and optional services all change the result. Capture per-turn usage and session duration, reconcile actual charges in Cost Management, then model realistic call length, concurrency, transfers, and retries.
Include operational cost as well: the managed feature removes voice-runtime code, not governance, identity, networking, or channel operations.
Criterion 6 — Latency
The caller notices one latency measure: after I stop talking, how long until the agent starts speaking? The model path sets the processing stages, but model choice, turn-detection settings, network distance, tools, retrieval, response length, and channel buffering determine the result.
What was measured
The companion harness reran the three earlier patterns and the new feature through one real-audio test:
- Realtime API direct and Voice Live with your own model used the same customer-deployed
gpt-realtime-1.5model. - The earlier Voice Live + prompt-agent path and the new cascaded voice agent both used
gpt-4o-mini, the same Azure neural voice, and no tools. This is the closest architecture-controlled pair. - The new feature also ran with managed
gpt-realtime-2.1and cascadedgpt-5.
Every path received the same 1.33-second “Say hello briefly” recording at real-time speed, the same short-answer instruction, and the same requested server-side turn-detection settings, including a 500-millisecond silence wait. Each configuration ran 50 measured turns across five fresh sessions, with one excluded warm-up per session and randomized sequential order. Knowledge and tools were disabled, and the first-class agents did not store conversations.
The headline clock starts when the caller finishes speaking and stops when the first response audio reaches the client:
- Typical turn is the median: half the turns were faster and half were slower.
- 95% by is the slow-end result: only one turn in twenty was slower.
Time from when the caller finishes speaking until response audio reaches the client. Lower is better; results are specific to this test environment.
| Architecture and configuration | Typical turn | 95% started by |
|---|---|---|
Earlier pattern — Realtime API direct, gpt-realtime-1.5 | 1.03 s | 1.10 s |
Earlier pattern — Voice Live with your own model, gpt-realtime-1.5 | 1.24 s | 1.43 s |
Earlier pattern — Voice Live + prompt agent, gpt-4o-mini | 1.90 s | 2.37 s |
New feature — managed gpt-realtime-2.1 | 1.47 s | 1.61 s |
New feature — cascaded gpt-4o-mini | 1.31 s | 1.52 s |
New feature — cascaded gpt-5 | 3.45 s | 5.32 s |
All 300 measured turns completed with response audio and the expected transcription.
The same-model-and-voice cascade comparison is the most informative result. The first-class gpt-4o-mini voice agent started 0.59 seconds sooner on the typical turn and 0.85 seconds sooner at the 95% boundary than Voice Live with the no-tool prompt agent. That is an observed end-to-end difference between these configurations; it does not reveal which internal stage produced it.
Realtime API direct was fastest in this environment. Voice Live with your own model was 0.21 seconds behind on the typical turn and 0.32 seconds behind at the 95% boundary. Both used the same realtime deployment, but the direct path used its model-native voice while Voice Live used an Azure neural voice. The difference therefore measures the configured stacks, not Voice Live overhead in isolation.
Within the new feature, managed gpt-realtime-2.1 and cascaded gpt-4o-mini were close enough that this run should not establish a permanent ordering between them. Changing the cascade to gpt-5 had a much larger effect, raising the typical result from 1.31 to 3.45 seconds. Model choice remains part of the architecture decision.
Why service monitoring can show sub-second results
The Overall latency chart in the voice-agent monitoring view in Foundry Agent Service measures time to first audio after voice activity detection decides the caller has finished. The benchmark chart starts earlier—when the caller actually stops speaking—and therefore includes the configured half-second silence wait.
Using that later service boundary:
- Managed speech-to-speech was typically 0.88 seconds, with 95% of turns starting by 1.02 seconds.
- The
gpt-4o-minicascade was typically 0.78 seconds, with 95% starting by 0.99 seconds. - The
gpt-5cascade was typically 2.92 seconds, with 95% starting by 4.78 seconds.
Both clocks are valid, but they answer different questions. The chart uses the caller's clock because it better represents the experience a user feels. The later service boundary explains how an operational metric can be sub-second even when caller-side time exceeds one second: the caller view adds the time required to detect the end of the turn and deliver audio across the network.
How this relates to the earlier measurements
The earlier article used ten warm, text-injected turns and excluded real speech and end-of-turn detection. Do not compare those values numerically with this chart. The unified rerun above is the cross-pattern comparison because every path uses the same audio, caller-side timing boundary, warm-up policy, sample count, region, network, and no-tool workload.
Test scope and reproduction
The test includes real speech, end-of-turn detection, model processing, and returned audio. It excludes device buffering, telephony, tools, retrieval, concurrency, private networking, and a production client, so its results describe this environment rather than a universal target.
For reproduction, see the unified benchmark harness, benchmark-only prompt-agent provisioner, first-class voice-agent provisioner, audio generator, raw turn data and method, and editable chart source.
Enterprise architecture checklist
For a workload decision, confirm:
- Platform fit: the region supports the selected model, voice, audio features, tools, and channel, and service quotas cover the expected concurrency.
- Data and network: every processing and storage location is recorded, and every route works from the intended client and network topology.
- Identity and authorization: callers are authenticated, retrieval respects their document permissions, and tool identity, delegation, approval, and duplicate-action protection are tested.
- Retention: conversation storage, trace capture, audio access, deletion, consent, and legal retention are approved.
- Performance: real audio, tools, target network, concurrency, and objectives for both typical turns and the slowest 5% are tested.
- Operations: quotas, errors, reconnects, handoff, dependency outages, monitoring, and rollback are rehearsed.
Verdict
For a new enterprise voice solution on Microsoft Foundry, voice agents in Foundry Agent Service are a strong primary candidate to evaluate when preview status is acceptable. They make voice a managed, versioned agent lifecycle and bring the runtime, channels, traces, monitoring, and evaluation together. That makes them a simpler starting point than rebuilding those layers by default.
This is not a claim that the feature is always faster or cheaper. Direct Realtime was fastest in this test, while the gpt-5 result showed how strongly model choice can dominate the outcome. Nor does the managed feature remove the surrounding trust boundaries: identity, authorization, data placement, networking, retention, and operational fitness still require an enterprise design.
The earlier patterns therefore remain valid, but for a new build they should be deliberate alternatives rather than the default starting point:
Begin by evaluating the voice-agent feature in Foundry Agent Service. Adopt it when its preview status, region, model, voice, channel, identity, data, network, cost, and measured performance fit the workload. Choose Realtime API or lower-level Voice Live patterns when a documented requirement needs control, capability, or a processing boundary that the managed definition does not provide.
For the lower-level alternatives and the trade-off between owning the runtime and consuming managed layers, see the earlier three-pattern architecture comparison. For Microsoft's feature overview and product direction, read the launch announcement.
Sources
- Introducing voice agents in Microsoft Foundry
- Agents in Microsoft Foundry
- Create a voice-based prompt agent
- Configure a voice agent
- Voice-agent design best practices
- Voice-agent tracing, monitoring, and evaluation
- Pricing for voice-based agents
- Foundry Agent Service regions and limits
- Private networking for Foundry Agent Service
- Agent identity concepts
- Data, privacy, and security for Voice Live
- Data, privacy, and security for Agent Service
- Publish and share a voice-based agent
- Integrate telephony channels with a voice agent
