Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
162346 stories
·
33 followers

Who Wins The AI Assistant Wars? + Why Didn’t Google Build Muse?

1 Share

M.G. Siegler is the author of Spyglass. Siegler joins Big Technology to discuss the escalating battle to build the dominant personal AI assistant. Tune in to hear why he thinks Meta’s Muse is surprisingly strong and what will ultimately determine who wins. We also cover Apple’s dark-horse position, Anthropic’s absence from the consumer assistant race, Google’s repeated product missteps, and how AI agents could reshape subscriptions, banking, shopping, and everyday computing. Hit play for a wide-ranging look at the companies fighting to become the AI assistant that runs your digital life.


---

Enjoying Big Technology Podcast? Please rate us five stars ⭐⭐⭐⭐⭐ in your podcast app of choice.

Want a discount for Big Technology on Substack + Discord? Here’s 25% off for the first year: https://www.bigtechnology.com/subscribe?coupon=0843016b

Learn more about your ad choices. Visit megaphone.fm/adchoices





Download audio: https://pdst.fm/e/tracking.swap.fm/track/t7yC0rGPUqahTF4et8YD/pscrb.fm/rss/p/traffic.megaphone.fm/AMPP4268505606.mp3
Read the whole story
alvinashcraft
17 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Voice agents in Foundry Agent Service: architecture, trust boundaries, and real-audio latency

1 Share

Building enterprise voice agents used to mean stitching together speech services, real-time models, turn detection, and tool orchestration—then owning every failure mode. Voice agents in Foundry Agent Service consolidate a multi-service voice solution into a versioned, managed agent. They reduce runtime engineering, but identity, authorization, data, networking, retention, and operations remain enterprise responsibilities.

Microsoft's launch announcement explains the product capabilities. This article asks a different question: should the feature be the starting point for an enterprise voice architecture? It applies the six criteria from my earlier architecture comparison and adds a 300-turn real-audio study across six configurations.

In this article, you will learn:

  • what Foundry manages and what remains the solution team's responsibility;
  • how to assess the data, network, identity, cost, and operational boundaries;
  • what a 300-turn real-audio benchmark reveals about latency across six configurations.

As of 29 September 2026, voice-based agents remain in public preview. Product statements were checked on that date; measurements were collected on 26 September 2026.


Executive assessment

A voice agent now packages the model, instructions, audio configuration, turn detection, tools, interim responses, and storage policy behind a managed real-time endpoint. The solution team can focus less on assembling the runtime and more on governing the complete workload.

Foundry owns much more of the voice runtime, while the customer remains accountable for identity, authorization, data placement, network design, retention, and operational fitness.

CriterionManaged or provided by MicrosoftCustomer responsibility
Features and effortManaged real-time voice runtime, speech and model orchestration, turn detection, agent versioning, tool orchestration, channels, tracing, monitoring, and evaluation surfaces.Client and channel integration, business workflows, tool implementation, caller controls, governance, regression testing, capacity planning, and incident management.
Data residencyPublished service and model geography commitments, platform encryption, and documented storage and retention behavior.Select deployment types and regions; map speech, model, storage, tools, telemetry, and channel processing; configure retention and document compliance.
Private networkingAgent Service network-isolation capabilities, private endpoints, and integration with customer-managed Storage, Azure AI Search, and Azure Cosmos DB.Client connectivity, private DNS, routing, tool access, telemetry and telephony egress, firewall policy, and end-to-end reachability testing.
Authentication and authorizationMicrosoft Entra authentication for supported access patterns, managed and agent identities, role-based access control, and delegated-identity primitives.Caller sign-in and admission, per-user data authorization, tool permissions, delegated access, approval and step-up controls, and duplicate-action protection.
CostService meters, usage signals, estimated-cost telemetry, and Azure Cost Management billing records.Select model hosting and capacity, measure the complete workload, set budgets and forecasts, reconcile charges, and optimize usage.
LatencyManaged runtime plus service-side latency metrics, traces, and monitoring views.Model, voice, turn-detection, tools, retrieval, network, and channel choices; real-audio and concurrency tests; user-facing latency objectives and startup experience.

Connection startup: Establishing a new voice-agent connection took 3.72–12.37 seconds in this test. Connect before the conversation begins, or show the user that the agent is still connecting.


What is architecturally new?

The new API represents the experience as a VoiceAgentDefinition. Its kind is voice, and every create or update operation produces an immutable agent version behind a stable agent name. The definition selects:

  • model hosting, model, instructions, and an optional greeting;
  • input audio, transcription, enhancement, and turn detection;
  • output audio, voice, language, speed, and optional avatar;
  • tools, tool policy, interim responses, and conversation storage.

The runtime is reached over a persistent real-time connection. Foundry Agent Service owns the agent lifecycle and tool orchestration; Voice Live supplies the speech runtime.

A managed voice runtime reduces implementation work. It does not merge the identity, data, networking, and retention boundaries around it.

One feature, two model paths

Foundry derives the voice architecture from the selected model.

Native speech-to-speech

caller audio -> realtime speech model -> spoken response

This path is intended for natural, low-latency conversation. The model consumes and emits audio directly. Available voice and audio controls depend on the selected model.

Cascaded text model

caller audio -> speech recognition -> text model -> speech synthesis -> spoken response

This path provides broader text-model and Azure voice choices, plus explicit transcription, phrase-list, and speech-synthesis controls. It also places serial stages on the critical latency path.

Service-hosted (managed) versus customer-deployed (self_deployed) is a separate decision. It controls where the model is hosted and metered. The model itself determines whether the resulting path is native or cascaded.


Criterion 1 — Features and implementation effort

This is the feature's clearest architectural value. Foundry combines real-time sessions, turn taking, transcription, voices, agent versions, models, knowledge, tools, channels, traces, monitoring, evaluation, and optional avatars. The portal supports no-code testing, while client libraries and Azure Developer CLI make definitions deployable from source control.

What remains customer-owned

An enterprise still owns the client or approved phone channel, caller admission and abuse controls, user-specific retrieval and action authorization, approval and duplicate-action protection for consequential tools, human handoff, retention policy, regression testing, capacity, and incident management.

Channel support is narrower than generic agent publishing. The current voice-based Channels experience exposes a Preview web app and phone-number integrations through Teams Phone extensibility or Twilio. Standard Microsoft Teams or Microsoft 365 agent publishing is not a substitute for Teams Phone extensibility.


Criterion 2 — Data residency

Data residency cannot be answered with one region field. A voice turn crosses multiple processing and storage surfaces:

SurfaceResidency question
Speech and modelWhere are turn detection, recognition, synthesis, and model inference processed? Is the model global, data-zone, or regional?
Agent state and conversationsWhere is service state stored? Is store enabled for transcripts and raw audio?
Knowledge and toolsWhere are files, indexes, retrieved passages, tool requests, and outputs processed?
ObservabilityWhere does Application Insights store telemetry, and is sensitive-content capture enabled?
ChannelWhat additional processing is introduced by Azure Communication Services, Teams Phone, or Twilio?

The Voice Live data-privacy documentation says that Voice Live itself does not retain customer data by default, but connected features can. If a customer opts into support logging, Microsoft can retain the relevant speech data in the resource region for up to 30 days. Agent Service documentation places data stored by stateful service features at rest in the Azure OpenAI resource geography; model inference and tools follow their own configuration and hosting location.

For a self_deployed configuration, model inference follows the selected Global, Data Zone, or regional deployment type; that choice does not determine the location of speech, storage, tools, or channels. A managed configuration uses a service-hosted model. The project region alone does not answer the full residency question; verify the current processing and storage commitments for the selected speech, model, storage, telemetry, and channel components in Microsoft documentation.


Criterion 3 — Private networking

Foundry Agent Service supports a network-secured Standard setup with a delegated subnet, private endpoints, customer-provided Azure Storage, Azure AI Search, and Azure Cosmos DB, and public network access disabled. The project managed identity receives data-plane access to those dependencies. That is the starting point, not proof that a voice solution is private end to end. Review these paths independently:

ConnectionWhat must be demonstrated
Client to voice-agent endpointPrivate DNS, successful real-time connection setup, reachability, and idle-timeout behavior from the deployed client environment, or an authenticated public edge
Voice runtime to selected modelThe exact managed or self-deployed route used by the chosen voice configuration
Agent to data and toolsPrivate endpoints, DNS, role-based access control, route, credentials, and successful runtime access
Telemetry and telephonyApproved Application Insights egress plus provider-specific media and signaling paths

The Agent Service private-networking guide documents the general private setup. It does not make a private endpoint on one resource proof of every voice-specific managed hop.


Criterion 4 — Authentication and authorization

Voice-agent security is easier to reason about as four identity layers:

  1. Caller identity. Access to the hosted web experience can be assigned to organizational users and groups. A custom client still needs sign-in, session admission, rate limiting, and business entitlements.
  2. Session invocation identity. Public SDK and portal access patterns use Microsoft Entra ID for access to the Foundry project and agent endpoint. Phone-channel integrations add provider-specific authentication and connection setup: Teams Phone extensibility uses Azure Communication Services, Event Grid, and app-registration setup, while Twilio uses a project connection backed by Twilio credentials. Applications should use an appropriate workload identity and must not put long-lived credentials in browser code.
  3. Agent identity. The Foundry agent-identity model also applies to voice agents and authenticates downstream tools. Publishing an agent as a general Agent Application creates a distinct identity, so project permissions do not automatically transfer.
  4. Delegated user identity. For attended scenarios, OAuth on-behalf-of flows let a downstream service evaluate both the agent identity and the user's delegated permissions.

A valid token proves identity; it does not prove that a requested action is safe. An agent might be allowed to query an index while the caller can see only some documents, or it might be able to call a tool that the caller cannot authorize. Apply the caller's document permissions before content reaches the model. For consequential actions, require application policy, step-up verification where needed, spoken confirmation, and duplicate-action protection so a reconnect cannot repeat a side effect. Do not treat a recognized voice or a phone number alone as proof of identity.


Criterion 5 — Cost

There is now official pricing guidance for voice-based agents, but it is more useful as billing anatomy than as a durable price table.

Five documented factors drive cost:

  1. Audio input and output tokens.
  2. Managed versus self-deployed model hosting.
  3. Connected session duration.
  4. Tool calls and the services behind them.
  5. Optional features such as avatars and stored conversations.

The complete architecture can also include speech recognition and synthesis, custom voice hosting, Application Insights, Azure Storage, Azure AI Search, Azure Cosmos DB, telephony, application ingress, identity, and network infrastructure.

With model_type: managed, usage is billed as part of the voice agent. With model_type: self_deployed, model usage is billed against the customer's own Foundry model deployment under that deployment's standard or provisioned capacity pricing model.

Voice traces can expose token usage and estimated-cost attributes while a system is being tuned. Microsoft explicitly states that these are estimates, not invoice records; Azure Cost Management remains the billing source of truth.

Why this article has no dollar estimate

A single per-conversation estimate would create false precision because region, agreement, model, audio volume, session duration, tools, and optional services all change the result. Capture per-turn usage and session duration, reconcile actual charges in Cost Management, then model realistic call length, concurrency, transfers, and retries.

Include operational cost as well: the managed feature removes voice-runtime code, not governance, identity, networking, or channel operations.


Criterion 6 — Latency

The caller notices one latency measure: after I stop talking, how long until the agent starts speaking? The model path sets the processing stages, but model choice, turn-detection settings, network distance, tools, retrieval, response length, and channel buffering determine the result.

What was measured

The companion harness reran the three earlier patterns and the new feature through one real-audio test:

  • Realtime API direct and Voice Live with your own model used the same customer-deployed gpt-realtime-1.5 model.
  • The earlier Voice Live + prompt-agent path and the new cascaded voice agent both used gpt-4o-mini, the same Azure neural voice, and no tools. This is the closest architecture-controlled pair.
  • The new feature also ran with managed gpt-realtime-2.1 and cascaded gpt-5.

Every path received the same 1.33-second “Say hello briefly” recording at real-time speed, the same short-answer instruction, and the same requested server-side turn-detection settings, including a 500-millisecond silence wait. Each configuration ran 50 measured turns across five fresh sessions, with one excluded warm-up per session and randomized sequential order. Knowledge and tools were disabled, and the first-class agents did not store conversations.

The headline clock starts when the caller finishes speaking and stops when the first response audio reaches the client:

  • Typical turn is the median: half the turns were faster and half were slower.
  • 95% by is the slow-end result: only one turn in twenty was slower.

Time from when the caller finishes speaking until response audio reaches the client. Lower is better; results are specific to this test environment.

Architecture and configurationTypical turn95% started by
Earlier pattern — Realtime API direct, gpt-realtime-1.51.03 s1.10 s
Earlier pattern — Voice Live with your own model, gpt-realtime-1.51.24 s1.43 s
Earlier pattern — Voice Live + prompt agent, gpt-4o-mini1.90 s2.37 s
New feature — managed gpt-realtime-2.11.47 s1.61 s
New feature — cascaded gpt-4o-mini1.31 s1.52 s
New feature — cascaded gpt-53.45 s5.32 s

All 300 measured turns completed with response audio and the expected transcription.

The same-model-and-voice cascade comparison is the most informative result. The first-class gpt-4o-mini voice agent started 0.59 seconds sooner on the typical turn and 0.85 seconds sooner at the 95% boundary than Voice Live with the no-tool prompt agent. That is an observed end-to-end difference between these configurations; it does not reveal which internal stage produced it.

Realtime API direct was fastest in this environment. Voice Live with your own model was 0.21 seconds behind on the typical turn and 0.32 seconds behind at the 95% boundary. Both used the same realtime deployment, but the direct path used its model-native voice while Voice Live used an Azure neural voice. The difference therefore measures the configured stacks, not Voice Live overhead in isolation.

Within the new feature, managed gpt-realtime-2.1 and cascaded gpt-4o-mini were close enough that this run should not establish a permanent ordering between them. Changing the cascade to gpt-5 had a much larger effect, raising the typical result from 1.31 to 3.45 seconds. Model choice remains part of the architecture decision.

Why service monitoring can show sub-second results

The Overall latency chart in the voice-agent monitoring view in Foundry Agent Service measures time to first audio after voice activity detection decides the caller has finished. The benchmark chart starts earlier—when the caller actually stops speaking—and therefore includes the configured half-second silence wait.

Using that later service boundary:

  • Managed speech-to-speech was typically 0.88 seconds, with 95% of turns starting by 1.02 seconds.
  • The gpt-4o-mini cascade was typically 0.78 seconds, with 95% starting by 0.99 seconds.
  • The gpt-5 cascade was typically 2.92 seconds, with 95% starting by 4.78 seconds.

Both clocks are valid, but they answer different questions. The chart uses the caller's clock because it better represents the experience a user feels. The later service boundary explains how an operational metric can be sub-second even when caller-side time exceeds one second: the caller view adds the time required to detect the end of the turn and deliver audio across the network.

How this relates to the earlier measurements

The earlier article used ten warm, text-injected turns and excluded real speech and end-of-turn detection. Do not compare those values numerically with this chart. The unified rerun above is the cross-pattern comparison because every path uses the same audio, caller-side timing boundary, warm-up policy, sample count, region, network, and no-tool workload.

Test scope and reproduction

The test includes real speech, end-of-turn detection, model processing, and returned audio. It excludes device buffering, telephony, tools, retrieval, concurrency, private networking, and a production client, so its results describe this environment rather than a universal target.

For reproduction, see the unified benchmark harness, benchmark-only prompt-agent provisioner, first-class voice-agent provisioner, audio generator, raw turn data and method, and editable chart source.


Enterprise architecture checklist

For a workload decision, confirm:

  • Platform fit: the region supports the selected model, voice, audio features, tools, and channel, and service quotas cover the expected concurrency.
  • Data and network: every processing and storage location is recorded, and every route works from the intended client and network topology.
  • Identity and authorization: callers are authenticated, retrieval respects their document permissions, and tool identity, delegation, approval, and duplicate-action protection are tested.
  • Retention: conversation storage, trace capture, audio access, deletion, consent, and legal retention are approved.
  • Performance: real audio, tools, target network, concurrency, and objectives for both typical turns and the slowest 5% are tested.
  • Operations: quotas, errors, reconnects, handoff, dependency outages, monitoring, and rollback are rehearsed.

Verdict

For a new enterprise voice solution on Microsoft Foundry, voice agents in Foundry Agent Service are a strong primary candidate to evaluate when preview status is acceptable. They make voice a managed, versioned agent lifecycle and bring the runtime, channels, traces, monitoring, and evaluation together. That makes them a simpler starting point than rebuilding those layers by default.

This is not a claim that the feature is always faster or cheaper. Direct Realtime was fastest in this test, while the gpt-5 result showed how strongly model choice can dominate the outcome. Nor does the managed feature remove the surrounding trust boundaries: identity, authorization, data placement, networking, retention, and operational fitness still require an enterprise design.

The earlier patterns therefore remain valid, but for a new build they should be deliberate alternatives rather than the default starting point:

Begin by evaluating the voice-agent feature in Foundry Agent Service. Adopt it when its preview status, region, model, voice, channel, identity, data, network, cost, and measured performance fit the workload. Choose Realtime API or lower-level Voice Live patterns when a documented requirement needs control, capability, or a processing boundary that the managed definition does not provide.

For the lower-level alternatives and the trade-off between owning the runtime and consuming managed layers, see the earlier three-pattern architecture comparison. For Microsoft's feature overview and product direction, read the launch announcement.


Sources

Read the whole story
alvinashcraft
17 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Prepare and submit your apps for iPhone Duo

1 Share

iPhone Duo will be available to customers starting October 23, 2026. To get your apps and games ready today, you can:

Build or recompile with Xcode

Recompile existing apps with Xcode 27.1, which adds development support for iPhone Duo. Use Device Hub to visualize how your apps behave on iPhone Duo across all poses and orientations.

Prepare your App Store product page assets

Updated screenshot and app preview specifications are now available for the latest Apple devices, including iPhone Duo, so you can showcase your app at its best across orientations. Explore design guidance and get downloadable templates to help you build your App Store assets, including new product page headers and search result assets, as well as your app previews, screenshots, and In-App Events. Visualize how your screenshots, product page headers, and other metadata look on your product page on iPhone Duo with the new preview tool in App Store Connect.

Submit your app

You can submit your iPhone Duo optimized apps and games in App Store Connect today. Starting April 2027, any apps or games submitted will need to include screenshots for iPhone Duo.

Learn about submitting

Read the whole story
alvinashcraft
18 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Showcase your apps with new assets on the App Store

1 Share
Two iPhone screens showing the App Store. On the left is the product page for The Coast, with an illustrated header of a lighthouse and a cargo ship, the app icon, a Get button, and a 4.7-star rating. On the right are search results for “hiking,” where Ridge Explorer appears with an Editors’ Choice badge and a screenshot that reads “Plan your hiking expedition” next to a map of nearby trails.

Make your apps and games stand out with new creative assets on the App Store and asset management tools in App Store Connect.

Market your app or game in new ways

Use new creative assets on your product page headers and in search result assets to highlight your brand, promote seasonal offerings, showcase new content, and more. These new image and video assets complement your existing app previews and screenshots to give your app additional visibility across the App Store. Take advantage of design guidance and downloadable templates to help you build captivating visuals.

Visualize your assets across the App Store

Use the new preview tool in App Store Connect to see how your creative assets, screenshots, and other metadata will appear on your product page and in search results before you submit them. With support for Dark Mode, as well as different devices and orientations, the preview tool helps ensure your app makes a great impression across the App Store.

Manage your assets in App Store Connect

Streamline your asset management with the new Asset Library in App Store Connect. Upload and organize your creative assets, screenshots, and app previews in one central place — even ahead of when you’d like to use them. For example, you can submit creative assets for approval today to use in a future seasonal campaign, then immediately use them once you’re ready.

Read the whole story
alvinashcraft
18 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

One MCP server used 18,000 tokens before doing anything. Here’s the workaround.

1 Share
stacks abstract

Pi spent much of the past year leaving MCP out of its coding agent, even as the protocol became standard across developer tools. Its creator, Mario Zechner, had a concrete reason for resisting it. When he measured the context cost of popular browser-automation servers, Chrome DevTools MCP alone took up roughly 18,000 tokens, or about 9% of a 200,000-token context window, before the agent had done anything useful.

MCP is now built into Pi, but connecting a server doesn’t put all of its tools in front of the model. Pi leaves those definitions out of the prompt and uses Codemode to find tools as needed, call them from JavaScript and return only the useful output.

Measuring MCP’s context cost

Zechner spelled out the problem last November. Playwright MCP needed about 13,700 tokens to describe its 21 tools, or 6.8% of a 200,000-token window, and each additional server added more overhead. He was also bothered by what happened after a tool was called. “MCP servers also aren’t composable,” he wrote. “Results returned by an MCP server have to go through the agent’s context to be persisted to disk or combined with other results.”

“Results returned by an MCP server have to go through the agent’s context to be persisted to disk or combined with other results.”

His answer was Bash and a handful of scripts. The model already knew how to use them, so there was little reason to teach it another large tool interface. His CLI-based browser tools needed only a 225-token README, and their output could be piped into another command, filtered or saved to disk without first passing through the model. Extensions such as pi-mcp-adapter brought MCP to Pi before it became a native feature in Pi 1.0.

Pi changed MCP instead

Earendil, which acquired Pi earlier this year, says MCP has matured enough to justify a second look, though the company’s explanation for the reversal points to something more practical. “The reason we brought MCP into the core is not just about how MCP has changed, but also because we found that the changes it would require were generally useful,” Earendil wrote.

“The reason we brought MCP into the core is not just about how MCP has changed, but also because we found that the changes it would require were generally useful,”

Pi was already using the same idea with Jev, putting Codemode — its version of the code mode pattern — between the model and its tools. MCP could use that architecture instead of exposing every tool directly to the model.

Codemode handles tool discovery

Codemode runs inside a QuickJS sandbox without Node APIs, file system, network access or timers. Scripts can call Pi’s tools and models, run operations concurrently and process the results before returning anything to the model.

By default, Pi keeps an MCP server’s tools out of the model’s context. The system prompt gets a one-line description of each server, while the agent uses Codemode to find and call the tools it needs.

Developers can change that behavior with toolExposure. On the same server, Pi can expose some tools directly, leave others behind Codemode or block them entirely. A GitHub setup, for example, might expose search_code, keep get_* behind Codemode and block delete_*.

Per-tool exposure controls

Codemode has a 3,000-token default budget for tool declarations; anything beyond it stays discoverable. Pi 1.0 also reduced Codemode’s overall footprint.

Codemode has a 3,000-token default budget for tool declarations; anything beyond it stays discoverable.

According to the release notes, a GPT-5.6 request with the default tools and Codemode fell from roughly 5,300 prompt tokens to 3,300 after Pi shortened the Codemode description, moved model API documentation out of the prompt and stopped repeating declarations for tools scripts could already access. MCP tools using the default exposure don’t count toward that 3,000-token budget.

Pi 1.0 doesn’t resolve Earendil’s broader complaints about MCP, including composability. Codemode gives Pi a way to support the protocol without adopting the approach Zechner objected to in the first place.

The post One MCP server used 18,000 tokens before doing anything. Here’s the workaround. appeared first on The New Stack.

Read the whole story
alvinashcraft
18 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Connect Azure Functions to more services with managed connectors

1 Share

Azure Functions can already connect to many Azure services through triggers and bindings. With managed connectors, your functions can access about 1,700 connectors across services such as Microsoft 365, Microsoft Teams, Dataverse, SharePoint, OneDrive, and third-party systems.

Connector triggers deliver events from these services to your function, while typed connector clients let your code take actions against them. You get this broader integration surface without writing the webhook registration code or managing the OAuth tokens required to connect to each service, so you can focus on your function’s business logic and let Azure Connector Namespace handle the connection.

Azure Functions integration with Connector Namespace is currently in public preview. It supports .NET isolated, Python, and Node.js. Review the managed connectors overview for current language, hosting plan, and regional availability.

To demonstrate how connector triggers and actions work together, this article follows a .NET sample that automates RFP intake across SharePoint, Azure Content Understanding, and Teams.

From an uploaded RFP to Teams notification

Consider an organization that receives requests for proposals (RFPs) in a shared SharePoint document library. Someone must read each document, identify the requested capabilities, determine which subject-matter experts should respond, and notify the right team.

The automated RFP intake sample turns that process into an event-driven workflow:

  • A customer uploads an RFP to a SharePoint document library.
  • A SharePoint connector trigger invokes an Azure Function when the file is created.
  • The function uses a typed SharePoint connector client to retrieve the file contents.
  • Azure Content Understanding extracts the document’s text and layout.
  • The function applies deterministic rules to identify the customer, required capabilities, and recommended subject-matter experts.
  • The function uses a typed Teams connector client to post the results as an Adaptive Card in a channel.

Connector Namespace manages the SharePoint and Teams connections. The function controls file processing, document analysis, routing rules, error handling, and notification content.

How the sample works

The .NET sample demonstrates both parts of the connector programming model: a connector trigger receives an event from SharePoint, and typed connector clients provided by the Connector SDKs perform actions against SharePoint and Teams.

The function starts when the SharePoint When a file is created trigger detects a new RFP. It declares the trigger using the ConnectorTrigger attribute and receives a typed payload containing the file’s properties:

[Function("OnNewFile")]
public async Task OnNewFile(
    [ConnectorTrigger] SharePointOnlineOnNewFileItemsTriggerPayload payload,
    CancellationToken cancellationToken)
{
    // Process the newly uploaded file.
}

Because the trigger provides file properties rather than its contents, the function uses a typed SharePoint client to retrieve the document:

byte[] response = await _sharePoint.GetFileContentAsync(
    Uri.EscapeDataString(siteAddress),
    fileIdentifier,
    cancellationToken: cancellationToken);

byte[] document = SharePointFileContent.Decode(response);

The sample registers the SharePoint and Teams clients through dependency injection. Each client uses the runtime URL of its Connector Namespace connection and authenticates with DefaultAzureCredential:

services.AddSingleton(
    new SharePointOnlineClient(
        new Uri(sharePointRuntimeUrl),
        credential));

services.AddSingleton(
    new TeamsClient(
        new Uri(teamsRuntimeUrl),
        credential));

The function sends the document to Content Understanding’s prebuilt-layout analyzer, which extracts the document’s text and structure. It then applies deterministic C# rules to identify the customer and required capabilities and map those capabilities to predefined subject-matter expert roles.

Finally, the function creates an Adaptive Card containing the results and posts it to the configured Teams channel with the typed Teams client:

await _teams.PostCardToConversationAsync(
    postAs,
    postIn,
    request,
    cancellationToken);

Connector Namespace handles the SharePoint and Teams connections, while the function controls the document analysis, routing logic, error handling, and notification content.

Try the sample

The RFP intake sample includes the function code, Bicep infrastructure, Azure Developer CLI configuration, and supporting scripts. Its README explains how to test the workflow locally and deploy it to Azure.

Common connector patterns

Managed connectors are useful when a function must react to events or perform operations in external systems. Common patterns include:

  • Event to action: React to an event in one service and take an action in another.
  • Event to enrich to action: Retrieve additional information related to an event before acting.
  • Event to document analysis to action: Extract text and structure from a document, apply application rules, and send the result through another connector.
  • Event to AI to action: Analyze event data with an AI service and write the result back through a connector.
  • Extend an existing function app: Add connector-based integrations alongside HTTP, timer, queue, Service Bus, Event Grid, or Durable Functions workloads.

The RFP sample combines several of these patterns. A SharePoint event starts the workflow, a SharePoint action retrieves the document, Content Understanding extracts its contents, application code enriches the result, and a Teams action sends the notification.

Closing thoughts

Managed connectors extend the external systems that can trigger your functions and the services your function code can act on. This brings services such as SharePoint, Teams, Microsoft 365, Dataverse, and many third-party systems into the Azure Functions programming model without requiring you to build the underlying webhook and OAuth infrastructure.

Choose Azure Functions with managed connectors when you want this broader integration surface in a code-first application and need custom branching, application libraries and SDKs, other Functions bindings, document or AI processing, or application-specific logic between the trigger and action.

If the workload primarily orchestrates connector operations, involves little custom code, and would benefit from a visual designer, Azure Logic Apps is usually the simpler choice.

Resources

Documentation

Connector SDK GitHub repositories

The post Connect Azure Functions to more services with managed connectors appeared first on Azure SDK Blog.

Read the whole story
alvinashcraft
18 minutes ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories