Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
161857 stories
·
33 followers

AI Crash Course: WebMCP

1 Share

WebMCP is a proposed web standard that would provide AI agents with a sort of mental model of a web app’s capabilities via a browser API, saving time and tokens.

One of the biggest things we’ve learned about AI is that its strengths aren’t necessarily related to what it can do “out of the box,” but rather in the tools we can give it to extend its default capabilities.

The major foundation models that we’re often working with are OK-but-not-great at most things. If we want them to be great (and we do), then we have to have to start layering on other approaches to broaden what they can do. That could look like prompt engineering, adding more (or better) context, keeping track of memories, writing skills, including retrieval-augmented generation, using MCP Servers and more.

However, all of those are things we add to the AI, itself, to extend its capabilities. What if, instead, we flipped the script and extended the capabilities of things we know the AI is going to interact with? That’s the core idea behind WebMCP.

WebMCP is a proposed web standard and browser API that allows web applications to expose functionality as structured tools that AI agents can discover and use. Instead of an agent having to figure out how your app works by inspecting the interface, guessing and iterating, what if we could provide all the meta information that it needed to know, upfront?

Giving Agents ‘Mental Models’

We rely on our users’ mental models quite a bit when designing interfaces. When we put a dropdown in front a human that says “Add to Cart,” we’re relying on the fact that they understand the normal steps in a checkout process, that they know they can find the Cart page by clicking the shopping cart icon in the header, so on and so forth. Our users build up this kind of familiarity based on the other similar experiences they’ve had browsing the web and navigating other user interfaces. But that’s not the case for AI.

AI doesn’t automatically build that same kind of high-level familiarity with interface patterns over time in the way humans do. Unless that information is explicitly provided through context, memory, training or tools, then each new interface can require the agent to “start over” with its understanding of how the UI works.

If you’ve ever watched an agentic web browsing workflow, you’ve probably seen the general pattern: the agent inspects the page, identifies the relevant element, determines how to interact with it and then performs an action. Depending on the agent and browser, that inspection might involve the DOM, accessibility information, screenshots or other browser APIs. If that causes a change on the page, then the whole process starts again. Even seemingly simple tasks like making a selection from a large dropdown, could take several minutes and many, many loops for an agent to actually accomplish.

Not only is this slow, inefficient and prone to lots of mistakes and backtracking, but with the cost of tokens today, it can also be pretty darn expensive (depending on the complexity of the page and the action it’s trying to take). WebMCP seeks to change this by allowing web applications to expose their capabilities as structured tools.

With WebMCP, the agent can receive a description of those capabilities from the browser—the purpose of the capability, what inputs it accepts, what values are allowed, etc.—rather than having to infer it from the rendered UI. Keeping in mind that WebMCP is a still a proposed API and the syntax isn’t finalized (in fact, it recently changed from navigator.modelContext to document.modelContext), our dropdown could offer the AI something like this:

await document.modelContext.registerTool({
  name: "filter_products",
  description: "Filter the product catalog by category.",
  inputSchema: {
    type: "object",
    properties: {
      category: {
        type: "string",
        enum: ["all", "shirts", "shoes", "accessories"],
        description: "The product category to display."
      }
    },
    required: ["category"]
  },
  execute: async ({ category }) => {
    filterProducts(category);

    return {
      success: true,
      category
    };
  }
});

Here, we can see important information like the input type, allowed values, required fields and descriptions all shared with the model upfront. In a way, it’s quite similar to semantic HTML: instead of forcing the browser to infer what an element means from how it looks or behaves, we explicitly provide information about its purpose and semantics. WebMCP applies a similar idea to application capabilities, giving an agent a structured description of what it can do and what inputs it expects.

This can also be expanded to include functionality. Modern web applications already contain functions representing the actions users can take. Our shopping example probably includes functions like addToCart() or updateProfile() that might currently be called in response to UI events and could also be exposed as agent-accessible capabilities.

What’s the Difference Between an MCP Server and WebMCP?

The primary difference is the involvement of the browser. Traditional MCP lets an AI application connect to an MCP Server that exposes tools and other capabilities. With WebMCP, those tools and capabilities are exposed via the browser. MCP and WebMCP aren’t competitors or replacements; they serve different purposes for different contexts. The thing they have in common is simply the fact that they’re surfacing relevant tools for AI, even though they do so in very different ways.

When we’re talking about an MCP Server (like the Progress Telerik and Kendo MCP Server), the flow looks (roughly) like:

AI App → MCP Client → MCP Server → External Service

For WebMCP, it looks more like:

AI Agent → Browser → Web App → App logic / API

For example, we use the Telerik MCP Server to give your AI coding assistant access to Kendo UI and Telerik-specific information and tools (like the Upgrade Assistant). On the other hand, our WebMCP support enables you to surface that additional information and functionality (as described above in the dropdown example) directly in our Telerik and Kendo UI components to AI agents interacting with applications that include them.

The high-level goal of both is the same—equipping AI with the extra tools and context it needs—but the approach is different.

Do I Need to Add WebMCP into My App Today?

Short answer: no, it’s not a thing yet.

Longer answer: not yet, but hopefully soon! And no matter how the syntax changes between now and whenever it officially launches (and I’m sure it will), the core idea is one that I think is worth keeping an eye on.

We’re now into the realm of personal opinion / future speculation / so-called “thought leadership,” so take this paragraph with a grain of salt. To me, WebMCP is one of the more exciting things to come along with AI because of all the ways it holds potential to improve the user experience beyond just AI. In accessibility, this is referred to as the “curb cut effect”—when a solution benefits not only the intended audience, but many more beyond. The more standardized it becomes for us to surface this kind of contextual information in our web applications, the more it could also be utilized by not only agents, but also potentially browsers, screen readers (and other assistive technologies), and alternate input devices. Are there agentic browsing possibilities that could help open up the web to more users?

Right now, some similar information (the name, description, role and state) already gets provided by elements to the DOM—and by extension, to the accessibility tree—via the use of semantic HTML elements and ARIA. Could WebMCP eventually offer a way for us to create an AI-friendly standardized syntax that provides not only that information, but also greater insight into associated functionality, typing and more? In the process of making the web friendlier for agents, I believe there’s a real opportunity for us to make it friendlier for humans as well.

Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

Delegates and events in C# explained

1 Share

Learn how delegates and events work in C#, including Action, Func, Predicate and custom EventArgs, with practical examples.

Read the whole story
alvinashcraft
11 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Podcast: The AI Revolution Fails Without Psychological Safety For Developers: A Conversation with Erin Doyle

1 Share

In this podcast, Michael Stiefel spoke to Erin Doyle about how artificial intelligence tools are transforming the software development process, and how they are affecting software engineers. Architecture skills are now critical because engineers need to understand how to build systems and handle the constraints and associated ambiguity in the requirements.

By Erin Doyle
Read the whole story
alvinashcraft
21 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

ACP Muse Agent Client Protocol and Muse Code

1 Share
<p>At <a href="https://slopcop.com/">work</a> we tend to try every new model and vendor we can so we know what the state of the art is. Enter <a href="https://dev.meta.ai/docs/muse-code">Muse Code</a>, the models are fast, cheap and pretty capable. For $50 a month you get more credits than I can spend, and for API usage there is a 92-95% discount letting them train on your data. TLDR this is a ton of value you get out of the models from <a href="https://about.meta.com/">Meta</a>. So I went ahead and created <a href="https://github.com/BrokkAi/muse-acp">Muse ACP</a> and my employer was nice enough to let me work on it at work time and now it is a part of several of our projects.</p> <h2 id="what-is-acp">What is ACP?</h2> <p><a href="https://www.agentclientprotocol.com">Agent Client Protocol</a> is an interaction standard to that one can wire up ACP clients (TUI, Text Editor, Batch Processes, Personal Assistants) to an AI agent that has implemented the same protocol. Many Agents provide this out of the box such as <a href="https://github.com/NousResearch/hermes-agent">Hermes</a> via <code>hermes acp</code>, <a href="https://opencode.ai">Opencode</a> via <code>opencode acp</code>, you get the pattern, however as of yet Muse Code does not ship one, so I made one.</p> <p>Because of this one can easily hookup Hermes, Opencode or any of the other dozens of clients in the <a href="https://agentclientprotocol.com/get-started/registry">registry</a> to <a href="https://www.jetbrains.com/idea/">IntelliJ</a>, <a href="https://zed.dev/">Zed</a>, or any number of the clients I have written like <a href="https://github.com/BrokkAi/micro-acp">Micro ACP</a>, <a href="https://github.com/foundev/belgr">Belgr</a>, or <a href="https://github.com/BrokkAi/mjolnir">Mjolnir</a>.</p> <p>Now there is a registry, but I have personally found it a real and total pain to get into, so I have chosen to ignore it for Muse ACP until I get enough popularity that someone else adds my ACP server to it (10 stars and counting). It is also easy to add custom clients and I have add an installer for IntelliJ and Zed with Muse ACP (muse-acp install) so that you can try this out with your favorite editor.</p> <h2 id="why-muse-acp-and-not-one-of-the-others">Why Muse ACP and not one of the others</h2> <p>When I started the project there were no mature ACP servers using <a href="https://meta-models.github.io/muse-code-sdk/next/guides/msp-concepts/">Muse Session Protocol</a> as the integration point. Also, it is written in Rust instead of Typescript and that is for some people a good reason to prefer the BrokkAi version of the rest.</p> <h2 id="conclusion">Conclusion</h2> <p>If you want access to a cheap capable model, go ahead and sign up for a Muse Code subscription plan and give Muse ACP a try, wire it up to your favorite editor or one of mine and get to work.</p>
Read the whole story
alvinashcraft
31 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

A .NET Agent Orchestration Framework - Introduction

1 Share

Introduction

In this new series of posts I will talk about a multi-agent orchestration framework that I've been developing in the last few months. This framework (name yet to be defined) is provider and model-agnostic and the responsibility to connect to the models of choice is outside its scope; anything works as long as it has a Microsoft.Extensions.AI wrapper. As with most of the code I write, it is being developed in .NET and this post explains some of the core concepts and technologies used to build it; I will expand on each concept (the orchestrator, the agents, the training, the knowledge management, etc) in subsequent posts. I also plan to eventually make the code available in GitHub and NuGet.

This is somewhat linked to my PhD research, but I won't get too academic here (I am not!), this is essentially a proof of concept that seems to work, and something I'm having fun with!

The posts in the series I plan to write are, as of now:

  • Introduction (this one)
  • Orchestrator (TBD)
  • Agents (TBD)
  • Agent training (TBD)
  • Agent scheduling (TBD)
  • Agent review and revise (TBD)
  • MCP tools (TBD)
  • Putting it all together (TBD)
  • Extensibility and next steps (TBD)

Do keep in mind that, as the framework is evolving, there may be changes. I will try to keep all posts updated.

Why a Multi-Agent Framework?

Let's start from here: why build a multi-agent framework - isn't a single agent good enough? Well, there are some good reasons, IMO:

  • Specialisation beats generalism: a single agent juggling database schema, UI conventions, .NET idioms, and business rules tends to produce shallower output in each domain than agents each prompted/tuned for one thing
  • Context window hygiene: each specialised agent only needs the context relevant to its domain, not everything. Keeps prompts smaller, more focused, and less prone to the model losing track of instructions buried in a huge system prompt
  • Parallelism: independent subtasks (e.g., drafting the DB schema while another agent drafts the UI component) can run concurrently instead of sequentially through one agent
  • Built-in review/critique: peer review between agents catches errors a single self-reviewing agent is more likely to miss, since it's a genuinely different "pass" rather than the same model re-reading its own output
  • Easier debugging/iteration: when something's wrong, you can inspect which agent's output was faulty and improve just that prompt/tool config, rather than re-tuning one monolithic prompt
  • Standardized agent/tool abstractions: IChatClient, function-calling middleware, etc. give us a consistent way to define what an agent is without reinventing message-passing, tool-call parsing, and streaming handling for each specialist
  • Orchestration primitives you'd otherwise rebuild: task distribution, result aggregation/synthesis, retry/timeout handling, and conversation-state threading between agents are all things a framework typically handles - and can get subtly wrong if hand-rolled (e.g., losing partial results on a failed agent call)
  • Observability/tracing hooks: multi-agent systems are hard to debug blind - which agent said what, in what order, with what tool calls. Frameworks generally wire up tracing (OpenTelemetry-style) so you can see the whole exchange, which matters a lot once you add peer review (now you have agent→agent traffic, not just agent→user)
  • Composability: if today it's DB/UI/.NET/business-logic and tomorrow you add a testing or security-review agent, a framework gives you a slot to plug it into rather than restructuring bespoke glue code

Let's see how this is addressed in what I built.

Basic Concepts

In this framework we have essentially agents and an orchestrator. There are, of course, lots of other smaller components, but these are the most important ones.

The orchestrator receives a prompt, dynamically decides how to distribute the work to be done by the different specialised agents it knows, and then collects and summarises the responses. It also takes care of logging, collecting metrics, etc. My implementation broadly follows the Orchestrator-Worker (Hierarchical/Magentic) pattern and the evaluator/critique pattern if you are curious. The pipeline is: plan → execute → (review → revise)* → synthesize.

Agents that can use whatever model they want (different agents can use different models, possibly picking the one that is the best fit for their specialisation) but are trained, specialised, for a specific technology or field of studies. These fields of studies are, for now, a closed set, and include:

  • .NET development
  • Database development
  • DevOps
  • Documentation writing
  • Testing
  • Planning
  • Synthesis
  • ...

Besides the specialisation, an agent has a distinct name, and some tags that more finely describe it (e.g., "SQLServer", "PostgreSQL, etc). It can make use of some services, of which I'll talk in a moment. Agents wrap an instance of IChatClient (possibly a different one for each agent), and are instantiated with different parameters. The actual training is performed by a different class. The agent's prompt is built from its original prompt and their training when the orchestrator executes it. After processing, the work of each agent is reviewed by another agent, and its feedback is used for revising it. Agents can load state from past invocations and, at the end of a successful processing, can save the new state too. They can also share state amongst all the agents and have access to a shared knowledge base, which can be used to enhance its prompt.

The following picture summarises this:

Core Assumptions

All of the services implement an interface that describes its contract, and there is always a minimum, default, implementation. It should be simple to provide new implementations. All relevant aspects can be configured, core services replaced or extended. Always Dependency Injection (DI)-friendly.

All of the steps - planning, execution, review, revise, synthesis - are executed by agents.

All of the calls are asynchronous and can be cancelled, thus terminating the entire orchestration process.

All of the agent work is implemented by LLM models, as provided by an IChatClient implementation, which must be built outside the scope of this framework. The usage of tokens is controlled and logged.

All work is logged using the standard .NET constructs, including the time it took, tokens consumed, and activity traces are also produced.

All work occurs in a session and is identified by a session id. All agent operations have a task id and exist in a session. 

Class Diagram

These are the essential interfaces and their implementing classes, plus the relationships between them. Some implementing classes are not shown.

Here's a brief summary of these types, there are many more, but I don't find them relevant enough to be listed here:

  • IAgent/SpecializedAgent: the core agent interface and its only provided implementation
  • IOrchestrator/Orchestrator: the agent orchestrator that coordinates the work of the different agents; it makes use of many different services
  • IAgentKnowledgeStore: a store for keeping the history of each agent's operations after they are done as well as loading it initially to build the agent prompt
  • IAgentScheduler/ParallelAgentScheduler/SequentialAgentScheduler: used by the Orchestrator to schedule the agents; there are two implementations: one for parallel and the other for sequential execution
  • IPlanner: used by the Orchestrator to plan the agent assignments
  • ISynthesizer: also used by the Orchestrator to synthesise the final agent's responses
  • IAgentSelector: the algorithms to use for selecting an agent for execution or for reviewing
  • IAgentTrainingProfile: the contract for the agent training; used by the SpecializedAgent to train the agent before actually doing any work
  • IKnowledgeStore: a general-purpose knowledge base that is indexed by a specialisation and stores string content and their corresponding vector embeddings
  • ISharedState: a key-value store bound to a session, which should be thread-safe
  • ISharedStateFactory: builds the shared state. Usually used by the ISharedStateManager implementation to build the shared state for a session when it does not already exist
  • ISharedStateManager: returns the shared state for a given session
  • IMcpToolsProvider: registers additional MCP tools for a SpecializedAgent or Orchestrator

Technologies Used

Besides .NET 10, the framework uses the Microsoft.Extensions.AI framework to provide a low-level, provider-agnostic abstraction. I probably could have used the Microsoft Agent Framework, I have played with it before, and at some point I will certainly have a go in porting my framework to it; they are both great choices, part of Microsoft's AI ecosystem. For now, it was a somewhat conscious choice which I can revisit later. I also make use of MCP for file operations and a sharing of state between agents. Finally, I also use the ONNX Runtime for generating vector embeddings but it doesn't have to be used if we provide an alternative implementation.

Conclusion

This is what I wanted to talk about now. Stay tuned for the next posts in this series. As always, feel free to reach out and ask any questions or make any comments! This is a very important topic to me, and sure would appreciate comments on this, as we go along!

Read the whole story
alvinashcraft
39 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Auto-accepting AI edits in VS Code

1 Share

When you let Copilot edit your code in agent mode, every change waits for a decision: Keep or Undo. For a small change that's fine. After ten prompts in a row, you're clicking Keep more than you're reading code.

VS Code has a setting to take that click away.

Why do we need this?

By default, edits from chat are pending until you accept or discard them. The pending state survives closing VS Code, so you can come back later and still decide.

That's a good default. But if you already review everything in a diff before you commit, the extra step per edit adds little.

Accept edits after a delay

Add this to your settings.json:

"chat.editing.autoAcceptDelay": 5

Edits are now accepted automatically after the delay. The default is 0, which means auto-accept is disabled.


You are not locked out. While the countdown runs, hover over the editor overlay controls to cancel it, and you can still Undo afterwards.


Keep approval for sensitive files

Auto-accept applies to edits after they are made. A separate setting decides which files Copilot may edit without asking you first: chat.tools.edits.autoApprove. It takes glob patterns.

This example allows everything except JSON files in .vscode and your .env files:

"chat.tools.edits.autoApprove": {
  "**/*": true,
  "**/.vscode/*.json": false,
  "**/.env": false
}

For those files you still get a diff and have to approve the edit.

Review before you commit

The VS Code docs are clear on this: if you automatically accept all edits, review the changes in source control before you commit them.

Source control also helps in the other direction. Staging your changes accepts any pending edits, and discarding them discards the pending edits too.

Tip: Newer Agent Host sessions skip the pending state entirely. The agent saves edits straight into the session's folder or an isolated Git worktree, so the diff before you commit is your only review step.

That's it! Fewer clicks, but the review step is still yours.

More information

Read the whole story
alvinashcraft
46 seconds ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories