WebMCP is a proposed web standard that would provide AI agents with a sort of mental model of a web app’s capabilities via a browser API, saving time and tokens.
One of the biggest things we’ve learned about AI is that its strengths aren’t necessarily related to what it can do “out of the box,” but rather in the tools we can give it to extend its default capabilities.
The major foundation models that we’re often working with are OK-but-not-great at most things. If we want them to be great (and we do), then we have to have to start layering on other approaches to broaden what they can do. That could look like prompt engineering, adding more (or better) context, keeping track of memories, writing skills, including retrieval-augmented generation, using MCP Servers and more.
However, all of those are things we add to the AI, itself, to extend its capabilities. What if, instead, we flipped the script and extended the capabilities of things we know the AI is going to interact with? That’s the core idea behind WebMCP.
WebMCP is a proposed web standard and browser API that allows web applications to expose functionality as structured tools that AI agents can discover and use. Instead of an agent having to figure out how your app works by inspecting the interface, guessing and iterating, what if we could provide all the meta information that it needed to know, upfront?
We rely on our users’ mental models quite a bit when designing interfaces. When we put a dropdown in front a human that says “Add to Cart,” we’re relying on the fact that they understand the normal steps in a checkout process, that they know they can find the Cart page by clicking the shopping cart icon in the header, so on and so forth. Our users build up this kind of familiarity based on the other similar experiences they’ve had browsing the web and navigating other user interfaces. But that’s not the case for AI.
AI doesn’t automatically build that same kind of high-level familiarity with interface patterns over time in the way humans do. Unless that information is explicitly provided through context, memory, training or tools, then each new interface can require the agent to “start over” with its understanding of how the UI works.
If you’ve ever watched an agentic web browsing workflow, you’ve probably seen the general pattern: the agent inspects the page, identifies the relevant element, determines how to interact with it and then performs an action. Depending on the agent and browser, that inspection might involve the DOM, accessibility information, screenshots or other browser APIs. If that causes a change on the page, then the whole process starts again. Even seemingly simple tasks like making a selection from a large dropdown, could take several minutes and many, many loops for an agent to actually accomplish.
Not only is this slow, inefficient and prone to lots of mistakes and backtracking, but with the cost of tokens today, it can also be pretty darn expensive (depending on the complexity of the page and the action it’s trying to take). WebMCP seeks to change this by allowing web applications to expose their capabilities as structured tools.
With WebMCP, the agent can receive a description of those capabilities from the browser—the purpose of the capability, what inputs it accepts, what values are allowed, etc.—rather than having to infer it from the rendered UI. Keeping in mind that WebMCP is a still a proposed API and the syntax isn’t finalized (in fact, it recently changed from navigator.modelContext to document.modelContext), our dropdown could offer the AI something like this:
await document.modelContext.registerTool({
name: "filter_products",
description: "Filter the product catalog by category.",
inputSchema: {
type: "object",
properties: {
category: {
type: "string",
enum: ["all", "shirts", "shoes", "accessories"],
description: "The product category to display."
}
},
required: ["category"]
},
execute: async ({ category }) => {
filterProducts(category);
return {
success: true,
category
};
}
});
Here, we can see important information like the input type, allowed values, required fields and descriptions all shared with the model upfront. In a way, it’s quite similar to semantic HTML: instead of forcing the browser to infer what an element means from how it looks or behaves, we explicitly provide information about its purpose and semantics. WebMCP applies a similar idea to application capabilities, giving an agent a structured description of what it can do and what inputs it expects.
This can also be expanded to include functionality. Modern web applications already contain functions representing the actions users can take. Our shopping example probably includes functions like addToCart() or updateProfile() that might currently be called in response to UI events and could also be exposed as agent-accessible capabilities.
The primary difference is the involvement of the browser. Traditional MCP lets an AI application connect to an MCP Server that exposes tools and other capabilities. With WebMCP, those tools and capabilities are exposed via the browser. MCP and WebMCP aren’t competitors or replacements; they serve different purposes for different contexts. The thing they have in common is simply the fact that they’re surfacing relevant tools for AI, even though they do so in very different ways.
When we’re talking about an MCP Server (like the Progress Telerik and Kendo MCP Server), the flow looks (roughly) like:
AI App → MCP Client → MCP Server → External Service
For WebMCP, it looks more like:
AI Agent → Browser → Web App → App logic / API
For example, we use the Telerik MCP Server to give your AI coding assistant access to Kendo UI and Telerik-specific information and tools (like the Upgrade Assistant). On the other hand, our WebMCP support enables you to surface that additional information and functionality (as described above in the dropdown example) directly in our Telerik and Kendo UI components to AI agents interacting with applications that include them.
The high-level goal of both is the same—equipping AI with the extra tools and context it needs—but the approach is different.
Short answer: no, it’s not a thing yet.
Longer answer: not yet, but hopefully soon! And no matter how the syntax changes between now and whenever it officially launches (and I’m sure it will), the core idea is one that I think is worth keeping an eye on.
We’re now into the realm of personal opinion / future speculation / so-called “thought leadership,” so take this paragraph with a grain of salt. To me, WebMCP is one of the more exciting things to come along with AI because of all the ways it holds potential to improve the user experience beyond just AI. In accessibility, this is referred to as the “curb cut effect”—when a solution benefits not only the intended audience, but many more beyond. The more standardized it becomes for us to surface this kind of contextual information in our web applications, the more it could also be utilized by not only agents, but also potentially browsers, screen readers (and other assistive technologies), and alternate input devices. Are there agentic browsing possibilities that could help open up the web to more users?
Right now, some similar information (the name, description, role and state) already gets provided by elements to the DOM—and by extension, to the accessibility tree—via the use of semantic HTML elements and ARIA. Could WebMCP eventually offer a way for us to create an AI-friendly standardized syntax that provides not only that information, but also greater insight into associated functionality, typing and more? In the process of making the web friendlier for agents, I believe there’s a real opportunity for us to make it friendlier for humans as well.
Learn how delegates and events work in C#, including Action, Func, Predicate and custom EventArgs, with practical examples.

In this podcast, Michael Stiefel spoke to Erin Doyle about how artificial intelligence tools are transforming the software development process, and how they are affecting software engineers. Architecture skills are now critical because engineers need to understand how to build systems and handle the constraints and associated ambiguity in the requirements.
By Erin DoyleIn this new series of posts I will talk about a multi-agent orchestration framework that I've been developing in the last few months. This framework (name yet to be defined) is provider and model-agnostic and the responsibility to connect to the models of choice is outside its scope; anything works as long as it has a Microsoft.Extensions.AI wrapper. As with most of the code I write, it is being developed in .NET and this post explains some of the core concepts and technologies used to build it; I will expand on each concept (the orchestrator, the agents, the training, the knowledge management, etc) in subsequent posts. I also plan to eventually make the code available in GitHub and NuGet.
This is somewhat linked to my PhD research, but I won't get too academic here (I am not!), this is essentially a proof of concept that seems to work, and something I'm having fun with!
The posts in the series I plan to write are, as of now:
Do keep in mind that, as the framework is evolving, there may be changes. I will try to keep all posts updated.
Let's start from here: why build a multi-agent framework - isn't a single agent good enough? Well, there are some good reasons, IMO:
Let's see how this is addressed in what I built.
In this framework we have essentially agents and an orchestrator. There are, of course, lots of other smaller components, but these are the most important ones.
The orchestrator receives a prompt, dynamically decides how to distribute the work to be done by the different specialised agents it knows, and then collects and summarises the responses. It also takes care of logging, collecting metrics, etc. My implementation broadly follows the Orchestrator-Worker (Hierarchical/Magentic) pattern and the evaluator/critique pattern if you are curious. The pipeline is: plan → execute → (review → revise)* → synthesize.
Agents that can use whatever model they want (different agents can use different models, possibly picking the one that is the best fit for their specialisation) but are trained, specialised, for a specific technology or field of studies. These fields of studies are, for now, a closed set, and include:
Besides the specialisation, an agent has a distinct name, and some tags that more finely describe it (e.g., "SQLServer", "PostgreSQL, etc). It can make use of some services, of which I'll talk in a moment. Agents wrap an instance of IChatClient (possibly a different one for each agent), and are instantiated with different parameters. The actual training is performed by a different class. The agent's prompt is built from its original prompt and their training when the orchestrator executes it. After processing, the work of each agent is reviewed by another agent, and its feedback is used for revising it. Agents can load state from past invocations and, at the end of a successful processing, can save the new state too. They can also share state amongst all the agents and have access to a shared knowledge base, which can be used to enhance its prompt.
The following picture summarises this:
All of the services implement an interface that describes its contract, and there is always a minimum, default, implementation. It should be simple to provide new implementations. All relevant aspects can be configured, core services replaced or extended. Always Dependency Injection (DI)-friendly.
All of the steps - planning, execution, review, revise, synthesis - are executed by agents.
All of the calls are asynchronous and can be cancelled, thus terminating the entire orchestration process.
All of the agent work is implemented by LLM models, as provided by an IChatClient implementation, which must be built outside the scope of this framework. The usage of tokens is controlled and logged.
All work is logged using the standard .NET constructs, including the time it took, tokens consumed, and activity traces are also produced.
All work occurs in a session and is identified by a session id. All agent operations have a task id and exist in a session.
These are the essential interfaces and their implementing classes, plus the relationships between them. Some implementing classes are not shown.
Here's a brief summary of these types, there are many more, but I don't find them relevant enough to be listed here:
Besides .NET 10, the framework uses the Microsoft.Extensions.AI framework to provide a low-level, provider-agnostic abstraction. I probably could have used the Microsoft Agent Framework, I have played with it before, and at some point I will certainly have a go in porting my framework to it; they are both great choices, part of Microsoft's AI ecosystem. For now, it was a somewhat conscious choice which I can revisit later. I also make use of MCP for file operations and a sharing of state between agents. Finally, I also use the ONNX Runtime for generating vector embeddings but it doesn't have to be used if we provide an alternative implementation.
This is what I wanted to talk about now. Stay tuned for the next posts in this series. As always, feel free to reach out and ask any questions or make any comments! This is a very important topic to me, and sure would appreciate comments on this, as we go along!