WebMCP is a proposed web standard that would provide AI agents with a sort of mental model of a web app’s capabilities via a browser API, saving time and tokens.
One of the biggest things we’ve learned about AI is that its strengths aren’t necessarily related to what it can do “out of the box,” but rather in the tools we can give it to extend its default capabilities.
The major foundation models that we’re often working with are OK-but-not-great at most things. If we want them to be great (and we do), then we have to have to start layering on other approaches to broaden what they can do. That could look like prompt engineering, adding more (or better) context, keeping track of memories, writing skills, including retrieval-augmented generation, using MCP Servers and more.
However, all of those are things we add to the AI, itself, to extend its capabilities. What if, instead, we flipped the script and extended the capabilities of things we know the AI is going to interact with? That’s the core idea behind WebMCP.
WebMCP is a proposed web standard and browser API that allows web applications to expose functionality as structured tools that AI agents can discover and use. Instead of an agent having to figure out how your app works by inspecting the interface, guessing and iterating, what if we could provide all the meta information that it needed to know, upfront?
Giving Agents ‘Mental Models’
We rely on our users’ mental models quite a bit when designing interfaces. When we put a dropdown in front a human that says “Add to Cart,” we’re relying on the fact that they understand the normal steps in a checkout process, that they know they can find the Cart page by clicking the shopping cart icon in the header, so on and so forth. Our users build up this kind of familiarity based on the other similar experiences they’ve had browsing the web and navigating other user interfaces. But that’s not the case for AI.
AI doesn’t automatically build that same kind of high-level familiarity with interface patterns over time in the way humans do. Unless that information is explicitly provided through context, memory, training or tools, then each new interface can require the agent to “start over” with its understanding of how the UI works.
If you’ve ever watched an agentic web browsing workflow, you’ve probably seen the general pattern: the agent inspects the page, identifies the relevant element, determines how to interact with it and then performs an action. Depending on the agent and browser, that inspection might involve the DOM, accessibility information, screenshots or other browser APIs. If that causes a change on the page, then the whole process starts again. Even seemingly simple tasks like making a selection from a large dropdown, could take several minutes and many, many loops for an agent to actually accomplish.
Not only is this slow, inefficient and prone to lots of mistakes and backtracking, but with the cost of tokens today, it can also be pretty darn expensive (depending on the complexity of the page and the action it’s trying to take). WebMCP seeks to change this by allowing web applications to expose their capabilities as structured tools.
With WebMCP, the agent can receive a description of those capabilities from the browser—the purpose of the capability, what inputs it accepts, what values are allowed, etc.—rather than having to infer it from the rendered UI. Keeping in mind that WebMCP is a still a proposed API and the syntax isn’t finalized (in fact, it recently changed from navigator.modelContext to document.modelContext), our dropdown could offer the AI something like this:
await document.modelContext.registerTool({
name: "filter_products",
description: "Filter the product catalog by category.",
inputSchema: {
type: "object",
properties: {
category: {
type: "string",
enum: ["all", "shirts", "shoes", "accessories"],
description: "The product category to display."
}
},
required: ["category"]
},
execute: async ({ category }) => {
filterProducts(category);
return {
success: true,
category
};
}
});
Here, we can see important information like the input type, allowed values, required fields and descriptions all shared with the model upfront. In a way, it’s quite similar to semantic HTML: instead of forcing the browser to infer what an element means from how it looks or behaves, we explicitly provide information about its purpose and semantics. WebMCP applies a similar idea to application capabilities, giving an agent a structured description of what it can do and what inputs it expects.
This can also be expanded to include functionality. Modern web applications already contain functions representing the actions users can take. Our shopping example probably includes functions like addToCart() or updateProfile() that might currently be called in response to UI events and could also be exposed as agent-accessible capabilities.
What’s the Difference Between an MCP Server and WebMCP?
The primary difference is the involvement of the browser. Traditional MCP lets an AI application connect to an MCP Server that exposes tools and other capabilities. With WebMCP, those tools and capabilities are exposed via the browser. MCP and WebMCP aren’t competitors or replacements; they serve different purposes for different contexts. The thing they have in common is simply the fact that they’re surfacing relevant tools for AI, even though they do so in very different ways.
When we’re talking about an MCP Server (like the Progress Telerik and Kendo MCP Server), the flow looks (roughly) like:
AI App → MCP Client → MCP Server → External Service
For WebMCP, it looks more like:
AI Agent → Browser → Web App → App logic / API
For example, we use the Telerik MCP Server to give your AI coding assistant access to Kendo UI and Telerik-specific information and tools (like the Upgrade Assistant). On the other hand, our WebMCP support enables you to surface that additional information and functionality (as described above in the dropdown example) directly in our Telerik and Kendo UI components to AI agents interacting with applications that include them.
The high-level goal of both is the same—equipping AI with the extra tools and context it needs—but the approach is different.
Do I Need to Add WebMCP into My App Today?
Short answer: no, it’s not a thing yet.
Longer answer: not yet, but hopefully soon! And no matter how the syntax changes between now and whenever it officially launches (and I’m sure it will), the core idea is one that I think is worth keeping an eye on.
We’re now into the realm of personal opinion / future speculation / so-called “thought leadership,” so take this paragraph with a grain of salt. To me, WebMCP is one of the more exciting things to come along with AI because of all the ways it holds potential to improve the user experience beyond just AI. In accessibility, this is referred to as the “curb cut effect”—when a solution benefits not only the intended audience, but many more beyond. The more standardized it becomes for us to surface this kind of contextual information in our web applications, the more it could also be utilized by not only agents, but also potentially browsers, screen readers (and other assistive technologies), and alternate input devices. Are there agentic browsing possibilities that could help open up the web to more users?
Right now, some similar information (the name, description, role and state) already gets provided by elements to the DOM—and by extension, to the accessibility tree—via the use of semantic HTML elements and ARIA. Could WebMCP eventually offer a way for us to create an AI-friendly standardized syntax that provides not only that information, but also greater insight into associated functionality, typing and more? In the process of making the web friendlier for agents, I believe there’s a real opportunity for us to make it friendlier for humans as well.
