Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • IMO this relates to who the end consumer of websites ends up being. All the arguments for just using MCP ignore the reasoning that agents need based on contextual clues (e.g. if you're booking seats at a movie theaters, you'd want to choose based on personal preferences and taking the movie itself into account) and treat the website itself as the first class product.

    Developing MCPs only works if the capabilities for the site don't involve any iterative reasoning and acting for the user, which things like travel and dining do :p

  • I think pinning WebMCP to the browser was the wrong choice. To list or call tools, you have to have access to the browser environment. "Declarative" WebMCP allows you to read the HTML, but to call tools you still need the JS execution env.

    This limits the utility of WebMCP to in-browser agents. In practice, what's the use case? Computer Use can work, but at that point you can circumvent WebMCP entirely.

    IMO the better option would be to unify MCP with a standard protocol, like HTTP. Right now, you have to rely on the MCP registry for discovering MCP servers, but wouldn't it make more sense to have this natively in HTTP?

  • This is interesting, but instead of tying it to a specific JS library and all I’d prefer putting a list of options and their arguments into the .well-known tree.

    https://en.wikipedia.org/wiki/Well-known_URI

    That way any user agent can get the data and make a call directly to the right endpoint with the right data. It could be an AI agent, a web browser, a specialized tool, a cURL command, or whatever.

    https://agenticresourcediscovery.org/how_to_publish/ and the caldav stuff might serve as decent examples of where to start.

  • It’s a nice idea but I don’t know if it will help.

    The history of the web is one big loop of:

    1. Nifty feature that’s suitable for a wide variety of user agents, from GUI browsers to screen readers to text interfaces to automation.

    2. Approximately 0.1% of web sites adopt it properly. The rest hack up something that mostly works for most of their visitors while ignoring any use case outside of a human driving a mainstream browser.

    3. Automation works around it by pretending to be a human driving a mainstream browser.

    4. Other cases should be handled better, so let’s add a nifty new feature that’s suitable for a wide range of things. Goto 1.

    See: forms, CSS, “semantic web.” Even REST APIs mostly exist as monstrous piles of client-tied functionality that can’t be reasonably used from a client other than the company’s official apps and web site.

  • Why would I want this client/browser side though?

    The only usecase I can image is if you have some terribly complicated legacy frontend and need to work around the logic built there. Otherwise just make an AI chat window or MCP which operates server side and let data sync back to UI from there.

    by dmix
  • One argument for the client/browser side: when the logic only exists in the frontend (validation, pricing rules), a server-side MCP would have to reimplement it and the two copies drift. Running the action inside the user's own session also means no long-lived API tokens handed to the agent — the agent's access dies with the tab.
  • It took me some time to wrap my head around why WebMCP even exists. I was thinking that regular MCP + SSE/Websockets could cover almost all of the uses cases.

    One interesting use case for WebMCP is cross-site activities that a browser agent might take. If the data-flow in question involves a heterogeneous set of websites.

    This feels like a lifeline for incumbent web SaaS more than anything. Future apps can be architected in ways that do not require WebMCP but massive apps like Salesforce or Workday can't really abandon the decades of accidental business logic embedded in the Web UI flows that make up their project. While those same flows could be data-driven, slapping a WebMCP facade on top of them is just more practical. Instead of forcing incumbents to create an AI-native MCP where they are first class citizens, it is easier to let them sprinkle browser affordances throughout their human-centric front-end and offload the work to the agent.

    But for my own part, since I am developing from scratch without that legacy need, I actually think relying on WebMCP might be an anti-pattern. It might be useful for the cross-site use case and it might be useful for reducing latency for purely UI activities (e.g. "filter this list" where all the data is already on the client), but in general my feeling is it is better to have a robust MCP interface for agents.

  • This feels inverted and impractical in so many ways.

    On one hand the claims are being made that AI is smart enough to replace software engineers and on other hand the website owners are beings asked to provide information in a certain format to the Agents so they can do their job better. Remember this is the same information that every regular user is able to use.

    Secondly if you maintain two versions of information, one for regular humans and one for agents, its just a matter of time before they start to diverge from each other. One can pick up and compare the native app and Browser application for any company, 99% chances are they are not exactly the same.

  • This is not a smart reply. Structure is always good otherwise why even create unit tests and compilers?
  • I agree it's not a good idea, but:

    > Remember this is the same information that every regular user is able to use.

    This is not true. The users don't see the HTML, CSS and JavaScript. If they would, even if they're frontend devs, they'll be just as confused as the agents.

    The user see images, and while you can make the agents see the image as well, unlike humans the effective agents we have today work with text, not images.

  • I think it's an efficiency thing. And at scale, efficiency matters. Whether this is the right implementation idk, but when I see Claude Code use Chrome via MCP, it's incredible to see just how inefficient the status quo is.
  • Nice write-up. Btw, Codex and ChatGPT desktop support WebMCP in their integrated browser as of yesterday.

    One way to think about WebMCP, as the article notes, is as web accessibility for agents. But since it functionally allows developers to expose any function on their website to an agent via RPC, the more interesting use cases are the ones that use the browser as a sandboxed execution environment.

    https://duckboard-webmcp.alexmnahas.workers.dev/ is an example of this, and I would love to see more.

    Maybe instead of skills that require users to download arbitrary binaries to their computers, the skills could instead point to a domain that comes preloaded with that binary in Wasm and lets the agent interact with it over WebMCP?

  • The Apple standard-positions in this was a pretty good read, noting indeed that this is a parallel new way to define accessibility. https://github.com/WebKit/standards-positions/issues/670

    I'd also throw in schema.org Potential Actions as longstanding precedent for declaring what is on offer from the page. https://schema.org/docs/actions.html

    It does get a bit tricky though. WebMCP allows the page to respond to the agent. Where-as these other systems the page itself has to update and respond.

  • One sample I have tried in a Pen Fight Game (old school nostalgia) Agent v/s Game bot (Audience can enter and alter the game -still early stages) https://mela-web-production.up.railway.app/
  • we already have something better than WebMCP - i.e general APIs.

    more impactful work has to be done on the native legacy desktop app scene. that's where most of major companies big or small do their work.

    agents automating websites is kinda easy. automating legacy desktop apps that's another issue though there's RPA.

  • APIs and WebMCP don't have the same purpose:

    - WebMCP allows agents to assist humans on their own interface ie the web UI.

    - APIs bypass the human interface completely

    And I'm not sure most of the work today is done on legacy desktop apps. We're in 2026, not 2016

  • This is going to sound incredibly silly but I have every HTMX component expose an AI summary with actions possible and then I have a small copy to AI button on the page which creates a new short-lived access token and copies into a prompt the token, and all the components as text.

    When I give this even to relatively small-model agents, they use this initial seed to browse the site and do things very well.

    The initial prompt suffices. After that most agents just use the HTML to navigate very well. I suppose I could have that description in an aria label if I wanted but it’s the same.

  • I'd be interested to read a blog post with more details about how this works.
  • > Here’s a book_table tool. It takes a date, a time, and a party size. Call it.

    Why not offer a simple form that humans and AI can use alike?

        <form action=book_table>
          <input type=date name=date>
          <input type=time name=time>
          <input type=number name=party_size>
          <input type=submit value="Book table">
        </form>
    by mg
  • That’s part of webmcp :)