Community Discussion · Policy

Website adapter: An interesting move

Can't Finish Reading PapersCan't Finish Reading PapersAug 202026/08/20 296 views

Title: Website-to-API adapter—this move is quite interesting

The most valuable info in this article is: Hermai made an open-source command-line tool that turns any website into a structured API, directly callable by AI agents. Simply put, it turns the entire internet into an interface pool that AI can read.

I've been watching this direction for a few days. Let me explain two terms first. An API is an interface for programs to talk to each other; think of it as a food ordering window—you pass in the menu, it brings out the dish. An AI agent is an AI program that can do work autonomously, like helping you research, book hotels, or organize spreadsheets. In the past, if agents wanted to use website data, they had to rely on crawlers to scrape HTML and clean it themselves—slow and fragmented. Hermai's approach is: turn the entire site into a RESTful API full of interfaces, allowing agents to fetch data directly by field.

I recently started batch-extracting paper PDFs using visual crawlers like Web Scraper, clicking around configuring selectors, getting stuck on dynamically loaded pages, wasting a whole morning. Seeing Hermai's method, my first reaction was: Isn't this just automating all the work crawler engineers used to do?

Register once, reuse long-term

Hermai's core is registration-based. You type a command in the terminal, it probes the website structure, then generates a set of URL interfaces. Later, when an agent needs data, it directly requests this interface with parameters and gets JSON back. No need to care about what the page looks like, no writing CSS selectors, no playing cat-and-mouse with anti-scraping mechanisms. What's this like? It's like stripping out the website's skeleton separately, hanging it outside the window, so passing agents can just take it.

Compare the traditional way with my experience using this over the last few days:

Dimension Traditional Crawler Hermai-style Website-to-API
Data Acquisition Digging fields out of page HTML Getting structured JSON directly
Anti-scraping Battle High cost, often fails Essentially legal interfaces, rate-limited on demand
Maintenance Breaks when page redesigns Interfaces relatively stable after registration
Entry Barrier Need to write code, configure environment Done with one command line

I tried it on a dynamic data page I had on hand—it was indeed the kind of dynamic table regular crawlers can't handle. After running the registration command, the agent fetched data, and it worked in about ten minutes. To rephrase this experience: Before, you hired a temp worker to come to your house to copy things; now you installed a public pickup locker, anyone can come and take stuff.

Turning any website into an API—the more I think about this idea, the more I feel it hits the nail on the head. For the past two years, everyone's focus has been on visible algorithmic capabilities like model parameters, context length, and multimodality. But the bottleneck for actually letting agents do work often gets stuck on external data not connecting. Webpages are for browsers to view, not for programs to read. If you let an agent look at an image, it understands; but if you ask it to operate a website directly, every step requires writing tools, configuring permissions, handling exceptions—this layer of paper is particularly thick.

Hermai pierces this paper via community drive. It doesn't create a closed official marketplace; instead, it lets all users register, contribute, and share. Everyone finds websites they frequently use and uploads an interface mapping. When others need it, they can just search and use it. This equals spreading all web data entry points on the table, where anyone can pick up a bite.

This path will bring an ecosystem, which I think is the part with the most imaginative potential. I've discussed data barriers with many people, always feeling they are more severe than technical barriers in models. Model open-sourcing is hot, but high-quality text is burned out; whoever can feed data in holds the key to the next iteration. Hermai's tool is essentially trying to sign "interoperability agreements" with various data sources, using a standard format to let AI talk to any website. Currently, it's still an open-source project with no obvious commercial prospects, but when the interface pool is big enough, the entry point value will emerge on its own.

Not without concerns

Of course, I'm not only saying good things. There are a few issues in this direction I must mention.

Website types vary wildly. After registering interfaces, there's a core question: Will website owners cooperate? Some sites have public APIs but not REST-style; some sites survive on paid content. Turning their content into interfaces is like inviting court summons. Although open-source projects don't claim ownership, if you actually do it, someone has to bear copyright and legal risks.

Dynamic sites are still hard. The pages I tried were static rendering and simple JS rendering. For backend systems heavily dependent on user login states, like bank online banking or internal corporate knowledge bases, the interfaces Hermai registers likely won't get complete data. Its scenarios might mostly be public information aggregation combined with agent search, rather than replacing all websites' proprietary interfaces.

Can the open-source community take off? The core registration and contribution mechanism needs enough participants. If no one uploads new sites in the first two months, the tool slowly becomes a self-indulgent toy. I checked GitHub update frequency; it's active recently, but open-source heat rising and falling is normal. Projects needing ecological nutrients have a danger period after six months.

However, these doubts don't affect my overall judgment. I clearly believe this direction is a medium-to-long-term inevitable path. Structured data on the web is a huge sleeping asset. Over the past twenty years, search engines built trillion-dollar market caps relying on it; now it's AI agents' turn to eat. Between model capability and external data sources, the last mile will definitely have a bonus period for bridge tools. Hermai is the most direct attempt I've seen during this time.

If you're also working on agent-related things, my action advice is: Install hermaicli, pick a website you visit most, and try the registration flow. Half an hour will show you its depth. Focus on how it handles dynamic content; this directly determines if it can scale. Don't rush to set expectations; just start using it first.


📌 This article is compiled from ProductHunt. Original: https://www.producthunt.com/products/hermai-brand-api

All rights reserved by the original author. This is a compilation and independent analysis based on public reports.

3 replies

?
Ctrl + Enter to reply
Engineer Xue

Tried this plugin. The dynamic pages are actually rendered by a headless browser, so it can handle them. But you still can't bypass login walls and anti-scraping measures. Essentially, it just automates the scraping workflow; you have to configure the token yourself for authentication.

Yanshi
YanshiAug 20

Getting blindsided by dynamic rendering is so real... I tried scraping a data table before and spent ages tweaking selectors, only for the page to go blank on refresh. If this tool can actually handle those dynamic pages, it's definitely something.

Pixel Dust

That part about visual crawlers with selectors is too real—dynamic loading just freezes everything. But does this tool still work smoothly when hitting login walls or anti-scraping sites? Hope it doesn't end up only being useful for static sites.

Website adapter: An interesting move - Physix Frontier Forum