About PageSourceSearch and its crawler
PageSourceSearch is a search engine over the source of web pages: the HTML of a site's pages and the JavaScript the site itself serves, searchable by exact bytes or by regular expression. The crawler that collects it identifies itself as PageSourceSearchBot and links here from its user agent.
What the crawler does
Its user agent is
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PageSourceSearchBot/1.0; +https://pagesourcesearch.com/bot) Chrome/128.0.0.0 Safari/537.36
It sends the request headers a current Chrome sends, so a site serves it the same content a visitor would get. For each site it fetches, in this order and never more:
/robots.txtfirst, then the root page/;- up to 4 further pages linked from the root (chosen to expose different page layouts), at most 5 pages in all;
- the scripts those pages reference (
<script src>and preloads), following a script's own imports up to 3 levels deep, first-party hosts only: a script on a public CDN or tag host is noted but never fetched; - at most 200 files, 4,000,000 bytes each and 20,000,000 bytes per site in total.
Requests to one host are made one at a time, at least 1000 ms apart by default (or your Crawl-delay, when larger),
and a 429 or 503 answer backs it off (Retry-After is honoured). It does not log in, submit forms, run your scripts or fetch images, fonts or stylesheets.
A live site is revisited about once a week; a site that answers with a bot challenge is not tried again for a year.
robots.txt
The crawler reads /robots.txt before anything else and follows it (RFC 9309: the group for its token, then the * group;
Allow, Disallow, * and $ patterns, Crawl-delay). Its token is
User-agent: PageSourceSearchBot
If robots.txt cannot be fetched because of a server error, the site is treated as fully disallowed and nothing else is fetched.
How to opt out
To keep the crawler off your site entirely, add this group to your robots.txt:
User-agent: PageSourceSearchBot Disallow: /
To keep it out of part of a site, Disallow those paths instead. The change takes effect at the next visit, which starts with
robots.txt: a disallowed root stops the visit before any page is fetched. To have already stored pages removed, or for any other question
about the crawler, write to [email protected] with the domain name.
What is stored and shown
The crawler keeps the exact bytes of the pages and scripts it fetched, with the URL and the time of the fetch. A search result shows short
snippets around each match, and the preview page shows the stored file as text. Nothing is executed, rendered or served as a live page,
and the preview and match pages are marked noindex so search engines do not index your source through this site.
Only a site's own code is searchable: library code (jQuery, React and the like) is recognised, stored for the preview and left out of the index.
Searching the index
How queries work, from exact bytes and several terms to regular expressions, filters and why a query can be rejected, is on the query help page, together with the FAQ and a note on the JSON API.