Human-reviewed websites

About Discovery Engine

Discovery Engine is a human-curated, searchable collection of websites worth spending time with.

In one sense, it is a public bookmark manager: people suggest links, a person decides which websites belong, and approved sites become part of a shared collection. It looks for originality, substance, usefulness, and integrity—not simply for what is already popular.

Unlike an ordinary bookmark list, Discovery Engine crawls a small number of pages from each approved website and indexes their visible contents. That makes it possible to find a website by the subjects and ideas it contains, rather than only by its name or manually assigned tags.

It is not intended to crawl the whole web or act as a general-purpose search engine. The website remains the unit of discovery; matching pages are shown as supporting references for why a result appeared.

Methodology

Anyone can suggest a URL. A human reviews the website and decides whether to approve or reject its domain. Approved sites are crawled from the submitted page, with a small fixed page limit and respect for each origin's robots policy. Visible page content is converted to Markdown and added to the search index after another explicit index build.

Search remains one result per approved website. When one crawled page fully matches a query, that page and a short matching passage appear as a reference beneath the website result. External websites discovered during a crawl return to the ordinary suggestion queue; discovery never counts as approval.

Suggestions and status

Suggesting a URL does not guarantee review, crawling, approval, or publication. Duplicate suggestions and suggestions for previously rejected domains may be accepted silently without creating another queue entry. This avoids exposing private review decisions through the suggestion form.

Approved and rejected websites may have public status pages after an index is published. “Approved” records a human decision about the website. “Indexed” separately means that its current retained pages are represented in the active searchable snapshot. Either status may become stale until the next operator-run crawl and index publication.

Privacy

The application stores suggested URLs and submission times, but no submitter account or profile. It stores a SHA-256 hash of the submitting network address in abuse-throttling records rather than storing that address with the suggestion. Records older than one day are pruned when later suggestions arrive. The hosting provider may retain ordinary web-server access and error logs separately.

Fetched source HTML and extracted Markdown remain in the operator's private crawl database. The public site uses a generated index that contains reviewed site information and searchable extracted text; its database files are not directly downloadable. The interface loads Bootstrap from jsDelivr and JetBrains Mono from Google Fonts, so visitors' browsers may contact those providers.

Limitations

This is a small human-curated collection, not a comprehensive or continuously updated search engine. Decisions are subjective and do not constitute an endorsement, security review, or guarantee that a website remains unchanged. Search is keyword-based and may miss relevant sites or rank imperfect matches.

Crawls are deliberately bounded, do not fetch images or downloads, and do not execute JavaScript. Sites that require browser rendering, reject the crawler, prohibit access through robots.txt, redirect outside the allowed origin, or are temporarily unavailable may have incomplete or no searchable content.