How Search Engines Classify Websites by Source Type
Website classification is the process a search engine uses to decide what a site is before judging what its pages say: expert or amateur source, commerce or content or service. Built from site-level representation vectors, it changes how every document on the domain is read, trusted and ranked.
On this page — 5 sections
What Is Website Classification?
Quick answer
The engine-level process of deciding what a website is: the Website Classification System builds site-level representations and sorts sources into classes — expert, apprentice or amateur; commerce, content or service. It exists because direct document understanding is minimal.
Website classification is the site-level complement to document understanding. A Website Classification System generates representations for websites and then classifies them — determining, in the documented implementation, whether a site resembles an "expert, apprentice, or amateur source" using "visual and layout-related embeddings and features," and distinguishing among "affiliate sites, aggregators, service providers, ecommerce sites, and SaaS platforms."
The reason this layer exists is the same reason behavioral signals exist: an engine's direct understanding of documents is minimal. Classifying a whole site once — by layout, components and function — is a cheap, stable substitute for re-understanding every page from scratch, and it supplies the context in which each page is read: what the site can do, whom it serves, and how much trust its statements start with.
Which Signals Feed the Classification?
Quick answer
Website representation vectors assembled from visual and layout-related embeddings: HTML tokens, the DOM tree, page layout and hierarchy, visual relationships, annotations, and the functional meaning of each structured information block — then compared against known site classes.
The core mechanism is the website representation vector: a composite representation assembled from feature vectors derived from the site's pages. The inputs named in the processing literature are mostly structural — "HTML tokens," the "DOM tree," "page layout," "hierarchy," "visual relationships," "annotations," and the "functional meaning of each structured information block." Models such as WebFormer incorporate the HTML layout into the document representation, and site-wide vocabulary features round out the picture.
- Layout and hierarchy: how pages segment into regions and which region dominates.
- Functional components: the interactive elements that say what the site does — purchasable, bookable, comparable, calculable.
- Site-wide vocabulary: the nouns and phrases that recur across the domain, which make a site labelable within its knowledge domain.
- Crawl attention: the crawl rate itself is read as a measure of how much the engine values the site.
Classification is comparative: a new site's representation is measured against composite representations of known classes, and the closer it sits to a class, the more that class's expectations transfer to it. There is no application form — the class is inferred from what the site shows.
Why Does the Same Content Rank Differently by Source Type?
Quick answer
Because engines classify websites by their type, not only their content quality: the same article on an affiliate site and on an ecommerce site meets different quality thresholds and different intent expectations, so the two sources are not interchangeable in ranking.
The sharpest documented consequence: "the same content can rank differently on an affiliate website than it does on an ecommerce website," because engines "classify websites by their type rather than their content quality." Different types carry different quality thresholds and different intent expectations. An ecommerce site is a task-completing commercial resource built for transactional intent; an affiliate site typically aggregates information to direct traffic outward, and it earns its place only by demonstrating real effort, originality and added value over the sources it draws on. One documented move — the same content and context relocated from an affiliate site to a commercial one — improved rankings for exactly that reason:
"The content itself didn't change. What changed was the function, context, and source type surrounding it."
— Semantic SEO source-type case study
Intent alignment completes the logic. E-commerce pages are expected to serve "Do" queries — purchasing, comparing, ordering — while affiliate-style content more naturally serves "Know" queries that precede the transaction. When the page's demonstrated capability matches the intent its source type implies, the engine spends less effort deciding whether the result fits.
| Source type | Built to serve | What the engine reads from it |
|---|---|---|
| E-commerce | Transactional "Do" intent: buying, ordering, comparing products | Functional components for purchasing, comparing, ordering, reviewing and filtering; product data in structured cards |
| Affiliate / content | Informational "Know" intent that may precede a purchase | Genuine effort, originality and added value — thin aggregation meets stricter thresholds |
| Aggregator | Collecting and organizing sources at scale | Clear added value over the underlying sources, or the classification turns against it |
| Service provider | Local or professional service intent | Task-completing pages: booking, contact, verifiable service information |
| SaaS platform | Product use and self-service | Working functionality the layout makes obvious, not descriptions of function |
What Does the Classification Decide?
Quick answer
Eligibility and trust: whether a source enters the candidate set at all, how site-level scores such as siteAuthority and Q* treat it, how queries in its domain are interpreted, and whether generative AI systems consider it worth citing.
Classification decides eligibility — whether a source is "even in the game" for a query class — before any position is contested. It feeds the site-level trust scores (siteAuthority and the broader Q*), helps the engine interpret queries by associating them with knowledge domains and site types, and it is integrated with the Helpful Content System, whose classifier leans on the same function-and-layout evidence. Generative systems use it too: classification informs which sources are appropriate to cite in AI answers.
The classification is also stable but not frozen: layout changes "can produce different vector representations, which may affect how a document is understood, classified, and retrieved." A site can drift between classes when its presentation drifts — which is why classification is a standing property to maintain, not a one-time verdict.
How Do You Work With Classification Instead of Against It?
Quick answer
Make the site's type legible: state the purpose the page serves, ship the functional components that type implies, keep visuals consistent with actual function, hold one coherent site-wide focus, and avoid the abuse patterns that force a negative classification.
None of the levers are exotic; they all make the site's type unambiguous:
- State and show the purpose: the page layout and its components should make the business model and the user task obvious without reading a word of marketing copy.
- Ship the function the type implies: commerce means working purchase, compare and filter paths; a service means booking and contact that work.
- Keep appearance honest with function: never imitate a function the page does not provide — misleading functionality is classified as spam.
- Hold one coherent focus site-wide: consistent topics and vocabulary make the site labelable; scattered topics are described as a self-inflicted eligibility problem.
- Stay out of the abuse classes: scaled production, borrowed reputation and recycled expired domains all end in negative classifications and demotion.
This article is part of the Search Engine Understanding & SEO series — How search engines read queries, pages, layout and user behavior, explained in plain terms with service-business examples.
About the author
Mohamed Youns
Semantic SEO Engineer · Author & system developer
Mohamed Youns writes about how search engines understand content — the same standards he applies when building semantic systems at Nut Hub. nut-hub.org