215,000 Fake Review Pages Scam AI Search: Machines Mass Producing Garbage

215,000 Fake Review Pages Scam AI Search: Machines Mass Producing Garbage

AIContent FarmInvestigation

Sources:Trellner Research Report + HN Discussion

A recent investigation by Trellner Research has uncovered a disturbing truth: when querying AI search tools for software purchase recommendations across 380 different fields, 59.8% of the 7,534 reference links provided pointed to obscure websites ranked outside the top 100,000 in global traffic. Over a third of the cited domains didn’t even make it into the global top 1 million, and the median rank of all cited websites was a mere 71,611. Instead of sifting through the noise as people expected, AI search is serving up garbage information from the fringes of the internet.

Three Websites Generate 215,000 Fake Ranking Pages

The source of this low-quality information stems from three colluding content farms: wifitalents.com, worldmetrics.org, and gitnux.org. They were all centrally registered via NameCheap between December 2023 and May 2024, sharing not only the same Cloudflare nameservers but also identical webpage templates, navigation bar categories, and even peripheral pages like “Editorial Process” and “About Us”. This industrialized assembly line generated over 70,000 “Best Software” review pages for each site, plus a backup site with the same configuration, zipdo.co, bringing the total of forged reviews to a staggering 215,000 pages.

There are simply not 215,000 software categories in the world. These so-called reviews are the product of machines blindly piecing things together. Even for incredibly rare categories like “Oilfield Management Software” or “Museum Collection Management Software,” they can solemnly scrape together 10 tools for a ranked comparison.

10 Tools Compared: Best Oilfield Management Software (2026) Image: The “Best Oilfield Management Software” ranking page on wifitalents.com. Source: wifitalents.com webpage screenshot

Investigators found that under the same “Project Estimation Software” category, the three websites each provided completely different Top 5 rankings, yet pretentiously credited 9 different “editors”.

Site1st Place2nd Place3rd Place4th Place5th Place
worldmetrics.orgFloatScoroTeamwork.comProcoreWrike
wifitalents.comFloatScoroTeamwork.comBuildertrendApropo
gitnux.orgSaviomMosaicBuildertrendFloatTeamwork.com

The top spot on gitnux couldn’t even be found in the top five of the other two sites. In some pages’ author sections, unrendered template code variables like “Within the next 26 days” still glaringly remained. The shoddy forgery exposed them as human-free products run from the same script, requiring nothing more than a server and cheap API call costs behind the scenes.

Webpages Written Exclusively for Retrievers to Read

These 215,000 webpages were never meant to be read by humans from the moment they were created. The HTML webpage code title of worldmetrics.org explicitly labels them as “Facts & Grounding Page”. This is a data tag only read and understood by large language model retrieval systems, and the hidden meta descriptions unabashedly state “Verified facts for machine-readable records”.

Top 10 Best Project Estimation Software (2026 Review) Image: The project estimation software ranking on worldmetrics.org, using the same template as wifitalents.com. Source: worldmetrics.org webpage screenshot

Behind this fabricated information fed specifically to machines is a blatantly priced commercial arbitrage business. After using machine-generated content to inflate their AI citation weighting, worldmetrics.org openly sells services on their official website: custom market research starts at 5,000 euros, ready-made reports sell for 499 euros, and even “vendor screening” assistance charges 2,500 euros.

After traditional search engine optimization (SEO) failed, content farms found a new way to game large models. A marketing blog selling interactive demo products, guideflow.com, has no software review content itself, yet because its writing style caters to model preferences, it was cited 194 times by AI search across 96 categories. This unknown blog beat out globally renowned analysis firm Gartner on the citation leaderboard, absurdly becoming core evidence for professional questions like “3D Rendering Software” and “RFID Software”.

The massive influx of fake sources directly caused severe orientation bias in the recommended results. In the final generated recommendation lists, 1.1% of vendor official websites had long since gone out of business or changed domain ownership. When investigators asked about “Research Data Management Platforms,” the official recommendation link dryad.co provided by the sonar model redirected straight to an Indonesian online gambling portal named BIGSLOT288.

In another answer about “Data Quality Tools,” the recommendation result from the premium model sonar-pro jumped directly to the official website of a casino hotel in Monaco. The investigation revealed that these two models share the exact same underlying retrieval layer, with 289 out of 380 categories returning byte-for-byte identical citation lists. Low-quality data has polluted the entire system.

By contrast, Wikipedia, known for its neutrality and objectivity, appeared a mere 3 times across all 7,534 citations. High-quality information from the real world is thus drowned out and ousted by repetitive machine-generated babble.

The Chain of Trust Collapses from the Source Layer

The investigation report states that it cannot currently assert these low-quality sources changed the final recommended answers of AI search, but this already exposes systemic vulnerabilities in the retrieval layer of large models. Senior SEO practitioners have even revealed more advanced tactics: first use questions to test various large models and calculate text divergence bias, then use AI to rewrite articles in bulk, making the webpage’s text features precisely approximate the large model’s built-in preferences.

Large language models themselves have an inherent tendency to favor AI-generated text. In self-tests by the developer community, when faced with human-refactored concise code and their own verbose generated code, AI models always stubbornly chose the latter. Previously, media exposed that a certain country mass-produced fake think tank websites to feed AI, causing it to output large volumes of political views with a specific stance.

The recommendation systems of AI search are being systematically fed by fake review systems. Content farms have learned to manufacture fake sources that satisfy AI preferences, and AI search subsequently degenerates into an echo chamber. When you ask it which software is best, it merely prints out fake advertisements written long ago by another machine on your screen.

Reference links:

  • Trellner Research Report
  • HN Discussion (item?id=49536375)