Meta Scrapes the Web Free While Rivals Negotiate With Google

▼ Summary
– Meta’s crawlers (Meta-ExternalAgent and Meta-WebIndexer) now carry the majority of AI agent traffic, with growth of 74% and 163% quarter-over-quarter, while GPTBot remains the most-blocked AI crawler.
– Publishers focus their blocking and licensing negotiations on Google because it historically drove traffic, whereas Meta never promised visits, so its heavy reading doesn’t register as a loss.
– If platforms keep users on-site, publishers have little to negotiate for beyond one-time licensing deals, which only benefit top brands like News Corp, CNN, and Fox News.
– The article suggests publishers should build direct audience connections and owned channels, moving toward a decentralized web, as the only honest way to win.
– Meta’s crawlers cost websites almost nothing directly, but the issue is Meta profiting from uncompensated work, and current access terms will shape future machine interactions.
The machine-access fight consuming the publishing world is aimed at the wrong target. Publishers are locked in high-stakes negotiations over whether to block Google’s AI or license content to it. Reddit devoted a significant portion of its July 30 earnings call to this exact dilemma, and every development at that table receives summit-level coverage. Yet the web’s largest machine reader belongs to Meta. It consumes more content than every blocked bot combined, returns virtually nothing in referral traffic, and the question of whether Zuckrawlers should be blocked barely registers in public discourse.
Meta’s crawlers now dominate AI agent traffic while GPTBot remains the most-blocked bot. DataDome, a bot-defense vendor, tracked 17.7 billion AI agent requests across its network during the second quarter of 2026, a 45% jump from the previous quarter. That growth did not originate from Google or OpenAI. Meta-ExternalAgent surged 74% quarter over quarter, while Meta-WebIndexer climbed 163%. Together, these two agents now carry the majority of AI agent traffic DataDome observes, with almost no referral traffic flowing back to publishers. A necessary caveat before continuing: DataDome sells bot protection, so its telemetry carries inherent commercial interest. Treat the exact figures as one network’s perspective. The broader trend is far harder to dismiss, particularly when examining robots.txt files. By the same vendor’s count, GPTBot remains the most-blocked AI crawler on the web. The bot the web built its defenses around is not the bot doing the heavy reading.
Google earned its seat at the negotiating table by paying in traffic; Meta never owed anyone a visit. This disparity stems from the historical roles these companies played. Google drove traffic. For two decades, the arrangement was straightforward: Google reads your site, Google sends you visitors. When Google’s AI summaries began retaining those visitors, publishers experienced it as a broken promise. That is why the response resembles a renegotiation, with blocking, licensing, and litigation as the primary levers.
Meta, by contrast, has always focused on keeping users inside its walled garden, a model Google now aspires to replicate. Nobody ever expected a visit from Meta, so when its crawlers became the heaviest readers of the open web, there was no promise to break. It simply does not register as a loss. Yet here is the uncomfortable reality: the company that perfected the keep-the-user model is now the biggest consumer of everyone else’s work, while the company being renegotiated with is busy imitating that very model.
If platforms keep the user, there is nothing left to negotiate for. Publishers hold no meaningful cards if platforms aim to retain users for themselves. The traffic that made the old arrangement viable is precisely what is being phased out, shrinking the bargaining chip to mere payment for content that trains and feeds AI systems. That is a losing position for nearly everyone, unless you are already massive. Meta does write checks, but only at the very top. In March 2026, it signed a licensing deal with News Corp valued at up to $50 million annually, covering both Meta AI answers and model training, alongside similar arrangements with CNN, Fox News, USA Today, and a select group of others. That table seats a few dozen of the biggest media brands while Meta’s crawlers read everyone. Reddit can put a licensing deal and a public maybe-we-walk on an earnings call. A store, a blog, or a trade publication cannot do that in 2026.
That leaves one honest path for publishers to win: “Fine, we can do this without you,” followed by actually building that reality. A stronger direct connection with the audience, owned channels the platforms cannot dilute, perhaps something resembling a step back toward the decentralized web and away from the platform-centered one. None of that is quick, and none of it runs through the meeting rooms where the Google negotiation is happening.
Meta’s reading costs you almost nothing, but that is not the point. Meta never sent websites traffic and never promised to, so its crawlers reading billions of pages takes nothing you ever had. It costs you almost nothing in direct terms. But Meta is building its business on the backs of people doing real work without compensating them. That is wrong, and it remains wrong whether or not any single website can feel the cost.
What DataDome reported is only the first behavior of this visitor class. The same machines reading your website today are the ones that will pay to read it or buy from it tomorrow. The access terms the web sets by brand recognition now are the terms it will live with then.
So what does someone with an ordinary website actually do about Meta? The honest answer is nothing today, but keep an eye on it. Keeping an eye on it means knowing who actually reads your website rather than which AI company dominates the headlines, because the two are diverging. The negotiation everyone can see is with Google. The reader almost no one seems to be watching is Meta’s. Your log files know the difference, even if the headlines do not. Pay attention to them, not the headlines.
(Source: Search Engine Journal)




