Topic: web scraping
-
Reddit revives its unusual DMCA battle over Google results
A federal judge allowed Reddit's lawsuit against SerpApi to proceed, ruling that Reddit plausibly alleged SerpApi conspired with Perplexity AI to bypass Google's access controls and harvest copyrighted Reddit content. This ruling contrasts with a recent dismissal of a similar Google case, where t...
Read More » -
Google Lost Its Scraping Case: The Open Web Forces a Choice
A federal judge ruled that Google's DMCA anti-circumvention claim against SerpApi for scraping search results fails because SearchGuard protects ad revenue, not a copyrighted work, dismissing the claim where no copyrighted content was involved. The ruling upholds the principle that content on the...
Read More » -
Court dismisses Google's DMCA lawsuit against SerpApi
A federal judge dismissed Google's DMCA lawsuit against SerpApi, ruling that search results (URLs, snippets, and index data) are public facts, not copyrighted works, so Google cannot use copyright law to block scraping of its results. Google has 21 days to refile on a narrower claim regarding lic...
Read More » -
Web Standards Set to Reshape AI Content Use
Content creators have faced unauthorized use of their work by AI models, with few existing mechanisms to protect intellectual property online. The IETF's AI Preferences Working Group is developing standardized rules to let website owners control how AI systems access and use their content. This i...
Read More » -
Google Updates SerpApi Lawsuit With Content Licensing Terms
Google filed an amended DMCA complaint against SerpApi on August 10, adding specific licensing contract language,including agreements with Reddit and a 2017 licensing partner,to address the court's earlier finding that Google lacked documented authorization for its SearchGuard access control. The...
Read More » -
AI giants discover the internet's harsh realities
AI companies like Anthropic, OpenAI, and Google are now complaining that competitors are using "distillation" to replicate their models, mirroring the same fair-use arguments they used to justify scraping content from websites without permission. The industry's logic is contradictory: they consid...
Read More » -
Google's SearchGuard: Bot Detection & the SerpAPI Lawsuit Exposed
A lawsuit reveals Google's advanced SearchGuard system, which uses real-time behavioral analysis and environmental fingerprinting to invisibly detect and block automated bots attempting to scrape search data. The case involves SerpAPI, a service accused of bypassing these protections, and highlig...
Read More » -
Web Intelligence Fuels Next-Gen AI Infrastructure
The web intelligence industry has evolved to handle massive multimodal data demands, exemplified by the Video Data API and High-Bandwidth Proxies that enable efficient processing of audio and video for AI training. Headless browsers have become crucial for AI agents to reliably access and interac...
Read More » -
Web Scraper Sues Google, Accuses It of Web Scraping
A legal dispute centers on whether Google's search results are copyrighted, with SerpApi arguing they are not, as Google aggregates public web data. SerpApi defends its web scraping by comparing it to Google's own practices, framing the lawsuit as an anti-competitive move to control data access. ...
Read More » -
Cloudflare Blocked 416 Billion AI Bot Requests Since July
Cloudflare has blocked over 416 billion AI bot requests since July, revealing the massive scale of web data harvesting for AI training and a significant imbalance in access, with Google's crawlers reaching far more webpages than competitors like OpenAI. The situation creates a dilemma for publish...
Read More » -
Cloudflare Blocked 416 Billion AI Bot Requests in 6 Months
Cloudflare blocked 416 billion AI bot requests in six months, highlighting a fundamental shift in data collection driven by large language models and its efforts to reshape the web's economics. The company's CEO argues AI is a "platform shift" altering the internet's core business model, forcing ...
Read More » -
Suno scraped millions of songs from YouTube, Genius, and Deezer
A hacker leaked Suno’s source code and training data, revealing that the AI music generator scraped millions of songs and lyrics from platforms like YouTube Music, Deezer, and Genius, supporting ongoing lawsuits that accuse the company of using copyrighted material without permission. The leaked ...
Read More » -
New App The Mall Unifies Online Shopping into One Feed
The Mall is a new app that lets users create a personalized virtual shopping hub by adding favorite brands, aggregating sales, restocks, and promotions into a single feed to combat fragmented online retail. The platform uses web scraping technology and large language models to track over 10,000 b...
Read More » -
SerpApi Fights Back: Seeks Dismissal of Google Scraping Lawsuit
SerpApi has moved to dismiss Google's lawsuit, arguing Google is misusing the DMCA to protect its advertising business model, not copyrighted works, by trying to control access to public search results. The company's defense asserts it accesses publicly visible data like any user and that Google ...
Read More » -
AI Bot Surge Ignites Online Arms Race
AI bots are projected to drive the majority of internet traffic, shifting the web from a human-centric to a bot-dominated landscape and introducing new challenges beyond copyright. A technological arms race is escalating as sophisticated AI bots increasingly circumvent website security to scrape ...
Read More » -
Google Loses Key DMCA Claims in SerpApi Scraping Lawsuit
A federal judge dismissed key DMCA claims by Google against data scraping service SerpApi on July 20, permanently tossing claims tied to search results lacking copyrighted material while giving Google 21 days to amend claims involving copyrighted content. The court found Google failed to provide ...
Read More » -
Code Formatting Sites Leak User Secrets and Credentials
Popular online code formatting platforms like JSONFormatter and CodeBeautify are leaking sensitive user data, including passwords and API keys, through publicly accessible links due to predictable URL patterns. Security researchers found over 80,000 exposed entries containing critical information...
Read More » -
Britannica and Merriam-Webster Sue Perplexity AI Over Copyright Claims
Encyclopedia Britannica and Merriam-Webster have sued Perplexity AI for copyright and trademark violations, alleging unauthorized content scraping and traffic diversion. The lawsuit claims Perplexity copies content verbatim, misuses brand names with inaccurate responses, and bypasses technical ba...
Read More » -
SerpApi Fights Google Over SERP Scraping Lawsuit
SerpApi is seeking to dismiss Google's lawsuit, arguing that Google cannot claim DMCA copyright protection for aggregated third-party content it merely displays, as the law is meant for actual copyright holders. The company disputes Google's legal standing and the characterization of its actions ...
Read More » -
Navigating Ethical Proxy Sourcing Challenges
Proxy servers are essential for AI infrastructure, enabling web scraping and data collection that trains large language models, but their growing use has raised critical concerns about ethical sourcing. Many proxy IPs are obtained from residential users or compromised devices without clear consen...
Read More »