Daily Tech Briefing
AI 科技速览
每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。
Ars Technica · 2026/7/27 20:12:42
“Google and Reddit do not own the Internet," web scraper says after court win
AI 中文解读
谷歌和Reddit输了一场关键官司,法院裁定它们不能阻止第三方爬虫抓取搜索结果并出售。核心亮点是:互联网不属于任何一家公司,数据公开抓取在特定条件下是合法的。
通俗来说,谷歌和Reddit一直想阻止别人用自动化工具批量抓取它们的网页内容,再转手卖给其他企业。这次它们起诉了一家名叫SerpApi的公司,理由是后者绕过了防爬技术,把搜索结果的“知识面板”内容打包出售。但法院认为,这些内容本身来自公开网络,谷歌和Reddit不能拿版权法当挡箭牌搞垄断。
这事对普通人的影响不小。如果你常用AI工具或第三方搜索引擎,它们背后往往依赖这种数据抓取来提供结果。如果谷歌赢了,很多免费或低价的替代服务可能消失;现在法院支持爬虫,意味着更多创新服务可以低成本获取公开数据。另外,这也给AI模型训练的数据合法性定了调——只要不直接复制受版权保护的核心内容,爬虫大致安全。不过谷歌已经表态要继续上诉,这场关于互联网开放性的拉锯战还没完。
After a big court loss last week, Google has confirmed that it won’t give up its fight to block AI bots from scraping its search results. And Reddit is weirdly along for the ride.
Curiously invoking the Digital Millennium Copyright Act (DMCA), Google sued SerpApi last December. The search giant accused the web scraper of circumventing its anti-scraping technology and then selling content scraped from Google search results through an unauthorized “Google Search API” software service.
According to Google, the anti-scraping tech was in place to protect copyrighted content in search results. Allegedly, SerpApi’s circumvention threatened to disrupt Google’s relationships with rights holders, including some who license content to Google to appear in so-called “knowledge panels” that are displayed in some search results for well-known people or entities.Read full article
Comments
分享
阅读原文 ↗