← All Posts Eloquent - Web Scraping Monthly Newsletter, August 2026

Eloquent - Web Scraping Monthly Newsletter, August 2026

· Updated 29 Jul 2026
Author
Nishant Choudhary

Founder of DataFlirt. Web scraping engineer helping data and engineering teams extract and operationalise web data at scale.

TL;DRQuick summary
  • The FBI, Google, Lumen, Shadowserver and IRS CI seized hundreds of NetNut domains on July 2. Google put the underlying device pool at more than 2 million.
  • Spur found residential proxy SDKs in over 42% of LG webOS apps and more than a quarter of Samsung Tizen apps. LG will suspend apps that do not remove them.
  • Cloudflare blocks Training and Agent bots by default on ad-supported pages from September 15, 2026, and judges mixed-use crawlers by the most restrictive rule.
  • The EDPB's draft Guidelines 03/2026 treat robots.txt, CAPTCHAs and login walls as signals feeding the legitimate interest test under GDPR.
  • Proxyway found residential rates stopped falling in early 2026, with IPRoyal roughly doubling its 50 GB rate year over year.
  • Oxylabs took $130 million from Warburg Pincus on July 9 at a $3.6 billion valuation.

On July 2, the FBI seized NetNut’s domains. Google put the underlying device pool above 2 million. Teams routing through whitelabel resellers found out mid-crawl. Three more shifts landed the same month. Each hit a different layer of the stack.

The residential proxy supply chain got audited in public

The takedown was coordinated. Google’s Threat Intelligence Group worked with the FBI, Lumen’s Black Lotus Labs, Shadowserver, and the IRS Criminal Investigation division. KrebsOnSecurity reported the seizure of hundreds of domains. NetNut is operated by Alarum Technologies, a publicly traded Israeli firm (NASDAQ: ALAR).

Seized properties included netnut.com and the subsidiary proxyjet.io. So did divinetworks.com. That one sourced static IPs through direct ISP contracts.

The scale is the part worth sitting with. Researchers track the underlying network as Popa. Google’s threat team counted 316 distinct threat clusters using suspected NetNut exit nodes in one week of June 2026. Most enrolled hardware was budget smart TVs and streaming boxes. This follows the IPIDEA disruption in January.

Google was explicit about what comes next. When one operator’s pool degrades, competitors sell them capacity. The reseller becomes the seller. SecurityWeek quoted Google on the fix: lasting disruption means targeting several interconnected providers at once. Expect more seizures.

Here is the operational problem. Google assessed with high confidence that many popular residential proxy brands were whitelabeling NetNut. Buying from a brand tells you nothing about whose pool carries your traffic.

Your proxy vendor’s sourcing is an uptime risk now, not a procurement footnote.

Where a residential IP actually comes from

Four supply paths feed a commercial residential pool. Consent and legal exposure differ sharply across them.

Security firm Spur made this concrete. Spur found residential proxy SDKs in more than 42% of apps on LG’s webOS store. Samsung’s Tizen store came in above 25%. The apps were screensavers, shovelware games and file utilities. Bright Data’s SDK accounted for the majority across both platforms. The company defends its network as consented and audited, citing a second independent PwC review.

LG moved within a month. Krebs reported that LG is working with developers to strip the proxy option from webOS apps. Non-compliant apps get suspended. Amazon and Roku banned the practice already. Samsung has not moved.

How an IP enters a residential proxy pool Four supply paths, one gateway, one exit node on a real home line SUPPLY PATHS Consented SDK in a free app User sees a prompt, developer gets paid Bandwidth-sharing app Honeygain, Pawns, JumpTask, Earn.fm Preinstalled on cheap hardware Smart TVs, streaming boxes, bundled SDKs Malware-enrolled devices Popa, Kimwolf, Badbox plugin components Provider IP pool Millions of consumer IPs, mostly Android devices Gateway applies geo filter, rotation and session length Exit node: a real home ISP line Someone's router, TV or phone carries the request Target site sees a residential IP Reputation checks pass, fingerprint checks still run The whitelabel layer sits between you and all of this Resellers buy pool capacity and sell it under their own brand. Your contract names the reseller, not the pool. Proxyway estimates only 10 to 15 companies run their own residential network, and about half share upstream sources.

That last figure comes from Proxyway’s Proxy Market Research 2026. Only 10 to 15 companies maintain their own residential network. Roughly half overlap upstream. The brand on your invoice and the pool your traffic exits from are often different things.

Vendors reacted fast. Some now require KYC to access ASN targeting. NetNut removed ASN targeting entirely. Bright Data extended KYC from specific target sites to residential access generally. Our Bright Data and Oxylabs comparison goes deeper on pool composition.

Why blocklists cannot fix this from the defense side

Site owners have the mirror-image problem. They cannot cleanly separate a proxy from a customer.

GreyNoise analyzed 4 billion malicious sessions. Its conclusion was blunt. IP signals and static blocklists cannot identify residential proxies on their own. The address belongs to a real household. There is nothing for a reputation feed to flag.

Detection moved up the stack instead. Browser fingerprinting, TLS fingerprinting and behavioral scoring do the work now. A pristine residential IP with a default headless fingerprint still fails.

The supply problem is not shrinking. Aisuru and its Android variant Kimwolf both monetized infected devices as residential proxies. Synthient found Kimwolf spreading through other proxy networks. It reached over 2 million devices by January 2026. Authorities in the US, Canada and Germany disrupted several of these botnets in March. Google separately shut down around 10 Chinese proxy brands linked to IPidea, including 922Proxy and LunaProxy. More than a dozen China-linked brands still trade.

The FBI also published a consumer bulletin about residential proxy networks. That is new for this category.

Machines now move more of the web than people

Load is the other half of the story. LWN revisited its 2025 piece on scraper bots this month. The finding was unwelcome. The problem is still growing.

The traffic shape is what makes it hard. LWN describes coordinated requests from millions of unique IP addresses over a few hours. Each address hits the site two or three times. Per-IP rate limiting does nothing against that. One bright spot: LWN noted scraper attack levels dropped somewhat after the NetNut takedown.

Four reports in the current cycle put numbers on the pressure.

ReportPublishedWhat it measuresHeadline figure
Cloudflare Radar, cited by CEO Matthew PrinceJune 3, 2026Bot share of HTML HTTP requests57.5%
Imperva Bad Bot Report 2026April 29, 2026Automated share of all web traffic in 202553%, up from 51%
HUMAN Security, State of AI Traffic 20262026Growth of agentic AI traffic, year over yearRoughly 7,851%
Proxyway Proxy Market Research 2026June 2026Proxy pricing, pools and benchmarks13 providers tested

Do not average the first two. Cloudflare counts HTML requests. Imperva counts all web traffic, including app and API calls. They use different denominators. The direction is identical.

The HUMAN figure explains the panic better than either. Agentic traffic is not a crawler indexing a catalog once a week. Matthew Prince framed the asymmetry at SXSW in March. A human shopping for a camera visits 5 websites. The agent doing the same job visits 5,000.

Cloudflare adds one more figure that explains everything downstream. 52% of crawler requests on its network now serve model training. That was 22% in spring 2025. Publishers are not blocking crawlers on principle. They are blocking them because the bandwidth bill arrived.

Cloudflare picked a date, and it is September 15

On July 1, Cloudflare split AI traffic into three named behaviors. Their changelog defines each one. Search indexes content to answer questions later. Agent acts in real time for a person. Training collects content to train or fine-tune a model. Each gets its own switch.

Then the defaults flip. Cloudflare’s blog states the date plainly. From September 15, 2026, Training and Agent are blocked by default on pages that display ads. Search stays allowed. TechCrunch reported the scope: new customers, new sites from existing customers, and all existing free-tier customers.

The mixed-use rule is the sharp edge. A crawler spanning Search and Training gets judged by the most restrictive rule that applies. That catches Googlebot, Applebot and BingBot on ad-monetized pages wherever Training blocks are on.

Cloudflare classWhat it coversDefault on ad pages from Sept 15
SearchIndexing content to answer questions laterAllowed
AgentReal-time fetching on a person’s behalfBlocked
TrainingCollection for model training or fine-tuningBlocked
Mixed-useAny crawler spanning two or more of the aboveMost restrictive rule wins

There is a payment path too. Cloudflare opened a Monetization Gateway waitlist built on the x402 protocol. An agent receives a 402 response, pays in stablecoins, then retries with settlement proof. HTTP 402 finally does the job it was specified for.

One tension stays unresolved. Googlebot bundles search indexing with training collection in a single crawler. A publisher who blocks Training therefore risks losing search visibility too. Cloudflare’s taxonomy names that conflict. It cannot fix it. Only Google can, by splitting the crawler.

For a scraper, the shift is classification, not IP quality. Your bot class now decides access across a large slice of the web. Ad-supported publishers and review sites feel it first. Check one thing before September. Find out which class your user agent and request pattern land in. Do it on the specific sites you depend on.

Europe wrote the rulebook for scraping into a model

The European Data Protection Board adopted draft Guidelines 03/2026 on July 7. Public consultation runs to October 30, 2026. Companion guidelines on anonymisation landed alongside.

Three findings change pipeline design. Reed Smith’s summary sets them out.

  1. Consent is off the table at scale. Article 6(1)(a) does not work for bulk scraping. Legitimate interest becomes the main route. It carries a three-part test.
  2. Technical signals now carry legal weight. The EDPB treats robots.txt, ai.txt, CAPTCHAs and login walls as evidence. They indicate what data subjects reasonably expect. Those signals feed the balancing test directly.
  3. Buyers are in scope. Sidley’s analysis is clear on this. The guidelines cover firms acquiring pre-scraped datasets from third parties, not only those running crawlers.

A page being public settles nothing about your lawful basis for collecting it.

Transparency survives scale, too. Individual notice may be excused where it is disproportionate. Detailed public privacy notices are not excused. Neither are data subject rights or pre-collection opt-out mechanisms.

This lands hardest on teams assembling AI training data from mixed sources. Product data, prices and listing attributes mostly stay clear of personal data. Reviews, profiles, seller identities and user-generated text do not. The line runs through the fields, not the site.

Two courtrooms, one question about intermediaries

The Google case resolved, briefly. A federal court granted SerpApi’s motion to dismiss on July 21. Techdirt covered the reasoning. The court refused to stretch DMCA Section 1201 into control over access to public pages. Google has said it will use the 21-day window to file an amended complaint. Treat this as a pause. Our earlier work on SERP scraping tools covers the extraction side of that market.

The Reddit case matters more to anyone buying data rather than collecting it. Reddit, Inc. v. SerpApi, LLC, Oxylabs UAB, AWMProxy and Perplexity AI, Inc. sits before Judge Paul A. Engelmayer in the Southern District of New York. The docket number is 1:25-cv-08736. Reddit alleges the intermediaries harvested its content through Google search results, masked their identities, then sold it downstream.

Oral argument ran on July 23. Bloomberg Law reported that the judge questioned Reddit’s authority to sue on behalf of its users. The hearing ran nearly three hours. The claims are DMCA anti-circumvention plus state-law unjust enrichment.

Here is why it matters. Three of the four defendants are not AI labs. They are a search-results API reseller and two proxy operators. A ruling that reaches downstream buyers turns vendor diligence from a procurement task into a legal one.

Residential proxy prices stopped falling

Proxyway’s annual report gives the clearest read. Its data was collected in March and April 2026. Residential rates contracted by up to 75% between 2023 and 2025. In early 2026 that reversed.

The mechanism was discount removal, not list-price hikes. IPRoyal, Decodo, Oxylabs and Massive dropped long-running 40 to 50% coupon codes. Decodo and Oxylabs then revised permanent plans to roughly 25% below original list. IPRoyal simply reverted to pre-coupon rates.

Entry points moved harder. Oxylabs, Massive and SOAX removed or hid pay-as-you-go plans. That raised the floor by up to 25 times for anyone testing at small volume.

Provider50 GB rate, 2026 vs 20251,000 GB rate, 2026 vs 2025
IPRoyal200%199%
Oxylabs133%125%
Decodo122%133%
RayobyteNot reported78%
Webshare53%50%

Source: Proxyway Proxy Market Research 2026, indexed to a 100% baseline.

Median pricing across the 13 tested providers sat at $3.75/GB at 5 GB. At 50 GB it was $3.00/GB. At 500 GB it reached $2.25/GB. Budget resellers went the other way. ProxyEmpire and Webshare halved their rates. The grey market runs under $0.50/GB and touches $0.15. That is exactly where seizure risk concentrates. Our residential proxy provider roundup tracks the legitimate end of that range.

Mobile inverted completely. Rayobyte cut rates by up to 98%. Bright Data discontinued its mobile proxy product in April 2026. Mobile pools now sometimes cost less per GB than residential ones. ISP proxies held steady, since IP leasing and bandwidth set a hard floor.

Revenue explains the confidence behind the repricing. Bright Data reported 50% year-over-year growth in November 2025, at $300 million annualized. NetNut grew 28%. Its position was thinner than that suggests. Six customers generated half of its $40 million 2025 revenue. AI companies became its largest customer segment, ahead of eCommerce.

Capital arrived on top. On July 9, Oxylabs announced a $130 million investment from Warburg Pincus at a $3.6 billion valuation. It was the first outside investment since the company’s founding in 2015. Nimble raised $47 million in February. Meanwhile Proxyway counted 56 new proxy companies between 2025 and March 2026. Around 60% sold residential proxies. Most were whitelabels.

Two things follow for anyone renewing a contract this quarter. First, the coupon you budgeted against is probably gone, so re-price the renewal rather than rolling it. Second, with pay-as-you-go removed at several vendors, cheap evaluation is over. Benchmark on a paid trial before you commit to a monthly tier. Measure success rate per target class, not the aggregate number in the sales deck.

What this changes inside a production crawler

The lesson from July is provenance. When a pool degrades overnight, you need to know which targets depended on it. Debugging selectors first wastes a day.

So stamp the routing metadata onto every record. Not into a side log. Into the record itself.

# Fetch one URL through a named provider and return the payload plus routing metadata.
# The provider and exit_ip fields turn a vendor migration into a config change
# instead of a week of blind re-testing.

import time
import httpx

UA = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 Chrome/126.0 Safari/537.36"

PROVIDERS = {                                        # routing table, never hardcoded calls
    "provider_a": "http://user:pass@gw.provider-a.example:7000",
    "provider_b": "http://user:pass@gw.provider-b.example:8000",
}

def fetch_with_provenance(url: str, provider: str) -> dict:
    """Return the response plus the exact route it travelled."""
    started = time.perf_counter()
    try:
        with httpx.Client(proxy=PROVIDERS[provider], timeout=20) as client:
            r = client.get(url, headers={"User-Agent": UA})   # 1. issue the request
        status, size = r.status_code, len(r.content)
        exit_ip = r.headers.get("x-exit-ip")                  # 2. gateway echoes the exit node
        error = None
    except httpx.HTTPError as exc:                            # 3. record failures, never swallow them
        status, size, exit_ip, error = None, 0, None, type(exc).__name__

    return {
        "url": url,
        "provider": provider,                                 # 4. the field that saves you later
        "exit_ip": exit_ip,
        "status": status,
        "bytes": size,
        "latency_ms": round((time.perf_counter() - started) * 1000),
        "error": error,
        "collected_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
    }

Sample input:

fetch_with_provenance("https://example-retailer.test/p/SKU-88421", "provider_a")

Sample output:

{
  "url": "https://example-retailer.test/p/SKU-88421",
  "provider": "provider_a",
  "exit_ip": "203.0.113.47",
  "status": 200,
  "bytes": 148213,
  "latency_ms": 812,
  "error": null,
  "collected_at": "2026-07-29T06:14:22Z"
}

Roll that up nightly. Vendor decisions become arithmetic instead of argument. Illustrative sample figures below, not benchmark results:

ProviderTarget classSuccess rateMedian latencyCost per 1,000 pages
provider_aCloudflare-protected retail94.1%812 ms$2.40
provider_aOpen listing pages99.3%410 ms$0.85
provider_bCloudflare-protected retail88.6%1,340 ms$1.60
provider_bOpen listing pages99.1%505 ms$0.55

The routing rule writes itself from there. Send hard targets like Amazon or Google Shopping through the higher-success pool. Send open Zillow and Indeed listing pages through the cheaper one. Cloudflare, Akamai and Imperva still decide the outcome after the IP passes. A better pool does not fix a bad fingerprint.

What changes between a test crawl and a recurring feed

A one-off pull tolerates a single provider. A daily feed does not.

Production means two providers configured at all times. It means a per-target routing table you can edit without a deploy. It means alerting on success rate by provider, not just overall. It also means keeping each vendor’s sourcing documentation and KYC posture on file. July turned that from a checkbox into a live question.

DataFlirt builds recurring pipelines as the default case, not a script someone reruns by hand. Deduplication, normalization and schema consistency happen before delivery. Output lands as CSV, JSON, a live API, or straight into S3, MongoDB, DynamoDB, Firebase or Supabase. When a pool gets seized on a Thursday, the eCommerce pricing or search results feed still lands Friday morning.

A crawler that survives a vendor disappearing 1. Classify Group targets by the defense they run, not by industry 2. Assign IP class Datacenter, ISP, residential or mobile per target class 3. Verify vendor Sourcing docs, KYC posture, whether the pool is whitelabelled 4. Route and stamp Provider and exit IP written onto every record collected 5. Deliver CSV, JSON, live API, or straight into S3, MongoDB, Supabase Failover trigger: success rate by provider drops below threshold Flip the affected target class to the secondary provider in the routing table. No selector rewrite. No redeploy. The stamped provenance already tells you which targets were affected, and from exactly which timestamp. Step 3 is the one July added Before the NetNut seizure, vendor sourcing was a procurement question. It is now an availability question and a legal one.

Vendor verification moved from a one-time step to a standing control. A pool can now vanish between two scheduled runs.

Scraping public data stays lawful in most jurisdictions, with real limits around personal data and access controls. The full breakdown sits in is web crawling legal.

Next Steps

Choose DataFlirt if:

  • Anti-bot resilience against Cloudflare, Akamai or Imperva is a hard requirement.
  • You want structured CSV, JSON or API delivery, not raw access and a proxy bill.
  • The feed is recurring, and a vendor seizure cannot be allowed to break it.
  • There is no in-house data engineering team to run provider migrations at 2am.

Look elsewhere if:

  • You need in-house infrastructure ownership. DataFlirt runs on cloud infrastructure and third-party proxy vendors rather than owning the pools. That is true of nearly every provider in this space, so it is not a differentiator against alternatives. It matters only if owning the IP supply chain is a hard requirement for you.

Send us the target list and the fields you need. We will come back with the crawler design, the delivery format, and the refresh cadence: dataflirt.com/contact.

More to read

Latest from the Blog

Services

Data Extraction for Every Industry

View All Services →