SYSTEM all green source douglas.de queue 12,943 pages p99 latency 218ms dataflirt.com · scraper/douglas-de
RUN · 42 active pipelines · douglas.de live

Douglas cosmetics data,
at warehouse scale.

We extract skincare catalogues, fragrance pricing, shade variants, ingredient lists, and customer reviews from Douglas.de. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Products extracted
142K /day
Price updates
314K /24h
Review records
1.2M /run
Active pipelines
42
Uptime
99.94%
Data Dictionary

Every field we extract from douglas.de

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Product Catalogues objects from douglas.de. All fields typed and schema-versioned.

skubrandnamecategorysub_categorypricelist_pricecurrencyvolume_mlingredientsdescriptionimage_urlsis_newis_limited
product_catalogues
● 200 OK
"sku": "1029384",
"brand": "Dior",
"name": "Sauvage Eau de Parfum",
"price": 98.99,
"list_price": 110.0,
"volume_ml": "100",
"currency": "EUR",
"is_new": false
# skubrandnamecategorysub_categoryprice
1
2
3

Complete list of extractable fields for Pricing & Promos objects from douglas.de. All fields typed and schema-versioned.

skubase_pricecurrent_pricediscount_pctbeauty_card_priceprice_per_100mlpromo_badgeis_salestock_statusscraped_at
pricing_& promos
● 200 OK
"sku": "1029384",
"current_price": 98.99,
"discount_pct": 10,
"beauty_card_price": 89.09,
"price_per_100ml": 98.99,
"promo_badge": "Sale",
"stock_status": "in_stock",
"scraped_at": "2026-10-12T10:00:00Z"
# skubase_pricecurrent_pricediscount_pctbeauty_card_priceprice_per_100ml
1
2
3

Complete list of extractable fields for Variants objects from douglas.de. All fields typed and schema-versioned.

parent_skuvariant_skuvariant_typevariant_valuehex_codepricestock_statusimage_url
variants
● 200 OK
"parent_sku": "993821",
"variant_sku": "993821-01",
"variant_type": "shade",
"variant_value": "01 Fair",
"hex_code": "#FAD6C3",
"price": 45.5,
"stock_status": "in_stock"
# parent_skuvariant_skuvariant_typevariant_valuehex_codeprice
1
2
3

Complete list of extractable fields for Reviews & Ratings objects from douglas.de. All fields typed and schema-versioned.

review_idskureviewer_namestar_ratingtitletextdateverified_purchasehelpful_votesskin_typeage_group
reviews_& ratings
● 200 OK
"review_id": "rev_9283",
"sku": "1029384",
"star_rating": 5,
"title": "Great scent",
"text": "Lasts all day.",
"date": "2026-09-15",
"verified_purchase": true,
"skin_type": "combination"
# review_idskureviewer_namestar_ratingtitletext
1
2
3

Complete list of extractable fields for Search & Categories objects from douglas.de. All fields typed and schema-versioned.

keywordcategory_pathpositionskubrandnamepriceratingreview_countis_sponsoredscraped_at
search_& categories
● 200 OK
"keyword": "mascara",
"category_path": "Makeup > Eyes",
"position": 1,
"sku": "883721",
"brand": "MAC",
"price": 28.0,
"rating": 4.7,
"review_count": 342
# keywordcategory_pathpositionskubrandname
1
2
3

Capabilities

Everything you need from Douglas, nothing you do not

Our Douglas scraper handles the React frontend, dynamic variant loading, and EU bot protection to deliver clean cosmetics data.

Full Beauty Catalogue

Extract SKUs, brands, categories, descriptions, and metadata across skincare, makeup, and fragrance.

Variant Mapping

Capture shade names, hex codes, and volume sizes (ML/OZ) mapped to parent SKUs.

Dynamic Pricing

Track base prices, Douglas Beauty Card member prices, and standardised per-100ml metrics.

Ingredient Extraction

Parse full INCI ingredient lists for compliance tracking and formulation analysis.

Review & Rating Mining

Extract star ratings, review text, and reviewer metadata like skin type and age group.

Stock & Availability

Monitor in-stock, low-stock, and out-of-stock statuses across all variants.

Category Rank Tracking

Track bestseller positions and category placements for competitive benchmarking.

Promo & Gift Badges

Identify Gift With Purchase (GWP), Sale, and Limited Edition promotional tags.

Scheduled Diffs

Configure continuous pipelines that only push changed records to reduce downstream load.

DE Localisation

Scrape via German residential IPs to ensure accurate VAT pricing and bypass geo-blocks.

// engagement pipeline

From SKU list to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide brand URLs, category paths, or keyword sets. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy / Playwright crawlers, German proxy rotation, and session management for douglas.de.

Validation & QA
d 4–6

Schema validation, null-rate checks, and price-outlier detection before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

Bypassing Douglas.de bot protections

Douglas employs strict WAF rules and dynamic frontend rendering. We handle the EU networking and JavaScript execution.

pipeline-monitor · douglas.de · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
EU Residential Proxies
German IPs for accurate pricing

Douglas geo-blocks aggressive traffic and alters pricing based on origin IP. We route requests through German residential proxies to ensure accurate EUR pricing and avoid WAF bans.

SPA Hydration
Playwright for React rendering

The Douglas frontend relies heavily on client-side React rendering. We execute full Playwright sessions to hydrate the DOM, ensuring dynamic pricing and variant data is fully loaded.

Variant Expansion
Iterating through shades and sizes

Cosmetics listings hide variant data behind JavaScript click events. Our crawlers simulate user interactions to expose every colour shade and bottle size attached to a parent SKU.

Rate Limit Evasion
Concurrency control and humanised delays

We manage request throughput with strict concurrency limits and randomised human-like delays, preventing Akamai from flagging our IP pools during deep catalogue crawls.

Schema Resilience
Fallback selectors for DOM updates

Douglas updates its frontend frequently during promotional seasons. We use multi-layered fallback selectors to maintain pipeline stability when HTML structures change.

Applications

Who uses Douglas data and how

Teams across industries use douglas.de data to build competitive products and smarter operations.

01
Competitor Price Monitoring

Beauty retailers track Douglas pricing, discounts, and Beauty Card offers to optimise their own pricing strategies.

02
Assortment & Gap Analysis

Brands analyse category depth and competitor brand presence to identify missing product lines or whitespace.

03
Ingredient Trend Analysis

Formulators extract INCI lists at scale to track trending active ingredients across top-selling skincare products.

04
Brand MAP Compliance

Premium beauty brands monitor Douglas to ensure their products are not discounted below Minimum Advertised Price agreements.

05
Review Sentiment Analysis

Marketing teams aggregate reviews to understand customer sentiment regarding specific formulations, scents, or packaging.

06
Promotional Tracking

Analysts track Gift With Purchase offers and seasonal sale events to map Douglas promotional calendars.

Why DataFlirt

"Douglas commands the European premium beauty market. Extracting their catalogue reveals exact brand positioning, variant pricing, and ingredient trends."

Scraping Douglas requires navigating aggressive bot protection, geo-fenced pricing, and complex React-based variant loading. DataFlirt manages the residential proxies and JavaScript hydration, delivering clean cosmetics data straight to your warehouse.

Technical Spec

Douglas scraper technical capabilities

Everything supported by our douglas.de scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions required for React hydration and dynamic pricing
Supported
DE Residential proxies
German ISP-grade IPs to bypass geo-blocks and capture accurate VAT
Supported
Shade and size mapping
Parent to child SKU mapping for all colour and volume variants
Supported
INCI ingredient lists
Extraction of full cosmetic ingredient declarations
Supported
Beauty Card pricing
Capture of member-specific discounted prices
Supported
Review pagination
Iteration through all customer reviews and ratings
Supported
Change detection
Hash-based diffs to only emit updated records
Supported
Webhook delivery
HTTP POST per record for real-time processing
Supported
User purchase history
Gated data requires individual account credentials
Partial
Beauty Card points balance
Private loyalty account data is not extracted
Partial
Infrastructure

Infrastructure powering the Douglas pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy orchestrates the crawl while Playwright handles JavaScript execution and DOM hydration for React-based product pages.

EU Proxy Infrastructure

We route requests through German residential proxy pools to ensure accurate EUR pricing and evade Akamai bot detection.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, dependency management, and SLA alerting.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested structures
CSV
Flat file with typed columns
XLS
Excel compatible export for business teams
Parquet
Columnar format for analytics engines
AWS S3
Direct bucket delivery
Webhook
HTTP POST per record
API
REST endpoints for data retrieval
BigQuery
Streamed directly into your dataset
Snowflake
Stage and COPY INTO workflow
Postgres
Upsert into your existing schema
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About douglas.de scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Douglas legal?

Scraping publicly available product, pricing, and review data from Douglas is generally permissible. DataFlirt targets only public data and does not extract personal user information or circumvent authentication walls.

How do you handle bot protection on douglas.de?

We use German residential proxies, Playwright browser sessions with realistic fingerprints, and strict concurrency limits to avoid triggering WAF blocks.

Can you extract all shade variations for makeup?

Yes. Our crawlers interact with the frontend to expose and extract every shade variant, hex code, and stock status associated with a parent SKU.

Do you capture Douglas Beauty Card member prices?

Yes. We extract both the standard base price and the discounted Beauty Card price displayed on product pages.

How fresh is the pricing data?

We configure pipelines to match your requirements, ranging from daily catalogue refreshes to high-frequency checks on specific SKUs.

Can you extract full INCI ingredient lists?

Yes. We parse the ingredient tabs on product pages to extract complete INCI declarations for compliance and formulation analysis.

Do you support other EU Douglas domains?

Yes. We can target douglas.at, douglas.ch, douglas.it, and other regional domains using appropriate localised proxies.

$ dataflirt scope --new-project --source=douglas.de ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off catalogue dump or continuous price monitoring across 100K beauty SKUs, we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in beauty and skincare

Services

Data Extraction for Every Industry

View All Services →