SYSTEM all green source selfridges.com queue 12,943 pages p99 latency 218ms dataflirt.com · scraper/selfridges-com
RUN * 42 active pipelines * selfridges.com live

Selfridges data,
at warehouse scale.

We extract designer catalogues, pricing signals, stock depth, and brand intelligence from Selfridges. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Products extracted
142K /day
Price updates
38K /24h
Brand catalogues
850 /run
Active pipelines
42
Uptime
99.98%
Data Dictionary

Every field we extract from selfridges.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Product Listings objects from selfridges.com. All fields typed and schema-versioned.

product_idtitlebrandcategory_pathpricecurrencydescriptionfabric_caresize_fitproject_earthimage_urlsurl
product_listings
● 200 OK
"product_id": "R03942188",
"title": "Le Chiquito leather top-handle bag",
"brand": "JACQUEMUS",
"price": 610.0,
"currency": "GBP",
"category_path": "Womens > Bags > Top handle bags",
"project_earth": false
# product_idtitlebrandcategory_pathpricecurrency
1
2
3

Complete list of extractable fields for Pricing & Stock objects from selfridges.com. All fields typed and schema-versioned.

product_idvariant_idsizecolourpriceoriginal_pricediscount_pctin_stocklow_stock_warningcurrencyscraped_at
pricing_& stock
● 200 OK
"product_id": "R03942188",
"variant_id": "V123456",
"colour": "Black",
"price": 610.0,
"in_stock": true,
"low_stock_warning": false,
"scraped_at": "2026-05-12T10:14:00Z"
# product_idvariant_idsizecolourpriceoriginal_price
1
2
3

Complete list of extractable fields for Brand & Designer objects from selfridges.com. All fields typed and schema-versioned.

brand_idbrand_nameboutique_urlproduct_countcategories_covereddescriptionhero_image_urlis_exclusive
brand_& designer
● 200 OK
"brand_name": "JACQUEMUS",
"boutique_url": "https://www.selfridges.com/GB/en/cat/jacquemus/",
"product_count": 214,
"is_exclusive": false,
"categories_covered": "['Bags', 'Clothing', 'Shoes']",
"brand_id": "B984"
# brand_idbrand_nameboutique_urlproduct_countcategories_covereddescription
1
2
3

Complete list of extractable fields for Sustainability Data objects from selfridges.com. All fields typed and schema-versioned.

product_idproject_earth_flagsustainability_criteriacertificationsmaterialsvegan_flagcruelty_freepackaging_details
sustainability_data
● 200 OK
"product_id": "R03942188",
"project_earth_flag": true,
"sustainability_criteria": "['For Nature', 'Better Materials']",
"materials": "['100% vegetable tanned leather']",
"vegan_flag": false,
"certifications": "['LWG Gold']"
# product_idproject_earth_flagsustainability_criteriacertificationsmaterialsvegan_flag
1
2
3

Complete list of extractable fields for Categories & Navigation objects from selfridges.com. All fields typed and schema-versioned.

category_idcategory_nameparent_categorydepartmenturlproduct_countfeatured_brandssort_options
categories_& navigation
● 200 OK
"category_name": "Top handle bags",
"parent_category": "Bags",
"department": "Womens",
"url": "/GB/en/cat/womens/bags/top-handle-bags/",
"product_count": 842,
"featured_brands": "['Prada', 'Gucci', 'Jacquemus']"
# category_idcategory_nameparent_categorydepartmenturlproduct_count
1
2
3

Capabilities

Extract luxury retail data with precision

Our Selfridges scraper handles complex product hierarchies, dynamic sizing grids, multi-region pricing, and sustainability metadata. Built for scale, delivered cleanly.

Designer Catalogue Extraction

Extract title, brand, description, fabric details, size and fit notes, and high-resolution imagery for every product across the site.

Multi-Region Pricing

Capture pricing in GBP, USD, EUR, and other supported currencies by simulating regional browsing sessions.

Stock & Variant Mapping

Map parent products to child variants. Track availability per size and colour, including low stock warnings.

Project Earth Metadata

Extract sustainability tags, material certifications, and Project Earth criteria attached to conscious products.

Brand Boutique Tracking

Monitor designer landing pages, exclusive drops, and assortment changes within specific brand boutiques.

Beauty & Grooming Specs

Capture volume, ingredients, cruelty-free flags, and specific beauty category metadata.

Search & Category Scraping

Extract product rankings, filter options, and facet counts across departments and search queries.

High-Frequency Updates

Monitor fast-moving luxury items, limited drops, and seasonal sales with sub-hourly crawl cadences.

Differential Sync

Maintain a hash index of last-seen values. We only push records when price, stock, or metadata changes.

// engagement pipeline

From luxury catalogue to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide categories, designer names, or search queries. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy / Playwright crawlers, proxy rotation, and session management for selfridges.com.

Validation & QA
d 4–6

Schema validation, null-rate checks, price-outlier detection, and variant mapping review before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our Selfridges pipeline handles the hard parts

Luxury retail sites employ strict rate limiting and dynamic frontends. Here is how we maintain data integrity.

pipeline-monitor · selfridges.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Anti-bot layer
Residential proxy rotation + fingerprint spoofing

Retailers block datacentre IPs and monitor request velocity. Our crawlers use residential ISP proxies with realistic browser fingerprints, randomised request timing, and full cookie session management to blend into normal traffic.

JavaScript rendering
Full Playwright execution for dynamic variants

Selfridges loads size grids, stock availability, and multi-angle images dynamically. We run full Playwright browser sessions with JavaScript execution to trigger these elements, capturing data that headless HTTP clients miss entirely.

Schema stability
Resilient selectors with fallback chains

DOM structures change during seasonal sales and site updates. Our selector strategy uses multiple fallback chains per field, including structured data extraction (LD+JSON), so a layout change does not break your data pipeline.

Change detection
Only re-scrape what has changed

For large catalogues, we maintain a hash index of last-seen values per field. Subsequent runs only push diffs, reducing compute cost, storage bloat, and downstream processing load. You get a clean changelog rather than full re-dumps.

Monitoring & alerting
24/7 pipeline health with anomaly detection

Every run emits structured logs to our observability stack. We alert on null-rate spikes, price outliers, schema drift, and coverage drops. We respond before you notice.

Applications

Who uses Selfridges data

Teams across industries use selfridges.com data to build competitive products and smarter operations.

01
Competitor Price Monitoring

Luxury retailers track pricing, markdown strategies, and promotional overlap across identical designer SKUs.

02
Assortment Planning

Merchandising teams analyse brand representation, category depth, and sizing availability to optimise their own buying strategies.

03
ESG & Sustainability Tracking

Analysts monitor the adoption of Project Earth criteria and sustainable materials across major fashion houses.

04
Visual AI Training

Machine learning teams use high-resolution product imagery and structured metadata to train computer vision models for fashion.

05
Brand MAP Compliance

Luxury brands monitor third-party retail channels to ensure Minimum Advertised Price compliance and correct brand positioning.

06
Trend Forecasting

Fashion analysts track out-of-stock velocity and new product introductions to predict seasonal colour and style trends.

Why DataFlirt

"Selfridges holds the blueprint for global luxury retail, but extracting that multi-region, variant-heavy catalogue requires precision engineering."

Most teams underestimate the investment required: reliable Selfridges scraping requires residential proxies, full JavaScript rendering for dynamic sizing grids, daily selector maintenance, and anomaly monitoring. DataFlirt absorbs that complexity so your engineers can focus on the analysis, not the infrastructure.

Technical Spec

Selfridges scraper - technical capabilities

Everything supported by our selfridges.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions required for dynamic sizing grids and stock flags
Supported
Residential proxy rotation
ISP-grade residential IPs from UK / US / EU pools
Supported
Multi-currency capture
Simulated regional sessions to extract GBP, USD, EUR pricing
Supported
Project Earth extraction
Capture sustainability tags, materials, and certifications
Supported
Variant mapping
Parent to child product relationships with all size/colour combinations
Supported
Change detection (diffs)
Hash-based diff: only emit records with changed fields since last run
Supported
Selfridges+ subscription data
Gated delivery tier pricing and member-exclusive offers
Partial
User wishlists
Private user account data requiring authentication
Partial
Infrastructure

Infrastructure powering the pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration, deduplication, and retry logic. Playwright handles JavaScript rendering, cookie sessions, and interaction flows.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies across UK and US regions. Rotation happens per-request with sticky sessions where required.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested array formatting
CSV
Flat file with typed columns for analytics
XLS
Excel compatible format for merchandising teams
Parquet
Columnar format for BigQuery, Snowflake, Athena
AWS S3
Direct bucket delivery compatible with any data lake
Webhook
HTTP POST per record for real-time downstream processing
API
Queryable REST endpoints for specific product lookups
BigQuery
Streamed directly into your dataset with schema auto-detect
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About selfridges.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Selfridges legal?

Scraping publicly available information from retail sites is generally permissible under applicable law in the UK and US. DataFlirt targets only public, non-authenticated product, pricing, and stock data. We do not extract personal data, circumvent authentication walls, or violate GDPR. Clients should consult legal counsel for specific use cases.

How do you handle rate limiting?

We use residential ISP proxies, full Playwright browser sessions with realistic fingerprints, and request timing modelled on human behaviour. We monitor for 403/503 rate spikes in real time and trigger pool rotation automatically.

Can you scrape specific regional pricing?

Yes. We configure pipelines to route through specific regional exit nodes and set appropriate location cookies to capture accurate local pricing and currency conversions.

How fresh is the stock data?

Real-time streaming pipelines achieve sub-60-minute latency for price and availability signals on a defined product set. Full catalogue refreshes at daily cadence complete within a 6-hour window.

What is the minimum viable engagement?

Our smallest packages start at a defined brand list or category subset with weekly delivery. For full-site catalogues or custom schema requirements, we price based on volume and delivery frequency.

Can I request a sample dataset before committing?

Absolutely. We provide a sample run of up to 500 products as part of the pre-engagement scoping process so you can validate schema fit, field completeness, and data quality before signing any contract.

$ dataflirt scope --new-project --source=selfridges.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off brand catalogue dump or a continuous price-monitoring feed across the entire site, we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in fashion and apparel

Services

Data Extraction for Every Industry

View All Services →