We extract FMCG catalogues, Nutri-Score ratings, dynamic pricing, and promotional mechanics from Carrefour.fr. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Product Listings objects from carrefour.fr. All fields typed and schema-versioned.
"ean": "3017620422003", "title": "Pâte à tartiner aux noisettes NUTELLA", "brand": "Nutella", "weight": "1kg", "price_per_kg": 6.49, "nutri_score": "E", "availability": true
| # | ean | title | brand | category | weight | price_per_kg |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Pricing & Promotions objects from carrefour.fr. All fields typed and schema-versioned.
"ean": "3017620422003", "base_price": 6.49, "promo_price": 4.99, "discount_pct": 23, "promo_type": "remise_immediate", "loyalty_bonus": 0.0, "store_id": "1234"
| # | ean | base_price | promo_price | discount_pct | promo_type | loyalty_bonus |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Nutritional Data objects from carrefour.fr. All fields typed and schema-versioned.
"ean": "3017620422003", "energy_kcal": 539, "fat_g": 30.9, "saturated_fat_g": 10.6, "sugars_g": 56.3, "protein_g": 6.3, "nutri_score": "E"
| # | ean | energy_kcal | fat_g | saturated_fat_g | carbs_g | sugars_g |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Store Availability objects from carrefour.fr. All fields typed and schema-versioned.
"store_id": "75013_ITALIE", "store_name": "Carrefour Market Paris Italie", "format": "Market", "postal_code": "75013", "drive_available": true, "next_slot": "2026-05-12T14:00:00Z", "delivery_available": true
| # | store_id | store_name | format | address | postal_code | drive_available |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Category Taxonomy objects from carrefour.fr. All fields typed and schema-versioned.
"category_id": "epicerie_sucree", "level_1": "Epicerie Sucrée", "level_2": "Petit déjeuner", "level_3": "Pâtes à tartiner", "product_count": 142, "scraped_at": "2026-05-12T09:14:33Z"
| # | category_id | parent_category | level_1 | level_2 | level_3 | url |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our Carrefour scraper handles the complexity of grocery retail: store-specific pricing, Datadome protection, promotional mechanics, and complex nutritional schemas.
Title, brand, weight, ingredients, allergens, EAN, and imagery scraped at the SKU level.
Capture official Nutri-Score, Eco-Score, and full macronutrient tables per 100g and per portion.
Extract localized pricing by injecting store IDs or postal codes. Track regional price variations across France.
Monitor immediate discounts, bulk offers, and Carrefour loyalty card mechanics.
Data unified across Carrefour Hyper, Market, City, and Contact store formats.
Track Carrefour Drive availability, pickup slots, and home delivery windows.
Map Carrefour listings to global EANs for exact cross-retailer product matching.
Normalised price per kg or litre to enable accurate competitor benchmarking.
Run one-off bulk exports or configure continuous pipelines with change-detection diffing.
Brief in. Clean data out.
Provide category URLs, EAN lists, or target store IDs. We design the extraction schema together.
We configure Scrapy / Playwright crawlers, proxy rotation, session management, and Datadome bypass.
Schema validation, null-rate checks, and price-outlier detection before full launch.
JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
French retail relies heavily on bot protection. Here is how we stay resilient.
Carrefour uses Datadome for bot mitigation. We route requests through French residential IPs and spoof TLS fingerprints to maintain high trust scores and avoid CAPTCHA blocks.
Pricing on Carrefour is not national. We inject specific store IDs into the session state, ensuring the prices extracted match the exact physical or Drive location you need.
Carrefour.fr relies on client-side rendering. We execute Playwright sessions to ensure dynamic price widgets, stock statuses, and promotional banners fully load before extraction.
Retail DOMs change often. Our selector strategy uses fallback chains and structured data extraction (LD+JSON) so a layout update does not break your data feed.
For daily price monitoring, we maintain a hash index of last-seen values. Subsequent runs only push diffs, reducing compute cost and downstream processing load.
Brands track retail pricing against recommended retail prices (RRP) and monitor competitor shelf prices.
Retailers compare their private label pricing and nutritional profiles against Carrefour brand products.
Trade marketing teams verify that negotiated promotions, discounts, and banners are executed correctly online.
Category managers analyse Carrefour range gaps, new product listings, and delisted items.
Analysts track basket inflation across basic commodities and FMCG categories over time.
Health tech applications ingest Nutri-Score, Eco-Score, and ingredient data to power consumer apps.
"Carrefour.fr holds the definitive digital catalogue of French grocery retail - but extracting store-level pricing requires bypassing aggressive Datadome protection."
Most teams underestimate the investment required: reliable Carrefour scraping requires French residential proxies, full JavaScript rendering for store-context hydration, Datadome CAPTCHA handling, and anomaly monitoring. DataFlirt absorbs that complexity so your engineers can focus on the analysis, not the infrastructure.
Everything supported by our carrefour.fr scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering, store session hydration, and interaction flows.
We maintain pools of French residential ISP proxies. Rotation happens per-request with sticky sessions where required to maintain store context.
Pipelines run on AWS Lambda and ECS. Airflow handles scheduling and dependency management. All state stored in managed Postgres.
Data delivered to where your team already works — no new tooling required.
About carrefour.fr scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available pricing and product information is generally permissible. DataFlirt targets only public, non-authenticated grocery data. We do not extract personal data or circumvent authentication walls. Clients should review Carrefour's ToS and consult legal counsel.
We use French residential ISP proxies, full Playwright browser sessions with realistic TLS fingerprints, and request timing modelled on human behaviour to avoid Datadome triggers.
Yes. We hydrate the browser session with the specific store ID or postal code before extraction, ensuring you get the exact local price and availability, not a national average.
Daily price monitoring pipelines complete within a 4-8 hour window depending on the catalogue size and store count required.
Our smallest packages start at a defined category or EAN list with weekly delivery. For full catalogue tracking across multiple stores, we price based on volume and frequency.
Yes. We provide a sample run of up to 500 products as part of the pre-engagement scoping process so you can validate schema fit and data quality.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off catalogue dump or a continuous price-monitoring feed across multiple French stores, we scope, build, and operate the pipeline. Tell us what you need.