SYSTEM all green source carrefour.fr queue 12,491 pages p99 latency 185ms dataflirt.com · scraper/carrefour-fr
RUN - 51 active pipelines - carrefour.fr live

Carrefour grocery data,
at warehouse scale.

We extract FMCG catalogues, Nutri-Score ratings, dynamic pricing, and promotional mechanics from Carrefour.fr. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Products extracted
182K /day
Price updates
415K /24h
Promotions tracked
28K /run
Active pipelines
51
Uptime
99.94%
Data Dictionary

Every field we extract from carrefour.fr

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Product Listings objects from carrefour.fr. All fields typed and schema-versioned.

eantitlebrandcategoryweightprice_per_kgnutri_scoreeco_scoreingredientsallergensimage_urlavailability
product_listings
● 200 OK
"ean": "3017620422003",
"title": "Pâte à tartiner aux noisettes NUTELLA",
"brand": "Nutella",
"weight": "1kg",
"price_per_kg": 6.49,
"nutri_score": "E",
"availability": true
# eantitlebrandcategoryweightprice_per_kg
1
2
3

Complete list of extractable fields for Pricing & Promotions objects from carrefour.fr. All fields typed and schema-versioned.

eanbase_pricepromo_pricediscount_pctpromo_typeloyalty_bonusprice_per_unitvalid_untilstore_id
pricing_& promotions
● 200 OK
"ean": "3017620422003",
"base_price": 6.49,
"promo_price": 4.99,
"discount_pct": 23,
"promo_type": "remise_immediate",
"loyalty_bonus": 0.0,
"store_id": "1234"
# eanbase_pricepromo_pricediscount_pctpromo_typeloyalty_bonus
1
2
3

Complete list of extractable fields for Nutritional Data objects from carrefour.fr. All fields typed and schema-versioned.

eanenergy_kcalfat_gsaturated_fat_gcarbs_gsugars_gprotein_gsalt_gfiber_gnutri_score
nutritional_data
● 200 OK
"ean": "3017620422003",
"energy_kcal": 539,
"fat_g": 30.9,
"saturated_fat_g": 10.6,
"sugars_g": 56.3,
"protein_g": 6.3,
"nutri_score": "E"
# eanenergy_kcalfat_gsaturated_fat_gcarbs_gsugars_g
1
2
3

Complete list of extractable fields for Store Availability objects from carrefour.fr. All fields typed and schema-versioned.

store_idstore_nameformataddresspostal_codedrive_availabledelivery_availablenext_slotopening_hours
store_availability
● 200 OK
"store_id": "75013_ITALIE",
"store_name": "Carrefour Market Paris Italie",
"format": "Market",
"postal_code": "75013",
"drive_available": true,
"next_slot": "2026-05-12T14:00:00Z",
"delivery_available": true
# store_idstore_nameformataddresspostal_codedrive_available
1
2
3

Complete list of extractable fields for Category Taxonomy objects from carrefour.fr. All fields typed and schema-versioned.

category_idparent_categorylevel_1level_2level_3urlproduct_countscraped_at
category_taxonomy
● 200 OK
"category_id": "epicerie_sucree",
"level_1": "Epicerie Sucrée",
"level_2": "Petit déjeuner",
"level_3": "Pâtes à tartiner",
"product_count": 142,
"scraped_at": "2026-05-12T09:14:33Z"
# category_idparent_categorylevel_1level_2level_3url
1
2
3

Capabilities

Extract grocery data with store-level precision

Our Carrefour scraper handles the complexity of grocery retail: store-specific pricing, Datadome protection, promotional mechanics, and complex nutritional schemas.

Full FMCG Extraction

Title, brand, weight, ingredients, allergens, EAN, and imagery scraped at the SKU level.

Nutri-Score & Eco-Score

Capture official Nutri-Score, Eco-Score, and full macronutrient tables per 100g and per portion.

Geo-Specific Pricing

Extract localized pricing by injecting store IDs or postal codes. Track regional price variations across France.

Promotion & Loyalty Tracking

Monitor immediate discounts, bulk offers, and Carrefour loyalty card mechanics.

Multi-Format Support

Data unified across Carrefour Hyper, Market, City, and Contact store formats.

Drive Slot Monitoring

Track Carrefour Drive availability, pickup slots, and home delivery windows.

EAN Matching

Map Carrefour listings to global EANs for exact cross-retailer product matching.

Price Per Unit

Normalised price per kg or litre to enable accurate competitor benchmarking.

Scheduled Diffs

Run one-off bulk exports or configure continuous pipelines with change-detection diffing.

// engagement pipeline

From category URL to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide category URLs, EAN lists, or target store IDs. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy / Playwright crawlers, proxy rotation, session management, and Datadome bypass.

Validation & QA
d 4–6

Schema validation, null-rate checks, and price-outlier detection before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our Carrefour pipeline handles the hard parts

French retail relies heavily on bot protection. Here is how we stay resilient.

pipeline-monitor · carrefour.fr · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Anti-bot layer
Datadome bypass and residential proxies

Carrefour uses Datadome for bot mitigation. We route requests through French residential IPs and spoof TLS fingerprints to maintain high trust scores and avoid CAPTCHA blocks.

Geo-location contexts
Store-specific session hydration

Pricing on Carrefour is not national. We inject specific store IDs into the session state, ensuring the prices extracted match the exact physical or Drive location you need.

SPA rendering
Handling React hydration

Carrefour.fr relies on client-side rendering. We execute Playwright sessions to ensure dynamic price widgets, stock statuses, and promotional banners fully load before extraction.

Schema stability
Resilient selectors

Retail DOMs change often. Our selector strategy uses fallback chains and structured data extraction (LD+JSON) so a layout update does not break your data feed.

Change detection
Only re-scrape what changes

For daily price monitoring, we maintain a hash index of last-seen values. Subsequent runs only push diffs, reducing compute cost and downstream processing load.

Applications

Who uses Carrefour data

Teams across industries use carrefour.fr data to build competitive products and smarter operations.

01
FMCG Price Monitoring

Brands track retail pricing against recommended retail prices (RRP) and monitor competitor shelf prices.

02
Private Label Benchmarking

Retailers compare their private label pricing and nutritional profiles against Carrefour brand products.

03
Promotional Compliance

Trade marketing teams verify that negotiated promotions, discounts, and banners are executed correctly online.

04
Assortment Intelligence

Category managers analyse Carrefour range gaps, new product listings, and delisted items.

05
Inflation Tracking

Analysts track basket inflation across basic commodities and FMCG categories over time.

06
Nutritional Analysis

Health tech applications ingest Nutri-Score, Eco-Score, and ingredient data to power consumer apps.

Why DataFlirt

"Carrefour.fr holds the definitive digital catalogue of French grocery retail - but extracting store-level pricing requires bypassing aggressive Datadome protection."

Most teams underestimate the investment required: reliable Carrefour scraping requires French residential proxies, full JavaScript rendering for store-context hydration, Datadome CAPTCHA handling, and anomaly monitoring. DataFlirt absorbs that complexity so your engineers can focus on the analysis, not the infrastructure.

Technical Spec

Carrefour scraper - technical capabilities

Everything supported by our carrefour.fr scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions required for React hydration and price widgets
Supported
Datadome bypass
Automated solver integration with residential proxy rotation
Supported
Residential proxy rotation
ISP-grade residential IPs from FR pools
Supported
Store-level pricing
Session injection for specific Carrefour store IDs
Supported
EAN extraction
Capture global trade item numbers for cross-retailer matching
Supported
Change detection (diffs)
Hash-based diff: only emit records with changed fields since last run
Supported
Loyalty card exclusive data
Personalised discounts requiring specific user loyalty accounts
Partial
User purchase history
Requires authenticated user sessions and violates privacy policies
Partial
Infrastructure

Infrastructure powering the Carrefour pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering, store session hydration, and interaction flows.

Residential Proxy Infrastructure

We maintain pools of French residential ISP proxies. Rotation happens per-request with sticky sessions where required to maintain store context.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling and dependency management. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested
CSV
Flat file with typed columns
XLS
Excel compatible export
Parquet
Columnar format for BigQuery, Snowflake
AWS S3
Direct bucket delivery
Webhook
HTTP POST per record
API
REST endpoint access
BigQuery
Streamed directly into your dataset
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About carrefour.fr scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Carrefour legal?

Scraping publicly available pricing and product information is generally permissible. DataFlirt targets only public, non-authenticated grocery data. We do not extract personal data or circumvent authentication walls. Clients should review Carrefour's ToS and consult legal counsel.

How do you handle Datadome protection?

We use French residential ISP proxies, full Playwright browser sessions with realistic TLS fingerprints, and request timing modelled on human behaviour to avoid Datadome triggers.

Can you extract prices for specific local stores?

Yes. We hydrate the browser session with the specific store ID or postal code before extraction, ensuring you get the exact local price and availability, not a national average.

How fresh is the data?

Daily price monitoring pipelines complete within a 4-8 hour window depending on the catalogue size and store count required.

What is the minimum viable engagement?

Our smallest packages start at a defined category or EAN list with weekly delivery. For full catalogue tracking across multiple stores, we price based on volume and frequency.

Can I request a sample dataset?

Yes. We provide a sample run of up to 500 products as part of the pre-engagement scoping process so you can validate schema fit and data quality.

$ dataflirt scope --new-project --source=carrefour.fr ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off catalogue dump or a continuous price-monitoring feed across multiple French stores, we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in food drink and kitchen

Services

Data Extraction for Every Industry

View All Services →