SYSTEM all green source hawesko.de queue 14,291 URLs p99 latency 218ms dataflirt.com · scraper/hawesko-de
RUN, 14 active pipelines, hawesko.de live

Hawesko wine data,
at warehouse scale.

We extract wine catalogues, pricing signals, vintage details, tasting notes, and expert ratings from Hawesko. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Wines extracted
12.4K /run
Price updates
4.1K /24h
Review records
98.4K /run
Active pipelines
14
Uptime
99.98%
Data Dictionary

Every field we extract from hawesko.de

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Wine Listings objects from hawesko.de. All fields typed and schema-versioned.

skunamevintagewineryregioncountrygrape_varietiespricebottle_sizeprice_per_litreurl
wine_listings
● 200 OK
"sku": "1234567",
"name": "Chateau Ripeau Grand Cru Classe",
"vintage": "2018",
"winery": "Chateau Ripeau",
"region": "Bordeaux",
"country": "France",
"price": 45.9,
"price_per_litre": 61.2
# skunamevintagewineryregioncountry
1
2
3

Complete list of extractable fields for Expert Ratings objects from hawesko.de. All fields typed and schema-versioned.

skurater_namescoremax_scoretasting_notedrinking_window_startdrinking_window_endrating_date
expert_ratings
● 200 OK
"sku": "1234567",
"rater_name": "Robert Parker",
"score": 94,
"max_score": 100,
"drinking_window_start": 2024,
"drinking_window_end": 2038,
"rating_date": "2021-04-15"
# skurater_namescoremax_scoretasting_notedrinking_window_start
1
2
3

Complete list of extractable fields for Pricing & Offers objects from hawesko.de. All fields typed and schema-versioned.

skucurrent_priceoriginal_pricediscount_pctbulk_discount_availablepackage_deal_idstock_statusdelivery_time_dayscurrency
pricing_& offers
● 200 OK
"sku": "1234567",
"current_price": 45.9,
"original_price": 55.0,
"discount_pct": 16.5,
"stock_status": "in_stock",
"delivery_time_days": "2-3",
"currency": "EUR"
# skucurrent_priceoriginal_pricediscount_pctbulk_discount_availablepackage_deal_id
1
2
3

Complete list of extractable fields for Product Details objects from hawesko.de. All fields typed and schema-versioned.

skualcohol_pctsweetnessacidityallergensdrinking_temperatureclosure_typeorganic_certifiedfood_pairings
product_details
● 200 OK
"sku": "1234567",
"alcohol_pct": 14.5,
"sweetness": "dry",
"acidity": 5.2,
"allergens": "contains sulfites",
"drinking_temperature": "16-18",
"closure_type": "natural cork"
# skualcohol_pctsweetnessacidityallergensdrinking_temperature
1
2
3

Complete list of extractable fields for Customer Reviews objects from hawesko.de. All fields typed and schema-versioned.

review_idskuratingauthorreview_datetitlebodyhelpful_votes
customer_reviews
● 200 OK
"review_id": "REV-98213",
"sku": "1234567",
"rating": 5,
"author": "WeinLiebhaber88",
"review_date": "2023-11-12",
"title": "Excellent Bordeaux",
"body": "Deep ruby colour, complex nose of dark berries and cedar."
# review_idskuratingauthorreview_datetitle
1
2
3

Capabilities

Everything you need from Hawesko, structured

Our Hawesko scraper handles the entire catalogue, capturing vintage rollovers, dynamic pricing, expert ratings, and tasting notes with full anti-bot circumvention built in.

Full Catalogue Extraction

Extract SKUs, names, vintages, and bottle sizes across all categories including red, white, rose, and sparkling wines.

Pricing & Discount Tracking

Capture base price, per-litre price, promotional discounts, and bulk purchase pricing across the entire assortment.

Expert Rating Aggregation

Parse scores and tasting notes from Falstaff, Luca Maroni, Robert Parker, and James Suckling directly from product pages.

Grape & Region Mapping

Extract structured data for terroir, appellation, country of origin, and specific grape variety blend percentages.

Tasting Notes & Pairings

Capture sommelier descriptions, flavour profiles, sweetness levels, acidity, and recommended food pairings.

Stock & Delivery Monitoring

Monitor stock availability flags, low-stock warnings, and estimated shipping times for inventory planning.

Package Deal Extraction

Map multi-bottle bundles and tasting packages back to their individual component SKUs for accurate value comparison.

Review Mining

Extract customer star ratings, review text, author names, and publication dates across all products.

Scheduled Change Detection

Track vintage rollovers and price changes over time with daily or weekly differential exports.

// engagement pipeline

From URL list to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide categories, search terms, or specific winery URLs. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy crawlers, German proxy rotation, session management, and CAPTCHA handling for hawesko.de.

Validation & QA
d 4–6

Schema validation, null-rate checks, price-outlier detection, and sample data review before full launch.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our Hawesko pipeline handles the hard parts

European eCommerce sites employ strict rate limiting and dynamic rendering. Here is how we maintain reliable extraction.

pipeline-monitor · hawesko.de · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Anti-bot layer
German residential proxy rotation

Hawesko employs geo-blocking and strict rate limits. Our crawlers use German residential ISP proxies with realistic browser fingerprints and full cookie session management to bypass perimeter defenses.

JavaScript rendering
Playwright execution for dynamic filters

Category pages and search filters rely on dynamic hydration. We run full Playwright browser sessions to trigger infinite scrolls and capture data that basic HTTP clients miss.

Schema stability
Resilient selectors for promotional layouts

Marketing campaigns frequently alter Hawesko's DOM structure. Our strategy uses multiple fallback chains per field, including structured JSON-LD data extraction, ensuring consistent output.

Change detection
Track vintage and price shifts

Wine catalogues change constantly as vintages sell out. We maintain a hash index of last-seen values per SKU, emitting clean diffs when prices change or new vintages arrive.

Monitoring & alerting
24/7 pipeline health tracking

Every run emits structured logs to our observability stack. We alert on null-rate spikes or coverage drops, resolving issues before they impact your downstream analytics.

Applications

Who uses Hawesko data, and how

Teams across industries use hawesko.de data to build competitive products and smarter operations.

01
Competitor Price Monitoring

European wine retailers track Hawesko pricing, discounts, and package deals to optimise their own pricing strategies.

02
Market & Trend Research

Analysts monitor shifts in grape varieties, regional popularity, and average price points to identify consumer trends.

03
Assortment Planning

Beverage distributors analyze the catalogue to identify gaps in their own portfolios and source new wineries.

04
Algorithmic Pricing

eCommerce teams feed competitor price points into dynamic repricing engines to maintain market positioning.

05
AI Sommelier Training

Machine learning teams use structured tasting notes, grape blends, and food pairings to train recommendation algorithms.

06
Brand Auditing

Wineries monitor their product representation, ensuring accurate tasting notes, correct vintages, and MAP compliance.

Why DataFlirt

"Hawesko holds the most structured dataset of European wine pricing and expert ratings, but accessing it requires navigating dynamic DOMs and strict rate limits."

Most teams underestimate the complexity of scraping European eCommerce targets. Reliable Hawesko extraction requires German residential proxies, Playwright for dynamic hydration, and strict schema validation to handle vintage rollovers. DataFlirt manages the infrastructure so your team can focus on market analysis.

Technical Spec

Hawesko scraper, technical capabilities

Everything supported by our hawesko.de scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions required for dynamic category filters and infinite scroll
Supported
CAPTCHA bypass
Automated solver integration with fallback to manual queue
Supported
German residential IPs
ISP-grade residential IPs from DE pools rotated per request
Supported
Vintage change detection
Diffing logic to detect when a SKU rolls over to a new vintage year
Supported
Expert rating parsing
Structured extraction of Parker, Falstaff, and Maroni scores
Supported
Package deal mapping
Resolving bundle SKUs to individual bottle components
Supported
Customer purchase history
Gated behind account login and loyalty program authentication
Partial
Age-verified checkout flows
Requires manual ID verification and age confirmation steps
Partial
Infrastructure

Infrastructure powering the Hawesko pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration and retry logic. Playwright manages JavaScript rendering and category pagination.

Residential Proxy Infrastructure

We maintain pools of German residential ISP proxies. Rotation happens per-request to prevent IP bans and rate limiting.

Cloud-Native Orchestration

Pipelines run on AWS infrastructure. Airflow handles scheduling and dependency management, ensuring reliable delivery.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested schema versioned per run
CSV
Flat file with typed columns for Excel compatibility
XLS
Formatted spreadsheet for non-technical stakeholders
Parquet
Columnar format optimised for BigQuery and Snowflake
AWS S3
Direct bucket delivery compatible with any data lake
Webhook
HTTP POST per record for real-time downstream processing
API
REST endpoint to query your extracted dataset
BigQuery
Streamed directly into your dataset with schema auto-detect
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About hawesko.de scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Hawesko legal?

Scraping publicly available pricing and catalogue information is generally permissible under applicable European law. DataFlirt extracts only public, non-authenticated data. We do not extract personal user data or circumvent authentication walls. Clients should review terms of service and consult legal counsel for specific use cases.

How do you handle Hawesko bot protection?

We use German residential ISP proxies, full Playwright browser sessions with realistic fingerprints, and request timing modelled on human behaviour. This prevents IP bans and ensures consistent access to the catalogue.

Can you track vintage changes?

Yes. Every pipeline run produces timestamped snapshots. We track vintage fields specifically and emit differential updates when a SKU transitions to a new year.

Do you extract expert ratings?

Yes. We parse structured scores and tasting notes from recognized critics like Robert Parker, Falstaff, and Luca Maroni directly from the product detail pages.

What is the delivery latency?

Full catalogue refreshes at a daily cadence complete within a 2-4 hour window. We can configure specific category pipelines for higher frequency monitoring if required.

Can I get a sample dataset before committing?

Absolutely. We provide a sample run of up to 500 SKUs as part of the pre-engagement scoping process so you can validate schema fit and data quality.

$ dataflirt scope --new-project --source=hawesko.de ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off catalogue dump or continuous price monitoring across thousands of wines, we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in food drink and kitchen

Services

Data Extraction for Every Industry

View All Services →