We extract product listings, tasting profiles, ABV metrics, vintage data, and pricing from The Whisky Exchange. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Product Listings objects from thewhiskyexchange.com. All fields typed and schema-versioned.
"sku": "014389", "name": "Lagavulin 16 Year Old", "distillery": "Lagavulin", "region": "Islay", "age": 16, "abv": 43.0, "volume_cl": 70, "price": 82.95, "stock_status": "In Stock"
| # | sku | name | distillery | region | age | abv |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Tasting Notes objects from thewhiskyexchange.com. All fields typed and schema-versioned.
"sku": "014389", "nose": "Intensely flavoured, peat smoke with iodine and seaweed and a rich, deep sweetness.", "palate": "Dry peat smoke fills the palate with a gentle but strong sweetness, followed by sea and salt with touches of wood.", "finish": "A long, elegant peat-filled finish with lots of salt and seaweed.", "character_tags": "['Peaty', 'Smoky', 'Maritime', 'Rich']", "reviewer": "The Whisky Exchange", "score": 92
| # | sku | nose | palate | finish | character_tags | reviewer |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Pricing & Stock objects from thewhiskyexchange.com. All fields typed and schema-versioned.
"sku": "014389", "price_gbp": 82.95, "price_ex_vat": 69.13, "currency": "GBP", "tax_status": "Inc. VAT", "delivery_restrictions": "['US', 'CA']", "in_stock": true, "stock_qty": 142
| # | sku | price_gbp | price_ex_vat | currency | tax_status | delivery_restrictions |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Cask & Vintage objects from thewhiskyexchange.com. All fields typed and schema-versioned.
"sku": "108492", "vintage_year": 1990, "bottling_year": 2021, "cask_type": "Sherry Butt", "cask_number": "4812", "number_of_bottles": 512, "bottler": "Signatory Vintage", "series": "Cask Strength Collection"
| # | sku | vintage_year | bottling_year | cask_type | cask_number | number_of_bottles |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Reviews & Ratings objects from thewhiskyexchange.com. All fields typed and schema-versioned.
"review_id": "REV-98412", "sku": "014389", "author": "PeatLover88", "rating": 5, "date": "2023-11-14", "title": "A classic Islay malt", "body": "Never disappoints. The benchmark for peated whisky.", "helpful_votes": 24
| # | review_id | sku | author | rating | date | title |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our scraper navigates age gates, extracts complex flavour profiles, maps distillery hierarchies, and tracks inventory for rare investment-grade bottles.
Extract ABV, volume, region, distillery, bottler, and age statements across all spirit categories.
Parse structured tasting notes separating nose, palate, and finish, alongside character tags and flavour profiles.
Capture ex-VAT and inc-VAT pricing, tracking currency conversions and regional tax adjustments.
Monitor vintage years, bottling dates, cask numbers, and limited edition bottle counts for investment analysis.
Extract character tags, style categories, and official flavour maps used for recommendation engines.
Track in-stock flags, low stock warnings, and availability changes across high-demand releases.
Scrape star ratings, user tasting notes, and review dates across the entire customer feedback corpus.
Automated session handling and cookie injection to reliably pass 18+ verification prompts.
Map shipping constraints and regional delivery blocks per SKU to understand market availability.
Brief in. Clean data out.
Provide target categories, distilleries, or specific SKUs. We map the extraction schema to your requirements.
We configure Scrapy crawlers, handle age-gate cookies, and set up residential proxies for thewhiskyexchange.com.
Schema validation, price-outlier detection, and tasting note completeness checks before full launch.
JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
Spirits eCommerce requires navigating age gates, regional pricing, and strict bot protection. Here is how we maintain reliable extraction.
The Whisky Exchange mandates age verification before displaying catalogue data. Our pipeline automatically handles session initiation and cookie injection to bypass the 18+ gateway without triggering bot flags.
Pricing changes based on user location and VAT status. We configure crawler sessions with specific regional headers and cookies to extract exact ex-VAT or inc-VAT prices based on your target market.
Rare bottles and limited releases sell out in minutes. We support high-frequency polling for specific SKUs to track stock availability and price changes in near real-time.
Spirits categorisation is deeply hierarchical. Our schema maps the exact relationship between parent regions, sub-regions, distilleries, and independent bottlers to maintain data integrity.
We utilise UK-based residential proxies and precise browser fingerprinting to navigate rate limits and bot protection, ensuring uninterrupted data flow from the catalogue.
Analysts track emerging trends in cask finishes, age statements, and regional popularity to inform product development.
Collectors and funds monitor pricing curves and stock depth for old and rare bottles to calculate asset appreciation.
Retailers track pricing parity, discount strategies, and shipping constraints to optimise their own merchandising.
Machine learning teams use structured tasting notes and character tags to train NLP models for recommendation engines.
Supply chain teams use stock availability signals across major retailers to predict supply shortages for specific distilleries.
Distilleries audit their representation, pricing consistency, and customer review sentiment across third-party retail platforms.
"The Whisky Exchange holds the definitive global catalogue of spirits and tasting profiles - but extracting structured flavour data requires a purpose-built pipeline."
Most teams fail at scraping spirits platforms because they trip over age gates, regional tax toggles, and strict bot protection. DataFlirt manages the residential proxies, session cookies, and schema mapping so your engineers can focus on the analysis - not the infrastructure.
Everything supported by our thewhiskyexchange.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration, deduplication, and retry logic. Playwright manages JavaScript rendering, age-gate cookies, and interaction flows.
We maintain pools of residential ISP proxies. Rotation happens per-request with sticky sessions for consistent regional pricing extraction.
Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.
Data delivered to where your team already works — no new tooling required.
About thewhiskyexchange.com scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available pricing, stock, and product specification data is generally permissible. DataFlirt targets only public, non-authenticated catalogue data. We do not extract personal user data or circumvent authentication walls.
Our crawlers are configured to automatically inject the required consent cookies at the start of every session, bypassing the 18+ prompt without manual intervention.
Yes. We configure the crawler sessions to select specific regional delivery settings, allowing us to capture exact ex-VAT pricing alongside standard inc-VAT retail prices.
Yes. The pipeline covers all categories, including limited editions, vintage releases, and the Old & Rare catalogue, extracting specific cask numbers and bottling dates.
We parse the raw text into structured JSON fields, separating nose, palate, and finish descriptions, and extract categorical character tags for easier database querying.
Pipelines can be configured for daily catalogue refreshes, or high-frequency hourly polling for specific high-demand SKUs to track stock availability.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a complete catalogue export or continuous price monitoring across rare vintages - we scope, build, and operate the pipeline. Tell us what you need.