SYSTEM all green source distiller.com queue 14,892 pages p99 latency 184ms dataflirt.com · scraper/distiller-com
RUN · 18 active pipelines · distiller.com live

Distiller data,
at warehouse scale.

We extract spirit catalogues, expert reviews, community ratings, and complex flavour profile matrices from Distiller. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Spirits extracted
124K /run
Community reviews
3.2M /total
Flavour matrices
98K /run
Active pipelines
18
Uptime
99.94%
Data Dictionary

Every field we extract from distiller.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Spirit Metadata objects from distiller.com. All fields typed and schema-versioned.

spirit_idnamebranddistillercategorysub_categoryabvagecask_typeprice_tierimage_urlurl
spirit_metadata
● 200 OK
"spirit_id": "lagavulin-16",
"name": "Lagavulin 16 Year",
"brand": "Lagavulin",
"category": "Whisky",
"abv": 43.0,
"age": 16,
"cask_type": "Ex-Bourbon"
# spirit_idnamebranddistillercategorysub_category
1
2
3

Complete list of extractable fields for Scores & Ratings objects from distiller.com. All fields typed and schema-versioned.

spirit_iddistiller_scorecommunity_ratingrating_countreview_countexpert_reviewer_nameexpert_review_daterecommendation_status
scores_& ratings
● 200 OK
"spirit_id": "lagavulin-16",
"distiller_score": 93,
"community_rating": 4.48,
"rating_count": 14205,
"review_count": 3102,
"expert_reviewer_name": "Stephanie Moreno"
# spirit_iddistiller_scorecommunity_ratingrating_countreview_countexpert_reviewer_name
1
2
3

Complete list of extractable fields for Flavour Profile objects from distiller.com. All fields typed and schema-versioned.

spirit_idsmokypeatyspicyherbaloilyfull_bodiedrichsweetbrinyvanillacaramelfloral
flavour_profile
● 200 OK
"spirit_id": "lagavulin-16",
"smoky": 85,
"peaty": 90,
"spicy": 40,
"sweet": 30,
"briny": 65,
"full_bodied": 80
# spirit_idsmokypeatyspicyherbaloily
1
2
3

Complete list of extractable fields for Tasting Notes objects from distiller.com. All fields typed and schema-versioned.

spirit_idnote_typeauthorcontentdate_publishednose_notespalate_notesfinish_notes
tasting_notes
● 200 OK
"spirit_id": "lagavulin-16",
"note_type": "expert",
"author": "Stephanie Moreno",
"nose_notes": "Intense peat smoke, iodine, seaweed.",
"palate_notes": "Rich, thick, sweet malt, massive peat.",
"finish_notes": "Long, spicy, roaring peat smoke."
# spirit_idnote_typeauthorcontentdate_publishednose_notes
1
2
3

Complete list of extractable fields for Community Reviews objects from distiller.com. All fields typed and schema-versioned.

review_idspirit_iduser_nameuser_profile_urlratingreview_textdate_postedhelpful_votes
community_reviews
● 200 OK
"review_id": "rev-84920",
"spirit_id": "lagavulin-16",
"user_name": "MaltMaster99",
"rating": 5.0,
"review_text": "The benchmark for Islay scotches.",
"date_posted": "2023-11-14"
# review_idspirit_iduser_nameuser_profile_urlratingreview_text
1
2
3

Capabilities

Extract the complete global spirits taxonomy

Our Distiller scraper navigates the entire platform: spirit catalogues, complex flavour matrices, expert tasting notes, and paginated community reviews. Built with JavaScript execution to capture dynamically rendered data.

Full Catalogue Extraction

Scrape Whisky, Rum, Tequila, Mezcal, Gin, Vodka, and Brandy categories with complete metadata including ABV, age, and cask type.

Flavour Profile Matrices

Extract the underlying 0-100 numeric values from Distiller radar charts, capturing peat, smoke, vanilla, and spice metrics.

Expert vs Community Scores

Separate the official Distiller Score from crowd-sourced community ratings, capturing review volume and rating distributions.

Tasting Notes Parsing

Structured breakdown of expert reviews into distinct nose, palate, and finish text blocks for NLP analysis.

Distillery & Brand Mapping

Link independent bottlings and specific labels back to their origin distilleries using structured platform metadata.

Pagination & Scroll Handling

Navigate infinite scroll on category pages and deep pagination on popular spirits to capture the entire review corpus.

Community Review Corpus

Extract user names, star ratings, review text, and helpful votes across millions of community entries.

Price Tier Tracking

Capture relative cost indicators and price tier classifications to correlate with quality scores.

Scheduled Updates

Track shifts in community ratings and review velocity over time with recurring pipeline executions.

Production Method Data

Isolate technical specifications including mash bill percentages, distillation methods, and maturation processes.

// engagement pipeline

From spirit category to structured dataset

Brief in. Clean data out.

Define Scope
d 0

Provide spirit categories, specific distilleries, or target URLs. We design the schema for flavour matrices and reviews.

Pipeline Build
d 2–4

We configure Playwright to render flavour charts and Scrapy to traverse the paginated review corpus.

Validation & QA
d 4–6

Schema validation, null-rate checks on tasting notes, and numeric verification of radar chart extraction.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on an agreed cadence.

Under the hood

How our Distiller pipeline handles the hard parts

Extracting data from Distiller requires handling dynamic rendering and deep pagination. Here is how we build resilient pipelines.

pipeline-monitor · distiller.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
JavaScript rendering
Playwright execution for flavour charts

Distiller flavour profiles are rendered dynamically via JavaScript canvas or complex DOM structures. We run full Playwright browser sessions to execute the rendering logic and extract the underlying 0-100 numeric values for each flavour axis.

Infinite scroll pagination
Traversing dynamic category lists

Spirit category pages use infinite scroll mechanisms. Our crawlers intercept the underlying API requests or simulate user scroll behaviour to ensure comprehensive catalogue coverage without missing items.

Anti-bot layer
Residential proxies for deep scraping

Scraping thousands of paginated community reviews triggers rate limits. We utilise residential ISP proxies and rotate fingerprints to maintain access during extensive historical review extraction.

Schema normalisation
Standardising disparate attributes

A Scotch whisky has different metadata fields than a Mezcal. We normalise these disparate attributes into a consistent schema, ensuring your downstream database receives clean, structured records regardless of the spirit type.

Change detection
Delta updates for community ratings

For ongoing pipelines, we track the last scraped review ID. Subsequent runs only extract new reviews and updated aggregate scores, reducing compute costs and delivering clean delta files.

Applications

Who uses Distiller data — and how

Teams across industries use distiller.com data to build competitive products and smarter operations.

01
Beverage Industry Market Research

Brands analyse flavour trends, category growth, and competitor profiles to guide new product development and maturation strategies.

02
Competitor Benchmarking

Distilleries track their Distiller Scores against peer products, monitoring community sentiment and rating shifts over time.

03
Retail & Distribution Assortment

Liquor retailers and distributors use expert scores and community ratings to select high-performing spirits for inventory.

04
Recommendation Engine Training

Machine learning teams use the 0-100 flavour profile matrices to train content-based filtering algorithms for spirit pairing apps.

05
Sentiment Analysis

NLP models process the vast corpus of community tasting notes to identify emerging consumer preferences and vocabulary.

06
Pricing Strategy

Analysts correlate price tier classifications with Distiller Scores to identify premiumisation opportunities or value-brand positioning.

Why DataFlirt

"Distiller holds the definitive taxonomy of global spirits and flavour profiles — but extracting radar chart matrices requires precise JavaScript execution and structural normalisation."

Most teams fail at scraping Distiller because the core value lies in the flavour profile charts, which are dynamically rendered. Reliable extraction demands full browser execution, session management, and careful parsing of the underlying data objects. DataFlirt handles the rendering and normalisation so you get clean, queryable matrices ready for analysis.

Technical Spec

Distiller scraper — technical capabilities

Everything supported by our distiller.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions to extract dynamically rendered data
Supported
Flavour profile extraction
Capture numeric values for all axes on the flavour radar charts
Supported
Review pagination
Traverse the complete history of community reviews per spirit
Supported
Search result scraping
Extract catalogue data based on specific keyword queries
Supported
Historical rating tracking
Time-series capture of community rating shifts over scheduled runs
Supported
Expert vs Community separation
Distinct fields for official Distiller Scores versus user averages
Supported
Webhook delivery
HTTP POST per record for real-time downstream integration
Supported
Distiller Pro user collections
Private user library and collection data requires authentication
Partial
Private tasting journals
Personal user notes gated behind account login walls
Partial
Infrastructure

Infrastructure powering the Distiller pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheusBigQuerySnowflake
Scrapy + Playwright Stack

Scrapy handles crawl orchestration and pagination logic. Playwright executes JavaScript to render flavour charts and extract underlying data objects.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies to handle deep pagination through millions of community reviews without triggering rate limits.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow manages scheduling and dependencies, ensuring reliable delivery of delta updates.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested arrays for complex flavour matrices
CSV
Flat file with typed columns for spreadsheet analysis
XLS
Excel compatible format for immediate business use
Parquet
Columnar format optimised for analytical queries
AWS S3
Direct bucket delivery compatible with modern data lakes
Webhook
HTTP POST per record for event-driven architectures
API
REST endpoints to query your extracted dataset
BigQuery
Streamed directly into your GCP environment
Snowflake
Stage and COPY INTO workflow for immediate availability
PostgreSQL
Direct database upserts with schema conflict resolution
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About distiller.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Distiller legal?

Scraping publicly available information from Distiller is generally permissible. DataFlirt targets only public spirit catalogues, expert reviews, and community ratings. We do not circumvent authentication walls to extract private user collections or tasting journals.

Can you extract the flavour profile radar charts?

Yes. While the charts are visually rendered, we extract the underlying 0-100 numeric values for each axis (e.g., peaty, smoky, sweet, floral) using JavaScript execution and DOM parsing.

How do you handle pagination on community reviews?

We traverse the full review corpus by intercepting background API requests or using Playwright to simulate interactions, ensuring we capture all historical reviews for a given spirit.

Can I get data on specific spirit categories only?

Yes. We can configure the pipeline to target specific categories like Whisky, Tequila, or Gin, or even restrict extraction to specific distilleries and brands.

How frequently can you update community ratings?

Pipelines can be configured for daily, weekly, or monthly cadences. We use delta extraction to only pull new reviews and updated aggregate scores, minimising processing overhead.

Do you scrape user profiles?

We extract public user names, profile URLs, and review histories associated with public community ratings. We do not extract private account details or gated collections.

Can you map independent bottlers to original distilleries?

Yes. We extract all available metadata fields, allowing you to link independent bottlings back to their origin distilleries based on platform taxonomy.

$ dataflirt scope --new-project --source=distiller.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a complete spirit catalogue dump or a continuous feed of community reviews and flavour profiles — we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in food drink and kitchen

Services

Data Extraction for Every Industry

View All Services →