SYSTEM all green source harney.com queue 1,842 pages p99 latency 218ms dataflirt.com · scraper/harney-com
RUN . 14 active pipelines . harney.com live

Harney & Sons data,
delivered structured.

We extract product listings, variant pricing, tasting notes, brewing instructions, and customer reviews from harney.com. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Products extracted
1,429 /run
Variant prices
6,104 /24h
Review records
84,912 /run
Active pipelines
14
Uptime
99.98%
Data Dictionary

Every field we extract from harney.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Product Listings objects from harney.com. All fields typed and schema-versioned.

product_idskutitlecategorytea_typecaffeine_levelbase_pricecurrencydescriptioningredientskosher_certifiedfair_tradeimage_urlsurl
product_listings
● 200 OK
"sku": "HT-101",
"title": "Hot Cinnamon Spice",
"tea_type": "Black Tea",
"caffeine_level": "40-60 milligrams",
"base_price": 10.5,
"currency": "USD",
"kosher_certified": true
# product_idskutitlecategorytea_typecaffeine_level
1
2
3

Complete list of extractable fields for Variants & Pricing objects from harney.com. All fields typed and schema-versioned.

variant_idproduct_idpackaging_typesizeweight_gramspricecompare_at_pricesubscription_pricein_stockskubarcode
variants_& pricing
● 200 OK
"variant_id": "314159265",
"packaging_type": "Classic Tin",
"size": "20 Sachets",
"price": 10.5,
"subscription_price": 9.45,
"in_stock": true,
"sku": "HT-101-TIN"
# variant_idproduct_idpackaging_typesizeweight_gramsprice
1
2
3

Complete list of extractable fields for Tasting & Brewing objects from harney.com. All fields typed and schema-versioned.

product_idaromabodyflavoursbrew_time_minutesbrew_temp_fahrenheitbrew_temp_celsiusliquor_colourinstructions_text
tasting_& brewing
● 200 OK
"product_id": "8923471",
"aroma": "Strong cinnamon and sweet clove",
"body": "Medium",
"flavours": "Spicy, sweet, cinnamon, orange",
"brew_time_minutes": "5",
"brew_temp_fahrenheit": "212",
"brew_temp_celsius": "100"
# product_idaromabodyflavoursbrew_time_minutesbrew_temp_fahrenheit
1
2
3

Complete list of extractable fields for Reviews & Ratings objects from harney.com. All fields typed and schema-versioned.

review_idproduct_idratingauthor_nameverified_buyerreview_titlereview_bodycreated_athelpful_votes
reviews_& ratings
● 200 OK
"review_id": "REV-99281",
"rating": 5,
"author_name": "Sarah M.",
"verified_buyer": true,
"review_title": "My daily staple",
"review_body": "The perfect blend of sweet and spicy without any added sugar.",
"created_at": "2026-02-14T08:30:00Z"
# review_idproduct_idratingauthor_nameverified_buyerreview_title
1
2
3

Complete list of extractable fields for Categories objects from harney.com. All fields typed and schema-versioned.

category_idnameslugparent_categorydescriptionproduct_counturlscraped_at
categories
● 200 OK
"category_id": "COL-42",
"name": "Earl Grey Teas",
"slug": "earl-grey-tea",
"parent_category": "Black Tea",
"product_count": 24,
"url": "https://www.harney.com/collections/earl-grey-tea",
"scraped_at": "2026-05-12T10:00:00Z"
# category_idnameslugparent_categorydescriptionproduct_count
1
2
3

Capabilities

Deep tea catalogue extraction

Our harney.com scraper handles the complexities of modern headless commerce storefronts, parsing nested variant structures, subscription pricing logic, and detailed product metadata.

Full Catalogue Extraction

Extract all active listings across loose leaf teas, sachets, teaware, and gifts. Captures every metadata field exposed on the product page.

Variant Mapping

Map parent products to all child variants including tin sizes, bulk bags, and sample pouches with accurate SKU and barcode data.

Tasting Note Parsing

Structure unstructured text into discrete aroma, body, liquor colour, and flavour profile fields for quantitative analysis.

Brewing Instruction Structuring

Normalise brewing temperatures across Fahrenheit and Celsius, alongside steeping times and water volume recommendations.

Review & Rating Mining

Extract full review text, star ratings, and verified buyer flags paginated across all customer feedback.

Subscription Price Tracking

Capture one-time purchase prices alongside Subscribe & Save discount tiers for every valid variant.

Stock & Availability

Monitor out-of-stock statuses at the variant level to track supply chain constraints and product popularity.

Ingredient & Allergen Data

Extract complete ingredient lists, caffeine levels, kosher certifications, and fair trade designations.

Scheduled Pipeline Modes

Run one-off bulk exports or configure continuous pipelines at daily or weekly cadences with change-detection diffing.

// engagement pipeline

From target URLs to warehouse records

Brief in. Clean data out.

Define Scope
d 0

Provide target collection URLs or specify a full-site crawl. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy crawlers, session management, and parsing logic specific to harney.com's frontend architecture.

Validation & QA
d 4–6

Schema validation, null-rate checks, and variant price-outlier detection before full launch.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

Handling modern commerce storefronts

Extracting structured data from dynamic storefronts requires handling state hydration and API rate limits. Here is how we ensure reliable delivery.

pipeline-monitor · harney.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
State hydration
Extracting hidden variant data

Modern commerce platforms often load variant pricing and inventory data via internal APIs or inline JSON state. We parse this underlying state directly rather than scraping the DOM, ensuring 100% accuracy for out-of-stock variants and subscription pricing.

Pagination limits
Navigating collection endpoints

Deep category pages often truncate results or use infinite scroll. Our crawlers interact with the underlying pagination APIs to ensure complete catalogue coverage without missing edge-case SKUs.

Anti-bot layer
Residential proxy rotation

We route requests through US-based residential proxies with realistic TLS fingerprints to prevent IP bans and ensure uninterrupted data extraction during bulk collection runs.

Change detection
Only re-scrape what changes

We maintain a hash index of last-seen values per field. Subsequent runs only push diffs. This reduces compute cost and downstream processing load for your engineering teams.

Monitoring
24/7 pipeline health checks

Every run emits structured logs to our observability stack. We alert on null-rate spikes, schema drift, and coverage drops. SLA uptime is contractual.

Applications

Who uses harney.com data

Teams across industries use harney.com data to build competitive products and smarter operations.

01
Competitor Price Tracking

Specialty tea retailers monitor Harney & Sons pricing strategies across variant sizes to optimise their own margins.

02
Market Research

F&B analysts track new product launches, seasonal blends, and discontinued lines to identify consumer trends.

03
Product Attribute Analysis

R&D teams extract tasting notes and ingredient combinations to inform new product development.

04
Inventory Forecasting

Supply chain analysts monitor out-of-stock variants to identify supply constraints in specific tea origins.

05
Sentiment Analysis

Marketing teams mine review corpora to understand consumer preferences regarding flavour profiles and packaging.

06
MAP Monitoring

Wholesale distributors track retail pricing to ensure compliance with minimum advertised price agreements.

Why DataFlirt

"Harney & Sons represents one of the most detailed structured datasets for premium teas, but extracting the variant-level pricing and tasting notes requires a dedicated pipeline."

Extracting e-commerce data from modern storefronts like harney.com requires handling dynamic variant hydration and rate limits. DataFlirt manages the underlying extraction infrastructure, outputting clean, normalised tea catalogues directly to your data warehouse so your analysts can focus on market intelligence.

Technical Spec

Harney & Sons scraper technical specifications

Everything supported by our harney.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Playwright sessions for dynamic variant loading and review widgets.
Supported
Residential proxy rotation
US-based ISP proxies rotated to prevent rate limiting.
Supported
Variant mapping
Parent to child SKU relationships with all packaging and size combinations.
Supported
Review pagination
Full review corpus extraction across all product pages.
Supported
Change detection (diffs)
Hash-based diffing to emit only updated records since the last run.
Supported
Webhook delivery
HTTP POST per record for real-time downstream processing.
Supported
Wholesale account pricing
B2B wholesale pricing requires authenticated account credentials.
Partial
Customer order history
Personalised order data and loyalty point balances are gated.
Partial
Infrastructure

Infrastructure powering the pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering and interaction flows. Combined via scrapy-playwright middleware.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies. Rotation happens per-request with sticky sessions where required. IP score monitoring prevents blacklisted pool contamination.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested arrays.
CSV
Flat file with typed columns.
XLS
Excel compatible tabular exports.
Parquet
Columnar format for data warehouses.
AWS S3
Direct bucket delivery.
Webhook
HTTP POST per record.
API
REST endpoints for on-demand querying.
PostgreSQL
Direct database upserts.
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About harney.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping harney.com legal?

Scraping publicly available information is generally permissible under applicable law. DataFlirt targets only public, non-authenticated product, pricing, and review data. We do not extract personal data or circumvent authentication walls.

How do you handle variant pricing?

We extract the underlying JSON state data to capture all variant combinations, including base prices, compare-at prices, and subscription discounts, mapped accurately to their respective SKUs.

How fresh is the data?

Full catalogue refreshes at daily or weekly cadences complete within a few hours. Historical snapshots are available from the day your pipeline is commissioned.

Do you support review extraction?

Yes. We extract the full review corpus, including star ratings, author names, verified buyer flags, and helpful votes, paginated across all product pages.

What is the minimum viable engagement?

Our packages start at defined collection lists with weekly delivery. For full-site catalogues or custom schema requirements, we price based on volume and delivery frequency.

Can I request a sample dataset?

Yes. We provide a sample run of up to 50 products as part of the pre-engagement scoping process so you can validate schema fit and data quality.

$ dataflirt scope --new-project --source=harney.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off catalogue dump or a continuous price-monitoring feed. We scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in food drink and kitchen

Services

Data Extraction for Every Industry

View All Services →