SYSTEM all green source watchbase.com queue 1,492 pages p99 latency 184ms dataflirt.com · scraper/watchbase-com
RUN · 18 active pipelines · watchbase.com live

Horological data,
at warehouse scale.

We extract watch reference numbers, caliber specifications, case dimensions, and family histories from Watchbase. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake.

Watches extracted
41.2K /run
Calibers mapped
4.8K /run
Brands tracked
134
Active pipelines
18
Uptime
99.98%
Data Dictionary

Every field we extract from watchbase.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Watch Models objects from watchbase.com. All fields typed and schema-versioned.

reference_numberbrandfamilynamelimited_editionproduced_yearcase_materialbezel_materialglasscase_backcase_shapediameter_mmheight_mmwater_resistance_mdial_colourdial_finishindexeshands
watch_models
● 200 OK
"reference_number": "116610LN",
"brand": "Rolex",
"family": "Submariner",
"case_material": "Stainless Steel",
"diameter_mm": 40.0,
"water_resistance_m": 300,
"dial_colour": "Black"
# reference_numberbrandfamilynamelimited_editionproduced_year
1
2
3

Complete list of extractable fields for Calibers objects from watchbase.com. All fields typed and schema-versioned.

caliber_referencebrandbase_movementmovement_typedisplaydate_complicationchronographjewelspower_reserve_hfrequency_vphdiameter_mmheight_mmtime_complicationhands
calibers
● 200 OK
"caliber_reference": "3135",
"brand": "Rolex",
"movement_type": "Automatic",
"jewels": 31,
"power_reserve_h": 48,
"frequency_vph": 28800,
"date_complication": true
# caliber_referencebrandbase_movementmovement_typedisplaydate_complication
1
2
3

Complete list of extractable fields for Brands objects from watchbase.com. All fields typed and schema-versioned.

brand_namefounded_yearfounderheadquarterswebsitedescriptionwatch_countcaliber_countfamily_countparent_company
brands
● 200 OK
"brand_name": "Omega",
"founded_year": 1848,
"founder": "Louis Brandt",
"headquarters": "Biel/Bienne, Switzerland",
"parent_company": "Swatch Group",
"watch_count": 3482,
"caliber_count": 184
# brand_namefounded_yearfounderheadquarterswebsitedescription
1
2
3

Complete list of extractable fields for Families objects from watchbase.com. All fields typed and schema-versioned.

brandfamily_namedescriptionwatch_countearliest_model_yearlatest_model_yearsignature_featuresub_families_count
families
● 200 OK
"brand": "Patek Philippe",
"family_name": "Nautilus",
"watch_count": 142,
"earliest_model_year": 1976,
"signature_feature": "Porthole case shape",
"sub_families_count": 4
# brandfamily_namedescriptionwatch_countearliest_model_yearlatest_model_year
1
2
3

Complete list of extractable fields for Market & Pricing objects from watchbase.com. All fields typed and schema-versioned.

reference_numbercurrencyretail_priceprice_dateavailability_statusdiscontinuedlimited_edition_countmarket_segment
market_& pricing
● 200 OK
"reference_number": "15202ST.OO.1240ST.01",
"currency": "USD",
"retail_price": 33200.0,
"price_date": "2023-10-14",
"discontinued": true,
"availability_status": "Out of Production",
"market_segment": "Luxury Sports"
# reference_numbercurrencyretail_priceprice_dateavailability_statusdiscontinued
1
2
3

Capabilities

Horological data extracted with precision

Our Watchbase scraper captures the highly relational nature of watch data, mapping base movements to specific calibers and tying them to thousands of individual reference numbers.

Full Reference Extraction

Capture case dimensions, materials, dial colours, and water resistance for every watch reference listed.

Caliber & Movement Specs

Extract jewel counts, power reserves, operating frequencies, and complication lists for every mapped movement.

Relational Mapping

Preserve the links between parent brands, watch families, specific references, and the calibers that power them.

Brand Taxonomy

Extract founding histories, parent company hierarchies, and total model counts per manufacturer.

Discontinued Status

Track production years and identify discontinued models for secondary market valuation models.

High-Resolution Images

Capture direct URLs for case, dial, and movement imagery for visual authentication and catalogue building.

Complication Parsing

Normalise complex complication strings into structured boolean fields (e.g., perpetual calendar, tourbillon, moonphase).

Retail Price Capture

Extract listed retail prices and currency data where available to establish baseline valuation metrics.

Scheduled Diffs

Run continuous pipelines to detect new reference releases and caliber updates without re-scraping the entire database.

// engagement pipeline

From brand list to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target brands, specific families, or request the entire Watchbase catalogue. We map the required schema.

Pipeline Build
d 2–4

We configure Scrapy / Playwright crawlers, proxy rotation, and session management for watchbase.com.

Validation & QA
d 4–6

Schema validation, null-rate checks, and relational integrity testing between watches and calibers.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

Handling horological data structures

Extracting Watchbase requires maintaining relational integrity across thousands of pages while navigating rate limits.

pipeline-monitor · watchbase.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Relational integrity
Mapping calibers to references

Watch data is inherently relational. A single base movement might be modified into ten different calibers, powering fifty different watches. Our pipeline maintains these foreign-key relationships during extraction, delivering normalised tables that join perfectly in your warehouse.

Rate limiting
Respectful crawl concurrency

Watchbase enforces strict request limits to protect their database. We utilise distributed proxy pools and configure low-concurrency, high-delay request profiles to ensure continuous extraction without triggering IP bans or degrading site performance.

DOM variability
Handling inconsistent spec sheets

Older watch references often have missing fields or non-standard specification formats compared to modern releases. Our parsers use fallback regex patterns and fuzzy matching to extract dimensions and materials even when the DOM structure deviates.

Taxonomy normalisation
Standardising horological terms

Case materials and dial colours are often described inconsistently (e.g., 'Pink Gold' vs 'Rose Gold'). We can apply post-extraction normalisation dictionaries to standardise these attributes for easier filtering and analysis.

Asset management
Image URL extraction

We extract the highest resolution image URLs available for each reference and caliber, avoiding thumbnails. These URLs are delivered alongside the metadata, ready for ingestion into your DAM or CDN.

Applications

Who uses Watchbase data — and how

Teams across industries use watchbase.com data to build competitive products and smarter operations.

01
Secondary Market Valuation

Grey market dealers use structured reference data and retail baselines to train pricing algorithms and detect arbitrage opportunities.

02
Authentication & Verification

Authentication platforms cross-reference case dimensions, caliber specs, and jewel counts to identify counterfeit watches.

03
Inventory Management

Retailers auto-populate their e-commerce catalogues with precise specifications by matching incoming stock to Watchbase references.

04
Market Research

Analysts track brand output, complication trends, and material usage over time to understand horological industry shifts.

05
Insurance Valuation

Underwriters use historical production dates and reference specifications to assess replacement values for scheduled property.

06
AI Training Data

Machine learning teams use the image URLs and associated metadata to train visual recognition models for watch identification.

Why DataFlirt

"Watchbase is the definitive digital encyclopaedia for horology, but extracting its highly relational caliber-to-reference data requires a precise extraction schema."

Watch data is inherently relational. A single caliber might power forty different references across multiple brands, each with distinct case materials and dial configurations. DataFlirt reconstructs this relational graph, handling rate limits and DOM variations so your engineers can focus on valuation models, not scraping infrastructure.

Technical Spec

Watchbase scraper — technical capabilities

Everything supported by our watchbase.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

Reference number extraction
Captures precise alphanumeric reference codes for exact model matching
Supported
Caliber specification mapping
Extracts detailed movement data including frequency and jewel count
Supported
Relational foreign keys
Maintains links between brands, families, watches, and calibers
Supported
High-res image URLs
Extracts direct links to the highest quality asset available
Supported
Complication parsing
Converts text-based complication lists into structured arrays
Supported
Historical production dates
Captures introduction and discontinuation years where listed
Supported
Change detection (diffs)
Only emits records with changed fields since last pipeline run
Supported
Direct database dump
Direct access to Watchbase's backend SQL database
Partial
Premium API endpoints
Extraction via Watchbase's paid commercial API
Partial
User collections
Scraping private user watch collections or wishlists
Partial
Infrastructure

Infrastructure powering the Watchbase pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy Orchestration

Scrapy handles the deep crawling required to traverse brand taxonomies down to individual reference pages, ensuring complete catalogue coverage.

Relational State Management

PostgreSQL maintains the mapping between discovered calibers and watches during the crawl, ensuring foreign keys are correctly assigned before export.

Cloud-Native Delivery

Airflow schedules regular diff-runs to detect new releases, pushing updated Parquet files directly to your S3 buckets.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Nested JSON preserving the watch-to-caliber relationship
CSV
Flat files separated by entity (Watches, Calibers, Brands)
XLS
Excel format for manual review by horological experts
Parquet
Columnar format optimised for analytical queries in BigQuery
AWS S3
Direct delivery to your cloud storage infrastructure
Webhook
HTTP POST for real-time alerts on new reference discoveries
API
Queryable REST endpoints for on-demand data retrieval
Postgres
Direct SQL inserts maintaining relational integrity
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About watchbase.com scraping, legality, and pipeline operations.

Ask us directly →
Can you extract data for specific brands only?

Yes. We can configure the pipeline to target specific brands (e.g., Rolex, Patek Philippe, Audemars Piguet) or specific families rather than crawling the entire Watchbase catalogue.

How do you handle the relationship between watches and calibers?

Our extraction schema treats watches and calibers as separate entities linked by a foreign key. This allows you to ingest the data into a relational database without duplicating caliber specifications across hundreds of watch models.

Are high-resolution images included?

We extract the direct URLs for all images associated with a reference or caliber. We do not host the images, but provide the URLs for your systems to download and ingest.

How frequently can the data be updated?

Given the relatively slow release cycle of new watch models, we typically recommend a weekly or monthly pipeline cadence to detect new references and calibers, though faster cadences are available.

Do you standardise the terminology used for materials?

By default, we extract the exact text displayed on Watchbase. However, we can implement custom post-processing scripts to normalise terms according to your internal taxonomy.

Is it possible to track discontinued models?

Yes. If Watchbase updates a model's status or lists a final production year, our change-detection pipeline will capture this update and flag the reference as discontinued in your next data delivery.

$ dataflirt scope --new-project --source=watchbase.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a complete dump of all mechanical calibers or continuous tracking of new reference releases — we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in watches

Services

Data Extraction for Every Industry

View All Services →