We extract live auction lots, bid velocities, expert estimates, and seller profiles from Catawiki. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Live Auctions objects from catawiki.com. All fields typed and schema-versioned.
"auction_id": "849201", "title": "Exclusive Vintage Bags Auction", "category": "Fashion", "end_time": "2026-08-14T18:00:00Z", "current_bid": 1250.0, "expert_estimate_min": 1500.0, "reserve_met": false
| # | auction_id | title | category | sub_category | end_time | current_bid |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Lot Details objects from catawiki.com. All fields typed and schema-versioned.
"lot_id": "9382716", "title": "Hermes Kelly 32 Black Togo Leather", "brand": "Hermes", "condition": "Very Good", "material": "Togo Leather", "dimensions": "32x23x12 cm", "shipping_costs": 45.0
| # | lot_id | auction_id | title | description | condition | material |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Bid History objects from catawiki.com. All fields typed and schema-versioned.
"lot_id": "9382716", "bidder_id": "User_8472", "bid_amount": 4200.0, "currency": "EUR", "timestamp": "2026-08-14T17:58:22Z", "auto_bid": true, "winning_bid": false
| # | lot_id | bid_id | bidder_id | bid_amount | currency | timestamp |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Expert Data objects from catawiki.com. All fields typed and schema-versioned.
"expert_id": "EXP_492", "expert_name": "Sophie Martin", "specialty": "Vintage Fashion & Bags", "total_curated_lots": 14290, "active_auctions": 12, "rating": 4.9, "review_count": 842
| # | expert_id | expert_name | specialty | bio | total_curated_lots | active_auctions |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Seller Profiles objects from catawiki.com. All fields typed and schema-versioned.
"seller_id": "SEL_9921", "username": "LuxuryVintageParis", "country": "France", "positive_feedback_pct": 99.4, "total_reviews": 3412, "member_since": "2018-03-12", "active_lots": 45
| # | seller_id | username | country | positive_feedback_pct | total_reviews | member_since |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our Catawiki scraper handles every layer of the platform: live bidding dynamics, expert valuations, seller metrics, and high-resolution lot imagery, with anti-bot circumvention built in.
Capture bid amounts, bidder IDs, timestamps, and auto-bid triggers in real time as auctions close.
Extract minimum and maximum expert valuations for every lot to calculate bid-to-estimate ratios.
Track whether reserve prices have been met, adjusted, or removed during the auction lifecycle.
Parse structured condition reports, material compositions, era tags, and authenticity guarantees.
Extract seller feedback scores, location data, language preferences, and historical lot volumes.
Traverse the entire taxonomy from Vintage Fashion and Luxury Watches to Classic Cars and Fine Art.
Capture all lot image URLs, including detail shots, hallmarks, and authenticity certificates.
Extract buyer protection fees, shipping costs by destination, and combined shipping rules.
Scrape expert bios, specialty areas, and curation histories for platform authority analysis.
Brief in. Clean data out.
Provide target categories, expert names, or seller IDs. We design the extraction schema together.
We configure Scrapy and Playwright crawlers, proxy rotation, and CAPTCHA handling for catawiki.com.
Schema validation, null-rate checks, bid-outlier detection, and sample lots before full launch.
JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
Catawiki protects its live auction data heavily. Here is how we stay resilient and maintain continuous extraction.
Catawiki uses advanced bot detection. Our crawlers use European residential ISP proxies with realistic browser fingerprints and full cookie session management.
Live auctions rely on WebSockets for bid updates. We intercept and decode these streams to capture bid velocities without constant HTTP polling.
Lot pages and image galleries are heavily JavaScript-rendered. We run full Playwright browser sessions to trigger lazy-loading and dynamic content.
Catawiki updates its UI frequently. Our selector strategy uses multiple fallback chains per field to ensure continuous data flow.
We manage request concurrency to avoid triggering IP bans while maintaining strict latency SLAs for closing auctions.
Appraisers and insurers use historical auction data and expert estimates to build accurate pricing models for vintage fashion and collectibles.
Analysts track category momentum, bid velocities, and sell-through rates to identify emerging trends in luxury goods.
Rival auction houses monitor Catawiki seller volumes, buyer premiums, and shipping structures to optimise their own platforms.
Dealers track lots with unmet reserve prices or low bid-to-estimate ratios to source inventory for secondary markets.
Brand protection teams monitor listings, seller profiles, and lot images to identify unauthorised dealers and counterfeit goods.
Alternative asset investors analyse historical returns on specific watch models, vintage bags, and art pieces.
"Catawiki holds the most structured dataset of European alternative assets and vintage luxury, but accessing it requires navigating complex live-bidding architectures."
Extracting live auction data requires reliable WebSocket interception, European residential proxies, and sub-second latency handling. DataFlirt absorbs that complexity so your analysts can focus on market trends, not infrastructure maintenance.
Everything supported by our catawiki.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering and WebSocket interception for live auctions.
We maintain pools of residential ISP proxies across European regions to bypass geo-restrictions and maintain high trust scores.
Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, while PostgreSQL stores state and deduplication hashes.
Data delivered to where your team already works — no new tooling required.
About catawiki.com scraping, legality, and pipeline operations.
Ask us directly →Yes. We intercept WebSocket connections to capture live bid increments, bidder IDs, and countdown extensions with sub-second latency during auction closing windows.
Yes. Every lot record includes the minimum and maximum expert valuations, along with the assigned expert's name and profile link.
We track the reserve price status indicator, capturing when a reserve is met or if a lot closes below the reserve threshold.
Yes. We can extract all image URLs or download the raw image files directly to your S3 bucket, including detail shots and hallmarks.
We support all categories, including Fashion, Watches, Classic Cars, Art, Stamps, and Coins. The schema adapts to category-specific fields like material, mileage, or provenance.
We utilise European residential proxies, Playwright for full browser rendering, and automated CAPTCHA solvers to maintain high success rates without triggering blocks.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need historical auction results for valuation models or live bid tracking for arbitrage, we scope, build, and operate the pipeline.