SYSTEM all green source gubi.com queue 1,492 pages p99 latency 318ms dataflirt.com · scraper/gubi-com
RUN · 14 active pipelines · gubi.com live

Gubi catalogue data,
structured for scale.

We extract designer profiles, product dimensions, fabric variants, and 3D asset metadata from Gubi. Delivered as clean JSON, CSV, or Parquet to S3 or BigQuery on your cadence.

Products extracted
842 /run
Fabric variants
4,192 /run
Designers
94 /run
Active pipelines
14
Uptime
99.98%
Data Dictionary

Every field we extract from gubi.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Products objects from gubi.com. All fields typed and schema-versioned.

skunamedesignercollectioncategorysub_categorybase_materialdescriptioncare_instructionswarranty
products
● 200 OK
"sku": "10023-01",
"name": "Beetle Dining Chair",
"designer": "GamFratesi",
"collection": "Beetle",
"category": "Seating",
"sub_category": "Dining Chairs",
"base_material": "Brass",
"warranty": "2 years"
# skunamedesignercollectioncategorysub_category
1
2
3

Complete list of extractable fields for Variants & Upholstery objects from gubi.com. All fields typed and schema-versioned.

variant_skuparent_skufabric_groupfabric_namecolour_codebase_finishprice_eurimage_urlin_stock
variants_& upholstery
● 200 OK
"variant_sku": "10023-01-F03",
"parent_sku": "10023-01",
"fabric_group": "Group 3",
"fabric_name": "Kvadrat Hallingdal 65",
"colour_code": "130",
"base_finish": "Antique Brass",
"price_eur": 895.0,
"in_stock": true
# variant_skuparent_skufabric_groupfabric_namecolour_codebase_finish
1
2
3

Complete list of extractable fields for Dimensions objects from gubi.com. All fields typed and schema-versioned.

skuheight_cmwidth_cmdepth_cmseat_height_cmweight_kgpackage_volume_m3package_weight_kgmetric_standard
dimensions
● 200 OK
"sku": "10023-01",
"height_cm": 87.0,
"width_cm": 56.0,
"depth_cm": 58.0,
"seat_height_cm": 45.0,
"weight_kg": 8.2,
"package_volume_m3": 0.34
# skuheight_cmwidth_cmdepth_cmseat_height_cmweight_kg
1
2
3

Complete list of extractable fields for Designers objects from gubi.com. All fields typed and schema-versioned.

designer_idnamebionationalityactive_yearsfamous_worksprofile_image_urlrelated_collectionsstudio_location
designers
● 200 OK
"designer_id": "D-GAMF",
"name": "GamFratesi",
"nationality": "Danish-Italian",
"active_years": "2006-Present",
"famous_works": "['Beetle Chair', 'Bat Chair', 'Epic Table']",
"studio_location": "Copenhagen",
"related_collections": "['Beetle', 'Bat', 'Epic']"
# designer_idnamebionationalityactive_yearsfamous_works
1
2
3

Complete list of extractable fields for Assets & Downloads objects from gubi.com. All fields typed and schema-versioned.

skuimage_urlslifestyle_imagesassembly_manual_urlcad_2d_urlcad_3d_urlrevit_urlproduct_presentation_pdfcare_guide_url
assets_& downloads
● 200 OK
"sku": "10023-01",
"assembly_manual_url": "https://gubi.com/assets/manuals/beetle_assembly.pdf",
"cad_3d_url": "https://gubi.com/assets/cad/beetle_3d.dwg",
"revit_url": "https://gubi.com/assets/bim/beetle.rfa",
"care_guide_url": "https://gubi.com/assets/guides/fabric_care.pdf",
"image_urls": "['https://gubi.com/img/1.jpg', 'https://gubi.com/img/2.jpg']"
# skuimage_urlslifestyle_imagesassembly_manual_urlcad_2d_urlcad_3d_url
1
2
3

Capabilities

Extract the complete Gubi design catalogue

Our Gubi scraper navigates complex variant matrices, capturing every fabric option, base finish, and technical specification required for interior design and procurement platforms.

Complete Catalogue Extraction

Extract seating, lighting, tables, and storage products with full hierarchical category mapping.

Variant & Fabric Mapping

Capture every upholstery option across price groups, including Kvadrat and Dedar fabric specifications.

Dimension Normalisation

Standardise height, width, depth, and seat height measurements into a consistent metric schema.

Designer Metadata

Link products to designer profiles, biographies, and studio locations across the Gubi ecosystem.

Asset URL Harvesting

Collect direct URLs for 3D DWG files, Revit models, assembly PDFs, and high-resolution lifestyle images.

Stockist Directory Scraping

Extract global showroom and partner locations, including geocoordinates and contact details.

Material Finishes

Catalogue base materials like unlacquered brass, black chrome, and oiled walnut with exact naming conventions.

Scheduled Cadence

Run pipelines weekly or monthly to capture new collection launches and discontinued variants.

Anti-Bot Circumvention

Bypass rate limits and request blocks using residential proxy rotation and realistic session fingerprints.

// engagement pipeline

From design catalogue to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Select specific collections, categories, or designer portfolios. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Playwright crawlers to handle dynamic fabric rendering and asset link extraction.

Validation & QA
d 4–6

Schema validation, null-rate checks on dimensions, and variant completeness testing before full launch.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket or API endpoint on the agreed schedule.

Under the hood

How our pipeline handles furniture data complexity

High-end furniture sites rely on heavy JavaScript to render thousands of fabric and finish combinations. Here is how we extract it accurately.

pipeline-monitor · gubi.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Dynamic variant hydration
Rendering complex upholstery matrices

Gubi products often feature hundreds of fabric and base combinations loaded dynamically. We use Playwright to execute JavaScript, triggering variant state changes to capture exact pricing and SKUs for every possible combination.

Asset link extraction
Capturing 3D models and technical PDFs

Technical assets are frequently gated behind interactive UI elements. Our crawlers simulate user interactions to expose and extract direct download links for CAD files, Revit models, and care instructions.

Nested category parsing
Maintaining collection hierarchies

Furniture catalogues use deep taxonomies (e.g., Seating > Lounge Chairs > Beetle Collection). We reconstruct this hierarchy in the final dataset, ensuring products are correctly categorised for downstream filtering.

Anti-bot layer
Residential proxy rotation

We route requests through European residential IPs to prevent rate limiting during deep catalogue crawls, ensuring complete extraction without IP bans.

Schema normalisation
Standardised dimensions and weights

We parse raw text strings into structured numerical fields for dimensions (cm) and weights (kg), making the data immediately queryable for logistics and spatial planning.

Applications

Who uses Gubi data — and how

Teams across industries use gubi.com data to build competitive products and smarter operations.

01
Interior Design Aggregation

Digital design platforms aggregate Gubi products alongside other brands to offer comprehensive 3D planning tools to architects.

02
B2B Procurement Catalogues

Corporate procurement teams maintain updated internal catalogues of approved furniture for office fit-outs.

03
Competitor Material Analysis

Furniture manufacturers analyse Gubi's fabric groups and material choices to forecast industry design trends.

04
3D Asset Libraries

ArchViz studios automate the ingestion of DWG and Revit files to populate their rendering asset libraries.

05
Retailer Synchronisation

Authorised dealers synchronise their e-commerce platforms with Gubi's latest product specifications and imagery.

06
Trend & Colour Forecasting

Design agencies track the introduction of new upholstery colours and base finishes across collections over time.

Why DataFlirt

"Gubi represents the pinnacle of modern Danish design, but integrating their complex fabric and finish variants into standard procurement systems requires deep schema normalisation."

Extracting high-end furniture data means handling thousands of nested upholstery options, base finishes, and 3D asset links. DataFlirt manages the JavaScript rendering and schema standardisation so your procurement and design teams get clean, structured tables ready for immediate use.

Technical Spec

Gubi scraper — technical capabilities

Everything supported by our gubi.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions required for dynamic fabric configuration and pricing
Supported
Fabric variant mapping
Extracts all combinations of upholstery groups, colours, and bases
Supported
3D asset link capture
Direct URLs for DWG, 3DS, and Revit files
Supported
Dimension normalisation
Parses text dimensions into structured metric floats
Supported
Designer profiles
Extracts biography and cross-references related collections
Supported
Stockist geodata
Extracts showroom addresses and coordinates from the store locator
Supported
High-res image extraction
Captures maximum resolution URLs for product and lifestyle imagery
Supported
B2B Trade Pricing
Trade-specific discounts require authenticated partner portal access
Partial
Partner-only CAD models
Certain technical assets gated behind dealer login walls
Partial
Infrastructure

Infrastructure powering the Gubi pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration while Playwright executes JavaScript to render complex upholstery configuration matrices.

Asset URL Pipeline

Dedicated extraction logic to identify and validate direct download URLs for 3D models and technical PDFs.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS, scheduled via Apache Airflow to ensure reliable data delivery on your required cadence.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Nested schema ideal for complex variant matrices
CSV
Flat file with denormalised variant rows
XLS
Excel format for direct procurement team usage
Parquet
Columnar format for data warehouse ingestion
AWS S3
Direct bucket delivery
Webhook
HTTP POST per record for real-time systems
API
REST endpoint to query extracted catalogue data
PostgreSQL
Direct database upserts with schema matching
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About gubi.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Gubi legal?

Scraping publicly available catalogue information is generally permissible. DataFlirt targets only public product data, dimensions, and public asset links. We do not extract authenticated B2B portal data or violate GDPR. Clients should consult legal counsel for specific commercial use cases.

How do you handle the complex fabric variants?

We use Playwright to interact with the product configuration UI, iterating through fabric groups, colours, and base finishes to capture the specific SKU, price, and image for every possible combination.

Can you extract the 3D CAD files?

We extract the direct URLs to the 3D files (DWG, Revit, etc.) hosted by Gubi. You can then script the downloading of these assets using the provided URLs.

How often is the data updated?

We typically run furniture catalogue pipelines on a weekly or monthly cadence, which is sufficient to capture new collection launches and price adjustments.

What is the minimum viable engagement?

Our minimum engagement covers the full extraction of the primary public catalogue (seating, lighting, tables) delivered monthly. Contact us for a scoped quote.

Can I request a sample dataset?

Yes. We provide a sample run covering a specific collection (e.g., the Beetle collection) so you can validate the variant schema and dimension normalisation before committing.

$ dataflirt scope --new-project --source=gubi.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off catalogue dump or continuous variant monitoring, we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in furniture

Services

Data Extraction for Every Industry

View All Services →