We extract technical specifications, photometric profiles, 3D CAD models, and designer collections from nemolighting.com. Delivered as clean JSON, CSV, or Parquet to your infrastructure.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Product Master objects from nemolighting.com. All fields typed and schema-versioned.
"sku": "NEM-CROWN-MAJOR", "name": "Crown Major", "collection": "Crown", "designer": "Jehs + Laub", "category": "Suspension", "materials": "Die-cast aluminium"
| # | sku | name | collection | designer | category | sub_category |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Technical Specs objects from nemolighting.com. All fields typed and schema-versioned.
"sku": "NEM-CROWN-MAJOR", "light_source": "halopin G9 QT-14", "power_wattage": "30x25W", "emission": "diffused", "switching": "dimmable", "tension": "110/230V", "ip_rating": "IP20"
| # | sku | light_source | power_wattage | emission | switching | tension |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Photometric & CAD objects from nemolighting.com. All fields typed and schema-versioned.
"sku": "NEM-CROWN-MAJOR", "ies_file_url": "https://nemolighting.com/assets/ies/crown_major.ies", "ldt_file_url": "https://nemolighting.com/assets/ldt/crown_major.ldt", "revit_file_url": "https://nemolighting.com/assets/bim/crown_major.rfa", "dwg_file_url": "https://nemolighting.com/assets/cad/crown_major.dwg", "pdf_spec_sheet": "https://nemolighting.com/assets/pdf/crown_major_tech.pdf"
| # | sku | ies_file_url | ldt_file_url | revit_file_url | dwg_file_url | 3dm_file_url |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Designer Data objects from nemolighting.com. All fields typed and schema-versioned.
"designer_id": "DES-042", "name": "Le Corbusier", "bio": "Charles-Edouard Jeanneret, known as Le Corbusier...", "collections_designed": "['Lampe de Marseille', 'Projecteur']", "nationality": "Swiss-French", "page_url": "https://nemolighting.com/designers/le-corbusier"
| # | designer_id | name | bio | image_url | collections_designed | active_years |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Variants & Finishes objects from nemolighting.com. All fields typed and schema-versioned.
"sku": "NEM-CROWN-MAJOR", "variant_id": "CRO-HLW-31", "finish_name": "Hand polished aluminium", "finish_code": "HLW", "material_type": "Aluminium", "availability_status": "In Production"
| # | sku | variant_id | finish_name | finish_code | ral_colour | material_type |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Architectural lighting requires deep structural extraction. We capture photometric profiles, 3D asset links, and precise technical specifications across the entire Nemo catalogue.
Capture both architectural and decorative lighting lines with complete category and sub-category hierarchies.
Extract precise metrics including CRI, luminous flux, colour temperature (Kelvin), IP ratings, and wattage.
Resolve and extract direct download URLs for IES and LDT photometric data files used in DIALux and Relux.
Capture links for BIM objects, Revit families (RFA), AutoCAD (DWG), and SketchUp files for every SKU.
Extract URLs for technical specification sheets and installation manuals directly from product pages.
Map complex variant matrices including body finishes, reflector colours, and material types to specific SKUs.
Extract parallel datasets for Italian and English descriptions, preserving technical terminology.
Link products to specific designers and architectural studios, maintaining relational integrity.
Monitor new product launches, finish additions, and technical specification revisions on your defined cadence.
Brief in. Clean data out.
Specify the collections, product types, or asset formats (BIM, IES) you need extracted from the catalogue.
We configure Scrapy spiders to traverse product hierarchies and resolve CDN links for heavy technical assets.
Schema validation ensures all technical fields like CRI and luminous flux are correctly typed and normalised.
Clean JSON or Parquet pushed to your S3 bucket or Snowflake instance, ready for your specifiers or engineers.
Lighting catalogues are asset-heavy. Here is how we extract engineering data reliably.
Photometric files and CAD models are often loaded dynamically via JavaScript or stored on separate asset domains. Our Playwright integration intercepts network requests to capture the true download URLs for all technical files.
A single lamp design might have 15 variants combining body finish, reflector type, and light source. We iterate through the DOM state to map exact technical specifications to the correct variant ID.
When critical data like beam angles or driver specifications only exist inside downloadable PDFs, we can route those documents through our OCR and parsing pipelines to extract structured text.
We extract both English and Italian variants of the site, ensuring technical terms like 'luminous flux' and 'flusso luminoso' map to a single unified database column.
Lighting manufacturers frequently update LED efficiency metrics. We hash technical specification blocks to detect when a product's lumens or CRI changes, delivering only the updated records.
Architects and specifiers aggregate Revit families and CAD models to build comprehensive 3D asset libraries for interior planning.
Engineering platforms ingest IES and LDT files directly to update their photometric databases for DIALux and Relux simulations.
Lighting manufacturers track Nemo's material choices, LED efficiency metrics, and designer collaborations to benchmark their own product lines.
Digital design tools import high-resolution imagery, dimensions, and finish options to populate virtual staging environments.
Authorised dealers synchronise product descriptions, technical specs, and imagery to maintain accurate eCommerce storefronts.
Analysts track the adoption of new materials, energy classes, and control protocols (like DALI) across high-end lighting collections.
"Architectural lighting data is heavily fragmented across PDFs, IES files, and CAD models. Extracting it requires more than simple HTML parsing; it requires structural awareness of engineering assets."
Most teams struggle with lighting catalogues because critical data lives inside downloadable assets rather than the DOM. DataFlirt extracts, parses, and normalises technical specifications, photometric profiles, and 3D models into a unified relational schema.
Everything supported by our nemolighting.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy manages the crawl state across the catalogue hierarchy, while Playwright renders dynamic variant selectors to expose precise technical metrics.
Network interception captures underlying CDN links for heavy 3D models and photometric files without requiring full downloads during the crawl phase.
Pipelines run on AWS infrastructure managed by Kubernetes. Airflow handles scheduling, ensuring your technical database stays synchronised with the live catalogue.
Data delivered to where your team already works — no new tooling required.
About nemolighting.com scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available product specifications and marketing data is generally permissible under standard web scraping legal frameworks. DataFlirt targets only public pages and does not bypass authentication systems to extract gated trade pricing or proprietary dealer data.
Yes. We extract the direct download URLs for all available photometric profiles associated with a product, allowing you to ingest them programmatically into your simulation software.
Our standard pipeline extracts the direct URL to the PDF asset. If you require text extraction from within the PDFs, we can route those files through a secondary OCR and parsing pipeline as a custom requirement.
We can configure the crawler to target specific language locales (e.g., English and Italian) and deliver separate or merged datasets, ensuring technical terminology is preserved accurately.
Yes. We iterate through the available product configurations to map specific finishes, reflector types, and material combinations to their respective technical specifications and variant IDs.
Given the update frequency of high-end lighting catalogues, most clients opt for weekly or monthly synchronisation runs. We can accommodate any schedule required by your engineering or design teams.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a complete extraction of photometric files or a continuous sync of technical specifications, we build and operate the pipeline. Tell us what you need.