We extract cosmetics listings, shade matrices, pricing signals, and brand catalogues from NewU. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Product Listings objects from newu.in. All fields typed and schema-versioned.
"sku": "NU-MUP-890103086", "title": "JaQuline USA Pro Stroke Liquid Eyeliner", "brand": "JaQuline USA", "category": "Makeup", "price": 299.0, "list_price": 399.0, "discount_pct": 25, "in_stock": true, "pack_size": "4.5 ml"
| # | sku | title | brand | category | sub_category | price |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Shade Variants objects from newu.in. All fields typed and schema-versioned.
"sku": "NU-LIP-890101112", "parent_sku": "NU-LIP-BASE-01", "shade_name": "Crimson Red 04", "hex_code": "#8A0303", "price": 450.0, "in_stock": true, "swatch_url": "https://newu.in/media/swatches/crimson_04.jpg"
| # | sku | parent_sku | shade_name | hex_code | price | in_stock |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Pricing & Offers objects from newu.in. All fields typed and schema-versioned.
"sku": "NU-SKN-890456123", "price": 599.0, "list_price": 799.0, "discount_abs": 200.0, "discount_pct": 25, "combo_offer": "Buy 2 Get 1 Free", "deal_badge": "Bestseller", "timestamp": "2026-05-12T10:15:00Z"
| # | sku | price | list_price | discount_abs | discount_pct | combo_offer |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Reviews objects from newu.in. All fields typed and schema-versioned.
"review_id": "REV-99812", "sku": "NU-MUP-890103086", "reviewer_name": "Priya S.", "rating": 5, "review_title": "Smudge proof and dark", "review_date": "2026-04-20", "verified": true
| # | review_id | sku | reviewer_name | rating | review_title | review_body |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Store Locator objects from newu.in. All fields typed and schema-versioned.
"store_id": "ST-BLR-04", "name": "NewU - Indiranagar", "city": "Bengaluru", "state": "Karnataka", "pincode": "560038", "phone": "+91-80-41123456", "latitude": 12.9784, "longitude": 77.6408
| # | store_id | name | address | city | state | pincode |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our NewU scraper captures the full complexity of beauty retail: parent-child shade matrices, nested ingredient lists, active combo offers, and inventory levels across thousands of SKUs.
Extract SKU, title, brand, ingredients, and categories across makeup, skincare, fragrance, and personal care.
Map parent products to child shade variants, capturing hex codes, swatch images, and variant-specific pricing.
Capture MRP, selling price, discount percentages, and active combo offers timestamped per run.
Track stock availability across product variations to identify supply chain gaps and restock patterns.
Structure raw ingredient text into queryable fields for formulation analysis and compliance checking.
Filter and extract specific brand portfolios like JaQuline USA, Lakme, or Maybelline for targeted audits.
Extract offline retail footprints including address, coordinates, and contact details for all NewU physical stores.
Collect customer ratings, review text, and verification status across the product catalogue.
Run continuous extraction at daily or weekly intervals with hash-based change detection.
Brief in. Clean data out.
Provide target categories, specific brands, or search terms. We design the extraction schema together.
We configure Scrapy and Playwright crawlers, proxy rotation, and session management for newu.in.
Schema validation, null-rate checks, price-outlier detection, and sample shade matrices before full launch.
JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
Extracting cosmetics data requires specific handling for dynamic variants and unstructured text. We manage the infrastructure so you get clean tables.
Beauty sites load shade matrices via asynchronous JavaScript. We deploy Playwright to hydrate the DOM and extract all colour variants, hex codes, and swatches without missing hidden SKUs.
Cosmetics data is notoriously unstructured. We normalise ingredient lists, volume metrics, and shade names into strict data types, converting raw HTML into queryable JSON arrays.
Retail sites employ standard rate limiting and bot protection. We use residential Indian proxies to distribute request volume, maintain healthy subnets, and avoid IP bans during catalogue sweeps.
We maintain a hash index of last-seen values per field. Subsequent runs only push diffs for price updates or stock changes, reducing compute cost and downstream processing load.
Every pipeline emits structured logs to our Grafana dashboards. We monitor null rates on critical fields like price and stock status, responding to layout changes before they affect your data.
Brands track NewU pricing against other platforms like Nykaa and Purplle to maintain parity.
Cosmetics manufacturers audit retail prices to ensure compliance with Minimum Advertised Price agreements.
Emerging D2C brands analyse category pricing tiers, combo offers, and discount frequency to position their own products.
Market researchers track shade popularity and new ingredient introductions across the catalogue.
Supply chain analysts monitor out-of-stock rates across specific brands or categories to identify retail demand spikes.
Conglomerates monitor brand visibility, product descriptions, and image accuracy across their retail partners.
"Beauty retail data requires deep variant mapping. A lipstick isn't one product; it's twenty distinct SKUs with individual prices, stock levels, and hex codes."
Most scraping tools fail on cosmetics sites because they cannot handle dynamic shade selectors or unstructured ingredient text. DataFlirt builds pipelines specifically designed to expand parent-child matrices and normalise beauty attributes into strict schemas. You receive clean data, ready for immediate analysis.
Everything supported by our newu.in scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration and deduplication. Playwright executes JavaScript to render shade selectors and dynamic combo offers.
We maintain pools of residential ISP proxies across India. Rotation happens per-request to bypass basic retail firewall rules.
Pipelines run on AWS infrastructure. Airflow handles scheduling and dependency management. All state is stored in managed Postgres.
Data delivered to where your team already works — no new tooling required.
About newu.in scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available product, pricing, and store data is generally permissible. DataFlirt targets only public retail information. We do not extract personal user data or circumvent authentication walls.
We map parent product URLs to all available child variants. The output schema includes the parent SKU alongside individual child SKUs, capturing variant-specific prices, hex codes, and stock levels.
Yes. The pipeline records binary stock status for every variant. Scheduled runs will track when an item drops out of stock and when it returns.
We configure pipelines to run daily or weekly based on your requirements. The data reflects the exact state of newu.in at the timestamp of extraction.
Engagements typically start at full-site category sweeps (e.g., all Makeup or all Skincare). Contact us with your target categories for a scoped quote.
Yes. We provide a sample run of up to 500 SKUs during the scoping phase so you can validate field completeness and shade mapping logic.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off catalogue dump or continuous price monitoring across all beauty categories, we build and operate the pipeline. Tell us what you need.