SYSTEM all green source newu.in queue 12,408 URLs p99 latency 215ms dataflirt.com · scraper/newu-in
RUN * 14 active pipelines * newu.in live

NewU retail data,
normalised for analysis.

We extract cosmetics listings, shade matrices, pricing signals, and brand catalogues from NewU. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Products tracked
18,492 /day
Price updates
42,105 /week
Shade variants
56,211 /run
Active pipelines
14
Uptime
99.98%
Data Dictionary

Every field we extract from newu.in

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Product Listings objects from newu.in. All fields typed and schema-versioned.

skutitlebrandcategorysub_categorypricelist_pricediscount_pctin_stockdescriptioningredientshow_to_usepack_sizeurl
product_listings
● 200 OK
"sku": "NU-MUP-890103086",
"title": "JaQuline USA Pro Stroke Liquid Eyeliner",
"brand": "JaQuline USA",
"category": "Makeup",
"price": 299.0,
"list_price": 399.0,
"discount_pct": 25,
"in_stock": true,
"pack_size": "4.5 ml"
# skutitlebrandcategorysub_categoryprice
1
2
3

Complete list of extractable fields for Shade Variants objects from newu.in. All fields typed and schema-versioned.

skuparent_skushade_namehex_codepricein_stockimage_urlswatch_url
shade_variants
● 200 OK
"sku": "NU-LIP-890101112",
"parent_sku": "NU-LIP-BASE-01",
"shade_name": "Crimson Red 04",
"hex_code": "#8A0303",
"price": 450.0,
"in_stock": true,
"swatch_url": "https://newu.in/media/swatches/crimson_04.jpg"
# skuparent_skushade_namehex_codepricein_stock
1
2
3

Complete list of extractable fields for Pricing & Offers objects from newu.in. All fields typed and schema-versioned.

skupricelist_pricediscount_absdiscount_pctcombo_offerdeal_badgetimestamp
pricing_& offers
● 200 OK
"sku": "NU-SKN-890456123",
"price": 599.0,
"list_price": 799.0,
"discount_abs": 200.0,
"discount_pct": 25,
"combo_offer": "Buy 2 Get 1 Free",
"deal_badge": "Bestseller",
"timestamp": "2026-05-12T10:15:00Z"
# skupricelist_pricediscount_absdiscount_pctcombo_offer
1
2
3

Complete list of extractable fields for Reviews objects from newu.in. All fields typed and schema-versioned.

review_idskureviewer_nameratingreview_titlereview_bodyreview_dateverified
reviews
● 200 OK
"review_id": "REV-99812",
"sku": "NU-MUP-890103086",
"reviewer_name": "Priya S.",
"rating": 5,
"review_title": "Smudge proof and dark",
"review_date": "2026-04-20",
"verified": true
# review_idskureviewer_nameratingreview_titlereview_body
1
2
3

Complete list of extractable fields for Store Locator objects from newu.in. All fields typed and schema-versioned.

store_idnameaddresscitystatepincodephonelatitudelongitudetimings
store_locator
● 200 OK
"store_id": "ST-BLR-04",
"name": "NewU - Indiranagar",
"city": "Bengaluru",
"state": "Karnataka",
"pincode": "560038",
"phone": "+91-80-41123456",
"latitude": 12.9784,
"longitude": 77.6408
# store_idnameaddresscitystatepincode
1
2
3

Capabilities

Cosmetics data extraction built for scale

Our NewU scraper captures the full complexity of beauty retail: parent-child shade matrices, nested ingredient lists, active combo offers, and inventory levels across thousands of SKUs.

Full Catalogue Extraction

Extract SKU, title, brand, ingredients, and categories across makeup, skincare, fragrance, and personal care.

Shade Matrix Mapping

Map parent products to child shade variants, capturing hex codes, swatch images, and variant-specific pricing.

Dynamic Price Tracking

Capture MRP, selling price, discount percentages, and active combo offers timestamped per run.

Inventory Monitoring

Track stock availability across product variations to identify supply chain gaps and restock patterns.

Ingredient List Parsing

Structure raw ingredient text into queryable fields for formulation analysis and compliance checking.

Brand Intelligence

Filter and extract specific brand portfolios like JaQuline USA, Lakme, or Maybelline for targeted audits.

Store Locator Scraping

Extract offline retail footprints including address, coordinates, and contact details for all NewU physical stores.

Review Aggregation

Collect customer ratings, review text, and verification status across the product catalogue.

Scheduled Pipelines

Run continuous extraction at daily or weekly intervals with hash-based change detection.

// engagement pipeline

From URL list to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target categories, specific brands, or search terms. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy and Playwright crawlers, proxy rotation, and session management for newu.in.

Validation & QA
d 4–6

Schema validation, null-rate checks, price-outlier detection, and sample shade matrices before full launch.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our NewU pipeline handles beauty retail complexity

Extracting cosmetics data requires specific handling for dynamic variants and unstructured text. We manage the infrastructure so you get clean tables.

pipeline-monitor · newu.in · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Dynamic shade selectors
Full Playwright execution for colour variants

Beauty sites load shade matrices via asynchronous JavaScript. We deploy Playwright to hydrate the DOM and extract all colour variants, hex codes, and swatches without missing hidden SKUs.

Schema normalisation
Structuring the unstructured

Cosmetics data is notoriously unstructured. We normalise ingredient lists, volume metrics, and shade names into strict data types, converting raw HTML into queryable JSON arrays.

Anti-bot layer
Residential proxy rotation

Retail sites employ standard rate limiting and bot protection. We use residential Indian proxies to distribute request volume, maintain healthy subnets, and avoid IP bans during catalogue sweeps.

Change detection
Only re-scrape what changes

We maintain a hash index of last-seen values per field. Subsequent runs only push diffs for price updates or stock changes, reducing compute cost and downstream processing load.

Observability
24/7 pipeline health monitoring

Every pipeline emits structured logs to our Grafana dashboards. We monitor null rates on critical fields like price and stock status, responding to layout changes before they affect your data.

Applications

Who uses NewU data

Teams across industries use newu.in data to build competitive products and smarter operations.

01
Price Monitoring

Brands track NewU pricing against other platforms like Nykaa and Purplle to maintain parity.

02
MAP Compliance

Cosmetics manufacturers audit retail prices to ensure compliance with Minimum Advertised Price agreements.

03
Competitor Analysis

Emerging D2C brands analyse category pricing tiers, combo offers, and discount frequency to position their own products.

04
Trend Forecasting

Market researchers track shade popularity and new ingredient introductions across the catalogue.

05
Inventory Tracking

Supply chain analysts monitor out-of-stock rates across specific brands or categories to identify retail demand spikes.

06
Brand Auditing

Conglomerates monitor brand visibility, product descriptions, and image accuracy across their retail partners.

Why DataFlirt

"Beauty retail data requires deep variant mapping. A lipstick isn't one product; it's twenty distinct SKUs with individual prices, stock levels, and hex codes."

Most scraping tools fail on cosmetics sites because they cannot handle dynamic shade selectors or unstructured ingredient text. DataFlirt builds pipelines specifically designed to expand parent-child matrices and normalise beauty attributes into strict schemas. You receive clean data, ready for immediate analysis.

Technical Spec

NewU scraper technical capabilities

Everything supported by our newu.in scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions for asynchronous shade matrices and pricing widgets
Supported
CAPTCHA bypass
Automated 2Captcha + CapSolver integration for rate-limit walls
Supported
Residential proxy rotation
ISP-grade residential IPs from IN pools
Supported
Shade mapping
Parent to child SKU relationships with all colour combinations
Supported
Store locator extraction
Geospatial data and contact info for physical retail locations
Supported
Ingredient parsing
Extraction of raw ingredient text blocks
Supported
Change detection
Hash-based diff logic to emit only changed records
Supported
User purchase history
Requires individual user authentication credentials
Partial
Loyalty points balance
NewU rewards data is gated behind user login walls
Partial
Infrastructure

Infrastructure powering the NewU pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration and deduplication. Playwright executes JavaScript to render shade selectors and dynamic combo offers.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies across India. Rotation happens per-request to bypass basic retail firewall rules.

Cloud-Native Orchestration

Pipelines run on AWS infrastructure. Airflow handles scheduling and dependency management. All state is stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested arrays for complex shade matrices
CSV
Flat file with typed columns for pricing analysis
XLS
Standard Excel format for business teams
Parquet
Columnar format for BigQuery, Snowflake, Athena
AWS S3
Direct bucket delivery compatible with any data lake
Webhook
HTTP POST per record for real-time stock alerts
API
REST endpoints to query your extracted catalogue
PostgreSQL
Direct database upsert with conflict resolution
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About newu.in scraping, legality, and pipeline operations.

Ask us directly →
Is scraping NewU legal?

Scraping publicly available product, pricing, and store data is generally permissible. DataFlirt targets only public retail information. We do not extract personal user data or circumvent authentication walls.

How do you handle shade variations?

We map parent product URLs to all available child variants. The output schema includes the parent SKU alongside individual child SKUs, capturing variant-specific prices, hex codes, and stock levels.

Can you track out-of-stock items?

Yes. The pipeline records binary stock status for every variant. Scheduled runs will track when an item drops out of stock and when it returns.

How fresh is the pricing data?

We configure pipelines to run daily or weekly based on your requirements. The data reflects the exact state of newu.in at the timestamp of extraction.

What is the minimum viable engagement?

Engagements typically start at full-site category sweeps (e.g., all Makeup or all Skincare). Contact us with your target categories for a scoped quote.

Can I request a sample dataset?

Yes. We provide a sample run of up to 500 SKUs during the scoping phase so you can validate field completeness and shade mapping logic.

$ dataflirt scope --new-project --source=newu.in ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off catalogue dump or continuous price monitoring across all beauty categories, we build and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in beauty and skincare

Services

Data Extraction for Every Industry

View All Services →