SYSTEM operational source igrmaharashtra.gov.in queue 14,892 queries p99 latency 3,450ms dataflirt.com · scraper/igrmaharashtra-gov
RUN : 14 active pipelines : igrmaharashtra.gov.in live

Maharashtra property records,
at warehouse scale.

We extract Index II records, transaction values, buyer and seller details, and stamp duty metrics from the IGR Maharashtra portal. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Index II records
18,450 /day
SRO offices tracked
518
Districts covered
36
Active pipelines
14
Uptime
98.45%
Data Dictionary

Every field we extract from igrmaharashtra.gov.in

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Index II Summary objects from igrmaharashtra.gov.in. All fields typed and schema-versioned.

document_numberregistration_yearsro_namedistricttalukavillageproperty_typeareaconsideration_amountmarket_valuestamp_dutyregistration_feeexecution_dateregistration_date
index_ii summary
● 200 OK
"document_number": "4892",
"registration_year": "2023",
"sro_name": "Andheri 4",
"district": "Mumbai Suburban",
"taluka": "Andheri",
"consideration_amount": 14500000.0,
"market_value": 14250000.0,
"registration_date": "2023-11-14"
# document_numberregistration_yearsro_namedistricttalukavillage
1
2
3

Complete list of extractable fields for Property Details objects from igrmaharashtra.gov.in. All fields typed and schema-versioned.

property_idsurvey_numbercts_numberplot_numberflat_numberbuilding_namefloorproject_namepin_codeboundaries
property_details
● 200 OK
"cts_number": "145/B",
"flat_number": "402",
"building_name": "Siddhivinayak Heights",
"floor": "4th Floor",
"pin_code": "400053",
"boundaries": "North: Road, South: Plot 14, East: CTS 146"
# property_idsurvey_numbercts_numberplot_numberflat_numberbuilding_name
1
2
3

Complete list of extractable fields for Party Details objects from igrmaharashtra.gov.in. All fields typed and schema-versioned.

party_typerolenameagepan_numberaddressverification_statusexecution_status
party_details
● 200 OK
"party_type": "Individual",
"role": "Purchaser",
"name": "Rajesh Kumar Sharma",
"age": 45,
"pan_number": "ABCDE1234F",
"address": "B-401, Vasant Vihar, Thane West"
# party_typerolenameagepan_numberaddress
1
2
3

Complete list of extractable fields for Financials objects from igrmaharashtra.gov.in. All fields typed and schema-versioned.

consideration_valuemarket_valuestamp_duty_paidregistration_fee_paiddocument_handling_chargespenaltytotal_paidpayment_crnpayment_date
financials
● 200 OK
"consideration_value": 14500000.0,
"market_value": 14250000.0,
"stamp_duty_paid": 870000.0,
"registration_fee_paid": 30000.0,
"document_handling_charges": 600.0,
"total_paid": 900600.0
# consideration_valuemarket_valuestamp_duty_paidregistration_fee_paiddocument_handling_chargespenalty
1
2
3

Complete list of extractable fields for Document Metadata objects from igrmaharashtra.gov.in. All fields typed and schema-versioned.

document_typesub_registrar_officevolume_numberpage_numbermicroform_numberstatusscanned_copy_availableremarksscraped_at
document_metadata
● 200 OK
"document_type": "Agreement to Sale",
"sub_registrar_office": "Bandra 2",
"volume_number": "4512",
"page_number": "12-45",
"status": "Registered",
"scraped_at": "2023-11-15T08:12:44Z"
# document_typesub_registrar_officevolume_numberpage_numbermicroform_numberstatus
1
2
3

Capabilities

Structured property data, extracted reliably

The IGR Maharashtra portal is notoriously difficult to scrape. We handle the CAPTCHAs, ASP.NET session states, and server timeouts to deliver clean, normalised Index II data.

Full Index II Extraction

Extract all fields from the Index II format including property descriptions, party details, and financial consideration values.

Automated Captcha Solving

Integration with CapSolver and 2Captcha to bypass the continuous image captchas required for every search query.

Search Parameter Iteration

Automated traversal of all districts, talukas, and villages to ensure complete geographical coverage.

Session State Management

Handling of heavy ASP.NET ViewState payloads and aggressive session timeouts to maintain uninterrupted extraction.

Marathi Normalisation

Processing and transliterating dual-language outputs to provide clean, queryable English strings.

Historical Backfills

Extract historical transaction records from available online archives dating back to 2002.

Daily Incremental Updates

Run daily pipelines to capture newly registered documents as soon as they appear in the public search portal.

SRO Office Mapping

Map transactions to specific Sub-Registrar Offices for accurate jurisdictional analysis.

Data Cleansing

Standardise addresses, correct malformed CTS numbers, and convert string currency values into strict numeric types.

// engagement pipeline

From search parameters to warehouse records

Brief in. Clean data out.

Define Scope
d 0

Specify target districts, talukas, villages, or specific date ranges for extraction.

Pipeline Build
d 2–4

We configure the Scrapy spiders, set up captcha solving workflows, and manage ASP.NET session states.

Validation & QA
d 4–6

Schema validation, null-rate checks on critical fields like consideration value, and data type enforcement.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on an agreed schedule.

Under the hood

Overcoming government portal infrastructure

Extracting data from IGR Maharashtra requires navigating outdated web architectures and aggressive rate limiting. Here is how we maintain pipeline stability.

pipeline-monitor · igrmaharashtra.gov.in · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Captcha walls
Continuous automated solving

The portal requires a captcha solution for every single search query. We route these challenges through automated solving APIs with fallback mechanisms to ensure the crawler never stalls.

Unstable infrastructure
Handling 502s and 503s

Government servers frequently drop connections or return gateway errors during peak hours. Our orchestration layer uses exponential backoff and intelligent retry logic to recover missing data during off-peak windows.

ASP.NET ViewState
Managing hidden form fields

The site relies heavily on ASP.NET ViewState for pagination and form submission. We parse and maintain these hidden payload states across requests to prevent session invalidation.

Data normalisation
Cleaning dirty text

User-entered data on the portal contains typos, varied date formats, and mixed Marathi/English text. We apply strict regex patterns and transliteration libraries to produce a normalised dataset.

Pagination limits
Bypassing query caps

Search results are often capped at a maximum number of records. We dynamically slice search parameters by narrower date ranges or specific property numbers to ensure total extraction without truncation.

Applications

Who uses IGR Maharashtra data

Teams across industries use igrmaharashtra.gov.in data to build competitive products and smarter operations.

01
PropTech Valuations

Automated Valuation Models require ground-truth transaction data to accurately price properties in specific micro-markets.

02
Title Verification

Legal firms and compliance teams automate initial title searches to verify ownership history and encumbrances.

03
Market Analytics

Real estate developers track absorption rates, pricing trends, and competitor project sales velocity.

04
Bank Underwriting

Financial institutions cross-reference declared property values against actual registered consideration amounts.

05
Lead Generation

Interior designers, brokers, and home service providers identify recent property buyers for targeted marketing.

06
Urban Planning

Researchers and policy makers analyse spatial transaction density and stamp duty revenue distribution.

Why DataFlirt

"The Maharashtra IGR portal holds the ground truth for real estate pricing, but its architecture actively resists automated extraction."

Extracting Index II records requires navigating aggressive session timeouts, complex ASP.NET state management, and constant CAPTCHA interruptions. DataFlirt handles the extraction infrastructure, delivering clean transaction data so your team can focus on market analysis rather than maintaining fragile web scrapers.

Technical Spec

IGR Maharashtra scraper technical capabilities

Everything supported by our igrmaharashtra.gov.in scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

Index II text extraction
Full parsing of HTML tables into structured JSON fields
Supported
Captcha bypass
Automated routing to solving services for search form access
Supported
District and Taluka iteration
Automated traversal of geographical dropdowns
Supported
Marathi to English transliteration
Basic normalisation of dual-language text fields
Supported
Incremental daily scraping
Capture only newly registered documents based on execution date
Supported
Residential proxy rotation
ISP-grade Indian IPs to avoid geographic blocking
Supported
Historical records pre-2002
Data prior to computerisation is not available on the public search portal
Partial
Certified copy PDF downloads
Requires authenticated login and individual fee payment per document
Partial
Infrastructure

Infrastructure powering the extraction

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy orchestrates the crawl logic while Playwright handles the complex JavaScript rendering and ASP.NET form submissions.

Captcha Solving Infrastructure

Asynchronous integration with multiple captcha solving APIs ensures high throughput even when the portal increases challenge difficulty.

Cloud-Native Orchestration

Pipelines are scheduled via Apache Airflow and executed on containerised infrastructure, allowing us to scale workers during off-peak hours.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited JSON for easy parsing
CSV
Flat files suitable for immediate spreadsheet analysis
XLS
Excel format for business teams
Parquet
Columnar storage optimised for analytical queries
AWS S3
Direct delivery to your cloud storage buckets
Webhook
HTTP POST delivery for immediate downstream processing
API
REST endpoint to query extracted records
PostgreSQL
Direct database insertion with upsert logic
BigQuery
Native streaming into Google Cloud data warehouses
Snowflake
Automated stage and copy workflows
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About igrmaharashtra.gov.in scraping, legality, and pipeline operations.

Ask us directly →
Is the data extracted from the public portal?

Yes. We exclusively extract data from the Free Search IGR Service (e-Search) which is publicly accessible. We do not bypass authentication walls or extract restricted internal documents.

How do you handle the constant captchas?

We route captcha images to automated solving services like CapSolver and 2Captcha. Our infrastructure handles the latency and retries automatically.

Can you extract historical data from 10 years ago?

Yes, provided the data exists in the online e-Search database. The portal generally holds digitised records from 2002 onwards for Mumbai and surrounding areas, though coverage varies by district.

How frequently can you deliver updates?

We typically configure pipelines to run daily or weekly, capturing newly indexed records. Real-time extraction is limited by the portal's own updating schedule.

Do you provide the actual PDF documents?

No. Downloading certified PDF copies requires a logged-in session and payment of fees per document. We extract the structured Index II text data visible on the search results pages.

How do you handle Marathi text?

We extract the text exactly as it appears on the portal. We can apply basic transliteration libraries to normalise names and addresses into English upon request.

What happens when the government site goes down?

Our pipelines detect 502/503 errors and session timeouts. They automatically pause and resume extraction using exponential backoff, ensuring complete data capture once the portal recovers.

$ dataflirt scope --new-project --source=igrmaharashtra.gov.in ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Stop fighting captchas and session timeouts. Tell us which districts and date ranges you need, and we will deliver clean Index II records directly to your database.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in real estate

Services

Data Extraction for Every Industry

View All Services →