SYSTEM all green source collegedekho.com queue 12,841 pages p99 latency 182ms dataflirt.com · scraper/collegedekho-com
RUN · 18 active pipelines · collegedekho.com live

Education data,
at warehouse scale.

We extract college directories, programme fees, entrance cut-offs, placement stats, and student reviews from CollegeDekho. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Colleges extracted
43.8K /run
Courses mapped
284K /24h
Review records
1.2M /run
Active pipelines
18
Uptime
99.94%
Data Dictionary

Every field we extract from collegedekho.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for College Profiles objects from collegedekho.com. All fields typed and schema-versioned.

college_idcollege_namelocation_citylocation_stateestablished_yearownership_typeaffiliationapprovalscampus_size_acrestotal_facultyfacilitiesgallery_urlslatitudelongitudepage_url
college_profiles
● 200 OK
"college_id": "CD-8492",
"college_name": "RV College of Engineering",
"location_city": "Bangalore",
"location_state": "Karnataka",
"established_year": 1963,
"ownership_type": "Private",
"affiliation": "Visvesvaraya Technological University",
"approvals": "['AICTE', 'NBA']"
# college_idcollege_namelocation_citylocation_stateestablished_yearownership_type
1
2
3

Complete list of extractable fields for Courses & Fees objects from collegedekho.com. All fields typed and schema-versioned.

college_idcourse_namedegree_levelstreamduration_yearsstudy_modetotal_fees_inrfirst_year_fees_inreligibility_criteriaaccepted_examstotal_seatssyllabus_url
courses_& fees
● 200 OK
"college_id": "CD-8492",
"course_name": "B.E. in Computer Science and Engineering",
"degree_level": "UG",
"stream": "Engineering",
"duration_years": 4,
"total_fees_inr": 386000,
"first_year_fees_inr": 96500,
"accepted_exams": "['KCET', 'COMEDK UGET']"
# college_idcourse_namedegree_levelstreamduration_yearsstudy_mode
1
2
3

Complete list of extractable fields for Cut-offs & Exams objects from collegedekho.com. All fields typed and schema-versioned.

college_idcourse_nameexam_nameexam_yearcounselling_roundcategoryquotaopening_rankclosing_rankclosing_scorecut_off_url
cut-offs_& exams
● 200 OK
"college_id": "CD-8492",
"course_name": "B.E. in Computer Science and Engineering",
"exam_name": "KCET",
"exam_year": 2024,
"counselling_round": "Round 1",
"category": "General",
"opening_rank": 45,
"closing_rank": 158
# college_idcourse_nameexam_nameexam_yearcounselling_roundcategory
1
2
3

Complete list of extractable fields for Placements objects from collegedekho.com. All fields typed and schema-versioned.

college_idplacement_yearhighest_package_inraverage_package_inrmedian_package_inrplacement_percentagetotal_recruiterstop_recruiterstotal_students_placedplacement_report_url
placements
● 200 OK
"college_id": "CD-8492",
"placement_year": 2023,
"highest_package_inr": 6200000,
"average_package_inr": 1450000,
"placement_percentage": 94.5,
"total_recruiters": 284,
"top_recruiters": "['Microsoft', 'Amazon', 'Cisco', 'Goldman Sachs']"
# college_idplacement_yearhighest_package_inraverage_package_inrmedian_package_inrplacement_percentage
1
2
3

Complete list of extractable fields for Student Reviews objects from collegedekho.com. All fields typed and schema-versioned.

review_idcollege_idreviewer_namecourse_enrolledgraduation_yearoverall_ratingplacement_ratinginfrastructure_ratingfaculty_ratinghostel_ratingreview_titlereview_textreview_date
student_reviews
● 200 OK
"review_id": "REV-928411",
"college_id": "CD-8492",
"course_enrolled": "B.E. in Computer Science",
"overall_rating": 4.5,
"placement_rating": 4.8,
"faculty_rating": 4.2,
"review_title": "Excellent placements but strict academics",
"review_date": "2025-08-14"
# review_idcollege_idreviewer_namecourse_enrolledgraduation_yearoverall_rating
1
2
3

Capabilities

Extract the complete Indian higher education catalogue

CollegeDekho structures a highly fragmented sector. Our pipeline extracts every nested table, dynamic filter, and paginated review corpus — normalising the data for immediate downstream analysis.

College Profiles & Infrastructure

Extract affiliations, approvals (AICTE, UGC), campus size, faculty counts, and facility lists across 43,000+ institutions.

Programme & Fee Structures

Map every course variant, duration, study mode, and detailed fee breakdown (first year vs total) across all streams.

Historical Cut-off Data

Capture opening and closing ranks across multiple exams (JEE, NEET, CAT, KCET), categorised by quota and counselling round.

Placement Intelligence

Extract highest, average, and median salary packages, along with recruiter lists and placement percentages per academic year.

Student Review Corpus

Paginate through thousands of student reviews, capturing granular ratings for placements, infrastructure, and faculty.

Admission Timelines

Track application start dates, exam schedules, and counselling deadlines for upcoming academic sessions.

Hostel & Accommodation

Extract hostel availability, gender-specific capacity, and annual fee structures where listed.

Dynamic Filter Extraction

Our Playwright infrastructure navigates complex state/city/stream filters to ensure zero data omission across the directory.

Admission Cycle Updates

Configure high-frequency pipelines during peak admission seasons to capture real-time cut-off and seat availability updates.

// engagement pipeline

From target parameters to warehouse records

Brief in. Clean data out.

Define Scope
d 0

Specify target streams (e.g., Engineering, Medical), states, exams, or specific college parameters. We design the schema.

Pipeline Build
d 2–4

We configure Scrapy and Playwright to handle CollegeDekho's dynamic DOM, accordions, and API endpoints.

Validation & QA
d 4–6

Schema validation, null-rate checks on critical fields like fees and cut-offs, and sample deliveries.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on your required cadence.

Under the hood

Handling the complexity of EdTech aggregators

CollegeDekho relies heavily on client-side rendering and nested accordions to display dense educational data. Here is how we extract it reliably.

pipeline-monitor · collegedekho.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Dynamic Tables
Parsing nested fee and cut-off accordions

Fee structures and cut-off ranks are often hidden behind JavaScript accordions and dynamic tabs. We use Playwright to simulate user interactions, expanding all nodes before parsing the DOM to ensure complete data extraction.

API Interception
Direct extraction from backend endpoints

Where CollegeDekho uses infinite scroll for reviews and college lists, our pipeline intercepts the underlying XHR/Fetch requests. This bypasses the DOM entirely, extracting clean JSON payloads for higher throughput and reliability.

Schema Normalisation
Standardising unstructured text

College descriptions and eligibility criteria are often free-text blocks. We apply post-processing regex and NLP to extract structured variables (e.g., extracting '50% in 10+2' into a clean integer field) before delivery.

Bot Mitigation
Residential proxies for rate limits

To prevent IP bans during high-volume directory crawls, we route requests through Indian residential proxy pools, rotating IPs per request and matching browser TLS fingerprints to appear as legitimate organic traffic.

Change Detection
Tracking fee and rank updates

During admission seasons, data changes daily. We maintain a hash index of last-seen values per college. Subsequent runs only push diffs, alerting you to updated fee structures or revised cut-off lists.

Applications

Who uses CollegeDekho data

Teams across industries use collegedekho.com data to build competitive products and smarter operations.

01
EdTech Platforms & Aggregators

Competitor platforms use directory data to backfill their own databases, ensuring parity in college coverage and fee accuracy.

02
Student Counselling Firms

Admissions consultants query historical cut-off data and placement records to build predictive admission models for students.

03
Financial Institutions

Banks and NBFCs use fee structures and placement metrics to underwrite education loans and determine college tier classifications.

04
Market Research

Analysts track the proliferation of specific courses (e.g., AI/ML specialisations) and fee inflation trends across private universities.

05
University Benchmarking

Institutions monitor competitor fee structures, facility offerings, and student review sentiment to optimise their own positioning.

06
Lead Generation

Service providers targeting educational institutions extract college contact details, affiliations, and faculty counts to build targeted outreach lists.

Why DataFlirt

"CollegeDekho aggregates the most fragmented sector in India — higher education — but standardising that data for downstream analysis requires dedicated infrastructure."

Most teams underestimate the complexity of extracting Indian education data: inconsistent fee tables, nested cut-off accordions, and dynamic filter states. DataFlirt manages the residential proxies, JavaScript rendering, and schema normalisation so your engineers can focus on product development and analysis.

Technical Spec

CollegeDekho scraper — technical specifications

Everything supported by our collegedekho.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Playwright integration for dynamic accordions and filter states
Supported
API interception
Direct extraction from background XHR requests for reviews and lists
Supported
Residential proxies
Indian ISP-grade proxies to bypass rate limiting
Supported
Nested table extraction
Structured parsing of complex fee and cut-off matrices
Supported
Historical cut-offs
Extraction of previous years' ranks where available on the portal
Supported
Review pagination
Full extraction of the student review corpus per college
Supported
OTP-gated counselling data
Personalised admission prediction features requiring user phone verification
Partial
Internal application forms
Common Application Form (CAF) backend logic and user application states
Partial
Infrastructure

Infrastructure powering the extraction

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy manages crawl logic and deduplication, while Playwright handles the client-side rendering required for CollegeDekho's dynamic components.

Proxy Infrastructure

We utilise Indian residential proxies to mimic legitimate domestic traffic, preventing blockades during large-scale directory crawls.

Cloud-Native Delivery

Pipelines are orchestrated via Apache Airflow on Kubernetes, ensuring reliable scheduling and automated delivery to your data warehouse.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Nested schema ideal for complex college profiles and course lists
CSV
Flat files for immediate analysis in Excel or Pandas
XLS
Legacy spreadsheet format for business stakeholders
Parquet
Columnar storage optimised for BigQuery and Snowflake
AWS S3
Direct upload to your cloud storage buckets
Webhook
HTTP POST for real-time data ingestion
API
REST endpoints to query your extracted datasets
PostgreSQL
Direct database upserts with conflict resolution
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About collegedekho.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping CollegeDekho legal?

Scraping publicly available directory information, fee structures, and reviews is generally permissible. DataFlirt extracts only public, non-authenticated data. We do not bypass OTP walls, extract student PII, or access internal application dashboards.

How do you handle CollegeDekho's dynamic tables?

We use Playwright to execute JavaScript, simulate clicks on accordions, and select different exam/category tabs before parsing the DOM. This ensures we capture all permutations of cut-offs and fees.

Can you track fee changes over time?

Yes. We maintain historical snapshots. By running pipelines at regular intervals (e.g., monthly or quarterly), we calculate diffs and flag fee increases or new course additions.

How fresh is the data during admission season?

We can configure high-frequency pipelines during peak periods (May to August) to scrape cut-off updates and seat availability daily.

What is the minimum viable engagement?

Engagements typically start with a defined subset (e.g., all Engineering colleges in South India) or a full directory baseline extract. Contact us to scope your specific volume requirements.

Can I get a sample dataset?

Yes. We provide a sample extract of 100-200 colleges, including nested course and placement data, during the scoping phase to validate schema fit.

$ dataflirt scope --new-project --source=collegedekho.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. From targeted state-level college lists to a full national directory extraction — we build and operate the pipeline. Tell us your requirements.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in education and courses

Services

Data Extraction for Every Industry

View All Services →