SYSTEM all green source collegescorecard.ed.gov queue 6,412 institutions p99 latency 218ms dataflirt.com · scraper/collegescorecard-ed.gov
RUN · 14 active pipelines · collegescorecard.ed.gov live

Higher education data,
at warehouse scale.

We extract institutional profiles, financial aid metrics, post-graduation earnings, and admission statistics from the US College Scorecard. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Institutions tracked
6,412
Degree programmes
241K
Earnings records
1.2M /run
Active pipelines
14
Uptime
99.98%
Data Dictionary

Every field we extract from collegescorecard.ed.gov

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Institution Overview objects from collegescorecard.ed.gov. All fields typed and schema-versioned.

unitidopeidinstitution_namecitystatezip_codewebsite_urlinstitution_typelocale_designationstudent_populationhbcu_statusreligious_affiliation
institution_overview
● 200 OK
"unitid": "166027",
"institution_name": "Harvard University",
"city": "Cambridge",
"state": "MA",
"institution_type": "Private nonprofit",
"student_population": 7240,
"hbcu_status": false
# unitidopeidinstitution_namecitystatezip_code
1
2
3

Complete list of extractable fields for Admissions & Test Scores objects from collegescorecard.ed.gov. All fields typed and schema-versioned.

unitidadmission_ratesat_math_25thsat_math_75thsat_read_25thsat_read_75thact_composite_25thact_composite_75thopen_admissions_policyapplication_countadmitted_count
admissions_& test scores
● 200 OK
"unitid": "166027",
"admission_rate": 0.04,
"sat_math_25th": 740,
"sat_math_75th": 800,
"sat_read_25th": 720,
"sat_read_75th": 780,
"open_admissions_policy": false
# unitidadmission_ratesat_math_25thsat_math_75thsat_read_25thsat_read_75th
1
2
3

Complete list of extractable fields for Cost & Financial Aid objects from collegescorecard.ed.gov. All fields typed and schema-versioned.

unitidavg_net_pricetuition_in_statetuition_out_statepct_pell_grantpct_federal_loanavg_grant_aidnet_price_income_0_30knet_price_income_30_48knet_price_income_75_110kcost_of_attendance
cost_& financial aid
● 200 OK
"unitid": "166027",
"avg_net_price": 18030,
"tuition_in_state": 54002,
"tuition_out_state": 54002,
"pct_pell_grant": 0.2,
"pct_federal_loan": 0.03,
"avg_grant_aid": 61000
# unitidavg_net_pricetuition_in_statetuition_out_statepct_pell_grantpct_federal_loan
1
2
3

Complete list of extractable fields for Graduation & Retention objects from collegescorecard.ed.gov. All fields typed and schema-versioned.

unitidretention_rate_ftretention_rate_ptgrad_rate_150_pctgrad_rate_pellgrad_rate_non_pelltransfer_out_ratecompletion_4yrcompletion_6yrcompletion_8yr
graduation_& retention
● 200 OK
"unitid": "166027",
"retention_rate_ft": 0.99,
"grad_rate_150_pct": 0.98,
"grad_rate_pell": 0.97,
"grad_rate_non_pell": 0.98,
"transfer_out_rate": 0.01,
"completion_4yr": 0.86
# unitidretention_rate_ftretention_rate_ptgrad_rate_150_pctgrad_rate_pellgrad_rate_non_pell
1
2
3

Complete list of extractable fields for Earnings & Debt objects from collegescorecard.ed.gov. All fields typed and schema-versioned.

unitidmedian_earnings_10yrpct_earning_above_hsmedian_debt_completersmedian_debt_noncompletersdefault_rate_3yrrepayment_rate_1yrmonthly_loan_paymentfield_of_study_codefield_of_study_earnings
earnings_& debt
● 200 OK
"unitid": "166027",
"median_earnings_10yr": 136700,
"pct_earning_above_hs": 0.92,
"median_debt_completers": 12000,
"default_rate_3yr": 0.0,
"monthly_loan_payment": 124,
"repayment_rate_1yr": 0.95
# unitidmedian_earnings_10yrpct_earning_above_hsmedian_debt_completersmedian_debt_noncompletersdefault_rate_3yr
1
2
3

Capabilities

Extract the complete higher education dataset

Our pipeline handles the deeply nested JSON structures and historical data schemas of the College Scorecard, delivering flat, queryable tables directly to your warehouse.

Institution Metadata

Extract core identifiers (UNITID, OPEID), location data, institution type, and demographic distributions across all 6,400+ active institutions.

Cost & Net Price

Capture published tuition rates, average net price by family income band, and overall cost of attendance metrics.

Earnings by Field of Study

Extract median post-graduation earnings mapped to specific CIP (Classification of Instructional Programs) codes and degree levels.

Student Debt Metrics

Track median debt loads for completers vs non-completers, cohort default rates, and estimated monthly loan payments.

Admissions Intelligence

Capture admission rates, yield rates, and SAT/ACT percentile distributions for incoming freshman cohorts.

Demographic Data

Extract student body composition by race, ethnicity, gender, and first-generation status.

Retention & Graduation

Track first-year retention rates and completion rates at 150 percent of normal time, segmented by Pell Grant status.

Historical Time-Series

Extract data across multiple academic years to track trends in tuition inflation and earnings growth.

Scheduled Updates

Automated pipeline runs sync new data releases from the Department of Education directly into your database.

// engagement pipeline

From government data to warehouse tables

Brief in. Clean data out.

Define Scope
d 0

Specify the metrics, academic years, and institution types you need. We design the target schema.

Pipeline Build
d 2–4

We configure extraction logic to handle the Scorecard's nested JSON responses and normalise historical field changes.

Validation & QA
d 4–6

Schema validation, null-rate checks, and data type enforcement before full launch.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage.

Under the hood

Overcoming government API complexities

The College Scorecard dataset is massive and deeply nested. Here is how we process it into usable formats.

pipeline-monitor · collegescorecard.ed.gov · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Nested JSON flattening
Transforming deep dictionaries into flat tables

The Scorecard API returns deeply nested JSON objects where metrics are buried under year, category, and sub-category keys. We flatten these structures into wide, columnar formats suitable for SQL querying.

Schema drift handling
Normalising historical field changes

The Department of Education frequently renames fields or changes metric definitions across academic years. Our pipeline maps legacy field names to a unified, version-controlled schema.

Rate limit management
Respecting government API quotas

Government endpoints enforce strict rate limits. We implement distributed token bucket algorithms and exponential backoff to ensure reliable extraction without triggering firewall blocks.

Pagination logic
Complete dataset extraction

We handle the API's pagination cursors to ensure all 6,400+ institutions and their historical records are extracted without missing pages or silent failures.

Data typing
Strict type enforcement

Government APIs often return mixed types (e.g., "PrivacySuppressed" strings in numeric fields). We cast data strictly, converting suppressed values to nulls and enforcing numeric types for downstream analytical use.

Applications

Who uses College Scorecard data

Teams across industries use collegescorecard.ed.gov data to build competitive products and smarter operations.

01
EdTech Platforms

College search engines and application portals enrich their platforms with authoritative cost and outcome data.

02
Policy Research

Think tanks and researchers analyse the correlation between student debt loads, institution types, and long-term earnings.

03
Student Loan Underwriting

Fintech lenders use institutional default rates and median earnings data to refine risk models for private student loans.

04
Institutional Benchmarking

Universities track competitor tuition rates, yield rates, and graduation metrics to inform strategic planning.

05
Career Counseling Tools

Platforms calculate the ROI of specific degree programmes by comparing upfront costs against 10-year median earnings.

06
Economic Development

State agencies track graduate retention and earnings to assess the regional economic impact of higher education institutions.

Why DataFlirt

"The College Scorecard contains the definitive dataset on higher education ROI in the US — but mapping its deeply nested structures into flat analytical tables requires serious engineering."

Most teams waste weeks untangling the Scorecard's nested JSON responses, handling rate limits from government endpoints, and normalising historical field changes. DataFlirt absorbs that complexity so your data engineers can focus on analysis, not pipeline maintenance.

Technical Spec

Scorecard extraction capabilities

Everything supported by our collegescorecard.ed.gov scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

Institution metadata
Core identifiers, location, and classification data
Supported
Field of study earnings
Earnings data mapped to specific CIP codes
Supported
Historical time-series
Data extraction across multiple academic years
Supported
Nested JSON flattening
Conversion of nested API responses to flat relational tables
Supported
Pagination handling
Automated cursor management for complete dataset extraction
Supported
Change detection
Hash-based diffing to identify updated records in new releases
Supported
Individual student PII
Names, addresses, or personal details of individual students
Partial
Internal FAFSA applications
Raw financial aid application data submitted by students
Partial
Individual loan default histories
Specific borrower default records (only cohort aggregates are available)
Partial
Infrastructure

Infrastructure powering the pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheusSnowflakeBigQuery
Automated Normalisation

Custom Python 3.12 processing layers flatten complex JSON dictionaries and cast mixed-type fields into strict warehouse-ready formats.

Rate Limit Management

Redis-backed distributed rate limiting ensures extraction stays within government API quotas while maximising throughput.

Cloud-Native Delivery

Airflow orchestrates extraction, transformation, and load sequences, pushing clean Parquet files directly to S3 or BigQuery.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Nested or flat structures depending on schema requirements
CSV
Flat files with strict column typing
Parquet
Columnar format optimized for analytical queries
AWS S3
Direct bucket delivery for data lake integration
Webhook
HTTP POST delivery for event-driven architectures
API
REST endpoint access to extracted datasets
PostgreSQL
Direct upsert into your relational schema
XLS
Excel format for business analyst teams
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About collegescorecard.ed.gov scraping, legality, and pipeline operations.

Ask us directly →
Is it legal to scrape collegescorecard.ed.gov?

Yes. The College Scorecard dataset is public government data provided by the US Department of Education. We extract only aggregate, publicly available metrics and respect API rate limits. No personal identifiable information (PII) is accessed or extracted.

How often is the College Scorecard data updated?

The Department of Education typically releases major updates annually, with occasional minor revisions throughout the year. We can configure pipelines to poll for changes or run on a defined schedule to capture updates as they occur.

Do you handle 'PrivacySuppressed' values?

Yes. The dataset frequently uses 'PrivacySuppressed' strings in numeric fields to protect small cohort identities. Our pipeline automatically converts these to NULL values to maintain strict numeric typing for your warehouse.

Can you extract data for specific academic years?

Yes. We can target the most recent reporting year or extract historical time-series data spanning multiple academic years to support trend analysis.

How do you handle the nested JSON structure of the API?

We build custom transformation layers that flatten the nested dictionaries. We map fields like earnings by CIP code or net price by income band into wide relational tables or distinct dimension tables based on your schema preference.

Can I get earnings data separated by degree type?

Yes. The dataset includes earnings metrics mapped to specific fields of study (CIP codes) and credential levels (e.g., Bachelor's vs Master's). We extract these relationships intact.

What format is best for this data?

Due to the width of the flattened dataset (often hundreds of columns), we strongly recommend Parquet format delivered to S3, BigQuery, or Snowflake for optimal query performance.

$ dataflirt scope --new-project --source=collegescorecard.ed.gov ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Stop wrestling with nested JSON and rate limits. We build and manage the extraction infrastructure so you can query clean College Scorecard data immediately.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in education and courses

Services

Data Extraction for Every Industry

View All Services →