SYSTEM all green source oyc.yale.edu queue 1,492 pages p99 latency 89ms dataflirt.com · scraper/oyc-yale.edu
RUN · 14 active pipelines · oyc.yale.edu live

Yale course data,
structured for ML.

We extract syllabi, lecture transcripts, reading assignments, and media links from Open Yale Courses. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake.

Courses extracted
42 /run
Lectures processed
1,382 /run
Transcripts parsed
1,382 /run
Active pipelines
14
Uptime
99.99%
Data Dictionary

Every field we extract from oyc.yale.edu

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Course Metadata objects from oyc.yale.edu. All fields typed and schema-versioned.

course_idtitledepartmentinstructorcourse_numberdescriptionsessions_countsyllabus_urlscraped_at
course_metadata
● 200 OK
"course_id": "PLSC-114",
"title": "Introduction to Political Philosophy",
"department": "Political Science",
"instructor": "Steven B. Smith",
"course_number": "PLSC 114",
"sessions_count": 24,
"scraped_at": "2026-05-12T09:14:00Z"
# course_idtitledepartmentinstructorcourse_numberdescription
1
2
3

Complete list of extractable fields for Lectures objects from oyc.yale.edu. All fields typed and schema-versioned.

lecture_idcourse_idsession_numbertitleoverviewtranscript_htmlvideo_urlaudio_urlscraped_at
lectures
● 200 OK
"lecture_id": "PLSC-114-01",
"course_id": "PLSC-114",
"session_number": 1,
"title": "Introduction: What is Political Philosophy?",
"video_url": "https://www.youtube.com/watch?v=...",
"audio_url": "https://oyc.yale.edu/audio/...",
"scraped_at": "2026-05-12T09:14:00Z"
# lecture_idcourse_idsession_numbertitleoverviewtranscript_html
1
2
3

Complete list of extractable fields for Transcripts objects from oyc.yale.edu. All fields typed and schema-versioned.

lecture_idcourse_idspeakertimestamptext_segmentsegment_indexfull_textsource_urlscraped_at
transcripts
● 200 OK
"lecture_id": "PLSC-114-01",
"course_id": "PLSC-114",
"speaker": "Professor Steven B. Smith",
"timestamp": "00:01:23",
"text_segment": "We begin today with the oldest and most fundamental question...",
"segment_index": 4,
"scraped_at": "2026-05-12T09:14:00Z"
# lecture_idcourse_idspeakertimestamptext_segmentsegment_index
1
2
3

Complete list of extractable fields for Reading Assignments objects from oyc.yale.edu. All fields typed and schema-versioned.

course_idlecture_idassignment_titleauthorpublicationpagesrequirednotesscraped_at
reading_assignments
● 200 OK
"course_id": "PLSC-114",
"lecture_id": "PLSC-114-02",
"assignment_title": "Apology",
"author": "Plato",
"publication": "Four Texts on Socrates",
"required": true,
"scraped_at": "2026-05-12T09:14:00Z"
# course_idlecture_idassignment_titleauthorpublicationpages
1
2
3

Complete list of extractable fields for Materials & PDFs objects from oyc.yale.edu. All fields typed and schema-versioned.

course_idmaterial_typetitlefile_urlformatsize_byteslecture_iddownload_statusscraped_at
materials_& pdfs
● 200 OK
"course_id": "PLSC-114",
"material_type": "Exam",
"title": "Midterm Examination",
"file_url": "https://oyc.yale.edu/sites/default/files/midterm.pdf",
"format": "PDF",
"download_status": "verified",
"scraped_at": "2026-05-12T09:14:00Z"
# course_idmaterial_typetitlefile_urlformatsize_bytes
1
2
3

Capabilities

Academic data extraction without the manual scraping

Our pipeline handles the legacy DOM structure of oyc.yale.edu, extracting text, PDFs, and media links into queryable relational formats.

Course Catalogues

Extract department listings, course numbers, titles, and descriptions across the entire Open Yale Courses repository.

Lecture Transcripts

Full text extraction from HTML transcript pages, capturing speaker attributions and segmenting by available timestamps.

Media Link Harvesting

Capture YouTube, audio, and low-bandwidth video URLs for every lecture session.

Syllabus Parsing

Convert unstructured syllabus pages into structured reading lists, exam dates, and assignment schedules.

Professor Metadata

Extract instructor names, titles, and biographical information linked to each course.

PDF Document Discovery

Identify and map URLs for exams, problem sets, and solution documents embedded within course pages.

Relational Mapping

Link lectures, transcripts, and reading assignments back to their parent course IDs for easy database ingestion.

Legacy DOM Handling

Navigate older HTML table structures and inconsistent markup common in legacy academic web platforms.

Text Normalisation

Clean transcript HTML formatting, strip extraneous tags, and normalise character encodings for NLP workloads.

// engagement pipeline

From course directory to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Select specific departments, individual courses, or request the entire oyc.yale.edu catalogue.

Pipeline Build
d 2–4

We configure Scrapy crawlers adapted to Open Yale Courses' specific HTML structure and asset paths.

Validation & QA
d 4–6

Schema validation, transcript completeness checks, and media URL verification before delivery.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage.

Under the hood

Navigating legacy academic infrastructure

Open Yale Courses relies on older web architectures. We structure this unstructured data reliably.

pipeline-monitor · oyc.yale.edu · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Legacy HTML parsing
Handling nested tables and inconsistent markup

Academic websites built in the 2000s often use deeply nested tables and inline styles. Our parsers normalise this legacy markup into clean, predictable JSON objects.

Transcript segmentation
Structuring continuous HTML text

We break long, poorly formatted transcript pages into logical text segments, preserving speaker attributions and paragraph structures for easier NLP processing.

Asset resolution
Normalising relative URLs

PDFs and audio files often use relative paths that break when scraped. We resolve all media links to absolute URLs, verifying endpoint availability during the crawl.

Rate limiting
Respecting university servers

We configure our crawlers with conservative concurrency limits and polite request delays to avoid overwhelming university infrastructure or triggering IP blocks.

Schema standardisation
Mapping varied syllabus formats

Different professors format their syllabi differently. We use heuristic matching to map varied reading lists and schedules into a single, unified schema.

Applications

Who uses Open Yale Courses data

Teams across industries use oyc.yale.edu data to build competitive products and smarter operations.

01
LLM Training

AI labs use high-quality, academically rigorous transcript corpora to fine-tune language models on complex reasoning and subject matter expertise.

02
EdTech Platforms

Course aggregators ingest metadata and syllabus details to build unified search engines for open courseware.

03
Accessibility Tools

Researchers use matched audio URLs and text transcripts to train and evaluate text-to-speech or automated captioning models.

04
Academic Research

Pedagogical researchers analyse syllabus construction, reading list diversity, and lecture structures across different departments.

05
RAG Applications

Enterprise teams build retrieval-augmented generation systems that query verified academic materials rather than the open web.

06
Content Curation

Digital libraries and academic institutions structure reading lists and material links to augment their own catalogues.

Why DataFlirt

"Open Yale Courses contains thousands of hours of Ivy League instruction, but extracting it from legacy HTML requires dedicated parsing logic."

Academic repositories often suffer from inconsistent formatting and outdated web standards. DataFlirt normalises this unstructured text, linking transcripts, reading lists, and video assets into a clean relational schema ready for ML training or database ingestion.

Technical Spec

OYC scraper - technical capabilities

Everything supported by our oyc.yale.edu scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

Transcript extraction
Full text parsing from HTML transcript pages
Supported
Media URL harvesting
Capture of YouTube, QuickTime, and MP3 links
Supported
PDF metadata parsing
Extraction of document titles and absolute URLs
Supported
Syllabus normalisation
Mapping unstructured reading lists to relational fields
Supported
Department traversal
Automated discovery of all listed courses
Supported
Legacy HTML parsing
Handling of deprecated tags and nested tables
Supported
Change detection
Hash-based diffs for updated syllabus content
Supported
Student enrollment data
Rosters and student counts are not public on OYC
Partial
Graded assignments
Accessing submitted coursework requires Yale student authentication
Partial
Infrastructure

Infrastructure powering the academic pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheusBeautifulSoup
HTML Parsing Stack

We use BeautifulSoup and Scrapy to reliably navigate and extract data from older DOM structures that modern headless browsers struggle to interpret efficiently.

Asset Pipeline

Custom middleware verifies the HTTP status of linked PDFs, audio files, and video endpoints, ensuring your database only contains valid URLs.

Cloud-Native Orchestration

Pipelines run on AWS Lambda for burst scaling, managed by Apache Airflow to handle scheduling, retries, and delivery to your data warehouse.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Nested structures ideal for transcript and syllabus data
CSV
Flat files for relational course and lecture metadata
XLS
Excel formats for manual review by academic researchers
Parquet
Columnar format for efficient querying in BigQuery or Athena
AWS S3
Direct delivery to your cloud storage buckets
Webhook
HTTP POST notifications upon pipeline completion
API
REST endpoints to query specific course data on demand
PostgreSQL
Direct database inserts with schema management
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About oyc.yale.edu scraping, legality, and pipeline operations.

Ask us directly →
Is it legal to scrape Open Yale Courses?

DataFlirt only extracts publicly available information from oyc.yale.edu. Open Yale Courses provides these materials freely to the public, often under Creative Commons licenses. Clients must review the specific terms of use and licensing agreements on the Yale website to ensure their intended use of the data complies with Yale's policies.

How do you handle the lecture transcripts?

We parse the HTML transcript pages, stripping out legacy web formatting while preserving structural elements like paragraphs and speaker changes. The output is clean, UTF-8 encoded text suitable for NLP tasks.

Do you download the video and audio files?

We extract and verify the absolute URLs pointing to the media files (e.g., YouTube links, MP3 paths). We do not download or host the multi-gigabyte media files ourselves, but deliver the links so your systems can fetch them if needed.

How often is the data updated?

Open Yale Courses is a relatively static repository. We typically configure pipelines to run monthly to detect newly added courses, corrected transcripts, or updated syllabus documents.

Do you parse the text inside the PDF documents?

Our standard pipeline extracts the metadata and URLs for PDFs (exams, problem sets). Full OCR or text extraction from inside the PDFs requires a custom processing step, which we can implement upon request.

Can I request data for only specific departments?

Yes. We can scope the crawl to target specific departments (e.g., Political Science, Physics) or even specific course IDs, rather than traversing the entire catalogue.

$ dataflirt scope --new-project --source=oyc.yale.edu ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a single department's syllabi or the entire transcript corpus for LLM training - we scope, build, and operate the pipeline.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in education and courses

Services

Data Extraction for Every Industry

View All Services →