We extract TEOS records, Form 990 metadata, AFR tables, and tax professional directories from irs.gov. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Form 990 (TEOS) objects from irs.gov. All fields typed and schema-versioned.
"ein": "12-3456789", "organization_name": "GLOBAL HEALTH INITIATIVE", "city": "WASHINGTON", "state": "DC", "tax_period": "2024-12", "form_type": "990", "asset_amount": 14500000.0, "pdf_url": "https://apps.irs.gov/pub/epubs/990/123456789_202412_990.pdf"
| # | ein | organization_name | doing_business_as | city | state | country |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Tax Professionals objects from irs.gov. All fields typed and schema-versioned.
"first_name": "JANE", "last_name": "DOE", "credentials": "['Enrolled Agent', 'CPA']", "city": "AUSTIN", "state": "TX", "zip_code": "78701", "enrollment_status": "Active", "disciplinary_actions": false
| # | preparer_id | first_name | last_name | credentials | city | state |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for AFR Tables objects from irs.gov. All fields typed and schema-versioned.
"revenue_ruling": "2025-04", "issue_date": "2025-03-15", "period": "April 2025", "short_term_afr": 4.12, "mid_term_afr": 3.98, "long_term_afr": 4.05, "section_382_rate": 4.15
| # | revenue_ruling | issue_date | period | short_term_afr | short_term_110 | short_term_120 |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Forms & Pubs objects from irs.gov. All fields typed and schema-versioned.
"product_number": "Form 1040", "title": "U.S. Individual Income Tax Return", "revision_date": "2024", "format": "PDF", "file_size_kb": 245, "download_url": "https://www.irs.gov/pub/irs-pdf/f1040.pdf", "language": "English"
| # | product_number | title | revision_date | category | format | file_size_kb |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for News & Bulletins objects from irs.gov. All fields typed and schema-versioned.
"document_number": "IR-2025-42", "title": "IRS announces tax relief for severe storm victims", "publication_date": "2025-02-14", "document_type": "News Release", "abstract": "The Internal Revenue Service today announced tax relief for individuals and businesses affected by severe storms...", "related_forms": "['Form 4684', 'Form 1040']", "internal_revenue_code_refs": "['Section 165(i)']"
| # | document_number | title | publication_date | document_type | abstract | full_text_url |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our irs.gov scraper navigates government WAFs, CAPTCHAs, and legacy HTML structures to extract clean, structured data from the Tax Exempt Organization Search, preparer directories, and regulatory publications.
Extract Form 990 metadata, asset amounts, income figures, and PDF links for millions of tax-exempt organisations.
Index the public directory of federal tax return preparers with credentials and select qualifications.
Monitor and extract monthly AFR tables, Section 382 rates, and low-income housing rates from Revenue Rulings.
Track revisions to tax forms, instructions, and publications. Download PDFs and extract metadata automatically.
Parse IRBs for new regulations, notices, announcements, and revenue procedures with cross-referenced tax codes.
Track organisations that have lost their tax-exempt status due to failure to file Form 990 for three consecutive years.
Run continuous pipelines to detect updates in tax manuals or procedural guidelines, outputting only the diffs.
Extract legacy Form 990s and historical AFR tables from IRS archives to build comprehensive time-series datasets.
Configure pipelines at monthly, weekly, or daily cadences to align with IRS publication schedules.
Extract text and form field metadata from IRS-published PDFs to normalise unstructured regulatory documents.
Brief in. Clean data out.
Specify target datasets: TEOS records by EIN list, AFR tables, or preparer directories by zip code.
We configure Scrapy crawlers, proxy rotation, and CAPTCHA solvers to navigate Akamai and IRS rate limits.
Schema validation, null-rate checks, and cross-referencing against public IRS indices before launch.
JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
Government websites are notorious for strict WAFs, aggressive rate limiting, and legacy architectures. Here is how we ensure reliable extraction.
The IRS uses enterprise WAFs that block standard datacenter IPs and headless browsers. We route requests through US-based residential proxies and apply realistic TLS fingerprints to maintain uninterrupted access.
The Tax Exempt Organization Search requires solving CAPTCHAs for bulk queries. Our pipeline integrates automated solver APIs to process high-volume EIN lookups without manual intervention.
Historical IRS bulletins and older forms are hosted on legacy HTML pages with inconsistent formatting. We maintain extensive fallback selectors to normalise data across decades of changing web standards.
Government servers drop connections under heavy load. We optimise concurrency limits and implement exponential backoff retry policies to respect server capacity while meeting delivery SLAs.
Public directories often restrict pagination depth. We use recursive geographic or chronological search parameters to bypass UI limits and extract the complete underlying dataset.
Philanthropic platforms aggregate Form 990 data to evaluate charity financials, executive compensation, and program efficiency.
Fintech and tax preparation software ingest monthly AFR tables and updated form instructions to keep calculation engines compliant.
Advisors analyse foundation filings and charitable contributions to identify high-net-worth individuals and corporate donors.
Corporate compliance teams track Internal Revenue Bulletins to monitor changes in tax codes and revenue procedures.
Economists and policy researchers build time-series databases of non-profit revenues and federal interest rates to study economic trends.
Background check providers query the Enrolled Agent directory to verify the credentials and disciplinary history of tax professionals.
"The IRS publishes massive datasets on non-profits, tax preparers, and federal rates, but the interfaces are built for single queries, not bulk extraction."
Extracting data from irs.gov requires navigating strict government WAFs, solving complex CAPTCHAs on the TEOS portal, and parsing inconsistent HTML formats from historical archives. DataFlirt handles the extraction, normalisation, and delivery so your analysts can focus on tax intelligence rather than infrastructure maintenance.
Everything supported by our irs.gov scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy manages concurrency and retry logic for static document archives, while Playwright handles JavaScript execution and CAPTCHA rendering on the TEOS portal.
Government sites aggressively block foreign and datacenter IPs. We route all irs.gov traffic through US-based residential ISP proxies to ensure consistent access.
Pipelines run on Kubernetes with Airflow handling scheduling for monthly AFR updates and daily TEOS refreshes. State is managed in PostgreSQL.
Data delivered to where your team already works — no new tooling required.
About irs.gov scraping, legality, and pipeline operations.
Ask us directly →Yes. DataFlirt extracts only publicly available information such as Form 990 filings, AFR tables, and public directories. We do not access authenticated systems, bypass ID.me, or extract non-public Personally Identifiable Information (PII). Scraping public government records is generally protected, though clients should ensure their specific use case complies with relevant regulations.
We integrate automated CAPTCHA solving services (CapSolver and 2Captcha) directly into our Playwright automation scripts. This allows us to process thousands of EIN lookups without manual intervention.
Yes. While the primary TEOS pipeline extracts the metadata and PDF URLs, we can configure secondary processing pipelines using OCR and PDF parsing libraries to extract specific line items from the documents themselves.
The IRS typically publishes AFRs around the 20th of each month for the following month. Our pipelines can be scheduled to run daily to detect the new Revenue Ruling and extract the tables within hours of publication.
Yes. We can configure backfill pipelines to extract historical Form 990s, past AFR tables, and legacy Internal Revenue Bulletins dating back to the earliest available records on the irs.gov archives.
No. Accessing individual or corporate tax transcripts requires authenticated access via ID.me and explicit taxpayer consent. DataFlirt strictly targets public, unauthenticated datasets.
Engagements typically start with a defined extraction scope, such as monitoring a list of 10,000 EINs for Form 990 updates or a one-off scrape of the entire tax professional directory. Contact us with your specific data requirements for a quote.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a complete dump of the tax professional directory or continuous monitoring of Form 990 filings, we build and operate the infrastructure. Tell us what you need.