Turning 750 PDFs into a Research Database

TL;DR: Medicaid’s rules for home and community-based care are public, but they are buried in hundreds of state waiver applications, each running to hundreds of pages. The Center for Retirement Research at Boston College extracted that policy detail from more than 750 applications into a multi-state, longitudinal database. We built the portal that publishes it: an interactive data browser, versioned downloads, and searchable documentation, all driven by one metadata model so the three never drift apart.

Most Americans who need long-term care would rather receive it at home than in a nursing facility. Medicaid pays for much of that care through home- and community-based services (HCBS), and states have wide latitude in deciding what to cover, for whom, and how generously. The rules are public, but they are buried in waiver applications that run to hundreds of pages each, filed state by state and amended year after year.

That makes a simple question surprisingly hard to answer: which states cover which services, and when did that change?

The Center for Retirement Research at Boston College (CRR) set out to fix this. With support from the Alfred P. Sloan Foundation, CRR researchers built a multi-state, longitudinal database of Medicaid HCBS policy. Gnaritas built the platform that publishes it: the Medicaid HCBS Waiver Dataset.

What is in the dataset

The first release covers 1915(c) waivers, the main vehicle states use to offer HCBS. It includes:

  • Nearly 100 waiver programs
  • 45 states and the District of Columbia
  • Policy detail dating back to 2006
  • Data extracted from more than 750 waiver applications submitted through the CMS Waiver Management System

Each application is broken into the sections researchers care about: the main application, requests for renewal and amendment, participant eligibility (Appendix B), covered services and provider rules (Appendix C), participant direction (Appendix E), financial responsibility (Appendix I), and cost-neutrality estimates (Appendix J). In all, the 1915(c) data spans more than 650 variables.

CRR plans to add state plan documents and CMS-372 reports later this year.

What we built

The site gives researchers three ways into the data.

An interactive data browser. Pick states, a range of years, and the variables you want, then view the results as a table. Variables can be browsed by topic or alphabetically, and each one shows its definition from the codebook. Researchers can answer a specific question in a minute or two without downloading anything.

Downloads, with versioned releases. The full dataset is available as a ZIP of joined CSV files, and every table can be downloaded on its own as CSV or JSON. Each release is a frozen snapshot, and earlier releases stay available, so an analysis run against last quarter’s data can be reproduced exactly.

Searchable documentation. Every variable is listed with its description and format, filterable by topic. The documentation is generated from the same metadata that drives the browser and the exports, so it cannot drift out of step with the data.

The interesting problems

The data has two shapes. Some facts describe a waiver as a whole: its cost limits, its eligibility groups. Others describe a single service within a waiver: who may provide it, what it is expected to cost. Combining these carelessly multiplies rows and produces wrong numbers. We built a query planner that understands the grain of each table and the relationships between them, so that whatever combination of variables a researcher selects, the result has the right rows.

The model is driven by metadata, not code. Table definitions, variable descriptions, join paths, and export column order all live in configuration. When the research team renames a variable or adds a table, the browser, the downloads, and the documentation update together.

Researchers own the pipeline. The CRR team maintains its working data in spreadsheets. We built a validation and export tool that checks a workbook against the approved structure and produces import-ready files, plus an admin workflow that stages each import, reports problems row by row, and publishes nothing until a person chooses to release it. Rows that cannot be tied back to a source document are rejected rather than published as orphans.

A standalone portal inside an existing site. The dataset has its own identity: its own header, navigation, and visual design, distinct from the main CRR website. Under the hood, though, it runs as a plugin inside CRR’s existing WordPress installation. Researchers get a focused portal built around the data, while CRR staff manage it from an admin area they already know, on hosting they already have.

Why it matters

Policy research depends on comparable data. Until now, anyone studying HCBS policy across states had to begin by reading the waivers themselves. With this dataset, a researcher can start from a clean, documented, versioned table and spend their time on the question instead.

Explore the dataset at crr.bc.edu/medicaid-hcbs-waiver-dataset. Questions about the data can be sent to crr@bc.edu.