Open NYC transit data
The data behind transitlab.nyc, free to download and query. Subway run times and slow zones, bus speeds and bus lanes, taxi and Uber/Lyft trips, and our archive of the MTA's live feeds, as Parquet files and one DuckDB database.
Everything on transitlab.nyc is built from public data, and we publish what we build. You can query it from your laptop with no account and no key.
Start in one line
Section titled “Start in one line”With DuckDB 1.1 or newer:
ATTACH 'https://data.transitlab.nyc/v1/transitlab.duckdb' AS transitlab (READ_ONLY);
-- The ten costliest subway slow zones since 2022SELECT segment, routes, start_date, end_date, rider_hours_lostFROM transitlab.subway_slow_zonesWHERE kind = 'slow zone'ORDER BY rider_hours_lost DESCLIMIT 10;DuckDB fetches only the parts of each file a query needs, so this runs in seconds even though the database points at gigabytes of trips. Rather not install anything? Use the SQL console in your browser.
What we publish
Section titled “What we publish”See all datasets for sizes and columns.
Three ways in
Section titled “Three ways in”- Parquet files at
https://data.transitlab.nyc/v1/. Small tables are one file (v1/<table>.parquet); large ones are one file a month or a day (v1/<table>/year=YYYY/month=MM/…).v1/catalog.jsonlists every table with its rows, size, source, license and columns. - One DuckDB database at
v1/transitlab.duckdbthat knows every file, soFROM transitlab.<table>just works. See DuckDB, Python, R and JavaScript. - JSON feeds that the site itself draws from, small enough for a web page: JSON feeds.
Using it
Section titled “Using it”Our data is free to use under CC BY 4.0: credit transitlab.nyc and the original source. The methods explain how each number is made and where it can mislead. The examples walk through four questions start to finish.