Skip to content
Docs

Recreate the slow zone chart

Rebuild the site's "time lost each week" chart and its "biggest since 2022" table, from the JSON feed and from the table.

The slow zones page opens with a chart of rider-hours lost each week. Here is that chart in about 15 lines of Python, then the table under it in SQL.

The site draws the chart from slow-zones.json, so the quickest copy reads the same file. The same series is in the subway_slow_zones_weekly table.

Python
import matplotlib.pyplot as plt
import pandas as pd
import requests
data = requests.get("https://transitlab.nyc/data/slow-zones.json", timeout=30).json()
weekly = pd.DataFrame(data["weekly"])
weekly["week"] = pd.to_datetime(weekly["week"])
weekly = weekly.iloc[1:-1] # the first and last weeks are partial
fig, ax = plt.subplots(figsize=(10, 4))
ax.fill_between(weekly["week"], weekly["riderHours"], color="#0039A6", alpha=0.12, lw=0)
ax.plot(weekly["week"], weekly["riderHours"], color="#0039A6", lw=1.2)
ax.set_ylim(0)
ax.set_title("Rider-hours lost to subway slow zones each week, weekdays 6 am to 10 pm", loc="left")
ax.spines[["top", "right"]].set_visible(False)
fig.savefig("slow_zones_weekly.png", dpi=150, bbox_inches="tight")

Divide a week’s riderHours by five for a typical weekday. The site’s headline number is that, averaged over the last four full weeks: summary.riderHoursPerWeekday.

“Biggest since 2022” ranks zones by total rider-hours lost. The feed holds the top 25; the subway_slow_zones table holds all of them, so you can rank any way you like.

SQL
ATTACH 'https://data.transitlab.nyc/v1/transitlab.duckdb' AS transitlab (READ_ONLY);
SELECT segment, routes, start_date, end_date, slow_days,
round(extra_s_median) AS extra_seconds,
riders_per_weekday,
rider_hours_lost
FROM transitlab.subway_slow_zones
WHERE kind = 'slow zone'
ORDER BY rider_hours_lost DESC
LIMIT 25;

The top row should be the E train between Jackson Heights-Roosevelt Av and Queens Plaza in spring 2023: about 50 extra seconds on some 141,000 riders a day, for nearly 124,000 rider-hours in all.

  • By line. routes is space-separated. unnest(string_split(routes, ' ')) gives a row per route.
  • By year. Group on year(start_date) to see if things are getting better.
  • Not just slow zones. Drop the kind filter to see disruptions and service changes too, and read the method for why they’re set aside.