Louisville Crime Data
Project: Data Lab · Code Louisville Data Analytics capstone · Completed: March 2026 · Data: Louisville Metro Police (LMPD) Open Data, U.S. Census ACS 5-year estimates · Tools: Python, pandas, SQL, SQLite, Census API, scikit-learn, matplotlib, seaborn
I combined 22 years of Louisville Metro Police incident data with U.S. Census data to answer a simple question: which neighborhood conditions track with crime across Louisville ZIP codes, and which ZIP codes break the pattern?
This is a self-directed analysis built for my Code Louisville Data Analytics capstone, not client work. Code, data sources and full methodology are on GitHub.
At a glance
- 1.7 million incident records combined from 22 yearly files (2004–2025)
- 32 Louisville ZIP codes matched to Census demographics
- 6 tables in a normalized SQLite database
- 48% of the variation in crime rate explained by five socioeconomic factors (R² = 0.482)
The questions
In 2003 Jefferson County and the City of Louisville merged into Louisville Metro, and the police departments merged with them. That gives one standardized crime dataset for the whole county from 2004 on, which makes a long-run, ZIP-level comparison possible. I set out to answer three questions:
- Which socioeconomic factors are most strongly associated with crime rates across Louisville ZIP codes over time?
- Are there ZIP codes where crime is low despite unfavorable conditions, or high despite favorable ones?
- Do national or local events line up with changes in Louisville's crime trends?
What I built
1. Combine 22 years of files
A Python function loads all 22 yearly CSVs into one 1.7M-row DataFrame and tags each row with its year. Column names changed between the 2004–2022 and 2023–2025 files, so headers were standardized first.
2. Clean what changed over time
In 2022 Louisville Metro expanded its offense classifications from 16 to 47. To keep a consistent 22-year timeline, I remapped the 47 newer categories back to the original 16. ZIP codes were standardized too: about 250 raw ZIP values were invalid, outside the county or data-entry errors, and were excluded.
3. Enrich with Census data
I pulled ACS 5-year estimates from the Census Bureau API for every Louisville ZIP code from 2012 to 2023: poverty rate, median household income, unemployment, education, single-parent households, residential mobility and rent burden.
4. Model it in SQL
The data lives in a six-table SQLite database. Lookup tables store offense types, locations and police divisions once, and a summary table pre-aggregates 1.7M incidents into about 12,000 rows by ZIP code, year and offense, so analytical queries run fast.

Three SQL queries drive the analysis:
- Core table (aggregation + join): crime rate per 1,000 residents for each ZIP code and year.
- Trends by offense (join + window function): each ZIP's rate next to the citywide average for that year, without collapsing the rows.
- ZIP profiles (two CTEs): labels every ZIP code each year as high or low poverty and high or low crime against the city average.
5. Test the relationships
A correlation matrix and a linear regression measure how each factor relates to crime rate. Two reusable functions let anyone look up a ZIP code's crime and demographic profile, or chart one ZIP against another.
What the data says
- Crime peaked in 2007–2008 and has declined since 2016, a drop that mirrors national FBI trends.
- Drug and alcohol violations and burglary fell the most between 2006–2015 and 2016–2025.
- Weapons offenses rose steadily from 2004 to 2021 before declining, and the 2021 homicide spike, part of a national pandemic-era pattern, is clearly visible.
- Poverty is the strongest positive predictor. Each percentage point of poverty adds about 15 crimes per 1,000 residents in the model (correlation +0.54).
- Education is the strongest protective factor. Each point of bachelor's-degree attainment lowers the rate by about 10 per 1,000.
- Unemployment barely registers (+0.18). Poverty explains far more than employment status alone.
- Several ZIP codes are high-poverty but low-crime. Conditions matter, but they aren't destiny.

Reading the numbers carefully
- Five factors explain 48%, not 100%. The rest likely comes from things the data doesn't measure: policing levels, neighborhood history, community programs and the built environment.
- Per-resident rates can mislead. Downtown (40202) and the expo and stadium area have few residents but heavy daytime and event traffic, so their per-capita rates run high. Residential mobility's strong negative correlation (−0.55) likely reflects the same bias.
- Census data covers 2012–2023 only, and ACS 5-year figures are rolling averages, so the joined analysis uses those overlapping years.
- Correlation isn't causation. Income and unemployment are so closely tied to poverty that they add little independent signal in the regression.
Why it's on a search consultant's site
Same method, different data. Joining records by ZIP code, cleaning categories that change over time and testing which factors really move the number is exactly how I audit ad accounts and search data. For a family law client, a ZIP-level join of spend and conversions showed that $10,517 of the first $70K had gone to ZIP codes that never produced a lead. Read the Brown Carrington case study.
