Installation
nycdogs is a data package. Install the package from GitHub with:
# install.packages("pak")
pak::pak("kjhealy/nycdogs")Alternatively, install it from my r-universe:
install.packages(
"nycdogs",
repos = c("https://kjhealy.r-universe.dev", "https://cloud.r-project.org")
)Including https://cloud.r-project.org ensures dependencies on CRAN are resolved automatically.
The nycdogs package contains two datasets, nyc_license and nyc_bites. They contain, respectively, data on all licensed dogs in New York city (current to 2026), and data on reported dog bites in New York city (2015 to 2025). Both tables have a breed_rc column of recoded breed names that is coded in the same way, so that breeds can be compared across them. It depends on nycmaps, which is attached automatically and provides the zip code table and map used to draw maps of the data.
Loading the data
The package works best with the tidyverse libraries and the simple features package for mapping.
library(tidyverse)
#> ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
#> ✔ dplyr 1.2.1 ✔ readr 2.2.0
#> ✔ forcats 1.0.1 ✔ stringr 1.6.0
#> ✔ ggplot2 4.0.3 ✔ tibble 3.3.1
#> ✔ lubridate 1.9.5 ✔ tidyr 1.3.2
#> ✔ purrr 1.2.2
#> ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
#> ✖ dplyr::filter() masks stats::filter()
#> ✖ dplyr::lag() masks stats::lag()
#> ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(sf)
#> Linking to GEOS 3.13.0, GDAL 3.8.5, PROJ 9.5.1; sf_use_s2() is TRUELoad the data:
To look at the tibble that contains the licensing data, do this:
nyc_license
#> # A tibble: 819,323 × 12
#> animal_name animal_gender animal_birth_year breed_name breed_rc zip_code
#> <chr> <chr> <int> <chr> <chr> <chr>
#> 1 Paige F 2014 American Pit B… Pit Bul… 10035
#> 2 Yogi M 2010 Boxer Boxer 10465
#> 3 Ali M 2014 Basenji Basenji 10013
#> 4 Queen F 2013 Akita Crossbre… Akita C… 10013
#> 5 Lola F 2009 Maltese Maltese 10028
#> 6 Ian M 2006 Unknown Unknown 10013
#> 7 Buddy M 2008 Unknown Unknown 10025
#> 8 Chewbacca F 2012 Labrador Retri… Labrado… 10013
#> 9 Heidi-bo F 2007 Dachshund Smoo… Dachshu… 11215
#> 10 Massimo M 2009 Bull Dog, Fren… French … 11201
#> # ℹ 819,313 more rows
#> # ℹ 6 more variables: zip <chr>, license_issued_date <date>,
#> # license_expired_date <date>, extract_year <int>, borough <chr>, city <chr>Licenses, dogs, and duplicate records
Two features of the data matter for any analysis.
First, dog licenses expire. Each row in nyc_license is a record of a time-limited license that was issued, not necessarily the record of a unique individual dog. A dog whose license is renewed appears once for each license period.
Second, the table contains several extracts of the licensing data, marked by the extract_year column. A license that was active in more than one extract year appears in each of them, so there are a substantial number of duplicated rows.
Any analysis of the table should try to de-duplicate the records. For example, arrange the data by extract year, find distinct records based on a number of the identifying columns, and keep only one of them (here, the earliest):
nyc_license |>
arrange(extract_year) |>
distinct(
animal_name,
animal_gender,
animal_birth_year,
breed_name,
zip,
license_issued_date,
license_expired_date,
.keep_all = TRUE
)
#> # A tibble: 664,044 × 12
#> animal_name animal_gender animal_birth_year breed_name breed_rc zip_code
#> <chr> <chr> <int> <chr> <chr> <chr>
#> 1 Paige F 2014 American Pit B… Pit Bul… 10035
#> 2 Yogi M 2010 Boxer Boxer 10465
#> 3 Ali M 2014 Basenji Basenji 10013
#> 4 Queen F 2013 Akita Crossbre… Akita C… 10013
#> 5 Lola F 2009 Maltese Maltese 10028
#> 6 Ian M 2006 Unknown Unknown 10013
#> 7 Buddy M 2008 Unknown Unknown 10025
#> 8 Chewbacca F 2012 Labrador Retri… Labrado… 10013
#> 9 Heidi-bo F 2007 Dachshund Smoo… Dachshu… 11215
#> 10 Massimo M 2009 Bull Dog, Fren… French … 11201
#> # ℹ 664,034 more rows
#> # ℹ 6 more variables: zip <chr>, license_issued_date <date>,
#> # license_expired_date <date>, extract_year <int>, borough <chr>, city <chr>This leaves one row per license. To count dogs and not licenses, drop the two license date columns from the call to distinct(). There is no dog identifier in the data, so this is approximate: two dogs with the same name, sex, birth year, and breed in the same zip code cannot be told apart.
Example
Where dogs with a particular name live. We want to count dogs and not licenses, so we first de-duplicate the table, leaving the license dates out of the identifying columns:
nyc_dogs <- nyc_license |>
arrange(extract_year) |>
distinct(
animal_name,
animal_gender,
animal_birth_year,
breed_name,
zip,
.keep_all = TRUE
)
nyc_dogs
#> # A tibble: 357,664 × 12
#> animal_name animal_gender animal_birth_year breed_name breed_rc zip_code
#> <chr> <chr> <int> <chr> <chr> <chr>
#> 1 Paige F 2014 American Pit B… Pit Bul… 10035
#> 2 Yogi M 2010 Boxer Boxer 10465
#> 3 Ali M 2014 Basenji Basenji 10013
#> 4 Queen F 2013 Akita Crossbre… Akita C… 10013
#> 5 Lola F 2009 Maltese Maltese 10028
#> 6 Ian M 2006 Unknown Unknown 10013
#> 7 Buddy M 2008 Unknown Unknown 10025
#> 8 Chewbacca F 2012 Labrador Retri… Labrado… 10013
#> 9 Heidi-bo F 2007 Dachshund Smoo… Dachshu… 11215
#> 10 Massimo M 2009 Bull Dog, Fren… French … 11201
#> # ℹ 357,654 more rows
#> # ℹ 6 more variables: zip <chr>, license_issued_date <date>,
#> # license_expired_date <date>, extract_year <int>, borough <chr>, city <chr>Then we find the share of dogs named Coco that live in each zip code, and map it:
boro_names <- c("Manhattan", "Queens", "Brooklyn", "Bronx", "Staten Island")
nyc_coco <- nyc_dogs |>
filter(borough %in% boro_names) |>
count(zip, animal_name) |>
complete(zip, animal_name, fill = list(n = 0)) |>
filter(animal_name == "Coco") |>
mutate(
freq = n / sum(n),
pct = round(freq * 100, 2)
)
nyc_coco
#> # A tibble: 197 × 5
#> zip animal_name n freq pct
#> <chr> <chr> <int> <dbl> <dbl>
#> 1 10001 Coco 21 0.00820 0.82
#> 2 10002 Coco 31 0.0121 1.21
#> 3 10003 Coco 12 0.00469 0.47
#> 4 10004 Coco 2 0.000781 0.08
#> 5 10005 Coco 4 0.00156 0.16
#> 6 10006 Coco 0 0 0
#> 7 10007 Coco 7 0.00273 0.27
#> 8 10009 Coco 32 0.0125 1.25
#> 9 10010 Coco 27 0.0105 1.05
#> 10 10011 Coco 28 0.0109 1.09
#> # ℹ 187 more rows
## nyc_zip_sf is from the nycmaps package, which is automatically loaded by nycdogs.
coco_map <- left_join(nyc_zip_sf, nyc_coco, by = join_by(zip))
coco_map |>
ggplot(mapping = aes(fill = pct)) +
geom_sf(color = "gray80", linewidth = 0.1) +
scale_fill_binned(guide = "bins", type = "viridis", option = "A") +
labs(
title = "Where's Coco?",
fill = "Percent of all NYC\ndogs named Coco"
) +
theme_void()
Distribution of Dogs Named Coco
Hex photo: Detail from Elliott Erwitt, “New York City, 1974”.
