ctOpenData provides a lightweight R interface to the Connecticut Open Data Portal.
The package allows users to search, filter, and download datasets from the Connecticut Open Data Portal directly into R without manually constructing API queries, handling JSON responses, or performing type conversion.
Designed for students, educators, researchers, journalists, civic
technologists, and analysts, ctOpenData reduces the
technical overhead required to begin working with municipal Open Data
while preserving access to the underlying Socrata infrastructure.
ctOpenData WorksThe package provides a streamlined interface to the Connecticut Open Data Portal’s Socrata API.
Internally, ctOpenData:
Most workflows begin with ct_list_datasets(), which
retrieves a live catalog of datasets available through the Connecticut
Open Data Portal.
Datasets can then be downloaded using either:
key"ffju-s5c5"The human-readable key is designed to improve readability and usability, while the UID is the stable identifier used by the Socrata platform.
The package provides three primary functions:
ct_list_datasets() retrieves a live catalog of
available Connecticut Open Data datasets, including human-readable keys,
Socrata UIDs, names, and other available metadata.
ct_pull_dataset() downloads cataloged datasets using
either a human-readable key or Socrata UID, with support for filtering,
ordering, date ranges, optional column-name cleaning, and optional type
coercion.
ct_any_dataset() downloads data directly from a
valid Socrata JSON endpoint without requiring the dataset to appear in
the package catalog.
Datasets retrieved through ct_pull_dataset() support
arguments including:
limitfiltersdatefromtodate_fieldwhereorderclean_namescoerce_typesAll functions return tibble outputs.
Advanced users may also provide raw SoQL conditions through the
where argument.
SoQL, or Socrata Query Language, is the query syntax used by Socrata-powered Open Data portals. Additional information is available from the Socrata developer documentation.
install.packages("ctOpenData")# install.packages("pak")
pak::pak("gomes-sh/ctOpenData")Alternatively:
# install.packages("remotes")
remotes::install_github("gomes-sh/ctOpenData")library(ctOpenData)
library(dplyr)
# Browse available datasets
catalog <- ct_list_datasets()
# Search for datasets containing a keyword
catalog |>
filter(grepl("spill", name, ignore.case = TRUE)) |>
select(key, uid, name)
# Pull a dataset using its UID
example_data <- ct_pull_dataset(
dataset = "ffju-s5c5",
limit = 100
)
# Pull the same dataset using its catalog key
example_data_by_key <- ct_pull_dataset(
dataset = "spill_incidents_from_july_1_2022_to_recent_for_download",
limit = 100
)
# Pull filtered data
filtered_data <- ct_pull_dataset(
dataset = "ffju-s5c5",
limit = 100,
filters = list(
incident_type = "Petroleum Incident"
)
)
The filters argument accepts a named list and
automatically constructs the corresponding SoQL filtering
conditions.
Multiple values may be supplied for one field:
filtered_data <- ct_pull_dataset(
dataset = "ffju-s5c5",
limit = 100,
filters = list(
incident_type = c("Biomedical Incident", "Dielectric Fluid Incident")
)
)
Multiple fields may also be combined:
filtered_data <- ct_pull_dataset(
dataset = "ffju-s5c5",
limit = 100,
filters = list(
incident_type = "Petroleum Incident",
township = "New Haven"
)
)
Date filtering is available for datasets containing date or datetime fields:
date_filtered_data <- ct_pull_dataset(
dataset = "ffju-s5c5",
from = "2023-01-01",
to = "2024-01-01",
date_field = "reported_date",
limit = 100
)
When a dataset is not available through
ct_list_datasets(), it can be downloaded directly using
ct_any_dataset().
endpoint_data <- ct_any_dataset(
json_link = "https://data.ct.gov/resource/ffju-s5c5.json",
limit = 100
)
Use ct_pull_dataset() for catalog-based workflows and
ct_any_dataset() when working directly with a Socrata JSON
endpoint.
A complete introductory workflow is available in the package vignette:
vignette("getting-started", package = "ctOpenData")The vignette demonstrates how to:
Complete documentation is available on the package website:
https://github.com/gomes-sh/ctOpenData
The website includes:
To run the package tests locally:
devtools::test()To rebuild the documentation:
devtools::document()To run a complete package check:
devtools::check()To rebuild the pkgdown website:
pkgdown::build_site()Contributions are welcome.
To report a bug, request a feature, or suggest an improvement, open an issue on GitHub:
https://github.com/gomes-sh/ctOpenData/issues
Pull requests are also welcome. Before submitting a pull request, please ensure that:
devtools::check() completes successfullyYOUR_NAME
Email:
gomessh@mailbox.org
GitHub: @gomes-sh
Because the package retrieves metadata dynamically from the live Connecticut Open Data catalog, newly published datasets may become available without requiring a package update.
Package updates may still be required when the portal changes its catalog structure, dataset metadata fields, or API behavior.
ctOpenData is an independent project and is not
affiliated with, endorsed by, or maintained by Connecticut or the
organization responsible for the Connecticut Open Data Portal.