Find the per-dataset parquet shards for a core table
Usage
core_shard_paths(
table,
root = ".",
parquet_dir = cc_stage_path("parquet"),
exclude = release_excluded_datasets(root)
)Arguments
- table
core table name (e.g.
"obs")- root
workflows repo root (contains
data/parquet/)- parquet_dir
directory holding the per-dataset output dirs. Defaults to the local staging root (see
cc_stage_dir()), where the bulk parquet lives; an absolute path is used as-is, a relative one is resolved againstroot. The JSON sidecars stay in the repo and are found separately.- exclude
dataset dir names to skip; defaults to the ingests that declare
calcofi.in_release: false(seerelease_excluded_datasets()), so an in-progress ingest's shards stay out of the release even though its parquet is on disk