For each object of each table in tables (read off catalog.json's objects[]),
one row per dataset holding rows in it:
Usage
publish_object_signatures(
con,
catalog,
tables,
owners = NULL,
cache = NULL,
base_url = "https://storage.googleapis.com/calcofi-db"
)Arguments
- con
a DuckDB connection with
httpfsloaded whenbase_urlis remote- catalog
the parsed (
simplifyVector = FALSE)catalog.json- tables
table names to sign
- owners
optional named list, table -> the dataset_keys that read it (from
datasets.json'stables[]); a table read by exactly one dataset is never scanned- cache
optional CSV path of the signature cache (created if absent)
- base_url
prefix joined to each object's
path(/-separated)
Value
A data frame: table, content_hash, dataset_key ("*" for a whole
object), signature. A table absent from the catalog gives one row with
content_hash = "<missing>".
Details
an object hive-partitioned by
dataset_key— its owncontent_hash, no read;an object of a table only one dataset reads (
owners), or with nodataset_keycolumn at all (a vocabulary such asmeasurement_type) —dataset_key = "*"and itscontent_hash: the whole object is an input of every dataset that reads the table;any other object — read once,
GROUP BY dataset_key, each group's row signature (the onefreeze_plan()uses to decideupload/copy), cached incacheunder the object'scontent_hashso an unchanged object is never read again.