15  The data management plan

The SIO-CalCOFI Data Management Plan is a two-year action plan (2026–2028) that organizes the program’s data work into five phases — Ingest, Publish, Integrate, Visualize, Synthesize — and names, for each action, a lead, a deliverable and the people doing it. This book is arranged along the same five phases, so every deliverable that is a document, a page or a product has a home here. The table below mirrors the team’s status sheet, refreshed with each quarterly status entry; the sheet itself is the working record.

Updated 2026-09-08 from the action sheet. Period: 2026-09-01 to 2028-08-31 (two years; Year 1 is the 2026-09 to 2027-08 Ocean Metrics SoW).

15.1 The tasks, by phase

Table 15.1 is the plan as the action sheet holds it, grouped by the five phases.

Table 15.1: The plan’s tasks (1–17 from the proposal; 18 onward added by the team), with the lead, the status the sheet reports, and where in this book the deliverable lives.
task action lead status in this book
Ingest
Task 1 · cast, profile and bottle data Migrate the bottle database into a non-proprietary relational database Rasmus Swalethorp awaiting finalization of the CTD profile database and training host:port:database:user:password — the last field is your CalCOFI database password, CTD QA/QC, start to release
Task 2 · cast, profile and bottle data Construct a CTD profile database and integrate it with the bottle database Rasmus Swalethorp PostgreSQL profile database built (2026-08-19), internal access set up, training started, the QA/QC protocol under discussion; weekly meetings with the SIO tech team; the QA/QC process being automated in Python CTD QA/QC, start to release, host:port:database:user:password — the last field is your CalCOFI database password
Task 22 · additional task Track which instruments and datasets have samples collected but no data yet Betty Huang & Mark Gold the holdings board on the CalCOFI Sheet and the holdings on calcofi.io/datasets are the start Metadata & the ingest loop
Publish
Task 3 · eDNA metabarcoding (12S MiFish, 12S MarVer1, d-loop, CO1) Ensure all future sampling follows FAIRe-formatted data collection Nastassia Patin the Fall 2026 cruise is the test case for real-time FAIRe sample metadata; starting, with the new eDNA technician
Task 4 · eDNA metabarcoding Format all past eDNA datasets and their environmental metadata to FAIRe Nastassia Patin a conversion pipeline for the existing metabarcoding datasets is built and validated
Task 5 · eDNA metabarcoding Upload eDNA datasets and metadata to OBIS Nastassia Patin one dataset ready for submission pending admin authorization; more expected in the coming months Portals
Task 6 · eDNA metabarcoding Assess the status of complementary CalCOFI-associated eDNA datasets (e.g. NCOG) Nastassia Patin not yet started
Task 18 · additional task The CalOOS inventory and portal Erin Satterthwaite & Betty Huang working with the portal team Portals
Task 19 · additional task Spin up a CalCOFI ERDDAP Ed Weber erddap.calcofi.io serves the release; NOAA's own server is Ed's Portals
Task 20 · additional task Publish the CDFW dataset and the marine mammal datasets Erin Satterthwaite the Dungeness crab megalopae dataset is in the release; its deposit with the UCSD Library and the CDFW portal link are in progress
Task 23 · additional task An internal eDNA data inventory with the FAIR status of every dataset Nastassia Patin starting
Integrate
Task 7 · zooplankton samples Explore where the SIO invertebrate tow and net samples match the ichthyoplankton samples Linsey Sala starts Fall 2026; waiting on input from the collection Keys and integrity
Task 8 · zooplankton samples Outline the matches and mismatches and an approach to optimizing the match Linsey Sala waiting on input from the collection
Task 9 · zooplankton samples Integrate the NOAA-SWFSC NetID (UUID) with incoming records and the invertebrate collection database Linsey Sala waiting on input from the collection; the database already keeps provider UUIDs as columns Keys and integrity
Task 10 · zooplankton samples Use the matching exercise and the naming conventions to decide where the database is normalized and restructured Linsey Sala waiting on input from the collection
Task 11 · cross-cutting A fast, cloud-based integrated database accessible internally and externally Ben Best & Betty Huang operational — versioned frozen releases on Google Cloud Storage, read from R, Python, ERDDAP and the browser; the internal PostgreSQL working store documented; the one access-guidance document is this book's Access and Server access chapters (Q1) Access the data, host:port:database:user:password — the last field is your CalCOFI database password, Releases
Task 12 · cross-cutting Standardized CalCOFI naming conventions for cruises, tows, times, coordinates, stations and lines, variables and units, mapped onto the existing datasets Ben Best & Betty Huang Betty's guide drafted (audience and purpose, mandatory / optional / best practice, an example table), reconciled against the registries 2026-09-08 and published as the provider guide; the database's conventions and the legacy-name crosswalk generated from the registries; the inventory of legacy names Q1, the guide circulated Q2, applied across datasets Q3 Naming conventions, Providing data to CalCOFI
Task 13 · cross-cutting Agree a standardized nomenclature for naming CalCOFI datasets Ben Best & Betty Huang dataset naming lives in each ingest's front matter (provider + dataset slug, names, category) and the provider registry, aligned with the CalOOS inventory; the governance record is the naming chapter's dataset section (Q3) Naming conventions
Visualize
Task 14 · cross-cutting A single authoritative CalCOFI data inventory and entry point, with variable-based discovery Ben Best & Betty Huang nearly done — the calcofi.io landing page, the dataset catalog with one page per dataset (in the database or not), the Explorer with six lenses, the schema and query explorers; next, variable-based (EOV) discovery folded into the portal (beta Q3, release Q4) and the federated-architecture documentation (Q4) Explore, Architecture
Task 21 · additional task A CTD viewer Betty Huang, Ben Best, Rasmus Swalethorp the CTD Explorer app and the Python profile explorer serve the team; the Explorer's sections lens serves the public CTD QA/QC, start to release, Explore
Synthesize
Task 15 · mid-term progress report (end of Year 1) A structured progress report — completed actions, challenges, next steps Mark Gold not yet due; the running record is the status chapter NA
Task 16 · final report (end of Year 2) A stand-alone, living document of the data management approach, what was done, what remains Mark Gold seeded by this book; the PDF of the book is the living document Start here, Architecture
Task 17 · guidance on the plan's actions Input, direction and review of how the actions are designed and implemented Mark Gold & Karen Stocks ongoing

15.2 Year 1: what the contractor delivers, by quarter

Ocean Metrics is the plan’s senior data-science contractor, and its Year 1 statement of work (2026-09-01 to 2027-08-31) puts each of its deliverables in a quarter (Table 15.2). The ones that are documentation land in this book:

Table 15.2: Year 1 of the Ocean Metrics statement of work, quarter by quarter.
quarter due deliverables
Q1 2026-11-30 the cloud database with internal and external access and its access guidance (Task 11: Access the data, Server access); CTD casts on ERDDAP; the bottle-migration schema documentation and the CTD profile database’s data dictionary (Tasks 1–2); the inventory of legacy names and a first draft of the naming conventions (Tasks 12–13: Naming conventions, Providing data); releases; the quarterly status entry
Q2 2027-02-28 bottle migration QA/QC against the release; CTD profile integration validated by the release gates; the naming guide and mapping tables circulated for review; NetID/UUID advisory (Task 9); the invertebrate–ichthyoplankton crosswalk (Tasks 7–8); webinar 1, the product showcase
Q3 2027-05-31 the naming guide and dataset nomenclature agreed and applied (Tasks 12–13); schema-normalization recommendations (Task 10); the eDNA-to-OBIS pipeline (Task 5); the data inventory with variable-based discovery, beta (Task 14)
Q4 2027-08-31 the inventory released, with the federated-architecture and data-access documentation (Task 14: Architecture); webinar 2, the technical deep-dives; contributions to the mid-term report and the seed of the living plan (Tasks 15–16)

15.3 How to read a status

A task’s status is the sentence the sheet carries, condensed; the sheet’s owner writes it, not the book. Where it lives is the chapter or product that holds the deliverable today, which is why some rows point at a chapter that is still being written: the plan is ahead of the record, as it should be. The status chapter is the dated log of what actually landed, by phase, each quarter.