| task | action | lead | status | in this book |
|---|---|---|---|---|
| Ingest | ||||
| Task 1 · cast, profile and bottle data | Migrate the bottle database into a non-proprietary relational database | Rasmus Swalethorp | awaiting finalization of the CTD profile database and training | host:port:database:user:password — the last field is your CalCOFI database password, CTD QA/QC, start to release |
| Task 2 · cast, profile and bottle data | Construct a CTD profile database and integrate it with the bottle database | Rasmus Swalethorp | PostgreSQL profile database built (2026-08-19), internal access set up, training started, the QA/QC protocol under discussion; weekly meetings with the SIO tech team; the QA/QC process being automated in Python | CTD QA/QC, start to release, host:port:database:user:password — the last field is your CalCOFI database password |
| Task 22 · additional task | Track which instruments and datasets have samples collected but no data yet | Betty Huang & Mark Gold | the holdings board on the CalCOFI Sheet and the holdings on calcofi.io/datasets are the start | Metadata & the ingest loop |
| Publish | ||||
| Task 3 · eDNA metabarcoding (12S MiFish, 12S MarVer1, d-loop, CO1) | Ensure all future sampling follows FAIRe-formatted data collection | Nastassia Patin | the Fall 2026 cruise is the test case for real-time FAIRe sample metadata; starting, with the new eDNA technician | |
| Task 4 · eDNA metabarcoding | Format all past eDNA datasets and their environmental metadata to FAIRe | Nastassia Patin | a conversion pipeline for the existing metabarcoding datasets is built and validated | |
| Task 5 · eDNA metabarcoding | Upload eDNA datasets and metadata to OBIS | Nastassia Patin | one dataset ready for submission pending admin authorization; more expected in the coming months | Portals |
| Task 6 · eDNA metabarcoding | Assess the status of complementary CalCOFI-associated eDNA datasets (e.g. NCOG) | Nastassia Patin | not yet started | |
| Task 18 · additional task | The CalOOS inventory and portal | Erin Satterthwaite & Betty Huang | working with the portal team | Portals |
| Task 19 · additional task | Spin up a CalCOFI ERDDAP | Ed Weber | erddap.calcofi.io serves the release; NOAA's own server is Ed's | Portals |
| Task 20 · additional task | Publish the CDFW dataset and the marine mammal datasets | Erin Satterthwaite | the Dungeness crab megalopae dataset is in the release; its deposit with the UCSD Library and the CDFW portal link are in progress | |
| Task 23 · additional task | An internal eDNA data inventory with the FAIR status of every dataset | Nastassia Patin | starting | |
| Integrate | ||||
| Task 7 · zooplankton samples | Explore where the SIO invertebrate tow and net samples match the ichthyoplankton samples | Linsey Sala | starts Fall 2026; waiting on input from the collection | Keys and integrity |
| Task 8 · zooplankton samples | Outline the matches and mismatches and an approach to optimizing the match | Linsey Sala | waiting on input from the collection | |
| Task 9 · zooplankton samples | Integrate the NOAA-SWFSC NetID (UUID) with incoming records and the invertebrate collection database | Linsey Sala | waiting on input from the collection; the database already keeps provider UUIDs as columns | Keys and integrity |
| Task 10 · zooplankton samples | Use the matching exercise and the naming conventions to decide where the database is normalized and restructured | Linsey Sala | waiting on input from the collection | |
| Task 11 · cross-cutting | A fast, cloud-based integrated database accessible internally and externally | Ben Best & Betty Huang | operational — versioned frozen releases on Google Cloud Storage, read from R, Python, ERDDAP and the browser; the internal PostgreSQL working store documented; the one access-guidance document is this book's Access and Server access chapters (Q1) | Access the data, host:port:database:user:password — the last field is your CalCOFI database password, Releases |
| Task 12 · cross-cutting | Standardized CalCOFI naming conventions for cruises, tows, times, coordinates, stations and lines, variables and units, mapped onto the existing datasets | Ben Best & Betty Huang | Betty's guide drafted (audience and purpose, mandatory / optional / best practice, an example table), reconciled against the registries 2026-09-08 and published as the provider guide; the database's conventions and the legacy-name crosswalk generated from the registries; the inventory of legacy names Q1, the guide circulated Q2, applied across datasets Q3 | Naming conventions, Providing data to CalCOFI |
| Task 13 · cross-cutting | Agree a standardized nomenclature for naming CalCOFI datasets | Ben Best & Betty Huang | dataset naming lives in each ingest's front matter (provider + dataset slug, names, category) and the provider registry, aligned with the CalOOS inventory; the governance record is the naming chapter's dataset section (Q3) | Naming conventions |
| Visualize | ||||
| Task 14 · cross-cutting | A single authoritative CalCOFI data inventory and entry point, with variable-based discovery | Ben Best & Betty Huang | nearly done — the calcofi.io landing page, the dataset catalog with one page per dataset (in the database or not), the Explorer with six lenses, the schema and query explorers; next, variable-based (EOV) discovery folded into the portal (beta Q3, release Q4) and the federated-architecture documentation (Q4) | Explore, Architecture |
| Task 21 · additional task | A CTD viewer | Betty Huang, Ben Best, Rasmus Swalethorp | the CTD Explorer app and the Python profile explorer serve the team; the Explorer's sections lens serves the public | CTD QA/QC, start to release, Explore |
| Synthesize | ||||
| Task 15 · mid-term progress report (end of Year 1) | A structured progress report — completed actions, challenges, next steps | Mark Gold | not yet due; the running record is the status chapter | NA |
| Task 16 · final report (end of Year 2) | A stand-alone, living document of the data management approach, what was done, what remains | Mark Gold | seeded by this book; the PDF of the book is the living document | Start here, Architecture |
| Task 17 · guidance on the plan's actions | Input, direction and review of how the actions are designed and implemented | Mark Gold & Karen Stocks | ongoing | |
15 The data management plan
The SIO-CalCOFI Data Management Plan is a two-year action plan (2026–2028) that organizes the program’s data work into five phases — Ingest, Publish, Integrate, Visualize, Synthesize — and names, for each action, a lead, a deliverable and the people doing it. This book is arranged along the same five phases, so every deliverable that is a document, a page or a product has a home here. The table below mirrors the team’s status sheet, refreshed with each quarterly status entry; the sheet itself is the working record.
Updated 2026-09-08 from the action sheet. Period: 2026-09-01 to 2028-08-31 (two years; Year 1 is the 2026-09 to 2027-08 Ocean Metrics SoW).
15.1 The tasks, by phase
Table 15.1 is the plan as the action sheet holds it, grouped by the five phases.
15.2 Year 1: what the contractor delivers, by quarter
Ocean Metrics is the plan’s senior data-science contractor, and its Year 1 statement of work (2026-09-01 to 2027-08-31) puts each of its deliverables in a quarter (Table 15.2). The ones that are documentation land in this book:
| quarter | due | deliverables |
|---|---|---|
| Q1 | 2026-11-30 | the cloud database with internal and external access and its access guidance (Task 11: Access the data, Server access); CTD casts on ERDDAP; the bottle-migration schema documentation and the CTD profile database’s data dictionary (Tasks 1–2); the inventory of legacy names and a first draft of the naming conventions (Tasks 12–13: Naming conventions, Providing data); releases; the quarterly status entry |
| Q2 | 2027-02-28 | bottle migration QA/QC against the release; CTD profile integration validated by the release gates; the naming guide and mapping tables circulated for review; NetID/UUID advisory (Task 9); the invertebrate–ichthyoplankton crosswalk (Tasks 7–8); webinar 1, the product showcase |
| Q3 | 2027-05-31 | the naming guide and dataset nomenclature agreed and applied (Tasks 12–13); schema-normalization recommendations (Task 10); the eDNA-to-OBIS pipeline (Task 5); the data inventory with variable-based discovery, beta (Task 14) |
| Q4 | 2027-08-31 | the inventory released, with the federated-architecture and data-access documentation (Task 14: Architecture); webinar 2, the technical deep-dives; contributions to the mid-term report and the seed of the living plan (Tasks 15–16) |
15.3 How to read a status
A task’s status is the sentence the sheet carries, condensed; the sheet’s owner writes it, not the book. Where it lives is the chapter or product that holds the deliverable today, which is why some rows point at a chapter that is still being written: the plan is ahead of the record, as it should be. The status chapter is the dated log of what actually landed, by phase, each quarter.