Phytoplankton
CalCOFI Phytoplankton (Venrick)
calcofi_phytoplanktonCalCOFICC0 · public domaindoi:10.6073/pasta/60edabfbfd85c623fce05822befaa071v2026.09.11ingested
- observations
- 159,804
- sampling events
- 409
- species
- 249
- depth
- 0–0 m
Overview
Abundance and species composition of phytoplankton (385 taxonomic categories) in the California Current, 1996-2022, by inverted-microscope counts (E. Venrick). Samples from the near-surface “second depth” are pooled across stations into four regions (NE, SE, Alley, Offshore; Hayward & Venrick 1998) before counting, so the grain is cruise x region, not per-station. Source: EDI knb-lter-cce.254.4.
calcofi.org page ↗source ↗ingest notebook ↗record JSON JSON-LD
Coverage
1 variable
phytoplankton_abundance cells/L
Species and taxa
- Dinophyceae dinoflagellates 10578
- Bacillariophyceae diatoms 10104
- Coccolithophyceae 7578
- Pterosperma 1263
- Nitzschia sicula 1263
- Nitzschia bicapitata 1263
- Oxytoxum laticeps 842
- Protoperidinium ovum 842
- Phaeocystis 842
- Chaetoceros 842
- Thalassionema nitzschioides 842
- Chaetoceros compressus 842
- Chaetoceros radicans 842
- Oxytoxum variabile 842
- Coronosphaera mediterranea 842
- Chaetoceros bacteriastroides 842
- Nitzschia americana 842
- Chaetoceros debilis 842
- Pyrocystis fusiformis 474
- 421
- Dinophysis 421
- Torodinium 421
- Kofoidinium 421
- Spatulodinium 421
- Ceratium 421
- Ceratocorys 421
- Gonyaulax 421
- Spiraulax 421
- Centrodinium 421
- Oxytoxum 421
- Blepharocysta 421
- Protoperidinium 421
- Pyrophacus 421
- Mesoporos 421
- Prorocentrum 421
- Dinophysis acuminata 421
- Dinophysis acuta 421
- Dinophysis caudata 421
- Dinophysis fortii 421
- Dinophysis mucronata 421
- Dinophysis tripos 421
- Actiniscus pentasterias 421
- Ptychodiscus noctiluca 421
- Pronoctiluca pelagica 421
- Noctiluca scintillans 421
- Ceratocorys armata 421
- Gonyaulax diegensis 421
- Gonyaulax digitale 421
- Gonyaulax fragilis 421
- Gonyaulax polygramma 421
Access
Every endpoint this dataset can be reached through, grouped by how you would use it. Each name is the link; the copy button beside it copies the address. Everything is in the release record and was answered when the release was cut.
Apps that read this dataset from the release. The icons after a name are the app’s lenses — the spatial grain it shows the data at.
Every way to have the bytes, by source. Nothing here asks you to register first.
Tables from the release (Parquet)
The release’s own tables, as the parquet objects it is frozen from — the same bytes every app and package below reads. A table this dataset shares with others holds every dataset’s rows, so filter on dataset_key; a partition holds only this dataset’s. since is the release whose rows these are: an unchanged table keeps its object.
CF netCDF
One self-describing file, for a tool that reads netCDF.
how far this file is CF
Fully CF: each row is an independent observation with its own time and position, which is a CF point collection. Rows are at the OCCURRENCE grain (event x taxon x life stage x depth), not the event grain, so a sample with many taxa contributes many points.
ERDDAP (erddap.calcofi.io)
One ERDDAP dataset per grain. Subset in the browser or query it from a script; for netCDF take the CF file above.
- observations
- one row per measurement — value, units and quality flag — joined to the sampling event it was taken on
- sampling events
- one row per cast, tow or transect: when, where and how it was sampled, with the effort that scales it
Legacy ids, still answering until 2026-12-04: calcofi_phytoplankton_old (replaced by calcofi_phytoplankton)
The same release, from a script or a browser SQL shell.
Packages
The whole release, pinned to a version, with the citation one call away.
con <- calcofi4r::cc_get_db()
calcofi4r::cc_cite("calcofi_phytoplankton")
con = calcofi4py.cc_get_db()
calcofi4py.cite("calcofi_phytoplankton")
DuckDB, anywhere
No CalCOFI package needed: each table above is a plain parquet object, readable by any DuckDB (or Arrow, pandas, Spark) from its URL. The path carries a content hash — a table whose rows did not change between releases keeps the same object, so nothing unchanged is stored or downloaded twice. Swap in any table above; one shared with other datasets needs WHERE dataset_key = 'calcofi_phytoplankton'. Every object of every release is listed in db-schema.
SELECT *
FROM read_parquet('https://storage.googleapis.com/calcofi-db/ducklake/tables/obs/dataset_key=calcofi_phytoplankton/4867e58b3ae6a03aa8b0868c/data_0.parquet')
LIMIT 100;
db-query, in the browser
__TBL:obs__ is db-query’s name for the pinned release’s obs object — the hashed path above, resolved for you — so the same SQL keeps working when a release changes the object.
-- calcofi_phytoplankton in the CalCOFI release v2026.09.11
SELECT *
FROM __TBL:obs__
WHERE dataset_key = 'calcofi_phytoplankton'
LIMIT 100;
Records about the data, in the standards each portal harvests.
Where this dataset is registered outside calcofi.io: each portal’s role, what it is for, the dataset’s status there and the identifier it is known by. The policy — which portal is the archive of record and why — is in Portals.
Policy Archive of record: CCE-LTER's EDI package knb-lter-cce.254.4, the ingest's source — never republished by CalCOFI. OBIS: planned (Darwin Core Archive built and staged each release; region-pooled events carry no date, so OBIS indexing is limited).
Everything here is open — nothing on calcofi.io asks you to register first. If this dataset ends up in something you build or publish, register your use so it can be credited and linked back; to hear when a release changes it, stay informed. Questions: data@calcofi.io.
Cite
This dataset's own citation. What to cite, and how, is a chapter of the docs book.
CalCOFI - Scripps Institution of Oceanography, California Current Ecosystem LTER, and E. Venrick. 2023. Temporal and spatial changes of the abundance and species composition of phytoplankton in the California Current from samples collected aboard CalCOFI cruises from summer 1996 through 2022. ver 4. Environmental Data Initiative. https://doi.org/10.6073/pasta/60edabfbfd85c623fce05822befaa071
License: CC0-1.0
DOI: https://doi.org/10.6073/pasta/60edabfbfd85c623fce05822befaa071
BibTeX
@misc{calcofi_phytoplankton,
title = {CalCOFI Phytoplankton (Venrick)},
howpublished = {CalCOFI - Scripps Institution of Oceanography, California Current Ecosystem LTER, and E. Venrick. 2023. Temporal and spatial changes of the abundance and species composition of phytoplankton in the California Current from samples collected aboard CalCOFI cruises from summer 1996 through 2022. ver 4. Environmental Data Initiative. https://doi.org/10.6073/pasta/60edabfbfd85c623fce05822befaa071},
year = {2023},
doi = {10.6073/pasta/60edabfbfd85c623fce05822befaa071},
url = {https://doi.org/10.6073/pasta/60edabfbfd85c623fce05822befaa071},
note = {License: CC0-1.0}
}
…and the release it came from
CalCOFI (2026). CalCOFI Integrated Database, release v2026.09.11 [Data set]. Scripps Institution of Oceanography, NOAA Fisheries, and California Department of Fish and Wildlife. https://calcofi.io/db-schema/?v=v2026.09.11
BibTeX
@misc{calcofi_release_v2026_09_11,
title = {CalCOFI Integrated Database, release v2026.09.11},
author = {CalCOFI},
year = {2026},
publisher = {Scripps Institution of Oceanography, NOAA Fisheries, and California Department of Fish and Wildlife},
url = {https://calcofi.io/db-schema/?v=v2026.09.11}
}
Source files
The original inputs the ingest notebook read, archived by the ingest to the public files bucket — 8 files, 1.5 MB.
- abund_1996_2012.csv422 KB2026-07-29
- abund_1996_2012.xlsx422 KB2026-07-29
- abund_2012_2018.csv192 KB2026-07-29
- abund_2012_2018.xlsx192 KB2026-07-29
- abund_2019_2022.csv119 KB2026-07-29
- abund_2019_2022.xlsx119 KB2026-07-29
- definitions.csv35.3 KB2026-07-29
- definitions.xlsx35.3 KB2026-07-29