Picoplankton & Bacteria
Picoplankton and Bacteria Abundance (CalCOFI Cruise)
cce-lter_picoplankton-bacteriaCCE-LTERunstatedv2026.09.11ingested
- years
- 2004–2023
- observations
- 60,802
- sampling events
- 16,011
- stations
- 64
- variables
- 4
- depth
- 1–270 m
Overview
Picophytoplankton (Prochlorococcus, Synechococcus, picoeukaryotes) and heterotrophic bacteria abundances analyzed by flow cytometry (FCM) from CCE-CalCOFI Augmented cruises in the California Current System, 2004-2023 (ongoing). Seawater collected from Niskin bottles at 3-8 depths per station; cells fixed shipboard with paraformaldehyde, stained with a DNA-specific dye, and enumerated on an Altra flow cytometer with dual argon-ion lasers.
Open with the provider
- Q06 license — proposed
Access
Every endpoint this dataset can be reached through, grouped by how you would use it. Each name is the link; the copy button beside it copies the address. Everything is in the release record and was answered when the release was cut.
Apps that read this dataset from the release. The icons after a name are the app’s lenses — the spatial grain it shows the data at.
Every way to have the bytes, by source. Nothing here asks you to register first.
Tables from the release (Parquet)
The release’s own tables, as the parquet objects it is frozen from — the same bytes every app and package below reads. A table this dataset shares with others holds every dataset’s rows, so filter on dataset_key; a partition holds only this dataset’s. since is the release whose rows these are: an unchanged table keeps its object.
CF netCDF
One self-describing file, for a tool that reads netCDF.
how far this file is CF
Fully CF: each row is an independent observation with its own time and position, which is a CF point collection. Rows are at the OCCURRENCE grain (event x taxon x life stage x depth), not the event grain, so a sample with many taxa contributes many points.
ERDDAP (erddap.calcofi.io)
One ERDDAP dataset per grain. Subset in the browser or query it from a script; for netCDF take the CF file above.
- observations
- one row per measurement — value, units and quality flag — joined to the sampling event it was taken on
- sampling events
- one row per cast, tow or transect: when, where and how it was sampled, with the effort that scales it
From the provider
The dataset as its provider publishes it, before CalCOFI ingested it.
The same release, from a script or a browser SQL shell.
Packages
The whole release, pinned to a version, with the citation one call away.
con <- calcofi4r::cc_get_db()
calcofi4r::cc_cite("cce-lter_picoplankton-bacteria")
con = calcofi4py.cc_get_db()
calcofi4py.cite("cce-lter_picoplankton-bacteria")
DuckDB, anywhere
No CalCOFI package needed: each table above is a plain parquet object, readable by any DuckDB (or Arrow, pandas, Spark) from its URL. The path carries a content hash — a table whose rows did not change between releases keeps the same object, so nothing unchanged is stored or downloaded twice. Swap in any table above; one shared with other datasets needs WHERE dataset_key = 'cce-lter_picoplankton-bacteria'. Every object of every release is listed in db-schema.
SELECT *
FROM read_parquet('https://storage.googleapis.com/calcofi-db/ducklake/tables/obs/dataset_key=cce-lter_picoplankton-bacteria/8164970a18e65412da071804/data_0.parquet')
LIMIT 100;
db-query, in the browser
__TBL:obs__ is db-query’s name for the pinned release’s obs object — the hashed path above, resolved for you — so the same SQL keeps working when a release changes the object.
-- cce-lter_picoplankton-bacteria in the CalCOFI release v2026.09.11
SELECT *
FROM __TBL:obs__
WHERE dataset_key = 'cce-lter_picoplankton-bacteria'
LIMIT 100;
Records about the data, in the standards each portal harvests.
Where this dataset is registered outside calcofi.io: each portal’s role, what it is for, the dataset’s status there and the identifier it is known by. The policy — which portal is the archive of record and why — is in Portals.
Policy Archive of record: CCE-LTER's EDI package knb-lter-cce.159; the ingest reads the DataZoo export of it. Never republished by CalCOFI. OBIS does not apply (abundances, not occurrences of named taxa).
Everything here is open — nothing on calcofi.io asks you to register first. If this dataset ends up in something you build or publish, register your use so it can be credited and linked back; to hear when a release changes it, stay informed. Questions: data@calcofi.io.
Cite
This dataset's own citation. What to cite, and how, is a chapter of the docs book.
Landry, M. (2004-2023). Picoplankton and Bacteria Abundance (CalCOFI Cruise). CCE LTER.
BibTeX
@misc{cce-lter_picoplankton-bacteria,
title = {Picoplankton and Bacteria Abundance (CalCOFI Cruise)},
howpublished = {Landry, M. (2004-2023). Picoplankton and Bacteria Abundance (CalCOFI Cruise). CCE LTER.},
year = {2004}
}
…and the release it came from
CalCOFI (2026). CalCOFI Integrated Database, release v2026.09.11 [Data set]. Scripps Institution of Oceanography, NOAA Fisheries, and California Department of Fish and Wildlife. https://calcofi.io/db-schema/?v=v2026.09.11
BibTeX
@misc{calcofi_release_v2026_09_11,
title = {CalCOFI Integrated Database, release v2026.09.11},
author = {CalCOFI},
year = {2026},
publisher = {Scripps Institution of Oceanography, NOAA Fisheries, and California Department of Fish and Wildlife},
url = {https://calcofi.io/db-schema/?v=v2026.09.11}
}
Source files
The original inputs the ingest notebook read, archived by the ingest to the public files bucket — 1 file, 2.1 MB.
- PicoplanktonandBacteriaAbundance.csv2.1 MB2026-07-29