Parse FCS (Flow Cytometry Standard) files v2.0-3.1. Extract events as NumPy arrays, read metadata/channels, convert to CSV/DataFrame, for flow cytometry data preprocessing.
SKILL.md
FlowIO
Purpose
Use FlowIO as a lightweight, low-level reader and writer for Flow Cytometry
Standard files. Examples in this skill target FlowIO 1.4.0, the current
stable release verified on 2026-07-23.
FlowIO is appropriate for:
Reading FCS 2.0, 3.0, and 3.1 files
Inspecting HEADER, TEXT, ANALYSIS, and channel metadata
Retrieving event data as a two-dimensional NumPy array
Reading legacy files that contain multiple datasets
Writing list-mode, single-precision FCS 3.1 files
Preparing data for pandas, machine-learning, or downstream cytometry tools
FlowIO does not perform compensation, logicle/biexponential transforms,
gating, clustering, or FlowJo workspace processing. Use FlowKit or another
analysis package for those tasks.
Install
Create or activate a Python environment, then install the verified release:
uv pip install "flowio==1.4.0"
Confirm the runtime version:
uv run python -c "import flowio; print(flowio.__version__)"
FlowIO 1.4.0 supports Python 3.9 through 3.13 and depends on NumPy.
Operating Workflow
Clarify the operation. Distinguish metadata inventory, event extraction,
file repair, conversion, and downstream biological analysis.
Inspect before loading events. Use only_text=True for metadata-only
work, especially with large or unfamiliar files.
Choose event semantics explicitly. Use as_array(preprocess=True) for
gain/log/time scaling from FCS metadata, or preprocess=False for values as
encoded in the DATA segment. Record the choice.
Keep parsing strict by default. Do not automatically suppress offset
errors. Relax checks only for a known vendor-format defect, and review the
resulting event data.
Treat metadata as potentially sensitive. FCS TEXT values can include
sample, subject, operator, and instrument identifiers. Export only fields
needed for the task.
Validate writes by reopening them. Check event/channel counts, labels,
metadata, and representative values after any FCS export.
Critical Semantics
TEXT keys are normalized
FlowData.text stores keys in lowercase and strips the leading $ from
standard FCS keywords:
Do not look up "$DATE", "$CYT", or other uppercase dollar-prefixed keys.
TEXT values remain strings. FlowIO 1.4.0 also removes every $ character from
the decoded TEXT segment, including $ characters inside values; preserve the
original file when exact metadata fidelity matters.
Events have two representations
flow.events is the unprocessed, flattened one-dimensional event array.
flow.as_array() returns shape (event_count, channel_count) as a NumPy
float64 array.
flow.as_array(preprocess=True) applies FCS gain, logarithmic, and time
scaling. It does not apply compensation or logicle/biexponential display
transforms.
flow.as_array(preprocess=False) reshapes the encoded event values without
those scaling steps.
as_array() creates another in-memory array. FlowIO does not provide chunked
or memory-mapped event access.
Channel numbering uses two conventions
NumPy columns and fluoro_indices, scatter_indices, and time_index use
zero-based indices.
flow.channels uses FCS parameter numbers beginning at 1.
null_channels contains the PnN label strings supplied through
null_channel_list, including supplied labels that were not found.
pns_labels always matches pnn_labels in length; missing optional PnS
labels appear as empty strings.
Writing is intentionally limited
create_fcs() requires:
An already-open binary file handle
Flattened one-dimensional event data in row-major event/channel order
One PnN name per channel
Optional PnS names and string-valued metadata via metadata_dict
It writes FCS 3.1 list-mode ($MODE=L) single-precision float
($DATATYPE=F) data. Required interpretation keywords are generated by
FlowIO and cannot be overridden through metadata.
Do not call as_array() on a metadata-only instance because its event data was
not loaded.
Prefer a path or Path over a caller-owned file handle. FlowData closes a
provided handle after parsing. In FlowIO 1.4.0,
read_multiple_data_sets(handle) can fail after the first dataset because the
handle has been closed; pass a filesystem path for multi-dataset files.
Quick Start: Read Multiple Datasets
Use the standalone helper rather than manually interpreting $NEXTDATA
offsets:
from flowio import read_multiple_data_sets
datasets = read_multiple_data_sets("legacy-multi-dataset.fcs")
for index, dataset in enumerate(datasets):
values = dataset.as_array(preprocess=True)
print(index, dataset.event_count, dataset.pnn_labels, values.shape)
The FCS 3.1 specification deprecated multiple datasets in one file, but FlowIO
can read legacy files that use them.
Metadata keys may be supplied in mixed case or with $, but lowercase keys
without $ match FlowIO's normalized representation and are less error-prone.
Metadata values must be strings.
Copy or Rewrite an Existing File
Use write_fcs() when the event data does not need to change:
from flowio import FlowData
flow = FlowData("source.fcs")
# Preserve selected source metadata (cyt, date, and spill/spillover when present).
flow.write_fcs("copy.fcs")
# Write only required metadata plus the custom fields supplied here.
flow.write_fcs("deidentified.fcs", metadata={"src": "Deidentified export"})
Passing metadata=None preserves FlowIO's selected defaults. Passing any
dictionary, including {}, replaces those defaults rather than merging with
them. write_fcs() always produces FCS 3.1 floating-point output; non-float
source events are preprocessed before writing. It opens the destination for
overwrite, so reject an existing output path before calling it unless
replacement is intentional. For floating-point sources it can preserve encoded
events while dropping PnG or timestep, changing later
as_array(preprocess=True) results. Validate both raw and preprocessed
round-trips.
Use create_fcs() instead when event values, event count, or channel layout
changes.
Bundled Inspector
scripts/inspect_fcs.py inventories one or more datasets without network
access. By default it reads metadata only, emits structural fields and channel
labels without full TEXT/ANALYSIS values, and refuses files above a
configurable size limit.
Set FLOWIO_SKILL_DIR to the installed skill directory. From this repository's
root, use skills/flowio:
FLOWIO_SKILL_DIR="skills/flowio"
# Metadata and channel inventory
uv run --no-project --with "flowio==1.4.0" \
python "$FLOWIO_SKILL_DIR/scripts/inspect_fcs.py" sample.fcs
# Include all normalized TEXT metadata; review output for identifiers
uv run --no-project --with "flowio==1.4.0" \
python "$FLOWIO_SKILL_DIR/scripts/inspect_fcs.py" sample.fcs --include-text
# Load events and compute finite-value statistics using FlowIO preprocessing
uv run --no-project --with "flowio==1.4.0" \
python "$FLOWIO_SKILL_DIR/scripts/inspect_fcs.py" sample.fcs --stats
# Compute statistics from encoded values instead
uv run --no-project --with "flowio==1.4.0" \
python "$FLOWIO_SKILL_DIR/scripts/inspect_fcs.py" sample.fcs --stats --raw
Use --help for output files, input/array memory limits, null-channel labels,
and controlled offset-recovery options.
References
Read only the reference needed for the current task:
references/api_reference.md — exact FlowIO 1.4.0 public API and signatures
references/workflows.md — inventory, DataFrame/CSV, batch, write, and
round-trip patterns