Assigns cell-type labels to clusters from a protein marker panel (multiplexed imaging such as MIBI, CODEX, IMC and CyCIF; mass and flow cytometry; CITE-seq protein), where the labels are judged against an expert reference. Within-dataset normalisation of cluster summaries, lineage first by positive defining markers, subtype only as far as the panel can determine it, marker-poor clusters assigned by exclusion, class-coverage sanity checks, and two independent annotators with adjudication. Use when the output is one label per cluster and the grader is agreement with an expert; for gating single cells in FCS data use flow-cytometry-analysis, and for transcript-based annotation use scanpy or scvi-tools.
75
94%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
An expert annotating clusters from a marker panel gets two things right that a first-pass reading gets wrong: the level at which each label is stated, and the clusters that carry almost nothing. Both directions cost equally under an expert-agreement grader: a subtype the panel cannot support scores like a wrong lineage, and a broad label where the panel determines the subtype scores like a miss. Everything below follows from that.
Assign the broad lineage from the markers that define it positively, in this order of precedence, each one excluding the ones below unless the panel says otherwise:
| Lineage | Defining positives | Typical negatives |
|---|---|---|
| Tumour / epithelial | the tumour's own markers in the panel: pan-cytokeratin, E-cadherin, EpCAM for carcinoma; SOX10, MelanA/MART-1, S100, HMB45 for melanoma; CD30 with CD15 and weak PAX5 for Hodgkin Reed-Sternberg cells; Ki-67 often high | CD45, CD3, CD20, CD68 |
| T cell | CD3 with CD4 or CD8; FoxP3 (and CD25) for regulatory; CD4 and CD8 together for double-positive | CD20, CD68 |
| B cell / plasma | CD20, CD19, PAX5; CD138 (and CD38, IRF4/MUM1) for plasma cells, which lose CD20 | CD3 |
| NK | CD56, CD57, granzyme B without CD3 | CD3 |
| Myeloid: macrophage / monocyte / DC | CD68, CD163, CD14, CD11b, CD11c, HLA-DR; CD11c with HLA-DR and without CD68 for dendritic cells; CD123 or CD303 for plasmacytoid DC | CD3, CD20, cytokeratin |
| Granulocyte / mast | CD15 or CD66b (neutrophil), CD117 or tryptase (mast), eosinophil peroxidase | CD3, CD68 |
| Endothelium | CD31, CD34, VWF; podoplanin or LYVE-1 for lymphatic | keratin |
| Fibroblast / stroma / smooth muscle | vimentin, collagen I, PDGFR-β, FAP; α-SMA (with desmin for muscle) | CD45, keratin |
Read the panel first and write down which of these markers it actually contains; the table above is a menu, not a checklist. A lineage whose defining marker is absent from the panel can only be inferred by exclusion (section 4).
A subtype label is correct only when the panel contains the marker that defines the subtype and the cluster is positive for it. The rule cuts both ways and both cuts are scored:
Every dataset has clusters that are low on everything. They are the clusters that decide a 0.95 agreement gate.
Under an expert-agreement grader, label each dataset twice, independently, by two workers who see the protocol, the vocabulary and the cluster-by-marker table and not each other's calls, with the rare classes named to both before they start. Then adjudicate every disagreement yourself by re-reading the table for that cluster against sections 2–4 — never by majority, never by splitting the difference. Record the deciding marker per cluster in an evidence file beside the output; a call you cannot justify from a marker is a call to revisit.
import numpy as np, pandas as pd
df = pd.read_csv(path) # one row per cell: cluster + marker columns
markers = [c for c in df.columns if c not in ("cellLabel", "cluster")]
means = df.groupby("cluster")[markers].mean()
z = (means - means.mean()) / means.std(ddof=0) # high for this dataset
cuts = {m: np.percentile(df[m], 90) for m in markers} # or a bimodal split per marker
frac = df.groupby("cluster")[markers].apply(lambda g: (g > pd.Series(cuts)).mean())
table = pd.concat({"mean": means, "z": z, "frac_pos": frac}, axis=1)
print(table.round(2).to_string())8d99f34
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.