Access the RCSB Protein Data Bank (PDB) to search, download, and programmatically retrieve 3D macromolecular structures and metadata; use when you need structure discovery (text/sequence/3D similarity) or automated structural data ingestion for structural biology and drug discovery workflows.
68
84%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Use this skill when you need to:
rcsb-api (latest recommended; provides rcsbapi.search and rcsbapi.data)requests>=2.0 (HTTP downloads)biopython>=1.80 (optional; parsing/analyzing PDB coordinates)Install (example):
uv pip install rcsb-api requests biopythonThe following script is end-to-end runnable: it searches for a target, fetches metadata, downloads coordinates, and parses the structure.
#!/usr/bin/env python3
import pathlib
import requests
from rcsbapi.search import TextQuery, AttributeQuery
from rcsbapi.search.attrs import rcsb_entry_info
from rcsbapi.data import fetch, Schema
from Bio.PDB import PDBParser
def download_text(url: str, out_path: pathlib.Path) -> None:
r = requests.get(url, timeout=60)
r.raise_for_status()
out_path.write_text(r.text, encoding="utf-8")
def main():
out_dir = pathlib.Path("pdb_out")
out_dir.mkdir(exist_ok=True)
# 1) Search: hemoglobin entries with resolution < 2.0 Å
q_text = TextQuery("hemoglobin")
q_res = AttributeQuery(
attribute=rcsb_entry_info.resolution_combined,
operator="less",
value=2.0,
)
query = q_text & q_res
pdb_ids = list(query())[:5]
if not pdb_ids:
raise SystemExit("No results found.")
pdb_id = pdb_ids[0]
print(f"Selected PDB ID: {pdb_id}")
# 2) Fetch entry metadata
entry = fetch(pdb_id, schema=Schema.ENTRY)
title = entry.get("struct", {}).get("title")
method = (entry.get("exptl") or [{}])[0].get("method")
resolution = (entry.get("rcsb_entry_info") or {}).get("resolution_combined")
deposit_date = (entry.get("rcsb_accession_info") or {}).get("deposit_date")
print("Metadata:")
print(f" Title: {title}")
print(f" Method: {method}")
print(f" Resolution: {resolution}")
print(f" Deposit date: {deposit_date}")
# 3) Download coordinates (PDB and mmCIF)
pdb_path = out_dir / f"{pdb_id}.pdb"
cif_path = out_dir / f"{pdb_id}.cif"
download_text(f"https://files.rcsb.org/download/{pdb_id}.pdb", pdb_path)
download_text(f"https://files.rcsb.org/download/{pdb_id}.cif", cif_path)
print(f"Downloaded: {pdb_path} and {cif_path}")
# 4) Parse PDB coordinates (example: count atoms)
parser = PDBParser(QUIET=True)
structure = parser.get_structure(pdb_id, str(pdb_path))
atom_count = sum(1 for _ in structure.get_atoms())
chain_ids = sorted({chain.id for chain in structure.get_chains()})
print("Parsed structure:")
print(f" Chains: {chain_ids}")
print(f" Atom count: {atom_count}")
if __name__ == "__main__":
main()evalue_cutoff: lower is more stringent (fewer, more confident hits).identity_cutoff: fraction identity threshold (e.g., 0.9 for near-identical).entry_id) as the geometric reference.query1 & query2 (AND)query1 | query2 (OR)~query (NOT), where supported by the clientSchema.ENTRY, Schema.POLYMER_ENTITY) is convenient for common objects and stable access patterns.Example GraphQL pattern:
from rcsbapi.data import fetch
query = """
{
entry(entry_id: "4HHB") {
struct { title }
exptl { method }
rcsb_entry_info { resolution_combined deposited_atom_count }
}
}
"""
data = fetch(query_type="graphql", query=query)Direct download endpoints:
https://files.rcsb.org/download/{PDB_ID}.pdbhttps://files.rcsb.org/download/{PDB_ID}.cifFor batch metadata retrieval, iterate over IDs and call fetch(pdb_id, schema=Schema.ENTRY); handle exceptions per-ID to keep pipelines robust. For large batches, consider rate limiting and caching to avoid repeated downloads.
If present in this repository, consult:
references/api_reference.md for advanced endpoint usage, query patterns, schema notes, rate limits, and troubleshooting.f5ef65b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.