Work with Domino Volumes for NetApp ONTAP - enterprise-grade, multi-terabyte storage with near-instant snapshots. Covers volume creation, snapshot versioning with commit messages, cross-project sharing, mount paths (/mnt/netapp-volumes/ or /domino/netapp-volumes/), and the NetApp Volumes REST API. Use when managing large-scale data storage, needing fast no-copy snapshots, or integrating existing NetApp ONTAP infrastructure with Domino.
65
78%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Fix and improve this skill with Tessl
tessl review fix ./skills/netapp-volumes/SKILL.mdThis skill helps users work with Domino Volumes for NetApp ONTAP — enterprise-grade storage that mounts NetApp ONTAP filesystems directly into Domino workloads, enabling multi-terabyte scale with near-instant snapshotting.
Activate this skill when users want to:
A Domino NetApp Volume is:
netapp-storage storage class| Feature | NetApp Volumes | Domino Datasets |
|---|---|---|
| Storage backend | External NetApp ONTAP | Domino-managed NFS/EFS |
| Data scale | Multi-terabyte and beyond | Up to ~1 TB |
| Snapshot speed | ~3 seconds (any size) | Scales with data size |
| Snapshot storage cost | No extra space (redirect-on-write) | Duplicates physical data |
| Commit messages on snapshots | Supported | Not supported |
| Create new volume from snapshot | Read-only clones | Can create editable dataset |
| Cross-project sharing | Yes | Yes |
| Admin prerequisite | Admin must register filesystems | None |
| REST API | Full dedicated API | Via Domino Python SDK |
Use NetApp Volumes when:
Use Domino Datasets when:
An admin must register at least one NetApp filesystem (Kubernetes PVC with the netapp-storage label) before users can create volumes. Contact your Domino administrator if no filesystems are available.
import os
import requests
# Auth token from in-cluster token service
token = requests.get("http://localhost:8899/access-token").text.strip()
headers = {
"Authorization": f"Bearer {token}",
"Content-Type": "application/json"
}
api_url = os.environ["DOMINO_API_HOST"]
remotefs_url = os.environ["DOMINO_REMOTE_FILE_SYSTEM_HOSTPORT"]
# Look up your user ID (username is available in the DOMINO_USER_NAME env var)
user_resp = requests.get(
f"{api_url}/v4/users?userName={os.environ['DOMINO_USER_NAME']}",
headers=headers
).json()
user_id = user_resp[0]["id"]
# Create a volume (capacity is in bytes; grants is required)
response = requests.post(
f"{remotefs_url}/remotefs/v1/volumes",
headers=headers,
json={
"name": "large-training-data",
"description": "Multi-TB training dataset for vision models",
"filesystemId": "<filesystem-id>",
"capacity": 5_000_000_000_000, # 5 TB in bytes
"grants": [
{"targetId": user_id, "targetRole": "VolumeOwner"}
]
}
)
volume = response.json()
print(f"Created volume ID: {volume['id']}")# Attach a volume to a project
requests.post(
f"{remotefs_url}/remotefs/v1/rpc/attach-volume-to-project",
headers=headers,
json={
"volumeId": "<volume-id>",
"projectId": "<project-id>"
}
)Mount paths depend on your project type. Check which exists in your execution to determine your project type.
| Volume Type | Mount Path |
|---|---|
| Live volume | /mnt/netapp-volumes/<volume-name>/ |
| Snapshot by number | /mnt/netapp-volumes/snapshots/<volume-name>/<snapshot-number>/ |
| Snapshot by tag | /mnt/netapp-volumes/snapshot-tags/<volume-name>/<tag-name>/ |
| Volume Type | Mount Path |
|---|---|
| Live volume | /domino/netapp-volumes/<volume-name>/ |
| Snapshot by number | /domino/netapp-volumes/snapshots/<volume-name>/<snapshot-number>/ |
| Snapshot by tag | /domino/netapp-volumes/snapshot-tags/<volume-name>/<tag-name>/ |
There are two ways to access snapshots, with an important behavioral difference:
/snapshots/<volume-name>/<number>/ — Accessed by snapshot number. When a new snapshot is taken while a workspace is running, the new numbered directory appears immediately in that workspace without a restart./snapshot-tags/<volume-name>/<tag-name>/ — Accessed by tag name. Each snapshot has at most one active tag path — a symlink to its numbered snapshot directory. If you apply multiple tags to the same snapshot, only the most recently applied tag creates a path; earlier tags for that snapshot are not accessible by path. Tag paths for new snapshots are also not visible in a running workspace — you must restart the workspace for them to appear.Use the numbered path when you need to access a fresh snapshot from within a live workspace. Use the tagged path for stable, named references in reproducible runs.
Important: Each snapshot exposes only one tag path at a time — the most recently applied tag. Older tags on the same snapshot do not have accessible paths.
Important: Renaming a volume changes its mount path. Update any hardcoded paths in your code after renaming.
import os
if os.path.exists("/domino/netapp-volumes"):
print("DFS Project")
netapp_root = "/domino/netapp-volumes"
elif os.path.exists("/mnt/netapp-volumes"):
print("Git-Based Project")
netapp_root = "/mnt/netapp-volumes"import pandas as pd
# Git-Based Project
df = pd.read_parquet("/mnt/netapp-volumes/large-training-data/features.parquet")
# DFS Project
df = pd.read_parquet("/domino/netapp-volumes/large-training-data/features.parquet")
# Read from a specific snapshot tag
df = pd.read_parquet("/mnt/netapp-volumes/snapshot-tags/large-training-data/v2.0/features.parquet")# Write directly to the live volume
df.to_parquet("/mnt/netapp-volumes/large-training-data/processed/output.parquet", index=False)
# List files
import os
files = os.listdir("/mnt/netapp-volumes/large-training-data/")A snapshot is a read-only, immutable record of the volume's data at a specific point in time. NetApp snapshots use redirect-on-write — no additional storage is consumed at creation time.
Advantages over Domino Dataset snapshots:
v2.0, production, 2024-Q4)# Create snapshot with description and one or more tags (tagNames is an array)
response = requests.post(
f"{remotefs_url}/remotefs/v1/snapshots",
headers=headers,
json={
"volumeId": "<volume-id>",
"description": "Added Q4 customer records — 50M new rows",
"tagNames": ["v2.0"]
}
)
snapshot = response.json()
print(f"Snapshot ID: {snapshot['id']}")requests.post(
f"{remotefs_url}/remotefs/v1/snapshots/{snapshot_id}/tags",
headers=headers,
json={"name": "production"}
)# Snapshot tied to a specific Domino run for reproducibility
requests.post(
f"{remotefs_url}/remotefs/v1/rpc/create-snapshot-from-run",
headers=headers,
json={
"volumeId": "<volume-id>",
"runId": "<domino-run-id>",
"userId": "<user-id>",
"description": "post-training-run-42"
}
)requests.post(
f"{remotefs_url}/remotefs/v1/rpc/restore-snapshot",
headers=headers,
json={"snapshotId": "<snapshot-id>"}
)| Role | API Value | Capabilities |
|---|---|---|
| Reader | VolumeReader | View files and snapshots; mount as read-only |
| Editor | VolumeEditor | All Reader capabilities + modify description, create/delete snapshots, manage shared access, manage users |
| Owner | VolumeOwner | All Editor capabilities + update volume grants, request deletion |
# targetRole values: "VolumeOwner", "VolumeEditor", "VolumeReader"
requests.put(
f"{remotefs_url}/remotefs/v1/volumes/{volume_id}/grants",
headers=headers,
json=[
{"targetId": "<user-id>", "targetRole": "VolumeEditor"},
{"targetId": "<other-user-id>", "targetRole": "VolumeReader"}
]
)requests.post(
f"{api_url}/api/jobs/v1/jobs",
headers=headers,
json={
"projectId": "<project-id>",
"runCommand": "python train.py",
"hardwareTierId": "<hardware-tier-id>",
"environmentId": "<environment-id>", # required — use DOMINO_ENVIRONMENT_ID env var
"netAppVolumeIds": ["<volume-id>"],
"snapshotNetAppVolumesOnCompletion": True # auto-snapshot mounted volumes when job finishes
}
)Set snapshotNetAppVolumesOnCompletion: true to automatically take a snapshot of all mounted NetApp volumes when the job completes. This is the recommended approach for training jobs — it captures the exact state of the volume at the end of the run without requiring a separate API call.
# List all volumes accessible to you
volumes = requests.get(
f"{remotefs_url}/remotefs/v1/volumes",
headers=headers
).json()
for v in volumes["data"]:
capacity_tb = v["capacity"] / 1_000_000_000_000
print(f"{v['name']} — {capacity_tb:.1f} TB — ID: {v['id']}")
# List snapshots for a volume
snapshots = requests.get(
f"{remotefs_url}/remotefs/v1/snapshots",
headers=headers,
params={"volumeId": "<volume-id>"}
).json()
for s in snapshots["data"]:
print(f"Snapshot {s['id']} v{s['version']} — {s.get('description', '')} — tags: {[t['name'] for t in s.get('tags', [])]}")| Data Type | Storage |
|---|---|
| Large data of any file type where the entire filesystem should be versioned as a unit, with intentional snapshots | NetApp Volume |
| Small output files — charts, reports, model binaries | Artifacts / DFS (auto-versioned per file, not suitable for large files) |
| Code | Git / Project files |
# Always snapshot before modifying large volumes
requests.post(
f"{remotefs_url}/remotefs/v1/snapshots",
headers=headers,
json={
"volumeId": "<volume-id>",
"description": "Pre-processing baseline snapshot",
"tagNames": ["pre-processing-2024-01"]
}
)
# Then run your data transformation
process_data()# Parquet for tabular data (faster reads, smaller storage)
df.to_parquet("/mnt/netapp-volumes/dataset/data.parquet")
# Feather for fast pandas I/O
df.to_feather("/mnt/netapp-volumes/dataset/data.feather")
# HDF5 for large numerical arrays
import h5py
with h5py.File("/mnt/netapp-volumes/dataset/arrays.h5", "w") as f:
f.create_dataset("features", data=features_array)/mnt/netapp-volumes/my-volume/
├── raw/
│ ├── 2024-Q1/
│ └── 2024-Q2/
├── processed/
│ ├── features.parquet
│ └── labels.parquet
└── metadata/
└── schema.jsonReference snapshot tags (not snapshot IDs) in your training scripts so that tagged paths remain stable across runs:
# Reproducible reference using a tag
TRAINING_DATA = "/mnt/netapp-volumes/snapshot-tags/dataset/v2.0/"
df = pd.read_parquet(f"{TRAINING_DATA}/features.parquet")import dask.dataframe as dd
# Lazy read — no data loaded until .compute()
df = dd.read_parquet("/mnt/netapp-volumes/dataset/large_data.parquet")
result = df.groupby("category").mean().compute()/snapshots/<name>/<number>/) appears immediately, but the tag symlink path only becomes visible after restarting the workspace.pd.read_parquet(path, columns=["col1", "col2"])Before writing or verifying any API call, use the cluster swagger to confirm current endpoint paths and field names. Use public docs for workflow context and field explanations.
Get the cluster base URL: $DOMINO_API_HOST (injected by Domino into every workspace, job, and app).
Fetch the NetApp Volumes swagger spec (requires bearer token):
TOKEN=$(curl -s http://localhost:8899/access-token)
# The swagger UI is only accessible via the external cluster URL (not $DOMINO_API_HOST).
# Derive it from the JWT iss claim — works in any workspace type.
CLUSTER_URL=$(echo $TOKEN | cut -d'.' -f2 | python3 -c "
import sys, base64, json, re
p = sys.stdin.read().strip()
p += '=' * (-len(p) % 4)
print(re.sub(r'/auth/realms/.*', '', json.loads(base64.b64decode(p))['iss']))
")
curl -H "Authorization: Bearer $TOKEN" "$CLUSTER_URL/domino-netapp-volumes/swagger/doc.json"
# Browser UI (must be logged in): $CLUSTER_URL/domino-netapp-volumes/swagger/index.htmlPublic docs (workflow context and field explanations):
92a240b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.