Datasets
Create datasets from folders or assets and manage Unpublished changes.
Create datasets from folders or assets and manage Unpublished changes.

The SDK creates the same named reusable membership exposed by Data Space › Datasets › New Dataset: name, description, and reviewed Folder or Asset sources.
Outcome: Create datasets from folders or assets and manage Unpublished changes.
Creation accepts optional folder and asset IDs. add_sources() changes the draft membership. A dataset is version-first: project automation should attach a published snapshot rather than assuming current draft membership.
Execution surface
Pinned unitlab==3.0.0 application environment on Python 3.10+.
Identity
Least-privilege API key supplied through approved configuration.
Target resolution
Stable resource IDs and an explicitly bounded target set.
Success evidence
Typed return fields, server-side state, and downstream acceptance of the result.
Authentication or authorization fails
Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check.
Validation or entitlement rejects the operation
Correct the input or entitlement; do not retry an unchanged request.
A request times out
Inspect remote state before repeating a mutation because the server may have accepted it.
Asynchronous processing exceeds its deadline
Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.
Only part of a batch succeeds
Keep successful identifiers, isolate failed rows, and retry only the corrected subset.
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
Continue with Unitlab: training-data curation workflows
dataset = client.datasets.create(
"Training set",
description="Approved July cohort",
folder_ids=["FOLDER_ID"],
)
dataset.add_sources(asset_ids=["ASSET_ID"])
print(dataset.unpublished_changes())
print(dataset.list_items())