> For the complete documentation index, see [llms.txt](https://docs.unitlab.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.unitlab.ai/documentation/datasets/create-a-dataset.md).

# Create a dataset

Create the dataset only after the candidate membership has been inspected. The dataset name and description should communicate why the cohort exists—not merely repeat its source folder.

### Before you make the change

* Confirm the membership question and owner.
* Resolve invalid, duplicate, or incomplete grouped items according to policy.
* Choose a naming and versioning convention that downstream teams can interpret.

![New Dataset dialog with name, description, Folders, Assets, search, and source selection](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F878wtKnyECkRNbmyCKhY%2Fdataset-create-modal.png?alt=media\&token=d09720b0-5aa9-4019-a940-d702f6a40e88)

*From Data Space › Datasets, choose New Dataset, name and describe the reusable cohort, then attach the exact Folders or Assets that passed curation. Source count is visible before creation.*

### Understand the product behavior

1. The user selects **New Dataset**.
2. The user enters a name and optional description.
3. The user chooses at least one folder or asset from the lazy-loaded, server-searched source picker.
4. Unitlab creates the mutable working draft.
5. The dataset detail page opens with its folders and assets.
6. The user publishes v1 before the dataset can be attached to a project.

Adding files to an existing dataset adds existing workspace folders/assets to the working draft. It is not an upload-directly-into-dataset operation.

### Create the reusable membership

{% stepper %}
{% step %}

#### 1. Open New Dataset

Go to Data Space › Datasets and choose New Dataset.
{% endstep %}

{% step %}

#### 2. Name the reusable cohort

Enter the purpose-oriented name and an optional description that explains its intended consumer.
{% endstep %}

{% step %}

#### 3. Choose Folders or Assets

Use the source tabs and search to select exact folders, individual assets, or grouped output; review the selected-source summary.
{% endstep %}

{% step %}

#### 4. Create and inspect

Choose Create Dataset, open the new dataset, and reconcile item count, data type, source provenance, and representative content.
{% endstep %}

{% step %}

#### 5. Control later additions

Use the dataset’s Add files or Attach Data surface; every accepted membership change creates or advances history rather than rewriting an earlier version silently.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision     | Production guidance                                                                                |
| ------------ | -------------------------------------------------------------------------------------------------- |
| Name         | Use a durable purpose-oriented name; keep environment and date in version metadata where possible. |
| Owner        | Assign a role accountable for membership and version changes.                                      |
| Source scope | Keep enough provenance to explain every member later.                                              |
| Groups       | Include the grouped unit when project work requires shared context.                                |

### Continue the operating flow

* Publish the first dataset version.
* Attach the exact version to a pilot project.
* Record future additions as a new version rather than silently changing history.
