# Welcome to Unitlab AI

Unitlab AI is an enterprise multimodal data platform for AI teams to curate, annotate, manage, version, and prepare training data at scale.

## Data

<table data-card-size="large" data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><h3><i class="fa-cloud-arrow-up" style="color:$primary;">:cloud-arrow-up:</i></h3></td><td><strong>Upload and connect data</strong></td><td>Ingest local files or governed cloud sources into Data Space.</td><td><a href="https://docs.unitlab.ai/documentation/data/data-upload">https://docs.unitlab.ai/documentation/data/data-upload</a></td></tr><tr><td><h3><i class="fa-folder-tree" style="color:$primary;">:folder-tree:</i></h3></td><td><strong>Organize assets</strong></td><td>Manage source assets, folders, lifecycle state, and metadata.</td><td><a href="https://docs.unitlab.ai/documentation/data/data-folders">https://docs.unitlab.ai/documentation/data/data-folders</a></td></tr><tr><td><h3><i class="fa-filter" style="color:$primary;">:filter:</i></h3></td><td><strong>Curate data</strong></td><td>Build precise cohorts with filters, embeddings, tags, and Collections.</td><td><a href="https://docs.unitlab.ai/documentation/data/data-curation">https://docs.unitlab.ai/documentation/data/data-curation</a></td></tr><tr><td><h3><i class="fa-layer-group" style="color:$primary;">:layer-group:</i></h3></td><td><strong>Version datasets</strong></td><td>Publish reusable, immutable dataset membership for production work.</td><td><a href="https://docs.unitlab.ai/documentation/datasets/dataset-versions-and-history">https://docs.unitlab.ai/documentation/datasets/dataset-versions-and-history</a></td></tr><tr><td><h3><i class="fa-object-group" style="color:$primary;">:object-group:</i></h3></td><td><strong>Build Data Groups</strong></td><td>Preserve related views and modalities as one structured unit of work.</td><td><a href="https://docs.unitlab.ai/documentation/data/data-groups-and-layouts">https://docs.unitlab.ai/documentation/data/data-groups-and-layouts</a></td></tr><tr><td><h3><i class="fa-circle-nodes" style="color:$primary;">:circle-nodes:</i></h3></td><td><strong>Explore embeddings</strong></td><td>Investigate similarity, duplicates, distribution, and outliers visually.</td><td><a href="https://docs.unitlab.ai/documentation/data/embeddings-similarity-and-outliers">https://docs.unitlab.ai/documentation/data/embeddings-similarity-and-outliers</a></td></tr></tbody></table>

## Platform

<table data-card-size="large" data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><h3><i class="fa-briefcase" style="color:$primary;">:briefcase:</i></h3></td><td><strong>Projects</strong></td><td>Configure data, instructions, ontology, workflow, QA, and releases in one operating boundary.</td><td><a href="https://docs.unitlab.ai/documentation/projects/projects-overview">https://docs.unitlab.ai/documentation/projects/projects-overview</a></td></tr><tr><td><h3><i class="fa-pen-to-square" style="color:$primary;">:pen-to-square:</i></h3></td><td><strong>Annotations</strong></td><td>Create and review image, video, text, document, audio, medical, and multimodal labels.</td><td><a href="https://docs.unitlab.ai/documentation/annotations/annotation-workbench">https://docs.unitlab.ai/documentation/annotations/annotation-workbench</a></td></tr><tr><td><h3><i class="fa-sitemap" style="color:$primary;">:sitemap:</i></h3></td><td><strong>Ontologies</strong></td><td>Define classes, properties, relations, validation, versions, and conditional logic.</td><td><a href="https://docs.unitlab.ai/documentation/ontologies/ontologies-overview">https://docs.unitlab.ai/documentation/ontologies/ontologies-overview</a></td></tr><tr><td><h3><i class="fa-diagram-project" style="color:$primary;">:diagram-project:</i></h3></td><td><strong>Workflows and queues</strong></td><td>Route human and model work through accountable stages, assignments, and exception paths.</td><td><a href="https://docs.unitlab.ai/documentation/workflows/workflows-overview">https://docs.unitlab.ai/documentation/workflows/workflows-overview</a></td></tr><tr><td><h3><i class="fa-building-shield" style="color:$primary;">:building-shield:</i></h3></td><td><strong>Workspaces and access</strong></td><td>Manage workspace boundaries, members, roles, security settings, API keys, and connected storage.</td><td><a href="https://docs.unitlab.ai/documentation/workspaces/workspaces-overview">https://docs.unitlab.ai/documentation/workspaces/workspaces-overview</a></td></tr><tr><td><h3><i class="fa-box-open" style="color:$primary;">:box-open:</i></h3></td><td><strong>Versioned releases</strong></td><td>Package approved annotations into reproducible outputs for downstream delivery.</td><td><a href="https://docs.unitlab.ai/documentation/releases/releases-overview">https://docs.unitlab.ai/documentation/releases/releases-overview</a></td></tr></tbody></table>

## Modalities

<table data-card-size="large" data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><h3><i class="fa-layer-group" style="color:$primary;">:layer-group:</i></h3></td><td><strong>Multimodal Annotation</strong></td><td>Keep related image, video, audio, text, document, and medical evidence in one governed work unit.</td><td><a href="https://docs.unitlab.ai/documentation/multimodal-annotations/multimodal-overview">https://docs.unitlab.ai/documentation/multimodal-annotations/multimodal-overview</a></td></tr><tr><td><h3><i class="fa-image" style="color:$primary;">:image:</i></h3></td><td><strong>Image Annotation</strong></td><td>Create boxes, polygons, masks, cuboids, landmarks, properties, and relations on still images.</td><td><a href="https://docs.unitlab.ai/documentation/annotations/image-annotation">https://docs.unitlab.ai/documentation/annotations/image-annotation</a></td></tr><tr><td><h3><i class="fa-video" style="color:$primary;">:video:</i></h3></td><td><strong>Video Annotation</strong></td><td>Build frame-accurate tracks, temporal events, timelines, and dynamic properties.</td><td><a href="https://docs.unitlab.ai/documentation/annotations/video-annotation">https://docs.unitlab.ai/documentation/annotations/video-annotation</a></td></tr><tr><td><h3><i class="fa-staff-snake" style="color:$primary;">:staff-snake:</i></h3></td><td><strong>Medical Annotation</strong></td><td>Annotate DICOM and medical imaging across synchronized multiplanar and sequence views.</td><td><a href="https://docs.unitlab.ai/documentation/annotations/medical-annotation">https://docs.unitlab.ai/documentation/annotations/medical-annotation</a></td></tr><tr><td><h3><i class="fa-microscope" style="color:$primary;">:microscope:</i></h3></td><td><strong>Pathology Annotation</strong></td><td>Label whole-slide images, tissue regions, cells, nuclei, biomarkers, and slide properties.</td><td><a href="https://docs.unitlab.ai/documentation/annotations/pathology-annotation">https://docs.unitlab.ai/documentation/annotations/pathology-annotation</a></td></tr><tr><td><h3><i class="fa-satellite" style="color:$primary;">:satellite:</i></h3></td><td><strong>Geospatial Annotation</strong></td><td>Annotate satellite, aerial, and large-raster imagery with spatially grounded geometry.</td><td><a href="https://docs.unitlab.ai/documentation/annotations/geospatial-annotation">https://docs.unitlab.ai/documentation/annotations/geospatial-annotation</a></td></tr><tr><td><h3><i class="fa-font" style="color:$primary;">:font:</i></h3></td><td><strong>Text Annotation</strong></td><td>Create entities, nested spans, classifications, structured properties, and relations.</td><td><a href="https://docs.unitlab.ai/documentation/annotations/text-annotation">https://docs.unitlab.ai/documentation/annotations/text-annotation</a></td></tr><tr><td><h3><i class="fa-file-pdf" style="color:$primary;">:file-pdf:</i></h3></td><td><strong>Document Annotation</strong></td><td>Label native PDF text, images, tables, regions, pages, and document-level values.</td><td><a href="https://docs.unitlab.ai/documentation/annotations/document-and-pdf-annotation">https://docs.unitlab.ai/documentation/annotations/document-and-pdf-annotation</a></td></tr><tr><td><h3><i class="fa-waveform" style="color:$primary;">:waveform:</i></h3></td><td><strong>Audio Annotation</strong></td><td>Label temporal events, speakers, transcripts, properties, waveform, and spectrogram context.</td><td><a href="https://docs.unitlab.ai/documentation/annotations/audio-annotation">https://docs.unitlab.ai/documentation/annotations/audio-annotation</a></td></tr></tbody></table>

## SDK/CLI

<table data-card-size="large" data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><h3><i class="fa-compass" style="color:$primary;">:compass:</i></h3></td><td><strong>Developer overview</strong></td><td>Choose the right automation surface and understand Unitlab resource boundaries.</td><td><a href="https://docs.unitlab.ai/api-sdk-cli/get-started/developer-overview">https://docs.unitlab.ai/api-sdk-cli/get-started/developer-overview</a></td></tr><tr><td><h3><i class="fa-python" style="color:$primary;">:python:</i></h3></td><td><strong>Python SDK</strong></td><td>Automate Unitlab with UnitlabClient and typed resource namespaces.</td><td><a href="https://docs.unitlab.ai/api-sdk-cli/sdks/unitlabclient-and-resource-namespaces">https://docs.unitlab.ai/api-sdk-cli/sdks/unitlabclient-and-resource-namespaces</a></td></tr><tr><td><h3><i class="fa-terminal" style="color:$primary;">:terminal:</i></h3></td><td><strong>CLI</strong></td><td>Run repeatable terminal workflows with human-readable or JSON output.</td><td><a href="https://docs.unitlab.ai/api-sdk-cli/clis/cli-overview-and-output-modes">https://docs.unitlab.ai/api-sdk-cli/clis/cli-overview-and-output-modes</a></td></tr><tr><td><h3><i class="fa-globe" style="color:$primary;">:globe:</i></h3></td><td><strong>HTTP APIs</strong></td><td>Integrate authenticated endpoints for projects, data, datasets, annotations, workflows, and more.</td><td><a href="https://docs.unitlab.ai/api-sdk-cli/apis/api-overview">https://docs.unitlab.ai/api-sdk-cli/apis/api-overview</a></td></tr><tr><td><h3><i class="fa-download" style="color:$primary;">:download:</i></h3></td><td><strong>Install SDK and CLI</strong></td><td>Install supported tooling and confirm the local Unitlab environment.</td><td><a href="https://docs.unitlab.ai/api-sdk-cli/get-started/install-the-python-sdk-and-cli">https://docs.unitlab.ai/api-sdk-cli/get-started/install-the-python-sdk-and-cli</a></td></tr><tr><td><h3><i class="fa-key" style="color:$primary;">:key:</i></h3></td><td><strong>Authentication</strong></td><td>Configure API credentials, environments, and secure automation access.</td><td><a href="https://docs.unitlab.ai/api-sdk-cli/get-started/authentication-and-configuration">https://docs.unitlab.ai/api-sdk-cli/get-started/authentication-and-configuration</a></td></tr></tbody></table>

## Solutions

<table data-card-size="large" data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><h3><i class="fa-industry" style="color:$primary;">:industry:</i></h3></td><td><strong>Manufacturing and mining</strong></td><td>Build governed visual inspection, safety, equipment, and defect programs.</td><td><a href="https://docs.unitlab.ai/solutions/industry-solutions-guides/manufacturing-and-mining">https://docs.unitlab.ai/solutions/industry-solutions-guides/manufacturing-and-mining</a></td></tr><tr><td><h3><i class="fa-heart-pulse" style="color:$primary;">:heart-pulse:</i></h3></td><td><strong>Medical AI</strong></td><td>Coordinate medical imaging, reports, ontologies, and specialist review.</td><td><a href="https://docs.unitlab.ai/solutions/industry-solutions-guides/medical-ai">https://docs.unitlab.ai/solutions/industry-solutions-guides/medical-ai</a></td></tr><tr><td><h3><i class="fa-camera" style="color:$primary;">:camera:</i></h3></td><td><strong>Multimodal operations</strong></td><td>Preserve multi-camera and multi-angle evidence as one unit of work.</td><td><a href="https://docs.unitlab.ai/solutions/multimodal-patterns-guides/multi-camera-and-multi-angle-work">https://docs.unitlab.ai/solutions/multimodal-patterns-guides/multi-camera-and-multi-angle-work</a></td></tr><tr><td><h3><i class="fa-wand-magic-sparkles" style="color:$primary;">:wand-magic-sparkles:</i></h3></td><td><strong>Human–AI annotation</strong></td><td>Combine model proposals, human judgment, review, and feedback with traceable accountability.</td><td><a href="https://docs.unitlab.ai/solutions/program-design-guides/human-ai-annotation-operations">https://docs.unitlab.ai/solutions/program-design-guides/human-ai-annotation-operations</a></td></tr><tr><td><h3><i class="fa-seedling" style="color:$primary;">:seedling:</i></h3></td><td><strong>Agriculture</strong></td><td>Build consistent crop, livestock, field, and multi-angle data programs.</td><td><a href="https://docs.unitlab.ai/solutions/industry-solutions-guides/agriculture">https://docs.unitlab.ai/solutions/industry-solutions-guides/agriculture</a></td></tr><tr><td><h3><i class="fa-car" style="color:$primary;">:car:</i></h3></td><td><strong>Autonomous systems</strong></td><td>Curate synchronized driving, robotics, temporal scenes, and rare events.</td><td><a href="https://docs.unitlab.ai/solutions/industry-solutions-guides/autonomous-systems">https://docs.unitlab.ai/solutions/industry-solutions-guides/autonomous-systems</a></td></tr></tbody></table>

***

> **Related Unitlab capability guides:** [multimodal data annotation](https://unitlab.ai/en/multimodal-annotation) · [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation) · [multimodal data curation](https://unitlab.ai/en/data-curation)


# Unitlab Product Documentation

Operate Unitlab from source ingestion through governed multimodal annotation and reproducible delivery.

This documentation is organized around the objects and responsibilities a production team operates: data, datasets, projects, annotations, multimodal context, ontologies, workflows, queues, releases, collaboration, security, and workspaces. It is grounded in the current Unitlab platform source dated July 29, 2026 and current product screens.

## Watch Unitlab AI in action

Unitlab helps AI teams curate, annotate, manage, version, and prepare multimodal training data at enterprise scale. This overview shows connected video, document, audio, image, workflow, review, and versioned dataset operations in one enterprise platform.

{% embed url="<https://cdn.prod.website-files.com/651fc1beafe23dfe4999151d/6a7610b174ee3346f6af02c1_unitlab-multimodal-data-platform-explainer-1080p.mp4>" %}

[Open the full-screen Unitlab AI explainer](https://cdn.prod.website-files.com/651fc1beafe23dfe4999151d/6a7610b174ee3346f6af02c1_unitlab-multimodal-data-platform-explainer-1080p.mp4).

## Follow the production lifecycle

```mermaid
flowchart LR
  A["Workspace"]
  B["Data"]
  C["Dataset version"]
  D["Project"]
  E["Annotate + review"]
  F["Release"]
  A --> B
  B --> C
  C --> D
  D --> E
  E --> F
```

## Documentation map

| Area                   | Use it to                                                                         |
| ---------------------- | --------------------------------------------------------------------------------- |
| Get Started            | Orient the team, understand the object model, and prove one end-to-end cycle.     |
| Data                   | Ingest, organize, curate, group, and explore durable source data.                 |
| Datasets               | Publish reusable membership versions and control project attachments.             |
| Projects               | Configure the operating boundary for policy, work, quality, and delivery.         |
| Annotations            | Use the shared Workbench and modality-native editors.                             |
| Multimodal Annotations | Preserve multiple views or modalities as one contextual work unit.                |
| Ontologies             | Govern classes, geometry, properties, relations, validation, and schema versions. |
| Workflows              | Design human, model, review, rework, and completion routes.                       |
| Queues                 | Allocate, prioritize, monitor, and recover individual and batch work.             |
| Releases               | Freeze approved output and validate it downstream.                                |
| Collaboration          | Coordinate members, roles, assignments, feedback, and ownership.                  |
| Security               | Protect identities, credentials, sensitive data, and privileged operations.       |
| Workspaces             | Govern the top-level resource, identity, storage, and capacity boundary.          |

{% hint style="warning" %}
Unitlab configuration is connected. Changes to source membership, Data Groups, ontology, workflow, permissions, models, or export mapping can affect active work and downstream data. Validate material changes on a controlled cohort and retain an operational record.
{% endhint %}

***

> **Related Unitlab capability guides:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# Welcome to Unitlab

Understand Unitlab’s role as an enterprise multimodal data-production platform.

Unitlab brings durable data, annotation policy, human and model work, review, and reproducible delivery into one operating model. The platform is designed for teams that need to produce trustworthy multimodal datasets without losing the context in which a labeling decision was made.

{% hint style="info" %}
**Use this area when:** you are evaluating the platform, onboarding a team, or deciding where a new production program belongs.
{% endhint %}

### How this area fits into production

```mermaid
flowchart LR
  A["Workspace"]
  B["Data"]
  C["Dataset version"]
  D["Project"]
  E["Annotate + review"]
  F["Release"]
  A --> B
  B --> C
  C --> D
  D --> E
  E --> F
```

![Unitlab projects overview in dark mode](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FOJlS4AoR0WSh2nHAeEDJ%2Fprojects-overview.png?alt=media\&token=cbc0a0a6-0c64-4bfd-a222-43d71439c575)

*Projects are the operating boundary that connects data, ontology, workflow, queues, annotation, review, and release delivery.*

### What this area controls

Unitlab connects four operating layers:

1. **Data Space** — ingest, connect, organize, filter, group, explore, and version raw or curated data.
2. **Annotation** — annotate images, video, audio, text, medical data, documents, and grouped multimodal cases.
3. **Quality operations** — define ontologies, route tasks through workflows, assign people, review outcomes, manage issues, and preserve instructions.
4. **AI and automation** — use interactive labeling assistance, tracking, model stages, external AI models, API keys, a Python SDK, and a CLI.

The platform is best understood as a lifecycle:

```
Raw files or cloud data
        ↓
Assets and folders
        ↓
Curation, filtering, embeddings, and multimodal grouping
        ↓
Versioned dataset
        ↓
Project + ontology + workflow
        ↓
Model assistance + human annotation + review
        ↓
Versioned release with annotations, files, metadata, and splits
        ↓
Training, evaluation, traceability, or another controlled project
```

This operating model matters because most training-data failures are not drawing-tool failures. They come from ambiguous label definitions, missing context, weak assignment rules, unreviewed model output, accidental dataset changes, and an inability to reproduce the exact data used by a model. Unitlab provides product surfaces for each part of that operating problem.

### Workspace navigation

The Unitlab workspace is organized into:

* **Annotation:** Projects, Workflows, Ontologies
* **Data Space:** Assets, Datasets, Releases
* **AI Suite:** My AI Models, Public AI Models
* **Workspace operations:** Documentation/Instructions, Members, Settings, cloud storage, roles, and API keys

Each area uses the same core concepts—projects, data, ontologies, workflows, queues, review, and releases—so teams can standardize how training data moves from source to production-ready output.

### Start with the right page

| Decision            | Production guidance                                      |
| ------------------- | -------------------------------------------------------- |
| Orient a new user   | Continue to Platform navigation.                         |
| Design a solution   | Read the Unitlab object model before creating resources. |
| Prove the lifecycle | Run the end-to-end quickstart with representative data.  |
| Approve scale       | Use the production-readiness review after the pilot.     |

### Operating boundary

* Data Space owns durable source resources and curation.
* Datasets own reusable, versioned membership.
* Projects own work, policy, routing, and quality state.
* Releases own frozen downstream delivery.

### A production-ready handoff

A new user can name the current workspace, identify the correct starting area for their role, and explain how one source item becomes reviewed release output.

***

> **Continue with Unitlab:** [cross-modal annotation workflows](https://unitlab.ai/en/multimodal-annotation) · [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# Platform navigation

Learn the workspace-level and project-level navigation model.

Navigation is part of operational safety. Before changing data, policy, workflow, or access, confirm the active workspace, project, lifecycle state, and role context.

![Projects list and workspace navigation](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FOJlS4AoR0WSh2nHAeEDJ%2Fprojects-overview.png?alt=media\&token=cbc0a0a6-0c64-4bfd-a222-43d71439c575)

*The workspace shell provides access to projects, Data Space, AI Suite, Releases, and administration; each project then exposes its own production controls.*

The Unitlab workspace is organized into:

* **Annotation:** Projects, Workflows, Ontologies
* **Data Space:** Assets, Datasets, Releases
* **AI Suite:** My AI Models, Public AI Models
* **Workspace operations:** Documentation/Instructions, Members, Settings, cloud storage, roles, and API keys

Each area uses the same core concepts—projects, data, ontologies, workflows, queues, review, and releases—so teams can standardize how training data moves from source to production-ready output.

### Use this in production

* Give each role one documented starting point.
* Use project navigation for work-specific decisions; use workspace navigation for shared resources and governance.
* Capture the workspace and project ID in operational records, not only the display name.

{% hint style="warning" %}
If an action appears missing, check the active workspace, project, workflow stage, selection, and permission before treating it as a product failure.
{% endhint %}

***

> **Explore related Unitlab capabilities:** [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# The Unitlab object model

Distinguish assets, groups, datasets, projects, tasks, and releases.

Most production mistakes begin by using the right capability at the wrong layer—for example, treating a folder as a dataset version or treating a project task as durable source data.

### The operating model

```mermaid
flowchart LR
  A["Asset"]
  B["Data Group"]
  C["Dataset version"]
  D["Project Data Unit"]
  E["Workflow task"]
  F["Release item"]
  A --> B
  A --> C
  B --> C
  C --> D
  D --> E
  E --> F
```

Several platform terms sound similar but serve different purposes. Keeping them distinct makes product explanations much clearer.

| Object          | Purpose                                                                                            | What changes over time                                                   |
| --------------- | -------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------ |
| **Asset**       | A source file or data item stored in or connected to Data Space                                    | Metadata, tags, folder placement, and curation state                     |
| **Folder**      | A source-oriented container for assets; it can also represent a cloud-backed location              | Contents, synchronization state, subfolders, and grouping                |
| **Data Group**  | A related set of files treated as one multimodal or multiview unit                                 | Group membership and tile layout                                         |
| **Dataset**     | A curated working collection assembled from folders or individual assets                           | Unpublished changes and explicit published versions                      |
| **Project**     | The operational environment where data is annotated and reviewed                                   | Attached sources, ontology copy, tasks, annotations, and status          |
| **Ontology**    | The reusable schema defining objects, properties, classifications, events, entities, and relations | Working edits, Live versions, nested logic, and project-specific history |
| **Workflow**    | The routing graph for Project, Annotate, Review, Model, Archive, and Complete stages               | Stage topology, assignments, and accepted/rejected paths                 |
| **Task**        | A workflow unit assigned to a person or made available to a queue                                  | Assignee, priority, state, review decision, and timeline                 |
| **Batch Queue** | A processing batch for uploaded or imported project data                                           | Processing, completion, and failure counts                               |
| **Release**     | A versioned annotation snapshot prepared for downstream use                                        | Version, split, export, annotation package, and associated files         |
| **AI model**    | A public or private model connected to annotation or workflow operations                           | Endpoint/configuration, validation, class mapping, and running state     |

### Data Units in the SDK

The SDK makes one additional distinction:

* A loose file becomes a `datasource` Data Unit.
* A Data Group becomes one `group` Data Unit whose `items` contain its tile summaries.
* Group member files are not duplicated as separate top-level project units.

This is an important design detail for multimodal cases. A four-camera inspection, a DICOM study with related views, or a video–document–audio case can remain one unit of work rather than four unrelated tasks.

### Use this in production

* Use folders for organization, datasets for reusable membership, and releases for delivery.
* Use Data Groups when multiple files must remain one work unit.
* Use stable resource IDs and explicit versions in automation and audit records.

***

> **Continue with Unitlab:** [multimodal data annotation](https://unitlab.ai/en/multimodal-annotation) · [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation) · [multimodal data curation](https://unitlab.ai/en/data-curation)


# End-to-end quickstart

Move a representative sample from ingestion to a validated release.

This quickstart is a controlled production rehearsal. Use a small cohort that contains normal cases, edge cases, invalid data, and any grouped or multimodal context required by real work.

### Before you make the change

* Name owners for source data, labeling policy, annotation operations, review, and downstream delivery.
* Prepare a representative sample; do not begin with full production volume.
* Confirm the required roles and an approved destination for the first release.

![Project Data page with attached content](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FsJhHsK7YkhabBtix4o95%2Fproject-data.png?alt=media\&token=06bd10a9-5d5a-4680-a807-d0fc697c37ac)

*Project Data is the control point for confirming exactly what entered the work system before annotation starts.*

### Understand the product behavior

**Problem:** The same item appears in four camera views. Treating each recording as an independent task can create inconsistent identity and defect decisions.

**Unitlab pattern:**

1. Ingest the four camera files with consistent identifiers.
2. Use auto-grouping to create one Data Group per inspected item.
3. Arrange four video tiles in a custom layout.
4. Define object classes and defect properties in the ontology.
5. Use tracking or interpolation within each view.
6. Route uncertain cases to a specialist-configured Review stage.
7. Publish a release that preserves the grouped case.

This workflow keeps the four views connected from curation through release, so annotators can make one case-level decision with all relevant visual context available.

### Pattern B: agriculture and repeated objects

**Problem:** A frame contains many similar cherries, and the same scene is captured from two angles.

**Unitlab pattern:**

1. Group the two camera views.
2. Annotate one trusted seed object.
3. Run Find Similar.
4. Adjust the confidence threshold and inspect candidates.
5. Accept only valid suggestions.
6. Continue tracking across frames if the data is video.

This workflow combines multiview context with seed-based assistance, reducing repetitive drawing while keeping acceptance under annotator control.

### Pattern C: medical study with contextual documentation

**Problem:** A finding must be understood across axial, sagittal, coronal, and 3D views, while a clinical document provides case context.

**Unitlab pattern:**

1. Upload or import the medical study and document.
2. Preserve study identity through grouping.
3. Use a custom layout with synchronized medical views and the document.
4. Annotate anatomy with geometry.
5. Record findings through conditional ontology properties.
6. Route the task through specialist review.
7. Export a controlled release after de-identification and governance checks.

This workflow keeps medical volumes and contextual documents together, allowing the specialist to review the case across synchronized views before release.

### Pattern D: customer interaction across text and audio

**Problem:** A written record contains entities and relations, while audio contains speakers, motion, noise, or transcription evidence.

**Unitlab pattern:**

1. Group the text/document and audio as one case.
2. Place them in a joint layout.
3. Annotate text entities and relations.
4. Mark audio event ranges or review a transcript.
5. Use item-level properties for case outcome or escalation status.
6. Validate completeness across both modalities before release.

This workflow preserves the relationship between written structure and temporal audio evidence within one task.

### Pattern E: model-assisted construction safety review

**Problem:** Many people must be tracked over time and classified by helmet status.

**Unitlab pattern:**

1. Use a Person class with a dynamic Helmet status property.
2. Seed or import initial detections.
3. Use Auto-Tracking across later frames.
4. Correct identity drift and status changes at keyframes.
5. Route to review; rejected tracks return to annotation.
6. Measure accepted tracks after rework, not raw predictions.

This workflow combines persistent tracks with properties that can change over time, allowing reviewers to distinguish identity errors from attribute changes.

### Run the first production cycle

{% stepper %}
{% step %}

#### 1. Ingest representative data

Upload or connect the source and wait for server-side processing to finish.
{% endstep %}

{% step %}

#### 2. Preserve context

Create Data Groups and a custom layout when related files must be seen together.
{% endstep %}

{% step %}

#### 3. Publish the contract

Align Instructions and ontology on classes, geometry, properties, ambiguity, and invalid-data rules.
{% endstep %}

{% step %}

#### 4. Build the operating path

Configure workflow stages, queues, assignment, review, rework, and escalation.
{% endstep %}

{% step %}

#### 5. Calibrate the team

Annotate and review the same hard examples until decisions are consistent.
{% endstep %}

{% step %}

#### 6. Deliver one release

Freeze the approved content, export it, and validate it in the downstream consumer.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision           | Production guidance                                                                    |
| ------------------ | -------------------------------------------------------------------------------------- |
| Pilot cohort       | Include representative failure and ambiguity cases, not only clean examples.           |
| Ontology state     | Test in Workbench before making the schema Live.                                       |
| Workflow exit      | Every accepted, rejected, invalid, failed, and escalated item needs a route and owner. |
| Release acceptance | The downstream consumer—not only the Unitlab UI—must validate the result.              |

{% hint style="warning" %}
Do not scale the pilot until the team can reproduce the same decision on hard examples and one release has passed downstream validation.
{% endhint %}

### Continue the operating flow

* Document the approved object model and ownership.
* Complete the Production-readiness review.
* Then expand source volume in controlled batches.

***

> **Related Unitlab capability guides:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation) · [multimodal data curation](https://unitlab.ai/en/data-curation)


# Production-readiness review

Make an evidence-backed go/no-go decision before production scale.

Production readiness is a review of connected controls, not a feature checklist. A project is ready only when access, data, policy, workflow, quality, automation, and delivery are all owned and exercised.

### The operating model

```mermaid
flowchart TB
  A["Access and ownership"]
  B["Data and grouping"]
  C["Instructions and ontology"]
  D["Workflow and queues"]
  E["Quality calibration"]
  F["Release validation"]
  A --> B
  B --> C
  C --> D
  D --> E
  E --> F
```

| Area                    | Unitlab capabilities                                                                                                                                                                                                                                       |
| ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Image annotation        | Bounding boxes, cuboids, polygons, masks, brush and eraser, keypoints, skeletons, lines, properties, relations, Magic Touch, Detect all objects, Find Similar, and Magic Crop                                                                              |
| Video annotation        | Frame and timeline navigation, persistent tracks, single- and multi-object Auto-Tracking, full bidirectional/forward/backward tracking, object and mask interpolation, dynamic class properties, and dynamic Item Properties                               |
| Audio annotation        | Waveform and spectrogram views, temporal events, segmentation, transcription-oriented workflows, and playback controls                                                                                                                                     |
| Text annotation         | Entities, relations, whole-item properties, configurable text windows, and source-text editing                                                                                                                                                             |
| PDF annotation          | Native multipage navigation, selectable PDF text, copied text values, bounding boxes from text or embedded image regions, page-aware history, and UUEF export                                                                                              |
| Medical annotation      | Grouped DICOM volumes, synchronized axial/sagittal/coronal and 3D views, clinical ontologies, and DICOM, NIfTI, and NRRD ingestion                                                                                                                         |
| Multiview Workbench     | Current File and Multiple Files modes, 1×1 through 4×4 grids, grouped custom layouts, resizable panels, and family-aware tools per tile                                                                                                                    |
| Multimodal data         | Data Groups, filename-based auto-grouping, custom layouts, mixed datasets, mixed releases, and grouped project work items                                                                                                                                  |
| AI assistance           | Magic Touch, Detect all objects, Find Similar, Magic Crop, AI captioning, Auto-Tracking, interpolation, batch auto-annotation, and workflow Model stages                                                                                                   |
| Ontologies              | Classes, class attributes, typed class properties, Item Properties, relations, annotation-side schema creation, required/dynamic fields, validation, unlimited conditional depth, Logic Map, Live switching, snapshots, and version history                |
| Workflows               | Project, Annotate, Review, Model, Archive, and Complete stages; approve/reject routes; specialist review; assignment; self-assignment; manager override; and task lifecycle controls                                                                       |
| Queues                  | Stage-based Task Queues, upload-oriented Batch Queues, assignment filters, priorities, claiming, submission, approval, rejection, and rework                                                                                                               |
| Quality operations      | Cohort filters, Grid/List/Embedding inspection, annotation-aware Display View controls, project instructions, item-level comments, owned issues, human review, notifications, project statistics, and member statistics                                    |
| Data curation           | Assets, folders, cloud folders, collections, activity history, metadata, advanced filters, Video/Frames browsing, filename-based auto-grouping, custom multimodal layouts, embeddings, natural-language search, duplicate detection, and outlier detection |
| Datasets                | Working drafts, immutable published versions, version history, restore, duplicate, exact-version project attachment, detachment, and restoration                                                                                                           |
| Releases                | Versioned annotation snapshots, public/private visibility, read-only viewing, cloning, train/validation/test splits, UUEF, Standard Bundle, family-specific formats, and stable tokenized URLs                                                             |
| AI models               | Public and private model catalogs, external-model registration, validation, input/output mapping, class mapping, and workflow integration                                                                                                                  |
| Accounts and workspaces | Email or Google authentication, email verification, password recovery, TOTP two-factor authentication, onboarding, workspace switching, settings, and quotas                                                                                               |
| Enterprise access       | Members, invitations, built-in and custom roles, granular permissions, and API-key management                                                                                                                                                              |
| Automation              | Python SDK and CLI access across projects, assets, datasets, ontologies, workflows, queues, embeddings, cloud storage, and releases                                                                                                                        |

### 25. Core product value

### One operating model across modalities

Image, video, audio, text, medical, PDF, and grouped multimodal cases share projects, ontologies, workflows, queues, review, and release concepts. Teams can standardize governance once and apply it across different forms of training data.

### Context-preserving annotation

Unitlab does more than accept multiple file formats. It groups related files into one work item and lets teams control how those sources appear together. Supported patterns include two- and four-camera views, synchronized medical views, document with audio, and video with PDF and audio.

### Reusable domain knowledge

Ontologies turn domain rules into reusable annotation structures. They support spatial objects, classifications, events, entities, relations, required properties, dynamic video properties, validation rules, and unlimited conditional nesting. Live versions, snapshots, history, Logic Map, and deleted-item restoration keep those structures manageable over time.

### Human and model work in one workflow

Project, Annotate, Review, Model, Archive, and Complete stages form forward and correction routes. Teams can add a second Review stage for specialist evaluation, configure eligible participants, and return rejected work to the appropriate annotation stage. Model output remains part of the same assignment, review, and release process as human-created annotations.

### Reproducible data handoffs

Datasets separate working changes from immutable published versions. Projects attach exact versions, and releases package annotation snapshots, source context, metadata, and data splits for downstream use. A model experiment can therefore reference the precise data, ontology, and release used to produce it.

### Integration with existing data operations

Cloud storage connections, public and private AI models, API keys, the Python SDK, the CLI, custom metadata, embeddings, and vector search allow Unitlab to fit into an existing AI data workflow while preserving the platform’s project, review, and release controls.

### Use this in production

* Record a named approver and evidence for every control area.
* Exercise rejection, invalid, failed, skipped, and escalation paths.
* Validate service identities, retries, model versions, and Batch Queue observability.
* Define the trigger for the next review: schema change, workflow change, new modality, new source, or material model update.

{% hint style="warning" %}
A material configuration change reopens the relevant part of the readiness review.
{% endhint %}

***

> **Continue with Unitlab:** [AI training-data annotation](https://unitlab.ai/en/data-annotation) · [multimodal data curation](https://unitlab.ai/en/data-curation)


# Data overview

Understand Data Space as the durable source and curation layer.

Data Space is where Unitlab keeps durable source resources independent from any one project. It supports ingestion, folders, lifecycle state, metadata, tags, grouping, filters, embeddings, similarity, duplicates, and outlier analysis.

{% hint style="info" %}
**Use this area when:** you need to bring data into Unitlab, organize it, understand its distribution, or prepare an exact cohort for a dataset.
{% endhint %}

### How this area fits into production

```mermaid
flowchart LR
  A["Upload or cloud"]
  B["Process"]
  C["Folders + Assets"]
  D["Curate + group"]
  E["Dataset membership"]
  A --> B
  B --> C
  C --> D
  D --> E
```

![Unitlab Data Space Asset library](https://storage.ghost.io/c/48/f4/48f4b614-5c29-430d-9cb3-e0b3f34395f3/content/images/2026/08/clean-assets-library.webp)

*The Asset library separates folders from individual assets and supports list, grid, and embedding views for different operational questions.*

### What this area controls

### Start with the right page

| Decision                   | Production guidance                                         |
| -------------------------- | ----------------------------------------------------------- |
| Bring in new data          | Use Data upload or Cloud providers.                         |
| Organize durable resources | Use Data folders and Data assets.                           |
| Prepare a cohort           | Use Data curation, metadata, tags, filters, and embeddings. |
| Preserve context           | Use Data Groups and custom layouts.                         |

### Operating boundary

* Data lifecycle state is not project workflow state.
* A saved filter is personal navigation state, not immutable dataset membership.
* Grouping changes the unit of work; validate it before project attachment.

### A production-ready handoff

The selected source cohort is processed, inspected, explainable, and ready to become a published dataset version.

### Product context

Unitlab helps AI teams curate, annotate, manage, version, and prepare multimodal training data at enterprise scale.

See [Unitlab’s multimodal data annotation platform](https://unitlab.ai/en/data-annotation) for the commercial overview and [Dataset Management](https://unitlab.ai/en/dataset-management) for managed downstream data operations.

***

> **Related Unitlab capability guides:** [cross-modal annotation workflows](https://unitlab.ai/en/multimodal-annotation) · [multimodal data curation](https://unitlab.ai/en/data-curation)


# Data upload

Upload current local files and monitor processing to a usable state.

A successful transfer is only the beginning of ingestion. Unitlab must recognize, process, and expose the resource before it is ready for curation or work.

### Before you make the change

* Confirm current file-family support and destination scope.
* Choose a small representative batch first.
* Remove secrets and regulated values from filenames and operational notes.

### Follow the interface

#### Upload to Data Space for reusable source data

![Data Space Assets page with the New Folder menu open](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FkaZggvW6dujbqKfqeO9r%2Fdata-space-create-menu.png?alt=media\&token=f486bf59-dc3b-4e26-8758-7b82c4ccd03d)

*From Data Space › Assets, open New Folder and choose Upload Data. The same menu also creates folders, connects cloud folders, and starts Auto-Groups.*

Use this route when the files should remain durable source resources that can be organized, curated, versioned in datasets, and reused across projects.

#### Upload inside a project for project-scoped work

![Project Data upload modal with Local, CLI, SDK, and Storage options](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FYIMKhOzBn8nrF0VZAPxY%2Fproject-upload-modal.png?alt=media\&token=3d793031-28d6-4529-bf7f-bfe87b2e98b8)

*From Project › Datasets, choose Upload to open the project uploader. Local, CLI, SDK, and Storage are distinct ingestion surfaces for the same project boundary.*

Use project upload only when the source belongs to that project’s operating context. Use Attach when the content already exists as an approved dataset or source that should retain its reusable lineage.

#### Qualify the processed result

![Data Space folders after ingestion](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F2KQcwfm4EbhGeUaipbJa%2Fassets-folders.png?alt=media\&token=3685bc54-6204-4a4e-a918-84c798aac689)

*After transfer and server-side processing, inspect the destination folder, file family, item count, representative viewers, and failed rows before downstream use.*

### Understand the product behavior

Data can be uploaded into the Asset library or directly to a Project. The SDK supports directories containing a mix of:

* images;
* video;
* audio;
* text;
* PDF documents;
* DICOM;
* NIfTI;
* NRRD.

A project upload creates one Batch Queue. The SDK exposes processing counts for total, completed, processing, and failed items. Completion means no items remain in processing; it does not imply that every item succeeded, so individual failures still need to be checked.

Medical uploads include a finalization step that groups related files into the user-facing medical volume before the data enters the project workflow.

### Common recognized file families

Current file detection includes:

| Family   | Common recognized extensions             |
| -------- | ---------------------------------------- |
| Image    | JPG, JPEG, PNG, GIF, WebP, BMP, ICO, SVG |
| Video    | MP4, AVI, MOV, WebM, MKV, M4V, WMV, FLV  |
| Audio    | MP3, WAV, OGG, AAC, FLAC, M4A            |
| Text     | TXT                                      |
| Medical  | DCM, NII, NII.GZ, NRRD                   |
| Document | PDF                                      |

The upload experience uses file detection to show previews and warnings, then confirms the saved family during ingestion. Unsupported files should appear as an explicit unsupported state rather than being forced into the wrong editor.

### Upload and qualify a batch

{% stepper %}
{% step %}

#### 1. Choose the destination

Open the intended Data Space folder or approved project upload entry point.
{% endstep %}

{% step %}

#### 2. Select the exact files

Confirm family, size, and expected item count before starting the transfer.
{% endstep %}

{% step %}

#### 3. Wait for processing

Separate transfer completion from server-side recognition and processing.
{% endstep %}

{% step %}

#### 4. Inspect the result

Open representative items and review type, metadata, dimensions, duration, status, and invalid state.
{% endstep %}

{% step %}

#### 5. Resolve exceptions

Use Batch Queue and row activity to identify failed or incomplete processing before continuing.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision        | Production guidance                                                                                                      |
| --------------- | ------------------------------------------------------------------------------------------------------------------------ |
| Where to upload | Use Data Space for durable reusable source; use project upload only when the source belongs to that operational context. |
| Batch size      | Prefer controlled batches that make failures attributable and recoverable.                                               |
| Ready state     | Do not treat a transferred file as usable until it opens correctly and processing is complete.                           |

### Continue the operating flow

* Organize the batch in Data folders.
* Add business context with metadata and tags.
* Create a dataset version only after validation.

***

> **Continue with Unitlab:** [training-data curation workflows](https://unitlab.ai/en/data-curation)


# Cloud providers

Connect approved cloud storage and import the exact source prefix.

Connected storage keeps large or governed source collections in their approved system of record while Unitlab registers and processes the content required for data production.

### Before you make the change

* Use an organization-owned cloud identity with least privilege.
* Identify the exact bucket, container, prefix, and ownership boundary.
* Agree how credentials are stored, rotated, revoked, and audited.

![Workspace Add Cloud Storage dialog](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FEceeKlJexxGEUrAINYfa%2Fcloud-storage-add.png?alt=media\&token=7f2c7aff-bb2a-476f-87cb-a8bd3a70540f)

*The connection is created in Workspace Settings › Cloud storage. After it shows Connected, return to Data Space › Assets and use New Folder › Add Cloud Folder to choose the connection and exact source scope.*

### Understand the product behavior

The current Add Cloud Storage dialog offers:

* Amazon S3;
* Google Cloud Storage;
* Azure Blob Storage;
* MinIO;
* DigitalOcean Spaces;
* Cloudflare R2;
* Wasabi;
* Backblaze B2;
* custom S3-compatible storage.

The exact credential fields vary by provider, but the current configuration model includes a display name, bucket or container, credentials, region where applicable, and an optional prefix or sub-path. A prefix can restrict Unitlab to a controlled part of a larger bucket.

The SDK can list safe storage metadata, browse a prefix, create a cloud-backed Unitlab folder, synchronize it, and import selected files or directory paths into a project. Cloud credentials are not returned by SDK list or browse operations.

### Cloud-folder user flow

Cloud folders are created from **New Folder ▾ → Add Cloud Folder**:

1. The user chooses a provider/integration.
2. The user selects an optional bucket sub-prefix and display name.
3. Unitlab creates a root folder representing that cloud location.
4. Opening it registers the first bucket level automatically if it has never been synchronized.
5. Immediate files appear as normal assets; child prefixes appear as child cloud folders.
6. The user can run **Sync** to refresh the current level.

Cloud folders reference customer storage; they do not copy the source bytes into Unitlab storage. Unitlab stores the object reference and a small preview where needed. Cloud folders display a provider badge, Synced state, and provider/resource path. Their contents are read-only mirrors for organization, so drag-move is disabled.

Cloud-folder Sync means “refresh this bucket prefix in Data Assets.” It is not project-to-dataset synchronization. A project receives cloud-backed content only through normal import/upload behavior or by attaching a frozen published dataset version.

### A safe cloud-connection pattern

For enterprise use, the connection should be treated as infrastructure, not as a convenient personal login:

1. Create a dedicated read-only or least-privilege identity for Unitlab.
2. Restrict it to the required bucket/container and prefix.
3. Avoid root or account-wide credentials.
4. Test with a small non-sensitive prefix.
5. Confirm that files, metadata, nested paths, and synchronization behave as expected.
6. Record the source system and connection owner.
7. Separate permission to manage cloud connections from permission to annotate data.

Workspace administrators can create, update, delete, test, and browse cloud connections from Workspace Settings. Connection administration requires cloud-storage permission and should remain separate from ordinary data browsing or annotation.

### Register a cloud source safely

{% stepper %}
{% step %}

#### 1. Create the connection

Select the provider and enter the approved connection details without exposing secrets in tickets or screenshots.
{% endstep %}

{% step %}

#### 2. Test access

Confirm Unitlab can enumerate only the intended storage scope.
{% endstep %}

{% step %}

#### 3. Browse to the prefix

Select the exact folder or object set; pay attention to provider-specific prefix behavior.
{% endstep %}

{% step %}

#### 4. Import the selection

Choose the destination and start registration or copy according to the connection mode.
{% endstep %}

{% step %}

#### 5. Reconcile results

Compare expected and processed counts, inspect representative items, and review Batch Queue failures.
{% endstep %}

{% step %}

#### 6. Record ownership

Document the source scope, identity owner, rotation process, and removal procedure.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision         | Production guidance                                                                             |
| ---------------- | ----------------------------------------------------------------------------------------------- |
| Identity         | Prefer a dedicated service identity over a personal user credential.                            |
| Scope            | Grant the smallest bucket or prefix required for the use case.                                  |
| Change control   | Treat provider, prefix, permission, and credential changes as production configuration changes. |
| Failure handling | Retry only after checking remote state so the operation is idempotent.                          |

{% hint style="warning" %}
Never place cloud secrets, signed URLs, or private source paths in public documentation, screenshots, comments, or logs.
{% endhint %}

### Continue the operating flow

* Validate recognized file families.
* Apply Data folder and metadata conventions.
* Review access when the source or owner changes.

***

> **Continue with Unitlab:** [Unitlab’s data curation platform](https://unitlab.ai/en/data-curation)


# Data folders

Organize durable source resources without confusing folders with versions.

Folders create an understandable source library. They are navigation and ownership boundaries—not immutable membership, workflow state, or downstream delivery.

![Data Space folder hierarchy](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F2KQcwfm4EbhGeUaipbJa%2Fassets-folders.png?alt=media\&token=3685bc54-6204-4a4e-a918-84c798aac689)

*Use folders to reflect stable source ownership or acquisition structure; use metadata and tags for dimensions that change more frequently.*

Data Space has three separate top-level pages:

* **Assets** — a Folders/Assets tabbed workspace library;
* **Datasets** — the standalone versioned dataset list;
* **Releases** — versioned project annotation snapshots.

The Assets page mirrors active tab and pagination in the URL. Folder detail presents one file-explorer list containing child folders, loose assets, and Data Groups. Breadcrumbs preserve the full path.

The Assets tab shows one source-of-truth row per asset rather than every clone created in every project. Source identity is reused when a published version is attached, so duplicate project copies do not inflate the workspace asset count.

### Lifecycle and bulk actions

Folders, assets, and datasets share three lifecycle views:

* **Active** — available for normal use;
* **Archived** — reversibly set aside;
* **Trash** — recoverable deletion state.

Hard deletion is available only from Trash. Selecting rows opens a floating action bar whose choices adapt to the current lifecycle. Depending on the object and state, actions include Move, Tags, Download, Archive, Unarchive, Restore, Delete, Delete forever, and Attach to project.

For Download, one selected workspace file downloads directly. Multiple selected files produce a manifest of secure file URLs rather than silently building an unbounded archive in the browser.

Deleting data that is still attached to projects is a two-phase flow: Unitlab first shows affected projects and requires the relevant source links to be detached. This reduces the risk that a workspace administrator deletes source data without seeing its downstream use.

### Use this in production

* Adopt a naming convention that communicates source and ownership without embedding sensitive data.
* Use folder depth sparingly; prefer searchable metadata for changing business dimensions.
* Before moving or archiving content, review dataset and project dependencies.
* Publish a dataset version when a folder-derived cohort must be reproduced.

***

> **Related Unitlab capability guides:** [Unitlab’s data curation platform](https://unitlab.ai/en/data-curation)


# Data assets

Inspect, search, select, and lifecycle individual source resources.

Assets are the durable source resources that Unitlab processes and presents across Data Space, datasets, projects, and releases. Asset operations should preserve traceability to the original source and current dependencies.

### Before you make the change

* Confirm the active lifecycle view: Active, Archived, or Trash.
* Use filters to resolve the exact target set before bulk action.
* Review dataset and project dependencies for any destructive lifecycle change.

![Unitlab Asset library list view](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FnjhYnLvSAjgnkvCf8QE9%2Fassets-library.png?alt=media\&token=e0b7397b-0f96-479d-8344-bf8a5f37d02b)

*List view is suited to operational details and bulk work; Grid is suited to visual inspection; Embedding is suited to distribution and similarity questions.*

### Understand the product behavior

The Asset library separates Folders and Assets and provides list, grid, and embedding views. It supports:

* upload;
* search;
* selection and bulk actions;
* list and grid presentation;
* sorting by name, size, modification date, and creator;
* folder organization;
* pagination;
* active, archived, and trash lifecycle states.

Search covers assets, folders, tags, and projects. The current list view makes source, size, modification date, and creator visible, which helps operational teams answer where data came from and who changed it.

Automation supports folder creation, nested subfolders, folder listing, asset upload, asset-level custom metadata, cloud-folder creation, and synchronization.

### Workspace Data Space navigation

Data Space has three separate top-level pages:

* **Assets** — a Folders/Assets tabbed workspace library;
* **Datasets** — the standalone versioned dataset list;
* **Releases** — versioned project annotation snapshots.

The Assets page mirrors active tab and pagination in the URL. Folder detail presents one file-explorer list containing child folders, loose assets, and Data Groups. Breadcrumbs preserve the full path.

The Assets tab shows one source-of-truth row per asset rather than every clone created in every project. Source identity is reused when a published version is attached, so duplicate project copies do not inflate the workspace asset count.

### Lifecycle and bulk actions

Folders, assets, and datasets share three lifecycle views:

* **Active** — available for normal use;
* **Archived** — reversibly set aside;
* **Trash** — recoverable deletion state.

Hard deletion is available only from Trash. Selecting rows opens a floating action bar whose choices adapt to the current lifecycle. Depending on the object and state, actions include Move, Tags, Download, Archive, Unarchive, Restore, Delete, Delete forever, and Attach to project.

For Download, one selected workspace file downloads directly. Multiple selected files produce a manifest of secure file URLs rather than silently building an unbounded archive in the browser.

Deleting data that is still attached to projects is a two-phase flow: Unitlab first shows affected projects and requires the relevant source links to be detached. This reduces the risk that a workspace administrator deletes source data without seeing its downstream use.

### Row actions and activity

Folder and asset row actions include Open, Rename, Move, Download, Archive, Tag, Activity, and Delete where the backing supports them. Opening **Activity** replaces the filter sidebar with a read-only, newest-first timeline showing who changed what, when, affected-file counts, and rename/move/metadata differences. Users can search or filter events and download the complete affected-file JSON for a large change.

The activity record remains available after the original subject is permanently deleted. Objects created before activity tracking was enabled can have an empty history because earlier events are not reconstructed.

### Inspect and operate an asset cohort

{% stepper %}
{% step %}

#### 1. Find the cohort

Search or filter by source, type, state, metadata, tag, size, dimensions, or dates.
{% endstep %}

{% step %}

#### 2. Choose the view

Use List for fields and bulk actions, Grid for visual QA, or Embedding for distribution.
{% endstep %}

{% step %}

#### 3. Inspect representative items

Open row activity and the native viewer before changing lifecycle state.
{% endstep %}

{% step %}

#### 4. Apply the intended action

Tag, archive, restore, or trash only the reviewed selection.
{% endstep %}

{% step %}

#### 5. Reconcile dependencies

Confirm expected dataset, project, and release behavior after the change.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision         | Production guidance                                                                                                      |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------ |
| Archive vs trash | Archive removes active operational visibility while retaining recoverable history; trash is a stronger lifecycle action. |
| Bulk scope       | The resolved item count is part of the approval decision.                                                                |
| Evidence         | Record the target filter, count, owner, and reason for material bulk changes.                                            |

### Continue the operating flow

* Curate with metadata and tags.
* Create Data Groups when files must remain together.
* Publish stable membership as a dataset version.

***

> **Related Unitlab capability guides:** [multimodal data curation](https://unitlab.ai/en/data-curation)


# Data curation

Build precise cohorts with filters, metadata, tags, collections, and frame granularity.

## Watch the curation workflow

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/data-curation-1.mp4>" %}

[Open the demo in a new tab](https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/data-curation-1.mp4).

Curation happens inside a Data Space folder. Explore resolves candidates with structured filters, visual views, and frame granularity; Collections save reviewed subsets inside that folder for later comparison, attachment, or dataset creation.

### Before you make the change

* Open the exact source folder—not the workspace-level Assets landing page.
* Write the cohort question and representative inclusion and exclusion examples.
* Decide whether the unit is a whole asset or an individual video frame before selecting files.

### Follow the interface

#### Curate inside the source folder

![Folder Explore view with List, Grid, Embedding, Collections, and advanced filters](https://storage.ghost.io/c/48/f4/48f4b614-5c29-430d-9cb3-e0b3f34395f3/content/images/2026/08/clean-data-curation-folder.webp)

*Inside a folder, Explore combines List, Grid, and Embedding views with asset type, granularity, source, tag, date, size, and dimension filters. The live result count is the scope of the next selection.*

Choose Video to curate whole video assets or Frames to expand videos into frame-level candidates. Filters narrow the current view; they do not create durable membership by themselves.

#### Save reviewed subsets as Collections

![Folder Collections tab and empty-state creation guidance](https://storage.ghost.io/c/48/f4/48f4b614-5c29-430d-9cb3-e0b3f34395f3/content/images/2026/08/clean-data-collections.webp)

*Collections belong to the current folder. Select files in Explore and use Add to Collections; then open Collections to review the saved subset and its item count.*

A Collection is useful for working cohorts and repeated inspection. Use a dataset version when membership must become an immutable, reusable production input.

#### Use embeddings as evidence, not as labels

![Embedding exploration view](https://storage.ghost.io/c/48/f4/48f4b614-5c29-430d-9cb3-e0b3f34395f3/content/images/2026/08/clean-data-curation.webp)

*Embedding view exposes clusters, gaps, duplicates, and outliers within the current folder. Human review determines whether proximity is relevant to the cohort decision.*

### Understand the product behavior

The **More Filters** sidebar is available from the Assets and Folders tabs and remains available inside folder detail. Its fixed sections are:

* **Asset Type:** Image, Video, Audio, Text, Medical, and Document;
* **Source:** Uploaded or Cloud storage;
* **Tags:** searchable tag selection;
* **Date:** asset creation date;
* **File Properties:** minimum/maximum size and minimum width/height.

The header badge counts active filter groups. The footer continuously reports the matching result count and displays **Updating…** while a new result set is loading. Added filters provide controls appropriate to their value: text operators, date or date-range inputs, boolean toggles, multiselect chips, color swatches, numeric range sliders, and searchable entity pickers for users, tags, folders, and datasets.

The complete additional-filter catalog is organized as follows:

| Category           | Filter groups and fields                                                                                                                                                                                                                            |
| ------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Data**           | File name, file type, extension, MIME type, asset ID; upload/created/modified/last-synced dates; width, height, resolution, aspect ratio, file size; video duration, frame count, FPS; medical modality, series count, study count, and slice count |
| **Metadata**       | Folder, subfolder, collection, dataset membership, sequence membership, group membership; storage provider, bucket, import source, data source; uploaded by, created by, updated by; asset and system tags                                          |
| **Embeddings**     | Similar assets/images/frames; text-to-image and natural-language search; embedding cluster and cluster membership; diverse and representative samples; similarity and distance thresholds                                                           |
| **Image Features** | Sharpness, noise level, exposure; brightness, contrast, saturation, dominant colors; entropy, complexity, texture density, edge density; orientation, foreground coverage, and background coverage                                                  |
| **Video & Frames** | Frame number, timestamp, keyframes only; scene changes, shot boundaries, segments; frame tags, frame attributes, and frame quality                                                                                                                  |
| **Data Quality**   | Exact and near duplicates; outliers and anomalies; blurry, low-resolution, corrupted, and incomplete assets; missing, invalid, and empty metadata                                                                                                   |
| **Search**         | Metadata search, tag search, natural-language query, and embedding search                                                                                                                                                                           |

Filters that operate on stored asset fields and computed metrics narrow the results immediately. These include extension, MIME type, asset ID, uploaded/created/modified dates, creator, orientation, low resolution, blurry assets, dominant color, duration, frame count, FPS, series/slice count, and numeric image-curation metrics. Catalog entries carrying a **Preview** badge can be configured in the interface but do not yet narrow the result set.

On the Folders tab, an active asset filter keeps a folder only when its subtree contains at least one matching asset. Filters run before duplicate source identities are collapsed, so matching behavior remains stable across project clones of the same underlying data.

Typical curation flows include finding low-resolution uploads, isolating studies from one source bucket and date range, locating long videos, selecting a brightness or sharpness range, finding duplicate-heavy regions, and building a balanced sample from embedding clusters.

Folders preserve source organization rather than label truth. A folder such as `warehouse-camera-07` describes provenance; a decision such as `forklift_present` belongs in an ontology or annotation.

### Folder Explore, Collections, and frame granularity

Folder detail includes **Explore** and **Collections**:

* **Explore** is the normal list/grid/embedding file browser.
* **Collections** stores static, folder-scoped curated file sets.

Selecting files or child folders in Explore enables **Add to Collections** or **Remove from Collections**. A new collection can be created with the current selection, or the selection can be added to an existing collection. Collection actions include View in explorer, Add to dataset, and Download. Removing or deleting a collection never deletes its files.

For video folders, Explore can switch between **Video** and **Frames** granularity:

* **Video** shows one entry per source file and keeps the standard List/Grid/Embedding view switcher.
* **Frames** expands videos with extracted-frame metadata into a paginated frame-card sequence. The breadcrumb reports the frame count and the normal view switcher is disabled while frame browsing is active.
* Native videos without extracted frame metadata appear as one poster entry.
* Selection remains file-level: selecting any frame selects its parent video and highlights all visible sibling frames. Collection, dataset, download, and bulk actions therefore operate on the complete video rather than an arbitrary frame subset.

Inside a collection-scoped Explore view, search and advanced filters continue to apply. A dismissible collection chip identifies the active scope.

### Custom metadata and tags

Tags and custom metadata allow source facts to travel with assets. Appropriate metadata includes capture site, sensor, acquisition date, device version, customer partition, consent state, or study identifier. It should not silently encode ground truth that annotators are expected to determine from the data.

The SDK supports setting custom metadata during upload and updating it later, including explicitly setting it to `null`.

### Build a reviewable cohort

{% stepper %}
{% step %}

#### 1. Open the folder

Navigate Data Space › Assets, open the folder that owns the source cohort, and stay on Explore.
{% endstep %}

{% step %}

#### 2. Choose asset or frame granularity

Keep Video for whole-video decisions or choose Frames when the collection is explicitly frame-level.
{% endstep %}

{% step %}

#### 3. Apply advanced filters

Combine asset type, source, tags, date, size, and dimensions; confirm the result count before selection.
{% endstep %}

{% step %}

#### 4. Inspect across views

Use List for fields and bulk scope, Grid for visual QA, and Embedding for distribution, similarity, and outliers.
{% endstep %}

{% step %}

#### 5. Create a Collection

Select the reviewed items in Explore, choose Add to Collections, create or select the named Collection, then verify it in the Collections tab.
{% endstep %}

{% step %}

#### 6. Promote durable membership

Create a dataset from the approved folder, Collection, assets, or groups when the cohort must be versioned and reused.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision          | Production guidance                                                                                                          |
| ----------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| Filter result     | Temporary query state; record the conditions when they matter to a review.                                                   |
| Collection        | Named folder-scoped working subset for review and reuse during curation.                                                     |
| Dataset version   | Immutable production membership for projects, automation, and provenance.                                                    |
| Frame granularity | Changes the unit selected from video; decide before creating the Collection.                                                 |
| Metadata or tag   | Use durable source/business context and controlled operational categories; keep annotation outputs in the ontology contract. |

### Continue the operating flow

* Create Data Groups before versioning when multiple files form one work unit.
* Create and inspect the dataset membership.
* Attach the explicit dataset version to a project.

## Distribution and quality analysis

### Duplicate and near-duplicate review

![Repeated frame groups with one retained representative](/files/dcHD8DDBB5yrv7WL57R9)

*Use duplicate signals to identify redundant regions, then review representatives before excluding data.*

### Metadata filters and saved views

![A multimodal sample set narrowed to a focused saved view](/files/ogkR9SJXTuSqLhvVWe7G)

*Combine source facts and computed characteristics into a reviewable cohort; save durable membership as a Collection or dataset version.*

### Dataset balancing

![An overrepresented road dataset changed into a balanced condition mix](/files/dP5x7xP7DPfntzyPomb2)

*Compare representation across conditions, classes, sources, and domains before annotation or training.*

### Outlier and quality detection

![Blurred, dark, corrupted, and domain-shifted samples isolated for review](/files/WM5aiRXoE8qWBobvkMgy)

*Outliers are review candidates, not automatic deletion decisions. Confirm whether each sample is erroneous, rare-but-valid, or strategically important.*

### Product context

Unitlab helps AI teams curate, annotate, manage, version, and prepare multimodal training data at enterprise scale.

See [Unitlab’s multimodal data annotation platform](https://unitlab.ai/en/data-annotation) for the commercial overview and [Dataset Management](https://unitlab.ai/en/dataset-management) for the downstream dataset workflow.

***

> **Explore related Unitlab capabilities:** [Unitlab’s data curation platform](https://unitlab.ai/en/data-curation)


# Data Groups and layouts

Preserve cross-view and cross-modal context as one unit of work.

## Watch multimodal annotation in one workspace

See how Unitlab preserves connected video, documents, audio, images, saved layouts, workflow state, review, and reproducible dataset context.

{% embed url="<https://cdn.prod.website-files.com/651fc1beafe23dfe4999151d/6a7610b174ee3346f6af02c1_unitlab-multimodal-data-platform-explainer-1080p.mp4>" %}

[Open the full-screen Unitlab AI explainer](https://cdn.prod.website-files.com/651fc1beafe23dfe4999151d/6a7610b174ee3346f6af02c1_unitlab-multimodal-data-platform-explainer-1080p.mp4).

Data Groups change the unit of work from one file to a related set of files. A custom layout then determines how those members are presented together in the Workbench.

### Before you make the change

* Define the real-world unit that must remain together.
* Identify the filename or metadata key that reliably joins members.
* Specify required, optional, duplicate, incomplete, and ambiguous group behavior.

### Follow the interface

#### Start Auto-Groups from the Asset library

![Assets page with Add Auto-Groups in the New Folder menu](https://storage.ghost.io/c/48/f4/48f4b614-5c29-430d-9cb3-e0b3f34395f3/content/images/2026/08/clean-data-space-create-menu.webp)

*Open Data Space › Assets or an eligible folder, open New Folder, then choose Add Auto-Groups. This starts a non-destructive grouping flow from the selected source scope.*

#### Configure grouping keys and tile match rules

![Auto-Groups Configure step with source folder, grouping keys, group name, and tile rules](https://storage.ghost.io/c/48/f4/48f4b614-5c29-430d-9cb3-e0b3f34395f3/content/images/2026/08/clean-auto-groups-configure.webp)

*Configure defines the source folder, grouping key, group-name template, expected tiles, match patterns, and optional exclusions. Auto-detect can propose a configuration from filenames.*

Treat every proposed regular expression as production logic. Inspect representative matches, naming collisions, missing members, and unexpected suffixes before advancing.

#### Set completeness rules and inspect the estimate

![Auto-Groups Rules step with required tiles and estimated valid groups](https://storage.ghost.io/c/48/f4/48f4b614-5c29-430d-9cb3-e0b3f34395f3/content/images/2026/08/clean-auto-groups-rules.webp)

*Rules controls the minimum matched tiles, required members, and incomplete-group behavior. The estimate and example outcomes show whether the policy produces valid groups before anything is created.*

#### Build the saved Workbench layout

![Auto-Groups Layout step with Grid, List, and Custom options and a three-panel custom builder](https://storage.ghost.io/c/48/f4/48f4b614-5c29-430d-9cb3-e0b3f34395f3/content/images/2026/08/clean-auto-groups-custom-layout.webp)

*Layout offers Grid, List, and Custom. In Custom, each expected tile becomes a draggable and resizable panel; Builder and JSON edit the same percentage-based layout.*

The saved layout follows the Data Group into dataset versions, project attachment, the grouped Workbench, and release context. Overlap is invalid and must be resolved before review.

### Understand the product behavior

Unitlab can automatically group related files into multiview or multimodal units using filename patterns. The current UI shows a four-step wizard:

1. **Configure**
2. **Rules**
3. **Layout**
4. **Review**

Available layout options are:

* **Grid** — related files arranged as a balanced grid;
* **List** — a primary tile with remaining files listed below;
* **Custom** — a draggable, resizable workspace for specialized arrangements.

A custom layout can place video, document, and audio tiles in one workspace. Layout determines what information an annotator can compare without leaving the task.

The SDK can:

* suggest a filename grouping pattern;
* estimate the outcome before creating groups;
* create groups from an explicit configuration;
* compile a literal filename template such as `{patient_id}_{view}`;
* map expected tile values such as `L_CC`, `R_CC`, `L_MLO`, and `R_MLO`.

### Auto-Grouping user flow

1. The user opens **Add auto-groups** from Assets or a folder.
2. **Configure:** choose the source folder. Unitlab analyzes filenames and can propose grouping keys, group-name template, and tiles; **Auto-detect** reruns the suggestion.
3. **Rules:** define the minimum matched tiles, required tiles, and incomplete-group handling. A live estimate shows valid, skipped, and conflicting examples. The user cannot continue without at least one valid estimated group.
4. **Layout:** arrange a panel per tile using Grid, List, or Custom/free-form placement. Panels can be dragged, resized, nudged, enlarged, or minimized. Overlap is highlighted and blocks continuation. An optional JSON view edits the same layout.
5. **Review:** inspect the final rules and static layout preview.
6. **Create:** Unitlab creates a new sibling `<source>_grouped` folder and leaves the source folder unchanged.

Every grouping run is non-destructive and creates a new grouped folder, even when the same configuration is run again. The saved percentage-based layout becomes the single layout used by preview, published dataset version, attached project group, and grouped Workbench.

For folders above 5,000 files or estimates above 1,000 groups, Unitlab starts an asynchronous grouping job and reports that it has begun. The SDK applies a stricter guard and may advise splitting the source before retrying.

### From Custom Layout to multimodal annotation

The layout is part of the Data Group, not a temporary preference inside the annotation screen. It moves with the group through the complete data lifecycle:

```
Source files
    ↓ Auto-group by filename rules
Data Group + saved tile layout
    ↓ Add to dataset and publish
Immutable dataset version with grouped tiles
    ↓ Attach to project
One grouped project work item
    ↓ Open in Workbench
Multimodal annotation in the saved layout
    ↓ Release
Grouped annotation context preserved in UUEF
```

Each configured tile retains its own data family and native viewer. A single layout can therefore combine video playback, PDF pages, audio waveforms, images, or medical views while the surrounding project supplies one ontology, workflow state, assignee, comment history, and review route for the grouped case.

The grouped Workbench follows the saved Grid, List, or Custom arrangement. Annotators activate one tile at a time for editing; the other tiles remain visible as read-only context. Previous/next navigation treats the complete group as one work item, and group members do not appear again as unrelated loose tasks.

Custom-layout UX is designed to prevent invalid arrangements before creation:

* each expected tile has one panel;
* panels can be dragged, resized, nudged, enlarged, or minimized;
* panel coordinates and dimensions are percentage-based so the layout scales with the Workbench;
* overlap is highlighted and blocks continuation;
* the visual editor and optional JSON view edit the same layout;
* Review shows the final rules and a static layout preview before groups are created.

### Create and qualify groups

{% stepper %}
{% step %}

#### 1. Open Add Auto-Groups

Start from Data Space › Assets or the intended source folder. Auto-Groups never rewrites the source folder.
{% endstep %}

{% step %}

#### 2. Configure

Choose the source folder, accept or correct suggested grouping keys, group naming, tile names, match rules, and exclusions.
{% endstep %}

{% step %}

#### 3. Rules

Set the minimum matched tiles, required tiles, and incomplete-group policy; continue only when the estimate contains valid, explainable groups.
{% endstep %}

{% step %}

#### 4. Layout

Choose Grid, List, or Custom. In Custom, place one panel per expected tile, resize or nudge panels, and clear all overlap errors.
{% endstep %}

{% step %}

#### 5. Review and create

Inspect the final rules and static layout preview. Creation writes a new sibling grouped folder and preserves the original source.
{% endstep %}

{% step %}

#### 6. Test in a pilot project

Publish grouped membership, attach it to a pilot, and verify the saved tile roles, active editor, navigation, workflow state, and release output.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision         | Production guidance                                                                        |
| ---------------- | ------------------------------------------------------------------------------------------ |
| Grouping key     | Prefer stable identifiers or metadata over fragile positional assumptions.                 |
| Missing member   | Decide whether the group remains usable, becomes invalid, or routes to exception handling. |
| Layout order     | Make the order meaningful and stable across users and releases.                            |
| Release contract | Confirm the downstream consumer can preserve or reconstruct group context.                 |

### Continue the operating flow

* Publish grouped membership as a dataset version.
* Configure the Grouped Workbench project flow.
* Train annotators on active versus passive panels.

### Product context

Unitlab helps AI teams curate, annotate, manage, version, and prepare multimodal training data at enterprise scale.

See [Unitlab’s multimodal data annotation platform](https://unitlab.ai/en/data-annotation) for the commercial overview of grouped and synchronized training-data workflows.

***

> **Continue with Unitlab:** [cross-modal annotation workflows](https://unitlab.ai/en/multimodal-annotation) · [Unitlab’s data curation platform](https://unitlab.ai/en/data-curation)


# Embeddings, similarity, and outliers

Investigate distribution, nearest neighbors, duplicates, and outliers.

Embedding tools help experts find the data worth inspecting. They do not replace ontology decisions, review, or source-aware judgment.

### Before you make the change

* Choose a source scope and compatible embedding space.
* Know whether the question is coverage, similarity, duplication, or anomaly discovery.
* Keep provenance and business metadata visible during interpretation.

![Find Similar result cohort](https://storage.ghost.io/c/48/f4/48f4b614-5c29-430d-9cb3-e0b3f34395f3/content/images/2026/08/clean-find-similar-results.webp)

*Nearest-neighbor results become useful only after a human checks source context, false positives, and the intended action on the cohort.*

### Understand the product behavior

The embedding view projects embeddable assets into a visual space and supports drag or crop-style region selection. Users can select a cluster and review its associated thumbnails before adding, excluding, or organizing the selected assets.

Useful curation patterns include:

* locating dense duplicate or near-duplicate regions;
* sampling from visually distinct clusters;
* finding outliers and rare conditions;
* comparing source domains;
* building a more diverse pilot dataset;
* selecting negative examples near a target class.

The two-dimensional plot is a navigation aid, not a guarantee that every nearby point is semantically identical. Operators should inspect the actual assets before turning a region into a dataset.

### Custom embedding spaces and vector search

The SDK can create named embedding spaces with a specified dimensionality and optional model name, upload vectors for an asset or a video frame, bulk-upsert vectors, search by a query vector, and optionally scope the search to a project or level.

This allows teams to use Unitlab’s curation interface while retaining control over the embedding model and vector space used for similarity operations.

### Search, duplicates, outliers, and UMAP

Unitlab includes visual similarity search, natural-language CLIP-style search, duplicate detection, outlier detection, quality-inconsistency checks, and UMAP visualization. These embedding and visual exploration experiences are available for image and video data.

### Turn vector evidence into a controlled cohort

{% stepper %}
{% step %}

#### 1. Inspect global distribution

Open Embedding view before zooming into a cluster so the local region has context.
{% endstep %}

{% step %}

#### 2. Choose a seed or region

Select a representative item, freehand region, or query vector.
{% endstep %}

{% step %}

#### 3. Review candidates

Compare neighbors, duplicates, or outliers against metadata and the native asset viewer.
{% endstep %}

{% step %}

#### 4. Separate signal from decision

Remove false positives and document the criterion used by the human reviewer.
{% endstep %}

{% step %}

#### 5. Version the accepted cohort

Tag, group, or publish the verified membership as required by the operating flow.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision           | Production guidance                                                                        |
| ------------------ | ------------------------------------------------------------------------------------------ |
| Embedding model    | Record the selected space because results can change across models or versions.            |
| Distance threshold | Use as a retrieval control, not as an unreviewed acceptance rule.                          |
| Duplicates         | Decide whether to remove, group, down-weight, or retain them for a documented reason.      |
| Outliers           | Review for rare value, corruption, distribution shift, or ingestion failure before acting. |

### Continue the operating flow

* Create a balanced dataset version.
* Use the cohort for annotation, review sampling, or model evaluation.
* Retain the embedding and source version in the decision record.

### Product context

Unitlab helps AI teams curate, annotate, manage, version, and prepare multimodal training data at enterprise scale.

See [Unitlab’s multimodal data annotation platform](https://unitlab.ai/en/data-annotation) for the commercial overview of curation and annotation in one workflow.

***

> **Explore related Unitlab capabilities:** [multimodal data curation](https://unitlab.ai/en/data-curation)


# Datasets overview

Understand datasets as reusable, versioned membership definitions.

A Unitlab dataset names a reusable membership definition. Publishing a version freezes the exact source membership so projects, automation, and downstream records can refer to the same cohort.

{% hint style="info" %}
**Use this area when:** a curated cohort must be reused, compared over time, attached to projects, or reproduced later.
{% endhint %}

### How this area fits into production

```mermaid
flowchart LR
  A["Curated assets or groups"]
  B["Dataset draft"]
  C["Published version"]
  D["Project attachment"]
  E["Release provenance"]
  A --> B
  B --> C
  C --> D
  D --> E
```

![Unitlab datasets overview](https://storage.ghost.io/c/48/f4/48f4b614-5c29-430d-9cb3-e0b3f34395f3/content/images/2026/08/clean-datasets-overview.webp)

*The dataset list is the starting point for ownership, current version, lifecycle, and reuse—not the place where project task state is managed.*

### What this area controls

### What a dataset is

A Unitlab dataset is a curated collection built from folders or individual assets. The live dataset table shows:

* name;
* version;
* asset count;
* size;
* data type;
* projects using the dataset;
* last modified date;
* creator.

Datasets can contain image, video, or multimodal data. Dataset lifecycle tabs include Active, Archived, and Trash.

### Version-first behavior

The SDK describes datasets as version-first:

1. Create a dataset from selected folders or assets.
2. Make edits, which appear as unpublished changes.
3. Publish a named version to freeze a snapshot.
4. Attach the latest or an exact version to a project.

This distinction prevents a project from silently changing when someone adds new assets to the working dataset. An exact attachment can reference, for example, dataset version 2 rather than “whatever the dataset contains today.”

The current model is:

```
Folders and assets
        ↓
Mutable dataset working draft
        ↓ Publish version
Immutable DatasetVersion
        ↓ Attach
Independent project copy
```

There is no **Sync**, **Ignore**, or **Auto-Sync** action between a project and an attached Data Assets source. Editing the source after attachment does not alter the project. To adopt the change, the user publishes a newer dataset version and attaches that version.

### Datasets-list UX

The standalone Datasets page includes search, **New Dataset**, Active/Archived/Trash views, pagination, row selection, and a table with:

* Name;
* Version;
* Assets;
* Size;
* Data Type;
* Used in;
* Last Modified;
* Created By;
* Actions.

The Version cell is an orange **Draft vN** pill when publishing would create the next version, or a blue **vN** pill when the working draft matches the latest published version. Mixed-family datasets display **Multimodal**.

Row actions adapt to lifecycle state. Active datasets can be renamed, have files added, attach a published version to a project, archive, or move to Trash. Archived datasets can be restored to Active or trashed. Trash supports Restore or Delete forever. Dataset lifecycle changes affect the dataset and its versions/memberships; they do not archive or delete the underlying folders and assets.

### Start with the right page

| Decision          | Production guidance                                   |
| ----------------- | ----------------------------------------------------- |
| Create membership | Create a dataset from an inspected Data Space cohort. |
| Freeze membership | Publish a dataset version.                            |
| Use in production | Attach the intended version to a project.             |
| Audit change      | Review dataset history and project dependencies.      |

### Operating boundary

* A dataset is not a folder; folders organize source.
* A dataset is not a project; projects operate work.
* A dataset is not a release; releases freeze downstream output.

### A production-ready handoff

The published version has a clear purpose, owner, exact membership, source provenance, and a reviewed relationship to every attached project.

### Product context

Unitlab helps AI teams curate, annotate, manage, version, and prepare multimodal training data at enterprise scale.

See [Unitlab Dataset Management](https://unitlab.ai/en/dataset-management) for the commercial overview and [multimodal data annotation](https://unitlab.ai/en/data-annotation) for upstream labeling workflows.

***

> **Explore related Unitlab capabilities:** [multimodal data curation](https://unitlab.ai/en/data-curation)


# Create a dataset

Create a named reusable cohort from validated assets or Data Groups.

Create the dataset only after the candidate membership has been inspected. The dataset name and description should communicate why the cohort exists—not merely repeat its source folder.

### Before you make the change

* Confirm the membership question and owner.
* Resolve invalid, duplicate, or incomplete grouped items according to policy.
* Choose a naming and versioning convention that downstream teams can interpret.

![New Dataset dialog with name, description, Folders, Assets, search, and source selection](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F878wtKnyECkRNbmyCKhY%2Fdataset-create-modal.png?alt=media\&token=d09720b0-5aa9-4019-a940-d702f6a40e88)

*From Data Space › Datasets, choose New Dataset, name and describe the reusable cohort, then attach the exact Folders or Assets that passed curation. Source count is visible before creation.*

### Understand the product behavior

1. The user selects **New Dataset**.
2. The user enters a name and optional description.
3. The user chooses at least one folder or asset from the lazy-loaded, server-searched source picker.
4. Unitlab creates the mutable working draft.
5. The dataset detail page opens with its folders and assets.
6. The user publishes v1 before the dataset can be attached to a project.

Adding files to an existing dataset adds existing workspace folders/assets to the working draft. It is not an upload-directly-into-dataset operation.

### Create the reusable membership

{% stepper %}
{% step %}

#### 1. Open New Dataset

Go to Data Space › Datasets and choose New Dataset.
{% endstep %}

{% step %}

#### 2. Name the reusable cohort

Enter the purpose-oriented name and an optional description that explains its intended consumer.
{% endstep %}

{% step %}

#### 3. Choose Folders or Assets

Use the source tabs and search to select exact folders, individual assets, or grouped output; review the selected-source summary.
{% endstep %}

{% step %}

#### 4. Create and inspect

Choose Create Dataset, open the new dataset, and reconcile item count, data type, source provenance, and representative content.
{% endstep %}

{% step %}

#### 5. Control later additions

Use the dataset’s Add files or Attach Data surface; every accepted membership change creates or advances history rather than rewriting an earlier version silently.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision     | Production guidance                                                                                |
| ------------ | -------------------------------------------------------------------------------------------------- |
| Name         | Use a durable purpose-oriented name; keep environment and date in version metadata where possible. |
| Owner        | Assign a role accountable for membership and version changes.                                      |
| Source scope | Keep enough provenance to explain every member later.                                              |
| Groups       | Include the grouped unit when project work requires shared context.                                |

### Continue the operating flow

* Publish the first dataset version.
* Attach the exact version to a pilot project.
* Record future additions as a new version rather than silently changing history.

***

> **Related Unitlab capability guides:** [training-data curation workflows](https://unitlab.ai/en/data-curation)


# Dataset versions and history

Publish immutable snapshots and understand version-first behavior.

Versioning turns a mutable working cohort into a reproducible input. The version—not the display name alone—belongs in project, automation, and release records.

### Before you make the change

* Reconcile membership count and source provenance.
* Inspect representative grouped and ungrouped items.
* Write a concise version note that explains the reason for change.

### Follow the interface

#### Select an explicit published version

![Dataset detail with the version selector open](https://storage.ghost.io/c/48/f4/48f4b614-5c29-430d-9cb3-e0b3f34395f3/content/images/2026/08/clean-dataset-version-selector.webp)

*Open a dataset and choose the current version badge. The selector shows the latest marker, version number, publication reason, item count, date, and a route to the complete history.*

#### Review the complete membership history

![Dataset history dialog showing version deltas and actions](https://storage.ghost.io/c/48/f4/48f4b614-5c29-430d-9cb3-e0b3f34395f3/content/images/2026/08/clean-dataset-history.webp)

*View all versions opens Dataset history. Expand a version to inspect who published it, its timestamp, item delta, activity, View version, and Duplicate as new dataset actions.*

### Understand the product behavior

The SDK describes datasets as version-first:

1. Create a dataset from selected folders or assets.
2. Make edits, which appear as unpublished changes.
3. Publish a named version to freeze a snapshot.
4. Attach the latest or an exact version to a project.

This distinction prevents a project from silently changing when someone adds new assets to the working dataset. An exact attachment can reference, for example, dataset version 2 rather than “whatever the dataset contains today.”

The current model is:

```
Folders and assets
        ↓
Mutable dataset working draft
        ↓ Publish version
Immutable DatasetVersion
        ↓ Attach
Independent project copy
```

There is no **Sync**, **Ignore**, or **Auto-Sync** action between a project and an attached Data Assets source. Editing the source after attachment does not alter the project. To adopt the change, the user publishes a newer dataset version and attaches that version.

### Datasets-list UX

The standalone Datasets page includes search, **New Dataset**, Active/Archived/Trash views, pagination, row selection, and a table with:

* Name;
* Version;
* Assets;
* Size;
* Data Type;
* Used in;
* Last Modified;
* Created By;
* Actions.

The Version cell is an orange **Draft vN** pill when publishing would create the next version, or a blue **vN** pill when the working draft matches the latest published version. Mixed-family datasets display **Multimodal**.

Row actions adapt to lifecycle state. Active datasets can be renamed, have files added, attach a published version to a project, archive, or move to Trash. Archived datasets can be restored to Active or trashed. Trash supports Restore or Delete forever. Dataset lifecycle changes affect the dataset and its versions/memberships; they do not archive or delete the underlying folders and assets.

### Create a dataset

1. The user selects **New Dataset**.
2. The user enters a name and optional description.
3. The user chooses at least one folder or asset from the lazy-loaded, server-searched source picker.
4. Unitlab creates the mutable working draft.
5. The dataset detail page opens with its folders and assets.
6. The user publishes v1 before the dataset can be attached to a project.

Adding files to an existing dataset adds existing workspace folders/assets to the working draft. It is not an upload-directly-into-dataset operation.

### Dataset detail and history

The default detail view is the live working draft. It includes a version dropdown, Grid/List browsing, **Publish version** when changes exist, and **Attach Data** for adding sources to the draft.

Selecting a published version opens a read-only snapshot. It preserves the frozen folder hierarchy, shows a **Read only** chip, and provides **Back to current**. No mutation or publish action appears in snapshot mode.

**View all versions** opens the single history modal. It contains the current working-draft card when dirty and one expandable card per published version. Available actions include:

* View version;
* Restore to working draft;
* Duplicate as new dataset.

Restore changes the working draft and requires a later Publish version action; it never creates a version immediately. Restore is unavailable for live-source-backed own-data or folder-tracking datasets because their draft follows the underlying source.

Changes that can trigger the Draft pill include adding/removing files, folder moves, renames, tag changes, and relevant dataset-name changes—not only item count.

### Operate published membership

{% stepper %}
{% step %}

#### 1. Open the dataset

From Data Space › Datasets, open the dataset name to reach its asset view.
{% endstep %}

{% step %}

#### 2. Inspect the active version

Open the version badge and verify latest status, publication reason, item count, and date.
{% endstep %}

{% step %}

#### 3. Open Dataset history

Choose View all versions to compare version-level counts, publishers, timestamps, and additions or removals.
{% endstep %}

{% step %}

#### 4. Inspect or duplicate a version

Use View version for read-only membership inspection or Duplicate as new dataset when a new lineage must begin from that snapshot.
{% endstep %}

{% step %}

#### 5. Use the explicit version

Attach or automate against the intended version ID; do not rely on the display name or whichever version is latest today.
{% endstep %}

{% step %}

#### 6. Treat membership change as history

Add or remove approved content through the dataset workflow so the change produces a new auditable version.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision         | Production guidance                                                                                 |
| ---------------- | --------------------------------------------------------------------------------------------------- |
| Version scheme   | Choose a scheme the team can advance consistently and interpret in audit records.                   |
| Change note      | State what changed and why, not only who clicked publish.                                           |
| Project behavior | Confirm whether an attached project follows a selected version or requires deliberate reattachment. |
| Recovery         | Know how to restore or duplicate a prior usable state without deleting evidence.                    |

### Continue the operating flow

* Attach the published version to the intended project.
* Record its ID in release provenance.
* Review history before any membership correction.

### Product context

Unitlab helps AI teams curate, annotate, manage, version, and prepare multimodal training data at enterprise scale.

See [Unitlab Dataset Management](https://unitlab.ai/en/dataset-management) for the commercial overview of governed dataset operations and versioned training data.

***

> **Explore related Unitlab capabilities:** [training-data curation workflows](https://unitlab.ai/en/data-curation)


# Attach datasets to projects

Connect the intended published membership to a project and inspect created work.

Attaching a dataset is the boundary where durable membership becomes project work. Validate the source version and resulting Data Units before annotation begins.

### Before you make the change

* Choose the exact dataset version and project.
* Confirm the project ontology and workflow are ready for the incoming modality and group structure.
* Understand whether attachment creates, reuses, or updates Data Units in the current project state.

![Project Data page after dataset attachment](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FsJhHsK7YkhabBtix4o95%2Fproject-data.png?alt=media\&token=06bd10a9-5d5a-4680-a807-d0fc697c37ac)

*After attachment, use Project Data to reconcile source membership with created work items and current workflow state.*

### Understand the product behavior

The project-side Attach Data modal has **Folders** and **Datasets** tabs; there is no loose Assets tab. Organize loose assets into a folder or dataset first.

1. The user opens **Attach Data**.
2. The user selects a workspace folder or a published dataset version.
3. Sources already fully attached are pre-checked, locked, and labeled Attached. This state is evaluated per version, so attaching v1 does not block v2.
4. Unitlab previews item and deduplication counts.
5. Commit clones the frozen version into the project as Unassigned data without re-uploading the files.
6. The newly attached items are selected automatically.
7. The Assign Members modal opens so annotator/reviewer assignment can continue immediately.

Attaching a folder first resolves it through a folder-backed dataset and frozen version. Attaching another project’s data includes only that project’s own source data; sources that were merely attached into that project do not cascade into the next project.

### Project dataset cards, View Source, and Detach

The project’s own data and every attached version appear as uniform dataset cards. Root-card actions include:

* **View Source** — opens the underlying project dataset, folder, or frozen version without mutation;
* **Detach** — removes that source from this project only.

Detach shows source name, asset count, and annotation count. By default it archives the project copies while preserving annotations. The user can explicitly choose **Clear annotations**, which requires a stronger confirmation. Re-attaching restores the archived project rows and kept histories instead of creating duplicates. Other projects and the Data Assets source remain unchanged.

### Attach and reconcile membership

{% stepper %}
{% step %}

#### 1. Select the project

Confirm project identity, purpose, ontology, workflow, and current operating state.
{% endstep %}

{% step %}

#### 2. Choose the dataset version

Use the explicit published version rather than a similarly named dataset.
{% endstep %}

{% step %}

#### 3. Attach the source

Start the operation and monitor any asynchronous processing.
{% endstep %}

{% step %}

#### 4. Reconcile Project Data

Compare expected item count, group behavior, modality, status, and source links.
{% endstep %}

{% step %}

#### 5. Open representative work

Verify the native Workbench and ontology load correctly before assigning volume.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision     | Production guidance                                                               |
| ------------ | --------------------------------------------------------------------------------- |
| Version      | Use the version approved for this project and record it in the change log.        |
| Duplicates   | Define whether already-attached items are skipped, reused, or require correction. |
| Grouped data | Confirm one group produces the intended project unit and layout.                  |
| Rollback     | Know whether detaching removes future availability only or affects active work.   |

### Continue the operating flow

* Write Project Instructions.
* Calibrate the Workbench on representative items.
* Assign work only after Project Data reconciliation.

***

> **Related Unitlab capability guides:** [cross-modal annotation workflows](https://unitlab.ai/en/multimodal-annotation) · [training-data curation workflows](https://unitlab.ai/en/data-curation)


# Dataset lifecycle and recovery

Duplicate, restore, detach, or retire datasets without losing provenance.

Dataset lifecycle operations can affect multiple projects and releases. Review dependencies first, preserve version history, and make the intended future behavior explicit.

The UI exposes rename, add files, attach to project, archive, and delete actions.

Automation adds:

* create from folder or asset IDs;
* add sources;
* inspect unpublished changes;
* publish a version with a title;
* list versions;
* list items for the current state or a specific version;
* preview attachment counts;
* attach the latest or an exact version;
* list and detach a project source.

Video attachments can require an explicit frames-per-second value. Both preview and commit validate this rule so an operator can discover the requirement before creating the project copy.

### Upload Data from a project

Project **Upload Data** uses the plain mixed-data uploader and has no destination-folder picker:

1. The user uploads one or more mixed files.
2. One upload action creates one Batch Queue.
3. Files land in the project’s own data under the project name.
4. The first upload lazily creates the project’s own folder and own-data dataset.
5. Unitlab waits for the batch to become quiet, then auto-publishes one new version for that upload batch.
6. The project Datasets view refreshes.

Project-side auto-publish applies only to the project’s own uploads. Uploading into a workspace Data Assets folder instead creates an unpublished change that still requires an explicit Publish version.

### Attach Data to a project

The project-side Attach Data modal has **Folders** and **Datasets** tabs; there is no loose Assets tab. Organize loose assets into a folder or dataset first.

1. The user opens **Attach Data**.
2. The user selects a workspace folder or a published dataset version.
3. Sources already fully attached are pre-checked, locked, and labeled Attached. This state is evaluated per version, so attaching v1 does not block v2.
4. Unitlab previews item and deduplication counts.
5. Commit clones the frozen version into the project as Unassigned data without re-uploading the files.
6. The newly attached items are selected automatically.
7. The Assign Members modal opens so annotator/reviewer assignment can continue immediately.

Attaching a folder first resolves it through a folder-backed dataset and frozen version. Attaching another project’s data includes only that project’s own source data; sources that were merely attached into that project do not cascade into the next project.

### Project dataset cards, View Source, and Detach

The project’s own data and every attached version appear as uniform dataset cards. Root-card actions include:

* **View Source** — opens the underlying project dataset, folder, or frozen version without mutation;
* **Detach** — removes that source from this project only.

Detach shows source name, asset count, and annotation count. By default it archives the project copies while preserving annotations. The user can explicitly choose **Clear annotations**, which requires a stronger confirmation. Re-attaching restores the archived project rows and kept histories instead of creating duplicates. Other projects and the Data Assets source remain unchanged.

### Dataset versus folder

A folder answers, “Where did these files come from?” A dataset answers, “Which controlled collection are we using for this experiment or project?” One source folder may feed several datasets; one dataset may combine multiple source folders and selected exceptions.

### Use this in production

* Duplicate when a new operating purpose needs an independently governed lineage.
* Restore when the existing lineage remains correct and a recoverable state should return to use.
* Detach only after reviewing active project tasks, assignees, comments, issues, and release dependencies.
* Retain stable IDs, version notes, affected project IDs, owner, reason, and reconciliation counts in the record.

{% hint style="warning" %}
Do not use lifecycle changes to erase evidence of a bad membership decision. Correct the lineage with an explicit new version or documented recovery action.
{% endhint %}

***

> **Related Unitlab capability guides:** [Unitlab’s data curation platform](https://unitlab.ai/en/data-curation)


# Projects overview

Understand projects as the operating system for annotation production.

A Unitlab project connects source membership, Instructions, ontology, workflow stages, queues, annotation, review, issues, statistics, settings, and releases. Projects are typeless so the same operating model can support current and future modalities.

{% hint style="info" %}
**Use this area when:** you need to turn reusable data into assigned, governed, reviewable production work.
{% endhint %}

### How this area fits into production

```mermaid
flowchart TB
  A["Project"]
  B["Data"]
  C["Instructions + ontology"]
  D["Workflow + queues"]
  E["Workbench + review"]
  F["Release"]
  A --> B
  A --> C
  A --> D
  B --> E
  C --> E
  D --> E
  E --> F
```

![Unitlab Projects overview](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FOJlS4AoR0WSh2nHAeEDJ%2Fprojects-overview.png?alt=media\&token=cbc0a0a6-0c64-4bfd-a222-43d71439c575)

*The project list is the administrative entry point; the project itself contains the controls that determine how data becomes reviewed output.*

### What this area controls

### Projects are typeless

A project is a container for annotation work, not a single-modality project type. New project creation asks for a name only. The creation modal may display informational chips for Image, Video, Audio, Text, Medical, and Document, but those chips are not selectable and do not constrain the project.

The resource establishes its own family during upload. A single project can therefore contain image, video, audio, text, medical, and PDF/document resources side by side. When the user opens an item, Unitlab selects the correct native editor from that item’s family and the active ontology’s markup capabilities.

Projects are typeless containers. A project can be used for a specific modality or combine several modalities, and the editor is selected from each resource’s data family rather than from a fixed project type.

### Start with the right page

| Decision          | Production guidance                                 |
| ----------------- | --------------------------------------------------- |
| Start a program   | Create and configure the project.                   |
| Bring in work     | Add project data from a published dataset version.  |
| Define policy     | Write Project Instructions and attach the ontology. |
| Inspect readiness | Use Project Data QA before assignment.              |
| Change operations | Review Project settings and lifecycle impact first. |

### Operating boundary

* Projects reference durable data; they do not replace Data Space.
* Workflow state belongs to project work, not the source asset lifecycle.
* A release is a deliberate frozen output, not simply the current project view.

### A production-ready handoff

A project is ready when its owners, source versions, Instructions, ontology, workflow, queues, review path, and release contract are all explicit.

***

> **Explore related Unitlab capabilities:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation) · [training-data curation workflows](https://unitlab.ai/en/data-curation)


# Create and configure a project

Establish ownership, operating intent, and initial controls.

Create the project around an operating purpose and delivery contract, not around a temporary file type. Modality-specific behavior is loaded by the Workbench from the attached data.

### Before you make the change

* Name the business or model outcome and accountable project owner.
* Identify the published dataset version, ontology owner, workflow owner, and reviewer role.
* Define the first release acceptance criteria.

![Projects list used to start a new project](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FOJlS4AoR0WSh2nHAeEDJ%2Fprojects-overview.png?alt=media\&token=cbc0a0a6-0c64-4bfd-a222-43d71439c575)

*Use a naming convention that distinguishes program, environment, or controlled period without relying on a modality label.*

### Understand the product behavior

A project is a container for annotation work, not a single-modality project type. New project creation asks for a name only. The creation modal may display informational chips for Image, Video, Audio, Text, Medical, and Document, but those chips are not selectable and do not constrain the project.

The resource establishes its own family during upload. A single project can therefore contain image, video, audio, text, medical, and PDF/document resources side by side. When the user opens an item, Unitlab selects the correct native editor from that item’s family and the active ontology’s markup capabilities.

Projects are typeless containers. A project can be used for a specific modality or combine several modalities, and the editor is selected from each resource’s data family rather than from a fixed project type.

### Project creation flow

1. The user selects **New Project**.
2. The user enters the project name.
3. Unitlab creates the project and binds the default workflow.
4. The project opens on its **Datasets** page.
5. The user uploads new data or attaches a published dataset version.
6. Each incoming resource is detected as Image, Video, Audio, Text, Medical, or Document.
7. The workflow creates work-item state and routes the item from Project into Annotate or a configured Model stage.
8. If the project has no Live ontology, the user can create/import one centrally or let the first annotation-side Quick Create Class lazily create the initial project ontology.

Project creation does not create an empty data folder, dataset, class, or ontology. Those objects appear when the user first performs the relevant action.

### Establish the project boundary

{% stepper %}
{% step %}

#### 1. Create the project

Give it a durable, purpose-oriented name and clear description.
{% endstep %}

{% step %}

#### 2. Assign ownership

Add the minimum administrators and project operators required for setup.
{% endstep %}

{% step %}

#### 3. Attach the source

Select the explicit dataset version and reconcile created Project Data.
{% endstep %}

{% step %}

#### 4. Add Instructions and ontology

Make the human guidance and structured schema express the same policy.
{% endstep %}

{% step %}

#### 5. Build the workflow

Connect stages, queues, assignment, review, rejection, escalation, and completion.
{% endstep %}

{% step %}

#### 6. Run a calibration cohort

Open representative items in each relevant editor before production assignment.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision      | Production guidance                                                                              |
| ------------- | ------------------------------------------------------------------------------------------------ |
| Project scope | One project should have a coherent policy, workflow, ownership model, and release contract.      |
| Modality      | Do not encode type in the project model; let attached data select the editor.                    |
| Ownership     | Separate platform administration, policy ownership, operations, review, and downstream approval. |
| Scale gate    | Do not assign full volume until the calibration cohort passes review.                            |

### Continue the operating flow

* Complete Project Instructions.
* Inspect Project Data and filter presets.
* Confirm workflow and queue behavior for every role.

***

> **Explore related Unitlab capabilities:** [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# Project data

Attach, inspect, filter, and reconcile the work population.

Project Data is the operational view of what entered the project and where it sits in the work lifecycle. Use it to catch source, grouping, filter, status, and rendering problems before they become assignment or quality problems.

### Before you make the change

* Know the expected dataset version and item count.
* Define representative normal, edge, grouped, and invalid examples.
* Confirm the intended workflow start state.

![Project Data view with filters and items](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FsJhHsK7YkhabBtix4o95%2Fproject-data.png?alt=media\&token=06bd10a9-5d5a-4680-a807-d0fc697c37ac)

*Use Project Data to reconcile source membership, Data Unit state, and the item population seen by annotators and reviewers.*

### Understand the product behavior

The project’s default Data page offers three URL-persisted views:

* **Grid** — dataset/source cards above item cards;
* **List** — dataset table at the root and annotation-oriented item table inside a dataset;
* **Embedding** — visual embedding scatter plus the filtered item grid.

At the root, attached datasets display their frozen version in the name, and the project’s own uploads appear as one dataset card named after the project. **View Source** and **Detach** are available from the source-card menu.

Inside a dataset, List view shows item Name, Type, workflow Status, Assigned to, and Priority. A Data Group appears as an expandable parent row whose children are indented tiles; the parent carries the group’s workflow status, assignee, priority, labeled progress, and grouped-Workbench link.

The toolbar includes Upload Data, Attach Data, search, Grid/List/Embedding view controls where applicable, card-size control, filters, and selection actions. Text and audio contexts can hide the visual embedding toggle.

### Project filter presets

Advanced Filters can be saved as private, user-scoped project presets. Supported conditions include status, annotators, reviewers, classes, issues, properties, item properties, and tags, using is-any-of or is-none-of logic. A preset can contain up to 20 conditions. Presets are personal UI state rather than auditable project data, so deleting one removes it immediately without changing project content.

### Project data QA and visual inspection

The project Datasets page is also a quality-inspection surface. It combines workflow filters, assignment filters, semantic filters, card-density controls, and annotation-aware display settings so a manager or reviewer can find a problematic cohort and inspect it consistently before opening individual items.

**QA entry points**

| Control                                  | What it helps the user inspect                                                               |
| ---------------------------------------- | -------------------------------------------------------------------------------------------- |
| Workflow stage and status                | New, in-annotation, in-review, processing, complete, archived, error, and invalid cohorts    |
| Assigned                                 | Work owned by a particular annotator or reviewer, or work that remains unassigned            |
| Classes, properties, and Item Properties | Items containing selected ontology content or missing the expected semantic coverage         |
| Issues and tags                          | Known exceptions, escalations, and project-specific quality categories                       |
| Search and Advanced Filters              | A precise subset defined by file, workflow, ontology, assignment, or saved-preset conditions |
| Grid, List, and Embedding views          | Visual inspection, operational table review, or distribution/outlier exploration             |
| Card-size slider                         | More items for rapid scanning or larger cards for closer visual inspection                   |
| Display View                             | Annotation-rendering controls applied consistently across the visible cards                  |

The **Display View** panel contains:

* **Show object names** — renders class/object names on the visible annotations;
* **Color by object ID** — assigns visual identity by instance rather than only by class, which is useful for distinguishing nearby or overlapping objects;
* **Crop view** — focuses each card on its annotated region when close inspection matters more than full-image context;
* **Additional zoom** — increases the inspection scale inside the crop;
* **Boundary thickness** — adjusts annotation-edge thickness;
* **Border opacity**, **Vector opacity**, and **Mask opacity** — independently control the visibility of outlines, vector geometries, and filled masks;
* **Classes** — searchable per-class visibility controls, with expandable **Properties** and **Attributes** visibility where those structures exist.

These are visual QA controls. They do not modify annotation geometry, ontology values, workflow state, or source media. Advanced Filters determine which items are in the result set; Display View determines how their annotations are rendered for inspection.

**Visual QA flow**

1. Open a project dataset and define the cohort with search, status, assignment, class/property, issue, tag, or saved-filter conditions.
2. Choose Grid for visual scanning, List for operational comparison, or Embedding for distribution and outlier inspection.
3. Adjust card size to balance cohort coverage against image detail.
4. Open **Display View** and enable the object names, object-ID coloring, crop, zoom, boundary, and opacity settings needed for the annotation type under review.
5. Search or isolate classes and, when necessary, expand their properties or attributes to remove unrelated overlays.
6. Inspect the visible cohort, then open a questionable item in the Workbench without losing the surrounding project context.
7. Use a comment for item-specific discussion, create or update an issue when ownership and follow-through are required, or use the current Review-stage action to approve or return the item for correction.

This flow supports cohort-first QA: the reviewer can first identify a repeated pattern across many items, then move into the exact annotation context where a correction or workflow decision is made.

### Qualify project data before assignment

{% stepper %}
{% step %}

#### 1. Reconcile attachments

Confirm every dataset card, version, source link, and detach control.
{% endstep %}

{% step %}

#### 2. Compare counts

Match expected source membership to created Data Units and grouped behavior.
{% endstep %}

{% step %}

#### 3. Use filters deliberately

Inspect status, source, modality, assignment, or saved preset cohorts.
{% endstep %}

{% step %}

#### 4. Switch visual modes

Use Grid, List, Embedding, and Display View according to the QA question.
{% endstep %}

{% step %}

#### 5. Open the native editor

Confirm representative items load the right Workbench, ontology, layout, and stage actions.
{% endstep %}

{% step %}

#### 6. Resolve exceptions

Correct invalid, failed, duplicate, ungrouped, or unreadable content before assignment.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision        | Production guidance                                                                       |
| --------------- | ----------------------------------------------------------------------------------------- |
| Count mismatch  | Stop and reconcile attachment, grouping, filters, duplicates, and processing failures.    |
| Saved preset    | Use for repeat operational views; document the underlying criteria for shared procedures. |
| Bad source item | Decide whether to correct, replace, mark invalid, archive, or route around it.            |
| Detach          | Review active work and downstream dependencies first.                                     |

### Continue the operating flow

* Publish Instructions and ontology.
* Test every modality and layout used by the project.
* Open queues only after the work population is approved.

***

> **Related Unitlab capability guides:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation) · [training-data curation workflows](https://unitlab.ai/en/data-curation)


# Project Instructions

Write the operational labeling policy visible inside work.

Instructions explain how experts should apply the ontology to real cases. They should resolve inclusion, exclusion, ambiguity, geometry, attributes, invalid data, review, and escalation with representative examples.

### Before you make the change

* Name the policy owner and reviewer.
* Collect representative easy, difficult, ambiguous, and invalid examples.
* Compare every instruction to the ontology and workflow so no route or required value is missing.

![Project Instructions editor](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F7jrWucEnNWPZV4fEMw0U%2Fproject-instructions.png?alt=media\&token=e5e4c58c-37e3-40d0-aaa8-47c01abb94cd)

*Instructions belong inside the project so annotators and reviewers can consult the current policy without leaving the work context.*

### Understand the product behavior

Instructions can contain:

* rich text/description;
* external URL;
* file attachment;
* PDF upload;
* PPT upload.

Instructions should be the current operational standard an annotator can consult while working. They should include definitions, decision rules, hard examples, counterexamples, uncertainty policy, and escalation steps.

### Publish an actionable policy

{% stepper %}
{% step %}

#### 1. Define the task

State the unit of work, intended output, and what counts as complete.
{% endstep %}

{% step %}

#### 2. Set inclusion and exclusion rules

Explain boundaries and show representative counterexamples.
{% endstep %}

{% step %}

#### 3. Define geometry and attributes

Describe how to draw, classify, relate, and fill required Item Properties.
{% endstep %}

{% step %}

#### 4. Handle ambiguity and invalid data

Give annotators a route that does not require inventing a label.
{% endstep %}

{% step %}

#### 5. Align review and escalation

State what reviewers accept, reject, return, or escalate.
{% endstep %}

{% step %}

#### 6. Calibrate and revise

Run the same difficult examples across roles and update the policy when disagreement reveals a gap.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision          | Production guidance                                                              |
| ----------------- | -------------------------------------------------------------------------------- |
| Example selection | Choose examples that expose the decision boundary, not only ideal annotations.   |
| Ontology coupling | Instructions and schema must change together when policy changes.                |
| Effective date    | Record when a material policy revision begins and how in-flight work is handled. |
| Questions         | Use comments for context and issues for owned correction or follow-through.      |

### Continue the operating flow

* Test the ontology in Workbench.
* Train annotators and reviewers on the same calibration set.
* Record material changes before resuming volume.

***

> **Related Unitlab capability guides:** [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# Project data QA

Inspect project work as a cohort with the project Embedding view, status filters, and Display View controls.

Project Data QA answers a different question from Data Space curation: **is the exact population attached to this project ready to be annotated, reviewed, and released?** Work from **Project › Data** so every filter, visual sample, and Workbench link refers to this project's Data Units and workflow state.

{% hint style="info" %}
The project Embedding view is not the Data Space embedding explorer. It is scoped to the current project and combines visual neighborhoods with project statuses, assignment, classes, and direct Workbench entry.
{% endhint %}

### Inspect the project population in Embedding view

![Project Data Embedding view with project filters and Workbench cards](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FwUyxNjghGhYtufsTnts2%2Fproject-embedding-view.png?alt=media\&token=f455d340-89a2-478a-bdde-e55b7fb596ea)

*Project Embedding view clusters the current project's Data Units while retaining project status, assignment, class filters, and Workbench access.*

Open the project, choose **Data**, then switch from Grid or List to **Embedding view**. Read the page in three layers:

1. **The plot** shows visual proximity between project items. A dense neighborhood may represent repeated acquisition conditions, similar content, or duplicated scenes; an isolated point may be a rare case, an error, or a valuable edge case.
2. **The item cards** identify the underlying Data Units. Open representative cards in Workbench before drawing a quality conclusion from the plot.
3. **The project filters** narrow the population by lifecycle state such as New, Processing, In annotation, In Review, Complete, Invalid, or Archived, and by assignment or class where available.

Embedding distance is an investigation signal, not a quality score. Confirm every suspected cluster, gap, or outlier against the source item, active Instructions, ontology, and workflow state.

### Control how annotations render with Display View

![Project Display View panel beside the Embedding view](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FMDqxnsZ7bWGXYdFl4cC0%2Fproject-display-view.png?alt=media\&token=e179818d-341a-448a-ae07-28772e94b9fc)

*Display View changes how project annotations appear in visual previews; it does not change the stored annotation, workflow state, or source media.*

Choose the **Display View** control beside the Grid, List, and Embedding selectors. The panel provides a shared rendering policy for the current project view:

| Control                          | What it changes                                         | When to use it                                                        |
| -------------------------------- | ------------------------------------------------------- | --------------------------------------------------------------------- |
| Show object names                | Adds class or object labels to rendered annotations.    | Check class confusion, naming consistency, or crowded scenes.         |
| Color by object ID               | Colors instances by identity instead of only by class.  | Inspect tracking identity, overlaps, or same-class instances.         |
| Crop view                        | Crops previews around annotations.                      | Compare boundary quality or small objects without opening every item. |
| Additional zoom                  | Adds crop magnification after Crop view is enabled.     | Review fine edges, small defects, or dense local regions.             |
| Boundary thickness               | Changes the rendered outline width.                     | Keep thin geometry visible without obscuring the source.              |
| Border, vector, and mask opacity | Balances annotation visibility against source evidence. | Inspect occlusion, segmentation leakage, and overlapping geometry.    |
| Classes                          | Shows or hides selected classes in previews.            | Isolate one ontology concept or compare class-specific coverage.      |

When Properties or Attributes are available for the rendered classes, use their Display View controls to isolate the values relevant to the QA question. Display settings are presentation-only; use Workbench to correct labels and workflow actions to route work.

### Run a cohort-level QA pass

{% stepper %}
{% step %}

#### 1. Define the question

Name the failure mode before filtering—for example unreviewed edge cases, rejected work, invalid media, missing class coverage, identity drift, or a recently changed ontology branch.
{% endstep %}

{% step %}

#### 2. Build the project cohort

Combine project status, assignment, class, source, and saved filters. Record the criteria when the cohort will be reviewed repeatedly.
{% endstep %}

{% step %}

#### 3. Inspect the distribution

Use Embedding view to find neighborhoods and outliers; use Grid for visual comparison and List for exact operational fields.
{% endstep %}

{% step %}

#### 4. Tune Display View

Show only the classes and rendering layers needed for the question. Use object-ID color, crop, zoom, thickness, and opacity deliberately.
{% endstep %}

{% step %}

#### 5. Open representative items

Enter Workbench from several typical, boundary, and outlier cards. Confirm the issue against the current item, annotations, timeline where applicable, Instructions, ontology, and stage.
{% endstep %}

{% step %}

#### 6. Route the correction

Correct item-level errors through the configured workflow. Create an Issue or change Instructions, ontology, workflow, model mapping, grouping, or source preparation when the pattern is systemic.
{% endstep %}

{% step %}

#### 7. Re-check the same cohort

Apply the original filter criteria again and verify the affected population—not only the example that exposed the problem.
{% endstep %}
{% endstepper %}

### Production QA cohorts

| Cohort                | Review focus                                                                                      |
| --------------------- | ------------------------------------------------------------------------------------------------- |
| New or Processing     | Attachment completeness, grouping, unreadable sources, and processing failures.                   |
| In annotation         | Missing coverage, policy ambiguity, assignment imbalance, and operator questions.                 |
| In Review or Rejected | Repeated error categories, reviewer consistency, and rework turnaround.                           |
| Invalid               | Whether the reason is source quality, unsupported format, policy, or incorrect routing.           |
| Complete              | Release eligibility, required values, class distribution, and downstream contract readiness.      |
| Recently changed      | Effect of a new ontology, Instructions revision, workflow route, model version, or source cohort. |

{% hint style="warning" %}
Do not use the workspace Data Space embedding screenshot or filters as evidence for project QA. The project view is authoritative for attached Data Units and project workflow state.
{% endhint %}

***

> **Related Unitlab capability guides:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation) · [Unitlab’s data curation platform](https://unitlab.ai/en/data-curation)


# Project settings and lifecycle

Rename, inspect, retire, or delete a project with explicit impact controls.

Project settings control the identity and terminal lifecycle of an operating boundary. Before changing a production project, inventory the attached dataset versions, Live ontology, workflow, open tasks, assignments, Issues, releases, API jobs, and downstream consumers that refer to it.

### Read the current project settings

![Project Settings page with project name and dangerous zone](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FJO2hX91LF97Fs8RBPXma%2Fproject-settings-lifecycle.png?alt=media\&token=b6842de6-2821-4b92-acc1-a53deefae4c0)

*Project Settings shows the current project identity and keeps irreversible deletion in a separate Dangerous zone.*

Open the project and choose **Settings**. The page shows the current project name and, for applicable text work, source-text editing controls. The **Dangerous zone** is intentionally separated because deleting a project is irreversible.

### Rename the project

![Edit project name dialog](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FPcxYgHAuWXkUcdyiTxze%2Fproject-settings-edit.png?alt=media\&token=b6e25a77-b61f-4db7-8a39-5d1eb9fd4a8b)

*Edit opens a focused name dialog. Confirm changes the display name; it does not replace the project's stable identity or reconfigure its data contract.*

1. In **Project name**, choose **Edit**.
2. Enter the durable operating name used by your team.
3. Choose **Confirm**.
4. Reopen the project list, queue links, and any human runbook that uses the display name.

A rename changes the human-facing label, not the meaning of the attached data, ontology, workflow, releases, or stable project ID. Automation should use the project ID rather than the display name.

### Change source text where supported

Text projects can expose source-text editing from Project Settings. Treat a source edit as a data change: identify affected spans, entities, relations, review decisions, and release output; make the smallest correction; then reopen representative work and revalidate downstream offsets and content.

### Retire a project without losing operational history

Use an explicit retirement runbook before considering deletion:

* stop new assignment and automated intake;
* resolve, reassign, or close open annotation and review work;
* close or transfer Issues and ownership;
* publish or cancel pending releases according to policy;
* record the final dataset, ontology, workflow, and release versions;
* disable project-specific service jobs and confirm no scheduled client still targets the project;
* retain the project when audit, provenance, or historical access is required.

### Delete a project

**Delete Project** in the Dangerous zone permanently removes the project boundary. Use it only when the owner has approved deletion, dependencies are reconciled, and required output or audit evidence has been retained elsewhere. The product confirmation is the last guardrail; it is not a substitute for the lifecycle review.

{% hint style="danger" %}
Deletion cannot be undone. Do not delete a project merely to hide completed work or clean up navigation; retain or archive operational history according to your governance policy.
{% endhint %}

### Change-impact record

For every material setting or lifecycle change, record the project ID, owner, reason, previous and new state, attached resource versions, in-flight work treatment, test evidence, downstream validation, and approver.

***

> **Related Unitlab capability guides:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# Annotation Workbench

Understand the shared shell, item state, controls, and navigation.

## Watch the multimodal Workbench in action

See how Unitlab brings modality-native editors into one operating model for ontology, item state, automation, review, and reproducible delivery.

{% embed url="<https://cdn.prod.website-files.com/651fc1beafe23dfe4999151d/6a7610b174ee3346f6af02c1_unitlab-multimodal-data-platform-explainer-1080p.mp4>" %}

[Open the full-screen Unitlab AI explainer](https://cdn.prod.website-files.com/651fc1beafe23dfe4999151d/6a7610b174ee3346f6af02c1_unitlab-multimodal-data-platform-explainer-1080p.mp4).

The Workbench is the shared execution environment for annotation and review. It keeps ontology, item state, instructions, comments, issues, navigation, and workflow actions consistent while loading the native editor required by the active data.

{% hint style="info" %}
**Use this area when:** you are training annotators or reviewers, troubleshooting missing actions, or designing a consistent multimodal operating flow.
{% endhint %}

## Choose an annotation workspace

Select the guide that matches the data and decision you need to produce.

<table data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td>🖼️</td><td><strong>Image Annotation</strong></td><td>Boxes, masks, polygons, cuboids, landmarks, properties, and relations.</td><td><a href="/pages/Rza6B4UB1DyVqnjNMlLi">/pages/Rza6B4UB1DyVqnjNMlLi</a></td></tr><tr><td>🎞️</td><td><strong>Video Annotation</strong></td><td>Frame-accurate tracks, events, timelines, and dynamic properties.</td><td><a href="/pages/ldEO0irjaq5OXCsVcvD0">/pages/ldEO0irjaq5OXCsVcvD0</a></td></tr><tr><td>🔤</td><td><strong>Text Annotation</strong></td><td>Entities, nested spans, classifications, and relations.</td><td><a href="/pages/xbWKrSfexlUCX5mO1L5B">/pages/xbWKrSfexlUCX5mO1L5B</a></td></tr><tr><td>📄</td><td><strong>Document &#x26; PDF</strong></td><td>Native text, images, tables, regions, pages, and document values.</td><td><a href="/pages/k6adwVRRrbaTLEWCjm14">/pages/k6adwVRRrbaTLEWCjm14</a></td></tr><tr><td>🎧</td><td><strong>Audio Annotation</strong></td><td>Temporal events, speakers, transcripts, waveforms, and spectrograms.</td><td><a href="/pages/1CvIkOdGLTP7r4QIVNtp">/pages/1CvIkOdGLTP7r4QIVNtp</a></td></tr><tr><td>🩻</td><td><strong>Medical Annotation</strong></td><td>DICOM, synchronized multiplanar views, clinical properties, and segmentation.</td><td><a href="/pages/wW2zOIgRhPZFntaQXyWg">/pages/wW2zOIgRhPZFntaQXyWg</a></td></tr><tr><td>🔬</td><td><strong>Pathology Annotation</strong></td><td>Whole-slide images, deep zoom, tissue, cells, nuclei, and biomarkers.</td><td><a href="/pages/4YVpEMMZm0Dq9V5qgliJ">/pages/4YVpEMMZm0Dq9V5qgliJ</a></td></tr><tr><td>🛰️</td><td><strong>Geospatial Annotation</strong></td><td>Satellite and aerial imagery, large rasters, coordinates, and spatial geometry.</td><td><a href="/pages/siuwyFzEzGdMVObDp2mX">/pages/siuwyFzEzGdMVObDp2mX</a></td></tr><tr><td>✨</td><td><strong>Detect Anything (SAM 1–SAM 3)</strong></td><td>Magic Touch, prompt labeling, Find Similar, tracking, batch automation, and models.</td><td><a href="/pages/QhQpGdq8eiX8qnGDA2Bo">/pages/QhQpGdq8eiX8qnGDA2Bo</a></td></tr></tbody></table>

### How this area fits into production

{% code collapsedlinecount="10" %}

```mermaid
flowchart TB
  A["Assigned task"]
  B["Workbench context"]
  C["Native editor"]
  D["Ontology values"]
  E["Save + stage action"]
  F["Review or completion"]
  A --> B
  B --> C
  B --> D
  C --> E
  D --> E
  E --> F
```

{% endcode %}

![Image annotation Workbench](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FJR6FTWW8jqHz0VT6Hh7F%2Fimage-workbench.png?alt=media\&token=bc9b4805-7ab8-4b94-a6b8-4a0863a388d8)

*The center canvas changes by modality; ontology, item status, contextual collaboration, navigation, and workflow actions remain part of the same operating shell.*

### What this area controls

Across modalities, Unitlab keeps a recognizable operating model:

* top-center previous/next work-item navigation;
* active class or annotation type;
* Multiview mode and layout controls;
* object/class or event/entity inspection;
* item properties;
* object properties and relations;
* comments;
* tags;
* appearance and visibility controls;
* save and version-history controls;
* workflow actions appropriate to the active stage;
* project instructions and issue context.

The active native editor supplies the modality-specific toolbar, timeline, player, page controls, and inspector content. The surrounding Workbench stays stable when navigation crosses data families.

### Annotation keyboard shortcuts

Press **H** inside the active annotation panel to open the context-aware **Annotation Shortcuts** dialog. It groups the shortcuts available for the current data family into Annotation Tools, Actions, and Views. Only the active Workbench panel receives shortcuts; passive Multiview panels remain read-only.

### Common actions and view controls

| Shortcut                                     | Action                                        | UX behavior                                                                                                       |
| -------------------------------------------- | --------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- |
| **H**                                        | Show shortcuts                                | Opens the shortcut dialog for the active modality                                                                 |
| **Ctrl/Cmd + Z**                             | Undo                                          | Reverses the last annotation edit                                                                                 |
| **Ctrl/Cmd + Y** or **Ctrl/Cmd + Shift + Z** | Redo                                          | Reapplies the last reversed edit                                                                                  |
| **Delete/Backspace**                         | Delete selected                               | Removes the currently selected annotation                                                                         |
| **Ctrl/Cmd + C**                             | Copy                                          | Copies the selected annotation; native text copy takes priority when text is selected                             |
| **Ctrl/Cmd + X**                             | Cut                                           | Cuts the selected annotation; native text editing takes priority in text fields                                   |
| **Ctrl/Cmd + V**                             | Paste                                         | Pastes the copied annotation into the active annotation context                                                   |
| **1–9**                                      | Select class                                  | Activates the project class assigned to that numeric hotkey                                                       |
| **+ / − / 0**                                | Zoom in / zoom out / reset                    | Changes or resets the active viewer zoom                                                                          |
| **Shift + H**                                | Home                                          | Returns from the annotation workspace to the project/home context                                                 |
| **R / K**                                    | Review / reject in compatible legacy contexts | Workflow-managed items use the explicit stage actions in the Workbench header, which guard these status shortcuts |

### Visual annotation tools

| Shortcut   | Tool                                | Available context                                                                            |
| ---------- | ----------------------------------- | -------------------------------------------------------------------------------------------- |
| **V**      | Pan/select/reposition               | Image, video, medical, document; also Pan mode in text                                       |
| **B**      | Bounding Box                        | Image, video, medical, document                                                              |
| **N**      | Cuboid                              | Image, video, document                                                                       |
| **F**      | Brush                               | Image, video, medical, document                                                              |
| **E**      | Eraser                              | Image, video, medical, document                                                              |
| **P**      | Polygon                             | Image, video, medical, document                                                              |
| **L**      | Polyline                            | Image, video, document                                                                       |
| **J**      | Skeleton                            | Image, video, document                                                                       |
| **U**      | Keypoint                            | Image, video, document                                                                       |
| **A**      | Add polygon point                   | Adds a point while editing polygon geometry                                                  |
| **M**      | Magic Touch                         | Image, video, medical, document; **Shift + Click** removes from the assisted mask            |
| **S**      | Detect all objects                  | Image and video for an active box, polygon, mask, or cuboid class                            |
| **T**      | Toggle crosshair                    | Image, video, medical, document                                                              |
| **C**      | Comment                             | Adds an annotation comment                                                                   |
| **D**      | Select PDF Text                     | Document only; selects, copies, or converts embedded PDF text into annotations               |
| **\[ / ]** | Decrease/increase brush size        | Image, video, medical, document                                                              |
| **Esc**    | Finish the active drawing operation | Completes the current segmentation/drawing interaction and returns to a stable editing state |

### Object ordering

| Shortcut | Action                                      |
| -------- | ------------------------------------------- |
| **W**    | Bring selected annotation to front          |
| **O**    | Bring selected annotation one level forward |
| **I**    | Send selected annotation one level backward |
| **Q**    | Send selected annotation to back            |

These ordering commands apply to overlapping canvas annotations. The same actions appear in the object right-click menu and become unavailable when the selected object is already at the relevant edge of the stack.

### Navigation and playback

| Context               | Shortcut                        | Action                                     |
| --------------------- | ------------------------------- | ------------------------------------------ |
| Image, audio, text    | **← / →**                       | Previous/next work item                    |
| Document              | **← / →**                       | Previous/next PDF page                     |
| Document              | **Shift + ← / Shift + →**       | Previous/next PDF document                 |
| Video, medical        | **← / →**                       | Previous/next frame or slice               |
| Video, medical        | **Shift + ← / Shift + →**       | Previous/next video or medical work item   |
| Video, medical, audio | **Space**                       | Play/pause                                 |
| Audio                 | **↑ / ↓**                       | Volume up/down                             |
| Audio                 | **Alt + → / Alt + ←**           | Increase/decrease playback speed           |
| Audio                 | **L**                           | Toggle loop                                |
| Audio                 | **P**                           | Toggle autoplay                            |
| Audio                 | **S**                           | Toggle spectrogram                         |
| Audio                 | **T**                           | Toggle timeline                            |
| Audio                 | **Ctrl/Cmd + O / Ctrl/Cmd + I** | Zoom waveform in/out                       |
| Text                  | **T / R / C**                   | Entity mode / Relation mode / Comment tool |

Shortcut meanings are modality-aware. For example, **T** is Crosshair on a visual canvas, Entity mode in text, and Timeline visibility in audio; **R** is the Relation tool in text and a review action only in compatible non-workflow contexts.

### Annotation View Settings

The View Settings panel controls how the active editor looks and responds without changing the source file. Available sections adapt to the resource family and Workbench mode.

| Section                     | Controls                                                                                                                 | Effect                                                                                                                                               |
| --------------------------- | ------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Canvas / Rendering**      | Pixel Perfect                                                                                                            | Preserves one-to-one pixel rendering for supported standalone image views; hidden where it does not apply, including medical and multi-panel layouts |
| **Annotation Display**      | Display object names; Show properties & attributes; object-label font size; selected-object opacity; object-edge opacity | Controls labels and visual emphasis without changing saved geometry                                                                                  |
| **Annotation Tools**        | Handle size; primitive keypoint sensitivity; show polyshape angles; ruler around cursor                                  | Adjusts editing precision, control-point size, angle visibility, and local measurement guidance                                                      |
| **Auto zoom**               | Auto zoom on timeline click; auto zoom on object-list click                                                              | Centers and enlarges the selected annotation when the corresponding navigation action is used                                                        |
| **Image Adjustments**       | Color Map; Invert Image; Brightness; Contrast; Image Saturation                                                          | Changes the inspection view only; source-image properties remain unchanged                                                                           |
| **Video**                   | Default annotation length; Jump frames                                                                                   | Sets the initial temporal span for a new annotation and the number of frames used by jump navigation                                                 |
| **Hounsfield unit presets** | Built-in preset selection; custom preset name; Save W/L; delete custom preset                                            | Applies and stores medical window/level presets                                                                                                      |
| **3D Viewer Settings**      | Threshold; Opacity                                                                                                       | Controls the medical 3D rendering                                                                                                                    |
| **Projection (MIP)**        | Single slice, MIP (max), MinIP (min), Average; slab thickness                                                            | Controls multi-slice projection when a medical volume is available                                                                                   |
| **Windows Levels**          | VOI LUT mode; histogram; window-width/level range; reset; **Tab + ←/→** for width and **Tab + ↑/↓** for level            | Controls the displayed intensity range for medical-image inspection                                                                                  |

Settings persist as annotation-view preferences. They alter rendering, navigation, or editing ergonomics; they do not rewrite the uploaded media.

### Work-item status

Item status is derived from the item’s current workflow stage, not maintained as an unrelated manual field. User-facing status buckets include:

* New;
* In annotation;
* In Review;
* AI Review where applicable;
* Processing;
* Complete;
* Archived;
* Error.

Task-level status can further show Reopened, Skipped, Pending, Dispatched, Running, Paused, Succeeded, or Failed. An **Invalid** sub-state appears when the latest saved history fails required-property or value validation. Validation is non-blocking: the save succeeds, the item is marked for correction, and the workflow can route it appropriately.

### 7. Multiview Workbench UX

Multiview is the default project annotation experience for all six data families. It is a persistent workspace containing the project header, work-item navigation, mode/layout selector, resizable panel grid, one active editor, and one or more passive inspection panels.

### Two modes

| Mode               | What the panels show                                       | Editing model                                                       | Default layout |
| ------------------ | ---------------------------------------------------------- | ------------------------------------------------------------------- | -------------- |
| **Current file**   | Multiple views of the same datasource                      | One active editor; sibling panels mirror its changes in real time   | 1×1            |
| **Multiple files** | Neighboring work items from the current queue/filter scope | One selected panel is editable; other items remain passive previews | 1×3            |

Layouts range from 1×1 to 4×4. Users can resize panel boundaries, fullscreen a panel, and switch modes from the Workbench header. Mode, layout, and panel sizes persist for the user.

### Active and passive panels

Only the active panel can mutate annotations. It receives the full native editor: tools, hotkeys, selection, object/event/entity editing, player or page controls, comments, classes, properties, history, and workflow actions.

Passive panels can show media, annotations, labels, pages, frames, waveforms, text, or medical projections. In Current file mode they receive the active panel’s live annotation changes, but they cannot originate edits, change the active selection, save, control playback, change a PDF page, edit text entities, or alter waveform regions.

When the user activates a passive panel:

1. Unitlab visually selects it immediately.
2. Any in-progress save in the outgoing panel is allowed to settle.
3. Unsaved changes are saved or safely snapshotted.
4. The outgoing panel becomes passive and pauses modality-specific playback.
5. The incoming native editor loads and restores its panel state.
6. The route updates only after the editor is ready.
7. If hydration fails, Unitlab restores the previous active panel instead of blanking the workspace.

### Current file flow

1. The user opens a work item from a dataset, Task Queue, Batch Queue detail, filtered grid, or direct link.
2. Unitlab creates multiple panel sessions for the same datasource without duplicating data or history.
3. One panel is active; the rest are read-only siblings.
4. Active edits are broadcast to siblings in real time.
5. The user can activate another panel to work from a different view, frame, page, or zoom state.
6. Saving refreshes the shared history and updates every sibling to the persisted result.

Current-file examples:

* **Image:** compare the same image at different zoom or inspection states.
* **Video:** inspect different frames of the same video while sharing annotations and downloaded frames.
* **Audio:** inspect different time regions while one waveform editor remains authoritative.
* **Text:** compare different windows of the same text while entity/relation changes remain synchronized.
* **Document:** compare different PDF pages without confusing page changes with work-item navigation.
* **Medical:** assign Axial, Sagittal, Coronal, and 3D views to separate slots with synchronized annotation state.

### Multiple files flow

1. The user selects **Multiple files** and a layout.
2. Unitlab fills the grid with the active item and neighboring items from the current visible set.
3. The visible set preserves queue, upload-session, archive, search, status, class, and assignment filters.
4. The user activates any ready panel.
5. If the item belongs to another data family, the correct editor loads within the same Workbench shell.
6. Previous/next navigation continues through the filtered work-item sequence.
7. Empty tail slots and inaccessible items appear as panel-level states rather than replacing the whole page with an error.

This mode supports mixed review—for example an image, video, PDF, and audio item in one 2×2 layout—while guaranteeing that only the selected panel is editable.

### Navigation hierarchy

Unitlab maintains three distinct navigation levels:

* **Top-center previous/next:** changes the project work item and can cross data families or move between loose items and Data Groups.
* **Panel activation:** changes which visible panel is editable without leaving the Workbench.
* **Within-item media navigation:** changes a video frame, audio time region, text window, medical slice/view, or PDF page inside the current work item.

Keeping these levels separate is essential for predictable UX. A PDF page change must never advance to another datasource, and a medical slice change must never appear as another project item.

### Grouped Workbench

A Data Group opens as one grouped Workbench whose layout is fixed by the Auto-Grouping builder. The header shows the group layout rather than the general mode/layout selector. Activating a tile never tears down the group route.

The project’s unified previous/next sequence treats each group as one work unit and excludes its member tiles from the loose-item sequence. Grouped saves remain attributable to the group.

### Start with the right page

| Decision                 | Production guidance                                                             |
| ------------------------ | ------------------------------------------------------------------------------- |
| Learn universal controls | Start with keyboard shortcuts, view settings, item state, and workflow actions. |
| Annotate one modality    | Continue to Image, Video, Text, Document & PDF, Audio, or Medical Annotation.   |
| Work across files        | Use Multimodal Annotations and the Multiview Workbench.                         |
| Improve speed            | Use AI-assisted annotation only inside the same quality contract.               |

### Operating boundary

* Saving annotation state is different from routing the task to another stage.
* Viewer navigation is different from project-item navigation.
* A missing action can be caused by role, stage, selection, resource state, or configuration.

### A production-ready handoff

An annotator can identify the active item, ontology, workflow stage, save state, navigation scope, and next valid action without leaving the task.

***

> **Explore related Unitlab capabilities:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation) · [multimodal data annotation](https://unitlab.ai/en/multimodal-annotation)


# Image Annotation

Create boxes, polygons, masks, cuboids, landmarks, properties, and relations on still images with AI-assisted precision.

Unitlab’s image Workbench combines pixel-accurate geometry, reusable ontologies, interactive AI assistance, and workflow review for computer-vision training data. Choose the least complex geometry that preserves the signal your model needs, then make the decision reproducible through instructions and ontology rules.

{% hint style="info" %}
**Use this guide when:** you are building detection, segmentation, keypoint, pose, classification, captioning, or visual-relation datasets from still images.
{% endhint %}

## See image annotation in action

The current demo shows image labeling and AI-assisted proposal creation inside the same editable Workbench.

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/auto-labeling-2.mp4>" %}

[Open the demo in a new tab](https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/auto-labeling-2.mp4).

## Before you begin

1. Create or select a project whose data and ontology match this modality.
2. Confirm the project instructions define the unit of annotation, boundary or timing policy, required properties, and review route.
3. Open the project and enter the assigned item from the project data view or queue. The Workbench loads the modality-native editor inside the shared Unitlab shell.

See [Annotation Workbench](/documentation/annotations/annotation-workbench) for navigation, saving, item state, comments, issues, and workflow actions.

## Understand the image work surface

The image occupies the central canvas. The toolbar provides pan, Magic Touch, prompt-based detection, bounding box, cuboid, brush, eraser, polygon, skeleton, line, keypoint, crosshair, comments, image settings, undo/redo, zoom, and annotation shortcuts. The ontology panel exposes classes, class properties, relations, and whole-image Item Properties; the workflow action reflects the item’s current stage.

![Magic Touch selects a cherry and Find Similar proposes matching masks](/files/NPEIFwgND8aE98GUCcku)

*Interactive assistance starts from a human-selected object. Proposals stay editable and must be reviewed before acceptance.*

## Supported annotation model

| Annotation type                     | Use it for                                                                           |
| ----------------------------------- | ------------------------------------------------------------------------------------ |
| **Bounding box**                    | Object detection when approximate extent is sufficient.                              |
| **Polygon or mask**                 | Pixel-level boundaries, irregular shapes, area, or occlusion-sensitive segmentation. |
| **Brush and eraser**                | Fine mask correction, holes, thin structures, and local cleanup.                     |
| **Cuboid**                          | Perspective-aware pseudo-3D extent represented by eight visible corners.             |
| **Keypoint or skeleton**            | Landmarks and connected pose structures.                                             |
| **Line or polyline**                | Roads, contours, paths, and elongated structures.                                    |
| **Classification or Item Property** | Whole-image labels, captions, quality, scene, or acquisition attributes.             |
| **Relation**                        | A governed connection between annotated objects.                                     |

The ontology—not the file type alone—determines which tools and values the annotator sees. Use Item Properties for facts about the complete image and class properties for facts about a specific object.

## Pixel-accurate labeling

Zoom to the level required by the boundary policy, create a polygon or mask, and refine it with brush and eraser. Review thin structures, holes, touching instances, reflections, shadows, truncation, and occlusion consistently. Do not demand pixel precision when the downstream task only needs object localization; unnecessary detail increases review cost without improving the target signal.

![Pixel-level mask editing beside the full annotated image](/files/diFD52zLu1JfnHv3grqW)

*The full object and the detailed boundary remain part of one labeling decision.*

## Nested ontologies, properties, and relations

A reusable ontology can combine visual classes with required attributes, nested options, Item Properties, and relations. This keeps class meaning stable across annotators and projects. For example, a Vehicle object can require type and occlusion properties, relate to a Road object, while image-level weather and capture conditions remain Item Properties.

![Image ontology with classes, attributes, relations, and Item Properties](/files/MfKk41BsNzanR6kzwUmZ)

*Use ontology structure to make visual labels machine-readable and reviewable, not just visually correct.*

## AI-assisted image labeling

Magic Touch creates an editable segmentation proposal from an interactive prompt. Prompt Auto-Labeling can detect supported objects on the current image, and Find Similar uses a selected box, polygon, or mask to propose visually similar instances on that same image. Accept only after checking missed objects, false positives, boundary quality, class mapping, and overlap with existing annotations.

## Annotate one production item

{% stepper %}
{% step %}

#### 1. Orient to the item

Read the instructions, confirm the active ontology and workflow stage, then inspect the complete image before zooming into the first target.
{% endstep %}

{% step %}

#### 2. Choose the class and geometry

Select the ontology class or numeric hotkey, then activate the geometry required by the task. Do not substitute a box for a required mask or create object geometry for a whole-image Item Property.
{% endstep %}

{% step %}

#### 3. Create the annotation

Draw the box, polygon, mask, cuboid, line, keypoint, or skeleton. Zoom and pan while keeping enough surrounding context to interpret the object correctly.
{% endstep %}

{% step %}

#### 4. Refine and describe it

Adjust vertices or mask pixels, then complete every required class property. Add relations only between the intended source and target objects.
{% endstep %}

{% step %}

#### 5. Inspect the complete image

Scan for missed instances, duplicate objects, inconsistent class choice, boundary drift, and invalid overlap. Use visibility and object-order controls when dense annotations obscure one another.
{% endstep %}

{% step %}

#### 6. Save and route the item

Save the current state, then use the stage action to submit, approve, reject, escalate, skip, or mark invalid according to the project workflow.
{% endstep %}
{% endstepper %}

## Quality review

| Review focus           | What to check                                                                            |
| ---------------------- | ---------------------------------------------------------------------------------------- |
| **Coverage**           | Inspect the entire image at a useful zoom; do not review only the first dense region.    |
| **Geometry**           | Check that the chosen shape matches the downstream task and the project instruction.     |
| **Boundary policy**    | Apply the same rule to occluded, truncated, touching, reflective, and ambiguous objects. |
| **Ontology values**    | Resolve required properties, Item Properties, and relation validation.                   |
| **Assisted proposals** | Measure false positives, misses, and boundary errors before expanding automation.        |

{% hint style="warning" %}
A saved annotation is not automatically a production-ready annotation. Required values, boundary or timing policy, cross-item consistency, and the configured review stage still apply.
{% endhint %}

## Move from labels to governed data

Accepted image annotations move through the configured review route and into versioned dataset or release outputs. Preserve the geometry in an export format that supports it: cuboids, masks, nested properties, and relations may require a richer format than a detection-only export.

![Integrated Unitlab workflow connecting model assistance, annotation, review, and quality assurance](/files/yPoYNby78Kd9mxWTVqxy)

*Use workflows to keep model output, human correction, review, and approval in one traceable operating path.*

## Next steps

* Use [Detect Anything (SAM 1–SAM 3)](/documentation/auto-labeling/detect-anything-sam-1-sam-3) to calibrate interactive and batch assistance.
* Use [Multimodal overview](/documentation/multimodal-annotations/multimodal-overview) when related files or views must stay in one task.
* Curate difficult cases and review cohorts in [Data curation](/documentation/data/data-curation).
* Read the current [image annotation product overview](https://unitlab.ai/en/image-annotation) for the feature overview and current media.


# Video Annotation

Create frame-accurate object tracks, temporal events, dynamic properties, and synchronized video annotations at scale.

Video annotation adds time, identity, and change to visual labeling. Unitlab combines frame-level geometry, full or directional auto-tracking, interpolation, synchronized audio, dynamic properties, and a complete timeline so teams can review what happened, when it happened, and which object remained the same.

{% hint style="info" %}
**Use this guide when:** you are building object-tracking, action, event, behavior, scene, or multimodal video datasets.
{% endhint %}

## See video annotation in action

The demo shows frame-aware visual annotation and AI assistance. Use the timeline—not playback alone—to inspect track continuity and temporal labels.

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/auto-labeling-2.mp4>" %}

[Open the demo in a new tab](https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/auto-labeling-2.mp4).

## Before you begin

1. Create or select a project whose data and ontology match this modality.
2. Confirm the project instructions define the unit of annotation, boundary or timing policy, required properties, and review route.
3. Open the project and enter the assigned item from the project data view or queue. The Workbench loads the modality-native editor inside the shared Unitlab shell.

See [Annotation Workbench](/documentation/annotations/annotation-workbench) for navigation, saving, item state, comments, issues, and workflow actions.

## Understand the video work surface

The video Workbench combines the visual canvas with playback controls, frame navigation, an annotation timeline, ontology controls, and item workflow actions. Geometry can be created on an exact frame, propagated with tracking or interpolation, and corrected at keyframes. Audio can remain synchronized to the video when sound is part of the decision.

![Three cyclist frames aligned to an exact keyframe timeline](/files/YUAPNOkZxH48yNYStvqf)

*Frame-accurate labeling keeps the visual state and the timeline state aligned.*

## Supported annotation model

| Annotation type                           | Use it for                                               |
| ----------------------------------------- | -------------------------------------------------------- |
| **Tracked box, polygon, mask, or cuboid** | Object localization that persists across frames.         |
| **Keypoint or skeleton track**            | Pose or landmark motion over time.                       |
| **Line or polyline**                      | Temporal paths and elongated structures.                 |
| **Temporal classification or event**      | An action, scene, state, or interval on the timeline.    |
| **Dynamic class property**                | A property of one tracked object that changes over time. |
| **Dynamic Item Property**                 | A whole-video state or event that changes over time.     |
| **Relation**                              | A governed connection between tracked objects.           |
| **Static Item Property**                  | A fact that applies to the complete sequence.            |

Tracking predicts object state across frames; interpolation fills geometry between explicit keyframes. They solve different problems and both require timeline review.

## Full, forward, backward, and multi-object tracking

Create a reliable seed annotation, then run the tracking direction that matches the visible interval. Full tracking covers both directions around the seed; forward or backward tracking limits the prediction range. Auto-Track All can propagate multiple selected objects together. Correct identity switches, drift, missed reappearances, and shape errors at the first frame where they occur.

![A van and cyclist selected and tracked together across video frames](/files/LoZ274Zk3xL0iRyw1GMm)

*Multi-object tracking accelerates propagation while the annotator remains responsible for identity and geometry.*

## Dynamic properties and Item Properties

Use dynamic class properties when the state belongs to one tracked object—for example vehicle motion or visibility. Use dynamic Item Properties when the state describes the complete video—for example scene condition or global event. Create explicit temporal segments and review their start and end frames. A property value outside its intended time range is a label error even when the geometry is correct.

## Timeline review for long sequences

Use the overview timeline to find tracks, keyframes, temporal segments, gaps, and dense regions. Zoom into a local interval for exact correction, then return to the full sequence to confirm continuity. Long videos should be reviewed at transition points, object entrances and exits, occlusion, shot changes, and every property boundary.

![Video moments connected to tracks, states, keyframes, and a shared playhead](/files/RHnKcRJrKnXjUI6nvDII)

*The timeline is the authoritative map of track continuity, temporal state, and review coverage.*

## Annotate one production item

{% stepper %}
{% step %}

#### 1. Find the first reliable frame

Navigate to a frame where the target is visible and unambiguous. Confirm the ontology class and the intended identity rule.
{% endstep %}

{% step %}

#### 2. Create the seed geometry

Draw the required box, polygon, mask, cuboid, keypoints, or skeleton and complete any properties that apply at the seed.
{% endstep %}

{% step %}

#### 3. Choose propagation

Use full, forward, or backward auto-tracking for model-based propagation, or add explicit keyframes and interpolation when controlled geometric transition is the correct method.
{% endstep %}

{% step %}

#### 4. Correct the first failure

Scrub the timeline and stop at the first drift, identity switch, occlusion error, or missed reappearance. Correct there before continuing so later predictions do not inherit the mistake.
{% endstep %}

{% step %}

#### 5. Add temporal meaning

Create dynamic class properties, dynamic Item Properties, events, or classifications with exact start and end frames. Add object relations where the ontology requires them.
{% endstep %}

{% step %}

#### 6. Review and route

Inspect the full timeline, key transitions, synchronized audio, required values, and final object identity; then save and use the configured stage action.
{% endstep %}
{% endstepper %}

## Quality review

| Review focus         | What to check                                                                      |
| -------------------- | ---------------------------------------------------------------------------------- |
| **Identity**         | One track must represent one real object; split or repair identity switches.       |
| **Frame boundaries** | Check the exact first and last valid frames for every track and event.             |
| **Geometry drift**   | Review changes in scale, pose, occlusion, motion blur, and camera movement.        |
| **Temporal values**  | Confirm dynamic properties and Item Properties change only on the intended frames. |
| **Audio context**    | When audio informs the label, review playback and waveform context together.       |

{% hint style="warning" %}
A saved annotation is not automatically a production-ready annotation. Required values, boundary or timing policy, cross-item consistency, and the configured review stage still apply.
{% endhint %}

## Move from labels to governed data

Video outputs should preserve track identity, keyframes, temporal ranges, dynamic properties, relations, and the source frame rate or timestamp basis. Route model-assisted results through the same review stage as manual work and version material changes to tracking or ontology policy.

![Integrated Unitlab workflow connecting model assistance, annotation, review, and quality assurance](/files/yPoYNby78Kd9mxWTVqxy)

*Use workflows to keep model output, human correction, review, and approval in one traceable operating path.*

## Next steps

* Use [Detect Anything (SAM 1–SAM 3)](/documentation/auto-labeling/detect-anything-sam-1-sam-3) to calibrate interactive and batch assistance.
* Use [Multimodal overview](/documentation/multimodal-annotations/multimodal-overview) when related files or views must stay in one task.
* Curate difficult cases and review cohorts in [Data curation](/documentation/data/data-curation).
* Read the current [video annotation product overview](https://unitlab.ai/en/video-annotation) for the feature overview and current media.


# Text Annotation

Create entities, nested spans, classifications, relations, and structured properties for NLP and LLM datasets.

Unitlab’s text Workbench keeps source text, span boundaries, entity identity, relations, classifications, and document-level context in one governed task. It supports precise named-entity recognition as well as nested and overlapping spans that simpler labeling interfaces cannot represent safely.

{% hint style="info" %}
**Use this guide when:** you are building NER, relation extraction, intent, sentiment, classification, information-extraction, or LLM training datasets.
{% endhint %}

## See text annotation in action

The demo shows text selection, entity labeling, and structured review within the current Unitlab interface.

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/text-annotation-1.mp4>" %}

[Open the demo in a new tab](https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/text-annotation-1.mp4).

## Before you begin

1. Create or select a project whose data and ontology match this modality.
2. Confirm the project instructions define the unit of annotation, boundary or timing policy, required properties, and review route.
3. Open the project and enter the assigned item from the project data view or queue. The Workbench loads the modality-native editor inside the shared Unitlab shell.

See [Annotation Workbench](/documentation/annotations/annotation-workbench) for navigation, saving, item state, comments, issues, and workflow actions.

## Understand the text work surface

The text editor keeps the source document readable while exposing ontology classes and structured values. Annotators select exact character spans, assign entity classes, connect entities with relations, and complete item-level classifications without rewriting the source content.

![Text with person, organization, location, topic, and arbitrary span annotations](/files/RHpBRqjDvFdP8tN6tf76)

*Entity labels preserve exact source spans while keeping surrounding language available for interpretation.*

## Supported annotation model

| Annotation type                  | Use it for                                                                               |
| -------------------------------- | ---------------------------------------------------------------------------------------- |
| **Named entity**                 | A typed span such as person, organization, location, product, date, or domain concept.   |
| **Arbitrary span**               | A project-defined phrase or token sequence that does not fit a standard entity taxonomy. |
| **Nested or overlapping entity** | Independent labels whose character ranges partially or fully overlap.                    |
| **Relation**                     | A directed or undirected connection between two annotated entities.                      |
| **Text classification**          | Intent, sentiment, topic, risk, or another label for the complete item.                  |
| **Entity property**              | Structured information that belongs to one entity.                                       |
| **Item Property**                | Document-level context, source, quality, language, or other whole-item value.            |

Span boundaries should follow a written policy for punctuation, articles, possessives, whitespace, and repeated mentions. The ontology defines meaning; visual highlight color alone does not.

## Relationships and entity linking

Select the source entity and create the ontology-defined relation to the intended target. Review relation direction, cardinality, and scope; a correct pair connected in the wrong direction is still structurally wrong. Use relations for facts such as works-for, located-in, refers-to, part-of, or any project-specific link.

![Directional relationships linking a person, organization, and location](/files/5OUMISWbCx0Tgum2xMcs)

*Relations turn isolated spans into a structured graph that downstream models can learn from.*

## Advanced text ontologies

Combine entity classes with nested properties, Item Properties, and controlled relations. Required values make incomplete annotations visible before submission. Reuse the ontology across projects when class definitions and value rules must remain stable.

![Text ontology with entities, attributes, relations, and Item Properties](/files/i9Kid1MFZlCHR5fhh3Fd)

*A structured ontology separates span location from semantic meaning and document-level context.*

## Nested, overlapping, and context-aware annotation

Create each valid label independently when entities overlap or one entity contains another. Preserve sentence, paragraph, and document context during review. For repeated mentions, follow the project’s coreference and mention policy rather than assuming the first label applies everywhere. AI-assisted proposals should be corrected against the source text, not accepted from confidence alone.

## Annotate one production item

{% stepper %}
{% step %}

#### 1. Read the full context

Review the item, instructions, language, and Item Properties before selecting the first span. Resolve whether the task is mention-level, entity-level, or document-level.
{% endstep %}

{% step %}

#### 2. Select the exact span

Highlight only the characters required by the boundary policy and choose the ontology class. Reopen the selection if punctuation or whitespace is wrong.
{% endstep %}

{% step %}

#### 3. Handle overlap intentionally

Create nested or overlapping entities as separate annotations when the policy permits them; do not merge distinct concepts to avoid overlap.
{% endstep %}

{% step %}

#### 4. Add entity structure

Complete required entity properties and create relations with the correct source, target, and direction.
{% endstep %}

{% step %}

#### 5. Complete whole-item labels

Add classifications and Item Properties that describe the complete text, not a single mention.
{% endstep %}

{% step %}

#### 6. Review and route

Scan every labeled span in context, inspect missed mentions and relation coverage, resolve validation, save, and use the current workflow action.
{% endstep %}
{% endstepper %}

## Quality review

| Review focus         | What to check                                                                       |
| -------------------- | ----------------------------------------------------------------------------------- |
| **Span boundary**    | Apply the same token, punctuation, whitespace, and article policy to every mention. |
| **Class meaning**    | Check the label definition against context, not surface wording alone.              |
| **Overlap**          | Preserve valid nested entities without creating accidental duplicates.              |
| **Relations**        | Check source, target, direction, and required relation coverage.                    |
| **Document context** | Review classification and Item Properties against the complete item.                |

{% hint style="warning" %}
A saved annotation is not automatically a production-ready annotation. Required values, boundary or timing policy, cross-item consistency, and the configured review stage still apply.
{% endhint %}

## Move from labels to governed data

Text exports should preserve character offsets, source text identity, nested spans, relations, properties, and document-level labels. If text normalization occurs downstream, maintain a traceable mapping back to the original offsets and version the ontology with the dataset.

![Integrated Unitlab workflow connecting model assistance, annotation, review, and quality assurance](/files/yPoYNby78Kd9mxWTVqxy)

*Use workflows to keep model output, human correction, review, and approval in one traceable operating path.*

## Next steps

* Use [Detect Anything (SAM 1–SAM 3)](/documentation/auto-labeling/detect-anything-sam-1-sam-3) to calibrate interactive and batch assistance.
* Use [Multimodal overview](/documentation/multimodal-annotations/multimodal-overview) when related files or views must stay in one task.
* Curate difficult cases and review cohorts in [Data curation](/documentation/data/data-curation).
* Read the current [text annotation product overview](https://unitlab.ai/en/text-annotation) for the feature overview and current media.


# Document & PDF Annotation

Annotate native PDF text, images, tables, regions, pages, relations, and document properties without losing document context.

Unitlab keeps a multipage document as one reviewable item. Annotators can work with native selectable PDF content, visual regions, tables, figures, page-aware history, structured OCR outputs, and document-level ontology values for Document AI datasets.

{% hint style="info" %}
**Use this guide when:** you are building invoice, receipt, form, contract, report, technical-document, OCR, or document-understanding datasets.
{% endhint %}

## See document annotation in action

The demo shows native document annotation inside the current multipage Workbench.

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/document-annotation-1.mp4>" %}

[Open the demo in a new tab](https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/document-annotation-1.mp4).

## Before you begin

1. Create or select a project whose data and ontology match this modality.
2. Confirm the project instructions define the unit of annotation, boundary or timing policy, required properties, and review route.
3. Open the project and enter the assigned item from the project data view or queue. The Workbench loads the modality-native editor inside the shared Unitlab shell.

See [Annotation Workbench](/documentation/annotations/annotation-workbench) for navigation, saving, item state, comments, issues, and workflow actions.

## Understand the document work surface

The document Workbench combines a native PDF viewer, page controls, text selection, region geometry, ontology values, annotation history, comments, and workflow actions. Page navigation changes the view inside the same work item; it must not be confused with project-item navigation.

![Three native PDF pages with structured text, table, and figure annotations](/files/yKGnTKQPl4erg0XDR7XP)

*Pages remain part of one document while annotations preserve page identity and source context.*

## Supported annotation model

| Annotation type            | Use it for                                                                               |
| -------------------------- | ---------------------------------------------------------------------------------------- |
| **Native text span**       | Selectable source text such as a clause, title, field, or table value.                   |
| **Bounding box or region** | Visual content, scanned text, signatures, stamps, figures, or layout regions.            |
| **Table structure**        | Rows, columns, cells, headers, and relationships defined by the project ontology.        |
| **Classification**         | Document type, clause type, status, or another page- or item-level category.             |
| **Relation**               | A connection such as field-to-value, key-to-cell, clause-to-party, or figure-to-caption. |
| **Property**               | Structured information attached to one annotation.                                       |
| **Item Property**          | A value that describes the complete document.                                            |

Use native text selection when the PDF exposes reliable text. Use geometry when the signal is visual, scanned, or layout-dependent. Do not convert a selectable PDF into a flattened image workflow unless the project requires it.

## Selectable PDF text, images, and tables

Select native text directly and assign the ontology class when exact textual content is the target. For embedded images, tables, signatures, stamps, or non-selectable content, draw the required region geometry. Keep the visual position and textual value connected through class structure or relations rather than duplicating meaning in free text.

![Selected PDF text, embedded chart, and table cells](/files/ARZgLGZp28NXdhcfRyZH)

*Native selection and visual regions can coexist in the same document task.*

## Page-aware history and review

Use page controls to move within the document and annotation history to understand what changed. Reviewers should inspect both the active page and any related pages before approving document-level classifications or relations. Comments and issues should identify the relevant page and annotation so rework is unambiguous.

![Invoice page with text selection, review, approval, and history lanes](/files/vuDBYCh8tYSoTBCDobtb)

*Page identity stays explicit throughout annotation, review, and history.*

## Long documents and structured extraction

Navigate large PDFs without breaking them into unrelated tasks. Use the ontology to define fields, clauses, line items, properties, and relations. For OCR-oriented projects, distinguish source text, extracted value, visual region, and normalized value so downstream systems can reproduce how the label was derived.

## Annotate one production item

{% stepper %}
{% step %}

#### 1. Open the document item

Confirm the file, page count, project instruction, ontology, and current workflow stage. Use within-document page controls to reach the target page.
{% endstep %}

{% step %}

#### 2. Choose native text or region geometry

Select exact PDF text when available. Draw a box or other supported region when the target is visual, scanned, or layout-dependent.
{% endstep %}

{% step %}

#### 3. Assign the ontology class

Choose the field, clause, table, figure, or document class and complete required annotation properties.
{% endstep %}

{% step %}

#### 4. Connect document structure

Create relations between keys and values, clauses and parties, figures and captions, or other ontology-defined pairs. Add Item Properties for whole-document facts.
{% endstep %}

{% step %}

#### 5. Review across pages

Check repeated fields, cross-page references, table continuation, page boundaries, and document-level values. Use history and comments when a correction needs context.
{% endstep %}

{% step %}

#### 6. Save and route

Resolve validation, save the document state, then submit, approve, reject, or escalate using the active workflow action.
{% endstep %}
{% endstepper %}

## Quality review

| Review focus        | What to check                                                                            |
| ------------------- | ---------------------------------------------------------------------------------------- |
| **Page identity**   | Every region and text span must resolve to the correct source page.                      |
| **Text boundary**   | Preserve exact source characters and avoid accidental whitespace or punctuation changes. |
| **Layout geometry** | Boxes and regions should cover the intended visual object without unrelated content.     |
| **Structure**       | Check table, field-value, clause, and cross-page relations.                              |
| **Document values** | Validate Item Properties against the complete document, not one visible page.            |

{% hint style="warning" %}
A saved annotation is not automatically a production-ready annotation. Required values, boundary or timing policy, cross-item consistency, and the configured review stage still apply.
{% endhint %}

## Move from labels to governed data

Document outputs should preserve the original file identity, page number, native text or character offsets, region coordinates, table or relation structure, properties, and whole-document labels. Publish reviewed membership through a dataset version or release so downstream extraction experiments remain reproducible.

![Integrated Unitlab workflow connecting model assistance, annotation, review, and quality assurance](/files/yPoYNby78Kd9mxWTVqxy)

*Use workflows to keep model output, human correction, review, and approval in one traceable operating path.*

## Next steps

* Use [Detect Anything (SAM 1–SAM 3)](/documentation/auto-labeling/detect-anything-sam-1-sam-3) to calibrate interactive and batch assistance.
* Use [Multimodal overview](/documentation/multimodal-annotations/multimodal-overview) when related files or views must stay in one task.
* Curate difficult cases and review cohorts in [Data curation](/documentation/data/data-curation).
* Read the current [document annotation product overview](https://unitlab.ai/en/document-annotation) for the feature overview and current media.


# Audio Annotation

Label temporal events, speakers, transcripts, classifications, properties, and relations on synchronized waveform and spectrogram views.

Audio annotation is a timing decision as much as a semantic decision. Unitlab aligns waveform, spectrogram, playback, transcript context, ontology values, and a complete annotation timeline so teams can create precise speech and sound datasets without losing temporal evidence.

{% hint style="info" %}
**Use this guide when:** you are building speech recognition, diarization, sound-event detection, intent, acoustic-scene, quality, or multimodal audio datasets.
{% endhint %}

## See audio annotation in action

The demo shows audio as part of the multimodal Workbench, with time-based annotation and contextual review.

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/multimodal-6.mp4>" %}

[Open the demo in a new tab](https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/multimodal-6.mp4).

## Before you begin

1. Create or select a project whose data and ontology match this modality.
2. Confirm the project instructions define the unit of annotation, boundary or timing policy, required properties, and review route.
3. Open the project and enter the assigned item from the project data view or queue. The Workbench loads the modality-native editor inside the shared Unitlab shell.

See [Annotation Workbench](/documentation/annotations/annotation-workbench) for navigation, saving, item state, comments, issues, and workflow actions.

## Understand the audio work surface

The audio Workbench exposes playback, zoom, a shared playhead, waveform and spectrogram context, temporal ranges, ontology controls, transcript-related values, and the current workflow action. Zoom changes inspection precision; it does not change the underlying timestamp basis.

![Synchronized waveform and spectrogram with one selected time range](/files/mv1yGFQv80UMnfyw2481)

*Waveform and spectrogram provide complementary evidence around one authoritative temporal selection.*

## Supported annotation model

| Annotation type        | Use it for                                                                     |
| ---------------------- | ------------------------------------------------------------------------------ |
| **Temporal event**     | A sound, action, acoustic state, or other interval with start and end time.    |
| **Speaker segment**    | A time range attributed to one speaker identity or role.                       |
| **Transcription span** | Text aligned to a clip or exact temporal interval.                             |
| **Classification**     | Intent, sentiment, acoustic scene, quality, or another item-level label.       |
| **Dynamic property**   | A value that changes over the recording timeline.                              |
| **Relation**           | A connection between speakers, events, transcripts, or contextual items.       |
| **Item Property**      | Language, source, channel, consent, quality, or other whole-recording context. |

Write an explicit boundary policy for silence, overlap, clipped speech, background noise, non-speech events, and minimum-duration segments before production starts.

## Time-aligned transcription review

Align transcript content with the intended speaker or temporal range. Check words at segment boundaries, overlapping speakers, false starts, filler words, unintelligible audio, and project-specific normalization rules. Keep source audio and transcript identity linked so corrections remain traceable.

![Speaker transcript spans aligned to a waveform and playhead](/files/utbFaZvfJHyrKDFLpGj4)

*Transcript content, speaker identity, and time range are reviewed together.*

## Complete audio timelines

Use the timeline to inspect event ranges, speaker turns, transcript coverage, dynamic properties, and quality segments across the complete recording. Long recordings require review at every transition, overlap, channel change, and abrupt acoustic event—not only at uniformly spaced samples.

![Audio moments connected to event, speaker, transcript, and quality ranges](/files/HWh0n8LoBWrvlCH0826T)

*The timeline reveals missing coverage, accidental overlap, and boundary inconsistency.*

## Ontology-driven audio structure

Use classes for speakers and events, properties for structured details, relations for connections, and Item Properties for facts about the entire recording. AI-assisted proposals can accelerate transcription or event discovery, but confidence should route review rather than replace it.

## Annotate one production item

{% stepper %}
{% step %}

#### 1. Listen before labeling

Play enough surrounding audio to understand the event, speaker, language, and acoustic context. Confirm the project’s overlap and silence policy.
{% endstep %}

{% step %}

#### 2. Zoom to the boundary

Use waveform and spectrogram detail to locate the exact start and end. Keep enough context visible to avoid cutting off leading or trailing signal.
{% endstep %}

{% step %}

#### 3. Create the temporal annotation

Drag the intended range, assign the ontology class, and adjust the handles until the interval matches the policy.
{% endstep %}

{% step %}

#### 4. Add transcript and structure

Enter or review time-aligned text, speaker identity, required properties, relations, and Item Properties.
{% endstep %}

{% step %}

#### 5. Review the timeline

Check gaps, overlaps, duplicated events, inconsistent speaker identity, and boundary drift across the full recording.
{% endstep %}

{% step %}

#### 6. Save and route

Resolve validation, save, and use the current workflow action to submit, approve, reject, or escalate.
{% endstep %}
{% endstepper %}

## Quality review

| Review focus   | What to check                                                                               |
| -------------- | ------------------------------------------------------------------------------------------- |
| **Timing**     | Confirm start and end against waveform, spectrogram, playback, and project tolerance.       |
| **Overlap**    | Represent simultaneous speakers or sounds exactly as the ontology and instruction define.   |
| **Transcript** | Apply the same casing, punctuation, filler, normalization, and unintelligible-audio policy. |
| **Identity**   | Keep speaker or event identity stable throughout the recording.                             |
| **Coverage**   | Review the complete timeline for missed intervals and accidental gaps.                      |

{% hint style="warning" %}
A saved annotation is not automatically a production-ready annotation. Required values, boundary or timing policy, cross-item consistency, and the configured review stage still apply.
{% endhint %}

## Move from labels to governed data

Audio outputs should preserve the source recording identity, time basis, segment boundaries, speaker or event identity, transcript text, properties, relations, and item-level context. Keep long-recording curation, review cohorts, dataset versions, and releases traceable to the original media.

![Integrated Unitlab workflow connecting model assistance, annotation, review, and quality assurance](/files/yPoYNby78Kd9mxWTVqxy)

*Use workflows to keep model output, human correction, review, and approval in one traceable operating path.*

## Next steps

* Use [Detect Anything (SAM 1–SAM 3)](/documentation/auto-labeling/detect-anything-sam-1-sam-3) to calibrate interactive and batch assistance.
* Use [Multimodal overview](/documentation/multimodal-annotations/multimodal-overview) when related files or views must stay in one task.
* Curate difficult cases and review cohorts in [Data curation](/documentation/data/data-curation).
* Read the current [audio annotation product overview](https://unitlab.ai/en/audio-annotation) for the feature overview and current media.


# Medical Annotation

Annotate DICOM and medical imaging with synchronized multiplanar views, clinical ontologies, timelines, tracking, and AI-assisted segmentation.

Unitlab’s medical Workbench keeps a study together across axial, coronal, sagittal, 3D, and sequence views. Specialists can create precise findings, inspect them from synchronized perspectives, apply structured clinical properties, and route work through review without fragmenting one case into unrelated images.

{% hint style="info" %}
**Use this guide when:** you are labeling DICOM studies, volumetric scans, medical sequences, lesions, organs, findings, measurements, or clinical imaging datasets.
{% endhint %}

## See medical annotation in action

The demo shows synchronized medical views and annotation assistance inside the current clinical Workbench.

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/auto-labeling-4.mp4>" %}

[Open the demo in a new tab](https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/auto-labeling-4.mp4).

## Before you begin

1. Create or select a project whose data and ontology match this modality.
2. Confirm the project instructions define the unit of annotation, boundary or timing policy, required properties, and review route.
3. Open the project and enter the assigned item from the project data view or queue. The Workbench loads the modality-native editor inside the shared Unitlab shell.

See [Annotation Workbench](/documentation/annotations/annotation-workbench) for navigation, saving, item state, comments, issues, and workflow actions.

## Understand the medical work surface

The medical Workbench combines synchronized slice position, native medical views, 3D context, annotation geometry, sequence timelines, clinical ontology values, and workflow actions. A change created in the active view remains tied to the same study and can be inspected from sibling views.

![Synchronized axial, coronal, sagittal, and 3D DICOM views](/files/mwB1JcWbVojXdn4tU3V3)

*Multiple clinical views preserve one finding and one study context while supporting specialist inspection.*

## Supported annotation model

| Annotation type                | Use it for                                                                   |
| ------------------------------ | ---------------------------------------------------------------------------- |
| **Segmentation mask**          | Organs, tissues, tumors, lesions, and other pixel- or voxel-aligned regions. |
| **Polygon or contour**         | Editable boundaries on a slice or supported projection.                      |
| **Bounding box**               | Coarse localization of a finding or region of interest.                      |
| **Point or keypoint**          | Landmarks, centers, and precise anatomical locations.                        |
| **Line or measurement**        | Distances, axes, and elongated structures where supported by the project.    |
| **Temporal or sequence label** | A finding or state on a medical sequence timeline.                           |
| **Clinical property**          | Structured values attached to a finding or anatomy class.                    |
| **Item Property or relation**  | Study-level context and governed connections between findings.               |

Clinical label policy must separate what is visible, what is inferred, and what is not assessable. Geometry agreement alone does not guarantee agreement on diagnosis, certainty, or acquisition quality.

## Multiplanar editing and synchronized views

Choose the active view for annotation, navigate to the relevant slice, and inspect the same finding in synchronized orthogonal and 3D context. Use view switching to detect leakage into adjacent anatomy, missed extent, or a contour that looks plausible in one plane but inconsistent in another.

![Multiplanar medical annotation with synchronized slice position and geometry](/files/LzoxfX5854v6IZXVD8Du)

*Edits remain part of one study while specialists inspect the finding from complementary planes.*

## Clinical ontologies and properties

Use nested classes, required properties, relations, and Item Properties to encode anatomy, finding type, certainty, grade, acquisition context, and project-specific clinical values. Keep visible morphology separate from inferred diagnosis when the project requires that distinction.

![Clinical ontology with lesion and liver classes, attributes, relations, and temporal segmentation](/files/KmRWIh4leBC3lfoxgStM)

*The ontology carries clinical meaning that geometry alone cannot represent.*

## AI-assisted segmentation and sequence tracking

Use supported model assistance to propose tissue or lesion masks, then refine them in the native editor and inspect the result across views. For medical sequences, auto-tracking can propagate contours across time or frames; correct the first drift and review every transition. Model output remains a proposal until specialist approval.

![AI-assisted medical segmentation with editable tissue masks](/files/W9me8egb0yEW7QZFNqO0)

*Model assistance accelerates delineation while clinical review controls the final label.*

## Annotate one production item

{% stepper %}
{% step %}

#### 1. Open the study

Confirm the patient-safe study identifier, series, modality, orientation, ontology, instructions, and current review stage.
{% endstep %}

{% step %}

#### 2. Synchronize the views

Navigate to the target slice or frame and inspect the finding in the available axial, coronal, sagittal, 3D, or sequence context.
{% endstep %}

{% step %}

#### 3. Create the finding

Choose the class and geometry, then draw the mask, contour, box, point, or measurement in the active view.
{% endstep %}

{% step %}

#### 4. Inspect extent and continuity

Move through adjacent slices or frames and compare sibling views. Correct leakage, discontinuity, inconsistent identity, or missed extent.
{% endstep %}

{% step %}

#### 5. Complete clinical structure

Fill required clinical properties, study Item Properties, relations, and temporal values without inferring unsupported facts.
{% endstep %}

{% step %}

#### 6. Specialist review and route

Review full-study context, validation, comments, and issue history; save and route to the configured medical review or approval stage.
{% endstep %}
{% endstepper %}

## Quality review

| Review focus               | What to check                                                                              |
| -------------------------- | ------------------------------------------------------------------------------------------ |
| **Study identity**         | All labels must remain attached to the correct study, series, slice, and frame.            |
| **Cross-view consistency** | Inspect the finding in synchronized views, not only the plane where it was drawn.          |
| **Clinical semantics**     | Separate observation, certainty, diagnosis, and non-assessable values according to policy. |
| **Segmentation extent**    | Check adjacent slices, small islands, holes, leakage, and partial-volume boundaries.       |
| **Specialist ownership**   | Use qualified reviewers and explicit escalation for ambiguous clinical findings.           |

{% hint style="warning" %}
A saved annotation is not automatically a production-ready annotation. Required values, boundary or timing policy, cross-item consistency, and the configured review stage still apply.
{% endhint %}

## Move from labels to governed data

Medical outputs should preserve study and series identity, slice or frame position, geometry, clinical properties, relations, reviewer provenance, and ontology version. Use controlled dataset versions and releases for training or validation cohorts, and apply organizational privacy, access, and deployment requirements to protected health information.

![Integrated Unitlab workflow connecting model assistance, annotation, review, and quality assurance](/files/yPoYNby78Kd9mxWTVqxy)

*Use workflows to keep model output, human correction, review, and approval in one traceable operating path.*

## Next steps

* Use [Detect Anything (SAM 1–SAM 3)](/documentation/auto-labeling/detect-anything-sam-1-sam-3) to calibrate interactive and batch assistance.
* Use [Multimodal overview](/documentation/multimodal-annotations/multimodal-overview) when related files or views must stay in one task.
* Curate difficult cases and review cohorts in [Data curation](/documentation/data/data-curation).
* Read the current [medical annotation product overview](https://unitlab.ai/en/medical-annotation) for the feature overview and current media.


# Pathology Annotation

Annotate whole-slide images with deep zoom, tissue and cellular labels, pathology ontologies, AI assistance, and expert review.

Unitlab supports whole-slide pathology workflows from tissue overview to cellular detail. Teams can label regions of interest, tissue compartments, tumor margins, cells, nuclei, biomarkers, and slide-level properties without losing the spatial context of the original slide.

{% hint style="info" %}
**Use this guide when:** you are building computational pathology, histopathology, tissue segmentation, cell or nuclei detection, tumor modeling, or biomarker datasets.
{% endhint %}

## See pathology annotation in action

The demo shows current whole-slide navigation and annotation behavior for pathology data.

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/pathology-1.mp4>" %}

[Open the demo in a new tab](https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/pathology-1.mp4).

## Before you begin

1. Create or select a project whose data and ontology match this modality.
2. Confirm the project instructions define the unit of annotation, boundary or timing policy, required properties, and review route.
3. Open the project and enter the assigned item from the project data view or queue. The Workbench loads the modality-native editor inside the shared Unitlab shell.

See [Annotation Workbench](/documentation/annotations/annotation-workbench) for navigation, saving, item state, comments, issues, and workflow actions.

## Understand the pathology work surface

The pathology experience applies Unitlab’s image annotation and ontology model to very large whole-slide imagery. Deep zoom keeps the slide continuous while the Workbench provides geometry, properties, relations, item context, comments, history, and workflow actions.

![Whole-slide pathology view with a focused region and cellular-detail inset](/files/b9YiwfFPMNG9S1HVxEjb)

*One continuous slide supports overview inspection, region selection, and cellular annotation.*

## Supported annotation model

| Annotation type                 | Use it for                                                                     |
| ------------------------------- | ------------------------------------------------------------------------------ |
| **Bounding box or ROI**         | Tissue, tumor, lesion, cell cluster, or review regions.                        |
| **Tissue segmentation mask**    | Pixel-level tissue, tumor, necrosis, lesion, or biomarker areas.               |
| **Polygon or tissue region**    | Irregular tissue compartments, glands, margins, and other editable regions.    |
| **Cell or nuclei instance**     | Distinct cellular objects for detection, counting, and morphology analysis.    |
| **Polyline or tissue boundary** | Margins, vessels, and elongated structures.                                    |
| **Point or cell marker**        | Cell centers, nuclei, glands, or microscopic landmarks.                        |
| **Classification or finding**   | Tissue type, grade, stain, biomarker status, or another project-defined label. |
| **Slide Item Property**         | Source, cohort, quality, stain, acquisition, or another whole-slide value.     |
| **Relation**                    | A contextual connection between findings, cells, and tissue regions.           |

Define minimum object size, edge handling, touching-instance policy, magnification requirements, and whether findings are exhaustive or sampled before annotators begin.

## Deep zoom and multi-resolution review

Start at the whole-slide overview to understand tissue distribution, then move through region and cellular detail without creating disconnected crops. Record which magnification level is required for each decision. Reviewers should return to the broader tissue context before approving high-magnification labels.

![One pathology specimen shown at overview, region, and cellular resolutions](/files/NIN9uNcxPVVzfA84uqSf)

*Multi-resolution viewing preserves the relationship between a cell-level label and its tissue context.*

## Pathology ontologies

Define tissue regions, findings, cell types, class properties, relations, and slide Item Properties in one reusable ontology. Required values make incomplete findings visible; relations can connect a cellular observation to the relevant tissue region or project-defined context.

![Pathology ontology for tissue regions, findings, attributes, relations, and slide properties](/files/EVIfJXsFVw0TzDhA60bf)

*Ontology structure keeps microscopic geometry and slide-level clinical context distinct but connected.*

## AI-assisted repetitive labeling

Use Magic Touch to create an editable mask and Find Similar to propose matching structures on the current image. For dense cells or nuclei, calibrate on representative fields before expanding volume. Check merge and split errors, boundary leakage, false positives in background tissue, and missed morphology variants.

![Pathology cells with matching masks and Find Similar assistance](/files/xZ2v5kXntY2WNIwJm9M6)

*AI assistance can accelerate repetitive structures, but expert review remains responsible for morphology and label meaning.*

## Annotate one production item

{% stepper %}
{% step %}

#### 1. Survey the whole slide

Inspect tissue coverage, artifacts, empty regions, stain variation, orientation, and the project’s required review areas.
{% endstep %}

{% step %}

#### 2. Navigate to the correct resolution

Zoom from overview to region and cellular detail. Confirm the required magnification for the target label.
{% endstep %}

{% step %}

#### 3. Choose class and geometry

Select the tissue, finding, cell, or nuclei class and use the required ROI, mask, polygon, line, or point tool.
{% endstep %}

{% step %}

#### 4. Create and refine the annotation

Trace the intended boundary or instance, correct holes and touching objects, and keep enough surrounding tissue visible to interpret the structure.
{% endstep %}

{% step %}

#### 5. Add pathology structure

Complete class properties, slide Item Properties, findings, relations, and uncertainty or quality values exactly as the ontology requires.
{% endstep %}

{% step %}

#### 6. Review across scales and route

Inspect dense regions, edge cases, and broader tissue context; resolve validation, save, and submit to the configured expert review stage.
{% endstep %}
{% endstepper %}

## Quality review

| Review focus            | What to check                                                                                       |
| ----------------------- | --------------------------------------------------------------------------------------------------- |
| **Magnification**       | Use the required resolution for each label and review the result at both detail and context levels. |
| **Instance separation** | Check touching cells, merged nuclei, fragments, and duplicate instances.                            |
| **Tissue boundary**     | Review holes, folds, tears, staining artifacts, necrosis, and uncertain margins.                    |
| **Slide context**       | Validate stain, cohort, source, quality, and whole-slide properties.                                |
| **Expert calibration**  | Measure agreement on representative fields before scaling annotation volume.                        |

{% hint style="warning" %}
A saved annotation is not automatically a production-ready annotation. Required values, boundary or timing policy, cross-item consistency, and the configured review stage still apply.
{% endhint %}

## Move from labels to governed data

Pathology outputs should preserve slide identity, coordinate system, magnification context, region or instance geometry, class and slide properties, relations, ontology version, and reviewer provenance. Promote approved cohorts through dataset versions and releases so training and evaluation remain reproducible.

![Integrated Unitlab workflow connecting model assistance, annotation, review, and quality assurance](/files/yPoYNby78Kd9mxWTVqxy)

*Use workflows to keep model output, human correction, review, and approval in one traceable operating path.*

## Next steps

* Use [Detect Anything (SAM 1–SAM 3)](/documentation/auto-labeling/detect-anything-sam-1-sam-3) to calibrate interactive and batch assistance.
* Use [Multimodal overview](/documentation/multimodal-annotations/multimodal-overview) when related files or views must stay in one task.
* Curate difficult cases and review cohorts in [Data curation](/documentation/data/data-curation).
* Read the current [pathology annotation product overview](https://unitlab.ai/en/pathology-annotation) for the feature overview and current media.


# Geospatial Annotation

Annotate satellite, aerial, and large-raster imagery with deep zoom, spatial coordinates, structured ontologies, and AI-assisted segmentation.

Unitlab supports geospatial annotation across large satellite, aerial, and drone imagery. Teams can move from area overview to object detail, preserve geospatial context, label land cover and infrastructure, and review model-assisted geometry within governed workflows.

{% hint style="info" %}
**Use this guide when:** you are building remote-sensing, land-cover, agriculture, mapping, infrastructure, environmental, or disaster-response datasets.
{% endhint %}

## See geospatial annotation in action

The demo shows current large-image navigation and spatial annotation for geospatial data.

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/geospatial-1.mp4>" %}

[Open the demo in a new tab](https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/geospatial-1.mp4).

## Before you begin

1. Create or select a project whose data and ontology match this modality.
2. Confirm the project instructions define the unit of annotation, boundary or timing policy, required properties, and review route.
3. Open the project and enter the assigned item from the project data view or queue. The Workbench loads the modality-native editor inside the shared Unitlab shell.

See [Annotation Workbench](/documentation/annotations/annotation-workbench) for navigation, saving, item state, comments, issues, and workflow actions.

## Understand the geospatial work surface

The geospatial experience applies Unitlab’s visual Workbench to large spatial rasters. Deep zoom preserves a continuous image while annotation geometry, ontology values, comments, history, and workflow actions remain available around the active view.

![Satellite mosaic shown at overview, regional, and object-detail scales](/files/ndKATVgbYAxrZ0oUu8QY)

*Large-image support keeps broad spatial context available while labeling small objects and boundaries.*

## Supported annotation model

| Annotation type       | Use it for                                                                            |
| --------------------- | ------------------------------------------------------------------------------------- |
| **Bounding box**      | Vehicles, structures, assets, and other localized objects.                            |
| **Segmentation mask** | Roads, water, vegetation, buildings, damage, and land-cover regions.                  |
| **Polygon**           | Parcels, rooftops, fields, sites, and irregular boundaries.                           |
| **Line or polyline**  | Roads, paths, utilities, coastlines, and other linear features.                       |
| **Point or keypoint** | Poles, signs, landmarks, and inspection targets.                                      |
| **Skeleton**          | Defined landmark structures where a connected point model is required.                |
| **Cuboid**            | Perspective-aware 3D-like extent for supported aerial targets.                        |
| **Item Property**     | Source, capture condition, sensor, scene, and quality context for the complete image. |
| **Relation**          | Connections among buildings, roads, vehicles, parcels, and other objects.             |

Write spatial rules for tile edges, partial objects, minimum mapping unit, occlusion, shadows, seasonal change, coordinate reference, and whether repeated features require exhaustive coverage.

## Coordinate-aware spatial context

Keep georeferencing and spatial coordinates attached to the large image and its annotations. Zoom or pan should change only the view, not the underlying coordinate meaning. Confirm coordinate expectations before export, especially when downstream systems require a specific reference or projection.

![Spatial context retained across overview, focus, and detail](/files/rJ2XT5gXQ0S1Rixd4S5y)

*The same labeled feature remains grounded in the source image’s spatial context.*

## Geospatial ontologies and relations

Use hierarchical classes for buildings, roads, land cover, crops, utilities, and project-specific targets. Add required properties such as type, condition, surface, confidence, or source; use Item Properties for capture and scene context; connect objects with governed relations such as adjacent-to or connected-to.

![Geospatial ontology with Building and Road classes, attributes, relations, and Item Properties](/files/NhVDZMsKXWfxTwOaEOo8)

*Structured ontology values turn shapes into consistent spatial training data.*

## AI-assisted masks and repetitive features

Use supported model assistance, Magic Touch, or Find Similar to propose visual labels, then inspect every boundary and object in source context. Repetitive rooftops, roads, fields, or vegetation can accelerate well, but seasonal variation, shadows, occlusion, small structures, and domain shift require human correction.

![AI-assisted geospatial masks with Find Similar, Magic Touch, and review](/files/EVBoCT3lEfShkE0iv6Ro)

*Assisted geospatial labels remain editable proposals inside the human review workflow.*

## Annotate one production item

{% stepper %}
{% step %}

#### 1. Survey the image

Inspect the full area, source, capture condition, orientation, spatial reference, and project coverage rule before labeling details.
{% endstep %}

{% step %}

#### 2. Navigate to the target scale

Use deep zoom to locate the region or object while keeping enough surrounding context to classify it correctly.
{% endstep %}

{% step %}

#### 3. Create spatial geometry

Choose the ontology class and draw the required box, mask, polygon, line, point, skeleton, or cuboid.
{% endstep %}

{% step %}

#### 4. Refine and connect

Correct boundaries at the specified resolution, complete properties, add Item Properties, and create ontology-defined spatial relations.
{% endstep %}

{% step %}

#### 5. Review at multiple scales

Inspect object detail, neighboring context, tile or image edges, missed instances, and class consistency across the wider area.
{% endstep %}

{% step %}

#### 6. Save, route, and export

Resolve validation, save, submit through the workflow, and export only after confirming coordinate and format requirements.
{% endstep %}
{% endstepper %}

## Quality review

| Review focus          | What to check                                                                             |
| --------------------- | ----------------------------------------------------------------------------------------- |
| **Spatial reference** | Preserve image identity, coordinate meaning, and expected export reference.               |
| **Coverage**          | Apply the same exhaustive or sampled labeling rule across the full area.                  |
| **Boundary policy**   | Check shadows, occlusion, partial objects, seasonal change, and the minimum mapping unit. |
| **Topology**          | Review connected lines, adjacent polygons, overlaps, gaps, and object relations.          |
| **Scale**             | Confirm geometry at detail resolution and semantic correctness at regional context.       |

{% hint style="warning" %}
A saved annotation is not automatically a production-ready annotation. Required values, boundary or timing policy, cross-item consistency, and the configured review stage still apply.
{% endhint %}

## Move from labels to governed data

Geospatial outputs should preserve source raster identity, spatial coordinates, geometry, ontology values, relations, version, and reviewer provenance. Use curated cohorts, dataset versions, and releases to separate geography, season, sensor, and domain conditions for reproducible model development.

![Integrated Unitlab workflow connecting model assistance, annotation, review, and quality assurance](/files/yPoYNby78Kd9mxWTVqxy)

*Use workflows to keep model output, human correction, review, and approval in one traceable operating path.*

## Next steps

* Use [Detect Anything (SAM 1–SAM 3)](/documentation/auto-labeling/detect-anything-sam-1-sam-3) to calibrate interactive and batch assistance.
* Use [Multimodal overview](/documentation/multimodal-annotations/multimodal-overview) when related files or views must stay in one task.
* Curate difficult cases and review cohorts in [Data curation](/documentation/data/data-curation).
* Read the current [geospatial annotation product overview](https://unitlab.ai/en/geospatial-annotation) for the feature overview and current media.


# Detect Anything (SAM 1–SAM 3)

Choose and operate Unitlab’s SAM-assisted segmentation, class-prompt detection, and supported geometry workflows.

Unitlab combines interactive segmentation, class-prompted detection, and temporal propagation under one governed annotation workflow. Choose the smallest operation that produces the geometry you need, then review the proposals before moving the item forward.

![Magic Touch and repeated-object assistance in the Unitlab Workbench](/files/wARddV8qJZlyIyZjxOIP)

*Current Unitlab product visual: one assisted selection becomes a set of editable object proposals.*

{% hint style="info" %}
The current Workbench exposes **SAM 1** and **SAM 3** for interactive segmentation. **Detect all objects** is the SAM 3 class-prompt operation. Temporal propagation belongs to [Bidirectional Auto-Tracking](/documentation/auto-labeling/bidirectional-auto-tracking); SAM 2 is not a selectable Detect-all generation in the current UI.
{% endhint %}

### Choose the operation

| Goal                                                | Operation                                                                               | Current output                                            |
| --------------------------------------------------- | --------------------------------------------------------------------------------------- | --------------------------------------------------------- |
| Isolate one object from a point or guided region    | Magic Touch with SAM 1 or SAM 3                                                         | Editable segmentation mask                                |
| Detect every visible instance of one class          | Detect all objects with SAM 3                                                           | Bounding box, polygon, mask, or cuboid                    |
| Find more objects that resemble a confirmed example | [Find Similar](/documentation/auto-labeling/find-similar)                               | Bounding box, polygon, mask, or cuboid proposals          |
| Describe the target in natural language             | [Prompt Labeling](/documentation/auto-labeling/prompt-labeling)                         | Class-bound SAM 3 proposals on the current image or frame |
| Propagate an object through a sequence              | [Bidirectional Auto-Tracking](/documentation/auto-labeling/bidirectional-auto-tracking) | Frame- or slice-aware object track                        |

### Supported scope

**Detect all objects** is available for an image and for the current video frame. Select an ontology class whose geometry is one of:

* bounding box;
* polygon;
* segmentation mask;
* cuboid / 3D box.

The operation is class-aware. Unitlab writes accepted output to the active class, so the prompt does not replace ontology governance.

### Before you start

* Open an image or a video frame in the [Annotation Workbench](/documentation/annotations/annotation-workbench).
* Confirm the intended class exists in the project ontology.
* Choose the geometry required by the downstream model.
* Read the project Instructions for inclusion, exclusion, truncation, and occlusion policy.
* Start on a representative item before processing dense or unusual scenes.

### Detect all objects

{% stepper %}
{% step %}

#### Select the class

Open **Classes** and select the exact ontology class. Detect all is unavailable when the active class uses an unsupported geometry.
{% endstep %}

{% step %}

#### Open Auto-labeling

Select the wand action in the vertical toolbar or press **S**. You can also open the class action for **Detect all objects of this class**.
{% endstep %}

{% step %}

#### Set the prompt

The prompt begins with the class name. Keep it when the ontology name is visually specific, or replace it with a clearer description. The current input accepts up to 300 characters.
{% endstep %}

{% step %}

#### Run detection

Select **Detect all objects**. Unitlab evaluates the current image or current frame and returns candidate instances in the active geometry.
{% endstep %}

{% step %}

#### Review and correct

Inspect false positives, misses, overlap, truncation, small objects, and boundary quality. Correct accepted geometry with the standard Workbench tools.
{% endstep %}
{% endstepper %}

### Geometry guidance

| Geometry        | Use it when                                              | Review closely                                      |
| --------------- | -------------------------------------------------------- | --------------------------------------------------- |
| Bounding box    | Coarse localization is sufficient                        | Tightness, truncation, overlap, tiny-object misses  |
| Polygon         | Boundary shape matters but a raster mask is not required | Vertex placement, holes, self-intersection          |
| Mask            | Pixel membership drives training or measurement          | Leakage, holes, thin structures, touching instances |
| Cuboid / 3D box | Orientation and spatial extent are part of the label     | Vanishing direction, depth edges, ground contact    |

### SAM selection

Use **SAM 1** when a stable interactive mask is sufficient and the operator wants a familiar click-guided segmentation path. Use **SAM 3** for the current concept-aware segmentation and class-prompted Detect-all experience. Validate either choice on the same representative sample before standardizing it for a team.

{% hint style="warning" %}
A model proposal is not ground truth. A qualified annotator must validate class, geometry, attributes, relations, temporal identity, and workflow outcome.
{% endhint %}

### Quality checklist

* The active ontology class is correct.
* Every proposal follows the project’s inclusion and occlusion policy.
* Duplicate and strongly overlapping proposals are removed.
* Small, partially visible, and edge-of-frame instances were inspected.
* Geometry was corrected at the zoom level required by Instructions.
* Required properties and relations are complete before submission.

### Troubleshooting

| Symptom                          | Check                                                                                           |
| -------------------------------- | ----------------------------------------------------------------------------------------------- |
| Wand action is unavailable       | Confirm the item is an image or video and the active class is box, polygon, mask, or cuboid     |
| No objects are returned          | Use a more concrete prompt, verify the object is visible, and test another representative frame |
| Too many unrelated objects       | Narrow the prompt with object type, visual context, or distinguishing state                     |
| Boundary quality is insufficient | Switch to mask or polygon output and correct with brush, eraser, or vertex tools                |
| Results drift across time        | Detect on a reliable frame, then use bidirectional tracking and review the timeline             |

### Related guides

* [Prompt Labeling](/documentation/auto-labeling/prompt-labeling)
* [Find Similar](/documentation/auto-labeling/find-similar)
* [Bidirectional Auto-Tracking](/documentation/auto-labeling/bidirectional-auto-tracking)
* [Image Annotation](/documentation/annotations/image-annotation)
* [Video Annotation](/documentation/annotations/video-annotation)


# Find Similar

Use one verified box, polygon, mask, or cuboid to find and review similar objects in the current image or frame.

Find Similar turns one verified object into a reviewed set of visually related proposals. It is designed for dense, repetitive scenes such as products, crops, cells, components, people, or vehicles.

![Magic Touch selects one object and Find Similar proposes matching instances](/files/wARddV8qJZlyIyZjxOIP)

### What Find Similar does

Find Similar is a contextual Workbench action. It appears after you select a compatible object and searches the **current image or current video frame**. It does not search an entire dataset.

| Seed geometry     | Proposal geometry | Supported |
| ----------------- | ----------------- | --------- |
| Bounding box      | Bounding box      | Yes       |
| Polygon           | Polygon           | Yes       |
| Segmentation mask | Segmentation mask | Yes       |
| Cuboid / 3D box   | Cuboid / 3D box   | Yes       |

The current experience exposes a confidence threshold, removes strong overlaps with committed annotations, and holds new results as pending proposals. Use **Clear** to discard the proposal set or **Accept all** after review.

{% hint style="info" %}
Find Similar is example-driven. [Prompt Labeling](/documentation/auto-labeling/prompt-labeling) is text-driven. Both produce proposals, but they solve different discovery problems.
{% endhint %}

### Before you start

* Choose a seed that is correctly classified and tightly annotated.
* Prefer a clear, representative instance rather than a heavily occluded edge case.
* Confirm repeated objects are visually similar enough for example-based retrieval.
* Zoom so the seed boundary can be inspected before search.
* Read the project policy for duplicates, partial objects, and minimum visible area.

### Find repeated objects

{% stepper %}
{% step %}

#### Create or select the seed

Draw a box, polygon, mask, or cuboid around one representative instance. Select the finished object in the canvas or Objects panel.
{% endstep %}

{% step %}

#### Start Find Similar

Choose **Find Similar** from the contextual header action. If it is missing, verify the seed geometry is supported and the object is selected.
{% endstep %}

{% step %}

#### Tune confidence

Start with a conservative threshold. Lower it when recall is too low; raise it when unrelated candidates dominate.
{% endstep %}

{% step %}

#### Inspect pending proposals

Review the complete canvas, not only the area around the seed. Compare each proposal with the class definition and check overlap with existing annotations.
{% endstep %}

{% step %}

#### Accept or clear

Choose **Accept all** only when the set is appropriate. Otherwise clear the set, improve the seed or threshold, and run again.
{% endstep %}

{% step %}

#### Finish annotation QA

Correct boundaries, complete required properties, and submit through the configured workflow.
{% endstep %}
{% endstepper %}

### Current auto-labeling demo

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/auto-labeling-2.mp4>" %}

*Live Unitlab demo used on the Video Annotation product page. The result remains editable and reviewable in the Workbench.*

### Threshold strategy

| Result pattern                          | Adjustment                                                          |
| --------------------------------------- | ------------------------------------------------------------------- |
| Many false positives                    | Raise confidence or choose a more distinctive seed                  |
| Similar objects are missed              | Lower confidence gradually and inspect the complete set             |
| One object receives duplicate proposals | Confirm the seed is committed and inspect overlap suppression       |
| Different states are mixed              | Use a more specific class or split the work by state/property       |
| Scale changes reduce recall             | Seed a second representative scale and review it as a separate pass |

### Review controls

For each proposal, verify:

* class and instance identity;
* geometry tightness or boundary precision;
* truncation and occlusion policy;
* duplicates and overlap with existing objects;
* required class properties;
* relations to other objects;
* consistency with nearby manually labeled examples.

{% hint style="warning" %}
Similarity is not semantic proof. A visually close result can still violate the ontology, and a valid instance can be visually different from the seed.
{% endhint %}

### When to use another tool

| Need                                                 | Better choice                                                                                  |
| ---------------------------------------------------- | ---------------------------------------------------------------------------------------------- |
| Find objects from a natural-language concept         | [Prompt Labeling](/documentation/auto-labeling/prompt-labeling)                                |
| Detect all visible instances of the active class     | [Detect Anything](/documentation/auto-labeling/detect-anything-sam-1-sam-3)                    |
| Propagate the same instance through frames or slices | [Bidirectional Auto-Tracking](/documentation/auto-labeling/bidirectional-auto-tracking)        |
| Search across many assets                            | [Embeddings, similarity, and outliers](/documentation/data/embeddings-similarity-and-outliers) |
| Process a project population before human review     | [Batch Auto-Labeling](/documentation/auto-labeling/batch-auto-labeling)                        |

### Operational guidance

Find Similar does not consume the standard AI-inference quota in the current implementation. Treat that as an execution detail, not a reason to skip review. Measure accepted proposals and correction rate on a representative cohort before using it as a standard labeling step.


# Prompt Labeling

Describe a visual concept in natural language and create editable, class-bound SAM 3 proposals.

Prompt Labeling lets an annotator describe a visual concept in natural language and create editable, class-bound proposals on the current image or frame. It is the prompt-authoring workflow behind the current SAM 3 **Detect all objects** action.

![Prompt Auto-Labeling identifies a described concept in the current image](/files/YITuAztNqIqU24WphHXp)

### Supported scope

| Dimension | Current behavior                                             |
| --------- | ------------------------------------------------------------ |
| Data      | Image and current video frame                                |
| Geometry  | Bounding box, polygon, segmentation mask, cuboid / 3D box    |
| Class     | Output is written to the active ontology class               |
| Prompt    | Defaults to the class name; up to 300 characters             |
| Result    | Editable candidate annotations                               |
| Quota     | One AI-inference unit is consumed after a successful request |

Prompt Labeling is not a dataset-wide text search and does not create a new ontology class. It uses the prompt to find instances, then applies the selected class and geometry.

### Write effective prompts

A good prompt names one visible concept and, only when necessary, adds a short disambiguator.

| Prompt quality | Example                       | Why                                                      |
| -------------- | ----------------------------- | -------------------------------------------------------- |
| Strong         | **yellow safety helmet**      | Concrete object and distinguishing state                 |
| Strong         | **white delivery van**        | Object plus visible attribute                            |
| Strong         | **tumor region in the liver** | Region plus anatomical context                           |
| Weak           | **all important things**      | Ambiguous and not visually testable                      |
| Weak           | **person, helmet, vehicle**   | Mixes several ontology classes                           |
| Weak           | **unsafe**                    | Describes a judgment rather than a stable visible object |

Use project Instructions to standardize prompts when several operators work on the same class. Record class-specific examples, known exclusions, and failure cases.

### Run Prompt Labeling

{% stepper %}
{% step %}

#### Select the target class

Choose the ontology class and its supported geometry. The class should represent the semantic label you intend to store.
{% endstep %}

{% step %}

#### Open Auto-labeling

Select the wand action or press **S**. The popup shows the active class and pre-fills the prompt with its name.
{% endstep %}

{% step %}

#### Refine the prompt

Keep the class name when it is precise. Otherwise add a short visible qualifier. Use the reset action to return to the class name.
{% endstep %}

{% step %}

#### Detect all objects

Run the request on the current image or frame. Wait for the proposal set to appear before navigating away.
{% endstep %}

{% step %}

#### Review the full result set

Check missed instances, false positives, boundary quality, duplicate overlap, and the effect of the chosen geometry.
{% endstep %}

{% step %}

#### Correct and submit

Edit proposals with the standard Workbench tools, complete required properties, and follow the project’s Annotate/Review workflow.
{% endstep %}
{% endstepper %}

### Prompt-to-geometry design

| Downstream task           | Recommended geometry | Prompt guidance                                                 |
| ------------------------- | -------------------- | --------------------------------------------------------------- |
| Object detection          | Bounding box         | Name one countable object                                       |
| Instance segmentation     | Mask                 | Name one object or region with visible boundaries               |
| Boundary-aware labeling   | Polygon              | Use the same semantic class as the polygon ontology             |
| Oriented spatial labeling | Cuboid / 3D box      | Name the physical object; review perspective and depth manually |

### Common failure modes

| Symptom                               | Likely cause                     | Recovery                                                              |
| ------------------------------------- | -------------------------------- | --------------------------------------------------------------------- |
| Correct concept, wrong class          | Active class was not changed     | Clear proposals, select the correct class, rerun                      |
| Good large objects, missed small ones | Scale or visibility limits       | Use a clearer frame, lower threshold where available, or add manually |
| Background regions are included       | Prompt is too broad              | Add visible context or use mask correction                            |
| Several concepts are mixed            | Prompt contains multiple classes | Run one class at a time                                               |
| Video results vary by frame           | Each request evaluates one frame | Use a reliable seed frame and continue with tracking                  |

{% hint style="warning" %}
Do not encode hidden business rules only in a prompt. Inclusion, exclusion, geometry, occlusion, and property policy belong in the project Instructions and ontology.
{% endhint %}

### Prompt governance

For production programs, retain a small prompt register with:

* project and ontology version;
* class and geometry;
* approved prompt;
* representative examples and exclusions;
* known domain failures;
* accepted-proposal and correction rate;
* owner and last validation date.

### Related guides

* [Detect Anything (SAM 1–SAM 3)](/documentation/auto-labeling/detect-anything-sam-1-sam-3)
* [Find Similar](/documentation/auto-labeling/find-similar)
* [Image Annotation](/documentation/annotations/image-annotation)
* [Project Instructions](/documentation/projects/project-instructions)


# Bidirectional Auto-Tracking

Track objects forward, backward, or through the full video, DICOM, or NIfTI sequence from one reliable seed.

Bidirectional Auto-Tracking propagates a verified object from one reliable seed through a temporal sequence. It supports full, forward, and backward runs so an annotator can start from the clearest frame or slice rather than the beginning.

![Multi-object Auto-Tracking in a video sequence](/files/KXWnxNuLB6LxREiHPU2F)

*Current Unitlab product visual: a selected group of objects is propagated while separate identities are preserved.*

![Bidirectional lesion tracking across a medical sequence](/files/LuhXIWfDfQ201dBcQQfc)

*Current Unitlab medical visual: the same region is tracked across the sequence and remains available for expert correction.*

### Supported data and actions

| Surface | Sequence unit                            | Current tracking actions           |
| ------- | ---------------------------------------- | ---------------------------------- |
| Video   | Frames                                   | Full annotation, forward, backward |
| DICOM   | Frames or slices in the medical sequence | Full annotation, forward, backward |
| NIfTI   | Volume slices / sequence positions       | Full annotation, forward, backward |

Supported temporal object geometries include bounding boxes, polygons, masks/brush annotations, and cuboids where the active editor and class allow them.

### Choose a direction

| Action                    | Range                                   | Use it when                                                          |
| ------------------------- | --------------------------------------- | -------------------------------------------------------------------- |
| **Track full annotation** | Both directions from the selected frame | The best seed is in the middle and the complete sequence is required |
| **Track forward**         | Selected frame to later frames          | The object first appears clearly at or after the seed                |
| **Track backward**        | Selected frame to earlier frames        | A later frame contains the clearest appearance                       |

Backward is unavailable on the first frame. Forward is unavailable when no later frame remains.

### Track one object

{% stepper %}
{% step %}

#### Choose a reliable seed

Navigate to the frame or slice where the object is clear, sufficiently large, and minimally occluded.
{% endstep %}

{% step %}

#### Create precise geometry

Draw the object using the project’s required class and geometry. A loose or incorrect seed propagates error.
{% endstep %}

{% step %}

#### Open the object action

Right-click the object or its selected timeline segment, choose **Auto track**, and select full, forward, or backward.
{% endstep %}

{% step %}

#### Monitor progress

The timeline displays the growing object track and keyframes. The object menu changes to **Stop Tracking** while a run is active.
{% endstep %}

{% step %}

#### Review discontinuities

Inspect occlusion, re-entry, camera cuts, fast motion, blur, scale changes, and crossings with similar objects.
{% endstep %}

{% step %}

#### Correct and continue

Edit the first unreliable frame, preserve identity, and rerun a shorter segment when necessary.
{% endstep %}
{% endstepper %}

### Track multiple objects

Hold **Command** on macOS or **Ctrl** on Windows/Linux and select several supported objects. Open **Auto track**, then choose the direction. Unitlab submits the selected objects together but writes each result to its own track.

{% hint style="info" %}
Multi-object tracking preserves separate identities. It does not merge selected objects into one annotation.
{% endhint %}

### Tracking versus interpolation

| Question                            | Auto-Tracking                              | Interpolation                               |
| ----------------------------------- | ------------------------------------------ | ------------------------------------------- |
| How intermediate labels are created | Model prediction                           | Geometry between manual keyframes           |
| Best fit                            | Complex but visually trackable motion      | Smooth change between known states          |
| Main risk                           | Identity drift or confident false geometry | Missed non-linear motion or topology change |
| Review focus                        | Occlusion, re-entry, crossings, drift      | Keyframe placement and transition shape     |

### Video demo

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/auto-labeling-2.mp4>" %}

### Medical guidance

For DICOM and NIfTI work:

* start from the slice where the finding or anatomy is most distinct;
* verify every anatomical plane or synchronized view used by the project;
* inspect entry and exit slices carefully;
* correct partial-volume boundary drift;
* route clinically significant output through expert review;
* avoid treating propagation as volumetric ground truth without slice-level inspection.

### Review a completed track

* Confirm the track begins and ends at the correct positions.
* Scrub through every transition around occlusion, re-entry, and motion change.
* Check identity when similar objects cross.
* Inspect geometry at representative zoom.
* Verify required dynamic properties at their change points.
* Confirm the object’s class and relations remain correct.
* Submit the item through the configured review route.

{% hint style="warning" %}
Full tracking is bidirectional propagation from one seed, not automatic approval. A reviewer must still verify the complete temporal range.
{% endhint %}

### Troubleshooting

| Symptom                      | Recovery                                                                 |
| ---------------------------- | ------------------------------------------------------------------------ |
| Tracking action is disabled  | Move away from the first/last boundary or select a supported object      |
| Track drifts after occlusion | Stop, correct the first unreliable frame, and restart from a new seed    |
| Two identities swap          | Split the run around the crossing and verify each object separately      |
| Mask degrades over time      | Add a corrected mask keyframe and rerun a shorter segment                |
| Medical boundary disappears  | Seed a clearer slice and review the entry/exit range under expert policy |

### Related guides

* [Video Annotation](/documentation/annotations/video-annotation)
* [Medical Annotation](/documentation/annotations/medical-annotation)
* [Validation and conditional logic](/documentation/ontologies/validation-and-conditional-logic)


# Batch Auto-Labeling

Run models at scale with workflow Model stages, human correction routes, queues, monitoring, and failure recovery.

Batch Auto-Labeling moves model inference from an annotator’s current item into a repeatable project workflow. Unitlab uses a **Model stage** to run the selected model, save its predictions into annotation history, and route the item to human annotation, review, another automated stage, or a terminal state.

![Unitlab workflow canvas with the Model stage available](/files/k5Ey2bhjx3nZVMH4sStR)

### Recommended production pattern

{% code collapsedlinecount="10" %}

```mermaid
flowchart LR
  P["Project"] --> M["Model stage"]
  M --> A["Annotate / correction"]
  A --> R["Review"]
  R -->|Approve| C["Complete"]
  R -->|Reject| A
  M -->|Failure| E["Visible error state"]
```

{% endcode %}

A reliable first rollout uses **Model → Annotate → Review**. Direct Model → Review routing is appropriate only after the model is calibrated on the target domain and the review team can detect systematic errors.

### Before you start

* A project with representative data and an approved ontology.
* A public or private AI model whose input and output contract matches the data.
* A class mapping from model outputs to ontology classes.
* Defined thresholds, failure ownership, and human acceptance criteria.
* Permission to edit and apply the project workflow.

### Configure batch auto-labeling

{% stepper %}
{% step %}

#### Select or integrate the model

Open **Public AI Models** or **My AI Models**. Confirm running state, supported data, output geometry, version, and owner. See [Bring your own Models](/documentation/auto-labeling/bring-your-own-models) for private endpoints.
{% endstep %}

{% step %}

#### Open the project workflow

From the project, open **Workflows**. Start from the project’s current graph or an approved reusable workflow.
{% endstep %}

{% step %}

#### Add a Model stage

Place **Model** after Project or another intended entry stage. Configure the model, generic data type, threshold controls, queue scope, and class mappings exposed for that model.
{% endstep %}

{% step %}

#### Add human control

Route successful predictions to **Annotate** for correction or **Review** for acceptance. Configure a rejection path back to the appropriate correction stage.
{% endstep %}

{% step %}

#### Define failure handling

Ensure model failures remain visible in an Error state with a named owner. Do not route an empty or malformed response directly to Complete.
{% endstep %}

{% step %}

#### Save and apply

Validate graph reachability and apply the workflow. Review the impact before replacing an active project workflow with items already in flight.
{% endstep %}

{% step %}

#### Add the batch population

Upload or attach the intended data. Each item enters the workflow and is dispatched when it reaches the Model stage.
{% endstep %}

{% step %}

#### Monitor and review

Use project queues and status filters to inspect Processing, Error, Annotate, Review, and Complete populations. Measure human correction before increasing volume.
{% endstep %}
{% endstepper %}

### Model-stage configuration

| Decision          | Production guidance                                                 |
| ----------------- | ------------------------------------------------------------------- |
| Model and version | Pin the approved integration and record its owner                   |
| Input data        | Match image, video, audio, text, or medical support                 |
| Output mapping    | Map every emitted class and geometry intentionally                  |
| Threshold         | Calibrate on the target domain; do not copy a generic default       |
| Queue scope       | Start with a representative batch or selected queue                 |
| Success route     | Prefer human correction or review before Complete                   |
| Failure route     | Keep failures visible and recoverable                               |
| Change control    | Re-test after endpoint, model, prompt, mapping, or ontology changes |

### Monitor the run

A Model-stage item shows **Processing** while inference runs. Successful predictions are saved as normal annotation history and advance through the configured route. A failed item moves to an explicit error state.

Track at least:

* total items entering the Model stage;
* completed, processing, and failed counts;
* empty-output rate;
* per-class proposal count;
* correction and deletion rate;
* reviewer rejection rate;
* latency and timeout rate;
* model and ontology version.

{% hint style="warning" %}
Batch throughput is not quality. Approve scale only after the correction rate, missed-instance rate, and failure behavior are stable on representative data.
{% endhint %}

### Safe rollout

| Phase            | Scope                     | Exit criterion                                             |
| ---------------- | ------------------------- | ---------------------------------------------------------- |
| Contract test    | A few known items         | Request, response, mapping, and failure states are valid   |
| Calibration      | Representative cohort     | Threshold and class behavior are acceptable                |
| Controlled batch | One queue or source slice | Human correction is stable and failures are owned          |
| Production       | Approved population       | Monitoring, review, rollback, and provenance are operating |

### Recovery

* Fix the model endpoint or mapping before retrying failed items.
* Inspect remote state before repeating a mutation to avoid duplicate work.
* Re-run only the affected cohort when possible.
* If a workflow change would reset in-flight work, review the impact count and schedule the change.
* Preserve model version and correction evidence in the release record.

### Related guides

* [Model stages](/documentation/workflows/model-stages)
* [Queues overview](/documentation/queues/queues-overview)
* [Bring your own Models](/documentation/auto-labeling/bring-your-own-models)
* [Create a release](/documentation/releases/create-a-release)


# Batch Classification

Classify project cohorts operationally by applying governed tags to selected items without confusing metadata with ontology labels.

Batch Classification applies one project tag to many selected items from the project data view. It is a fast way to create operational cohorts such as **needs-review**, **night**, **domain-a**, **priority-source**, or **holdout**.

![Project data view where operators select the items to classify with tags](/files/hVajfP3eULufUpzYOzGc)

*Open a project’s data view, select items, then use the selection action bar.*

{% hint style="info" %}
Project tags are metadata for filtering and operations. They are not ontology classifications. Use an **Item Property** when the value must be part of the annotation schema, versioned ground truth, or export.
{% endhint %}

### Choose tags or Item Properties

| Need                                                 | Use                    |
| ---------------------------------------------------- | ---------------------- |
| Build a temporary operational cohort                 | Project tag            |
| Filter and assign a batch                            | Project tag            |
| Mark source, campaign, or intake state               | Project tag            |
| Train a model on an image-level class                | Ontology Item Property |
| Require a reviewer to validate the value             | Ontology Item Property |
| Export the value as governed annotation ground truth | Ontology Item Property |

### Apply a tag to selected items

{% stepper %}
{% step %}

#### Open the project

Open **Projects**, choose the project, then open its **Data / Datasets** view.
{% endstep %}

{% step %}

#### Narrow the population

Use status, assignee, class, source, file, tag, or advanced filters so the visible set represents the intended cohort.
{% endstep %}

{% step %}

#### Select items

Select individual rows or cards. Use **Select all** only after confirming whether it targets the current page or the complete filtered result set.
{% endstep %}

{% step %}

#### Open Tags

The bulk action bar appears after selection. Choose **Tags**.
{% endstep %}

{% step %}

#### Choose or create a tag

Select an existing project tag or create a new name. Tag names are limited to 50 characters in the current UI.
{% endstep %}

{% step %}

#### Apply and verify

Unitlab adds the tag to the selected items. Filter by the new tag and confirm the count and sample before using the cohort in assignment, review, or release decisions.
{% endstep %}
{% endstepper %}

### Naming conventions

Prefer stable, machine-readable names:

| Pattern           | Example                    |
| ----------------- | -------------------------- |
| Domain            | **domain-construction**    |
| Capture condition | **condition-night**        |
| QA state          | **qa-needs-expert-review** |
| Source            | **source-partner-a**       |
| Experiment        | **exp-vehicle-v3-holdout** |

Avoid tags such as **good**, **final**, or **test** without a documented owner and meaning. They become ambiguous across teams and time.

### Bulk-action safeguards

Before applying a tag to a large population:

* clear unrelated selections left from another view;
* inspect active filters and the selected count;
* confirm archived items are excluded unless intentionally targeted;
* use a small sample when introducing a new naming convention;
* avoid encoding sensitive information in a tag name;
* record the cohort definition when it affects training or evaluation.

### Verify the cohort

1. Filter the project by the tag.
2. Compare the result count with the selection count.
3. Open representative items from different sources and statuses.
4. Confirm no unintended archived or failed items are included.
5. If the cohort feeds a dataset or release, record the filter and tag definition.

### Common mistakes

| Mistake                                          | Consequence                                       | Correction                                                     |
| ------------------------------------------------ | ------------------------------------------------- | -------------------------------------------------------------- |
| Using a tag as ground truth                      | Value may not follow ontology/review/export rules | Create an Item Property and route through review               |
| Applying to all filtered results unintentionally | Cohort becomes too broad                          | Inspect selection scope and undo/correct before downstream use |
| Reusing an ambiguous tag                         | Different teams interpret it differently          | Rename by convention and document the definition               |
| Encoding PHI or secrets in tags                  | Sensitive data leaks into operational metadata    | Use approved identifiers and governance controls               |

### Related guides

* [Project data](/documentation/projects/project-data)
* [Data curation](/documentation/data/data-curation)
* [Properties, relations, and Item Properties](/documentation/ontologies/properties-relations-and-item-properties)
* [Attach datasets to projects](/documentation/datasets/attach-datasets-to-projects)


# Bring your own Models

Register, validate, map, secure, and operate private HTTP inference models inside Unitlab workflows.

Bring Your Own Model (BYOM) connects a private HTTP inference endpoint to Unitlab so proprietary, fine-tuned, or domain-specific models can participate in annotation and workflow automation.

![Current My AI Models catalog and Integrate External Model entry point](/files/2FEzY11WampBN2pZmac3)

![Bring your model into the Unitlab annotation workflow](/files/Lw6yquVO4YIvhXQr3l7H)

### Operating model

Unitlab stores the integration configuration and invokes the approved endpoint. Your team owns the model runtime, capacity, availability, version, authentication, and change control. Integrated models are private to the workspace and appear under **My AI Models**.

### Supported contracts

| Dimension        | Current support                                            |
| ---------------- | ---------------------------------------------------------- |
| Inputs           | Image, video, audio, text, medical                         |
| Visual outputs   | Bounding box, polygon, mask, skeleton, line, point, cuboid |
| Audio outputs    | Event and optional speech-recognition transcript           |
| Text outputs     | Entity                                                     |
| Lifecycle states | Running, Stopped, Integration unfinished                   |
| Workflow use     | Model stage and supported assisted/batch operations        |

### Before you start

* Deploy a reachable HTTPS endpoint that accepts POST requests.
* Identify the model owner, version, and on-call owner.
* Prepare authentication headers without exposing secrets in documentation.
* Define the exact input and output schema.
* Prepare representative validation data, including failure cases.
* Create or approve the destination ontology classes and integer mappings.
* Confirm the endpoint can handle the intended batch concurrency.

{% hint style="warning" %}
Do not paste live API keys, bearer tokens, or private endpoints into screenshots, tickets, or public documentation. Use an approved secret-management and rotation process.
{% endhint %}

### Integration wizard

{% stepper %}
{% step %}

#### Registration

Open **My AI Models** and choose **Integrate External Model**. Select the generic input type and output data type, then enter the model name, description, endpoint, headers, and parameters.
{% endstep %}

{% step %}

#### Validation

Upload a representative sample. Unitlab calls the endpoint and displays the validation status and raw response. Continue only when the endpoint is reachable and the response matches the selected schema.
{% endstep %}

{% step %}

#### Integration

Add organizational tags and define output classes. Map each integer class value to the intended class name, color, and annotation geometry.
{% endstep %}

{% step %}

#### Confirmation

Review the complete contract: model identity, input type, endpoint, validation status, tags, classes, and mappings. Confirm to create the private model.
{% endstep %}
{% endstepper %}

### Representative image request

The exact payload depends on the configured data type. A typical image request contains a signed source URL and optional crop context:

{% code collapsedlinecount="10" %}

```json
{
  "src": "https://signed-source-url.example/image.jpg",
  "coordinates": [[120, 80], [920, 680]],
  "rotation": 0
}
```

{% endcode %}

| Field           | Meaning                                                        |
| --------------- | -------------------------------------------------------------- |
| **src**         | Time-limited source URL the endpoint downloads                 |
| **coordinates** | Optional crop bounds; omitted or null for full-image inference |
| **rotation**    | Source orientation in degrees                                  |

A representative bounding-box response maps each result to an integer class:

{% code collapsedlinecount="10" %}

```json
{
  "bboxes": [[
    {
      "point": [[100, 50], [300, 50], [300, 240], [100, 240]],
      "class": 0
    }
  ]],
  "classes": ["person"]
}
```

{% endcode %}

Treat these examples as a contract starting point. The Validation step is authoritative for the selected input and output type.

### Class mapping

For every emitted class, define:

* stable integer value;
* human-readable name;
* destination geometry;
* destination ontology class;
* confidence interpretation;
* behavior for unknown or unmapped classes.

Never silently coerce an unsupported class into another ontology label. Reject or quarantine unmapped output.

### Production readiness

| Control           | Acceptance evidence                                            |
| ----------------- | -------------------------------------------------------------- |
| Endpoint security | HTTPS, approved authentication, secret rotation owner          |
| Availability      | Health checks, timeout, retry, capacity plan                   |
| Schema            | Successful and malformed-response tests                        |
| Mapping           | Every output has an intentional destination or rejection rule  |
| Calibration       | Threshold validated on target-domain data                      |
| Human control     | Annotate/Review route and correction policy                    |
| Observability     | Correlation ID, model version, latency, status, redacted error |
| Change control    | Revalidation after model, endpoint, schema, or mapping change  |

### Use the model in a workflow

1. Open the project workflow.
2. Add or select a **Model stage**.
3. Choose the integrated private model.
4. Configure thresholds, generic type, queue scope, and class mappings.
5. Route success to Annotate or Review.
6. Keep failure visible with a named owner.
7. Save and apply on a controlled cohort.
8. Monitor correction and failure rates before scaling.

### Manage the integration

From **My AI Models**, operators can:

* inspect Running, Stopped, or unfinished status;
* continue an unfinished integration;
* update endpoint or configuration;
* review tags and supported output;
* stop or retire a model under change control.

Re-run validation after material changes. Record model version and mapping version in the release provenance used for training or evaluation.

### Failure handling

| Failure                   | Response                                                    |
| ------------------------- | ----------------------------------------------------------- |
| Non-200 endpoint response | Check availability, authentication, and server logs         |
| 200 with invalid schema   | Compare the raw response with the selected output contract  |
| Empty predictions         | Distinguish a valid abstention from a model/runtime failure |
| Timeout                   | Inspect remote task state before retrying                   |
| Unknown class integer     | Stop routing and correct the class mapping                  |
| Capacity saturation       | Reduce concurrency or scale the endpoint                    |

### Related guides

* [Batch Auto-Labeling](/documentation/auto-labeling/batch-auto-labeling)
* [Model stages](/documentation/workflows/model-stages)
* [API keys and service identities](/documentation/security/api-keys-and-service-identities)
* [Cloud credential governance](/documentation/security/cloud-credential-governance)


# Multimodal overview

Understand grouped context, layouts, and one-active-editor behavior.

Multimodal annotation keeps related evidence together without pretending every panel is the same kind of editor. Data Groups define the unit; custom layouts define the presentation; one active panel defines where new annotation is created.

## Watch multimodal annotation

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/multimodal-1.mp4>" %}

[Open the demo in a new tab](https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/multimodal-1.mp4).

## Core multimodal capabilities

### Unified labeling interface

Unitlab can keep video, image, PDF, audio, text, medical, and structured evidence inside one work unit. Each tile uses its native viewer while the group preserves identity, ontology, workflow, and delivery context.

![Unified multimodal workspace with video, document, and audio evidence](/files/JQmJNJVVmr06fbCAGqin)

*The active tile is editable; the remaining tiles provide synchronized or contextual evidence.*

### Dynamic layouts

Choose the grid, list, or custom presentation required by the task. Resize panels, enlarge the current evidence, and preserve the layout with the Data Group so every annotator and reviewer sees the same operating context.

![Custom multimodal layout with grid, list, document, image, and audio views](/files/Z7AMj4XHa9FhoCcw3Vxi)

*A custom layout is part of the work-unit design, not a decorative preference.*

### Auto-Grouping

Auto-Grouping turns related filenames into governed Data Groups. Define the source folder, grouping keys, group-name pattern, tile rules, exclusions, and layout; review the preview before creating groups.

![MP4, CSV, PDF, and MP3 evidence grouped into one quality case](/files/PMDdjgvDNjBkvSihEseF)

*Grouping rules determine which files become one annotation unit and how they appear in the Workbench.*

### Multi-camera and cross-modal annotation

Use one Data Group for synchronized camera views of the same event or connect a visual finding to thermal, audio, text, or document evidence. Shared ontologies and relations keep the meaning governed across modalities.

![Four synchronized camera views with aligned cuboids](/files/j3bcITcPmAT1cxWXFXSv)

*Multi-camera layouts preserve event identity while each view retains its own perspective.*

![Package defect linked to thermal, audio, and inspection-document evidence](/files/sH8vZbEefsHes7gwM1p2)

*Cross-modal relations turn co-located evidence into an explicit, reviewable training-data structure.*

{% hint style="info" %}
**Use this area when:** a decision requires multiple cameras, modalities, documents, audio, or clinical views to remain together.
{% endhint %}

### How this area fits into production

```mermaid
flowchart TB
  A["Data Group"]
  B["Custom layout"]
  C["Active panel"]
  D["Passive context panels"]
  E["One workflow task"]
  F["Context-preserving release"]
  A --> B
  B --> C
  B --> D
  C --> E
  D --> E
  E --> F
```

![Multiview grouped Workbench](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FuWV2FdAcECE2UrevUndn%2Fmultiview-video-workbench.png?alt=media\&token=77ad56ff-14dd-430b-9d53-cb23ba63956d)

*The layout preserves synchronized context while keeping the active annotation source visually distinct from passive evidence panels.*

### What this area controls

Multiview is the default project annotation experience for all six data families. It is a persistent workspace containing the project header, work-item navigation, mode/layout selector, resizable panel grid, one active editor, and one or more passive inspection panels.

### Two modes

| Mode               | What the panels show                                       | Editing model                                                       | Default layout |
| ------------------ | ---------------------------------------------------------- | ------------------------------------------------------------------- | -------------- |
| **Current file**   | Multiple views of the same datasource                      | One active editor; sibling panels mirror its changes in real time   | 1×1            |
| **Multiple files** | Neighboring work items from the current queue/filter scope | One selected panel is editable; other items remain passive previews | 1×3            |

Layouts range from 1×1 to 4×4. Users can resize panel boundaries, fullscreen a panel, and switch modes from the Workbench header. Mode, layout, and panel sizes persist for the user.

### Active and passive panels

Only the active panel can mutate annotations. It receives the full native editor: tools, hotkeys, selection, object/event/entity editing, player or page controls, comments, classes, properties, history, and workflow actions.

Passive panels can show media, annotations, labels, pages, frames, waveforms, text, or medical projections. In Current file mode they receive the active panel’s live annotation changes, but they cannot originate edits, change the active selection, save, control playback, change a PDF page, edit text entities, or alter waveform regions.

When the user activates a passive panel:

1. Unitlab visually selects it immediately.
2. Any in-progress save in the outgoing panel is allowed to settle.
3. Unsaved changes are saved or safely snapshotted.
4. The outgoing panel becomes passive and pauses modality-specific playback.
5. The incoming native editor loads and restores its panel state.
6. The route updates only after the editor is ready.
7. If hydration fails, Unitlab restores the previous active panel instead of blanking the workspace.

### Current file flow

1. The user opens a work item from a dataset, Task Queue, Batch Queue detail, filtered grid, or direct link.
2. Unitlab creates multiple panel sessions for the same datasource without duplicating data or history.
3. One panel is active; the rest are read-only siblings.
4. Active edits are broadcast to siblings in real time.
5. The user can activate another panel to work from a different view, frame, page, or zoom state.
6. Saving refreshes the shared history and updates every sibling to the persisted result.

Current-file examples:

* **Image:** compare the same image at different zoom or inspection states.
* **Video:** inspect different frames of the same video while sharing annotations and downloaded frames.
* **Audio:** inspect different time regions while one waveform editor remains authoritative.
* **Text:** compare different windows of the same text while entity/relation changes remain synchronized.
* **Document:** compare different PDF pages without confusing page changes with work-item navigation.
* **Medical:** assign Axial, Sagittal, Coronal, and 3D views to separate slots with synchronized annotation state.

### Multiple files flow

1. The user selects **Multiple files** and a layout.
2. Unitlab fills the grid with the active item and neighboring items from the current visible set.
3. The visible set preserves queue, upload-session, archive, search, status, class, and assignment filters.
4. The user activates any ready panel.
5. If the item belongs to another data family, the correct editor loads within the same Workbench shell.
6. Previous/next navigation continues through the filtered work-item sequence.
7. Empty tail slots and inaccessible items appear as panel-level states rather than replacing the whole page with an error.

This mode supports mixed review—for example an image, video, PDF, and audio item in one 2×2 layout—while guaranteeing that only the selected panel is editable.

### Navigation hierarchy

Unitlab maintains three distinct navigation levels:

* **Top-center previous/next:** changes the project work item and can cross data families or move between loose items and Data Groups.
* **Panel activation:** changes which visible panel is editable without leaving the Workbench.
* **Within-item media navigation:** changes a video frame, audio time region, text window, medical slice/view, or PDF page inside the current work item.

Keeping these levels separate is essential for predictable UX. A PDF page change must never advance to another datasource, and a medical slice change must never appear as another project item.

### Grouped Workbench

A Data Group opens as one grouped Workbench whose layout is fixed by the Auto-Grouping builder. The header shows the group layout rather than the general mode/layout selector. Activating a tile never tears down the group route.

The project’s unified previous/next sequence treats each group as one work unit and excludes its member tiles from the loose-item sequence. Grouped saves remain attributable to the group.

### Start with the right page

| Decision                       | Production guidance                                              |
| ------------------------------ | ---------------------------------------------------------------- |
| Define grouped data            | Start with Data Groups and layouts.                              |
| Choose one view at a time      | Use Current file mode.                                           |
| Compare or synchronize context | Use Multiple files mode.                                         |
| Build a domain pattern         | Use multi-camera, document-plus-media, or medical study layouts. |

### Operating boundary

* Project-item navigation and viewer navigation are distinct.
* Only the active editor creates annotation at a given moment; passive panels provide context.
* The release consumer must preserve or reconstruct group membership and layout meaning.

### A production-ready handoff

Users can identify the active panel, switch context intentionally, navigate within and between grouped items correctly, and explain how group structure appears downstream.

***

> **Explore related Unitlab capabilities:** [Unitlab’s multimodal annotation platform](https://unitlab.ai/en/multimodal-annotation) · [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# Multiview Workbench

Operate Current file, Multiple files, and Grouped Workbench modes.

Multiview is a navigation system as much as a layout. Users must know whether a control changes a viewer, a group member, or the next project work item.

## See the workspace model

![Unified multimodal workspace with modality-native tiles](/files/JQmJNJVVmr06fbCAGqin)

*One active native editor creates annotations while the other panels preserve context.*

![Custom grid, list, and resized multimodal layouts](/files/Z7AMj4XHa9FhoCcw3Vxi)

*Choose and resize the layout around the review task; do not squeeze the data into a generic grid.*

### Before you make the change

* Confirm the group membership and slot order.
* Know which panel is authoritative for each annotation type.
* Define synchronization, playback, and navigation expectations.

![Multiple video panels inside Unitlab Workbench](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FuWV2FdAcECE2UrevUndn%2Fmultiview-video-workbench.png?alt=media\&token=77ad56ff-14dd-430b-9d53-cb23ba63956d)

*Use visual emphasis and a stable slot order so the team can identify the active source and the contextual panels immediately.*

### Understand the product behavior

Multiview is the default project annotation experience for all six data families. It is a persistent workspace containing the project header, work-item navigation, mode/layout selector, resizable panel grid, one active editor, and one or more passive inspection panels.

### Two modes

| Mode               | What the panels show                                       | Editing model                                                       | Default layout |
| ------------------ | ---------------------------------------------------------- | ------------------------------------------------------------------- | -------------- |
| **Current file**   | Multiple views of the same datasource                      | One active editor; sibling panels mirror its changes in real time   | 1×1            |
| **Multiple files** | Neighboring work items from the current queue/filter scope | One selected panel is editable; other items remain passive previews | 1×3            |

Layouts range from 1×1 to 4×4. Users can resize panel boundaries, fullscreen a panel, and switch modes from the Workbench header. Mode, layout, and panel sizes persist for the user.

### Active and passive panels

Only the active panel can mutate annotations. It receives the full native editor: tools, hotkeys, selection, object/event/entity editing, player or page controls, comments, classes, properties, history, and workflow actions.

Passive panels can show media, annotations, labels, pages, frames, waveforms, text, or medical projections. In Current file mode they receive the active panel’s live annotation changes, but they cannot originate edits, change the active selection, save, control playback, change a PDF page, edit text entities, or alter waveform regions.

When the user activates a passive panel:

1. Unitlab visually selects it immediately.
2. Any in-progress save in the outgoing panel is allowed to settle.
3. Unsaved changes are saved or safely snapshotted.
4. The outgoing panel becomes passive and pauses modality-specific playback.
5. The incoming native editor loads and restores its panel state.
6. The route updates only after the editor is ready.
7. If hydration fails, Unitlab restores the previous active panel instead of blanking the workspace.

### Current file flow

1. The user opens a work item from a dataset, Task Queue, Batch Queue detail, filtered grid, or direct link.
2. Unitlab creates multiple panel sessions for the same datasource without duplicating data or history.
3. One panel is active; the rest are read-only siblings.
4. Active edits are broadcast to siblings in real time.
5. The user can activate another panel to work from a different view, frame, page, or zoom state.
6. Saving refreshes the shared history and updates every sibling to the persisted result.

Current-file examples:

* **Image:** compare the same image at different zoom or inspection states.
* **Video:** inspect different frames of the same video while sharing annotations and downloaded frames.
* **Audio:** inspect different time regions while one waveform editor remains authoritative.
* **Text:** compare different windows of the same text while entity/relation changes remain synchronized.
* **Document:** compare different PDF pages without confusing page changes with work-item navigation.
* **Medical:** assign Axial, Sagittal, Coronal, and 3D views to separate slots with synchronized annotation state.

### Multiple files flow

1. The user selects **Multiple files** and a layout.
2. Unitlab fills the grid with the active item and neighboring items from the current visible set.
3. The visible set preserves queue, upload-session, archive, search, status, class, and assignment filters.
4. The user activates any ready panel.
5. If the item belongs to another data family, the correct editor loads within the same Workbench shell.
6. Previous/next navigation continues through the filtered work-item sequence.
7. Empty tail slots and inaccessible items appear as panel-level states rather than replacing the whole page with an error.

This mode supports mixed review—for example an image, video, PDF, and audio item in one 2×2 layout—while guaranteeing that only the selected panel is editable.

### Navigation hierarchy

Unitlab maintains three distinct navigation levels:

* **Top-center previous/next:** changes the project work item and can cross data families or move between loose items and Data Groups.
* **Panel activation:** changes which visible panel is editable without leaving the Workbench.
* **Within-item media navigation:** changes a video frame, audio time region, text window, medical slice/view, or PDF page inside the current work item.

Keeping these levels separate is essential for predictable UX. A PDF page change must never advance to another datasource, and a medical slice change must never appear as another project item.

### Grouped Workbench

A Data Group opens as one grouped Workbench whose layout is fixed by the Auto-Grouping builder. The header shows the group layout rather than the general mode/layout selector. Activating a tile never tears down the group route.

The project’s unified previous/next sequence treats each group as one work unit and excludes its member tiles from the loose-item sequence. Grouped saves remain attributable to the group.

### Work through a grouped item

{% stepper %}
{% step %}

#### 1. Confirm the group

Read group identity, member slots, project state, and workflow stage.
{% endstep %}

{% step %}

#### 2. Choose the mode

Use Current file for focused work or Multiple files for simultaneous context.
{% endstep %}

{% step %}

#### 3. Set the active panel

Click or select the source on which new annotation should be created.
{% endstep %}

{% step %}

#### 4. Use passive panels as evidence

Navigate or synchronize them according to layout rules without annotating the wrong source.
{% endstep %}

{% step %}

#### 5. Move within the group

Use viewer or member controls before using project-level previous/next navigation.
{% endstep %}

{% step %}

#### 6. Save one task state

Complete required values and route the grouped work item through its current workflow action.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision       | Production guidance                                                                         |
| -------------- | ------------------------------------------------------------------------------------------- |
| Current file   | Best when one source must receive focused annotation while other members remain selectable. |
| Multiple files | Best when simultaneous comparison or synchronized context is required.                      |
| Active panel   | Must be obvious, stable, and aligned to the ontology tools being used.                      |
| Navigation     | Train users on viewer, group-member, and project-item scopes separately.                    |

### Continue the operating flow

* Test the layout on complete and incomplete groups.
* Calibrate reviewers on cross-panel consistency.
* Validate group context in the release output.

***

> **Related Unitlab capability guides:** [multimodal data annotation](https://unitlab.ai/en/multimodal-annotation) · [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# Multi-camera video

Annotate synchronized camera views of the same event as one governed work unit.

Multi-camera annotation keeps several views of one event together while preserving the perspective of each camera. Use a Data Group and a deliberate layout so annotators can maintain object identity, compare occlusion, and review geometry without opening unrelated tasks.

## See multi-camera annotation

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/multimodal-1.mp4>" %}

![Four synchronized camera views with perspective-aligned cuboids](/files/j3bcITcPmAT1cxWXFXSv)

*Each tile represents its own camera perspective; the Data Group represents the shared event.*

## Design the group before annotation

Define:

* the file or camera naming rule;
* the grouping key that identifies one shared event;
* one tile per expected camera or sensor;
* the layout and tile order;
* whether timestamps are synchronized exactly or within a documented tolerance;
* the object-identity policy across views;
* the ontology geometry and properties required per view;
* the review route for missing, late, corrupted, or unsynchronized feeds.

Use [Data Groups and layouts](/documentation/data/data-groups-and-layouts) to create or audit the groups before attaching them to a project.

## Annotate one event

{% stepper %}
{% step %}

#### 1. Open the grouped item

Confirm the group identity, expected camera tiles, timestamps, and current workflow stage. Missing or duplicate views should be resolved before labeling.
{% endstep %}

{% step %}

#### 2. Choose the active camera

Activate the view with the clearest seed geometry. Passive views remain visible for identity and occlusion context.
{% endstep %}

{% step %}

#### 3. Create the annotation

Select the ontology class and draw the required box, mask, polygon, keypoints, skeleton, or cuboid in the active view.
{% endstep %}

{% step %}

#### 4. Compare sibling views

Inspect the same object in the other cameras. Preserve shared identity according to the project rule while keeping perspective-specific geometry distinct.
{% endstep %}

{% step %}

#### 5. Review time and occlusion

Scrub synchronized video where available. Check entrances, exits, occlusion, reappearance, camera delay, track continuity, and dynamic properties.
{% endstep %}

{% step %}

#### 6. Save and route the group

Resolve required values and issues, save the active annotation state, and submit the grouped work item through the configured review stage.
{% endstep %}
{% endstepper %}

## Quality controls

| Review focus    | Production guidance                                                                           |
| --------------- | --------------------------------------------------------------------------------------------- |
| Synchronization | Define and test timestamp tolerance before production.                                        |
| Identity        | One real object should keep the intended cross-view identity.                                 |
| Geometry        | Do not copy perspective-specific coordinates directly between cameras.                        |
| Missing views   | Surface absent or corrupted feeds as an explicit group state.                                 |
| Delivery        | Preserve group membership, tile role, camera identity, time basis, and annotation provenance. |

{% hint style="warning" %}
A multi-camera layout provides context; it does not automatically calibrate cameras or convert one view’s geometry into another view’s coordinates.
{% endhint %}

## Next steps

* Configure the general [Multiview Workbench](/documentation/multimodal-annotations/multiview-workbench).
* Review [Video Annotation](/documentation/annotations/video-annotation) for tracking and timeline behavior.
* Use [Multimodal overview](/documentation/multimodal-annotations/multimodal-overview) for cross-modal groups.


# Media and document layouts

Combine image, video, audio, text, PDF, and structured evidence in one custom Data Group layout.

A media-and-document task should preserve one business or research case even when its evidence uses different file types. Unitlab Data Groups bring native viewers into one custom layout so an annotator can create a label in the active tile while consulting the remaining evidence.

![Video, PDF, and audio evidence in one multimodal workspace](/files/JQmJNJVVmr06fbCAGqin)

*The group—not an individual file—defines the reviewable case.*

## Choose the evidence model

Use this pattern for examples such as:

* manufacturing video + inspection report + vibration audio;
* image + OCR text + supporting PDF;
* customer recording + transcript + case document;
* field image + sensor export + technician note;
* clinical image + report or study documentation.

Keep each source as its native file whenever possible. Flattening a PDF, waveform, or video into screenshots removes page, time, and content behavior that reviewers need.

## Build the layout

{% stepper %}
{% step %}

#### 1. Define one case

Choose the stable grouping key shared by every file in the work unit—for example inspection ID, encounter ID, capture session, or case number.
{% endstep %}

{% step %}

#### 2. Create or auto-generate the Data Group

Map filename patterns or metadata to tile roles. Add exclusions for files that must not join a group and preview the matches before creation.
{% endstep %}

{% step %}

#### 3. Design the presentation

Choose grid, list, or custom layout; set tile order; enlarge the primary evidence; and keep reference panels large enough to read without distortion.
{% endstep %}

{% step %}

#### 4. Attach the group to a project

Use an ontology whose classes, properties, Item Properties, and relations describe the decisions that span the evidence.
{% endstep %}

{% step %}

#### 5. Annotate from the active tile

Activate the correct native editor, create the label, and use passive tiles for context. Switch active tiles intentionally when another file requires its own annotation.
{% endstep %}

{% step %}

#### 6. Review and deliver the case

Check file membership, tile roles, cross-modal relations, required values, and workflow outcome before publishing the group through a dataset version or release.
{% endstep %}
{% endstepper %}

![Grid, list, and custom multimodal layouts](/files/Z7AMj4XHa9FhoCcw3Vxi)

*Custom layouts control how a repeated case is presented to every annotator and reviewer.*

## Auto-Grouping

![Related MP4, CSV, PDF, and MP3 files auto-grouped into one case](/files/PMDdjgvDNjBkvSihEseF)

Use Auto-Grouping when filenames or metadata consistently encode case membership. Configure the source folder, grouping keys, name pattern, tile rules, exclusions, and layout. Review unmatched and duplicate files before creating groups; an incorrect grouping rule creates a data-quality problem upstream of annotation.

## Cross-modal relations

![A package defect linked to thermal, audio, and document evidence](/files/sH8vZbEefsHes7gwM1p2)

Relations can connect labels across the ontology when the downstream dataset needs explicit evidence structure. Define direction and meaning before production—for example finding-supported-by-report or defect-correlates-with-audio-event. Do not assume proximity in a layout is an implicit relation.

## Production checks

* Every group has the expected files and no unintended duplicates.
* Tile roles and order are stable.
* The active editor is visually unambiguous.
* Page, frame, time, and source-file identity remain separate.
* Reviewers can reconstruct why a label was created from the visible evidence.
* Dataset and release outputs preserve group membership and relation semantics.


# Medical study and clinical document

Review synchronized medical imaging and clinical documentation as one governed case.

Use a medical Data Group when imaging, synchronized projections, sequences, and supporting clinical documentation must stay together. The active medical editor creates geometry and clinical values; document panels preserve report, page, and contextual evidence.

## See the medical workflow

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/auto-labeling-4.mp4>" %}

![Synchronized axial, coronal, sagittal, and 3D DICOM views](/files/mwB1JcWbVojXdn4tU3V3)

*One study remains available across complementary views while the active editor controls annotation.*

## Design the case

Include only the sources allowed by the project’s privacy and access policy. Define:

* the study and series identity;
* which DICOM views or sequences belong together;
* which report or clinical document provides supporting context;
* whether the document is evidence, ground truth, or a source for separate labels;
* the clinical ontology and required properties;
* the specialist review and escalation path;
* redaction or de-identification requirements.

## Work through the group

{% stepper %}
{% step %}

#### 1. Confirm case identity

Verify the study, series, document, and project-safe identifiers before viewing or annotating protected data.
{% endstep %}

{% step %}

#### 2. Inspect synchronized imaging

Navigate axial, coronal, sagittal, 3D, or sequence views to understand the finding in full context.
{% endstep %}

{% step %}

#### 3. Consult the clinical document

Open the relevant report page or text window. Keep a clear distinction between a visible finding and an interpretation stated in the document.
{% endstep %}

{% step %}

#### 4. Create the clinical annotation

Activate the medical view, choose the ontology class, create geometry, and complete visible-morphology properties.
{% endstep %}

{% step %}

#### 5. Add governed context

Use Item Properties or relations for study-level and cross-source information. Do not copy free text into structured values when the ontology requires a controlled choice.
{% endstep %}

{% step %}

#### 6. Specialist review

Inspect cross-view extent, sequence continuity, document consistency, uncertainty, validation, and provenance before routing the group to approval.
{% endstep %}
{% endstepper %}

## Clinical quality boundary

{% hint style="warning" %}
Supporting documentation can inform review, but the project must define whether a label represents image evidence, a clinician’s interpretation, or a combined adjudicated result.
{% endhint %}

Preserve study and series identity, source document and page, geometry, clinical properties, relations, reviewer provenance, ontology version, and group membership in downstream dataset versions and releases.

## Next steps

* Read [Medical Annotation](/documentation/annotations/medical-annotation) for DICOM and synchronized-view tools.
* Read [Document & PDF Annotation](/documentation/annotations/document-and-pdf-annotation) for native document behavior.
* Use [Multimodal overview](/documentation/multimodal-annotations/multimodal-overview) for group and active-panel rules.


# Ontologies overview

Navigate workspace and project ontologies, understand Live state, and inspect the complete semantic contract.

An ontology is Unitlab's versioned semantic contract for annotation. It defines classes, geometry or annotation type, instance properties, relations, Item Properties, required values, conditional questions, validation, operator guidance, and the schema that reaches releases and integrations.

### Start from the workspace ontology catalog

![Workspace Ontologies list with Live status and schema counts](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F5SPPXUzaE7ep9xPCoVzq%2Fontologies-list-live.png?alt=media\&token=8b258dd3-96a7-487e-b2d8-8345a4bb6d5c)

*The workspace catalog makes ontology identity, Live state, class/property/relation counts, update time, and row actions visible before an editor is opened.*

Open **Ontologies** from the workspace sidebar. Use the catalog to:

* search by ontology name;
* filter by lifecycle status;
* compare class, property, and relation counts;
* identify the current **Live** schemas;
* open row actions for an existing ontology;
* choose **New Ontology** to start a new workspace schema.

Treat names as human labels and IDs plus versions as the durable contract. Two ontologies can have the same display name while representing different resources or histories.

### Understand workspace, Global, and Project Ontologies

#### Project Ontologies

![Project Ontologies page with a Live ontology](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FjOjE1y7iPqXilO3mp6Qu%2Fproject-ontologies-current.png?alt=media\&token=1085aad6-b926-4e13-827c-dcecf81bd01d)

*A project can maintain several ontology copies, but only one is Live for active annotation at a time.*

Open a project and choose **Ontologies** to create, switch, and maintain the schemas owned by that project. Use the row action to make the intended copy Live after testing. Replacing the Live selection does not delete the previous ontology.

#### Global Ontologies

![Global Ontologies catalog opened from a project](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FKIDCx8TOBuPPZZI2OEeI%2Fproject-global-ontologies.png?alt=media\&token=0263d87e-19ce-48cc-86da-90d7363ccef4)

*Global Ontologies is the reusable workspace catalog. Importing a catalog ontology creates an independent project copy.*

From Project Ontologies, choose **Global Ontologies** to browse reusable schemas. Import the approved schema into the project, then review and test the resulting project copy. Import is not a permanent live-sync relationship: later global changes do not silently rewrite the project's contract.

Some teams call reusable catalog schemas public or shared ontologies. The current product surface uses **Global Ontologies**. Project Ontologies are the project-owned working copies used for project annotation.

### Read the ontology editor

![Live ontology editor with entity tree, schema canvas, and summary inspector](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FWRnC5HAgbJZLgLj2P6lM%2Fontology-live-editor.png?alt=media\&token=d9b23baa-2b46-4515-869b-cd51677c7a24)

*The editor makes lifecycle state, schema hierarchy, class configuration, Summary/JSON inspection, version history, logic map, and export available in one workspace.*

The editor is organized into three working areas:

1. **Entity tree** — create or select classes, Item Properties, and relations; search the schema; move between concepts.
2. **Schema canvas** — inspect and edit class properties, options, nested conditions, relation definitions, and ordering.
3. **Inspector** — review class or property identity, description, color, hotkey, geometry, Required state, thumbnails, counts, Summary or JSON, **View logic map**, and export.

The header shows the ontology name, **Version history**, and current lifecycle action or state such as **Publish draft** or **Live**.

### The ontology contract

| Structure       | Meaning in annotation output                                              | Typical example                                                     |
| --------------- | ------------------------------------------------------------------------- | ------------------------------------------------------------------- |
| Class           | The concept being labeled.                                                | Person, Vehicle, Defect, Organization.                              |
| Annotation type | How the concept is represented in its native editor.                      | Bounding box, polygon, mask, entity, event, keypoint, line, cuboid. |
| Property        | Typed information about one annotation instance.                          | Severity, material, helmet status.                                  |
| Relation        | A named connection between annotation instances.                          | Person operates Machine; Entity refers to Entity.                   |
| Item Property   | A label for the complete asset, sequence, document, study, or Data Group. | Weather, procedure phase, overall quality.                          |
| Condition       | A rule that reveals a child property only for a relevant parent answer.   | Ask glass color only when container material is Glass.              |
| Validation      | A type-aware constraint and visible invalid state.                        | Numeric range, text length, regex, choice count.                    |

### Lifecycle at a glance

| State or view   | What it means                                                                                   |
| --------------- | ----------------------------------------------------------------------------------------------- |
| Draft           | Editable schema work that must be reviewed and tested before production use.                    |
| Live            | The active contract selected for production annotation in its scope.                            |
| Version history | Published versions plus change events, publisher, date, snapshot, and restore-as-draft actions. |
| Snapshot        | Read-only inspection of a published version.                                                    |
| Logic Map       | Read-only visual graph of classes, properties, options, and conditions.                         |

### Choose the next guide

| Goal                                                    | Continue to                                 |
| ------------------------------------------------------- | ------------------------------------------- |
| Create classes and geometry                             | Classes and annotation types.               |
| Add instance detail or whole-item labels                | Properties, relations, and Item Properties. |
| Add Required rules, validation, or conditional branches | Validation and conditional logic.           |
| Inspect Draft, Live, versions, snapshots, or Logic Map  | Draft, Live, and version history.           |
| Confirm operator behavior and project wiring            | Test an ontology in Workbench.              |

{% hint style="warning" %}
Instructions explain how to apply policy; the ontology constrains the structured output. Publish them as one coordinated change whenever meaning, required values, workflow routing, or downstream schema is affected.
{% endhint %}

***

> **Explore related Unitlab capabilities:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# Classes and annotation types

Create reusable concepts and choose geometry from the downstream task backward.

A class names the concept; the annotation type defines how that concept is expressed. Choose the representation that the downstream model, analysis, or quality decision actually requires.

### Before you make the change

* Write the decision the label must support.
* Collect positive, negative, borderline, and invalid examples.
* Identify every modality and editor in which the class will appear.

![Ontology builder class configuration](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FH7BipjdCSKnZXmzpu8s4%2Fontology-builder.png?alt=media\&token=e88eadbf-aab2-41bc-a67b-46eb2cc58d82)

*Class names, descriptions, colors, hotkeys, annotation types, and required state should form one understandable operator experience.*

### Understand the product behavior

**Workbench ontology entities:**

* item properties: single-choice, multi-choice, text;
* spatial classes: bounding box, cuboid, polygon, mask, keypoint, line, skeleton;
* text entity;
* audio event;
* relation.

**SDK ontology structures:**

* shapes: bounding box, polygon, point, skeleton, polyline, bitmask, cuboid, time range, text;
* attributes: radio, checklist, text, numeric;
* whole-item/global classification;
* required attributes;
* dynamic attributes for values that can change across video keyframes;
* conditional attributes nested beneath selected options.

The UI and SDK use slightly different vocabulary for closely related concepts—for example Mask/bitmask, Event/time range, Entity/text, single-choice/radio, and multi-choice/checklist.

### Class configuration

A class can include:

* name;
* description;
* color;
* numeric hotkey;
* geometry;
* attributes/properties;
* single-select, multi-select, and text properties;
* relations.

Geometry is selected when the class is created and is read-only afterward. This prevents an existing class from silently changing from one annotation representation to another. The class summary shows the stable class ID, description, color, hotkey, geometry, required state, thumbnails, statistics, maximum property depth, logic map, and JSON export.

### Define the top-level schema

{% stepper %}
{% step %}

#### 1. Name the concept

Use a stable domain term with a precise description and examples.
{% endstep %}

{% step %}

#### 2. Choose annotation types

Select classification, bounding box, polygon, mask, line, keypoint, relation, segment, span, or other current representation as required.
{% endstep %}

{% step %}

#### 3. Configure operator cues

Set color, hotkey, ordering, and required state to reduce ambiguity in Workbench.
{% endstep %}

{% step %}

#### 4. Add only meaningful structure

Avoid classes or geometry that do not change downstream interpretation.
{% endstep %}

{% step %}

#### 5. Test in each native editor

Confirm the class behaves correctly for every attached modality and layout.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision       | Production guidance                                                                          |
| -------------- | -------------------------------------------------------------------------------------------- |
| Class scope    | Prefer a coherent reusable concept over project-specific wording when the meaning is stable. |
| Geometry       | Use the minimum representation that preserves the required downstream information.           |
| Required state | Make only true acceptance requirements mandatory.                                            |
| Change         | Treat renamed, removed, or retyped concepts as schema changes with migration impact.         |

### Continue the operating flow

* Add typed properties and relations.
* Define Item Properties for work-unit outputs.
* Test the complete ontology before publishing.

***

> **Continue with Unitlab:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# Properties, relations, and Item Properties

Create typed schema content from the Ontology builder or directly from Annotation Workbench.

Properties, relations, and Item Properties all become annotation output, but they attach to different owners. Decide the owner first; then choose type, required state, dynamic behavior, validation, conditions, and operator guidance.

### Choose the correct structure

| Structure       | Attach it to                                              | Use it for                                                    | Example                                   |
| --------------- | --------------------------------------------------------- | ------------------------------------------------------------- | ----------------------------------------- |
| Property        | One annotation object or span.                            | Structured detail about that instance.                        | Person → Helmet status.                   |
| Relation        | Two or more annotation instances.                         | A named semantic connection with defined targets.             | Person → operates → Forklift.             |
| Item Property   | The full asset, sequence, document, study, or Data Group. | Whole-item annotation output independent of one object.       | Scene → Weather; Study → Overall quality. |
| Source metadata | The source resource outside annotation output.            | Acquisition or business context used to find or prepare data. | Camera ID, source path, acquisition date. |

### Path 1: create schema content in the Ontology builder

![Ontology entity-type menu with Item Properties, object detection, text, audio, and relation types](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FJwNKpXsKwqxWM4RkOYmM%2Fontology-entity-types.png?alt=media\&token=ffe784be-db31-4844-a48d-86f171450971)

*Select the entity type before naming a new class, Item Property, event, entity, or relation. The available groups reflect the supported annotation structures.*

1. Open the workspace ontology or the project's ontology.
2. Open **Select type** in the entity tree.
3. Choose an Item Property type, an object-detection geometry, a text **Entity**, an audio **Event**, or **Relation**.
4. Enter the new entity name and choose the add action.
5. Select the resulting class or property in the tree.
6. For a class, use **Add Property** in the schema canvas; choose Single choice, Multiple choice, or Text and the scalar type needed by the contract.
7. Add description, options, Required state, Dynamic state where supported, default value, validation, help text, and conditional dependency.
8. Review Summary, JSON, statistics, and Logic Map before testing in Workbench.

Object-detection options include Bounding box, Cuboid, Polygon, Mask, Keypoint, Line, and Skeleton. Geometry is part of the class contract; do not change representation simply because another tool is faster.

### Build properties and conditional children

![Ontology with three levels of conditional single-choice properties](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F5WGTh7i4QrEJfLGt2ZjV%2Fontology-nested-properties.png?alt=media\&token=03c67080-6bdb-4616-a711-8e3b8a50b5e8)

*The builder shows the parent option, WHEN condition, child property, and deeper child branch directly in the schema canvas.*

Use controlled choice properties for stable enumerations and Text for string, number, boolean, date, datetime, URL, email, or object-ID values supported by the current builder. A child property belongs beneath the parent option that makes it relevant. The illustrated schema asks a second question when **Bottle** is selected, then asks for a glass color only when **Glass** is selected.

Deep conditional nesting is supported, but every branch adds operator and downstream complexity. Keep only questions that change meaning or acceptance.

### Path 2: extend the ontology from Annotation Workbench

![Image Annotation Workbench with Classes panel and item-level controls](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FXDWBnDB8WFTU6efWwzzN%2Fworkbench-classes-item-properties.png?alt=media\&token=6689950e-874b-493e-bba4-0bc8729fd62f)

*Workbench exposes the active project's classes beside the source item so ontology changes can be tested in the same operator context in which they will be used.*

Workbench provides two levels of schema authoring:

#### Quick class and class management

* Use **Quick Create Class** when a compatible class is missing for the active editor.
* Open the class-management entry and choose **Manage Classes** for the lightweight manager.
* Select the annotation type, search for an existing class or create one, and choose **Add**.
* Configure name, color, description, Required state, and hotkey. Geometry remains fixed after creation.
* Use the same manager to create lightweight Item Properties and relations.
* Choose **Manage Ontologies** when the complete builder, JSON, logic map, versions, or advanced schema controls are required; Unitlab opens the full builder in a new tab.

#### Add a property from an annotation object

1. Create or select an annotation object in Workbench.
2. In the object inspector, choose **Add property**.
3. Enter name and description.
4. Set **Required** and, for supported video or medical work, **Dynamic**.
5. Choose the property type and scalar type, then configure options where applicable.
6. Choose **Add property**, set the value on the selected object, and inspect the result in the timeline when Dynamic is enabled.

#### Add a relation

Select an object, choose **Add relation**, name and describe the relation, define the allowed target classes, save it, then create the relationship between the intended source and target annotations. Test direction and target restrictions; a relation label without direction is ambiguous downstream.

#### Add an Item Property

Open **Item properties** at the top of the inspector, choose **Add item properties**, enter name and description, configure Required and Dynamic, choose the type and options, then add and set the value. Item Properties are independent of one annotation object.

### Static and Dynamic properties

A static property applies one value to the object or item. In video and medical annotation, a Dynamic property stores frame-aware states while the owning identity remains the same.

* **Dynamic class-property example:** one Person track keeps the same identity while Helmet status changes from Not visible to Present to Absent.
* **Dynamic Item Property example:** the sequence-level Weather changes from Clear to Rain, or Procedure phase changes from Preparation to Inspection.

In Workbench, the first value creates a keyframe or range. Changing the value later creates another state. Expand the parent object or Item Property row in the timeline to inspect child property ranges and transition points; adjust temporal boundaries only after reviewing the source frames.

{% hint style="warning" %}
For an existing property, the Dynamic switch can be locked. When the product requires recreation, assess the schema and historical-label impact before deleting and rebuilding the property.
{% endhint %}

### Before publishing

* confirm every field owner, type, ID, default, help text, Required state, Dynamic state, validation, and condition;
* traverse every nested branch in Workbench;
* create representative relations and verify source, target, and direction;
* inspect static and dynamic Item Properties in the correct editor;
* compare JSON and sample release output to the downstream schema.

***

> **Explore related Unitlab capabilities:** [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# Validation and conditional logic

Configure Required state, type-aware validation, help text, and nested conditional branches.

Validation defines whether a supplied value satisfies the data contract. Conditional logic controls when a question is relevant. **Required** is separate: a value may be required and also subject to a type-aware constraint.

### Configure a selected property

![Selected ontology property with Required, Dynamic, default, validation, help text, and dependency controls](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FlXf333HEHEVKy8p3tnXJ%2Fontology-property-validation.png?alt=media\&token=2a5bb77e-666a-47b9-9d72-e41e0cbbe92d)

*Select a property in the schema canvas to configure its identity and additional settings in the right inspector.*

1. Open the ontology and select the class.
2. Select the property card in the schema canvas.
3. In **Selected property**, confirm the property name, description, and stable ID.
4. Enable **Required** only when an absent value makes the annotation incomplete.
5. Enable **Dynamic** at property creation for supported video or medical time-varying states.
6. Set a default only when it is semantically true for new annotations; a convenience default can silently bias labels.
7. Open **Validation** and choose the rule appropriate to the scalar type.
8. Add **Help text** that explains the accepted form or boundary to the annotator.
9. Review **Depends on** when the property is conditional.

### Type-aware validation

| Property value           | Compact validation supported by the current contract       |
| ------------------------ | ---------------------------------------------------------- |
| String                   | Length or regular expression.                              |
| Number                   | Numeric range or regular expression.                       |
| Date or datetime         | Allowed date/time range.                                   |
| URL, email, or object ID | Regular expression appropriate to the identifier contract. |
| Multiple choice          | Minimum or maximum selection count.                        |

Single-choice properties already constrain the value to one configured option. Keep Required state, controlled options, and any available type-specific validation conceptually separate.

### Create conditional questions

![Conditional ontology tree with Bottle, Glass, and color branches](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F5WGTh7i4QrEJfLGt2ZjV%2Fontology-nested-properties.png?alt=media\&token=03c67080-6bdb-4616-a711-8e3b8a50b5e8)

*The schema canvas keeps each WHEN condition next to the option that reveals its child property, including deeper nested branches.*

The illustrated logic can be read as:

1. Ask the parent property.
2. If **Bottle** is selected, reveal the child material property.
3. If **Glass** is selected, reveal the nested glass-color property.
4. Otherwise, do not ask the irrelevant child question.

To add the branch, create the parent choice property and its options, add the child property beneath the relevant option, then continue nesting only where another answer changes what must be collected. Use **View Logic Map** before publication to confirm all branches are reachable and named clearly.

### What happens when a value is invalid

Unitlab keeps invalid values visible rather than silently replacing them. The operator receives a message, the value remains in annotation history, and the item can surface through the project's **Invalid** path or QA cohort. Correct the value when evidence supports a valid answer; otherwise use the configured issue, escalation, invalid-data, or rework route.

{% hint style="warning" %}
Do not invent a value merely to clear validation. Add Unknown or Not applicable only when those are legitimate domain values and define how downstream systems interpret them.
{% endhint %}

### Acceptance test for validation and logic

* enter values at, below, and above every boundary;
* test valid and invalid string formats;
* select the minimum and maximum allowed choice counts;
* leave every Required field empty once and confirm the visible invalid state;
* traverse every parent option and verify only relevant children appear;
* test nested branches with keyboard and pointer input;
* inspect the JSON and a sample release so IDs, types, missing values, and nested answers match the downstream contract.

***

> **Related Unitlab capability guides:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# Draft, Live, and version history

Publish, inspect, restore, and compare ontology versions without silently mutating production history.

Ontology lifecycle controls which schema is editable, which schema is active, and which historical version produced existing labels. Treat publication as a data-contract change, not a cosmetic save.

### Draft and Live

| State | Editability                | Production meaning                                                              |
| ----- | -------------------------- | ------------------------------------------------------------------------------- |
| Draft | Editable.                  | Proposed schema changes under review; not yet the approved production contract. |
| Live  | Active production version. | The schema selected for current annotation in its workspace or project scope.   |

In the builder header, **Publish draft** indicates unpublished changes. **Live** identifies the active published version. Before publishing, inspect the visual tree, Summary, JSON, IDs, required values, validation, nested conditions, relation targets, Logic Map, Workbench behavior, and sample output.

### Inspect version history

![Ontology history dialog with Live version, change timeline, snapshot, and restore controls](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FtNGEgNCHnksn0hIL3POs%2Fontology-version-history.png?alt=media\&token=b88769b8-62ae-47f6-a684-8e9f1b8ac95d)

*Version history records published versions, lifecycle state, publisher and time, semantic change events, and the actions View snapshot and Restore as draft.*

Choose **Version history** in the ontology header. Search or filter the history when several versions exist. Expand a version to review the semantic event timeline: which class, property, relation, or option changed, when it changed, and who made the change.

**Restore as draft** does not rewrite the historical version. It creates editable Draft work based on that published snapshot so the restored schema can be reviewed, tested, and published deliberately.

### Open a read-only snapshot

![Read-only ontology snapshot](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F2QABJkTatxRAPLozxoIf%2Fontology-snapshot.png?alt=media\&token=457efe97-180a-4e4e-a7d0-86f39006654d)

*Snapshot · Read only disables schema mutation while preserving search, hierarchy inspection, Summary/JSON, Logic Map, and export for comparison.*

Choose **View snapshot** on the version you need to inspect. Use the read-only state to compare class IDs, geometry, property types, options, conditions, validation, relations, and JSON without risking edits. Record the version used by the labels or release under investigation.

### Inspect the Logic Map

![Ontology Logic Map dialog with zoom and fit controls](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FY5b2tJZqEf2RhriKaG5b%2Fontology-logic-map.png?alt=media\&token=e3efd5f5-bb22-4b71-8033-a4f863f31bd1)

*Logic Map is a read-only graph of ontology structure. Use zoom, Fit view, and branch inspection to review the semantic path before publication.*

Choose **View logic map** from the ontology inspector. Use it to find unreachable branches, duplicated concepts, deep conditional paths, ambiguous names, or relations whose targets are difficult to understand from the canvas alone. The map is an inspection tool; return to the Draft to make changes.

### Publish a controlled version

{% stepper %}
{% step %}

#### 1. Inventory dependencies

Identify projects, Live selection, active tasks, Instructions, workflow routes, model mappings, API clients, releases, and downstream schemas affected by the change.
{% endstep %}

{% step %}

#### 2. Review the Draft

Inspect hierarchy, Summary, JSON, IDs, geometry, Required and Dynamic states, validation, conditions, relations, defaults, help text, and Logic Map.
{% endstep %}

{% step %}

#### 3. Test in Workbench

Exercise every native editor, custom layout, role, conditional branch, invalid state, dynamic timeline behavior, and representative relation.
{% endstep %}

{% step %}

#### 4. Approve the contract

Record owner, reason, compatibility decision, in-flight work treatment, and downstream acceptance evidence.
{% endstep %}

{% step %}

#### 5. Publish

Publish the Draft and confirm the intended version shows Live in the relevant catalog and project.
{% endstep %}

{% step %}

#### 6. Monitor

Review invalid-state cohorts, rejections, Issues, automation errors, and the next release for schema drift.
{% endstep %}
{% endstepper %}

{% hint style="warning" %}
Historical labels must remain attributable to the ontology version under which they were produced. Never use a silent edit or a renamed display label as a substitute for controlled version history.
{% endhint %}

***

> **Related Unitlab capability guides:** [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# Test an ontology in Workbench

Attach or import the project ontology, make the intended copy Live, and exercise it in native annotation editors.

A schema can be structurally valid and still fail the operator. Workbench testing proves that the intended project ontology is actually available, Live, understandable, editable by the correct roles, and compatible with every source modality and release consumer.

### Add the ontology to the project

#### Create or switch a Project Ontology

![Project Ontologies list with a Live ontology](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FjOjE1y7iPqXilO3mp6Qu%2Fproject-ontologies-current.png?alt=media\&token=1085aad6-b926-4e13-827c-dcecf81bd01d)

*Project Ontologies owns the independent ontology copies used by this project and identifies the single Live selection.*

Open the project and choose **Ontologies**. Use **New Ontology** to create a project-owned schema, or open the row action on an existing copy to make the tested ontology Live. A project can retain several copies for controlled evolution, but annotation uses the Live selection.

#### Import from Global Ontologies

![Global Ontologies catalog inside a project](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FKIDCx8TOBuPPZZI2OEeI%2Fproject-global-ontologies.png?alt=media\&token=0263d87e-19ce-48cc-86da-90d7363ccef4)

*Importing a Global Ontology creates an independent project copy that can be tested and governed without silently following later catalog changes.*

Choose **Global Ontologies**, find the approved workspace schema, import it, return to Project Ontologies, inspect the copy, and make it Live only after the acceptance test.

### Confirm the ontology in Workbench

![Annotation Workbench Classes panel](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FXDWBnDB8WFTU6efWwzzN%2Fworkbench-classes-item-properties.png?alt=media\&token=6689950e-874b-493e-bba4-0bc8729fd62f)

*The Classes panel is the operator-facing proof that the project's Live ontology loaded into the native editor. The source canvas, class list, item controls, and workflow action must agree.*

1. Attach a representative dataset version to the project and wait for Project Data to be ready.
2. Open **Project › Data** and enter a representative item in Workbench.
3. Confirm the **Classes** panel contains the expected names, order, colors, geometry icons, required classes, and hotkeys.
4. Create one annotation for every class and supported annotation type.
5. Select an object and set every property; use **Add property** only when the test includes Workbench-side schema authoring.
6. Open **Item properties**, set each static and Dynamic Item Property, and verify the correct item or timeline scope.
7. Create each relation and confirm source, target, direction, and target restrictions.
8. Traverse every conditional branch and trigger every validation error intentionally.
9. Save, route to review, reject or return representative errors, and confirm role-specific visibility and editability.

### Test temporal properties

For video or medical work, use a sequence containing a real state change:

* create a track and set a Dynamic class property at its first authoritative frame;
* move later and change the value;
* expand the object row and verify property ranges and transition points;
* add a Dynamic Item Property, change it later, and confirm it appears as its own timeline row rather than beneath one object;
* adjust a range boundary and verify the intended frames only;
* confirm static properties remain one value across the item.

### Workbench-side schema authoring

Use **Quick Create Class** for a missing compatible class. Use **Manage Classes** to add or configure lightweight classes, Item Properties, and relations. Use **Manage Ontologies** to open the full builder when the change needs advanced types, JSON, Logic Map, version history, or broader review. Changes made during testing are still schema changes and must re-enter the Draft review and publication process.

### Acceptance matrix

| Area                 | Evidence required                                                                                    |
| -------------------- | ---------------------------------------------------------------------------------------------------- |
| Project wiring       | Intended ontology copy is present and the correct one is Live.                                       |
| Classes and geometry | Every class is visible and produces the required native annotation type.                             |
| Properties           | Types, defaults, help text, Required state, Dynamic state, and values behave as designed.            |
| Conditional logic    | Every branch is reachable; irrelevant children stay hidden.                                          |
| Validation           | Invalid input remains visible, understandable, and recoverable.                                      |
| Relations            | Direction and allowed targets are correct in Workbench and output.                                   |
| Item Properties      | Static and dynamic whole-item values use the correct scope.                                          |
| Roles and workflow   | Annotators and reviewers see only valid controls and routes.                                         |
| Downstream contract  | Sample JSON or release output matches field IDs, types, coordinates, relations, and temporal ranges. |

Record each defect with the ontology version, project, source item, role, expected behavior, observed behavior, owner, correction, and retest evidence. Publish only after the acceptance matrix passes.

***

> **Continue with Unitlab:** [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# Workflows overview

Understand stages, routes, ownership, and valid actions.

A Unitlab workflow is the operational state machine for project work. It determines where a task is, who can act, which actions are valid, and where accepted, rejected, failed, or escalated work goes next.

{% hint style="info" %}
**Use this area when:** you are designing production operations, adding review, introducing a model, or changing how work is assigned and completed.
{% endhint %}

### How this area fits into production

```mermaid
flowchart LR
  A["Project"]
  B["Annotate"]
  C["Review"]
  D["Rework"]
  E["Complete"]
  F["Model or specialist"]
  A --> B
  B --> C
  C -->|accept| E
  C -->|reject| D
  D --> C
  B --> F
  F --> C
```

![Unitlab workflow canvas](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FAjCMPQRpjcxTkfYgo8FQ%2Fworkflow-canvas.png?alt=media\&token=941ab2b3-7827-408e-978c-b7cf80252715)

*The canvas makes stage ownership and routing visible; a production workflow should expose normal, rejected, failed, invalid, and escalated paths clearly.*

### What this area controls

### Workflow canvas

A new project receives the default workflow:

```
Project → Annotate → Review → Complete
```

The six stage types currently released in the editor are:

* **Project** — the single entry point;
* **Annotate** — human labeling;
* **Review** — human quality review with approve and reject routes;
* **Model** — automated inference on entry;
* **Archive** — terminal set-aside state;
* **Complete** — terminal successful state and release-readiness destination.

The released stage catalog contains Project, Annotate, Review, Model, Archive, and Complete. A team that needs specialist or expert review adds and names another Review stage, then configures its eligible reviewers and routing rules.

The canvas supports:

* adding stages;
* drawing connections;
* naming the workflow;
* binding the reusable workflow to a project;
* configuring eligible annotators or reviewers per human stage;
* allowing or preventing self-assignment;
* allowing manager override;
* hiding unassigned work where required;
* configuring whether the stage can be skipped;
* selecting a model, thresholds, generic type, queue scope, and class mappings for Model stages;
* accepted and rejected review branches;
* Save and Apply with impact validation;
* zoom, reset, and canvas navigation.

The workspace Workflows page displays reusable workflow cards and a Create Workflow action. Opening a project’s Workflows tab goes directly to the editor for that project’s active workflow.

In the full-page editor, users drag released stages from **Add Stages** onto the canvas, connect or remove edges, select a node to edit its configuration, and use one explicit **Save & Apply** action. The Project entry node embeds project progress. Stage names are editable except Project; duplicate display names are numbered consistently in both the canvas and Task Queue.

Replacing an active project workflow while items are in flight prompts with the number of items that would reset to Project. An unsaved-changes guard protects navigation away from the editor. Workflows can also be duplicated, archived, and restored.

### Start with the right page

| Decision               | Production guidance                                              |
| ---------------------- | ---------------------------------------------------------------- |
| Build a graph          | Configure stages and routes.                                     |
| Define human action    | Set assignment and stage actions.                                |
| Add quality control    | Configure review, rework, and escalation.                        |
| Introduce automation   | Use a Model stage with explicit failure and human-review routes. |
| Change live operations | Run a workflow impact review first.                              |

### Operating boundary

* Rejection is a route back to corrective work, not an unexplained terminal state.
* Invalid annotation state is different from workflow rejection or processing failure.
* Queue visibility and actions are consequences of stage, role, assignment, and task state.

### A production-ready handoff

Every state and transition has a name, owner, valid action, success condition, exception route, and observable queue representation.

***

> **Explore related Unitlab capabilities:** [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# Stages and routes

Build the production path from project intake to completion.

A stage should represent a meaningful responsibility or automation boundary. A route should explain what decision moved the work and who owns it next.

### Before you make the change

* Write the normal path and every exception path in plain language.
* Name the human or service owner for each stage.
* Define entry criteria, valid actions, exit criteria, and observability.

![Workflow canvas with connected stages](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FAjCMPQRpjcxTkfYgo8FQ%2Fworkflow-canvas.png?alt=media\&token=941ab2b3-7827-408e-978c-b7cf80252715)

*Use the canvas to make routes legible; a visually impressive graph is not useful if operators cannot explain why an item follows each edge.*

### Understand the product behavior

A new project receives the default workflow:

```
Project → Annotate → Review → Complete
```

The six stage types currently released in the editor are:

* **Project** — the single entry point;
* **Annotate** — human labeling;
* **Review** — human quality review with approve and reject routes;
* **Model** — automated inference on entry;
* **Archive** — terminal set-aside state;
* **Complete** — terminal successful state and release-readiness destination.

The released stage catalog contains Project, Annotate, Review, Model, Archive, and Complete. A team that needs specialist or expert review adds and names another Review stage, then configures its eligible reviewers and routing rules.

The canvas supports:

* adding stages;
* drawing connections;
* naming the workflow;
* binding the reusable workflow to a project;
* configuring eligible annotators or reviewers per human stage;
* allowing or preventing self-assignment;
* allowing manager override;
* hiding unassigned work where required;
* configuring whether the stage can be skipped;
* selecting a model, thresholds, generic type, queue scope, and class mappings for Model stages;
* accepted and rejected review branches;
* Save and Apply with impact validation;
* zoom, reset, and canvas navigation.

The workspace Workflows page displays reusable workflow cards and a Create Workflow action. Opening a project’s Workflows tab goes directly to the editor for that project’s active workflow.

In the full-page editor, users drag released stages from **Add Stages** onto the canvas, connect or remove edges, select a node to edit its configuration, and use one explicit **Save & Apply** action. The Project entry node embeds project progress. Stage names are editable except Project; duplicate display names are numbered consistently in both the canvas and Task Queue.

Replacing an active project workflow while items are in flight prompts with the number of items that would reset to Project. An unsaved-changes guard protects navigation away from the editor. Workflows can also be duplicated, archived, and restored.

### Rejection is a route, not a failure state

Review has both approve and reject edges. This allows a correction path to be designed before production starts.

A production workflow might be:

```
Project
├── Model → Annotate → Review
│                    ├── Accepted → Complete
│                    └── Rejected → Annotate
└── Difficult/low-confidence → Specialist Review
                               ├── Accepted → Complete
                               └── Rejected → Annotate
```

Here, Specialist Review is another stage of the released **Review** type with a specialist eligibility list, not a different stage type.

### Build the operating state machine

{% stepper %}
{% step %}

#### 1. Place the project entry

Define how newly attached work becomes available.
{% endstep %}

{% step %}

#### 2. Add execution stages

Create human annotation, model, or other current stages according to responsibility.
{% endstep %}

{% step %}

#### 3. Add review

Define what the reviewer can accept, reject, return, or escalate.
{% endstep %}

{% step %}

#### 4. Connect every outcome

Draw explicit routes for success, rework, invalid data, failure, archive, specialist escalation, and completion.
{% endstep %}

{% step %}

#### 5. Test stage actions

Use each role to move representative tasks through every path.
{% endstep %}

{% step %}

#### 6. Save with impact awareness

Record the graph version and treatment of in-flight work.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision       | Production guidance                                                                         |
| -------------- | ------------------------------------------------------------------------------------------- |
| Stage boundary | Create a stage when responsibility, permission, automation, or acceptance criteria changes. |
| Route label    | Use outcome language the operator can interpret.                                            |
| Terminal state | Reserve completion or archive for explicitly accepted outcomes.                             |
| Exception      | Keep failed or escalated work observable and owned.                                         |

### Continue the operating flow

* Configure assignment and queue behavior.
* Exercise review and rework.
* Validate the graph before applying it to production volume.

***

> **Continue with Unitlab:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# Assignment and stage actions

Control who can claim, start, save, submit, reject, or move work.

Workflow actions are permissioned state transitions. Operators should see only the actions valid for their role, stage, assignment, and current task state.

![Task Queue with assignment controls](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FNuZoKkKYJ9eoN6ckUKWT%2Ftask-queue.png?alt=media\&token=698abe91-80e4-461c-9837-9874a5b22451)

*Queue assignment and Workbench stage actions should express the same operating rules so users do not need side-channel coordination.*

Saving and moving through the workflow are separate actions. Saving appends annotation history inside the current stage. A stage action changes the item’s route.

The Workbench header builds available actions from the current item and can show:

* **Send to Review** or **Send to `<stage>`**;
* **Mark as Complete**;
* **Reject** as a danger action;
* **Restart Workflow** for a manager viewing a Complete item;
* item timeline.

Automated-stage items open read-only while the model or automation owns them. After a successful stage action, Unitlab saves dirty work, advances to the next item in the current queue/filter context, and returns to the project Datasets page with a Queue complete message when no work remains.

### Use this in production

* Define whether work is manually assigned, self-claimed, or automatically allocated.
* Separate the ability to annotate, review, move tasks, change priority, and administer the project.
* Make save, submit, reject, skip, invalid, and escalate behavior distinct.
* Test unassigned, assigned-to-self, assigned-to-other, completed, and reopened states for every role.

***

> **Continue with Unitlab:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# Review, rework, and escalation

Route quality decisions with explicit ownership and context.

Review should turn a quality decision into a clear next state. Rejected work needs an actionable reason and destination; specialist escalation needs an owner and return path.

### Before you make the change

* Define acceptance criteria and reviewer authority.
* Separate item-level correction from systemic policy or schema correction.
* Name the specialist or governance owner for escalated cases.

### Understand the product behavior

Review has both approve and reject edges. This allows a correction path to be designed before production starts.

![Unitlab AI workflow canvas connecting project, annotation, and review stages](https://storage.ghost.io/c/48/f4/48f4b614-5c29-430d-9cb3-e0b3f34395f3/content/images/2026/08/clean-workflow-canvas.webp)

A production workflow might be:

```
Project
├── Model → Annotate → Review
│                    ├── Accepted → Complete
│                    └── Rejected → Annotate
└── Difficult/low-confidence → Specialist Review
                               ├── Accepted → Complete
                               └── Rejected → Annotate
```

Here, Specialist Review is another stage of the released **Review** type with a specialist eligibility list, not a different stage type.

### Human review

Review stages require Approve and Reject outcomes. Approve follows the forward edge; Reject returns the item to the configured rework stage and can generate a rework notification. Workbench actions come from the item’s current stage rather than a fixed button set. Review should test the task’s quality policy, not merely confirm that an annotation exists.

### Workbench stage actions

Saving and moving through the workflow are separate actions. Saving appends annotation history inside the current stage. A stage action changes the item’s route.

The Workbench header builds available actions from the current item and can show:

* **Send to Review** or **Send to `<stage>`**;
* **Mark as Complete**;
* **Reject** as a danger action;
* **Restart Workflow** for a manager viewing a Complete item;
* item timeline.

![Unitlab AI video Workbench with frame timeline and Send to Review control](https://storage.ghost.io/c/48/f4/48f4b614-5c29-430d-9cb3-e0b3f34395f3/content/images/2026/08/clean-video-annotation-workers.webp)

Automated-stage items open read-only while the model or automation owns them. After a successful stage action, Unitlab saves dirty work, advances to the next item in the current queue/filter context, and returns to the project Datasets page with a Queue complete message when no work remains.

### Operate a quality decision

{% stepper %}
{% step %}

#### 1. Inspect the full task

Review the relevant frames, pages, panels, timeline, properties, relations, and source context.
{% endstep %}

{% step %}

#### 2. Make the decision

Accept, reject, or escalate against the published Instructions and ontology.
{% endstep %}

{% step %}

#### 3. Explain corrective work

Identify the instance, interval, page, field, or policy rule that needs attention.
{% endstep %}

{% step %}

#### 4. Route to the owner

Move the task to the configured rework or specialist stage.
{% endstep %}

{% step %}

#### 5. Re-review the correction

Confirm the specific defect and any related cohort risk are resolved.
{% endstep %}

{% step %}

#### 6. Escalate systemic gaps

Update Instructions, ontology, workflow, or training through controlled change.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision       | Production guidance                                                              |
| -------------- | -------------------------------------------------------------------------------- |
| Accept         | The work satisfies current policy and required values.                           |
| Reject         | The correction is understood and can return to an owner.                         |
| Escalate       | The decision exceeds the current role or policy and requires a named specialist. |
| Systemic issue | Inspect the affected cohort and control, not only the triggering item.           |

### Continue the operating flow

* Track repeated rejection reasons.
* Update calibration examples for recurring defects.
* Confirm queue state and ownership after every transition.

### Product context

Unitlab helps AI teams curate, annotate, manage, version, and prepare multimodal training data at enterprise scale.

See [Unitlab’s multimodal data annotation platform](https://unitlab.ai/en/data-annotation) for the commercial overview of enterprise annotation and quality workflows.


# Model stages

Integrate model inference as an observable, reviewable workflow responsibility.

A Model stage makes inference part of the same state machine as human work. It must have a versioned contract, mapped outputs, visible failure behavior, and a human-controlled acceptance route.

### Before you make the change

* Define model owner, version, endpoint or runtime, supported modality, input, output, and ontology mapping.
* Prepare a calibration cohort and expected failure cases.
* Define timeout, retry, empty-output, malformed-output, and low-confidence routes.

![Unitlab AI Models catalog](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FISFVz8ybYzJM3OxSHs4v%2Fai-models.png?alt=media\&token=e7c6b0f2-1222-4c1b-8a07-b6d1d55e8a09)

*Treat every public or private model as a versioned production dependency whose mapping and failure behavior are part of workflow design.*

### Understand the product behavior

* **Model → Annotate → Review:** a model proposes labels, an annotator corrects them, and a reviewer verifies the result.
* **Annotate with local assistance → Review:** Magic Touch, prompt detection, Find Similar, or tracking helps the annotator within the task.
* **Model → Review with rejection to annotation:** high-confidence output moves directly to review; rejected work returns to a human correction stage.
* **Annotate → Review → specialist Review:** ambiguous or high-risk cases move to a second Review stage configured for specialists.

### Workflow save and change impact

Workflows are reusable workspace definitions that can be bound to multiple projects. Saving an active graph is therefore not a cosmetic edit. Before applying a change, Unitlab validates reachability, required edges, terminal stages, and the effect on in-flight items.

Important UX rules:

* exactly one Project and one Complete stage;
* Project has no incoming edge and must lead to an entry stage;
* Review has one Approve and one Reject route;
* terminal stages have no outgoing routes;
* every stage is reachable from Project;
* duplicate outgoing action names are rejected;
* a stage holding active items cannot be silently deleted;
* sensitive automated-stage configuration cannot change while occupied without resolving the impact.

When an active workflow change would strand work, Unitlab returns an apply-impact conflict instead of silently moving or losing items.

### Model-stage user experience

Entering a Model stage automatically dispatches the configured model. The item shows Processing while inference runs. Predictions are saved as normal annotation history, then the item follows the Model stage’s default edge to another Model stage, Annotate, Review, or a terminal stage. A failure moves the item to Error so it can be diagnosed rather than disappearing from the workflow.

### Add a model to production routing

{% stepper %}
{% step %}

#### 1. Approve the contract

Document inputs, outputs, classes, geometry, confidence, version, and ownership.
{% endstep %}

{% step %}

#### 2. Map to ontology

Verify every model output has an intentional destination or rejection rule.
{% endstep %}

{% step %}

#### 3. Place the stage

Connect model input, success, human review, failure, and escalation routes.
{% endstep %}

{% step %}

#### 4. Exercise failure modes

Test timeouts, empty output, malformed output, partial success, and unavailable runtime.
{% endstep %}

{% step %}

#### 5. Review corrected output

Measure human correction by source, class, and model version.
{% endstep %}

{% step %}

#### 6. Version change safely

Re-run calibration and workflow impact review for material model or mapping updates.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision                      | Production guidance                                                            |
| ----------------------------- | ------------------------------------------------------------------------------ |
| Pre-label vs autonomous route | Choose according to risk and the required human acceptance gate.               |
| Confidence                    | Use for prioritization or routing only when calibrated on the relevant domain. |
| Retry                         | Inspect remote task state before repeating a mutation.                         |
| Ownership                     | A failed model task must land in a visible queue with a named owner.           |

### Continue the operating flow

* Monitor Batch Queue and task failures.
* Inspect systematic correction cohorts.
* Record model version in release provenance.

***

> **Continue with Unitlab:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# Change a live workflow

Update routing without orphaning or misrouting in-flight work.

A live workflow change can alter assignment, permissions, queue visibility, review paths, automation, and release timing. Treat it as an operational migration.

Workflows are reusable workspace definitions that can be bound to multiple projects. Saving an active graph is therefore not a cosmetic edit. Before applying a change, Unitlab validates reachability, required edges, terminal stages, and the effect on in-flight items.

Important UX rules:

* exactly one Project and one Complete stage;
* Project has no incoming edge and must lead to an entry stage;
* Review has one Approve and one Reject route;
* terminal stages have no outgoing routes;
* every stage is reachable from Project;
* duplicate outgoing action names are rejected;
* a stage holding active items cannot be silently deleted;
* sensitive automated-stage configuration cannot change while occupied without resolving the impact.

When an active workflow change would strand work, Unitlab returns an apply-impact conflict instead of silently moving or losing items.

### Use this in production

* Capture the current graph, stage IDs, task counts, assignees, priorities, open Issues, and automation dependencies.
* Define how every in-flight stage maps to the new graph.
* Test each role and route in a controlled cohort before applying the change.
* Monitor queues, failed work, unassigned tasks, reviewer availability, and release timing after rollout.
* Record owner, reason, affected population, migration result, and rollback decision.

{% hint style="warning" %}
Do not edit a live graph only to improve its appearance. Every stage and edge can be an operational dependency.
{% endhint %}

***

> **Continue with Unitlab:** [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# Queues overview

Understand Task Queue, Batch Queue, assignment, priority, and observability.

Queues are the operational surface of the workflow. Task Queue shows individual work availability and ownership; Batch Queue provides processing and grouped operational context.

{% hint style="info" %}
**Use this area when:** you need to allocate work, inspect backlog, prioritize cohorts, recover failures, or understand why a user cannot start a task.
{% endhint %}

### How this area fits into production

```mermaid
flowchart TB
  A["Workflow stage"]
  B["Task Queue"]
  C["Assignment + priority"]
  D["Workbench action"]
  E["Next stage"]
  F["Batch Queue processing"]
  A --> B
  B --> C
  C --> D
  D --> E
  F -. supports .-> B
```

![Unitlab Task Queue](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FNuZoKkKYJ9eoN6ckUKWT%2Ftask-queue.png?alt=media\&token=698abe91-80e4-461c-9837-9874a5b22451)

*The queue should make stage, status, priority, assignee, and valid actions visible enough for operators to work without a separate spreadsheet.*

### What this area controls

The project Queue page has two URL-persisted tabs with different purposes:

| Queue           | Default? | Unit represented                          | Primary action                  |
| --------------- | -------- | ----------------------------------------- | ------------------------------- |
| **Task Queue**  | Yes      | One workflow work item waiting in a stage | Manage Workflow / open work     |
| **Batch Queue** | No       | One upload or import action               | Upload data / inspect ingestion |

### Start with the right page

| Decision                             | Production guidance                                                        |
| ------------------------------------ | -------------------------------------------------------------------------- |
| Operate individual work              | Use Task Queue.                                                            |
| Monitor grouped or asynchronous work | Use Batch Queue.                                                           |
| Allocate fairly                      | Configure assignment, claiming, and priority.                              |
| Change many items                    | Use filters and bulk actions only after reviewing the resolved target set. |
| Recover exceptions                   | Inspect failed, invalid, unassigned, or stalled cohorts.                   |

### Operating boundary

* Queue state reflects workflow state; it does not replace the workflow graph.
* Invalid annotation, rejected review, and processing failure are different conditions.
* Priority should express an approved business rule, not hidden personal urgency.

### A production-ready handoff

Every visible backlog has an owner, priority rule, service expectation, valid action, exception path, and reconciled count.

***

> **Related Unitlab capability guides:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# Task Queue

Assign, claim, prioritize, filter, and start individual work.

Task Queue is where workflow state becomes executable work. What a user sees depends on the project, stage, role, assignment, filters, and item status.

### Before you make the change

* Confirm the intended workflow stage and role.
* Know the assignment strategy and priority policy.
* Define how unassigned, skipped, rejected, invalid, or reopened work is handled.

![Task Queue with stage and action controls](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FNuZoKkKYJ9eoN6ckUKWT%2Ftask-queue.png?alt=media\&token=698abe91-80e4-461c-9837-9874a5b22451)

*Read the queue as an operating view: stage and filters define the population; assignment and priority define order; task actions change ownership or state.*

### Understand the product behavior

The Task Queue is a two-pane operational surface.

**Stage rail**

* one row per visible stage in the bound workflow;
* tinted stage-type icon and stage name;
* copyable stage identifier;
* live item count;
* selected stage persisted in the URL.

The Project entry stage is not shown as a work queue even though it anchors routing.

**Task table**

* select checkbox;
* inline numeric Priority editor for managers;
* media thumbnail and short task ID;
* derived Status;
* Assigned to;
* Actions.

The default order is highest priority first. The explicit row action is **Open**, which launches the correct native editor in the Workbench while preserving queue scope. Assign and Release live in the Assigned-to dropdown. Moving work between stages is a bulk action, not a casual per-row shortcut.

When one or more tasks are selected, the bulk action bar offers Priority, Move, Unassign, and Assign. Bulk Move first calculates a plan and asks for confirmation.

Data Groups appear as one mixed work item with group name, tile count, aggregate status, assignee, and priority. Their member tiles do not appear as duplicate tasks.

Manager/admin users can view and manage all stage queues according to permission. Member-family users see only stages for which their annotator/reviewer position is eligible and only work that is assigned, claimable, or intentionally visible as unassigned.

Task queues answer operational questions:

* What work is waiting at each stage?
* Which tasks are unassigned?
* Who owns the oldest or highest-priority items?
* Which modality or dataset is causing a backlog?
* How much work was rejected back to annotation?

### Operate the individual backlog

{% stepper %}
{% step %}

#### 1. Select the stage

Open the queue population that corresponds to the current responsibility.
{% endstep %}

{% step %}

#### 2. Apply filters

Resolve status, source, assignee, modality, issue state, or other cohort criteria.
{% endstep %}

{% step %}

#### 3. Review priority and ownership

Confirm the order and whether items are unassigned, assigned to self, or assigned elsewhere.
{% endstep %}

{% step %}

#### 4. Assign or claim

Use the configured operating model; avoid side-channel assignment.
{% endstep %}

{% step %}

#### 5. Start the task

Open the Workbench and confirm the same project, stage, and item context.
{% endstep %}

{% step %}

#### 6. Reconcile after action

Check the task moved to the expected owner or next stage.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision           | Production guidance                                                                   |
| ------------------ | ------------------------------------------------------------------------------------- |
| Self-claim         | Useful for elastic workforces when stage eligibility and fairness controls are clear. |
| Manual assignment  | Useful for specialists, calibrated cohorts, or explicit ownership.                    |
| Priority           | Document the business rule and review for starvation or hidden overrides.             |
| Unavailable action | Check role, stage, assignment, selection, task status, and processing state.          |

### Continue the operating flow

* Use saved queue views for recurring operational questions.
* Inspect Batch Queue when processing context is missing.
* Escalate stalled or orphaned cohorts to the workflow owner.

***

> **Explore related Unitlab capabilities:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# Batch Queue

Monitor grouped processing, status, failures, and recovery context.

Batch Queue exposes asynchronous or grouped operations that cannot be understood from a single item alone. Use it to separate request acceptance from actual processing completion.

### Before you make the change

* Record the operation, expected item count, owner, and correlation or batch ID.
* Know whether retry is safe and idempotent.
* Define the failure destination and reconciliation method.

### Understand the product behavior

Each Batch Queue row represents one upload/import action and shows:

* cover thumbnail or type fallback;
* name and import date;
* data-type chips;
* item count;
* progress bar;
* queue status;
* assignee avatars aggregated from the imported items;
* running automation summary where applicable.

Search matches queue name and metadata. Filters include queue status, assignee, and one or more data types; selecting several types uses OR behavior. Opening a row shows the project data grid scoped to that upload session. Uploading again from inside this detail view reuses the same Batch Queue; uploading from the top-level project page creates a new queue.

The SDK exposes total, completed, processing, and failed counts plus item-level data. A queue can be finished processing and still contain failures, so operators should inspect both the overall state and individual rows.

### Monitor and reconcile a batch

{% stepper %}
{% step %}

#### 1. Open the batch

Identify the originating operation, project, source, owner, and expected population.
{% endstep %}

{% step %}

#### 2. Read status and counts

Separate queued, running, succeeded, failed, cancelled, or other current states.
{% endstep %}

{% step %}

#### 3. Inspect failures

Review representative errors and determine whether the cause is source data, validation, permission, integration, or platform processing.
{% endstep %}

{% step %}

#### 4. Check remote state before retry

Confirm which items already succeeded so a retry does not duplicate work.
{% endstep %}

{% step %}

#### 5. Recover the cohort

Correct the cause, retry only the safe scope, or route failures to an owned exception path.
{% endstep %}

{% step %}

#### 6. Close the record

Reconcile expected and final counts and retain the batch ID and outcome.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision        | Production guidance                                                       |
| --------------- | ------------------------------------------------------------------------- |
| Retry           | Use only after inspecting completed and partial state.                    |
| Partial success | Preserve successful work and isolate the failed cohort.                   |
| Timeout         | A client timeout does not prove the remote operation failed.              |
| Ownership       | Every non-terminal batch needs a named operator and escalation threshold. |

### Continue the operating flow

* Return successful items to their Task Queue stage.
* Inspect repeated failures as a systemic source or integration issue.
* Include important batch IDs in the operational record.

***

> **Related Unitlab capability guides:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# Assignment and priority

Choose a transparent allocation model for human work.

Assignment design balances fairness, specialization, throughput, review independence, and operational control. The queue should express that design visibly.

Unitlab supports direct assignment and “Anyone” availability. A mature operating model combines:

* pooled work for routine tasks;
* direct assignment for accountable ownership;
* expert routing for difficult cases;
* priority values for urgent or high-value units;
* reviewer independence where the risk justifies it.

Assignment belongs to the workflow item state. Workspace role determines broad authority; project position and stage configuration determine whether a user can work as an annotator or reviewer in that queue.

### Use this in production

* Use self-claim for broad eligible pools and manual assignment for specialization or controlled calibration.
* Keep reviewer independence and conflict rules explicit.
* Define priority inputs, tie-breaking, aging, and starvation prevention.
* Monitor unassigned, long-running, repeatedly rejected, and reopened cohorts.
* Change the allocation rule through workflow and role review—not through side-channel spreadsheets.

***

> **Related Unitlab capability guides:** [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# Queue failures and recovery

Diagnose missing actions, stalled work, invalid states, and partial processing.

Queue recovery starts by identifying the layer that failed: permission, workflow stage, assignment, annotation validation, source processing, model inference, batch processing, or downstream release creation.

### The operating model

```mermaid
flowchart TB
  A["Symptom"]
  B["Role + stage"]
  C["Assignment + task state"]
  D["Source + validation"]
  E["Batch/model processing"]
  F["Owned recovery"]
  A --> B
  B --> C
  C --> D
  D --> E
  E --> F
```

### Use this in production

* A missing button usually starts with role, stage, selection, or resource-state checks.
* Invalid annotations remain saved but require property or value correction before acceptance.
* Rejected work needs a corrective route; failed processing needs operational recovery.
* For partial batches, preserve succeeded items and retry only after remote-state inspection.
* Escalate systemic failure patterns with affected IDs, counts, timestamps, versions, and redacted errors.

***

> **Related Unitlab capability guides:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# Releases overview

Understand release content, provenance, formats, splits, and lifecycle.

A release is a deliberate frozen delivery of approved data and annotations. It records what was included, how it was encoded, how it was split, which source files were bundled, and which project context produced it.

{% hint style="info" %}
**Use this area when:** approved project work must become a reproducible training, evaluation, analytics, or audit input.
{% endhint %}

### How this area fits into production

```mermaid
flowchart LR
  A["Approved project work"]
  B["Release scope"]
  C["Format + splits"]
  D["Processing"]
  E["Inspect + download"]
  F["Downstream validation"]
  A --> B
  B --> C
  C --> D
  D --> E
  E --> F
```

![Unitlab Releases overview](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F9VKDEzTnJLl3oTFNKrSw%2Freleases-overview.png?alt=media\&token=9459389d-f245-46cd-b624-9e2bdcda0810)

*The release list is the delivery ledger: name, version, project, status, and ownership should make each handoff understandable before opening it.*

### What this area controls

### What a release contains

A release is a versioned annotation snapshot with:

* visibility such as Private or Public;
* modality;
* version number;
* creation date;
* item counts;
* data preview;
* annotation preview;
* Clone Release;
* Overview, Data, and Settings tabs;
* item type, data ID, preview, ground truth, and metadata.

A release can contain video, document, and audio items together. Metadata is represented as structured content: video metadata can include frame, video, and audio facts, while document metadata includes PDF and page information.

### Start with the right page

| Decision           | Production guidance                                 |
| ------------------ | --------------------------------------------------- |
| Plan a handoff     | Define release scope and acceptance first.          |
| Create delivery    | Choose format, splits, and source inclusion.        |
| Inspect the result | Review Overview, Data, and Settings.                |
| Move downstream    | Download and validate in the target consumer.       |
| Reproduce later    | Retain versioned provenance and stable identifiers. |

### Operating boundary

* A dataset version freezes source membership; a release freezes delivered project output.
* Release creation is asynchronous until processing reaches a terminal state.
* The release is not accepted until the downstream consumer validates it.

### A production-ready handoff

The release has a named owner, explicit scope, source and schema provenance, format and split contract, processing result, downstream acceptance, and retained stable ID.

***

> **Related Unitlab capability guides:** [multimodal data curation](https://unitlab.ai/en/data-curation)


# Create a release

Freeze approved content with an explicit export contract.

Release creation starts inside the project—not from the workspace My Releases gallery. The project Releases page scopes the eligible annotated data and exposes the queue, data-type, format, distribution, and export-token controls used for that delivery.

### Before you make the change

* Resolve or explicitly exclude rejected, invalid, failed, escalated, or unreviewed work.
* Choose the downstream consumer and schema requirements.
* Name the release owner and validation owner.

![Project Release dialog showing annotated data, queue scope, data types, format, distribution, and token URL option](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FHLzEBrP9ddsrTUbQ4ovz%2Frelease-create-modal.png?alt=media\&token=70baced4-91a7-4c40-bf7b-417177d75b90)

*Open Project › Releases and choose Release. The dialog shows the eligible annotated item count, selected queue scope, included data types, export format, train/validation/test distribution, and optional export token URL before creation.*

### Understand the product behavior

Releases are created from a project’s Releases page:

1. The user opens the export dialog.
2. The user chooses the source scope: the full project or one/more Batch Queues.
3. The user selects one or more data families inside that scope.
4. Unitlab previews releasable counts, annotation/review progress, format compatibility, and recommended format.
5. The user chooses an export format or per-family Standard Bundle formats.
6. The user selects a license when required.
7. The user enters nonnegative integer train, validation, and test percentages that sum to 100.
8. The user optionally enables stable tokenized item URLs.
9. Unitlab creates the versioned release.

Invalid latest histories are excluded from ordinary releasable counts. Cloud-storage-sourced releases are forced private.

### Create the controlled delivery

{% stepper %}
{% step %}

#### 1. Open the project Releases page

Enter the approved project and choose Releases in project navigation. The workspace My Releases page is the cross-project gallery, not the creation entry point.
{% endstep %}

{% step %}

#### 2. Open Release

Choose Release and reconcile the eligible Annotated data count before configuring delivery.
{% endstep %}

{% step %}

#### 3. Choose queue scope

Use Whole project or select the intended queues so the release contains the approved operational cohort.
{% endstep %}

{% step %}

#### 4. Choose data types

Include only the Image, Video, Audio, Document, Text, Medical, or grouped types required by the consumer.
{% endstep %}

{% step %}

#### 5. Choose the format

Select UUEF or another supported format that preserves every required geometry, property, relation, temporal field, and group context.
{% endstep %}

{% step %}

#### 6. Configure distribution

Set train, validation, and test percentages with the sliders; keep related entities and Data Groups in one split when leakage control requires it.
{% endstep %}

{% step %}

#### 7. Decide token behavior

Enable Include export token URL only when the approved downstream access model requires it; treat generated tokens and URLs as secrets.
{% endstep %}

{% step %}

#### 8. Create and monitor

Choose Create release, retain the release ID, and wait for terminal processing before opening the release detail and downloading output.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision      | Production guidance                                                                     |
| ------------- | --------------------------------------------------------------------------------------- |
| Queue scope   | Whole project is correct only when every eligible item is approved for this handoff.    |
| Content scope | Make data types, inclusions, and exclusions explicit and reviewable.                    |
| Format        | Choose for downstream fidelity, not convenience alone.                                  |
| Splits        | Use a reproducible policy that prevents leakage across related items or groups.         |
| Token URL     | Enable only for an approved access workflow; never publish or log the resulting secret. |
| Sources       | Balance reproducibility, storage, access, and sensitive-data requirements.              |

### Continue the operating flow

* Inspect sample output and counts.
* Download into a controlled validation environment.
* Record downstream acceptance or required correction.

***

> **Explore related Unitlab capabilities:** [training-data curation workflows](https://unitlab.ai/en/data-curation)


# Export formats and source inclusion

Choose a representation that preserves the required annotation contract.

Export format determines which Unitlab structures remain native and which require mapping. Source inclusion determines how a consumer locates media, documents, groups, and protected data.

Working converters include:

| Data family  | Supported formats          |
| ------------ | -------------------------- |
| Image        | COCO, YOLOv8, YOLOv5, UUEF |
| Video        | COCO, YOLOv8, YOLOv5, UUEF |
| Medical      | COCO, YOLOv8, YOLOv5, UUEF |
| Document/PDF | UUEF only                  |
| Text         | JSONL, UUEF                |
| Audio        | Audio JSON, RTTM, UUEF     |

Native single-family formats apply to one compatible family. A release containing more than one family must use **UUEF** or **Standard Bundle**.

* **UUEF** is the universal full-fidelity format. It preserves properties, attributes, relations, item properties, tags, page/frame context, and the richer annotation graph.
* **Standard Bundle** creates one ZIP containing a native output per selected family plus a manifest describing the release, source scope, selected types, split ratios, written files, and omitted empty combinations.

Default Standard Bundle choices are COCO for image/video/medical, JSONL for text, and Audio JSON for audio. Document data uses UUEF.

### Use this in production

* Inventory classes, geometry, masks, keypoints, segments, spans, relations, properties, Item Properties, temporal state, and Data Group context before choosing a format.
* Prefer the platform-native representation when another format would lose required structure.
* Document any lossy mapping or post-processing step.
* Choose source bundling, references, stable URLs, and tokens according to downstream runtime and security requirements.
* Validate a sample with the real parser before releasing full volume.

***

> **Continue with Unitlab:** [multimodal data curation](https://unitlab.ai/en/data-curation)


# Release splits and grouped data

Create train, validation, and test membership without leakage.

Split policy is a modeling and provenance decision. Related views, sequences, documents, patients, acquisitions, or Data Groups may need to remain in one split to prevent leakage.

A release is a versioned annotation snapshot with:

* visibility such as Private or Public;
* modality;
* version number;
* creation date;
* item counts;
* data preview;
* annotation preview;
* Clone Release;
* Overview, Data, and Settings tabs;
* item type, data ID, preview, ground truth, and metadata.

A release can contain video, document, and audio items together. Metadata is represented as structured content: video metadata can include frame, video, and audio facts, while document metadata includes PDF and page information.

### Use this in production

* Choose the split unit before choosing percentages.
* Keep related Data Group members, study series, time-adjacent sequences, or same-entity samples together when required.
* Record the random seed or deterministic membership rule when available.
* Inspect class, source, modality, and difficulty distribution across splits.
* Treat split changes as a new release version.

***

> **Related Unitlab capability guides:** [Unitlab’s multimodal annotation platform](https://unitlab.ai/en/multimodal-annotation) · [Unitlab’s data curation platform](https://unitlab.ai/en/data-curation)


# Inspect, download, and validate

Prove release integrity in the downstream consumer.

Unitlab processing success confirms the release was created; downstream validation confirms the release is useful. Both are required for acceptance.

### Before you make the change

* Know the expected item count, schema, geometry conventions, group behavior, splits, and source path model.
* Use the same parser or consumer that production will use.
* Define acceptance tolerances and the owner of a failed validation.

![Release overview and data inspection](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FeMSyrUXzrEfCeiYryFcW%2Frelease-detail.png?alt=media\&token=9a228bde-4a68-4e78-bf9d-ec843ad3415f)

*Inspect release metadata and sample content before download, then reproduce the same checks in the actual downstream environment.*

### Understand the product behavior

Data Space contains **My Releases** and **Public Releases**, toggled from the page header. A release detail contains Overview, Data, and—on private workspace releases—Settings.

Opening a release item uses a separate read-only annotation viewer for Image, Video, Audio, Text, Medical, or Document. It reuses the familiar modality surface but disables mutation. This lets teams inspect exactly what a release contains without accidentally changing the project history that produced it.

Making a release public requires license selection. Release deletion is permission- and subscription-gated. If the latest release version is deleted, the previous version becomes latest.

### Export formats

Working converters include:

| Data family  | Supported formats          |
| ------------ | -------------------------- |
| Image        | COCO, YOLOv8, YOLOv5, UUEF |
| Video        | COCO, YOLOv8, YOLOv5, UUEF |
| Medical      | COCO, YOLOv8, YOLOv5, UUEF |
| Document/PDF | UUEF only                  |
| Text         | JSONL, UUEF                |
| Audio        | Audio JSON, RTTM, UUEF     |

Native single-family formats apply to one compatible family. A release containing more than one family must use **UUEF** or **Standard Bundle**.

* **UUEF** is the universal full-fidelity format. It preserves properties, attributes, relations, item properties, tags, page/frame context, and the richer annotation graph.
* **Standard Bundle** creates one ZIP containing a native output per selected family plus a manifest describing the release, source scope, selected types, split ratios, written files, and omitted empty combinations.

Default Standard Bundle choices are COCO for image/video/medical, JSONL for text, and Audio JSON for audio. Document data uses UUEF.

### SDK export behavior

The SDK can:

* create a release from a project;
* specify `export_type="UUEF"`;
* define split ratios such as 80% train and 20% test;
* list and retrieve releases;
* download annotations for a selected split;
* download associated files to a destination folder.

The SDK supports UUEF creation and split-aware annotation and file downloads. The complete release format matrix is shown above.

### Accept the downstream handoff

{% stepper %}
{% step %}

#### 1. Inspect in Unitlab

Review status, counts, version, project, sources, splits, format, settings, and representative items.
{% endstep %}

{% step %}

#### 2. Download deliberately

Use the approved destination and source-handling method.
{% endstep %}

{% step %}

#### 3. Parse the output

Load annotations, properties, relations, temporal data, groups, and source references with the production parser.
{% endstep %}

{% step %}

#### 4. Check semantics and geometry

Compare representative output to the Workbench and ontology version.
{% endstep %}

{% step %}

#### 5. Reconcile counts and splits

Confirm every expected item and no unintended item is present.
{% endstep %}

{% step %}

#### 6. Record acceptance

Store release ID, version, source dataset versions, ontology, workflow, model versions, format, split policy, exclusions, and validation result.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision   | Production guidance                                                                                                                            |
| ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
| Mismatch   | Stop consumption, isolate the scope, and determine whether the cause is source, project state, format mapping, processing, or parser behavior. |
| Correction | Create a new controlled release rather than altering a delivered artifact silently.                                                            |
| Retention  | Retain enough metadata and artifacts to reproduce the handoff under the organization policy.                                                   |
| Consumer   | Acceptance belongs to the real downstream owner.                                                                                               |

### Continue the operating flow

* Promote the accepted release to its intended training or analysis pipeline.
* Create a new version for any correction.
* Use the release record during audit or model-reproduction work.

***

> **Continue with Unitlab:** [training-data curation workflows](https://unitlab.ai/en/data-curation)


# Collaboration overview

Understand how workspace and project responsibilities fit together.

Enterprise annotation is collaborative by design. Workspace membership establishes the identity boundary; roles and permissions control capability; project assignment and workflow stages control current responsibility; comments, Issues, Instructions, and notifications preserve context.

{% hint style="info" %}
**Use this area when:** you are onboarding a team, separating duties, operating review, or preparing a project for an internal or external workforce.
{% endhint %}

### How this area fits into production

```mermaid
flowchart TB
  A["Workspace member"]
  B["Workspace role"]
  C["Project assignment"]
  D["Workflow stage"]
  E["Task ownership"]
  F["Comment, Issue, or review decision"]
  A --> B
  B --> C
  C --> D
  D --> E
  E --> F
```

![Roles and permissions settings](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2Fae9ur8IvUdY0Oj98kTAD%2Froles-permissions.png?alt=media\&token=605e99b1-c6d9-4642-8d2f-b724c50790ab)

*The permission matrix should be read alongside project assignment and workflow-stage eligibility; one layer alone does not describe effective access.*

### What this area controls

### Authentication and account recovery

Unitlab supports email/password sign-up and sign-in, Google authentication, email verification, invitation-token access, password reset, and TOTP two-factor authentication.

Two-factor authentication includes QR/secret setup, verification, ten single-use backup codes shown once, login challenge, disable, and backup-code regeneration. When 2FA is enabled, password change, password-reset completion, account deletion, and workspace destruction require an appropriate second factor.

If a user refreshes during the temporary 2FA login challenge, the challenge is cleared and the user returns to login rather than leaving reusable sensitive state in the browser.

### First-workspace onboarding

1. The user authenticates.
2. If no workspace exists, Unitlab opens the workspace wizard.
3. The user selects a purpose: Work, Education, or Personal.
4. The user enters a workspace name.
5. The user can optionally invite teammates.
6. Unitlab creates the workspace, makes the user Owner, provisions the free subscription, and creates the initial workspace API key.
7. The application switches into the new workspace and guides the user toward projects.

Guided quick-start actions include creating a project, creating or cloning a release, inviting members, integrating a model, configuring reviewer or custom-model projects, trying batch/crop auto-annotation or Magic Touch, and opening project/member statistics.

### Workspace switching and settings

The workspace area includes:

* workspace list and switcher;
* general settings for name, purpose, and logo;
* account security and 2FA;
* billing and pricing portal;
* usage and quota visibility;
* members and member statistics;
* Roles & Permissions editor;
* API keys;
* cloud storage connections;
* user profile.

Usage can report datasource, image, video, medical, token, audio-duration, AI-inference, and member consumption. The interface warns when a downgraded plan would be exceeded.

Workspace destruction is intentionally different from leaving a workspace. It is Owner-only, requires the exact workspace name, and requires a second factor when the Owner has 2FA.

### Members

Member management supports search, filtering, invitations, role changes, member actions, analytics, and Active, Pending, Disabled, and Rejected states. Pending invitations can be resent.

### Start with the right page

| Decision            | Production guidance                                    |
| ------------------- | ------------------------------------------------------ |
| Bring in a person   | Invite and manage members.                             |
| Define capability   | Use role-based access and permission groups.           |
| Assign current work | Use project membership, workflow stages, and queues.   |
| Preserve context    | Use Instructions, comments, Issues, and notifications. |
| Remove access       | Follow the offboarding runbook.                        |

### Operating boundary

* Workspace role, project role, workflow-stage action, and current assignment are distinct layers.
* Comments preserve context; Issues create owned follow-through; workflow review changes task state.
* External workforces should receive the minimum project and data access required.

### A production-ready handoff

Every person and service identity has a named sponsor, minimum role, intended projects, current responsibilities, review date, and offboarding path.

***

> **Related Unitlab capability guides:** [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# Members and invitations

Onboard and remove people through an organization-owned identity process.

Membership is the entry point to the Unitlab trust boundary. Invite only known users for an approved purpose and remove access as soon as that purpose ends.

### Before you make the change

* Confirm the sponsor, business purpose, expected projects, role, and review date.
* Use the organization-approved email and identity verification process.
* Separate human identities from service identities.

![Workspace Members page with the Invite Member dialog](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FEDjhQt2nQzIfnIWpcDYU%2Fmember-invite-modal.png?alt=media\&token=4608b940-5f0e-49a3-bfcc-9b876c2fad1e)

*Go to Workspace Settings › Members and choose Invite to workspace. Enter the approved email, select the minimum Member role, and send the invitation. The Members table then shows Pending until acceptance and provides Resend invitation and member actions.*

### Understand the product behavior

Member management supports search, filtering, invitations, role changes, member actions, analytics, and Active, Pending, Disabled, and Rejected states. Pending invitations can be resent.

### Onboard a member

{% stepper %}
{% step %}

#### 1. Open Members

From the active workspace, open Settings › Members and confirm seat capacity, current status counts, and the intended workspace.
{% endstep %}

{% step %}

#### 2. Open the invitation dialog

Choose Invite to workspace.
{% endstep %}

{% step %}

#### 3. Enter the approved identity

Enter the organization-approved email address; do not use a shared mailbox as a human identity.
{% endstep %}

{% step %}

#### 4. Select the minimum role

Choose the Member role that covers the immediate responsibility. Prefer Annotator or Reviewer when administrative access is not required.
{% endstep %}

{% step %}

#### 5. Send and monitor

Choose Invite, confirm the member appears as Pending, and use Resend invitation only after verifying the address and sponsor.
{% endstep %}

{% step %}

#### 6. Verify account setup

After acceptance, confirm Active status, authentication controls, project assignment, workflow eligibility, and important denied actions.
{% endstep %}

{% step %}

#### 7. Offboard completely

Reassign tasks, reviews, Issues, integrations, and owned credentials before disabling or removing membership.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision           | Production guidance                                                   |
| ------------------ | --------------------------------------------------------------------- |
| Workspace role     | Grants broad platform capability and should remain minimal.           |
| Project assignment | Limits current operational scope inside a project.                    |
| Temporary access   | Use a defined end date and scheduled review.                          |
| Removal            | Reconcile owned work before deprovisioning so tasks are not orphaned. |

### Continue the operating flow

* Confirm role-based access.
* Review custom permissions only when built-in roles are insufficient.
* Schedule a periodic membership review.

***

> **Continue with Unitlab:** [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# Role-based access

Use built-in roles to separate ownership, management, annotation, and review.

Built-in roles provide understandable starting points for least privilege. Effective access still depends on workspace, project, workflow, assignment, and resource state.

### Before you make the change

* List the actions each persona needs and does not need.
* Separate administrative ownership, operations management, annotation, and independent review.
* Identify any sensitive data or model-management restrictions.

![Unitlab role-based access settings](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2Fae9ur8IvUdY0Oj98kTAD%2Froles-permissions.png?alt=media\&token=605e99b1-c6d9-4642-8d2f-b724c50790ab)

*Review the effective capability across permission groups rather than relying on a role name alone.*

### Understand the product behavior

Built-in roles include:

* Owner;
* Manager;
* Member;
* Annotator;
* Reviewer.

Administrators can also create custom roles for workspace-specific access patterns.

Roles can be configured to permit assignment as an Annotator or Reviewer.

Workspace role and project position are different:

* **Workspace role** controls tenant-wide capabilities.
* **Project position** determines eligibility for annotator or reviewer workflow stages.

Owner, Manager, and custom roles use the administrative branch by default. Member, Annotator, and Reviewer are assignment-scoped and see only the stage queues and work items for which they are eligible.

### Assign the minimum role

{% stepper %}
{% step %}

#### 1. Choose the closest built-in role

Start with Owner, Manager, Member, Annotator, or Reviewer according to current responsibilities.
{% endstep %}

{% step %}

#### 2. Add project scope

Assign only the projects needed for the work.
{% endstep %}

{% step %}

#### 3. Confirm workflow eligibility

Check the stages and actions available to the role.
{% endstep %}

{% step %}

#### 4. Test with the real account

Verify both expected access and important denied actions.
{% endstep %}

{% step %}

#### 5. Review periodically

Revalidate when responsibility, project scope, or employment status changes.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision  | Production guidance                                                  |
| --------- | -------------------------------------------------------------------- |
| Owner     | Reserve for accountable workspace administration and recovery.       |
| Manager   | Use for operational management without unnecessary ownership rights. |
| Annotator | Limit to assigned annotation work and applicable collaboration.      |
| Reviewer  | Preserve independent review capability and avoid hidden conflicts.   |

### Continue the operating flow

* Create a custom role only for a durable gap.
* Document project and stage assignments.
* Include role review in offboarding and production-readiness checks.

***

> **Explore related Unitlab capabilities:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# Permissions and custom roles

Configure permission groups for durable enterprise responsibilities.

Custom permissions are powerful because they can create exactly the access an organization needs. They also make access harder to explain unless the role has a durable purpose, owner, and review process.

![Permission groups and custom role controls](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2Fae9ur8IvUdY0Oj98kTAD%2Froles-permissions.png?alt=media\&token=605e99b1-c6d9-4642-8d2f-b724c50790ab)

*Review Workspace, Projects, Annotation & Review, and Models & Data permission groups as a whole before saving a custom role.*

Granular permissions include:

**Workspace**

* manage workspace settings;
* manage billing;
* manage API keys;
* manage cloud storage;
* manage members.

**Projects and schemas**

* manage projects;
* manage ontologies;
* assign project members;
* manage releases;
* view ontology;
* view instructions;
* manage instructions.

**Annotation and review**

* view labeling interface;
* create labels;
* comment;
* view statistics.

**Models and data**

* manage data;
* manage AI models;
* manage workflows;
* manage augmentation.

### Use this in production

* Use built-in roles when they meet the responsibility.
* Name custom roles by durable job function, not a temporary person or ticket.
* Document allowed and deliberately denied actions.
* Test project, workflow, queue, model, data, and release behavior with a representative account.
* Assign an owner and review date; remove roles that no longer have active members or purpose.

***

> **Continue with Unitlab:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# Comments, Issues, and notifications

Preserve context, create ownership, and keep work moving.

Collaboration tools serve different purposes: Instructions publish policy, comments preserve contextual discussion, Issues create owned follow-through, notifications create awareness, and workflow actions change the task state.

### The operating model

```mermaid
flowchart LR
  A["Policy question"]
  B["Comment"]
  C["Issue with owner"]
  D["Correction or policy change"]
  E["Workflow review"]
  F["Notification"]
  A --> B
  B --> C
  C --> D
  D --> E
  E --> F
```

Comments live inside the annotation workbench and are appropriate for context attached to a specific item or label. They should not become the only place where a repeated rule is documented; recurring decisions belong in the instructions or ontology.

### Issues

Project issues include content, status, responsible member, creator, and created date. Issue links can return the user to the exact annotation context and load the correct editor for that item’s data family. An issue is appropriate when a problem requires ownership and follow-through beyond one annotation comment.

### Review and specialist escalation

Review is an explicit workflow decision. Specialist escalation can be modeled as another Review stage with a restricted eligible-member list. A useful closed loop is:

```
Annotation or policy error
    ↓
Reviewer correction or rejection
    ↓
Issue categorized
    ↓
Ontology, instruction, model, or assignment change
    ↓
New tasks measured for recurrence
```

The most important quality metric is not raw labeling speed. It is accepted, usable units per hour after correction, rework, and downstream failures are included.

### Notifications

Workflow-aware notifications include stage-ready, task-assigned, task-claimed, rework, comment, mention, automation, and conflict events. The notification list supports bulk actions. Rework notifications are particularly important because they connect a reviewer’s rejection to the annotator who must correct it.

### Statistics and team visibility

Statistics appear at project and member levels. Current surfaces include:

* project overview charts;
* monthly progress;
* daily time series;
* overall progress;
* average time per item;
* total working time;
* issue counts;
* annotator and reviewer breakdowns;
* member heatmaps and summaries;
* workspace member statistics pages.

Some statistics and premium role experiences depend on the workspace subscription. Statistics should be interpreted with workflow context—for example, faster annotation can coexist with higher reviewer rejection and should not be reported as improved productivity in isolation.

### Use this in production

* Use comments for item context that helps another person understand the decision.
* Create an Issue when correction, investigation, or policy work needs an owner and status.
* Use workflow review to accept, reject, or escalate the task itself.
* Configure notifications for awareness; do not use notification receipt as proof of ownership.
* Review recurring Issues as signals of source, Instructions, ontology, workflow, or training problems.

***

> **Continue with Unitlab:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# Security overview

Apply least privilege and explicit ownership across the Unitlab operating model.

Unitlab security is a set of connected identity, access, credential, data-handling, and operational controls. The organization remains responsible for configuring those controls according to its data classification and regulatory obligations.

{% hint style="info" %}
**Use this area when:** you are onboarding a workspace, connecting cloud storage, issuing API keys, handling sensitive data, or reviewing production access.
{% endhint %}

### How this area fits into production

```mermaid
flowchart TB
  A["Human or service identity"]
  B["Authentication"]
  C["Workspace role"]
  D["Project + workflow scope"]
  E["Data/model/release action"]
  F["Review + revocation"]
  A --> B
  B --> C
  C --> D
  D --> E
  E --> F
```

![Personal Account security page with password and two-factor authentication controls](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FJrstdUsQmqNdGlH4P2Mp%2Faccount-security.png?alt=media\&token=8a678b4a-9d4c-463b-b462-17986926e8fa)

*Personal Settings › Account security is the human sign-in control surface. It shows password management, 2FA status, the Enable 2FA action, and the account danger zone; role-based authorization is configured separately at workspace and project scope.*

### What this area controls

Unitlab supports email/password sign-up and sign-in, Google authentication, email verification, invitation-token access, password reset, and TOTP two-factor authentication.

Two-factor authentication includes QR/secret setup, verification, ten single-use backup codes shown once, login challenge, disable, and backup-code regeneration. When 2FA is enabled, password change, password-reset completion, account deletion, and workspace destruction require an appropriate second factor.

If a user refreshes during the temporary 2FA login challenge, the challenge is cleared and the user returns to login rather than leaving reusable sensitive state in the browser.

### Start with the right page

| Decision               | Production guidance                                                                                             |
| ---------------------- | --------------------------------------------------------------------------------------------------------------- |
| Protect sign-in        | Change compromised passwords and enable TOTP two-factor authentication in Personal Settings › Account security. |
| Control access         | Use workspace roles, permission groups, project assignment, and workflow eligibility.                           |
| Protect automation     | Create, store, rotate, disable, and delete API keys from Workspace Settings › API keys.                         |
| Connect source systems | Add cloud storage with a dedicated least-privilege identity and exact prefix.                                   |
| Handle sensitive data  | Apply organization policy to screenshots, exports, logs, source paths, and releases.                            |
| Remove access          | Use the offboarding runbook and verify denial after ownership transfer.                                         |

### Operating boundary

* Do not place secrets, signed URLs, regulated content, or private source paths in public docs or tickets.
* A personal user credential is not a service identity strategy.
* Access reviews must include in-flight tasks, integrations, model endpoints, keys, and release destinations.

### A production-ready handoff

Every privileged identity and connection has an owner, minimum scope, approved secret location, rotation and revocation path, review date, and tested failure behavior.

***

> **Explore related Unitlab capabilities:** [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# Authentication and account recovery

Protect human accounts and recovery paths.

Account controls protect the entry point to every workspace, project, source, model, and release a user can access. Use organization-owned identities and a recovery process that does not depend on one unavailable person.

![Account security page showing password and 2FA status](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FJrstdUsQmqNdGlH4P2Mp%2Faccount-security.png?alt=media\&token=8a678b4a-9d4c-463b-b462-17986926e8fa)

*Open Personal Settings › Account security. Unitlab identifies whether two-factor authentication is enabled and explains that sign-in will require a six-digit code from a TOTP-compatible authenticator app.*

Unitlab supports email/password sign-up and sign-in, Google authentication, email verification, invitation-token access, password reset, and TOTP two-factor authentication.

Two-factor authentication includes QR/secret setup, verification, ten single-use backup codes shown once, login challenge, disable, and backup-code regeneration. When 2FA is enabled, password change, password-reset completion, account deletion, and workspace destruction require an appropriate second factor.

If a user refreshes during the temporary 2FA login challenge, the challenge is cleared and the user returns to login rather than leaving reusable sensitive state in the browser.

### Use this in production

* Require verified organization-owned email addresses and strong unique credentials according to policy.
* To enable 2FA, choose Enable 2FA, scan the QR code or enter the setup key in an approved TOTP authenticator, choose Next, and verify the generated six-digit code.
* Never capture, publish, or paste the QR code, setup key, one-time code, password, or recovery material.
* Test the next sign-in and retain approved recovery information outside tickets and public documentation.
* Keep at least two accountable workspace owners where policy permits so account recovery does not depend on one person.
* Investigate unexpected sign-in, recovery, or notification activity promptly and rotate affected passwords, API keys, and connected credentials.

***

> **Explore related Unitlab capabilities:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# API keys and service identities

Issue, store, rotate, and revoke automation credentials safely.

Automation credentials should represent a named service responsibility with the minimum required scope. They must never be embedded in code, screenshots, public documentation, logs, or support tickets.

### Before you make the change

* Name the service owner, workload, environment, required resources, and review date.
* Choose an approved secret manager and rotation process.
* Define how to disable the workload safely before revoking the key.

![Workspace Settings API Keys header and Create new key action](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F327PnSQHyZTbfbLyJSpN%2Fapi-keys-entry.png?alt=media\&token=dba8de6a-29ac-4957-a68f-74e283028da4)

*Open Workspace Settings › API keys and choose Create new key. The key is added to the table with a masked value, Show and Copy controls, an Enabled switch, and Delete. The screenshot intentionally excludes key identifiers and secret values.*

### Understand the product behavior

API-key management includes create, masked display, reveal, copy, enable/disable, and delete controls.

The Unitlab Python SDK accepts an API key and optional API URL directly, through environment variables, or through CLI configuration. Version 3.0.0 requires Python 3.10 or newer.

### Manage a service credential

{% stepper %}
{% step %}

#### 1. Open the correct workspace

Confirm the workspace name and ID, then open Settings › API keys. Keys belong to that workspace boundary.
{% endstep %}

{% step %}

#### 2. Create a new key

Choose Create new key. Treat this as immediate credential issuance and be prepared to store the value before using any other surface.
{% endstep %}

{% step %}

#### 3. Copy into the secret manager

Use Copy key and place the value directly in the approved secret store. Do not reveal it for screenshots, documentation, or troubleshooting.
{% endstep %}

{% step %}

#### 4. Configure the workload

Inject the secret at runtime through the environment or secret manager expected by the pinned SDK, CLI, or HTTP client.
{% endstep %}

{% step %}

#### 5. Test expected access

Run a read-only identity or list operation first, then one scoped representative workflow. Record stable resource IDs and redacted errors.
{% endstep %}

{% step %}

#### 6. Disable before replacement

Use the Enabled switch to stop a key during investigation or a controlled rotation while preserving the row for reconciliation.
{% endstep %}

{% step %}

#### 7. Rotate

Create and deploy the replacement, verify dependent jobs, then disable and delete the superseded key.
{% endstep %}

{% step %}

#### 8. Delete deliberately

Use Delete only after inventorying all dependent workloads and completing the product confirmation. Deleted keys cannot be restored.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision         | Production guidance                                                                                                                                |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| Service identity | Unitlab issues the workspace API key; your operating model should assign that key to a named non-human workload owner rather than a shared person. |
| Environment      | Use separate credentials and resource scope per production, staging, and development environment.                                                  |
| Enabled state    | Disable first when you need a reversible containment step; delete after dependencies are reconciled.                                               |
| Logs             | Record resource IDs and redacted errors, never the secret or signed response values.                                                               |
| Incident         | Create a replacement and disable the suspected key immediately; inspect affected operations before permanent deletion.                             |

{% hint style="warning" %}
Never paste an API key into GitBook, source control, a screenshot, a chat, or a ticket.
{% endhint %}

### Continue the operating flow

* Document the workload owner and resource scope.
* Schedule rotation and access review.
* Use stable IDs and explicit versions in automated operations.

***

> **Continue with Unitlab:** [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# Cloud credential governance

Control source-system identities, prefixes, rotation, and revocation.

Cloud connections extend the Unitlab trust boundary into another system. Scope the identity to the minimum bucket or prefix, monitor its ownership, and define how Unitlab behavior changes when access is removed.

For enterprise use, the connection should be treated as infrastructure, not as a convenient personal login:

1. Create a dedicated read-only or least-privilege identity for Unitlab.
2. Restrict it to the required bucket/container and prefix.
3. Avoid root or account-wide credentials.
4. Test with a small non-sensitive prefix.
5. Confirm that files, metadata, nested paths, and synchronization behave as expected.
6. Record the source system and connection owner.
7. Separate permission to manage cloud connections from permission to annotate data.

Workspace administrators can create, update, delete, test, and browse cloud connections from Workspace Settings. Connection administration requires cloud-storage permission and should remain separate from ordinary data browsing or annotation.

### Use this in production

* Use a dedicated organization-owned cloud identity.
* Grant the minimum read or write actions and exact storage scope required.
* Store secrets or role configuration in approved systems and rotate on schedule.
* Test enumeration and import against the intended prefix only.
* Before revocation, understand whether already-registered data, source references, and releases remain usable.

{% hint style="warning" %}
Do not publish provider account IDs, secret names, bucket paths, signed URLs, or regulated filenames in public documentation.
{% endhint %}

***

> **Continue with Unitlab:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation) · [Unitlab’s data curation platform](https://unitlab.ai/en/data-curation)


# Sensitive data and offboarding

Remove access without orphaning regulated work or privileged dependencies.

Offboarding is an operational handoff followed by access removal. Reassign work and privileged ownership first; then revoke access and verify the result.

### The operating model

```mermaid
flowchart LR
  A["Identify scope"]
  B["Reassign tasks + reviews"]
  C["Transfer Issues + integrations"]
  D["Rotate keys"]
  E["Remove membership"]
  F["Verify access denied"]
  A --> B
  B --> C
  C --> D
  D --> E
  E --> F
```

### Use this in production

* Inventory workspaces, projects, tasks, review duties, Issues, datasets, releases, cloud connections, model endpoints, API keys, and ownership roles.
* Reassign active operational responsibility before removing the identity.
* Rotate or revoke any secret the person could access.
* Remove workspace and project access according to policy.
* Verify the account and dependent integrations can no longer access protected resources, then retain the completion record.

***

> **Related Unitlab capability guides:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# Workspaces overview

Understand the top-level boundary for people, projects, data, models, and releases.

A workspace is the top-level operational boundary in Unitlab. It contains members, roles, projects, Data Space, datasets, ontologies, models, releases, storage configuration, and subscription or quota context.

{% hint style="info" %}
**Use this area when:** you are creating an environment, separating organizations or programs, managing shared resources, or governing access and capacity.
{% endhint %}

### How this area fits into production

```mermaid
flowchart TB
  A["Workspace"]
  B["Members + roles"]
  C["Data + datasets"]
  D["Projects + workflows"]
  E["Ontologies + models"]
  F["Releases + settings"]
  A --> B
  A --> C
  A --> D
  A --> E
  A --> F
```

![Workspace switcher open from the Unitlab sidebar](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F4jGdNCVAjx3AXBzWerLu%2Fworkspace-switcher.png?alt=media\&token=3874b3f5-b671-44c8-aefe-4851e9d06723)

*The workspace switcher is the boundary selector for projects, Data Space, datasets, models, members, releases, and settings. The item count beside each workspace helps confirm context before entry; New Workspace starts a separate boundary.*

### What this area controls

The workspace area includes:

* workspace list and switcher;
* general settings for name, purpose, and logo;
* account security and 2FA;
* billing and pricing portal;
* usage and quota visibility;
* members and member statistics;
* Roles & Permissions editor;
* API keys;
* cloud storage connections;
* user profile.

Usage can report datasource, image, video, medical, token, audio-duration, AI-inference, and member consumption. The interface warns when a downgraded plan would be exceeded.

Workspace destruction is intentionally different from leaving a workspace. It is Owner-only, requires the exact workspace name, and requires a second factor when the Owner has 2FA.

### Start with the right page

| Decision                    | Production guidance                                  |
| --------------------------- | ---------------------------------------------------- |
| Create or enter a workspace | Use workspace creation and switching.                |
| Manage shared configuration | Use workspace settings.                              |
| Control people              | Use members, roles, and permission groups.           |
| Control source connections  | Use storage and cloud settings.                      |
| Plan capacity               | Review quotas, subscription, and operational limits. |
| Review governance           | Run periodic workspace operations and access review. |

### Operating boundary

* Use separate workspaces when identity, ownership, data policy, or administrative control must be isolated.
* Project permissions do not replace workspace governance.
* Workspace settings changes can affect every project and integration in the boundary.

### A production-ready handoff

The workspace has accountable owners, a defined purpose, approved membership, resource naming standards, shared configuration ownership, capacity monitoring, and an offboarding process.

***

> **Related Unitlab capability guides:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# Create and switch workspaces

Establish the correct top-level boundary and avoid cross-workspace mistakes.

Workspace identity should be obvious before a user changes data, access, ontology, workflow, model, or release state. The sidebar workspace control switches the complete operating boundary; New Workspace starts a separate onboarding flow.

### Before you make the change

* Define why the new boundary cannot be represented by a project inside an existing workspace.
* Name accountable owners, identity policy, data policy, environment, and capacity owner.
* Choose a name that distinguishes the organization or program without exposing sensitive details.

### Follow the interface

#### Open the workspace switcher

![Workspace switcher in the Unitlab sidebar](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F4jGdNCVAjx3AXBzWerLu%2Fworkspace-switcher.png?alt=media\&token=3874b3f5-b671-44c8-aefe-4851e9d06723)

*Choose the active workspace name at the bottom of the sidebar. Select another workspace to switch context, or choose New Workspace to begin creation.*

Switching changes the workspace-scoped projects, Data Space, datasets, ontologies, models, releases, members, roles, API keys, cloud storage, billing, and settings visible to the user.

#### Create the new workspace boundary

![Create your workspace screen with workspace name and Personal or Team usage](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FgFlTnhlxK6re6ODNlD4T%2Fworkspace-create.png?alt=media\&token=42d2589d-5acb-40af-a60a-150679258efc)

*Enter the Workspace Name, choose Personal or Team usage, then continue through the current onboarding flow. Usage describes the intended operating context; it does not replace role and permission design.*

### Understand the product behavior

1. The user authenticates.
2. If no workspace exists, Unitlab opens the workspace wizard.
3. The user selects a purpose: Work, Education, or Personal.
4. The user enters a workspace name.
5. The user can optionally invite teammates.
6. Unitlab creates the workspace, makes the user Owner, provisions the free subscription, and creates the initial workspace API key.
7. The application switches into the new workspace and guides the user toward projects.

Guided quick-start actions include creating a project, creating or cloning a release, inviting members, integrating a model, configuring reviewer or custom-model projects, trying batch/crop auto-annotation or Magic Touch, and opening project/member statistics.

### Create and enter a workspace

{% stepper %}
{% step %}

#### 1. Open the switcher

Choose the active workspace name in the lower-left sidebar.
{% endstep %}

{% step %}

#### 2. Choose New Workspace

Use the dedicated action at the end of the workspace list; do not create a workspace for a temporary task that belongs in a project.
{% endstep %}

{% step %}

#### 3. Name the boundary

Enter the durable organization, environment, or program name.
{% endstep %}

{% step %}

#### 4. Choose usage

Select Personal for an individual operating boundary or Team for a team or organization context, then choose Continue.
{% endstep %}

{% step %}

#### 5. Complete onboarding

Confirm the purpose and enter the new workspace before adding shared data or inviting members.
{% endstep %}

{% step %}

#### 6. Establish ownership and controls

Add accountable owners, configure roles, enable personal 2FA, and define API-key and cloud-connection ownership.
{% endstep %}

{% step %}

#### 7. Switch safely

When returning through the switcher, verify the workspace name and relevant item count before any privileged or bulk action.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision             | Production guidance                                                                                                                              |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ |
| Workspace vs project | Use a new workspace for a durable identity, ownership, policy, billing, or data-isolation boundary; use a project for work inside that boundary. |
| Personal vs Team     | Choose the operating context that matches ownership; team access still requires explicit member and role configuration.                          |
| Owners               | Use more than one accountable owner where policy permits to avoid recovery dependence on one person.                                             |
| Automation           | Bind jobs to explicit workspace and resource IDs, never to whichever workspace a human last opened.                                              |

### Continue the operating flow

* Invite approved members and assign minimum roles.
* Configure cloud storage and API keys only after ownership is recorded.
* Create the first project inside the verified workspace.

***

> **Explore related Unitlab capabilities:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# Workspace settings

Manage shared configuration with impact awareness.

Workspace settings influence people and resources across projects. Treat changes to ownership, permissions, storage, models, integrations, and subscription context as shared configuration changes.

### Before you make the change

* Identify affected projects, integrations, service identities, and owners.
* Record the current state and reason for change.
* Define a test and rollback or containment path.

### Understand the product behavior

The workspace area includes:

* workspace list and switcher;
* general settings for name, purpose, and logo;
* account security and 2FA;
* billing and pricing portal;
* usage and quota visibility;
* members and member statistics;
* Roles & Permissions editor;
* API keys;
* cloud storage connections;
* user profile.

Usage can report datasource, image, video, medical, token, audio-duration, AI-inference, and member consumption. The interface warns when a downgraded plan would be exceeded.

Workspace destruction is intentionally different from leaving a workspace. It is Owner-only, requires the exact workspace name, and requires a second factor when the Owner has 2FA.

### Members

Member management supports search, filtering, invitations, role changes, member actions, analytics, and Active, Pending, Disabled, and Rejected states. Pending invitations can be resent.

### Change shared configuration safely

{% stepper %}
{% step %}

#### 1. Open the correct workspace

Confirm workspace ID, owner, environment, and current settings.
{% endstep %}

{% step %}

#### 2. Identify the control

Determine whether the change affects identity, storage, models, data, projects, releases, or subscription behavior.
{% endstep %}

{% step %}

#### 3. Review dependencies

List the people, jobs, projects, and source systems that rely on the current setting.
{% endstep %}

{% step %}

#### 4. Apply the smallest change

Use the minimum scope needed to achieve the intended result.
{% endstep %}

{% step %}

#### 5. Test effective behavior

Verify expected access and resource operations with representative users or workloads.
{% endstep %}

{% step %}

#### 6. Record and monitor

Capture owner, reason, affected scope, result, and next review trigger.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision  | Production guidance                                                                  |
| --------- | ------------------------------------------------------------------------------------ |
| Scope     | Prefer project-level configuration when the requirement is not truly workspace-wide. |
| Ownership | Every shared integration or setting needs an accountable role.                       |
| Timing    | Schedule high-impact changes around active tasks and automated workloads.            |
| Recovery  | Know how to restore service without reintroducing excessive access.                  |

### Continue the operating flow

* Review membership and roles.
* Re-test cloud and model integrations when shared settings change.
* Update production-readiness evidence for material changes.

***

> **Explore related Unitlab capabilities:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# Storage and cloud settings

Govern reusable source connections and storage behavior at workspace scope.

Cloud Storage is configured once at workspace scope and then used from Data Space to create a connected cloud folder. Separate connection administration from the folder-level source scope that data operators actually import or synchronize.

### Before you make the change

* Create a dedicated organization-owned cloud identity with the minimum bucket and prefix permissions required.
* Record the connection owner, environment, rotation schedule, revocation procedure, and approved secret store.
* Confirm the current Unitlab connection mode and downstream source-availability requirements.

![Add Cloud Storage dialog for Amazon S3](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FEceeKlJexxGEUrAINYfa%2Fcloud-storage-add.png?alt=media\&token=7f2c7aff-bb2a-476f-87cb-a8bd3a70540f)

*Open Workspace Settings › Cloud storage and choose Add cloud storage. Select the provider and enter the display name, bucket, access key ID, secret access key, region, and optional prefix. Credentials in this product capture are intentionally blank.*

### Understand the product behavior

The current Add Cloud Storage dialog offers:

* Amazon S3;
* Google Cloud Storage;
* Azure Blob Storage;
* MinIO;
* DigitalOcean Spaces;
* Cloudflare R2;
* Wasabi;
* Backblaze B2;
* custom S3-compatible storage.

The exact credential fields vary by provider, but the current configuration model includes a display name, bucket or container, credentials, region where applicable, and an optional prefix or sub-path. A prefix can restrict Unitlab to a controlled part of a larger bucket.

The SDK can list safe storage metadata, browse a prefix, create a cloud-backed Unitlab folder, synchronize it, and import selected files or directory paths into a project. Cloud credentials are not returned by SDK list or browse operations.

### Cloud-folder user flow

Cloud folders are created from **New Folder ▾ → Add Cloud Folder**:

1. The user chooses a provider/integration.
2. The user selects an optional bucket sub-prefix and display name.
3. Unitlab creates a root folder representing that cloud location.
4. Opening it registers the first bucket level automatically if it has never been synchronized.
5. Immediate files appear as normal assets; child prefixes appear as child cloud folders.
6. The user can run **Sync** to refresh the current level.

Cloud folders reference customer storage; they do not copy the source bytes into Unitlab storage. Unitlab stores the object reference and a small preview where needed. Cloud folders display a provider badge, Synced state, and provider/resource path. Their contents are read-only mirrors for organization, so drag-move is disabled.

Cloud-folder Sync means “refresh this bucket prefix in Data Assets.” It is not project-to-dataset synchronization. A project receives cloud-backed content only through normal import/upload behavior or by attaching a frozen published dataset version.

### Connect and use cloud storage

{% stepper %}
{% step %}

#### 1. Open Cloud Storage settings

Confirm the active workspace, then open Settings › Cloud storage.
{% endstep %}

{% step %}

#### 2. Add the provider

Choose Add cloud storage and select the supported provider, such as Amazon S3.
{% endstep %}

{% step %}

#### 3. Enter the approved connection

Provide the display name, bucket name, access key ID, secret access key, region, and optional prefix. The prefix restricts the connection to a sub-path.
{% endstep %}

{% step %}

#### 4. Save and confirm Connected

Choose Add, wait for validation, and confirm the connection card shows the correct provider, bucket, region, creation date, and Connected status.
{% endstep %}

{% step %}

#### 5. Create a connected cloud folder

Go to Data Space › Assets, open New Folder, choose Add Cloud Folder, select the approved connection and exact source prefix, then create the folder.
{% endstep %}

{% step %}

#### 6. Reconcile source processing

Compare expected and processed counts, open representative assets, and review Batch Queue failures before dataset creation.
{% endstep %}

{% step %}

#### 7. Operate and rotate

Use Storage actions for approved lifecycle changes. Test a replacement credential before disabling or removing the old connection.
{% endstep %}

{% step %}

#### 8. Validate downstream behavior

Before changing the connection, confirm how existing assets, source references, stable URLs, and releases will behave.
{% endstep %}
{% endstepper %}

### Decisions that affect production

| Decision     | Production guidance                                                                                              |
| ------------ | ---------------------------------------------------------------------------------------------------------------- |
| Display name | Use a durable source-system and environment name that operators can distinguish.                                 |
| Prefix       | Restrict the connection to the smallest approved sub-path.                                                       |
| Identity     | Use a dedicated service identity, not a personal cloud credential.                                               |
| Lifecycle    | Changing or removing a connection can affect ingestion and later source retrieval; reconcile dependencies first. |
| Secrets      | Never place access keys, signed URLs, private bucket paths, or regulated filenames in documentation or logs.     |

{% hint style="warning" %}
The screenshot shows the field structure only. Never populate or publish a real access key ID, secret access key, private prefix, or signed URL in documentation.
{% endhint %}

### Continue the operating flow

* Create the connected cloud folder in Data Space.
* Curate the processed content inside that folder.
* Review credential ownership and rotation on schedule.

***

> **Explore related Unitlab capabilities:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation) · [Unitlab’s data curation platform](https://unitlab.ai/en/data-curation)


# Capacity, limits, and workspace operations

Monitor shared scale, queues, ownership, and periodic governance.

Workspace operations keep growth visible. Capacity is not only storage or subscription—it also includes queue backlog, review availability, model throughput, failed batches, release processing, and privileged-owner coverage.

Statistics appear at project and member levels. Current surfaces include:

* project overview charts;
* monthly progress;
* daily time series;
* overall progress;
* average time per item;
* total working time;
* issue counts;
* annotator and reviewer breakdowns;
* member heatmaps and summaries;
* workspace member statistics pages.

Some statistics and premium role experiences depend on the workspace subscription. Statistics should be interpreted with workflow context—for example, faster annotation can coexist with higher reviewer rejection and should not be reported as improved productivity in isolation.

### 19. Accounts, workspaces, members, roles, and permissions

### Authentication and account recovery

Unitlab supports email/password sign-up and sign-in, Google authentication, email verification, invitation-token access, password reset, and TOTP two-factor authentication.

Two-factor authentication includes QR/secret setup, verification, ten single-use backup codes shown once, login challenge, disable, and backup-code regeneration. When 2FA is enabled, password change, password-reset completion, account deletion, and workspace destruction require an appropriate second factor.

If a user refreshes during the temporary 2FA login challenge, the challenge is cleared and the user returns to login rather than leaving reusable sensitive state in the browser.

### First-workspace onboarding

1. The user authenticates.
2. If no workspace exists, Unitlab opens the workspace wizard.
3. The user selects a purpose: Work, Education, or Personal.
4. The user enters a workspace name.
5. The user can optionally invite teammates.
6. Unitlab creates the workspace, makes the user Owner, provisions the free subscription, and creates the initial workspace API key.
7. The application switches into the new workspace and guides the user toward projects.

Guided quick-start actions include creating a project, creating or cloning a release, inviting members, integrating a model, configuring reviewer or custom-model projects, trying batch/crop auto-annotation or Magic Touch, and opening project/member statistics.

### Workspace switching and settings

The workspace area includes:

* workspace list and switcher;
* general settings for name, purpose, and logo;
* account security and 2FA;
* billing and pricing portal;
* usage and quota visibility;
* members and member statistics;
* Roles & Permissions editor;
* API keys;
* cloud storage connections;
* user profile.

Usage can report datasource, image, video, medical, token, audio-duration, AI-inference, and member consumption. The interface warns when a downgraded plan would be exceeded.

Workspace destruction is intentionally different from leaving a workspace. It is Owner-only, requires the exact workspace name, and requires a second factor when the Owner has 2FA.

### Members

Member management supports search, filtering, invitations, role changes, member actions, analytics, and Active, Pending, Disabled, and Rejected states. Pending invitations can be resent.

### Built-in and custom roles

Built-in roles include:

* Owner;
* Manager;
* Member;
* Annotator;
* Reviewer.

Administrators can also create custom roles for workspace-specific access patterns.

Roles can be configured to permit assignment as an Annotator or Reviewer.

Workspace role and project position are different:

* **Workspace role** controls tenant-wide capabilities.
* **Project position** determines eligibility for annotator or reviewer workflow stages.

Owner, Manager, and custom roles use the administrative branch by default. Member, Annotator, and Reviewer are assignment-scoped and see only the stage queues and work items for which they are eligible.

### Permission groups

Granular permissions include:

**Workspace**

* manage workspace settings;
* manage billing;
* manage API keys;
* manage cloud storage;
* manage members.

**Projects and schemas**

* manage projects;
* manage ontologies;
* assign project members;
* manage releases;
* view ontology;
* view instructions;
* manage instructions.

**Annotation and review**

* view labeling interface;
* create labels;
* comment;
* view statistics.

**Models and data**

* manage data;
* manage AI models;
* manage workflows;
* manage augmentation.

### Governance pattern

Separate high-impact administration from production work. An annotator rarely needs permission to manage API keys, cloud storage, releases, or ontologies. A reviewer should be able to make review decisions without necessarily changing the schema. External contributors should receive a minimal custom role and access only to the projects they need.

### Use this in production

* Track active projects, member growth, source volume, dataset versions, queue backlog, invalid and failed work, model processing, and releases.
* Define thresholds and owners for capacity, failure, and subscription or quota signals.
* Review privileged roles, unused custom roles, service identities, cloud connections, and stale projects on a schedule.
* Archive or retire resources through documented lifecycle paths rather than deleting operational history.
* Use stable IDs and periodic reports so trends remain comparable across display-name changes.

***

> **Related Unitlab capability guides:** [AI training-data annotation](https://unitlab.ai/en/data-annotation) · [multimodal data curation](https://unitlab.ai/en/data-curation)


# API, SDK & CLI

Build production Unitlab automations with the Python SDK, CLI, and authenticated HTTP API.

```mermaid
flowchart LR
  APP["Applications and workers"] --> SDK["Python SDK"]
  OPS["Shell, CI, and runbooks"] --> CLI["CLI"]
  EXT["Other runtimes"] --> HTTP["Authenticated HTTP API"]
  SDK --> PLATFORM["Unitlab platform API"]
  CLI --> PLATFORM
  HTTP --> PLATFORM
  PLATFORM --> RES["Projects · Data · Workflows · Queues · Releases"]
```

*Choose the integration surface by runtime and operating model. All three surfaces address the same authenticated Unitlab resources.*

{% hint style="info" %}
**Outcome:** Build production Unitlab automations with the Python SDK, CLI, and authenticated HTTP API.
{% endhint %}

This reference is generated from the current `unitlab-sdk-codebase` package at version **3.0.0**, its public source, executable command tree, and tests. It does not reuse previous documentation.

### Choose a surface

| Need                           | Recommended surface    |
| ------------------------------ | ---------------------- |
| Typed application integration  | Python SDK             |
| Shell, CI, or operator runbook | CLI with `--json`      |
| Non-Python runtime             | Authenticated HTTP API |

### Coverage

Projects, Data Units, Batch Queues, Assets, folders, cloud storage, grouping, datasets and versions, embedding spaces, ontologies, workflow stages and tasks, releases, errors, and production automation patterns.

{% hint style="warning" %}
The installed SDK and CLI help are authoritative for the pinned package version. Test upgrades against representative create, upload, attach, workflow, embedding, and release operations before changing production jobs.
{% endhint %}

### Operating contract

| Concern           | Required behavior                                                                |
| ----------------- | -------------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` application environment on Python 3.10+.                 |
| Identity          | Least-privilege API key supplied through approved configuration.                 |
| Target resolution | Stable resource IDs and an explicitly bounded target set.                        |
| Success evidence  | Typed return fields, server-side state, and downstream acceptance of the result. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Explore related Unitlab capabilities:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# Developer overview

Choose the Python SDK, CLI, or authenticated HTTP interface for a Unitlab automation.

### One platform, three automation surfaces

```mermaid
flowchart LR
  APP["Applications and workers"] --> SDK["Python SDK"]
  OPS["Shell, CI, and runbooks"] --> CLI["CLI"]
  EXT["Other runtimes"] --> HTTP["Authenticated HTTP API"]
  SDK --> PLATFORM["Unitlab platform API"]
  CLI --> PLATFORM
  HTTP --> PLATFORM
  PLATFORM --> RES["Projects · Data · Workflows · Queues · Releases"]
```

*Choose the integration surface by runtime and operating model. All three surfaces address the same authenticated Unitlab resources.*

{% hint style="info" %}
**Outcome:** Choose the Python SDK, CLI, or authenticated HTTP interface for a Unitlab automation.
{% endhint %}

### Supported surfaces

| Surface    | Best fit                                         | Contract                                          |
| ---------- | ------------------------------------------------ | ------------------------------------------------- |
| Python SDK | Applications, notebooks, workers, data pipelines | Typed resource handles and exceptions             |
| CLI        | Shell workflows, CI jobs, operator runbooks      | Human output or machine-readable JSON             |
| HTTP API   | Runtimes where Python and the CLI cannot be used | Authenticated JSON requests to the SDK API routes |

The Python SDK and CLI cover projects, Data Units, Batch Queues, Assets, folders, cloud storage, Data Groups, datasets and versions, custom embedding spaces, ontologies, workflow tasks, and releases. The current package requires Python 3.10 or newer.

### Resource model

```mermaid
flowchart LR
  C["UnitlabClient"] --> P["Projects"]
  C --> A["Assets & folders"]
  C --> D["Datasets & versions"]
  C --> E["Embedding spaces"]
  C --> O["Ontologies"]
  C --> R["Releases"]
  C --> S["Cloud storage"]
  P --> W["Workflow tasks"]
  P --> B["Batch Queues"]
```

Use the highest-level public method that expresses the operation. The resource handles normalize identifiers, translate product terms, page list responses, and map HTTP failures to typed exceptions.

### Operating contract

| Concern           | Required behavior                                                                |
| ----------------- | -------------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` application environment on Python 3.10+.                 |
| Identity          | Least-privilege API key supplied through approved configuration.                 |
| Target resolution | Stable resource IDs and an explicitly bounded target set.                        |
| Success evidence  | Typed return fields, server-side state, and downstream acceptance of the result. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Related Unitlab capability guides:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# Install the Python SDK and CLI

Install Unitlab 3.0.0 on Python 3.10 or newer and verify the SDK and CLI.

{% hint style="info" %}
**Outcome:** create a reproducible Unitlab environment and prove that both developer entry points resolve to the same installation.
{% endhint %}

### Prerequisites

* Python 3.10 or newer.
* An isolated environment owned by the application or automation job.
* Access to the Python Package Index through the organization’s approved dependency manager.
* A lock file or image manifest that can preserve the selected version.

### Install the package

Add the package named `unitlab` at version `3.0.0` to the approved dependency manifest, then resolve that manifest inside the isolated environment. The distribution provides both the Python module and the `unitlab` console entry point.

Do not place a Unitlab API key in a dependency file, image layer, shell history, or build log. Installation does not require a production credential.

### Verify both entry points

1. Inspect the active environment and confirm that it reports Unitlab version 3.0.0.
2. Import `unitlab` from the application runtime.
3. Run `unitlab --version` from the same environment.
4. Run `unitlab --help` and confirm that the command groups required by the automation are present.
5. If module and console results differ, correct the environment before authenticating.

### Pin production environments

Commit the dependency lock or immutable image reference. Test an upgrade in a non-production workspace against the same create, upload, attach, workflow, and release operations used by the real job.

### Acceptance criteria

* Python and the CLI resolve to one pinned Unitlab installation.
* The package version is recorded in deployment evidence.
* No production credential was used during installation.
* A read-only authenticated check succeeds only after the environment has passed verification.

{% hint style="warning" %}
Treat a package upgrade as an application change. Preserve the prior lock or image so the automation can be rolled back if typed responses, validation behavior, command output, or long-running operations change.
{% endhint %}

***

> **Related Unitlab capability guides:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# Authentication and configuration

Resolve API keys and the API base URL through explicit arguments, environment variables, or CLI configuration.

### Create the credential in Unitlab

![Workspace API key creation entry point](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F327PnSQHyZTbfbLyJSpN%2Fapi-keys-entry.png?alt=media\&token=dba8de6a-29ac-4957-a68f-74e283028da4)

*Workspace Settings › API keys is the credential issuance surface. Create a key there, copy it directly into an approved secret manager, and never reveal it in documentation, logs, command history, or screenshots.*

{% hint style="info" %}
**Outcome:** Resolve API keys and the API base URL through explicit arguments, environment variables, or CLI configuration.
{% endhint %}

### Credential precedence

`UnitlabClient` resolves configuration in this order:

1. Explicit `api_key` or `api_url` constructor argument.
2. `UNITLAB_API_KEY` or `UNITLAB_API_URL` environment variable.
3. Values written by `unitlab configure`.

If no API key is found, the constructor raises `AuthenticationError` before making a request.

```bash
export UNITLAB_API_KEY="YOUR_API_KEY"
export UNITLAB_API_URL="https://api.unitlab.ai"
```

```python
from unitlab import UnitlabClient

client = UnitlabClient()
try:
    print(client.projects.list())
finally:
    client.close()
```

For a developer workstation, the CLI can persist configuration:

```bash
unitlab configure --api-key "$UNITLAB_API_KEY"
unitlab configure --api-url "https://api.unitlab.ai"
```

Use a workspace-owned service identity where possible. Rotate keys on an explicit schedule and immediately after suspected exposure.

### Operating contract

| Concern           | Required behavior                                                                |
| ----------------- | -------------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` application environment on Python 3.10+.                 |
| Identity          | Least-privilege API key supplied through approved configuration.                 |
| Target resolution | Stable resource IDs and an explicitly bounded target set.                        |
| Success evidence  | Typed return fields, server-side state, and downstream acceptance of the result. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Continue with Unitlab:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# Python SDK quickstart

Create a project, upload multimodal data, wait for processing, and inspect the resulting Data Units.

{% hint style="info" %}
**Outcome:** Create a project, upload multimodal data, wait for processing, and inspect the resulting Data Units.
{% endhint %}

```python
from unitlab import UnitlabClient

client = UnitlabClient()
project = client.projects.create("Medical review")

batch = project.upload("./multimodal-data")
status = batch.wait(timeout=1800)

print({
    "project_id": project.id,
    "batch_queue_id": batch.batch_queue_id,
    "uploaded": batch.uploaded,
    "failed": status.failed,
})

for unit in project.data_units():
    print(unit.id, unit.kind, unit.name, unit.status)

client.close()
```

A directory may mix images, video, audio, text, PDFs, DICOM, NIfTI, and NRRD. One upload call creates one Batch Queue. `wait()` ends when server-side processing reaches zero active items; completion can still include individual failures, so inspect `status.failed` and the batch data before the next state-changing step.

### Definition of success

* The project ID is recorded.
* The Batch Queue reaches no active processing.
* Failed files are explained or routed for correction.
* Data Units have the intended loose or grouped structure.
* The project workflow and ontology are ready before tasks are assigned.

### Operating contract

| Concern           | Required behavior                                                                |
| ----------------- | -------------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` application environment on Python 3.10+.                 |
| Identity          | Least-privilege API key supplied through approved configuration.                 |
| Target resolution | Stable resource IDs and an explicitly bounded target set.                        |
| Success evidence  | Typed return fields, server-side state, and downstream acceptance of the result. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Continue with Unitlab:** [AI training-data annotation](https://unitlab.ai/en/data-annotation)


# CLI quickstart

Configure the CLI, create a project, upload data, wait for processing, and emit JSON for automation.

{% hint style="info" %}
**Outcome:** Configure the CLI, create a project, upload data, wait for processing, and emit JSON for automation.
{% endhint %}

```bash
unitlab configure --api-key YOUR_API_KEY
unitlab project create "Road review" --json
unitlab project list
unitlab project upload PROJECT_ID ./data --fps 2.0 --json
unitlab batch-queue wait PROJECT_ID BATCH_QUEUE_ID --json
unitlab project data-units PROJECT_ID --json
```

Use `--json` when another process will parse the result. Human output is concise and can change for readability; JSON output is the machine interface. A project upload exits non-zero on partial file failure, even when some files succeeded. Capture output and inspect the created Batch Queue before deciding whether to retry.

Every command group supports `--help`. Use the installed command’s help as the authoritative option list for the pinned SDK version.

### Operating contract

| Concern           | Required behavior                                                           |
| ----------------- | --------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` command in the intended Python 3.10+ environment.   |
| Machine contract  | Use `--json`; human-readable output is not an automation interface.         |
| Target resolution | Resolve stable IDs with a read command before a state-changing command.     |
| Success evidence  | Exit status, JSON result, returned IDs, and a read-after-write state check. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Explore related Unitlab capabilities:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# UnitlabClient and resource namespaces

Construct and close the client and navigate its public resource namespaces.

{% hint style="info" %}
**Outcome:** Construct and close the client and navigate its public resource namespaces.
{% endhint %}

```python
from unitlab import UnitlabClient

client = UnitlabClient()
print(client.projects.list())
print(client.assets.folders())
print(client.datasets.list())
print(client.embedding_spaces.list())
print(client.ontologies.list())
print(client.releases.list())
print(client.cloud_storages.list())
client.close()
```

The client owns one HTTP session and exposes namespace objects for each resource family. Resource methods accept a handle or ID where documented; the SDK converts handles through their stable identifiers. Close long-lived clients during application shutdown and close short-lived clients in a `finally` block.

Convenience lookups include `get_workflow_task(task_id)`, `get_data_unit(project_id, unit_id)`, and `attach_dataset(project, dataset, version=...)`.

### Operating contract

| Concern           | Required behavior                                                                |
| ----------------- | -------------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` application environment on Python 3.10+.                 |
| Identity          | Least-privilege API key supplied through approved configuration.                 |
| Target resolution | Stable resource IDs and an explicitly bounded target set.                        |
| Success evidence  | Typed return fields, server-side state, and downstream acceptance of the result. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Continue with Unitlab:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# Projects

List, create, retrieve, update, and soft-delete project resources.

{% hint style="info" %}
**Outcome:** List, create, retrieve, update, and soft-delete project resources.
{% endhint %}

```python
project = client.projects.create("Road scenes")
project = client.projects.get(project.id)
project.update(name="Road scenes v2", description="Night and rain")

for item in client.projects.list():
    print(item.id, item.name)
```

`projects.create(name, ontology_hash=None)` can attach a Workspace Ontology at creation. Project creation itself does not upload data or wait for processing. `project.delete()` soft-deletes the project and schedules backend cleanup; resolve the exact ID and downstream release dependencies before calling it.

### Operating contract

| Concern           | Required behavior                                                                |
| ----------------- | -------------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` application environment on Python 3.10+.                 |
| Identity          | Least-privilege API key supplied through approved configuration.                 |
| Target resolution | Stable resource IDs and an explicitly bounded target set.                        |
| Success evidence  | Typed return fields, server-side state, and downstream acceptance of the result. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Related Unitlab capability guides:** [Unitlab’s data annotation platform](https://unitlab.ai/en/data-annotation)


# Data Units

List and retrieve loose datasource or grouped project work units with stable filters.

{% hint style="info" %}
**Outcome:** List and retrieve loose datasource or grouped project work units with stable filters.
{% endhint %}

```python
units = project.data_units(
    search="study-001",
    data_type="medical",
    status="annotate",
    kind="group",
)

unit = client.get_data_unit(project.id, units[0].id)
print(unit.kind, unit.items)
```

A loose file is a `datasource` Data Unit. A Data Group is a `group` Data Unit whose `items` contain member tile summaries; members are not duplicated in the top-level list. Available filters are `search`, `data_type`, `status`, `batch_queue`, and `kind`.

Use Data Unit IDs for task-scoped automation and retain group membership when downstream work depends on context.

### Operating contract

| Concern           | Required behavior                                                                |
| ----------------- | -------------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` application environment on Python 3.10+.                 |
| Identity          | Least-privilege API key supplied through approved configuration.                 |
| Target resolution | Stable resource IDs and an explicitly bounded target set.                        |
| Success evidence  | Typed return fields, server-side state, and downstream acceptance of the result. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Continue with Unitlab:** [multimodal data curation](https://unitlab.ai/en/data-curation)


# Project uploads

Upload local multimodal files into one project Batch Queue and handle partial failure explicitly.

### Where this operation appears in the product

![Project upload dialog with Local, CLI, SDK, and Storage tabs](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FYIMKhOzBn8nrF0VZAPxY%2Fproject-upload-modal.png?alt=media\&token=3d793031-28d6-4529-bf7f-bfe87b2e98b8)

*Project › Datasets › Upload exposes the same project-scoped ingestion boundary used by the SDK and CLI examples on this page.*

{% hint style="info" %}
**Outcome:** Upload local multimodal files into one project Batch Queue and handle partial failure explicitly.
{% endhint %}

```python
batch = project.upload("./data", fps=2.0, batch_size=100)
print(batch.uploaded, batch.failed)

# File-transfer failure is available immediately.
batch.raise_on_failure()

# Server processing is asynchronous.
status = batch.wait(timeout=1800, show_progress=True)
print(status.total, status.completed, status.failed)
```

`fps` defaults to 1.0 and is used for video ingestion. `batch_size` controls files per upload batch. Local validation detects recognized extensions; server processing can still fail after transfer. Treat upload transfer and Batch Queue processing as two separate checkpoints.

### Operating contract

| Concern           | Required behavior                                                                |
| ----------------- | -------------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` application environment on Python 3.10+.                 |
| Identity          | Least-privilege API key supplied through approved configuration.                 |
| Target resolution | Stable resource IDs and an explicitly bounded target set.                        |
| Success evidence  | Typed return fields, server-side state, and downstream acceptance of the result. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Explore related Unitlab capabilities:** [Unitlab’s data curation platform](https://unitlab.ai/en/data-curation)


# Batch Queues and processing

Inspect, wait for, and diagnose asynchronous server processing.

{% hint style="info" %}
**Outcome:** Inspect, wait for, and diagnose asynchronous server processing.
{% endhint %}

```python
queues = project.batch_queues()
queue = project.batch_queue("BATCH_QUEUE_ID")

print(queue.refresh())
print(queue.status())
print(queue.data())
final = queue.wait(timeout=1800)
```

`ProcessingStatus` exposes `status`, `total`, `completed`, `processing`, and `failed`. Completion means `processing == 0`; it does not mean every row succeeded. `ProcessingTimeoutError` includes the last observed status so a caller can log progress and resume monitoring without guessing.

### Operating contract

| Concern           | Required behavior                                                                |
| ----------------- | -------------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` application environment on Python 3.10+.                 |
| Identity          | Least-privilege API key supplied through approved configuration.                 |
| Target resolution | Stable resource IDs and an explicitly bounded target set.                        |
| Success evidence  | Typed return fields, server-side state, and downstream acceptance of the result. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Related Unitlab capability guides:** [training-data curation workflows](https://unitlab.ai/en/data-curation)


# Attach and detach project sources

Preview source attachment, attach folders or dataset versions, and detach only after impact review.

{% hint style="info" %}
**Outcome:** Preview source attachment, attach folders or dataset versions, and detach only after impact review.
{% endhint %}

```python
preview = project.attach_preview(
    dataset_versions=[{
        "dataset_id": "DATASET_ID",
        "version_number": 3,
    }]
)
print(preview.resolved_asset_count, preview.requires_fps)

result = project.attach(
    dataset_versions=[{
        "dataset_id": "DATASET_ID",
        "version_number": 3,
    }],
    fps=2.0,
)

for source in project.attached_sources():
    print(source.id, source.name)
    print(source.detach_preview())
```

`AttachResult` reports created, already-attached, resolved, and unassigned counts; created datasource and Data Group IDs; attachment IDs; the project dataset version; and an optional Batch Queue ID. If a Batch Queue ID is present, monitor it before assigning work.

### Operating contract

| Concern           | Required behavior                                                                |
| ----------------- | -------------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` application environment on Python 3.10+.                 |
| Identity          | Least-privilege API key supplied through approved configuration.                 |
| Target resolution | Stable resource IDs and an explicitly bounded target set.                        |
| Success evidence  | Typed return fields, server-side state, and downstream acceptance of the result. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Related Unitlab capability guides:** [enterprise data annotation workflows](https://unitlab.ai/en/data-annotation)


# Assets and custom metadata

Upload durable assets, select a folder, apply tags, and replace custom metadata intentionally.

{% hint style="info" %}
**Outcome:** Upload durable assets, select a folder, apply tags, and replace custom metadata intentionally.
{% endhint %}

```python
result = client.assets.upload(
    "./images",
    folder="Raw data",
    tags=["incoming"],
    custom_metadata={"source": "camera-7", "shift": "night"},
)
asset = result.assets[0]
asset.update_custom_metadata({"source": "camera-7", "approved": True})
```

Use either `folder` or `folder_id` to select a destination. `path` can scope the upload path. `AssetUploadResult` exposes the folder, uploaded count, file failures, parsed Asset handles, and raw responses.

Custom metadata updates replace the metadata object; pass the complete intended object. Passing `None` clears custom metadata.

### Operating contract

| Concern           | Required behavior                                                                |
| ----------------- | -------------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` application environment on Python 3.10+.                 |
| Identity          | Least-privilege API key supplied through approved configuration.                 |
| Target resolution | Stable resource IDs and an explicitly bounded target set.                        |
| Success evidence  | Typed return fields, server-side state, and downstream acceptance of the result. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Continue with Unitlab:** [Unitlab’s data curation platform](https://unitlab.ai/en/data-curation)


# Folders

Create, traverse, and inspect local or cloud-backed folder resources.

{% hint style="info" %}
**Outcome:** Create, traverse, and inspect local or cloud-backed folder resources.
{% endhint %}

```python
root = client.assets.create_folder("Raw data")
child = root.create_subfolder("2026-07")

for folder in client.assets.folders(parent=root):
    print(folder.id, folder.name)

for folder in client.assets.all_folders():
    print(folder.id, folder.name)

items = child.list_items()
```

`folders(parent=...)` scopes one level; `all_folders()` walks the complete hierarchy. `folder(id)` retrieves one handle. A cloud folder also supports `sync_cloud()`; treat synchronization as a state-changing discovery operation and inspect the result.

### Operating contract

| Concern           | Required behavior                                                                |
| ----------------- | -------------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` application environment on Python 3.10+.                 |
| Identity          | Least-privilege API key supplied through approved configuration.                 |
| Target resolution | Stable resource IDs and an explicitly bounded target set.                        |
| Success evidence  | Typed return fields, server-side state, and downstream acceptance of the result. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Related Unitlab capability guides:** [training-data curation workflows](https://unitlab.ai/en/data-curation)


# Cloud storage

List safe cloud-storage metadata, browse paginated entries, and import approved paths into a project.

### Configure the provider before automation

![Add Cloud Storage configuration dialog](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2FEceeKlJexxGEUrAINYfa%2Fcloud-storage-add.png?alt=media\&token=7f2c7aff-bb2a-476f-87cb-a8bd3a70540f)

*Workspace Settings › Cloud storage creates the approved provider connection used by SDK and CLI cloud-storage operations. Credentials are never returned by the SDK.*

{% hint style="info" %}
**Outcome:** List safe cloud-storage metadata, browse paginated entries, and import approved paths into a project.
{% endhint %}

```python
storage = next(s for s in client.cloud_storages.list() if s.name == "Production")

for entry in storage.browse(prefix="incoming/", page_size=500):
    print(entry.type, entry.name, entry.size, entry.is_extension_allowed)

batch = project.import_cloud(storage, ["incoming/study-001/"])
batch.wait(timeout=1800)
```

Directory paths end in `/`. Browse results expose safe provider metadata; cloud credentials are never returned by the SDK. `browse()` follows response pagination until no next token remains. An optional project argument applies project-specific validation.

### Operating contract

| Concern           | Required behavior                                                                |
| ----------------- | -------------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` application environment on Python 3.10+.                 |
| Identity          | Least-privilege API key supplied through approved configuration.                 |
| Target resolution | Stable resource IDs and an explicitly bounded target set.                        |
| Success evidence  | Typed return fields, server-side state, and downstream acceptance of the result. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Related Unitlab capability guides:** [Unitlab’s data curation platform](https://unitlab.ai/en/data-curation)


# Data Group automation

Suggest, estimate, and create Data Groups from a folder with explicit grouping configuration.

### The product state this automation creates

![Auto-Groups custom layout editor](https://292810646-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FGjVLUz4wthGkGlRKM6rM%2Fuploads%2F8eyEMKLe8RbL8q4uIsNW%2Fauto-groups-custom-layout.png?alt=media\&token=810552ad-bdf3-44b8-b959-3f3afe77847d)

*The SDK grouping configuration controls the same grouping keys, completeness rules, tile identities, and saved Grid, List, or Custom layout exposed by the Auto-Groups wizard.*

{% hint style="info" %}
**Outcome:** Suggest, estimate, and create Data Groups from a folder with explicit grouping configuration.
{% endhint %}

```python
folder = client.assets.folder("FOLDER_ID")
suggestion = folder.suggest_grouping()

config = {
    "group_by": "study_id",
    "layout": {"columns": 2, "rows": 2},
}
print(folder.estimate_grouping(config))
result = folder.auto_group(config)
print(result.group_count, result.grouped_folder().id)
```

Use `suggest_grouping()` to inspect a proposed strategy and `estimate_grouping(config)` to preview its effect. `auto_group()` creates groups and returns a `GroupingResult`. Validate incomplete and ambiguous groups before attaching the result to production work. `tiles_from_template(...)` is exported for constructing template-based tile definitions.

### Operating contract

| Concern           | Required behavior                                                                |
| ----------------- | -------------------------------------------------------------------------------- |
| Execution surface | Pinned `unitlab==3.0.0` application environment on Python 3.10+.                 |
| Identity          | Least-privilege API key supplied through approved configuration.                 |
| Target resolution | Stable resource IDs and an explicitly bounded target set.                        |
| Success evidence  | Typed return fields, server-side state, and downstream acceptance of the result. |

### Failure and recovery boundary

| Condition                                       | Response                                                                                                     |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Authentication or authorization fails           | Stop, correct the service identity or access model, rotate exposed credentials, and rerun a read-only check. |
| Validation or entitlement rejects the operation | Correct the input or entitlement; do not retry an unchanged request.                                         |
| A request times out                             | Inspect remote state before repeating a mutation because the server may have accepted it.                    |
| Asynchronous processing exceeds its deadline    | Preserve the Batch Queue or release ID, continue bounded monitoring, and inspect item-level failures.        |
| Only part of a batch succeeds                   | Keep successful identifiers, isolate failed rows, and retry only the corrected subset.                       |

{% hint style="warning" %}
Pin and test the production SDK version. Use stable IDs, keep secrets out of logs and command history, and record the correlation ID, target IDs, counts, final state, and redacted error details for material state changes.
{% endhint %}

***

> **Explore related Unitlab capabilities:** [cross-modal annotation workflows](https://unitlab.ai/en/multimodal-annotation) · [Unitlab’s data curation platform](https://unitlab.ai/en/data-curation)




---

[Next Page](/llms-full.txt/1)

