> For the complete documentation index, see [llms.txt](https://docs.unitlab.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.unitlab.ai/documentation/annotations/audio-annotation.md).

# Audio Annotation

Audio annotation is a timing decision as much as a semantic decision. Unitlab aligns waveform, spectrogram, playback, transcript context, ontology values, and a complete annotation timeline so teams can create precise speech and sound datasets without losing temporal evidence.

{% hint style="info" %}
**Use this guide when:** you are building speech recognition, diarization, sound-event detection, intent, acoustic-scene, quality, or multimodal audio datasets.
{% endhint %}

## See audio annotation in action

The demo shows audio as part of the multimodal Workbench, with time-based annotation and contextual review.

{% embed url="<https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/multimodal-6.mp4>" %}

[Open the demo in a new tab](https://homepage-files.s3.us-east-2.amazonaws.com/hero-videos/hero/multimodal-6.mp4).

## Before you begin

1. Create or select a project whose data and ontology match this modality.
2. Confirm the project instructions define the unit of annotation, boundary or timing policy, required properties, and review route.
3. Open the project and enter the assigned item from the project data view or queue. The Workbench loads the modality-native editor inside the shared Unitlab shell.

See [Annotation Workbench](/documentation/annotations/annotation-workbench.md) for navigation, saving, item state, comments, issues, and workflow actions.

## Understand the audio work surface

The audio Workbench exposes playback, zoom, a shared playhead, waveform and spectrogram context, temporal ranges, ontology controls, transcript-related values, and the current workflow action. Zoom changes inspection precision; it does not change the underlying timestamp basis.

![Synchronized waveform and spectrogram with one selected time range](/files/mv1yGFQv80UMnfyw2481)

*Waveform and spectrogram provide complementary evidence around one authoritative temporal selection.*

## Supported annotation model

| Annotation type        | Use it for                                                                     |
| ---------------------- | ------------------------------------------------------------------------------ |
| **Temporal event**     | A sound, action, acoustic state, or other interval with start and end time.    |
| **Speaker segment**    | A time range attributed to one speaker identity or role.                       |
| **Transcription span** | Text aligned to a clip or exact temporal interval.                             |
| **Classification**     | Intent, sentiment, acoustic scene, quality, or another item-level label.       |
| **Dynamic property**   | A value that changes over the recording timeline.                              |
| **Relation**           | A connection between speakers, events, transcripts, or contextual items.       |
| **Item Property**      | Language, source, channel, consent, quality, or other whole-recording context. |

Write an explicit boundary policy for silence, overlap, clipped speech, background noise, non-speech events, and minimum-duration segments before production starts.

## Time-aligned transcription review

Align transcript content with the intended speaker or temporal range. Check words at segment boundaries, overlapping speakers, false starts, filler words, unintelligible audio, and project-specific normalization rules. Keep source audio and transcript identity linked so corrections remain traceable.

![Speaker transcript spans aligned to a waveform and playhead](/files/utbFaZvfJHyrKDFLpGj4)

*Transcript content, speaker identity, and time range are reviewed together.*

## Complete audio timelines

Use the timeline to inspect event ranges, speaker turns, transcript coverage, dynamic properties, and quality segments across the complete recording. Long recordings require review at every transition, overlap, channel change, and abrupt acoustic event—not only at uniformly spaced samples.

![Audio moments connected to event, speaker, transcript, and quality ranges](/files/HWh0n8LoBWrvlCH0826T)

*The timeline reveals missing coverage, accidental overlap, and boundary inconsistency.*

## Ontology-driven audio structure

Use classes for speakers and events, properties for structured details, relations for connections, and Item Properties for facts about the entire recording. AI-assisted proposals can accelerate transcription or event discovery, but confidence should route review rather than replace it.

## Annotate one production item

{% stepper %}
{% step %}

#### 1. Listen before labeling

Play enough surrounding audio to understand the event, speaker, language, and acoustic context. Confirm the project’s overlap and silence policy.
{% endstep %}

{% step %}

#### 2. Zoom to the boundary

Use waveform and spectrogram detail to locate the exact start and end. Keep enough context visible to avoid cutting off leading or trailing signal.
{% endstep %}

{% step %}

#### 3. Create the temporal annotation

Drag the intended range, assign the ontology class, and adjust the handles until the interval matches the policy.
{% endstep %}

{% step %}

#### 4. Add transcript and structure

Enter or review time-aligned text, speaker identity, required properties, relations, and Item Properties.
{% endstep %}

{% step %}

#### 5. Review the timeline

Check gaps, overlaps, duplicated events, inconsistent speaker identity, and boundary drift across the full recording.
{% endstep %}

{% step %}

#### 6. Save and route

Resolve validation, save, and use the current workflow action to submit, approve, reject, or escalate.
{% endstep %}
{% endstepper %}

## Quality review

| Review focus   | What to check                                                                               |
| -------------- | ------------------------------------------------------------------------------------------- |
| **Timing**     | Confirm start and end against waveform, spectrogram, playback, and project tolerance.       |
| **Overlap**    | Represent simultaneous speakers or sounds exactly as the ontology and instruction define.   |
| **Transcript** | Apply the same casing, punctuation, filler, normalization, and unintelligible-audio policy. |
| **Identity**   | Keep speaker or event identity stable throughout the recording.                             |
| **Coverage**   | Review the complete timeline for missed intervals and accidental gaps.                      |

{% hint style="warning" %}
A saved annotation is not automatically a production-ready annotation. Required values, boundary or timing policy, cross-item consistency, and the configured review stage still apply.
{% endhint %}

## Move from labels to governed data

Audio outputs should preserve the source recording identity, time basis, segment boundaries, speaker or event identity, transcript text, properties, relations, and item-level context. Keep long-recording curation, review cohorts, dataset versions, and releases traceable to the original media.

![Integrated Unitlab workflow connecting model assistance, annotation, review, and quality assurance](/files/yPoYNby78Kd9mxWTVqxy)

*Use workflows to keep model output, human correction, review, and approval in one traceable operating path.*

## Next steps

* Use [Detect Anything (SAM 1–SAM 3)](/documentation/auto-labeling/detect-anything-sam-1-sam-3.md) to calibrate interactive and batch assistance.
* Use [Multimodal overview](/documentation/multimodal-annotations/multimodal-overview.md) when related files or views must stay in one task.
* Curate difficult cases and review cohorts in [Data curation](/documentation/data/data-curation.md).
* Read the current [audio annotation product overview](https://unitlab.ai/en/audio-annotation) for the feature overview and current media.
