For the complete documentation index, see llms.txt. This page is also available as Markdown.

Audio Annotation

Label temporal events, speakers, transcripts, classifications, properties, and relations on synchronized waveform and spectrogram views.

Audio annotation is a timing decision as much as a semantic decision. Unitlab aligns waveform, spectrogram, playback, transcript context, ontology values, and a complete annotation timeline so teams can create precise speech and sound datasets without losing temporal evidence.

Use this guide when: you are building speech recognition, diarization, sound-event detection, intent, acoustic-scene, quality, or multimodal audio datasets.

See audio annotation in action

The demo shows audio as part of the multimodal Workbench, with time-based annotation and contextual review.

Open the demo in a new tab.

Before you begin

  1. Create or select a project whose data and ontology match this modality.

  2. Confirm the project instructions define the unit of annotation, boundary or timing policy, required properties, and review route.

  3. Open the project and enter the assigned item from the project data view or queue. The Workbench loads the modality-native editor inside the shared Unitlab shell.

See Annotation Workbench for navigation, saving, item state, comments, issues, and workflow actions.

Understand the audio work surface

The audio Workbench exposes playback, zoom, a shared playhead, waveform and spectrogram context, temporal ranges, ontology controls, transcript-related values, and the current workflow action. Zoom changes inspection precision; it does not change the underlying timestamp basis.

Synchronized waveform and spectrogram with one selected time range

Waveform and spectrogram provide complementary evidence around one authoritative temporal selection.

Supported annotation model

Annotation type
Use it for

Temporal event

A sound, action, acoustic state, or other interval with start and end time.

Speaker segment

A time range attributed to one speaker identity or role.

Transcription span

Text aligned to a clip or exact temporal interval.

Classification

Intent, sentiment, acoustic scene, quality, or another item-level label.

Dynamic property

A value that changes over the recording timeline.

Relation

A connection between speakers, events, transcripts, or contextual items.

Item Property

Language, source, channel, consent, quality, or other whole-recording context.

Write an explicit boundary policy for silence, overlap, clipped speech, background noise, non-speech events, and minimum-duration segments before production starts.

Time-aligned transcription review

Align transcript content with the intended speaker or temporal range. Check words at segment boundaries, overlapping speakers, false starts, filler words, unintelligible audio, and project-specific normalization rules. Keep source audio and transcript identity linked so corrections remain traceable.

Speaker transcript spans aligned to a waveform and playhead

Transcript content, speaker identity, and time range are reviewed together.

Complete audio timelines

Use the timeline to inspect event ranges, speaker turns, transcript coverage, dynamic properties, and quality segments across the complete recording. Long recordings require review at every transition, overlap, channel change, and abrupt acoustic event—not only at uniformly spaced samples.

Audio moments connected to event, speaker, transcript, and quality ranges

The timeline reveals missing coverage, accidental overlap, and boundary inconsistency.

Ontology-driven audio structure

Use classes for speakers and events, properties for structured details, relations for connections, and Item Properties for facts about the entire recording. AI-assisted proposals can accelerate transcription or event discovery, but confidence should route review rather than replace it.

Annotate one production item

1

1. Listen before labeling

Play enough surrounding audio to understand the event, speaker, language, and acoustic context. Confirm the project’s overlap and silence policy.

2

2. Zoom to the boundary

Use waveform and spectrogram detail to locate the exact start and end. Keep enough context visible to avoid cutting off leading or trailing signal.

3

3. Create the temporal annotation

Drag the intended range, assign the ontology class, and adjust the handles until the interval matches the policy.

4

4. Add transcript and structure

Enter or review time-aligned text, speaker identity, required properties, relations, and Item Properties.

5

5. Review the timeline

Check gaps, overlaps, duplicated events, inconsistent speaker identity, and boundary drift across the full recording.

6

6. Save and route

Resolve validation, save, and use the current workflow action to submit, approve, reject, or escalate.

Quality review

Review focus
What to check

Timing

Confirm start and end against waveform, spectrogram, playback, and project tolerance.

Overlap

Represent simultaneous speakers or sounds exactly as the ontology and instruction define.

Transcript

Apply the same casing, punctuation, filler, normalization, and unintelligible-audio policy.

Identity

Keep speaker or event identity stable throughout the recording.

Coverage

Review the complete timeline for missed intervals and accidental gaps.

Move from labels to governed data

Audio outputs should preserve the source recording identity, time basis, segment boundaries, speaker or event identity, transcript text, properties, relations, and item-level context. Keep long-recording curation, review cohorts, dataset versions, and releases traceable to the original media.

Integrated Unitlab workflow connecting model assistance, annotation, review, and quality assurance

Use workflows to keep model output, human correction, review, and approval in one traceable operating path.

Next steps