Media and document layouts
Combine image, video, audio, text, PDF, and structured evidence in one custom Data Group layout.

Choose the evidence model
Build the layout

Auto-Grouping

Cross-modal relations

Combine image, video, audio, text, PDF, and structured evidence in one custom Data Group layout.
A media-and-document task should preserve one business or research case even when its evidence uses different file types. Unitlab Data Groups bring native viewers into one custom layout so an annotator can create a label in the active tile while consulting the remaining evidence.

The group—not an individual file—defines the reviewable case.
Use this pattern for examples such as:
manufacturing video + inspection report + vibration audio;
image + OCR text + supporting PDF;
customer recording + transcript + case document;
field image + sensor export + technician note;
clinical image + report or study documentation.
Keep each source as its native file whenever possible. Flattening a PDF, waveform, or video into screenshots removes page, time, and content behavior that reviewers need.

Custom layouts control how a repeated case is presented to every annotator and reviewer.

Use Auto-Grouping when filenames or metadata consistently encode case membership. Configure the source folder, grouping keys, name pattern, tile rules, exclusions, and layout. Review unmatched and duplicate files before creating groups; an incorrect grouping rule creates a data-quality problem upstream of annotation.

Relations can connect labels across the ontology when the downstream dataset needs explicit evidence structure. Define direction and meaning before production—for example finding-supported-by-report or defect-correlates-with-audio-event. Do not assume proximity in a layout is an implicit relation.
Every group has the expected files and no unintended duplicates.
Tile roles and order are stable.
The active editor is visually unambiguous.
Page, frame, time, and source-file identity remain separate.
Reviewers can reconstruct why a label was created from the visible evidence.
Dataset and release outputs preserve group membership and relation semantics.