TranscriptWord

Schema

Fields

· 5
word string required

A single word with punctuation if applicable

start number required

Time in original media when this word utterance has started.

Can be off by a few milliseconds sometimes due to media format and bitrate conversions for transcription. Original media with fixed bitrates tend to do better.

end number required

Time in original media when this word utterance has finished.

Can be off by a few milliseconds sometimes due to media format and bitrate conversions for transcription. Original media with fixed bitrates tend to do better.

confidence number required

Confidence of the model when detecting this word. Ranges between 0.0 to 1.0

speaker number required

Index of the speaker.

Our models estimate the total number of speakers in the given audio and gives each distinct voice an index starting 0 to total number of speakers - 1.

For example:
If we detect 3 distinct speakers in the audio. The speaker index can be any of 0, 1 or 2.

NOTE: Additional speakers may be added manually while editing the transcript in our dashboard.

Used by

· 2 operations