Guide
Transcription Guidelines — the seven core standards
A reference doc covering Cognegica's seven transcription standards: native-linguist transcribers, orthographic accuracy, timestamping, diarization, non-speech events, dialect handling and formatting.
This reference documents the seven core standards behind every Cognegica transcription delivery. They exist so that transcripts are accurate, consistent and machine-readable across languages and scripts.
- Native-linguist transcribers for each language and script.
- Orthographic accuracy to each script's conventions.
- Timestamping and segmentation at a defined granularity.
- Speaker identification and diarization across multi-speaker recordings.
- Non-speech event annotation: [noise], [laughter], [music], [overlapping speech].
- Dialect and accent considerations handled by native linguists.
- Data formatting and consistency: UTF-8, standardized punctuation, consistent spacing and project annotation tags.
Together these feed the multi-layer QA framework and downstream training and evaluation.
Using these standards
Questions about the transcription standards
- Why native-linguist transcribers?
Orthography, dialect and accent edge cases need native judgement that generic typists can't provide. It's the first of the seven standards for that reason.
- How are non-speech events handled?
With a fixed vocabulary — [noise], [laughter], [music], [overlapping speech] — applied consistently so downstream models can account for them.
- What formatting do deliveries use?
UTF-8 with standardized punctuation, consistent spacing and project annotation tags, applied uniformly across the delivery.