Its sources are gathered under selection and preprocessing rules that define the resulting distribution. Documentation should record coverage, time range, provenance, permissions, and known gaps.
A corpus is a collected body of text, speech, images, or other material used for analysis or model development.
Its sources are gathered under selection and preprocessing rules that define the resulting distribution. Documentation should record coverage, time range, provenance, permissions, and known gaps.