# Choose a corpus for Romanian texts

Start with domain and the text form you need: segments, documents or annotations.

## 1. Choose translation segments

The translation memory provides aligned units. Evaluate a sample before use and retain edition and language pair.

[DGT-TM translation memory](https://mariuscomper.uk/cunoastere-deschisa/en/resurse/dgt-tm/)

## 2. Retain document context

When structure and paragraphs matter, consult the document corpus. Check overlap with other Acquis collections before splitting training and evaluation sets.

[DGT-Acquis — aligned documents and paragraphs](https://mariuscomper.uk/cunoastere-deschisa/en/resurse/dgt-acquis/)

## 3. Choose annotations for language analysis

For syntax, consult annotation conventions and corpus version. Legislative texts do not represent every register of Romanian.

[UD Romanian RRT](https://mariuscomper.uk/cunoastere-deschisa/en/resurse/ud-rrt/)

## Caution

Duplicates across collections can distort model evaluation. One collection’s licence does not establish rights for every other collection.

A corpus selection with domain, version, language, terms and an overlap-checking procedure.
