Practical guide
Choose a corpus for Romanian texts
Start with domain and the text form you need: segments, documents or annotations.
Choose translation segments
The translation memory provides aligned units. Evaluate a sample before use and retain edition and language pair.
DGT-TM translation memory ↗Retain document context
When structure and paragraphs matter, consult the document corpus. Check overlap with other Acquis collections before splitting training and evaluation sets.
DGT-Acquis — aligned documents and paragraphs ↗Choose annotations for language analysis
For syntax, consult annotation conventions and corpus version. Legislative texts do not represent every register of Romanian.
UD Romanian RRT ↗
A pitfall to avoid
Duplicates across collections can distort model evaluation. One collection’s licence does not establish rights for every other collection.
What you can produce
A corpus selection with domain, version, language, terms and an overlap-checking procedure.
Browse this topic →