Open knowledge
Romania
Practical guide

Choose a corpus for Romanian texts

Start with domain and the text form you need: segments, documents or annotations.

  1. Choose translation segments

    The translation memory provides aligned units. Evaluate a sample before use and retain edition and language pair.

    DGT-TM translation memory ↗
  2. Retain document context

    When structure and paragraphs matter, consult the document corpus. Check overlap with other Acquis collections before splitting training and evaluation sets.

    DGT-Acquis — aligned documents and paragraphs ↗
  3. Choose annotations for language analysis

    For syntax, consult annotation conventions and corpus version. Legislative texts do not represent every register of Romanian.

    UD Romanian RRT ↗

A pitfall to avoid

Duplicates across collections can distort model evaluation. One collection’s licence does not establish rights for every other collection.

What you can produce

A corpus selection with domain, version, language, terms and an overlap-checking procedure.

Browse this topic →

Build a dossier with the guide’s sources →

Top ↑