- Tytuł:
- Elektroniczny Korpus Tekstów Polskich z XVII i XVIII w. – problemy teoretyczne i warsztatowe
- Autorzy:
-
Gruszczyński, Włodzimierz
Adamiec, Dorota
Bronikowska, Renata
Wieczorek, Aleksandra - Powiązania:
- https://bibliotekanauki.pl/articles/1630441.pdf
- Data publikacji:
- 2020
- Wydawca:
- Towarzystwo Kultury Języka
- Tematy:
-
electronic text corpus
historical corpus
17th-18th-century Polish
natural language processing - Opis:
- This paper presents the Electronic Corpus of 17th- and 18th-century Polish Texts (KorBa) – a large (13.5-million), annotated historical corpus available online. Its creation was modelled on the assumptions of the National Corpus of Polish (NKJP), yet the specifi c nature of the historical material enforced certain modifi cations of the solutions applied in NKJP, e.g. two forms of text representation (transliteration and transcription) were introduced, the principle of designating foreign-language fragments was adopted, and the tagset was adapted to the description of the grammatical structure of the Middle Polish language. The texts collected in KorBa are diversified in chronological, geographical, stylistic, and thematic terms although, due to e.g. limited access to the material, the postulate of representativeness and sustainability of the corpus was not fully implemented. The work on the corpus was to a large extent automated as a result of using natural language processing tools.
- Źródło:
-
Poradnik Językowy; 2020, 777, 8; 32-51
0551-5343 - Pojawia się w:
- Poradnik Językowy
- Dostawca treści:
- Biblioteka Nauki