Browsing by Subject "text processing"
Now showing 1 - 5 of 5
- Results Per Page
- Sort Options
Item type:Thesis, Access status: Restricted , Klasyfikacja danych tekstowych przy pomocy rekurencyjnych sieci neuronowych(Data obrony: 2018-01-23) Łyko, Tomasz
Wydział Informatyki, Elektroniki i TelekomunikacjiItem type:Thesis, Access status: Restricted , Przetwarzanie danych tekstowych na platformie Apache Spark(Data obrony: 2018-09-21) Grygiel, Łukasz
Wydział Elektrotechniki, Automatyki, Informatyki i Inżynierii BiomedycznejItem type:Article, Access status: Open Access , Recognizing non-translatable symbols in a multi-lingual computer-assisted translation system for DTP documents(Wydawnictwa AGH, 2010) Grabowski, Szymon; Draus, Cezary; Bieniecki, WojciechThe paper is devoted to the problem of computer-assisted translation of catalogues and advertising brochures (DTP documents), where the text to translate consists of many short separated snippets. One of the issues that may facilitate the translation proces is to recognize the phrases which should be copied verbatim, no matter what the target language is. These include technical data with units of measurement, abbreviations, numbers etc. but also trademark symbols. As for the first problem, the presented algorithm uses statistical analysis of the characters inside a character sequence within each segment of the considered phrase, where segments boundaries are marked by special characters, like hyphens or slashes. If at least one of the segmented is labeled »non-symbol«, the whole phrase should be handled by the human translator, otherwise it is considered non-translatable and copied verbatim, hence saving the translator's work. For the trademark start boundary recognition problem, we proposed a simple but seemingly robust solution based on similarity of word suffixes preceding ® and similar characters in a given phrase, together with heuristic rules based on character case of those words.Item type:Article, Access status: Open Access , Web pages content analysis using browser-based volunteer computing(Wydawnictwa AGH, 2013) Turek, Wojciech; Nawarecki, Edward; Dobrowolski, Grzegorz; Krupa, Tomasz; Majewski, PrzemysławExisting solutions to the problem of finding valuable information on the Web suffers from several limitations like simplified query languages, out-of-date in- formation or arbitrary results sorting. In this paper a different approach to this problem is described. It is based on the idea of distributed processing of Web pages content. To provide sufficient performance, the idea of browser-based volunteer computing is utilized, which requires the implementation of text processing algorithms in JavaScript. In this paper the architecture of Web pages content analysis system is presented, details concerning the implementation of the system and the text processing algorithms are described and test results are provided.Item type:Thesis, Access status: Restricted , Wykrywanie anomalii w dużych zbiorach danych za pomocą sieci neuronowych(Data obrony: 2018-01-23) Piwowarczyk, Wojciech
Wydział Informatyki, Elektroniki i Telekomunikacji
