Repository logo
Article

Ekstrakcja spójnych tekstów z Internetu na potrzeby algorytmów lingwistycznych

Loading...
Thumbnail Image

Date

Presentation Date

Editor

Other contributors

Access rights

Access: otwarty dostęp
Rights: AGH Licence
AGH Licence - Fair Use

Licencja AGH - Fair use of copyrighted works

Other title

Extraction of coherent text from the Internet for use in natural language processing

Resource type

Version

wersja wydawnicza
Item type:Journal Issue,
Automatyka
2008 - T. 12 - Nr 2

Pagination/Pages:

s. 423-431

Research Project

Event

Description

Abstract

Computer Linguistic is aimed to develop and improve text information extraction methods. Internet becomes a very extensive source of text, yet it is overloaded by thematically incoherent texts grouped by one presentation context (e.g. WWW page). This fact determines difficulties with usage of such texts as text corpuses for NLP processing (especially statistics based algorithms). Presented work is aimed to develop methods of extraction coherent texts from Web pages, that can improve quality of information extraction.


Lingwistyka komputerowa dąży do wytworzenia coraz lepszych algorytmów ekstrakcji informacji z tekstu. Bardzo obszernym źródłem tekstu jest obecnie Internet. Jest on jednak przeładowany informacjami nie skojarzonymi ze sobą tematycznie, a pojawiającymi się w jednym kontekście (np. na jednej stronie WWW). Powoduje to duże trudności w użyciu tych tekstów jako korpusów tekstu do przetwarzania lingwistycznego (szczególnie dla metod statystycznych). Celem stworzenia prezentowanych algorytmów była próba ekstrahowania tekstów spójnych tematycznie ze stron WWW, tak by teksty te mogły stanowić dobry korpus dla prac nad ekstrakcją informacji.

Access rights

Access: otwarty dostęp
Rights: AGH Licence
AGH Licence - Fair Use

Licencja AGH - Fair use of copyrighted works