Repository logo
Article

Ekstrakcja spójnych tekstów z Internetu na potrzeby algorytmów lingwistycznych

creativeworkseries.issn1429-3447
dc.contributor.authorDorosz, Krzysztof
dc.date.available2017-08-23T09:33:47Z
dc.date.issued2008
dc.description.abstractComputer Linguistic is aimed to develop and improve text information extraction methods. Internet becomes a very extensive source of text, yet it is overloaded by thematically incoherent texts grouped by one presentation context (e.g. WWW page). This fact determines difficulties with usage of such texts as text corpuses for NLP processing (especially statistics based algorithms). Presented work is aimed to develop methods of extraction coherent texts from Web pages, that can improve quality of information extraction.en
dc.description.abstractLingwistyka komputerowa dąży do wytworzenia coraz lepszych algorytmów ekstrakcji informacji z tekstu. Bardzo obszernym źródłem tekstu jest obecnie Internet. Jest on jednak przeładowany informacjami nie skojarzonymi ze sobą tematycznie, a pojawiającymi się w jednym kontekście (np. na jednej stronie WWW). Powoduje to duże trudności w użyciu tych tekstów jako korpusów tekstu do przetwarzania lingwistycznego (szczególnie dla metod statystycznych). Celem stworzenia prezentowanych algorytmów była próba ekstrahowania tekstów spójnych tematycznie ze stron WWW, tak by teksty te mogły stanowić dobry korpus dla prac nad ekstrakcją informacji.pl
dc.description.placeOfPublicationKraków
dc.description.versionwersja wydawnicza
dc.identifier.eissn2353-0952
dc.identifier.issn1429-3447
dc.identifier.nukatdd2009317037
dc.identifier.urihttps://repo.agh.edu.pl/handle/AGH/46011
dc.language.isopol
dc.publisherWydawnictwa AGH
dc.relation.ispartofAutomatyka
dc.rightsAGH Licence - Fair Use
dc.rights.accessotwarty dostęp
dc.rights.urihttps://repo.uci.agh.edu.pl/info/licence-agh
dc.subjecttext extractionen
dc.subjectInterneten
dc.subjecttexten
dc.subjectekstrakcja tekstówpl
dc.subjectInternetpl
dc.subjectDOMen
dc.subjectspójność tekstupl
dc.subjectDOMpl
dc.subjectHTMLen
dc.subjectHTMLpl
dc.titleEkstrakcja spójnych tekstów z Internetu na potrzeby algorytmów lingwistycznychpl
dc.title.alternativeExtraction of coherent text from the Internet for use in natural language processingen
dc.title.relatedAutomatyka
dc.typeartykuł
dspace.entity.typePublication
publicationissue.issueNumberZ. 2
publicationissue.paginations. 423-431
publicationvolume.volumeNumberT. 12
relation.isJournalIssueOfPublicationbaa86304-7559-4325-93e9-65fd4ccdc169
relation.isJournalIssueOfPublication.latestForDiscoverybaa86304-7559-4325-93e9-65fd4ccdc169
relation.isJournalOfPublicationb16a3604-d334-41d9-9446-dfef1368171d

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Auto20.pdf
Size:
631.03 KB
Format:
Adobe Portable Document Format
Description:
Artykuł z czasopisma