Modelos para a identificação de tópicos em notícias extraídas da web e filtradas por qualidade de dados
Carregando...
Arquivos
Data
Título da Revista
ISSN da Revista
Título de Volume
Editor
Universidade Federal do Rio de Janeiro
DOI
Resumo
The increase in the number of users online, connectivity rates and mobile devices has created a new dynamic for the publication and dissemination of online content. Due to different publication formats, the capture of information by automatic agents is a complex task and prone to the introduction of errors in the extracted content. In this way, the search for relevant information becomes more complex, which leads to the development of text processing techniques that are capable of extracting relevant information from a large volume of documents with possible data quality issues. Topic Modeling is a set of techniques that aims to summarize, explore and cat- egorize a set of documents using unsupervised learning. As challenges in the area are the interpretability of the clusters, as well as the choice of the best number of topics. This work evaluates the use of coherence and stability measures for choosing the number of topics, in order to guarantee the groups show semantic coherence, in non- annotated databases, with news with data quality issues.To this end, data quality dimensions and criteria are defined to be met by the documents, and coherence and stability measures are evaluated for different levels of noise. As a result, filtering the news using data quality criteria increased the consistency of topic extraction, while the measure of coherence and stability helped to narrow the range of choice for the number of topics. However, a way of combining the number of topics, stability and coherence to choose between a more generalist or more detailed extraction has not yet been found.
Descrição
Palavras-chave
Citação
Coleções
Avaliação
Revisão
Suplementado Por
Referenciado Por
Direitos e licensiamento
Acesso Aberto