Explorations in Automatic Thesaurus Discovery

Explorations in Automatic Thesaurus Discovery
Author :
Publisher : Springer Science & Business Media
Total Pages : 313
Release :
ISBN-10 : 9781461527107
ISBN-13 : 1461527104
Rating : 4/5 (104 Downloads)

Book Synopsis Explorations in Automatic Thesaurus Discovery by : Gregory Grefenstette

Download or read book Explorations in Automatic Thesaurus Discovery written by Gregory Grefenstette and published by Springer Science & Business Media. This book was released on 2012-12-06 with total page 313 pages. Available in PDF, EPUB and Kindle. Book excerpt: Explorations in Automatic Thesaurus Discovery presents an automated method for creating a first-draft thesaurus from raw text. It describes natural processing steps of tokenization, surface syntactic analysis, and syntactic attribute extraction. From these attributes, word and term similarity is calculated and a thesaurus is created showing important common terms and their relation to each other, common verb--noun pairings, common expressions, and word family members. The techniques are tested on twenty different corpora ranging from baseball newsgroups, assassination archives, medical X-ray reports, abstracts on AIDS, to encyclopedia articles on animals, even on the text of the book itself. The corpora range from 40,000 to 6 million characters of text, and results are presented for each in the Appendix. The methods described in the book have undergone extensive evaluation. Their time and space complexity are shown to be modest. The results are shown to converge to a stable state as the corpus grows. The similarities calculated are compared to those produced by psychological testing. A method of evaluation using Artificial Synonyms is tested. Gold Standards evaluation show that techniques significantly outperform non-linguistic-based techniques for the most important words in corpora. Explorations in Automatic Thesaurus Discovery includes applications to the fields of information retrieval using established testbeds, existing thesaural enrichment, semantic analysis. Also included are applications showing how to create, implement, and test a first-draft thesaurus.

Explorations in Automatic Thesaurus Discovery Related Books

Explorations in Automatic Thesaurus Discovery
Language: en
Pages: 313
Authors: Gregory Grefenstette
Categories: Computers
Type: BOOK - Published: 2012-12-06 - Publisher: Springer Science & Business Media

GET EBOOK

Explorations in Automatic Thesaurus Discovery presents an automated method for creating a first-draft thesaurus from raw text. It describes natural processing s
Language: en
Pages: 4947
Authors:
Categories:
Type: BOOK - Published: - Publisher: IOS Press

GET EBOOK

Research and Development in Intelligent Systems XXIII
Language: en
Pages: 421
Authors: Frans Coenen
Categories: Computers
Type: BOOK - Published: 2010-05-30 - Publisher: Springer Science & Business Media

GET EBOOK

The papers in this volume are the refereed technical papers presented at AI-2006, the Twenty-sixth SGAI International Conference on Innovative Techniques and Ap
Foundations of Statistical Inference
Language: en
Pages: 227
Authors: Yoel Haitovsky
Categories: Mathematics
Type: BOOK - Published: 2012-12-06 - Publisher: Springer Science & Business Media

GET EBOOK

This volume is a collection of papers presented at a conference held in Shoresh Holiday Resort near Jerusalem, Israel, in December 2000 organized by the Israeli
Machine Learning and Data Mining in Pattern Recognition
Language: en
Pages: 709
Authors: Petra Perner
Categories: Computers
Type: BOOK - Published: 2005-08-25 - Publisher: Springer

GET EBOOK

We met again in front of the statue of Gottfried Wilhelm von Leibniz in the city of Leipzig. Leibniz, a famous son of Leipzig, planned automatic logical inferen