Author
Simpson, Heather, author.
Publication Date
2014
Publication Place
-
[Philadelphia, Pa.] Linguistic Data Consortium, [2014].©2014
Subject
Text files -- Data processing.
XML (Document markup language) -- Databases.
Data warehousing.
Information retrieval -- Databases.
Data warehousing.
Information retrieval.
XML (Document markup language)
Databases.
Type
Other
Language
ara,eng
Digital
Yes
Manuscript
No
Physical Dimensions
|
Library
University of Chicago
Record ID
10373231
Date
2014
Notes
LDC2014T16Title from disc label.Access restricted to University of Sydney staff and students. Educational use only.System requirements: CD-ROM drive; web browser; other requirements not specified.English.. 1 computer disc (CD-ROM) : sound, colour ; 12 cm.. "TAC KBP Reference Knowledge Base was developed by the Linguistic Data Consortium (LDC) in support of the NIST-sponsored TAC-KBP evaluation series. It is a knowledge base built from English Wikipedia articles and their associated infoboxes and covers over 800,000 entities. TAC (Text Analysis Conference) is a series of workshops organized by NIST (the National Institute of Standards and Technology) to encourage research in natural language processing and related applications by providing a large test collection, common evaluation procedures, and a forum for researchers to share their results. TAC's KBP track (Knowledge Base Population) encourages the development of systems that can match entities mentioned in natural texts with those appearing in a knowledge base and extract novel information about entities from a document collection and add it to a new or existing knowledge base. Consult the LDC TAC-KBP project page for further information about LDC's resource development for the TAC-KBP program. The source data (Wikipedia infoboxes and articles) was taken from an October 2008 snapshot of Wikipedia. TAC KBP Reference Knowledge Base contains a set of entities, each with a canonical name and title for the Wikipedia page, an entity type, an automatically parsed version of the data from the infobox in the entity's Wikipedia article, and a stripped version of the text of the Wiki article. Each entity is assigned one of four types: PER (person), ORG (organization), GPE (geo-political entity) and UKN (unknown). All data files are presented as UTF-8 encoded XML. "-- LDC catalogue https://catalog.ldc.upenn.edu/LDC2014T16 (viewed October 30, 2014).
Diğer yazarlar / katkıda bulunanlar
Ellis, Joe, author.
Parker, Robert, author.
Linguistic Data Consortium.
ISBN
15856368519781585636853
Ses özellikleri
digital optical stereo
Video özellikleri
laser optical
Dijital dosya özellikleri
program file