Publication Place
-
[Philadelphia, PA] : Linguistic Data Consortium, c2012.
Subject
Arabic language -- Written Arabic -- Data processing.
Arabic language -- Machine translating.
Arabic language -- Translating into English.
Machine translating.
Machine translating.
Arabic language -- Translating into English.
Arabic language -- Machine translating.
Type
Other
Language
Arabic
Digital
Yes
Manuscript
No
Physical Dimensions
|
Library
University of Chicago
Record ID
8926012
Notes
Title from disc label.Data type: Text.Data sources: Newsgroups, newswire, weblogs.Applications: Handwriting recognition, machine translation."LDC2012T15".Authors: David Lee, Safa Ismael, Stephen Grimes, Dave Doermann, Stephanie Strassel, Zhiyi Song.Arabic. 2 DVD ; 4 3/4 in.. "MADCAT (Multilingual Automatic Document Classification Analysis and Translation) phase 1 training set contains all training data created by the Linguistic Data Consortium (LDC) to support Phase 1 of the DARPA MADCAT Program. The material in this release consists of handwritten Arabic documents, scanned at high resolution and annotated for the physical coordinates of each line and token. Digital transcripts and English translations of each document are also provided, with the various content and annotation layers integrated in a single MADCAT XML output. The goal of the MADCAT program is to automatically convert foreign text images into English transcripts." -- LDC online catalogue.
Başlığın Farklı Biçimleri
Title in LDC online catalogue: MADCAT phase 1 training set
Diğer yazarlar / katkıda bulunanlar
Lee, David.
Linguistic Data Consortium.
ISBN
15856362319781585636235