3) Quality of OCR. This varies from corpus to corpus as described above. For English, we spent a great deal of time examining the data by hand as an additional check on its reliability. The other corpora may not be as reliable. 4) Quality of Metadata. Again, the English language...
Results for “Corpus Christi”
Search across the indexed text of every released document.
Names that match “Corpus Christi”
32 documents found
… whom correspondence should be addressed. E-mail: [email protected] (J.B.M.); [email protected] (E.A.). We constructed a corpus of digitized texts containing about 4% of all books ever printed. Analysis of this corpus enables us to investigate cultural trends quantitatively. We su...
…put in the paper. In particular, the primary object of study in this paper is the English language from 1800-2000; this corpus during this period is therefore the most carefully curated of the datasets. However, to encourage further research, we are releasing all available datase...
000 barrels a day moving through the pipe is loaded onto barges at Corpus Christi and towed toward
…tates of America” (a 5-gram). We restricted n to 5, and limited our study to n-grams occurring at least 40 times in the corpus. Usage frequency is computed by dividing the number of instances of the n-gram in a given year by the total number of words in the corpus in that year....
…g Fields 1D, 16,24 2B,C i. Counting Historical N-gram a Cleaned li Construction Gre a Source N-gram ; Aggregation 3A Corpus by Date 3B 3A, Corpus Fig. $3. Outline of n-gram corpus construction. The numbering corresponds to sections of the text. HOUSE_OVERSIGHT_017042
…ooks into ‘base corpora’ using such metadata fields as language, country of publication, and subject. 3. For each base corpus, construct a massive numerical table that lists, for each n-gram (often a word or phrase), how often it appears in the given base corpus in every single...
how often it appears in the given base corpus in every single year between 1550
the likelihood of error in a randomly sampled book from the final corpus is
while retaining the vast majority of the books in the original corpus.
the English language corpus was checked very carefully and
we note that earlier portions of the Hebrew corpus contain a large
we report that our corpus contains about 4% of all books ever published. Obtaining this
then our 5 million book corpus encompasses a little more than 5% of all books ever
then our corpus would constitute a little less than 3%. We report an