Full-text search

Results for “Corpus Christi”

Search across the indexed text of every released document.

Names that match “Corpus Christi”

32 documents found

The computations required to generate these corpora were performed at Google using the MapReduce
document IMAGES-004-HOUSE_OVERSIGHT_017020.txt House Oversight Committee — Epstein Estate Records (Nov 2025)

…put in the paper. In particular, the primary object of study in this paper is the English language from 1800-2000; this corpus during this period is therefore the most carefully curated of the datasets. However, to encourage further research, we are releasing all available datase...

(“3.14159”) and typos (“excesss”). An n-gram is sequence of
document IMAGES-004-HOUSE_OVERSIGHT_016997.txt House Oversight Committee — Epstein Estate Records (Nov 2025)

…tates of America” (a 5-gram). We restricted n to 5, and limited our study to n-grams occurring at least 40 times in the corpus. Usage frequency is computed by dividing the number of instances of the n-gram in a given year by the total number of words in the corpus in that year....

Language Selection
document IMAGES-004-HOUSE_OVERSIGHT_017042.txt House Oversight Committee — Epstein Estate Records (Nov 2025)

…g Fields 1D, 16,24 2B,C i. Counting Historical N-gram a Cleaned li Construction Gre a Source N-gram ; Aggregation 3A Corpus by Date 3B 3A, Corpus Fig. $3. Outline of n-gram corpus construction. The numbering corresponds to sections of the text. HOUSE_OVERSIGHT_017042

II. Construction of Historical N-grams Corpora
document IMAGES-004-HOUSE_OVERSIGHT_017013.txt House Oversight Committee — Epstein Estate Records (Nov 2025)

…ooks into ‘base corpora’ using such metadata fields as language, country of publication, and subject. 3. For each base corpus, construct a massive numerical table that lists, for each n-gram (often a word or phrase), how often it appears in the given base corpus in every single...