periodicals such as newspapers). This final database contains bibliographic information for each of
Epstein Suite indexes the text; the original document lives at its official source. We don't host the original file — view it on the official release to read it in full.
View the original on the official releaseDocument text
Text is machine OCR and may contain errors. Confirm against the original source above.
periodicals such as newspapers). This final database contains bibliographic information for each of these
129 million editions (Ref. $1). The country of publication is known for 85.3% of these editions, authors for
87.8%, publication dates for 92.6%, and the language for 91.6%. Of the 15 million books scanned, the
country of publication is known for 91.5%, authors for 92.1%, publication dates for 95.1%, and the
language for 98.6%.
I.2. Digitization
We describe the way books are scanned and digitized. For publisher-provided books, Google removes
the spines and scans the pages with industrial sheet-fed scanners. For library-provided books, Google
uses custom-built scanning stations designed to impose only as much wear on the book as would result
from someone reading the book. As the pages are turned, stereo cameras overhead photograph each
page, as shown in Figure $1.
One crucial difference between sheet-fed scanners and the stereo scanning process is the flatness of the
page as the image is captured. In sheet-fed scanning, the page is kept flat, similar to conventional flatbed
scanners. With stereo scanning, the book is cradled at an angle that minimizes stress on the spine of the
book (this angle is not shown in Figure $1). Though less damaging to the book, a disadvantage of the
latter approach is that it results in a page that is curved relative to the plane of the camera. The curvature
changes every time a page is turned, for several reasons: the attachment point of the page in the spine
differs, the two stacks of pages change in thickness, and the tension with which the book is held open
may vary. Thicker books have more page curvature and more variation in curvature.
This curvature is measured by projecting a fixed infrared pattern onto each page of the book,
subsequently captured by cameras. When the image is later processed, this pattern is used to identify the
location of the spine and to determine the curvature of the page. Using this curvature information, the
scanned image of each page is digitally resampled so that the results correspond as closely as possible
to the results of sheet-fed scanning. The raw images are also digitally cropped, cleaned, and contrast
enhanced. Blurred pages are automatically detected and rescanned. Details of this approach can be
found in U.S. Patents 7463772 and 7508978; sample results are shown in Figure $2.
Finally, blocks of text are identified and optical character recognition (OCR) is used to convert those
images into digital characters and words, in an approach described elsewhere (Ref. $2). The difficulty of
applying conventional OCR techniques to Google’s scanning effort is compounded because of variations
in language, font, size, paper quality, and the physical condition of the books being scanned.
Nevertheless, Google estimates that over 98% of words are correctly digitized for modern English books.
After OCR, initial and trailing punctuation is stripped and word fragments split by hyphens are joined,
yielding a stream of words suitable for subsequent indexing.
1.3. Structure Extraction
After the book has been scanned and digitized, the components of the scanned material are classified
into various types. For instance, individual pages are scanned in order to identify which pages comprise
the authored content of the book, as opposed to the pages which comprise frontmatter and backmatter,
such as copyright pages, tables of contents, index pages, etc. Within each page, we also identify
repeated structural elements, such as headers, footers, and page numbers.
Using OCR results from the frontmatter and backmatter, we automatically extract author names, titles,
ISBNs, and other identifying information. This information is used to confirm that the correct consensus
record has been associated with the scanned text.
HOUSE_OVERSIGHT_017012
Have a question about what this document contains?
Ask the documents