A READING ROOM IN PROGRESS

Finding the stories within the page.

Old newspapers preserve more than the news. Advertisements, community notices and everyday announcements offer a view of the places and people who made them. Ink makes those pages easier to explore.

Start with the original

Ink builds on the preservation and digitization work of libraries, publishers and archival partners. Readers can browse an issue, inspect individual clippings and compare the supplied machine-readable text with the original scan. Each issue retains its source attribution.

The scan matters because machine transcription is imperfect. A faint letter may be misread, a heading overlooked, or unrelated text joined together. Keeping the image within reach lets readers check what was actually printed.

A clipping is a beginning

A page region is not necessarily a complete article. A headline, subheadline and body may appear as separate clippings; a story may continue in another column or on another page. Some collections already provide complete articles, while others remain a collection of source regions.

The gallery makes this material accessible while that work continues. Tiny fragments are hidden by default to make browsing clearer, with an option to show every block. The underlying source records are preserved.

Try a page

Select a clipping, inspect its position on the scan, then open “Read OCR” to compare the text with the print.

What we hope to improve

We are exploring better ways to recover difficult text and connect the pieces that belong together. Advertisements and small local notices deserve attention alongside editorial articles. Their language, appearance and recurrence can support questions about community life and change over time.

These are research directions, not claims of completed results. We want to evaluate improvements on a manageable set of newspapers with people who understand and use the material, before expanding further.

Evidence, credit and uncertainty

Ink is informed by work such as Harvard’s American Stories project, which brings together newspaper layout analysis, text recognition and article association. Our current examples draw on several sources; they are not all outputs of Harvard’s models.

Coverage and transcription quality vary. A browsable collection does not imply complete newspaper coverage, verified article boundaries or error-free text. Our aim is to make useful material accessible while keeping its source and uncertainties visible.

Read about the collections and their sources