GetItWebbed
0%
Historical Newspaper OCR Dataset
GetItWebbed
dat-019

Historical Newspaper OCR Dataset

Text strings mapped to bounding boxes on images — Download scans from library archives; extract text with Tesseract; manually correct errors. Stored as JSON / XML / PNG.

OCRPythonText DataResearch

What you get

  • Source Code
  • Documentation
  • PPT
  • Dataset Sample

Technical Details

Data Type

Text strings mapped to bounding boxes on images

Hardware / Tools

Python + Tesseract OCR + Public-domain newspaper scans

Collection Method

Download scans from library archives; extract text with Tesseract; manually correct errors

Storage Format

JSON / XML / PNG

How it works.

Common questions about ordering, delivery, and support.

More from Data Collection.

View all
Academic Paper Abstract Database
View Project
GetItWebbed
dat-008

Academic Paper Abstract Database

Scientific text (titles, abstracts, authors, publication dates) — Query APIs with keywords to fetch research papers and map publication trends. Stored as JSON / SQLite.

APIPythonResearchText Data
View Details
Subreddit Sentiment Archive
View Project
GetItWebbed
dat-006

Subreddit Sentiment Archive

Unstructured text (social media posts and comments) — Pull daily posts/comments from finance, tech, or pop-culture subreddits via Reddit API. Stored as JSON / CSV.

NLPPythonAPIText Data
View Details
Customer Review Aggregator
View Project
GetItWebbed
dat-007

Customer Review Aggregator

Text reviews + metadata (ratings, dates, verified status) — Extract customer reviews from Amazon, Yelp, or Trustpilot using browser automation. Stored as CSV / JSON.

Web ScrapingPythonNLPText Data
View Details

dat-019

Historical Newspaper OCR Dataset

Get This Project