Existing collections
A ready-to-license collection of 500 varied French-language PDFs is immediately available, representing more than 140 million words.
French-language data for AI, OCR & research
We provide digitized French books and custom historical corpora for AI companies, LLM teams, OCR developers, NLP researchers and digital humanities projects.
Data built around your use case
Historical books contain specialist vocabulary, long-form reasoning, technical terminology, period language, structured knowledge and typography that are often underrepresented in modern web corpora.
A ready-to-license collection of 500 varied French-language PDFs is immediately available, representing more than 140 million words.
Tell us the subject, period, volume and format you need. We can draw on 30,000 books in our own stock and source from a wider pool of approximately 600,000 books.
We can digitize suitable physical books specifically for AI, OCR, document understanding or research workflows at up to 600 dpi, in grayscale or black & white.
Why our scans are different
Many historical collections are digitized non-destructively with the book held open. This can introduce gutter curvature, shadows, warped text and geometric distortion. Our workflow is designed for image quality first.
When appropriate, books are carefully disbound so each page can be scanned flat. This significantly reduces gutter curvature, page shadow and text deformation compared with open-book scanning.
Pages are digitized at 600 dpi in grayscale or black and white, producing high-quality source images suitable for OCR, document AI, layout analysis and multimodal training.
A significant part of our source material consists of damaged, incomplete, mismatched or otherwise unsellable books acquired from private libraries — volumes that might otherwise be discarded or pulped.
Preservation through publishing
When an original work is rare, historically significant or worth preserving as a physical publication, we may create a small print re-edition alongside the digital archive. In this way, digitization can help preserve the work and return it to circulation.
Flexible output
Every project is different. Datasets can be delivered as raw digitizations or as richer, structured packages designed for downstream processing.
Complete digital reproductions preserving original layout, illustrations, tables and typography. Our standard digitization workflow supports 600 dpi grayscale or black & white scanning.
Individual page files suitable for OCR training, computer vision and multimodal document models.
Machine-readable text for selected collections, with quality depending on source material and project scope.
Author, title, date, publisher, place of publication, subject, edition, language and other available fields.
Subject coverage
Why work with us
Our background combines antiquarian books, publishing and digitization. That means we understand the source material before it becomes a file.
We can help identify relevant editions, historical periods, subject areas and suitable source material, then prepare the collection in a format compatible with your technical workflow. Our ability to work from physical books also gives us more control over scan quality, provenance and source selection than a workflow based only on third-party digital libraries.
Use cases
French and historical long-form text.
Scanned pages, typography and difficult layouts.
Multimodal understanding of printed material.
Diachronic language and specialist vocabulary.
Domain-specific historical knowledge bases.
Evaluation sets for OCR and language models.
Structured corpora for scholarly analysis.
Entities, topics and semantic indexing.
Responsible sourcing
When we acquire private libraries, part of the collection is often unsuitable for the antiquarian or second-hand market: damaged books, incomplete sets, isolated volumes or copies in poor condition.
Instead of allowing this material to be discarded, we can transform it into high-quality digital source data. We do not position valuable collectible copies as disposable raw material: the workflow is designed to recover knowledge from books that are otherwise difficult or impossible to sell in their current state.
Rights & licensing
The legal status of historical works varies by author, publication date, territory and intended use. We can prioritize public-domain material and define licensing and permitted uses for each commercial dataset. Final legal assessment remains project-specific.
Request a dataset
Send us your target subject, historical period, approximate number of books or pages, preferred formats, OCR requirements, metadata needs and any image-quality constraints.
We can propose our immediately available 500-PDF collection or prepare a custom digitization project from our own stock and wider sourcing network.