French-language data for AI, OCR & research

Historical French books, prepared for machine learning.

We provide digitized French books and custom historical corpora for AI companies, LLM teams, OCR developers, NLP researchers and digital humanities projects.

500 PDFs / 140M+ words immediately available 600 dpi digitization 30,000 books in stock

Data built around your use case

French book datasets for models that need more than web text.

Historical books contain specialist vocabulary, long-form reasoning, technical terminology, period language, structured knowledge and typography that are often underrepresented in modern web corpora.

01

Existing collections

A ready-to-license collection of 500 varied French-language PDFs is immediately available, representing more than 140 million words.

02

Custom corpora

Tell us the subject, period, volume and format you need. We can draw on 30,000 books in our own stock and source from a wider pool of approximately 600,000 books.

03

Digitization projects

We can digitize suitable physical books specifically for AI, OCR, document understanding or research workflows at up to 600 dpi, in grayscale or black & white.

Why our scans are different

Flat pages. Cleaner images. Better source data.

Many historical collections are digitized non-destructively with the book held open. This can introduce gutter curvature, shadows, warped text and geometric distortion. Our workflow is designed for image quality first.

01

Destructive flat-page digitization

When appropriate, books are carefully disbound so each page can be scanned flat. This significantly reduces gutter curvature, page shadow and text deformation compared with open-book scanning.

02

600 dpi source quality

Pages are digitized at 600 dpi in grayscale or black and white, producing high-quality source images suitable for OCR, document AI, layout analysis and multimodal training.

03

Books given a second life

A significant part of our source material consists of damaged, incomplete, mismatched or otherwise unsellable books acquired from private libraries — volumes that might otherwise be discarded or pulped.

Preservation through publishing

Rare works can remain available in print.

When an original work is rare, historically significant or worth preserving as a physical publication, we may create a small print re-edition alongside the digital archive. In this way, digitization can help preserve the work and return it to circulation.

Flexible output

From physical books to machine-readable assets.

Every project is different. Datasets can be delivered as raw digitizations or as richer, structured packages designed for downstream processing.

High-resolution PDF

Complete digital reproductions preserving original layout, illustrations, tables and typography. Our standard digitization workflow supports 600 dpi grayscale or black & white scanning.

Page images

Individual page files suitable for OCR training, computer vision and multimodal document models.

OCR text

Machine-readable text for selected collections, with quality depending on source material and project scope.

Bibliographic metadata

Author, title, date, publisher, place of publication, subject, edition, language and other available fields.

Subject coverage

A broad range of French intellectual and technical heritage.

Literature History Medicine Law Science Engineering Agriculture Architecture Geography Travel Fine arts Crafts Dictionaries Encyclopaedias Technical manuals Local history Memoirs Bibliography

Why work with us

Book expertise first. Data delivery second.

Our background combines antiquarian books, publishing and digitization. That means we understand the source material before it becomes a file.

We can help identify relevant editions, historical periods, subject areas and suitable source material, then prepare the collection in a format compatible with your technical workflow. Our ability to work from physical books also gives us more control over scan quality, provenance and source selection than a workflow based only on third-party digital libraries.

500 PDFs available immediately
140M+ words in the ready collection
600 dpi grayscale or black & white scanning
30,000 books in our own stock
600,000 books accessible through our sourcing network
20+ years of experience

Use cases

Designed for teams working on language and documents.

LLM training

French and historical long-form text.

OCR

Scanned pages, typography and difficult layouts.

Document AI

Multimodal understanding of printed material.

NLP research

Diachronic language and specialist vocabulary.

RAG corpora

Domain-specific historical knowledge bases.

Benchmarking

Evaluation sets for OCR and language models.

Digital humanities

Structured corpora for scholarly analysis.

Knowledge extraction

Entities, topics and semantic indexing.

Responsible sourcing

Digitizing material that often has little or no resale value.

When we acquire private libraries, part of the collection is often unsuitable for the antiquarian or second-hand market: damaged books, incomplete sets, isolated volumes or copies in poor condition.

Instead of allowing this material to be discarded, we can transform it into high-quality digital source data. We do not position valuable collectible copies as disposable raw material: the workflow is designed to recover knowledge from books that are otherwise difficult or impossible to sell in their current state.

Request a dataset

Tell us what your model needs.

Send us your target subject, historical period, approximate number of books or pages, preferred formats, OCR requirements, metadata needs and any image-quality constraints.

We can propose our immediately available 500-PDF collection or prepare a custom digitization project from our own stock and wider sourcing network.

Direct enquiries: fred.douin@gmail.com