Government & Public SectorDocument IntelligenceOn-Premise

The European Patent Office accelerates innovation with Mistral

European Patent OfficeMistral AI OCR

The European Patent Office partnered with Mistral AI on a three-month proof of concept to automate patent document transcription. A 1B-parameter OCR model was fine-tuned on over 150,000 PDFs (50,000 unique patents) and built into a pipeline converting complex patent documents (with formulas, chemical structures, tables, and figures across multiple European languages) from Markdown to HTML to ST36-compliant XML, deployed on-premises. The system achieves 400,000 pages/day throughput with a character recognition error rate below 1%, cutting lead time from up to 5 days down to a few minutes.

Overview

The European Patent Office partnered with Mistral AI on a three-month proof of concept to automate patent document transcription. A 1B-parameter OCR model was fine-tuned on over 150,000 PDFs (50,000 unique patents) and built into a pipeline converting complex patent documents (with formulas, chemical structures, tables, and figures across multiple European languages) from Markdown to HTML to ST36-compliant XML, deployed on-premises. The system achieves 400,000 pages/day throughput with a character recognition error rate below 1%, cutting lead time from up to 5 days down to a few minutes.

This entry has 14 published fields tied to exact passages in an immutable source capture.

Inspect the highlighted source

The challenge

Patent documents frequently embed mathematical formulas, chemical structures, detailed images, and intricate tables within text across multiple European languages, elements that are notoriously difficult for standard digital tools to interpret, often leading to incomplete or inaccurate data structuring.

The solution

The EPO partnered with Mistral AI on a three-month proof of concept, deploying a 1B OCR model fine-tuned on over 150,000 PDFs representing 50,000 unique patents, built into a pipeline that converts complex patent documents from Markdown to HTML and finally into the required ST36 XML format, ready for on-premises deployment.

Document Intelligence

Reported business value

The system achieved a throughput of 400,000 pages transcribed per day while maintaining a character recognition error rate of less than 1%, cutting lead time from up to 5 days down to a few minutes.

Sources

Open any source and check the claim yourself — that is the point of the register.

This record was researched and written with AI assistance, and its claims were checked against the sources above. (EU AI Act art. 50 transparency notice.)

Related entries

Other government & public sector entries in the register.

All entries
Government & Public SectorIntelligent AutomationPublic Cloud

GovTech unlocks data insights to improve nationwide services with Databricks

Singapore's Government Technology Agency (GovTech) migrated from an on-premises dashboarding system to the Databricks Data + AI Platform on AWS with Unity Catalog and Delta Lake, cutting dashboard creation time from 90 to 30 days, democratizing data across 50% of corporate divisions in the first year, and saving 8,000 labor hours annually.

96/100HighPrimary source
GovTech· SingaporeDelta Lake · Lakeflow Connect · Lakeflow Jobs +1
Government & Public SectorConversational AIPublic Cloud

Estonia rolls out Bürokratt, an AI-guided virtual assistant network for public services

Bürokratt is a network of chatbots deployed on Estonian public sector institutions' websites, letting people obtain information from institutions and use public and information services via virtual assistants. It is a state-created, AI-based digital assistant that helps institutions deliver modern, efficient, around-the-clock customer service using large language models.

96/100HighPrimary source
Estonian Government (RIA - State Information System Authority)· EstoniaLarge Language Models (LLM) · Retrieval-Augmented Generation (RAG)
Government & Public SectorGenerative AIPublic Cloud

VA Advances Healthcare Insights With AI

The U.S. Department of Veterans Affairs leverages Databricks Data Intelligence to modernize healthcare analytics for millions of veterans, unifying massive distributed datasets into a single secure environment and streaming petabytes of health data in real time, reducing processes that once took hours to seconds. Databricks enables AI and large language models to detect risk early and enhance governance.

96/100HighPrimary source
U.S. Department of Veterans AffairsDatabricks
Government & Public SectorNatural Language ProcessingUnknown

Austrian Academy of Sciences unlocks Ancient Greek with Mistral

The Austrian Academy of Sciences (OeAW), together with its Austrian Archaeological Institute, partnered with Mistral and services partner Reply to build Apollo, described as the first advanced large language model for Ancient Greek. Apollo is trained on a specialized corpus of 600 million words of historical Greek text plus tens of thousands of published inscriptions and papyri, helping researchers reconstruct damaged texts and identify thematic connections across collections. The OeAW reports Apollo turns work that once took years into hours, addressing over one million unread Greek papyri worldwide, with future phases planned for semantic search and handwritten inscription decipherment.

92/100HighPrimary source
Austrian Academy of Sciences· AustriaMistral

Was this helpful?

Your feedback helps us improve our use case database