Government & Public SectorLarge Language Models

Choctaw Nation preserves its language using a fine-tuned Llama AI translation system on Oracle AI infrastructure

Choctaw Nation of Oklahoma· United StatesOCI Data Science · OCI AI Infrastructure · Oracle APEX +2

The Choctaw Nation of Oklahoma, whose native language has dwindled to about 300 first-language speakers, built an AI-powered translation application using Oracle Cloud Infrastructure (OCI) Data Science, OCI AI infrastructure, and Oracle APEX. Tribal linguists fine-tuned Meta's Llama large language models to validate translations between English and Choctaw. When a translation request comes in, an AI agent uses the fine-tuned Llama model with retrieval-augmented generation to consult dictionary resources and validated translation databases, generating a proposed translation that human linguists then review and refine. The application can extract and validate translation pairs from historical documents such as 19th-century newspapers and interview transcripts. Using Oracle AI infrastructure and model fine-tuning tools, tribal linguists can process double the translation requests per day compared to the prior manual process, and sensitive language data stays within Choctaw-controlled OCI environments for data residency and cultural protection.

Overview

The Choctaw Nation of Oklahoma, whose native language has dwindled to about 300 first-language speakers, built an AI-powered translation application using Oracle Cloud Infrastructure (OCI) Data Science, OCI AI infrastructure, and Oracle APEX. Tribal linguists fine-tuned Meta's Llama large language models to validate translations between English and Choctaw. When a translation request comes in, an AI agent uses the fine-tuned Llama model with retrieval-augmented generation to consult dictionary resources and validated translation databases, generating a proposed translation that human linguists then review and refine. The application can extract and validate translation pairs from historical documents such as 19th-century newspapers and interview transcripts. Using Oracle AI infrastructure and model fine-tuning tools, tribal linguists can process double the translation requests per day compared to the prior manual process, and sensitive language data stays within Choctaw-controlled OCI environments for data residency and cultural protection.

This entry has 12 published fields tied to exact passages in an immutable source capture.

Inspect the highlighted source

The challenge

After generations of pressured assimilation, use of the Choctaw language dwindled to just 300 native speakers, prompting urgency to preserve the endangered language, and the tribe's manual translation processes could not keep up with demand from tribal programs.

The solution

The tribe used Oracle Cloud Infrastructure (OCI) Data Science, Oracle AI infrastructure, and an Oracle APEX-built application to fine-tune Meta's Llama large language models for English-Choctaw translation. When a translation request is submitted, an AI agent uses the fine-tuned Llama model with retrieval-augmented generation to consult dictionary resources and validated translation databases and generate a proposed translation, which human linguists then review and refine.

Large Language ModelsRetrieval-Augmented GenerationNatural Language Processing

Reported business value

Using Oracle AI infrastructure and model fine-tuning tools, tribal linguists can process double the translation requests per day, helping meet demand from tribal programs, while sensitive language data stays within Choctaw-controlled OCI environments for data residency and cultural protection.

Sources

Open any source and check the claim yourself — that is the point of the register.

This record was researched and written with AI assistance, and its claims were checked against the sources above. (EU AI Act art. 50 transparency notice.)

Related entries

Other government & public sector entries in the register.

All entries
Government & Public SectorIntelligent AutomationPublic Cloud

GovTech unlocks data insights to improve nationwide services with Databricks

Singapore's Government Technology Agency (GovTech) migrated from an on-premises dashboarding system to the Databricks Data + AI Platform on AWS with Unity Catalog and Delta Lake, cutting dashboard creation time from 90 to 30 days, democratizing data across 50% of corporate divisions in the first year, and saving 8,000 labor hours annually.

96/100HighPrimary source
GovTech· SingaporeDelta Lake · Lakeflow Connect · Lakeflow Jobs +1
Government & Public SectorConversational AIPublic Cloud

Estonia rolls out Bürokratt, an AI-guided virtual assistant network for public services

Bürokratt is a network of chatbots deployed on Estonian public sector institutions' websites, letting people obtain information from institutions and use public and information services via virtual assistants. It is a state-created, AI-based digital assistant that helps institutions deliver modern, efficient, around-the-clock customer service using large language models.

96/100HighPrimary source
Estonian Government (RIA - State Information System Authority)· EstoniaLarge Language Models (LLM) · Retrieval-Augmented Generation (RAG)
Government & Public SectorGenerative AIPublic Cloud

VA Advances Healthcare Insights With AI

The U.S. Department of Veterans Affairs leverages Databricks Data Intelligence to modernize healthcare analytics for millions of veterans, unifying massive distributed datasets into a single secure environment and streaming petabytes of health data in real time, reducing processes that once took hours to seconds. Databricks enables AI and large language models to detect risk early and enhance governance.

96/100HighPrimary source
U.S. Department of Veterans AffairsDatabricks
Government & Public SectorNatural Language ProcessingUnknown

Austrian Academy of Sciences unlocks Ancient Greek with Mistral

The Austrian Academy of Sciences (OeAW), together with its Austrian Archaeological Institute, partnered with Mistral and services partner Reply to build Apollo, described as the first advanced large language model for Ancient Greek. Apollo is trained on a specialized corpus of 600 million words of historical Greek text plus tens of thousands of published inscriptions and papyri, helping researchers reconstruct damaged texts and identify thematic connections across collections. The OeAW reports Apollo turns work that once took years into hours, addressing over one million unread Greek papyri worldwide, with future phases planned for semantic search and handwritten inscription decipherment.

92/100HighPrimary source
Austrian Academy of Sciences· AustriaMistral

Was this helpful?

Your feedback helps us improve our use case database