Choctaw Nation preserves its language using a fine-tuned Llama AI translation system on Oracle AI infrastructure
The Choctaw Nation of Oklahoma, whose native language has dwindled to about 300 first-language speakers, built an AI-powered translation application using Oracle Cloud Infrastructure (OCI) Data Science, OCI AI infrastructure, and Oracle APEX. Tribal linguists fine-tuned Meta's Llama large language models to validate translations between English and Choctaw. When a translation request comes in, an AI agent uses the fine-tuned Llama model with retrieval-augmented generation to consult dictionary resources and validated translation databases, generating a proposed translation that human linguists then review and refine. The application can extract and validate translation pairs from historical documents such as 19th-century newspapers and interview transcripts. Using Oracle AI infrastructure and model fine-tuning tools, tribal linguists can process double the translation requests per day compared to the prior manual process, and sensitive language data stays within Choctaw-controlled OCI environments for data residency and cultural protection.
Overview
The Choctaw Nation of Oklahoma, whose native language has dwindled to about 300 first-language speakers, built an AI-powered translation application using Oracle Cloud Infrastructure (OCI) Data Science, OCI AI infrastructure, and Oracle APEX. Tribal linguists fine-tuned Meta's Llama large language models to validate translations between English and Choctaw. When a translation request comes in, an AI agent uses the fine-tuned Llama model with retrieval-augmented generation to consult dictionary resources and validated translation databases, generating a proposed translation that human linguists then review and refine. The application can extract and validate translation pairs from historical documents such as 19th-century newspapers and interview transcripts. Using Oracle AI infrastructure and model fine-tuning tools, tribal linguists can process double the translation requests per day compared to the prior manual process, and sensitive language data stays within Choctaw-controlled OCI environments for data residency and cultural protection.
This entry has 12 published fields tied to exact passages in an immutable source capture.
Inspect the highlighted sourceThe challenge
After generations of pressured assimilation, use of the Choctaw language dwindled to just 300 native speakers, prompting urgency to preserve the endangered language, and the tribe's manual translation processes could not keep up with demand from tribal programs.
The solution
The tribe used Oracle Cloud Infrastructure (OCI) Data Science, Oracle AI infrastructure, and an Oracle APEX-built application to fine-tune Meta's Llama large language models for English-Choctaw translation. When a translation request is submitted, an AI agent uses the fine-tuned Llama model with retrieval-augmented generation to consult dictionary resources and validated translation databases and generate a proposed translation, which human linguists then review and refine.
Reported business value
Using Oracle AI infrastructure and model fine-tuning tools, tribal linguists can process double the translation requests per day, helping meet demand from tribal programs, while sensitive language data stays within Choctaw-controlled OCI environments for data residency and cultural protection.
Sources
Open any source and check the claim yourself — that is the point of the register.
This record was researched and written with AI assistance, and its claims were checked against the sources above. (EU AI Act art. 50 transparency notice.)
Other government & public sector entries in the register.
GovTech unlocks data insights to improve nationwide services with Databricks
Singapore's Government Technology Agency (GovTech) migrated from an on-premises dashboarding system to the Databricks Data + AI Platform on AWS with Unity Catalog and Delta Lake, cutting dashboard creation time from 90 to 30 days, democratizing data across 50% of corporate divisions in the first year, and saving 8,000 labor hours annually.
Estonia rolls out Bürokratt, an AI-guided virtual assistant network for public services
Bürokratt is a network of chatbots deployed on Estonian public sector institutions' websites, letting people obtain information from institutions and use public and information services via virtual assistants. It is a state-created, AI-based digital assistant that helps institutions deliver modern, efficient, around-the-clock customer service using large language models.
VA Advances Healthcare Insights With AI
The U.S. Department of Veterans Affairs leverages Databricks Data Intelligence to modernize healthcare analytics for millions of veterans, unifying massive distributed datasets into a single secure environment and streaming petabytes of health data in real time, reducing processes that once took hours to seconds. Databricks enables AI and large language models to detect risk early and enhance governance.
Austrian Academy of Sciences unlocks Ancient Greek with Mistral
The Austrian Academy of Sciences (OeAW), together with its Austrian Archaeological Institute, partnered with Mistral and services partner Reply to build Apollo, described as the first advanced large language model for Ancient Greek. Apollo is trained on a specialized corpus of 600 million words of historical Greek text plus tens of thousands of published inscriptions and papyri, helping researchers reconstruct damaged texts and identify thematic connections across collections. The OeAW reports Apollo turns work that once took years into hours, addressing over one million unread Greek papyri worldwide, with future phases planned for semantic search and handwritten inscription decipherment.
Was this helpful?
Your feedback helps us improve our use case database
