Kantar Worldpanel fine-tunes GenAI models on Databricks to generate market-insight training data faster
Kantar Worldpanel used the Databricks Data + AI Platform and MLflow to experiment with Llama, Mistral, GPT-4 and GPT-3.5 for a proof of concept linking receipt descriptions to product barcode names. GPT-4 produced the most accurate outputs (94%), which the team used to automatically generate a training dataset of about 120,000 receipt-to-barcode pairs in a couple of hours to fine-tune a smaller production model.
Overview
Kantar Worldpanel used the Databricks Data + AI Platform and MLflow to experiment with Llama, Mistral, GPT-4 and GPT-3.5 for a proof of concept linking receipt descriptions to product barcode names. GPT-4 produced the most accurate outputs (94%), which the team used to automatically generate a training dataset of about 120,000 receipt-to-barcode pairs in a couple of hours to fine-tune a smaller production model.
This entry has 13 published fields tied to exact passages in an immutable source capture.
Inspect the highlighted sourceThe challenge
Kantar Worldpanel's legacy systems were inflexible, resource-intensive to maintain, and required specialized, outdated programming skillsets, limiting data democratization and experimentation with new AI-driven use cases.
The solution
Kantar Worldpanel used the Databricks Data + AI Platform and MLflow to manage the ML lifecycle, experimenting with Llama, Mistral, GPT-4 and GPT-3.5 to fine-tune a model linking receipt descriptions to product barcode names, downloading models via Databricks Marketplace, exploring Databricks AI Search for description comparisons, and using Unity Catalog to govern data sharing across teams.
Reported business value
Kantar Worldpanel automatically generated a training dataset of about 120,000 receipt-to-barcode description pairs at 94% accuracy in just a couple of hours, letting manual coding teams focus on discrepant results and freeing engineering resources for core development, while streamlining data scientists' workflows.
Sources
Open any source and check the claim yourself — that is the point of the register.
This record was researched and written with AI assistance, and its claims were checked against the sources above. (EU AI Act art. 50 transparency notice.)
Other technology & software entries in the register.
HP crafts marketing campaigns that resonate with customers using Databricks and Uniphore
HP centralized first-party customer data on the Databricks Data + AI Platform with Delta Lake and Unity Catalog, and connected it to Uniphore's HybridCompute for federated query pushdown, cutting campaign setup from 2 weeks to 2 hours and processing 400 million records in seconds.
Transforming Weather Forecasting with Lakeflow Jobs
AccuWeather migrated from on-premises infrastructure to Databricks and Lakeflow Jobs, working with Datadog for observability, to unify diverse weather data formats and orchestrate 4,500+ weekly jobs. Lakeflow Jobs coordinates the ingestion of multiple weather models, triggers machine learning processes that weight and blend different forecasts, and manages complex job dependencies for reinforcement training workflows used in AccuWeather's proprietary forecasting engine. AccuWeather reports 3x faster dataset development (three months to one month per dataset), a 50% reduction in unactionable alerts, and 50% cost savings on serverless job usage.
Adobe brings creativity to life with Databricks
Adobe uses the Databricks Data + AI Platform for end-to-end data management that unifies all data and AI at scale, with 20% faster performance. Databricks equips over 92 teams at Adobe to unify data from financials, sales, products, customers and employees so they can drive personalized experiences across Adobe's digital platforms with AI.
Supermetrics: Helping Marketers Redefine Efficiency with AI-Powered Data Analysis
Supermetrics, a Finland-based marketing intelligence platform serving 15,000+ customers across 132 countries, built an AI agent on Google Cloud using Vertex AI Agent Builder and the Agent Development Kit (ADK) that autonomously manages data connections, fixes pipeline errors, and analyzes campaign performance in real time, suggesting new creative options using Imagen. The agent automates the weekly marketing reporting cycle that previously took performance marketers up to four hours, reclaiming over 15 hours per month per marketer for strategy and creative testing. The system uses a central AI agent that interprets natural language requests and delegates tasks to sub-agents, and stores 'core memories' of user preferences for personalized context.
Was this helpful?
Your feedback helps us improve our use case database

