← Back to Model Experimentation Made Easier

Source proof for Model Experimentation Made Easier

Source-bound proof

Verified source excerpts for every supported field

Each colour maps a published value to the exact source passage used to support it. Only bounded excerpts are public; administrators can inspect the complete captured source.

11 fields supported

Title

Model Experimentation Made Easier

quote · high
…emo Login Contact Us Try Databricks Customer Stories / Refuel.AI CUSTOMER STORY Model Experimentation Made Easier 86% Improvement in data label quality using a custom LLM 3 Days to build a mode…

Description

Refuel.AI built Refuel LLM, a purpose-built model for data labeling and enrichment, by instruction-tuning a Llama-v2-13b base model on more than 5 billion tokens using Databricks Training infrastructure. The initial training run produced a 78% increase in label quality, and subsequent fine-tuning on a cluster of 8x H100s added a further 16% performance gain, outperforming trained human annotators and several other LLMs on a 15-dataset text labeling benchmark.

derived · high
…aining runs, each about three days long, on Databricks Training infrastructure. The end result was a 78% increase in label quality. The next step was fine-tuning to further improve the performance. Thanks to Dat…
…arget domain, further improving performance and TCO by reducing prompt lengths. Refuel.AI was able to fine-tune the model on a cluster of 8x H100s within the Databricks Training environment, for an additional 16% performance gain. (See Figure 2) Figure 2: Label Quality Improvement From Fine-tuning Refuel-LLM…

Company

Refuel.AI

quote · high
…rity and Trust Ready to get started? Get a Demo Login Contact Us Try Databricks Customer Stories / Refuel.AI CUSTOMER STORY Model Experimentation Made Easier 86% Improvement in data label…

Industry

Technology & Software

classification · high
…recently completed the next generation of Refuel LLM . Share this post Details Industry : Technology and Software Use Case : Artificial Intelligence Product : Agent Bricks Ready to get started?…

Problem

Data labeling requires considerable resources and time; manual labeling is costly and vulnerable to human errors, and human-in-the-loop methods only marginally speed up the process and increase accuracy.

derived · high
…ne platforms and improved product recommendations on e-commerce sites. However, data labeling requires considerable resources and time. Manual labeling is costly and vulnerable to human errors. Human-in-the-loop methods only marginally speed up the process and increase accuracy. Large language models (LLMs) can enable vast improvements to this workflow. Fas…

Solution

Refuel.AI built Refuel LLM, a purpose-built model for data labeling and enrichment tasks, instruction-tuned on more than 5 billion tokens across more than 2,500 unique tasks on top of a Llama-v2-13b base model. The team trained close to 50 models over almost three months on Databricks Training infrastructure, with initial training runs of about three days each, then fine-tuned the model on a cluster of 8x H100s within Databricks Training to further improve performance and reduce prompt lengths.

derived · high
…e-built model for data labeling and enrichment tasks. Launched in October 2023, the model was instruction-tuned on more than 5 billion tokens (comprising more than 2,500 unique tasks) on top of a Llama-v2-13b base model. It outperforms trained human annotators (80.4%), GPT-3-5-turbo (81.3%), PaLM-2…

Business value

Refuel LLM outperformed trained human annotators and other leading LLMs (GPT-3.5-turbo, PaLM-2, Claude) across a benchmark of 15 text labeling data sets, and the initial release attracted over ten thousand users accessing the Refuel LLM cloud or playground.

derived · high
…(comprising more than 2,500 unique tasks) on top of a Llama-v2-13b base model. It outperforms trained human annotators (80.4%), GPT-3-5-turbo (81.3%), PaLM-2 (82.3%), and Claude (79.3%) across a benchmark of 15 text labeling data sets. (See Figure 1) Figure 1: Evaluating Label Quality Across LLMs The Refuel.AI tea…

AI capabilities

Large Language Models, Machine Learning

classification · high
…ut with LLMs as data annotators. Feedback cycles can be up to 100 times faster. Refuel LLM is a purpose-built model for data labeling and enrichment tasks. Launched in October 2023, the model was instruction-tuned on more than 5 billio…

Technology

Refuel LLM, Llama-v2-13b, Databricks Training, Agent Bricks

classification · high
…t Details Industry : Technology and Software Use Case : Artificial Intelligence Product : Agent Bricks Ready to get started? Try Databricks for free Learn more about our product Talk…

Implementation approach

Trained the model iteratively using Databricks Training's auto-scheduling to queue runs in advance, with graceful resumption, data streaming and dynamic memory usage, running close to 50 experimental models over almost three months before fine-tuning.

derived · high
…l.AI can afford to run experiments and train their LLM in an iterative process. One of the factors that made experimentation easier was the auto-scheduling of training runs. The team was able to queue up runs in advance and avoid hassles with GPU availability or node failures. The platform offers “set it and forget it” capabilities like graceful resumption, data streaming, and dynamic memory usage. The Refuel.AI team was able to leverage the optimized infrastructure and compre…

Headline outcome

derived · high
…s Customer Stories / Refuel.AI CUSTOMER STORY Model Experimentation Made Easier 86% Improvement in data label quality using a custom LLM 3 Days to build a model from scratch with Databricks Training Access to high-qu…
Capture details
Captured
16 Sept 2026, 06:05 UTC
Extractor
fetch-strip@1
Snapshot hash
6d72838926c6cdfb063173a6ef83f7de93688f9d6dd2f7da1d79a8b55e651f99