{"slug":"multitudes-builds-code-review-quality-feature-in-2-months-using-3-llms-on-amazon-bedrock","url":"https://findausecase.com/use-cases/multitudes-builds-code-review-quality-feature-in-2-months-using-3-llms-on-amazon-bedrock","title":"Multitudes Builds Code Review Quality Feature in 2 Months Using 3 LLMs on Amazon Bedrock","description":"Multitudes, a New Zealand-based engineering analytics startup, used Amazon Bedrock to build a code review quality feature, testing over 10 large language models across roughly 1,000 code reviews before choosing Amazon Nova Pro for bot detection, Anthropic Claude for feedback specificity and prompt-injection detection, and Mistral for sentiment analysis, orchestrated with Amazon Elastic Container Service. The feature increased monthly active users by 44 percent within two months of launch and reduced severe misclassification rates from 20 percent to under 1 percent.","company":"Multitudes","industry":"Technology & Software","country":"New Zealand","aiCapabilities":["Natural Language Processing"],"technology":["Amazon Bedrock","Amazon Nova Pro","Anthropic Claude","Mistral","Amazon Elastic Container Service"],"deployment":"Public Cloud","problemStatement":"Multitudes had always measured code review activity by the number of reviews or comments, but customers wanted insight into the quality of those reviews. Traditional natural language processing and machine learning models couldn't deliver the level of accuracy needed to build a feature customers could trust.","solutionApproach":"Multitudes used Amazon Bedrock to build a code review quality feature, evaluating nearly 1,000 code reviews and manually creating a labeled ground truth dataset across three dimensions: feedback specificity, tone/sentiment, and bot-generated activity. It tested over 10 large language models, then used different models for each dimension -- Amazon Nova Pro for bot detection, Anthropic Claude for feedback specificity and prompt-injection detection, and Mistral for sentiment analysis -- with Amazon Elastic Container Service used to orchestrate the data pipeline.","businessValue":"Within two months of launch, the new code review quality feature drove a 44 percent increase in monthly active users. Model accuracy improved significantly, with severe misclassification rates falling from 20 percent to under 1 percent. The feature quickly became one of the platform's top five most-used capabilities.","evidence":{"band":"high"},"sourceUrl":"https://aws.amazon.com/solutions/case-studies/multitudes-case-study","dates":{"publishedAt":"2026-08-18T12:18:23.317Z","publishedAtSource":"pipeline","updatedAt":"2026-08-18T12:18:23.317Z"},"license":"Open for reading and citing with a link to https://findausecase.com/use-cases/multitudes-builds-code-review-quality-feature-in-2-months-using-3-llms-on-amazon-bedrock. Bulk republication requires permission."}