Media & EntertainmentLarge Language ModelsComputer VisionRecommendation & Personalization
StoryGraph scales self-hosted LLM and ML models to 300 million monthly requests
StoryGraph
Book recommendation platform StoryGraph, run by a two-person team, scaled AI/ML infrastructure to handle 300 million monthly requests by self-hosting large language models rather than using external APIs like ChatGPT. Self-hosting was driven by data privacy and cost: using GPT-4 at their scale was calculated to cost approximately $770,000 per month, versus a fraction of that self-hosted. The company processes roughly 1 million LLM requests per day using Redis-based queuing across GPU servers and migrated its database to a 60-server YugabyteDB cluster.