Good Start Labs
Good Start Labs builds game-based reinforcement learning environments and public leaderboards to train and evaluate frontier AI models. It serves AI labs such as OpenAI, Cohere, and Arcee with evaluation services, RL training data, and humor-alignment benchmarks.
- Company typePrivate
- Founded2025
- HeadquartersNew York, United States
- Headcount1–10
- GTM typeB2B
- OfferingSoftware
What Good Start Labs does
Good Start Labs (legal entity Good Start, Inc.) is a New York-based, venture-backed company founded in 2025 that builds game-based AI training and evaluation environments for frontier AI labs. The company operates three core products: AI Diplomacy — a multi-agent strategic negotiation environment that trains and tests models on long-horizon reasoning, honesty, and multi-agent coordination; Diplomacy Arena — a public leaderboard evaluating 23 AI models from Google, OpenAI, xAI, Anthropic, Meta, and DeepSeek on overall performance, betrayal tendency, and steerability; and LOL Arena — a humor alignment benchmark built on the Bad Cards platform that evaluates 17 models against human-curated preferences. Underlying the products is a reinforcement learning environment stack designed to generate generalizable skills that transfer from simulated gameplay to real-world agent tasks.
Revenue is generated primarily through enterprise evaluation services and collaborative RL gym development with AI labs (e.g., OpenAI's GPT-5 pre-release evaluation and Arcee's Trinity model family partnership), supplemented by referral distribution via integration with Bad Cards' 2 million+ user consumer platform, which runs 20,000+ AI games weekly and supplies live human feedback. Go-to-market combines direct enterprise sales through website inquiry forms with product-led growth through public leaderboards, research publications (including a NeurIPS workshop-accepted paper), and Twitch demonstration content. The company is privately held, spun out of Every, with 1–10 employees and a $3.6 million seed round co-led by Inovia Capital and General Catalyst in October 2025.
Good Start Labs firmographics
Firmographics- Name
- Good Start Labs
- Legal name
- Good Start, Inc.
- Website
- https://goodstartlabs.com
- Company type
- Private
- Founded year
- 2025
- Operating status
- Operating
- Headcount range
- 1–10 employees
- Short description
- Good Start Labs builds game-based reinforcement learning environments and public leaderboards to train and evaluate frontier AI models. It serves AI labs such as OpenAI, Cohere, and Arcee with evaluation services, RL training data, and humor-alignment benchmarks.
- Ownership category
- akta.pro rank
Good Start Labs industry classification
Industry- Product category
- AI Training and Evaluation Software
- NAICS
- Technical and Trade Schools (6115)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Programming Services (7371)
- akta.pro primary industry
- Safety & Alignment Evaluation (red-teaming, harmful capability testing) (HDAAAMAL)
- akta.pro secondary industries
- Safety, Alignment & Content Moderation (Guardrails, Red Teaming) (HDAAACAI), Gaming & Interactive Entertainment Personalization (HDAAAGAF), Simulation-Based Learning & Virtual Labs (EDAGAEAG)
Keywords
Where Good Start Labs is headquartered
LocationHeadquarters
- HQ city
- New York
- HQ country
- United States
- HQ region
- North America
Markets served
Good Start Labs business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Infrastructure, Operations, Marketing or Sales
Revenue model
- AI Model Training Services: Providing AI labs with reinforcement learning environments, training data, and evaluation services. Partnered with Arcee for Trinity model family training, evaluation, and data sourcing.
- Model Evaluation and Benchmarking: Evaluation services for AI labs including pre-release model assessment. Evaluated GPT-5 ahead of release to inform model decisions, identifying high steerability and prompt sensitivity. Labs can submit models for transparent, reproducible evaluation.
- Targeted RL Gym Environments: Collaborative development of custom RL gym environments for AI labs seeking targeted training data and evaluation frameworks.
Go-to-market motion2 records
Distribution channels3 records
Marketing channels5 records
Good Start Labs product offering
Product offeringCore offering
Good Start Labs trains and evaluates AI models through game-based reinforcement learning environments. The company builds multi-agent game simulations—primarily AI Diplomacy and a Bad Cards integration—where AI models play full games to generate training data and produce public benchmark leaderboards (Diplomacy Arena, LOL Arena) that measure long-horizon reasoning, negotiation, tool use, theory of mind, honesty, steerability, and humor alignment with human preferences.
Product overview
Good Start Labs operates as a game-based AI training and evaluation platform offering three interconnected products: (1) AI Diplomacy - a strategic negotiation game environment for training AI models on long-horizon reasoning and multi-agent coordination; (2) Diplomacy Arena - a benchmark leaderboard evaluating 23 models on overall performance, betrayal tendency, and steerability in Diplomacy; (3) LOL Arena - a humor alignment benchmark evaluating 17 models on their ability to match human preferences for humor through Bad Cards gameplay. The platform also provides reinforcement learning environments where models learn generalizable skills from millions of game interactions. All products interconnect through shared game mechanics and evaluation methodologies, with human feedback from Bad Cards informing RL training while leaderboards provide transparent model comparison across reasoning, strategy, honesty and alignment dimensions.
Differentiator
Problem solved
Functional benefit
Brands
- AI Diplomacy: Multi-agent AI training environment based on the board game Diplomacy, measuring long-horizon reasoning, honesty, and multi-agent coordination.
- Diplomacy Arena
- LOL Arena
Products and services
- AI Diplomacy
- Diplomacy Arena
- LOL Arena
- Reinforcement Learning Environments
Quantifiable outcome
- Identified high steerability and prompt sensitivity in GPT-5, enabling OpenAI to provide specific prompting instructions and a prompt optimizer
- +3 more outcomes
Companies that use Good Start Labs
Customer profileNamed customers4 records
Segments3 records
Ideal customer profiles3 records
Good Start Labs technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
Integration1 record
AI capability5 records
Feature5 records
Good Start Labs partnerships and signals
Strategic signalPartnerships
Four partnerships are on record, tiered core and flagship.
- Bad CardscorePartnership where Bad Cards uses game feedback from their 2 million+ users to improve AI humor and alignment. Bad Cards' AI agents run 20,000+ games weekly, providing live human feedback. Their users easily add AI agents to games with friends, creating a scalable feedback mechanism.
- ArceecorePartnership with Arcee to support training, evaluation, and data sourcing efforts of the Trinity model family. Arcee is also listed as a key partner with logo placement on the homepage.
- OpenAIflagshipGood Start Labs evaluated GPT-5 ahead of release, identifying high steerability and prompt sensitivity. This evaluation enabled OpenAI to create specific prompting instructions and a prompt optimizer cookbook for GPT-5 users.
- CoherecoreCohere is listed as a key partner and customer, using Good Start Labs' game environments for AI training and evaluation. Featured on homepage alongside other major AI labs.
Scale indicators7 records
Recent moves7 records
Expansion highlights5 records
Good Start Labs competitors and assessment
Company assessmentEmerging players
- Prime Intellect: Distributed AI training and RL environment startup. Comparable because both build RL environments for training frontier models, though Prime Intellect focuses on distributed compute and synthetic environments rather than game-based behavioral eval.
- Apart Research: AI safety research organization producing alignment evaluations and benchmarks. Comparable because both build novel AI evaluation methodologies and engage the AI safety research community, though Apart is research-only and non-commercial.
- Labelbox: AI data labeling and evaluation platform serving enterprise AI teams. Comparable as a training-data and evaluation vendor to AI developers, with overlap in RLHF/preference data workflows, though Labelbox is broader and more enterprise-software oriented.
- Tonic.ai: Synthetic data and evaluation platform for AI development. Comparable as a provider of training data and model evaluation infrastructure to AI teams, with some overlap in serving RL training needs, though focused on synthetic rather than interactive data.
Broad incumbents
- Hugging Face: Open-source AI platform hosting models, datasets, and community benchmarks. Comparable because both surface public model evaluations/leaderboards to the AI ecosystem, though Hugging Face is a much broader platform rather than a focused evaluation vendor.
- Snorkel AI: Data-centric AI platform providing programmatic data labeling and evaluation pipelines for enterprise AI. Comparable because both sell training data and evaluation tooling to AI developers; Snorkel is broader and enterprise-focused.
- Scale AI: Large data-labeling and evaluation platform serving most frontier AI labs. Comparable because both sell training data, RL environments, and model evaluation services to AI labs, though Scale is much broader and not game-based.
- Surge AI: Human-in-the-loop data and evaluation provider for AI labs. Comparable because it supplies RL training data and qualitative model evaluations (preference, alignment) to frontier model developers, though it does not specialize in game-based environments.
Direct peers
- METR (Model Evaluation and Threat Research): Non-profit research organization that builds rigorous evaluations of frontier AI capabilities, including agentic task suites. Most directly comparable peer because both organizations sell third-party AI model evaluation services to labs, researchers, and policymakers.
- Apollo Research: AI safety research lab focused on evaluating frontier models for deception, scheming, and alignment failures. Directly comparable in that both organizations provide behavioral/alignment evaluations of frontier AI models to labs and external stakeholders.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights6 records
Customer concentration
Good Start Labs social profiles
Digital presenceGood Start Labs financial estimates
Financial estimateRevenue estimate
Valuation estimate
Good Start Labs leadership team
Management profileNumber of profiles
Profiles2 records
Good Start Labs funding detail
Funding detailFunding overview
Funding rounds1 record
Investors3 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Good Start Labs M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Good Start Labs
What does Good Start Labs do?
Good Start Labs trains and evaluates AI models through game-based reinforcement learning environments. The company builds multi-agent game simulations—primarily AI Diplomacy and a Bad Cards integration—where AI models play full games to generate training data and produce public benchmark leaderboards (Diplomacy Arena, LOL Arena) that measure long-horizon reasoning, negotiation, tool use, theory of mind, honesty, steerability, and humor alignment with human preferences.
Is Good Start Labs a public or private company?
Good Start Labs is a private company. It is classified as venture growth investor backed and is currently operating.
When was Good Start Labs founded?
Good Start Labs was founded in 2025. It employs 1 to 10 people.
Where is Good Start Labs based?
Good Start Labs is headquartered in New York, United States, in the North America region.
How does Good Start Labs make money?
Three revenue lines are on record. AI Model Training Services are the primary driver. The others are model Evaluation and Benchmarking and targeted RL Gym Environments.
Who are Good Start Labs's main competitors?
Emerging players on record are Prime Intellect, Apart Research, Labelbox and Tonic.ai. Broad incumbents are Hugging Face, Snorkel AI, Scale AI and Surge AI. Direct peers are METR (Model Evaluation and Threat Research) and Apollo Research.
Does Good Start Labs have an API?
No public API is recorded for Good Start Labs.
What industry is Good Start Labs in?
Good Start Labs's product category is AI Training and Evaluation Software. Its primary akta.pro industry code is HDAAAMAL, Safety & Alignment Evaluation (red-teaming, harmful capability testing), with a secondary code of HDAAACAI, Safety, Alignment & Content Moderation (Guardrails, Red Teaming). Its NAICS code is 6115 and its SIC code is 7372.