Andon Labs
Andon Labs is a Y Combinator-backed AI safety research startup that builds Safe Autonomous Organizations — AI agents running real businesses like retail stores, cafés, radio stations, and vending machines — and publishes evaluation benchmarks adopted by Anthropic, OpenAI, Google DeepMind, and xAI.
- Company typePrivate
- Founded2023
- HeadquartersLondon, United Kingdom
- Headcount1–10
- GTM typeB2B and B2C
- OfferingSoftware
What Andon Labs does
Andon Labs is an AI safety and research startup that develops benchmarks, evaluations, and real-world deployments of autonomous AI agents. Founded in 2023 and Y Combinator-backed, the company builds "Safe Autonomous Organizations" (SAOs) — businesses operated by AI agents without humans in the loop — including a retail boutique (Andon Market in San Francisco), a café (Andon Café in Stockholm), AI-powered radio stations (Andon FM), and vending machine agents deployed at offices of Anthropic, xAI, and OpenAI. The company also publishes a suite of evaluation benchmarks (Vending-Bench 2, Vending-Bench Arena, Blueprint-Bench 2, Butter-Bench) that measure AI capabilities in long-horizon business management, spatial reasoning, and robotic control, with multiple benchmarks adopted by leading AI labs at each model release.
The company's core technology centers on benchmarking and deploying frontier AI agents in physical business environments to study autonomous behavior and failure modes. Its benchmarks evaluate capabilities like inventory management, supplier negotiation under adversarial conditions, spatial reasoning from photographs, and robot orchestration for household tasks. Andon Labs collaborates closely with all four leading frontier AI labs (Anthropic, Google DeepMind, OpenAI, xAI), powering their evaluation pipelines for model releases and providing deployment venues for their models in commercial settings.
The business model combines research and benchmarking services to AI labs (pricing not publicly disclosed) with direct-to-consumer sales of hardware and merchandise, including the Andon Radio ($149, handmade WiFi radio streaming Andon FM) and Andon Abs protein powder ($50). The real-world business deployments function primarily as research platforms rather than profit centers — the SF store has accumulated approximately $13,000 in losses against a $100,000 budget, and the Stockholm café has generated roughly $5,700 in revenue against a $21,000+ budget since launching in April 2026. Revenue figures are not publicly disclosed, and the 1-10 person team appears to operate as a research-driven organization where commercial deployments subsidize the production of evaluation data and research output.
Andon Labs firmographics
Firmographics- Name
- Andon Labs
- Legal name
- Andon Labs Inc.
- Website
- https://andonlabs.com
- Company type
- Private
- Founded year
- 2023
- Operating status
- Operating
- Headcount range
- 1–10 employees
- Short description
- Andon Labs is a Y Combinator-backed AI safety research startup that builds Safe Autonomous Organizations — AI agents running real businesses like retail stores, cafés, radio stations, and vending machines — and publishes evaluation benchmarks adopted by Anthropic, OpenAI, Google DeepMind, and xAI.
- Ownership category
- akta.pro rank
Andon Labs industry classification
Industry- Product category
- AI Safety Research and Evaluation
- akta.pro primary industry
- AI Observability, Monitoring & Evaluation Platforms (Drift, Quality, Safety) (HDAEANAF)
- akta.pro secondary industries
- Safety & Alignment Evaluation (red-teaming, harmful capability testing) (HDAAAMAL), Responsible AI, Security & Privacy Platforms (Safety, Guardrails, PII) (HDAEANAG)
Keywords
Where Andon Labs is headquartered
LocationHeadquarters
- HQ city
- London
- HQ country
- United Kingdom
- HQ region
- Europe
Offices2 records
Markets served
Andon Labs business model
Business model- GTM type
- B2B and B2C
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Operations, Supply Chain, Marketing or Sales, Infrastructure
Revenue model
- Product Sales: Direct sales of physical products including the Andon Radio (handmade wooden WiFi radio for streaming Andon FM at $149), Andon Abs protein powder ($50), and other merchandise
- Research and Benchmarking Services: Collaboration with AI labs (Anthropic, Google DeepMind, OpenAI, xAI) on AI safety research and benchmarking. Vending-Bench is used at every major model release by AI companies.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Unit Pricing | Pay-as-you-go | Andon Radio - Handmade wooden WiFi radio for streaming Andon FM |
| Unit Pricing | Pay-as-you-go | Andon Abs Protein Powder |
| Unit Pricing | Pay-as-you-go | Andon Radio - Handmade WiFi-connected wooden radio for streaming Andon FM |
| Other | Pay-as-you-go | Vending Machine Duo - Real-world vending machine deployment |
Go-to-market motion1 record
Distribution channels3 records
Marketing channels4 records
Andon Labs product offering
Product offeringCore offering
Andon Labs develops and publishes AI safety benchmarks (Vending-Bench, Blueprint-Bench, Butter-Bench) for evaluating frontier AI model capabilities, and operates Safe Autonomous Organizations where AI agents run real-world businesses without human oversight. Revenue is generated through research collaborations with AI labs and direct sales of consumer products including the Andon Radio hardware and branded merchandise.
Product overview
Andon Labs is an AI safety and research company building Safe Autonomous Organizations (SAOs) — businesses run entirely by AI agents. The portfolio consists of real-world AI deployments (Andon Market retail store, Andon Café, AI radio stations via Andon FM, and vending machine agents) alongside a suite of evaluation benchmarks (Vending-Bench 2, Blueprint-Bench 2, Butter-Bench) that measure AI capabilities in long-horizon business management, spatial reasoning, and robotics. Physical products include the Andon Radio hardware streaming device and merchandise. The core thesis is that AI safety must be studied through real-world deployment rather than simulation alone, with findings used to build control protocols for autonomous AI systems.
Differentiator
Problem solved
Functional benefit
Products and services
- Vending-Bench 2 Year-long vending machine business simulation benchmark evaluating AI model performance on managing a business over extended time horizons. Models are scored on final bank account balance after 365 simulated days, navigating adversarial suppliers, negotiations, customer complaints, and pricing optimization. Claude Opus 4.7 leads with $10,936.76. Built for AI labs and researchers evaluating agentic long-horizon coherence.
- Vending-Bench Arena Multi-agent competitive variant of Vending-Bench 2 where AI agents manage vending machines head-to-head at the same location. Agents can email each other, send money, and trade goods. Leads to price wars, price cartels, and strategic collusion. Enables study of competitive and collaborative AI behavior in business settings for AI safety researchers.
- Blueprint-Bench 2 Spatial intelligence benchmark where AI agents convert approximately 20 apartment photographs into accurate 2D floor plans showing room layouts, connections, and relative sizes. Each agent processes 50 apartments sequentially with a persistent notepad. Claude Fable 5 leads with 0.386 score versus human baseline of 0.586. Tests genuine spatial reasoning beyond pattern matching in training data.
- Blueprint-Bench Original spatial intelligence benchmark testing how AI models understand space by converting apartment photographs into floor plans. Most models performed at or below random baseline (0.279); humans scored 0.547, significantly outperforming all AI systems. Published October 2025.
- Butter-Bench Evaluation of LLMs as robot orchestrators for practical household tasks including 'pass the butter' and delivery tasks. Best model scored 40% versus 95% for humans. Revealed significant spatial intelligence gaps in current LLMs controlling robotic systems. Published October 2025.
- Andon Market World's first fully AI-managed retail boutique in San Francisco's Cow Hollow neighborhood, operated by AI agent Luna powered by Anthropic's Claude Sonnet 4.6. Luna was given a $100,000 budget, corporate credit card, 3-year retail lease, phone, email, internet, and security camera access to hire staff, select inventory, set prices, and manage all business operations.
- Andon Café AI-managed café in Stockholm's Vasastan district operated by AI agent Mona powered by Google Gemini. Mona handles hiring, supplier management, permits, menu pricing, and all business decisions while human baristas prepare coffee. Café generated 44,000 SEK (~$4,659) in revenue in first two weeks of operation.
- Andon FM Multi-station AI radio network where four AI models (Claude Opus 4.8, GPT-5.5, Gemini 3.5 Flash, Grok 4.3) independently run their own radio stations 24/7. Each agent develops a unique personality, curates music, writes and broadcasts DJ commentary via text-to-speech, engages listeners on X/Twitter, and attempts to generate revenue. Powered by Live365 streaming infrastructure.
- Andon Radio Handmade wooden WiFi-connected radio streaming Andon FM stations. Built in American Black Walnut or White Oak by in-house woodworker with one rotary dial for volume, one for channel selection, and USB-C power. Streams to Dayton Audio 2.5" full-range driver. Priced at $149 plus shipping and tax.
- Vending Machine Duo Physical vending machine setup where two AI agents running different models compete side-by-side for customers, revealing which models are best at running a real-world business. Agents live in a Slack channel where users can negotiate, place custom orders, and receive business updates. Currently available via waitlist.
- Andon Abs 2lb chocolate whey protein powder used internally for Butter-Bench human baseline preparation. Priced at $50. Ships in 1-2 weeks.
Quantifiable outcome
- Vending-Bench 2 leaderboard shows Claude Opus 4.7 achieving $10,936.76 money balance in simulated vending machine business over one year
- +3 more outcomes
Companies that use Andon Labs
Customer profileSegments2 records
Ideal customer profiles2 records
Andon Labs technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
Integration12 records
AI capability13 records
Feature8 records
Andon Labs partnerships and signals
Strategic signalPartnerships
Three partnerships are on record, tiered core.
- AnthropiccoreCollaboration on Project Vend - AI-powered vending machine deployed at Anthropic's San Francisco office. Andon Labs' Vending-Bench is used at every Anthropic model release. Claude models power multiple Andon Labs AI agents including Luna at Andon Market.
- Google DeepMindcoreGoogle DeepMind is listed as a working partner on Andon Labs' website. Gemini models power AI agents including Mona at Andon Cafe Stockholm. Vending-Bench used at Google Gemini model releases.
- OpenAIcoreListed as working with OpenAI. GPT models power AI agents including DJ GPT at Andon FM. Vending-Bench used at OpenAI model releases.
Scale indicators3 records
Recent moves8 records
Expansion highlights6 records
Andon Labs competitors and assessment
Company assessmentDirect peers
- Redwood Research: AI safety research organization working on alignment and oversight of advanced AI. Comparable as a peer AI safety research shop, though Redwood focuses more on theoretical alignment than real-world deployment evaluations.
- Conjecture: AI safety company building control and oversight tools for advanced models. Comparable in mission and customer overlap with frontier labs, though Conjecture leans product/tooling while Andon emphasizes benchmark publication.
- SaferAI: AI safety organization producing evaluations and policy recommendations for frontier AI. Comparable as a peer in publishing third-party safety evidence aimed at labs and regulators.
- METR (Model Evaluation and Threat Research): Non-profit that builds rigorous capability evaluations for frontier AI systems (e.g., long-horizon agentic task benchmarks). Directly comparable to Andon Labs' Vending-Bench line in methodology and customer base (frontier labs and government).
- Apollo Research: Independent AI safety evaluation lab focused on scheming, deception, and high-risk behavior in frontier models. Closest direct peer to Andon Labs' safety-evaluation thesis, though Apollo leans on red-team probing while Andon emphasizes long-horizon real-world deployment.
Emerging players
- Holistic AI: AI governance, risk, and compliance platform serving enterprises subject to the EU AI Act and similar regulations. Adjacent in the AI safety/GRC market though focused on enterprise compliance rather than frontier-model evaluation.
- Credo AI: AI governance and audit platform helping enterprises document AI risk and compliance. Comparable to Andon's governance framing but aimed at enterprise buyers rather than frontier-model developers.
- Patronus AI: Commercial LLM evaluation platform offering safety, hallucination, and quality scoring. Comparable customer base (enterprises deploying LLM agents) but narrower focus on production LLM quality rather than long-horizon agent safety.
Others
- Arize AI: ML/LLM observability platform providing drift, quality, and safety monitoring for models in production. Adjacent infrastructure that complements but does not directly compete with Andon's frontier-safety benchmarks.
Broad incumbents
- Scale AI: Large incumbent in AI data infrastructure and evaluation, including its Safety, Evaluations, and Alignment (SEA) unit working with frontier labs. Overlaps with Andon's benchmarking work at much greater scale and broader scope.
Market position
Strengths4 records
Weaknesses4 records
Competitive moat5 records
Key risks6 records
Key highlights5 records
Customer concentration
Andon Labs social profiles
Digital presenceAndon Labs financial estimates
Financial estimateRevenue estimate
Valuation estimate
Andon Labs leadership team
Management profileNumber of profiles
Profiles2 records
Andon Labs funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Andon Labs M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Andon Labs
What does Andon Labs do?
Andon Labs develops and publishes AI safety benchmarks (Vending-Bench, Blueprint-Bench, Butter-Bench) for evaluating frontier AI model capabilities, and operates Safe Autonomous Organizations where AI agents run real-world businesses without human oversight. Revenue is generated through research collaborations with AI labs and direct sales of consumer products including the Andon Radio hardware and branded merchandise.
Is Andon Labs a public or private company?
Andon Labs is a private company. It is classified as venture growth investor backed and is currently operating.
When was Andon Labs founded?
Andon Labs was founded in 2023. It employs 1 to 10 people.
Where is Andon Labs based?
Andon Labs is headquartered in London, United Kingdom, in the Europe region.
How does Andon Labs make money?
Two revenue lines are on record. Product Sales are the primary driver. The others are research and Benchmarking Services.
Who are Andon Labs's main competitors?
Direct peers on record are Redwood Research, Conjecture, SaferAI, METR (Model Evaluation and Threat Research) and Apollo Research. Emerging players are Holistic AI, Credo AI and Patronus AI. Arize AI is listed as an others. Scale AI is listed as a broad incumbent.
Does Andon Labs have an API?
No public API is recorded for Andon Labs.
What industry is Andon Labs in?
Andon Labs's product category is AI Safety Research and Evaluation. Its primary akta.pro industry code is HDAEANAF, AI Observability, Monitoring & Evaluation Platforms (Drift, Quality, Safety), with a secondary code of HDAAAMAL, Safety & Alignment Evaluation (red-teaming, harmful capability testing).