Quesma
- Company typePrivate
- Founded2023
- HeadquartersWarsaw, Poland
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
Quesma firmographics
Firmographics- Name
- Quesma
- Legal name
- Quesma Poland Sp. z o.o.
- Website
- https://quesma.com
- Company type
- Private
- Founded year
- 2023
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Ownership category
- akta.pro rank
Quesma industry classification
Industry- Product category
- AI Agent Evaluation and Benchmarking
- NAICS
- Software Publishers (5132), Custom Computer Programming Services (541511), Other Computer Related Services (541519)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Programming, Data Processing, Etc. (7370)
- akta.pro primary industry
- End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management) (HDAEANAA)
- akta.pro secondary industries
- Model Security Testing & Red Teaming (adversarial ML, jailbreaks) (HDAAAKAC), Model Development & Training Platforms (AutoML, Notebooks, Feature Stores) (HDAEANAB)
Keywords
Where Quesma is headquartered
LocationHeadquarters
- HQ city
- Warsaw
- HQ country
- Poland
- HQ region
- Europe
Offices1 record
Markets served
Quesma business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Marketing or Sales, Infrastructure
Revenue model
- AI Agent Evaluation Services: Provides independent verification and benchmarking services for AI labs and enterprise buyers seeking to evaluate AI agent capabilities. Likely subscription or project-based pricing.
- Quesma Charts (B2C Product): AI-powered chart generation tool using ggplot2 for professional data visualization. Alpha version launched for individual users.
Go-to-market motion1 record
Distribution channels2 records
Marketing channels5 records
Quesma product offering
Product offeringCore offering
Quesma provides independent evaluation and training services for the AI agent ecosystem, built on open-source benchmarks such as BinaryAudit, OTelBench, CompileBench, Tau², and SWE-Bench Pro. Its offering combines realistic multi-hour simulation environments with cheat-proof reward functions for reinforcement learning training of AI agents. The company also operates Quesma Charts, an open-source AI-powered data visualization tool.
Product overview
Quesma is a technology company that develops a suite of open-source AI benchmarking and testing tools. The portfolio includes OTelBench for evaluating LLMs on OpenTelemetry instrumentation, CompileBench for testing AI on compilation tasks, BinaryAudit for security analysis of binaries, Quesma Charts for AI-powered data visualization, Tau² for AI agent evaluation, and SWE-Bench Pro for testing AI on real GitHub issues. The company focuses on creating evaluation frameworks to assess AI model capabilities across software engineering, security, and instrumentation tasks.
Differentiator
Problem solved
Functional benefit
Brands
- Quesma Charts: AI-powered data visualization tool that turns CSV, Excel, or SQL data into clean, accurate, and professional visuals using prompts.
Products and services
- OTelBench Open-source benchmark that evaluates large language models on OpenTelemetry instrumentation tasks, testing 14 state-of-the-art models across 23 real-world tasks in 11 programming languages. Used by AI app developers and frontier labs to measure model capability on observability code generation.
- BinaryAudit Open-source benchmark testing AI agents on detecting hidden backdoors in compiled binaries of real open-source servers, proxies, and network infrastructure. Uses Ghidra and Radare2 for reverse engineering and static analysis. Aimed at AI security researchers and frontier labs.
- CompileBench Open-source benchmark testing LLMs on real-world compilation tasks such as dependency resolution, Makefile analysis, legacy toolchain support, cross-compilation, and resurrecting old code. Used to evaluate AI agent capability on software engineering build challenges.
- Tau²
Quantifiable outcome
- Claude Opus 4.6 achieved 49% detection rate on BinaryAudit (finding backdoors in binaries)
- +2 more outcomes
Companies that use Quesma
Customer profileSegments3 records
Ideal customer profiles3 records
Quesma technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability9 records
Feature4 records
Quesma partnerships and signals
Strategic signalScale indicators1 record
Recent moves6 records
Expansion highlights5 records
Quesma competitors and assessment
Company assessmentBroad incumbents
- Scale AI: Scale AI runs SWE-Bench Pro (which Quesma uses) and offers evaluation services for frontier AI models. Both companies provide AI benchmarking and enterprise verification, but Scale AI is far larger and operates as a broad data/eval infrastructure incumbent.
- Hugging Face: Hugging Face operates the Open LLM Leaderboard and extensive model evaluation infrastructure. Both companies publish open-source AI benchmarks and serve the developer community, though Hugging Face is a much broader model-and-data platform.
- Weights & Biases: Weights & Biases provides MLOps and model evaluation tooling as part of its broader platform. Both companies support AI developers in tracking, benchmarking, and improving AI model performance.
Direct peers
- Patronus AI: Patronus AI provides automated evaluation and testing for LLM applications, including agent evaluation. Both companies offer independent verification of AI capabilities for enterprise buyers and AI developers building production agents.
- Galileo (Galileo AI): Galileo provides AI evaluation and observability tools for LLM-powered applications and agents. Both Quesma and Galileo target AI app developers and enterprise buyers needing to measure and improve AI agent quality.
- Arize AI: Arize AI offers evaluation, observability, and monitoring for LLM applications. Both companies serve AI developers and enterprise teams deploying AI agents, providing independent quality measurement and performance tracking.
- Braintrust: Braintrust provides an AI evaluation platform for LLM applications with scoring, tracing, and benchmarking. Both companies enable AI app developers to measure quality and pick optimal models in a fast-changing landscape.
- Vellum AI: Vellum offers LLM evaluation, prompt testing, and benchmarking tooling for AI developers and enterprises. Both companies provide AI quality measurement and model selection capabilities.
Emerging players
- Langfuse: Langfuse offers open-source LLM observability and evaluation tooling. Both Quesma and Langfuse target developers building AI agents with open-source, developer-led distribution models.
- WhyLabs: WhyLabs provides AI observability and monitoring for LLM applications. Both companies focus on independent measurement of AI system behavior and serve enterprise buyers deploying AI in production.
Market position
Strengths4 records
Weaknesses5 records
Competitive moat4 records
Key risks6 records
Key highlights6 records
Customer concentration
Quesma social profiles
Digital presenceQuesma financial estimates
Financial estimateRevenue estimate
Valuation estimate
Quesma leadership team
Management profileNumber of profiles
Profiles6 records
Quesma funding detail
Funding detailFunding overview
Funding rounds1 record
Investors2 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Quesma M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Quesma
What does Quesma do?
Quesma provides independent evaluation and training services for the AI agent ecosystem, built on open-source benchmarks such as BinaryAudit, OTelBench, CompileBench, Tau², and SWE-Bench Pro. Its offering combines realistic multi-hour simulation environments with cheat-proof reward functions for reinforcement learning training of AI agents. The company also operates Quesma Charts, an open-source AI-powered data visualization tool.
Is Quesma a public or private company?
Quesma is a private company. It is classified as venture growth investor backed and is currently operating.
When was Quesma founded?
Quesma was founded in 2023. It employs 11 to 50 people.
Where is Quesma based?
Quesma is headquartered in Warsaw, Poland, in the Europe region.
How does Quesma make money?
Two revenue lines are on record. AI Agent Evaluation Services are the primary driver. The others are quesma Charts (B2C Product).
Who are Quesma's main competitors?
Broad incumbents on record are Scale AI, Hugging Face and Weights & Biases. Direct peers are Patronus AI, Galileo (Galileo AI), Arize AI, Braintrust and Vellum AI. Emerging players are Langfuse and WhyLabs.
Does Quesma have an API?
No public API is recorded for Quesma.
What industry is Quesma in?
Quesma's product category is AI Agent Evaluation and Benchmarking. Its primary akta.pro industry code is HDAEANAA, End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management), with a secondary code of HDAAAKAC, Model Security Testing & Red Teaming (adversarial ML, jailbreaks). Its NAICS code is 5132 and its SIC code is 7372.