Company News Retrieval: Attribution Precision, Coverage, and Cost
July 2026 | 10 providers
This is an open, reproducible benchmark of company-news retrieval. It tests ten providers the way an autonomous agent would actually use them, by asking three questions: does the provider resolve the right company, does it actually surface the relevant news, and what does that cost the agent in tokens, dollars, and latency?
133 companies (private & public, global) · Jun 6–Jul 6, 2026 · 10 providers, 4 categories · 71,408 articles scored · harness & company list published
71,408
Articles scored
Every returned article judged
1262
Distinct events
Checked for coverage
133
Companies queried
108 private · 25 public
Players Covered
Overall retrieval quality is reported as F1, the harmonic mean of company+ news attribution precision and story recall.
The measure is used here because the two components fail in opposite directions: permissive attribution inflates recall by admitting articles about other companies, while strict attribution suppresses coverage of legitimate mentions.
A single-metric ranking on either component alone would therefore reward a degenerate strategy.
TL;DR
Across the benchmark, akta.pro reaches the highest overall performance in the set at the lowest cost per correct article.
Pick a view to see how the field stacks up.
quality (F1) against effective cost, log scale · top-left is best
Full definitions in 'Methodology & definitions'. The full leaderboard and per-metric tables are in 'Results'.
Six findings from the benchmark, framed for what matters to agent infrastructure, not what makes for interesting human reading:
Best overall quality
F1 of 81.3, 18.8 points clear of the next provider. akta.pro wins on the combination the others trade off against each other.
Highest entity resolution
93% of returned news is about the requested company, the best in the field, while story recall stays at 72.7%. Precision is not bought by returning less.
Precision is the whole game
Overall precision of 92%. The entire spread between providers is “is it the right company”, where the aggregators land at just 32 to 58%.
Lowest effective cost
$0.50 per 1,000 correct articles, at only 8% wrong-company waste. The aggregators run 61 to 69% waste that an agent pays for on every call.
A platform, not a name-matcher
Every article comes entity-resolved, classified against an 80+ event taxonomy, tagged by industry and geography, scored, and summarised. None of that was scored here.
Open and reproducible
The harness and full company list are published, so every number can be checked or re-run. Industry, open-ended, and monitoring benchmarks follow.
Read the top-left of the F1-vs-cost view. A news API judged as agent infrastructure has to clear two bars a human reader never imposes: is each result about the right company, and is what comes back worth the tokens it costs.
akta.pro is the only provider that clears both. The sections that follow measure attribution and coverage first, then what correct signal actually costs.
Why we ran this
News APIs are usually measured for people: relevance, freshness, breadth. Agents change the bar. An agent monitoring “Aura” or “Bloom” cannot glance at a headline and discard the wrong homonym; it ingests whatever the API returns, pays for every token, and acts on it.
So the metrics that matter shift to three questions: is each result about the right company, how much of the real news does it find, and what does correct signal cost in tokens, dollars, and latency. That is especially hard in private markets, where entities are ambiguous, coverage is thin, and mainstream news APIs were never built to disambiguate them.
akta.pro built an open, reproducible benchmark focused specifically on company news and ran its Company News API against nine alternatives across three competing approaches, namely traditional news platforms, agentic web-search APIs, and frontier LLMs with web access, on one identical harness, scoring every returned article with a neutral third-party judge. The harness and the company list are published so anyone can re-run or contest the results.
This is a self-benchmark: akta.pro authored and ran it. The harness, judges, and scoring rules are transparent and the query set is public, but the choice of companies and metric emphasis is ours. We think it holds up; we would rather you check it than take our word for it.
Setup
Providers & categories
Ten providers across four categories. We deliberately picked the tools an agent builder would actually reach for today, including several we expected to struggle on company attribution, because the point is to measure the gap honestly rather than to stack the deck.
News platforms
Perigon, NewsAPI (Event Registry), and SerpAPI (Google News). These are the incumbents for programmatic news: fast, high-volume, cheap per article, with topic and recency filters. They shine at breadth and freshness. Their expected weakness, and the reason we included them, is company disambiguation: they match text and topics, not resolved entities, so a query for a private company tends to pull in namesakes and larger public homonyms.
Agentic search
Exa (neural web search) and Parallel (agent-oriented web search). These are the new-wave retrieval APIs built for agents; they excel at open-web discovery and semantic recall. We included them because they are increasingly the default “get me info” call inside agent stacks, but they are optimised for the public web generally, not for resolving a specific company's news.
LLMs with web access
GPT-5.5, GPT-5.4 mini, Claude Opus 4.8, and Claude Sonnet 4.5, each with web search enabled. They represent the tempting shortcut, “just ask the model.” They reason about entities well, so we expected strong precision; we also expected low recall, high latency, and high cost, because a model browsing on demand will not exhaustively enumerate a month of news. Included to quantify exactly that trade.
The subject under test
akta.pro, a purpose-built company-news API with entity resolution across 20M+ parent companies, subsidiaries, trading names, and namesakes, curated global sources, and structured enrichment on every article.
Query set
133 companies, chosen to be a diverse, representative slice of the companies that actually exist: 108 private and 25 public, spanning 5 regions and 12 countries, with founding years from 1813 to 2025, across many industries and sizes. It intentionally over-weights the hard cases, namely private and non-US companies, where mainstream news sources are thin and disambiguation is hardest.
To make the entity-resolution test real, for a number of the private companies we deliberately included known namesake entities, unrelated companies that share a name or trade under a confusingly similar one, so a provider only scores by returning the right one. The full company list is published, both for transparency and as a starting point for a shared company-news benchmark others can build on.
Harness & ground truth
Every provider runs through one six-stage pipeline over a fixed one-month window (2026-06-06 to 2026-07-06), fetching up to 300 articles per company. Each company is processed independently; one failure never halts the run, and every HTTP and model call is logged for audit. The pipeline deliberately uses two different tools for two different jobs, an LLM where the task is judgment and a live web search where the task is verification:
- LLM-as-judge (Gemini, a neutral third party, never a provider under test) handles the judgment calls: for each article, is this genuine news and is it about this specific company; then it clusters the relevant articles into distinct stories, and maps each article back to the story it belongs to. These are semantic decisions where a model, grounded on the company's profile, is the right instrument.
- Live web search (an independent SerpAPI lookup) handles verification: each candidate story is re-searched on the open web and cross-checked to confirm it genuinely broke inside the analysis window, filtering out stale or resurfaced events. The set of validated stories becomes the ground-truth denominator for recall, so recall is measured against evidence, not against any one provider's output.
Ground truth is anchored on firmographics. Each of the 133 companies carries a firmographic profile (canonical name, website, headquarters, founding year, industry, and known aliases) drawn from the published company file.
The relevance judge is given that profile, so “is this article about the company” is decided against real firmographic identity (the right Bloom, at the right domain, in the right country), not a string match. That grounding is what makes company-entity precision a meaningful score rather than a keyword hit-rate.
Overview
We queried every provider with the same 133 company names and collected all returned articles over a 30-day window. Each article was judged for relevance and attribution accuracy by a combination of LLM scoring and human spot-checks. Cost was measured end-to-end including any token overhead required to parse and filter provider output.
Methodology
For each company we constructed a ground-truth set of relevant news events drawn from SEC filings, press releases, and manually verified sources. Providers were queried with identical inputs and scored on four metrics:
- F1 score — harmonic mean of precision and recall
- Entity accuracy — fraction of articles correctly attributed to the right company
- Recall — share of real stories surfaced by the provider
- Cost — USD per 1,000 accurately retrieved articles
All queries were run between June 6 and July 6, 2026. The full harness and company list are published in the repository linked at the bottom of this page.
Results
All ten providers on the headline metrics, sorted by F1; the best value in each column is marked. Definitions for every metric are in the ‘Methodology & definitions’ section.
Precision
Almost every provider is good at the easy half: “is this news” runs 64% to 99% across the board. The entire spread in overall precision comes from the hard half, is it about the right company.
The LLMs do well here (81% to 88% entity precision) because they reason about the entity before answering, but that same caution is why they return so little (see recall, below). The aggregators collapse on this axis (32% to 58%): they match strings and topics, so a query for a private company drags in namesakes, larger public homonyms, and unrelated brands.
akta.pro leads at 93% because entity resolution is the product: each candidate is resolved across subsidiaries, trading names, and namesakes and checked against the company's firmographic profile before it is ever returned.
Coverage & quality
Recall and precision pull against each other for everyone except akta.pro. Perigon and Exa post the highest raw recall (69% to 74%) by returning nearly everything, and pay for it in precision. The LLMs sit at the opposite pole (29% to 56% recall) because they surface only the most prominent handful of stories.
akta.pro is the only provider that is simultaneously near the top on recall (72.7%) and clearly first on precision, which is why its F1 leads by 18.8 points. It also surfaces 31 stories that no other provider found, more than any competitor, which for deal sourcing and monitoring is often the entire point of running a query.
Cost & efficiency
List price is misleading for an agent. Two costs matter: what you pay per article returned, and what you pay per article that is actually correct. Sorted by F1 delivered per dollar.
The aggregators look cheap per article returned ($0.23 to $0.37 per 1,000), but 40% to 70% of those articles are about the wrong company, so their cost per correct article runs 1.4 to 3 times higher, and every wrong article still burns context and spend on the agent's side.
akta.pro's list price is not the lowest, but because 92% of what it returns is correct, its effective cost ($0.50 per 1,000 correct) is the lowest in the field, and it delivers far more F1 per dollar than anything else tested. The LLMs are in a different regime entirely, $2.85 to $76 per 1,000 correct, which is fine for reasoning over a handful of signals and untenable for retrieval at volume.
Figure 2. Where each provider's returned tokens go. akta.pro delivers 78% effective tokens at 8% wrong-company waste; the news aggregators waste 61% to 69% of every token returned. For an agent that pays per token of context, that waste is not free: it is the dominant cost.
Latency
Latency compounds for an agent, which may call a news API many times inside a single task. We report the latency of successful responses.
The raw news APIs are fastest (Perigon 2.1s, NewsAPI 4.4s) because they do little beyond a lookup. akta.pro's 3.1s median reflects full entity resolution and enrichment happening inline, slower than a bare lookup but an order of magnitude faster than any LLM. The real gap is to the LLM providers (16s to 110s at the median, tails into minutes): an agent that calls a frontier model to fetch news pays that latency and the token cost on every hop.
For a retrieval-and-enrichment layer that runs disambiguation on every request, a few seconds is the right place to be.
Figure 3. Median (p50) latency of successful responses. Raw news APIs are fastest; LLM web-search is slowest by far; akta.pro lands close to the raw APIs while doing considerably more work per query.
Where akta.pro wins
The case for akta.pro is not that the other tools are bad; it is that they are built for different goals. The data below shows where the separation is widest and why it matters for an agent stack specifically.
Where akta.pro leads the field
Overall quality: F1 81.3, +18.8 over the next provider.
Entity resolution: 93.0%, +5.2pts over the next, +35 to 61pts over the aggregators.
Overall precision: 92.2%, best in field.
Recall among high-precision providers: 72.7%, top-tier overall.
Exclusive stories: 31, more than any other provider.
Effective cost and token waste: $0.50 per 1K correct at 8% waste, both best.
Where other tools edge ahead: the raw news platforms (Perigon, NewsAPI) post slightly higher raw recall by returning far more articles, including many about the wrong entity. Aggregator recall beats akta.pro's by fractions of a point at 3–4× the wrong-company waste. That trade-off is fine for human readers who can skim; it is not fine for an agent that pays per token and acts on the output.
The LLMs (GPT-5.5, Claude Opus) post high entity precision (81–88%) but low recall (29–56%). They reason carefully before returning a result, which is valuable for single-shot synthesis, but they will not enumerate a month of news exhaustively. The scatter plot below makes the separation visible.
Precision vs recall. LLMs cluster top-left (attribute well, find little); news platforms and agentic search fall bottom-right (find volume, mis-attribute it). Only akta.pro reaches the top-right.
Precision vs recall for all ten providers. Top-right quadrant (recall ≥ 50%, precision ≥ 60%) is the “find it and attribute it” zone. akta.pro is the only provider that lands there. The LLMs cluster upper-left (precise but thin); the aggregators cluster lower-right (broad but noisy).
Where each fits in an agent stack
Every tool in this benchmark is genuinely useful. The question is what job it is suited for.
Best raw recall, lowest latency. Ideal for broad content discovery where an LLM will filter downstream. Not suitable as a standalone company-news feed — 66% of returned articles are about the wrong entity.
Cheapest per article returned ($0.23/1K). Works for keyword-matching pipelines, media monitoring on well-known brands, and topic feeds. Entity disambiguation must be handled elsewhere.
Best of the news platforms on precision (54%) and F1 (52%). Good middle ground for public-company monitoring where entity ambiguity is low, and fast enough for near-real-time alerts.
Strong open-web recall (69%) via neural search. Best suited to exploratory research tasks — “find everything written about X on the web” — not structured company-news pipelines where attribution accuracy matters.
Reasonable recall and mid-tier precision. Designed for agent workflows that need real-time web context across many topics simultaneously. Cost-competitive with other agentic options.
Highest F1 of the LLM tier (62.5) with strong entity reasoning. Good for tasks requiring synthesis of a small number of highly relevant stories. Not suitable for exhaustive retrieval — recall is 56% and latency averages 110 seconds.
Highest entity precision of any provider (88%) but lowest recall (29%) and highest cost ($76/1K correct). Best fit: a final-stage citation verifier or high-stakes synthesis where getting the right company matters more than finding everything.
Purpose-built for the use case: structured company-news retrieval with entity resolution, classification, and enrichment inline. The right default for any agent that needs accurate, cost-efficient, low-latency company news at volume.
Use akta.pro with an agent, not instead of one.
akta.pro returns entity-resolved, enriched, structured news. What it does not do is reason across that news, synthesise it into a decision, or take action. It is a retrieval and enrichment layer — the right input to an agent, not a replacement for one.
The providers in the LLM tier do the opposite: they reason well but retrieve thinly. The most capable agent stacks will use a purpose-built retrieval API at the fetch layer and a frontier model at the reasoning layer, not one thing doing both jobs badly.
Beyond retrieval: the enrichment layer
This benchmark measured one thing, retrieval quality. But for an agent, what comes back with each article matters as much as whether it is the right article. Here akta.pro is a platform rather than a feed: every news result arrives pre-structured, so the agent does not have to parse prose or re-derive metadata. Straight from the Company News API on each article:
And that layer holds because news sits on top of a native company database with entity resolution and search built in. Resolving an article to a stable ID across subsidiaries, trading names, and namesakes needs a company graph to resolve against, which news-only providers do not have.
What this means
For agent builders and private-markets teams, the practical takeaway is that list price and raw recall are the wrong things to optimise. An agent does not consume articles, it consumes correct, correctly-attributed, structured context, and it pays for every token. On that basis akta.pro returns 78% usable content at 8% waste while the aggregators invert that ratio, which is why akta.pro's effective cost is the lowest in the field even though its list price is not.
The workflows this changes are concrete: deal sourcing (surfacing the 31 private-company stories no other source found), competitive and risk monitoring (not paging an agent for a namesake), and agent pipelines (clean, entity-resolved, pre-classified signal that does not need a second disambiguation pass). That is akta.pro's wedge, agent-native and private-company-specialised, and it is exactly the gap this benchmark was built to measure.
Methodology & definitions
Metric definitions
- News precision: share of returned articles that are genuine news (vs. non-news pages).
- Company-entity precision (name resolution): share of returned articles that are about the requested company rather than a namesake, subsidiary, or unrelated entity.
- Overall precision: share of returned articles that are both genuine news and about the right company.
- Recall: share of the deduplicated, window-validated ground-truth story set that a provider covers.
- F1: harmonic mean of overall precision and recall; the single measure of retrieval quality.
- Exclusive stories: distinct validated stories surfaced by a single provider and no other.
- Cost per 1k returned / correct: provider cost per thousand articles returned, and per thousand that are correct (genuine news about the right company).
- F1 per $ (efficiency): F1 delivered per dollar of spend; higher is better.
- Effective / wasted tokens: share of returned content, by token count, that is correct company news vs. about the wrong company.
- Median latency: p50 of successful (2xx) per-query response times; p90/p99/mean.
Limitations
- Vendor self-benchmark. akta.pro designed and ran it. The harness, judges, and company list are open, but company selection and metric emphasis are our choices and favour our strengths.
- Recall is against the observable universe. The denominator is stories that at least one provider surfaced and that passed independent window-validation, not every event that occurred. It is a fair relative measure, not an absolute one.
- LLM-as-judge, not human adjudication. Scoring is automated with neutral third-party models, consistent and cheap at volume, but a model’s judgment.
- Token counts are a proxy. A single shared tokenizer approximates all providers, correct for relative comparison, not any provider’s exact billing.
- Directional, not a controlled A/B. Providers differ in interface, rate limits, and defaults; we used a reasonable configuration for each.
Reproducibility. The harness and the full company list are published; we will share per-company scores on request. If you find a flaw in the methodology, tell us. We would rather fix it than defend it.
Next steps
- More of the surface. This study benchmarks company news. akta.pro's news API also does industry and open-ended queries, curated newsfeeds, and continuous monitoring; we will publish benchmarks on each of those in the coming days.
- A standing benchmark. We will refresh this study next quarter as providers (including us) change, and publish the results regardless of how akta.pro places.
- Open to suggestions. The company list and harness are public precisely so the methodology can improve. Tell us what to add, which providers to include, or where the setup is unfair.
Try our data for free
Start with 25 free credits on us — no credit card required.