English news API
Query English coverage by keyword, meaning or both. Written in Latin . Pass eng as the language filter.
What the coverage looks like
English is the largest stream in the index by a wide margin, and that size is the thing to be careful about rather than the thing to celebrate. Because so many outlets publish in English, syndication is heavier here than anywhere else: a single wire report can appear under hundreds of bylines within an hour, and a naive query returns all of them ranked as independent results.
| Articles observed | 20k |
|---|---|
| Distinct outlets | 2920 |
| Articles per outlet | 7 |
| ISO-639-3 code | eng |
| Script | Latin |
Outlet spread is the number worth watching here, not volume. 2920 distinct publishers at roughly 7 articles each describes a very different information environment than the same total produced by a handful of prolific sources — and it is the difference between reaching independent reporting and reading one newsroom repeatedly.
In this sample the most active publisher is iheart.com, carrying
about 2% of English articles on its own, and the three busiest
— iheart.com, indiatimes.com, themarketsdaily.com — account for roughly 6% between
them. That is a comparatively flat distribution, which means volume here translates fairly directly into independent reporting rather than one newsroom publishing repeatedly.
Script and matching behaviour
Latin script with weak inflection and reliable whitespace boundaries, which is why keyword matching works better in English than in most languages and why so much tooling is silently built on the assumption that it always works. The remaining difficulty is vocabulary rather than morphology — British, American, Indian and Nigerian outlets describe the same event with materially different word choices.
How to query it
This is the one language where keyword mode is routinely defensible. It is also the language where story grouping matters most, because syndication volume is highest. If you take one action on English coverage, collapse to stories before doing anything else with the results.
Bound every request with a time range regardless of mode. Meaning-based retrieval always returns its best guesses, so an unbounded query produces a confident-looking result set anchored to nothing in particular. If results feed a model rather than a person, collapse to stories as well — one event covered by thirty outlets should cost you one item, not thirty.
Recent headlines in English
Pulled from the index on 2026-07-31, shown in their original script so you can see what the API actually returns rather than a translated approximation of it.
| Headline | Outlet |
|---|---|
| Cabinet approves PM-Surya Sarovar Yojana, extends PM-Kisan for 5 years, clears revamped Khelo India scheme worth Rs 36,441 crore | indiatimes.com |
| Total Voting Rights | Company Announcement | investegate.co.uk |
| Federal arts body can't terminate artists over conduct | singletonargus.com.au |
| Who's driving a bookmobile full of Epstein files? The quiz knows! | wysu.org |
| New York Sues Kalshi, Calls It Unlicensed Gambling Market | iheart.com |
| No to divide | thehitavada.com |
Outlets publishing in English
The most active publishers in this language, by article count. Use these as an outlet filter when you want to pin a particular market or editorial position rather than taking the whole language.
iheart.com · indiatimes.com · themarketsdaily.com · dailypolitical.com · aol.co.uk · yahoo.com · tickerreport.com · finanznachrichten.de · globalsecurity.org · investegate.co.uk · express.co.uk · manilatimes.net
What catches people out
The trap in English is mistaking syndication for corroboration. Because the wire services feed hundreds of English-language outlets, a single agency report can produce a wall of results that looks like independent confirmation from many newsrooms and is in fact one story republished. Teams building credibility scoring on source counts get this backwards on a regular basis, rating a heavily syndicated wire item above an exclusive from a single serious newsroom. Count distinct stories and distinct originating outlets, never raw article volume.
A worked case
For coverage of a company announcement, query in hybrid mode over a two-day window and collapse to stories. A typical result is one story carrying forty or more outlets, which is the shape you want: the event once, with the evidence of how widely it travelled attached. Running the same query without collapsing returns those forty as separate rows, and anything downstream that weights by frequency will now believe this was forty times more important than it was.
Where this language leads
Strong across every subject area, with unmatched depth in technology, finance and international affairs. The weakness is the inverse of the strength: English-language coverage of a non-English-speaking country is usually thinner, later and more summarised than that country's own press.
Who queries it, and why
Queried by almost everyone, which is exactly why it should rarely be queried alone. If your product serves users outside the anglophone world, English coverage of their region is a summary written for outsiders, and treating it as the full picture builds a blind spot into the product from day one.
Reading these numbers
These figures are provisional. They are measured from the public source the index is built from rather than from the index itself, because the query endpoint is not yet reachable from the build. They are the right order of magnitude and the outlets and headlines are real, but they should not be quoted as coverage guarantees until a snapshot taken against the API replaces them.
A page exists for this language because the index carries enough English coverage from enough distinct outlets to answer a real question about it. Languages that do not clear that bar deliberately have no page rather than a page that implies coverage which is not there.