Home / Product / Language coverage
Language coverage
Language is a filter, not an afterthought. Source language is recorded per article using ISO-639-3 codes, and retrieval behaviour is designed for scripts where exact matching does not work.
English-first design has a visible cliff
Most news APIs were built for English and extended outward. You can usually see where the extension stopped: filters exist for many languages, but the retrieval quality behind them degrades sharply, because the tokenizer, the ranking and the evaluation set were all tuned on English text.
The failure is quiet. A query in Thai or Hindi returns results, they look plausible, and nothing indicates that most of the relevant coverage was never a candidate. Teams typically discover this when a major story in a market they cover fails to trigger anything, and then cannot tell whether the problem is their query, their filter or the index.
Why script determines which retrieval mode to use
Keyword matching depends on assumptions that hold in English and fail elsewhere, so the right mode is a property of the language rather than a preference.
| Script property | Example languages | What breaks |
|---|---|---|
| No whitespace word boundaries | Japanese, Chinese, Thai | There is nothing for a keyword matcher to split on; token boundaries must be inferred. |
| Rich inflection and agglutination | Turkish, Finnish, Russian | A dictionary form matches few of the surface forms actually published. |
| Templatic roots, optional diacritics | Arabic, Hebrew | One root generates dozens of forms; prefix matching produces both false positives and negatives. |
| Unstable transliteration | Hindi, Urdu, Korean | Proper nouns appear in multiple scripts and spellings within a single news cycle. |
The short version: prefer semantic or hybrid mode outside English, and reserve keyword mode for identifiers that publishers render consistently, such as tickers.
What a language count actually tells you
Advertised language counts across this category vary by an order of magnitude for corpora that are probably comparable, and no vendor publishes how the number is computed. A language with two outlets and a language with two hundred both increment the count by one.
So we publish the distribution instead. Per-language pages carry article volume, distinct outlet count and observed history, regenerated from the index each publication cycle. Languages without enough outlets to answer a real question get no page, rather than a page implying coverage that is not there.
View as table
| Language | Articles | Outlets | Per outlet |
|---|---|---|---|
| Indonesian | 2,541 | 33 | 77 |
| Chinese | 6,415 | 159 | 40 |
| Turkish | 3,393 | 126 | 27 |
| Greek | 2,823 | 109 | 26 |
| Vietnamese | 1,741 | 67 | 26 |
| Serbian | 1,127 | 51 | 22 |
| Hindi | 875 | 39 | 22 |
| German | 5,814 | 301 | 19 |
| Romanian | 1,766 | 119 | 15 |
| Russian | 4,253 | 296 | 14 |
| Croatian | 677 | 48 | 14 |
| Spanish | 5,216 | 450 | 12 |
| Ukrainian | 1,323 | 118 | 11 |
| Portuguese | 1,108 | 100 | 11 |
| Arabic | 941 | 94 | 10 |
| Italian | 4,402 | 503 | 9 |
| Albanian | 845 | 91 | 9 |
| French | 2,362 | 303 | 8 |
| English | 19,792 | 2920 | 7 |
| Polish | 887 | 125 | 7 |
| Dutch | 502 | 81 | 6 |
Outlet spread is the number to watch. Volume tells you how much was published; spread tells you how many independent voices you are actually reaching. A language served by three prolific publishers will out-produce one served by fifty modest ones and give you a far narrower picture.
Languages with a coverage page
A language appears here when the index carries enough of it, from enough distinct outlets, to answer a real question — and when we have something to say about querying it beyond the numbers. Languages that clear only one of those two bars deliberately have no page.
| Language | Script | Articles | Outlets | Code |
|---|---|---|---|---|
| English · English | Latin | 20k | 2920 | eng |
| Chinese · 中文 | Han (Simplified and Traditional) | 6.4k | 159 | zho |
| German · Deutsch | Latin | 5.8k | 301 | deu |
| Spanish · Español | Latin | 5.2k | 450 | spa |
| Italian · Italiano | Latin | 4.4k | 503 | ita |
| Russian · Русский | Cyrillic | 4.3k | 296 | rus |
| Turkish · Türkçe | Latin with diacritics | 3.4k | 126 | tur |
| Greek · Ελληνικά | Greek | 2.8k | 109 | ell |
| Indonesian · Bahasa Indonesia | Latin | 2.5k | 33 | ind |
| French · Français | Latin | 2.4k | 303 | fra |
| Romanian · Română | Latin with diacritics | 1.8k | 119 | ron |
| Vietnamese · Tiếng Việt | Latin with tone marks | 1.7k | 67 | vie |
| Ukrainian · Українська | Cyrillic | 1.3k | 118 | ukr |
| Serbian · Српски / Srpski | Cyrillic and Latin | 1.1k | 51 | srp |
| Portuguese · Português | Latin | 1.1k | 100 | por |
| Arabic · العربية | Arabic | 941 | 94 | ara |
| Polish · Polski | Latin with diacritics | 887 | 125 | pol |
| Hindi · हिन्दी | Devanagari | 875 | 39 | hin |
| Albanian · Shqip | Latin | 845 | 91 | sqi |
| Croatian · Hrvatski | Latin | 677 | 48 | hrv |
| Dutch · Nederlands | Latin | 502 | 81 | nld |
Figures are provisional — measured from the public source the index is built from rather than from the index itself, pending a snapshot against the API.