Home / Product / Language coverage

Language coverage

Language is a filter, not an afterthought. Source language is recorded per article using ISO-639-3 codes, and retrieval behaviour is designed for scripts where exact matching does not work.

English-first design has a visible cliff

Most news APIs were built for English and extended outward. You can usually see where the extension stopped: filters exist for many languages, but the retrieval quality behind them degrades sharply, because the tokenizer, the ranking and the evaluation set were all tuned on English text.

The failure is quiet. A query in Thai or Hindi returns results, they look plausible, and nothing indicates that most of the relevant coverage was never a candidate. Teams typically discover this when a major story in a market they cover fails to trigger anything, and then cannot tell whether the problem is their query, their filter or the index.

Latin Elections Devanagari चुनाव Arabic انتخابات Han 选举 one filter, one query surface — direction and tokenization handled per language
Four scripts, one query surface. Direction and tokenization are handled per language.

Why script determines which retrieval mode to use

Keyword matching depends on assumptions that hold in English and fail elsewhere, so the right mode is a property of the language rather than a preference.

Script propertyExample languagesWhat breaks
No whitespace word boundaries Japanese, Chinese, Thai There is nothing for a keyword matcher to split on; token boundaries must be inferred.
Rich inflection and agglutination Turkish, Finnish, Russian A dictionary form matches few of the surface forms actually published.
Templatic roots, optional diacritics Arabic, Hebrew One root generates dozens of forms; prefix matching produces both false positives and negatives.
Unstable transliteration Hindi, Urdu, Korean Proper nouns appear in multiple scripts and spellings within a single news cycle.

The short version: prefer semantic or hybrid mode outside English, and reserve keyword mode for identifiers that publishers render consistently, such as tickers.

What a language count actually tells you

Advertised language counts across this category vary by an order of magnitude for corpora that are probably comparable, and no vendor publishes how the number is computed. A language with two outlets and a language with two hundred both increment the count by one.

So we publish the distribution instead. Per-language pages carry article volume, distinct outlet count and observed history, regenerated from the index each publication cycle. Languages without enough outlets to answer a real question get no page, rather than a page implying coverage that is not there.

Articles per outlet, by language Higher means volume is concentrated in fewer newsrooms. Lower means the same coverage comes from more independent publishers.
0 19 39 58 77 Indonesian: 77 articles per outlet (2,541 articles, 33 outlets) Indonesian 77 Chinese: 40 articles per outlet (6,415 articles, 159 outlets) Chinese 40 Turkish: 27 articles per outlet (3,393 articles, 126 outlets) Turkish 27 Greek: 26 articles per outlet (2,823 articles, 109 outlets) Greek 26 Vietnamese: 26 articles per outlet (1,741 articles, 67 outlets) Vietnamese 26 Serbian: 22 articles per outlet (1,127 articles, 51 outlets) Serbian 22 Hindi: 22 articles per outlet (875 articles, 39 outlets) Hindi 22 German: 19 articles per outlet (5,814 articles, 301 outlets) German 19 Romanian: 15 articles per outlet (1,766 articles, 119 outlets) Romanian 15 Russian: 14 articles per outlet (4,253 articles, 296 outlets) Russian 14 Croatian: 14 articles per outlet (677 articles, 48 outlets) Croatian 14 Spanish: 12 articles per outlet (5,216 articles, 450 outlets) Spanish 12 Ukrainian: 11 articles per outlet (1,323 articles, 118 outlets) Ukrainian 11 Portuguese: 11 articles per outlet (1,108 articles, 100 outlets) Portuguese 11 Arabic: 10 articles per outlet (941 articles, 94 outlets) Arabic 10 Italian: 9 articles per outlet (4,402 articles, 503 outlets) Italian 9 Albanian: 9 articles per outlet (845 articles, 91 outlets) Albanian 9 French: 8 articles per outlet (2,362 articles, 303 outlets) French 8 English: 7 articles per outlet (19,792 articles, 2920 outlets) English 7 Polish: 7 articles per outlet (887 articles, 125 outlets) Polish 7 Dutch: 6 articles per outlet (502 articles, 81 outlets) Dutch 6
View as table
LanguageArticlesOutletsPer outlet
Indonesian 2,541 33 77
Chinese 6,415 159 40
Turkish 3,393 126 27
Greek 2,823 109 26
Vietnamese 1,741 67 26
Serbian 1,127 51 22
Hindi 875 39 22
German 5,814 301 19
Romanian 1,766 119 15
Russian 4,253 296 14
Croatian 677 48 14
Spanish 5,216 450 12
Ukrainian 1,323 118 11
Portuguese 1,108 100 11
Arabic 941 94 10
Italian 4,402 503 9
Albanian 845 91 9
French 2,362 303 8
English 19,792 2920 7
Polish 887 125 7
Dutch 502 81 6

Outlet spread is the number to watch. Volume tells you how much was published; spread tells you how many independent voices you are actually reaching. A language served by three prolific publishers will out-produce one served by fifty modest ones and give you a far narrower picture.

Languages with a coverage page

A language appears here when the index carries enough of it, from enough distinct outlets, to answer a real question — and when we have something to say about querying it beyond the numbers. Languages that clear only one of those two bars deliberately have no page.

LanguageScriptArticlesOutletsCode
English · English Latin 20k 2920 eng
Chinese · 中文 Han (Simplified and Traditional) 6.4k 159 zho
German · Deutsch Latin 5.8k 301 deu
Spanish · Español Latin 5.2k 450 spa
Italian · Italiano Latin 4.4k 503 ita
Russian · Русский Cyrillic 4.3k 296 rus
Turkish · Türkçe Latin with diacritics 3.4k 126 tur
Greek · Ελληνικά Greek 2.8k 109 ell
Indonesian · Bahasa Indonesia Latin 2.5k 33 ind
French · Français Latin 2.4k 303 fra
Romanian · Română Latin with diacritics 1.8k 119 ron
Vietnamese · Tiếng Việt Latin with tone marks 1.7k 67 vie
Ukrainian · Українська Cyrillic 1.3k 118 ukr
Serbian · Српски / Srpski Cyrillic and Latin 1.1k 51 srp
Portuguese · Português Latin 1.1k 100 por
Arabic · العربية Arabic 941 94 ara
Polish · Polski Latin with diacritics 887 125 pol
Hindi · हिन्दी Devanagari 875 39 hin
Albanian · Shqip Latin 845 91 sqi
Croatian · Hrvatski Latin 677 48 hrv
Dutch · Nederlands Latin 502 81 nld

Figures are provisional — measured from the public source the index is built from rather than from the index itself, pending a snapshot against the API.