Skip to main content
Nasjonalbiblioteket

N-grams from NBdigital

  • Datasets
  • Public access 

    Publicly available to everyone. Access may still require registration and an API key request, as long as anyone can request such registration and/or API keys.

    Read more about access levels here

  • Open data 

    The dataset is classified as public access and has at least one distribution with an approved open license.

Description

This resource contains n-grams - i.e. unigrams, bigrams and trigrams - from all books and newspapers that had been digitized at the National Library of Norway up to September 2013. The n-grams have been extracted from a material consisting of approximately 220,000 books and 540,000 newspapers.

The n-grams are available in two formats, CSV and SQlite: CSV is probably the most interesting format for most developers, because it is very easy to import these files into standard applications. The SQLite files contain indexed databases, which are used in the service NB N-gram. Users who want to contribute to the development of NB N-gram can download the source code on GitHub, and the SQLite databases from this page.

A word count by source (books/newspapers) and language variety (Bokmål/Nynorsk) is given in the json file.


Similar datasets

NST Pronunciation Lexicon for SwedishNasjonalbiblioteket
Public access
Grapheme-to-Phoneme Models for NorwegianNasjonalbiblioteket
Public access
SCARRIE LexiconNasjonalbiblioteket
Public access
ONOMASTICA Pronunciation LexiconNasjonalbiblioteket
Public access
N-grams from NBdigital 2021Nasjonalbiblioteket
Public access