Skip to content

Who uses BlackLab?

A number of institutes and companies have chosen BlackLab for their corpus or text analysis projects.

If you're using BlackLab, let us know; we'd like to mention your project here!

  • The Dutch Language Institute (INT) is leading BlackLab development and uses it for their corpora, such as its Corpus of Contemporary Dutch (3+ billion tokens public version, 4.5+ billion internal), the historical Letters as Loot and Corpus Gysseling and various other corpora.
  • OpenSonar, a Dutch corpus, by the University of Tilburg in collaboration with the Dutch Language Institute. The look and feel of our corpus frontend was developed as part of this collaboration as well, for which we'd like to thank Martin Reynaert and Matje van de Camp.
  • Frisian Corpora, Frisian corpora ranging from runes to modern Frisian, mostly unrestricted material, Mid Frisian contains detailed linguistic annotations. Eduard Drenth of the Fryske Akademy has made some valuable contributions as well, such as initial support for the Saxon parser.
  • In Search of the Drowned: Testimonies and Testimonial Fragments of the Holocaust, edited and built by Gabor M. Toth in collaboration with the Yale Digital Humanities Laboratory.
  • CRIME (The Corpus of Recorded Investigative, Media, and Evidence-based Proceedings), Coats, Steven and Dana Roemling
  • CoANZSE Audio (The Corpus of Australian and New Zealand Spoken English)
  • CO.RA.PAN Radio Corpus (Corpus Radiofónico Panhispánico), Marburg University
  • COCO@NWU (Corpus Cooperative), North-West University, South-Africa
  • VIVA Korpusportaal, Virtuele Instituut Vir Afrikaans (South Africa)
  • SADiLaR corpus portal, South African Centre for Digital Language Resources
  • EarlyPrint, corpus of the early English print record from 1473 to the early 1700s, Northwestern University and Washington University in St. Louis.
  • Cosycat, corpus query and annotation interface used for the Mind-Bending Grammars project, University of Antwerp.
  • IKE, a knowledge extraction tool, developed at the Allen Institute for Artificial Intelligence. Allen AI has made very useful contributions as well, improving performance and fixing bugs.
  • Corpus of Spoken Hindi (Japan)
  • Alpheios Latin lemmatized texts
  • Arabic Digital Humanities. The right-to-left support in our corpus frontend came about thanks to this project.

Past users/contributors:

  • Lexion AI used BlackLab in their contract management system before being acquired by DocuSign. Lexion had an interestingly different use case from INT, with high query volume on many smaller indexes. We are grateful for their valuable contributions to the project.

Apache license 2.0