| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
A collection of corpora for named entity recognition (NER) and entity recognition tasks. These annotated datasets cover a variety of languages, domains and entity types.
Data repository for pretrained NLP models and NLP corpora.
A collaborative catalog of NLP resources for Indic languages
微信公众号语料库
Links to Russian corpora + Python functions for loading and parsing
Official source for spanish Language Models and resources made @ BSC-TEMU within the "Plan de las Tecnologías del Lenguaje" (Plan-TL).
A web-based engine for creating and annotating textual corpora
Open Korean NLP Dataset Curation for the Users All Around the Globe
CrossNER: Evaluating Cross-Domain Named Entity Recognition (AAAI-2021)
The Self-dialogue Corpus - a collection of self-dialogues across music, movies and sports
Unannotated Spanish 3 Billion Words Corpora
Automatic categorization of documents, consists in assigning a category to a text based on the information it contains. We'll follow different approach of Supervised Machine Learning.
A curated list of resources dedicated to Natural Language Processing (NLP) of Cantonese | 粵語 NLP
An advanced, extensible web front-end for the Manatee-open corpus search engine
An R package for dynamic exploration of text collections
[NLPCC 2023] CCAE: A Corpus of Chinese-based Asian Englishes
Named Entity Recognition for biomedical entities
Tools for filtering and cleaning parallel and monolingual corpora for machine translation and other natural language processing tasks.
A comprehensive list of annotated training datasets classified by use case.
Add a description, image, and links to the corpora topic page so that developers can more easily learn about it.
To associate your repository with the corpora topic, visit your repo's landing page and select "manage topics."
| Back | FazBrowse Home | New Git URL |