Showing posts with label ressources. Show all posts
Showing posts with label ressources. Show all posts

Sunday, March 29, 2026

Romanian Language NLP Datasets

 Table of contents

Unlabeled text Corpora

The FuLG dataset is a comprehensive Romanian language corpus comprising
150 billion tokens, carefully extracted from Common Crawl. 

arXiv

Part of a large multilanguage corpus originated from Common Crawl.
It's a raw, unannotated corpus. It has roughly 50 GB of Romanian text
in 4.5 million documnets. For details check its homepage 
and the paper

arXiv Homepage

 Similar to Oscar, part of a multilanguage corpus also based on Common Crawl
 from 2018. Romanian text is 16GB large

arXiv Homepage

  Romanian language wikipedia dump. 
  A collection of varoius unannotated corpora collected around 2018-2019.
  Includes books, scraped newspapers and juridical documents  
  A collection of written and spoken text from various
  sources: Articles, Fairy tales, Fiction, History, Theatre, News
 Romanian national legilation from  1881 to 2021. The corpus
 includes mainly: governmental decisions, ministerial orders,
 decisions, decrees and laws.
 Automatically annotated for Named Entities

ACL Homepage

Mega-COV is a billion-scale dataset from Twitter for studying COVID-19. It is available in over 100+ languages, Romanian being one of them. Tweets need to be rehydrated

arXiv Medium

A corpus of Romanian tweets related to COVID and vaccination against COVID, created and collected between January 2021 and February 2022. It contains 19319 tweets.

Minutes of the Sittings of the Chamber of Deputies of Romania (2016-2018)
Unannotated corpus
contains 500k+ instances of speech from the parliament podium from
1996 to 2018. Sentence splitting and deduplication onm sentence level
have been applied as processing steps
Unannotated corpus
Romanian presidential discouses (1990-2020) split in 4 files
one for each president. Unannotated corpus
Monolingual Romanian corpus, including content from public websites related to culture

Monolingual (ron) corpus, containing 38063991 tokens and 854096 lexical types in the law domain.

Monolingual Romanian corpus, containing 360833 sentences (9064764 words) in the public administration domain.

The New Civil Procedure Code in Romanian (monolingual) comprising 297888 words.

The Romanian updated criminal code: text with law content.

news articles dataset from romanian newssites title, summary and article

multi-language corpus from online available news sources. It contains also 43mil words in Romanian language from Twitter, Blogs and Newspapers

Homepage

The Romanian novel collection for ELTeC, the European Literary Text Collection Sources: Biblioteca Metropolitana din Bucuresti, Biblioteca Universitara "Mihai Eminescu" din Iasi, Biblioteca Judeteana din Botosani, personal micro-collections uploaded on Zenodo under the following labels: "Hajduks Library"; "RomanianNovel Library"; "CityMysteries Library"; "BibliotecaDHL_Iasi"

Public dataset of 1447 manually annotated Romanian business-oriented emails. The corpus is annotated with 5 token-related labels, as well as 5 sequence-related classes

MDPI

The corpus consists of texts written by Romanian authors between 19th century and present, representing stories, short-stories, fairy tales and sketches. The current version contains 19 authors, 1263 full texts and 12516 paragraphs of around 200 words each, preserving paragraphs integrity.

A dataset containing 400 Romanian texts written by 10 authors The dataset contains stories, short stories, fairy tales, novels, articles, and sketches written by Ion Creangă, Barbu Ştefănescu Delavrancea, Mihai Eminescu, Nicolae Filimon, Emil Gârleanu, Petre Ispirescu, Mihai Oltean, Emilia Plugaru, Liviu Rebreanu, Ioan Slavici.

MDPI

891 Cooking Recipes in Romanian Language

Semantic Textual Similarity / Paraphrasing

Semantic Textual Similarity dataset for the Romanian language RO-STS contains 8,628 sentence pairs with their similarity scores

NeurIPS

A paraphrase corpus created from 10 different Romanian language Bible versions. The final dataset contains 904,815 similar records and 218,977 non matching records, totaling 1,123,927

Around ~100k examples of paraphrases. No clear explanation on how the dataset was built

A multi-language paraphrase corpus for 73 languages extracted from the Tatoeba database. It has ~ 2000 romanian phrases totaling 941 paraphrase groups.

ACL Homepage

Natural Language Inference

We introduce the first Romanian NLI corpus (RoNLI) comprising 58K training sentence pairs, which are obtained via distant supervision, and 6K validation and test sentence pairs, which are manually annotated with the correct labels. ACL

The repository seems to be just an attempt at starting to build the dataset

Summarization

Around ~72k Full texts and their summary. Source seems to be news websites. No description or explanation available

Dialect and regional speech identification

varied compilation of speech samples from five distinct regions of Romania, covering both urban and rural environments. Around 2800 records labeled with age, gender and type of dialect

arXiv

MOROCO: The Moldavian and Romanian Dialectal Corpus The MOROCO data set contains Moldavian and Romanian samples of text collected from the news domain. The samples belong to one of the following six topics: culture, finance, politics, science, sports, tech totaling over 32.000 labeled records

arXiv

Named Entity Recognition (NER)

The dataset contains 323k tokens of text, covering more than half of the 19th century (i.e., 1817) until the late part of the 20th century (i.e., 1990). The samples belong to one of the following four historical regions of Romania, namely Bessarabia, Moldavia, Transylvania, and Wallachia.

arXiv

Autorship Attribution

Sentiment Analysis

Dependency Parsing

Diacritics Restoration / Grammar Correction

Fake News / Clickbait / Satirical News

Offensive Language

manually annotated 4,052 comments on a Romanian local news website into one of the following classes: non-offensive, targeted insults, racist, homophobic, and sexist.

arXiv

4455 organic generated comments from Facebook live broadcasts annotated not binary offensive language detection tasks and for fine-grained offensive language detection

IEEE

4800 Romanian comments annotated with offensive text spans Offensive span detection

MDPI

3860 labeled hate speech records

Dataset consists of 5000 tweets, from which 924 were labeled as offensive (18.48 %) and 4076 tweets as non-offensive.

ACL

The corpus contains 39 245 tweets, annotated by multiple annotators, following the sexist label set of a recent study.

ACL

Questions and Answers

This dataset is just the translation of the gsm8k dataset. GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. There is no information on the quality of the translation

RoCode, a competitive programming dataset, consisting of 2,642 problems written in Romanian, 11k solutions in C, C++ and Python and comprehensive testing suites for each problem. The purpose of RoCode is to provide a benchmark for evaluating the code intelligence of language models trained on Romanian / multilingual text as well as a fine-tuning set for pretrained Romanian models.

arXiv

Romanian IT Dataset (RoITD) resembling SQuAD 1.1. RoITD consists of 9575 Romanian QA pairs formulated by crowd workers. QA pairs are based on 5043 articles from Romanian Wikipedia articles describing IT and household products. Of the total number of questions, 5103 are possible (i.e. the correct answer can be found within the paragraph) and 4472 are not possible (i.e. the given answer is a "plausible answer" and not correct)

The dataset comprises 102,646 high-quality QA pairs from real-world clinical records of 1,011 oncology patients (796 patients with breast cancer and 215 patients with lung cancer). The QA pairs are the results of a manual annotation process carried out by physicians specialized in oncology and radiotherapy RoMedQA includes 76,416 QA pairs about breast cancer patients and 26,230 about lung cancer patients, with questions grounded in medical case summaries (epicrises).

arXiv

Romanian legal MCQA dataset, comprising 10,836 questions from three examinations. Each entry essentially consists of a body in which a theoretical question is posed regarding a legal aspect, along with three possible answer choices labeled A, B, and C, out of which at mosttwo answers are correct.

ACL

Spelling, Dictionaries and Gramatical Errors

Synthetic dataset with ~1.9M records. Altered and correct statement as columns

Romanian Archaisms Regionalisms Lexicon containing ~ 1940 Word definitions

Romanian Rules for Dialects - 1940 regionalisms, meanings and the region of provenience

The dataset was developed mainly for speech processing applications, yet its applicability extends beyond this domain. RoLEX includes over 330,000 curated entries with information regarding lemma, morphosyntactic description, syllabification, lexical stress and phonemic transcription.

Cambridge

Automatic Speech Recognition

Underrepresented Speech Dataset from Open Data. Duration of the dataset is 4h 18m 55s. Distribution according to platforms: 83% of the content comes from YouTube, 12% from SoundCloud, and 5% from Vimeo Dataset covers primarily under-represented speech groups (outside the 19–29 male category)

Source: https://github.com/AndyTheFactory/romanian-nlp-datasets 

Thursday, July 25, 2024

Zotero Translators

ABC News Australia
ACLS Humanities EBook
ACLWeb
ACM Digital Library
ACS Publications
ADS Bibcode
AEA Web
AGRIS
AIP
AMS Journals
AMS MathSciNet (Legacy)
AMS MathSciNet
APA PsycNET
APN.ru
APS-Physics
APS
ARTFL Encyclopedie
ARTnews
ARTstor
ASCE
ASCO Meeting Library
ASTIS
ATS International Journal
Ab Imperio
Access Engineering
Access Medicine
Access Science
Adam Matthew Digital
Agencia del ISBN
Ahval News
Air University Journals
Airiti
Alexander Street Press
AllAfrica
Alsharekh
AlterNet
Aluka
Amazon
American Archive of Public Broadcasting
American Institute of Aeronautics and Astronautics
American Prospect
Ancestry.com US Federal Census
Annual Reviews
Antikvarium.hu
AquaDocs
Archeion
Archiv fuer Sozialgeschichte
Archive Ouverte en Sciences de l'Information et de la Communication (AOSIC)
Archives Canada
Ariana News
Art Institute of Chicago
Artefacts Canada
Artforum
Atlanta Journal-Constitution
Atypon Journals
AustLII and NZLII
Australian Dictionary of Biography
BAILII
BBC Genome
BBC
BIBSYS
BOCC
BOE
BOFiP-Impots
Baidu Scholar
Bangkok Post
Baruch Foundation
Beobachter
Bezneng Gajit
BibLaTeX
BibTeX
Biblio.com
Bibliontology RDF
Biblioteca Nacional de Maestros
Bibliotheque et Archives Nationale du Quebec (Pistard)
Bibliotheque et Archives Nationales du Quebec
Bibliotheque nationale de France
BioMed Central
BioOne
Bioconductor
Blaetter
Blogger
Bloomberg
Bloomsbury Food Library
BnF ISBN
Bookmarks
Bookshop.org
Boston Review
Bosworth Toller's Anglo-Saxon Dictionary Online
Bracero History Archive
Brill
Brukerhandboken
Bryn Mawr Classical Review
Bundesgesetzblatt
Business Standard
CABI - CAB Abstracts
CAOD
CBC
CCfr (BnF)
CERN Document Server
CEUR Workshop Proceedings
CFF References
CFF
CIA World Factbook
CLACSO
CLASE
CNKI
COBISS
COinS
CQ Press
CROSBI
CSIRO Publishing
CSL JSON
CSV
Cairn.info
CalMatters
Calisphere
Camara Brasileira do Livro ISBN
Cambridge Core
Cambridge Engage Preprints
CanLII
Canada.com
Canadian Letters and Images
Canadiana.ca
Cascadilla Proceedings Project
Cell Press
Central and Eastern European Online Library Journals
Champlain Society - Collection
Christian Science Monitor
Chronicling America
CiNii
Citavi 5 XML
CiteSeer
Citizen Lab
Civilization.ca
Climate Change and Human Health Literature Portal
Clinical Key
Code4Lib Journal
Colorado State Legislature
Columbia University Press
Common-Place
Computer History Museum Archive
Copernicus
Cornell LII
Cornell University Press
CourtListener
Crossref Unixref XML
Crossref-REST
Current Affairs
DABI
DAI-Zenon
DART-Europe
DBLP Computer Science Bibliography
DBpia
DEPATISnet
DOAJ
DOI Content Negotiation
DOI
DPLA
DSpace Intermediate Metadata
Dagens Nyheter
Dagstuhl Research Online Publication Server
Dar Almandumah
Data.gov
Databrary
Datacite JSON
Dataverse
Daum News
De Gruyter
Defense Technical Information Center
Delpher
Demographic Research
Denik CZ
Der Freitag
Der Spiegel
Desiring God
Deutsche Fotothek
Deutsche Nationalbibliothek
Dialnet
Die Zeit
DigiZeitschriften
Digital Humanities Quarterly
Digital Spy
Dimensions
Douban
Dreier Neuerscheinungsdienst
DrugBank.ca
Dryad Digital Repository
Duke University Press Books
E-periodica Switzerland
EBSCO Discovery Layer
EBSCOhost
EIDR
EPA National Library Catalog
ERIC
ESpacenet
EUR-Lex
Eastview
Edinburgh University Press Journals
Education Week
El Comercio (Peru)
El Pais
Electronic Colloquium on Computational Complexity
Elicit
Elsevier Health Journals
Elsevier Pure
Embedded Metadata
Emerald Insight
Encyclopedia of Chicago
Encyclopedia of Korean Culture
Endnote XML
Engineering Village
Epicurious
Erudit
Euclid
EurasiaNet
EurogamerUSgamer
Europe PMC
Evernote
F1000 Research
FAO Publications
FAZ.NET
Fachportal Padagogik
Factiva
Failed Architecture
Fairfax Australia
Fatcat
Figshare
Financial Times
Finna
Flickr
Foreign Affairs
Foreign Policy
FreeCite
FreePatentsOnline
Frieze
Frontiers
GMS German Medical Science
GPO Access e-CFR
Gale Databases
GaleGDC
Galegroup
Gallica
Game Studies
GameSpot
GameStar GamePro
Gasyrlar Awazy
Gemeinsamer Bibliotheksverbund ISBN
Gene Ontology
GitHub
Globes
Gmail
Goodreads
Google Books
Google Patents
Google Play
Google Presentation
Google Research
Google Scholar
Gulag Many Days, Many Lives
HAL Archives Ouvertes
HCSP
HLAS (historical)
HUDOC
Haaretz
Handelszeitung
Hanrei Watch
Harper's Magazine
Harvard Business Review
Harvard Caselaw Access Project
Harvard University Press Books
HathiTrust
HeinOnline
Heise
Herder
HighBeam
HighWire 2.0
HighWire
Hindawi Publishers
Hispanic-American Periodical Index
Homeland Security Digital Library
Huff Post
Human Rights Watch
IBISWorld
IDEA ALM
IEEE Computer Society
IEEE Xplore
IETF
IGN
IMDb
INSPIRE
IPCC
ISTC
Idref
In These Times
InfoTrac
Informationssystem Medienpaedagogik
IngentaConnect
Inside Higher Ed
Insignia OPAC
Institute of Contemporary Art
Institute of Physics
Integrum
Intellixir
Inter-Research Science Center
International Nuclear Information System
Internet Archive Scholar
Internet Archive Wayback Machine
Internet Archive
InvenioRDM
Isidore
J-Stage
JETS
JISC Historical Texts
JRC Publications Repository
JSTOR
Jahrbuch
Japan Times Online
Journal of Electronic Publishing
Journal of Extension
Journal of Machine Learning Research
Journal of Religion and Society
JurPC
Juricaf
Juris
K10plus ISBN
KStudy
Kanopy
Khaama Press
KitapYurdu.com
Kommersant
Korean National Library
L'Annee Philologique
LA Times
LIBRIS ISBN
LIVIVO
La Croix
La Nacion (Argentina)
La Presse
La Republica (Peru)
Lagen.nu
Landesbibliographie Baden-Wurttemberg
Lapham's Quarterly
Le Devoir
Le Figaro
Le Maitron
Le Monde
Le monde diplomatique
Legifrance
Legislative Insight
Lexis+
LexisNexis
Libraries Tasmania
Library Catalog (Aleph)
Library Catalog (Amicus)
Library Catalog (Aquabrowser)
Library Catalog (BiblioCommons)
Library Catalog (Blacklight)
Library Catalog (Capita Prism)
Library Catalog (DRA)
Library Catalog (Dynix)
Library Catalog (Encore)
Library Catalog (InnoPAC)
Library Catalog (Koha)
Library Catalog (Mango)
Library Catalog (OPALS)
Library Catalog (PICA)
Library Catalog (PICA2)
Library Catalog (Pika)
Library Catalog (Polaris)
Library Catalog (Quolto)
Library Catalog (RERO ILS)
Library Catalog (SIRSI eLibrary)
Library Catalog (SIRSI)
Library Catalog (SLIMS)
Library Catalog (TIND ILS)
Library Catalog (TLCYouSeeMore)
Library Catalog (VTLS)
Library Catalog (Visual Library 2021)
Library Catalog (Voyager 7)
Library Catalog (Voyager)
Library Genesis
Library Hub Discover
Library of Congress ISBN
LingBuzz
Lippincott Williams and Wilkins
Literary Hub
LiveJournal
London Review of Books
LookUs
Lulu
MAB2
MARC
MARCXML
MCV
MDPI Journals
MEDLINEnbib
METS
MIDAS Journals
MIT Press Books
MODS
MPG PuRe
Mailman
Mainichi Daily News
Mastodon
Matbugat.ru
Max Planck Institute for the History of Science Virtual Laboratory Library
Medium
MetaLib
Microbiology Society Journals
Microsoft Academic
Mikromarc
Milli Kutuphane
Musee du Louvre
NASA ADS
NASA NTRS
NCBI Nucleotide
NPR
NRC Research Press
NRC.nl
NTSB Accident Reports
NYPL Menus
NYPL Research Catalog
NYTimes.com
NZZ.ch
Nagoya University OPAC
National Academies Press
National Agriculture Library
National Archives of Australia
National Archives of South Africa
National Bureau of Economic Research
National Diet Library Catalogue
National Gallery of Art - USA
National Gallery of Australia
National Library of Australia (new catalog)
National Library of Belarus
National Library of Norway
National Library of Poland ISBN
National Post
National Technical Reports Library
National Transportation Library ROSA P
Nature Publishing Group
Neural Information Processing Systems
New Left Review
New Zealand Herald
Newlines Magazine
News Corp Australia
NewsBank
NewsnetTamedia
Noor Digital Library
Note HTML
Note Markdown
Notre Dame Philosophical Reviews
OAPEN
OCLC WorldCat FirstSearch
OECD
ORCID
OSF Preprints
OSTI Energy Citations
OVID Tagged
OZON.ru
OhioLINK
Old Bailey Online
Open Conf
Open Knowledge Repository
Open WorldCat
OpenEdition Books
OpenEdition Journals
Optical Society of America
Optimization Online
Ovid
Oxford Dictionaries Premium
Oxford English Dictionary
Oxford Music and Art Online
Oxford Reference
Oxford University Press
PC Gamer
PC Games
PEI Archival Information Network
PEP Web
PKP Catalog Systems
PLoS Journals
PRC History Review
Pajhwok Afghan News
Papers Past
Paris Review
Pastebin
Patents - USPTO
Peeters
Perceiving Systems
Perlego
PhilPapers
Philosopher's Imprint
Pleade
Polygon
Potsdamer Neueste Nachrichten
Preprints.org
Primo 2018
Primo Normalized XML
Primo
ProMED
ProQuest Ebook Central
ProQuest PolicyFile
ProQuest
Probing the Past
Project Gutenberg
Project MUSE
Protein Data Bank
PubFactory Journals
PubMed Central
PubMed XML
PubMed
PubPub
Publications Office of the European Union
Publications du Quebec
PyPI
Qatar Digital Library
R-Packages
RAND
RDF
REDALYC
RIS
RSC Publishing
Radio Free Europe Radio Liberty
RePEc - Econpapers
RePEc - IDEAS
Rechtspraak.nl
RefWorks Tagged
ReferBibIX
Regeringskansliet
Research Square
ResearchGate
Retsinformation
Reuters
Rock, Paper, Shotgun
Roll Call
Russian State Library
SAE Papers
SAGE Journals
SAGE Knowledge
SAILDART
SALT Research Archives
SFU IPinCH
SIPRI
SIRS Knowledge Source
SLUB Dresden
SORA
SSOAR
SSRN
SVT Nyheter
Sacramento Bee
Safari Books Online
Scholars Portal Journals
Scholia
Schweizer Radio und Fernsehen SRF
SciELO
ScienceDirect
Scopus
Semantic Scholar
Silverchair
Slate
SlideShare
Springer Link
Stack Exchange
Standard Ebooks
Stanford Encyclopedia of Philosophy
Stanford University Press
State Records Office of Western Australia
Stitcher
Store norske leksikon
Stuff.co.nz
Substack
Sud Ouest
Sueddeutsche.de
Summon 2
Superlib
Svenska Dagbladet
Sveriges radio
TEI
TV by the Numbers
TVNZ
Tagesspiegel
Talis Aspire
TalisPrism
Tatknigafund
Tatpressa.ru
Taylor & Francis eBooks
Taylor and Francis+NEJM
Tesis Doctorals en Xarxa
The Art Newspaper
The Atlantic
The Boston Globe
The Chronicle of Higher Education
The Daily Beast
The Economic Times - The Times of India
The Economist
The Free Dictionary
The Globe and Mail
The Guardian
The Hamilton Spectator
The Hindu
The Independent
The Intercept
The Met
The Microfinance Gateway
The Nation
The National Archives (UK)
The New Republic
The New York Review of Books
The New Yorker
The Open Library
The Straits Times
The Telegraph
The Times and Sunday Times
TheMarker
Theory of Computing
ThesesFR
Thieme
Time.com
TimesMachine
Tony Blair Institute for Global Change
Toronto Star
Transportation Research Board
Treesearch
Trove
Tumblr
Twitter
UChicago VuFind
UNZ Print Archive
UPCommons
US National Archives Research Catalog
Ubiquity Journals
University Press Scholarship
University of California Press Books
University of Chicago Press Books
University of Wisconsin-Madison Libraries Catalog
Unqualified Dublin Core RDF
UpToDate References
Vanity Fair
Verniana-Jules Verne Studies
Verso Books
Vice
Victoria & Albert Museum
Vimeo
VoxEU
WHO
WIPO
Wall Street Journal
Wanfang Data
Washington Monthly
Washington Post
Web of Science Nextgen
Web of Science Tagged
Web of Science
Welt Online
WestLaw UK
WikiLeaks PlusD
Wikidata QuickStatements
Wikidata
Wikimedia Commons
Wikipedia Citation Templates
Wikipedia
Wikisource
Wikiwand
Wiktionary
Wildlife Biology in Practice
Wiley Online Library
Wilson Center Digital Archive
Winnipeg Free Press
Wired
Womennews
World Digital Library
World History Connected
World Shakespeare Bibliography Online
WorldCat Discovery Service
XML ContextObject
YPSF
Ynet
YouTube
ZIPonline
ZOBODAT
Zotero RDF
ZoteroBib
arXiv Vanity
arXiv.org
artnet
beck-online
clinicaltrials.gov
dLibra
dejure.org
deleted.txt
dhistory
digibib.net
eLibrary.ru
eLife
eMJA
eMedicine
ePrint IACR
ebrary
etatar.ru
feb-web.ru
fishpond.co.nz
fr-online.de
govinfo
index.d.ts
informIT database
io-port
jsconfigon
jurion
mEDRA
magazines.russ.ru
medes
newshub.co.nz
newspapers.com
openJur
package-lockon
packageon
reddit
sbn.it
scinapse
semantics Visual Library
taz.de
unAPI
wiso
zbMATH
zotero.org