Conference Program

Venue: Transilvania University of Brașov with a permanently open online session via WEBEX meeting: https://uaic.webex.com/meet/daniela.gifu All hours are EET = CET+1

11 December – Consortium Meeting Day 1
9:15 – 11:00 Connections

11:00 – 12:25 ConsILR Meeting Session 1

5 to 20 minutes speeches:

– Anca Vasilescu – Department of Mathematics-Informatics, Transilvania University of Brașov

– Nicolae-Victor Zamfir – Vice-president of the Romanian Academy

– Svetlana Cojocaru – Vice-president of the Sciences Academy of the Republic of Moldova

– Dan Tufiș – president of the Informatisation Commission for the Romanian Language

– Ruxandra Cosma – Department of German Language and Literature, University of Bucharest

– Dan Cristea – Institute of Computer Science, Romanian Academy, Iași

– Corneliu Burileanu – Faculty of Electronics, Telecommunications and Information Technology, National University Politehnica of Bucharest

– Mircea Giurgiu – Faculty of Electronics, Telecommunications and Information Technology, Technical University of Cluj Napoca

– Marius Popescu & Radu Ionescu – Faculty of Mathematics and Computer Science, University of Bucharest

– Elena Isabelle Tamba – Institute of Romanian Philology “Alexandru Philippide”, Iași

12:25 – 13: 15 Invited talk: Gabriela Haja – Director of the Institute of Romanian Philology “Alexandru Philippide”, Iași

13:15 – 14:30 – Lunch break

14:30 – 16:40 ConsILR Meeting Session 2

5 to 15 minutes speeches:

– Ștefan Trășan-Matu – Faculty of Computer Engineering, National University Politehnica of Bucharest

– Vasile Păiș – Research Institute for Artificial Intelligence, Romanian Academy, Bucharest

– Victoria Bobicev – Technical University of Moldova

– Alice Toma – Institute of Romanian Language, Université libre de Bruxelles

– Roxana Rogobete, Ana-Maria Bucur, Mădălina Chitez – West University of Timișoara

– Mincu Eugenia – Institute of Romanian Philology “B. P. Hașdeu”, Chișinău

– Mihaela Mocanu – Department of Social Sciences and Humanities, Institute of Interdisciplinary Research, “Alexandru Ioan Cuza” University of Iași

– Daniela Gîfu, Diana Trandabăț – Institute of Computer Science and Faculty of Computer Science of the A.I.Cuza University, Iași

– Oana Niculescu – “Iorgu Iordan – Alexandru Rosetti” Linguistics Institute of the Romanian Academy, Bucharest

– Horia Cucu – Faculty of Electronics, Telecommunications and Information Technology, National University Politehnica of Bucharest

– Galaction Verebceanu – Institute of Romanian Philology “B. P. Hașdeu”, Chișinău

– Silviu-Ioan Bejenariu – Institute of Computer Science of the Romanian Academy, Iași

– Anca-Diana Bibiri – Department of Social Sciences and Humanities, Institute of Interdisciplinary Research, “Alexandru Ioan Cuza” University of Iași

– Marc Frâncu – West University of Timișoara

– Alexandr Parahonco – Vladimir Andrunachievici Institute of Mathematics and Computer Science, “Alecu Russo” Bălți State University, Republic of Moldova

– Rareș Arvinte – Faculty of Computer Science, “Alexandru Ioan Cuza” University of Iași

– Gheorghe Tecuci – George Mason University and Romanian Academy

_________________________________________________________________________ 

12 December – ConsILR Conference Day 1

8:30 – 9:00 Registration and connections
9:00 – 9:10 Welcome message

9:10 – 10:00 Keynote Session 1 – chaired by Ștefan Trăușan-Matu
         Marius Popescu: AD-NLP: A Benchmark for Anomaly Detection in Natural Language Processing

Abstract:

Deep learning models have reignited the interest in Anomaly Detection research in recent years. Methods for Anomaly Detection in text have shown strong empirical results on ad-hoc anomaly setups that are usually made by downsampling some classes of a labeled dataset. This can lead to reproducibility issues and models that are biased toward detecting particular anomalies while failing to recognize them in more sophisticated scenarios. In the present work, we provide a unified benchmark for detecting various types of anomalies, focusing on problems that can be naturally formulated as Anomaly Detection in text, ranging from syntax to stylistics. In this way, we are hoping to facilitate research in Text Anomaly Detection. We also evaluate and analyse two strong shallow baselines, as well as two of the current state-of-the-art neural approaches, providing insights into the knowledge the neural models are learning when performing the anomaly detection task.

Presentations Session 1: eLearning Systems – chaired by Daniela Gîfu
10:00 – 10:20 Alexandr Parahonco and Mircea Petic: Educational text content generation for eLearning systems - DOI:10.47743ConsILR2023.04
10:20 – 10:40 Olesea Caftanatov, Ludmila Malahov, Inga Țitchiev, Dan Tălămbuță, Daniela Caganovschi and Eugenia Mincu: Metaphor Mastery: Augmented Reality Dictionary Flashcards for E-Learning - DOI:10.47743ConsILR2023.05

10:40-11:00 Olesea Caftanatov, Ludmila Malahov, Inga Țitchiev, Dan Tălămbuță, Tudor Bumbu, Ana Nastasiu and Victoria Bobicev: Enhancing Language Learning through Augmented Reality: a Focus on Multiword Expressions - DOI:10.47743ConsILR2023.06

11:00 – 11:30 Coffee break

11:30 – 12:20 Keynote Session 2 – chaired by Dan Tufiș
          Radu Ion: The Romanian language in the age of Standard and Large Language Models

 Abstract:

We have entered a decade in which large neural networks, with billions of parameters, can be trained in a matter of weeks, and that can deal with multi-modal information in a way that is unprecedented. These networks perform many tasks related to text comprehension, can produce naturally sounding speech or can extract a lot of information from images or videos. Furthermore, the generative neural networks of different types can produce text, images or speech that can fool almost anyone as to who is the source of these productions: human or machine. In this context, we take a bird’s eye view on how the Romanian language is covered by the text encoding and generation large neural networks (called Standard Language Models, if they have fewer than a billion parameters and Large Language Models, if they have more than a billion parameters) and we open the discussion on whether and how Romanian can have a LLM of its own, trained only on Romanian texts.

Presentations Session 2: GDPR and Summarisation – chaired by Diana Trandabăț
12:20 – 12:40 Vasile Păiș, Verginica Barbu Mititelu, Elena Irimia, Radu Ion, Valentin Badea, Dan Tufis and Cosmin Sterea-Grossu: Towards anonymizing the Romanian jurisprudence - DOI:10.47743ConsILR2023.07
12:40 – 13:00 Victoria Alexei and Victoria Bobicev: Comparison of Manual and Automatic Summarization - DOI:10.47743ConsILR2023.08

13:00 – 15:00 Lunch break

17:00 – City tour

19:00 Welcome party
_________________________________________________________________________ 

13 December – ConsILR Conference Day II
8:30 – 9:00 Connections

Presentations Session 3: Language Models – chaired by Vasile Păiș
9:00 – 9:20 Bogdan Florea and Adrian Iftene: The News in Brief – Leveraging Machine Learning and Artificial Intelligence in News Clustering, Summarization and Evaluation - DOI:10.47743ConsILR2023.09

9:20 – 9:40 Bhone Myint Swe, Min Thaw Phyo, Chan L. Min, Aung Sann Thit, Aung Hlaing Moe, Kaung Khant Min, Phyo Thu Htet and Aye Su Phyo: Myanmar Sign Language Recognition using Deep Learning - DOI:10.47743ConsILR2023.10
9:40 – 10:00 Radu Ion: A Romanian Bert Model for Linguistic Analysis - DOI:10.47743ConsILR2023.11

10:00 – 10:30 Coffee break

10:30 – 11:20 Keynote Session 3 – chaired by Dan Cristea
          Liviu Dinu: Computational Approaches in Historical Linguistics

Abstract:

Natural languages are living eco-systems, they are constantly in contact and, by consequence, they change continuously. Traditionally, the main Historical Linguistics problems (How are languages related? How do languages change across space and time?) have been investigated with comparative linguistics instruments. The main idea of the comparative method is to perform a property-based comparison of multiple sister languages to infer properties of their common ancestor. It is a time-consuming manual process that required a large amount of intensive work.

We propose here computer-assisted methods for identifying cognates, for discriminating between related words, and for protoword reconstruction. The identification of cognates is a fundamental process in historical linguistics, on which any further research is based. Even though there are several cognate databases for Romance languages, they are rather scattered, incomplete, noisy, contain unreliable information, or have uncertain availability. We introduced a comprehensive database of Romance cognates and borrowings based on the etymological information provided by the dictionaries (the largest known database of this kind, in our best knowledge). We extracted pairs of cognates between any two Romance languages by parsing electronic dictionaries of Romanian, Italian, Spanish, Portuguese and French. Based on this resource, we proposed a strong benchmark for the automatic detection of cognates, by applying machine learning and deep learning-based methods on any two pairs of Romance languages. Beside the largest database of this kind, we find also that automatic identification of cognates is possible with accuracy averaging around 94% for the more difficult task formulations. We also proposed methods for automatic discrimination between related words (cognates vs loanwords, inherited vs loanwords, derivative vs cognates). Further we developed a methodology for protoword reconstruction and missing Romanian cognates reconstruction. Given words in Romance modern languages, the task is to automatically reconstruct the Latin proto-words from which the modern words evolved. We applied the method for producing related words based on sequence labeling, aiming to fill in the gaps in incomplete cognate sets in Romance languages with Latin etymology (producing Romanian cognates that are missing).

Presentations Session 4: Discourse and Humor – chaired by Anca Vasilescu
11:20 – 11:40    Tudor Voicu and Verginica Mititelu: Annotating Discourse Relations in the Romanian Reference Treebank - DOI:10.47743ConsILR2023.12

11:40 – 12:00 Eugenia Mincu: Specialized Language: Teaching and Computerization
12:00 – 12:20 Daniela Gîfu and Dan Stoica: This Makes Me Laugh. Does It Make You Laugh, too? - DOI:10.47743ConsILR2023.13

12:20 – 14:00 Lunch break

14:00 – 14:50 Keynote Session 4 – chaired by Verginica Barbu Mititelu

Radu Tudor Ionescu: Recent Text and Audio Resources for the Romanian Language

 Abstract:

In this talk, I will present two recent resources for the Romanian language. First, we will introduce a novel Romanian Clickbait Corpus (RoCliCo) comprising 8,313 news samples which are manually annotated with clickbait and non-clickbait labels. We present preliminary results with various machine learning methods. Among the considered models, we will present a novel BERT-based contrastive learning model that learns to encode news titles and contents into a deep metric space. Next, we will present RoDia, the first dataset for Romanian dialect identification from speech. RoDia includes a varied compilation of speech samples from five distinct regions of Romania, covering both urban and rural environments, totalling 2 hours of manually annotated speech data. Along with RoDia, we introduce a set of competitive models to be used as baselines for future research. Moreover, we present empirical evidence showing that Automatic Speech Recognition on dialectal speech is more challenging.

 14:50 – 15:20 Coffee break

Presentations Session 5: Romanian Speech – chaired by Corneliu Burileanu
15:20 – 15:40 Florin-Teodor Olariu, Alexandru Laurențiu Cohal, Luminița Botoșineanu, Veronica Olariu, Ramona Luca and Silviu Ioan Bejinariu: Making linguistic resources accessible. The audio archive of the New Romanian Linguistic Atlas by Regions. Moldova and Bukovina - DOI:10.47743ConsILR2023.15
15:40 – 16:00 Irina Vasilița and Adrian Iftene: Towards Speech-Based Web App Development - DOI:10.47743ConsILR2023.16

16:00 – 16:20 Carol-Luca Gasan and Vasile Păiș: Investigation of Romanian Speech. Recognition Improvement by Incorporating Italian Speech Data - DOI:10.47743ConsILR2023.17

16:20 ConsILR Conference Closing

Evening: Cultural event (to be announced)

_________________________________________________________________________ 

14 December – Consortium Meeting Day 2

9:00 – 10:30 Meeting of the Romanian Academy Commission for the Informatisation of the Romanian Language

10:30 – 11:00 Coffee break

11:00 – 12:00 Meeting of the Romanian Association for Computational Linguistics

12:00 – Conclusions