• Skip to primary navigation
  • Skip to main content
  • Skip to primary sidebar
  • Skip to footer

Center for Artificial Intelligence and Cybersecurity – AIRI

  • Home
  • About Us
    • Vision, Mission and Goals
    • Center Activities
    • Center Faculty
    • Steering Committee
    • Press
  • Research
    • Scientific Projects
    • Research Papers
  • Laboratories
    • Machine Learning
    • Natural Speech & Language Processing
    • Blockchain Technology
    • Information Processing & Pattern Recognition
    • AI in Medicine
    • Data Mining
    • Computer Vision
    • Complex Networks
    • Human-Computer Interaction
    • Maritime Cybersecurity
    • Autonomous Navigation
    • AI in Mechatronics
    • AI in Education
    • Hybrid Computational Methods
    • Drug Design
    • Legal Aspects of AI
    • Ethically Aligned AI
    • Cultural Complexity
  • Collaboration
    • Industry Collaboration
    • Industry Projects
    • International Collaboration
  • News
  • Contact

Evaluation of Language Models over Croatian Newspaper Texts

28.11.2017

Statistical language modeling involves techniques and procedures that assign probabilities to word sequences or, said in other words, estimate the regularity of the language. This paper presents basic characteristics of statistical language models, reviews their use in the large set of speech and language applications, explains their formal definition and shows different types of language models. Detailed overview of n-gram and class-based models (as well as their combinations) is given chronologically, by type and complexity of models, and in aspect of their use in different NLP applications for different natural languages. The proposed experimental procedure compares three different types of statistical language models: n-gram models based on words, categorical models based on automatically determined categories and categorical models based on POS tags. In the paper, we propose a language model for contemporary Croatian texts, a procedure how to determine the best n-gram and the optimal number of categories, which leads to significant decrease of language model perplexity, estimated from the Croatian News Agency articles (HINA) corpus. Using different language models estimated from the HINA corpus, we show experimentally that models based on categories contribute to a better description of the natural language than those based on words. These findings of the proposed experiment are applicable, except for Croatian, for similar highly inflectional languages with rich morphology and non-mandatory sentence word order.

Authors:
Slobodan Beliga, Ivo Ipšić, Sanda Martinčić-Ipšić.
Journal:
Information Technology and Control
Publishing date:
15.11.2017
View original article

Primary Sidebar

Latest Projects

ABsistemDCiCloud

Machine Learning for Knowledge Transfer in Medical Radiology

Estimating River Discharges in Highly Stratified Estuaries

Multilayer Framework for the Information Spreading Characterization in Social Media during the COVID-19 Crisis (InfoCoV)

European Network for assuring food integrity using non-destructive spectral sensors

Latest Research Papers

Neural Natural Language Generation: A Survey on Multilinguality, Multimodality, Controllability and Learning

Entropy-Based Concentration and Instantaneous Frequency of TFDs from Cohen’s, Affine, and Reassigned Classes

Coupled encoding methods for antimicrobial peptide prediction: How sensitive is a highly accurate model?

The Choice of Time–Frequency Representations of Non-Stationary Signals Affects Machine Learning Model Accuracy: A Case Study on Earthquake Detection from LEN-DB Data

Improved Parametrized Multiple Window Spectrogram with Application in Ship Navigation Systems

Latest News

Assoc. prof. Jonatan Lerga received the Croatian Academy of Sciences and Arts award

Dr. Sc. Nikola Lopac successfully defended his doctoral dissertation

Presentation at the conference “Digital Innovation and Technology for People”

Assoc. prof. dr. sc. Jonatan Lerga presented AIRI Center at the IEEE Rijeka : Computer Society Congress 2021

Prof. dr. sc. Ana Mestrovic participated at the Panel on perspectives and real-life applications of AI organized by IEEE Technology and Engineering Management Society

We provide the expertise for solving real world problems using AI

If your company wants to implement artificial intelligence in your products or services, or increase your level of cybersecurity, our multidisciplinary team of scientists is your ideal partner.

Contact us

Footer

Center for Artificial Intelligence and Cybersecurity
  • jlerga@airi.uniri.hr
  • +385 51 406 500

University of Rijeka

University of Rijeka

About the Center

  • About Us
  • News
  • Privacy Policy
  • Contact

Center Activities

  • Laboratories
  • Scientific Projects
  • Industry Projects
  • Research Papers
  • Industry Collaboration
  • International Collaboration

Footer bottom left

© 2020 Center for Artificial Intelligence and Cybersecurity, all rights reserved.

Designed & developed by Nela Dunato Art & Design