Now showing 1 - 10 of 33
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    Capturing the Varieties of Natural Language Inference: A Systematic Survey of Existing Datasets and Two Novel Benchmarks
    Transformer-based Pre-Trained Language Models currently dominate the field of Natural Language Inference (NLI). We first survey existing NLI datasets, and we systematize them according to the different kinds of logical inferences that are being distinguished. This shows two gaps in the current dataset landscape, which we propose to address with one dataset that has been developed in argumentative writing research as well as a new one building on syllogistic logic. Throughout, we also explore the promises of ChatGPT. Our results show that our new datasets do pose a challenge to existing methods and models, including ChatGPT, and that tackling this challenge via fine-tuning yields only partly satisfactory results.
    Type:
    Journal:
    Scopus© Citations 20
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    A Canonical Context-Preserving Representation for Open IE: Extracting Semantically Typed Relational Tuples from Complex Sentences
    Modern systems that deal with inference in texts need automatized methods to extract meaning representations (MRs) from texts at scale. Open Information Extraction (IE) is a prominent way of extracting all potential relations from a given text in a comprehensive manner. Previous work in this area has mainly focused on the extraction of isolated relational tuples. Ignoring the cohesive nature of texts where important contextual information is spread across clauses or sentences, state-of-the- art Open IE approaches are thus prone to generating a loose arrangement of tuples that lack the expressiveness needed to infer the true meaning of complex assertions. To overcome this limitation, we present a method that allows existing Open IE systems to enrich their output with additional meta information. By leveraging the semantic hierarchy of minimal propositions generated by the discourse-aware Text Simplification (TS) approach presented in Niklaus et al. (2019), we propose a mechanism to extract semantically typed relational tuples from complex source sentences. Based on this novel type of output, we introduce a lightweight semantic representation for Open IE in the form of normalized and context-preserving relational tuples. It extends the shallow semantic representation of state-of-the-art approaches in the form of predicate-argument structures by capturing intra-sentential rhetorical structures and hierarchical relationships between the relational tuples. In that way, the semantic context of the extracted tuples is preserved, resulting in more informative and coherent predicate-argument structures which are easier to interpret. In addition, in a comparative analysis, we show that the semantic hierarchy of minimal propositions benefits Open IE approaches in a second dimension: the canonical structure of the simplified sentences is easier to process and analyze, and thus facilitates the extraction of relational tuples, resulting in an improved precision (up to 32%) and recall (up to 30%) of the extracted relations on a large benchmark corpus.
    Type:
    Journal:
    Issue:
    Scopus© Citations 4
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    Supporting Cognitive and Emotional Empathic Writing of Students
    We present an annotation approach to capturing emotional and cognitive empathy in student-written peer reviews on business models in German. We propose an annotation scheme that allows us to model emotional and cognitive empathy scores based on three types of review components. Also, we conducted an annotation study with three annotators based on 92 student essays to evaluate our annotation scheme. The obtained inter-rater agreement of α = 0.79 for the components and the π = 0.41 for the empathy scores indicate that the proposed annotation scheme successfully guides annotators to a substantial to moderate agreement. Moreover, we trained predictive models to detect the annotated empathy structures and embedded them in an adaptive writing support system for students to receive individual empathy feedback independent of an instructor, time, and location. We evaluated our tool in a peer learning exercise with 58 students and found promising results for perceived empathy skill learning, perceived feedback accuracy, and intention to use. Finally, we present our freely available corpus of 500 empathy-annotated, student-written peer reviews on business models and our annotation guidelines to encourage future research on the design and development of empathy support systems.
    Type:
    Journal:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    BioMistral-Clinical: A Scalable Approach to Clinical LLMs via Incremental Learning and RAG
    (The Asian Federation of Natural Language Processing and The Association for Computational Linguistics, 2025-12)
    Chen, Ziwei
    ;
    ;
    The integration of large language models (LLMs) into clinical medicine represents a major advancement in natural language processing (NLP). We introduce BioMistral-Clinical 7B, a clinical LLM built on BioMistral-7B (Labrak et al., 2024), designed to support continual learning from unstructured clinical notes for real-world tasks such as clinical decision support. Using the augmented-clinical notes dataset provided by Hugging Face (2024), we apply prompt engineering to transform unstructured text into structured JSON capturing key clinical information (symptoms, diagnoses, treatments, outcomes). We employ selfsupervised continual learning (SPeCiaL) (Caccia and Pineau, 2021) to achieve efficient incremental training. Evaluation on MedQA (Jin et al., 2021) and MedMCQA (Pal et al., 2022) shows that BioMistral-Clinical 7B improves accuracy on MedMCQA by nearly 10 points (37.4% vs. 28.0%) over the base model, while maintaining comparable performance on MedQA (34.8% vs. 36.5%). Building on this, we propose the BioMistral-Clinical System, which integrates Retrieval-Augmented Generation (RAG) (Lewis et al., 2020) to enrich responses with relevant clinical cases retrieved from a structured vector database. The full system enhances clinical reasoning by combining domain-specific adaptation with contextual retrieval.
    Type:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    ARTIST: A Learning Support System for Fostering Students’ Argumentative Writing Skills
    We present ARTIST, a learning support system that can help students assess their argumentative writing and provide automated, individual feedback, thus improving their writing performance. It analyzes student-written argumentative texts by identifying argument components and their relationships. The resulting argumentative discourse structure is displayed in an interactive interface. In that way, the ARTIST tool provides immediate and personalized visual feedback on the quality of students’ texts, supporting self-monitoring and reflection on how to improve their texts.
    Type:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    CLEAR: A Comprehensive Linguistic Evaluation of Argument Rewriting by Large Language Models
    (Association for Computational Linguistics, 2025) ;
    While LLMs have been extensively studied on general text generation tasks, there is less research on text rewriting, a task related to general text generation, and particularly on the behavior of models on this task. In this paper we analyze what changes LLMs make in a text rewriting setting. We focus specifically on argumentative texts and their improvement, a task named Argument Improvement (ArgImp). We present CLEAR: an evaluation pipeline consisting of 57 metrics mapped to four linguistic levels: lexical, syntactic, semantic and pragmatic. This pipeline is used to examine the qualities of LLM-rewritten arguments on a broad set of argumentation corpora and compare the behavior of different LLMs on this task and analyze the behavior of different LLMs on this task in terms of linguistic levels. By taking all four linguistic levels into consideration, we find that the models perform ArgImp by shortening the texts while simultaneously increasing average word length and merging sentences. Overall we note an increase in the persuasion and coherence dimensions.
    Type:
    Journal:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    LLMs meet Bloom’s Taxonomy: A Cognitive View on Large Language Model Evaluations
    (Association for Computational Linguistics, 2025-01) ;
    Current evaluation approaches for Large Language Models (LLMs) lack a structured approach that reflects the underlying cognitive abilities required for solving the tasks. This hinders a thorough understanding of the current level of LLM capabilities. For instance, it is widely accepted that LLMs perform well in terms of grammar, but it is unclear in what specific cognitive areas they excel or struggle in. This paper introduces a novel perspective on the evaluation of LLMs that leverages a hierarchical classification of tasks. Specifically, we explore the most widely used benchmarks for LLMs to systematically identify how well these existing evaluation methods cover the levels of Bloom’s Taxonomy, a hierarchical framework for categorizing cognitive skills. This comprehensive analysis allows us to identify strengths and weaknesses in current LLM assessment strategies in terms of cognitive abilities and suggest directions for both future benchmark development as well as highlight potential avenues for LLM research. Our findings reveal that LLMs generally perform better on the lower end of Bloom’s Taxonomy. Additionally, we find that there are significant gaps in the coverage of cognitive skills in the most commonly used benchmarks.
    Type:
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    Exploring the Usefulness of Open and Proprietary LLMs in Argumentative Writing Support
    In this article, we present the results of an exploratory study conducted with our self-developed tool Artist. The goal of the tool is to give formative feedback to develop students' argumentation skills. We compare the feedback that two different LLMs, an open-sourced one by META and one of OpenAI's fully proprietary ones, give to students' argumentative writing. We find that, overall, students find the feedback provided by both LLMs helpful (7.51 vs. 7.65 on a scale from 1 to 10), and they rate the quality of the feedback as good to very good. We take this as a very encouraging provisional result that invites larger and more extensive studies on the topic.
    Type:
    Scopus© Citations 8