scieee AI-readable full text Open interactive document viewer

RoBERTaSense-FACIL: A Technical Report and Model Selection Study for Meaning Preservation in Easy-to-Read Spanish Texts

Diab, Isam; Suárez-Figueroa, Mari Carmen

Abstract

RoBERTaSense-FACIL is a Spanish Transformer-based model fine-tuned to evaluate meaning preservation in Easy-to-Read (E2R) text adaptations. The model builds upon RoBERTa-base-bne and incorporates a balanced dataset of expert-validated E2R adaptations together with automatically generated hard negatives designed to introduce structural, semantic, and cross-textual distortions. This technical report describes the full methodology used to construct the dataset, the hard negative generation framework, the fine-tuning process, and a comparative evaluation of three models: MeaningBERT, RoBERTa-base-bne, and a BERTScore-based regression variant. Results show that the fine-tuned RoBERTa-base-bne, referred to as RoBERTaSense-FACIL, achieves the most robust and reliable performance for binary meaning-preservation classification in Spanish E2R texts. Data availability:The datasets and intermediate scripts used in this work cannot be made publicly available due to privacy and copyright restrictions. However, access may be granted upon reasonable request for academic research purposes. Model availability:The RoBERTaSense-FACIL model is publicly available on Hugging Face:https://huggingface.co/oeg/RoBERTaSense-FACIL

Full text

RoBERTaSense-FACIL: A Technical Report and Model Selection Study for Meaning Preservation in Easy-to-Read Spanish Texts Isam Diab-Lozano 1and Mari Carmen Su´arez-Figueroa 1 1Ontology Engineering Group (OEG), Universidad Polit´ecnica de Madrid, Spain Abstract This technical report presents RoBERTaSense-FACIL, a Spanish model based on RoBERTa designed to evaluate meaning preservation in Easy-toRead (E2R) text adaptations. The report includes a comparative study of three approaches to determine the most reliable architecture for the task. Based on the results, RoBERTa-base-bne fine-tuned on a balanced dataset of positives and hard negatives achieves the best performance and is adopted as the final model, hereafter referred to as RoBERTaSenseFACIL. The report documents the dataset construction, negative generation strategies, fine-tuning pipeline, evaluation metrics, and error analysis, providing a complete description of the model and its training process. 1 Introduction Ensuring accessible information for people with cognitive disabilities is a crucial component of inclusive communication. Equal opportunities and universal access to information are recognised as fundamental rights1. However, certain groups, particularly those with cognitive or intellectual disabilities, experience significant difficulties in reading comprehension. Enhancing cognitive accessibility is therefore essential to promote active participation in domains such as politics, education, employment, and culture. To support this goal, the Easy-to-Read (E2R) methodology was developed and formalised in standards such as the Spanish UNE 153101:2018 [1] and other European guidelines [2, 3]. E2R provides linguistic and design recommendations to improve comprehension and goes beyond simply simplifying vocabulary or 1Convention on the Rights of Persons with Disabilities (United Nations, 2006). Available at: https://www.ohchr.org/en/instruments-mechanisms/instruments/ convention-rights-persons-disabilities 1 summarising content. It allows structural and lexical transformations and may introduce supporting elements that are not present in the original text [4]. A key challenge in E2R is ensuring that adaptations preserve the intended meaning of the source text. Meaning preservation [5, 6, 7] refers to the extent to which an adapted version conveys the same overall message and communicative intent as the original. Although structural and lexical changes are allowed, the adaptation must still retain the core ideas. Reliable tools for evaluating meaning preservation are limited, particularly for Spanish and for accessibility-oriented text adaptation. To address this gap, this report presents a comparative study of three modelbased approaches applied to Spanish. We fine-tuned and evaluated three different architectures on a dataset of original and E2R-adapted text pairs, with the goal of determining which model best captures semantic equivalence in the context of cognitive accessibility. The models evaluated are: •MeaningBERT [8]: a model trained to assess semantic similarity and meaning preservation, originally designed for English. •RoBERTa-base-bne [9]: a monolingual Spanish RoBERTa model finetuned using a methodology inspired by MeaningBERT. •RoBERTa-base-bne with BERTScore fine-tuning: a Spanish adaptation of BERTScore [7] for evaluating text similarity. Based on the comparative evaluation presented in this report, the bestperforming model —RoBERTa-base-bne fine-tuned on a balanced dataset of positive and hard-negative pairs— is selected as the final system for assessing meaning preservation2. Throughout this report, we refer to this final model as RoBERTaSense-FACIL. Before fine-tuning, we use the name RoBERTa-basebne; after fine-tuning, RoBERTaSense-FACIL denotes the resulting model ready for practical use. 2 State of the Art In the context of Automatic Text Simplification3(ATS), both human and automatic evaluation methods have been explored, including metrics commonly used in machine translation and summarization tasks, such as BLEU [11], ROUGE [12], SARI [13] and METEOR [14]. Nevertheless, these metrics often fail to capture the semantic richness and subtle changes in meaning introduced during the adaptation process. 2The datasets and scripts used in this work cannot be made publicly available due to privacy restrictions. Access may be granted upon request. 3Text adaptation always aims to transform texts to meet the needs of a specific audience, while text simplification tends to reduce textual complexity and does not always consider the target reader’s profile [10]. 2 Recent surveys on ATS [15, 16] have shown that BLEU and SARI are the most widely used automatic metrics. However, BLEU has been found to correlate negatively with textual simplicity, making it unsuitable for ATS, while SARI focuses mostly on lexical simplifications and minor reordering, failing to capture deeper structural changes. ROUGE, despite being useful in summarisation, is rarely used in ATS. Readability formulas such as Flesch Reading Ease [17] or Flesch-Kincaid Grade Level [18] also present significant limitations: they are language-dependent, disregard layout and user-related variables, and are not designed for cognitive accessibility. Furthermore, regarding automatic text adaptation into E2R, there is currently no standardised and universally accepted evaluation system. This gap makes it difficult to compare adaptation methods or to reliably assess which version of a text best meets accessibility standards. As a result, researchers often combine multiple metrics. Manual evaluations, typically based on Likert scales assessing grammar, simplicity, and meaning preservation, remain a common approach, but they are time-consuming, subjective, and often conducted by experts rather than target users. Moreover, in the E2R context, it is crucial to involve validators, that is, individuals with reading comprehension difficulties, in order to ensure that adapted texts effectively serve their intended audience [4]. In terms of semantic evaluation, BERTScore [7] has gained traction for its use of contextual embeddings to estimate similarity between original and simplified texts. It is, however, unsupervised and trained for English. Meanwhile, supervised models like MeaningBERT [8], fine-tuned on human similarity ratings, show promising results in English but lack Spanish counterparts. This points to a research gap in the development of language-specific models for supervised evaluation of meaning preservation, particularly in accessibility-focused adaptations. 3 Model Selection The initial choice to fine-tune MeaningBERT was based on its specialisation in meaning preservation, which aligns closely with the objective of this study: evaluating whether automatically adapted Easy-to-Read texts in Spanish maintain the core meaning of their original versions. MeaningBERT was originally developed as a fine-tuned BERT model specifically designed to measure semantic similarity with a focus on meaning retention. However, its training was conducted exclusively on English datasets, which significantly impacted its performance when applied to Spanish text pairs. A critical limitation emerged from the use of the bert-base-uncased tokenizer, which is tightly coupled to English lexical and morphological patterns. During preliminary evaluations, we observed that this tokenizer failed to handle Spanish inputs adequately: many common Spanish words were fragmented into multiple subwords or misrepresented altogether. As a result, the model generated weak semantic representations and produced low similarity scores, even 3 when text pairs were near-identical in meaning. This finding highlights the importance of ensuring language alignment not only at the model level but also at the tokenization level when applying pretrained architectures cross-lingually. To address this issue, a second experiment was conducted using RoBERTabase-bne, a Spanish language model pretrained on large-scale Spanish corpora. This model was fine-tuned using a binary classification setup similar to MeaningBERT, but with a tokenizer specifically optimized for the Spanish language. The shift to a native Spanish model substantially improved the tokenization quality, which in turn enhanced the model’s ability to detect fine-grained semantic equivalence. The improved scores obtained in this setting validated the hypothesis that language-specific pretraining and tokenization are essential for tasks involving subtle meaning comparison in non-English texts. Finally, to explore alternative evaluation strategies beyond binary classification, a third experiment was designed following the principles of the BERTScore metric. Traditionally, BERTScore is an unsupervised evaluation method that computes cosine similarity between contextual token embeddings, typically relying on English models like roberta-base. Instead, our approach implemented a supervised regression framework built upon RoBERTa-base-bne, in which the model was trained to predict continuous content preservation scores. The training objective used a Mean Squared Error (MSE) loss. 4 Dataset Construction The expert-annotated dataset used for fine-tuning consisted exclusively of positive pairs, each containing an original Spanish text and its corresponding E2R adaptation. These adaptations were validated by experts to ensure maximum meaning preservation (label = 1). Although this data set provides high-quality positive examples, it is insufficient for supervised training of binary or regression models, as no negative instances are available from expert annotation alone. To enable robust supervised learning, it was therefore necessary to automatically generate hard negatives. These are artificial pairs that maintain surface similarity to legitimate E2R adaptations, but introduce structural or semantic distortions that alter the meaning. Hard negatives force the model to distinguish between subtle cases of meaning preservation and meaning alteration. Our design draws on prior work in contrastive learning and data augmentation [19, 20, 21, 22] to ensure diversity and controlled difficulty. 5 Hard Negative Generation and Final Dataset To enable supervised learning, the expert-annotated dataset, containing only positive E2R adaptations, was extended with automatically generated hard negatives. These negatives resemble valid adaptations at the surface level while introducing structural or semantic distortions. The goal is to force models to distinguish subtle meaning changes rather than relying on trivial lexical cues. 4 5.1 Types of Hard Negatives We define five categories of hard negatives, each representing a distinct form of structural or semantic distortion. These categories follow established perturbation families in contrastive learning and semantic augmentation: •Sentence Shuffle: A structural distortion in which the sentences of an adaptation appear in an incorrect order. Although the lexical content remains intact, the narrative coherence is disrupted and the meaning is altered. Shuffling-based distortions are widely used in contrastive learning to weaken structural cues [19, 22]. •Sentence Dropout: A deletion-based distortion where one or more sentences are removed from the adaptation. This reduces information content while preserving surface-level fluency, producing subtle losses of meaning. Deletion-based perturbations are common in augmentation frameworks such as EDA [22] and ConSERT [19]. •Mismatch: A cross-textual distortion where an original text is paired with the adaptation of a different story. Although both texts remain independently coherent, their semantic correspondence is broken. This follows derangement-based negative sampling used in sentence similarity modelling [21]. •Paraphrase-based Negatives: A semantic distortion in which the adaptation remains lexically simple and fluent but meaning is altered through omissions, polarity shifts, light contradictions, or changes in quantitative information. Paraphrase-based perturbations are widely used in text simplification and translation augmentation [23]. •Natural Language Inference (NLI) Contradictions: A meaninglevel distortion where the adaptation expresses a proposition that contradicts the content or intent of the original text. These cases preserve surface similarity while inverting core semantic information, following practices in hard-negative mining for natural language inference [20]. Together, these strategies generate structural (shuffle), dropout), semantic (paraphrasing, NLI), and cross-textual (mismatch) distortions, forming a diverse and challenging negative space. 5.2 Data Generation Pipeline The complete negative-generation process was implemented in Python 3.11. The pipeline followed these steps: 1. Load positives: Expert-adapted original/E2R pairs were imported from Excel files, labelled as 1, and tagged as positive. 5 2. Paraphrase generation: For each E2R adaptation, paraphrase-based negatives were generated with OpenAI GPT-4.1-mini4(temperature 0.7, n= 1), keeping lexical simplicity while introducing controlled meaning distortions. Outputs were cached for reproducibility. 3. NLI contradiction mining: The model somosnlp-hackathon-2022/ bertin-roberta-base-zeroshot-esnli5classified candidates as entailment,neutral, or contradiction. Pairs with contradiction probability ≥0.5 were retained. 4. Surface perturbations: •Shuffle: The Natural Language Processing (NLP) library spaCy6 (es core news lg) was used for sentence segmentation, followed by random permutation. •Dropout: Each sentence was removed with probability p= 0.2. •Mismatch: A derangement algorithm ensured that each original text was paired with the adaptation of another story. 5. Balanced sampling: For each positive, a matching negative was selected, and negative types were equalised across categories. 6. Final assembly: Positives and negatives were concatenated, shuffled with a fixed seed (42), and exported to Excel with metadata (Label, neg type). The final dataset was fully balanced. 5.3 Quality Filtering Results To ensure the negatives were sufficiently challenging, lexical similarity with the original E2R texts was measured using BLEU and ROUGE-L. Extremely lowsimilarity pairs (BLEU <0.05, ROUGE-L <0.25) were removed. Negative Type BLEU ROUGE-L Sentence Dropout 0.0770 0.3044 Domain Mismatch 0.1175 0.3537 NLI Contradiction 0.0895 0.2948 Improper Paraphrasing 0.1123 0.3465 Sentence Shuffle 0.1191 0.2623 Table 1: Lexical similarity (BLEU, ROUGE-L) of hard negatives relative to their original E2R adaptations. Higher values indicate surface overlap despite meaning distortion. 4https://platform.openai.com/docs/models/gpt-4.1-mini 5https://huggingface.co/somosnlp-hackathon-2022/bertin-roberta-base-zeroshot-esnli 6https://spacy.io/ 6 6 Model Architectures and Fine-Tuning Three different architectures were fine-tuned and evaluated to determine which approach best captures meaning preservation in Spanish E2R adaptations. All models share a 12-layer Transformer architecture with a hidden size of 768, but differ in their training objectives, language coverage, and intended tasks. •MeaningBERT: A 12-layer, 110M-parameter model originally trained to predict semantic similarity and meaning preservation. Although conceptually aligned with the task, it was developed for English, raising concerns about cross-lingual robustness. •RoBERTa-base-bne: A 12-layer, 125M-parameter monolingual Spanish RoBERTa model. It was fine-tuned using a binary classification setup (0/1 meaning preservation), following a methodology comparable to MeaningBERT. •RoBERTa-BERTScore: A Spanish RoBERTa-base-bne model fine-tuned using a regression objective. Instead of binary labels, it predicts a continuous meaning preservation score based on BERTScore similarity. The main fine-tuning hyperparameters for each model are shown in Table 2. Model Epochs Learning Rate Batch Size Loss Function MeaningBERT 5 2e-5 8 CrossEntropyLoss RoBERTa-base-bne 5 2e-5 8 CrossEntropyLoss RoBERTa-BERTScore 5 2e-5 8 Mean Squared Error (MSE) Table 2: Fine-tuning hyperparameters for the three evaluated models. 7 Results The performance of the three fine-tuned models was evaluated using multiple metrics to capture different dimensions of meaning preservation. Table 3 summarises the global performance across classification and regression settings. Model Eval Loss Accuracy F1 Score ROC-AUC Pearson MSE MeaningBERT 0.6916 0.5310 0.6936 0.536 - - RoBERTa-base-bne 0.4381 0.8064 0.8408 0.825 - - RoBERTa-BERTScore 0.1406 - - - 0.660 0.141 Table 3: Evaluation metrics for the three fine-tuned models. Dashes indicate metrics not applicable to regression-based models. To complement the global metrics, we conducted a deeper quantitative and qualitative analysis of the error patterns observed in the three models. This includes confusion matrices and an examination of error distribution across hardnegative categories. 7 As shown in Figure 1, displayed as confusion matrices: •MeaningBERT shows strong confusion between the two classes, especially misclassifying meaning-preserving pairs (label 1) as negatives. •RoBERTa-base-bne exhibits near-perfect classification with very few false positives or false negatives. •RoBERTa-BERTScore performs better on intermediate meaning levels but struggles when forced into strict binary classification. (a) MeaningBERT (b) RoBERTa-base-bne (c) RoBERTa-BERTScore Figure 1: Confusion matrices for the three fine-tuned models. 8 To better understand how hard negatives affected model behaviour, we examined false positives and false negatives across categories. •MeaningBERT displays low variance but systematic bias. The model collapses toward the negative class: it produces 0 false positives and 12 false negatives, i.e., it fails to detect all positives (FNR = 100%; FPR = 0%). Errors are concentrated on the same class, indicating poor crosslingual transfer and weak sensitivity to meaning preservation in Spanish. •RoBERTa-base-bne shows higher variance but balanced behaviour. It produces 14 false positives out of 60 negatives (FPR ≈23.3%) and 0 false negatives (FNR = 0%). Most false positives arise from mismatch and difficult paraphrase cases, where surface-level similarity misleads the model. •RoBERTa-BERTScore presents medium variance with a slightly more permissive decision boundary. It outputs 18 false positives (FPR ≈30%) and 0 false negatives. Predictions cluster between 0.50 and 0.70, stabilising positives but increasing ambiguity in mismatch and some dropout/rephrase items. In general, MeaningBERT’s errors are systematic and caused by poor generalisation to Spanish, while RoBERTa-base-bne and RoBERTa-BERTScore exhibit distributed errors dominated by false positives on structurally deceptive hard negatives. Among the two Spanish models, RoBERTa-base-bne is the most conservative (lower FPR), while the regression-based RoBERTa-BERTScore is more permissive. Figure 2: Distribution of prediction errors across hard-negative categories. 9