Full text
AvengER: Ensembling and Fine-Tuning LLMs for SELECT Prompts in Entity Resolution Alexandros Zeakis1,2, George Papadakis1, Dimitrios Skoutas2, and Manolis Koubarakis1 1National and Kapodistrian University of Athens, Greece {alzeakis,gpapadis,koubarak}@di.uoa.gr 2Athena Research Center, Greece {azeakis,dskoutas}@athenarc.gr Abstract. Link Discovery plays a vital role in enhancing connections between data sources within the Linked Open Data Cloud. A key aspect of this process is Entity Resolution (ER), which focuses on identifying owl:sameAs relationships between different entity descriptions representing the same real-world object. Recent works on ER have investigated the use of Large Language Models (LLMs) for Entity Matching, showing promising results. In the simplest case, a pair of entities is given as input to an LLM, asking whether these entities match or not. A recent approach introduced SELECT prompts, which include the query entity along with multiple candidates generated by a Blocking method. However, this increases the complexity of the questions posed to LLMs, while being susceptible to position bias among the presented candidates. To address these issues, we introduce AvengER, a novel approach to SELECT prompts for LLM-based Matching that effectively handles both unsupervised and supervised settings. For the former, we introduce a hybrid approach that utilizes an ensemble of open-source medium-size models (8b) and also selectively leverages (in just 12% of the cases) an external larger-size Judge (32b), thus balancing high accuracy with computational efficiency. For the latter, we use a dataset containing data from multiple domains to fine-tune a medium-size model so that it surpasses both its pre-trained version and pre-trained large-size models (GPT-3.5). Keywords: LLMs ·Entity Matching ·Model ensembles ·Fine-tuning 1 Introduction Entity Resolution (ER) facilitates the integration of independent data sources by identifying their duplicates or matches, i.e., different entity descriptions that refer to the same real-world object [2]. ER constitutes a crucial and challenging task for realizing the vision of Linked Data, facilitating a diverse set of applications, from common data analytics tasks to advanced question answering [3, 4]. The Linked Open Data (LOD) Cloud3forms a central part of the Semantic Web, 3https://lod-cloud.net/#about
2 Al. Zeakis et al. Do these two entity descriptions refer to the same real-world entity? (1) [COL] name [VAL] Sony Turntable - PSLX350H [COL] description [VAL] Sony Turntable - PSLX350H/ Belt Drive System/ 33-1/3 and 45 RPM Speeds/ Servo Speed Control/ Supplied Moving Magnet Phono Cartridge/ Bonded Diamond Stylus/ Static Balance Tonearm/ Pitch Control (2) [COL] name [VAL] Sony Picture Station Digital Photo Printer - DPPFP95 LLM Response: No MATCH Prompt Select an entity description from the following candidates that refers to the same real-world entity as the given entity description. Answer with the corresponding entity description number surrounded by "[]" or "[0]" if there is none. Given entity record: [COL] name [VAL] Sony Turntable - PSLX350H [COL] description [VAL] Sony Turntable - PSLX350H/ Belt Drive System/ 33-1/3 and 45 RPM Speeds/ Servo Speed Control/ Supplied Moving Magnet Phono Cartridge/ Bonded Diamond Stylus/ Static Balance Tonearm/ Pitch Control Candidate records: [1] [COL] name [VAL] Denon DP-29F Analog Record Turntable - DP29F [COL] description [VAL] Belt Drive [2] [COL] name [VAL] Sony PS-LX350H Belt-Drive Turntable [3] [COL] name [VAL] Sony PSLX300USB USB Record Turntable [COL] description [VAL] Belt Drive [COL] price [VAL] 119.72 [4] [COL] name [VAL] Sony Picture Station Digital Photo Printer - DPPFP95 [5] [COL] name [VAL] Denon DP-300F Analog Record Turntable - DP300F [COL] description [VAL] Belt Drive [COL] price [VAL] 328.99 LLM Response: [2] SELECT Prompt Fig. 1: Prompts for different tasks. The first decides for a pair of entities, the second selects from a given list of choices. expanding from 570 datasets in 2014 to 1,573 in 2024. However, its interconnectivity remains limited, with only 17,882 links recorded as of December 2024. On average, each dataset connects to just 11 others, which accounts for roughly 1% of all possible links. To improve this connectivity, Link Discovery automates the identification of relationships between entity descriptions across datasets. A typical end-to-end ER pipeline consists of two steps: Blocking and Matching [3]. The first step identifies candidate pairs that are likely to match, while the second applies a more precise method to determine which pairs are genuine matches. The last few years, Blocking has been studied with supervised settings [1] and Matching has been dominated by approaches that combine pretrained transformer models with Deep Learning, such as DITTO [6] and Unicorn [16], while more recent studies leverage Large Language Models (LLMs) [12, 15, 18]. Most of the latter approaches feed LLMs with MATCH prompts, which ask whether two particular entities are matching [12, 15]. Yet, these prompts are impractical and time-consuming in the context of end-to-end ER pipelines, where Blocking typically yields kcandidates per entity [11, 9]. To address this, SELECT prompts are proposed in [18], including all Blocking candidates in a single prompt. This transforms Matching from a binary decision (“Yes / No”) into a “select which of the following” response, as shown in Figure 1. A major challenge with SELECT prompts is position bias, where LLM responses depend on the duplicate’s position among kcandidates [18]. This is indirectly addressed through a post-Blocking refinement step that reduces candidates using k intermediate MATCH prompts, trading efficiency for effectiveness. In this work, we propose AvengER, a comprehensive approach that enhances the SELECT prompts by minimizing the impact of position bias in a more efficient and effective way. Rather than relying on intermediate MATCH prompts,
AvengER 3 AvengER applies post-Blocking refinement through reciprocity: two entities are candidates only if each one is among the top-k most similar entities for the other. This strategy significantly reduces the number of candidates per entity, while conveying a minor decrease in Blocking effectiveness. To further enhance Matching performance, AvengER employs multiple LLMs, combining their outputs through diverse voting strategies: internal ones, where the ensemble of models decides independently, external ones, where a separate model acts as a judge, and hybrid ones, where a judge decides in case the ensemble is not certain. Despite the higher run-times, our thorough experimental analysis demonstrates that AvengER’s combination of reciprocity and voting schemes offers a better balance between effectiveness and efficiency under the unsupervised settings. Under the supervised settings, where annotated data are available, AvengER fine-tunes medium-sized models, improving the performance of SELECT prompts to such an extent that position bias is mitigated without the need to reduce the candidate pairs generated by Blocking. Their performance is also significantly higher than their pre-trained version and than large-sized models. Overall, we make the following contributions: –We introduce AvengER, a novel LLM-based Matching approach that leverages SELECT prompts, handling both unsupervised and supervised settings. –Under unsupervised settings, AvengER suggests a categorization of existing Post-Blocking strategies and employs a new one that drastically reduces candidate pairs based on reciprocity. –AvengER further improves ER performance under unsupervised settings through voting strategies that combine the results of an ensemble of LLMs. –Under the supervised settings, AvengER fine-tunes medium-sized LLMs to maximize the robustness of SELECT prompts and their mitigate position bias, significantly outperforming their pre-trained version and large-sized LLMs. –We demonstrate the superiority of AvengER compared to baselines and state-of-the-art through an extensive experimental study that uses several real-world datasets. 2 Related Work We organize the recent works on LLMs for ER in the following categories: Blocking. The state-of-the-art Blocking methods detect the kmost likely matches per query entity [9, 11]. Higher values for ktypically increase Blocking recall, at the cost of more candidate pairs and, thus, lower Blocking precision. For this reason, the goal of Blocking is to achieve a good balance between recall and precision. To improve this balance, a hybrid post-Blocking refinement strategy is proposed in [18]: it constructs “MATCH” prompts, which return a “yes / no” answer, and “COMPARE” prompts, which choose the most likely match among two candidates for the query entity at hand. Both prompts are applied to a
4 Al. Zeakis et al. Matching Candidate Generation Voting Model Inference Standard Blocking Post-Blocking Refinement Prompt Construction D1 D2 C Fine-Tuning M Fig. 2: The phases and steps of AvengER. Fine-Tuning the model is optional (dashed). medium-sized model, yielding a ranking of options, which then produce the topNresults, where Nis a parameter defined by experimentation. The overhead, though, is much larger than traditional Blocking methods, due to the large number of prompts. Matching. In Matching we have two distinct categories. In Prompt Engineering, typically, a matching prompt consists of three parts [12]: (a) the task description, which describes Matching as the task of determining whether two entity descriptions refer to the same real-world object; (b) a serialization of the entity descriptions of the given candidate pair; (c) optionally, a set of examples, with each one comprising a candidate pair along with its label (match / non-match). These examples yield few-shot prompts, whereas their absence yields zero-shot prompts. A thorough investigation of variants of both types is presented in [12]. Prompts containing multiple candidates, instead of a single pair, have been investigated in [18]. Fine-Tuning improves the accuracy of a pre-trained LLM on a specific task by updating its parameters on a new dataset [15]. To this end, the LLM is partially retrained using <input, output> pairs of representative examples of the desired behavior. A thorough study of the impact of this approach on Matching is presented in [15], which emphasizes the generalization across different domains and the impact of training dataset size. A related approach leverages an LLM to generate explanations and then uses these explanations to fine-tune a PLM (FlanT5-base) to perform the matching [17]. Augmentation strategies to increase the amount of training data have also been considered in [19]. Finally, a finetuned model using transfer learning can outperform competitors on zero-shot prompts, while reducing inference time by nearly 4,000 times [22]. 3AvengER Components AvengER consists of two main phases, as shown in Figure 2: Candidate Generation and Matching. The former receives as input a pair of data sources, D1and D2, and returns as output a set of candidate pairs C, such that Blocking recall and precision are maximized. The second phase receives as input the set of candidate pairs Cand returns as output a set of matching pairs M, such that ER recall and precision are maximized. Putting it all together, Algorithm 1 presents the main steps of AvengER in pseudo-code. Initially, Blocking is performed and
AvengER 5 Algorithm 1 AvengER’s end-to-end algorithm. Require: Query entity q, blocking parameter k, post-blocking refinement parameter N, List of models models, collection C Ensure: Final result result 1: candidates ←blocking(q, k, C) 2: candidates ←post_blocking_refinement(q, N, candidates) 3: prompt ←construct_prompt(q, candidates) 4: responses ← {} 5: for each model in models do 6: responses.add(model.inference(prompt)) 7: end for 8: result ←voting(responses) refined with the post-Blocking approaches presented in Section 3.1 (Lines 1-2). Then, the prompt is constructed by one of the methods discussed in Section 3.2 (Line 3). The prompt is then passed to Model Inference and then, the responses of the individual models are aggregated (Lines 4-7). Finally, a voting scheme is optionally applied (Line 8). Fine-Tuning, if needed, is a process done offline. Below, we explain their functionality in more detail. 3.1 Candidate Generation Based on [21], AvengER applies a state-of-the-art approach to Standard Blocking: it uses S-GTR-T5 to vectorize the serialized entity descriptions and then performs approximate nearest neighbor (NN) search with FAISS4in order to efficiently detect the kmost similar descriptions per query entity. S-GTR-T5, which extends GTR [10], is a dual encoder that transforms two pieces of text into two dense vectors. GTR models are built on top of T5 [13], an encoder-decoder model that aims to unify all natural language processing tasks under a single model. They also use the Siamese architecture in [14] for efficient and scalable similarity computations. Typically, Standard Blocking achieves high recall by generating large sets of candidates that involve repeated and non-matching pairs. This results in low Blocking precision. AvengER aims to significantly increase Blocking precision at an insignificant cost in recall through the Post-Blocking Refinement step. We suggest a taxonomy of existing and new Post-Blocking strategies, that involves three categories for removing repeated and non-matching pairs: 1. Ranking: Assuming that every candidate pair is associated with a matching probability, this strategy simply retains the top-N weighted ones; N(≫k) sets the maximum number of candidates forwarded to Matching, with its actual value set arbitrarily or according to experimental results. AvengER conveys two implementations of this strategy: (i) COS-ST5 uses the cosine similarity scores between embedding vectors already computed during the 4https://github.com/facebookresearch/faiss
6 Al. Zeakis et al. NN search of Standard Blocking. (ii) COMPARE prompts involve a query entity along with two of the candidates determined by Standard Blocking, requesting the LLM to rank them from the most to the least similar one [18]. This requires k 2invocations per query entity. 2. Filtering: This strategy reduces the number of candidates by pruning the pairs that are more unlikely to be matching. AvengER conveys two implementations of this strategy: (i) Reciprocal pruning performs NN search for the entity descriptions of both D1and D2, retaining only the pairs (ei, ej), where ejis among the top-kcandidates of eiand vice versa for ei. (ii) MATCH prompts assess whether every candidate pair involves duplicate descriptions and prunes those that do not [18]. In both cases, the number of candidates differs among the input entities (there might even be cases with empty lists). 3. Hybrid: This strategy combines the other two, retaining only the top-N weighted candidate pairs generated by Filtering. In the case of MATCH prompts, this can be accomplished by first requesting that each pair is assigned a matching probability and then ranking the results accordingly. 3.2 Matching We present the steps of AvengER’s Matching in the order they are applied. Prompt Construction. We now elaborate on the structure of the prompt used as a first step for AvengER’s Matching. Basic Prompt. The standard prompt for EM, first introduced in [12], consists of two core parts. The first one is the task description, which specifies that the goal is to find entities that refer to the same real-world entity. This can be formulated as a “Yes/No” question [12, 15, 18]. When given a pair of entities, the prompt can be formulated as “select which of the following” [18]. The second part pertains to the entity serialization. Many approaches have been proposed in the literature, but since an entity consists of triples < subject, predicate, object >, two are the most common options: the schema-agnostic one, where only the objects are concatenated into a single sentence [21], and the schema-aware one, where the format is [COL]predicate[V AL]object [6]. We propose extensions for both parts. In the task description, we focus exclusively on SELECT prompts, asking LLMs not only to respond to the query, but also to explain their choice. As experimentally shown in Section 4, this leads to a better performance for some models. We also tried another prompt, where we instructed the model to provide a confidence estimation for its choice [20]. Instead of a numeric score, we requested a qualitative response among three options: “Certain”, “Moderately-certain” and “Uncertain”. Note that in any prompt type, the model can respond that it cannot decide, returning the special answer of [0], which indicates that none of the choices is correct. Both prompts can be found in Figures 5 and 6. For the entity serialization part, we also examine a semantically richer description that is generated through a pre-processing step that provides the LLM
AvengER 7 with the serialized entity and requests a summary in natural language. This is demonstrated in Figure 7. Then, inside the EM prompt, we replace the entity serializations with their summary. In-context Learning. Adding examples before the main question in prompts helps LLMs to better understand the given task [12, 15]. This has been widely exploited in EM literature, but only for MATCH prompts [18]. To the best of our knowledge, this work is the first to explore enriching SELECT prompts with examples in an experimental setting, yielding few-shot EM SELECT prompts. We have selected examples randomly, but there is room to investigate more advanced strategies in the future. Model Inference. For model inference, some works [12, 18, 5] use pre-trained models, particularly the large ones like GPT-3.5. In other works [12, 15], FineTuning shows to improve the performance of a model. For example, in [12] a fine-tuned GPT-3.5 can outperform a pre-trained GPT-4 on certain benchmarks. In our task, the problem is transformed from binary classification to selecting the index position, which is heavily dependent on the order of the choices within the prompt. To enhance robustness, we introduce prompt permutations to train the model on variations in choice positions, thus reducing position bias. Voting. All existing works on LLM-based EM, rely on individual LLMs. In this work, we go beyond them by considering ensembles of LLMs. We actually examine multiple strategies for combining the individual decisions of multiple LLMs. These strategies depend on whether the ensemble determines the vote internally or relies on an external judge to make the final decision. Internal Vote. We introduce the following internal voting strategies: –Majority-Vote. Each model produces one answer, and the most-voted answer is selected. If no majority is reached, the ensemble yields an indecisive vote, similar to [0], when the true answer is not among the options. –Weighted-Vote. Each model’s contribution is weighted based on its performance, with precomputed global weights derived from metrics like F1-score (see Section 4). The answer with the highest total weight is selected. –Confidence-Vote. Instead of relying on pre-computed weights, this method assigns weights dynamically based on the confidence scores generated by each model’s response. By utilizing a confidence prompt, we derive a unique weight for each model per response. This approach adapts to the context of each query, making it more flexible than static weight assignments. We introduce 3 certainty levels, {Certain, Moderately Certain, Uncertain}, which are evaluated to {0.5, 0.75, 1}, resembling Shannon Entropy. –Aggregated Majority-Vote. This approach involves asking each model Ntimes with different permutations of answer choices in each iteration. Instead of relying on mvotes from the original models, this process generates m×N votes. By aggregating these votes, the ensemble improves the overall robustness, as models demonstrating consistent predictions across iterations have a stronger contribution to the final decision.
8 Al. Zeakis et al. D1D2D3D4D5D6D7D8 Dat1Abt Amz ACM IMDb IMDb TMDb Wmt DBLP Dat2Buy GPr. DBLP TMDb TVDB TVDB Amz Scholar |R1|1,076 1,354 2,294 5,118 5,118 6,056 2,554 2,516 |R2|1,076 3,039 2,616 6,056 7,810 7,810 22,074 61,353 |P1|3 4 4 13 13 30 6 4 |P2|3 4 4 30 9 9 6 4 |D|1,076 1,104 2,224 1,968 1,072 1,095 853 2,308 Table 1: Statistics of the real ER datasets used in the experiments, showing the number of resources (|Rx|), distinct predicates (|Px|) and duplicates (|D|). Regarding computational differences, notice that Weighted-vote requires training, whereas Confidence-vote does not; for inference, there is no significant difference in cost. Notice also that Confidence-Vote expresses each model’s confidence qualitatively, while Aggregated Majority-Vote quantifies it. External Vote. We introduce the following external voting strategies: –Judge-vote. The idea of using an LLM as a judge for evaluating other LLMs on open-ended questions was proposed in [23]. We apply this by merging individual LLM responses into a prompt for the judge model, since each LLM accompanies its response with an explanation regarding the given answer. Due to the large prompt size, the judge model requires an LLM with a large context window, involving more parameters and resources. –Hybrid-vote. To mitigate the extra cost of asking a heavier judge model, we also propose a hybrid voting strategy. In this approach, we assume that the top-1 answer accumulates a score Σ, which depends on the internal voting strategy (majority, weighted, etc). We also define a threshold θ∈[0,1]. If Σ < θ, we invoke a Judge model to make a decision. In this way, we save computational cost, by reducing the cases where the Judge model is invoked. 4 Evaluation 4.1 Experimental Setup Datasets. We use 8 real-world ER datasets that are widely used in the literature [8, 18]. Our approach assumes that each entity comprises a set of <predicate, object> pairs, which are then serialized into a single textual description. This approach accommodates both relational and Knowledge Graph data. For example, the movie datasets (D4, D5, D6) stem from the MovieGraphBenchmark5, using URIs as predicates. Table 1 shows the statistics of these datasets. Following [18], we sampled 300 annotated resources per dataset and 100 non-annotated entities (when possible) for testing. For training, we sampled again 300 labeled and 100 unseen pairs, which were also used in the Weighted-vote. For in-context learning, we used 1 positive example, where the answer was among the provided options, 5https://github.com/ScaDS/MovieGraphBenchmark
AvengER 9 and 1 negative example, where the answer was not included in the provided options. Both examples were randomly selected. Models. We used five open-source LLMs: Llama-3.1:8b, Mistral-v0.3:7b, Orca2:7b, Qwen-2.5:32b6, and Flan-T5-XL7. These models were selected to ensure comparability with state-of-the-art works while leveraging their specific strengths. Llama and Mistral were used in almost all experiments and were also utilized in [18], facilitating direct comparisons. Orca was included in the ensemble due to its strong reasoning capabilities [7], which are beneficial for decision-making in multiple-choice settings. Qwen, being a larger model, served as a Judge in External Voting, while Flan was used for MATCH prompts in ComEM, as in [18]. For brevity, we omit their parameters in subsequent references. Baseline LLM-based EM. The top performing LLM-based approach to EM is presented in [18]. It utilizes Sparkly [11] for Blocking and Hybrid for PostBlocking Refinement. The constructed prompt is based on the basic SELECT format with no examples and the model is a pre-trained GPT-3.5-turbo-0613. Evaluation Metrics. For all experiments, we report F-Measure (F1). For the Voting Strategies, we also report Precision and Recall, whereas for Blocking, we also report the corresponding Blocking Precision, Recall and F1. We also report the average response size in characters. Settings. For Standard Blocking, we used S-GTR-T5 for entity serialization in combination with FAISS for indexing and searching the top-N candidate pairs per entity [21]. For the pre-trained LLMs, we used models provided by Ollama8, while for their fine-tuning we used Unsloth9. All of our code, datasets and finetuned models are publicly available10. All experiments were executed on a server with Ubuntu 20.04, AMD Ryzen Threadripper 3960X 24-Core processor, 256 GB RAM and an RTX 4090 GPU. 4.2 Candidate Generation Strategy D1 D2 D3 D4 D5 D6 D7 D8 Mean Ranking@10 0.84 0.61 0.83 0.69 0.60 0.63 0.71 0.70 0.70 Rec. Pruning 0.87 0.59 0.86 0.74 0.64 0.67 0.72 0.76 0.73 Hybrid@4 0.47 0.41 0.84 0.88 0.64 0.82 0.30 0.75 0.64 Table 2: Comparison on F1 for different Post-Blocking Refinement approaches for SELECT prompts and using Llama-3.1 for inference. In this experiment, we assess the effect of Post-Blocking Refinement on the performance of AvengER. We actually measure the evolution of Blocking effectiveness with regard to the Blocking parameter k∈[1,10] with a step of 1. 6https://ollama.com/library/{X} llama3.1:8b, mistral:7b, orca2:7b, qwen2.5:32b 7https://huggingface.co/google/flan-t5-xl 8https://ollama.com/ 9https://unsloth.ai/ 10 https://github.com/alexZeakis/AvengER
16 Al. Zeakis et al. References 1. Brinkmann, A., Shraga, R., Bizer, C.: Sc-block: Supervised contrastive blocking within entity resolution pipelines. In: ESWC (1). Lecture Notes in Computer Science, vol. 14664, pp. 121–142. Springer (2024) 2. Christophides, V., Efthymiou, V., Palpanas, T., Papadakis, G., Stefanidis, K.: An overview of end-to-end entity resolution for big data. ACM CSUR 53(6), 127:1– 127:42 (2021) 3. Christophides, V., Efthymiou, V., Stefanidis, K.: Entity Resolution in the Web of Data. Morgan & Claypool (2015) 4. Dong, X.L., Srivastava, D.: Big data integration. PVLDB 6(11), 1188–1189 (2013) 5. Fan, M., Han, X., Fan, J., Chai, C., Tang, N., Li, G., Du, X.: Cost-effective incontext learning for entity resolution: A design space exploration. In: ICDE. pp. 3696–3709. IEEE (2024) 6. Li, Y., Li, J., Suhara, Y., Doan, A., Tan, W.: Deep entity matching with pre-trained language models. Proc. VLDB Endow. 14(1), 50–60 (2020) 7. Mitra, A., Corro, L.D., Mahajan, S., Codas, A., Simões, C., Agrawal, S., Chen, X., Razdaibiedina, A., Jones, E., Aggarwal, K., Palangi, H., Zheng, G., Rosset, C., Khanpour, H., Awadallah, A.: Orca 2: Teaching small language models how to reason. CoRR abs/2311.11045 (2023) 8. Mudgal, S., Li, H., Rekatsinas, T., Doan, A., Park, Y., Krishnan, G., Deep, R., Arcaute, E., Raghavendra, V.: Deep learning for entity matching: A design space exploration. In: SIGMOD Conference. pp. 19–34. ACM (2018) 9. Neuhof, F., Fisichella, M., Papadakis, G., Nikoletos, K., Augsten, N., Nejdl, W., Koubarakis, M.: Open benchmark for filtering techniques in entity resolution. VLDB J. 33(5), 1671–1696 (2024) 10. Ni, J., Qu, C., Lu, J., Dai, Z., Ábrego, G.H., Ma, J., Zhao, V.Y., Luan, Y., Hall, K.B., Chang, M., Yang, Y.: Large dual encoders are generalizable retrievers. In: EMNLP. pp. 9844–9855. Association for Computational Linguistics (2022) 11. Paulsen, D., Govind, Y., Doan, A.: Sparkly: A simple yet surprisingly strong TF/IDF blocker for entity matching. Proc. VLDB Endow. 16(6), 1507–1519 (2023) 12. Peeters, R., Bizer, C.: Entity matching using large language models. CoRR abs/2310.11244 (2023) 13. Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J.: Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 140:1–140:67 (2020) 14. Reimers, N., Gurevych, I.: Sentence-bert: Sentence embeddings using siamese bertnetworks. In: EMNLP/IJCNLP (1). pp. 3980–3990. Association for Computational Linguistics (2019) 15. Steiner, A., Peeters, R., Bizer, C.: Fine-tuning large language models for entity matching. CoRR abs/2409.08185 (2024) 16. Tu, J., Fan, J., Tang, N., Wang, P., Li, G., Du, X., Jia, X., Gao, S.: Unicorn: A unified multi-tasking model for supporting matching tasks in data integration. Proc. ACM Manag. Data 1(1), 84:1–84:26 (2023) 17. Wadhwa, S., Krishnan, A., Wang, R., Wallace, B.C., Kong, C.: Learning from natural language explanations for generalizable entity matching. CoRR abs/2406.09330 (2024) 18. Wang, T., Lin, H., Chen, X., Han, X., Wang, H., Zeng, Z., Sun, L.: Match, compare, or select? an investigation of large language models for entity matching. CoRR abs/2405.16884 (2024)
AvengER 17 19. Xia, Y., Chen, J., Li, X., Gao, J.: Aprompt4em: Augmented prompt tuning for generalized entity matching. CoRR abs/2405.04820 (2024) 20. Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., Hooi, B.: Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms. In: ICLR. OpenReview.net (2024) 21. Zeakis, A., Papadakis, G., Skoutas, D., Koubarakis, M.: Pre-trained embeddings for entity resolution: An experimental analysis. Proc. VLDB Endow. 16(9), 2225– 2238 (2023) 22. Zhang, Z., Groth, P., Calixto, I., Schelter, S.: Anymatch - efficient zero-shot entity matching with a small language model. CoRR abs/2409.04073 (2024) 23. Zheng, L., Chiang, W., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E.P., Zhang, H., Gonzalez, J.E., Stoica, I.: Judging llm-as-a-judge with mt-bench and chatbot arena. In: NeurIPS (2023)
18 Al. Zeakis et al. A B C D E Ranking - COS A D C E B A D Ranking - Compare D C E A B D C Filtering - Reciprocal A B C D E D Filtering - Matching A B C D E D E D LLM LLM F D k=2 k=2 k≤2 Hybrid A B C D E LLM k≤2 Original Candidates Fig. 4: Post-Blocking Strategies. Select a record from the following candidates that refers to the same real-world entity as the given record. Answer with the corresponding record number surrounded by "[]" or "[0]" if there is none and explain briefly why. Given entity record: [COL] name [VAL] Sony Turntable - PSLX350H [COL] description [VAL] Sony Turntable - PSLX350H/ Belt Drive System/ 33-1/3 and 45 RPM Speeds/ Servo Speed Control/ Supplied Moving Magnet Phono Cartridge/ Bonded Diamond Stylus/ Static Balance Tonearm/ Pitch Control Candidate records: [1] [COL] name [VAL] Denon DP-29F Analog Record Turntable - DP29F [COL] description [VAL] Belt Drive [2] [COL] name [VAL] Sony PS-LX350H Belt-Drive Turntable [3] [COL] name [VAL] Sony PSLX300USB USB Record Turntable [COL] description [VAL] Belt Drive [COL] price [VAL] 119.72 [4] [COL] name [VAL] Sony Picture Station Digital Photo Printer - DPPFP95 [5] [COL] name [VAL] Denon DP-300F Analog Record Turntable - DP300F [COL] description [VAL] Belt Drive [COL] price [VAL] 328.99 LLM Response: [2] The candidate record [2] refers to the same real-world entity as the given record, specifically a Sony turntable model (PS-LX350H). The names and descriptions in both records match exactly. EXPLAIN Prompt Fig. 5: Prompt for EXPLAIN. Select a record from the following candidates that refers to the same real-world entity as the given record. Answer with the corresponding record number surrounded by "[]" or "[0]" if there is none. Accompany your answer by one of the following confidence measures: {Certain},{Moderately-certain}, {Uncertain}. Given entity record: [COL] name [VAL] Sony Turntable - PSLX350H [COL] description [VAL] Sony Turntable - PSLX350H/ Belt Drive System/ 33-1/3 and 45 RPM Speeds/ Servo Speed Control/ Supplied Moving Magnet Phono Cartridge/ Bonded Diamond Stylus/ Static Balance Tonearm/ Pitch Control Candidate records: [1] [COL] name [VAL] Denon DP-29F Analog Record Turntable - DP29F [COL] description [VAL] Belt Drive [2] [COL] name [VAL] Sony PS-LX350H Belt-Drive Turntable [3] [COL] name [VAL] Sony PSLX300USB USB Record Turntable [COL] description [VAL] Belt Drive [COL] price [VAL] 119.72 [4] [COL] name [VAL] Sony Picture Station Digital Photo Printer - DPPFP95 [5] [COL] name [VAL] Denon DP-300F Analog Record Turntable - DP300F [COL] description [VAL] Belt Drive [COL] price [VAL] 328.99 LLM Response: [Certain] [2] The given entity record refers to a Sony turntable model, and candidate record [2] is the exact match for this real-world entity. CONFIDENCE Prompt Fig. 6: Prompt for CONFIDENCE. Appendix Fine-Tuning Configuration: For the training data described in Section 4.1, we used zero-shot SELECT prompts with schema-aware serialization and provided
AvengER 19 Given this concatenated entity description: [COL] name [VAL] Sony Turntable - PSLX350H [COL] description [VAL] Sony Turntable - PSLX350H/ Belt Drive System/ 33-1/3 and 45 RPM Speeds/ Servo Speed Control/ Supplied Moving Magnet Phono Cartridge/ Bonded Diamond Stylus/ Static Balance Tonearm/ Pitch Control Generate a meaningful entity description in natural language. In the output give only the new description. LLM Response: The Sony PSLX350H is a belt-driven turntable that supports 33-1/3 and 45 RPM playback, featuring servo speed control and a built-in moving magnet phono cartridge with bonded diamond stylus, as well as a static balance tonearm and pitch control for precise audio reproduction. Summary Prompt Fig. 7: Prompts for creating a Summary for dataset description. D1 D2 D3 D4 D5 D6 D7 D8 Datasets 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 F1 (a) Pre-trained llama3.1 D1 D2 D3 D4 D5 D6 D7 D8 Datasets 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 F1 (b) Fine-Tuned Mistral-v0.3instruct Fig. 8: Measuring variance in performance for permutated prompts between the best of pre-trained and fine-tuned models . 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 Left-to-Right Right-to-Left Reciprocal 2 4 6 8 10 Blocking k 0.00 0.25 Blocking F1 D1 2 4 6 8 10 Blocking k 0.0 0.1 Blocking F1 D2 2 4 6 8 10 Blocking k 0.00 0.25 Blocking F1 D3 2 4 6 8 10 Blocking k 0.00 0.25 Blocking F1 D4 2 4 6 8 10 Blocking k 0.00 0.25 Blocking F1 D5 2 4 6 8 10 Blocking k 0.0 0.2 Blocking F1 D6 2 4 6 8 10 Blocking k 0.0 0.1 Blocking F1 D7 2 4 6 8 10 Blocking k 0.0 0.2 Blocking F1 D8 Fig. 9: Comparing Post-Blocking Refinement strategies for different values of k. the answer in brackets. For each entity we created 3 such prompts, with the answers shuffled, to cover the issue of position bias. For training, we used LoRA with rank 16, α= 16, and no dropout for efficient adaptation while managing memory constraints. Training was conducted for three epochs with a learning rate of 2·10−4, gradient accumulation steps of 4, and the AdamW 8-bit optimizer. To reduce memory usage, we applied 4-bit quantization to non-instruction-tuned models.