scieee AI-readable full text Open interactive document viewer

Discuit – a tool for dividing items into equal sets

de Kok, Dörte

Abstract

In psycho- and neurolinguistic research, items often need to be partitioned into a specified number of subsets that are as comparable as possible regarding multiple variables. This enables a reliable comparison between experimental conditions. With Discuit (De Kok, 2023), we created a tool that can do this task automatically, using clustering algorithms. The tool takes a spreadsheet as input and generates the required number of sets, matched on the provided data. These data can be categorical (e.g. ‘word class’) and/or continuous (e.g. ‘word frequency’). Finally, one variable can be marked as “primary”. The items will be split perfectly even with regard to this variable (e.g. ‘accuracy in pre- treatment assessment’). We have piloted the tool in a recent aphasia treatment study to split the diagnostic items into to-be- trained and not-to-be-trained items for each of the two participants. In this presentation or poster, I will explain the underlying mechanism of the tool, illustrate how it can be used for different purposes and discuss our pilot study as an example use case.

Full text

1| faculty of arts neurolinguistics 30-05-2023 1| faculty of arts neurolinguistics 30-05-2023 Discuit – a tool for dividing items into equal sets Dörte de Kok [email protected] eScience Center Fellow 2022-2023 2| faculty of arts neurolinguistics 30-05-2023 Problem ›For experiments, treatment studies etc you need 2 (or more) sets of items that are balanced §For various factors, e.g. word frequency, concreteness, word class ›Solution so far §Make 1 long list of items §Split by hand 3| faculty of arts neurolinguistics 30-05-2023 Problem ›But what if this needs to be done for each participant? §Because 1 variable is performance on material in pretest ›Also: often reduction of continuous variables into categories §Highvs low-frequent rather than original values 4| faculty of arts neurolinguistics 30-05-2023 Proposed solution ›Automatic splitting into sets ›Python package: Discuit1 §Open source license §Available through github/pypi/zenodo 5| faculty of arts neurolinguistics 30-05-2023 Discuit - functionality ›Based on k-means/k-mode clustering ›Division into sets based on §Categorical variables (e.g. word class, animacy) §Continuous variables (e.g. word frequency, AoA, concreteness) ›Input list of items (.csv file) is split into desired number of sets based on variables (also in file) ›Statistics to check item division 8| faculty of arts neurolinguistics 30-05-2023 Example use case ›Investigation of treatment efficacy (PhD-research Pauline Cuperus2) ›For each participant: divide material in treated & untreated items ›Example for participant DTR: §107 experimental items (filling in verbs in past & future tense) §Variables: Log10 frequency3, AoA ratings4, concreteness ratings5, transitivity, instrumentality, regularity of past tense form, accuracy in 2 sessions before treatment (for both past and future) 9| faculty of arts neurolinguistics 30-05-2023 Use case – input file Verb Log10 freq3 AoA4 Concr.5 Instr. Transit. Irreg. Past Pre1 Past Pre1 Future Pre2 Past Pre2 Future Box 2.2 4.6 3.81 0 0 0 1 0 1 1 Break 2.35 5.94 3.71 0 1 1 0 0 1 1 Brush 1.15 4.12 4.54 1 1 0 0 0 0 1 Build 2.47 4.64 3.71 0 1 1 1 0 1 1 Catch 2.32 5.2 4.11 0 1 1 0 0 0 1 Cheer 1.35 5.75 2.77 0 0 0 0 0 1 0 Clean 1.84 4.22 3.81 1 1 0 1 0 0 1 Cook 2.15 4.88 4.32 1 0 0 0 1 0 0 … 11| faculty of arts neurolinguistics 30-05-2023 Discuit - Use ›pip install discuit ›python3 discuit/run_discuit.py example/dtr.csv 2 --columns l n n n c c c c c a c Input file with items Number of sets Specify type of variable --runs 3 Specify number of runs 12| faculty of arts neurolinguistics 30-05-2023 Use case – output Verb Log10 freq3 AoA4 Concr.5 Instr. Transit. Irreg. Past Pre1 Past Pre1 Future Pre2 Past Pre2 Future Set Box 2.2 4.6 3.81 0 0 0 1 0 1 1 2 Break 2.35 5.94 3.71 0 1 1 0 0 1 1 2 Brush 1.15 4.12 4.54 1 1 0 0 0 0 1 1 Build 2.47 4.64 3.71 0 1 1 1 0 1 1 1 Catch 2.32 5.2 4.11 0 1 1 0 0 0 1 2 Cheer 1.35 5.75 2.77 0 0 0 0 0 1 0 1 Clean 1.84 4.22 3.81 1 1 0 1 0 0 1 1 Cook 2.15 4.88 4.32 1 0 0 0 1 0 0 2 …