Language Models from the Sweatshop? Helping Researchers Avoid Ethical and Legal Issues With Off-The-Shelf Software
Author(s): Hendrik Erz, Sebastian Gießler
Friday 16 | 11:20-11:40
Room: TP45
Session: Pretrained models and sociology: ethical, methodological and theoretical considerations
Social scientists increasingly make use of language models to analyse the growing amount of textual data with the help of computers (Grimmer et al., 2022; Bonikowski and Nelson, 2022; Knight, 2022).
Today’s computationally expensive Large Language Models (LLMs) (Bender et al., 2021) are “pre-trained” by a few institutions that possess the necessary compute power (Whittaker, 2021). In this “pre-training” process, LLMs are trained on large but generic text corpora. These models can then be “fine-tuned” for specific use-cases.
By using these models, researchers waive control over methods and data they use, taking on the role of consumers of off-the-shelf software. One cannot assume, however, that training data fulfils ethical standards and that the resulting models do not pose legal risks. Additionally, LLMs come with methodological and theoretical caveats that remain underexplored.
Model architecture and data sets of many commercial models are inaccessible and companies utilise underpaid workers (Xiang, 2023) or copyrighted material (Vincent, 2022). Furthermore, normative values encoded into LLMs can be mismatched with the theoretical assumptions of researchers (“model alignment”; Norvig and Davis 2010). These value-judgments and design decisions can harm research subjects (Biddle 2022).
Researchers should therefore explicitly vet LLMs to avert harm for both research subjects and themselves. Biases can affect study results (Akter et al., 2021); as can “supply-chain attacks” (Szegedy et al., 2014) or “poisoning” the training data (Carlini et al., 2023).
This work develops a checklist to verify and increase trust in pre-trained language models. It does not require researchers to fully understand a model. Instead, it ensures that LLMs can be used confidently, minimising the ethical and legal risks that arise with the use of improperly curated off-the-shelf models.
The checklist covers ethical and legal (accountability, adversarial attacks, consent), methodological (algorithmic and data bias, machine reasoning), and theoretical (alignment) issues.