Skip to main content

Showing 1–26 of 26 results for author: Röttger, P

Searching in archive cs. Search in all archives.
.
  1. arXiv:2406.14508  [pdf, other

    cs.CL cs.AI cs.CY cs.HC

    Evidence of a log scaling law for political persuasion with large language models

    Authors: Kobi Hackenburg, Ben M. Tappin, Paul Röttger, Scott Hale, Jonathan Bright, Helen Margetts

    Abstract: Large language models can now generate political messages as persuasive as those written by humans, raising concerns about how far this persuasiveness may continue to increase with model size. Here, we generate 720 persuasive messages on 10 U.S. political issues from 24 language models spanning several orders of magnitude in size. We then deploy these messages in a large-scale randomized survey ex… ▽ More

    Submitted 20 June, 2024; originally announced June 2024.

    Comments: 16 pages, 4 figures

  2. arXiv:2405.09482  [pdf, other

    cs.CL

    Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts

    Authors: Donya Rooein, Paul Rottger, Anastassia Shaitarova, Dirk Hovy

    Abstract: Using large language models (LLMs) for educational applications like dialogue-based teaching is a hot topic. Effective teaching, however, requires teachers to adapt the difficulty of content and explanations to the education level of their students. Even the best LLMs today struggle to do this well. If we want to improve LLMs on this adaptation task, we need to be able to measure adaptation succes… ▽ More

    Submitted 6 June, 2024; v1 submitted 15 May, 2024; originally announced May 2024.

  3. arXiv:2404.17874  [pdf, other

    cs.CL

    From Languages to Geographies: Towards Evaluating Cultural Bias in Hate Speech Datasets

    Authors: Manuel Tonneau, Diyi Liu, Samuel Fraiberger, Ralph Schroeder, Scott A. Hale, Paul Röttger

    Abstract: Perceptions of hate can vary greatly across cultural contexts. Hate speech (HS) datasets, however, have traditionally been developed by language. This hides potential cultural biases, as one language may be spoken in different countries home to different cultures. In this work, we evaluate cultural bias in HS datasets by leveraging two interrelated cultural proxies: language and geography. We cond… ▽ More

    Submitted 27 April, 2024; originally announced April 2024.

    Comments: Accepted at WOAH (NAACL 2024)

  4. arXiv:2404.17047  [pdf, other

    cs.LG

    Near to Mid-term Risks and Opportunities of Open-Source Generative AI

    Authors: Francisco Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schroeder de Witt, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Botos Csaba, Fabro Steibel, Fazl Barez, Genevieve Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph Marvin Imperial, Juan A. Nolazco-Flores, Lori Landay, Matthew Jackson, Paul Röttger, Philip H. S. Torr, Trevor Darrell, Yong Suk Lee, Jakob Foerster

    Abstract: In the next few years, applications of Generative AI are expected to revolutionize a number of different areas, ranging from science & medicine to education. The potential for these seismic changes has triggered a lively debate about potential risks and resulted in calls for tighter regulation, in particular from some of the major tech companies who are leading in AI development. This regulation i… ▽ More

    Submitted 24 May, 2024; v1 submitted 25 April, 2024; originally announced April 2024.

    Comments: Accepted to ICML'24 as a position paper

  5. arXiv:2404.16019  [pdf, other

    cs.CL

    The PRISM Alignment Project: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models

    Authors: Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, Bertie Vidgen, Scott A. Hale

    Abstract: Human feedback plays a central role in the alignment of Large Language Models (LLMs). However, open questions remain about the methods (how), domains (where), people (who) and objectives (to what end) of human feedback collection. To navigate these questions, we introduce PRISM, a new dataset which maps the sociodemographics and stated preferences of 1,500 diverse participants from 75 countries, t… ▽ More

    Submitted 24 April, 2024; originally announced April 2024.

  6. arXiv:2404.12241  [pdf, other

    cs.CL cs.AI

    Introducing v0.5 of the AI Safety Benchmark from MLCommons

    Authors: Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed, Victor Akinwande, Namir Al-Nuaimi, Najla Alfaraj, Elie Alhajjar, Lora Aroyo, Trupti Bavalatti, Max Bartolo, Borhane Blili-Hamelin, Kurt Bollacker, Rishi Bomassani, Marisa Ferrara Boston, Siméon Campos, Kal Chakra, Canyu Chen, Cody Coleman, Zacharie Delpierre Coudert, Leon Derczynski, Debojyoti Dutta, Ian Eisenberg, James Ezick, Heather Frase, Brian Fuller , et al. (75 additional authors not shown)

    Abstract: This paper introduces v0.5 of the AI Safety Benchmark, which has been created by the MLCommons AI Safety Working Group. The AI Safety Benchmark has been designed to assess the safety risks of AI systems that use chat-tuned language models. We introduce a principled approach to specifying and constructing the benchmark, which for v0.5 covers only a single use case (an adult chatting to a general-pu… ▽ More

    Submitted 13 May, 2024; v1 submitted 18 April, 2024; originally announced April 2024.

  7. arXiv:2404.08382  [pdf, other

    cs.CL cs.AI

    Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think

    Authors: Xinpeng Wang, Chengzhi Hu, Bolei Ma, Paul Röttger, Barbara Plank

    Abstract: Multiple choice questions (MCQs) are commonly used to evaluate the capabilities of large language models (LLMs). One common way to evaluate the model response is to rank the candidate answers based on the log probability of the first token prediction. An alternative way is to examine the text output. Prior work has shown that first token probabilities lack robustness to changes in MCQ phrasing, an… ▽ More

    Submitted 12 April, 2024; originally announced April 2024.

  8. arXiv:2404.05399  [pdf, other

    cs.CL cs.AI

    SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety

    Authors: Paul Röttger, Fabio Pernisi, Bertie Vidgen, Dirk Hovy

    Abstract: The last two years have seen a rapid growth in concerns around the safety of large language models (LLMs). Researchers and practitioners have met these concerns by introducing an abundance of new datasets for evaluating and improving LLM safety. However, much of this work has happened in parallel, and with very different goals in mind, ranging from the mitigation of near-term risks around bias and… ▽ More

    Submitted 8 April, 2024; originally announced April 2024.

  9. arXiv:2403.19559  [pdf, other

    cs.CL

    Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset

    Authors: Janis Goldzycher, Paul Röttger, Gerold Schneider

    Abstract: Hate speech detection models are only as good as the data they are trained on. Datasets sourced from social media suffer from systematic gaps and biases, leading to unreliable models with simplistic decision boundaries. Adversarial datasets, collected by exploiting model weaknesses, promise to fix this problem. However, adversarial data collection can be slow and costly, and individual annotators… ▽ More

    Submitted 28 March, 2024; originally announced March 2024.

    Comments: Accepted at NAACL 2024 (main conference)

  10. arXiv:2403.03814  [pdf, other

    cs.CL cs.AI

    Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ

    Authors: Carolin Holtermann, Paul Röttger, Timm Dill, Anne Lauscher

    Abstract: Large language models (LLMs) need to serve everyone, including a global majority of non-English speakers. However, most LLMs today, and open LLMs in particular, are often intended for use in just English (e.g. Llama2, Mistral) or a small handful of high-resource languages (e.g. Mixtral, Qwen). Recent research shows that, despite limits in their intended use, people prompt LLMs in many different la… ▽ More

    Submitted 18 July, 2024; v1 submitted 6 March, 2024; originally announced March 2024.

  11. arXiv:2402.16786  [pdf, other

    cs.CL cs.AI

    Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models

    Authors: Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Rose Kirk, Hinrich Schütze, Dirk Hovy

    Abstract: Much recent work seeks to evaluate values and opinions in large language models (LLMs) using multiple-choice surveys and questionnaires. Most of this work is motivated by concerns around real-world LLM applications. For example, politically-biased LLMs may subtly influence society when they are used by millions of people. Such real-world concerns, however, stand in stark contrast to the artificial… ▽ More

    Submitted 5 June, 2024; v1 submitted 26 February, 2024; originally announced February 2024.

    Comments: Accepted at ACL 2024 (Main Conference)

  12. arXiv:2402.14499  [pdf, other

    cs.CL

    "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models

    Authors: Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber-Genzel, Paul Röttger, Frauke Kreuter, Dirk Hovy, Barbara Plank

    Abstract: The open-ended nature of language generation makes the evaluation of autoregressive large language models (LLMs) challenging. One common evaluation approach uses multiple-choice questions (MCQ) to limit the response space. The model is then evaluated by ranking the candidate answers by the log probability of the first token prediction. However, first-tokens may not consistently reflect the final r… ▽ More

    Submitted 4 July, 2024; v1 submitted 22 February, 2024; originally announced February 2024.

    Comments: ACL 2024 Findings

  13. arXiv:2311.08370  [pdf, other

    cs.CL

    SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

    Authors: Bertie Vidgen, Nino Scherrer, Hannah Rose Kirk, Rebecca Qian, Anand Kannappan, Scott A. Hale, Paul Röttger

    Abstract: The past year has seen rapid acceleration in the development of large language models (LLMs). However, without proper steering and safeguards, LLMs will readily follow malicious instructions, provide unsafe advice, and generate toxic content. We introduce SimpleSafetyTests (SST) as a new test suite for rapidly and systematically identifying such critical safety risks. The test suite comprises 100… ▽ More

    Submitted 16 February, 2024; v1 submitted 14 November, 2023; originally announced November 2023.

  14. arXiv:2310.07629  [pdf, other

    cs.CL cs.CY

    The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values

    Authors: Hannah Rose Kirk, Andrew M. Bean, Bertie Vidgen, Paul Röttger, Scott A. Hale

    Abstract: Human feedback is increasingly used to steer the behaviours of Large Language Models (LLMs). However, it is unclear how to collect and incorporate feedback in a way that is efficient, effective and unbiased, especially for highly subjective human preferences and values. In this paper, we survey existing approaches for learning from human feedback, drawing on 95 papers primarily from the ACL and ar… ▽ More

    Submitted 11 October, 2023; originally announced October 2023.

    Comments: Accepted for the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP, Main)

  15. arXiv:2310.02457  [pdf, other

    cs.CL cs.CY

    The Empty Signifier Problem: Towards Clearer Paradigms for Operationalising "Alignment" in Large Language Models

    Authors: Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, Scott A. Hale

    Abstract: In this paper, we address the concept of "alignment" in large language models (LLMs) through the lens of post-structuralist socio-political theory, specifically examining its parallels to empty signifiers. To establish a shared vocabulary around how abstract concepts of alignment are operationalised in empirical datasets, we propose a framework that demarcates: 1) which dimensions of model behavio… ▽ More

    Submitted 15 November, 2023; v1 submitted 3 October, 2023; originally announced October 2023.

    Comments: Socially Responsible Language Modelling Research (SoLaR) @ NeurIPs 2023

  16. arXiv:2309.07875  [pdf, other

    cs.CL

    Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

    Authors: Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger, Dan Jurafsky, Tatsunori Hashimoto, James Zou

    Abstract: Training large language models to follow instructions makes them perform better on a wide range of tasks and generally become more helpful. However, a perfectly helpful model will follow even the most malicious instructions and readily generate harmful content. In this paper, we raise concerns over the safety of models that only emphasize helpfulness, not harmlessness, in their instruction-tuning.… ▽ More

    Submitted 19 March, 2024; v1 submitted 14 September, 2023; originally announced September 2023.

  17. arXiv:2308.01263  [pdf, other

    cs.CL cs.AI

    XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

    Authors: Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, Dirk Hovy

    Abstract: Without proper safeguards, large language models will readily follow malicious instructions and generate toxic content. This risk motivates safety efforts such as red-teaming and large-scale feedback learning, which aim to make models both helpful and harmless. However, there is a tension between these two objectives, since harmlessness requires models to refuse to comply with unsafe prompts, and… ▽ More

    Submitted 1 April, 2024; v1 submitted 2 August, 2023; originally announced August 2023.

    Comments: Accepted at NAACL 2024 (Main Conference)

  18. arXiv:2306.11559  [pdf, other

    cs.CL

    The Ecological Fallacy in Annotation: Modelling Human Label Variation goes beyond Sociodemographics

    Authors: Matthias Orlikowski, Paul Röttger, Philipp Cimiano, Dirk Hovy

    Abstract: Many NLP tasks exhibit human label variation, where different annotators give different labels to the same texts. This variation is known to depend, at least in part, on the sociodemographics of annotators. Recent research aims to model individual annotator behaviour rather than predicting aggregated labels, and we would expect that sociodemographic information is useful for these models. On the o… ▽ More

    Submitted 20 June, 2023; originally announced June 2023.

    Comments: ACL2023 Camera-Ready

  19. arXiv:2303.05453  [pdf, ps, other

    cs.CL cs.CY

    Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback

    Authors: Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, Scott A. Hale

    Abstract: Large language models (LLMs) are used to generate content for a wide range of tasks, and are set to reach a growing audience in coming years due to integration in product interfaces like ChatGPT or search engines like Bing. This intensifies the need to ensure that models are aligned with human preferences and do not produce unsafe, inaccurate or toxic outputs. While alignment techniques like reinf… ▽ More

    Submitted 9 March, 2023; originally announced March 2023.

    Comments: 19 pages, 1 table

  20. arXiv:2303.04222  [pdf, other

    cs.CL cs.CY

    SemEval-2023 Task 10: Explainable Detection of Online Sexism

    Authors: Hannah Rose Kirk, Wenjie Yin, Bertie Vidgen, Paul Röttger

    Abstract: Online sexism is a widespread and harmful phenomenon. Automated tools can assist the detection of sexism at scale. Binary detection, however, disregards the diversity of sexist content, and fails to provide clear explanations for why something is sexist. To address this issue, we introduce SemEval Task 10 on the Explainable Detection of Online Sexism (EDOS). We make three main contributions: i) a… ▽ More

    Submitted 8 May, 2023; v1 submitted 7 March, 2023; originally announced March 2023.

    Comments: SemEval-2023 Task 10 (ACL 2023)

  21. arXiv:2210.11359  [pdf, other

    cs.CL

    Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages

    Authors: Paul Röttger, Debora Nozza, Federico Bianchi, Dirk Hovy

    Abstract: Hate speech is a global phenomenon, but most hate speech datasets so far focus on English-language content. This hinders the development of more effective hate speech detection models in hundreds of languages spoken by billions across the world. More data is needed, but annotating hateful content is expensive, time-consuming and potentially harmful to annotators. To mitigate these issues, we explo… ▽ More

    Submitted 20 October, 2022; originally announced October 2022.

    Comments: Accepted at EMNLP 2022 (Main Conference)

  22. arXiv:2206.09917  [pdf, other

    cs.CL

    Multilingual HateCheck: Functional Tests for Multilingual Hate Speech Detection Models

    Authors: Paul Röttger, Haitham Seelawi, Debora Nozza, Zeerak Talat, Bertie Vidgen

    Abstract: Hate speech detection models are typically evaluated on held-out test sets. However, this risks painting an incomplete and potentially misleading picture of model performance because of increasingly well-documented systematic gaps and biases in hate speech datasets. To enable more targeted diagnostic insights, recent research has thus introduced functional tests for hate speech detection models. H… ▽ More

    Submitted 20 June, 2022; originally announced June 2022.

    Comments: Accepted at WOAH (NAACL 2022)

  23. arXiv:2112.07475  [pdf, other

    cs.CL

    Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks

    Authors: Paul Röttger, Bertie Vidgen, Dirk Hovy, Janet B. Pierrehumbert

    Abstract: Labelled data is the foundation of most natural language processing tasks. However, labelling data is difficult and there often are diverse valid beliefs about what the correct data labels should be. So far, dataset creators have acknowledged annotator subjectivity, but rarely actively managed it in the annotation process. This has led to partly-subjective datasets that fail to serve a clear downs… ▽ More

    Submitted 29 April, 2022; v1 submitted 14 December, 2021; originally announced December 2021.

    Comments: Accepted at NAACL 2022 (Main Conference)

  24. arXiv:2108.05921  [pdf, other

    cs.CL cs.CY

    Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-based Hate

    Authors: Hannah Rose Kirk, Bertram Vidgen, Paul Röttger, Tristan Thrush, Scott A. Hale

    Abstract: Detecting online hate is a complex task, and low-performing models have harmful consequences when used for sensitive applications such as content moderation. Emoji-based hate is an emerging challenge for automated detection. We present HatemojiCheck, a test suite of 3,930 short-form statements that allows us to evaluate performance on hateful language expressed with emoji. Using the test suite, we… ▽ More

    Submitted 6 May, 2022; v1 submitted 12 August, 2021; originally announced August 2021.

    Journal ref: 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL 2022)

  25. arXiv:2104.08116  [pdf, other

    cs.CL

    Temporal Adaptation of BERT and Performance on Downstream Document Classification: Insights from Social Media

    Authors: Paul Röttger, Janet B. Pierrehumbert

    Abstract: Language use differs between domains and even within a domain, language use changes over time. For pre-trained language models like BERT, domain adaptation through continued pre-training has been shown to improve performance on in-domain downstream tasks. In this article, we investigate whether temporal adaptation can bring additional benefits. For this purpose, we introduce a corpus of social med… ▽ More

    Submitted 8 September, 2021; v1 submitted 16 April, 2021; originally announced April 2021.

    Comments: Accepted at EMNLP 2021 (Findings)

  26. HateCheck: Functional Tests for Hate Speech Detection Models

    Authors: Paul Röttger, Bertram Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, Janet B. Pierrehumbert

    Abstract: Detecting online hate is a difficult task that even state-of-the-art models struggle with. Typically, hate speech detection models are evaluated by measuring their performance on held-out test data using metrics such as accuracy and F1 score. However, this approach makes it difficult to identify specific model weak points. It also risks overestimating generalisable model performance due to increas… ▽ More

    Submitted 27 May, 2021; v1 submitted 31 December, 2020; originally announced December 2020.

    Comments: Accepted at ACL 2021 (Main Conference)