Abstract
Scientific discourse, as seen on social media and in news articles, has been shown to compromise the accuracy of scientific findings. Complex scientific claims are uttered in the form of short, digestible, and often inaccurate snippets. This phenomenon has led to online scientific debates being uninformed and has conduced to controversy and polarization. Examples include social media discussions about global health pandemics or climate change.To address this challenge, this thesis focuses on the study of scientific online discourse processing, an emerging research field at the intersection of Natural Language Processing (NLP) and Information Retrieval (IR). It is developed in four parts:The first part motivates the necessity for robust definitions, ground-truth corpora and methods for scientific online discourse processing. We present a survey of existing literature where we highlight key remaining challenges and how they are addressed by this thesis.The second part lays the foundation for developing and evaluating robust methods for scientific online discourse processing. We provide the first hierarchical and domain-agnostic definition, annotation framework, and expert-annotated corpus that includes various forms of scientific online discourse, including claims and references.. We also provide the first task formalization and baseline models for the detection and differentiation between scientific claims, references, and research contexts.The third part focuses on scientific web claims, a subcategory of scientific online discourse. We present the first in-depth analysis of the linguistic characteristics of scientific web claims. We also run the first empirical evaluation of the performance of language models on scientific web claims in multiple fact-checking-related tasks.The fourth part focuses on scientific citations from the web, another subcategory of scientific online discourse. We provide the first task formalization and baseline models for (1) flagging social media posts that formulate scientific claims without citing corresponding references, and (2) retrieving the original scientific publications informally referred to by social media posts.Through a unified definition, multiple ground-truth corpora, empirical task formalizations, and baseline methods and models, this thesis aims to lay the foundation for the study of scientific online discourse processing as a distinct, well-defined field of NLP/IR research.