This specialized course is designed for social scientists seeking to deepen their skills in web scraping and learning algorithm auditing using R. Participants will learn how to efficiently scrape complex and dynamic web data, including handling JavaScript-heavy sites, overcoming common scraping problems and working with APIs.
The course begins with a refresher on basic web scraping concepts, ensuring all participants are on the same page. Following this, we will cover advanced scraping methods, covering scraping of dynamically loaded JavaScript-heavy websites and techniques for bypassing common scraping obstacles (e.g., avoiding getting detected and blocked by websites; creating user profiles with specific characteristics, etc). A significant portion of the course is dedicated to algorithm auditing. This includes understanding the fundamentals of this technique that allows evaluating the functionality and impact of black-box algorithms used in social media, search engines, LLM-based chatbots and other platforms.
Participants will learn how to collect and analyze data to audit these algorithms, uncovering potential biases and analyzing related ethical implications. The course will cover practical examples and case studies using advanced web scraping techniques and working with APIs. Throughout the course, hands-on exercises and projects will allow participants to apply their learning in real-world scenarios. Ethical considerations, such as respecting privacy and adhering to legal frameworks, will be a recurrent theme. By the end of the course, participants will be equipped with advanced skills in web scraping and a solid foundation in algorithm auditing, enabling them to conduct sophisticated data-driven research in the social sciences.