drwebdomain.blog
LIFE Summer Research Grant Reflections: How Language Shapes Conflict: A Computational Approach to Media Framing of Iran
This is the first blog post in a series written by the 2025 recipients of the Duke University Libraries Summer Research Fellowship for LIFE Students. Rose Simons is a junior majoring in Computer Science. Why are we less surprised about conflicts between certain countries than others? Is it accepted that there are some groups that … Continue reading LIFE Summer Research Grant Reflections: How Language Shapes Conflict: A Computational Approach to Media Framing of Iran → The post LIFE Summer Research Grant Reflections: How Language Shapes Conflict: A Computational Approach to Media Framing of Iran appeared first on Duke University Libraries Blogs. Read More This is the first blog post in a series written by the 2025 recipients of the Duke University Libraries Summer Research Fellowship for LIFE Students. Rose Simons is a junior majoring in Computer Science. Why are we less surprised about conflicts between certain countries than others? Is it accepted that there are some groups that are more contentious, or is there something smaller having larger effects than we imagine? As someone who has always been interested in politics and the interdisciplinary applications of computer science tools, I came up with the following question to help identify how the way we talk about Iran in media frames it in an adversarial light: How has American media constructed Iran as a geopolitical adversary from 1979 to the present, and how have these representations shifted across major political events? Throughout the summer I was able to explore several theoretical frameworks. This included Edward Said’s Orientalism, which posited the West’s history of constructing the ‘other’ through language and literary means. I also explored framing theory: studying what is emphasized versus omitted, and how repetition normalizes things. These frameworks were very important to study as I sought to design this research question, and they will be even more vital when I try to analyze and synthesize the data that we receive from the methods. Newspapers are an important tool for looking at a more longitudinal sample of media studies, as there are many consistent news outlets founded before the 1970s which produce media on the conflicts of interest. Public perception is also shaped at a large scale by newspapers, due to the respect these outlets have gained and their longevity. For these reasons, newspapers served as our source for data collection. One of the major questions was how one quantifies this study of media, and what tools are available in the computer science field that offer us the capability to look at such a complex issue. The answer to this is word embeddings, a technique in which words from a corpus are vectorized—turned into mathematical representations—and placed in a 3D space trained on specific data, allowing us to calculate distances between terms and ideas. The accepted premise is that terms closer in distance are more closely related than those farther apart, with these calculations based on cosine similarity. There are many different word embedding models to choose from. After performing a literature review, I found that several articles point to the benefits of Word2Vec models for analyzing semantic shifts over long periods of time (Zhang et al.; Hamilton et al., 2016), leading us to choose Word2Vec as the model best suited to answer our research question. The exploration of this project over this summer meant that I worked with my mentors to constantly brainstorm different methods to best answer the research question. We ended on the idea of spending most of the summer collecting data from respected sources such as the New York Times, the Washington Post, and the Wall Street Journal, utilizing databases from the Duke Research Libraries, such as ProQuest and Nexis Uni. These tools and subscriptions allowed us to download mass amounts of news articles to build a large enough dataset for Word2Vec. My mentors also helped me design a timeline focusing on key events of conflict throughout U.S. and Iranian history, starting just before the Hostage Crisis of 1979 and ending in 2011—a period bookended by two of the most significant moments of tension between the two countries in recent history. The next step was to write scripts to clean the data, train models on the text, and then code a script to identify terms describing “Iran” or “Iranian,” classify those terms as adversarial or otherwise, and measure how closely related they are to “Iran” using cosine similarity. Computation was also an important component of the research question, allowing us to scale the project over a long period of time (32 years of coverage across three outlets)—an undertaking that would be nearly impossible without computation and technology. Moreover, computational methodologies can reveal semantic drifts that human readers cannot detect consistently from article to article. Unfortunately, the cleaning of data took longer than anticipated, and we were not able to complete the analysis portion of this project. However, I plan to continue this work next semester. In the end, I hope that I can uncover more about how language affects conflict on a larger scale, how it shapes perception, and how we can push against dangerous or preconceived notions, and increase media literacy and understanding of subtleties and bias. This summer has taught me a lot about the narrowing down of research questions, how to find the right methods for the research question, as well as how many tools and resources are available to me at Duke for me to freely take advantage of. I hope to continue this project in the future, and possibly present findings at a conference. I would like to offer a special thank you to DukeLife and Duke Libraries for this amazing opportunity for the summer, as well as my mentor Rukimani Prathivadhibhayankaram for extensively helping to frame my research question, offer continued support throughout the program duration, as well as help with some of the more tedious and technical aspects of research. Thank you to Hannah Jacobs, my librarian mentor who worked with me through extremely differing time zones and was an amazing resource and fount of knowledge specializing in Digital Humanities. Thank you as well to Heidi Madden, Librarian for Western European & Medieval Renaissance Studies and Head of the Global Studies Department, and Dr. Mira Xenia Schwerda, Assistant Professor of History of Photography and New Media, for consulting with me on this project and lending me their advice and expertise in their respective fields. Works Cited: Hamilton, William L., Jure Leskovec, and Dan Jurafsky. “Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change.” Zhang, R., Nie, L., Zhao, C., and Chen, Q. “Achieving Semantic Consistency: Contextualized Word Representations for Political Text Analysis.” The post LIFE Summer Research Grant Reflections: How Language Shapes Conflict: A Computational Approach to Media Framing of Iran appeared first on Duke University Libraries Blogs.