Sign in

Rohan

@rohandas.net
429 followers 407 following 32 posts

CS PhD at CU Boulder · NLP · Narratives and Discourse · Knowledge Discovery and Retrieval www.rohandas.net

PostsRepliesMedia
Reposted by Rohan
Boulder NLP @bouldernlp.bsky.social · 29/06/2026
Boulder NLP has 19 papers accepted to #ACL2026 across the main conference, findings and workshops! 🚀 We'll be presenting work on morphology, machine translation, narratives, low-resource languages, CSS, HCI, LLMs, fairness/bias, and more. See you in San Diego! 🌊 ☀️ 🌮
List of main conference papers at ACL 2026 from Boulder NLP.List of main conference and findings papers at ACL 2026 from Boulder NLP.List of findings and workshop papers at ACL 2026 from Boulder NLP.List of workshop papers at ACL 2026 from Boulder NLP.
0147
Rohan @rohandas.net · 11/05/2026
Paper: arxiv.org/abs/2604.10368 Code: github.com/blast-cu/str... This work was done with my brilliant colleagues at the BLAST group at CU Boulder. with @advaitdeshmukh.com, @alexxandria-l.bsky.social, Zohar Naaman, I-Ta Lee, and @mlpacheco.bsky.social 🧵10/10
arxiv.org
A Structured Clustering Approach for Inducing Media Narratives
Media narratives wield tremendous power in shaping public opinion, yet computational approaches struggle to capture the nuanced storytelling structures that communication theory emphasizes as central ...
060
Rohan @rohandas.net · 11/05/2026
At the domain level, immigration is driven by character portrayals (who is being portrayed and how) while gun control is driven by judicial conflict and institutional enforcement schemas (what is happening), suggesting the two domains are organized around different narrative dimensions. 🧵9/10
We find that for immigration, characters play a central role in driving frame prediction, particularly immigrants. In contrast, gun control frame predictions are driven by narrative clusters, particularly those dealing with judicial and legal schemas.
120
Rohan @rohandas.net · 11/05/2026
At the frame level, we found narrative signatures vary meaningfully across policy frames. For example, the Legality/Constitutionality frame in gun control is dominated by judicial conflict schemas, suggesting legal framing emphasizes courtroom battles while political actors remain peripheral. 🧵8/10
Legal framing of gun control coverage tends to emphasize courtroom battles rather than legislative politics.
110
Rohan @rohandas.net · 11/05/2026
The SHAP analysis operates at three levels: instance, frame, and domain. At the instance level, narrative schema features capture equivalent information to RoBERTa embeddings, while offering more succinct and interpretable representations. 🧵7/10
SHAP feature importance comparison between RoBERTa embeddings and narrative features for frame prediction on a single test instance, showing both approaches capture similar semantic patterns.
110
Rohan @rohandas.net · 11/05/2026
Main findings: 1. Structured clustering outperforms k-means on frame prediction and cluster purity, indicating our schemas carry high predictive signal. 2. Narrative features based on induced schemas match black box models in predictive performance, while offering greater interpretability. 🧵6/10
110
Rohan @rohandas.net · 11/05/2026
We evaluated 4,000 news articles across immigration and gun control. Schema quality was validated by a trained linguist, with 94-96% rated high quality, consistently capturing Entman's framing elements. We also conducted a SHAP analysis using induced schemas as features for frame prediction. 🧵5/10
Examples of high-quality generated schemas, evaluated on their overall coherence, as well as on striking a balance between coverage of and specificity to the sentences included in the context.
120
Rohan @rohandas.net · 11/05/2026
Textual similarity alone can group narratives that look alike but represent different framings of the same events. Our structured clustering method uses character-role configurations as cannot-link constraints to ensure narratives with conflicting framings are appropriately separated. 🧵4/10
120
Rohan @rohandas.net · 11/05/2026
We present a framework for unsupervised, domain-agnostic narrative schema induction that scales to any large corpora. Our method extracts causal event chains, assigns character roles (Hero/Threat/Victim) to entities, and uses these as constraints in a structured clustering framework. 🧵3/10
For a large scale news corpus, we first construct narrative event chains and obtain character and role annotations for them. We then cluster these narrative chains using the character and role information as constraints. The generated narrative clusters are representative of fine-grained and nuanced narrative schemas.
110
Rohan @rohandas.net · 11/05/2026
Existing NLP approaches often build narratives bottom-up from extractable atomic units like predicate argument structures or entities. While highly scalable, these methods seldom capture the evaluative and ideological dimensions central to how meaning is constructed in the media. 🧵2/10
Law enforcement spending framed as worker protection versus government waste across policy domains.
120
Rohan @rohandas.net · 11/05/2026
Computational approaches to media narrative analysis either miss nuanced storytelling patterns through coarse-grained analysis, or require domain-specific taxonomies that limit scalability. We show joint event and character modeling can address this gap. Details in our #ACL2026 (Main) paper. 🧵1/10
Paper Title: A Structured Clustering Approach for Inducing Media Narratives

Authors: Rohan Das, Advait Deshmukh, Alexandria Leto, Zohar Naaman, I-Ta Lee, Maria Leonor Pacheco
12411
Rohan @rohandas.net · 10/05/2026
Paper: arxiv.org/abs/2408.09030 Code: github.com/blast-cu/int... This work was done with my brilliant colleagues at the BLAST group at CU Boulder. Led by Alvin Chen, with Dananjay Srinivas and @mlpacheco.bsky.social 🧵7/7
arxiv.org
Effects of Collaboration on the Performance of Interactive Theme Discovery Systems
NLP-assisted solutions to support qualitative data analysis have gained considerable traction. However, no unified evaluation framework exists which can account for the many different settings in whic...
000
Rohan @rohandas.net · 10/05/2026
Main Takeaways: 1. Synchronous collaboration produces more consistent and cohesive themes, especially on semantically homogeneous data. 2. Systems with richer user control benefit more from real-time deliberation. 3. Dataset characteristics shape outcomes as much as collaboration setting. 🧵6/7
110
Rohan @rohandas.net · 10/05/2026
We evaluated 3 NLP-assisted coding systems (topic model, relational, LLM-based) under synchronous vs. asynchronous collaboration, on 2 corpora: 85K COVID vaccine tweets and 5.5K climate change ads. 🧵5/7
111
Rohan @rohandas.net · 10/05/2026
Our evaluation framework operates across 3 dimensions: 1. Consistency: Do coders across modalities surface the same themes? 2. Cohesiveness/Distinctiveness: Are docs within a theme more similar to each other than to docs in other themes? 3. Correctness: Are automated assignments accurate? 🧵4/7
100
Rohan @rohandas.net · 10/05/2026
Our contributions: 1. A 3-pronged evaluation framework for interactive qualitative coding systems. 2. Experimental results comparing synchronous and asynchronous coding. 🧵3/7
In this study, we measure the quality of coded themes using different interactive systems under different coding configurations.
110
Rohan @rohandas.net · 10/05/2026
We tasked researchers to use NLP tools to identify themes from a dataset of 5.5k Facebook ads on climate change. Here are two of the resulting code-books. Do you think one code-book is better than the other? 🧵2/7
Themes as they belong to two different qualitative codebooks derived from the same dataset.
100
Rohan @rohandas.net · 10/05/2026
What’s the best way to analyze online discourse on any given topic? Is there a right way to use NLP tools to sift through massive datasets? To find out, we tested several tools across different collaboration settings and report findings in an #ACL2026 (Main) paper: arxiv.org/abs/2408.09030 🧵1/7
Paper Title: Effects of Collaboration on the Performance of Interactive Theme Discovery Systems

Authors: Alvin Po-Chun Chen, Rohan Das, Dananjay Srinivas, Alexandra Barry, Maksim Seniw, Maria Leonor Pacheco
1125
Rohan @rohandas.net · 05/05/2026
Paper: arxiv.org/abs/2408.09030 Code and Data: github.com/blast-cu/int... This work was done with my brilliant colleagues at the BLAST group at CU Boulder. Led by Alvin Chen, with Dananjay Srinivas and @mlpacheco.bsky.social 🧵5/5
arxiv.org
Effects of Collaboration on the Performance of Interactive Theme Discovery Systems
NLP-assisted solutions to support qualitative data analysis have gained considerable traction. However, no unified evaluation framework exists which can account for the many different settings in whic...
000
Rohan @rohandas.net · 05/05/2026
Main Takeaways: 1. Synchronous collaboration produces more consistent and cohesive themes, especially on semantically homogeneous data 2. Systems with richer user control benefit more from real-time deliberation 3. Dataset characteristics shape outcomes as much as collaboration setting does 🧵4/5
100
Rohan @rohandas.net · 05/05/2026
We introduce an evaluation framework across 4 dimensions: 1. Consistency: Do coders across modalities surface the same themes? 2. Cohesiveness: Are documents within a theme similar? 3. Distinctiveness: Are themes different from each other? 4. Correctness: Are automated assignments accurate? 🧵3/5
100
Rohan @rohandas.net · 05/05/2026
We evaluated 3 NLP-assisted coding systems (topic model, relational, LLM-based) under synchronous vs. asynchronous collaboration, on 2 corpora: 85K COVID vaccine tweets and 5.5K climate change ads. The study involved 33 researchers across 2 universities, spanning 30 coding experiments. 🧵2/5
In this study, we measure the quality of coded themes using different interactive systems under different coding configurations.
110
Reposted by Rohan
Boulder NLP @bouldernlp.bsky.social · 03/11/2025
Strong showing from Boulder NLP at #EMNLP2025! 🚀 Come find us at these sessions to discuss morphology, semantics, LLMs, cultural analytics, HCI, and more. See you in Suzhou! 🇨🇳
0104
Rohan @rohandas.net · 01/11/2025
Also check out Grameen Bank which is a microfinance org that caters to the poor and is owned by the borrowers. Wiki: en.wikipedia.org/wiki/Grameen...
010
Rohan @rohandas.net · 01/11/2025
In the Indian context, Amul stands out as a shining example of the co-operative model that revolutionized the dairy industry. Wiki: en.wikipedia.org/wiki/Amul
120
Rohan @rohandas.net · 16/08/2025
Consider registering the bike on the CU Bike Index (just in case) - www.colorado.edu/police/crime...
colorado.edu
Bike Registration
Bike Theft: Realities & RisksLike any university campus, CU Boulder students are not immune to the crime of bike theft. Neither are the residents of Boulder, a
010
Reposted by Rohan
Sundance Institute @sundance.org · 27/03/2025
Big news! Sundance Film Festival reveals Boulder, Colorado as our new location starting in 2027 and beyond. The annual event by the nonprofit Sundance Institute will continue to entertain and inspire audiences through independent film. 🎬 Read more: sndnc.org/boulder2027 #SundanceFilmFestival
129138
Reposted by Rohan
Lauren Klein @laurenfklein.bsky.social · 03/03/2025
Already a time capsule, but back in January a bunch of us working at the intersection of the humanities and AI/ML came together to sketch out eight provocations from the humanities for genAI research. Here's a 🧵 1/ arxiv.org/abs/2502.19190
arxiv.org
Provocations from the Humanities for Generative AI Research
This paper presents a set of provocations for considering the uses, impact, and harms of generative AI from the perspective of humanities researchers. We provide a working definition of humanities res...
511444
Reposted by Rohan
chenhaotan.bsky.social @chenhaotan.bsky.social · 24/01/2025
Spent a great day at Boulder meeting new students and old colleagues. I used to take this view every day. Here are the slides for my talk titled "Alignment Beyond Human Preferences: Use Human Goals to Guide AI towards Complementary AI": chenhaot.com/talks/alignm...
0165
Rohan @rohandas.net · 01/01/2025
IDE: I can't recommend PyCharm enough for Python heavy workflows. The good folks at JetBrains truly deserve some flowers, and I think everything they build is chef's kiss! I love PyCharm's debugger and python shell implementations (I dislike notebooks with a vengeance.)
010
Rohan @rohandas.net · 01/01/2025
Terminal Emulator: A switch to Termius last year massively sped up my somewhat complex workflow that interfaces with CU's HPC cluster. It has mostly everything that Warp offers (I think?) + smart SSH and FTP client, port forwarding, and automation routines.
120
Rohan @rohandas.net · 01/01/2025
dodgy solution for CU's implementation of O365 Mail/Exchange Online.
110
Rohan @rohandas.net · 01/01/2025
Email: Having tried tons of different email clients over the years, I like Spark Mail the best (which is what I use on my phone/tablet, and probably one of the few services I would pay a subscription fee for. But they don't support Linux, and so on my laptop I am stuck with Thunderbird with a
110
Rohan @rohandas.net · 01/01/2025
Browser: I switched to Brave last year after having used Firefox forever. It is chromium-based, (supposedly) privacy-centric, and tabs live indefinitely by default. However, I think Vivaldi is the most like for like replacement for Arc - never tried it since it is an overkill for my needs.
110
Rohan @rohandas.net · 19/11/2024
Bonus: You get to work out of a lab with not only windows but also gorgeous views of the Rockies! ⛰️
010
Rohan @rohandas.net · 19/11/2024
If you're applying to NLP PhD programs, consider Maria’s new group at CU Boulder - she’s amazing, and the larger Boulder NLP community is a great place to learn, grow, and collaborate! Feel free to reach out if you want to chat about the CS PhD program or life in Boulder! My DMs are open.
220
Reposted by Rohan
Joe Stacey @joestacey.bsky.social · 18/11/2024
After going to NAACL, ACL and #EMNLP2024 this year, here are a few tips I’ve picked up about attending #NLP conferences. Would love to hear any other tips if you have them! This proved very popular on another (more evil) social media platform, so sharing here also 🙂 My 10 tips:
148316
Rohan @rohandas.net · 14/11/2024
Happy to report that the event was a disaster. Took us an hour to even get into the venue. Long winding queues everywhere and barely any room to stand. 🤡
110
Reposted by Rohan
Boulder NLP @bouldernlp.bsky.social · 11/11/2024
📢 Check out the lineup of papers our students will be showcasing at #EMNLP2024 in Miami next week! 🌴 We'll be presenting new work on morphology, Q&A, and narratives.🔍
067