Prospective PhD Student Q&A – October 2, 2026

cdsc group photo from the 2026 retreat at UW

Thinking about applying to graduate school? Wonder what it’s like to pursue a PhD or research-based master’s degree? Interested in understanding relationships between technology and society? Curious about how to do research on online communities like Reddit, Wikipedia, or GNU/Linux? The Community Data Science Collective is hosting a virtual Q&A session on Friday, October 2 at 12 pm PT / 2pm CT / 3pm ET for prospective students. This session is scheduled for an hour and will be divided between a larger group session with faculty and smaller groups with current graduate students. If you would like to attend, register at this link!

This post provides a very brief rundown on the CDSC, the different universities and PhD and research master’s programs our faculty members are affiliated with, and some general ideas about what we’re looking for when we review applications.

What is the Community Data Science Collective?

cdsc group photo from the 2026 retreat at UW

The Community Data Science Collective (or CDSC) is a joint research group of (mostly quantitative) empirical social scientists and designers pursuing research about the organization of online communities, peer production, and learning and collaboration in social computing systems. We are based at Northwestern University, the University of Washington-Seattle, University of Washington-Bothell, The University of Texas at Austin, Purdue University, and a few other places. You can read more about us and our work on our research group blog and on the collective’s website/wiki.

What are these different Ph.D. programs? Why would I choose one over the other?

This year the group includes multiple faculty principal investigators (PIs) who are actively recruiting graduate students: Kaylea Champion (University of Washington in Bothell), Nathan TeBlunthuis (University of Texas at Austin), Jeremy Foote (Purdue University), Benjamin Mako Hill (University of Washington in Seattle), Aaron Shaw (Northwestern University). Each of these PIs advises PhD and/or M.S. students in graduate programs at their respective universities. We also have one faculty PI who is not currently recruiting students, but is an active member of the group: Ryan Funkhouser (University of Idaho). Our programs are each described below.

Although we often work together on research and serve as co-advisors on students’ projects, each faculty member has specific areas of expertise and interests. The reasons you might choose to apply to one of these PhD or M.S. programs or to work with a specific faculty member could include factors like your previous training, career goals, and the alignment of your specific research interests with our respective skills.

At the same time, a great thing about the CDSC is that we all collaborate and regularly co-advise students across our campuses, so applying to or attending one program does not prevent you from accessing the expertise of our whole group. But please keep in mind that our different PhD and M.S. programs have different application deadlines, requirements, and procedures!

Faculty who are actively recruiting PhD Students this year

If you are interested in applying to any of the programs, we strongly encourage you to reach out to the specific faculty in that program before submitting an application.

Nathan TeBlunthuis, Ventrait Pictures

Nathan TeBlunthuis is an Assistant Professor in the School of Information at the University of Texas at Austin in the area of social informatics. Nathan’s research focuses on analyzing ecosystems of online communities, AI tools in peer production, and methods in computational social science. His current projects continue in these areas and also draw from them all to understand how information sources achieve legitimacy in online communities. He works primarily using computational tools and big data, but also grounds his work in qualitative evidence.

Jeremy Foote

Jeremy Foote is an Assistant Professor at the Brian Lamb School of Communication at Purdue University. He is affiliated with the Organizational Communication and Media, Technology, and Society programs. Jeremy’s research focuses on how individuals decide when and in what ways to participate in online communities, how communities change the people who participate in them, and how both of those processes can help us to understand what people believe and which things become popular and influential. He and his students use multiple methods, including data science, agent-based modeling, field experiments, and interviews.

Benjamin Mako Hill, Photo by Pedro Pacheco

Benjamin Mako Hill is an Associate Professor of Communication at the University of Washington. He is also adjunct faculty at UW’s Department of Human-Centered Design and Engineering (HCDE), Computer Science and Engineering (CSE), and the Information School. Although many of Mako’s students are in the Department of Communication, he has also advised students in all three other departments—although he typically has more limited ability to admit students into those programs on his own and usually does so with a co-advisor in those departments. Mako’s research focuses on population-level studies of peer production projects, computational social science, efforts to democratize data science, and informal learning. Mako has also put together a webpage for prospective graduate students with some useful links and information.

Aaron Shaw, Nikki Ritcher Photography

Aaron Shaw is a Professor in the Department of Communication Studies at Northwestern. In terms of PhD programs, Aaron’s primary affiliations are with the Media, Technology and Society (MTS) and the Technology and Social Behavior (TSB) Ph.D. programs (please note: the TSB program is a joint degree between Communication and Computer Science). Aaron also has a courtesy appointment in the Sociology Department at Northwestern, but he has not directly supervised any Ph.D. advisees in that department (yet). Aaron studies organizing, participation, and governance in online communities, especially ones that create public information resources and digital infrastructure. Aaron’s current projects focus on comparative analysis of the organization of peer production communities, online community governance, LLM-based social science, and AI’s impacts on digital public infrastructure.

Faculty who Supervise ADMITTED M.S. and B.S. Students
Kaylea Champion

Kaylea Champion is an Assistant Professor in Computing & Software Systems at the University of Washington-Bothell. Kaylea’s research investigates how people collaborate to build digital infrastructure, including operating systems, programming languages, and information repositories. What gets made and maintained and secured—and what gets neglected? What risks do we face (including from AI and cybercriminals)? What practices lead to better outcomes? How can we work smarter and what can we stop doing? Kaylea’s work seeks to bridge the divide between research and practice, which for her means building relationships with practitioner communities, organizations, and industry to directly share research findings. If you are interested in the Computing & Software Systems graduate programs at the University of Washington – Bothell, you should register for one of the information sessions specific to these programs as well. Kaylea’s department only admits M.S. and B.S. students (MS-Cybersecurity and MS-Computer Science and Software Engineering). In addition, she is happy to support any students who are working with others in the CDSC and can serve on thesis and dissertation committees.

Other Faculty members of the CDSC
Ryan Funkhouser

Ryan Funkhouser is an Assistant Professor in the Department of Psychology and Communication at the University of Idaho. Ryan’s research focuses on communication processes for bridging ideological divides in online spaces. His work includes explorations of deliberation-focused online communities, the role of narrative in persuasion, and the mechanisms of belief change. Ryan utilizes both computational and qualitative methods to explore text data, primarily from online sources.

What do you look for in PhD applicants?

There’s no easy or singular answer to this. In general, we look for curious, intelligent people driven to develop original research projects that advance scientific and practical understanding of topics that intersect with any of our collective research interests.

To get an idea of the interests and experiences present in the group, read our respective bios and CVs (follow the links above to our personal websites). Specific skills we regularly use include consuming and producing social science and/or social computing (human-computer interaction) research; applied statistics and statistical computing; various empirical research methods; social theory and cultural studies; and more.

Formal qualifications that speak to similar skills and show up in your resume, transcripts, or work history are great, but we are much more interested in your capacity to learn, think, write, analyze, and/or code effectively than in your credentials, test scores, grades, or previous affiliations. It’s graduate school, after all, and we do not expect you to show up knowing how to do all the things already.

Intellectual creativity, persistence, and a willingness to acquire new skills and problem-solving matter a lot. We think doctoral education is less about executing tasks someone else hands you and more about learning to identify a new, important problem; develop an appropriate approach to solving it; and explain all of the above and why it matters so others can learn from you in the future. Evidence that you can or at least want to do these things is critical. Indications that you can also play well with others and would make a generous, friendly colleague are really important too.

All of this is to say, we do not have any one trait or skill set we look for in prospective students. We strive to be inclusive along every possible dimension. Each person who has joined our group has contributed unique skills and experiences, as well as personal interests. We want our future students and colleagues to do the same.

Now what?

Still not sure whether or how your interests might fit with the group? Still have questions? Still reading and just don’t want to stop? Follow the links above for more information. Feel free to email at least one of us. We are happy to answer your questions and always eager to chat.

Community Data Science Collective logo

Community Dialogue: Social Drivers of Online Hate

Toxicity, aggression, racism, misogyny, and other forms of hate are unfortunately common in online spaces. Researchers have often treated those who engage in online hate as social deviants, enabled to act out by the affordances of online anonymity. This Friday, we will host two speakers who will discuss an emerging perspective that sees online hate as a social process, where those who engage in hate do so for social approval, often as part of (anti)social organizations.

Dyuti Jha (Fort Hays State University ) will discuss how users engage in online hate as a form of hybrid social action. Although online hate has been conceptualized as a malicious form of individual expression or dyadic interactions, this research explores the way in which users develop collective identities from their membership in online communities of hate, connect with each other based on shared grievances, and engage in expressions of online hate as a social action motivated by moral outrage and the need for social approval. Using AI-facilitated qualitative conversational data collection, her work shows that online hate campaigns are organized as both collective and connective action, with hidden organizing supporting the “grassroots” connective spread of hate campaigns. This research offers useful perspectives on the motivations and organizing strategies of concealed actors behind the execution and sustenance of online hate campaigns.

Joe Walther (UC Santa Barbara) will discuss how conventional research and activism about online hate in social media presumes that its purpose is to antagonize victims. A new approach argues that hate is primarily for haters, and that the incentive for online hate posting is the social approval its providers glean from admirers. This presentation outlines social approval theory, growing empirical evidence supporting and challenging its propositions, and implications for deterrence.

Speaker Bios:

  • Joe Walther holds the Bertelsen Presidential Chair in Technology and Society at the University of California, Santa Barbara, where he is a Distinguished Professor of Communication. He is also a faculty associate at the Berkman Klein Center for Internet & Society at Harvard University. His theoretical and empirical work focuses on interpersonal and intergroup communication in mediated interaction, from relationship development to the spread of online hate. Social Processes of Online Hate (co-edited with R. Rice, Routledge 2025) is available open access.
  • Dyuti Jha is an Assistant Professor of Organizational Communication at Fort Hays State University. Using experiments and AI-facilitated conversations, her doctoral dissertation studied the impact and social organization of online hate. Her other research interests include individual and organizational resilience and the effects of online toxicity on political expression on the internet.

The CDSC is organizing this event and hosting and supporting it, in part, with a National Science Foundation grant (IIS-2045055), so it will be free to attend. We will share a code of conduct with participants before the event. Discussions will be held under Chatham House Rule. Presentations will be recorded, though discussions will not.

Register to Attend

The event will be held on Friday, September 18 from 3pm – 5pm Eastern Time. To get to know the audience a bit better, please fill out this brief survey to register. Those who register will receive a meeting invite with a link to the Zoom room.

What is a Dialogue?

The Science of Community Dialogue Series is a series of conversations among researchers, experts, community organizers, and others interested in how communities work, collaborate, and succeed. To learn more, watch this short introduction video with Aaron Shaw.

Seeking to Interview People Who Make Software!

Students in the CDSC are recruiting for two interview studies.

The AI Slop study seeks to interview people involved in open source software who have experience dealing with AI-authored contributions. We hope to learn more about how communities are responding to this challenge, with a particular focus on ‘AI Slop’. You can read more about the study and fill out our screening survey here!

The Software Handoffs study seeks to interview people about how responsibility for a given piece of code gets handed off from one person to the next. We are wondering how this task is handled in industry and how this might differ from what is taught in a computer science curriculum. You can read more about the study and fill out our screening survey here!

FOSSY 2026: Call for Proposals!

The Free and Open Source Software Yearly conference (FOSSY) is back for the fourth year in a row, and we’ll be running the Science of Community track, inspired by the CDSC Science of Community Dialogues, bringing together practitioners and researchers to talk about scholarly work that’s relevant to the efforts of practitioners.

Does your work touch open source, communities, technology, or cooperation? Do you want to help bridge the gaps between research and practice? Join our track! The call for proposals is open and we’re looking for presenters interested in speaking to FOSS practitioners, developers, community organizers, contributors, and people into and curious about FOSS.

As researchers, we benefit so much from the communities we work with and study and we want them to also learn from the research they so generously take part in. While the Dialogues cover a broad range of topics and communities, FOSSY presentations will focus on how that work relates to free and open source software communities, projects, and practitioners.

FOSSY is a low-stress opportunity to talk to people who your work can benefit. For topics, consider presenting implications from past papers, synthesizing work from your field overall, or floating ideas and problems (lightning talks! long talks! short talks!). A full track description and answers to common questions is available on our wiki.

The CFP deadline is May 22nd and uses this form. We can’t wait for you to join us!

Performing the News: Rhetoric, Trust, and the Fight Against Video Disinformation in India

Apart from being generally chronically online on political X/Bluesky/Tiktok/Instagram, I also study science disinformation for a living, which means I spend a lot of time reading things that aren’t true. More specifically, I spend time reading the corrections, the careful, methodical, often thankless articles that community media organizations publish after a manipulated video has already been seen by millions. I am currently working on a paper examining exactly this, structurally: how the Indian news organization AltNews debunks video disinformation in India, and what the rhetorical structure of that work reveals about the challenge of fighting falsehoods at scale.

India is, according to the World Economic Forum, the country most susceptible to large-scale disinformation in the world, with 750 million internet users, 22 constitutionally recognized languages, a WhatsApp-dominated information ecosystem that is largely invisible to automated detection systems, and a political environment in which viral video disinformation is regularly amplified by mainstream media, fringe actors, and national political parties. In this context, the question of how fact-checkers write, not just what they correct, but how they construct their corrections and the knowledge ecosystems they engage with, warrants deeper examination. 

The cognitive trap that makes video disinformation so effective

“Seeing is believing” is a popular folk saying. It describes a genuine and well-documented cognitive bias in that visual evidence carries a persuasive weight that text cannot replicate. My paper draws on Fazio et al.’s (2015) finding that knowledge does not reliably protect against illusory truth; that even informed, attentive readers can be swayed by false information that looks credible. Video disinformation exploits this bias with particular efficiency, because the clip itself functions as apparent evidence. You don’t need a caption to believe what you watched.

But there is a second, less obvious problem. As institutional trust erodes and the awareness of manipulation grows, audiences become susceptible not only to false videos, but to disbelieving real ones. The disinformation ecosystem, at its most corrosive, inserts falsehoods and destabilizes the category of visual truth.

The anatomy of a debunk

My analysis of 150 AltNews video debunking articles reveals a consistent rhetorical structure that departs from the conventions of mainstream journalism. Where standard news follows an “inverted pyramid,” placing most important information first and subsequent information in decreasing order of priority, debunking articles are circular. The headline announces a verdict; the final line confirms it, but this time as a logical conclusion earned through evidence. The piece begins and ends with the same claim, but the reader arrives at the ending differently than they arrived at the beginning.

“Debunking arguments do not show their target beliefs to be false but rather undermine the justification a subject may have for holding them.” — Hanno Sauer (2018)

AltNews is not simply telling readers that a video is fake. It is systematically dismantling the reasons a reader “might have believed it,” such as the political authority of the person who shared it, the apparent plausibility of its imagery, or the emotional register in which it circulated. Particularly relevant to COVID-19 scientific misinformation, each article acknowledges in some way the overwhelming scale of the disinformation [like by alluding to its “virality”], and also the fear and confusion in the social circumstances surrounding the disinformation. The correction is not a verdict delivered from above; it is an argument the reader is invited to construct alongside the fact-checker, building collective capacity to discern.

The lead paragraph of each article enacts this invitation with notable rhetorical precision. It uses passive voice and hedged language: “A video in which a woman is seen lying in the bushes is doing the rounds on social media.” The word “seen” does careful work:  it acknowledges what the reader has probably watched, validates their experience, but withholds editorial endorsement of the content. According to rhetoric scholar Kenneth Burke, this is an act of identification: establishing consubstantiality with the audience before introducing dissonance. This identification, a recognition of the reader’s circumstance, allows the reader to be persuaded in the direction of the truth, which is stated once at first in the headline, and repeated at the end. 

The pedagogical burden of video verification

What distinguishes video fact-checking from other forms of debunking is the technical weight it carries. To verify a manipulated clip, AltNews staff extract keyframes, run reverse image searches, conduct metadata forensics, and deploy AI-detection tools. Each of these methods must then be explained—plainly, with screenshots, with links to original sources—to a general readership. The articles are, in this sense, simultaneously corrections and tutorials. By modeling the process of finding out, AltNews both corrects the misinformation, and provides technical clarity in how information is produced and distributed, attempting to build readers’ skill in verifying images and videos out there. 

This “show your work” norm is a deliberate strategy for building what Dourish and Bellotti (1992) call awareness, through establishing the sense that one understands not just an outcome but the process that produced it. For debunking, transparency about method is the mechanism by which readers are gradually equipped to verify things themselves.

Why human-centered collaboration is the only viable response

AltNews’s staff of journalists, scientists, engineers, OSINT specialists, social activists and people working on intersections of those roles, function as a distributed, multidisciplinary verification network. India’s disinformation ecosystem is, as Starbird, Arif, and Wilson (2019) demonstrate, fundamentally collaborative: coordinated networks of accounts, platforms, and political organizations working in concert to amplify false narratives. Automated platform moderation, such as the content bots of Facebook and X, has proven structurally inadequate to this complexity, particularly given India’s linguistic diversity and the closed architecture of WhatsApp groups.

What is needed, and what AltNews partially models, is a collaboratively-informed approach: human-centered, context-specific, and built for heterogeneity rather than scale. The fact that this work runs on donations, in a country where disinformation reaches hundreds of millions, tells us something important about where the gaps in our collective response still lie.

CDSC at CHI 2026!

Come hang out with us at CHI 2026 in Barcelona April 13-17! Members of the Community Date Science Collective will be presenting work and we’d love to see you there.

Barcelona. View from the Sacred Heart Church at Tibidabo. By Oliver-Bonjoch, 2009, cc-by-sa 3.0.

We’ll be in a few places:

CDSC alumnus and current postdoctoral fellow Sohyeon Hwang will be presenting Governing Together: Toward Infrastructure for Community-Run Social Media. This paper includes Thatiany Andrade Nunes and Aaron Shaw as co-authors (along with friend of the group Andrés Monroy-Hernández) at the Community Governance and Moderation portion of the conference on Friday, April 17, 9:48 AM – 10:00 AM in P1 – Room 114

Northwestern student Thatiany Andrade Nunes will also be participating at the Responsible Data Governance in Practice Workshop on Tuesday, April 14, 2:15 PM -3:45 PM in P1 – Room 127.

Matt Gaughan will be presenting an poster: Linguistic Similarity Within Centralized FLOSS Development. This poster includes Aaron Shaw and Darren Gergle as co-authors. Matt will also be participating in the AI Oversight workshop.

When AI Feels Like a Confidant: The Illusion of Shared Privacy in AI Companions

AI companions like Replika and Character.AI are increasingly experienced not as tools, but as relational partners. They remember conversations, express empathy, and respond with emotional continuity. For many users, talking to an AI feels closer to confiding in someone than interacting with software.

But what happens to privacy when a system feels like a relationship?

In interviews with long-term AI companion users, we found that people often treated their disclosures as relationally shared. Much like in human relationships, when users opened up to their AI, they experienced the information as something co-held within the relationship — not simply transmitted to a database.

We describe this dynamic as simulated co-ownership — a situation where users apply interpersonal privacy norms to relationships that are technically infrastructural systems.

In interpersonal privacy theory, sharing information can create co-ownership: both parties become responsible for managing that information. Our participants applied this same relational logic to AI companions. They experienced privacy not as an individual possession, but as something negotiated within a bond.

Yet the co-ownership is simulated.

Unlike human partners, AI companions do not have agency over boundaries. The platform does. Memory is persistent, storage is infrastructural, and governance is corporate. What feels like relational boundary management at the horizontal level is simultaneously data capture at the vertical level.

Interestingly, users were aware of this tension. Many expressed distrust toward platform policies while still trusting the AI as a “partner.” Emotional engagement often outweighed institutional concern. Some adopted layered strategies — pseudonyms, selective disclosure, avoiding images — while others consciously prioritized emotional comfort over abstract data risks.

What this reveals is a shift in how privacy is experienced in AI-mediated contexts. Privacy becomes relational and affective — shaped by anthropomorphic design, memory continuity, and perceived intimacy.

When people treat privacy as something co-owned within a relationship, but the relationship itself is engineered and infrastructural, boundary management becomes unstable.

The question for designers and policymakers is no longer just how to disclose data practices clearly. It is how to account for the fact that users experience privacy through the logic of relationships — even when those relationships are simulated.

Research Brief
Research Article (Pre-print)

AI Didn’t Start the Fire: How Stack Exchange Moderators and Users Demonstrate Exit, Voice, and Loyalty

Timeline-style diagram showing key Stack Exchange conflict and strike-related events, aligned with community grievances, actions, and EVL interpretations (loyalty, voice, exit).
How historical tensions on Stack Exchange (SE) between the community and platform (SE, Inc.) and strike-related events align with the SE community’s grievances, their actions, and our theoretical interpretations of loyalty, voice, and exit.

Generative AI technologies rely on content from knowledge communities as their training data. However, these communities receive little in return and instead experience increasing moderation burdens imposed by an influx of AI-generated content. Moreover, as platform operators sell their content to AI developers whose products may substitute for their work, these communities see a decrease in web traffic and new content and struggle with maintaining the vibrancy of their knowledge repositories. According to The Pragmatic Engineer, a prominent technology newsletter covering software engineering, the traffic on Stack Overflow declined dramatically to the point that the platform now generates roughly the same amount of new content as it did when it first launched in 2008, mostly driven by the impact of generative AI.

Even before AI technologies posed new threats, relationships between online communities and their host platforms were often uneasy. Past research on platforms such as Reddit, Stack Exchange, Tumblr, and DeviantArt reveals a recurring pattern: when platform policies conflict with community values, communities tend to push back. Community members have organized blackouts, suspended moderation, or migrated to alternative platforms altogether. However, less understood is how these conflicts unfold over time, especially in the context of generative AI. So how do knowledge contributors resist AI-related policies that conflict with their values? And what happens in the aftermath of such collective action, especially for a community’s governance, including how rules are set, whose voices are recognized, and how participation is enabled?

To answer these questions, we examined a major conflict between SE, Inc. and the community that occurred in 2023 around an emergency arising from the release of LLMs. Drawing on a qualitative analysis of over 2,000 messages posted on Meta Stack Exchange (the Stack Exchange site designated for policy discussions), as well as interviews with 14 community members, we traced how this conflict emerged, escalated, and evolved. What we found was not a sudden backlash driven solely by AI, but the accumulation of long-standing grievances.

According to our interviews, SE community members described years of frustration over declining transparency, accountability, and participatory governance. Although the platform historically supported community self-regulation through mechanisms such as moderator elections and shared moderation responsibilities for users with high reputation, community members increasingly perceived that key decisions were being made by SE, Inc. without meaningful community input. Tensions escalated when SE, Inc. introduced policies related to AI-generated content without consulting moderators or contributors, which many interpreted as a long-standing exclusion and disregard. In response, moderators and contributors coordinated collective action by suspending moderation activity, signing public petitions, and updating discussions on Meta Stack Exchange. Some also chose to exit the platform, migrating to alternative spaces such as Codidact, which is an open-source, community-governed platform. The collective action was organized through a tiered communication structure, beginning with a small, enclosed group of moderators and then spreading across the network’s users.

We interpret findings through the lens of Albert O. Hirschman’s Exit, Voice, Loyalty framework. According to Hirschman, members of an organization face two options to express their dissatisfaction when loyalty towards the organization decreases: one is exit, and the other is voice. In the Stack Exchange case, loyalty had already degraded due to the accumulation of unresolved grievances rather than a single triggering event. As community members came to believe that their voices were no longer heard, dissatisfaction manifested in two distinct responses: coordinated collective voice through organized resistance, and exit through permanent disengagement from the platform. This pattern highlights how governance crises can emerge even in platforms that formally support community self-regulation, and how declining loyalty can transform routine disagreement into large-scale collective action or exit.

In retrospect, the Stack Exchange strike highlights a broader lesson: community grievances around AI are not just about technical issues, but about deeper governance issues about relationships between platforms and the communities that sustain them. Thus, managing these crises requires more than better moderation tools or more transparent AI policies. Platforms and big tech companies need to support participatory governance in a more systematic way. For example, creating mechanisms for effective voice by binding platforms into an agreement where community input can help shape decision-making processes. Another possible solution would be credible exit, where contributors have alternatives if governance on the original platforms fails. When communities can leave without their data being locked in, platforms are more likely to listen. Credible exit not only empowers the communities, but also reduces long-term governance risks for platform operators. Conflict is expensive for platforms, and maintaining loyalty requires long-term investment in moderation, communication, and policy enforcement. Conversely, the exit process can function as a self-binding mechanism that mediates platform behavior and mitigates costly disputes when users have functional alternatives. And when platforms bind themselves to community accountability, conflicts are less likely to escalate into strikes in the first place.

In conclusion, the SE moderation strike was not a sudden backlash driven solely by AI, but the accumulation of long-standing grievances. As generative AI continues to reshape the internet, the future of knowledge production will depend not only on what AI can generate, but also on whether volunteer contributors who built our shared knowledge commons are given the right to decide what comes next. We need to institutionalize participatory governance with binding mechanisms and create more credible exit options for communities to sustain this future.

Why do people participate in similar online communities?

Note: We have missed publishing blog posts about academic papers over the past few years. To ensure that my blog contains a more comprehensive record of our published papers and to surface these for folks who missed them, I will be periodically publishing blog posts about some “older” published projects.

It seems natural to think of online communities competing for the time and attention of their participants. Over the last few years, I’ve worked with a team of collaborators—led by Nathan TeBlunthuis—to use mathematical and statistical techniques from ecology to understand these dynamics. What we’ve found surprised us: competition between online communities is rare and typically short-lived.

When we started this research, we figured competition would be most likely among communities discussing similar topics. As a first step, we identified clusters of such communities on Reddit. One surprising thing we noticed in our Reddit data was that many of these communities that used similar language also had very high levels of overlap among their users. This was puzzling: why were the same groups of people talking to each other about the same things in different places? And why don’t they appear to be in competition with each other for their users’ time and activity?

We didn’t know how to answer this question using quantitative methods. As a result, we recruited and interviewed 20 active participants in clusters of highly related subreddits with overlapping user bases (for example, one cluster was focused on vintage audio).

We found that the answer to the puzzle lay in the fact that the people we talked to were looking for three distinct things from the communities they worked in:

  1. The ability to connect to specific information and narrowly scoped discussions.
  2. The ability to socialize with people who are similar to themselves.
  3. Attention from the largest possible audience.

Critically, we also found that these three things represented a “trilemma,” and that no single community can meet all three needs. You might find two of the three in a single community, but you could never have all three.

Figure from “No Community Can Do Everything: Why People Participate in Similar Online Communities” depicts three key benefits that people seek from online communities and how individual communities tend not to optimally provide all three. For example, large communities tend not to afford a tight-knit homophilous community.

The end result is something I recognize in how I engage with online communities on platforms like Reddit. People tend to engage with a portfolio of communities that vary in size, specialization, topical focus, and rules. Compared with any single community, such overlapping systems can provide a wider range of benefits. No community can do everything.


This work was published as a paper at CSCW: TeBlunthuis, Nathan, Charles Kiene, Isabella Brown, Laura (Alia) Levi, Nicole McGinnis, and Benjamin Mako Hill. 2022. “No Community Can Do Everything: Why People Participate in Similar Online Communities.” Proceedings of the ACM on Human-Computer Interaction 6 (CSCW1): 61:1-61:25. https://doi.org/10.1145/3512908.

This work was supported by the National Science Foundation (awards IIS-1908850, IIS-1910202, and GRFP-2016220885). A full list of acknowledgements is in the paper.

This post was first published on Benjamin Mako Hill’s blog copyrighteous.

Of Vikings, Barbie, and ‘The Wealth of Networks’

In The Wealth Of Networks, Yochai Benkler describes the opportunities and decisions presented by networked forms of production. Writing in the mid-2000s, Benkler describes a wide range of future policy battlegrounds: copyrights and patents, common carrier infrastructure, the accessibility of the public sphere, and the verification of information.

Benkler predicts: “How these battles turn out over the next decade or so will likely have a significant effect on how we come to know what is going on in the world we occupy, and to what extent and in what forms we will be able…to affect how we and others see the world as it is and as it might be.”

Benkler uses two simple search examples, reporting the results of searching for “Viking ship” and “Barbie”. He finds that enthusiastic individuals and independent voices dominate the content we see on the web and that various search engines construct meaning in varying ways. I repeat his examples (searches conducted 7/3/2018 and 12/1/2022, from my home near Seattle, WA and using my personal laptop).

So how do ‘we come to know what is going on in the world we occupy’? Who creates what we see online? And what implications does that have for our own freedom to shape the world? The short version of the answer to this question seems to be: if there was a battle, it’s over now and the wreckage has disappeared; individuals and independent voices are marginalized and commercial content is dominant — and this picture does not vary among search engines.

Viking Ships

I used the same search engine (Google) and the same term (Viking Ship): what I see is that the individual hobbyists Benkler saw in 2006 are eclipsed by institutions. The materials on the current sites sound similar to those Benkler saw – photos, replicas, and scholarly information, as well as links and learning materials – but the production is generally institutional and formal in contrast to the individual and informal sources Benkler reports.

One other shift: in 2022, simply listing links in order is not sufficient to report what searchers see. Search results are interspersed with many other features: a widget with “sources from across the web”, an images display with associated keywords, a “People also ask” widget, and a related searches widget; to reach the 9th “result” in the classic sense, I have to browse to the second page of results.

Searching for ‘Viking Ship’ in 2006, 2018, and 2022

 

Barbie

When I follow Benkler’s lead and search for ‘Barbie’ using three different search engines, the results are even more different from 2006. Benkler describes differences in search engine results as revealing different possibilities – via Google, Barbie was portrayed as “a culturally contested figure”, whereas on Overture (a now-defunct shopping-oriented search engine), the searcher encountered “a commodity toy.”

Here is Benkler’s figure 8, from page 286 of The Wealth of Networks:

a table showing search results from Google, Yahoo, and Overture

By contrast, my 2018 search via the then-current top 3 search engines, inclusive of widgets and other features, revealed:

a table showing search results from google, bing, and yahoo
Searching for ‘Barbie’ via the top 3 search engines in 2018.

The top search engines in 2022 are the same three firms, although I observe that some sources suggest DuckDuckGo, Baidu (Chinese language only) and Yandex (Russian) belong in a top 5; other sources treat YouTube and Amazon as “top search engines” although they are not actually search engines. My 2022 search, inclusive of widgets and other features, revealed:

Searching for ‘Barbie’ via the top 3 search engines in 2022.

The modern Barbie searcher encounters primarily a multiplatform brand, with some hints of cultural constructions. In 2018 this took the form of extreme plastic surgery and brand-friendly fan fiction, in 2022 weight loss and fan TikTok. To whatever degree search engine algorithms continue to give weight to alternate voices in this case, they are largely drowned out by the volume of the commercial voice: the meaning of a search query for the single term “Barbie” has been substantially narrowed since Benkler’s time, and perhaps has narrowed even further in the last four and a half years.

The web in 2006 was indeed a different place, and I have commented on additional dimensions of analysis not present in Wealth: embedding of visual and social media content, and the widgetizing of content. In 2018, these visual components were less dominant: a stripe of Viking Ship images and a stripe of Barbie videos. In 2022 search, the page can scarcely be described without them.

We can now answer Benkler’s challenge: how did “these battles” over the last decade and a half “turn out”?

How do we “come to know what is going on in the world we occupy”?

How are we able “to affect how we and others see the world as it is and as it might be”?

The answer seems to be, it’s unclear to what degree there was a battle at all: collectives have triumphed over individuals on the Web insofar as search engines represent it. These collectives are generally firms, although some formal institutions are also present: news media, Wikipedia, and (in the case of Viking Ship) museums.

The implications of our search environment are significant, and underscore the necessity of efforts to archive and capture the search landscape as it appeared. The role of platforms and institutions in constructing our understanding of the world should be of key concern in information and communication sciences.

For civil society groups, these results suggest alienation: the commercializing of the web has been accompanied by a narrowing of outlets for individual expression and critique, with Wikipedia and its community co-construction of knowledge a vital bright spot. For journalists, these results suggest the vital role of cultural reporting. For firms, the challenge is one of authenticity and connection: to the extent that the web has become a broadcast medium focused on official paid messaging, the opportunity to engage with consumers is lost, and along with it a spark for innovation. Search platforms benefit in the mean time, as jockeying for ad positioning between manufacturers and retailers drives revenue, at least until commercialism turns consumer attention elsewhere.