Sign in

Collective Intelligence Project

@cip.org
553 followers 44 following 152 posts

We're on a mission to steer transformative technology for the collective good. cip.org

PostsRepliesMedia
Collective Intelligence Project @cip.org · 21/01/2026
Full 2025 Global Dialogues Index Report cip.org/2025gdindex
cip.org
2025 Global Dialogues Index — The Collective Intelligence Project
020
Collective Intelligence Project @cip.org · 21/01/2026
Put together, these findings reveal early glimpses of the ways in which AI is reorganizing trust, intimacy, and work, providing a picture of how the world now lives with AI
100
Collective Intelligence Project @cip.org · 21/01/2026
New report drop! After seven rounds of Global Dialogues with more than 6000 people across 70 countries in 2025, we are releasing the 2025 Global Dialogues Index Report. blog.cip.org/2025gdindex
273
Collective Intelligence Project @cip.org · 15/08/2025
Apple: podcasts.apple.com/us/podcast/a... Spotify: open.spotify.com/episode/6UDj...
podcasts.apple.com
Audrey Tang and Divya Siddarth on Outfitting Democracy for the AI Era
Podcast Episode · Possible · 08/13/2025 · 52m
000
Collective Intelligence Project @cip.org · 15/08/2025
- Why "uncommon ground" beats common ground every time - Sci-fi book recommendations - And much more
100
Collective Intelligence Project @cip.org · 15/08/2025
- Our work bringing 100K+ people into AI development through globaldialogues.ai - How we're building evaluation benchmarks from lived experiences, not just lab tests - Digital twins that could represent your values without taking up all your evenings
globaldialogues.ai
Global Dialogues
Exploring humanity's vision for artificial intelligence through global conversations and collective intelligence.
100
Collective Intelligence Project @cip.org · 15/08/2025
What you'll find in this episode: - How Taiwan crowdsourced anti-deepfake legislation in 24 hours (and it worked) - Why 1 in 3 adults now use AI for daily emotional support, and what that means for democracy
210
Collective Intelligence Project @cip.org · 15/08/2025
@divya.bsky.social and @audreyt.org joined @reidhoffman.bsky.social and Aria Finger, hosts of the Possible Podcast, to talk about how democracy and AI can bring out the best of each other. Apple: podcasts.apple.com/us/podcast/a... Spotify: open.spotify.com/episode/6UDj...
podcasts.apple.com
Audrey Tang and Divya Siddarth on Outfitting Democracy for the AI Era
Podcast Episode · Possible · 08/13/2025 · 52m
120
Collective Intelligence Project @cip.org · 26/05/2025
We're asking the a global sample of the world: "𝖯𝖾𝗋𝗌𝗈𝗇𝖺𝗅𝗅𝗒, 𝗐𝗈𝗎𝗅𝖽 𝗒𝗈𝗎 𝖾𝗏𝖾𝗋 𝖼𝗈𝗇𝗌𝗂𝖽𝖾𝗋 𝗁𝖺𝗏𝗂𝗇𝗀 𝖺 𝗋𝗈𝗆𝖺𝗇𝗍𝗂𝖼 𝗋𝖾𝗅𝖺𝗍𝗂𝗈𝗇𝗌𝗁𝗂𝗉 𝗐𝗂𝗍𝗁 𝖺𝗇 𝖠𝖨, 𝗂𝖿 𝗍𝗁𝖾 𝖠𝖨 𝗐𝖺𝗌 𝖺𝖽𝗏𝖺𝗇𝖼𝖾𝖽 𝖾𝗇𝗈𝗎𝗀𝗁?" Prediction time: What % do you think will say yes? Tell us your response in the comments!
231
Collective Intelligence Project @cip.org · 23/05/2025
10/10: Read the piece to learn more about this under-explored issue. It includes specific strategies to address these biases and provides access to the full Github suite. www.cip.org/blog/llm-jud...
cip.org
LLM Judges Are Unreliable — The Collective Intelligence Project
When Large Language Models are used as judges for decision-making across various sensitive domains, they consistently exhibit unpredictable and hidden measurement biases, making their verdicts unrelia...
020
Collective Intelligence Project @cip.org · 23/05/2025
9/10: We built a Github suite to systematically test and quantify these biases. It lets you:
110
Collective Intelligence Project @cip.org · 23/05/2025
8/10: To improve reliability: Neutralize labels, vary order, empirically validate all prompt components, and optimize scoring mechanics. Diversify your model portfolio and critically evaluate human baselines.
100
Collective Intelligence Project @cip.org · 23/05/2025
7/10: These aren't just minor quirks. LLMs lack the mechanistic precision of traditional software. Their architecture means system prompts and input material exist in the same context, leading to unpredictable interactions.
100
Collective Intelligence Project @cip.org · 23/05/2025
6/10: Rubric-based scoring is also affected. We observed 'recency bias' where criteria scored later received lower averages. Holistic vs. isolated evaluation dramatically shifted scores too.
100
Collective Intelligence Project @cip.org · 23/05/2025
5/10: For example, in pairwise choices, LLMs favored "Response B" 60-69% of the time, a significant deviation from random. Even explicit "de-biasing" prompts sometimes increased bias.
100
Collective Intelligence Project @cip.org · 23/05/2025
4/10: LLMs exhibit cognitive biases similar to humans: serial position, framing, anchoring. Our tests across frontier models from Google, Mistral, Anthropic, and OpenAI consistently show these biases in judgment contexts.
120
Collective Intelligence Project @cip.org · 23/05/2025
3/10: "Prompt engineering" often relies on untested folklore. We found even minor prompt changes, like "Response A" vs. "Response B" labeling, significantly bias LLM choices.
100
Collective Intelligence Project @cip.org · 23/05/2025
2/10: This is important because LLMs are increasingly deployed for evaluation tasks, ranking, decision-making, and judgement in many critical domains.
100
Collective Intelligence Project @cip.org · 23/05/2025
1/10: LLM Judges Are Unreliable. Our latest blog post from @j11y.io shows that positional preferences, order effects, and prompt sensitivity fundamentally undermine the reliability of LLM judges.
120
Reposted by Collective Intelligence Project
seher @heyitsseher.bsky.social · 21/05/2025
The Collective Intelligence Project @cip.org has launched the Global Dialogues Challenge, an open call to explore global perspectives on the future of artificial intelligence. A $10,000 prize fund will be distributed among the winning entrants. www.cip.org/challenge
cip.org
Global Dialogues Challenge — The Collective Intelligence Project
162
Reposted by Collective Intelligence Project
James Padolsey @j11y.io · 20/05/2025
We're really thrilled to be able to have such a juicy prize fund. If you're feeling a sassiness with data and want to build something small to explore or inspire better AI for humans, take a look and enter. cip.org/challenge Step 1. Grab the data. Step 2. Build something cool. <3
cip.org
Global Dialogues Challenge — The Collective Intelligence Project
121
Collective Intelligence Project @cip.org · 19/05/2025
Details and how to apply: cip.org/challenge
cip.org
Global Dialogues Challenge — The Collective Intelligence Project
100
Collective Intelligence Project @cip.org · 19/05/2025
Submissions will be judged by an amazing panel: @audreyt.org (Cyber Ambassador-at-large for Taiwan) @nabiha.bsky.social (Executive Director of @mozilla.org ) Zoe Hitzig (Research Scientist at OpenAI and Poet)
342
Collective Intelligence Project @cip.org · 19/05/2025
The challenge runs from Monday, May 19th through Friday, July 11th. A $10,000 prize fund will be distributed among the winning submissions.
100
Collective Intelligence Project @cip.org · 19/05/2025
This is an open call to explore global perspectives on AI using the public datasets sourced from our globaldialogues.ai project. Participants can submit benchmarks, visualizations, artistic responses, or analytical reflections.
globaldialogues.ai
Global Dialogues
Exploring humanity's vision for artificial intelligence through global conversations and collective intelligence.
120
Collective Intelligence Project @cip.org · 19/05/2025
We're officially launching the Global Dialogues Challenge!
253
Collective Intelligence Project @cip.org · 14/03/2025
www.technologyreview.com/2025/03/11/1...
technologyreview.com
These new AI benchmarks could help make models less biased
They could offer a more nuanced way to measure AI’s bias and its understanding of the world.
110
Collective Intelligence Project @cip.org · 14/03/2025
“We have been sort of stuck with outdated notions of what fairness and bias means for a long time,” says @divya.bsky.social, “we have to be aware of differences, even if that becomes somewhat uncomfortable.” Read the full @technologyreview.com article on new approaches to evaluating AI ⬇️
150
Collective Intelligence Project @cip.org · 03/03/2025
5/ "When One LLM Drools, Multi-LLM Collaboration Rules" argues that single LLMs underrepresent real-world diversity. The authors propose multi-LLM collaboration to address reliability, democratization, and pluralism. arxiv.org/abs/2502.04506
arxiv.org
When One LLM Drools, Multi-LLM Collaboration Rules
This position paper argues that in many realistic (i.e., complex, contextualized, subjective) scenarios, one LLM is not enough to produce a reliable output. We challenge the status quo of relying sole...
220
Collective Intelligence Project @cip.org · 03/03/2025
4/ Vending-Bench tests LLMs' long-term coherence and capital acquisition—a capability relevant to AI risk scenarios. Top models struggle with simple business tasks, and breakdowns don't stem from memory limits. arxiv.org/abs/2502.15840
arxiv.org
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
While Large Language Models (LLMs) can exhibit impressive proficiency in isolated, short-term tasks, they often fail to maintain coherent performance over longer time horizons. In this paper, we prese...
210
Collective Intelligence Project @cip.org · 03/03/2025
3/ @robinsloan.com's "Is it okay?" provides a framework for AI trained on the commons of human writing: "If an AI application delivers some profound public good, it's probably okay. If it simply replicates Everything, it's probably not okay." www.robinsloan.com/lab/is-it-ok...
robinsloan.com
Is it okay?
Squaring up to the foundational question for language models.
110
Collective Intelligence Project @cip.org · 03/03/2025
2/ How Costa Rica is using bioacoustics to save ecosystems: Collective listening reveals forest health, monitors biodiversity, and measures conservation success. A beautiful example of using technology to advocate for nature. www.wired.com/story/costa-...
wired.com
Costa Rica Is Saving Forest Ecosystems by Listening to Them
Monitoring the noises within ecosystems reveals their health—allowing researchers to monitor changes in biodiversity, detect threats, and measure the effectiveness of conservation strategies.
110
Collective Intelligence Project @cip.org · 03/03/2025
1/ @alondra's "Three Fallacies" speech given in Paris: "The second fallacy...is that AI requires a tradeoff – between safety and progress, between competition and collaboration, and between rights and innovation. But each of these is a false choice." www.techpolicy.press/three-fallac...
techpolicy.press
Three Fallacies: Alondra Nelson's Remarks at the Elysée Palace on the Occasion of the AI Action Summit | TechPolicy.Press
Dr. Nelson was an invited speaker at a dinner hosted by French President Emmanuel Macron at the Palais de l'Élysée on February 10, 2025.
110
Collective Intelligence Project @cip.org · 03/03/2025
The future of AI isn't about single brilliant models, but collective intelligence that serves humanity. Our latest roundup explores how collaboration—between models, species, and nations—is reshaping AI governance and development. Our Weekly Collections:
291
Reposted by Collective Intelligence Project
B! 🐝 Cavello (they/them) @b-cavello.bsky.social · 27/02/2025
YESSSSS! 🔥 Grateful to be working with brilliant collaborators like Brandon Jackson at @metagov.bsky.social and @divya.bsky.social at @cip.org to do a better job of centering what people ACTUALLY want and need these tools to be! www.aspendigital.org/project/ai-b...
aspendigital.org
Community-Aligned AI Benchmarks
Reimagining the technical machine learning benchmarks that drive model development to reflect and encode public values.
161
Collective Intelligence Project @cip.org · 25/02/2025
Whose vision shapes our AI future? Visit globaldialogues.ai and join the effort to define an AI trajectory that serves humanity.
globaldialogues.ai
Global Dialogues
Exploring humanity's vision for artificial intelligence through global conversations and collective intelligence.
110
Collective Intelligence Project @cip.org · 25/02/2025
Proud to have @b-cavello.bsky.social and @aspendigital.bsky.social as partners on our Global Dialogues project. Check out @b-cavello.bsky.social's remarks to the UN on how Global Dialogues can advance AI governance and the goals of the #globaldigitalcompact.
131
Collective Intelligence Project @cip.org · 23/02/2025
8/ Google's AI co-scientist built on Gemini 2.0 combines multi-agent systems with scientific methodology to generate novel research directions. research.google/blog/acceler...
research.google
Accelerating scientific breakthroughs with an AI co-scientist
110
Collective Intelligence Project @cip.org · 23/02/2025
7/ @arcinstitute.org releases Evo 2, open-source genomics tool decoding 100,000+ species, integrating with NVIDIA BioNeMo for pattern analysis in genetic research. arcinstitute.org/news/blog/evo2
arcinstitute.org
AI can now model and design the genetic code for all domains of life with Evo 2 | Arc Institute
Arc Institute develops the largest AI model for biology to date in collaboration with NVIDIA, bringing together Stanford University, UC Berkeley, and UC San Francisco researchers
210
Collective Intelligence Project @cip.org · 23/02/2025
6/ Massive new report from Cooperative AI Foundation details failure modes and possible risks in multi-agent AI systems. arxiv.org/abs/2502.14143
arxiv.org
Multi-Agent Risks from Advanced AI
The rapid development of advanced AI agents and the imminent deployment of many instances of these agents will give rise to multi-agent systems of unprecedented complexity. These systems pose novel an...
120
Collective Intelligence Project @cip.org · 23/02/2025
5/ Recent study from @intuit.bsky.social, @matijafranklin.bsky.social and Rupal Jain examines market mechanisms—insurance, auditing, due diligence—as complements to traditional regulation for enhancing AI safety and accountability. arxiv.org/abs/2501.17755
arxiv.org
AI Governance through Markets
This paper argues that market governance mechanisms should be considered a key approach in the governance of artificial intelligence (AI), alongside traditional regulatory frameworks. While current go...
110
Collective Intelligence Project @cip.org · 23/02/2025
4/ Research models children's playful goals as reward programs, showing how humans and machines could share similar goal representations—a crucial step for AI alignment. via @guydav.bsky.social www.nature.com/articles/s42...
nature.com
Goals as reward-producing programs - Nature Machine Intelligence
To enable artificial agents to generate human-like goals, a model must capture the complexity and diversity of human goals. Davidson et al. model playful goals from a naturalistic experiment as reward...
120
Collective Intelligence Project @cip.org · 23/02/2025
3/ New literature review from Atoosa Kasirzadeh and @gbalint.bsky.social proposes a pluralistic framework integrating existential risk analysis with practical concerns like adversarial robustness and interpretability. arxiv.org/abs/2502.09288
arxiv.org
AI Safety for Everyone
Recent discussions and research in AI safety have increasingly emphasized the deep connection between AI safety and existential risk from advanced AI systems, suggesting that work on AI safety necessa...
120
Collective Intelligence Project @cip.org · 23/02/2025
2/ In a new paper co-authored by @divya.bsky.social , @audreyt.org, @glenweyl.bsky.social and others, they argue for redesigning social media platforms to build cohesion, not divisions. Read about it in @time.com time.com/7258238/soci...
time.com
Social Media Fails Many Users. Experts Have an Idea to Fix It
In a new paper, digital activist Audrey Tang and others emphasize the need for context and community over clicks.
273
Collective Intelligence Project @cip.org · 23/02/2025
1/ Weekly Collections: Key developments in prosocial media, AI safety frameworks, governance models, and breakthrough technologies reshaping AI's impact.
240
Collective Intelligence Project @cip.org · 16/02/2025
8/ OpenAI’s latest model spec showcases its Deliberative Alignment—a blend of chain-of-thought reasoning, process-based feedback, and reward modeling. Can we build this chain-of-command structure to capture a spectrum of diverse moral regimes? model-spec.openai.com/2025-02-12.h...
model-spec.openai.com
OpenAI Model Spec
The Model Spec specifies desired behavior for the models underlying OpenAI's products (including our APIs).
010
Collective Intelligence Project @cip.org · 16/02/2025
7/ New research on long chain-of-thought reasoning shows that scaling compute and careful reward design unlocks better error correction and complex problem solving in LLMs. Crafting reward signals that drive ethical chain-of-thought reasoning—and thwart reward hacks—is key. arxiv.org/abs/2502.03373
arxiv.org
Demystifying Long Chain-of-Thought Reasoning in LLMs
Scaling inference compute enhances reasoning in large language models (LLMs), with long chains-of-thought (CoTs) enabling strategies like backtracking and error correction. Reinforcement learning (RL)...
110
Collective Intelligence Project @cip.org · 16/02/2025
6/ IssueBench deploys 2.49M realistic prompts to measure bias in large language models. The findings reveal persistent issue bias that shapes public debate—an urgent call for robust AI evaluation. From @paul-rottger.bsky.social, et al. arxiv.org/abs/2502.08395
arxiv.org
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
Large language models (LLMs) are helping millions of users write texts about diverse issues, and in doing so expose users to different ideas and perspectives. This creates concerns about issue bias, w...
110
Collective Intelligence Project @cip.org · 16/02/2025
5/ In Southeast Asia, developers are crafting language models that capture local dialects and values. Yet, in a region as diverse as this, the politics of language raise tough questions about representation and influence. From @carnegieendowment.org carnegieendowment.org/research/202...
carnegieendowment.org
Speaking in Code: Contextualizing Large Language Models in Southeast Asia
Southeast Asia’s developers have sought to democratize AI by building language models that better represent the region’s languages, worldviews, and values. Yet, language is deeply political in a regio...
110
Collective Intelligence Project @cip.org · 16/02/2025
4/ @anthropic.com has published the Economic Index to track how AI is actually reshaping work, distinguishing between automation and augmentation in job tasks. www.anthropic.com/news/the-ant...
anthropic.com
The Anthropic Economic Index
Announcing our new Economic Index, and the first results on AI use in the economy
210