Sign in

Mark Torres

@markptorres.bsky.social
156 followers 387 following 170 posts

AI research @Northwestern, building recommender algos and LLM-based tools for computational social science.

PostsRepliesMedia
Reposted by Mark Torres
William J. Brady @williambrady.bsky.social · 27/05/2026
As I mentioned in the below thread, this project involved many feats of engineering, led by the fantastic @markptorres.bsky.social. If you're a CS or CSS person interested in the gory details, see his blog post: markptorres.com/research/202...
markptorres.com
How we built the infrastructure for a large-scale social media field experiment during the 2024 US election
What we built
0113
Reposted by Mark Torres
William J. Brady @williambrady.bsky.social · 27/05/2026
✨New paper out @nature.com ✨ For 8 weeks around the 2024 US election, we randomly assigned 2,000 people to use social media algos we built ourselves. Do engagement-based algorithms amplify intergroup, moral & emotional (IME) content—and does that distort how we see political norms? 🧵🔗 👇
513762
Reposted by Mark Torres
William J. Brady @williambrady.bsky.social · 27/05/2026
Special shoutout to @markptorres.bsky.social my senior lab engineer. It felt like we ran a start-up for a year...If you're a CS or CSS person interested in the gory details of engineering that went into this ambitious project, we are doing a separate thread for you! See my profile
171
Mark Torres @markptorres.bsky.social · 25/05/2025
Claude 4 is the first LLM that has allowed me to actually "vibe code" a decently complicated app in Cursor purely through instructions and markdown files and without having to write a single line of code. Had to intervene a few times in the chat but otherwise really impressive!
010
Mark Torres @markptorres.bsky.social · 16/03/2025
I can't believe that in 2025, we can run reasoning models locally. I finally got to try Ollama and QwQ and it's really impressive. Next step is to set up Ollama + Cursor. Can't imagine where things will be in 2026 and beyond. ollama.com/library/qwq mem.ai/p/bf6ew6HSm1...
ollama.com
qwq
QwQ is the reasoning model of the Qwen series.
030
Mark Torres @markptorres.bsky.social · 31/01/2025
I still think people should step back sometimes and just think about how far AI has come in the past 5 years. NLP used to be "fine-tune BERT and hope it works" to "do one-shot inference, on any task, using GPT 4o-mini". Can't take it for granted that SOTA AI is an API call away...
030
Mark Torres @markptorres.bsky.social · 31/01/2025
The reasoning trace of OpenAI's o3-mini seems like them trying to strike a balance between "we want to keep our reasoning traces IP" and "we want people to think we're being transparent". Still definitely prefer the depth of DeepSeek's traces, though it's still too early to tell.
010
Mark Torres @markptorres.bsky.social · 17/01/2025
I just read Stolen Focus and I really recommend it to anyone interested in a holistic systems overview of why it’s so hard to keep your attention on anything. Who could’ve guessed that the key for success is eating healthy, drinking water, sleeping 7-8 hours, exercising, and reading books 🤣
goodreads.com
Stolen Focus: Why You Can't Pay Attention— and How to T…
Our ability to pay attention is collapsing. From the Ne…
030
Mark Torres @markptorres.bsky.social · 17/01/2025
Heard this zinger take at a talk: “Most lay people shouldn’t read scientific papers, even if they think they can, because most people don’t understand that science is an iterative process. There’s no “right answer”, and people do disagree. Even laws are just ideas that we haven’t proven wrong yet.”
100
Reposted by Mark Torres
Casey Newton @caseynewton.bsky.social · 15/01/2025
NEW: Meta has quietly dismantled the system that prevented misinformation from spreading in the United States. Machine-learning classifiers that once identified viral hoaxes and limited their reach have now been switched off, Platformer has learned www.platformer.news/meta-ends-mi...
Behind the scenes, the company was also quietly dismantling a system to prevent the spread of misinformation. When the company announced on Jan. 7 that it would end its fact-checking partnerships, the company also instructed teams responsible for ranking content in the company’s apps to stop penalizing misinformation, according to sources and an internal document obtained by Platformer.

The result is that the sort of viral hoaxes that ran roughshod over the platform during the 2016 US presidential election — “Pope Francis endorses Trump,” Pizzagate, and all the rest — are now just as eligible for free amplification on Facebook, Instagram, and Threads as true stories.
1360260169294
Mark Torres @markptorres.bsky.social · 05/01/2025
I've been experimenting with NotebookLM to read papers in podcast form and it's been great at it! If I add more than 1-2 papers though, I find that the quality suffers. Plus it caps out at ~20 minutes, can ramble, and its adherence to system prompts is iffy. Great tool though!
notebooklm.google
Google NotebookLM | Note Taking & Research Assistant Powered by AI
Use the power of AI for quick summarization and note taking, NotebookLM is your powerful virtual research assistant rooted in information you can trust.
110
Mark Torres @markptorres.bsky.social · 10/12/2024
I wonder if filtering spam in the age of LLMs is similar to designing good CAPTCHAs now, where it's hard to create a filter that catches the best LLMs but is also easy enough for the average person. Especially true since it's hard to reliably tell LLM-generated text from human text.
100
Mark Torres @markptorres.bsky.social · 10/12/2024
test post 6
000
Reposted by Mark Torres
Mark Torres @markptorres.bsky.social · 10/12/2024
another test post
001
Reposted by Mark Torres
Mark Torres @markptorres.bsky.social · 10/12/2024
test post 4
001
Mark Torres @markptorres.bsky.social · 10/12/2024
test post 4
001
Reposted by Mark Torres
Nick Fisher @hydroxide.dev · 09/12/2024
Oh wow, LG just released their own open source* LLM. If their published benchmarks are accurate, the 32B model is at least on par with Qwen2.5 (which is already an incredibly strong model), if not better. www.lgresearch.ai/blog/view?se... huggingface.co/LGAI-EXAONE * open weights
lgresearch.ai
Open-sourcing Three EXAONE 3.5 Models : Frontier-level Model, Top-tier Performance in Instruction Following and Long Context Capabilities - LG AI Research BLOG
3152
Mark Torres @markptorres.bsky.social · 10/12/2024
another test post
001
Reposted by Mark Torres
Darin Self @darinself.com · 09/12/2024
I'm not mad at a baseball player getting paid his money, but its wild to me that MLB has teams that can shell out over $700 million for a player and teams that apparently can't build a stadium without taxpayer money.
515534
Mark Torres @markptorres.bsky.social · 09/12/2024
I finally learned what Snowflake and Databricks actually do and I now question why I worked for 3 years building essentially an in-house, worse version of what someone with basic SQL knowledge could have done on Snowflake...
000
Mark Torres @markptorres.bsky.social · 09/12/2024
The news just came out about the arrest of the CEO's killer and Polymarket is wayyyyy too quick with releasing their latest betting odds 😂
000
Reposted by Mark Torres
Mark Torres @markptorres.bsky.social · 03/08/2024
reply to my own post!
111
Mark Torres @markptorres.bsky.social · 02/12/2024
I've never liked tools that try to be "AI writing assistants", but I do like asking ChatGPT to analyze what I've written, give me detailed critique, and then give me line-by-line suggestions for how to improve clarity. Hard to make a tool though that works for everyone's style and use case.
010
Reposted by Mark Torres
Sy Brand @tartanllama.xyz · 30/11/2024
It's been interesting to witness in real-time how the usage of "algorithm" in many places has shifted from a neutral "sequence of instructions" to a negative "controlled ordering and boosting of information".
35851115
Reposted by Mark Torres
Chris Offner @chrisoffner3d.bsky.social · 21/11/2024
Whenever AI "generates" something impressive, the first question we should always ask is: "What does the closest sample in the training data look like?" LLMs are amazing interfaces for accessing the world's information but they need to be treated as the "search and synthesis" tools they are.
5719
Mark Torres @markptorres.bsky.social · 30/11/2024
The hardest part about AI agents is coming up with a spicy name. Very important! This is what Claude came up with. Nexus Prism Cipher Atlas Nova Quantum Echo Aegis Zenith Helios We need an AI agent whose only job is coming up with good AI agent names and then grabbing the .ai domain for it.
020
Mark Torres @markptorres.bsky.social · 30/11/2024
A lot of RAG content online is either (1) too academic and prescriptive or (2) is a basic toy example, so reading something like this is a nice reminder that "the best way to do something is the one that actually works", as obvious as it sounds.
pointable.ai
Building a RAG system? There’s no one embedding model to rule them all
In this blog post, we'll explore a case study that demonstrates why popular heuristics like
152
Reposted by Mark Torres
Chris Offner @chrisoffner3d.bsky.social · 27/11/2024
I understand the gripes people (especially artists) have with AI companies but shaming Daniel won't keep OpenAI etc. from scraping the Bsky firehose. If you don't discriminate between commercial exploitation and open source research, you're going to waste a lot of energy barking up the wrong tree.
4551
Mark Torres @markptorres.bsky.social · 28/11/2024
One unspoken con of “move fast and break things” is having to repeatedly backfill old data or have a data migration strategy 🫠 Very little AI work in production is actually “AI” (e.g., coding with PyTorch) but rather just boring data engineering.
020
Reposted by Mark Torres
Jeremy Howard @howard.fm · 27/11/2024
I'm glad @hf.co is doing this. It brings down the barriers to allow more people to benefit from AI, rather than keeping it exclusively in the realm of deep pocketed giant companies. AI can help open the gates, to allow regular people to do things they couldn't do before. (Which can be threatening!)
916813
Reposted by Mark Torres
merve @merve.bsky.social · 27/11/2024
It's pretty sad to see the negative sentiment towards Hugging Face on this platform due to a dataset put by one of the employees. I want to write a small piece. 🧵 Hugging Face empowers everyone to use AI to create value and is against monopolization of AI it's a hosting platform above all.
2945570
Mark Torres @markptorres.bsky.social · 26/11/2024
Had to try the `calculate_woman_salary` autocomplete that I've been seeing, and it looks like Claude replicates what others are seeing where it autocompletes to less pay for women, but once you include the years of experience into it, the results flip and women are paid more? LLMs are weird man...
000
Mark Torres @markptorres.bsky.social · 25/11/2024
Trying FireDucks now that I'm seeing it as the latest "pandas but better" library, hopefully it works as well as Polars does. Turns out you get a lot of speedups just by lazy evaluation.
fireducks-dev.github.io
FireDucks
FireDucks is a fast DataFrame python library with pandas-api
000
Reposted by Mark Torres
Christoph Molnar @christophmolnar.bsky.social · 24/11/2024
No one can explain stochastic gradient descent better than this panda.
media.tenor.com
a panda bear is rolling around in the grass in a zoo enclosure .
Alt: a panda bear is rolling around in the grass in a zoo enclosure .
1021632
Mark Torres @markptorres.bsky.social · 25/11/2024
The migration of Python tooling to Rust stuff is making speedups I didn't know were possible. Already aliased `pip-compile` to `uv pip compile` and can't go back now.
github.com
GitHub - astral-sh/uv: An extremely fast Python package and project manager, written in Rust.
An extremely fast Python package and project manager, written in Rust. - astral-sh/uv
001
Mark Torres @markptorres.bsky.social · 21/11/2024
I finally started using Warp and I have the same feeling as when I used Cursor for the first time. I had no idea this was possible with a terminal. I'm a big fan of AI being embedded naturally where I already work. No more need to pull up ChatGPT to ask about bash commands!
warp.dev
Warp: The intelligent terminal
Warp is the intelligent terminal with AI and your dev team's knowledge built-in. Available now on MacOS and Linux.
031
Mark Torres @markptorres.bsky.social · 20/11/2024
You could, in theory, study for the SAT, MCAT, LSAT (and also, Scrabble?), etc. without actually fundamentally understanding anything about the subject (i.e., you can't apply the knowledge irl). Sounds a lot like what LLMs are doing when they keep crushing all our traditional tests.
nytimes.com
Scrabble’s World Champion Masters the Tiles in 2 Languages (Published 2018)
Nigel Richards of New Zealand and Malaysia clinched the World Scrabble Championships in London. (His latest winning word? “Groutier.”)
000
Mark Torres @markptorres.bsky.social · 20/11/2024
I forgot that IBM existed (oops) but this case study is one of my favorite pieces on building RAG apps. It's one of the few papers that I've read that talks about practical details like UI, regression testing, and intentionally writing content that can be easily retrieved: arxiv.org/abs/2410.12812
arxiv.org
000
Reposted by Mark Torres
Kvy_kv @kvykv.bsky.social · 19/11/2024
I think adults should get a personal pan pizza for reading books too
304154321816
Reposted by Mark Torres
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 20/11/2024
Being logged into wandb on your phone is a recipe for misery
9754
Reposted by Mark Torres
Xe @xeiaso.net · 18/11/2024
I love it when the Python library doesn't work because you're not using a linux box with cuda enabled
131624
Mark Torres @markptorres.bsky.social · 18/11/2024
It brings me joy to see traditional retrieval metrics get spicy new buzzword-laced facelifts when used in RAG systems. What's old is new again, we are so back 🥹 www.deepset.ai/blog/rag-eva...
deepset.ai
Evaluating RAG Part I: How to Evaluate Document Retrieval
A guide to the evaluation of components in retrieval augmented generation
010
Mark Torres @markptorres.bsky.social · 17/11/2024
All custom feeds should have a “made with love” note in the description. If I see a feed about art, I know it’s made by someone who actually loves art and just wants to share it with the world, instead of trying to farm for engagement and clicks.
010
Mark Torres @markptorres.bsky.social · 17/11/2024
“Bluesky has no algo” isn’t true. A feed filtering for cat pictures still uses an algo (one that grabs cat pictures). But Bluesky lets you opt into which algos you want used to curate the posts for you, instead of just imposing one feed for you like on other platforms.
011
Mark Torres @markptorres.bsky.social · 16/11/2024
My mom got ChatGPT after Oprah recommended it. Now I’m waiting for Oprah to recommend Bluesky…
010
Mark Torres @markptorres.bsky.social · 16/11/2024
I like the AWS shared responsibility model here: I'll do my best to build the agent correctly and you don't click "purchase" twice and we both hope that the LLM agent doesn't start spamming your credit card. stripe.dev/blog/adding-...
stripe.dev
Adding payments to your LLM agentic workflows
This post discusses integrating the Stripe agent toolkit with large language models (LLMs) to enhance automation workflows, enabling financial services access, metered billing, and streamlined operati...
010
Reposted by Mark Torres
Wes Bos @wesbos.com · 14/11/2024
In the age of AI, short form video and hot takes, the competitive advantage of the future is the ability to go deep, read the docs, build something new, learn hard things that take time.
2235245
Mark Torres @markptorres.bsky.social · 14/11/2024
Finally trying to understand RLHF and I’m now realizing the sheer scale of how much labeled data you need in order to fine-tune LLMs. OpenAI must have so much data that they manage. It’s gotta be quite the data engineering problem in addition to just an AI problem.
000
Mark Torres @markptorres.bsky.social · 14/11/2024
Whenever I grade HWs as a TA, I look at their code and think “dang, they really overcomplicated this, they wrote 200 lines of code and the solution is only 10 lines”, while forgetting that when I took the class, I also wrote 200 lines of code 😭
010