Sign in

Ryan Marcus

@ryanmarcus.discuss.systems.ap.brid.gy
4 followers 0 following 40 posts

Asst prof computer science @ UPenn. Machine learning for systems. Databases. He/him. [bridged from discuss.systems/@ryanmarcus on the fediverse by fed.brid.gy ]

PostsRepliesMedia
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 02/10/2026
You'll hop on the L for a quick trip into the city -- The announcer tells you that trains arrive every 8 minutes. Five trains come through going the other way. Fifteen minutes later, a train you want arrives looking like a mosh pit, smelling like a locker room, and with the visibility of a […]
discuss.systems
Original post on discuss.systems
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 14/09/2026
Putting a joke in your PhD / research application is shooting the moon. I've interviewed 100% of the students with SoPs that made me laugh, but very few of the ones that were obviously trying to make me laugh but didn't.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 05/09/2026
When you do a pushup, you mostly move yourself away from the Earth -- but you move the Earth a little bit too! I just wrapped up VLDB 2026. My first VLDB was 2016, so this was a soft "10 year anniversary" for me (I only attended 7 of the 10 VLDBs). It was […] [Original post on discuss.systems]
Screenshot of the learned query optimization session at VLDB 26.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 03/08/2026
I heard a song on the radio today that rhymed "Massachusetts" with "mass excuses," so I'm now calling for a temporary moratorium on new songs in order to establish appropriate safety measures and regulatory bodies.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 22/07/2026
CAREER submitted! I'm now moving from databases to my second passion, eating carrot cake.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 19/07/2026
Things I did during my PhD (2014 - 2019) for which I am now extremely grateful, but didn't seem like a big deal at the time: Zotero Roth IRA Rust Internships Teaching (once) TAing (all but 2 semesters)
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 10/07/2026
Blog post -- "How not to complain about peer review." I finally collected my thoughts about how I moved from "peer review rage" to a healthier (IMHO) perspective. rmarcus.info/blog/2026/07/10/peer-r…
A cartoon of a scientist's shiny machine and a beat-up paper. Caption: I can’t believe reviewer 2 didn’t comment on the flux initiators! Designed and captioned by me, rendered by Gemini.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 07/07/2026
In retrospect, one of the most valuable lessons I learned in grad school is that, even if you try really hard, you can still fail. Experiencing my own fallibility in the absence of any possible excuse gave me the opportunity to understand my own limitations, and taught me how to improve myself […]
discuss.systems
Original post on discuss.systems
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 16/06/2026
The "read every paper you cite" discourse doesn't seem to mesh with the reality of the wide range of what "reading a paper" can mean. 1️⃣ I can read a paper to understand if it's directly or tangentially relevant to a particular topic I'm working on in about 30 seconds. I do this dozens of […]
discuss.systems
Original post on discuss.systems
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 10/06/2026
Crazy idea: NSF GRFP decisions released prior to the PhD application deadline. (I'd love to learn the reason why this isn't the case, if anyone knows!)
001
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 28/05/2026
After painstakingly building a custom pipeline for database papers (SIGMOD, PODS, VLDB, CIDR) to replace the now-defunct public APIs from ACM and Semantic Scholar, DBScholar is live with the complete DB citation graph, semantically similar papers, PageRank […] [Original post on discuss.systems]
Screenshot of a “Database Paper Browser” web page showing details for the paper “How Good Are Query Optimizers, Really?” The page includes a short summary, paper metadata such as Paper ID 11313, venue VLDB, year 2016, pagerank, overall rank, and DOI, plus an authors list with six linked names. On the right, a bar chart titled “Incoming Non-self Citations Over Time” shows yearly citation counts from 2016 to 2026, peaking in 2025.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 15/04/2026
Model routers have taught me that if I don't see "thinking..." for 5-20 seconds before getting a response, I'm likely looking at garbage. I give it 3 months before we see artificially added "thinking" delays.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 12/04/2026
From "Implementation Techniques For Main Memory Database Systems" by DeWitt et al., 1984 -- one of the most influential database papers.
“Throughout the past decade main memory prices have plummeted and are expected to continue to do so. At the present time, memory for super-minicomputers such as the VAX 11/780 costs approximately $1,500 a megabyte. By 1990, 1 megabit memory chips will be commonplace and should further reduce prices by another order of magnitude. Thus, in 1990 a gigabyte of memory should cost less than $200,000. If 4 megabit memory chips are available, the price might be as low as $50,000.”
013
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 02/03/2026
Folks at "elite" universities have an extremely skewed vision of what the "college experience" looks like for the vast majority of Americans. According to data from Drafty ( drafty.cs.brown.edu/csprofessors ), there are: (100%) 377 Ivy League CS Professors (53%) 201 Have undergraduate […]
discuss.systems
Original post on discuss.systems
004
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 20/01/2026
RE: discuss.systems/@ryanmarcus/1157874… Update: Our paper received the CIDR Best Paper Award! Our lab is extremely grateful for the honor.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 19/01/2026
Seems like all the data agents, when faced with a massive input set, either synthesize traditional filters (like keyword search), or call an LLM for every row. There's a performance (both accuracy and latency) chasm between the two. Our ScaleLLM paper is one clue that there has to be something […]
discuss.systems
Original post on discuss.systems
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 26/12/2025
Most database teams optimize what they see in workload logs. But those very optimizations change what users choose to run! In our CIDR paper, we argue that industrial workloads exhibit 𝐬𝐮𝐫𝐯𝐢𝐯𝐨𝐫𝐬𝐡𝐢𝐩 𝐛𝐢𝐚𝐬: logs reflect a negotiation between users and the […] [Original post on discuss.systems]
A four-step cycle diagram showing feedback between database users and engineers.
① User submits their workload to the system: A square grid of colored squares represents the workload (3 green, 2 purple).
② Engineers observe properties of the workload: A database cylinder leads to a chart showing workload composition: purple 30%, green 60%. A speech bubble from an engineer character says, “Most queries are green. I’ll trade lower purple performance for higher green performance.”
③ Engineers identify hotspots and optimize their system.
④ Users optimize their workloads based on their platform: A speech bubble from a user character says, “Our platform is good at green, but bad at purple — send more green!” An arrow shows users adjusting the workload and feeding it back into the system.

Overall, the diagram illustrates a feedback loop where system optimizations influence user behavior, which then shapes the workload engineers observe.A side-by-side pair of charts.

Left chart: “Repeated Query Rate.”
A scatter plot shows median percentage of weekly repeated queries (y-axis, ranging roughly 57.5%–75%) versus query startup time in milliseconds (x-axis, 0–5500 ms). Eight points form an upward trend: higher startup times are associated with more repeated queries. A red regression line slopes upward across the points, labeled R² = 0.643. A small annotation at the bottom right reads p = 0.017, N = 8.

Right chart: “Feature Usage.”
A line chart shows the percentage of customers using a particular feature (y-axis, 0–10%) over time measured in weeks since an optimization release (x-axis, –15 to +17). Before release (negative weeks), usage fluctuates around ~2%. After week 0, usage gradually rises, reaching around 6% by week 10–15, with small variations.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 15/12/2025
Neo, a query optimizer powered by deep RL, is my first paper to hit 600 citations! While the cites are great, I'm most proud of the impact Neo had on the QO research community. I wrote a short retrospective on Neo, including how lessons learned led us to our next system, Bao […]
discuss.systems
Original post on discuss.systems
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 22/11/2025
SoCC '25 is a wrap! Shockingly, @dev was the only person to plug this Mastodon instance. Clearly, we have work to do!
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 03/11/2025
@Talie and I's Halloween costume might've been too deep a cut. But all the PL people got it immediately!
Two adults in costumes stand smiling inside a home. The person on the left is dressed as a janitor, wearing a dark blue coverall with a name tag and holding a broom. The person on the right is dressed as an Expo whiteboard marker, wearing a white apron with the Expo logo and a pink rolled paper “marker cap” on their head. They are standing close together in a doorway with an arm around each other.
010
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 05/10/2025
During my PhD, I got paid about 22% of my highest-declined job offer. During my postdoc, I got paid about 30% of my highest-declined job offer. During my assistant professorship, I'm getting paid about 15% of my highest-departed job's salary. If you want to make it in academia, you can't just […]
discuss.systems
Original post on discuss.systems
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 15/09/2025
Now that I've been the PC chair for a conference, I will never submit a late review again.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 31/08/2025
This is a class of bugs that can only exist in cyberphysical systems. Plane thinks it's on the ground, but it's actually in the air. Hard to imagine being on this support call -- "The plane says I'm on the ground." "Are you?" "Nope." "Huh. That's not good." Salient quote: "At that point, the […]
discuss.systems
Original post on discuss.systems
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 05/08/2025
If you still use Twitter, there's a sale on Premium today. I went through the checkout process, closed the window at the Stripe payment page, then got an email 10 minutes later with an offer for a year of free Premium. I have a stupid blue check next to my name but there are fewer ads! "Cancel […]
discuss.systems
Original post on discuss.systems
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 06/07/2025
I thought database folks were always hyperbolic and overly resistant to change in the peer review process, but I just saw a NeurIPS AC compare desk rejecting papers submitted by absentee reviewers to familial extermination / collective punishment during the Qin dynasty, so maybe the database […]
discuss.systems
Original post on discuss.systems
020
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 12/06/2025
This is your yearly reminder that anyone who publishes CS papers should have a personal website that lists their current position, research interests, publications, and email address. If you don't, it's basically impossible for me to invite you to a PC […] [Original post on discuss.systems]
A meme featuring Bernie Sanders standing outdoors in a winter coat, speaking directly to the camera. The caption reads, “I am once again asking PhD students to make a damn website."
3224
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 03/06/2025
OLAP workloads are dominated by repetitive queries -- how can we optimize them? A promising direction is to do 𝗼𝗳𝗳𝗹𝗶𝗻𝗲 query optimization, allowing for a much more thorough plan search. Two new SIGMOD papers! ⬇️ LimeQO (by Zixuan Yi), a 𝑤𝑜𝑟𝑘𝑙𝑜𝑎𝑑-𝑙𝑒𝑣𝑒𝑙 […] [Original post on discuss.systems]
Infographic describing LimeQO, a workload-level, offline, learned query optimizer. On the left, it shows a workload consisting of multiple queries (q₁ to q₄), each with a default execution time (3s, 9s, 12s, 22s respectively). On the right, alternate plans (h₁, h₂, h₃) show varying execution times for each query, with some entries missing (represented by question marks). For example, q₁ takes 1s under h₂, much faster than the 3s default. A specific callout highlights that for q₃, plan h₃ reduced the time from 12s to 3s, but took 18s to find, resulting in a benefit of 9s gained / 18s search. The image poses the question: “Where should we explore next to maximize benefit?” The image credits Zixuan Yi et al., SIGMOD '25, and provides a link: https://rm.cab/limeqoInfographic describing BayesQO, an offline, multi-iteration learned query optimizer. On the left, it shows a Variational Autoencoder (VAE) being pretrained to reconstruct query plans from vectors, using orange-colored plan diagrams. The decoder part of the VAE is retained. In the center and right, the image shows Bayesian optimization being performed in the learned vector space: new vectors are decoded into query plans, tested for latency, and refined iteratively. At the bottom, a library of optimized query plans is used to train a robot labeled “LLM,” which can then generate new plans directly. The caption reads: "We get a fast query, but also a library of high-quality plans. We can train an LLM to speed up the process for next time!" The image credits Jeff Tao et al., SIGMOD '25, and links to https://rm.cab/bayesqo
010
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 01/06/2025
The abstract deadline for SoCC '25 is about a month away! This year's event is fully online. acmsocc.org/2025/papers.html
acmsocc.org
2025 ACM Symposium on Cloud Computing
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 29/05/2025
Kept writing bad code today, an expert had to take over and guide my hand.
A small, colorful bird, a Bourke's parakeet, with a mix of pastel pink, blue, green, and gray feathers is perched on a person's arm while they type on a keyboard. The setting appears to be a workspace with a computer mouse, mouse pad, and other desk items visible in the background.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 24/05/2025
"Database venues are just about LLMs now!" -- LLMs are certainly on the rise, but this claim isn't supported by the data. Papers about LLMs are certainly growing very quickly, but last year there were: * Almost 3x more papers about indexing, * Almost 4x […] [Original post on discuss.systems]
A table and a line chart depict trends in database research papers from conferences such as VLDB, SIGMOD, CIDR, and PODS.

The table on the left shows the number of papers from 2017 to 2024, categorized by topic: all papers, query optimization, transaction processing, indexing, and large language models (LLMs). It shows that overall paper counts vary each year, peaking at 778 in 2023. Query optimization papers consistently appear in large numbers, while LLM-related papers increase dramatically—from 0 in 2017 to 57 in 2024—indicating growing interest in LLMs in database research.

The line chart on the right shows the historical trend from 1970 to 2024 for each category. Total papers increase exponentially over time. Query optimization, transaction processing, and indexing papers also rise gradually. LLM papers remain flat until around 2018, then begin a sharp upward trend, reflecting their recent emergence in the field.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 08/05/2025
When modern analytic databases process `GROUP BY` queries, they tend to use a partitioning strategy for parallelism. The conventional wisdom is that partitioning has better scalability due to lower contention. But is this wisdom still true in 2025? Penn […] [Original post on discuss.systems]
A diagram showing a two-stage parallel processing system involving key-value pairs. On the left, a column labeled "K V" holds key-value pairs divided into morsels (small batches) for processing. Stage 1 is labeled "ticketing", where keys are matched against a shared hash table (middle column labeled "K T") to obtain a ticket (index). These tickets help arrange the data into a new table (right-middle column labeled "T K V"). Stage 2, labeled "update", uses the ticket to update values into a shared result vector (rightmost column labeled "V") at the index specified by the ticket. Blue and red colors distinguish different morsels processed concurrently.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 02/05/2025
A meme featuring the "two buttons" format. In the top panel, a hand hovers hesitantly between two red buttons. One button is labeled "diatribe: peer review is broken!" and the other "our lab is proud to present...". In the bottom panel, a distressed man in a superhero outfit wipes sweat from his forehead, clearly anxious. The caption reads: "academics when 1 paper gets accepted and 1 paper get rejected". The meme highlights the conflicting emotions researchers face when dealing with peer review outcomes.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 17/04/2025
The NSF GRFP, a training grant awarded to promising American students at public and private colleges looking to earn a PhD, was cut in half this year. When faculty admit a PhD student, they are committing to raising ~$500k to fund that student. The GRFP […] [Original post on discuss.systems]
Stacked bar chart titled "NSF GRFP Recipients" showing the number of recipients from private and public institutions for each year from 2015 to 2025. Each year is represented by a stacked bar with blue indicating private institutions and orange indicating public institutions. The total number of recipients remains relatively stable around 2000–2100 per year, with a noticeable peak in 2023 reaching above 2500 and a sharp drop in 2025 to about 1000. The distribution between private and public institutions varies slightly each year, with public institutions generally having a slightly higher share.
010
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 24/03/2025
The deadline for aiDM 2025 -- the SIGMOD workshop on Exploiting Artificial Intelligence Techniques for Data Management -- has been extended to this Friday, March 28th. If you were on the fence about a submission, now is your change to make it! aidm-conf.org/#dates
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 12/03/2025
Just a reminder that the names of your bibtex citations get included in the PDF (both as the link anchor name and in the metadata), so if you name a paper `morons_who_copied_us`, reviewers and readers will be able to see that...
1516
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 06/03/2025
A very nice paper from UMD about catalog storage on data lakes. While I'm not totally sold on their solution (I have some doubts about the hierarchical data model), the discussion of various tradeoffs and design principles is top notch. I think there's clear space for major innovation in the […]
discuss.systems
Original post on discuss.systems
001
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 15/02/2025
Pair(akeet) programming.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 03/02/2025
I made a simple tool to look for related database papers (VLDB, SIGMOD, CIDR, PODS) given a new paper's title and abstract using vector embeddings. It highlights authors who are currently in the reviewer pool. I'm sure my horrible Python hack will break at […] [Original post on discuss.systems]
A screenshot of the linked webpage.
000
Ryan Marcus @ryanmarcus.discuss.systems.ap.brid.gy · 12/01/2025
My hot take of the day is that we're pretty good at evaluating PhD applicants (at least better than random), but the number of qualified applicants greatly exceeds the number of available slots. So there exists pairs of students (X, Y) where X is admitted and Y isn't admitted, but the […]
discuss.systems
Original post on discuss.systems
010