Sign in

Jonathan Cheng

@jonathancheng.bsky.social
4.8K followers 1.9K following 1.3K posts

English Literature PhD turned ML Researcher. I convert cold brews into noisy experiments. Foundation Models @ Apple NYC 🏳️‍🌈

PostsRepliesMedia
Jonathan Cheng @jonathancheng.bsky.social · 26/09/2026
Poll 😂:
010
Jonathan Cheng @jonathancheng.bsky.social · 19/09/2026
The phrase “I’ll finally get a break after <mask> is done” needs to be removed from my training data. It really shouldn’t pass any factuality or quality filter.
030
Reposted by Jonathan Cheng
Kieran Healy @kjhealy.co · 18/09/2026
Which all just reeks of what happens when a status order gets upended by some new technique. Basically everyone who is under threat runs around shouting “Never mind what I appeared to value before, what I *really* value is this incredibly intangible thing I possess that also preserves my status”.
57613
Reposted by Jonathan Cheng
Maria Antoniak @mariaa.bsky.social · 10/09/2026
New from our lab! #COLM2026 When people generate stories, they don't just write one prompt. Instead, they explore narrative space via branching edits 🌱 We reconstruct 24k of these edit trees 🌳 from chat logs and map edit types, story formats, how they relate to tree depth, and more!
The Garden of Forking Prompts: How Users Explore Narrative
Space in Story Generation
Advait Deshmukh♣ Nora Benedict♠ Melanie Walsh♡ Maria Antoniak♣
♣University of Colorado Boulder ♠University of Georgia ♡University of Washington

Abstract

Large language models (LLMs) have changed the way people engage
with stories. Drawing on public chatbot logs, we can see that when users
generate stories, they iteratively edit their prompts to explore narrative
possibilities, adjusting characters, redirecting plots, and swapping fictional universes. As aggregated data, these prompts represent rich traces of creative preference at scale. Yet story generation evaluation benchmarks rely on static, one-shot prompts that cannot capture this exploratory behavior. In this work, we study how users revise consecutive story prompts in the wild. Using a dataset of naturally occurring user-chatbot conversations, we construct WildStories, a sample of 275,635 story generation prompts (labeled with story format, prompt components, and explicitness), and WildEdits, a collection of 24,291 edit trees that model how users iteratively edit base story prompts and explore branching story possibilities. From these trees we develop a framework of edit types crossing four directions (adding, removing, changing, and extending) with fourteen targets (e.g., plot, character, genre). We then use our datasets and this framework to
analyze user behavior in navigating narrative space via LLMs. Finally, we
show how automated permutations based on the framework can be used
for story generation benchmarking. Content Warning: This paper works with “wild” chatbot logs, which often include toxic and sexually explicit themes.
29529
Jonathan Cheng @jonathancheng.bsky.social · 11/09/2026
My facial expression when I wake up to wandb graphs that *do not* match my expectations.
041
Jonathan Cheng @jonathancheng.bsky.social · 09/09/2026
Like, which rl env produced this behavior? Is there an etaBench I’m unaware of?!
110
Jonathan Cheng @jonathancheng.bsky.social · 09/09/2026
Anytime a model gives me a weird time estimate, like “this code change will take ~1 day,” i lose a little sanity.
120
Jonathan Cheng @jonathancheng.bsky.social · 08/09/2026
And, by fun, I mean it haunts me in every interaction I have with that person.
030
Jonathan Cheng @jonathancheng.bsky.social · 08/09/2026
Having a coworker with 18 more years of experience lets me do a fun approximation: if I added the equivalent of three PhD programs/years of study, would I be performing at their level?
140
Jonathan Cheng @jonathancheng.bsky.social · 08/09/2026
Or that your intellectual interests can survive while being incredibly not-well defined. What I’m trying to say is every other word in your abstract better be ‘aporia’ or ‘liminal.’
040
Jonathan Cheng @jonathancheng.bsky.social · 08/09/2026
starting to think about what happened to border drawing when humans realized you could put cannons *on wheels*
140
Jonathan Cheng @jonathancheng.bsky.social · 08/09/2026
Relatedly, if we imagine that there could be related research groups within a tech company, one might imagine the motto is increasingly “all well-defined research is up for grabs.” Drawing borders of intellectual territory is hard when everyone can move *fast.*
1100
Jonathan Cheng @jonathancheng.bsky.social · 07/09/2026
This paper and its predecessor have just been a joy to read.
021
Jonathan Cheng @jonathancheng.bsky.social · 31/08/2026
I just happened upon Babymetal and the song Ratatata. And the people in my life are not ready for this to become my whole personality.
050
Jonathan Cheng @jonathancheng.bsky.social · 21/08/2026
It’s been a long week, but my god. Silent hill returning to its New England-Rust Belt Gothic roots. Still not sure why Japanese renderings of Centralia was such a thrill for my pre-teen self 😂 youtu.be/df9v2l1kEx8?...
youtu.be
Silent Hill Townfall - Official Gameplay Trailer
YouTube video by IGN
010
Jonathan Cheng @jonathancheng.bsky.social · 18/08/2026
“Girlfriend, the model is not diverging. It’s just exploring the loss landscape — let a girl live 🌸”
010
Jonathan Cheng @jonathancheng.bsky.social · 18/08/2026
Someone asked me what my ideal work environment looks like. And all I can imagine is me swiveling my chair, with a valley girl accent, saying “you already tried to increase the batch size, don’t do it, girlfriend!🌸”
160
Jonathan Cheng @jonathancheng.bsky.social · 17/08/2026
Claude just now: “the file you remember exists — as a tombstone.” …Jesus, Claude. Surely my files are not *so* dire?
050
Jonathan Cheng @jonathancheng.bsky.social · 16/08/2026
TIL that my dad, who wasn’t close with biological parents, was wrong about his genealogy. He always thought he was 1/2 Spanish, 1/2 Chinese. TIL I learned that I’m 1/4 Peruvian! And I have a bunch of cousins in Peru and *Wales* 😂
0100
Jonathan Cheng @jonathancheng.bsky.social · 15/08/2026
Rename tax *chef’s kiss*
000
Jonathan Cheng @jonathancheng.bsky.social · 14/08/2026
“Is reading fiction more an act of sft or rl or opd?” “Should we explain the benefits of opd through close reading?”
020
Jonathan Cheng @jonathancheng.bsky.social · 14/08/2026
Going to ask them: * their top period survey * their favorite author specific course * Robert Browning or Brownian motion?
120
Jonathan Cheng @jonathancheng.bsky.social · 14/08/2026
Just learned that a much smarter member of my team did their BA in English. And, boy oh boy, I’m about to become *very* annoying.
170
Jonathan Cheng @jonathancheng.bsky.social · 10/08/2026
The passage describing the wall of contorted, angry statues was foreshadowing…of bsky’s displeasure towards someone not liking Piranesi. So it’s more like a Hall. In the southwest wing, in the year the albatrosses arrived.
080
Jonathan Cheng @jonathancheng.bsky.social · 10/08/2026
Something, something rl gym environments.
071
Jonathan Cheng @jonathancheng.bsky.social · 05/08/2026
My favorite part about these model collaboration setups is that I get to feel like im working for both of them as a data broker — getting each other distillation data 🤣
020
Jonathan Cheng @jonathancheng.bsky.social · 03/08/2026
I’m disappointed in myself 😭😭
010
Jonathan Cheng @jonathancheng.bsky.social · 03/08/2026
I really wanted to love it. — Those whose judgement I trust have written of this Hall with great feeling, and I wished to feel it too. — The writing style is very fun to emulate but not very fun to read (for me).
010
Jonathan Cheng @jonathancheng.bsky.social · 03/08/2026
Got around to reading Piransei. Anyone who saw me on the subway, would’ve seen me *really* trying to get into it. My take: The Book is a Hall of great Beauty in which very little happens. I walked its length twice and found the same Statues, and the Tide, when it came, did not carry me anywhere.
460
Jonathan Cheng @jonathancheng.bsky.social · 01/08/2026
I’m upset that I’ve never heard this phrase before — or thought of it
010
Jonathan Cheng @jonathancheng.bsky.social · 01/08/2026
“Wabi sabi of integers” is so good
110
Jonathan Cheng @jonathancheng.bsky.social · 30/07/2026
Oh sure, if *I* watch videos in bed for 30 mins, it’s a “bad habit.” But if a model watches pedabytes of videos it’s “sft?” Life is so unfair.
060
Jonathan Cheng @jonathancheng.bsky.social · 23/07/2026
About to pose the classic weeb dilemma to my partner, would you rather watch: * a deeply serious and beautifully animated film about an aging Japanese actress * or… an anime about an idiosyncratic teenager using his power to usurp an American empire — also, robot fights!
020
Jonathan Cheng @jonathancheng.bsky.social · 17/07/2026
TIL Stonington Burrow contains one of TSwift’s homes and, perhaps more importantly, a nautical art museum! I’ve learned that “haunted galley” is generally what I’m looking for.
010
Jonathan Cheng @jonathancheng.bsky.social · 15/07/2026
*boat-building museum
020
Jonathan Cheng @jonathancheng.bsky.social · 15/07/2026
Saw a bunch of these ship heads at a boat building in Mystic today. Thinking things like, “first arcade fighter?” “Which character would I choose?”
120
Reposted by Jonathan Cheng
Naomi Saphra @nsaphra.bsky.social · 15/07/2026
if you think students will be routinely collaborating with LLMs after graduation, #1 priority is to study rhetoric because you need to recognize BS and critique the rigor of an argument that *looks* good. the elite users are going to be like, philosophy, classics, and english majors.
511014
Jonathan Cheng @jonathancheng.bsky.social · 15/07/2026
You sure it’s not the 3 books on R? :p
010
Jonathan Cheng @jonathancheng.bsky.social · 14/07/2026
I stumbled into the Yale dh library — filled with cheery thoughts. And then I stumbled into this and went, “ah right, I forgot that it all comes with some baggage.” 😊 Woof, that article. Big yikers.
010
Jonathan Cheng @jonathancheng.bsky.social · 14/07/2026
@sramsay2.bsky.social your book, well represented in Yale!
130
Jonathan Cheng @jonathancheng.bsky.social · 14/07/2026
Yale, I’m not gonna lie, you’ve got it real good. 😂
150
Jonathan Cheng @jonathancheng.bsky.social · 12/07/2026
Just watched Six the musical, and I’ve discovered: 1) the closer a musical is to a concert, the more I prefer it 2) Kirstin Maldonado from Pentatonix is one of the cast members right now!!! I think my favorite musical experience I’ve ever had?
010
Jonathan Cheng @jonathancheng.bsky.social · 08/07/2026
Wistful sigh! I miss it sometimes! Please enjoy all the food and cultural sites! /jealous 😂
020
Jonathan Cheng @jonathancheng.bsky.social · 06/07/2026
Tonight marks the first night my partner and I are cheering for very different teams tonight. Lord save this household 😂
030
Reposted by Jonathan Cheng
David Ho @davidho.bsky.social · 04/07/2026
😄
A world map filled with the flag of Cape Verde, featuring blue, red, and white stripes, along with yellow stars. Argentina is distinctly illustrated with its own flag colors and symbol.
32086576
Jonathan Cheng @jonathancheng.bsky.social · 03/07/2026
Upsides of heatwave, sneaking into usually busy restaurant, staring into David Bowie’s eyes. Downsides, *faints*
020
Jonathan Cheng @jonathancheng.bsky.social · 02/07/2026
Yes, the interest rates is the other big thing. I wasn’t aware of the tax credit part though!
110
Jonathan Cheng @jonathancheng.bsky.social · 02/07/2026
I haven’t seen too much convincing evidence that tech layoffs are not mostly signs of hiring over-eagerness from the covid era. Middle managers at various tech companies went unsustainably *bananas* during that period. And I think we’re seeing the delayed effects of that. Mostly anecdotal take tho
460
Reposted by Jonathan Cheng
Eryk @ambisinister.planetbanatt.net · 28/06/2026
Got some Liminal Space Dim Sum. The Lim Sum.
0263
Reposted by Jonathan Cheng
Bella Fascendini @bellafascendini.bsky.social · 26/06/2026
New paper! w/ @cocoscilab.bsky.social🧵Can large language models reason flexibly, or have they learned what reasoning looks like? We introduce a new paradigm to test this question—the riddle riddle—and find that humans and LLMs show opposite patterns of performance. 📜
619776