Sign in

Philipp Leitner

@philippleitner.net
312 followers 125 following 188 posts

Associate Professor @ Chalmers University of Technology icet-lab.eu

PostsRepliesMedia
Philipp Leitner @philippleitner.net · 13/02/2026
Of course the LLM will reinforcement-learn its way towards cheating your test suite somehow, so you'll need to stay vigilant of this. Maybe something like what we do in automated student assessment - e.g., a combination of seen and unseen test cases.
210
Philipp Leitner @philippleitner.net · 13/02/2026
(Some form of) TDD seems to be a fairly natural fit for vibe coding. You and the LLM both need a way to validate what the AI has implemented. A good, complete, and ideally simple test suite written upfront seems to be a natural way to provide that.
210
Philipp Leitner @philippleitner.net · 11/02/2026
I got interested and asked Claude to do a topic map of ICSE 2016 vs 2026:
100
Philipp Leitner @philippleitner.net · 29/01/2026
The older I get, the more convinced I become that our entire way of organizing capitalism (through stock markets and large international corporations) is fundamentally flawed and will, soon, flop on its belly. I just hope we are not standing beyond it when it does (but we likely will).
120
Reposted by Philipp Leitner
Victor Hogemann @victor.hogemann.com · 29/01/2026
“AI job losses“ is just a code word for “Global recession driven by the USA criminally unregulated financial market job losses“.
274
Philipp Leitner @philippleitner.net · 30/12/2025
Some holiday news - paper accepted in Future Generation Computing Systems (FGCS): doi.org/10.1016/j.fu...
doi.org
Redirecting
010
Philipp Leitner @philippleitner.net · 12/12/2025
Don't get me wrong, it's hard not to root for Expedition 33 and their super humble and nice team, but some of these awards are a bit sus. Is it even an indie game? Best art direction in a year with freaking Silksong? www.pcgamer.com/games/clair-...
pcgamer.com
Clair Obscur: Expedition 33 takes home an absurd 9 wins at The Game Awards, more than Baldur's Gate 3 in 2023
This year's Game Awards GOTY (and almost everything else) is the popular French RPG Clair Obscur: Expedition 33.
010
Philipp Leitner @philippleitner.net · 12/12/2025
Still, I found it fun to observe how PC Gamer dealt with their reviewing snafu over time. In the weeks after release, when it became clear that E33 will be a *big deal*, they took it with humor, but as time went on they adopted more of a policy of "the review shall not be mentioned henceforth".
000
Philipp Leitner @philippleitner.net · 12/12/2025
(Not actual criticism of PC Gamer. Reviewing is inherently subjective, and honestly Expedition 33 is kind of a weird and difficult-to-review game.)
100
Philipp Leitner @philippleitner.net · 12/12/2025
Remember when @pcgamer.com gave #expedition33 70% in their review, calling the game "rarely as fun as it looks"? Yeah, that didn't age great. #expedition33 is now officially the most decorated game at the Game Awards, ever. www.pcgamer.com/games/rpg/cl...
pcgamer.com
Clair Obscur: Expedition 33 review
Clair Obscur: Expedition 33 is a stylish riff on the JRPG, but its real-time-infused combat is rarely as fun as it looks.
100
Philipp Leitner @philippleitner.net · 11/12/2025
Reminder - I still have an opening for a postdoc in my lab (closing date is in one week): www.chalmers.se/en/about-cha...
chalmers.se
Vacancies
000
Philipp Leitner @philippleitner.net · 09/12/2025
It goes against common sense, everything we know about economics tells us it shouldn't work, there is no serious data that suggests it works, and yet it forms the basis of all economic decision making in the west. Why? Because it would be awfully convenient for the people in power if it *did* work.
030
Philipp Leitner @philippleitner.net · 09/12/2025
At some point in the future, people will read about trickle-down economics and have the same confused reaction that we have when learning how universally accepted catholic indulgences were in the Middle Ages.
173
Philipp Leitner @philippleitner.net · 03/12/2025
I'm nowhere close to a financial expert, but how these things usually go is that nobody is *obviously* overleveraged, but everyone depends on everyone and once the first domino pieces start to fall it triggers a chain reaction at whose end the bank's "safe investments" suddenly appear hazardous.
000
Philipp Leitner @philippleitner.net · 25/11/2025
That's a fairly common "playing both sides" argument. Productivity+++, but really nobody needs to worry for their jobs. These things are not likely to be true at the same time.
020
Philipp Leitner @philippleitner.net · 19/11/2025
I have a new job ad for a postdoc out: www.chalmers.se/en/about-cha... Application deadline: Dec. 18th Find out more about the work of my lab: icet-lab.eu
011
Philipp Leitner @philippleitner.net · 06/11/2025
I heard the term "spec-based programming" from a colleague for the paradigm where you really only provide and refine requirements, and do not care at all about the code. I don't think the tools I am using are there yet.
010
Philipp Leitner @philippleitner.net · 06/11/2025
IDK. My definition of vibe coding is "coding based almost exclusively on prompts, without or with minimal manual editing afterwards". Not sure if that is a standard definition, but it feels right.
110
Philipp Leitner @philippleitner.net · 06/11/2025
(8) An interesting mind shift happens when you vibe code a lot. Code turns into a kind of transient artifact that you just aren't very attached to. Is the code messy? Who cares (as long as it works), you aren't looking very much at it anyway. This has strong implications for security, safety, etc.
100
Philipp Leitner @philippleitner.net · 06/11/2025
(7) Overall, the final system turns kind of messy, but realistically so did all other research prototypes I implemented by hand. But now, nobody, not even me, really understands the messy system.
100
Philipp Leitner @philippleitner.net · 06/11/2025
(6) Planning mode is great. Claude is surprisingly good at creating, updating, and evaluating a plan of what to do. Complex changes became much more feasible once I started working more with planning mode upfront.
110
Philipp Leitner @philippleitner.net · 06/11/2025
(5) For non-trivial code, you'll still need decent understanding of the solution space. I feel like some of the more hairy implementation issues I could only solve because I implemented similar systems in the past, and could prompt the AI with *very* fine-grained designs.
110
Philipp Leitner @philippleitner.net · 06/11/2025
(4) Somewhat relatedly, AI loves to generate tests alongside changes (good) but they are often not very useful. They often stub out all business logic, turning them into classic "Python isn't broken" kind of tests. Getting it to write (and keep!) useful end-to-end tests seems surprisingly hard.
110
Philipp Leitner @philippleitner.net · 06/11/2025
(3) Do.Not. Trust. the AI when it declares success. Whether something actually *works* you need to check yourself. I'll leave this example here - the AI broken 25 tests, and decided after fixing one of the failures that the rest probably isn't their fault.
120
Philipp Leitner @philippleitner.net · 06/11/2025
(2) Validation is king, but also very hard. Again, these tools produce a lot of code. I quickly realized that reviewing it line-by-line is unrealistic. It may be more realistic when doing small changes in an established system, but in greenfield dev you have to go with the flow.
120
Philipp Leitner @philippleitner.net · 06/11/2025
Lesson Learned (1): you feel more productive than you truly are. These tools produce *a lot* of code in short time, but if you take a step back after a few weeks you notice than a fair bit of it wasn't actually that useful. It still takes time to build something that actually works, and not just 75%
120
Philipp Leitner @philippleitner.net · 06/11/2025
I initially used Gemini (in the console), but eventually moved on to Claude Console. They seem similar, but results from Claude where subjectively better, the tooling seems more mature, and the rates allowed me to work without much interruption. I am using the Pro subscription for USD 25 / month.
110
Philipp Leitner @philippleitner.net · 06/11/2025
For the last couple of weeks I have been trying to vibe-code a relatively complicated research system in the area of Java microbenchmarking in my spare time. I am slowly reaching the point where the system does something useful, so here are some initial impressions:
110
Philipp Leitner @philippleitner.net · 05/11/2025
People talk a lot about echo chambers on here, but I think it's important to remember that you are not entitled to anybody's attention, independently of how important you or your cause are.
010
Reposted by Philipp Leitner
Adolfo Neto @adolfoneto.elixiremfoco.com · 30/10/2025
People are saying that AI will transform the way we teach and learn. It has already transformed the way students cheat and, to my surprise, how they apologize for cheating.
0226
Philipp Leitner @philippleitner.net · 29/10/2025
27:0 ist aber auch im American Football eine richtige Klatsche.
000
Philipp Leitner @philippleitner.net · 13/10/2025
Note that this work does not judge correctness of papers, but only innovation. A crackpot theory is certainly very innovative.
010
Reposted by Philipp Leitner
Ian Foster @ianfoster42.bsky.social · 13/10/2025
Fascinating paper by Zhen Zhang & James Evans: arxiv.org/pdf/2509.05591  Analyzing 2M papers published immediately following the training of five prominent open LLMs, we show that ... the most perplexing are disproportionately represented among the most celebrated ... and also the most discounted.
arxiv.org
231
Philipp Leitner @philippleitner.net · 10/10/2025
If you want to see ChatGPT have a stroke in real-time just ask it "Is there a seahorse emoji?".
110
Philipp Leitner @philippleitner.net · 09/10/2025
How do software development companies think about LLM policies? New paper accepted in IEEE Software, Special Issue on AIware in the Foundation Models Era. Congratulations to Ranim Khojah, Mazen Mohamad, Linda Erlenhov, and Francisco Gomes Oliveira Neto. Preprint: arxiv.org/abs/2510.06718
arxiv.org
LLM Company Policies and Policy Implications in Software Organizations
The risks associated with adopting large language model (LLM) chatbots in software organizations highlight the need for clear policies. We examine how 11 companies create these policies and the factor...
051
Philipp Leitner @philippleitner.net · 07/10/2025
I do agree that formal peer review as we know it is a dead idea walking, but mostly because we have long swooped past a breaking point where the effort justified the systemic gains. (if there ever were systemic gains to start with, which isn't super clear to me)
110
Philipp Leitner @philippleitner.net · 07/10/2025
Everything is cyclical. Some 25 years ago European institutions pushed for quantitative assessment because previous qualitative, subjective assessments meant that departments were riddled with nepotism.
110
Philipp Leitner @philippleitner.net · 07/10/2025
What even was that? youtu.be/D5_xwX_jxB0?...
youtu.be
Kansas City Chiefs vs Jacksonville Jaguars Game Highlights | 2025 NFL Season Week 5
YouTube video by NFL
000
Philipp Leitner @philippleitner.net · 07/10/2025
Fuck's sake, the @chiefs.bsky.social lost the game via the most ugly touchdown I have ever seen. Heartbreak.
120
Philipp Leitner @philippleitner.net · 01/10/2025
"The American military will follow lawful orders and disobey unlawful ones." Will it? So far the track record of long-standing institutions pushing back isn't great. www.theatlantic.com/ideas/archiv...
theatlantic.com
Pete Hegseth Is Living the Dream
A man who retired as a major lectures hundreds of generals about the need to meet his standards.
020
Philipp Leitner @philippleitner.net · 01/10/2025
As a sidenote, I don't truly think that teaching students how to use AI effectively is all that important. These tools aren't hard to use, and the "tricks of the trade" are ephemeral. Teaching students how to prompt feels a lot like teaching SEO or googling.
010
Philipp Leitner @philippleitner.net · 01/10/2025
That said, you still need to know some form of programming / algorithmic thinking / problem decomposition when programming with AI. How to teach that when students have access to Gemini, ChatGPT, and Cursor day 1 I honestly have no idea.
200
Philipp Leitner @philippleitner.net · 01/10/2025
Every day that I use Gemini and similar, and every conversation I have with others, leads me closer to seeing insistence on "coding yourself" more similar to the "real programmers use assembly" mindset I abhorred myself when I was a student.
100
Philipp Leitner @philippleitner.net · 01/10/2025
It's an interesting question what "the fundamentals of programming" are going to be in an AI age. Two months ago I would have agreed that being able to program yourself, without AI, line-by-line, will remain crucial for the foreseeable future. Today, I'm much less sure.
110
Philipp Leitner @philippleitner.net · 28/09/2025
New paper accepted by Huaifeng Zhang, Mohannad Alhanahnah, YT, and Ahmed Ali El Din: BLAFS: A Bloat-Aware Container File System (accepted at the ACM Symposium on Cloud Computing) Preprint: arxiv.org/abs/2305.04641 Tool: github.com/negativa-ai/... Congratulations to Huaifeng and the team!
arxiv.org
The Cure is in the Cause: A Filesystem for Container Debloating
Containers have become a standard for deploying applications due to their convenience, but they often suffer from significant software bloat-unused files that inflate image sizes, increase provisionin...
040
Philipp Leitner @philippleitner.net · 19/09/2025
I am emphatically in favor of this new type of "open source ish" license: If you’re a little guy, do whatever you want with my work. If you’re a big guy, fuck you pay me.
041
Philipp Leitner @philippleitner.net · 12/09/2025
Slides for yesterday's talk at the 2025 WASP Software Engineering cluster meeting: www.icet-lab.eu/news/2025090...
icet-lab.eu
WASP Software Engineering and Technology Cluster Workshop Talk | Internet Computing and Emerging Technologies lab (ICET-lab)
Welcome to the Internet home of the the Internet Computing and Emerging Technologies lab at Chalmers and the University of Gothenburg
030
Philipp Leitner @philippleitner.net · 11/09/2025
Eh.
020
Philipp Leitner @philippleitner.net · 11/09/2025
Ich fürchte die traurige Wirklichkeit ist dass die USA seit Trump I ein entgleister Zug in Slow-Motion ist, dem man nur schockiert zusehen kann wie eine erwartbare Eskalation auf die Andere folgt.
1180
Philipp Leitner @philippleitner.net · 09/09/2025
I find this equal parts fascinating and weird.
010