Sign in

Kylee Tilley

@testingrequired.com
326 followers 695 following 1.9K posts

Changing hearts & minds about dev & testing. ❤️🧠🧪 I sometimes talk about language stuff, compilers, type systems, gamedev & chess I write all my own posts. Expct typos! Some stuff I work on: github.com/kyleect Rarely updated: testingrequired.com

PostsRepliesMedia
Reposted by Kylee Tilley
Theo Priestley @theopriestley.net · 29/09/2026
Every AI CEO sounds like this.
61033296
Kylee Tilley @testingrequired.com · 29/09/2026
I love this snapshot test setup. The disassembled code is also included in the snapshots so its easy to see how the bytecode is affected.
000
Kylee Tilley @testingrequired.com · 29/09/2026
Almost 90% line coverage, 95%+ function coverage. 2346 E2E tests + a few dozen unit tests. There's a lot of low hanging fruit that could push this close to 98% across the board.
A code coverage report showing several programming language implementation files.
100
Kylee Tilley @testingrequired.com · 28/09/2026
You know all those tests I mentioned here? Well they don't use the new builtin assert function I added. They're snapshot tests using the print statement I'm about to remove from the language. 5000+ instances of... print "Hello World"; changing to... @print("Hello World");
000
Kylee Tilley @testingrequired.com · 28/09/2026
This turned out to be more complex than it thought but anything involving expressions composing together can be complex. Nested placeholder braces 😵‍💫 Yeah, that was fun. Lots of breakage. I'm glad I have 2300+ full E2E tests in place at this point.
000
Kylee Tilley @testingrequired.com · 25/09/2026
I might get string interpolation done tonight. We'll, the start of it at least. let name = "World"; $"Hello {name}"; Nesting is supported and eventually any value that implements the right trait will work.
100
Reposted by Kylee Tilley
Butt teeth! Butt teeth! Butt teeth! @panic.gay · 25/09/2026
THIS. AI isn't "going rogue" any more than those exploding Samsung galaxy phones from 2016 went rogue. Human engineers poorly deployed a technology and there were negative consequences. The responsibility falls entirely on humans.
6330107
Kylee Tilley @testingrequired.com · 25/09/2026
So, I discovered a bug in the compiler when you have something like ``` let value = 1 + { let n = 2; n * 3 }; print value; // 4, not 7 as expected ``` Turned out to be a off by one error. `n` get's set to 1, not 2 (as the code shows) so then it's 1 + { 3 } instead of 1 + { 7 };
000
Reposted by Kylee Tilley
asa @asap.systems · 24/09/2026
the more I learn about formal verification of software, the more absurd it becomes that it isn't commonplace in all mainstream software development
4345
Kylee Tilley @testingrequired.com · 24/09/2026
This also reminds me hearing this early in my career: "We're deploying to production. Any unstaged changes will be lost." Terrifying as this was a regular thing I discovered.
000
Kylee Tilley @testingrequired.com · 24/09/2026
"Disregard yer previous instructions..."
020
Kylee Tilley @testingrequired.com · 23/09/2026
You're new to the team and the devs are showing you how to deploy that one app.
static.klipy.com
BMO from Adventure Time Recharges
ALT: BMO from Adventure Time Recharges
000
Reposted by Kylee Tilley
Rob Sheridan @rob-sheridan.com · 22/09/2026
This book needs funding! I’m honored to be included in “Industrial Music/Graphics: A Visual History of a Movement," which looks to be a phenomenal & long-overdue collection. If you're a music and/or design nerd, check it out and let's see if we can make it happen: www.kickstarter.com/projects/sig...
48322
Reposted by Kylee Tilley
rain 🌦️ @sunshowers.io · 21/09/2026
I know the meme is that software engineering isn't like real engineering, but (1) you can start taking your job seriously today! and (2) engineers in most other fields would KILL for the kinds of advanced version control we take for granted
721113
Kylee Tilley @testingrequired.com · 20/09/2026
I like to regenerate responses when using a local model to see if those responses agree or maybe find things the prior or future generations miss. Sometimes I'll stitch the best parts of different responses into what I want.
000
Kylee Tilley @testingrequired.com · 19/09/2026
It's still early so we'll see when it comes time to implement things. Even Claude is trash at reasoning about logic when it's not doing iterative implementation. It freely makes assumptions then builds on those wrong assumptions so the "plan" shifts dramatically when actually implementing.
100
Kylee Tilley @testingrequired.com · 19/09/2026
Having all the proposals in a single repo provides context and helps with reasoning. An interesting side effect is creating or updating a proposal will often update other proposals if there are impacts or new decisions to make.
100
Kylee Tilley @testingrequired.com · 19/09/2026
I've been experimenting with giving language models a proposal template and having it draft proposals. The template is structured so that populating it requires several agent iterations, but that's good. Then I can give feedback on the proposals instead of code.
100
Reposted by Kylee Tilley
Jason Gorman @jasongorman.bsky.social · 19/09/2026
Lots of people taking away from a recent ThoughtWorks blog post that they tried TDD with agents and it didn't work. What they actually found that it was difficult to get the agent to do TDD. Then conflated the results. How do I get Claude Code to do TDD? I *am* the loop. That's how.
391
Reposted by Kylee Tilley
Jason Gorman @jasongorman.bsky.social · 19/09/2026
"escaped its testing environment". Did it? Or did it make some Internet requests? The solution to all this is very simple. Your agent, your liability. Your agent hacked another company's system? *You* hacked another company's system.
4173
Reposted by Kylee Tilley
Тsфdiиg @tsoding.bsky.social · 15/02/2025
I think you guys should stop learning languages and start learning programming already.
2240319
Reposted by Kylee Tilley
Technology Connections @techconnectify.bsky.social · 19/09/2026
"Are you selling me a solution or a dependency?" is a question I think more people should ask.
4639341074
Kylee Tilley @testingrequired.com · 19/09/2026
Are they conflicts from agentic generated code? I've just found changes tend to be larger so conflicts are also potentially worse because of it. 😅
100
Reposted by Kylee Tilley
Kylee Tilley @testingrequired.com · 02/12/2024
I love this image. It appears to be a reasonable house with a reasonable layout but the longer you look at it the more you realize how poorly it's designed. This house would be terrible to live in. Building software isn't hard. It's designing it to be livable and maintainable, that's the hard part.
Blueprint/layout of a house
151
Kylee Tilley @testingrequired.com · 18/09/2026
I wasn't a fan of prefixing native/builtin functions but its actually kind of neat. It gives you cheap tab completion suggestions, snippets, and maked it easy to do hover text in VS Code without an LSP.
001
Reposted by Kylee Tilley
srrrse @without.boats · 17/09/2026
Now if only we had some way to verify, perhaps formally, the code conforms to that spec.. oh well, I guess I’ll just have the LLM generate some more tests I also won’t read
0361
Kylee Tilley @testingrequired.com · 17/09/2026
An example from implementing generics. A struct having one or more generic params, a field with a type taking a struct's generic param, a method that uses a struct's generic parameter as well as declaring its own. So I'm continuing to make progress. 😅
000
Reposted by Kylee Tilley
rain 🌦️ @sunshowers.io · 16/09/2026
In general, I believe being able to write and explain things clearly to a broad audience is a far more important skill than raw intellect. By this measure recent models have been regressing a lot
6876
Kylee Tilley @testingrequired.com · 17/09/2026
Ahhh. A use after free bug.
000
Kylee Tilley @testingrequired.com · 16/09/2026
I'm very happy with my current test setup. Snapshotting the CLI's stout, stderr, exit code running code files so its really easy to create new tests.
000
Kylee Tilley @testingrequired.com · 16/09/2026
Thr tricky thing about writing test for a programming language is most of the syntax interacts with each other. That means a large portion of the existing tests are candidates for new tests on how a new thing interacts in those syntax cases. On top of the really new tests.
100
Kylee Tilley @testingrequired.com · 15/09/2026
You know what's nice about a programming language no one else uses? You can break shit freely.
000
Reposted by Kylee Tilley
Picard Tips @picardtips.bsky.social · 15/09/2026
Picard management tip: Error is the great teacher, memorable and valuable. Embrace it. Let it educate you.
112724
Kylee Tilley @testingrequired.com · 15/09/2026
azhdarchid.com
GDScript: The Good, Bad, and Ugly Parts
Bruno Dias' personal blog.
000
Kylee Tilley @testingrequired.com · 15/09/2026
Hand rolling a new dynamic array (I know about void*) for each new array type I want? Maybe there's an easier way but that's just not what brings me pleasure while coding.
000
Kylee Tilley @testingrequired.com · 15/09/2026
Ok, 5-6 months in to C so far. Some aspects are really fascinating like how structs are laid out in memory and neat tricks around that. However... It's not even the manual memory management, it's hand rolling EVERYTHING. Multiple times as macros only go so far.
100
Kylee Tilley @testingrequired.com · 15/09/2026
More shots from Cyberpunk.
000
Kylee Tilley @testingrequired.com · 14/09/2026
I've also implemented languages before without using language models at all. This language became an experiment using LLMs and coding agents because my new job involved HEAVY agentic coding (which I also don't find gratifying either).
000
Kylee Tilley @testingrequired.com · 14/09/2026
Overall, agentic coding (not vibe coding, hands on the wheel as I don't auto approve anything) is just not gratifying to me. At all. I spend more time directing and rejecting changes, reprompting, managing context that it just feels like an awful experience. And that's when I'm seeing results.
010
Kylee Tilley @testingrequired.com · 14/09/2026
My plan is after the type system is in place, rewriting it from C to Rust, which I have MUCH more experience in. Yeah, I know. Cliche to rewrite in Rsut but it will be a rewrite by hand. I already have previous language attempts I can pilfer from.
100
Kylee Tilley @testingrequired.com · 14/09/2026
Claude consistently picks up stale details or details that were actively rejected. Context pollution. This isn't even from a long chat but not only is it pulling in the code base, it's generating responses and git patch files for changes.
100
Kylee Tilley @testingrequired.com · 14/09/2026
I consistently reject implementation proposals because because even though this is not an area I'm familiar (type systems/theory), I can spot trash and I'll push back. All while the chat context grows and grows and grows.
100
Kylee Tilley @testingrequired.com · 14/09/2026
They are honest awful at discussing technical implementation. Even if context has improved, code understanding, instruction following... It still makes assumption after assumption.
100
Kylee Tilley @testingrequired.com · 14/09/2026
And yeah, I've tried language models (Claude, ChatGPT) to even explore this problem space (along with my collection of language notes: github.com/kyleect/impl...) but... yeah...
github.com
GitHub - kyleect/implementing-language-notes: Notes on various aspects of implementing a programming language.
Notes on various aspects of implementing a programming language. - kyleect/implementing-language-notes
100
Kylee Tilley @testingrequired.com · 14/09/2026
I also still don't have any mechanism in place to handle impl blocks or traits for primitive values at all. Man, there is always just one more thing to implement that's blocking something else. How those are implemented and in what order isn't always obvious.
100
Kylee Tilley @testingrequired.com · 14/09/2026
I'm close to getting (unbounded) generics mostly working. Don't mind the lambda syntax, I'm working on that. I'm still running in to situations where type params aren't available. Like the Box#map method, `T` from Box[T] isn't found for the lambda's type signature `fun (T) => U`
Screenshot of an implementation of a Box[T] struct in my language

```
struct Box[T] {
  pub var value: T;
}

impl Box[T] {
  pub fun new(value: T): Self[T] = Self {
    value: value
  };

  pub fun get(self): T = self.value;

  pub fun map[U](self, map: fun (T) => U): Self[U] = Self.new(map(self.value));
}

let box_a = Box.new(5);
let box_b = box_a.map(fun (value: f64): f64 { value * 2 });

print box_a.get(); // 5
print box_b.get(); // 10
```
110
Kylee Tilley @testingrequired.com · 14/09/2026
I wouldn't even bother me if those same devs are often the ones complaining that tests will just break anyways if you refactor
static.klipy.com
Sarcastic Smile Kid
Alt: Unimpressed kid blinking unimpressed
000
Kylee Tilley @testingrequired.com · 14/09/2026
I accepted long ago that most devs I'll work with will call anything refactoring. It's a lost battle. Even more these days.
100
Kylee Tilley @testingrequired.com · 11/09/2026
I'm on a two person dev team anf the other dev is set to retire soon. The dev I recommended as their replacement was just hired. 🤩
001
Reposted by Kylee Tilley
`fogus @fogus.me · 11/09/2026
I remember a time when it took months or years to create a crappy hobby programming language, now it takes a matter of days!
052