Sign in

brumbus

@brumbus.bsky.social
115 followers 192 following 717 posts

growing the brumbus brand

PostsRepliesMedia
brumbus @brumbus.bsky.social · 3h
yeah boston rips
010
brumbus @brumbus.bsky.social · 05/10/2026
“they stole” ok then why do you still have it
000
brumbus @brumbus.bsky.social · 05/10/2026
I use AI an immense amount, every day, and I still do not understand at all how or why people have dozens of persistent agents constantly working. How many emails do you receive a day? How many projects have you amassed? Don’t you finish tasks? idgi. Maybe this is related to my not keeping tabs open
000
brumbus @brumbus.bsky.social · 05/10/2026
update: this is going very well
000
brumbus @brumbus.bsky.social · 05/10/2026
110
brumbus @brumbus.bsky.social · 05/10/2026
absolutely crushed our AI cost by pushing folks toward openai models for subagents. we let folks use opus or astra as they prefer, but insist on sol and luna subagents. we’ve gone from a nearly exponential trend to nearly flat, despite headcount growth
000
brumbus @brumbus.bsky.social · 02/10/2026
Doomsday, or the 120 days of Marvel
000
brumbus @brumbus.bsky.social · 02/10/2026
bedrock's reliability issues are killing me
000
brumbus @brumbus.bsky.social · 02/10/2026
a mood house?
000
brumbus @brumbus.bsky.social · 02/10/2026
i'm gonna cut your hair, and put it in a bowl
100
brumbus @brumbus.bsky.social · 26/09/2026
joining the "well I guess I will just maintain a prime-agent fork" club
000
brumbus @brumbus.bsky.social · 26/09/2026
a bloo bloo bloo
110
brumbus @brumbus.bsky.social · 26/09/2026
“it’s pathetic to ask what i mean”
210
brumbus @brumbus.bsky.social · 24/09/2026
barely on raw score, massively on cost
000
brumbus @brumbus.bsky.social · 24/09/2026
Sol 6 actually beats Astra on our internal company benchmark
100
brumbus @brumbus.bsky.social · 22/09/2026
(it was already important when we were still doing it manually, but it's moreso now)
000
brumbus @brumbus.bsky.social · 22/09/2026
the economics of testing code is changing, imo. 2 minutes of testing to 60 minutes of coding makes total sense, but 2 minutes of testing to 30 seconds of coding is different. testing is more important than ever as we take our hands off the wheel, but which tests and when is important.
100
brumbus @brumbus.bsky.social · 22/09/2026
and this is with the "all queued messages arrive at once" mode. don't get me started on the default one-at-a-time
000
brumbus @brumbus.bsky.social · 22/09/2026
I feel like I'm hacking the gibson when it launches 10 subagents and half of those launch their own, but is it efficient? I don't think it would have taken 1 astra 90 minutes to do what I asked of it. time isn't the only axis to care about, but it's one of them
100
brumbus @brumbus.bsky.social · 22/09/2026
it's super neat that sibling agents can communicate easily in prime-agent but I don't feel like it actually helps. heavy coordination slows things down a lot, plus it's expensive. each agent message is two calls.
110
brumbus @brumbus.bsky.social · 21/09/2026
imo system 1 models are clearly an exciting and promising concept, but Jev itself is kinda mid
000
brumbus @brumbus.bsky.social · 21/09/2026
Woo-hoo! Look at those tokens fly!
020
brumbus @brumbus.bsky.social · 21/09/2026
same. mine always thinks (5, 6, 7) is a win, or really any sequence of three consecutive numbers
110
brumbus @brumbus.bsky.social · 21/09/2026
I tried a few representations too and none really improved things. Even in a narrower "find the winning move" test it only got 16/24 correct.
110
brumbus @brumbus.bsky.social · 21/09/2026
whoa I literally just posted the same thing
100
brumbus @brumbus.bsky.social · 21/09/2026
But still, I think it's pretty reasonable to be disappointed that it can't consistently identify optimal tic-tac-toe moves, at least in the (I think reasonable) harness I built
000
brumbus @brumbus.bsky.social · 21/09/2026
I think it's that Doom play can look proficient even when mildly sloppy, whereas rigid discrete games like tic-tac-toe and chess aggressively punish even a single bad move.
100
brumbus @brumbus.bsky.social · 21/09/2026
how can Jev play Doom when I can't even get it to play tic tac toe very well?
100
brumbus @brumbus.bsky.social · 21/09/2026
it's not great at it. 0-58-2 against stockfish on skill 0. but it gets about half of the chess puzzles I've given it right.
000
brumbus @brumbus.bsky.social · 21/09/2026
having Jev play some chess
101
brumbus @brumbus.bsky.social · 21/09/2026
Astra, what the hell are you talking about
000
brumbus @brumbus.bsky.social · 19/09/2026
the older i get the more i understand
010
brumbus @brumbus.bsky.social · 19/09/2026
fixed something about my subwoofer setup. newfound levels of beefy
000
brumbus @brumbus.bsky.social · 18/09/2026
hmm
010
brumbus @brumbus.bsky.social · 18/09/2026
incredible
040
brumbus @brumbus.bsky.social · 18/09/2026
I know Jev can’t generate text but what would happen if you gave it a prompt, the alphabet as choices, plus what it had typed so far, in a loop? brb
100
brumbus @brumbus.bsky.social · 17/09/2026
Jev!
000
brumbus @brumbus.bsky.social · 17/09/2026
Jev is not great at MTG draft
000
brumbus @brumbus.bsky.social · 17/09/2026
Jev pulls the lever in the trolley problem every time
110
brumbus @brumbus.bsky.social · 17/09/2026
adversarial is the way!
011
brumbus @brumbus.bsky.social · 17/09/2026
i’m at the bar though so no testing til later!
000
brumbus @brumbus.bsky.social · 17/09/2026
just got Jev! feeling jevved up
100
brumbus @brumbus.bsky.social · 16/09/2026
ok yeah suno is pretty fun
000
brumbus @brumbus.bsky.social · 16/09/2026
the best one was definitely when Luna thought "Gatsby’s unattainable dream of winning Daisy and achieving the American Dream", acceptable only verbatim, was an appropriate answer for a bar trivia quiz
010
brumbus @brumbus.bsky.social · 16/09/2026
generally it would just pick the wrong part of the question to make unknown, like picking the year of a discovery instead of the scientist who did it. but it also liked to completely spoil the question sometimes too (which board game was AlphaGo good at?)
110
brumbus @brumbus.bsky.social · 16/09/2026
worked on a little trivia question generator last night with astra. we built a good pipeline (weighted sample of wiki pages, feed a few dozen to small models, bigger model downselects), but no model, including astra, was able to consistently gauge the interestingness or difficulty of a question.
120
brumbus @brumbus.bsky.social · 14/09/2026
i love making graphs
000
brumbus @brumbus.bsky.social · 10/09/2026
Finding local model performance to be impressive but also finding some negatives in the actual user experience. Cloud AI means I can work for hours and hours and still have tons of battery left. I don't like my laptop hot and loud. Mac Studio and other small desktop boxes seem like the way.
000
brumbus @brumbus.bsky.social · 10/09/2026
holy shit the 256GB mac studios don't arrive until January? we're getting one for work and I really don't feel like waiting that long!
000
brumbus @brumbus.bsky.social · 10/09/2026
“Everyone Is Cheating Their Way Through College” uh everyone already was decades ago, it sucks, that’s part of why i dropped out
000