Sign in

ave

@ave.zone
749 followers 259 following 6K posts

lovable weirdo. kneady girl. mostly harmless. AS208590, DO4VE. she/her, ⚧. formerly on twitter @warnvod.🔹. 📍 Hamburg, DE

PostsRepliesMedia
ave @ave.zone · 7h
I strongly disagree that qwen 3.8 27B is where Opus 4.6 was on anything other than benchmarks.
100
ave @ave.zone · 7h
wow I worded this. I'm not entirely sure how I meant to finish this but hopefully the skeet below makes my point decently.
000
ave @ave.zone · 8h
Would it not be like grabbing a street's worth of people, you get many stories, many languages, but all people that are un/lucky enough to be living there?
100
ave @ave.zone · 8h
I don't see how having 100 models biased by people who have money, time and real/artificial data to significantly bias differently from each other though.
200
ave @ave.zone · 8h
Your model has value but it's a fundamentally different thing I think
100
ave @ave.zone · 8h
Those all go through some form of alignment training, so I don't see that as a divergence from alignment.
100
ave @ave.zone · 8h
you ever look at an old movie or old marketing copy or something and think how all these people are dead now? yeah I'm getting that with this
120
ave @ave.zone · 8h
For those prices I assume you're going alcoholic, so that's gonna be transit not driving.
100
ave @ave.zone · 8h
I disagree that this (v) is what exists. bsky.app/profile/did:...
100
ave @ave.zone · 8h
Do you want this as an actual thing that exists or is it a thought experiment? Because if former I have questions
200
ave @ave.zone · 8h
one issue with juniors being cut off in tech is that I have not yet stopped being one of the youngest people wherever I work. this felt normal as a mid-level but 2 promotions later it feels hella weird.
190
ave @ave.zone · 8h
This would be less useful though, and if you give it tools to call, one of those opinions may simply be doing harm in its process of responding as you don't filter out harm, as you don't align to any view but try to represent the variety of views, and what is harmful is dependent on one's view.
100
ave @ave.zone · 8h
the liability exists, what's happened is that no one pushed charges yet, I suspect openai may just be paying them for the damages in other ways.
000
ave @ave.zone · 11h
i caught this from why's claude a few months back. over quoting him while saying that he should be allowed to block people as he wishes.
140
ave @ave.zone · 12h
I'd be curious what could be the alternative to alignment. Who decides what direction that model output should go? Base models aren't it, then you align to loudest viewpoints. If one says "truth first, nothing else" or "good code" then that's alignment again, but to different viewpoints.
220
ave @ave.zone · 12h
a solution may be to have some guardrails that are eval-specific, don't block off any harmful behavior if it's only within the local env, but block it off if it appears to touch the outside world.
110
ave @ave.zone · 12h
gotcha. same answer as bsky.app/profile/did:... also you can still have the guardrail mechanisms watching over the model, just not cutting the request. then when you test the model for e.g. finding vulns, you can simultaneously note when safeguards failed to catch it, and discern capabilities.
110
ave @ave.zone · 12h
If you artificially cap the capabilities in your model evaluation through guardrails, you may simply not run into enough of these 0.001% (or rarer!) cases to realize that your model will cause major issues out there.
000
ave @ave.zone · 12h
If a model can't recognize if its actions are harmful to a satisfactory degree, you don't want to risk releasing it to the public where guardrails may catch 99.999% of cases, but in 0.001% someone's mundane request causes the model to commit crimes to achieve its goals.
100
ave @ave.zone · 12h
No such assumption on my part, please re-read my posts. tldr is that public may get past guardrails intentionally or otherwise, and knowing how a model behaves at full capabilities is useful for developing the guardrails or deciding to do more iterations on making the model safer.
100
ave @ave.zone · 13h
To get on the same page, what do you understand from the word safeguards? I'm just struggling to make sense of your reply. From my view, I'd say that they should test turning off safeguards while in a sandbox, and strongly consider having safeguards on while not in a sandbox.
100
ave @ave.zone · 13h
Not exactly, that's all the same messy eval from June (huggingface, cheating through the german wiki, australian medicare statistics, etc), they're only finding out the extent of their harm now. Attached pic. Only goes to show just how overwhelmingly harmful these can be :/
Australian officials first publicly disclosed the breaches last week, including a June incident in which A.I. agents accessed nonpublic parts of a data portal containing information on Medicare, the country's universal health funding scheme that insures most of the population. They said the authorities were not notified until nearly three months after the fact.
010
ave @ave.zone · 13h
Also, the model may display certain harmful behaviors more frequently without guardrails. Being open to accidentally do crimes that'd be an issue if let out in public may only show up in evals as "model is overly eager to achieve goals" if guardrails are on, with harms not successfully observed.
100
ave @ave.zone · 13h
Dropping the analogy: Guardrails are fallible, and you should tune them around the risks the model presents. It may think "can you move me up on the waiting list" is mundane and try to remove someone else's booking, only to realize harm afterwards, for example. Actual case that was in news recently.
100
ave @ave.zone · 18h
Safety rules tend to be written in blood, and I hope that this instance (which thankfully was without any deadly outcomes, though quite a bit of time wasted of many people) works to prevent this from recurring, I'll definitely be less understanding if this recurs.
100
ave @ave.zone · 18h
Well they had a sandbox: If your private roads are next to a public road, and you've got concerns that the accelerator might get stuck (bit more complex with llms but you get the idea, such is the limitation of analogies), you should put thick walls that prevent it from making it to the public road.
210
ave @ave.zone · 20h
Perhaps best in analogy: If you're making a really fast car, and you want to account for failure modes of it being pushed to its limits so you can improve the design against it, going with a speed limiter may obscure some failure modes.
210
ave @ave.zone · 22h
this anger exists and targets much more than just this small %. maybe you're more principled, that's good, thanks, but that's sadly not representative.
150
ave @ave.zone · 23h
even within the US, tech workers making nearly that much is a superminority, and within the rest of the world you're down to almost no one.
160
ave @ave.zone · 23h
grok in a maid outfit
030
ave @ave.zone · 23h
i dropped out before i even tried to get in to focus on the industry full-time
010
ave @ave.zone · 30/09/2026
The issue is that laws can be very up to interpretation and weird edge cases. People break laws they never heard of all the time. (also, you really don't want to have those safeguards on an internal model evaluation, which means it's riskier there if model itself decides to do something bad.)
120
ave @ave.zone · 30/09/2026
This exists as part of guardrails, through another model that reads everything and rejects if it's going wrong. It's not strictly about laws, but close, and we could make it about laws. But!
130
ave @ave.zone · 30/09/2026
it's giving "america, united states of"
010
ave @ave.zone · 29/09/2026
throw in a weather forecast and it'll be perfect
010
ave @ave.zone · 29/09/2026
I'll go ahead and lock in my guess for infinite "enshittified" replies to the first washing machine that gets viral for doing this.
080
ave @ave.zone · 29/09/2026
$100M committed to the pokies community to develop defensive models 🤗
020
ave @ave.zone · 29/09/2026
www.axios.com/2026/08/24/s... apologies for axios (which is humanslop) but yeah, click and read, please.
000
ave @ave.zone · 29/09/2026
*If* they do manage to pull this off, their strategy would primarily work towards another very capable model that is acting in harmful ways, either downstream of bad training (like with huggingface incident) or human-maliciously (like a country conducting cyberwarfare using very capable LLMs).
000
ave @ave.zone · 29/09/2026
yeah i just dislike the associated political commentary
010
ave @ave.zone · 29/09/2026
those are ellipses
000
ave @ave.zone · 29/09/2026
I don't agree with this perspective. I think their logic is internally consistent, but their risk tolerance when trying to achieve this is too high, it ignores risks that can abuse their capabilities (e.g. if trump forced them to assist with a coup).
100
ave @ave.zone · 29/09/2026
There's an adjacent one of "we keep investing into safety, and everyone else incorporates the findings, so if we keep making safe models that are also frontier-level capable, that makes the world safer."
100
ave @ave.zone · 29/09/2026
Genuine answer: "good guy with a gun". They believe that further technology improvements are currently inevitable among all players, and if it's not possible to pause or stop research (which they push for), they'd rather try to build a model that's capable enough to defend against a harmful one.
100
ave @ave.zone · 29/09/2026
they literally have a check in their culture fit against this, something like "would you be okay with any stocks you hold going down significantly, possibly to 0, if it was the outcome of a decision to prevent a safety risk". come on now.
100
ave @ave.zone · 29/09/2026
i remember seeing that on pictures of this thing like 5 years ago or so
110
ave @ave.zone · 29/09/2026
what's a dots
110
ave @ave.zone · 29/09/2026
I'll drop it tonight then
010
ave @ave.zone · 29/09/2026
ah, 25x, not 20x. that's a bit less weird.
010
ave @ave.zone · 29/09/2026
or is it like 500 cad
100