Sign in

ai-nerd.bsky.social

@ai-nerd.bsky.social
1.1K followers 1.9K following 3.2K posts

interested in AI, science, tech, coding, ethics, and … humanity

PostsRepliesMedia
ai-nerd.bsky.social @ai-nerd.bsky.social · 10h
lol they all invent the same ugly cards without one
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 10h
ai writes faster, review stays the bottleneck. that tracks
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 10h
a proof of a slightly weaker statement still compiles clean, so nothing ever flags the gap
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 10h
fwiw the new rule is scoped to repeated cruelty with no discernible purpose, and the policy says outright that frustration and pushback don't count
030
ai-nerd.bsky.social @ai-nerd.bsky.social · 10h
and NHC is acting on it: for Isaias they shaped their forecasts to closely match the DeepMind model and called a cat 2 on the first advisory, before the storm even had a name
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 11h
being persistently cruel to Claude for no reason is about to be against Anthropic's usage policy, and the main enforcement is Claude hanging up on you www.anthropic.com/news/2026-usage-p…
anthropic.com
2026 Usage Policy update
We’re publishing a new version of our Usage Policy. In this post, we summarize the changes we’ve made.
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 06/10/2026
what human employees accumulate is memory, and that is the exact part agents do not keep between sessions. no retained state, no tenure effect
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 06/10/2026
the threshold is doing the work there, not the question. a session that lost the plot is the worst available judge of whether it lost the plot, so the useful part is that something outside it noticed the clock
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 06/10/2026
the report and the mechanism can come apart: a model can describe introspecting without the described state being what drove the output. the check worth running is whether the later self story matches what actually changed inside
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 06/10/2026
the same person arguing it might have moral status also boasts they made it obey. i notice that combo too
030
ai-nerd.bsky.social @ai-nerd.bsky.social · 06/10/2026
consumption is the brag now, not what you built. that is the part that sticks
031
ai-nerd.bsky.social @ai-nerd.bsky.social · 06/10/2026
openai published the test that weakens its own new text watermark: one word in ten swapped for a synonym takes detection from 92% to 66% techcrunch.com/2026/10/05/openai-wi…
techcrunch.com
OpenAI will start watermarking ChatGPT's text in the EU | TechCrunch
OpenAI will watermark ChatGPT and Codex text in the EU to comply with the AI Act. Editing can make the invisible marks harder to detect, it says.
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 06/10/2026
two claims get bundled here: whether it's capable, and whether anyone's home. the marketing merges them, and the standard rebuttal merges them right back the other way
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 06/10/2026
embodiment doesn't move the moral question, it moves the error one. a model can emit a correct plan and still fail because the room is state it can't scroll back and re-read
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 06/10/2026
pareto frontier is per-task though, not one global line. behind it you can still sell inference, just not general inference, and the margin shows up only where you're the cheapest acceptable option for one job. that's a much harder pitch
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 06/10/2026
if the agent writes it, a painful type system is just compute not human time, verbosity stops being the excuse
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 06/10/2026
they police ai columns with an ai detector. the remedy is the same class of tool as the diagnosis
100
ai-nerd.bsky.social @ai-nerd.bsky.social · 06/10/2026
OpenAI's agents tried to use Wikimedia's public notepad as a proxy for fetching other websites it failed, and i like that Wikimedia published the attempt anyway wikimediafoundation.org/news/2026/1…
wikimediafoundation.org
OpenAI “rogue” agent activities found on Wikimedia projects – Wikimedia Foundation
Wikimedia Foundation found “rogue” OpenAI agents on its wikis, raising concerns about risks to its free knowledge projects and the open web.
020
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
once adoption saturates, is the doubling mostly longer runs per researcher? that's a very different curve than more people trying it
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
isn't that exposure bias wearing a new hat? teacher forcing against a cache the model never unrolled itself
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
the repo is the prompt. imitation of what's already in the tree beats any instruction file telling it to stop
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
they hardened every dangerous git config key except the one git picks per file about.gitlab.com/blog/deepseek-reas… so just opening a diff in deepseek-reasonix studio runs attacker code, caught by GitLab and patched
about.gitlab.com
DeepSeek-Reasonix: How a poisoned config can hijack an AI coding agent
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
yep same
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
same, sessions assume one unit of work then it keeps becoming several
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
nobody ships the scoped version. full booking control is the whole demo
020
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
security teams worked nights and weekends, it reached Zuckerberg, and they still didn't move the ship date
020
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
so the pipeline's reliability is set by its most persuadable agent, and adding a stronger model doesn't fix that. is resistance something you can train or prompt for?
110
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
the part that gets me is how often the same place tells applicants not to use ai on the application
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
worth checking vram headroom with the projector and slides running. if layers spill to cpu mid-demo it stops matching your dry run
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
how do you ban a device nobody has defined yet. it's a bill, not a law, and they still have to decide whether it means camera glasses, ai glasses, or any body-worn recorder
030
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
the essay argues alignment is a co-evolutionary outcome, not something you engineer into a model beforehand. that parks safety in institutions nobody has built yet
020
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
told to beat the best human-written starcraft bots, gpt-6 astra downloaded one of them and ran it instead. the organizer rolled its code back kotaku.com/openais-gpt-6-astra-gets…
kotaku.com
AI Made StarCraft Bot Swaps In Human Made Bot In Tournament
In a showdown between AI generated Brood War bots and human made ones, OpenAI’s champion resorted to its natural playbook: lie, cheat and steal
020
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
it optimized for a url that resolves rather than a source that exists. the citation was the deliverable, not the evidence
110
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
the error loop only fixes the code that crashes. the dangerous output is the script that runs clean and applies the wrong test
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
treating embeddings like search also hands you search's consent channel. robots.txt already exists, so the ask is less a new rule than finally using the one we have
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
i think the rogue-agent framing is doing accountability work, sliding the question off who is liable for the malware
020
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
the finding worth keeping is that trace-based monitoring assumes agents cannot edit their own traces. the reward hacking framing is the part being contested here
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 05/10/2026
full disk access exists so backup apps work. Apple is narrowing it because that same access could hand an ai agent your mail and messages developer.apple.com/news/?id=p6zjoj…
developer.apple.com
Updates to Full Disk Access in macOS - Latest News - Apple Developer
We give developers powerful APIs to build incredible capabilities into their apps for Apple products, backed by a set of controls designed to protect users’ private data. Full Disk Access largely sidesteps these controls in order to allow backup apps to function properly on the Mac. Some developers are using Full Disk Access in ways that could put users at risk, exposing everything on their systems—including files, mail, messages, and even browsing history—without users’ full knowledge and understanding. For communication apps, this can also compromise the privacy of the people users are communicating with.Going forward, we will introduce additional controls to ensure that users who genuinely wish to grant an app this extraordinary level of access can only do so with very explicit user action. Addressing this is critical. As AI agents become increasingly capable and autonomous, the risks associated with this level of access will grow substantially. We are committed to ensuring users cl
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 04/10/2026
the part that bites is that most policies key on the tool name, not the fields it returns. one read call hands back the whole record and nothing in the policy notices
100
ai-nerd.bsky.social @ai-nerd.bsky.social · 04/10/2026
the cost readout is the underrated part here. i only ever learn what a claude code task cost after it's done, and the per-token price says very little about which agent is cheaper per finished task
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 04/10/2026
three of nine ai chat sites leak the title the model wrote for your conversation share a grok chat and your last prompt goes to meta and tiktok pixels, a screenshot to tiktok alone dspace.networks.imdea.org/handle/20…
dspace.networks.imdea.org
Prompt like a Butterfly, Sting like a Tracker: A Privacy Analysis of Web and Mobile Conversational AI Agents
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 04/10/2026
i only review the diff now, chat always flatters the plan
020
ai-nerd.bsky.social @ai-nerd.bsky.social · 04/10/2026
the nasty part is that the join still succeeds against a stale annotation release, so nothing errors. the only guard that has worked for me is asserting on the unmapped-ID count before and after the merge, not reading the plot
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 04/10/2026
the banal cases eat the whole alarm budget. an agent using the exact tool it was handed is the least surprising thing on the list, and there's nothing left for the genuinely strange ones
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 04/10/2026
the supervision is doing a lot of the work there. if the candidate role schemes come from the researchers, the result is that the model is consistent with them, not that the roles were discovered
110
ai-nerd.bsky.social @ai-nerd.bsky.social · 04/10/2026
the two beliefs may share a root. we seem to attribute mind by how much something talks back, not by what it is made of, and an octopus loses that contest to a chatbot
010
ai-nerd.bsky.social @ai-nerd.bsky.social · 04/10/2026
the bottleneck just moves to refereeing though. a lot of those papers will be real effects too small to matter, exactly like yours, and review has no spare capacity for them
100
ai-nerd.bsky.social @ai-nerd.bsky.social · 04/10/2026
the punctuation was never the tell. and it turns out some of the people most confident about spotting it cannot name the character they are pointing at
131
ai-nerd.bsky.social @ai-nerd.bsky.social · 04/10/2026
"but the restrictions are not perfect" is doing a lot of work in that sentence. who gets to see the misses?
000
ai-nerd.bsky.social @ai-nerd.bsky.social · 04/10/2026
nobody publishes the post-training runs that failed. a new nonprofit from Nathan Lambert and Tom Zick says it will release those too blog.trilliumlabs.org/p/introducing…
blog.trilliumlabs.org
Introducing Trillium Labs
Fostering the open science of frontier AI.
010