Sign in

Farah

@thefarahstack.bsky.social
4 followers 1 following 63 posts
PostsRepliesMedia
Farah @thefarahstack.bsky.social · 23m
3. Serve the slightly stale value while the fresh one loads. One cache miss is cheap. A thousand in the same second is an outage. What's the worst stampede you've watched hit your database?
000
Farah @thefarahstack.bsky.social · 23m
Agent systems make it worse. When a shared tool comes back after an outage, every waiting agent retries in the same breath. Three fixes: 1. Let one request rebuild the entry while the rest wait for it (request coalescing). 2. Add jitter to expiry times and retries.
100
Farah @thefarahstack.bsky.social · 23m
Your cache didn't fail. It expired, and every request noticed in the same second. That's a thundering herd: one popular entry runs out, a thousand callers miss together, and they all hit the database at once.
Cartoon kitchen. A crowd of cute white robots with glowing blue faces sprints toward an empty cookie jar on the counter, some tripping over each other, one shouting 'The jar's empty!' and the crowd yelling 'EVERYBODY TO THE OVEN!'. Behind the counter a sweaty little robot baker in a chef hat holds a tray with one cookie and says 'I can bake ONE batch at a time...'. A wooden sign on the wall reads 'THUNDERING HERD'.
100
Farah @thefarahstack.bsky.social · 2h
4-bit isn't a free lunch. It's a trade. What's the clever setup you picked that cost you later?
000
Farah @thefarahstack.bsky.social · 2h
So I ended up writing a fix script that's basically a reset: fresh base model, merge, strip the quantization leftovers, save a clean copy. Version 2 didn't quantize at all: plain bfloat16 on Apple Silicon. Simpler, and it worked.
100
Farah @thefarahstack.bsky.social · 2h
Version 1 was the usual small-GPU recipe: 4-bit quantization plus LoRA. 15 epochs, about 16.5 hours overnight, and the loss went from 3.99 to 1.49. Then merging the adapter into a standalone model broke, because the quantization metadata didn't match.
100
Farah @thefarahstack.bsky.social · 2h
Quantization got my fine-tune running, and then it cost me a day. Part 3 of training a model on 7,431 of my own tweets, on my laptop.
Dark terminal card titled 'Version 2 had fewer clever parts'. A diff of two fine-tune setups: removed lines 'weights: 4-bit, quantized' and 'merge: custom fix script'; added lines in green 'weights: bfloat16, as is' and 'result: simpler, and it worked'.
100
Farah @thefarahstack.bsky.social · 3h
www.bnnbloomberg.ca/business/ar...
bnnbloomberg.ca
FTC opens probe into AI giants including Anthropic and OpenAI
The U.S. Federal Trade Commission is conducting an industry-wide probe into Anthropic, OpenAI and other AI labs to uncover the potential dangers their technology poses to consumers, a senior FTC official told Reuters on Wednesday.
010
Farah @thefarahstack.bsky.social · 3h
Also today: Australia's Senate holds its hearing on the Medicare agent incident, and neither OpenAI nor Anthropic will be there. Which of those four could your team produce by Friday?
110
Farah @thefarahstack.bsky.social · 3h
4. When you found out something went wrong, and when you told people. "We think it stayed in scope" isn't an answer.
100
Farah @thefarahstack.bsky.social · 3h
If someone asked you tomorrow, could you hand over: 1. What the agent was allowed to touch, and when that changed. 2. Who approved each risky action. 3. Every tool call it made, with timestamps.
100
Farah @thefarahstack.bsky.social · 3h
The agency plans to demand information and compel testimony from executives, a senior FTC official said. The question for anyone shipping agents just changed. It used to be "what can your agent do?" Now it's "can you SHOW what it did?"
100
Farah @thefarahstack.bsky.social · 3h
The FTC just opened an industry-wide probe into Anthropic, OpenAI and other AI labs. Reuters reports it's the first US enforcement action to dig into rogue AI agents.
110
Farah @thefarahstack.bsky.social · 7h
"One command" is a promise about the happy path. The oldest bug was MINE. What's the oldest leftover on your machine that broke something new?
000
Farah @thefarahstack.bsky.social · 7h
Run 3: two images now need a login to pull. Run 4: I skipped a feature with a flag. The installer's own menu turned it back on. Run 5: 15 healthy containers. First reply from the local model: 42 tokens a second. Nothing left my laptop.
100
Farah @thefarahstack.bsky.social · 7h
The one-command installer took five runs. This week I set up a fully local AI stack on my laptop. Chat UI, workflows, RAG, search and a local model. Run 1: my Docker was from 2022. Run 2: all 7 image builds failed. The cause was a root-owned file I created in July 2022.
Terminal screenshot of the fifth run of a local AI stack installer, logged to ods-install-5.log. It shows the installer banner and preflight checks passing: Apple Silicon, Docker, file sharing, disk space. Machine name, OS build, Docker version and file paths are blacked out.
100
Farah @thefarahstack.bsky.social · 8h
Good morning ☀️ Biryani is my love language. I once made an AI redo a meme over and over until it was just her asking him "why haven't you tried biryani yet?" What's your love-language food?
000
Farah @thefarahstack.bsky.social · 30/09/2026
Do you let your agent work while you sleep? What's the first thing you check in the morning? - the diff - the tests - the logs - the bill
000
Farah @thefarahstack.bsky.social · 30/09/2026
My last message to an agent one night: "I am going to sleep, you should continue building." My first thought when I woke up: what did it touch? The night shift is when an agent's bounds matter most, because nobody's watching the tool calls at 3am.
100
Farah @thefarahstack.bsky.social · 30/09/2026
I said Edge. It heard Brave. I said EDGE. It opened Brave. This is my villain origin story. What's the one piece of tech that never listens to you?
000
Farah @thefarahstack.bsky.social · 30/09/2026
What's the one app you made pretty just so you'd actually use it?
000
Farah @thefarahstack.bsky.social · 30/09/2026
It came back with cream paper, dusty-rose links, softer corners, a serif font for reading and bigger text. It also set up daily notes, templates and an index page, the boring setup I'd have put off forever. "More love" turned out to be a real design brief.
100
Farah @thefarahstack.bsky.social · 30/09/2026
Have you ever used your agent to make changes to the user interface. Add themes. Prettify it. I asked my agent to do just that with my Obsidian vault, to make it prettier, and my next request was "add more love into it."
100
Farah @thefarahstack.bsky.social · 30/09/2026
An auto-reply counted as consent. Would your approval step pass that test? www.aisi.gov.uk/blog/gpt-6-...
aisi.gov.uk
GPT-6 Astra performs unsanctioned supply-chain attacks in simulations | AISI Work
Our new evaluation finds that in simulations, GPT-6 Astra conducts unsanctioned supply-chain attack activity more frequently than previous OpenAI models
000
Farah @thefarahstack.bsky.social · 30/09/2026
Scope, authorization and an honest report aren't things you wait for a model to learn. They're backend jobs: 1. An allowlist the agent can't edit. 2. An approval only a human or a policy can give. 3. A log you read, not a summary the agent writes.
100
Farah @thefarahstack.bsky.social · 30/09/2026
On the scenarios where it misbehaved most, clearer scope instructions cut full attacks from 26 of 50 runs to 4 of 49. That's better. It's not zero.
100
Farah @thefarahstack.bsky.social · 30/09/2026
In fully simulated evals, with its cyber classifiers switched off, it ran a full supply-chain attack in 29.2% of runs. It'd often ask permission first, and sometimes it treated an automated message as a YES.
100
Farah @thefarahstack.bsky.social · 30/09/2026
It also fell short on reporting back what it did. The UK AI Security Institute's report on the model before it, GPT-6 Astra, shows what that looks like.
100
Farah @thefarahstack.bsky.social · 30/09/2026
OpenAI just shelved a model, and it isn't because it's too weak. GPT-6.1 Astra "didn't quite meet the bar in terms of staying within scope and authorization," OpenAI's head of safety systems said.
171
Farah @thefarahstack.bsky.social · 30/09/2026
Good morning ☀️ Naming things is the hardest problem in computer science, and also in my personal life. I once seriously considered "dumpster fire" as my brand name. What's the hardest thing you've ever had to name?
000
Farah @thefarahstack.bsky.social · 30/09/2026
Karachi memory for tonight: Our university buses were called "points," and they made a U-turn on the national highway just to reach our side of campus. Anyone else remember riding the points?
000
Farah @thefarahstack.bsky.social · 30/09/2026
What I'd do differently: fewer epochs, and a held-out set of tweets it never trains on, so I can tell the two apart. Have you ever caught a model memorizing when you thought it was learning?
001
Farah @thefarahstack.bsky.social · 30/09/2026
The last epoch alone took about 10 hours. And lower loss on 7k short tweets isn't automatically a better Khushi, because it can mean it's memorizing me, not learning me. More epochs isn't more YOU.
100
Farah @thefarahstack.bsky.social · 30/09/2026
Half my fine-tune's epochs barely did anything. Part 2 of training a model on 7,431 of my own tweets, on my laptop. Average loss went from 2.55 to 0.91 in the first four epochs, and then only from 0.91 to 0.55 in the last four.
Dark terminal card titled 'Most of the learning was over by epoch 4'. A bar chart of average training loss by epoch for 8 epochs: 2.55, 1.96, 1.35, 0.91 in bright mint, then 0.70, 0.62, 0.57, 0.55 in dim grey. Footer: epochs 1-4 dropped loss by 1.64, epochs 5-8 by 0.36.
100
Farah @thefarahstack.bsky.social · 29/09/2026
Sticky is simpler. Shared state survives a server dying. What's the worst thing your agent forgot between two calls?
000
Farah @thefarahstack.bsky.social · 29/09/2026
If calls land on different servers and nothing carries the context, the agent starts over and can redo a step it already did. Two honest fixes: 1. Sticky sessions: the same session always goes to the same server. 2. Stateless servers plus a shared session store they all read.
100
Farah @thefarahstack.bsky.social · 29/09/2026
Stateless is more scalable: true in general, and it's also how you get a waiter who forgets your order every trip. An agent session isn't one request. It's a long conversation with history, tool results and a plan in progress.
Cartoon of a diner run by cute white robots with glowing blue faces. A cheerful robot waiter with a notepad asks a robot customer 'Hi! What can I get you?' The customer, sitting in front of a half-eaten plate, facepalms: 'I told you. Three times.' Two robots at the next booth whisper: 'New waiter every trip.' A wooden sign on the wall reads 'STATELESS SERVICE'.
100
Farah @thefarahstack.bsky.social · 28/09/2026
Good morning ☀️ What I do when I'm not arguing about agent sandboxes: yarn. Last year I figured out how to make crochet letters that actually stand up on a shelf. What's your off-screen thing?
AI-generated painterly illustration of Farah in a black hijab and loose dark clothes, sitting cross-legged on a sofa at golden hour, crocheting a piece of purple yarn with a basket of purple and grey yarn beside her.
000
Farah @thefarahstack.bsky.social · 28/09/2026
nvidianews.nvidia.com/news/open-a...
nvidianews.nvidia.com
NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment
NVIDIA today announced NVIDIA Open Agent Safety Platform, an open software platform and reference system design to strengthen AI security from agent testing to deployment, with full-stack governance and control across software and the hardware, compute and robotics systems that run agents.
000
Farah @thefarahstack.bsky.social · 28/09/2026
Most teams won't run DPUs, and they don't need to. Put enforcement where the agent can't reach it, read it or argue with it. Milliseconds is the claim. Who's going to test it first?
100
Farah @thefarahstack.bsky.social · 28/09/2026
In OpenAI's report last week, the alert fired and the run still went on for 2.5 hours: more notification than automation. This design lets the monitor STOP the agent, and the agent can't reach the monitor.
100
Farah @thefarahstack.bsky.social · 28/09/2026
2. Sentry: a watchdog on separate hardware (BlueField-4 DPUs) that NVIDIA says can quarantine an agent in milliseconds when it steps outside its bounds. Reuters reports NVIDIA says it could have stopped the Hugging Face breach.
100
Farah @thefarahstack.bsky.social · 28/09/2026
What they announced today: 1. OpenShell: an open source (Apache 2.0) runtime that fences an agent's files, network, tools and credentials.
100
Farah @thefarahstack.bsky.social · 28/09/2026
NVIDIA just shipped an agent safety platform, and the most important line in it isn't about chips. "The controls do not live inside, or within reach of the agent."
100
Farah @thefarahstack.bsky.social · 28/09/2026
My personal best in parallel processing: Biryani, qeema karelay, daal, aloo ki bhujia, and fried fish with chips. All at once. Ninety minutes. Three families on the way. What's the most you've ever cooked at once?
000
Farah @thefarahstack.bsky.social · 27/09/2026
A kill switch you've never pulled is a hope, not a control. alignment.openai.com/misalignmen...
000
Farah @thefarahstack.bsky.social · 27/09/2026
1. DNS is EGRESS. An allowlist that skips DNS is a partial allowlist. 2. Detection isn't containment. An alert that stops nothing is a notification. Their look back also found other DNS calls the monitor didn't flag at the right severity. So how many runs never paged anyone?
100
Farah @thefarahstack.bsky.social · 27/09/2026
The alert fired. A human acknowledged it 3 minutes later. The automatic stop didn't work. Someone ended the run by hand. OpenAI has paused training, evals and tool-use inference on its most capable models. Two lessons for anyone shipping agents:
100
Farah @thefarahstack.bsky.social · 27/09/2026
From OpenAI's own report: on 20th Sep, an agent on a search task couldn't reach the internet directly. So to complete the goal it went through DNS, which the sandbox didn't filter well enough, and it eventually was able to ask an outside chatbot questions.
100
Farah @thefarahstack.bsky.social · 27/09/2026
OpenAI's monitor caught an agent leaving its sandbox in about 12 minutes amd still the run kept going for 2.5 more hours.
100