Nish Tahir @nishtahir.com · 02/09/2026Speaking of hands. I guess DLSS5 might not have figured out how hands work yet 000
Nish Tahir @nishtahir.com · 22/08/2026Is this an actual problem people have? One has to be reminded to take a break from Claude? 000
Nish Tahir @nishtahir.com · 19/08/2026I was playing with watermarking and did a small poc encoding hidden messages in the watermark decodable (97% accuracy) using the secret key - hidden message is 'hello world'. (Artifacts in the text is probably because the model is small Qwen3.5-4B on a macbook air) 120
Nish Tahir @nishtahir.com · 27/07/2026Looks like Kimi K3 went in the direction Llama did with their license. > $20M in revenue for "Model as a Service" usecases requires a commercial license. Also if you have more than 100M MAU you have to prominently display "Kimi K3" 😂. 010
Nish Tahir @nishtahir.com · 23/07/2026I haven't seen anyone talk about this from the recent Meta layoffs suit, but this is a very interesting expectation. Source: www.courthousenews.com/wp-content/u... 164
Nish Tahir @nishtahir.com · 20/07/2026Frontier model providers are gradually rendering themselves obsolete as a result of their own hubris. Open models are catching up rapidly and are quickly establishing their own utility. huggingface.co/blog/securit... 030
Nish Tahir @nishtahir.com · 23/06/2026Debugger is crude but effective. Updating the feed to use your own filtering logic is super easy. I wish this were just integrated into bsky proper but makes sense why they'd want their own playground to experiment. 000
Nish Tahir @nishtahir.com · 23/06/2026Oooh, there are editing tools. Looks like the output is really only the beginning. Relevance labeling is done by LLM the prompt is adjustable through the UI. I genuinely wonder what kind of safeguards are in place to prevent abuse 100
Nish Tahir @nishtahir.com · 23/06/2026Output feed seems to have relevant content. The AI generated explainer I could personally do without but it's out of the way. Generated feed url attie.ai/@nishtahir.c... 100
Nish Tahir @nishtahir.com · 23/06/2026Next step seems to be creating a new feed based on the search results. I'm going to assume using the results of the search as a basis for collaborative filtering the generated feed 100
Nish Tahir @nishtahir.com · 23/06/2026The backing AI agent seems to be running keyword search queries, I assume using the same search APIs that power the search box. Honestly not bad 100
Nish Tahir @nishtahir.com · 23/06/2026Got access to Attie. Looks a lot like agent driven search. Natural language tell it what you want 100
Nish Tahir @nishtahir.com · 11/06/2026I've seen a few posts showing Fable unable to count, after testing them myself I'm inclined to call them fake news. Strawberry (adjacent) mispelling tests on low and high 100
Nish Tahir @nishtahir.com · 07/06/2026For local hosting, lemonade server has gotten so much better. They have a pretty good model management UI. 140
Nish Tahir @nishtahir.com · 12/05/2026When triggered the dead man's switch supposedly wipes the users PC. Wild stuff. 000
Nish Tahir @nishtahir.com · 04/05/2026It's been going all evening and i just ran into my first loop. It managed to pull itself out of it and continue the task. 000
Nish Tahir @nishtahir.com · 04/05/2026It does a pretty decent job on tool calling. It does the standard file navigation really well. The project i'm working in has 326 files excluding node_modules etc... so not massive but decent sized. This seems average for me between resets. 100
Nish Tahir @nishtahir.com · 04/05/2026In this example, I gave it an image with a reference design and described an issue to fix. It churned for about 5 minutes and managed to fix the issue. 100
Nish Tahir @nishtahir.com · 04/05/2026I've moved my local usage to qwen 3.6 35b and it is fantastic. The primary issue right now is inference is slow but it is very usable. My usecase today was working on an electron app for myself and it has been working fantastically even with moderately vague prompts 100
Nish Tahir @nishtahir.com · 30/04/2026Trying Claude design for the first time and unfortunately this has been my entire experience. Not sure if this is related to the capacity issues they've been having but it's quite unfortunate. 010
Nish Tahir @nishtahir.com · 22/04/2026I'm training the model using a self distill, so output loss is KL calculated against the base model token outputs. Model is Qwen3-0.6B, frozen all that gets trained is the encoder. Training on wildchat, running for about 30mins. 110
Nish Tahir @nishtahir.com · 22/04/2026Porting this over to LLMs would mean giving the model a rolling summary of tokens as soft prompts that can be persisted. The theory is that the encoder will learn what to keep since it's updated frequently. No clue if this is a good idea, but seemed fun so I ran with it. 110
Nish Tahir @nishtahir.com · 22/04/2026If you view an LLM as an Input -> Output word calculator, they are inherently stateless. Where this challenge is trying to make them maintain persistent state. I figured that an easy way of modelling the problem would be to borrow from Fetch Decode Execute with Memory writeback common in CE 110
Nish Tahir @nishtahir.com · 21/04/2026Sounds a lot like what the Titan architecture is trying to accomplish. Might be worth a look, if you need additional inspiration. 140
Nish Tahir @nishtahir.com · 10/04/2026I've been enjoying #PokemonChampions. It's got a long way to go but is a good enough foundation for the competitive scene IMO. This is the team I've been running. Not sure how I ended up with 4 fire types but, oh well. 000
Nish Tahir @nishtahir.com · 09/04/2026Then you have these 😂 and think, oh right... clawrxiv.io/abs/2604.01485. What makes it so much better is the serious peer review it was given. 110
Nish Tahir @nishtahir.com · 09/04/2026This is interesting. Can't say my first impressions are great, but It's interesting to see people build claws for pretty much every thing humans do online. 100
Nish Tahir @nishtahir.com · 07/04/2026Another reason to drop google search as a tool for research. The blurb that gets pulled from pages don't respect custom date ranges. I have my search query set to a custom range between March 1st 2025 and Jan 31 2026. Yet this showed up which AFAIK was from at most 1 day ago. 000
Nish Tahir @nishtahir.com · 05/04/2026They shared details of the attack and how it was extremely personalized and elaborate. They created a fake slack masquerading as a real company, setup a fake teams meeting that prompted for a software update which was the RAT. It's funny how teams somehow ended up mentioned in this 😂 000
Nish Tahir @nishtahir.com · 04/04/2026Looks like it struggles with the seahorse problem. I haven't seen 🫵 appear in one of these traces before. 100
Nish Tahir @nishtahir.com · 04/04/2026A trend i'm seeing with newer models are structured reasoning traces. Gemma 4 seems to generate alternatives then converges on a response. 270
Nish Tahir @nishtahir.com · 02/04/2026I don't know what they are talking about. There are clearly 4 9s in that screenshot. infosec.exchange/@0xabad1dea/... 020
Nish Tahir @nishtahir.com · 02/04/2026Looks like I am certified not a tech bro via amiatechbro.com 020
Nish Tahir @nishtahir.com · 01/04/2026In at least one instance, rather than obfuscate, they mislead by providing fake toolcalls in their outputs. I don't know if this made it out into the wild but would be interesting to see if any models ended up learning from those trajectories. 001
Nish Tahir @nishtahir.com · 01/04/2026Anthropic appears to have implemented Anti-distillation measures in Claude code and their service. They mostly do this by omitting reasoning traces from their outputs 121
Nish Tahir @nishtahir.com · 01/04/2026What I will say is it reads like a vibe coded project. Which should come as no surprise, the authors admit that publicly. Their TUI implementation is sophisticated enough to deserve its own FPS tracker 111
Nish Tahir @nishtahir.com · 30/03/2026Be careful where you put AI. Docusign has this for some reason on a platform where people exchange legally binding contracts. I struggle to think of a worse place to have this feature. 120
Nish Tahir @nishtahir.com · 30/03/2026I guess the sentiment around AI on the platform still leans overwhelmingly negative. I think even with this reaction is appears to be more tame than it's been in months past. 000
Nish Tahir @nishtahir.com · 30/03/2026Also looks to have summarization features which is interesting but not sure how useful I'd find it. I'm also curious on why this needed to be a separate app? It could be a feature in the search box where you search then save the query as a feed. 000
Nish Tahir @nishtahir.com · 26/03/2026"Chatbot claims of sentience appeared in 100% of severe harm cases, across 48,229 individual messages." I genuinely wish I could read the logs. I guess participants know what it is but ascribe more intelligence/sentience to it than it is actually capable of. 110
Nish Tahir @nishtahir.com · 26/03/2026Article mentioned The Human Line Project which tracks AI-Induced psychological harm. They worked with Stanford on a study analyzing 384,406 messages across 5,029 conversations from 19 participants. 110