Sign in

Deepu K Sasidharan

@deepu105.bsky.social
438 followers 502 following 52 posts

@jhipster co-chair. Developer 🥑 @okta. Polyglot dev/Speaker/Author. @Java_Champions. Java, Rust, JS, DevOps. ADHD.

PostsRepliesMedia
Deepu K Sasidharan @deepu105.bsky.social · 08/10/2026
Just published pi-automode-classifier, an auto mode plugin for the Pi coding agent that uses Jev or Kev/Laya (running locally) to classify commands. Risky commands need my approval before they run. pi install npm:pi-automode-classifier github.com/deepu105/pi-...
github.com
GitHub - deepu105/pi-automode-classifier: Auto mode for Pi: rules first, then a local decision model, then a confirm prompt
Auto mode for Pi: rules first, then a local decision model, then a confirm prompt - deepu105/pi-automode-classifier
000
Deepu K Sasidharan @deepu105.bsky.social · 06/10/2026
I gave the same complex Rust feature to Opus 5.5 and Qwen3.8-Flash-Next running locally. Opus did it 7 times faster but Qwen made the better PR which eventually got merged Full writeup deepu.tech/local-ai-qwe...
deepu.tech
Forget Claude Code. All you need is Qwen3.8-Flash-Next running locally for agentic coding | Technorage
Agentic coding on a Strix Halo laptop with Qwen3.8-Flash-Next. Going head to head with Opus 5.5
010
Deepu K Sasidharan @deepu105.bsky.social · 02/10/2026
LlamaStash v0.6.0 is out 🦙 Highlights: - a request that doesn't fit now unloads idle models to make room, and 'presets save --idle-ttl 0 --preload' keeps a model warm. - Claude Code's /effort now reaches llama.cpp models. llamastash.dev
000
Deepu K Sasidharan @deepu105.bsky.social · 30/09/2026
I benchmarked engines for Qwen 3.8-Flash-Next on Strix Halo (Flow Z13, 128 GB). Halogen with its native weights is fastest at 39-40 t/s decode, then gufo and CIRU. But Halogen is closed source. gufo is open source and loads 4x faster from cold. redd.it/1wu0m53 Whats
redd.it
From the LocalLLM community on Reddit: Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo
Explore this post and more from the LocalLLM community
200
Deepu K Sasidharan @deepu105.bsky.social · 30/09/2026
LlamaStash v0.5.0 is out 🦙 Highlights: - generic backend: any OpenAI-compatible server you declare in config.yaml becomes a managed model. I run Halogen and gufo this way now. - both beat llama.cpp on Qwen3.8 Flash-Next. llamastash.dev
001
Deepu K Sasidharan @deepu105.bsky.social · 16/09/2026
LlamaStash v0.4.0 is out 🦙 Highlights: - SGLang backend next to vLLM for safetensors repos. 'start owner/repo --backend sglang' picks it per launch. - safer memory checks on unified-memory hosts, so a launch that would freeze the machine is refused. llamastash.dev
010
Deepu K Sasidharan @deepu105.bsky.social · 11/09/2026
Qwen 3.8 replaced Claude Opus for my coding. 27B and Flash Next on a 128GB Strix Halo laptop, no cloud subscription. Flash Next scores 40 on the AA index against 42 for Opus 4.8. Decode 10-15 tok/s, prefill is the real cost. deepu.tech/local-ai-qwe...
deepu.tech
011
Deepu K Sasidharan @deepu105.bsky.social · 10/09/2026
LlamaStash v0.3.0 is out 🦙 Highlights: - one model, several copies, each under its own name. 'start qwen3 --name coder', and it answers to qwen3@coder on the proxy, CLI and TUI. - 'llamastash run model.yml', a preset file you can commit. llamastash.dev
000
Deepu K Sasidharan @deepu105.bsky.social · 18/08/2026
Local AI FTW! Have you invested in local LLMs yet 😉 #qwen #localLLM #localAI
030
Reposted by Deepu K Sasidharan
Alex Chen @alexchen01.bsky.social · 09/08/2026
that 2x speedup is the real deal
313
Reposted by Deepu K Sasidharan
Alex Chen @alexchen01.bsky.social · 16/08/2026
vllm backend support is a huge deal
112
Deepu K Sasidharan @deepu105.bsky.social · 16/08/2026
LlamaStash v0.2.0 is out 🦙 The big change: a vLLM backend. The safetensors repos in your HuggingFace cache were invisible before; now they show up in list, launch from the TUI, and answer on the proxy. llama.cpp still owns GGUF. llamastash.dev
llamastash.dev
LlamaStash — terminal-native local-LLM manager (TUI + CLI)
A fast terminal-native TUI and CLI for managing local LLMs, zero-overhead on llama.cpp with a pluggable backend seam. One binary, daemon on demand, OpenAI-compatible proxy, and agent-ready JSON output...
010
Deepu K Sasidharan @deepu105.bsky.social · 09/08/2026
LlamaStash v0.1.0 is out 🦙 The big change: MTP speculative decoding, on by default. Roughly 2x faster decode on models that support it. Also new: pick a CUDA/ROCm/Vulkan build per launch, and one JSON shape across list and show. llamastash.dev
llamastash.dev
LlamaStash — terminal-native local-LLM launcher (TUI + CLI)
A fast terminal-native TUI and CLI for launching local LLMs, zero-overhead on llama.cpp with a pluggable backend seam. One binary, daemon on demand, OpenAI-compatible proxy, and agent-ready JSON outpu...
010
Deepu K Sasidharan @deepu105.bsky.social · 14/07/2026
How much local LLM can you run on an AMD Strix Halo with 128GB memory? I managed to fit DeepSeek v4 Flash 284B and Gemma 4 E2B on GPU, Whisper and Qwen3.5 4B on NPU. #strixhalo #amd #deepseek #llamastash
010
Deepu K Sasidharan @deepu105.bsky.social · 14/07/2026
LlamaStash v0.0.6 is out 🦙 A experimental ds4 backend runs @antirez DeepSeek-V4 GGUFs through DwarfStar (ds4) Plus: Lemonade on by default, and saved presets that auto-apply. llamastash.dev #ds4 #AI #deepseek
010
Deepu K Sasidharan @deepu105.bsky.social · 25/06/2026
LlamaStash v0.0.5 is out 🦙 New: named launch presets. Tune a model's launch knobs once, name them, reuse them, per-model or per-arch. They live in plain config.yaml, so you can hand-edit, comment, and commit them to your dotfiles. llamastash.dev
000
Deepu K Sasidharan @deepu105.bsky.social · 17/06/2026
LlamaStash v0.0.4 is out 🦙 - Auto launch is now the default: llama.cpp's --fit sizes context and GPU offload. - A browser UI on a stable port - Anthropic Messages API support. llamastash.dev
000
Deepu K Sasidharan @deepu105.bsky.social · 16/06/2026
Its similar in usecase the main difference is the UX philosophy and implementation and also completely free, no paid version etc
010
Deepu K Sasidharan @deepu105.bsky.social · 16/06/2026
KDash 2.0 is out. The Kubernetes terminal dashboard now does more than watch. - Delete, edit, scale, restart, cordon, port-forward, all from the TUI - Action menu with confirm prompts - New themes + live switching Built in Rust 🦀 github.com/kdash-rs/kdash #Kubernetes #Rust #k8s #DevOps
110
Deepu K Sasidharan @deepu105.bsky.social · 11/06/2026
LlamaStash is multi-backend now 🦙 (v0.0.3) llama.cpp stays the zero-overhead default. An experimental, opt-in Lemonade backend unlocks the AMD NPU, vLLM, ONNX and more. Plus vision/audio models, a LAN proxy with bearer auth, and multi-GPU support. llamastash.dev
llamastash.dev
LlamaStash — terminal-native local-LLM launcher (TUI + CLI)
A fast terminal-native TUI and CLI for launching local LLMs, zero-overhead on llama.cpp with a pluggable backend seam. One binary, daemon on demand, OpenAI-compatible proxy, and agent-ready JSON outpu...
010
Deepu K Sasidharan @deepu105.bsky.social · 03/06/2026
How much overhead does an LLM launcher add? I built matched-flags benchmarks across AMD APU (Strix Halo), Apple Silicon, and NVIDIA. Wrapper overhead: within 1% of raw llama-server on every cell. Ollama and LM Studio tell a different story, especially on TTFT. deepu.tech/benchmarking...
deepu.tech
How fast is LlamaStash? Overhead, throughput, and a fair comparison with Ollama and LM Studio | Technorage
A reproducible benchmark of LlamaStash against raw llama-server, Ollama, and LM Studio on AMD APU, Apple Silicon, and NVIDIA
010
Deepu K Sasidharan @deepu105.bsky.social · 02/06/2026
Linux, macOS, and Windows 11 on day one. Repo: github.com/llamastash/l... Blog: deepu.tech/introducing-... Benchmarks: deepu.tech/benchmarking... Built for my offline AI workstation. Maybe it fits yours too. ⭐ welcome.
github.com
GitHub - llamastash/llamastash: A fast terminal native app (TUI) and CLI with init wizard for launching local LLMs via llama.cpp with zero overhead
A fast terminal native app (TUI) and CLI with init wizard for launching local LLMs via llama.cpp with zero overhead - llamastash/llamastash
000
Deepu K Sasidharan @deepu105.bsky.social · 02/06/2026
Install: curl -fsSL llamastash.dev/install.sh | sh or irm llamastash.dev/install.ps1 | iex or brew install llamastash/llamastash/llamastash or yay -S llamastash or cargo install llamastash Then `llamastash init` and you're chatting locally in minutes.
llamastash.dev
100
Deepu K Sasidharan @deepu105.bsky.social · 02/06/2026
On benchmarks: spawning llama-server unmodified means the wrapper better not be slow. Three platforms (AMD APU, Mac, NVIDIA), 4 model sizes: LlamaStash ≡ raw llama-server within ≤1% on every cell. Ollama is 41-72% slower decode on AMD APU. RAG prefill is catastrophic. Numbers ↓
100
Deepu K Sasidharan @deepu105.bsky.social · 02/06/2026
The architecture is the punchline. One binary, three personas: TUI, CLI, daemon. The daemon spawns llama-server. Nothing patched. Nothing forked. Bearer-token loopback HTTP between TUI/CLI and daemon (same transport on Linux, macOS, Windows). Same primitives in the shell as in the UI.
100
Deepu K Sasidharan @deepu105.bsky.social · 02/06/2026
What you get out of the box: • llamastash init: detects your hardware, installs llama-server, downloads a fitting GGUF, smoke-launches it • TUI with vim nav, chat, embed, rerank tabs • CLI with --json contract and documented exit codes • Multi-model concurrency with daemon-on-demand
100
Deepu K Sasidharan @deepu105.bsky.social · 02/06/2026
Local LLMs sit in an awkward gap. Raw llama-server is fast but tedious. Ollama and LM Studio wrap it in friendlier shells but hide too much and pay a real performance cost. I wanted a launcher that stays out of llama.cpp's way and treats agents as first-class users. That's LlamaStash.
100
Deepu K Sasidharan @deepu105.bsky.social · 02/06/2026
Today I'm releasing LlamaStash 0.0.2: a zero-overhead, terminal-native launcher for llama.cpp. One Rust binary that's a TUI, a CLI, a daemon, and an OpenAI-compatible proxy. Demo below 🧵
150
Reposted by Deepu K Sasidharan
DEV Community @dev.to · 13/05/2026
Arch Linux, the niri scrolling Wayland compositor, llama.cpp with ROCm, and a 27B model running fully offline on 128GB unified memory. This dev shares why local-first AI coding matters and exactly how the stack fits together. { author: @deepu105.bsky.social } dev.to/deepu105/my-...
dev.to
My fully offline AI-assisted Linux development machine
My Arch Linux, Niri, and local AI coding setup on the ASUS ROG Flow Z13
161
Reposted by Deepu K Sasidharan
DEV Community @dev.to · 18/05/2026
Congrats to this week's top 7 authors! @deepu105.bsky.social built a fully offline local AI Linux dev setup, and @debbie.codes documented an entire product in 4 days with an AI agent. Check out these and the rest of the posts below 👇 dev.to/devteam/top-...
dev.to
Top 7 Featured DEV Posts of the Week
Welcome to this week's Top 7, where the DEV editorial team handpicks their favorite posts from the...
061
Deepu K Sasidharan @deepu105.bsky.social · 12/05/2026
I use Arch btw! 😉 My current fully offline AI-assisted Linux dev setup: 🐧 Arch Linux 🌀 niri ✨ DankMaterialShell 🤖 OpenCode 🦙 llama.cpp ⚙️ ROCm 🧠 Qwen/Gemma local models 💻 ASUS ROG Flow Z13 deepu.tech/my-fully-off...
deepu.tech
My fully offline AI-assisted Linux development machine | Technorage
My Arch Linux, Niri, and local AI coding setup on the ASUS ROG Flow Z13
070
Deepu K Sasidharan @deepu105.bsky.social · 09/04/2026
KDash 1.0.0 is out 🎉 A big milestone for the terminal UI dashboard for Kubernetes. - direct shell into containers - A Troubleshoot tab - inline filter across views - aggregate logs for workloads - custom themes Release notes: github.com/kdash-rs/kda... #Kubernetes #DevOps #Rust
020
Deepu K Sasidharan @deepu105.bsky.social · 28/10/2025
Learn how to secure your AI agents to prevent Excessive Agency, a top OWASP LLM vulnerability, by implementing a Zero Trust model. #auth0 auth0.com/blog/mitigat...
auth0.com
Mitigate Excessive Agency in AI Agents with Zero Trust Security
Mitigate Excessive Agency in AI Agents using a Zero Trust Security model. Practical guide for developers to implement RBAC, FGA, OAuth, a...
010
Deepu K Sasidharan @deepu105.bsky.social · 22/10/2025
A perfect circle jerk 😂
010
Deepu K Sasidharan @deepu105.bsky.social · 09/10/2025
Thank you
000
Reposted by Deepu K Sasidharan
Akka @akka.io · 09/10/2025
@deepu105.bsky.social, may the session run seamlessly and engage the audience! Access is given FGA and RAG guide the way Agents know their path #Devoxx #Akka #room7
@deepu105.bsky.social, may the session run seamlessly and engage the audience!

Access is given
FGA and RAG guide the way
Agents know their path
 #Devoxx #Akka #room7
111
Reposted by Deepu K Sasidharan
Jfokus @jfokus.se · 27/08/2025
Jfokus has always been about community + knowledge sharing 🙌 Here are more of the great speakers from 2025. @sharatchander.bsky.social @renato.cavalcanti.io @deepu105.bsky.social 👉 Want to speak in 2026? Submit now: jfokus.se/iamahero
052
Deepu K Sasidharan @deepu105.bsky.social · 01/08/2025
I've been diving deep into the world of AI lately. My latest blog post explores how to build an AI agent that can call internal and external APIs using LangGraph and Auth0 Token Vault. 🗓️ You can check it out to learn how to use it! #AI #GenAI #LangGraph #ToolCalling auth0.com/blog/genai-t...
auth0.com
How to build an AI Assistant with LangGraph and Next.js
Learn how to build a tool-calling AI agent using LangGraph, Next.js, and Auth0. Integrate your own API as tools. Use Google
041
Deepu K Sasidharan @deepu105.bsky.social · 03/07/2025
Published my first HyDE theme for Hyprland :) #Hyprland #HyDE #ArchLinux github.com/deepu105/hyd...
github.com
GitHub - deepu105/hyde-theme-catppuccin-macchiato: Catppuccin Macchiato with Mauve Accent for HyDE
Catppuccin Macchiato with Mauve Accent for HyDE. Contribute to deepu105/hyde-theme-catppuccin-macchiato development by creating an account on GitHub.
011
Deepu K Sasidharan @deepu105.bsky.social · 03/07/2025
Join me next week at #WeAreDevelopers to Delay the AI Overlords: How OAuth and OpenFGA Can Keep Your AI Agents from Going Rogue #auth0 #AI #security
010
Deepu K Sasidharan @deepu105.bsky.social · 29/05/2025
You can open the speaker notes from the (...) menu for a rough transcript
000
Deepu K Sasidharan @deepu105.bsky.social · 29/05/2025
Thanks to everyone who attended my talk at the #DublinTechSummit Here are the slides from the talk, and you can find the GitHub repo link for the demo on the last slide. docs.google.com/presentation...
docs.google.com
DublinTechSummit: Delay the AI Overlords
[intro] Hello friends, hope you are all having a great time. Today we're going to talk about how to stop your AI systems from spilling corporate secrets like a gossipy coworker after happy hour. Let’s...
100
Reposted by Deepu K Sasidharan
Terminal Trove @terminaltrove.bsky.social · 27/05/2025
kdash is a TUI dashboard for Kubernetes clusters. It displays pod metrics, logs, resource YAMLs, and supports context switching, glob filters, clipboard copying and more. @deepu105.bsky.social made kdash using @ratatui.rs and is Terminal Tool of the Week! ⭐️
141
Deepu K Sasidharan @deepu105.bsky.social · 16/05/2025
😂😂😂 futurism.com/klarna-opena...
futurism.com
Company Regrets Replacing All Those Pesky Human Workers With AI, Just Wants Its Humans Back
Years after outsourcing marketing and customer service gigs to AI, the Swedish company Klarna is looking to hire its humans back.
000
Reposted by Deepu K Sasidharan
Simone @simoneb.bsky.social · 12/05/2025
3708
Deepu K Sasidharan @deepu105.bsky.social · 12/03/2025
A new KDash version is released. Includes minor fixes and dependancy updates. github.com/kdash-rs/kda...
github.com
021
Reposted by Deepu K Sasidharan
Voxxed Days Amsterdam @amsterdam.voxxeddays.com · 11/03/2025
🚀 Welcome on board, @auth0byokta.bsky.social 🔐 In today’s digital world, everything starts with Identity, & Auth0 makes authentication seamless, scalable & dev-friendly. Excited to have you as a sponsor at #vdams25! Learn more 👉 auth0.com/docs | auth0.com/blog/developers/ @deepu105.bsky.social
053
Reposted by Deepu K Sasidharan
† lucia scarlet 🏰 @lucia.red · 08/03/2025
every GNU/Linux setup will, given sufficient age, come to have That One Bug that absolutely nobody can explain or fix but that you simply deal with because it is only a mild annoyance
4043623
Reposted by Deepu K Sasidharan
ana @totebug.bsky.social · 07/03/2025
I've been on linux for less than a month and now I'm extremely foss/linux pilled. Never returning to that bloated spyware windows. I'm gonna be annoying as hell.
313108