Sign in

Suvash Thapaliya

@suva.sh
630 followers 122 following 82 posts

programming & etcéteras

PostsRepliesMedia
Suvash Thapaliya @suva.sh · 28/10/2025
Super late to the "Gigabit at home" party, but recently updated to it, also moved from Synology to Ubiquiti Router+AP setup, Ethernet where possible. Finally I can max out on downloading these bulky models from HF & Ollama store. 😅
130
Suvash Thapaliya @suva.sh · 09/01/2025
Ah man, I really loled at this one. On an abstract level, I really appreciate the 'Building effective agents' post by Anthropic. Still have to fully read the smolagents post by HF, but I already like this "Agency level" table.
150
Suvash Thapaliya @suva.sh · 21/12/2024
I probably read this blog multiple times today. You can really tell that it's written by a team that cares about helping engineers build simple & reliable systems. Thanks again to the authors Erik Schluntz and Barry Zhang. I wish more of the @anthropic.com team would be here on Bsky.
Acknowledgements
Written by Erik Schluntz and Barry Zhang. This work draws upon our experiences building agents at Anthropic and the valuable insights shared by our customers, for which we're deeply grateful.
101
Suvash Thapaliya @suva.sh · 21/12/2024
Finally, the blog summarizes by once again warning against complexity. LLM augmented systems are already hard enough to evaluate, it's only sensible to keep things simple outside the LLM boundaries.
Summary
Success in the LLM space isn't about building the most sophisticated system. It's about building the right system for your needs. Start with simple prompts, optimize them with comprehensive evaluation, and add multi-step agentic systems only when simpler solutions fall short.
111
Suvash Thapaliya @suva.sh · 21/12/2024
And then, there's a pretty good coverage on building agents. Once again, after a good amount of warning that most systems don't need to be designed this way. Sure, agentic systems sound more fun. But, from what I see, most real world problems don't need autonomous plan-attempt-verify loops.
Agents
Agents are emerging in production as LLMs mature in key capabilities—understanding complex inputs, engaging in reasoning and planning, using tools reliably, and recovering from errors. Agents begin their work with either a command from, or interactive discussion with, the human user. Once the task is clear, agents plan and operate independently, potentially returning to the human for further information or judgement. During execution, it's crucial for the agents to gain “ground truth” from the environment at each step (such as tool call results or code execution) to assess its progress. Agents can then pause for human feedback at checkpoints or when encountering blockers. The task often terminates upon completion, but it’s also common to include stopping conditions (such as a maximum number of iterations) to maintain control.Agents can handle sophisticated tasks, but their implementation is often straightforward. They are typically just LLMs using tools based on environmental feedback in a loop. It is therefore crucial to design toolsets and their documentation clearly and thoughtfully. We expand on best practices for tool development in Appendix 2 ("Prompt Engineering your Tools").When to use agents: Agents can be used for open-ended problems where it’s difficult or impossible to predict the required number of steps, and where you can’t hardcode a fixed path. The LLM will potentially operate for many turns, and you must have some level of trust in its decision-making. Agents' autonomy makes them ideal for scaling tasks in trusted environments.

The autonomous nature of agents means higher costs, and the potential for compounding errors. We recommend extensive testing in sandboxed environments, along with the appropriate guardrails.

Examples where agents are useful:

The following examples are from our own implementations:

A coding Agent to resolve SWE-bench tasks, which involve edits to many files based on a task description;
Our “computer use” reference implementation, where Claude uses a computer to accomplish tasks.
351
Suvash Thapaliya @suva.sh · 21/12/2024
And then, there's slightly complex Orchestrator-worker & Evaluator-optimizer workflows. I've been thinking of using something like Evaluator-optimizer for steps that could fail or produce unwanted results, but haven't gotten there yet. Either ways, it's useful to have names for these patterns.
Workflow: Orchestrator-workers
In the orchestrator-workers workflow, a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results.Workflow: Evaluator-optimizer
In the evaluator-optimizer workflow, one LLM call generates a response while another provides evaluation and feedback in a loop.
101
Suvash Thapaliya @suva.sh · 21/12/2024
I appreciate that they've introduced names for some common workflows. I've already been using Prompt Chaining & Routing, considering Parallelization next. More importantly, notice that they're not new hypey names, these flows already existed in the data engineering world.
Workflow: Prompt chaining
Prompt chaining decomposes a task into a sequence of steps, where each LLM call processes the output of the previous one. You can add programmatic checks (see "gate” in the diagram below) on any intermediate steps to ensure that the process is still on track.Workflow: Routing
Routing classifies an input and directs it to a specialized followup task. This workflow allows for separation of concerns, and building more specialized prompts. Without this workflow, optimizing for one kind of input can hurt performance on other inputs.Workflow: Parallelization
LLMs can sometimes work simultaneously on a task and have their outputs aggregated programmatically. This workflow, parallelization, manifests in two key variations:

Sectioning: Breaking a task into independent subtasks run in parallel.
Voting: Running the same task multiple times to get diverse outputs.
101
Suvash Thapaliya @suva.sh · 21/12/2024
...where they walk readers cleanly through a lot of this noise (right now & probably all of 2025). I really love the fact that they keep iterating for simplicity. When starting out, there's no good reason to inherit complexity(langchain etc.), esp. if you plan to have end-to-end control over it.
When (and when not) to use agents
When building applications with LLMs, we recommend finding the simplest solution possible, and only increasing complexity when needed. This might mean not building agentic systems at all. Agentic systems often trade latency and cost for better task performance, and you should consider when this tradeoff makes sense.

When more complexity is warranted, workflows offer predictability and consistency for well-defined tasks, whereas agents are the better option when flexibility and model-driven decision-making are needed at scale. For many applications, however, optimizing single LLM calls with retrieval and in-context examples is usually enough.When and how to use frameworks
There are many frameworks that make agentic systems easier to implement, including:

LangGraph from LangChain;
Amazon Bedrock's AI Agent framework;
Rivet, a drag and drop GUI LLM workflow builder; and
Vellum, another GUI tool for building and testing complex workflows.
These frameworks make it easy to get started by simplifying standard low-level tasks like calling LLMs, defining and parsing tools, and chaining calls together. However, they often create extra layers of abstraction that can obscure the underlying prompts ​​and responses, making them harder to debug. They can also make it tempting to add complexity when a simpler setup would suffice.

We suggest that developers start by using LLM APIs directly: many patterns can be implemented in a few lines of code. If you do use a framework, ensure you understand the underlying code. Incorrect assumptions about what's under the hood are a common source of customer error.
111
Suvash Thapaliya @suva.sh · 21/12/2024
From my own experience building LLM augmented systems this year, I've been telling a lot of my colleagues that most use cases don't really need an "agent". Instead, they need a well defined product workflow with models(llms etc.) in the mix. It's reassuring to read the recent @anthropic.com blog...
What are agents?

"Agent" can be defined in several ways. Some customers define agents as fully autonomous systems that operate independently over extended periods, using various tools to accomplish complex tasks. Others use the term to describe more prescriptive implementations that follow predefined workflows. At Anthropic, we categorize all these variations as agentic systems, but draw an important architectural distinction between workflows and agents:

Workflows are systems where LLMs and tools are orchestrated through predefined code paths.
Agents, on the other hand, are systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks.
141