jamiemaguire.net
Semantic Kernel: Implementing 100% Local RAG Using Phi-3 With Local Embeddings
In an earlier blog post, we saw how to run the small language model, Phi-3, on your local machine using the ONNX Runtime.
This let you create a simple agent hosted within a console application.
The agent had no access to specific data and would use existing generative AI capabilities to answer human prompts.
To increase the usefulness of your AI agents, you can ground them in your own custom data.
A pattern often used to help achieve this is Retrieval-Augmented Generation, or RAG for short.
RAG modifies interactions with a language model, and lets the model responds to human prompts with references to your own data.
Most of the examples we’ve seen in recent months show how to implement RAG by consuming cloud services such as OpenAI or Azure Search. That might not be suitable for use cases -it wasn’t for me on certain projects.
In this blog post we see how to implement a 100% local RAG solution using Semantic Kernel. We also learn about some basic RAG concepts such as vectors and embeddi