Retrieval-Augmented Generation (RAG)

Retrieve relevant context from your own documents, then answer with that context

Table of contents

  1. Setup
  2. Document Model with Embeddings
  3. Retrieval Tool
  4. Answering Agent
  5. Next Steps

After reading this guide, you will know:

  • How RAG fits as one step in a larger workflow.
  • How to store document embeddings with neighbor and pgvector.
  • How to generate embeddings automatically when a document changes.
  • How to expose semantic search to an agent as a retrieval tool.
  • How to build an answering agent that cites its sources.

RAG is often just one step in a larger workflow: retrieve relevant context, then answer with that context. You embed your documents once, search them by similarity at query time, and feed the closest matches to the model as grounding. For the mechanics of turning text into vectors, see Embeddings.

Setup

# Gemfile
gem 'neighbor'
gem 'ruby_llm'

# Generate migration for pgvector
bin/rails generate neighbor:vector
bin/rails db:migrate

class CreateDocuments < ActiveRecord::Migration[7.1]
  def change
    create_table :documents do |t|
      t.text :content
      t.string :title
      t.vector :embedding, limit: 1536 # OpenAI embedding size
      t.timestamps
    end

    add_index :documents, :embedding, using: :hnsw, opclass: :vector_l2_ops
  end
end

Document Model with Embeddings

class Document < ApplicationRecord
  has_neighbors :embedding

  before_save :generate_embedding, if: :content_changed?

  private

  def generate_embedding
    response = RubyLLM.embed(content)
    self.embedding = response.vectors
  end
end

Retrieval Tool

class DocumentSearch < RubyLLM::Tool
  description "Searches knowledge base for relevant information"
  parameter :query, description: "Search query"

  def execute(query:)
    embedding = RubyLLM.embed(query).vectors

    documents = Document.nearest_neighbors(
      :embedding,
      embedding,
      distance: "euclidean"
    ).limit(3)

    documents.map do |doc|
      "#{doc.title}: #{doc.content.truncate(500)}"
    end.join("\n\n---\n\n")
  end
end

Answering Agent

class SupportWithDocsAgent < RubyLLM::Agent
  tools DocumentSearch
  instructions "Search for context before answering. Cite sources."
end

agent = SupportWithDocsAgent.new
response = agent.ask("What is our refund policy?").content

Next Steps

  • Embeddings - Turn text into vectors for similarity search.
  • Agentic Workflows - Compose retrieval into larger orchestrations.
  • Tools - Build the retrieval tool and other capabilities.
  • Agents - Define the answering agent class.