class RubyLLM::Evaluation
Defines reusable evaluations with a dataset, semantic criteria, and Ruby assertions. Datasets default to app/evals/<class_name>.yml, .yaml, .json, or .jsonl. Without declared criteria, evaluates correctness against each case’s expected_output.
class SupportEvaluation < RubyLLM::Evaluation def perform(input) SupportAgent.new.ask(input) end end SupportEvaluation.run.save("tmp/support.json")
Attributes
Returns the current case’s input, reference output, and metadata, respectively.
Returns the current case’s input, reference output, and metadata, respectively.
Returns the current case’s input, reference output, and metadata, respectively.
Returns the original value returned by perform.
Public Class Methods
Source
# File lib/ruby_llm/evaluation.rb, line 74 def adapt(type, &block) raise ArgumentError, 'An adapter needs a class and a block' unless type.is_a?(Module) && block adapters[type] = block end
Converts a custom result into evaluation evidence. The original stays available as result.
adapt Invoice do |invoice| invoice.attributes.slice("total", "currency") end
# File lib/ruby_llm/evaluation.rb, line 82 def cases(dataset: nil, only: nil) loaded = Dataset.load(dataset || self.dataset, name: name) return loaded if only.nil? names = Array(only).map(&:to_s) missing = names - loaded.map(&:name) raise ArgumentError, "Unknown evaluation cases: #{missing.join(', ')}" if missing.any? || names.empty? loaded.select { |test_case| names.include?(test_case.name) } end
Returns the dataset cases without running the application or its evaluators. Supply only to select one or more case names, raising when any name is missing.
# File lib/ruby_llm/evaluation.rb, line 30 def dataset(source = nil, &block) return @dataset if source.nil? && !block raise ArgumentError, 'Pass a dataset or a block, not both' if source && block @dataset = block || source end
Sets the dataset path, enumerable of cases, or block returning cases. With no argument, returns the configured source. The default is discovered by class name.
# File lib/ruby_llm/evaluation.rb, line 58 def evaluation(name, instructions = nil, minimum: nil, evaluator: nil) validate_criterion(name, minimum) name = name.to_sym @declared_names ||= [] raise ArgumentError, "Duplicate evaluation: #{name}" if @declared_names.include?(name) @declared_names << name backend = Evaluator.new(evaluator) if evaluator definitions[name] = { name:, instructions: instructions&.dup&.freeze, minimum:, evaluator: backend }.freeze end
Declares a semantic criterion. A minimum applies to a native probability or score; without one, numeric decisions are measured but do not count as passes. Omit instructions to set a minimum on a question already defined by a Judge. Declared criteria replace implicit correctness. Use evaluation :correctness without instructions to include the built-in reference comparison explicitly. Supply evaluator to override the class evaluator for this criterion.
# File lib/ruby_llm/evaluation.rb, line 42 def evaluator(target = nil, **options) if target.nil? && options.empty? @evaluator = Evaluator.new if @evaluator.nil? return @evaluator end raise ArgumentError, 'A disabled evaluator cannot have model options' if target == false && options.any? @evaluator = target == false ? false : Evaluator.new(target, **options) end
# File lib/ruby_llm/evaluation.rb, line 98 def run(dataset: nil, only: nil, repetitions: 1, id: nil) validate_run(repetitions) id = (id || SecureRandom.uuid).to_s raise ArgumentError, 'An evaluation run id cannot be empty' if id.empty? cases = self.cases(dataset:, only:) groups = evaluation_groups(cases) definitions = Judge::Data.copy(groups.flat_map do |backend, criteria| criteria.map { |criterion| criterion.merge(evaluator: backend.description) } end) started_at = Time.now.utc name = self.name || 'Anonymous evaluation' payload = { evaluation_id: id, evaluation_name: name, total: cases.size * repetitions, completed: 0, started_at: } RubyLLM.instrument('evaluation.ruby_llm', payload) do |event| trials = cases.flat_map do |test_case| Array.new(repetitions) do |index| trial = run_trial(test_case, index + 1, groups, event) yield trial if block_given? trial end end event[:report] = Report.new(name:, trials:, id:, started_at:, definitions:) end end
Executes fresh instances for every case and repetition and returns an Evaluation::Report. A supplied dataset overrides discovery for this run. Configuration errors raise; task, assertion, and evaluator failures are recorded per case. Yields each completed Trial before starting the next. Exceptions from the block propagate. Supply id to correlate reports and instrumentation with an application record or job.
Public Instance Methods
# File lib/ruby_llm/evaluation.rb, line 227
Asserts that condition is truthy. Delegates to Minitest::Assertions and records its assertion count. Requires the minitest gem, which Rails applications already include. The other Minitest assert_* and refute_* methods use the same contract. Call these from assertions.
# File lib/ruby_llm/evaluation.rb, line 242
Asserts equality through Minitest::Assertions. See assert.
# File lib/ruby_llm/evaluation.rb, line 248
Asserts that the collection includes value. See assert.
Source
# File lib/ruby_llm/evaluation.rb, line 204 def assertions; end
Runs Ruby assertions after perform. Override to use assert, refute, and the Minitest assertion family.
Source
# File lib/ruby_llm/evaluation.rb, line 218 def messages @evidence.messages end
Source
# File lib/ruby_llm/evaluation.rb, line 213 def output @evidence.output end
Returns the primary answer or value extracted from result.
Source
# File lib/ruby_llm/evaluation.rb, line 199 def perform(_input) raise NotImplementedError, 'Define perform(input) in your evaluation' end
Runs the application under evaluation. Override in your evaluation class.
# File lib/ruby_llm/evaluation.rb, line 236
Asserts that condition is false or nil. See assert.
Source
# File lib/ruby_llm/evaluation.rb, line 207 def setup; end
Runs before perform on each fresh case instance. Override for application fixtures.
Source
# File lib/ruby_llm/evaluation.rb, line 210 def teardown; end
Runs after each case, even when setup, perform, or an assertion fails.
Source
# File lib/ruby_llm/evaluation.rb, line 223 def tool_calls @evidence.tool_calls end
Returns tool calls recorded in the returned conversation or message.