Class: Aireview::LlmClient

Inherits:
Object
  • Object
show all
Defined in:
lib/aireview/llm_client.rb

Overview

One request to one model with one key through RubyLLM. Retries, keys and reserves belong to LlmRouter: the built-in RubyLLM/Faraday retries (3 by default) are off, otherwise every router attempt would turn into four HTTP requests and burn quota before the error reaches the classifier.

Defined Under Namespace

Classes: Prompt

Constant Summary collapse

NO_TEMPERATURE_MODELS =

RubyLLM 2 sends the temperature as given, while 1.x replaced it with 1.0 for OpenAI reasoning models (o1, o3, gpt-5…) and dropped it for the search ones (gpt-4o-search-preview…), which take no other: such a request is a BadRequest, fatal for the router, and the stage would not even reach a fallback model. The RubyLLM registry knows which models take a temperature; where it does not (a server of your own, assume_model_exists, search models with an empty flag) the name decides, as in 1.x.

%r{\A(?:openai/)?(?:o\d|gpt-5)|-search}

Class Method Summary collapse

Instance Method Summary collapse

Constructor Details

#initialize(config:, logger: Logger.new($stderr)) ⇒ LlmClient

Returns a new instance of LlmClient.



30
31
32
33
34
# File 'lib/aireview/llm_client.rb', line 30

def initialize(config:, logger: Logger.new($stderr))
  @config = config
  @logger = logger
  @contexts = {}
end

Class Method Details

.content(response) ⇒ Object

The answer as RubyLLM 1.x gave it under a schema: the parsed JSON when the text is JSON, the text itself otherwise (the pipeline repairs it). RubyLLM 2 always returns the text, and a Hash that breaks the schema would go to a repair request instead of the next model. An empty answer stays an empty String (Message#parsed would turn it into nil); JSON null becomes nil, as in 1.x.



72
73
74
75
76
77
78
79
# File 'lib/aireview/llm_client.rb', line 72

def self.content(response)
  content = response.content
  return content unless content.is_a?(String) && !content.empty?

  response.parsed
rescue JSON::ParserError
  content
end

Instance Method Details

#prepare(prompt, candidate:, key:, key_index: 0) ⇒ Object

The request as it will go, without sending it: the chat with the model, key, schema, instructions and temperature set. chat.render builds it — that is how the specs check the real request without the network.



54
55
56
57
58
59
60
61
62
63
64
# File 'lib/aireview/llm_client.rb', line 54

def prepare(prompt, candidate:, key:, key_index: 0)
  load_ruby_llm
  stage = prompt.stage.to_s
  chat = build_chat(context: context(stage, candidate, key, key_index), stage: stage,
                    model: candidate.model, provider: candidate.provider)
  chat = configure_reasoning(chat: chat, model: candidate.model, provider: candidate.provider)
    .with_temperature(temperature_for(chat, prompt, candidate))
    .with_schema(prompt.schema)
  chat.with_instructions(prompt.system)
  chat
end

#request(prompt, candidate:, key:, timeout:, key_index: 0) ⇒ Object

Returns the RubyLLM answer; read it with LlmClient.content. A request error is re-raised as is — LlmFailure classifies it.



38
39
40
41
42
43
44
45
46
47
48
49
# File 'lib/aireview/llm_client.rb', line 38

def request(prompt, candidate:, key:, timeout:, key_index: 0)
  stage = prompt.stage.to_s
  model = candidate.model
  chat = prepare(prompt, candidate: candidate, key: key, key_index: key_index)
  @logger.info("LLM #{stage} request started (model=#{model}, temperature=#{chat.temperature || 'model default'})")
  response = Timeout.timeout(timeout) { chat.ask(prompt.user) }
  @logger.info("LLM #{stage} request completed (model=#{model}#{token_counts(response)})")
  response
rescue Timeout::Error
  @logger.warn("LLM #{stage} request timed out after #{timeout.round} seconds (model=#{model})")
  raise
end