Class: Aireview::LlmRouter
- Inherits:
-
Object
- Object
- Aireview::LlmRouter
- Defined in:
- lib/aireview/llm_router.rb
Overview
Walks the models and keys of a stage. An overload is a property of the model, a quota is a property of "key + model", so a 503 switches the model and a quota switches the key. An overloaded model goes into quarantine and is skipped until it expires; once the chain has been walked, the walk goes round again over the models whose quarantine has expired. Two limits: the time budget of the run and the number of requests sent per model per stage (ModelState).
Defined Under Namespace
Classes: Carry, Delay, Route, Slot, Visit
Constant Summary collapse
- SWITCH_HINT =
'Try again later or switch model via --generate-model/--critique-model.'- KIND_LABELS =
{ daily_quota: 'daily quota exhausted', rate_limit: 'rate limited', overloaded: 'overloaded', timeout: 'timed out', unavailable: 'model unavailable' }.freeze
- MAX_ATTEMPTS_PER_MODEL =
ModelState::MAX_REQUESTS_PER_STAGE
- SHORT_RETRY_DELAY =
The first failure of a visit to a model gets one short retry, the second one a quarantine and the next model.
30.0- SHORT_RETRY_JITTER_RANGE =
0.85..1.15
- RATE_LIMIT_BASE_DELAY =
2.0- RATE_LIMIT_JITTER_RANGE =
2.0..5.0
- PROVIDER_RETRY_DELAY_MULTIPLIER_RANGE =
2.0..2.4
- RETRY_WAIT_LOG_FORMAT =
'LLM %<stage>s request will sleep %<delay>.1fs before retry%<source>s ' \ '(request %<next_request>d/%<max_requests>d of the model in this stage, model=%<model>s)'
Instance Method Summary collapse
-
#answered(stage) ⇒ Object
The model that answered last in the stage and its place in the chain.
-
#call(stage:, request_chars:, pinned: false, &request) ⇒ Object
The block receives a route and a request timeout, makes the request and returns the answer.
-
#critique_weaker? ⇒ Boolean
Critique answered with a model below Generate in the pool, for the report.
-
#exclude_answered(stage:, reason:) ⇒ Object
Excludes the model that answered until the end of the stage: its result is invalid (not JSON, not the schema) even after the repair.
-
#fallback_models ⇒ Object
Stages answered by a model other than the primary one, for the report.
-
#initialize(config:, logger:, routing: nil, clock: nil, sleeper: nil) ⇒ LlmRouter
constructor
clock and sleeper are injected in tests: the schedule is checked without real waiting.
- #remaining_time ⇒ Object
Constructor Details
#initialize(config:, logger:, routing: nil, clock: nil, sleeper: nil) ⇒ LlmRouter
clock and sleeper are injected in tests: the schedule is checked without real waiting.
63 64 65 66 67 68 69 70 71 72 73 |
# File 'lib/aireview/llm_router.rb', line 63 def initialize(config:, logger:, routing: nil, clock: nil, sleeper: nil) @config = config @routing = routing || config.routing @logger = logger @clock = clock || -> { Process.clock_gettime(Process::CLOCK_MONOTONIC) } @sleeper = sleeper || ->(seconds) { sleep(seconds) } @models = Hash.new { |states, name| states[name] = ModelState.new } @cursor = {} @used = {} @deadline = nil end |
Instance Method Details
#answered(stage) ⇒ Object
The model that answered last in the stage and its place in the chain.
104 105 106 107 108 109 |
# File 'lib/aireview/llm_router.rb', line 104 def answered(stage) route = @used[stage.to_s] return nil unless route "#{route.candidate} (#{route.candidate_index + 1}/#{chain_for(stage.to_s).size})" end |
#call(stage:, request_chars:, pinned: false, &request) ⇒ Object
The block receives a route and a request timeout, makes the request and returns the answer. An error raised by the block is classified, then comes a retry, another key, another model, or ApiError once the routes are exhausted. pinned — only the model that answered last in this stage (the JSON repair): its failure is RouteExhaustedError, and the stage restarts on another model.
81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 |
# File 'lib/aireview/llm_router.rb', line 81 def call(stage:, request_chars:, pinned: false, &request) @deadline ||= now + @config.llm_time_budget visits = [] noted = Set.new carry = nil loop do slots = available_slots(stage, request_chars, pinned, visits, noted) slot = slots.empty? ? nil : ready_slot(stage, slots, visits) raise exhausted_error(stage, visits, pinned) unless slot status, value = try_candidate(stage, slot, visits, carry, &request) return value if status == :ok carry = value end end |
#critique_weaker? ⇒ Boolean
Critique answered with a model below Generate in the pool, for the report.
112 113 114 115 116 117 118 |
# File 'lib/aireview/llm_router.rb', line 112 def critique_weaker? generate = @used['generate'] critique = @used['critique'] return false unless generate && critique @routing.weaker?(critique.candidate, generate.candidate) end |
#exclude_answered(stage:, reason:) ⇒ Object
Excludes the model that answered until the end of the stage: its result is invalid (not JSON, not the schema) even after the repair. The next request of the stage goes to another model. Returns the excluded model; nil — nothing to exclude.
124 125 126 127 128 129 130 131 |
# File 'lib/aireview/llm_router.rb', line 124 def exclude_answered(stage:, reason:) route = @used[stage.to_s] return nil unless route state(route.candidate).exclude_for_stage(stage, reason) @logger.warn("LLM #{stage}: #{route.candidate} excluded for this stage: #{reason}") route.candidate end |
#fallback_models ⇒ Object
Stages answered by a model other than the primary one, for the report.
99 100 101 |
# File 'lib/aireview/llm_router.rb', line 99 def fallback_models @used.select { |_, route| route.fallback? }.transform_values { |route| route.candidate.to_s } end |
#remaining_time ⇒ Object
133 134 135 |
# File 'lib/aireview/llm_router.rb', line 133 def remaining_time @deadline ? @deadline - now : @config.llm_time_budget.to_f end |