Class: Aireview::JevCritic
- Inherits:
-
Object
- Object
- Aireview::JevCritic
- Defined in:
- lib/aireview/jev_critic.rb
Overview
Jev as a critic. Every candidate becomes a few questions (prompts/jev_questions.yml) over one state: the MR and Jira sections, the candidates with the hunks they point at, the diff. The answers turn into keep, reject or unverifiable by thresholds, then duplicates among the kept ones are dropped. Jev cannot rewrite a finding, so there is no refinement. Used as the critique engine (JevStage) and in shadow mode.
Defined Under Namespace
Classes: Asked, Assessment, Request, Result
Constant Summary collapse
- QUESTIONS =
YAML.safe_load_file(File.('prompts/jev_questions.yml', __dir__)).freeze
- DECISION_QUESTIONS =
The questions whose answers decide; severity only goes to the log. Their templates go into the review key whole, not as asked about a stub candidate: a single stub gets no duplicate_of question.
%w[real_issue enough_context version_claim duplicate_of].freeze
- STATE_WITH_LONGEST_QUESTION_TOKENS =
Jev limits (docs.typesafe.ai/models): the state plus the longest question, and the state plus all questions together. Questions count in both, so the state budget is what they leave.
32_000- REQUEST_TOKENS =
64_000- CHARS_PER_TOKEN =
There is no Jev tokenizer here: characters are converted pessimistically (code and Cyrillic take more tokens than English prose) with a margin. Every request logs the real input_tokens to compare against.
2.5- SAFETY =
0.9- CANDIDATE_FIELDS =
%w[id file line quoted_code category severity problem why suggestion note].freeze
- TRUNCATED =
'[truncated to fit the Jev request]'- NOT_SHOWN =
'(the file is not shown in the diff)'- TASK =
'Second pass of a merge request review: check the candidate findings of the first pass ' \ 'against the diff and the requirements.'
- SEVERITY_RANK =
{'critical' => 0, 'major' => 1, 'minor' => 2}.freeze
- SIZE_ERROR =
The docs only say that a 422 body names the offending field; a size rejection is recognized by an explicit phrase about going over a limit. Bare words like "context" are not enough: they occur in question names (enough_context_C1). Any other 422 is a malformed request, and asking again would not help.
/ \bexceed(?:s|ed|ing)?\b | \btoo\s+(?:long|large|many\s+tokens)\b | \bmaximum\s+(?:context|length|tokens?)\b | \b(?:token|context|length)\s+limit\b /ix
Class Method Summary collapse
-
.build(config:, scrub:, logger:) ⇒ Object
Kept candidates in generate order: an edge comes from a duplicate_of_
answer that names another kept id with enough confidence. - .decision_templates ⇒ Object
-
.duplicate_answer(id, ids, answers, threshold) ⇒ Object
The other kept id a duplicate_of answer names with enough confidence.
- .duplicate_groups(ids, answers, threshold) ⇒ Object
- .duplicates(kept, answers, threshold:) ⇒ Object
- .field(candidate, key) ⇒ Object
- .join_groups(group_of, id, other) ⇒ Object
Instance Method Summary collapse
-
#assess(context:, candidates:, dedupe: true) ⇒ Object
dedupe — false when the caller merges the verdicts with the LLM ones first and drops duplicates over the whole set (JevStage).
-
#initialize(client:, thresholds:, review_instructions: nil, scrub: ->(text) { text }, logger: Logger.new($stderr)) ⇒ JevCritic
constructor
scrub — the secret scrubber of the context: the candidates are LLM text and are scrubbed like the candidates JSON of the Critique prompt.
-
#preview(context:, candidates:) ⇒ Object
The first request as it would be sent, for --dry-run.
Constructor Details
#initialize(client:, thresholds:, review_instructions: nil, scrub: ->(text) { text }, logger: Logger.new($stderr)) ⇒ JevCritic
scrub — the secret scrubber of the context: the candidates are LLM text and are scrubbed like the candidates JSON of the Critique prompt.
75 76 77 78 79 80 81 82 83 |
# File 'lib/aireview/jev_critic.rb', line 75 def initialize(client:, thresholds:, review_instructions: nil, scrub: ->(text) { text }, logger: Logger.new($stderr)) @client = client @thresholds = thresholds @scrub = scrub instructions = Aireview::Utils.presence(review_instructions.to_s.strip) @review_instructions = instructions && scrub.call(instructions) @logger = logger end |
Class Method Details
.build(config:, scrub:, logger:) ⇒ Object
Kept candidates in generate order: an edge comes from a
duplicate_of_
91 92 93 94 |
# File 'lib/aireview/jev_critic.rb', line 91 def self.build(config:, scrub:, logger:) new(client: JevClient.new(config: config, logger: logger), thresholds: config.jev_thresholds, review_instructions: config.review_instructions, scrub: scrub, logger: logger) end |
.decision_templates ⇒ Object
96 97 98 |
# File 'lib/aireview/jev_critic.rb', line 96 def self.decision_templates QUESTIONS.slice(*DECISION_QUESTIONS) end |
.duplicate_answer(id, ids, answers, threshold) ⇒ Object
The other kept id a duplicate_of answer names with enough confidence.
128 129 130 131 132 133 134 |
# File 'lib/aireview/jev_critic.rb', line 128 def self.duplicate_answer(id, ids, answers, threshold) answer = answers["duplicate_of_#{id}"] return nil unless answer.is_a?(Hash) && answer['confidence'].to_f >= threshold other = answer['choice'] other if other != id && ids.include?(other) end |
.duplicate_groups(ids, answers, threshold) ⇒ Object
111 112 113 114 115 116 117 118 |
# File 'lib/aireview/jev_critic.rb', line 111 def self.duplicate_groups(ids, answers, threshold) group_of = ids.to_h { |id| [id, [id]] } ids.each do |id| other = duplicate_answer(id, ids, answers, threshold) join_groups(group_of, id, other) if other end group_of.values.uniq(&:object_id).select { |group| group.size > 1 } end |
.duplicates(kept, answers, threshold:) ⇒ Object
100 101 102 103 104 105 106 107 108 109 |
# File 'lib/aireview/jev_critic.rb', line 100 def self.duplicates(kept, answers, threshold:) ids = kept.map { |candidate| field(candidate, 'id') } rank = kept.to_h do |candidate| [field(candidate, 'id'), SEVERITY_RANK.fetch(field(candidate, 'severity'), SEVERITY_RANK.size)] end duplicate_groups(ids, answers, threshold).each_with_object({}) do |group, dropped| winner = group.min_by { |id| [rank[id], ids.index(id)] } (group - [winner]).each { |id| dropped[id] = winner } end end |
.field(candidate, key) ⇒ Object
136 137 138 |
# File 'lib/aireview/jev_critic.rb', line 136 def self.field(candidate, key) (candidate[key] || candidate[key.to_sym]).to_s.strip end |
.join_groups(group_of, id, other) ⇒ Object
120 121 122 123 124 125 |
# File 'lib/aireview/jev_critic.rb', line 120 def self.join_groups(group_of, id, other) return if group_of[id].equal?(group_of[other]) merged = group_of[id] + group_of[other] merged.each { |member| group_of[member] = merged } end |
Instance Method Details
#assess(context:, candidates:, dedupe: true) ⇒ Object
dedupe — false when the caller merges the verdicts with the LLM ones first and drops duplicates over the whole set (JevStage).
142 143 144 145 146 147 148 149 150 |
# File 'lib/aireview/jev_critic.rb', line 142 def assess(context:, candidates:, dedupe: true) candidates = candidates.map { |candidate| normalize(candidate) } sections = diff_sections(context.diff_text) asked = Asked.new(answers: {}, unfit: [], requests: 0, model: nil) plan_requests(context, candidates, sections).each { |request| ask(request, asked, context, sections) } assessments = candidates.map { |candidate| decide(candidate, asked.answers, asked.unfit) } assessments = drop_duplicates(assessments, candidates, asked.answers) if dedupe Result.new(assessments: assessments, answers: asked.answers, requests: asked.requests, model: asked.model) end |
#preview(context:, candidates:) ⇒ Object
The first request as it would be sent, for --dry-run.
153 154 155 156 |
# File 'lib/aireview/jev_critic.rb', line 153 def preview(context:, candidates:) candidates = candidates.map { |candidate| normalize(candidate) } plan_requests(context, candidates, diff_sections(context.diff_text)).first end |