Class: Aireview::JevCritic

Inherits:
Object
  • Object
show all
Defined in:
lib/aireview/jev_critic.rb

Overview

Jev as a critic. Every candidate becomes a few questions (prompts/jev_questions.yml) over one state: the MR and Jira sections, the candidates with the hunks they point at, the diff. The answers turn into keep, reject or unverifiable by thresholds, then duplicates among the kept ones are dropped. Jev cannot rewrite a finding, so there is no refinement. Used as the critique engine (JevStage) and in shadow mode.

Defined Under Namespace

Classes: Asked, Assessment, Request, Result

Constant Summary collapse

QUESTIONS =
YAML.safe_load_file(File.expand_path('prompts/jev_questions.yml', __dir__)).freeze
DECISION_QUESTIONS =

The questions whose answers decide; severity only goes to the log. Their templates go into the review key whole, not as asked about a stub candidate: a single stub gets no duplicate_of question.

%w[real_issue enough_context version_claim duplicate_of].freeze
STATE_WITH_LONGEST_QUESTION_TOKENS =

Jev limits (docs.typesafe.ai/models): the state plus the longest question, and the state plus all questions together. Questions count in both, so the state budget is what they leave.

32_000
REQUEST_TOKENS =
64_000
CHARS_PER_TOKEN =

There is no Jev tokenizer here: characters are converted pessimistically (code and Cyrillic take more tokens than English prose) with a margin. Every request logs the real input_tokens to compare against.

2.5
SAFETY =
0.9
CANDIDATE_FIELDS =
%w[id file line quoted_code category severity problem why suggestion note].freeze
TRUNCATED =
'[truncated to fit the Jev request]'
NOT_SHOWN =
'(the file is not shown in the diff)'
TASK =
'Second pass of a merge request review: check the candidate findings of the first pass ' \
'against the diff and the requirements.'
SEVERITY_RANK =
{'critical' => 0, 'major' => 1, 'minor' => 2}.freeze
SIZE_ERROR =

The docs only say that a 422 body names the offending field; a size rejection is recognized by an explicit phrase about going over a limit. Bare words like "context" are not enough: they occur in question names (enough_context_C1). Any other 422 is a malformed request, and asking again would not help.

/
  \bexceed(?:s|ed|ing)?\b |
  \btoo\s+(?:long|large|many\s+tokens)\b |
  \bmaximum\s+(?:context|length|tokens?)\b |
  \b(?:token|context|length)\s+limit\b
/ix

Class Method Summary collapse

Instance Method Summary collapse

Constructor Details

#initialize(client:, thresholds:, review_instructions: nil, scrub: ->(text) { text }, logger: Logger.new($stderr)) ⇒ JevCritic

scrub — the secret scrubber of the context: the candidates are LLM text and are scrubbed like the candidates JSON of the Critique prompt.



75
76
77
78
79
80
81
82
83
# File 'lib/aireview/jev_critic.rb', line 75

def initialize(client:, thresholds:, review_instructions: nil, scrub: ->(text) { text },
               logger: Logger.new($stderr))
  @client = client
  @thresholds = thresholds
  @scrub = scrub
  instructions = Aireview::Utils.presence(review_instructions.to_s.strip)
  @review_instructions = instructions && scrub.call(instructions)
  @logger = logger
end

Class Method Details

.build(config:, scrub:, logger:) ⇒ Object

Kept candidates in generate order: an edge comes from a duplicate_of_ answer that names another kept id with enough confidence. In a group of duplicates the most severe one stays, on a tie the one generate listed first. Returns id => kept id. The client and the critic as the config sets them up; scrub — the secret scrubber of the context.



91
92
93
94
# File 'lib/aireview/jev_critic.rb', line 91

def self.build(config:, scrub:, logger:)
  new(client: JevClient.new(config: config, logger: logger), thresholds: config.jev_thresholds,
      review_instructions: config.review_instructions, scrub: scrub, logger: logger)
end

.decision_templates ⇒ Object



96
97
98
# File 'lib/aireview/jev_critic.rb', line 96

def self.decision_templates
  QUESTIONS.slice(*DECISION_QUESTIONS)
end

.duplicate_answer(id, ids, answers, threshold) ⇒ Object

The other kept id a duplicate_of answer names with enough confidence.



128
129
130
131
132
133
134
# File 'lib/aireview/jev_critic.rb', line 128

def self.duplicate_answer(id, ids, answers, threshold)
  answer = answers["duplicate_of_#{id}"]
  return nil unless answer.is_a?(Hash) && answer['confidence'].to_f >= threshold

  other = answer['choice']
  other if other != id && ids.include?(other)
end

.duplicate_groups(ids, answers, threshold) ⇒ Object



111
112
113
114
115
116
117
118
# File 'lib/aireview/jev_critic.rb', line 111

def self.duplicate_groups(ids, answers, threshold)
  group_of = ids.to_h { |id| [id, [id]] }
  ids.each do |id|
    other = duplicate_answer(id, ids, answers, threshold)
    join_groups(group_of, id, other) if other
  end
  group_of.values.uniq(&:object_id).select { |group| group.size > 1 }
end

.duplicates(kept, answers, threshold:) ⇒ Object



100
101
102
103
104
105
106
107
108
109
# File 'lib/aireview/jev_critic.rb', line 100

def self.duplicates(kept, answers, threshold:)
  ids = kept.map { |candidate| field(candidate, 'id') }
  rank = kept.to_h do |candidate|
    [field(candidate, 'id'), SEVERITY_RANK.fetch(field(candidate, 'severity'), SEVERITY_RANK.size)]
  end
  duplicate_groups(ids, answers, threshold).each_with_object({}) do |group, dropped|
    winner = group.min_by { |id| [rank[id], ids.index(id)] }
    (group - [winner]).each { |id| dropped[id] = winner }
  end
end

.field(candidate, key) ⇒ Object



136
137
138
# File 'lib/aireview/jev_critic.rb', line 136

def self.field(candidate, key)
  (candidate[key] || candidate[key.to_sym]).to_s.strip
end

.join_groups(group_of, id, other) ⇒ Object



120
121
122
123
124
125
# File 'lib/aireview/jev_critic.rb', line 120

def self.join_groups(group_of, id, other)
  return if group_of[id].equal?(group_of[other])

  merged = group_of[id] + group_of[other]
  merged.each { |member| group_of[member] = merged }
end

Instance Method Details

#assess(context:, candidates:, dedupe: true) ⇒ Object

dedupe — false when the caller merges the verdicts with the LLM ones first and drops duplicates over the whole set (JevStage).



142
143
144
145
146
147
148
149
150
# File 'lib/aireview/jev_critic.rb', line 142

def assess(context:, candidates:, dedupe: true)
  candidates = candidates.map { |candidate| normalize(candidate) }
  sections = diff_sections(context.diff_text)
  asked = Asked.new(answers: {}, unfit: [], requests: 0, model: nil)
  plan_requests(context, candidates, sections).each { |request| ask(request, asked, context, sections) }
  assessments = candidates.map { |candidate| decide(candidate, asked.answers, asked.unfit) }
  assessments = drop_duplicates(assessments, candidates, asked.answers) if dedupe
  Result.new(assessments: assessments, answers: asked.answers, requests: asked.requests, model: asked.model)
end

#preview(context:, candidates:) ⇒ Object

The first request as it would be sent, for --dry-run.



153
154
155
156
# File 'lib/aireview/jev_critic.rb', line 153

def preview(context:, candidates:)
  candidates = candidates.map { |candidate| normalize(candidate) }
  plan_requests(context, candidates, diff_sections(context.diff_text)).first
end