Module: Ace::Review::Atoms::TokenEstimator
- Defined in:
- lib/ace/review/atoms/token_estimator.rb
Overview
Pure function for estimating token counts from text
Uses chars/4 for ASCII and counts each non-ASCII UTF-8 byte as one potential token. The latter deliberately errs high for multilingual review packets rather than letting a Unicode-heavy prompt exceed its cap. This is intentionally simple and fast - actual tokenization would require model-specific tokenizers which adds complexity and latency.
Constant Summary collapse
- CHARS_PER_TOKEN =
Average characters per token for the heuristic Most modern tokenizers average around 3-5 chars/token 4 is a reasonable middle ground
4
Class Method Summary collapse
-
.estimate(text) ⇒ Integer
Estimate token count from a string using chars/4 heuristic.
-
.estimate_file(path) ⇒ Integer
Estimate token count from a file.
-
.estimate_files(paths) ⇒ Integer
Estimate token count from multiple files.
-
.estimate_many(texts) ⇒ Integer
Estimate token count from multiple strings.
Class Method Details
.estimate(text) ⇒ Integer
Estimate token count from a string using chars/4 heuristic
35 36 37 38 39 40 41 42 43 44 |
# File 'lib/ace/review/atoms/token_estimator.rb', line 35 def self.estimate(text) return 0 if text.nil? || text.empty? return (text.length.to_f / CHARS_PER_TOKEN).ceil if text.ascii_only? return text.bytesize unless text.valid_encoding? ascii_chars = text.each_codepoint.count { |codepoint| codepoint < 128 } non_ascii_bytes = text.bytesize - ascii_chars (ascii_chars.to_f / CHARS_PER_TOKEN).ceil + non_ascii_bytes end |
.estimate_file(path) ⇒ Integer
Estimate token count from a file
56 57 58 59 |
# File 'lib/ace/review/atoms/token_estimator.rb', line 56 def self.estimate_file(path) content = File.read(path) estimate(content) end |
.estimate_files(paths) ⇒ Integer
Estimate token count from multiple files
84 85 86 87 88 |
# File 'lib/ace/review/atoms/token_estimator.rb', line 84 def self.estimate_files(paths) return 0 if paths.nil? || paths.empty? paths.sum { |path| estimate_file(path) } end |
.estimate_many(texts) ⇒ Integer
Estimate token count from multiple strings
69 70 71 72 73 |
# File 'lib/ace/review/atoms/token_estimator.rb', line 69 def self.estimate_many(texts) return 0 if texts.nil? || texts.empty? texts.sum { |text| estimate(text) } end |