Module: BrazilianUtils::TextUtils
- Defined in:
- lib/brazilian-utils/text-utils.rb
Overview
General-purpose text helpers used across the other domains (and useful on their own) for handling names, company names and addresses written in Brazilian Portuguese.
Constant Summary collapse
- DEFAULT_PREPOSITIONS =
Prepositions and articles that stay lower case between two words.
%w[de da do das dos e].freeze
- DEFAULT_DESIGNATIONS =
Company designations/abbreviations that are always upper-cased. Deliberately conservative: common words that also happen to be valid abbreviations (e.g. "me", the reflexive pronoun; "sa", the surname "Sá") are left out on purpose, see
text.capitalize's description. %w[LTDA EPP MEI EIRELI CNPJ].freeze
- ROMAN_NUMERALS =
Roman numerals commonly used in Brazilian company/entity names (e.g. "Fundação XXI"). Only applied when the original token was already written fully upper case, to avoid mistaking short Portuguese words (like "vi", "mim") for numerals.
%w[ I II III IV V VI VII VIII IX X XI XII XIII XIV XV XVI XVII XVIII XIX XX ].freeze
- WORD_OR_SEPARATOR_REGEX =
/[\p{L}\p{N}]+|[^\p{L}\p{N}]+/.freeze
Class Method Summary collapse
-
.capitalize(value, options = {}) ⇒ String
Capitalizes the first letter of each word the way a Brazilian name, company name or address is written.
-
.remove_accents(value) ⇒ String
Removes diacritical marks (accents, tildes, cedillas) from a string.
Class Method Details
.capitalize(value, options = {}) ⇒ String
Capitalizes the first letter of each word the way a Brazilian name, company name or address is written.
40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 |
# File 'lib/brazilian-utils/text-utils.rb', line 40 def self.capitalize(value, = {}) return '' unless value.is_a?(String) return '' if value.strip.empty? prepositions = ([:prepositions] || ['prepositions'] || DEFAULT_PREPOSITIONS).map(&:downcase) designations = ([:designations] || ['designations'] || DEFAULT_DESIGNATIONS).map(&:upcase) collapsed = value.strip.gsub(/\s+/, ' ') tokens = collapsed.scan(WORD_OR_SEPARATOR_REGEX) word_indices = tokens.each_index.select { |i| tokens[i].match?(/\A[\p{L}\p{N}]+\z/) } first_word_idx = word_indices.first last_word_idx = word_indices.last word_indices.each do |i| token = tokens[i] if token.match?(/\A\d/) tokens[i] = token.downcase next end upcase_token = token.upcase if token == upcase_token && ROMAN_NUMERALS.include?(upcase_token) tokens[i] = upcase_token next end if designations.include?(upcase_token) tokens[i] = upcase_token next end downcase_token = token.downcase next_token = tokens[i + 1] followed_by_punctuation = !next_token.nil? && next_token != ' ' if prepositions.include?(downcase_token) && i != first_word_idx && i != last_word_idx && !followed_by_punctuation tokens[i] = downcase_token else tokens[i] = token[0].upcase + token[1..-1].to_s.downcase end end tokens.join end |
.remove_accents(value) ⇒ String
Removes diacritical marks (accents, tildes, cedillas) from a string.
Every character is decomposed (Unicode NFD) and every combining mark is dropped.
99 100 101 102 103 |
# File 'lib/brazilian-utils/text-utils.rb', line 99 def self.remove_accents(value) return '' unless value.is_a?(String) value.unicode_normalize(:nfd).gsub(/[̀-ͯ]/, '') end |