Module: BrazilianUtils::TextUtils

Defined in:
lib/brazilian-utils/text-utils.rb

Overview

General-purpose text helpers used across the other domains (and useful on their own) for handling names, company names and addresses written in Brazilian Portuguese.

Constant Summary collapse

DEFAULT_PREPOSITIONS =

Prepositions and articles that stay lower case between two words.

%w[de da do das dos e].freeze
DEFAULT_DESIGNATIONS =

Company designations/abbreviations that are always upper-cased. Deliberately conservative: common words that also happen to be valid abbreviations (e.g. "me", the reflexive pronoun; "sa", the surname "Sá") are left out on purpose, see text.capitalize's description.

%w[LTDA EPP MEI EIRELI CNPJ].freeze
ROMAN_NUMERALS =

Roman numerals commonly used in Brazilian company/entity names (e.g. "Fundação XXI"). Only applied when the original token was already written fully upper case, to avoid mistaking short Portuguese words (like "vi", "mim") for numerals.

%w[
  I II III IV V VI VII VIII IX X XI XII XIII XIV XV XVI XVII XVIII XIX XX
].freeze
WORD_OR_SEPARATOR_REGEX =
/[\p{L}\p{N}]+|[^\p{L}\p{N}]+/.freeze

Class Method Summary collapse

Class Method Details

.capitalize(value, options = {}) ⇒ String

Capitalizes the first letter of each word the way a Brazilian name, company name or address is written.

Examples:

capitalize("esponja vegetal")   #=> "Esponja Vegetal"
capitalize("fulano de tal")     #=> "Fulano de Tal"
capitalize("JOAQUIM JOSÉ")      #=> "Joaquim José"

Parameters:

  • value (String) —

    The text to capitalize.

  • options (Hash) (defaults to: {}) —

    :prepositions and :designations replace the default lists.

Returns:

  • (String) —

    The capitalized text, or an empty string for empty input.



40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
# File 'lib/brazilian-utils/text-utils.rb', line 40

def self.capitalize(value, options = {})
  return '' unless value.is_a?(String)
  return '' if value.strip.empty?

  prepositions = (options[:prepositions] || options['prepositions'] || DEFAULT_PREPOSITIONS).map(&:downcase)
  designations = (options[:designations] || options['designations'] || DEFAULT_DESIGNATIONS).map(&:upcase)

  collapsed = value.strip.gsub(/\s+/, ' ')
  tokens = collapsed.scan(WORD_OR_SEPARATOR_REGEX)

  word_indices = tokens.each_index.select { |i| tokens[i].match?(/\A[\p{L}\p{N}]+\z/) }
  first_word_idx = word_indices.first
  last_word_idx = word_indices.last

  word_indices.each do |i|
    token = tokens[i]

    if token.match?(/\A\d/)
      tokens[i] = token.downcase
      next
    end

    upcase_token = token.upcase

    if token == upcase_token && ROMAN_NUMERALS.include?(upcase_token)
      tokens[i] = upcase_token
      next
    end

    if designations.include?(upcase_token)
      tokens[i] = upcase_token
      next
    end

    downcase_token = token.downcase
    next_token = tokens[i + 1]
    followed_by_punctuation = !next_token.nil? && next_token != ' '

    if prepositions.include?(downcase_token) && i != first_word_idx && i != last_word_idx && !followed_by_punctuation
      tokens[i] = downcase_token
    else
      tokens[i] = token[0].upcase + token[1..-1].to_s.downcase
    end
  end

  tokens.join
end

.remove_accents(value) ⇒ String

Removes diacritical marks (accents, tildes, cedillas) from a string.

Every character is decomposed (Unicode NFD) and every combining mark is dropped.

Examples:

remove_accents("São Paulo")  #=> "Sao Paulo"
remove_accents("Açaí")       #=> "Acai"

Parameters:

  • value (String) —

    The text to strip accents from.

Returns:

  • (String) —

    The text without diacritics.



99
100
101
102
103
# File 'lib/brazilian-utils/text-utils.rb', line 99

def self.remove_accents(value)
  return '' unless value.is_a?(String)

  value.unicode_normalize(:nfd).gsub(/[̀-ͯ]/, '')
end