Module: DS::Util::Strings

Included in:
DS::Util
Defined in:
lib/ds/util/strings.rb

Constant Summary collapse

TERMINAL_PUNCT_REGEX =

TERMINAL_PUNCT_REGEX matches strings terminated by any of .,;:?!

%r{\s*([.,;:!]+)("?)$}
ELLIPSIS_REGEX =

ELLIPSIS_REGEX matches strings terminated by ...

%r{\.\.\."?$}
ABBREV_REGEX =

ABBREV_REGEX matches values like 'N.T.', 'O.T.'

%r{\W[A-Z]\.$}
FINAL_QUESTION_REGEX =

Final ? regex

%r{\s*\?(\s*"?\s*)$}

Instance Method Summary collapse

Instance Method Details

#clean_string(string, terminator: nil, force: false) ⇒ String

This method calls

  • convert_mets_superscript
  • remove_brackets
  • fix_double_periods
  • escape_pipes
  • normalize_string

If terminator is non-nil, the method removes any trailing punctuation and whitespace and appends terminator.

Set terminator to `` (empty string) to remove trailing punctuation.

Parameters:

  • string (String) —

    the string to clean

  • terminator (String) (defaults to: nil) —

    the terminator to use, if any

  • force (Boolean) (defaults to: false) —

    use exact termination with terminator

Returns:

  • (String) —

    the cleaned string



26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
# File 'lib/ds/util/strings.rb', line 26

def clean_string string, terminator: nil, force: false
  normal = normalize_string(
    escape_pipes(
      fix_double_periods(
        remove_brackets(
          convert_mets_superscript(string.to_s)
        )
      )
    )
  )

  return normal if terminator.nil?

  cleaned = terminate normal, terminator: terminator, force: force
  # keep cleaning until no changes are made
  return clean_string cleaned unless cleaned == string
  cleaned
end

#clean_white_space(value) ⇒ Object



140
141
142
# File 'lib/ds/util/strings.rb', line 140

def clean_white_space value
  value.to_s.strip.gsub(%r{\s+}, ' ')
end

#convert_mets_superscript(value) ⇒ Object

converts encoded DS 1.0 encoded superscripts to parenthetical values; e.g., 'XVI#^4/4#' is converted to 'XVI(4/4)'



129
130
131
# File 'lib/ds/util/strings.rb', line 129

def convert_mets_superscript value
  value.to_s.gsub(%r{#\^([^#]+)#}, '(\1)')
end

#escape_pipes(value) ⇒ Object

Escape pipe characters in source strings so split operations can avoid splitting on them.



136
137
138
# File 'lib/ds/util/strings.rb', line 136

def escape_pipes value
  value.gsub('|', '\|')
end

#fix_double_periods(value) ⇒ String

Replace any sequence of two '..' with a single period. Ellipses, that is, sequences of three periods '...', are ignored.

fix_double_periods('....')       #  => "...."
fix_double_periods('.. ..')      #  => ". ."
fix_double_periods('... ..')     #  => "... ."
fix_double_periods('... a..')    #  => "... a."
fix_double_periods('a... a..')   #  => "a... a."

Parameters:

  • value (String) —

    the string to process

Returns:

  • (String)


185
186
187
# File 'lib/ds/util/strings.rb', line 185

def fix_double_periods value
  value.to_s.gsub(%r{(?<!\.)\.\.(?!\.)}, '.')
end

#is_url?(value) ⇒ Boolean

Returns:

  • (Boolean)


189
190
191
# File 'lib/ds/util/strings.rb', line 189

def is_url? value
  value.to_s =~ URI::regexp
end

#normalize_string(value) ⇒ String

Strip and replace all sequences of white space with single spaces and apply Unicode normalization. NFC normalization is used for all strings except URLs, to which NFKC normalization is applied. See RFC 3987:

https://datatracker.ietf.org/doc/html/rfc3987#section-5.3.2.2

Parameters:

  • value (String) —

    the string to normalize

Returns:

  • (String) —

    the normalized string



117
118
119
120
121
122
123
124
# File 'lib/ds/util/strings.rb', line 117

def normalize_string value
  form = is_url?(value) ? :nfkc : :nfc
  escape_pipes(
    clean_white_space(
      unicode_normalize(value, form)
    )
  )
end

#remove_brackets(value) ⇒ Object



168
169
170
# File 'lib/ds/util/strings.rb', line 168

def remove_brackets value
  value.to_s.strip.delete_prefix('[').delete_suffix(']')
end

#terminate(str, terminator: '.', force: false) ⇒ String

Add termination to string if it lacks terminal punctuation. Terminal punctuation is one of

. , ; : ? !

When :terminator is '' or nil, trailing punctuation isalways removed.

Strings ending with ellipsis, '...' or '..."' are returned unaltered. This behavior cannot be overridden with :force.

Parameters:

  • str (String) —

    the string to terminate

  • terminator (String) (defaults to: '.') —

    the terminator to use; default: .

  • force (Boolean) (defaults to: false) —

    use exact termination with terminator

Returns:

  • (String)


72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
# File 'lib/ds/util/strings.rb', line 72

def terminate str, terminator: '.', force: false
  str.strip!
  # DE 2022.08.12 Note the \s* to match and replace whitespace before
  #     punctuation; this addresses a bug where some strings were returned
  #     with trailing whitespace: 'value :' => 'value '
  # TODO: Refactor? Two functions: strip_punctuation(), terminate() ??

  # don't strip ellipses
  return str if str =~ ELLIPSIS_REGEX
  # don't strip final periods for strings like "N.T."
  return str if str =~ ABBREV_REGEX

  # don't strip final question marks
  return str if str =~ FINAL_QUESTION_REGEX

  # if :terminator is '' or nil, remove any terminal punctuation
  return str.sub TERMINAL_PUNCT_REGEX, '\2' if terminator.blank?

  # str is already terminated
  return str if str.end_with? terminator
  return str if str.end_with? %Q{#{terminator}"}

  # if string ends with '?', don't add terminator
  return str if str.end_with? '?'

  # str lacks terminal punctuation; add it;
  #  \\1 => keep final '"' (double-quote)
  return str.sub %r{("?)$}, "#{terminator}\\1" if str !~ TERMINAL_PUNCT_REGEX
  # str has to have exact terminal punctuation
  #  \\1 => keep final '"' (double-quote)
  return str.sub TERMINAL_PUNCT_REGEX, "#{terminator}\\2" if force
  # string has some terminal punctuation; return it
  str
end

#unicode_normalize(value, form = :nfc) ⇒ String

Return the string using unicode normalization form form. Use NFC normalization by default. NFC normalization is recommended best practice. See

https://www.honeybadger.io/blog/ruby-unicode-normalization/

In short: NFC should be used for most strings, but NFKC for URLs. See RFC 3987:

https://datatracker.ietf.org/doc/html/rfc3987#section-5.3.2.2

Wikibase uses NFC normalization:

https://doc.wikimedia.org/Wikibase/REL1_28/php/classWikibase_1_1Repo_1_1Parsers_1_1WikibaseStringValueNormalizer.html

Parameters:

  • value (String) —

    the string to normalize

  • form (Symbol) (defaults to: :nfc) —

    the normalization form: :nfc, :nfkc. :nfd, or :nfkd; default: :nfc

Returns:

  • (String) —

    the normalized string



164
165
166
# File 'lib/ds/util/strings.rb', line 164

def unicode_normalize value, form = :nfc
  value.to_s.unicode_normalize form
end